From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DA00A3DD85E; Sun, 16 Aug 2026 22:46:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920382; cv=none; b=KRjSUDvBL2DMXlcqzPjVNW8idU0aS3jUOyCnb0TInWlojAJtRz1dGAXb5LdXJbInjTPenv3L3YDMyha0DJ/+b+tDBCO7kUxkoG50OzqXD3zHUep8zF6O68VKDQu0syR/dTey0bl0hFDb/haGzzOf00LriYUqr46NrTGm6+v6xOk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920382; c=relaxed/simple; bh=uXb/VQqMucbwkmnH4eFidoCTDJAJTdRCAblD3zMVaKk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lp7RHL4NQFZMsQ9gwoL0hcVLN2CBVUsJeFZ2SM/6spGlyshFTXsf8KOZX7UtzXDNKnR9Jm5ShDihJkwAl7FXuTfzj2QNemD3rEEn9TBvBIXyOqdSndYIu6jZG0gfBFBXEe1KNjKqesu02todyEPAdHL1RLT9rPGqol0NQeqYfs8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=RILNZ87d; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=G7qFUhb6; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="RILNZ87d"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="G7qFUhb6" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 06489EC023D; Sun, 16 Aug 2026 18:46:20 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:20 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920380; x= 1787006780; bh=sX7WB2eZz/VvYH3aSWITi7MLh+xZLicHniz1/7X+rUM=; b=R ILNZ87d+g1/9s1nbp82zB0mI4PrI+aktWEf19zi+pWENoo504d2e67BbgSEUH6Nh +NBa4zJFPaG12uZGaPW4XrJXO0eEX+8Lg1Cy9Bp7+wV1/07dGfmCj0S6ipGlpApy WybuVu60UGDbCSiUplMpZNqy45VWbEVnD25wsuy2KT2CPoA4pQLtwRU8X78rsSoe mP7YNcx4NV5pgX4otjGJ4Y7EICNj2IZ2LvoTvZFbifmEKIC5cZJ43SVA5j6zcvGD 4Jpj5kdf8Gc09S3GmC921A5yjLuKYtydZ6Jfu3OarGHGS8GVsV3ZmWLCGYyExUns 0oWdnpQgbBMs8cBIiGIIw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920380; x=1787006780; bh=s X7WB2eZz/VvYH3aSWITi7MLh+xZLicHniz1/7X+rUM=; b=G7qFUhb6XtMIDY+Dr IW8SlbfgnPxQU0iMmKxXfLrX3fjUTufNyA8tSPbN91rcCuLXb6JNRrJYcpPkPWtL 1v8r3662BieAIRUlD+PHAssmKR5pI3lmKJ0ejf+o89hwRyeor/7twEGVyjgLMshP 5QoEHyPKeI+EMWyxu7ycHjmChYeh87Sa4/AvY84Ha5aKqo5xWjd0NXz/FWQBKQv9 vG/jdqKR74QwZkCE5PDm/ciExeiC5xlO3XWDa3Qk3wj9vaMIQQpA2Lai3fihj7Aw 5DeKbhWSC6fJBYTBZZp/cfRYSiDMlb1qpjq9OuGRiT32njRrgWFU9EaHpLqdjCQu orl5g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtH3 TF5+acMZn3V2kyzNvRs++91DDKXew9DYXOyEQix+jcX2Yij4aOpcXX1d7+cQnG8rqBXUZ5 POE9MfSwpz6o4uAqku2Fk/mui8mjbSq4sgrrfJ+vkfJ6Gp+5/Bhn9p5+re7C/IRKkDU64g 4gbmTiwk8a0UUO+c+nDCpdr//c0KoiHipv/SLkGJBqgqUoywz38yFC3hmxePXhxAL6C77x TfWIxre0e4oICcETnXmhzjHmjO1hlQKMhyWrsIDHJfM6eV52F0dmFRugg7H7QTaq+6wnvn 8Y57Fn2WiXpOTE/0ip6GJbvB75qJpCu/Yb2+Z3gdqHLrT5wlQqTjYKWRTCaQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:18 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 01/57] mm: add pte_folio() Date: Sun, 16 Aug 2026 23:45:13 +0100 Message-ID: <20260816224609.308019-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Callers that want the folio behind a present PTE spell it out as page_folio(pte_page(pte)). Add pte_folio() as the folio companion to pte_page(), and convert the callers in fs/proc/task_mmu.c and mm/hugetlb.c. Preparation for the anonymous collapse engine, which reads the folio behind a PTE in several places. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Rik van Riel --- fs/proc/task_mmu.c | 4 ++-- include/linux/mm.h | 14 ++++++++++++++ mm/hugetlb.c | 8 ++++---- 3 files changed, 20 insertions(+), 6 deletions(-) diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c index 5c54aebe2118..459c779b8141 100644 --- a/fs/proc/task_mmu.c +++ b/fs/proc/task_mmu.c @@ -1277,7 +1277,7 @@ static int smaps_hugetlb_range(pte_t *pte, unsigned l= ong hmask, ptl =3D huge_pte_lock(hstate_vma(vma), walk->mm, pte); ptent =3D huge_ptep_get(walk->mm, addr, pte); if (pte_present(ptent)) { - folio =3D page_folio(pte_page(ptent)); + folio =3D pte_folio(ptent); present =3D true; } else { const softleaf_t entry =3D softleaf_from_pte(ptent); @@ -2227,7 +2227,7 @@ static int pagemap_hugetlb_range(pte_t *ptep, unsigne= d long hmask, ptl =3D huge_pte_lock(hstate_vma(vma), walk->mm, ptep); pte =3D huge_ptep_get(walk->mm, addr, ptep); if (pte_present(pte)) { - struct folio *folio =3D page_folio(pte_page(pte)); + struct folio *folio =3D pte_folio(pte); =20 if (!folio_test_anon(folio)) flags |=3D PM_FILE; diff --git a/include/linux/mm.h b/include/linux/mm.h index 0829e0d3b2d1..eb44e3dfee09 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -2681,6 +2681,20 @@ static inline pte_t mk_pte(const struct page *page, = pgprot_t pgprot) return pfn_pte(page_to_pfn(page), pgprot); } =20 +/** + * pte_folio - Return the folio mapped by a present PTE. + * @pte: A present page table entry. + * + * The folio companion to pte_page(); only meaningful for a present PTE + * that maps a struct-page-backed folio. + * + * Return: The folio containing the page @pte maps. + */ +static inline struct folio *pte_folio(pte_t pte) +{ + return page_folio(pte_page(pte)); +} + /** * folio_mk_pte - Create a PTE for this folio * @folio: The folio to create a PTE for diff --git a/mm/hugetlb.c b/mm/hugetlb.c index dded1768193a..bceab8e14118 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -5280,7 +5280,7 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, s= truct vm_area_struct *vma, * are about to unmap is the actual folio of interest. */ if (folio_provided) { - if (folio !=3D page_folio(pte_page(pte))) { + if (folio !=3D pte_folio(pte)) { spin_unlock(ptl); continue; } @@ -5291,7 +5291,7 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, s= truct vm_area_struct *vma, */ set_vma_resv_flags(vma, HPAGE_RESV_UNMAPPED); } else { - folio =3D page_folio(pte_page(pte)); + folio =3D pte_folio(pte); } =20 pte =3D huge_ptep_get_and_clear(mm, address, ptep, sz); @@ -5514,7 +5514,7 @@ static vm_fault_t hugetlb_wp(struct vm_fault *vmf) return 0; } =20 - old_folio =3D page_folio(pte_page(pte)); + old_folio =3D pte_folio(pte); =20 delayacct_wpcopy_start(); =20 @@ -6189,7 +6189,7 @@ vm_fault_t hugetlb_fault(struct mm_struct *mm, struct= vm_area_struct *vma, * checks whether we can re-use the folio exclusively * for us in case we are the only user of it. */ - folio =3D page_folio(pte_page(vmf.orig_pte)); + folio =3D pte_folio(vmf.orig_pte); if (folio_test_anon(folio) && !folio_trylock(folio)) { need_wait_lock =3D true; goto out_ptl; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 195283E4C98; Sun, 16 Aug 2026 22:46:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920384; cv=none; b=ANDNHTebJrmVtqjFDtDsfjxTH3T3XsEKyQgv0IcYVijIEXcG3Xbl2DeCJ1fINq1WemNUVaJslXJX1z3M63+Xq9DMQK/DKKySOUruzUuH/meuYYaeBMnkTGZttuzTG36HcourbzJC51Q3RiC8oQHewGhKqLr2iyAYKuNYPpMkQGY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920384; c=relaxed/simple; bh=byyeadpcr8jQv/zrgoLcgfBdxmdzK41O0wjm6ljIN14=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FC0NdTpBafBbzuDDnaBHMmU3rjsBgpCMlD8LMdA1HgWiV/W2/wzfBYox7kUW3epTgicoFBxF2Iar8T3QwdIJhbBcl3rX/Dy8Zv8eV0aWcmRf/AYKD3MzDSJM2oW7XnWwU9llfPM4o5a7RLgF79XpH6x+/crMpHh9T0/R8KyMilY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Nhg7kGbZ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=OTOxpCzs; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Nhg7kGbZ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="OTOxpCzs" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfhigh.phl.internal (Postfix) with ESMTP id 0C34D14000EB; Sun, 16 Aug 2026 18:46:22 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-03.internal (MEProxy); Sun, 16 Aug 2026 18:46:22 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920382; x= 1787006782; bh=BHyFDjIBobrd0qUaPZVYZ0a7uxaNjqw7aneS2s20bwA=; b=N hg7kGbZbS1oZ/HPIqYWFxcYj6+ij+hMYs5k9zdJjhEJjcU/ixl6HguJDYdWdoQ3+ Nt6jnTOJvRz8DUgxKyfoMIHlryxvXrOb4j9v0P1kkEcpWRdYXg3j9gQwJ2KP1iKC Sq2iH4w6qvEVoCK/QJaUO2Z0Art2GmZgbABBJGytDptwdY5HZB6GQsiRajt5p/TK g/ffFmTeNYdP4t6TMiJLn70tdLsUSDTVFgGiPjsNuE88L0Y2dAmSkVLNVskBgqwx cAw/smsPGiDR7mk/8TaOH2griIwparA2ex60KTO71BK1glk9RokcEbb37RxkNc33 yRRl7VpXXc5qin+OzG0dQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920382; x=1787006782; bh=B HyFDjIBobrd0qUaPZVYZ0a7uxaNjqw7aneS2s20bwA=; b=OTOxpCzsBncl6WlW4 ealynY9yom3qEKh/pQa57Bk/qtxumvlHUJ14ThdDJn3Z8RgwBIIQjvx3c3fH8Gi1 IuKRnf9LeXs/f76M0OT7egu5SahMgoSxSuwtuCw691gjKkdLRViihx6Axp0vd0yr TOBmXqmYErCZ5dZVALBuPpDJmnxtfRqLPtycNTDm25KUgEf/DcWIujfGddUF/wXQ CRL4Za9072eZQBpzNgLo8MeFQ1+i2deso6Xh181A/vsI4+9OwSfbmBzUYkheDXAS CmTtriT9i8WZ4lAfp3Ug8fT39TuRVQzhDJ+uMo45oUM9C+CLqkr1U76wKlFFpNbD ewAlA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtOm 978x7RbxHsQpwvbWpMsc0V7jUKZyBIK4Ghgn899w66TrjufmyMrNztYeOV/7aH+D06vSn2 fkYq1Xxl4Ay6MsKxqBmUPt0ZykwUxbvAwpCsewJwqFzXlAMjJq9UinsYw26tdhBhGHRVMP 24ruPq5aepnUmarI0NLCHTFHzAVmtYZHpuoq960eg6QHGBrLQRGlCZFJ4ACdQaf9TQA7JF IvIGuBlht/53MNG0pdSU/Y1d4iP+TKdJ1NQMcdleTlUuM1ajYvmJK4vnPYEhwTvaAqWOw0 gXGr0c98E7XTShO9e4Oi1oOA5heJo91XV4uKcA/woltUTTyD97DvZIT80mYA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:21 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 02/57] mm: add pte_none_or_zero() Date: Sun, 16 Aug 2026 23:45:14 +0100 Message-ID: <20260816224609.308019-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A PTE that is none and one that maps the shared zeropage both stand for a page of zeroes the mapping does not own. Code that cares only about the contents can treat the two alike. Move khugepaged's local helper for that test to pgtable.h, below the is_zero_pfn() it is built on. migrate_vma_insert_page() open-codes the same test on the slot it is about to fill. Convert it. It still tells none from the zeropage, but only to decide whether there is an old mapping to flush. No functional change intended. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/linux/pgtable.h | 17 +++++++++++++++++ mm/khugepaged.c | 7 ------- mm/migrate_device.c | 9 ++------- 3 files changed, 19 insertions(+), 14 deletions(-) diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h index 8c093c119e5a..bbee6d31f015 100644 --- a/include/linux/pgtable.h +++ b/include/linux/pgtable.h @@ -2064,6 +2064,23 @@ static inline struct page *_zero_page(unsigned long = addr) =20 #ifdef CONFIG_MMU =20 +/** + * pte_none_or_zero - Does this PTE map nothing, or the shared zeropage? + * @pte: The page table entry to test. + * + * A PTE that is none and one that maps the shared zeropage both stand for= a + * page of zeroes the mapping does not own, so code that only cares about = the + * contents can treat them alike. + * + * Return: %true if @pte is none or maps the shared zeropage. + */ +static inline bool pte_none_or_zero(pte_t pte) +{ + if (pte_none(pte)) + return true; + return pte_present(pte) && is_zero_pfn(pte_pfn(pte)); +} + #ifndef CONFIG_TRANSPARENT_HUGEPAGE static inline int pmd_trans_huge(pmd_t pmd) { diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 5a06e3942e88..5f7126cf42f5 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -348,13 +348,6 @@ struct attribute_group khugepaged_attr_group =3D { }; #endif /* CONFIG_SYSFS */ =20 -static bool pte_none_or_zero(pte_t pte) -{ - if (pte_none(pte)) - return true; - return pte_present(pte) && is_zero_pfn(pte_pfn(pte)); -} - /** * collapse_max_ptes_none - Calculate maximum allowed empty PTEs or PTEs m= apping * the shared zeropage for the given collapse operation. diff --git a/mm/migrate_device.c b/mm/migrate_device.c index 9a346162c688..60afa556b994 100644 --- a/mm/migrate_device.c +++ b/mm/migrate_device.c @@ -1067,14 +1067,9 @@ static void migrate_vma_insert_page(struct migrate_v= ma *migrate, if (check_stable_address_space(mm)) goto unlock_abort; =20 - if (pte_present(orig_pte)) { - unsigned long pfn =3D pte_pfn(orig_pte); - - if (!is_zero_pfn(pfn)) - goto unlock_abort; - flush =3D true; - } else if (!pte_none(orig_pte)) + if (!pte_none_or_zero(orig_pte)) goto unlock_abort; + flush =3D pte_present(orig_pte); =20 /* * Check for userfaultfd but do not deliver the fault. Instead, --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1CCD3E51E4; Sun, 16 Aug 2026 22:46:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920386; cv=none; b=TBSeJEBlbE7S39ZfcMpYw5Z79AcNjU/xEwprMigfbVA1BRABCf96XspkMB1Zb7GkxC4OwTFjjGfElG2IwIrYGnnhpUCrE4IlYGenKB0TGXC2nEl0qIxiuBEB1N0Lynayuj8jhhhjYNVZa+YVhHL2rztTomGPI9oIx4tmmTYdVwQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920386; c=relaxed/simple; bh=HPOCNx4enohxIzT7Fhfygk3tYV5wiLuL+4DNZUVHhGo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=R3y9VMDbArhi99vc/1MmjwkSkkgyKTN0ASvavUxS2QLN97mXXoQHXm2aDxWfjPCxcHzDA2/+Iw8CoXG1i8MGfk7SvrdoGoh+F3wXoP6nklKEZeoxBK++RCSctoV3G1Z/8Nalab1W8ccei5f3d2ckm+p4UivYZ/bvFRr83u6+1Os= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=NVMztfIJ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=QbF1DrsC; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="NVMztfIJ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="QbF1DrsC" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id CBDA2EC0235; Sun, 16 Aug 2026 18:46:23 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:23 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920383; x= 1787006783; bh=dznHIYuQ1ajyQBks66dmRoan+CjULGV0RuvhMnG/G2s=; b=N VMztfIJHgpufE0l08lyVHehrOv6tgNPkbVqrKdNJiWgC1/DN6tCXD+b3K652oQv3 rUPk1XJdMIWbUlYOHW5t3EAUnp+FeqGPPbqacTy2Kf/iBI1IDjSFUWmtQjSxlBRg XeSvsuaPmFGDv6SvP+KBjfBj/089rqB0rSOMpkgmF2LeiyJLK9va1ALJgv36XkoB 6gaP44OjEOnRPDfKuJUCNIuEwaMcKddLWPdcCrEka7gyg2krDJ1dpp5iEX4k4xJk A86uTeE21xObCvG/GHEST4rn+47mfQh0/pT1992r4vmNVe322qu0loavXD9nr01E 0IVHOl+5M8ILE3psRI79w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920383; x=1787006783; bh=d znHIYuQ1ajyQBks66dmRoan+CjULGV0RuvhMnG/G2s=; b=QbF1DrsCUsaenM/zP 75ODtb3Wg9zYCtU/3GE/5VL/FqEdTYOoS1GCyO/KrhfRDNMNxdsk79MgNRV61Apt sxtmsLkpNSzfLgQFyFFh3d0ssLsB3DmdnbiM5hhuDUYTw9lIaA9hYU8JfNVrIMwX G1TUUtCxbkzhATrGpfpoGHET+8UXNUROHIahAwkqdidRDV1g/Hd4QqfMGx9zFdGk hWJVF6IgwHjXYwE0ctmGl6nwi36qNtDdFmaupGgyaH8O/ZISHzBARJ+3H/1NDQ7s y7pJ6GgWxpg5vqicq0TUe+9I9UY+6MsGmNUZJ9DtJNA3tITwSwZFEuJYxW1yn/AN go3kg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtow zRcsF02Rw/s88ciL+Ju4C2PxiXq8tPwu+8Qdo1MrZebmhpblmIj686SZmknQU+yMUOhU1l t9NuITAmqXyZDmpmrtLZj+XoUkG3+0KvG83OUa5Yww/LEDfhZtqn4v9WO3J2AGIB6HEnq5 VYcp0R+H7kX0WKHf7r87JCH51mvuDSLM8YQEEUE/nxj37Y+vVO2vT3W+DqI+jkRevSJNLo DDJIZE4VkgrwzHArnxvHkDSm0mvPCgPYD9SDMYzOR5q3nvaLKSZ/Y1Du6Rf+jn2NkIiJRJ lTQ05f2gfc02JHZZ5xdgaNmnLcXtlUmISaBbPVItb11H6yaSx6wrc2XCMBUw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:23 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 03/57] mm/collapse: add collapse.h for the shared collapse state Date: Sun, 16 Aug 2026 23:45:15 +0100 Message-ID: <20260816224609.308019-4-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Preparation for building the new collapse engine in its own file. The engine and khugepaged.c need to agree on what a collapse result is and what state a scan carries. Move enum scan_result and struct collapse_control into a new mm/collapse.h. No functional change intended. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 60 +++++++++++++++++++++++++++++++++++++++++++++++++ mm/khugepaged.c | 52 +----------------------------------------- 2 files changed, 61 insertions(+), 51 deletions(-) create mode 100644 mm/collapse.h diff --git a/mm/collapse.h b/mm/collapse.h new file mode 100644 index 000000000000..26dbac7beddd --- /dev/null +++ b/mm/collapse.h @@ -0,0 +1,60 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef __MM_COLLAPSE_H +#define __MM_COLLAPSE_H + +#include +#include +#include + +enum scan_result { + SCAN_FAIL, + SCAN_SUCCEED, + SCAN_NO_PTE_TABLE, + SCAN_PMD_MAPPED, + SCAN_EXCEED_NONE_PTE, + SCAN_EXCEED_SWAP_PTE, + SCAN_EXCEED_SHARED_PTE, + SCAN_PTE_NON_PRESENT, + SCAN_PTE_UFFD, + SCAN_PTE_MAPPED_HUGEPAGE, + SCAN_LACK_REFERENCED_PAGE, + SCAN_PAGE_NULL, + SCAN_SCAN_ABORT, + SCAN_PAGE_COUNT, + SCAN_PAGE_LRU, + SCAN_PAGE_LOCK, + SCAN_PAGE_ANON, + SCAN_PAGE_LAZYFREE, + SCAN_PAGE_COMPOUND, + SCAN_ANY_PROCESS, + SCAN_VMA_NULL, + SCAN_VMA_CHECK, + SCAN_ADDRESS_RANGE, + SCAN_DEL_PAGE_LRU, + SCAN_ALLOC_HUGE_PAGE_FAIL, + SCAN_CGROUP_CHARGE_FAIL, + SCAN_TRUNCATED, + SCAN_PAGE_HAS_PRIVATE, + SCAN_STORE_FAILED, + SCAN_COPY_MC, + SCAN_PAGE_FILLED, + SCAN_PAGE_DIRTY_OR_WRITEBACK, +}; + +struct collapse_control { + bool is_khugepaged; + + /* Num pages scanned per node */ + u32 node_load[MAX_NUMNODES]; + + /* Num pages scanned (see khugepaged_pages_to_scan) */ + unsigned int progress; + + /* nodemask for allocation fallback */ + nodemask_t alloc_nmask; + + /* Each bit represents a single occupied (!none/zero) page. */ + DECLARE_BITMAP(mthp_present_ptes, MAX_PTRS_PER_PTE); +}; + +#endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 5f7126cf42f5..804b1d35f52a 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -26,45 +26,11 @@ #include =20 #include +#include "collapse.h" #include "internal.h" #include "page_alloc.h" #include "mm_slot.h" =20 -enum scan_result { - SCAN_FAIL, - SCAN_SUCCEED, - SCAN_NO_PTE_TABLE, - SCAN_PMD_MAPPED, - SCAN_EXCEED_NONE_PTE, - SCAN_EXCEED_SWAP_PTE, - SCAN_EXCEED_SHARED_PTE, - SCAN_PTE_NON_PRESENT, - SCAN_PTE_UFFD, - SCAN_PTE_MAPPED_HUGEPAGE, - SCAN_LACK_REFERENCED_PAGE, - SCAN_PAGE_NULL, - SCAN_SCAN_ABORT, - SCAN_PAGE_COUNT, - SCAN_PAGE_LRU, - SCAN_PAGE_LOCK, - SCAN_PAGE_ANON, - SCAN_PAGE_LAZYFREE, - SCAN_PAGE_COMPOUND, - SCAN_ANY_PROCESS, - SCAN_VMA_NULL, - SCAN_VMA_CHECK, - SCAN_ADDRESS_RANGE, - SCAN_DEL_PAGE_LRU, - SCAN_ALLOC_HUGE_PAGE_FAIL, - SCAN_CGROUP_CHARGE_FAIL, - SCAN_TRUNCATED, - SCAN_PAGE_HAS_PRIVATE, - SCAN_STORE_FAILED, - SCAN_COPY_MC, - SCAN_PAGE_FILLED, - SCAN_PAGE_DIRTY_OR_WRITEBACK, -}; - #define CREATE_TRACE_POINTS #include =20 @@ -103,22 +69,6 @@ static struct kmem_cache *mm_slot_cache __ro_after_init; =20 #define KHUGEPAGED_MIN_MTHP_ORDER 2 =20 -struct collapse_control { - bool is_khugepaged; - - /* Num pages scanned per node */ - u32 node_load[MAX_NUMNODES]; - - /* Num pages scanned (see khugepaged_pages_to_scan) */ - unsigned int progress; - - /* nodemask for allocation fallback */ - nodemask_t alloc_nmask; - - /* Each bit represents a single occupied (!none/zero) page. */ - DECLARE_BITMAP(mthp_present_ptes, MAX_PTRS_PER_PTE); -}; - /** * struct khugepaged_scan - cursor for scanning * @mm_head: the head of the mm list to scan --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 69DD53E3152; Sun, 16 Aug 2026 22:46:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920388; cv=none; b=eHEI5bOf19EmH2Ixn33lZYNMFe6640HsfCml1s5W5xJAVp6xIN5G7NUMaFRvNFUwe8+TtWhgiFNkFdPty1WiNvJ7M6DdRQFuZV/C311HA0Ei1zK4yfBtewAZ04oRGMaiqOD4iWtsB13uoBuEyxGsmcW38plI6kXsXW45/AwZLME= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920388; c=relaxed/simple; bh=Syuy4JD6Oq0GAssmdTBBD6hlSmTyaCJxJFgG6ptkOhA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TfytTgd5ErwMyF19/eCoGjqKCc6l6OO5EIamxcf7juQaxS2+vWca6Lw3Kmo0h5V9pD7lcuPdvf8QiihGv1pmy0otj+F03h63tdlizByGaAMrt5MzwDblukfpcR6KN1yk3dVL63VcOebl9rHz3wQkNpKydHoaWmZWSHUSlrvwouc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=D3Ni3tbJ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=cEPsT0B+; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="D3Ni3tbJ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="cEPsT0B+" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfout.phl.internal (Postfix) with ESMTP id 91497EC0242; Sun, 16 Aug 2026 18:46:25 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-03.internal (MEProxy); Sun, 16 Aug 2026 18:46:25 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920385; x= 1787006785; bh=2u63L0ngEfZ0oZc5UJMeb03US//762Pv3O79FLFlVN8=; b=D 3Ni3tbJD9MHvRzagaj9i5rQ+7ZHnppGXYfpjw2FMYRVFugWxwF2lndP58rlS8Tkg i8BmzBwEv4BnalDdR6aqfy7bXtflnh8yvRW6wN64fynJJmH5QtlpBVF8DFmbl22n 5CZwd5ICB9pds79B5RXAf1QCr2PVbAF+uGOSrl9bve/wnBQCWKIrnvm6ftSygTN6 TsNRaGsZ/xSNFSz5/dOPQRxbxcy4WpAP1XQIvtFJ1sNjZHLg7B9Neyz7wCiH86zo /jIuqyMr58LpUUuk9WXLnERRSfgWF/F1bFjkKrJXQfIEOqD9e9UUvleg+5Rvrhl5 3zmzJEeCGEJQhUrwPgT0w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920385; x=1787006785; bh=2 u63L0ngEfZ0oZc5UJMeb03US//762Pv3O79FLFlVN8=; b=cEPsT0B+Xjq4pdpqh LRcDAkqohAJ+F3LI8aRpaHyCnmDweNb34uftCX1p0uIJvCRUpmJD0rSOE/RJM2mK Phf5Fn/xR9li1nkJDl2nsiHD7b/Xl9WcW1ux9pX5OQ+IHheVJ/WaDVIO9fZ/Sj1o pA8bbp883MtVJa8WGyFN1dvm6H8gpFxNYkXNJPZetaNoQ89MHIyzd0uSKjxcDscw ACrlaQFWotIay+ZxEl0oGs6DvFLAJlgB0P9Fl0ce87hCDdPzriuoXGfigcurwBcR sFzkKM/KGO1leZpLI2e0DVTEnOAIDkEg7z7CtMt3Z8EZZYiufT7pEts8YZboMYUO Z5SdA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTOe O2T+e6i2Gb6iXpx6u856YCM0Twl2zjRv6jsf0mu3PZPrhPjoJp43+TcgdmbLSiZaIa3ttT Xtv3uffV61LDU9Bxx8ngiLZEpCG/jF7lw3dGmsuX6sYHJKahwa5GXZ9nam8ziMSeI1kra1 48PX9SyKC5GvF0rAs+rjraJErRymPbrT3FHL6+nlhWTPQ1cNs+YTdOTsO+Dl3vMKOgjwoZ DqQuTq/btPOXP6J7W2Wo+83UdwvraRj7Hw3QYH/7wgDoOk1mjJLYwF0YAfuJAnZePL/9ld 8D3vhPzRY4A8pdb3x4tFK7UN0/lbov4pDTOoRfRpRgcix+lPNecjcwqamMUA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:24 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 04/57] mm/collapse: rename mthp_present_ptes to eligible_ptes Date: Sun, 16 Aug 2026 23:45:16 +0100 Message-ID: <20260816224609.308019-5-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Neither half of the name holds. A bit is set only after the PTE has passed every check the scan makes: uffd, lazyfree, anonymity and sharing among them. Presence is the first of those criteria, not the whole of it. mthp_collapse() then reads the bitmap starting at the PMD order, so the bitmap is not specific to mTHP either. Name the bitmap for what a set bit means: the scan accepted that PTE as a collapse source. No functional change. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 4 ++-- mm/khugepaged.c | 13 ++++++------- 2 files changed, 8 insertions(+), 9 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index 26dbac7beddd..9c82e71533df 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -53,8 +53,8 @@ struct collapse_control { /* nodemask for allocation fallback */ nodemask_t alloc_nmask; =20 - /* Each bit represents a single occupied (!none/zero) page. */ - DECLARE_BITMAP(mthp_present_ptes, MAX_PTRS_PER_PTE); + /* Each bit marks a PTE the scan accepted as a collapse source */ + DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE); }; =20 #endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 804b1d35f52a..a12aafae8d9c 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -576,7 +576,7 @@ static void collapse_control_init_scan(struct collapse_= control *cc) { memset(cc->node_load, 0, sizeof(cc->node_load)); nodes_clear(cc->alloc_nmask); - bitmap_zero(cc->mthp_present_ptes, MAX_PTRS_PER_PTE); + bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); } =20 static void release_pte_folio(struct folio *folio) @@ -1437,8 +1437,8 @@ static unsigned int max_order_from_offset(unsigned in= t offset) * mthp_collapse() consumes the bitmap that is generated during * collapse_scan_pmd() to determine what regions and mTHP orders fit best. * - * Each bit in cc->mthp_present_ptes represents a single occupied (!none/z= ero) - * page. We start at the PMD order and check if it is eligible for collaps= e; + * Each bit in cc->eligible_ptes marks a PTE the scan accepted as a collap= se + * source. We start at the PMD order and check if it is eligible for colla= pse; * if not, we check the left and right halves of the PTE page table we are * examining at a lower order. * @@ -1469,12 +1469,12 @@ static enum scan_result mthp_collapse(struct mm_str= uct *mm, goto next_order; =20 max_ptes_none =3D collapse_max_ptes_none(cc, NULL, order); - nr_occupied_ptes =3D bitmap_weight_from(cc->mthp_present_ptes, offset, + nr_occupied_ptes =3D bitmap_weight_from(cc->eligible_ptes, offset, offset + nr_ptes); =20 /* * Swap PTEs accepted during the scan are counted in @unmapped, - * not in the present-PTE bitmap. Account them for the PMD-order + * not in the eligible bitmap. Account them for the PMD-order * candidate. */ if (is_pmd_order(order)) @@ -1682,8 +1682,7 @@ static enum scan_result collapse_scan_pmd(struct mm_s= truct *mm, } } =20 - /* Set bit for occupied pages */ - __set_bit(i, cc->mthp_present_ptes); + __set_bit(i, cc->eligible_ptes); /* * Record which node the original page is from and save this * information to cc->node_load[]. --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 926953DD85E; Sun, 16 Aug 2026 22:46:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920390; cv=none; b=iw8frNlkcnbVH40Wv091xnMA4A1qSOmJ9/RFwKuI9eF3NjrOGfTNQpUwISyPQIBJktAIatIO1LN2RHh12TgUFIT9xIsbQSJsau+gDOUKk9G3H0vTDg1GufUbHf7Cs256xh0KLunG3xeFUdTm/0jl9Rfr7Z8qLKJEnKYvmc3cSkk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920390; c=relaxed/simple; bh=Wjvh31Qq1kLIw7hbnmrSdVbHpIWK5viqRhWuswNZaqc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=f2vnr+DkcRlSqdb7b/Ac7or11410JWvPFOaD4JJqwAc3fAf9diTaFc5fsKY+G2BDABg1wnUxin7X9F//qDUOJy6w5be1YKbuSm38b6tKfgwB8+xoLxKmR/ki3csxOun6ydKWlSj+F1LLgtLAsAd3q4QHYwt9W3HGAg48TnbfTi4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=U7g7kOhf; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DjD3EFIf; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="U7g7kOhf"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DjD3EFIf" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.phl.internal (Postfix) with ESMTP id A171514000EE; Sun, 16 Aug 2026 18:46:27 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:46:27 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920387; x= 1787006787; bh=3Yz+MQeog4f8alVx8miCtyDi54X9wM7Fdg32w1gvpyM=; b=U 7g7kOhfMJqjOAi+dbGwJXfRRqaaXLjJwtXsQeEA9D3Y2uNZcCziic5QZxRY/7nIJ BGeUVnTiY+UcoWWmKUL8cCJ0Ys/+DeDBTF22018QH1W3ubgIaHqKmPEZhSVOqkms FSofqfl0QsiKKsMI3TkQLxmmBy0gF7ae4OUZWdYAT8HMMMLGfIgNh/zkh4dvMa96 FkSlQKO2zF0bP2pfp4FClK+PqwGUcL0O/s7EFBqVQTcoYL0ixsi3S47I7g3r0qfN SAvCq1CMQqtzP1NXMPouyU7CWh5p7gxsCz8caXl2SIdBhB2HjF60a8IGIwe9as0k 5xR4VAvU9Ar9jBgcP6/pA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920387; x=1787006787; bh=3 Yz+MQeog4f8alVx8miCtyDi54X9wM7Fdg32w1gvpyM=; b=DjD3EFIf9ognUhVez d6IWPQyTzU9K9kmpvkqWj3QeuLpMn7oldc8X/gWqyqpSFGlueHKUlhLhuB+jBWpK zXejcBu58l3E03ru64kqjwSu8RqwosIr1mbfFsnQz/a4arRvctPVdWdXPFUd7vfY 49tX5O3xG9X7xsQX9EqmalwJlmg6Rm0Tn6OW+Iydx2iE30sSD/CwYrabWu8JKboH 1N45kGs6xD/Zuth05e5jHVRjwzH0F0+IeD+iAyZC951UHUBYSrGNrrudLezaidLc r4PEJZfPfZLjCtneL/qU3fxjAs0z5q/PlVz6jbXcqek0AEhW9DauAO+BL/Wa54pT PdaxQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGj1BuhETEStCjbN/L6JrGmiaqzRpeX6rRNWhHsAzwzRlY9AKzqCvfWqke6XA7GOE A6FFYRRSvteD9S264L5lew1YUpPhGI3TlhRe1mYLYRUsBMnidhJP2r2PMm5Lpq2ztO8wmr yF1ddTX1W5NBA6qs0+EDXUaWwTs6UQPsEiicZ6SmY6z0vOVnZaAe+3gsuR1dg3g0gP4peZ U6KBKKlAxat5APGrEFg0r1ubDi9svGZRyISInygLcwTo8cXqakBbn35auVMrKeg7PKmswS qkEQEJGI+h4SGQWw8a8xHcXO10CUTBpRtdxlbIKv4nbtwWiZJ1jrVLIQttQS09NS+fWjNn 6HzLMzm9DNDw61B2nSsR9GukLd62od1iLfT7YKnHgaC3UwalD7TMYHiRMPWLJfwOpz+p08 K4F9yn8szmSmyQ8k4xp1xbjp++Cw7ifd/qAdDCX5wybC8EKHnqs9um7DxJaIJta+yjJwhG QNHw6zuuolynwtlWYUtWgrLd/Nv1uZYCe0w4LWJI4ojicPtFU41tlRVntG3HDsjNLEpZhY 450m9elMctgV5MCBh248tYELoW9T9UxaeEUMZ1cVYuTVWVMMyEwqIZGoLAFxD5CM5cQxek nby3yM48SCOmGjJs1Qxbz7simVtLvPZ6zgoMfgiL5rjhL4HsQMmlGT25p9+A X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:26 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 05/57] mm/collapse: state what a collapse may do in the policy Date: Sun, 16 Aug 2026 23:45:17 +0100 Message-ID: <20260816224609.308019-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Tests scattered through the collapse path decide what a collapse is allowed to do by asking whether khugepaged started it. Between them they settle: - which VMAs are eligible, and how hard to try for a folio; - how many empty, swapped-out or shared PTEs a window may contain, and whether a sub-PMD window is held to a stricter rule than a PMD; - whether a range has to look used, and whether a MADV_FREE'd page is left alone; - whether the PMD is mapped as part of the request, and whether dirty pages are worth writing back and retrying. None of those is a fact about khugepaged. Each is something the caller decided before asking, and the collapse code should not have to look up who called to find out. Add struct collapse_policy for the caller to fill: khugepaged from its own settings, MADV_COLLAPSE from the fact that a user asked explicitly. Every test becomes a read of a field. khugepaged fills the policy once per scan pass, MADV_COLLAPSE once per call. That is the one change in behaviour: the tunables are sampled once per pass rather than on every call, so a table scanned early in a pass and one scanned late are judged alike. cc->is_khugepaged stays, with a single reader left: the daemon's collapse counter, which is bookkeeping and not policy. collapse_file() also drops a NULL check on the collapse_control. It has one call site, reached only from collapse_single_pmd(), which dereferences cc unconditionally, so the check was already dead. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 47 ++++++++++++++++++++++ mm/khugepaged.c | 105 +++++++++++++++++++++++++++--------------------- 2 files changed, 107 insertions(+), 45 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index 9c82e71533df..44f52ea5bbb8 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -2,6 +2,7 @@ #ifndef __MM_COLLAPSE_H #define __MM_COLLAPSE_H =20 +#include #include #include #include @@ -41,7 +42,53 @@ enum scan_result { SCAN_PAGE_DIRTY_OR_WRITEBACK, }; =20 +/* + * What a collapse is allowed to do, decided by whoever asked for it, so t= he + * code doing it need not ask who its caller is: khugepaged fills this in = from + * its own settings, MADV_COLLAPSE from the fact that a user asked explici= tly. + */ +struct collapse_policy { + /* Limits, stated per PMD; HPAGE_PMD_NR means "no limit" */ + unsigned int max_ptes_none; + unsigned int max_ptes_swap; + unsigned int max_ptes_shared; + + /* + * Hold a sub-PMD window to a stricter rule than a PMD: no swapped-out + * and no shared PTEs at all, and max_ptes_none as + * collapse_max_ptes_none() scales it. khugepaged holds mTHP collapse + * to that; an explicit request does not. + */ + bool strict_sub_pmd; + + /* + * Collapse only where it looks worth doing: require some sign the range + * is in use, and leave clean lazyfree folios for reclaim rather than + * collapsing them into a folio that is not lazyfree. A user who asked + * for a collapse gets one either way. + */ + bool skip_lazyfree; + bool require_referenced; + + /* + * Finish the job rather than leaving it half done for a fault to pick + * up: map the PMD over a file collapse before returning, and write + * dirty pages back and retry once instead of refusing them. Both cost + * latency the caller has asked to pay. + */ + bool install_pmd; + bool writeback_dirty; + + /* How hard to try for a destination folio */ + gfp_t gfp; + + /* Which VMAs are eligible, as thp_vma_allowable_orders() spells it */ + enum tva_type tva_type; +}; + struct collapse_control { + struct collapse_policy policy; + bool is_khugepaged; =20 /* Num pages scanned per node */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index a12aafae8d9c..eebc044a930e 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -310,15 +310,12 @@ struct attribute_group khugepaged_attr_group =3D { static unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order) { - const unsigned int max_ptes_none =3D khugepaged_max_ptes_none; + const unsigned int max_ptes_none =3D cc->policy.max_ptes_none; =20 if (vma && userfaultfd_armed(vma)) return 0; - /* for MADV_COLLAPSE, allow any empty/shared zeropage PTEs */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* for PMD collapse, respect the user defined maximum */ - if (is_pmd_order(order)) + /* The limit as given, at the PMD order and wherever it is not capped */ + if (is_pmd_order(order) || !cc->policy.strict_sub_pmd) return max_ptes_none; /* * for mTHP collapse with the sysctl value set to KHUGEPAGED_MAX_PTES_LIM= IT, @@ -350,19 +347,12 @@ static unsigned int collapse_max_ptes_shared(struct c= ollapse_control *cc, unsigned int order) { /* - * For MADV_COLLAPSE, do not restrict the number of PTEs that map shared - * anonymous pages. + * A sub-PMD window held to the strict rule takes no shared page at all: + * an mTHP is not worth the CoW-breaking. */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* - * for mTHP collapse do not allow collapsing anonymous memory pages that - * are shared between processes. - */ - if (!is_pmd_order(order)) + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) return 0; - /* for PMD collapse, respect the user defined maximum */ - return khugepaged_max_ptes_shared; + return cc->policy.max_ptes_shared; } =20 /** @@ -378,16 +368,12 @@ static unsigned int collapse_max_ptes_swap(struct col= lapse_control *cc, unsigned int order) { /* - * For MADV_COLLAPSE, do not restrict the number PTEs entries or - * pagecache entries that are non-present. + * A sub-PMD window held to the strict rule takes nothing non-present: + * reading pages back to build an mTHP is not worth the latency. */ - if (!cc->is_khugepaged) - return HPAGE_PMD_NR; - /* for mTHP collapse do not allow any non-present PTEs or pagecache entri= es */ - if (!is_pmd_order(order)) + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) return 0; - /* for PMD collapse, respect the user defined maximum */ - return khugepaged_max_ptes_swap; + return cc->policy.max_ptes_swap; } =20 int hugepage_madvise(struct vm_area_struct *vma, @@ -686,7 +672,7 @@ static enum scan_result __collapse_huge_page_isolate(st= ruct vm_area_struct *vma, * If the vma has the VM_DROPPABLE flag, the collapse will * preserve the lazyfree property without needing to skip. */ - if (cc->is_khugepaged && !(vma->vm_flags & VM_DROPPABLE) && + if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && folio_test_lazyfree(folio) && !pte_dirty(pteval)) { result =3D SCAN_PAGE_LAZYFREE; goto out; @@ -775,12 +761,12 @@ static enum scan_result __collapse_huge_page_isolate(= struct vm_area_struct *vma, if (folio_test_large(folio)) list_add_tail(&folio->lru, compound_pagelist); next: - if (cc->is_khugepaged && + if (cc->policy.require_referenced && folio_pte_referenced(folio, vma, addr, pteval)) referenced++; } =20 - if (unlikely(cc->is_khugepaged && !referenced)) { + if (unlikely(cc->policy.require_referenced && !referenced)) { result =3D SCAN_LACK_REFERENCED_PAGE; } else { result =3D SCAN_SUCCEED; @@ -984,6 +970,36 @@ static inline gfp_t alloc_hugepage_khugepaged_gfpmask(= void) return khugepaged_defrag() ? GFP_TRANSHUGE : GFP_TRANSHUGE_LIGHT; } =20 +/* khugepaged collapses on its own initiative, so it obeys its own setting= s. */ +static void collapse_policy_khugepaged(struct collapse_policy *p) +{ + p->max_ptes_none =3D READ_ONCE(khugepaged_max_ptes_none); + p->max_ptes_swap =3D READ_ONCE(khugepaged_max_ptes_swap); + p->max_ptes_shared =3D READ_ONCE(khugepaged_max_ptes_shared); + p->strict_sub_pmd =3D true; + p->skip_lazyfree =3D true; + p->require_referenced =3D true; + p->install_pmd =3D false; + p->writeback_dirty =3D false; + p->gfp =3D alloc_hugepage_khugepaged_gfpmask(); + p->tva_type =3D TVA_KHUGEPAGED; +} + +/* MADV_COLLAPSE was asked for explicitly, so it is not held to those. */ +static void collapse_policy_forced(struct collapse_policy *p) +{ + p->max_ptes_none =3D HPAGE_PMD_NR; + p->max_ptes_swap =3D HPAGE_PMD_NR; + p->max_ptes_shared =3D HPAGE_PMD_NR; + p->strict_sub_pmd =3D false; + p->skip_lazyfree =3D false; + p->require_referenced =3D false; + p->install_pmd =3D true; + p->writeback_dirty =3D true; + p->gfp =3D GFP_TRANSHUGE; + p->tva_type =3D TVA_FORCED_COLLAPSE; +} + #ifdef CONFIG_NUMA static int collapse_find_target_node(struct collapse_control *cc) { @@ -1021,8 +1037,7 @@ static enum scan_result hugepage_vma_revalidate(struc= t mm_struct *mm, unsigned l struct collapse_control *cc, unsigned int order) { struct vm_area_struct *vma; - enum tva_type type =3D cc->is_khugepaged ? TVA_KHUGEPAGED : - TVA_FORCED_COLLAPSE; + enum tva_type type =3D cc->policy.tva_type; =20 if (unlikely(collapse_test_exit_or_disable(mm))) return SCAN_ANY_PROCESS; @@ -1205,8 +1220,7 @@ static enum scan_result __collapse_huge_page_swapin(s= truct mm_struct *mm, static enum scan_result alloc_charge_folio(struct folio **foliop, struct m= m_struct *mm, struct collapse_control *cc, unsigned int order) { - gfp_t gfp =3D (cc->is_khugepaged ? alloc_hugepage_khugepaged_gfpmask() : - GFP_TRANSHUGE); + gfp_t gfp =3D cc->policy.gfp; int node =3D collapse_find_target_node(cc); struct folio *folio; =20 @@ -1559,7 +1573,7 @@ static enum scan_result collapse_scan_pmd(struct mm_s= truct *mm, const unsigned int max_ptes_shared =3D collapse_max_ptes_shared(cc, HPAGE= _PMD_ORDER); const unsigned int max_ptes_swap =3D collapse_max_ptes_swap(cc, HPAGE_PMD= _ORDER); unsigned int max_ptes_none =3D collapse_max_ptes_none(cc, vma, HPAGE_PMD_= ORDER); - enum tva_type tva_flags =3D cc->is_khugepaged ? TVA_KHUGEPAGED : TVA_FORC= ED_COLLAPSE; + enum tva_type tva_flags =3D cc->policy.tva_type; pmd_t *pmd; pte_t *pte, *_pte, pteval; int i; @@ -1658,7 +1672,7 @@ static enum scan_result collapse_scan_pmd(struct mm_s= truct *mm, * If the vma has the VM_DROPPABLE flag, the collapse will * preserve the lazyfree property without needing to skip. */ - if (cc->is_khugepaged && !(vma->vm_flags & VM_DROPPABLE) && + if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && folio_test_lazyfree(folio) && !pte_dirty(pteval)) { result =3D SCAN_PAGE_LAZYFREE; goto out_unmap; @@ -1716,11 +1730,11 @@ static enum scan_result collapse_scan_pmd(struct mm= _struct *mm, goto out_unmap; } =20 - if (cc->is_khugepaged && + if (cc->policy.require_referenced && folio_pte_referenced(folio, vma, addr, pteval)) referenced++; } - if (cc->is_khugepaged && + if (cc->policy.require_referenced && (!referenced || (unmapped && referenced < HPAGE_PMD_NR / 2))) { result =3D SCAN_LACK_REFERENCED_PAGE; @@ -2572,11 +2586,11 @@ static enum scan_result collapse_file(struct mm_str= uct *mm, unsigned long addr, xas_unlock_irq(&xas); =20 /* - * Remove pte page tables, so we can re-fault the page as huge. - * If MADV_COLLAPSE, adjust result to call try_collapse_pte_mapped_thp(). + * Remove pte page tables, so we can re-fault the page as huge. A caller + * that wants the PMD mapped now is told to go and do that. */ retract_page_tables(mapping, start); - if (cc && !cc->is_khugepaged) + if (cc->policy.install_pmd) result =3D SCAN_PTE_MAPPED_HUGEPAGE; folio_unlock(new_folio); =20 @@ -2760,11 +2774,8 @@ static enum scan_result collapse_single_pmd(unsigned= long addr, retry: result =3D collapse_scan_file(mm, addr, file, pgoff, cc); =20 - /* - * For MADV_COLLAPSE, when encountering dirty pages, try to writeback, - * then retry the collapse one time. - */ - if (!cc->is_khugepaged && result =3D=3D SCAN_PAGE_DIRTY_OR_WRITEBACK && + /* Dirty pages are worth a writeback and one more try, if asked for */ + if (cc->policy.writeback_dirty && result =3D=3D SCAN_PAGE_DIRTY_OR_WRITEB= ACK && !triggered_wb && mapping_can_writeback(file->f_mapping)) { const loff_t lstart =3D (loff_t)pgoff << PAGE_SHIFT; const loff_t lend =3D lstart + HPAGE_PMD_SIZE - 1; @@ -2781,7 +2792,7 @@ static enum scan_result collapse_single_pmd(unsigned = long addr, result =3D SCAN_ANY_PROCESS; else result =3D try_collapse_pte_mapped_thp(mm, addr, - !cc->is_khugepaged); + cc->policy.install_pmd); if (result =3D=3D SCAN_PMD_MAPPED) result =3D SCAN_SUCCEED; mmap_read_unlock(mm); @@ -2931,6 +2942,9 @@ static void khugepaged_do_scan(struct collapse_contro= l *cc) =20 lru_add_drain_all(); =20 + /* One policy for the whole pass, so every table is judged the same */ + collapse_policy_khugepaged(&cc->policy); + cc->progress =3D 0; while (true) { cond_resched(); @@ -3159,6 +3173,7 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, if (!cc) return -ENOMEM; cc->is_khugepaged =3D false; + collapse_policy_forced(&cc->policy); cc->progress =3D 0; =20 mmgrab(mm); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 337433E7151; Sun, 16 Aug 2026 22:46:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920391; cv=none; b=OaGmH/K56A4vtnL4Of/0ms0ppvE67pIxQGRLQUYe+RwaKGk6QgkC1CHiWyrtgMF1S0vSrm75Gbx0H2sFQjuJEy0+GPUgjK97w9VjmJm3QigfnapiSTACnE33ZF7vj4b2bV7SjOO/DJ9oIUATqilzTv36F1JZPCUf+59qbCroma0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920391; c=relaxed/simple; bh=fVVigGbJs6r6fUgXMmmASKEboV4jeGdkVi4xoCZtqcw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UFddAcL93earOLF6eQBfB5XljGWP3F3vO9nP09DIbjGVGh84xgNQn+4ioebHcD5OY6irQDPdDnhFHaxBYpft5eeJ+JKCp4IolOya5nVJr4VnOUmc1nLGXkVpdTFffPBTGoJd1TH/z9TQ6hqz72bhVooL8+51zsQAoxkJQkzltIs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=aYM3CqxM; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=c55YAllF; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="aYM3CqxM"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="c55YAllF" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.phl.internal (Postfix) with ESMTP id 57CE514000F8; Sun, 16 Aug 2026 18:46:29 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:46:29 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920389; x= 1787006789; bh=mHXoaFr3Xl7m0kg4SnYsmksRc4op7lmfQDtkE+LuXY4=; b=a YM3CqxMWC4w8Dbescu6PN7PKi8swtEa0wj40y1BxGNKVLIC8LIkPUUZWJl3pOr9r msff05iZkLAaHhIOo/2wMh+Mi3ElwpB767fs6iPJsVcbY+pkVpR16896Qo9qLPLQ LIJKGkfrEp6TZlYGJJAiJgKViA6HXiPSk6NUJCwbGkkO1K0X93NL5kKeVqJfXE5N AuX0SflMysadufZy0zdx6azcQVtGLNh09aHN5A6RuXlHGtM8wQVsqtn8Z2BJyLq9 DazIskgVX6xE9q8CsHMjPlb9F5MaQS/je0MFrwJFHHX1EkoiK4P61wg43FFYuDTs rYICP8xqwmnUm0JOJGO7A== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920389; x=1787006789; bh=m HXoaFr3Xl7m0kg4SnYsmksRc4op7lmfQDtkE+LuXY4=; b=c55YAllFLUgt4HnS4 SvDhPrbmHQ4iv3qKX0c2kkTJNECgf0+UoXVmAELh9yNotkkECH66scAHGi4xerT2 o6+FOnt5TbdCS9n4e73bDTXeddHCiNB9AD6q4VHbo3h31k4XheCOw4aufxmg3+Sb Ld9ODcFi3vQxb0TLNjT6CzBkH8qieQLcjNQc3RL/ps5SQd7nLUv+p54wUGiqqyIe c/wJMpRKbhdblmkR7D8kqKA/djKYK1tQ9LaxXLJ+oAyXkWM1VINdja0pPfgUpyh5 AMpRYhG5f3zTKN604St16D5mdW7QLHKuWQe31YXfmuiRujsezJ6BHWtpwoEWUOTA +kZ/w== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGj1BuhETEStCjbN/L6JrGmiaqzRpeX6rRNWhHsAzwzRlY9AKzqCvfWqke6XA7GOE A6FFYRRSvteD9S264L5lew1YUpPhGI3TlhRe1mYLYRUsBMnidhJP2r2PMm5Lpq2ztO8wmr yF1ddTX1W5NBA6qs0+EDXUaWwTs6UQPsEiicZ6SmY6z0vOVnZaAe+3gsuR1dg3g0gP4peZ U6KBKKlAxat5APGrEFg0r1ubDi9svGZRyISInygLcwTo8cXqakBbn35auVMrKeg7PKmswS qkEQEJGI+h4SGQWw8a8xHcXO10CUTBpRtdxlbIKv4nbtwWiZJ1jrVLIQttQS09NS+fWjmx kJfkCJarA+CKWoFELrMO+xFyN9skUAkWMIJFCoZs0ww52p03xN5269uqaLdXJZhhJnw51P pJAiV6KgxOayjzcdbjHznmIWAtiXC/kL1MmpHj1sKlILUfzdm+nj9opRb7cSWJO6WPUn87 TpYeNZUjHs66wSPOBwVossQcukDWjK1YuqqPDPxP05a5pGxfw/QiM8DBKJqlZMeN1rT2T7 u3O+T/BPulsGv+glBoLmjhm4blSdkZpIliMw28KCjwtdcq1eMRy6TLLfQMjDwHcYdEoLf5 a7Ry8XeDa8Y/LOz3sWyJsySoSKZd6miYbNeg3WRmbo3sBb9qvpeES3J9zOfQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:28 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 06/57] mm/collapse: move the smallest collapse order to collapse.h Date: Sun, 16 Aug 2026 23:45:18 +0100 Message-ID: <20260816224609.308019-7-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The floor on the order a collapse will build is a property of collapse, not of khugepaged. MADV_COLLAPSE reaches the same code and is held to the same floor. Move KHUGEPAGED_MIN_MTHP_ORDER to the shared header as COLLAPSE_MIN_MTHP_ORDER. Preparation for the collapse engine, which sizes its per-window arrays from the same floor. No functional change intended. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 3 +++ mm/khugepaged.c | 6 ++---- 2 files changed, 5 insertions(+), 4 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index 44f52ea5bbb8..1e969292edcb 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -7,6 +7,9 @@ #include #include =20 +/* The smallest order a collapse will build, and so the finest window it c= uts */ +#define COLLAPSE_MIN_MTHP_ORDER 2 + enum scan_result { SCAN_FAIL, SCAN_SUCCEED, diff --git a/mm/khugepaged.c b/mm/khugepaged.c index eebc044a930e..f31689bf75a6 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -67,8 +67,6 @@ static DEFINE_READ_MOSTLY_HASHTABLE(mm_slots_hash, MM_SLO= TS_HASH_BITS); =20 static struct kmem_cache *mm_slot_cache __ro_after_init; =20 -#define KHUGEPAGED_MIN_MTHP_ORDER 2 - /** * struct khugepaged_scan - cursor for scanning * @mm_head: the head of the mm list to scan @@ -1540,8 +1538,8 @@ static enum scan_result mthp_collapse(struct mm_struc= t *mm, * any smaller order enabled. When at the smallest order * we must always move to the next offset. */ - if (order > KHUGEPAGED_MIN_MTHP_ORDER && - (enabled_orders & GENMASK(order - 1, 0))) { + if (order > COLLAPSE_MIN_MTHP_ORDER && + (enabled_orders & GENMASK(order - 1, 0))) { order--; continue; } --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0FD9C3E5594; Sun, 16 Aug 2026 22:46:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920394; cv=none; b=ZRtT4ylVb+I8C5p9qI0RqdhiOr3SnUzNQP4jqQuTtJMAhXjwmHqAc1ptdyZYmjVgmPNy+5oFJEtPYo5cQ6Fy+66Hvvgmy7olVpuO+6OvMWtyolgdkb96hJ/ahlqnotTqN96aWcvqomaVPcUHB+xGz8G2h5/T7QU8U+MtRrJeevE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920394; c=relaxed/simple; bh=geS5fk/U+OuUoV7oZhk3PqAPMQGhLqX4ExImyJ+KQGc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pM8i2XJ+f+R9kkh2Xb1LQyfuMHl5EEOhm3SS3Gq+/zFso/GR0n6lMZULB+pNRLMWRGqvDZHsG7vhLGcTO5eAo96d0xGYZGg9PVfm+DGbdne4UPNPGCcNs6WK+XoLjssRoAY35MsxutlzSrmpjxgeBSniwjeJ/Ua5t/u1FuqopM8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=gAb0TOVp; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=V9uGK6Lo; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="gAb0TOVp"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="V9uGK6Lo" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 3088BEC0243; Sun, 16 Aug 2026 18:46:31 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:31 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920391; x= 1787006791; bh=7VxCBv+v1ZXEJKNvfPy2aqiMmOmTEgUap7KvFAJc7uU=; b=g Ab0TOVpn+9qoUMsUan8zShJ2eYFdfYlNN48hFz7V7rEO6adbiMjWHyXwoc9+B3D+ hEgy73wLfFs/OKh7MXYMvZHEppDFJ7EL82dC4HldTYTGnY9wTtUrM94lz36xZo+A dLPF5/h3pH+dUnyJhVBrdntA79JQiU46MNqF/zWywBW7Ku1ON6cfrFnNieNxrg2x CaWyZa6GDa0PdOgtx/G+gP6udYy1ILjXH/+Nxb/JtwrdFYjR22UCs//sshwbpQ3R KY0ZQhA4JDx87QEczUD1KGu6UHQJA97P9l4WU156i78r5ULxxpwlPnKdbURvOQdN N9USjpiEWUpxwy5ZPNGNg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920391; x=1787006791; bh=7 VxCBv+v1ZXEJKNvfPy2aqiMmOmTEgUap7KvFAJc7uU=; b=V9uGK6LoUNikpiRjD CTRKPEM+RkNxg4AL9IZRimPQcHj0+Y/s89JbzBlqwOrHQ/88GlKAg2GdMr1NXvFn ZGeEIaBa94w9uXJBz6AAv7PFvR9TyNzdZqLG7cBc/iTIUatH2V4ReQZ3sX1bXrtq r87h3/tA0BPU901YuxbyyfH0h+qA7CzO60qAPIVHiQ2L3ZHoxmXQ016qObtnUUvc z88x10Xi7mgucQNiIIHixKQz6bKbZS3zIgVo6TXleiOwEmw+hqMlLc40SmcL50yU 8WUCOygX0YHUHPgvoYf00XGMSY95qDV7uZmVD2pEqf5nXrvn9s71sQVrx8BsI/z2 9p12Q== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7uwKlsd4UyzlmYX1bnlzO4+qp5TmvoTS/qzkFKfMuj25xYssUPl6UU5Ie9xRba Y0bWuaC/AePVMhqrkCnU3j3fo4+kQ4lrNr0qcDw4ZM2xDmZI+m2r8oZbGDZU/DcLlRMJvk pxa8BemzxDXpf0p5hqW+oOweEwIQBvnr/hc1O3uCrLqxLkUCX+QMvT/AmKStBz4XcOBl17 ynvfoFkCi5Qu6ZndF8QgBKMwimBrechIaqu56LTRaFtWn2N2FSoSPFDDGPAClh8o3Wii3h WmXtU4UIpdlwyanZt6DS7CEKxM1G2yBGNUvAI/9KSggaAyLfyuWlS9a0iep84iGUQaZXBZ FT7lUXICRewRR2hApjAGIT2CFP2aecTYW5WBOWdkAu3PZGJMwVOaTCXCbpyyAw5uDutR9N cM1vfgm+C6HgAfygQSu8EUBMJo7+djMMGvy1w1JiB0r7oXIzcj8YaeHLzsFDQuCP7Tk3JE VYcUFIxgd6tXhGEDYtL3d75Uwq7YHKTP8lYdadtR5e+r7ixAFvfU/UibDxJtY36j4/ENrG 7O8W7AyTPTaJ/cCibLitvEYey+a9uFu63HUfggyD3oir7BbLzhjWVTsIkYheuXSmJhUyub ho7En+8+81Adj/Pfu7T+uFcVIloLHtGz/Skb7mjPPkRiSkCjwpI39xh5Zgqw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:30 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 07/57] mm/collapse: sketch the new anonymous collapse engine Date: Sun, 16 Aug 2026 23:45:19 +0100 Message-ID: <20260816224609.308019-8-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Anonymous collapse is held back in two ways that its current shape cannot be patched out of. Functionally, an mTHP is collapsed only when the whole PMD-sized window qualifies: collapse_scan_pmd() reaches mthp_collapse() only on SCAN_SUCCEED. One PTE that disqualifies itself -- uffd-armed, not anonymous, clean lazyfree, off the LRU, pinned -- takes the whole table with it. A 2M range with a single such page yields nothing, even where the half beside it would collapse perfectly well. For scalability, collapse_huge_page() holds mmap_write across a collapse, and anon_vma_lock_write() with it, then does the same again for the next huge page. Every collapse stops every fault in the address space, one huge page at a time. The new engine addresses both. It quiesces its sources the way migration does: migration entries in their PTEs, then a frozen refcount. That is enough to make the copy safe without the exclusive lock, so the engine runs under mmap_read throughout. It also carries a batch of windows through each step together, rather than one window through all of them. Its verdict is per window rather than per table, so a table that cannot become one huge page still yields the largest windows inside it. Lay the design out first and fill it in afterwards. What arrives here is the shape of the engine: the two entry points, and a comment at the top of mm/collapse.c mapping the whole call tree. Every step below them is a stub, filled in before the anonymous path is pointed at the engine. collapse_scan_anon_pmd() finds the table, asks the VMA which orders it allows, scans it, and leaves in the collapse_control what a collapse could use. collapse_anon_pmd() takes that range and cuts it into windows. The line between the two is where the lock goes. A scan only reads a VMA and a page table, so it keeps the mmap_lock it was called under and hands it back. A collapse allocates, copies and flushes, so it is called without the lock and takes it again for each round of its own. Nothing then has to tell a caller whether the lock survived a call, and a caller that finds nothing to collapse never gives the lock up at all. Nothing calls either entry point yet; the anonymous path keeps using the mechanism the engine replaces. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/Makefile | 2 +- mm/collapse.c | 165 ++++++++++++++++++++++++++++++++++++++++++++++++ mm/collapse.h | 24 +++++++ mm/khugepaged.c | 4 +- 4 files changed, 192 insertions(+), 3 deletions(-) create mode 100644 mm/collapse.c diff --git a/mm/Makefile b/mm/Makefile index e7245cb88c66..2bef749a5c21 100644 --- a/mm/Makefile +++ b/mm/Makefile @@ -98,7 +98,7 @@ obj-$(CONFIG_MEMTEST) +=3D memtest.o obj-$(CONFIG_MIGRATION) +=3D migrate.o obj-$(CONFIG_NUMA) +=3D memory-tiers.o obj-$(CONFIG_DEVICE_MIGRATION) +=3D migrate_device.o -obj-$(CONFIG_TRANSPARENT_HUGEPAGE) +=3D huge_memory.o khugepaged.o +obj-$(CONFIG_TRANSPARENT_HUGEPAGE) +=3D collapse.o huge_memory.o khugepage= d.o obj-$(CONFIG_PAGE_COUNTER) +=3D page_counter.o obj-$(CONFIG_LIVEUPDATE_MEMFD) +=3D memfd_luo.o obj-$(CONFIG_MEMCG_V1) +=3D memcontrol-v1.o diff --git a/mm/collapse.c b/mm/collapse.c new file mode 100644 index 000000000000..0e6c3c68b44c --- /dev/null +++ b/mm/collapse.c @@ -0,0 +1,165 @@ +// SPDX-License-Identifier: GPL-2.0 +#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt + +#include +#include +#include +#include /* x86 flush_tlb_range() uses hstate_vma() */ +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include "collapse.h" +#include "internal.h" + +/* + * Anonymous collapse, in rounds. + * + * The folios mapped across a window of PTEs become one folio of that wind= ow's + * order, with the sources quiesced by the two barriers migration uses -- + * migration entries in their PTEs, then a frozen refcount -- so the copy = itself + * needs no lock. The engine runs under mmap_read throughout. + * + * A round carries a batch of candidate windows through the passes togethe= r, + * rather than carrying one window through the whole collapse. [ptl] and + * [pmd lock] mark a pass that takes that lock and drops it again, so no + * page-table lock is ever held across passes; the source folio locks are = the + * exception, held from freeze to putback. [rcu] marks a pass that takes = no + * page-table lock at all and reads the table racily, which only the scan = does. + * + * Allocation happens on both sides of the freeze, and which side comes fi= rst + * matters. collapse_provision() tries first, inside the window and after= the + * freeze, with reclaim masked out of the gfp: the sources are frozen by t= hen, + * so a faulter on one of them waits for this allocation. A candidate the + * allocator has nothing ready for is not failed -- it goes back to select= ion, + * and collapse_reserve() allocates for it before the next round freezes + * anything, outside the window, where reclaim costs khugepaged its own + * progress and nobody else's wait. + * + * That second chance needs the policy's gfp to allow reclaim at all. Whe= n it + * does not, a retry would miss the same way, so the first miss is the ans= wer. + * + * collapse_scan_anon_pmd() judge one PTE table's worth of a VMA + * `- collapse_scan_table() [rcu] a bit per PTE a collapse can = use + * + * collapse_anon_pmd() cut windows from those bits, run th= em + * |- collapse_next_candidate() the next window worth attempting + * `- collapse_run_batch() run the batch, then classify it + * |- collapse_round() below + * `- collapse_classify_result() carry on / lower / abandon + * `- collapse_push_retry() queue it for a lower order + * + * collapse_round() one batch of candidates + * |- collapse_reserve() second try for what the last round + * | missed, with reclaim; sleeps + * |- collapse_deposit() a page table per PMD-order candidate + * |- collapse_revalidate() check the VMA and the table survived + * |- collapse_faultin() make the sources present and exclus= ive; + * | sleeps, and may leave the lock drop= ped + * |- collapse_freeze() raise the barriers [ptl], flush the= TLB + * |- collapse_provision() first try for every other destinati= on, + * | without reclaim: the sources are fr= ozen + * |- collapse_copy() copy into the destinations; sleeps + * |- collapse_install() publish them [ptl], or [pmd lock] a= nd a + * | second TLB flush at the PMD order + * |- collapse_putback() lower the barriers, in order + * `- collapse_finish() settle whatever the round reached + * + * Every slot of a candidate is a real source, a hole (pte_none, zero-fill= ed + * and re-verified still-none at install), or the zeropage (cleared at fre= eze, + * zero-filled). Sources come in "spans" -- consecutive PTEs mapping + * consecutive pages of one folio -- so partially mapped and compound sour= ces + * collapse too: any order below the window's is a source, and a PMD candi= date + * takes even a PTE-mapped THP of its own order. + * + * Nothing calls any of this yet: the anon path still uses the mechanism it + * replaces, and is switched over once both halves are complete. + */ + +/* + * Scan the PTEs between @start and @end and record what a collapse could = use: a + * bit in cc->eligible_ptes for every PTE that may be a source. Returns + * SCAN_SUCCEED when every PTE in the range qualified, otherwise the reaso= n one + * did not, and narrows cc->select_orders to what is still worth trying he= re. + */ +static enum scan_result collapse_scan_table(struct vm_area_struct *vma, + pmd_t *pmd, unsigned long start, + unsigned long end, + struct collapse_control *cc) +{ + return SCAN_SUCCEED; +} + +/* Everything a table is judged on starts empty for each table */ +static void collapse_anon_scan_init(struct collapse_control *cc) +{ + bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); + memset(cc->node_load, 0, sizeof(cc->node_load)); + nodes_clear(cc->alloc_nmask); + + cc->select_orders =3D 0; + cc->nr_collapsed =3D 0; +} + +/* + * Judge one table's worth of @vma, leaving in @cc what a collapse could u= se: + * which orders are still worth attempting, and why the table was turned d= own if + * some order was. Holds mmap_lock throughout -- it only reads -- and a c= aller + * that acts on what it found hands the range to collapse_anon_pmd() after= wards, + * without the lock. + */ +static enum scan_result __maybe_unused +collapse_scan_anon_pmd(struct vm_area_struct *vma, unsigned long start, + unsigned long end, struct collapse_control *cc) +{ + const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; + struct mm_struct *mm =3D vma->vm_mm; + pmd_t *pmd; + + /* One table's worth at most, not empty, and inside the VMA */ + VM_WARN_ON_ONCE(end > pmd_addr + HPAGE_PMD_SIZE || start >=3D end); + VM_WARN_ON_ONCE(start < vma->vm_start || end > vma->vm_end); + + cc->scan_refusal =3D find_pmd_or_thp_or_none(mm, pmd_addr, &pmd); + if (cc->scan_refusal !=3D SCAN_SUCCEED) { + cc->progress++; + return cc->scan_refusal; + } + + /* Cleared only once a table has turned out to be there */ + collapse_anon_scan_init(cc); + + cc->select_orders =3D collapse_possible_orders(vma, vma->vm_flags, + cc->policy.tva_type); + if (!cc->select_orders) { + cc->scan_refusal =3D SCAN_VMA_CHECK; + return cc->scan_refusal; + } + + /* The scan narrows select_orders to whatever is left worth trying */ + cc->scan_refusal =3D collapse_scan_table(vma, pmd, start, end, cc); + + return cc->scan_refusal; +} + +/* + * Cut the table into candidate windows and collapse what fits, from the + * largest order downwards. Returns what the table yielded: a collapse, or + * the reason it did not. + */ +static enum scan_result __maybe_unused +collapse_anon_pmd(struct mm_struct *mm, unsigned long start, unsigned long= end, + struct collapse_control *cc) +{ + return SCAN_FAIL; +} diff --git a/mm/collapse.h b/mm/collapse.h index 1e969292edcb..e2af4c47cb60 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -105,6 +105,30 @@ struct collapse_control { =20 /* Each bit marks a PTE the scan accepted as a collapse source */ DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE); + + /* Orders still worth attempting in the table being scanned */ + unsigned long select_orders; + + /* PTEs collapsed in it so far */ + unsigned int nr_collapsed; + + /* + * Why the scan would not take all of the table, or SCAN_SUCCEED if it + * took every order it was offered. Not the opposite of what the scan + * selected: a table can be worth collapsing at one order and refused at + * another, so a scan that found work still has a reason to report, and + * the collapse reports it when it salvages nothing. + */ + enum scan_result scan_refusal; }; =20 +/* + * Defined in khugepaged.c, which still uses them itself. + * TODO: move each into collapse.c once its last khugepaged.c user is gone. + */ +unsigned long collapse_possible_orders(struct vm_area_struct *vma, + vm_flags_t vm_flags, enum tva_type tva_flags); +enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, + unsigned long address, pmd_t **pmd); + #endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index f31689bf75a6..26d25093260b 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -498,7 +498,7 @@ void __khugepaged_enter(struct mm_struct *mm) * Check what orders are possible based on the vma and collapse type. * This is used to determine if mTHP collapse is a viable option. */ -static unsigned long collapse_possible_orders(struct vm_area_struct *vma, +unsigned long collapse_possible_orders(struct vm_area_struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags) { unsigned long orders; @@ -1090,7 +1090,7 @@ static inline enum scan_result check_pmd_state(pmd_t = *pmd) return SCAN_SUCCEED; } =20 -static enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, +enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, unsigned long address, pmd_t **pmd) { *pmd =3D mm_find_pmd(mm, address); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B12CA3E7BCB; Sun, 16 Aug 2026 22:46:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920396; cv=none; b=UYhKbc0q4tswI8K9ORkauXZy8LmZ/y2GF3trYqT+6S7N6HQl9xlD9M1k7Mtq+qFw8bztKfKJoRa/xRE+dLcRGbWQGzDtuivKZ1zJ3BXgKg74N3KH4uN+xyiCRGS97NKv1tUq0NvR8G3KYOdibqrc/Vcxg0E5DOgrbHs6ylBAT4A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920396; c=relaxed/simple; bh=8WkLGlcjcLk0GOM4RsXBUVwghE+fp5OQUIpf2WE6QHA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ovE0chv4q5FvOeyHOe/I8syV5Jx6DeA/ev8ggCu3F+RsR7AIXAT4qQy70e6pth+GwrvJapY4GL4DO0jd0tM5bYZ7U0QStNMABL0r0s0o1CWYzQxeCWTbKhOwWEIfdax3L6UI0B23J6Kn1GnrIbieUMyT94JGhK6hdIb05ahC3JM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=p4USo8N9; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Qhx8ktES; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="p4USo8N9"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Qhx8ktES" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfout.phl.internal (Postfix) with ESMTP id 10E7BEC0241; Sun, 16 Aug 2026 18:46:33 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:46:33 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920393; x= 1787006793; bh=8XhPXvryxg5fpsuZkpieQ1Un5AH038hbB+W7T9Aoa/s=; b=p 4USo8N9/+B5/dIGGB8UZK3mvOa5tz05SY2zXme+4ezgXIb5PS9VjAjOicaRhtCW5 95dF04FTD/SLLZHhVVpSvn87U24l1Ii/ViYCLIvUUoHwNLsyHGUcnU8+0vYmuZnV 502aXAXVXT7bAfLCGrZ1y9np0XapvqWNaOaKRhAWPB/pxdOxlnOfBhdwdfyq5epK H1hR8/siemyvyEUrUetJvOnYU3jRYFp2R86xQbPwGeTH1pqveyStgZJ9z42kAtbq WQ8/jHAk6PzTfCCzNEYX41l5GTH8o6KnY9konMzLX30on9Fn7U7umcuhjzlJk3Q+ njpbqTTf8UseR0OPVIWrg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920393; x=1787006793; bh=8 XhPXvryxg5fpsuZkpieQ1Un5AH038hbB+W7T9Aoa/s=; b=Qhx8ktESQtz/mufMN pHMdTjo47SG2Zg3B4+ysZ+o26OxTe5JUrqwUF6Gf751BtExE1XB4/4ANROgtIqaa mxJtXU5cwmm9KAuQW8PtKm4jrNH3C5zXNbUPNA8VOD75+eVgTVLq0hU6O6kOeiSV NouKdzqyAABI3+bT+8b9pfN6dZBZEovMcCyuAkUquAwUPohgaytbkUhr+F5DG6j/ bhRVd3ehj/dPhTVzJwnC9qBe8MNYFmwn+CRvjmgg6S8PUmugc+fpBsmNU5Alb4tW Gv0jcu54m04uBcU75FO6N+FVboxZbucchY7e5Y6pqOaQfLP4fUa2vH9smOICFDUB qbBMw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtfW NEpLu0J2+1AtombAx/oKQAUo4FdAqIyUBygKONCIsIuzHn0ehaewfPdQc9JRA72ZVjUcvT NN0yVeiW9iRHbQYdFkF2m3zJxcH+dpN3/rKfytNUy69Sg+2VSptJPgM26K/k2U4bFEl7EU 690xUCzQ2LJrxar1DAPsq7zqzyyluJwuIcx9Z841gYPLj1lJoDte6VWlNqmUvxKKp1BwlP FLEH3ulX8mo8toqp8ngqwN6XZKd6wxD+06z8GoGpI2LzwWCTbBUI/s9vlIEsR3N3tn75UR b2ROg4E02kVetP8D1V/XH9Rhum6U9QIMIrGTLJxUGQ368Ahd/W1KjD8NSZPA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:32 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 08/57] mm/collapse: scan a table for what a collapse could use Date: Sun, 16 Aug 2026 23:45:20 +0100 Message-ID: <20260816224609.308019-9-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the scan. Walk the range and set a bit in cc->eligible_ptes for every PTE a collapse may take as a source: present, anonymous, not uffd-armed, on the LRU and unlocked. The bit is set last, so a PTE that failed anything leaves it clear. The walk takes no page table lock. What it produces is advice: the freeze settles every question the scan asks, by re-reading the table under the lock and freezing each source to the count it expects. A racy read can only cost a candidate the freeze then refuses, or miss one the next pass finds. What it buys is that a fault in the range does not wait for a walk of the whole table. pte_offset_map() holds rcu_read_lock() until pte_unmap(), which keeps the table from being freed underneath the walk. mmap_lock keeps the VMA attached, without which free_pgtables() could free it without waiting for RCU at all. The verdict is two-sided, which is the point: - A PTE that disqualifies only itself leaves the bitmap clear there and drops the PMD order, since a PMD candidate needs the whole table. Selection still gets the smaller windows that avoid it. - What refuses the table as a unit -- a limit the whole range exceeds, or sources spread across nodes too distant for one folio to serve -- leaves no order eligible at all. Limits on swapped-out and shared PTEs are stated per PMD and scaled to what was actually scanned, so a partial table is held to the same density as a whole one. A folio whose reference count its mappings do not account for -- a GUP pin, say -- is left to the freeze rather than refused here. folio_expected_ref_count() wants a folio that cannot change order while it is read. This walk holds no page table lock and no folio lock, so a folio splitting underneath it would have its count read for the wrong size. A reference of its own would not help: that stops a folio being freed, not split. Whether a range has to look used at all is the caller's policy, so only a caller that asks gathers the young/referenced evidence. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 239 +++++++++++++++++++++++++++++++++++++++++++++++- mm/collapse.h | 7 ++ mm/khugepaged.c | 8 +- 3 files changed, 249 insertions(+), 5 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 0e6c3c68b44c..66931ef6a6d0 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -86,6 +86,20 @@ * replaces, and is switched over once both halves are complete. */ =20 +/* + * Is @count past a limit stated per PMD, when only part of a table was sc= anned? + * Scale the comparison to the table so a partial scan is held to the same + * density as a whole one. + */ +static bool collapse_exceeds_limit(unsigned int count, unsigned int max_pe= r_pmd, + unsigned long start, unsigned long end) +{ + const unsigned long nr_scanned =3D (end - start) >> PAGE_SHIFT; + + return (unsigned long)count * HPAGE_PMD_NR > + (unsigned long)max_per_pmd * nr_scanned; +} + /* * Scan the PTEs between @start and @end and record what a collapse could = use: a * bit in cc->eligible_ptes for every PTE that may be a source. Returns @@ -97,7 +111,230 @@ static enum scan_result collapse_scan_table(struct vm_= area_struct *vma, unsigned long end, struct collapse_control *cc) { - return SCAN_SUCCEED; + const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; + unsigned int max_ptes_none, max_ptes_swap, max_ptes_shared; + int none_or_zero =3D 0, shared =3D 0, referenced =3D 0, unmapped =3D 0; + enum scan_result result, pmd_result =3D SCAN_SUCCEED; + unsigned int first_offset; + unsigned long addr; + pte_t *pte; + int i; + + max_ptes_none =3D collapse_max_ptes_none(cc, vma, HPAGE_PMD_ORDER); + max_ptes_swap =3D collapse_max_ptes_swap(cc, HPAGE_PMD_ORDER); + max_ptes_shared =3D collapse_max_ptes_shared(cc, HPAGE_PMD_ORDER); + + /* + * No page table lock: what this builds is advice, and the freeze settles + * every question it asks by re-reading the table under the lock and + * freezing each source to the count it expects. A racy read can only + * cost a candidate that the freeze then refuses, or miss one that the + * next pass finds. What it buys is that a fault in this range does not + * wait for a scan of the whole table. + * + * pte_offset_map() holds rcu_read_lock() until pte_unmap(), which is + * what keeps the table itself from being freed underneath the walk; + * mmap_lock keeps the VMA attached, without which free_pgtables() could + * free it without waiting for RCU at all. Nothing below here sleeps. + */ + pte =3D pte_offset_map(pmd, start); + if (!pte) { + cc->progress++; + result =3D SCAN_NO_PTE_TABLE; + goto out_no_table; + } + + /* + * The bitmap and the selection offsets stay relative to the table: + * natural-alignment math needs the table-absolute position, not the + * position within an arbitrarily placed VMA. + */ + first_offset =3D (start - pmd_addr) >> PAGE_SHIFT; + for (i =3D first_offset, addr =3D start; addr < end; + i++, addr +=3D PAGE_SIZE) { + pte_t pteval =3D ptep_get(pte + (i - first_offset)); + struct folio *folio; + struct page *page; + int node; + + cc->progress++; + + if (pte_none_or_zero(pteval)) { + if (++none_or_zero > max_ptes_none && + pmd_result =3D=3D SCAN_SUCCEED) { + pmd_result =3D SCAN_EXCEED_NONE_PTE; + count_vm_event(THP_SCAN_EXCEED_NONE_PTE); + count_mthp_stat(HPAGE_PMD_ORDER, + MTHP_STAT_COLLAPSE_EXCEED_NONE); + } + continue; + } + if (!pte_present(pteval)) { + unmapped++; + if (collapse_exceeds_limit(unmapped, max_ptes_swap, + start, end)) { + result =3D SCAN_EXCEED_SWAP_PTE; + count_vm_event(THP_SCAN_EXCEED_SWAP_PTE); + count_mthp_stat(HPAGE_PMD_ORDER, + MTHP_STAT_COLLAPSE_EXCEED_SWAP); + goto out_table_refused; + } + /* Swap entries armed with uffd-wp are refused too */ + if (pte_swp_uffd_any(pteval) && + pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PTE_UFFD; + continue; + } + if (pte_uffd(pteval)) { + /* + * The huge PMD could be marked write protected when any + * of the small ones is, but that could deliver + * userfaults outside the registered range. Keep it + * simple and refuse the PTE. + */ + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PTE_UFFD; + continue; + } + + page =3D vm_normal_page(vma, addr, pteval); + if (unlikely(!page) || unlikely(is_zone_device_page(page))) { + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PAGE_NULL; + continue; + } + folio =3D page_folio(page); + + /* + * A VM_DROPPABLE VMA keeps the lazyfree property across the + * collapse, so there is nothing to preserve by skipping. + */ + if (cc->policy.skip_lazyfree && + !(vma->vm_flags & VM_DROPPABLE) && + folio_test_lazyfree(folio) && !pte_dirty(pteval)) { + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PAGE_LAZYFREE; + continue; + } + + if (!folio_test_anon(folio)) { + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PAGE_ANON; + continue; + } + + /* + * A page counts as shared if any part of its folio is, which + * bounds the cost of CoW-breaking rather than the count of it: + * collapse_faultin() unshares on !PageAnonExclusive(), a broader + * test -- a page whose fork co-mapper has exited is + * single-mapped, so not counted here, yet stays non-exclusive + * until a write reuses it. Those are the cheap ones, reused in + * place. A page that has to be copied is one this test catches, + * so the limit does bound the copying it is there to bound. + */ + if (folio_maybe_mapped_shared(folio)) { + shared++; + if (collapse_exceeds_limit(shared, max_ptes_shared, + start, end)) { + result =3D SCAN_EXCEED_SHARED_PTE; + count_vm_event(THP_SCAN_EXCEED_SHARED_PTE); + count_mthp_stat(HPAGE_PMD_ORDER, + MTHP_STAT_COLLAPSE_EXCEED_SHARED); + goto out_table_refused; + } + } + + /* + * Which node the sources are on decides where the destination is + * allocated: the one with the most of them wins. + */ + node =3D folio_nid(folio); + if (collapse_scan_abort(node, cc)) { + result =3D SCAN_SCAN_ABORT; + goto out_table_refused; + } + cc->node_load[node]++; + + /* + * Usually a folio somebody else is already isolating, whose + * reference the freeze would refuse anyway. Not exact: one + * still on a per-CPU add batch reads the same, and the freeze + * drains those before it starts. + */ + if (!folio_test_lru(folio)) { + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PAGE_LRU; + continue; + } + if (folio_test_locked(folio)) { + if (pmd_result =3D=3D SCAN_SUCCEED) + pmd_result =3D SCAN_PAGE_LOCK; + continue; + } + + /* + * A folio whose reference count its mappings do not account for + * -- a GUP pin, say -- is refused by the freeze, not here. + * folio_expected_ref_count() wants a folio that cannot change + * order while it is read, and this walk holds no page table lock + * and no folio lock, so a folio splitting underneath it would + * have the count read for the wrong size. A reference of our + * own would not help: it stops the folio being freed, not split. + * + * So leave it to the freeze, which reads the table under the + * lock and settles the question by freezing each source to the + * count it expects. What it costs is a window selected here and + * refused there. + */ + + /* + * Every check passed: this PTE can be a collapse source. The + * bit is set last, so a disqualified PTE leaves it clear. + */ + __set_bit(i, cc->eligible_ptes); + + /* + * Whether a range has to look used at all is the caller's + * policy, so only a caller that asks gathers the evidence. + */ + if (cc->policy.require_referenced && + (pte_young(pteval) || folio_test_young(folio) || + folio_test_referenced(folio) || + mmu_notifier_test_young(vma->vm_mm, addr))) + referenced++; + } + + if (cc->policy.require_referenced && + (!referenced || (unmapped && referenced < HPAGE_PMD_NR / 2))) + result =3D SCAN_LACK_REFERENCED_PAGE; + else + result =3D pmd_result; + pte_unmap(pte); + goto out; + +out_table_refused: + /* + * The table is refused as a unit -- a limit the whole range exceeds, or + * pages on nodes too distant for one folio to serve them all -- so no + * window inside it is eligible either. + */ + pte_unmap(pte); +out_no_table: + cc->select_orders =3D 0; +out: + /* + * A PMD candidate needs the whole table, so anything that disqualified a + * single PTE rules it out. Smaller windows that avoid the offending + * PTEs are still collapsible, so drop just that order and leave the rest + * to selection -- dropping it also lowers the order selection roots its + * windows at. MADV_COLLAPSE has no other order enabled, so it is left + * with none. + */ + if (result !=3D SCAN_SUCCEED) + cc->select_orders &=3D ~BIT(HPAGE_PMD_ORDER); + + return result; } =20 /* Everything a table is judged on starts empty for each table */ diff --git a/mm/collapse.h b/mm/collapse.h index e2af4c47cb60..ad88b91d9a72 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -130,5 +130,12 @@ unsigned long collapse_possible_orders(struct vm_area_= struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags); enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, unsigned long address, pmd_t **pmd); +bool collapse_scan_abort(int nid, struct collapse_control *cc); +unsigned int collapse_max_ptes_none(struct collapse_control *cc, + struct vm_area_struct *vma, unsigned int order); +unsigned int collapse_max_ptes_swap(struct collapse_control *cc, + unsigned int order); +unsigned int collapse_max_ptes_shared(struct collapse_control *cc, + unsigned int order); =20 #endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 26d25093260b..9823884a83c9 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -305,7 +305,7 @@ struct attribute_group khugepaged_attr_group =3D { * * Return: Maximum number of empty/shared zeropage PTEs for the collapse o= peration */ -static unsigned int collapse_max_ptes_none(struct collapse_control *cc, +unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order) { const unsigned int max_ptes_none =3D cc->policy.max_ptes_none; @@ -341,7 +341,7 @@ static unsigned int collapse_max_ptes_none(struct colla= pse_control *cc, * Return: Maximum number of PTEs that map shared anonymous pages for the * collapse operation */ -static unsigned int collapse_max_ptes_shared(struct collapse_control *cc, +unsigned int collapse_max_ptes_shared(struct collapse_control *cc, unsigned int order) { /* @@ -362,7 +362,7 @@ static unsigned int collapse_max_ptes_shared(struct col= lapse_control *cc, * Return: Maximum number of non-present PTEs or the maximum allowed non-p= resent * pagecache entries for the collapse operation. */ -static unsigned int collapse_max_ptes_swap(struct collapse_control *cc, +unsigned int collapse_max_ptes_swap(struct collapse_control *cc, unsigned int order) { /* @@ -934,7 +934,7 @@ static struct collapse_control khugepaged_collapse_cont= rol =3D { .is_khugepaged =3D true, }; =20 -static bool collapse_scan_abort(int nid, struct collapse_control *cc) +bool collapse_scan_abort(int nid, struct collapse_control *cc) { int i; =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 312A13E9C03; Sun, 16 Aug 2026 22:46:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920398; cv=none; b=LxXgc484XGgFzlmSUlhKnNJjL0k0M6doZBCtd2BD/oiQksljpEHaJpLBc9CiAM10NbWSV1CpY+q0/9rixDZjvaVtw2Zi3czZxta2zpN76/B0sZ/BuDYvGHpCt/h1fg/4/fC6cUNqijQ/8BagNV3GSzkpbFBJbA2GGhL5Scvk0kQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920398; c=relaxed/simple; bh=DCEgpTVokzjFZ/ojFLnJ5WWKdCp2LPXg/Sv0MouFvy8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pG38PoWcQ2Gr/kXSUR2hA1i4asPavjypcLWlvjpsURRCp0pz8yiYHbrfzB4Y2v8ffFfVHELMZNkHgNdg/DEPL0zAupwHO7UihePxtDLazJu/JRLAN+k0WYFA7GO+oAVQW+o6Z+n4yzNnAxHzjPUtXRWLlgnXx0i1AMqOegkooUk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=bY7a2Sw+; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=K4YmaUyL; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="bY7a2Sw+"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="K4YmaUyL" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.phl.internal (Postfix) with ESMTP id 3AC3C14000FA; Sun, 16 Aug 2026 18:46:35 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:46:35 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920395; x= 1787006795; bh=CN8cWHp7kEAssTf3TdEJt2sw3VugR1AhfU4UisSCp8o=; b=b Y7a2Sw+PAG1xudQsxe9IGkUnqymZoX2Cq1eNOP04Yy2NUZ+C+Wkgvw/TTYv+X9L0 bTIPxLBj3YSQKI5wPX0UCTW3kvRQokWtIAsuKRHmG7/f4nMrA3rowp1SejoQY3vq vL7R2RuJwZ9Mg/y17i0ueFGqj+/Iqbe/SnP6LNkN9DgyPO0cKtNkfBEpUQ4cZNRW CJzYSTvOciq2OCRW+yxmvYuz8ybnlpOHoek71IhEV8J9346OIytu75F4P5qdWvVW bvWfqbE+G4Or6OM4bUrGjDLwZsnDudWwokkRvnB6V/cf1iFF33buH6AePHJftuYj nQUYimQIlApAcnnuZYJBQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920395; x=1787006795; bh=C N8cWHp7kEAssTf3TdEJt2sw3VugR1AhfU4UisSCp8o=; b=K4YmaUyLkLOHPK1ix 7zo/Kse1Y5W5wYkKJ34DydDerQpHAfbawbISXBZxhWMSECKn0jhy8LH+rB/uNK8x bpN3HH27CyxXBncBUyw3BPPv2isG7iRAlKurw00e2BqR721uJF3YNXSaifZHU8r7 r4xNshww1hWzdQNICdHdksD1Bn+OM6hsZVryj2vDYDBAV7RXTKkPiHct54oOj7Tc XbWqalJr3gE2mJak39cwbwhmdjFUuxsgBhDN84kzYryeLtlxznHabxambeIcjCDh 04jJDZByaZf04YXGHIp5PkvJXrAFRCm66/VVj0oYWWCk4gg7LMsT+8wooFRykqh2 MC4Xg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtgo 5zg/3qDFuCLK/R4oQt+pvC0QjgMhDXb5EXIIFWtX7kNVP0YSW/7JBU5omCjUzM7elgoUGw UYgSR9+whjAWV+aVhXHXy0vKt61g56d+GtkunQxxbCr1+YORbs4nmAAJkjtQwjXlNP9QB0 tUUv+dhhvQmjNmt2ZN67pTE/IPOoWUHNvjAWz0p4ASmnGGj/aQTyXNnbyqDa8qYzTzSG6S Kn5NwNoH7Lf0tH8MPzHEvXuZxRUcX/tRsg0aT0y6dmAkkoI3BRUor7gvB/8xwSQP7qXPQX CxoSG0dqiVxgJLqTEM9syNbIXyhJihGbchjHU3rbyaRL2T/i8tCDOzBcitiA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:34 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 09/57] mm/collapse: collect candidate windows into a round Date: Sun, 16 Aug 2026 23:45:21 +0100 Message-ID: <20260816224609.308019-10-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_anon_pmd() is the half of a table's collapse that follows the scan: cut windows out of the PTEs the scan accepted, and run them. It runs them a round at a time, so a round needs somewhere to be collected. Add that array to collapse_control. Its size is the number of windows one table holds at the smallest order a collapse builds, or as many as the byte cap allows, whichever is fewer. A round is capped because it holds destination folios that are allocated but not yet installed, and because a faulter on any source inside it waits for the round to finish. A dense table is collapsed as several rounds rather than one. The array is too large for the stack. khugepaged takes it when the daemon starts, so an allocation failure is reported to the sysfs write that enabled khugepaged rather than surfacing inside the daemon; MADV_COLLAPSE takes one per call. collapse_anon_pmd() then gets its shape. Take candidates from selection until the round is full or selection is done, run the round, and stop once selection has nothing left and the round is empty. A candidate the full round could not take stays pending for the next one, so nothing is dropped at the boundary. Selection and the batch run are stubs here, so the loop collects nothing and the range still yields nothing. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 156 +++++++++++++++++++++++++++++++++++++++++++++++- mm/collapse.h | 9 +++ mm/khugepaged.c | 24 +++++++- 3 files changed, 185 insertions(+), 4 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 66931ef6a6d0..6dae5e35e61d 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -86,6 +86,52 @@ * replaces, and is switched over once both halves are complete. */ =20 +/* + * Cap on the memory a round may hold in flight: destination folios alloca= ted + * but not yet installed, the fault latency of anything inside a candidate= being + * collapsed, and memcg charge pressure all scale with it. A dense table = is + * collapsed as several rounds rather than one. + */ +#define COLLAPSE_BATCH_BYTES SZ_32M + +/* Windows in one table at the finest order collapse cuts */ +#define COLLAPSE_TABLE_WINDOWS (HPAGE_PMD_NR >> COLLAPSE_MIN_MTHP_ORDER) + +/* + * How many candidates a round can hold, fixed by the table geometry: the = byte + * cap decides it at the smallest order collapse builds, but never more th= an the + * windows one table has at that order. + */ +#define COLLAPSE_MAX_CANDIDATES \ + min(COLLAPSE_BATCH_BYTES >> (PAGE_SHIFT + COLLAPSE_MIN_MTHP_ORDER), \ + COLLAPSE_TABLE_WINDOWS) + +/* + * A candidate is an (addr, order) window selected for collapse. Selection + * counts in PTE offsets -- the bitmap it reads and the alignment it honou= rs are + * indexed that way -- while the passes that run a candidate work in addre= sses, + * like the page tables and VMAs they touch. This is where the two meet. + */ +struct collapse_candidate { + unsigned long addr; + unsigned int order; +}; + +void collapse_control_release(struct collapse_control *cc) +{ + kfree(cc->candidates); + cc->candidates =3D NULL; +} + +int collapse_control_init(struct collapse_control *cc) +{ + cc->nr_candidates =3D 0; + cc->candidates =3D kmalloc_objs(*cc->candidates, COLLAPSE_MAX_CANDIDATES); + if (!cc->candidates) + return -ENOMEM; + return 0; +} + /* * Is @count past a limit stated per PMD, when only part of a table was sc= anned? * Scale the comparison to the table so a partial scan is held to the same @@ -389,6 +435,75 @@ collapse_scan_anon_pmd(struct vm_area_struct *vma, uns= igned long start, return cc->scan_refusal; } =20 +/* Point the selection cursor at [start, end) of the table, in PTE offsets= */ +static void collapse_selection_init(struct collapse_control *cc, + unsigned int start, unsigned int end) +{ +} + +/* + * The next window worth attempting, as an (offset, order) pair. False wh= en + * selection is exhausted, which is what ends the range. + * + * A candidate is only ever an (offset, order) pair: the scan that recorde= d the + * eligible PTEs has dropped the ptl, so anything else -- folio pointers in + * particular -- would be stale by construction. + */ +static bool collapse_next_candidate(struct collapse_control *cc, + unsigned int *offset, unsigned int *order) +{ + return false; +} + +/* + * Run and classify the collected batch. Returns false when a candidate's + * outcome abandons the table. + */ +static bool collapse_run_batch(struct mm_struct *mm, unsigned long pmd_add= r, + struct collapse_control *cc) +{ + /* collapse_anon_pmd() only runs a round it has put something in */ + VM_WARN_ON_ONCE(!cc->nr_candidates); + + cc->nr_candidates =3D 0; + return true; +} + +/* + * One more candidate of @order would either overflow the array or push wh= at the + * round holds past the byte cap. An empty round takes whatever it is off= ered: + * a single candidate is above the cap all by itself once a PMD is (512M w= ith + * 64K pages), and refusing it would collapse nothing at all. + */ +static bool collapse_batch_full(struct collapse_control *cc, + unsigned long bytes, unsigned int order) +{ + if (!cc->nr_candidates) + return false; + + return cc->nr_candidates =3D=3D COLLAPSE_MAX_CANDIDATES || + bytes + (PAGE_SIZE << order) > COLLAPSE_BATCH_BYTES; +} + +/* + * Take the next array slot for the window at @addr. A slot may still hol= d a + * previous round's values, so every field is set here. + */ +static void collapse_add_candidate(struct collapse_control *cc, + unsigned long addr, unsigned int order) +{ + struct collapse_candidate *cand; + + /* collapse_batch_full() has already made room */ + if (WARN_ON_ONCE(cc->nr_candidates >=3D COLLAPSE_MAX_CANDIDATES)) + return; + + cand =3D &cc->candidates[cc->nr_candidates]; + cc->nr_candidates++; + cand->addr =3D addr; + cand->order =3D order; +} + /* * Cut the table into candidate windows and collapse what fits, from the * largest order downwards. Returns what the table yielded: a collapse, or @@ -398,5 +513,44 @@ static enum scan_result __maybe_unused collapse_anon_pmd(struct mm_struct *mm, unsigned long start, unsigned long= end, struct collapse_control *cc) { - return SCAN_FAIL; + const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; + unsigned int offset, order; + unsigned long bytes =3D 0; + bool pending =3D false; + bool cont =3D true; + + collapse_selection_init(cc, (start - pmd_addr) >> PAGE_SHIFT, + (end - pmd_addr) >> PAGE_SHIFT); + + while (cont) { + if (!pending) + pending =3D collapse_next_candidate(cc, &offset, &order); + + if (!pending || collapse_batch_full(cc, bytes, order)) { + /* + * Selection is exhausted and the round is empty: the + * range is done. Without this a flush of an empty + * round would return, collect nothing, and come + * straight back here. + */ + if (!cc->nr_candidates) + break; + + cont =3D collapse_run_batch(mm, pmd_addr, cc); + bytes =3D 0; + continue; + } + + /* + * The round holds no resources until it is run, so + * collecting costs nothing but the array slot. A candidate the + * full round could not take is kept pending for the next one. + */ + collapse_add_candidate(cc, pmd_addr + offset * PAGE_SIZE, order); + + bytes +=3D PAGE_SIZE << order; + pending =3D false; + } + + return cc->nr_collapsed ? SCAN_SUCCEED : SCAN_FAIL; } diff --git a/mm/collapse.h b/mm/collapse.h index ad88b91d9a72..1159ed39b9eb 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -10,6 +10,8 @@ /* The smallest order a collapse will build, and so the finest window it c= uts */ #define COLLAPSE_MIN_MTHP_ORDER 2 =20 +struct collapse_candidate; + enum scan_result { SCAN_FAIL, SCAN_SUCCEED, @@ -120,8 +122,15 @@ struct collapse_control { * the collapse reports it when it salvages nothing. */ enum scan_result scan_refusal; + + /* The candidate windows collected for the current round */ + struct collapse_candidate *candidates; + unsigned int nr_candidates; }; =20 +int collapse_control_init(struct collapse_control *cc); +void collapse_control_release(struct collapse_control *cc); + /* * Defined in khugepaged.c, which still uses them itself. * TODO: move each into collapse.c once its last khugepaged.c user is gone. diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 9823884a83c9..43f6107c953a 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -3079,12 +3079,22 @@ int start_stop_khugepaged(void) guard(mutex)(&khugepaged_mutex); if (hugepage_enabled()) { if (!khugepaged_thread) { - struct task_struct *new_thread =3D kthread_run(khugepaged, - NULL, - "khugepaged"); + struct task_struct *new_thread; + int err; =20 + /* + * The engine collapses out of its candidate array, so + * take it before starting the thread that needs it: a + * failure surfaces here rather than in the daemon. + */ + err =3D collapse_control_init(&khugepaged_collapse_control); + if (err) + return err; + + new_thread =3D kthread_run(khugepaged, NULL, "khugepaged"); if (IS_ERR(new_thread)) { pr_err("khugepaged: kthread_run(khugepaged) failed\n"); + collapse_control_release(&khugepaged_collapse_control); return PTR_ERR(new_thread); } =20 @@ -3096,6 +3106,7 @@ int start_stop_khugepaged(void) } else if (khugepaged_thread) { kthread_stop(khugepaged_thread); khugepaged_thread =3D NULL; + collapse_control_release(&khugepaged_collapse_control); } set_recommended_min_free_kbytes(); return 0; @@ -3154,6 +3165,7 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, enum scan_result last_fail =3D SCAN_FAIL; int thps =3D 0; bool mmap_unlocked =3D false; + int err; =20 BUG_ON(vma->vm_start > start); BUG_ON(vma->vm_end < end); @@ -3173,6 +3185,11 @@ int madvise_collapse(struct vm_area_struct *vma, uns= igned long start, cc->is_khugepaged =3D false; collapse_policy_forced(&cc->policy); cc->progress =3D 0; + err =3D collapse_control_init(cc); + if (err) { + kfree(cc); + return err; + } =20 mmgrab(mm); lru_add_drain_all(); @@ -3231,6 +3248,7 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, out_nolock: mmap_assert_locked(mm); mmdrop(mm); + collapse_control_release(cc); kfree(cc); =20 return thps =3D=3D ((hend - hstart) >> HPAGE_PMD_SHIFT) ? 0 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2A9023E5EC0; Sun, 16 Aug 2026 22:46:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920399; cv=none; b=lRzLyFKv13CuKQyrq+Q4Rgyb4BQvd+CLWR1+4+SnNMhu+gQF0MUFnYpREaxRwnrDkE53nG5waIL7Vqkxfk5PkDtRpnE+zZCkgBy2QbsU9hSUEJFA2KrfQYsECB5MeaHNgvme9Cr145mXV4s206stk+E2NH8icikWEqP2sUn+Tn8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920399; c=relaxed/simple; bh=ZKykw6oGV9m9mVc5sFT/N3XDmLaiZXCrjH8hXh72Hio=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Q7rGFp1WSdPoHhAw7ZMWkTuWQCU6jhwMs1UE2UakkAm8mHswvvVB27QHXuEGPCpglWrzbZ/ezd22tKQ9LQN9Q3tVP0QWcBHPcUwjVMB5KoTTXAWCyJLXJt/XAmoCStRlRkV5dDgNZntAYiodsHxh7IL3CH3P/RMeEqP/ce9tPD4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=mc1G/v9u; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=fOTtMaWM; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="mc1G/v9u"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="fOTtMaWM" Received: from phl-compute-07.internal (phl-compute-07.internal [10.202.2.47]) by mailfout.phl.internal (Postfix) with ESMTP id 66121EC0244; Sun, 16 Aug 2026 18:46:37 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-07.internal (MEProxy); Sun, 16 Aug 2026 18:46:37 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920397; x= 1787006797; bh=6GvdX/z9pV4lxPc5SRlNT52Nk8AS9W2M2/m8d2d4Is4=; b=m c1G/v9uXgZYFsICqr+QjexNh2FmtgLrSlr6dPVqX0XI3dImjsBSP1R60zY41Eoaj 5aRJnqRUNOV/2Ws+I0YNMdLWWNFMBoPavhEa15M9Reqq/cBAJGzVVG7BANQlctcQ Ei887na/bVCIyuCv54DNliN85QSOUnoY7nIhSy1UhtPPhpd0InWAtRNcifvMSAJg ZLvl/+nZtAa4Cl1H6KXdQ9VzpdYdCAR1KFN3j1aGJITx2ArNKl3lxRQTEUQNKtfO 8RbvJLiXrVKTED9/bZB4RdM9DAx0G7tYye4mch1AH6wtvm9UzuJZ0IUXtrJaDQSC VX7nN2uk7kO5tu6QZ+K5g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920397; x=1787006797; bh=6 GvdX/z9pV4lxPc5SRlNT52Nk8AS9W2M2/m8d2d4Is4=; b=fOTtMaWM84K4Muz6r PzZ0Lal+gYVeumdOkmn7iZovsSiLzZPJVthtDVYytfcbi+dZBHqz8BmCdN+qRzeK r2PEUUMbakRN8e97AVEPFbRJ5ledgJlT52d4v2BIYtIrX/jeQqmlAYLBdSGer+CC VzS+09RzmGYz4oB772oWkJqDqqZryZxDXwpzDRNYIVKphAA38rVJM+QENVEFMDpE 1TrYhQ6vDcoj1KaRuRkVbkGDG5ZJY0xWeY28Yw+GhsK3mgH3NkXln9+fJfkK2Tb/ pdBY437ADNxHaltx70MYrrHfelsGTRoy9Cj0ZEI6SRkXE9H8ojAaYfH+ZXmtfX96 x9CvA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtt2 jeK1puiUf8/lXHw9run/+wq4ZMkal3NqfCGVfHNZQeqzajSoM0xs91hLNdfkGJVMEPEo50 EPe+Ncx4BTBm7MyOsoX0Yv+7dODbPcL1P8bwxI8Ww0Vd+OJKkNSE+UeUiLjokmOTqE3x/g NVb6tth4Fc2XxFG8ViDfPqUuCtPt+2/drARpFjlBg0GZTT3TNpKGPRy/wAl24096Jq8fMk 0IMtaDT4p2a5c8LGBnQHrXa+BH7i8dsceJq7w4HtmfDRlPOGvjk72Mh1kFWRA+NJT+Ycq8 hlwNa62Gk1NaJ1js+PYaOu4IOTxXbZOkzWhgmhlQmOvIcjzJo6//3LBCCChg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:36 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 10/57] mm/collapse: run a round and feed the outcomes back Date: Sun, 16 Aug 2026 23:45:22 +0100 Message-ID: <20260816224609.308019-11-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A candidate a round attempts either collapsed or did not, and if it did not there is a reason. Selection needs those outcomes to decide what comes next: carry on past the window, try the same region at a lower order, or give the table up. Fill in collapse_run_batch(): run the round, then walk the batch handing each candidate's result to classification. The walk covers the whole batch. A pass that refuses one candidate marks it and carries on rather than truncating the round, so every candidate has a result of its own to hand back. Only an outcome that condemns the table cuts the walk short, and then nothing of that table re-enters selection. The round and the classification it feeds are both stubs, so nothing is attempted and nothing is decided. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 53 insertions(+) diff --git a/mm/collapse.c b/mm/collapse.c index 6dae5e35e61d..ad9e5a447854 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -115,8 +115,16 @@ struct collapse_candidate { unsigned long addr; unsigned int order; + enum scan_result result; }; =20 +/* Where a candidate sits in the table, in the PTE offsets selection count= s in */ +static unsigned int candidate_offset(const struct collapse_candidate *cand, + unsigned long pmd_addr) +{ + return (cand->addr - pmd_addr) >> PAGE_SHIFT; +} + void collapse_control_release(struct collapse_control *cc) { kfree(cc->candidates); @@ -132,6 +140,17 @@ int collapse_control_init(struct collapse_control *cc) return 0; } =20 +/* + * Carry one batch of candidates through the passes. Every candidate come= s back + * with a result of its own: the passes before the freeze mark what they r= efuse + * and carry on, each pass after it works on what the last left, so no fai= lure + * truncates the round. + */ +static void collapse_round(struct mm_struct *mm, unsigned long pmd_addr, + struct collapse_control *cc) +{ +} + /* * Is @count past a limit stated per PMD, when only part of a table was sc= anned? * Scale the comparison to the table so a partial scan is held to the same @@ -455,6 +474,18 @@ static bool collapse_next_candidate(struct collapse_co= ntrol *cc, return false; } =20 +/* + * Feed one candidate's outcome back into selection: its region is done, it + * re-enters the retry store at a lower order, or the table is abandoned. + * Returns false in that last case. + */ +static bool collapse_classify_result(struct collapse_control *cc, + unsigned int offset, unsigned int order, + enum scan_result result) +{ + return true; +} + /* * Run and classify the collected batch. Returns false when a candidate's * outcome abandons the table. @@ -462,9 +493,30 @@ static bool collapse_next_candidate(struct collapse_co= ntrol *cc, static bool collapse_run_batch(struct mm_struct *mm, unsigned long pmd_add= r, struct collapse_control *cc) { + unsigned int i; + /* collapse_anon_pmd() only runs a round it has put something in */ VM_WARN_ON_ONCE(!cc->nr_candidates); =20 + collapse_round(mm, pmd_addr, cc); + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + unsigned int offset =3D candidate_offset(cand, pmd_addr); + + if (!collapse_classify_result(cc, offset, cand->order, + cand->result)) { + /* + * The table is abandoned: the candidates behind this one + * keep their results and are left unclassified, so + * nothing more of this table enters selection, and the + * abandoning result clears what earlier ones left there. + */ + cc->nr_candidates =3D 0; + return false; + } + } + cc->nr_candidates =3D 0; return true; } @@ -502,6 +554,7 @@ static void collapse_add_candidate(struct collapse_cont= rol *cc, cc->nr_candidates++; cand->addr =3D addr; cand->order =3D order; + cand->result =3D SCAN_FAIL; } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B757A3E5EF2; Sun, 16 Aug 2026 22:46:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920402; cv=none; b=aoiU+uy7JWTwUsmsw4v+p+g6/RTo4MthNY0fmO/PTh+O7991SuX3Lmf5Z0TrYbfnvMinycwpewj5ZGSrcH1j1j3cGexqfQghboi8OukXIf+fbOTlWiE5yMgMbbGDUy33w4los2WdZPQYBTeh8tmnlHvetpzVZpLAO+C1S3pL/qY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920402; c=relaxed/simple; bh=+/hcrqZo62vBvoTQQ5P0qEK7WD6KqidCD5edv4xqgzI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dj/7jnztwcNl7InUMa6OLgo++55CGj5tz8WrOzX2VWk3HMetlDUEJLHeLC7pZ2LwEAd4U+DcGKYGYFzIazyV8TaszdNKt9iNmxH+SfYaPvATrPn6fvpGEziIgfJKsNu4pKBotQ1d/u1/ZWH9cv5+HJUl/vIslQhi18cdAbcxSts= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=CifAacI5; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=ke+Pz5jQ; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="CifAacI5"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="ke+Pz5jQ" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.phl.internal (Postfix) with ESMTP id ED44C14000F8; Sun, 16 Aug 2026 18:46:39 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:46:39 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920399; x= 1787006799; bh=J+VwmvOWkQQIRGq2yrrL0lOcYIpS4iEV3Q9I35NYUqY=; b=C ifAacI5NLADSIgETp9aUgPDTj/HPjrbGohJHGfsGwfpEo9KdvcvoAoUBdqDMvbB4 xBHuRNjkopP/EfuEPaypBjHIDKETtXEZAGyqd/q+tX4bYiC1RJQ3JGUZIco+2ufd LISfGTTlwLXNr0LahSfa4Z2RExpcVUG1z7jqK7IvCb71plD0WLOoyd34xfXrQhZX UptGTHX09vbD78EAXlP7ZFgPNQT1fyWATp/HH5BnewSLIktm5sz+zG/zj3t7IaGB PCDwf4Blk2wZ1K4I41MgOP4zeOkq7rpc8JklO9EBWEFVWLemaObCzQUuUVkAMsrq 1iAjbCB9Sj2YxPLkrWKmQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920399; x=1787006799; bh=J +VwmvOWkQQIRGq2yrrL0lOcYIpS4iEV3Q9I35NYUqY=; b=ke+Pz5jQQ1F8Pj6qK 5LqN1WRiReWrEOxswy3084QMQEc9S9UOGbNnSgNbiMuZprqcu4Vk4ZZ6DxDtxzY6 i+DCfT2KI3lwdO9QteVxv9V2FVnE/mf2eN633nRA/D2qMUJfAHzEzCxLH902x/OW LtNDJlCrqLIdltJ4HqofveFmukXb6qwMzdOWG4iSfOKQPeZ4CbDkvwpeOdA+yPk5 inw2Oi4Y29mKPx+oUTe6bjDrgO6MBirWz1ZYNT7fmkO0sdDyQ0lagVaL1veqUKC5 VDJCnJsn7boa9sormKAoDDs4xq+ik4MhMWHr1+Q38BuJsrCeaOjc/d7NvZbdy27/ FWFLw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTE1oit9MZbgdPNUA4HHWsMULK6Cz5sdpDL6Wk/wD7kfNXGu3e/EXf5M9JXawVpDPI Kd9lJ1GZVafs565rM9S+cZw+n/cMd1S9oSxi3Xv7RrTuiH9afQzGtYFTyBhKsmYEaht9J5 vV7dyNSnAuCNScX81AQnXoyHXD4pg8pag19vnm592IJQcIf6BATwad+RD51AIfid8gba17 gKSwUtsEMegipv8zNYIHYPR4e68lUdiPrbHwyMm74ZzHlAmAfr/QUkf9Bs0a0GPARTfnrV SpFrZ9EiV9esVNjz/LhavbueoSNqwHuOcKcj3vQwMSxA4okLdcH8vRa2PTuYlMcoV2XUMw iPTPisQJJ3aMD2hQvaOlWvC36TobhKHzVMIEx3pDiySZgO1/07ljaQzRO6Zpje+xuTKxkE ZKzA3xepHraairMh3zaBFf4LK64p7T+TbqDRH2+EsmE/ss3KoXhGPe2C/P+SzYzKECJzBk sqx70BhANkCN6E47RQvptQNsJ606DlBKGaRnIAUe3hEbKuOipgIW7nC5bMTZkAZA1+KGYa da+N4OnGE3uVHiChxofFWQqm4T/uX3/79S7l1N9TBgu5F6fBvk9VO8iNUMlDDL8G0RzFNf hoa4Si9vWNYj/p4zzPWh9da7jH4bUOZQtdWKOMHfFgYp2+E7s2pLchRukpYA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:39 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 11/57] mm/collapse: sketch the passes of a round Date: Sun, 16 Aug 2026 23:45:23 +0100 Message-ID: <20260816224609.308019-12-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A round is a sequence of passes over the same batch, and their order is most of the design. It falls into three parts: - before any lock, the allocations that may sleep: destination folios, and the page table a PMD-order candidate deposits; - under mmap_read, revalidation and fault-in, which may have to give the lock up; - from the freeze onwards, a stretch that has to run to completion. Each pass in that last part works on what the one before it left, and every barrier raised has to be lowered again. The destinations still missing are asked for there too, without reclaim: a faulter on a frozen source would wait for the allocation. Lay that sequence out, with every pass a stub but one. collapse_round() takes mmap_read for the middle of it: a collapse is called without the lock, and takes its own for each round. It brackets the frozen window in one mmu-notifier invalidate over the whole batch. That invalidate needs a span before any pass has a body, so collapse_revalidate() settles it from the start, in cc->batch_start and cc->batch_end. The span is taken over the candidates rather than off the ends of the array: a region refused at one order can re-enter selection at a lower one, so a round is not address-ordered and candidates[0] need not be the lowest. A fault-in that had to sleep comes back with the lock dropped, reported as SCAN_LOCK_DROPPED. The round takes the lock again and runs the pass afresh, a bounded number of times, rather than sending the batch back to selection. The stubs do nothing, so the round does nothing. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 1 + mm/collapse.c | 207 +++++++++++++++++++++++++++++ mm/collapse.h | 9 ++ 3 files changed, 217 insertions(+) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 5a48c5406cce..778f5a56956c 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -24,6 +24,7 @@ EM( SCAN_PAGE_COUNT, "not_suitable_page_count") \ EM( SCAN_PAGE_LRU, "page_not_in_lru") \ EM( SCAN_PAGE_LOCK, "page_locked") \ + EM( SCAN_LOCK_DROPPED, "lock_dropped") \ EM( SCAN_PAGE_ANON, "page_not_anon") \ EM( SCAN_PAGE_LAZYFREE, "page_lazyfree") \ EM( SCAN_PAGE_COMPOUND, "page_compound") \ diff --git a/mm/collapse.c b/mm/collapse.c index ad9e5a447854..25c0f72a9a68 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -94,6 +94,14 @@ */ #define COLLAPSE_BATCH_BYTES SZ_32M =20 +/* + * How many times a round runs the fault-in pass. A fault that has to wai= t drops + * the lock, and running the pass again costs a walk of the batch but buys= at + * least one completed swap-in; readahead brings a cluster in at a time, s= o this + * covers a PMD's default max_ptes_swap. + */ +#define COLLAPSE_FAULTIN_PASSES 8 + /* Windows in one table at the finest order collapse cuts */ #define COLLAPSE_TABLE_WINDOWS (HPAGE_PMD_NR >> COLLAPSE_MIN_MTHP_ORDER) =20 @@ -118,6 +126,21 @@ struct collapse_candidate { enum scan_result result; }; =20 +static unsigned long candidate_start(const struct collapse_candidate *cand) +{ + return cand->addr; +} + +static unsigned long candidate_size(const struct collapse_candidate *cand) +{ + return PAGE_SIZE << cand->order; +} + +static unsigned long candidate_end(const struct collapse_candidate *cand) +{ + return candidate_start(cand) + candidate_size(cand); +} + /* Where a candidate sits in the table, in the PTE offsets selection count= s in */ static unsigned int candidate_offset(const struct collapse_candidate *cand, unsigned long pmd_addr) @@ -140,6 +163,134 @@ int collapse_control_init(struct collapse_control *cc) return 0; } =20 +/* + * The scan and the allocation both dropped mmap_lock, so nothing seen bef= ore it + * can be trusted: find the VMA and the PTE table again, and check they st= ill + * allow every provisioned candidate. + * + * This is also where the batch's span is settled, for the invalidate the = round + * issues over it. + */ +static enum scan_result collapse_revalidate(struct vm_area_struct *vma, + unsigned long pmd_addr, + struct collapse_control *cc, + pmd_t **pmdp) +{ + unsigned int i; + + cc->batch_start =3D ULONG_MAX; + cc->batch_end =3D 0; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + + cc->batch_start =3D min(cc->batch_start, candidate_start(cand)); + cc->batch_end =3D max(cc->batch_end, candidate_end(cand)); + } + + return SCAN_SUCCEED; +} + +/* + * Make every source the round needs present and exclusively owned by this= mm, + * by faulting it in as an ordinary access would. Sleeps, and drops mmap_= lock on + * failure, since a fault may have to be retried with it released. + * + * Anything faulted in lands on a per-CPU LRU batch, holding a reference t= he + * freeze cannot account for, so the freeze drains those batches before it + * starts. + */ +static enum scan_result collapse_faultin(struct vm_area_struct *vma, + struct collapse_control *cc, + pmd_t *pmd) +{ + return SCAN_SUCCEED; +} + +/* + * Raise the two barriers on the sources of every candidate: migration ent= ries in + * their PTEs, then a frozen refcount. Takes the table's ptl once for the= whole + * batch, and flushes the TLB once before dropping it. A candidate whose = sources + * moved is dropped here. + */ +static void collapse_freeze(struct vm_area_struct *vma, + struct collapse_control *cc, pmd_t *pmd) +{ +} + +/* + * Allocate ahead of the freeze for the candidates whose light allocation = missed + * last round. This is where reclaim belongs: nothing is held or frozen, = so a + * long compaction costs only khugepaged's own progress -- which is why the + * mechanism this replaces allocated here too. Having asked the allocator= to try + * hard, a miss now is a failure. + */ +static void collapse_reserve(struct mm_struct *mm, struct collapse_control= *cc) +{ +} + +/* + * Secure the page table the PMD terminal layer deposits. This stays ahea= d of the + * freeze because pte_alloc_one() allocates with GFP_PGTABLE_USER and take= s no gfp + * to strip: order-0 or not, it may reclaim and sleep, which is what the w= indow + * exists to keep out. The destination folio has a light gfp to fall back= on and + * so can be deferred; this has none. + */ +static void collapse_deposit(struct mm_struct *mm, struct collapse_control= *cc) +{ +} + +/* + * Give the frozen candidates that still need one a destination folio, wit= hout + * reclaim: a faulter on their sources would wait for it. + * + * A miss here is not a failure, as long as a retry could do better: the + * candidate keeps its freeze and asks for the reclaiming gfp, which + * collapse_reserve() uses before the next round freezes anything. When t= he + * policy forbids reclaim there is nothing better to retry with, so the mi= ss is + * the answer, and a smaller order over the same region is the better next= move. + */ +static void collapse_provision(struct mm_struct *mm, + struct collapse_control *cc) +{ +} + +/* + * Copy the frozen sources into their destinations. Nothing can reach eit= her + * side, so this needs no page-table lock, and it sleeps. + */ +static void collapse_copy(struct vm_area_struct *vma, + struct collapse_control *cc) +{ +} + +/* Publish each destination folio in place of the sources it replaces */ +static void collapse_install(struct vm_area_struct *vma, + struct collapse_control *cc, pmd_t *pmd) +{ +} + +/* + * Lower the barriers the freeze raised, on the sources of an installed ca= ndidate + * and on those of one that got no further. + */ +static void collapse_putback(struct vm_area_struct *vma, + struct collapse_control *cc) +{ +} + +/* + * Settle whatever the round reached: account what was installed, release = what + * was not, and give every candidate the result selection will classify. = Returns + * how many candidates were installed. + */ +static unsigned int collapse_finish(struct mm_struct *mm, + struct collapse_control *cc, + enum scan_result result) +{ + return 0; +} + /* * Carry one batch of candidates through the passes. Every candidate come= s back * with a result of its own: the passes before the freeze mark what they r= efuse @@ -149,6 +300,62 @@ int collapse_control_init(struct collapse_control *cc) static void collapse_round(struct mm_struct *mm, unsigned long pmd_addr, struct collapse_control *cc) { + unsigned int passes =3D COLLAPSE_FAULTIN_PASSES; + struct mmu_notifier_range range; + struct vm_area_struct *vma; + enum scan_result result; + pmd_t *pmd; + + collapse_reserve(mm, cc); + collapse_deposit(mm, cc); + +retry: + mmap_read_lock(mm); + + vma =3D find_vma(mm, pmd_addr); + if (!vma) { + result =3D SCAN_VMA_NULL; + goto out_unlock; + } + + result =3D collapse_revalidate(vma, pmd_addr, cc, &pmd); + if (result !=3D SCAN_SUCCEED) + goto out_unlock; + + result =3D collapse_faultin(vma, cc, pmd); + /* + * A fault dropped the lock to wait, as a swap-in does. The swap-in it + * started is still running and the walk skips whatever has arrived, so + * take the lock again rather than send the batch back to selection. The + * VMA and the table are looked up afresh: both may have changed. + */ + if (result =3D=3D SCAN_LOCK_DROPPED && --passes) + goto retry; + if (result !=3D SCAN_SUCCEED) + goto out; /* the callee released mmap_lock */ + + /* One invalidate window spans the batch, as collapse_revalidate() left i= t */ + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, + cc->batch_start, cc->batch_end); + mmu_notifier_invalidate_range_start(&range); + + /* + * None of these can fail as a whole: the freeze takes the sources it + * can and drops the candidates it cannot, and each pass after it works + * on what the one before left, so every barrier raised is lowered again. + */ + collapse_freeze(vma, cc, pmd); + collapse_provision(mm, cc); + collapse_copy(vma, cc); + collapse_install(vma, cc, pmd); + collapse_putback(vma, cc); + + mmu_notifier_invalidate_range_end(&range); + +out_unlock: + mmap_read_unlock(mm); +out: + collapse_finish(mm, cc, result); } =20 /* diff --git a/mm/collapse.h b/mm/collapse.h index 1159ed39b9eb..c61db86dc6c2 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -29,6 +29,7 @@ enum scan_result { SCAN_PAGE_COUNT, SCAN_PAGE_LRU, SCAN_PAGE_LOCK, + SCAN_LOCK_DROPPED, SCAN_PAGE_ANON, SCAN_PAGE_LAZYFREE, SCAN_PAGE_COMPOUND, @@ -126,6 +127,14 @@ struct collapse_control { /* The candidate windows collected for the current round */ struct collapse_candidate *candidates; unsigned int nr_candidates; + + /* + * What the candidates the round still means to freeze span, settled by + * collapse_revalidate() as it walks them. A round is not necessarily + * address-ordered, so this cannot be read off the ends of the array. + */ + unsigned long batch_start; + unsigned long batch_end; }; =20 int collapse_control_init(struct collapse_control *cc); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4132D3E51F7; Sun, 16 Aug 2026 22:46:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920404; cv=none; b=SBCZrkbLz7Bdp9jR06dvAuNrnhVmLckHzM6LIrS+bAcDqiaHjAYEh2vqGoWI4ZLs8u9o9ahiGptXU2BuqtS3V7LtTT7NxBaFoCBnRnQ/hzHGgg9s24MZMUCLzB7roMY0NeMCnwE1pj+SW3bq6Es6L3LmTj2yZaYHbOp5BPV6iZo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920404; c=relaxed/simple; bh=jomEwovj9b8k417t0IRW9Y93tpW80uP6vIda7Gb1KT4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o4UJyP4YMiCOV3wE5aqp4HnwnyQDyxszbTOg1VhNs4m5c9eLflIJl/SbpO4gaE/2JQyIhL9zLe6hYn3bNM4iIQu9uXwbB64dpOsxcfgV1pJRSmNrGFiVJMcaDCbHHafyVaplPyfAvIXHXEBV4Jf3lVmcFrBcs9VJ4VDxZ5J7Z/4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=dMgjTlrs; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=j7J0K0ur; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="dMgjTlrs"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="j7J0K0ur" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.phl.internal (Postfix) with ESMTP id A7F3714000FD; Sun, 16 Aug 2026 18:46:42 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:46:42 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920402; x= 1787006802; bh=NZOPrMQr1FnE5Kz+0VaDbhYd69V7nE+v+WGyJ8jLpqs=; b=d MgjTlrsA2TAN2Gz5Z208Bh7RwwxK4hCZ9Xp6qOXmN8dtzhGXvQ7fQZnjNFqzv5Jw kXfZKB2tY9yvRg7yblTrPv0aW6UT6t5F8gyVGkI8Si58SuFnE/tdQ5FrgBk76Fw7 cZhh+dVKWLjD0hQ/W0GQj/YgWw+zNCFJWSlSllQg5p6kxgDg3+d+xPxi2ua0+6C7 DCnp0KWYKAqNVMyrZFfrd6L9nknXvlIXL4i0Oqho9U5n2kibbaP4LNBmK0MOcFNw uQ4ipd8hFvpAiSo8yzX5ynoJdKckAoUKFa/oUo+tOxXfTX58LSFfg42o7bbTHtvl G5i6KeH+YbxggOiW3310A== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920402; x=1787006802; bh=N ZOPrMQr1FnE5Kz+0VaDbhYd69V7nE+v+WGyJ8jLpqs=; b=j7J0K0urCx9gvHOAE fuXPkD8y2DAoHplqbZBj3Aa3ZUsV6wBtLdv/eKXRz6ov2B+PWXyrNpMK9OySr0hz KrNDP1xSt/qEwgmlp+00k8C9B7tSByxiDN3zqS+mYIYmF1+wOpq0srdNOeDtXjrp DfRvz5/b9v33YzZELj6CrMA62mPobzpfhyzmOegLk0zNfziW9v2T8ilIjx+UDe+f cIIPHqmGih8w/bQKiIYQINbL2yeqkq47y+gsD8ZxlKIy3ZcjhhE/AnKDYQShGb0x G9Jofik0aSdiYr6QbbbqeB1BFxqIUn+d4UGfLM/8hulSqudEI+b8ajm5wENvQFpr qxgIg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTQ2 FelKSTgHEKVgvkY8tHcuDNVEAPCoxnB1STILl3mSOMqFTZSay6/ZYJqGaXIjf15nNBLQkS eU9mF5lLSVPLb2CPUcn9gJUbtqVZ/RCOIIdYfQK21+XHkehy1IcjeQ9DMmCuHy0yD6l4Bo YX/xWt+2giJuXoJnKCkcQdlCkYPKsf6Ju+WbotE2ZKZ7otu5xY8ea1wm66+wob/k5mrVv2 8kDZAoiYUG5B+oXDxbOgKDaeXqBa405IAf2lrH5h5KhqIh01oOJwJJ+HkLfWi1VRc36FAg hVEvqxZZK3G9CgHUJk6MTFpk8O74DQPxp8Vc6G5mmmQON0G5YZoi+6Yfhobw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:41 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 12/57] mm/collapse: allocate a destination per candidate Date: Sun, 16 Aug 2026 23:45:24 +0100 Message-ID: <20260816224609.308019-13-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the allocation, which happens on both sides of the freeze. A destination is a folio of the candidate's order, charged to the memcg, with the memcg's deferred-split list entry taken up front while sleeping is still allowed: the PMD-order install would otherwise need one under the pmd lock. collapse_alloc() does all of that for one candidate with the gfp it is handed, and counts nothing when it fails: what a miss means is up to the caller. collapse_provision() is the caller inside the window. The sources are frozen by then and a faulter on any of them is waiting, so it asks without __GFP_DIRECT_RECLAIM: reclaim entered there would be paid for by that faulter. A candidate the allocator cannot spare one for is declined rather than failed. It keeps its freeze and records SCAN_ALLOC_LIGHT_MISS, which asks for the reclaiming gfp so a later round can allocate for it before freezing anything. Where the policy forbids reclaim there is nothing better to retry with, so the miss is the verdict: the real result is recorded and the failure counters fire. collapse_reserve() honours those requests, before the round takes any lock. This is where reclaim belongs: nothing is held or frozen, so a long compaction costs only khugepaged's own progress, which is why the mechanism being replaced allocated here too. Having asked the allocator to try hard, a miss there is a failure. Nothing sets cand->reclaim yet, so collapse_reserve() has nothing to do. The request comes from the selection side, which queues a region refused at one order for another attempt. The page table a PMD-order candidate deposits cannot be deferred the same way. pte_alloc_one() allocates with GFP_PGTABLE_USER and takes no gfp to strip, so it may reclaim and sleep whatever the order asked for. collapse_deposit() secures it ahead of the freeze and refuses the candidate when it cannot. A PMD-order window is a whole table, so there is at most one such candidate and it is the first. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 3 +- mm/collapse.c | 126 +++++++++++++++++++++++++++++ mm/collapse.h | 2 + mm/khugepaged.c | 4 +- 4 files changed, 132 insertions(+), 3 deletions(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 778f5a56956c..68693eba82ef 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -40,7 +40,8 @@ EM( SCAN_STORE_FAILED, "store_failed") \ EM( SCAN_COPY_MC, "copy_poisoned_page") \ EM( SCAN_PAGE_FILLED, "page_filled") \ - EMe(SCAN_PAGE_DIRTY_OR_WRITEBACK, "page_dirty_or_writeback") + EM( SCAN_PAGE_DIRTY_OR_WRITEBACK, "page_dirty_or_writeback") \ + EMe(SCAN_ALLOC_LIGHT_MISS, "alloc_light_miss") =20 #undef EM #undef EMe diff --git a/mm/collapse.c b/mm/collapse.c index 25c0f72a9a68..58c8d83f3468 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -17,6 +17,7 @@ #include #include #include +#include =20 #include #include "collapse.h" @@ -114,6 +115,12 @@ min(COLLAPSE_BATCH_BYTES >> (PAGE_SHIFT + COLLAPSE_MIN_MTHP_ORDER), \ COLLAPSE_TABLE_WINDOWS) =20 +/* How far a candidate got, and so what a failure has to undo for it */ +enum collapse_candidate_state { + CAND_SELECTED, /* collected; nothing held on its behalf yet */ + CAND_SKIPPED, /* refused; nothing of it left to undo */ +}; + /* * A candidate is an (addr, order) window selected for collapse. Selection * counts in PTE offsets -- the bitmap it reads and the alignment it honou= rs are @@ -123,7 +130,12 @@ struct collapse_candidate { unsigned long addr; unsigned int order; + /* The light allocation missed last round: this one may reclaim for it */ + bool reclaim; + enum collapse_candidate_state state; enum scan_result result; + struct folio *new_folio; + pgtable_t deposit; /* PMD order: fresh table to deposit */ }; =20 static unsigned long candidate_start(const struct collapse_candidate *cand) @@ -218,6 +230,43 @@ static void collapse_freeze(struct vm_area_struct *vma, { } =20 +/* + * Allocate one candidate's destination with @gfp: a folio of its order, c= harged, + * with the memcg's deferred-split list heads in place so the install cann= ot need + * to allocate under the pmd lock. Those heads cost only the first collap= se in a + * memcg. + * + * A failure counts nothing and changes nothing: what a miss means is the = caller's + * policy. + */ +static enum scan_result collapse_alloc(struct mm_struct *mm, + struct collapse_control *cc, + struct collapse_candidate *cand, + gfp_t gfp) +{ + struct folio *folio; + + folio =3D __folio_alloc(gfp, cand->order, collapse_find_target_node(cc), + &cc->alloc_nmask); + if (!folio) + return SCAN_ALLOC_HUGE_PAGE_FAIL; + + if (unlikely(mem_cgroup_charge(folio, mm, gfp)) || + folio_memcg_alloc_deferred(folio)) { + folio_put(folio); + return SCAN_CGROUP_CHARGE_FAIL; + } + + if (is_pmd_order(cand->order)) { + count_vm_event(THP_COLLAPSE_ALLOC); + count_memcg_folio_events(folio, THP_COLLAPSE_ALLOC, 1); + } + count_mthp_stat(cand->order, MTHP_STAT_COLLAPSE_ALLOC); + cand->new_folio =3D folio; + + return SCAN_SUCCEED; +} + /* * Allocate ahead of the freeze for the candidates whose light allocation = missed * last round. This is where reclaim belongs: nothing is held or frozen, = so a @@ -227,6 +276,31 @@ static void collapse_freeze(struct vm_area_struct *vma, */ static void collapse_reserve(struct mm_struct *mm, struct collapse_control= *cc) { + unsigned int i; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + enum scan_result result; + + if (!cand->reclaim) + continue; + cand->reclaim =3D false; + + result =3D collapse_alloc(mm, cc, cand, cc->policy.gfp); + if (result =3D=3D SCAN_SUCCEED) + continue; + + if (result =3D=3D SCAN_ALLOC_HUGE_PAGE_FAIL) { + /* Asked the allocator to try hard and it still missed */ + if (is_pmd_order(cand->order)) + count_vm_event(THP_COLLAPSE_ALLOC_FAILED); + count_mthp_stat(cand->order, + MTHP_STAT_COLLAPSE_ALLOC_FAILED); + } + + cand->state =3D CAND_SKIPPED; + cand->result =3D result; + } } =20 /* @@ -235,9 +309,28 @@ static void collapse_reserve(struct mm_struct *mm, str= uct collapse_control *cc) * to strip: order-0 or not, it may reclaim and sleep, which is what the w= indow * exists to keep out. The destination folio has a light gfp to fall back= on and * so can be deferred; this has none. + * + * A round is one table and a PMD-order window is the whole of it, so such= a + * candidate cannot share a round: if there is one it is the only one, and= it is + * candidates[0]. This secures one page table, never a batch of them. */ static void collapse_deposit(struct mm_struct *mm, struct collapse_control= *cc) { + struct collapse_candidate *cand =3D &cc->candidates[0]; + + if (!is_pmd_order(cand->order)) + return; + + VM_WARN_ON_ONCE(cc->nr_candidates !=3D 1); + + if (cand->state !=3D CAND_SELECTED) + return; + + cand->deposit =3D pte_alloc_one(mm); + if (!cand->deposit) { + cand->state =3D CAND_SKIPPED; + cand->result =3D SCAN_ALLOC_HUGE_PAGE_FAIL; + } } =20 /* @@ -253,6 +346,35 @@ static void collapse_deposit(struct mm_struct *mm, str= uct collapse_control *cc) static void collapse_provision(struct mm_struct *mm, struct collapse_control *cc) { + const gfp_t gfp =3D cc->policy.gfp & ~__GFP_DIRECT_RECLAIM; + const bool may_retry =3D gfp !=3D cc->policy.gfp; + unsigned int i; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + enum scan_result result; + + if (cand->state !=3D CAND_SELECTED || cand->new_folio) + continue; + + result =3D collapse_alloc(mm, cc, cand, gfp); + if (result =3D=3D SCAN_SUCCEED) + continue; + + if (may_retry) { + /* A charge miss too: charging may reclaim when allowed */ + cand->result =3D SCAN_ALLOC_LIGHT_MISS; + } else { + /* The gfp a retry would use, so this is the answer */ + if (result =3D=3D SCAN_ALLOC_HUGE_PAGE_FAIL) { + if (is_pmd_order(cand->order)) + count_vm_event(THP_COLLAPSE_ALLOC_FAILED); + count_mthp_stat(cand->order, + MTHP_STAT_COLLAPSE_ALLOC_FAILED); + } + cand->result =3D result; + } + } } =20 /* @@ -761,7 +883,11 @@ static void collapse_add_candidate(struct collapse_con= trol *cc, cc->nr_candidates++; cand->addr =3D addr; cand->order =3D order; + cand->reclaim =3D false; + cand->state =3D CAND_SELECTED; cand->result =3D SCAN_FAIL; + cand->new_folio =3D NULL; + cand->deposit =3D NULL; } =20 /* diff --git a/mm/collapse.h b/mm/collapse.h index c61db86dc6c2..feb2e0d57339 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -46,6 +46,7 @@ enum scan_result { SCAN_COPY_MC, SCAN_PAGE_FILLED, SCAN_PAGE_DIRTY_OR_WRITEBACK, + SCAN_ALLOC_LIGHT_MISS, }; =20 /* @@ -148,6 +149,7 @@ unsigned long collapse_possible_orders(struct vm_area_s= truct *vma, vm_flags_t vm_flags, enum tva_type tva_flags); enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, unsigned long address, pmd_t **pmd); +int collapse_find_target_node(struct collapse_control *cc); bool collapse_scan_abort(int nid, struct collapse_control *cc); unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order); diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 43f6107c953a..50b520961b9b 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -999,7 +999,7 @@ static void collapse_policy_forced(struct collapse_poli= cy *p) } =20 #ifdef CONFIG_NUMA -static int collapse_find_target_node(struct collapse_control *cc) +int collapse_find_target_node(struct collapse_control *cc) { int nid, target_node =3D 0, max_value =3D 0; =20 @@ -1018,7 +1018,7 @@ static int collapse_find_target_node(struct collapse_= control *cc) return target_node; } #else -static int collapse_find_target_node(struct collapse_control *cc) +int collapse_find_target_node(struct collapse_control *cc) { return 0; } --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 63DF73E6DC9; Sun, 16 Aug 2026 22:46:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920406; cv=none; b=FNm/7sSrILdJ8Vkg51H/+W6n16uhtZCNBBri8un/exMJ8NbGo8w5npZ9jJqEe2qo1/rW+CwTtCxfyz/NiplHTE8vTai1SwevwDRAYfV13D/oUmJSqy9hOWudN5qtOvUCCEjOsDJBgbrjq5n7u9jKzpQOfqn8dXYCTZ6Trw/pXIw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920406; c=relaxed/simple; bh=/jL/Xj9S814uUwoxK2MxMU+JfxxbtFZppiYbnEyQZHo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tRqFwMiW2jutWuFFL3RmXpME4gdJkzuUvIH+tp6seARnlW+2fPbiM6+BW5ViKtTFq8V3sNAUOKW8IQbgmrSJYKDaDOmGsIzh2nVR74EYMLLAm/ztI2abzvqXkWHcHLwyY5FHvxhwtqry7kQh+sNxTPs/nBMxQFGjb//K8jAtN4g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=fok10aMN; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=TrtxYGws; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="fok10aMN"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="TrtxYGws" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfout.phl.internal (Postfix) with ESMTP id 82FECEC0242; Sun, 16 Aug 2026 18:46:44 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-03.internal (MEProxy); Sun, 16 Aug 2026 18:46:44 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920404; x= 1787006804; bh=N4gTJnRBaxE38n6SDCeLxvRHbIPeapcnn6D1U93IbHc=; b=f ok10aMNy/39spNh7b/1myx2RlfH2nwhR07eiGD3+8tdIAOvvp3r1NxQL3CQm9gX9 exPEu3E6Pdlh9OKlRymUvcO5Cmzlbeb4+y4Q5NIOYltWeSt52FARVhNOTlfX7wzq YEvXjYdtm0PmZR9mttTLi9jSYcyCAofqu4iS0gLe2HHbVtd7O5EPtSfv0Lgp7zgb im8tAy638F1YmQLRlL2DQ/UgB5oMXnMSThTlhpDR1LeuB5uUj8WRn0+PhYANlcT+ 0MtN3ZUfNaWZltx9paljgwQGHxLvQCLT/HBu+VZK+RNZHXnsbhgQ/uZ8AmfuvG9c 3e9isqXU1mY3210uyeMzQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920404; x=1787006804; bh=N 4gTJnRBaxE38n6SDCeLxvRHbIPeapcnn6D1U93IbHc=; b=TrtxYGwsmRjm2TbNX bH2PZ53nziOXdJwzUfYOv3I5MDfiZh8pAx4e7XFlTEijRD/9jZ2tlDasrBaZJ2lw 7m2Bdn1SUrYy1yi3vhc7u7SrQn3p7JHEClnBiDja+yYiAERfEuhje4gme7veNGVz 2Ac1taUOpSrRYF4Fb+1DT/lieHcmZhBqIssHPMVrclS0RTUWRxPPJ7cxzcZyNdTY mZ1ZRJW+k4SZmWpnZ4aJvzZIQQokDaLNmY03++LEyqwCFjn2EW0ZH4V93vG8yYDG Sfj0MJeDoaQoL7GpTBY+vgsEntLkdfHY0WxYIBbVm8xSk/k+k10VYvy/JwGgHgz+ 6fSzw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7uwKlsd4UyzlmYX1bnlzO4+qp5TmvoTS/qzkFKfMuj25xYssUPl6UU5Ie9xRba Y0bWuaC/AePVMhqrkCnU3j3fo4+kQ4lrNr0qcDw4ZM2xDmZI+m2r8oZbGDZU/DcLlRMJvk pxa8BemzxDXpf0p5hqW+oOweEwIQBvnr/hc1O3uCrLqxLkUCX+QMvT/AmKStBz4XcOBl17 ynvfoFkCi5Qu6ZndF8QgBKMwimBrechIaqu56LTRaFtWn2N2FSoSPFDDGPAClh8o3Wii3h WmXtU4UIpdlwyanZt6DS7CEKxM1G2yBGNUvAI/9KSggaAyLfyuWlS9a0iep84iGUQaZXq0 6iNUpsqBIpwRdfrQa0A+/xTO+yk5knsx0U/h5BqCJ/u1O6PL0q6GeYKIPmzYx5SPzHQ90K GAQl6gjHLUnYI9oTVvITQFNmfjJjcdgjAyhs8xQcvwJyiAhikc+UIMAcAhfI3zKYiIIkcE WWpUS08yMoVNcqLgXdHgmzFrsyxbCXZdYepPx80Bk8ZUo4Cdg7gHz6+AsOQTtQ9KnRJVSf zc0rFuEWWRtO/usObjtA2Xve2VvTlmFG4Do92vEmazEoslwtpBFG6WA8NGs1wLiiUhFvIP MwKHN3wYSFQyJeGVeoSpJLqXArC8iLMvgyAGWm0y3Soa8UTQkYadjlc9KJgw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:43 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 13/57] mm/collapse: revalidate a round against the VMA Date: Sun, 16 Aug 2026 23:45:25 +0100 Message-ID: <20260816224609.308019-14-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the pass that re-establishes what the round is working on. Selection ran under mmap_lock and the allocation ran without it, so by the time the round takes the lock back the address space may have changed underneath it. Check that the mm is not exiting and has not had THP disabled, that the VMA the round looked up is still anonymous with an anon_vma, and find the PTE table again in case it became a huge PMD or went away. Then re-check each candidate on its own. A VMA that shrank, or was replaced by a smaller one, may no longer hold a window that fitted when it was selected, and per-size enablement may have been turned off for its order since. Such a candidate is dropped and the rest of the round goes on without it. The check is per candidate rather than over the batch because a window is aligned to its own order: thp_vma_suitable_order() on each one is the containment check in full. What the walk leaves is what the round goes on to freeze, so it also settles the batch's span, in cc->batch_start and cc->batch_end, for the one invalidate the round issues. A candidate the walk dropped is not in the span, and a round left with no candidates has no span and nothing to run. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 49 ++++++++++++++++++++++++++++++++++++++++++++----- mm/collapse.h | 11 +++++++++++ mm/khugepaged.c | 11 ----------- 3 files changed, 55 insertions(+), 16 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 58c8d83f3468..1367ade721f7 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -177,18 +177,36 @@ int collapse_control_init(struct collapse_control *cc) =20 /* * The scan and the allocation both dropped mmap_lock, so nothing seen bef= ore it - * can be trusted: find the VMA and the PTE table again, and check they st= ill - * allow every provisioned candidate. + * can be trusted: check the VMA the round just looked up and the PTE table + * again, and that they still allow every provisioned candidate. * - * This is also where the batch's span is settled, for the invalidate the = round - * issues over it. + * The VMA was found by address, so it need not be the one the scan saw, n= or + * still cover everything the round collected -- thp_vma_suitable_order() = asks + * that of each candidate, since a window is aligned to its own order. A = VMA + * that shrank under a candidate therefore refuses that candidate and no m= ore, + * like every other pass. + * + * What survives is what the round goes on to freeze, so this is also wher= e the + * batch's span is settled, for the invalidate the round issues over it. */ static enum scan_result collapse_revalidate(struct vm_area_struct *vma, unsigned long pmd_addr, struct collapse_control *cc, pmd_t **pmdp) { - unsigned int i; + struct mm_struct *mm =3D vma->vm_mm; + enum scan_result result; + unsigned int i, nr_live =3D 0; + + if (unlikely(collapse_test_exit_or_disable(mm))) + return SCAN_ANY_PROCESS; + + if (!vma->anon_vma || !vma_is_anonymous(vma)) + return SCAN_PAGE_ANON; + + result =3D find_pmd_or_thp_or_none(mm, pmd_addr, pmdp); + if (result !=3D SCAN_SUCCEED) + return result; =20 cc->batch_start =3D ULONG_MAX; cc->batch_end =3D 0; @@ -196,10 +214,31 @@ static enum scan_result collapse_revalidate(struct vm= _area_struct *vma, for (i =3D 0; i < cc->nr_candidates; i++) { struct collapse_candidate *cand =3D &cc->candidates[i]; =20 + if (cand->state !=3D CAND_SELECTED) + continue; + + /* + * The window has to still fit the VMA, which may have shrunk or + * been replaced, and its order to still be one the VMA allows. + */ + if (!thp_vma_suitable_order(vma, cand->addr, cand->order) || + !thp_vma_allowable_orders(vma, vma->vm_flags, + cc->policy.tva_type, + BIT(cand->order))) { + cand->state =3D CAND_SKIPPED; + cand->result =3D SCAN_VMA_CHECK; + continue; + } + cc->batch_start =3D min(cc->batch_start, candidate_start(cand)); cc->batch_end =3D max(cc->batch_end, candidate_end(cand)); + nr_live++; } =20 + /* Nothing the VMA still allows: no span to invalidate, nothing to run */ + if (!nr_live) + return SCAN_VMA_CHECK; + return SCAN_SUCCEED; } =20 diff --git a/mm/collapse.h b/mm/collapse.h index feb2e0d57339..0d6f77a7233b 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -138,6 +138,17 @@ struct collapse_control { unsigned long batch_end; }; =20 +static inline int collapse_test_exit(struct mm_struct *mm) +{ + return atomic_read(&mm->mm_users) =3D=3D 0; +} + +static inline int collapse_test_exit_or_disable(struct mm_struct *mm) +{ + return collapse_test_exit(mm) || + mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); +} + int collapse_control_init(struct collapse_control *cc); void collapse_control_release(struct collapse_control *cc); =20 diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 50b520961b9b..1244e161beae 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -421,17 +421,6 @@ void __init khugepaged_destroy(void) kmem_cache_destroy(mm_slot_cache); } =20 -static inline int collapse_test_exit(struct mm_struct *mm) -{ - return atomic_read(&mm->mm_users) =3D=3D 0; -} - -static inline int collapse_test_exit_or_disable(struct mm_struct *mm) -{ - return collapse_test_exit(mm) || - mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); -} - static inline bool anon_hpage_enabled(void) { if (READ_ONCE(huge_anon_orders_always)) --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 179873E6DEB; Sun, 16 Aug 2026 22:46:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920408; cv=none; b=RdRkakeMwXdPxdK8Jzr/Z7i/uGk2dTKvS8m4iQbKJgWGOnrhO7v9jexXgArBDyxJXnCzKU/CgHc0MPI0PC8rmY/l0HbpoZb/yJGNu1RWrbT2ifpAajtXNQVB8VJ/PWhfq+FAnFDPJAX2X9Q8bAodulaR0CJ4XveImplVB0pGBX8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920408; c=relaxed/simple; bh=/K3P6R1c6Qu5CInX6Cpfv/h92w0mnXi7XhwPhrHaGMw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=EMLMmhT/cDI2txVFIdvR9QEuHLiVWucgDeB4QvB3B87+xYbRDfM7Pm10geXkmniUfnv8FHkDLqYNiBKGg78OwSlbBLb4JDbu8cpoPDa62P/Cnd5nSMh32J2Gqpg+u1rY2nTtLJPOZnQ1bNO0ZXg6vy26tayd4LB2CYE9+xwbXK0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=QnZ21ipj; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=R4kNWgV5; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="QnZ21ipj"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="R4kNWgV5" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfout.phl.internal (Postfix) with ESMTP id 565A2EC0235; Sun, 16 Aug 2026 18:46:46 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:46:46 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920406; x= 1787006806; bh=Zb85UJztHw3JA3M5nAcJSta5reetpMNqLHuRgU9BdLM=; b=Q nZ21ipjWyUd1eTzzivOv25KHiP7Eif0HV83n0n76FZgkJBH6+QvBb4xClRCGXMnS /1usaiI3/SVFdUYQL/Y0lBSKJXTrOtsyDZbrMkmvbSCfKBZ58m4FkBDECYTznHzW Ls0UUc9z4e1xXRqRjuMH7XFS2WnnPTGiLfcO3RzvifEKwKTxIZIkE+dczzErsp1M c83n2PCDsCn3l0FWooOP79bWMzGi0P3H1QJG5P3tYVXSdn2rrvxQwc0j5Fn7bMEr 4/MrU7C8fsA4QNRGJJHOkBQKlXVdD+Hh/u3hdv8YRK2PjCbg1kkua2a3ICnKe1Wz hSPehFdTlVHRQAwbXWegQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920406; x=1787006806; bh=Z b85UJztHw3JA3M5nAcJSta5reetpMNqLHuRgU9BdLM=; b=R4kNWgV5fZmKgUlpr VjmQn+JP7FviQCcDn28QUyNuwhM85o4fAd0O/fHnZuoI0sSQixUZcOcmEL/KQWkM fwRXn33cAJs/eephI5MhnL4XxHxfKAGXfEc434m/am9BL4I32nFrUvIRUauRf6i6 VjFEFqoxOQ7BdsPVOykihOBblKf9u5dF8K3vH6QoNOSTFXi+Fa+OzL7/2RFeRmMS pubFQZ/vEE6gyXeEJKhI+84AKASvOQyvCIhRe48rRddiwCzUVaDQDOF/TRSjFYZV 84BThpWufY3vwLpJ0HT+tVr6J1PJA5MeMFNWLCJWOsXUhK4/8zsCB6atewy2N4BY wUnHw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGZUkmTslnKCs7to/74HclCJPP3VJxScbqVRegQUfOCFxFw5+bolWK/BzCtVpImiL qUoqqipYizGRZCW/QqKQsGuW1r9sawcQpPF5VVSsi50rSl4WY0FgsqPmgyX/AxTqnJEEYU 8Aue+P7bmaQGut7CVp7DLPD3XZ1358Jxa/HMzcUO0c1V9QHT5rZ20To/M+ZhpArhbl8m0P Kq0bqT8wsLV5YFfKRlZGuKUSK+HqUMNqsCZaBBL3myYMmQW6NzFzYcenzms/0bYbquZQOj 9ZI8n0ygmC7ceV93VHQH0DqBIBjKEWMiIDR7rGwkZoIdDy7gnvFabP0Z+vKHmdDDitqQqZ tnV91u4h5K8w3szoc0CCzwSVWOhiy31U0B+5LNf/lg0QTvO5tj8OdzdB8c1mzTelhoi0DE xMiHzh181hdhcz6Ku3Nes2jBg2hmcAGZG7+cYFgyNqYMHAJ0OLXvd7YW/0gQqp3qiPOoJx jxvd3DDT+er8jdLuKwds5D5U1bOBtrVdWkEsRc1yBLDqJIYO5inh9Xwist/oOStMUH4gzs sC/8orncApMqn+fYpCDusfC6dnceyI+4OYMdo9lPMUkgKfWy4bJjozpfCzIUvCXW/4evxn 1kyBXjfE1ghcPqncM3U/d9NxCLKqAYWbIW2WxC/1Qff2w+GR1fXobzTK7EOg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:45 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 14/57] mm/collapse: fault the sources in before the freeze Date: Sun, 16 Aug 2026 23:45:26 +0100 Message-ID: <20260816224609.308019-15-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the pass that makes the sources fit to freeze. The freeze takes the PTEs as it finds them and cannot fault, so before it starts every slot a candidate covers has to be a hole, the zeropage, or a present page this mm owns exclusively. Walk each candidate slot by slot and let the fault path do the work: a swap entry is read back in, a page shared with a fork child is unshared in place. Exclusivity is tested with PageAnonExclusive() rather than by asking whether the folio looks shared. They are not the same test. A page whose fork co-mapper has exited is mapped once and looks unshared, but stays non-exclusive until some write reuses it, so a sharing test would skip the unshare on exactly the pages that need one. Each address gets a few tries, since an unshare can lose a race with a co-mapper and a swap read can be interrupted; the PTE is re-read after every fault. What is still unfit after that is left to the freeze, which refuses it. Swap-in is refused outright below the PMD order, where reading pages back to build an mTHP is not worth the latency. That verdict is against the one candidate, which is skipped so the rest of the batch can go on. The pass sleeps, and on failure it returns with mmap_lock already dropped, since the fault path may drop it and the caller cannot tell which case happened. A fault that has to wait for a swap read is one of those: it drops the lock and returns VM_FAULT_RETRY, which says nothing about the window it was working on. Report it as SCAN_LOCK_DROPPED, so the round can take the lock and run the pass again rather than treat it as a refusal. Reporting it as SCAN_PAGE_LOCK made it indistinguishable in a trace from a folio someone else had locked. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 130 +++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 129 insertions(+), 1 deletion(-) diff --git a/mm/collapse.c b/mm/collapse.c index 1367ade721f7..4ec02071f588 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -153,6 +153,11 @@ static unsigned long candidate_end(const struct collap= se_candidate *cand) return candidate_start(cand) + candidate_size(cand); } =20 +static unsigned int candidate_nr_pages(const struct collapse_candidate *ca= nd) +{ + return 1U << cand->order; +} + /* Where a candidate sits in the table, in the PTE offsets selection count= s in */ static unsigned int candidate_offset(const struct collapse_candidate *cand, unsigned long pmd_addr) @@ -242,6 +247,93 @@ static enum scan_result collapse_revalidate(struct vm_= area_struct *vma, return SCAN_SUCCEED; } =20 +/* + * Faults one address may take before the freeze is left to judge it. Mor= e than + * one because the unshare can race a co-mapper re-sharing the page, and a= swap + * read can be interrupted; each try re-reads the PTE to see where it stan= ds. + */ +#define COLLAPSE_FAULTIN_TRIES 3 + +/* + * Bring one address to a state the freeze will accept: present, and exclu= sive if + * it is anonymous. Returns with mmap_lock dropped on every failure, beca= use the + * fault path may drop it and the caller cannot tell which case it is in. + * + * SCAN_EXCEED_SWAP_PTE is the exception: it is a verdict on this candidate + * rather than on the round, nothing was faulted to reach it, and it keeps= the + * lock so the caller can refuse this candidate and carry on with the rest. + */ +static enum scan_result collapse_faultin_addr(struct vm_area_struct *vma, + struct collapse_candidate *cand, + pmd_t *pmd, unsigned long addr) +{ + struct mm_struct *mm =3D vma->vm_mm; + const unsigned int flags =3D FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_UNSHARE | + (mm !=3D current->mm ? FAULT_FLAG_REMOTE : 0); + unsigned int tries; + + for (tries =3D 0; tries <=3D COLLAPSE_FAULTIN_TRIES; tries++) { + struct page *page; + pte_t ptent, *pte; + vm_fault_t ret; + + pte =3D pte_offset_map(pmd, addr); + if (!pte) { + mmap_read_unlock(mm); + return SCAN_NO_PTE_TABLE; + } + ptent =3D ptep_get_lockless(pte); + pte_unmap(pte); + + /* A hole or the zeropage is population's business */ + if (pte_none_or_zero(ptent)) + break; + + if (pte_present(ptent)) { + page =3D vm_normal_page(vma, addr, ptent); + + /* + * PageAnonExclusive is the invariant the freeze relies + * on, and the only exact test for it. Testing sharing + * with folio_maybe_mapped_shared() is not the same: a + * page whose fork co-mapper has gone away is + * single-mapped, yet stays non-exclusive until a write + * reuses it, so sharing would skip the unshare on + * exactly the pages that need it. Unsharing one of + * those is cheap -- it reuses the page in place and + * just sets the bit. + */ + if (!page || !folio_test_anon(page_folio(page)) || + PageAnonExclusive(page)) + break; /* already exclusive */ + } else if (!is_pmd_order(cand->order)) { + /* Sub-PMD collapse does not fault swap in */ + count_mthp_stat(cand->order, + MTHP_STAT_COLLAPSE_EXCEED_SWAP); + return SCAN_EXCEED_SWAP_PTE; + } + + if (tries =3D=3D COLLAPSE_FAULTIN_TRIES) + break; /* the freeze refuses it if still unfit */ + + /* Only swap or shared PTEs reach here; the rest broke out */ + ret =3D handle_mm_fault(vma, addr, flags, NULL); + /* + * Not a verdict on this window: the fault dropped the lock to + * wait, which is what a swap-in normally does. Distinct from + * SCAN_PAGE_LOCK, a folio someone else holds locked. + */ + if (ret & VM_FAULT_RETRY) + return SCAN_LOCK_DROPPED; + if (ret & VM_FAULT_ERROR) { + mmap_read_unlock(mm); + return SCAN_FAIL; + } + } + + return SCAN_SUCCEED; +} + /* * Make every source the round needs present and exclusively owned by this= mm, * by faulting it in as an ordinary access would. Sleeps, and drops mmap_= lock on @@ -255,7 +347,43 @@ static enum scan_result collapse_faultin(struct vm_are= a_struct *vma, struct collapse_control *cc, pmd_t *pmd) { - return SCAN_SUCCEED; + enum scan_result result =3D SCAN_SUCCEED; + unsigned int i; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + unsigned long addr; + unsigned int j; + + if (cand->state !=3D CAND_SELECTED) + continue; + + for (j =3D 0, addr =3D cand->addr; + j < candidate_nr_pages(cand); + j++, addr +=3D PAGE_SIZE) { + enum scan_result r; + + r =3D collapse_faultin_addr(vma, cand, pmd, addr); + /* + * The one failure that judges this candidate rather + * than the round, and so the one that leaves the lock + * in our hands: refuse it and go on to the next. + * Failing the round here would lower the order of every + * candidate it carries, a verdict nobody reached. + */ + if (r =3D=3D SCAN_EXCEED_SWAP_PTE) { + cand->state =3D CAND_SKIPPED; + cand->result =3D r; + break; + } + if (r !=3D SCAN_SUCCEED) { + result =3D r; + goto out; + } + } + } +out: + return result; } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 040F53EE1EA; Sun, 16 Aug 2026 22:46:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920410; cv=none; b=FZQw/CfDJN8GhA5W9HZseqJw3nquWsUsP9JQc02lcd1kqWaMDi9iYnjwJGzrCWeG5vQDaso3Mi54+MmLhB9wu+A17o21NFTKqDawwcqkEqZLQd+vpu85DO4jvGUjdzmIOPTEqiyyNiUoNL4iBISc+UjbO5xaA5ANwMjd0yyw7js= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920410; c=relaxed/simple; bh=0HtoTUgdKn4MCSLvYCOYck5ILC37qSnuMNZoMgnMGyo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kiTJkJ6Wmvhb+vKKdrN3NOfOL2VgkXEKxYzSXxnFGI+KdhFaPdUSTXAL7uElXKLKa9ANuPnQpLUjXGv41XVVAD5PMudNuHqkfM7m8LNbFMsni++rjd9dMFNGchTYrYUKe7w/RU8spDxOGiWlypLqPRo3b+x3kNpa4CwuD1Epnkw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=u2eiQZS+; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=MorGnxWR; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="u2eiQZS+"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="MorGnxWR" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfhigh.phl.internal (Postfix) with ESMTP id 2012814000FB; Sun, 16 Aug 2026 18:46:48 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-03.internal (MEProxy); Sun, 16 Aug 2026 18:46:48 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920408; x= 1787006808; bh=rUR3Zpc/XHDoRz+IWnP+uEJLJCI5tm2EqK26zu94310=; b=u 2eiQZS+PPcnEmgXisrRYA6y9wwXZOsne4YxKehtUO3tQk3gZLV5Uqbmm8zPeW72v ycl5cv/UWjL+l1OglOPprBxLB6LbFW/UNZLzaHNDQY2C+IimYiVcYleAX10hq1vR 2Ro2s5VBVSs71o46dPqcocLLQn4/OglqLYNoZ9xFpwmpf0v2bMdg0o/8UNWN5+XX WK494ewheYgmhL284NzvvHrek66pOVjDnufTOnOOoPQJO26UbBPSccDPAajf+PMq TjtMeKxRpvEhO+MzLM3RvryGrk9egyKNVmST/6HgP1UL2VjVmEWN0JiE1uq5ogHT pOS+/jCAxNrh1Y9NUppRg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920408; x=1787006808; bh=r UR3Zpc/XHDoRz+IWnP+uEJLJCI5tm2EqK26zu94310=; b=MorGnxWRAxibSJlMq CuFoPjBKqc9op52AmmngkdTrU5U13L6qUMiC7OGDS0Abr0MvAavpBbH4wC1asHAd xOqiF4n6Pb2hZlZXUO4aCQJ/oP8N5r5yMVvvdM6rmwc8PS35dlh0NOUbK3Hm1+ua W59YwUdgcKCryrzT8giPQ88sqxPniiOPdq6ZtAmVnpAgGOQ3cu5ko21eO7BPNy+2 ii603Z6dtn47t80Q+8irE5o6H+NS8rtXCIXafEBGzC1e8xlGOnZKVGGaPQse7Ozu fYL2zNBZ+iCNfJldyQyhniN1u8ZzY5gwKVByvy/vhKF9mGYIFBWIX4d+7mAoSOve 96hQQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7Np 6woFLzhF1hDHTk9QeVB0U2SY8KNjMUPT0K/OFRCJ5NqYQs5V8G3gGWet0vT437YNR6DToW /gIXHEXzbuC4u+J7MRVGXHvi6R37UNnaUIBS1CWMVm3C9IWXgDMNOz1H19R6kIfoYFHHNM WnDie5YX6XYhQ12+CGbhXjbEWTvlVvRGLTpuTDTSL8zLja5nBpRTet3lSYowUoPCOAdex6 zpObjG0Nn4tYu6EmkcOqvXVkDv9wJkReKSomq/wNVDF7nLboDGymkyErZb9niUdqfxvI6g LV1S6ltfp5cEKg49QpYiq0df5oTJC9zv+GRSx7+nPicgLRGUqtfWdxfHONDA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:47 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 15/57] mm/collapse: check what a candidate would freeze Date: Sun, 16 Aug 2026 23:45:27 +0100 Message-ID: <20260816224609.308019-16-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The freeze takes folio locks, rewrites PTEs and flushes the TLB, and any of that has to be undone slot by slot if the candidate turns out unfit -- while faulters on those sources wait. So it decides first and acts second. This is the deciding half: walk every slot a candidate covers, under the table's ptl, and answer whether all of it can be frozen. It touches nothing, so a refusal costs the round only the walk. The walk goes in source spans, a span being consecutive PTEs mapping consecutive pages of one folio. No layout is refused for its shape: where a span ends, the next slot starts one of its own, which is what lets partially mapped and compound sources collapse. A slot may also be a hole or the zeropage, both of which the destination just zero-fills. What a span has to satisfy, beyond being present, anonymous and not uffd-armed: - Every live mapping of its folio is this span. The freeze is whole-folio, so a live PTE anywhere else would race a zap whose folio_put() underflows the frozen count. Under the ptl this is exact, since fork -- the only way an exclusive anon folio gains mappings -- takes mmap_write. - Every page of it is PageAnonExclusive(). A shared folio has no refcount the freeze can pin down without the other mappers' ptls. - It is not MADV_FREE'd, unless the caller asked for the collapse. Copying a lazyfree page into a folio that is not lazyfree would quietly make memory the user offered up undroppable again, which is why the policy carries that choice. Sub-PMD candidates also refuse folios already at or above their own order, there being nothing to gain; a PMD candidate takes them, that being the PTE-mapped-THP re-collapse case. SCAN_PAGE_NOT_EXCLUSIVE joins enum scan_result and the trace symbol list. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 1 + mm/collapse.c | 177 +++++++++++++++++++++++++++++ mm/collapse.h | 1 + 3 files changed, 179 insertions(+) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 68693eba82ef..ff938ac9c43c 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -41,6 +41,7 @@ EM( SCAN_COPY_MC, "copy_poisoned_page") \ EM( SCAN_PAGE_FILLED, "page_filled") \ EM( SCAN_PAGE_DIRTY_OR_WRITEBACK, "page_dirty_or_writeback") \ + EM( SCAN_PAGE_NOT_EXCLUSIVE, "page_not_exclusive") \ EMe(SCAN_ALLOC_LIGHT_MISS, "alloc_light_miss") =20 #undef EM diff --git a/mm/collapse.c b/mm/collapse.c index 4ec02071f588..c75d91cb9d48 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -386,6 +386,145 @@ static enum scan_result collapse_faultin(struct vm_ar= ea_struct *vma, return result; } =20 +/* + * How many slots a source span starting at @first may cover: the pages le= ft in + * its folio, capped at @max. Every freeze-side walker bounds spans with = this, + * so per-span batching of clears, locks and freezes cannot reach a slot t= he span + * does not cover. + */ +static unsigned int collapse_span_max(pte_t first, unsigned int max) +{ + struct page *page =3D pte_page(first); + struct folio *folio =3D page_folio(page); + unsigned int left =3D folio_nr_pages(folio) - folio_page_idx(folio, page); + + return min(max, left); +} + +/* + * Can this candidate's sources be frozen? Every slot is checked and noth= ing is + * touched, so a refusal costs the round nothing but the walk. + * + * The walk is in source spans: a span is consecutive PTEs mapping consecu= tive + * pages of one folio, and it ends wherever the next PTE stops being the f= olio's + * next page. No layout is refused for its shape -- the next slot simply = starts + * its own span -- so partially mapped and compound sources collapse too. + * + * Caller holds mmap_read and the table's ptl. + */ +static enum scan_result collapse_check_candidate(struct vm_area_struct *vm= a, + struct collapse_control *cc, + struct collapse_candidate *cand, + pte_t *pte) +{ + const unsigned int nr_pages =3D candidate_nr_pages(cand); + unsigned long addr; + unsigned int i; + + for (i =3D 0, addr =3D cand->addr; i < nr_pages;) { + pte_t ptent =3D ptep_get(pte + i); + unsigned int nr, nr_max, k; + struct folio *folio; + struct page *page; + + if (!pte_present(ptent)) { + /* Holes are population; swap and markers are not */ + if (pte_none(ptent)) { + i++; + addr +=3D PAGE_SIZE; + continue; + } + return SCAN_PTE_NON_PRESENT; + } + if (pte_uffd(ptent)) + return SCAN_PTE_UFFD; + + /* The zeropage zero-fills like a hole, and has no normal page */ + if (is_zero_pfn(pte_pfn(ptent))) { + i++; + addr +=3D PAGE_SIZE; + continue; + } + page =3D vm_normal_page(vma, addr, ptent); + if (!page || unlikely(is_zone_device_page(page))) + return SCAN_PAGE_NULL; + + folio =3D page_folio(page); + if (!folio_test_anon(folio)) + return SCAN_PAGE_ANON; + + /* + * Collapsing a MADV_FREE'd page would copy it into a folio that + * is not lazyfree, quietly making memory the user offered up + * undroppable again. + */ + if (cc->policy.skip_lazyfree && + !(vma->vm_flags & VM_DROPPABLE) && + folio_test_lazyfree(folio) && !pte_dirty(ptent)) + return SCAN_PAGE_LAZYFREE; + + /* + * A sub-PMD candidate refuses folios of its own order and above: + * collapsing those would gain nothing. A PMD candidate accepts + * every order up to its own -- the PTE-mapped-THP re-collapse + * class. + */ + if (folio_order(folio) >=3D cand->order && + !is_pmd_order(cand->order)) + return SCAN_PTE_MAPPED_HUGEPAGE; + + /* + * Exclusive anon only: the expected refcount of a shared folio + * cannot be pinned down without its other mappers' ptls. + * Swapcache membership is fine -- folio_expected_ref_count() + * accounts those references. + */ + if (folio_maybe_mapped_shared(folio)) + return SCAN_PAGE_NOT_EXCLUSIVE; + + nr_max =3D collapse_span_max(ptent, nr_pages - i); + for (nr =3D 1; nr < nr_max; nr++) { + pte_t tail =3D ptep_get(pte + i + nr); + + if (!pte_present(tail) || + pte_pfn(tail) !=3D pte_pfn(ptent) + nr) + break; + if (pte_uffd(tail)) + return SCAN_PTE_UFFD; + } + + /* + * Every live mapping of the folio must be this span: the freeze + * is whole-folio, and a live PTE left anywhere else loses to a + * racing zap -- its rmap drop is paired with a folio_put() that + * would underflow the frozen count. The check is race-free + * under our ptl: in-window PTEs are ours, fork (the only way + * exclusive anon gains mappings) takes mmap_write, and a folio + * whose mappings all sit under this ptl cannot lose one either. + * This also refuses a folio scattered across several spans of + * the window, whose mapcount exceeds any single span. + */ + if (folio_mapcount(folio) !=3D nr) + return SCAN_PAGE_COUNT; + + /* + * Every page of the span must be exclusive: the freeze accounts + * only references it can see, and a non-exclusive page may be + * unshared under us. collapse_faultin() should have arranged + * this; enforce it here, where it is depended on. + */ + for (k =3D 0; k < nr; k++) { + if (!PageAnonExclusive(pte_page(ptep_get(pte + i + k)))) + return SCAN_PAGE_NOT_EXCLUSIVE; + } + + i +=3D nr; + addr +=3D nr * PAGE_SIZE; + } + + return SCAN_SUCCEED; +} + /* * Raise the two barriers on the sources of every candidate: migration ent= ries in * their PTEs, then a frozen refcount. Takes the table's ptl once for the= whole @@ -395,6 +534,44 @@ static enum scan_result collapse_faultin(struct vm_are= a_struct *vma, static void collapse_freeze(struct vm_area_struct *vma, struct collapse_control *cc, pmd_t *pmd) { + struct mm_struct *mm =3D vma->vm_mm; + pte_t *pte, *table; + spinlock_t *ptl; + unsigned int i; + + pte =3D pte_offset_map_lock(mm, pmd, cc->candidates[0].addr, &ptl); + if (!pte) { + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + + if (cand->state !=3D CAND_SELECTED) + continue; + cand->state =3D CAND_SKIPPED; + cand->result =3D SCAN_NO_PTE_TABLE; + } + return; + } + + /* + * Index each candidate from the table base, not relative to + * candidates[0]: a round is not necessarily address-ordered, so + * candidates[0] need not be the lowest. They all share one table. + */ + table =3D pte - pte_index(cc->candidates[0].addr); + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + pte_t *cand_pte =3D table + pte_index(cand->addr); + + if (cand->state !=3D CAND_SELECTED) + continue; + + cand->result =3D collapse_check_candidate(vma, cc, cand, cand_pte); + if (cand->result !=3D SCAN_SUCCEED) + cand->state =3D CAND_SKIPPED; + } + + pte_unmap_unlock(pte, ptl); } =20 /* diff --git a/mm/collapse.h b/mm/collapse.h index 0d6f77a7233b..747168104a72 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -46,6 +46,7 @@ enum scan_result { SCAN_COPY_MC, SCAN_PAGE_FILLED, SCAN_PAGE_DIRTY_OR_WRITEBACK, + SCAN_PAGE_NOT_EXCLUSIVE, SCAN_ALLOC_LIGHT_MISS, }; =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 94F313E764F; Sun, 16 Aug 2026 22:46:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920412; cv=none; b=UPOh8wS22OFz7a4D4UHgj2IhSolZt/T4KUMIt+iC+2B2NJozszaxuHqc6RRDsFJjBV1oTR92O9IDWVYUNmIBXJZTwncLpZqUqMUePp+w06W6pZs3YTdZffI2gljs33MiXrRXWTtDplTxnHcF10qch70pDo6S+lVvcpQF785hK9s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920412; c=relaxed/simple; bh=oFoPuWXxu3i4BN+ux57sLwk+W/FDUJ70ewm5/juGERc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=O+ERwfi40RlZ/vjbOTyXvs6Z/Cnksb/vrDu0b9PI7WRjdGq1OlCMzjEONq1mtaE7pOFvVPCNRDcZsgpakXspz/yDNNaMqFPID4GV38fq81znAIWj6dnJBzRZYU70Xv2RbSBlMSJv25+9tXpTugSGo27kNqfsyXuNVsAwyKl1HSA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=E069G+Do; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=KntytyFr; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="E069G+Do"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="KntytyFr" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfout.phl.internal (Postfix) with ESMTP id E9702EC0074; Sun, 16 Aug 2026 18:46:49 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:46:49 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920409; x= 1787006809; bh=2ywJVJxC7MTikatPnYdIRmEVTRQZLN6oVNDZix3roGc=; b=E 069G+DoIYd+WI0gOOhBCViUF8G821/Ar0ln7d4a8iNI21rU26YxRjhxNFbMtFjQh cv8WOu/uaHbRqGQspQZFAdpZPm/l3pQIYrAGApx4Ps7mWlDilw8BalVVB4uVNoBu CgAq/RoxKLCOWMcsNrshl6idHZ3jjQ6pzJyI4aic82mNYoNPZA8QSd0kbdsEApz0 fdknEXAdVZdgUaZQ0P1K43Q4S2WfbEf0nJYU7DU0ALObbv8dlqsLyeJyo2Wa5wVT 5NaTF6BWpEh4myesfbPy/U1b0W1aXoly0JCb1/oib9R7K8TnYC/kqW75fKQ8W7/f /haFW09Y4cf6g3N7xAxbA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920409; x=1787006809; bh=2 ywJVJxC7MTikatPnYdIRmEVTRQZLN6oVNDZix3roGc=; b=KntytyFr+cnbA8Fp1 nE3IT98uGb8FecuX6Oi9fH2eUS5cOK+OhKoM7BKBAejvC7cQTVOWn6rMAhA7ieNr yGtiKztcbtHmpxLjs1okF+pzv4qsxaE7ycrysMzswxHrJbCxo6GJG9UFO/qKrcZY 8vEkR2/DH8lhZIQ8bY9XzPRON2zTTo2O0uRSaabMvpqQMiKvAU9zUmhkY4EBtyw3 X3e5tCvyxQ/4iT2aRIdRHrB/ZPjsj8fv+rOH6Q7KqN9/PgsPZqEviz2Ji0RfYyTN NqUM/mlfHl9N+NPU0kem8cZBSo6N4/FNA1VwbgYuRMFkhblATvg8mJ9aINlrzCTX N9edA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7uwKlsd4UyzlmYX1bnlzO4+qp5TmvoTS/qzkFKfMuj25xYssUPl6UU5Ie9xRba Y0bWuaC/AePVMhqrkCnU3j3fo4+kQ4lrNr0qcDw4ZM2xDmZI+m2r8oZbGDZU/DcLlRMJvk pxa8BemzxDXpf0p5hqW+oOweEwIQBvnr/hc1O3uCrLqxLkUCX+QMvT/AmKStBz4XcOBl17 ynvfoFkCi5Qu6ZndF8QgBKMwimBrechIaqu56LTRaFtWn2N2FSoSPFDDGPAClh8o3Wii3h WmXtU4UIpdlwyanZt6DS7CEKxM1G2yBGNUvAI/9KSggaAyLfyuWlS9a0iep84iGUQaZXYv /Kitamsvub2Ns1F5CATJjmh/MwtuDVaVCUSOL+ReX2n/nS6/v8j7ZLyQvo1bPe9fvWARNN 1GR7OlnsagNqnFjSx+CpisG+sJTSVBoCjRelw724v+Plqd7NKxMQMfzQQMqfXxeSl8X09O tGE42hHsNTh4oM3hZ4RbwZNz9Ozf5SsyXL0bVSRUeVrgxMT6Ah9BfHH+FuYMZVLN3HUPlJ rXAvkqE085ShHgBKFrQgk5/v8UnTjEndj+ycM/lbi6F3pLD8CHnrgfNvFMJCTOpalHL+ey yU+2VzoySjHyuq+Imx2QXeWPjwppnEDRQ6ADg2jdj92om8VH5PgbGz/CrJJQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:49 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries Date: Sun, 16 Aug 2026 23:45:28 +0100 Message-ID: <20260816224609.308019-17-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The acting half. For every span the checking half accepted: 1. take a reference on the source folio and lock it; 2. replace its PTEs with migration entries. This closes the userspace side: faults and GUP-slow now wait on that folio lock, taken before the first entry becomes visible; 3. freeze the folio to folio_expected_ref_count() + 1. This closes the kernel side: folio_try_get() fails. The two barriers rise in that order because it is reachability order, and together they are what lets the copy run with no lock at all. Writeback is the one case the freeze cannot catch, so such a folio is refused up front. PG_writeback holds no reference of its own, so the frozen count is exactly right and the freeze succeeds -- then folio_end_writeback() takes a reference outright and BUGs on it, or frees it under the copy. From the freeze to the putback the round holds every source folio's lock at once, and folio locks have no global order. The engine only ever folio_trylock()s, and unfreezes rather than blocks on refusal, so it is never the waiting edge of a cycle. One ranged TLB flush covers everything that froze, before the ptl is dropped. Until it completes a CPU with a stale entry could still write a source through the old mapping, and that path never consults a refcount. A candidate that cannot finish restores what it displaced, unfreezes, unlocks and drops out; its neighbours carry on. The restore needs no flush of its own: what it puts back is identical to whatever a stale entry holds. It restores slot by slot. The PTEs of one span agree on the PFN and nothing else, so a partial CoW, or a clear_refs write-protect undone one page at a time, leaves permissions the first PTE cannot stand for. Dirty accumulated over the span goes to the folio instead, the way unmap does. The displaced values live in a pool sized to a whole table, since one candidate can displace that much and the byte cap bounds a round, not a candidate. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 335 ++++++++++++++++++++++++++++++++++++++++++++++++-- mm/collapse.h | 3 + 2 files changed, 330 insertions(+), 8 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index c75d91cb9d48..cf2b9b3640ae 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -115,10 +115,20 @@ min(COLLAPSE_BATCH_BYTES >> (PAGE_SHIFT + COLLAPSE_MIN_MTHP_ORDER), \ COLLAPSE_TABLE_WINDOWS) =20 +/* + * The saved-PTE pool spans a whole table. The byte cap bounds what a rou= nd + * holds, but not what one candidate does: a sub-PMD order goes up to + * HPAGE_PMD_NR/2 pages -- 256M at order 12 with 64K pages -- and displace= s all + * of its PTEs in one shot regardless. So the pool has to fit the largest= span + * of displaced PTEs a table can hold, which is the table itself. + */ +#define COLLAPSE_SAVED_PTES HPAGE_PMD_NR + /* How far a candidate got, and so what a failure has to undo for it */ enum collapse_candidate_state { CAND_SELECTED, /* collected; nothing held on its behalf yet */ CAND_SKIPPED, /* refused; nothing of it left to undo */ + CAND_FROZEN, /* sources displaced and frozen */ }; =20 /* @@ -136,6 +146,7 @@ struct collapse_candidate { enum scan_result result; struct folio *new_folio; pgtable_t deposit; /* PMD order: fresh table to deposit */ + pte_t *saved_ptes; /* its slice of collapse_control::saved_ptes */ }; =20 static unsigned long candidate_start(const struct collapse_candidate *cand) @@ -168,15 +179,20 @@ static unsigned int candidate_offset(const struct col= lapse_candidate *cand, void collapse_control_release(struct collapse_control *cc) { kfree(cc->candidates); + kfree(cc->saved_ptes); cc->candidates =3D NULL; + cc->saved_ptes =3D NULL; } =20 int collapse_control_init(struct collapse_control *cc) { cc->nr_candidates =3D 0; cc->candidates =3D kmalloc_objs(*cc->candidates, COLLAPSE_MAX_CANDIDATES); - if (!cc->candidates) + cc->saved_ptes =3D kmalloc_objs(*cc->saved_ptes, COLLAPSE_SAVED_PTES); + if (!cc->candidates || !cc->saved_ptes) { + collapse_control_release(cc); return -ENOMEM; + } return 0; } =20 @@ -401,6 +417,99 @@ static unsigned int collapse_span_max(pte_t first, uns= igned int max) return min(max, left); } =20 +/* + * Length of the source span at slot @i, read from the saved PTEs rather t= han the + * table: once frozen the slots hold migration entries, so a rollback re-d= erives + * the freeze's spans from what it displaced. + */ +static unsigned int collapse_saved_span_len(struct collapse_candidate *can= d, + unsigned int i, unsigned int bound) +{ + pte_t first =3D cand->saved_ptes[i]; + unsigned int nr, nr_max; + + nr_max =3D collapse_span_max(first, bound - i); + for (nr =3D 1; nr < nr_max; nr++) { + pte_t saved =3D cand->saved_ptes[i + nr]; + + if (pte_none_or_zero(saved) || + pte_pfn(saved) !=3D pte_pfn(first) + nr) + break; + } + return nr; +} + +/* + * Undo a freeze that could not complete: restore the displaced PTE values= over + * the candidate's migration entries, then unfreeze, unlock and release the + * source folios. + * + * How far the freeze got: + * + * - @nr_saved slots were displaced, in PTEs; + * - @nr_frozen of those belong to folios that were also frozen. + * + * Each slot restores by class: a hole was never modified, a cleared zerop= age is + * stored back plainly, and a source's saved value goes back as it was. A= ll are + * plain stores -- writing over a non-present entry has no hardware A/D ra= ce. + * + * Slot by slot, not one set_ptes() over the span: the PTEs of one folio n= eed + * not agree on more than the PFN, so the first one's permissions are not = the + * span's. + * + * Deliberately no TLB flush: the restored translation is identical to any= thing + * a stale TLB entry may hold, so every stale entry is benign. This reads= like + * a missing flush; it is not. + * + * Caller holds the table's ptl -- the same uninterrupted hold the freeze = ran + * under. + */ +static void collapse_unfreeze_candidate(struct mm_struct *mm, + struct collapse_candidate *cand, + pte_t *pte, unsigned int nr_saved, + unsigned int nr_frozen) +{ + unsigned long addr =3D cand->addr; + unsigned int i =3D 0; + + while (i < nr_saved) { + pte_t saved =3D cand->saved_ptes[i]; + struct folio *folio; + unsigned int nr, k; + + if (pte_none(saved)) { + /* Hole: nothing was touched */ + i++; + addr +=3D PAGE_SIZE; + continue; + } + if (is_zero_pfn(pte_pfn(saved))) { + /* Cleared zeropage: plain non-present -> present store */ + set_pte_at(mm, addr, pte + i, saved); + i++; + addr +=3D PAGE_SIZE; + continue; + } + + folio =3D pte_folio(saved); + nr =3D collapse_saved_span_len(cand, i, nr_saved); + + for (k =3D 0; k < nr; k++) { + set_pte_at(mm, addr + k * PAGE_SIZE, pte + i + k, + cand->saved_ptes[i + k]); + } + if (i < nr_frozen) { + folio_ref_unfreeze(folio, + folio_expected_ref_count(folio) + 1); + } + folio_unlock(folio); + folio_put(folio); + + i +=3D nr; + addr +=3D nr * PAGE_SIZE; + } +} + /* * Can this candidate's sources be frozen? Every slot is checked and noth= ing is * touched, so a refusal costs the round nothing but the walk. @@ -525,6 +634,193 @@ static enum scan_result collapse_check_candidate(stru= ct vm_area_struct *vma, return SCAN_SUCCEED; } =20 +/* + * Freeze one candidate's sources, span by span, raising both quiescence + * barriers in reachability order: + * + * 1. the span's PTEs become migration entries. Faults and GUP-slow now = wait + * on the source folio's lock, taken before the first entry is visible. + * 2. the folio is frozen to its expected reference count, so folio_try_g= et() + * fails for anyone taking a speculative reference. + * + * All or nothing: a failure part way through unwinds what it displaced and + * leaves the table as it was found. + * + * Neither barrier deflects a path that takes its reference outright rather + * than speculatively. Such a source has to be refused before the freeze,= not + * survive it -- see the writeback test below. + * + * A round holds every source folio's lock at once, from freeze to putback= , and + * folio locks have no global order. That cannot deadlock: folio_trylock(= ) is + * the engine's only acquisition and a refusal unfreezes instead of blocki= ng, so + * the engine is never the waiting edge of a cycle. Nothing between freez= e and + * putback waits on anything that could wait on us -- allocation and charg= ing + * happen earlier, and the copy only copies. The install does take the ptl + * while holding these folio locks, which is the safe order: a faulter on = one of + * our migration entries cannot sleep on the folio lock under a spinlock, = so it + * drops the ptl first. Do not add a blocking lock or a sleeping allocati= on + * between freeze and putback. + * + * On entry: + * + * - mmap_read is held, and the table's ptl for the whole freeze; + * - collapse_check_candidate() has accepted the candidate under that sam= e ptl + * hold; + * - the round is covered by an mmu_notifier_invalidate_range_start() iss= ued + * outside the ptl. + * + * collapse_freeze() issues the ranged TLB flush over everything that froze + * before dropping the ptl. No copy may run before it completes. + */ +static enum scan_result collapse_freeze_candidate(struct mm_struct *mm, + struct collapse_candidate *cand, pte_t *pte) +{ + const unsigned int nr_pages =3D candidate_nr_pages(cand); + unsigned int nr_saved =3D 0, nr_frozen =3D 0; + enum scan_result result; + struct folio *folio; + unsigned long addr; + unsigned int i; + + for (i =3D 0, addr =3D cand->addr; i < nr_pages;) { + pte_t ptent =3D ptep_get(pte + i); + unsigned int nr, nr_max, k; + pte_t rep; + + if (pte_none(ptent)) { + /* Hole: nothing to freeze; install verifies it stayed one */ + cand->saved_ptes[i] =3D ptent; + nr_saved =3D ++i; + addr +=3D PAGE_SIZE; + continue; + } + if (is_zero_pfn(pte_pfn(ptent))) { + /* + * Clear the zeropage mapping now, covered by the round's + * ranged flush: overwriting a live PTE at install would + * be a valid->valid transition, breaking arm64's + * break-before-make. The zeropage has neither rmap nor + * per-map references -- the saved value alone undoes it. + */ + cand->saved_ptes[i] =3D + ptep_get_and_clear(mm, addr, pte + i); + nr_saved =3D ++i; + addr +=3D PAGE_SIZE; + continue; + } + + folio =3D pte_folio(ptent); + + /* + * A folio revisited by a second span of this round is already + * ours and frozen at its first span: folio_get() on a zero count + * is a bug, and try-get fails cleanly. Scrambled layouts + * (mremap) construct this; nothing else can hold a folio frozen + * while its PTE is live under our ptl, so it is not transient. + */ + if (!folio_try_get(folio)) { + result =3D SCAN_PAGE_COUNT; + goto unfreeze; + } + if (!folio_trylock(folio)) { + folio_put(folio); + result =3D SCAN_PAGE_LOCK; + goto unfreeze; + } + + /* + * Never freeze a folio under writeback. PG_writeback holds no + * reference of its own -- the swapcache reference keeps the folio + * alive, and everything that would drop it waits for the flag -- + * so folio_end_writeback() plain folio_get()s a folio it may + * assume is alive: a BUG on a frozen one, or with + * CONFIG_DEBUG_VM off, a free under our copy. + * + * Unlike every other hazard here, the freeze does not catch it. + * folio_expected_ref_count() counts the swapcache reference, so + * the count is exactly right and the freeze succeeds. Nor can + * "is it in the swapcache" stand in for this test: that would + * refuse the pages the fault-in pass just swapped in. + * + * Reachable even though writeback starts on an unmapped folio: a + * re-fault from the swapcache maps it back before the bio + * completes, and folio_free_swap() will not drop the cache entry + * under writeback. Testing once is enough -- writeback starts + * only under the folio lock, which we hold from here through + * putback. + */ + if (folio_test_writeback(folio)) { + folio_unlock(folio); + folio_put(folio); + result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; + goto unfreeze; + } + + /* Each slot's own value: a span agrees on the PFN, not the rest */ + cand->saved_ptes[i] =3D ptent; + nr_max =3D collapse_span_max(ptent, nr_pages - i); + for (nr =3D 1; nr < nr_max; nr++) { + pte_t tail =3D ptep_get(pte + i + nr); + + if (!pte_present(tail) || + pte_pfn(tail) !=3D pte_pfn(ptent) + nr) + break; + cand->saved_ptes[i + nr] =3D tail; + } + + /* + * The clear is the GUP-fast linearization point: a grab landing + * before it elevates the refcount and the freeze below fails + * (the candidate unfreezes); one landing after fails its PTE + * re-read and retries. Clear and store sit adjacent under one + * uninterrupted ptl hold, batched per span + * (get_and_clear_full_ptes() unfolds contpte), so the transient + * none window is invisible to installers, which all take the ptl. + */ + rep =3D get_and_clear_full_ptes(mm, addr, pte + i, nr, 0); + + /* + * Dirty from the clear -- including any the hardware set since + * the reads above -- goes to the folio, the way unmap does, + * rather than onto PTEs that never had it. Young needs no such + * care: a migration entry drops it either way. + */ + if (pte_dirty(rep)) + folio_mark_dirty(folio); + + for (k =3D 0; k < nr; k++) { + pte_t saved =3D cand->saved_ptes[i + k]; + swp_entry_t entry; + pte_t swp_pte; + + entry =3D make_readable_migration_entry(pte_pfn(saved)); + swp_pte =3D swp_entry_to_pte(entry); + if (pte_soft_dirty(saved)) + swp_pte =3D pte_swp_mksoft_dirty(swp_pte); + set_pte_at(mm, addr + k * PAGE_SIZE, pte + i + k, + swp_pte); + } + nr_saved =3D i + nr; + + if (!folio_ref_freeze(folio, + folio_expected_ref_count(folio) + 1)) { + result =3D SCAN_PAGE_COUNT; + goto unfreeze; + } + nr_frozen =3D nr_saved; + + i +=3D nr; + addr +=3D nr * PAGE_SIZE; + } + + cand->state =3D CAND_FROZEN; + return SCAN_SUCCEED; + +unfreeze: + collapse_unfreeze_candidate(mm, cand, pte, nr_saved, nr_frozen); + return result; +} + /* * Raise the two barriers on the sources of every candidate: migration ent= ries in * their PTEs, then a frozen refcount. Takes the table's ptl once for the= whole @@ -534,11 +830,15 @@ static enum scan_result collapse_check_candidate(stru= ct vm_area_struct *vma, static void collapse_freeze(struct vm_area_struct *vma, struct collapse_control *cc, pmd_t *pmd) { + unsigned long flush_start =3D ULONG_MAX, flush_end =3D 0; struct mm_struct *mm =3D vma->vm_mm; pte_t *pte, *table; spinlock_t *ptl; unsigned int i; =20 + /* Pending per-CPU folio batches hold references that fail the freeze */ + lru_add_drain(); + pte =3D pte_offset_map_lock(mm, pmd, cc->candidates[0].addr, &ptl); if (!pte) { for (i =3D 0; i < cc->nr_candidates; i++) { @@ -562,15 +862,27 @@ static void collapse_freeze(struct vm_area_struct *vm= a, for (i =3D 0; i < cc->nr_candidates; i++) { struct collapse_candidate *cand =3D &cc->candidates[i]; pte_t *cand_pte =3D table + pte_index(cand->addr); + enum scan_result result; =20 if (cand->state !=3D CAND_SELECTED) continue; =20 - cand->result =3D collapse_check_candidate(vma, cc, cand, cand_pte); - if (cand->result !=3D SCAN_SUCCEED) + result =3D collapse_check_candidate(vma, cc, cand, cand_pte); + if (result =3D=3D SCAN_SUCCEED) + result =3D collapse_freeze_candidate(mm, cand, cand_pte); + + cand->result =3D result; + if (result !=3D SCAN_SUCCEED) { cand->state =3D CAND_SKIPPED; + continue; + } + + flush_start =3D min(flush_start, candidate_start(cand)); + flush_end =3D max(flush_end, candidate_end(cand)); } =20 + if (flush_end) + flush_tlb_range(vma, flush_start, flush_end); pte_unmap_unlock(pte, ptl); } =20 @@ -698,7 +1010,7 @@ static void collapse_provision(struct mm_struct *mm, struct collapse_candidate *cand =3D &cc->candidates[i]; enum scan_result result; =20 - if (cand->state !=3D CAND_SELECTED || cand->new_folio) + if (cand->state !=3D CAND_FROZEN || cand->new_folio) continue; =20 result =3D collapse_alloc(mm, cc, cand, gfp); @@ -1200,13 +1512,14 @@ static bool collapse_run_batch(struct mm_struct *mm= , unsigned long pmd_addr, * a single candidate is above the cap all by itself once a PMD is (512M w= ith * 64K pages), and refusing it would collapse nothing at all. */ -static bool collapse_batch_full(struct collapse_control *cc, +static bool collapse_batch_full(struct collapse_control *cc, unsigned int = slots, unsigned long bytes, unsigned int order) { if (!cc->nr_candidates) return false; =20 return cc->nr_candidates =3D=3D COLLAPSE_MAX_CANDIDATES || + slots + (1U << order) > COLLAPSE_SAVED_PTES || bytes + (PAGE_SIZE << order) > COLLAPSE_BATCH_BYTES; } =20 @@ -1215,7 +1528,8 @@ static bool collapse_batch_full(struct collapse_contr= ol *cc, * previous round's values, so every field is set here. */ static void collapse_add_candidate(struct collapse_control *cc, - unsigned long addr, unsigned int order) + unsigned long addr, unsigned int order, + pte_t *saved_ptes) { struct collapse_candidate *cand; =20 @@ -1232,6 +1546,7 @@ static void collapse_add_candidate(struct collapse_co= ntrol *cc, cand->result =3D SCAN_FAIL; cand->new_folio =3D NULL; cand->deposit =3D NULL; + cand->saved_ptes =3D saved_ptes; } =20 /* @@ -1246,6 +1561,7 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned long= start, unsigned long end, const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; unsigned int offset, order; unsigned long bytes =3D 0; + unsigned int slots =3D 0; bool pending =3D false; bool cont =3D true; =20 @@ -1256,7 +1572,7 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned long= start, unsigned long end, if (!pending) pending =3D collapse_next_candidate(cc, &offset, &order); =20 - if (!pending || collapse_batch_full(cc, bytes, order)) { + if (!pending || collapse_batch_full(cc, slots, bytes, order)) { /* * Selection is exhausted and the round is empty: the * range is done. Without this a flush of an empty @@ -1267,6 +1583,7 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned long= start, unsigned long end, break; =20 cont =3D collapse_run_batch(mm, pmd_addr, cc); + slots =3D 0; bytes =3D 0; continue; } @@ -1276,8 +1593,10 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned lon= g start, unsigned long end, * collecting costs nothing but the array slot. A candidate the * full round could not take is kept pending for the next one. */ - collapse_add_candidate(cc, pmd_addr + offset * PAGE_SIZE, order); + collapse_add_candidate(cc, pmd_addr + offset * PAGE_SIZE, order, + cc->saved_ptes + slots); =20 + slots +=3D 1U << order; bytes +=3D PAGE_SIZE << order; pending =3D false; } diff --git a/mm/collapse.h b/mm/collapse.h index 747168104a72..3256c45ee228 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -137,6 +137,9 @@ struct collapse_control { */ unsigned long batch_start; unsigned long batch_end; + + /* PTE values the round displaced, carved up between its candidates */ + pte_t *saved_ptes; }; =20 static inline int collapse_test_exit(struct mm_struct *mm) --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 848103EFFC3; Sun, 16 Aug 2026 22:46:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920414; cv=none; b=cu70oyhamweUIiGVQj/C9E9NG4EgRAqV68mGeLF0CmyguTKlyk18gHz6fM2xB+zUFQIVkVRhLX1ecYf1VZjgjnqcmX11jO7PN2y4YTDWnu9rg5JElCymrw6lhnLHhssfOfFsj1/H6mTN9hV2CrSDjTGngZbzNTm/a5s+8KuNJGQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920414; c=relaxed/simple; bh=AAv8HDLt8yoFB9hlA4V4PCnVkSVozL46bnU1yZP1v8k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=U1lp/bzpyrqEAsCfeZgBUuIEr5jAdbs9JakHke4xIojNQ1Q345U4cN/jSLzc5l9YjNRiX/WzxABM/uMKhzRPicctgHrolf3YEukeA2AlU0Q2JiF+L0Va5C0/KjIdpdv5mMvNbDDFJSKu38yH7oTDyIFgayf2Pk7lrutUm1vuUyI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=jLMSO2l2; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=AN+XvE2s; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="jLMSO2l2"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="AN+XvE2s" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id BC4A214000F8; Sun, 16 Aug 2026 18:46:51 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:51 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920411; x= 1787006811; bh=R4W786WotZ/WyKgbYqzGFWEVUdrtVlmGMAxyFOyzC0s=; b=j LMSO2l2Tkn4IFDjiBFITPmcvC6gilrwaaIUKad64wuKspAqfgKmPLxgHAtwpa0hd Cowwi8p0wmo+w5zBptY8Ar+2SYh6//rz9dekEMx4YM1ZaVOriFY3an9e9pymOv+1 YofpPsXPcvoYHiOzqZWKpKup6k/TZkYlYUaNCacAi8LxUdpP84uzZCGZrkr+0mAd R0Tpi0EsdDPp+GtcEbsqHxPGWXADj/Hevv9aOrFShhPfPkhfBhzdpm0Vd5HdTI/4 K9WAEMZMBV2TOqaPEKjVHO3x2pLVEqMrKyZKOdr3KPoPezK+jKFoMzHLT9jRv4gD bHzZ2gjchcdyi3WtcFZWw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920411; x=1787006811; bh=R 4W786WotZ/WyKgbYqzGFWEVUdrtVlmGMAxyFOyzC0s=; b=AN+XvE2soozc02UQ6 wFmVllwTbQ5Hiyh809UHjROYJ/bfJWm0R+R1YrJ6BMm1fga1TsbOfKwzaMdWKko4 GmFoiO1TDchv7Gaoi7vs/pE+GqZKxfdE3ZAkzxZh2HWUjleFBFkBqYQcEdGSKJXg qhhaIHDp4B/RE+5pid3tGDoMxXtqwW/4drWZ90ZiIvDFthbkSjAH8kybytaO2mTG aGqQg01TnMpLOtoV4fY/qTjS6OPwALS7iCEccaVQ1Wlf68hGyCAEW6lyO0CEjGCc cQjpQNncbF3uVOUbTvHYDwk1TX/R5KYv+zCq3gc4cmZagzB5xnLGenLWzaQYQfTE GU4BA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7Uq ItwriLCF+51TbxcwytB+TuK/breu65MYkkuLOzLs+ujqeaqn/pevMWVPqxiCoRFuWTt1yI TxkkSazc/sdb4Gg8JwC+VYHvOTdGOHidP4yQvZ0WzHzRxfXn3JCvWJI5/w7HnYISzKfmf4 lx2xOWgTIQD4t8lfY1ZHKlwyJBrXBk7m1l4H44gDmSDmr01V9EJgxP9vje8PjesT29WW9U Wn1swMWva1/FrKKdWnLsKPlXMe2VAw53yMdNbCjFGXO01sPl+lPJIorm93GO0kVZaGXukD Y0IOMQTB7hYFxrqc7rQT1O08VIOytVqdHlzChlJ5Y0HyJoZFd8c5dXbdQXxg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:51 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 17/57] mm/collapse: copy the sources into the destinations Date: Sun, 16 Aug 2026 23:45:29 +0100 Message-ID: <20260816224609.308019-18-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the copy. For each candidate that both froze and has a destination folio, the page mapped at slot k of its window becomes page k of that folio. A slot with no source -- a hole, or a zeropage the freeze cleared -- is zero-filled instead. Neither condition implies the other. A reserve before the lock can leave a folio on a candidate that then fails to freeze, and the provision inside the window can decline one for a candidate that froze. Nothing else is carried across: the destination's PTE bits are not derived from the sources, and the install rebuilds them the way a fault would. A machine check reading a source is the only failure the copy can report, and only where the architecture provides an MC-safe copy. The candidate records SCAN_COPY_MC and the install undoes it, that being where the ptl the undoing needs is held. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 40 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 40 insertions(+) diff --git a/mm/collapse.c b/mm/collapse.c index cf2b9b3640ae..a3882d897d11 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1040,6 +1040,46 @@ static void collapse_provision(struct mm_struct *mm, static void collapse_copy(struct vm_area_struct *vma, struct collapse_control *cc) { + unsigned int i; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + const unsigned int nr_pages =3D candidate_nr_pages(cand); + unsigned long addr =3D cand->addr; + unsigned int k; + + /* A folio does not imply a freeze: reserve runs before the lock */ + if (cand->state !=3D CAND_FROZEN) + continue; + + /* A freeze does not imply a folio: provision may have declined */ + if (!cand->new_folio) + continue; + + /* Each source lands where its address puts it: slot k, page k */ + for (k =3D 0; k < nr_pages; k++, addr +=3D PAGE_SIZE) { + struct page *dst =3D folio_page(cand->new_folio, k); + struct page *src; + + /* No source: a hole, or a zeropage the freeze cleared */ + if (pte_none_or_zero(cand->saved_ptes[k])) { + clear_user_highpage(dst, addr); + continue; + } + + src =3D pte_page(cand->saved_ptes[k]); + + /* + * A machine check on a source is the only way this + * fails, and the install is what undoes the candidate: + * that is where the ptl the undoing needs is held. + */ + if (copy_mc_user_highpage(dst, src, addr, vma)) { + cand->result =3D SCAN_COPY_MC; + break; + } + } + } } =20 /* Publish each destination folio in place of the sources it replaces */ --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 25D333E8C64; Sun, 16 Aug 2026 22:46:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920416; cv=none; b=UlgZyArzQ++iBDQKTiPKXrUD6bscJ0YmoKUSMk39Yu/lVuVlpMMzJXEeRzyuMI8k2bK/ICO0Il2oLX5t2lgZ76hXxF6h01NRW3gMzCt34Siv7jRSCaM4zqsrBKSwV17zDBLmadLbvy5GEpDqXgIBuE1mNSbuFm+wPF4tmZ+JYyI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920416; c=relaxed/simple; bh=p5aYzYvIN3rJ2V10UBoXgQO98iLdbTLdlyC1OIFJX0k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hp53ATL6Sldg5pIQWnBXBrs/qnYSm/m7873nkEV2SwmzUoEjzTlvrfnAo01P2vb9TJjSKRvXZxDH3BsYUXx4RjfC8CHcIB0LPk1KNJYiHWL5CBe4vLFqCQvZi3yGUsMTAvUfMqiWRkj27TeuzS9iZx4sKX5l//BbTZ3kIdGyPcU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=A7nHQUNW; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=VJe0v8d8; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="A7nHQUNW"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="VJe0v8d8" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 783DB14000FD; Sun, 16 Aug 2026 18:46:53 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:53 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920413; x= 1787006813; bh=bt/1DtqAT/rmlcUJsAd7Y7VJ/qj4oL01k7OngiisgMs=; b=A 7nHQUNWeeK6LwXx1dXKoJVpZq9SxvzD7JjuDWIqNRZJgGwdlcZNogrZDHOc62aCm 9X49B4enpi/nVLJ3FV3Fy8fnVTIZcysdQfXO8+mJ6dBjpnxG5Y7POYIlAsiV6DMp lVhA8e7CVGx2v+TMZ3DgPX2G7SDmptIB28D05F6U7W45AIGuDAcHrtXV89uUSjtl 9iPunvkzUZ06yv+ZGc+hguZmFClKzJ5rPTCh/MfYmWeepS4i3/pw1KSWf2M1ZfT2 eWlT7YQ7BDVbeE4yN47SJkaiU7c//FwFZ1sJ0SAfryFqcQL1itb6ZAjjVrpuB+3e 00PzrWkEPExcogV7Z8CUA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920413; x=1787006813; bh=b t/1DtqAT/rmlcUJsAd7Y7VJ/qj4oL01k7OngiisgMs=; b=VJe0v8d8Lfz19Q2jd Ykl0QKJwzAPimFIJxtf6xYOw8WywgdGNuj4tHaI9N6/FV9fA0Y82OsW2R5uFb0vx uDe+PeJOIxoF4PFK6u6+NSJUMifLREsRpeWMfCc9yg7n9cgwo4zfZGtyp3aarzHc THu0RgDq4vqmU6akcsDPps+Xz1l3KNLyrroXRVBKXUXFa9V0yZP5rJmnpm8rA1vO Uv3ZhhHJuxnPU3PnA+WRJsGK3JY/11yJSRIrmY7DsDdRXKagyemHNvj+t3B/+vRt XFHbVrMGiWBCnKqWPMNWUTWOEiG7HyWaW/Nk6pfSmJXGWbaHrvMrRmwka3nK/OuO 2eS6A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7Sa oM4s3kdxff6tSMV/KSNRoNHe1vOGvdBaQHDPRfcXMYtPRiFroRQvfuDcmb5v1yCxJWEnfO /WmyTZQvMyyPBnEmifuOiBwpGWBSUhY0WR8/MJ34EsN3oddj+bXWoDdXMw2D/cHFi1bnI0 CO7GHKJTqzIij4V1D0Oodu6+BZfTerMwggOGzyrZ8NL1lA7ilN6fLJXYiGSQQkRbIi2Var GE2q6NHWdzcvB+r+8KC06irB/u6W/vRiy7FZ8SIO3J2jy+CYXdsvFAEhb1H07a0TRm/j3G lZks2zzWHZy7xeuWtHbYdLFyCgPb2iXFm3eFsl984HbzWvAOJ3nUMnAJVWPQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:52 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 18/57] mm/collapse: install the destinations at PTE level Date: Sun, 16 Aug 2026 23:45:30 +0100 Message-ID: <20260816224609.308019-19-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the install for sub-PMD candidates, and with it the two things every install needs: a verify, and an abort. The verify decides whether the window is still the round's to publish. Under the ptl, every slot that had a source must still hold the round's migration entry, and every slot that had none must still be none. A hole some fault refilled is not the round's to overwrite. Verify and install share one ptl hold, so a verified candidate cannot lose a slot before it is published. The abort undoes a frozen candidate that cannot be published: the copy took a machine check, the provision pass could not spare it a destination, or the verify refused it. Slot by slot: - a slot still holding the round's migration entry is restored from the saved value. No TLB flush: the translation is identical. - a slot that does not is left exactly as found, since restoring it would resurrect memory the user zapped. Its rmap is dropped here, because the zapper fixed up rss for what it cleared but could not drop the rmap a frozen source keeps. The table itself going away is that same rule at whole-table scale. A racing MADV_DONTNEED over the whole table, and the empty-table reclaim behind it, can free the table between freeze and install. Every slot then reads as foreign, and what is left to undo is exactly the half of the teardown a zapper cannot do for a frozen source. Publishing is the fault path's own helper, so the destination gets rmap and LRU insertion before its PTEs, fresh bits from the VMA, contpte painting, and no TLB flush -- every transition is non-present to present. Slots that had no source become anon memory no zap ever accounted for, so rss is corrected by hand. PMD-order candidates need a terminal layer of their own, a stub here. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 226 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 226 insertions(+) diff --git a/mm/collapse.c b/mm/collapse.c index a3882d897d11..842adc30aeb0 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -129,6 +129,7 @@ enum collapse_candidate_state { CAND_SELECTED, /* collected; nothing held on its behalf yet */ CAND_SKIPPED, /* refused; nothing of it left to undo */ CAND_FROZEN, /* sources displaced and frozen */ + CAND_INSTALLED, /* the destination is mapped */ }; =20 /* @@ -1082,10 +1083,235 @@ static void collapse_copy(struct vm_area_struct *v= ma, } } =20 +/* + * Undo one frozen slot: restore the saved PTE if our migration entry is s= till + * there, or drop the rmap the freeze took if a racing zap already replace= d it. + * Returns true when the slot was zapped -- its mapping reference is then = ours to + * release. + */ +static bool collapse_abort_slot(struct vm_area_struct *vma, struct folio *= folio, + pte_t *slot, unsigned long addr, pte_t saved) +{ + /* No table left: the slot cannot still be holding our entry */ + if (slot) { + softleaf_t entry =3D softleaf_from_pte(ptep_get(slot)); + + if (softleaf_is_migration(entry) && + softleaf_to_pfn(entry) =3D=3D pte_pfn(saved)) { + set_pte_at(vma->vm_mm, addr, slot, saved); + return false; + } + } + folio_remove_rmap_pte(folio, pte_page(saved), vma); + return true; +} + +/* + * Abort one frozen candidate at install time: it took a machine check dur= ing the + * copy, or some of its slots no longer hold our migration entries. mmap_= read + * (held freeze..putback) blocks fork, mremap and munmap, and faults wait = on the + * migration entries -- but madvise-class operations run under mmap_read t= oo, so a + * concurrent MADV_DONTNEED may have zapped frozen slots, and a fault may = have + * refilled a zapped one. + * + * Slots still holding our entries are restored from the saved values (no = TLB + * flush: identical translation). Foreign slots are left exactly as found= -- + * restoring them would resurrect memory the user zapped -- but their rmap= is + * dropped here: the zapper fixed up rss for the slots it cleared, yet cou= ld not + * drop the rmap a frozen source keeps, unlike a migrating one, which unma= ps at + * freeze time. Slots with no source follow the same rule with no rmap to= drop: a + * cleared zeropage is restored only while its slot is still none, and a h= ole was + * never touched at all. + * + * @pte is NULL when the table itself is gone: a racing whole-table MADV_D= ONTNEED + * zapped every entry, frozen slots included, and the empty-table reclaim + * (CONFIG_PT_RECLAIM) freed it, clearing the pmd under the pmd lock and t= he pte + * ptl, neither of which excludes it between our freeze and install. Ever= y slot + * then reads as foreign, which is exactly right: nothing of ours survives= to + * restore or verify, and what is left is the half of the teardown the zap= per + * cannot perform for a frozen source -- the kept rmap, the freeze, the fo= lio + * locks and the references. The caller holds no page-table lock in that = case, + * there being no table to lock. + */ +static void collapse_abort_candidate(struct vm_area_struct *vma, + struct collapse_candidate *cand, + pte_t *pte) +{ + const unsigned int nr_pages =3D candidate_nr_pages(cand); + struct mm_struct *mm =3D vma->vm_mm; + unsigned long addr =3D cand->addr; + unsigned int i, nr; + + for (i =3D 0; i < nr_pages; i +=3D nr, addr +=3D nr * PAGE_SIZE) { + pte_t saved =3D cand->saved_ptes[i]; + unsigned int k, nr_dropped; + struct folio *folio; + + nr =3D 1; /* skip stride; a span overrides it */ + if (pte_none(saved)) + continue; + if (is_zero_pfn(pte_pfn(saved))) { + if (pte && pte_none(ptep_get(pte + i))) + set_pte_at(mm, addr, pte + i, saved); + continue; + } + + folio =3D pte_folio(saved); + nr =3D collapse_saved_span_len(cand, i, nr_pages); + + /* + * Unfreeze before any rmap drop: rmap removal munlocks under + * VM_LOCKED, and munlock_folio() takes a reference a frozen folio + * forbids. The expected count still holds every slot's mapping + * reference; restored slots keep theirs, and the zapped slots' + * references become ours to drop with the rmap. + */ + folio_ref_unfreeze(folio, folio_expected_ref_count(folio) + 1); + + nr_dropped =3D 0; + for (k =3D 0; k < nr; k++) { + pte_t *slot =3D pte ? pte + i + k : NULL; + + nr_dropped +=3D collapse_abort_slot(vma, folio, slot, + addr + k * PAGE_SIZE, + cand->saved_ptes[i + k]); + } + + folio_unlock(folio); + folio_put_refs(folio, nr_dropped + 1); + } + + /* + * Not installed; collapse_finish() releases the destination, which has to + * wait for the ptl to be dropped. + */ + cand->state =3D CAND_SKIPPED; +} + +/* + * Nothing may have shifted under the round: every source slot must still = hold our + * migration entry, and every slot with no source must still be none -- th= e freeze + * cleared the zeropage ones, so both read as none by then, and a slot som= e fault + * has refilled, with a page or with a zeropage, is not ours to overwrite. + * @nr_populated returns how many source-less slots the install is about t= o make + * present, which is rss no zap ever accounted for. + */ +static bool collapse_verify_candidate(struct collapse_candidate *cand, + pte_t *pte, unsigned int *nr_populated) +{ + const unsigned int nr_pages =3D candidate_nr_pages(cand); + unsigned int k, populated =3D 0; + + for (k =3D 0; k < nr_pages; k++) { + pte_t live =3D ptep_get(pte + k); + softleaf_t entry; + + if (pte_none_or_zero(cand->saved_ptes[k])) { + if (!pte_none(live)) + return false; + populated++; + continue; + } + + entry =3D softleaf_from_pte(live); + if (!softleaf_is_migration(entry) || + softleaf_to_pfn(entry) !=3D pte_pfn(cand->saved_ptes[k])) + return false; + } + *nr_populated =3D populated; + return true; +} + +/* + * The PMD terminal layer: verify, detach the table, deposit a fresh one a= nd + * install the leaf, as one atomic section under the pmd lock. A pmd_none= window + * never exists -- faults stay held at pte level by the migration entries + * throughout -- which is what lets PMD collapse run under mmap_read like = the rest + * of the engine. + */ +static void collapse_install_pmd(struct vm_area_struct *vma, + struct collapse_control *cc, pmd_t *pmd) +{ +} + /* Publish each destination folio in place of the sources it replaces */ static void collapse_install(struct vm_area_struct *vma, struct collapse_control *cc, pmd_t *pmd) { + struct mm_struct *mm =3D vma->vm_mm; + pte_t *pte, *table; + spinlock_t *ptl; + unsigned int i; + + if (is_pmd_order(cc->candidates[0].order)) { + /* A PMD candidate fills the slot pool: always alone */ + VM_WARN_ON_ONCE(cc->nr_candidates !=3D 1); + collapse_install_pmd(vma, cc, pmd); + return; + } + + pte =3D pte_offset_map_lock(mm, pmd, cc->candidates[0].addr, &ptl); + if (!pte) { + /* + * Table gone under us (see collapse_abort_candidate() on @pte). + * Tear down every frozen candidate -- stranding them would leak + * frozen, locked sources. + */ + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + + if (cand->state !=3D CAND_FROZEN) + continue; + + cand->result =3D SCAN_NO_PTE_TABLE; + collapse_abort_candidate(vma, cand, NULL); + } + return; + } + table =3D pte - pte_index(cc->candidates[0].addr); + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + pte_t *cand_pte =3D table + pte_index(cand->addr); + unsigned int nr_populated; + + if (cand->state !=3D CAND_FROZEN) + continue; + + if (cand->result !=3D SCAN_SUCCEED) { + /* Machine check during the copy */ + collapse_abort_candidate(vma, cand, cand_pte); + continue; + } + + /* No destination: the provision pass could not spare one */ + if (!cand->new_folio) { + collapse_abort_candidate(vma, cand, cand_pte); + continue; + } + + if (!collapse_verify_candidate(cand, cand_pte, &nr_populated)) { + cand->result =3D SCAN_PTE_NON_PRESENT; + collapse_abort_candidate(vma, cand, cand_pte); + continue; + } + + /* + * The smp_wmb() in __folio_mark_uptodate() orders the copied + * data before the set_ptes() that publishes it. + */ + __folio_mark_uptodate(cand->new_folio); + map_anon_folio_pte_nopf(cand->new_folio, cand_pte, vma, + cand->addr, /*uffd_wp=3D*/ false); + + /* Slots with no source gain anon memory that no zap accounted */ + if (nr_populated) + add_mm_counter(mm, MM_ANONPAGES, nr_populated); + cand->new_folio =3D NULL; /* ownership: the mappings */ + cand->state =3D CAND_INSTALLED; + } + + pte_unmap_unlock(pte, ptl); } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2DF2A3E8351; Sun, 16 Aug 2026 22:46:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920418; cv=none; b=bddbJHmQzxEAd7CRYlQoFuHoKcF66QkkDvSjj/b9pDg2ZsjkGNT6rCUE7VFM20YfkIoQNR9JSD4q0DjPNmJ87fTAEIy5QQaXUjTtgCsqgw44LMFua3CLjeYk0BTu+Frbiq3ldUeujRfna3QdgErqjKK53lkxYV362R1jOpg0Yfc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920418; c=relaxed/simple; bh=2TaYaTaJG3DUEb+2Yt76ZzAIWLcNA7Ing2Udv2Pjyjc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uE+hf0FtC8/05lViWnPDQGUYt3RKH24/lJ0qKObl+bVhyNhuFcF9VCdqvvUK0IJBiAPJCB5xWpSVGs5ah69XpPWstKSeli5UN8ieeL7lAor8IQEohk9+vW9XK4Bm3+rdXMQX6eKpDGbpvjxeAEzJ9Fiweb7osYS7/iTLCRE7xYo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=sKocDFEW; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Pn8FC+cK; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="sKocDFEW"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Pn8FC+cK" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.phl.internal (Postfix) with ESMTP id 3831214000FE; Sun, 16 Aug 2026 18:46:55 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:46:55 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920415; x= 1787006815; bh=QHqKPBFVnTs29f21vr16ac6OAqWm4F4NnweEXoqIRdA=; b=s KocDFEWxrud+mkbmUPVtPbiu+43Yn2FKamFIEZsI8h1/8zMn+TOqdH179ImZq0r6 cWYLXuGvF1fkFRAMw4ywLXI+WzpZcgVRCKxenmyXvhuDoXTXJXsA9giezG4P5EfC kvOWiCXvcK3ShC6M7rvmIKTUWXLmjF1JntkGSBWhMTWESJRubIGqMQMcqtRWRaL8 9peyASLhyFQinM7O7SL7VGdgkjfHx4lPSK3KXyx4bBUYNyFmzhrQmx6Cg5Dr+aMh rDPQ4jPr/FWiaQ6xmdcaIRMUYuCycuG1ESU2mCHfYEeod50x5/CwPCT0BCaTj61k ZTB3rAabkvnyWmH3sM+Bg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920415; x=1787006815; bh=Q HqKPBFVnTs29f21vr16ac6OAqWm4F4NnweEXoqIRdA=; b=Pn8FC+cK/ZMRgElW7 Y2ryQGC/MlYol41rOcraOqvuS9tz98JDGzwqcbpAN6+cKxWG1j2SGamTMSQVKThD La/KjzbQXBrOfCGFhuvPaOiD1uZYMpPEo7m86OchlgTRnbeuayAKCcOOEpJpPFt5 XI7W8XMlZg21XxckGOwcigBZBtAGfIwjtrc4vORVl2Aro+aPusKxCJzoULKTIW76 sHInn6i8JGuaQDvSZGcuDEMUgyU1tNZKrtiJosDCJeUrSXwF41nK8zrM8l32Sclh 9P/gAxclYtIliRvNM4e/pMEN8wUTumUoIl5D5Jqu3yChwWC7rMlWJaA1/7BhzbE+ uKeQg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7CF ULG6OtC6a1s/NmdMZEoDsIJcrpBkdRdhXB11gNV50z2d4E/iJeXml3T3h/o0H7gU+64vOP PqPDa4EdcFe/rGNcm71i1IAqWlLhpil33moV0JBcgZQm9qfgBsqf0h59iIN1pPwwFe7xeZ P7twIKhdh+uX1rwSO+SB2Ua/tk/l9fjmZe+S/3gFTdptk2aaquAFLmgWmNltN4ZZAqkppG LpcqV7O4OtJwJTFWHt0SRtr6RBcJbb1AY6je6ZAylypaQl8WN8RmScIh8gpwmMGiCcQ2r8 fuizZBKpxbm2PItZCAiLKBTcdO/z5AopY1nNAlm7k36xkZfTwfcgPCHgB6jQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:54 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 19/57] mm/collapse: install a PMD leaf as the terminal layer Date: Sun, 16 Aug 2026 23:45:31 +0100 Message-ID: <20260816224609.308019-20-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the PMD install. Under the pmd lock, with the pte ptl nested inside it: verify, detach the table with pmdp_collapse_flush(), deposit a fresh one and map the leaf. That is one atomic section, so no pmd_none() window ever exists: faults stay held down at pte level by the migration entries throughout. It is what lets PMD collapse run under mmap_read like everything else here. Two things force that nesting, which is the one the tree already uses to reinstall a table. A racing zap of a frozen entry takes the pte ptl, so the verify has to hold it. And the table must not come apart between verify and detach, which is the pmd lock's job. Nothing leaves the section early, aborts included. An abort only restores PTEs and would need no pmd-level exclusion of its own, except that its pte pointer came from pte_offset_map_rw_nolock(), whose caller must establish that the pmd is stable. The deposited table is the freshly allocated one, never the table just detached. A deposited table has to be quiescent, because whoever withdraws it frees it immediately with nothing to hold a lockless walker off first, and a table that has never been reachable is quiescent by construction. The detached one is not: GUP-fast and RCU pte walks that read the old PMD may still be inside it, and on broadcast-TLBI architectures the flush expels nobody. Quiescing it would need an IPI, which has nowhere to go here -- outside the pmd lock it opens the pmd_none() window this design does not have, inside it is a broadcast under a spinlock. So the detached table goes to pte_free_defer(), which holds the free until those walkers finish. One transient table page per PMD collapse is the cost. No anon_vma_lock_write() is taken, unlike the mechanism being replaced: - rmap walks on the sources are unreachable, their refcounts frozen and their folio locks held from freeze to putback; - non-rmap pte walkers see migration entries; - pmd-level observers see either the old table or the leaf, never an intermediate; - fork, mremap and munmap take mmap_write, which the mmap_read held here excludes. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 118 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 118 insertions(+) diff --git a/mm/collapse.c b/mm/collapse.c index 842adc30aeb0..ab7476471b8d 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1232,6 +1232,124 @@ static bool collapse_verify_candidate(struct collap= se_candidate *cand, static void collapse_install_pmd(struct vm_area_struct *vma, struct collapse_control *cc, pmd_t *pmd) { + struct collapse_candidate *cand =3D &cc->candidates[0]; + struct mm_struct *mm =3D vma->vm_mm; + spinlock_t *pmd_ptl, *pte_ptl; + pgtable_t old_table =3D NULL; + unsigned int nr_populated; + pmd_t old_pmd, pmdval; + pte_t *pte; + + if (cand->state !=3D CAND_FROZEN) + return; + + /* No destination: the provision pass could not spare one */ + if (!cand->new_folio) { + pte =3D pte_offset_map_lock(mm, pmd, cand->addr, &pte_ptl); + collapse_abort_candidate(vma, cand, pte); + if (pte) + pte_unmap_unlock(pte, pte_ptl); + return; + } + + /* + * The pte ptl nests inside the pmd lock, the nesting the tree already + * uses for reinstalling a table: a racing zap of a frozen entry takes + * the pte ptl, so the verify must hold it, and the table must not come + * apart between verify and detach. pmd_same() rechecks are unnecessary, + * the pmd lock being held across the whole section. + */ + pmd_ptl =3D pmd_lock(mm, pmd); + pte =3D pte_offset_map_rw_nolock(mm, pmd, cand->addr, &pmdval, &pte_ptl); + if (!pte) { + /* Table gone under us; see collapse_abort_candidate() on @pte */ + spin_unlock(pmd_ptl); + cand->result =3D SCAN_NO_PTE_TABLE; + collapse_abort_candidate(vma, cand, NULL); + return; + } + if (pte_ptl !=3D pmd_ptl) + spin_lock_nested(pte_ptl, SINGLE_DEPTH_NESTING); + + /* + * Every exit is inside that section, the aborts as much as the install. + * An abort needs no pmd-level exclusion of its own; it only restores + * PTEs. But the table it works on came from pte_offset_map_rw_nolock(), + * which leaves its caller to establish that the pmd is stable, and the + * held pmd lock is what does that here. + */ + if (cand->result !=3D SCAN_SUCCEED) { + /* Machine check during the copy */ + collapse_abort_candidate(vma, cand, pte); + goto out_unlock; + } + + if (!collapse_verify_candidate(cand, pte, &nr_populated)) { + cand->result =3D SCAN_PTE_NON_PRESENT; + collapse_abort_candidate(vma, cand, pte); + goto out_unlock; + } + + /* + * Nothing fallible sits past here. No anon_vma_lock_write either: rmap + * walks on the sources are unreachable -- refcounts frozen, folio locks + * held from freeze to putback -- non-rmap pte walkers see migration + * entries, pmd-level observers see the old table or the leaf and never an + * intermediate, and fork, mremap and munmap take mmap_write, which our + * mmap_read excludes. + * + * The flush inside pmdp_collapse_flush() is the round's second over this + * range: the freeze displaced every leaf here and flushed before dropping + * the ptl, and the verify above proved nothing has been mapped since. + * What it covers is the paging-structure caches -- a CPU may still hold + * the pmd-to-table link, for a table that is about to be freed -- which + * is why the helper shoots down a pte range rather than a pmd. + */ + old_pmd =3D pmdp_collapse_flush(vma, cand->addr, pmd); + old_table =3D pmd_pgtable(old_pmd); + + /* + * The smp_wmb() in __folio_mark_uptodate() orders the copied data before + * the install below publishes it. + */ + __folio_mark_uptodate(cand->new_folio); + + /* + * Deposit a freshly allocated table, not the one just detached: a + * deposited table has to be quiescent, because whoever withdraws it frees + * it immediately (zap_huge_pmd()) with nothing to hold a lockless walker + * off first. A table that has never been reachable is quiescent by + * construction, which is why collapse_alloc() secured one. + * + * The detached table is not. GUP-fast and RCU pte walks that read the + * old PMD before pmdp_collapse_flush() may still be inside it, and on + * broadcast-TLBI arches that flush expels nobody. Quiescing it would + * take an IPI (tlb_remove_table_sync_one()), which has nowhere to go + * here: outside the pmd lock it opens a pmd_none window a fault can fill, + * inside it is a broadcast under a spinlock. So it goes to + * pte_free_defer(), which holds the free until those walkers finish, as + * retract_page_tables() does. One transient table page per PMD collapse + * is what that costs. + */ + pgtable_trans_huge_deposit(mm, pmd, cand->deposit); + map_anon_folio_pmd_nopf(cand->new_folio, pmd, vma, cand->addr); + + /* Slots with no source gain anon memory that no zap accounted */ + if (nr_populated) + add_mm_counter(mm, MM_ANONPAGES, nr_populated); + cand->deposit =3D NULL; + cand->new_folio =3D NULL; /* ownership: the mapping */ + cand->state =3D CAND_INSTALLED; + +out_unlock: + if (pte_ptl !=3D pmd_ptl) + spin_unlock(pte_ptl); + pte_unmap(pte); + spin_unlock(pmd_ptl); + + /* The deposit balanced the detached table, so the count is already right= */ + if (old_table) + pte_free_defer(mm, old_table); } =20 /* Publish each destination folio in place of the sources it replaces */ --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E4CF33EFFC3; Sun, 16 Aug 2026 22:46:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920419; cv=none; b=Pr5lXkRSklYBvh9IVoQ+FJlV2Gduzz1Si5Rbzcx8sp7tsiLNTIxEVDhRrhyNFt3i+5pYVv46909+5KJYockAmIITJEKUMpA2aTaeZwwqWk2YQV/rlZN1P9BirXUwwxVC0h6HGd/H9VaaX0EajYQOTyKLWjyFHxXuOHN4tgy5tKc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920419; c=relaxed/simple; bh=rkYUDi4omK53ksGFAid26gCCWZwhOAmjSe1EwS0KNXQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hI1DZixn+39bF5GgqruAl9+BkX1xMjxsdzkvF8XtkaKXs0iKUngd6NKhUOj3D61eQfYlI2hok+noSR6dI9BhWbXA5lpfM6rrtn3mXc2oMf7CxVkpELJevWQWWpdD1lgL2d1bccGyGa9T7PgyU3mCJyb0t6UJ2+Ed8uPfTzCZrpY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Ump2JR5L; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=VJUR1o9r; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Ump2JR5L"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="VJUR1o9r" Received: from phl-compute-12.internal (phl-compute-12.internal [10.202.2.52]) by mailfhigh.phl.internal (Postfix) with ESMTP id 44C3E1400100; Sun, 16 Aug 2026 18:46:57 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-12.internal (MEProxy); Sun, 16 Aug 2026 18:46:57 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920417; x= 1787006817; bh=EQu6T+lgU+FFDqaM3vM2kEMbX9/iIUBJ0iT3p7UidZk=; b=U mp2JR5Lu9hjjdwu6/ZQWvKVFtdYnlCBZG+uicCmp+SUg/360NTTX1awYM5jaoqq2 rBY/rWk3tYkopKExR3nFX6CXwEIDu3GqAYgcvF5HSAAr9p0jKEgvP06bLA870RJO C/V99/ZpKmrxhDz+JR9cuFEYDKptM6bu5e/aNUTftGd6h4I5v00obUcT6eAs3ZKl 6KEHWUv7yieowRBiie8GeA+ivB03mFSfnzN7C5Y6ssAW98kuIVppF/t6ZU1ODjbQ f//EjGkc2FNOpg4mdb3pZYgkLl7K+EAHmWWb6wdBQXci1w7Z7MF+3aihS3zjHmBj vMVxRKxXiMYB69T1wx43Q== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920417; x=1787006817; bh=E Qu6T+lgU+FFDqaM3vM2kEMbX9/iIUBJ0iT3p7UidZk=; b=VJUR1o9rvJbGtzh3V VBs2bRwDDvy0aLsr2PnsD192uCbrJwf4UDphyR3f4nds3mKvDMF/Qkn2Dv8qopVU sggwudKzLaKJnsjzhDdPBMkEtZUZVPTxEhNQkQVVjHDHvMtHh4kPROHRGHshkJTJ PFgAv6Qp82RR6b5GylgrOe8Sf7OJ0eeSMdoFqB7bGAVZJetzgrFciY5o8bjBvRo0 GfjAJsrkScVza4YoCtb1PA6UR3CvIDX7iT5cKq/1c/V5G7HPY0hvFL7pGPGfWfjp EkHJP84c0Q6bQY8V2p7bwKZLv0Wcye+gZ7GoqcBheq5L4c4yIiFqFvJwNTOf4V6S PhD1Q== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGj1BuhETEStCjbN/L6JrGmiaqzRpeX6rRNWhHsAzwzRlY9AKzqCvfWqke6XA7GOE A6FFYRRSvteD9S264L5lew1YUpPhGI3TlhRe1mYLYRUsBMnidhJP2r2PMm5Lpq2ztO8wmr yF1ddTX1W5NBA6qs0+EDXUaWwTs6UQPsEiicZ6SmY6z0vOVnZaAe+3gsuR1dg3g0gP4peZ U6KBKKlAxat5APGrEFg0r1ubDi9svGZRyISInygLcwTo8cXqakBbn35auVMrKeg7PKmswS qkEQEJGI+h4SGQWw8a8xHcXO10CUTBpRtdxlbIKv4nbtwWiZJ1jrVLIQttQS09NS+fWjUx bbU18ME9VB5GU2gZOkOiPzSzfvLBB2gKm3VPT4Jvs4WJlASKQCrcJfkt1KsUDLDw5GtOlj 5h1b8gVcH+y8R4+XAP4fykvX9OA01NzIsYP1flpH2DxLcDCL/Z74+wcUZ3/rUdosCzCYPa 4QOUrJOowaj8VVpAyzhIuwB6AzFzUdB5E0IHAQgBsmOZddb9BEIKS+sLlBCjYjgErEhYF+ UTKZzgA+PYxdRSZgiTUWwWkTaCL28GJwHzQ3SUyJ0fGZmD2knP1midXv4+VODV1YSbH3HT dK/V64rcrhNHP3t/5YY3nBM7Po8YEWEv2Ay2k37+bp3BgltPHNNAT6tlhQuQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:56 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 20/57] mm/collapse: put the sources back Date: Sun, 16 Aug 2026 23:45:32 +0100 Message-ID: <20260816224609.308019-21-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the putback: for every installed candidate, lower the barriers the freeze raised, span by span. This is also what wakes the faulters the collapse held up. They sleep on a source folio's lock; once it is dropped they refault and find present PTEs pointing at the new folio. The order within a span is important: - Unfreeze first. Rmap removal munlocks under VM_LOCKED, and munlock_folio() takes a reference a frozen folio forbids. - Then drop the rmap. Until it is gone the expected count still holds the span's mapping references; afterwards they belong to the round, so every folio_remove_rmap_ptes() is paired with a folio_put_refs() for the same slots. - Then unlock, which is the wake. Holding the lock until here keeps lock-taking rmap walkers out, and the window it leaves -- a live folio with no PTEs -- is one any teardown of a mapped folio passes through. - Drop the references strictly last, the round's included. Waiters wait without a reference of their own, so the round's has to outlive the unlock. The stale swapcache entry goes too: the copy has replaced what it described. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 50 insertions(+) diff --git a/mm/collapse.c b/mm/collapse.c index ab7476471b8d..f65f413339bf 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1439,6 +1439,56 @@ static void collapse_install(struct vm_area_struct *= vma, static void collapse_putback(struct vm_area_struct *vma, struct collapse_control *cc) { + unsigned int i; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + const unsigned int nr_pages =3D candidate_nr_pages(cand); + unsigned int k =3D 0; + + if (cand->state !=3D CAND_INSTALLED) + continue; + + while (k < nr_pages) { + struct folio *folio; + unsigned int nr; + + /* A slot with no source has nothing to put back */ + if (pte_none_or_zero(cand->saved_ptes[k])) { + k++; + continue; + } + + folio =3D pte_folio(cand->saved_ptes[k]); + nr =3D collapse_saved_span_len(cand, k, nr_pages); + + /* + * Unfreeze before the rmap drop: rmap removal munlocks + * under VM_LOCKED, and munlock_folio() takes a reference + * a frozen folio forbids. The expected count still + * holds the span's mapping references; once the rmap is + * gone they are ours to drop, so every + * folio_remove_rmap_ptes() is paired with a + * folio_put_refs() for the same slots. The folio lock + * is held until the wake below, so lock-taking rmap + * walkers stay excluded, and the stale-rmap window this + * leaves -- live folio, no PTEs -- is one any teardown of + * a mapped folio passes through. + */ + folio_ref_unfreeze(folio, + folio_expected_ref_count(folio) + 1); + folio_remove_rmap_ptes(folio, + pte_page(cand->saved_ptes[k]), + nr, vma); + folio_unlock(folio); + + /* The copy replaced it; drop the stale swap entry */ + free_swap_cache(folio); + folio_put_refs(folio, nr + 1); + + k +=3D nr; + } + } } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA1FE3F3285; Sun, 16 Aug 2026 22:46:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920421; cv=none; b=BKVGM2Tnq00eYAU/0j1NFPCIyyQfSgss/9w9elmXLi9C/3idPhTh2ZymtIdeojzsrWW5pE9C70S2O2ij1hUb319jM75wgGfDnEcyYE9YpS4OQrkpZK+jelSmqcrXgEK8P+dg6F6FG4TYdGq4CA+qMr5V1+C3vIGh0mhC1W9ZSPc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920421; c=relaxed/simple; bh=qxezClCXR1oALlFl7oDnUW813LKJ437GJTADJ3u48dY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=auvk4JLudOlQIefWbQdl51J4OZkQTTDrzO7ma4Yze19tpYV9PKxEKnWT9Ak/QorW+BMceZLyHsSwsDckfDbJrTUm7BiI2qseB6VuVkEdBWAY+DBLliE7BQhTDYL8JiP1rJL76/bvokpr582Iq4zfnWUsPjQgXLGc81kDvhClxCU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=d7WK9YCR; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=GJUIGH96; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="d7WK9YCR"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="GJUIGH96" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id EF0B3EC0242; Sun, 16 Aug 2026 18:46:58 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:46:58 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920418; x= 1787006818; bh=mo9/YyuZ+27uiSkLEdlk6rzsMwO0ifUZXBUSyhK0NkI=; b=d 7WK9YCRZJRHPwU0pAQUS10+wHOnpDVzzsLoMHT4KYOSJTernWLvI5zEDM0UkfP6l sDfoB0P4LSOIrmvVWaMZU5WgPwdOZ26zGOGxOQnDXF7zH9//nENd0dCVsng3FqEi OfNRivgTRghzdRb6B7imJ3QTM3oow78pgAp7RWFlnqlPDWREvgcw7lTuNGFIMucN HHliw5wwSFtsS8pIfixDSCSHfsp2yV+7TtjaDqxBOCz9kYofyi62Najmtyg7aEcI FSR1kV7BgaHkM7t0KShfBPl5TJt3QZi4a4CzjmLD/tozU/ZJZUyg8b/oNakVgcKt 9arG4B0r60J5KjF9IK80g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920418; x=1787006818; bh=m o9/YyuZ+27uiSkLEdlk6rzsMwO0ifUZXBUSyhK0NkI=; b=GJUIGH96wLw4Ae4b/ KotBYngP01kE9QSEmFZ40LRhjCHLB0RR2+4JbCP2BK0l/nh2BSKndFsJe3TmuVr/ Uvk1WU/Z8DtRr0wK4C0ws33aKKox37XCxk05NgPvaKBYWbrmrDtzPktez79HbNmi JHctxjj5epxyWTBslwvYTB7Ti9pz2NQHCXgz9OYdXjGXA+aTU8TsreCpgNHaMWK5 LTznPv6P4TCcJjocAe0tKvz25zwiJRIW6II/5pxqD06ZMfbFHk5GatQ2wvvlCz5S qXRwGf68w9y0T4WgtoHJxrkkgdWtZuIOOEzoHb92rsS9hg/YC+ZdTZWQ4m4J5+Zn ahRpQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkhNC7YI+Ei0MrF6zu90W9yR7z714JmmtstcBEn9fnZ0JEx8KxNru7XjOIPijUAd VrNDumTIvQCtLJHAbrPKMRdbJsCYKY1AWxa28xNEDdkQaDqy3PIOJVA9R+OjsTN2M4dnB4 WQk97ejO2lmWwXR5mw6146Dl5x2vSRCf1uehpA4Z3DNP/uvzS8DmGwE+/LsFmt9BQO6ZtJ eWd0dPZ2F0WrWW1tiYo0zV8d4Xgh8ztMbcUaDgwCK69ZeQi0ebesOWu50OSj19NNGHzec3 w5R1etyOsXXy4NxsI/YxbKwNbzdBv+aWoLN5/PFLP+n740pr735uP0KJxEPqjezkxGUEYw zcznv5swWAyw6XmESrxIGM8A1bTTyyCb8TGgw+qEShLJxmw0H1QgFlfOeuxPMpjGoxHDtp PCjigTxTQQ2SQYH+g1Zjow5z90rNIRRpTNdfes5gZJF781hKtVjYAOe4if7ktPzJJA4yNg SoS74FodWm5MlganZ52ThBjBAI0zsQCzKw1kMNWmaHUySWEGqzr+6R499T3SjVD8uPcKN/ 4SLidXdUPP7GCHgXfbqC5pStWuEVCtQYOB96tZz4WWjhPZl2k+r4upntklA4ZEhZgnxts8 Lxz66rlvmgByz0pVLkrls4qWwn+V7sbwBmY/Z6EOmYn3bFnzh2lepzJSEjtg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:46:58 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 21/57] mm/collapse: settle whatever the round reached Date: Sun, 16 Aug 2026 23:45:33 +0100 Message-ID: <20260816224609.308019-22-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the last pass. A candidate that never froze is recorded as given up on, with whatever result ended the round. Anything no pass took ownership of -- a destination folio, a table meant for deposit -- goes back. The count of installed candidates is what the round reports. Holding no lock here is the point. Every refusal before this happens under a page-table lock: the freeze unwinds under the ptl it took, and the install aborts under that ptl or the pmd lock. Dropping the last reference to a folio, and the memcg uncharge behind it, is not spinlock work. So a refused candidate keeps its folio and its table until the round is over, and this is where they are released. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 25 ++++++++++++++++++++++++- 1 file changed, 24 insertions(+), 1 deletion(-) diff --git a/mm/collapse.c b/mm/collapse.c index f65f413339bf..2da1f8ddcca8 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1500,7 +1500,30 @@ static unsigned int collapse_finish(struct mm_struct= *mm, struct collapse_control *cc, enum scan_result result) { - return 0; + unsigned int i, nr_installed =3D 0; + + for (i =3D 0; i < cc->nr_candidates; i++) { + struct collapse_candidate *cand =3D &cc->candidates[i]; + + /* Never froze: the round gave up before it got that far */ + if (cand->state =3D=3D CAND_SELECTED) { + cand->state =3D CAND_SKIPPED; + cand->result =3D result; + } + + if (cand->new_folio) { + folio_put(cand->new_folio); + cand->new_folio =3D NULL; + } + if (cand->deposit) { + pte_free(mm, cand->deposit); + cand->deposit =3D NULL; + } + if (cand->state =3D=3D CAND_INSTALLED) + nr_installed++; + } + + return nr_installed; } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A09DE3F4114; Sun, 16 Aug 2026 22:47:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920423; cv=none; b=Z1QTx0OmproPd8NVrnF/QGBzzVZOPEuNobMasxnsth+x8XKRK5iz8PHfKaXOD7n2wGMpXBCMMDnzpzeZhGtgR8SZVW8IvtqTSAQ/nuhJqDLohG9/edduyGjyYdTonlIem/J3rbOgT8BjwlBe1CMaVzHd81USxOXvUTmY7H8TsJk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920423; c=relaxed/simple; bh=+hXqI4nBcQjO/KoCZ14TLbLqKBYVJhTdirQSx5Z51js=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Jbs3wJNqEaUVIRDakVHiQI1rL/w5ctWqLR28zj9jYA7XdJimTySwwN6akxPaIfxNejj/Zk8d3eF2omY5ogTzyB43k+P5gmXx/gzhcfJXMQdntoREiVMaLhg3V4mlgcpJElOJUf7CnP2jyOpHYM9RFRTBK+MMc8poTqF5OuJQVA8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Tf8ZBSBA; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=AB0csPAg; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Tf8ZBSBA"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="AB0csPAg" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.phl.internal (Postfix) with ESMTP id E7A64EC0241; Sun, 16 Aug 2026 18:47:00 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Sun, 16 Aug 2026 18:47:00 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920420; x= 1787006820; bh=J1+9D9mcuyt/yxdgXW36/YFxLEOMLmPJXeDS+6QRV2M=; b=T f8ZBSBA3TQRK3cD9jO4SD+8+5wk4KAvSmI88xPbHNQQ5uy6iP3VbK+kdBjeI3lQH 4Uc+i4JN1TGKJHDaJqLp87se4FThEzjG+MY6YJqo1ZSRYbIz54PGd597teurTpcZ TiL4MzYQNoRWKr9j4akAwU9Y7zNfm9IOVaGc69Gf4WVWgDFoVOlL/UHDdSwBTWAG DV+LyaFOKVXAheken+gJw4on7+C9BRs2RK5WwAHys4l0r/HxDuy7VKJandTVEiX3 ruO6aLt/NZYWUivJ26MQrQklgCuYO1mYxyANBTQEDz14R+sRW0tUVto6uBCgy5BN zswoaiGjXIld8YV4WXptg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920420; x=1787006820; bh=J 1+9D9mcuyt/yxdgXW36/YFxLEOMLmPJXeDS+6QRV2M=; b=AB0csPAgdvYAGp2gQ PiZ8s9TB0lKptBLgc6G8XjReoWX3VQPWod9imkjpE79kqUKaLGn80ihiDFB1OIz0 B/CNlftWgGIIm35vtT2TvU17FV5LvtxlkCrlwYLrflNfzrpEnUZyxTBsQkgCqPL2 PxTKH2EqVOV2hhQEYiwpEtRnx8n0K44p0beGRM9qHVcrNgKXqNM3Hw8V6zKg2uUf 8hSVfY18P+DiMRqCfRwFEuEUDYj3UmnUGcyXBtVHhW7CsEZ5+169BbNgrcdk8gbu X9vFIZ1sM4q2VI2IkvwLa0hg6UQKBvJIxlY979/8XvTxKQfQuRkUX35QffVinQIK OX61g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtc7 YPFXNzpcj4NYLt23UmZkN6vYCsvNuqRLXW2/cXefqbxj2Ku/EoFycRIBFGawf0PxryN+Ss t7WEunz71VUfOY4zX967A9vNhnBU7zoN7JwtQVp8mdlYrYolVL1fcnlt4bdpNKdHuwZ3+v ZxaZVO8TeJhVBOw1hQ1lmw0SyaRFM0DgZbVpmtF16qIJ/Y2yOFAOIDaAzF1wjmqhEnnqjk IOf+1hbIQ3Wdy1ombl+1sqs1iZoiq11q6UdwsaD8e5fjodHmOzN/3y40+8nouAVZJxY6iF aj0zkUU68uxiRWlFTw0RZwSEXhpsRX558YVTFYZaKy+qpBE9hovTY7LMtYzA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:00 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 22/57] mm/collapse: walk a table with a selection cursor Date: Sun, 16 Aug 2026 23:45:34 +0100 Message-ID: <20260816224609.308019-23-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the half of selection that emits candidates: a cursor over the table, handing out the largest window that fits where it stands. Two things bound the order at any point. A huge page has to be naturally aligned, so the cursor's own offset caps it -- at offset 4 nothing above order 2 can start -- and the largest enabled order caps it too. A window qualifies when enough of it is eligible: the scan's bits counted over the window, against the max_ptes_none limit for that order. One that does not qualify drops to the next enabled order below, which need not be half of it, since a sparse set of enabled sizes may skip several. When no smaller order is left, the cursor steps over the region. Only the scan's bitmap is read, so a clear bit is either a hole or a PTE the scan disqualified. Occupancy here means what a collapse could use, not what is present. Non-present PTEs the scan accepted are the exception. They are counted apart, in cc->scan_unmapped, and added back only for a PMD candidate, which faults them in; a smaller window leaves them as holes, sub-PMD collapse not reading swap. The cursor advances at emission and never rewinds. A round is collected before it is run, so within a round every attempt is assumed to succeed. Nothing here gives a refused region a second chance. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 126 ++++++++++++++++++++++++++++++++++++++++++++++++ mm/collapse.h | 10 ++++ mm/khugepaged.c | 2 +- 3 files changed, 137 insertions(+), 1 deletion(-) diff --git a/mm/collapse.c b/mm/collapse.c index 2da1f8ddcca8..258bb9cc32c5 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1841,6 +1841,7 @@ static enum scan_result collapse_scan_table(struct vm= _area_struct *vma, if (result !=3D SCAN_SUCCEED) cc->select_orders &=3D ~BIT(HPAGE_PMD_ORDER); =20 + cc->scan_unmapped =3D unmapped; return result; } =20 @@ -1852,6 +1853,7 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) nodes_clear(cc->alloc_nmask); =20 cc->select_orders =3D 0; + cc->scan_unmapped =3D 0; cc->nr_collapsed =3D 0; } =20 @@ -1896,10 +1898,116 @@ collapse_scan_anon_pmd(struct vm_area_struct *vma,= unsigned long start, return cc->scan_refusal; } =20 +/* + * Selection cuts the table into candidate windows and feeds them to round= s. A + * window is cut at the largest enabled order that fits and qualifies -- t= he PMD + * order, when the whole table qualified -- and a region that does not qua= lify is + * probed at the next enabled order below, which need not be half of it: a= sparse + * set of enabled sizes may skip several. + * + * Only cc->eligible_ptes is read, so a clear bit is either a hole or a PT= E the + * scan disqualified: a window's occupancy is what a collapse could use, n= ot what + * is present. + */ + +/* + * Largest order a window may be rooted at: the largest enabled one. + * select_orders is fixed for the table, and the caller checked it is not = empty, + * so this is well-defined for the whole walk. + */ +static unsigned int collapse_root_order(struct collapse_control *cc) +{ + return __fls(cc->select_orders); +} + +/* + * The next enabled order below @order, or 0 when there is none. select_o= rders + * never carries an order below COLLAPSE_MIN_MTHP_ORDER -- THP_ORDERS_ALL_= ANON + * masks orders 0 and 1 -- so __fls() honours that floor by itself. Order= 0 has + * no bits below it to mask and has to answer 0 outright: a walk that asce= nded + * instead would emit a window at an offset it is not aligned for. + */ +static unsigned int collapse_lower_order(struct collapse_control *cc, + unsigned int order) +{ + unsigned long lower; + + if (!order) + return 0; + + lower =3D cc->select_orders & GENMASK(order - 1, 0); + return lower ? __fls(lower) : 0; +} + /* Point the selection cursor at [start, end) of the table, in PTE offsets= */ static void collapse_selection_init(struct collapse_control *cc, unsigned int start, unsigned int end) { + cc->select_start =3D start; + cc->select_end =3D end; + cc->select_offset =3D start; + cc->select_order =3D min(max_order_from_offset(start), + collapse_root_order(cc)); +} + +/* + * Advance past the region [select_offset, select_offset + nr_ptes) and de= termine + * the highest order that can be attempted next. Since huge pages must be + * naturally aligned, it is limited by the alignment of the new offset: af= ter an + * order-2 mTHP at offset 0 the offset becomes 4, and __ffs(4) =3D=3D 2, s= o the next + * attempt starts at order 2. + */ +static void collapse_selection_advance(struct collapse_control *cc, + unsigned int nr_ptes) +{ + cc->select_offset +=3D nr_ptes; + cc->select_order =3D min(max_order_from_offset(cc->select_offset), + collapse_root_order(cc)); +} + +/* + * The window at the cursor did not qualify. Drop to the next smaller ena= bled + * order over the same region, or -- when no smaller order remains -- give= the + * region up and advance the cursor past it. + */ +static void collapse_selection_reject(struct collapse_control *cc) +{ + unsigned int lower =3D collapse_lower_order(cc, cc->select_order); + + if (lower) + cc->select_order =3D lower; + else + collapse_selection_advance(cc, 1U << cc->select_order); +} + +/* Is the window at @offset one a collapse of @order should be attempted o= n? */ +static bool collapse_window_eligible(struct collapse_control *cc, + unsigned int offset, unsigned int order) +{ + unsigned int nr_ptes =3D 1U << order; + unsigned int max_ptes_none, nr_eligible_ptes; + + if (!test_bit(order, &cc->select_orders)) + return false; + + /* The window must lie inside the scanned range */ + if (offset < cc->select_start || offset + nr_ptes > cc->select_end) + return false; + + max_ptes_none =3D collapse_max_ptes_none(cc, NULL, order); + nr_eligible_ptes =3D bitmap_weight_from(cc->eligible_ptes, offset, + offset + nr_ptes); + + /* + * Swap PTEs the scan accepted are counted in cc->scan_unmapped, not in + * the bitmap. collapse_faultin() reads them in for a PMD candidate, so + * there they do become sources; a smaller window leaves them as holes, + * sub-PMD collapse not faulting swap in. + */ + if (is_pmd_order(order)) + nr_eligible_ptes +=3D cc->scan_unmapped; + + return nr_eligible_ptes >=3D nr_ptes - max_ptes_none; } =20 /* @@ -1913,6 +2021,24 @@ static void collapse_selection_init(struct collapse_= control *cc, static bool collapse_next_candidate(struct collapse_control *cc, unsigned int *offset, unsigned int *order) { + while (cc->select_offset < cc->select_end) { + if (!collapse_window_eligible(cc, cc->select_offset, + cc->select_order)) { + collapse_selection_reject(cc); + continue; + } + + /* + * The cursor advances past the window at emission: a round is + * collected before it is run, so within a round every attempt is + * assumed to succeed. + */ + *offset =3D cc->select_offset; + *order =3D cc->select_order; + collapse_selection_advance(cc, 1U << cc->select_order); + return true; + } + return false; } =20 diff --git a/mm/collapse.h b/mm/collapse.h index 3256c45ee228..94b796271843 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -114,6 +114,15 @@ struct collapse_control { /* Orders still worth attempting in the table being scanned */ unsigned long select_orders; =20 + /* Non-present PTEs the scan accepted, which no bitmap bit marks */ + unsigned int scan_unmapped; + + /* Where selection has got to in the table, and at what order */ + unsigned int select_start; + unsigned int select_end; + unsigned int select_offset; + unsigned int select_order; + /* PTEs collapsed in it so far */ unsigned int nr_collapsed; =20 @@ -166,6 +175,7 @@ enum scan_result find_pmd_or_thp_or_none(struct mm_stru= ct *mm, unsigned long address, pmd_t **pmd); int collapse_find_target_node(struct collapse_control *cc); bool collapse_scan_abort(int nid, struct collapse_control *cc); +unsigned int max_order_from_offset(unsigned int offset); unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order); unsigned int collapse_max_ptes_swap(struct collapse_control *cc, diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 1244e161beae..c7c933e819e2 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -1426,7 +1426,7 @@ static enum scan_result collapse_huge_page(struct mm_= struct *mm, unsigned long s } =20 /* Return the highest naturally aligned order that fits at @offset within = a PMD. */ -static unsigned int max_order_from_offset(unsigned int offset) +unsigned int max_order_from_offset(unsigned int offset) { if (offset =3D=3D 0) return HPAGE_PMD_ORDER; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6C7B63F44E4; Sun, 16 Aug 2026 22:47:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920425; cv=none; b=E0/97t3Sl/njVNkTBo0Ob8Qhr3Tuu3VWmxiTps/Dqe+H4Zkb0JAvqzxgDBBp3w3G5dNcvt68mboeV+a7zPNLDGyMtyh6pIpJz1PQF6VIkdkQQzo4CUUdVaFx1Vqm+j23LihjflBb+JD6gPc94jBgdKkHeDoUIO9xbm92tLm9iBA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920425; c=relaxed/simple; bh=wdR3ztM3/OJI5oytLfnBwUORH9CS8Cr/Xx6y/1c4hN8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=SpF+e9ZXY2xIBbOo7CFFCJlB9UTfnHnvttXZaLxbgWIAtjZc+mbjrxzQqwMMPuvFntFa6TFD1AWjswMvpS9KBLQAHIyj43klKy9wQLR58wx1BEV59/Tw4w4dl/FPQBIBhrUgoXxqb7pdoV/JQLThiaQkk3Pg2ZsPQJ7wfdKMI/Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ny2uZjQ3; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=QHQBe79j; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ny2uZjQ3"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="QHQBe79j" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.phl.internal (Postfix) with ESMTP id A1ACAEC0246; Sun, 16 Aug 2026 18:47:02 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Sun, 16 Aug 2026 18:47:02 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920422; x= 1787006822; bh=qzlfiKIcvEbT8Dr/8J+UmVaUqw8Y/2tyxTnz6oo7+IU=; b=n y2uZjQ39Z45oxzCucWHqHZWv5YZDRetLzNNV5GCNHCDEGC9wlzSI5jzbCF7uR7fk 1GvWKXurQgug48KbUeFublWC6BH7EEj6V3z5VvNideshKDSaJQVEmhzWHxg5SrVD nKNs3PrXYkGNjn9tKOAX/VGCSQAeRzjhDuamghEBlzdKF31lvQY+NTDygccAxTNK 1iEeAwXTrcq4l9CTMRTp4drqBkBizQb2NDyXmjFwwQ9Mwgr+2X8Ae2wzwJzLs2ZY fGbQb3Gm7ElB8nBZ0j3PpBlWEsMY3IpZrusavoC32Ja/Z8V+/IAWwQEJsWzDBSKw l2jGNopytQGRqwoyzLYDA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920422; x=1787006822; bh=q zlfiKIcvEbT8Dr/8J+UmVaUqw8Y/2tyxTnz6oo7+IU=; b=QHQBe79jZ5W+AjIsH 8jI3at8052DccDT1GJdP0OC5UE9md2mgTewL+KWmZuZCnOLTPFGApWNq8Yq48taZ 4I3s2KdKjUdsBEPJmNr+3e11seON2EDhVFQPTe1TgzYB/OPKRUESJVWP8qoXp/9r H6PpynctpSGgy58eeRFXP9hD2gypGH/dX8sLRQxiNWbkzx4xnvEuAnP4Ldq9xb8K KyQpazIflBi6BCR3vGBi/np3WAp0OkUdZNZ4SVDnLvrD2Bmbj1xo4yCx81TV0HDz iGgVpBWhpjrzFpFvKUIdDeaWRmbiHsPbSsgGsAPddNGvlPlqxdDXEqAQct9SeAOv GtFdw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzToM MZcmzdxzEDBChv7E0c2Xjsinj9zJe8Q5aa6n4OlJqqveLcQGvaXqWzd7QR9JsdEtU27rOw JUcvqeU5GST0gG0WMSxAMvGj4+YO6ChL3A4m737E7UOdycrEYBdvuyCiO1B1fN4ildGW75 9DhSN/aJ4aniV6bisp3pLIVgHYSzBzKjs0JMAg4qCqUXig9qixxwzzHICDHHPzIQYQi7WP ImdDWbUrXQYfXkXPLMpK+OxzBZQu5Wm+uObzndajfcgGbkFpakF3J0YAW0YeNiCPQKZxRr XNeomXEHQ/V11EoDQUpSqlBfnVMS7YlgywXks3dovS90F97XkDJJiYmecVSg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:01 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 23/57] mm/collapse: give a refused region a second chance Date: Sun, 16 Aug 2026 23:45:35 +0100 Message-ID: <20260816224609.308019-24-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Fill in the rest of selection: a store of regions to re-enter, and the classification that decides what goes in it. Each attempted candidate's outcome is one of four: - it collapsed, or was already a huge page. The region is done; the cursor stepped past it at emission. - it was refused for something a smaller window might avoid. The whole region goes back at the next enabled order down: selection cannot tell which slot refused, so it re-probes the region rather than guessing. - only its in-window allocation missed. The region goes back at the same order, marked so the round that picks it up allocates with reclaim before freezing anything. - the outcome condemns the table. Selection stops there and drops everything queued. A queued region is walked exactly as the table is: the largest order its offset's alignment allows, capped by the region's own, descending through the enabled orders until one qualifies, then stepping past what it emitted. Walking rather than shrinking one window is what keeps the tail of a region in play, which is often where the collapsible part is. The store is a stack, and the classify loop feeding it walks the batch in emission order, so entries pop in the order they were refused rather than by address. A round drawn from two of them is not address-ordered. Selection terminates because the pushes that tile a region strictly descend, and the one that keeps the order cannot repeat for a region: it comes back asking for reclaim, and a miss the allocator was asked to work for is a failure, which descends. An allocation failure is only reported as one when nothing smaller is left to try. The caller answers such a failure by backing off for a while, and a failure at a large order is no reason to: one PMD is 512M with 64K pages, so that attempt fails as a matter of course, while the order the region settles for allocates fine. collapse_anon_pmd() can now say what the table yielded, in order of precedence: - a collapse; - an allocation failure, which the caller answers by backing off. It outranks both refusals below, being the only result acted on; - the scan's own refusal, when selection never got as far as refusing a window; - the last reason a window was refused. All are per table, so the reset that opens a table clears them, the retry store included. A dropped lock is classified with the outcomes that abandon the table, not with those that try a smaller order. The fault-in pass has already taken the lock again as many times as it may, so what is left says nothing about any window, and demoting every candidate the round was carrying would be a verdict no pass reached. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 201 +++++++++++++++++++++++++++++++++++++++++++++++--- mm/collapse.h | 11 +++ 2 files changed, 203 insertions(+), 9 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 258bb9cc32c5..9b73ebff1103 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -124,6 +124,16 @@ */ #define COLLAPSE_SAVED_PTES HPAGE_PMD_NR =20 +/* + * Capacity of the retry store: the most regions a table can hold at once.= Live + * entries cover disjoint regions -- a region is one candidate's extent, a= nd an + * extent is consumed from the cursor or from one entry, never from two --= and + * the smallest a producer pushes is one window at the smallest order. Not + * bounded by what a round pushes: the stack is drained from the top, so a= n entry + * below a live one outlives the round that pushed it. + */ +#define COLLAPSE_RETRY_STORE_SIZE COLLAPSE_TABLE_WINDOWS + /* How far a candidate got, and so what a failure has to undo for it */ enum collapse_candidate_state { CAND_SELECTED, /* collected; nothing held on its behalf yet */ @@ -132,6 +142,21 @@ enum collapse_candidate_state { CAND_INSTALLED, /* the destination is mapped */ }; =20 +/* + * A region queued to re-enter selection, walked like the table itself: @o= ffset is + * the next window to probe, @end one past the region, and @order the larg= est to + * try -- below the order that just failed, so the same window cannot be e= mitted + * twice. Walking the region rather than shrinking one window keeps its t= ail, + * which is often where the collapsible window is. + */ +struct collapse_retry { + unsigned int offset; + unsigned int end; + unsigned int order; + /* The light allocation missed here: the next attempt may reclaim */ + bool reclaim; +}; + /* * A candidate is an (addr, order) window selected for collapse. Selection * counts in PTE offsets -- the bitmap it reads and the alignment it honou= rs are @@ -181,16 +206,20 @@ void collapse_control_release(struct collapse_control= *cc) { kfree(cc->candidates); kfree(cc->saved_ptes); + kfree(cc->retries); cc->candidates =3D NULL; cc->saved_ptes =3D NULL; + cc->retries =3D NULL; } =20 int collapse_control_init(struct collapse_control *cc) { cc->nr_candidates =3D 0; + cc->nr_retries =3D 0; cc->candidates =3D kmalloc_objs(*cc->candidates, COLLAPSE_MAX_CANDIDATES); cc->saved_ptes =3D kmalloc_objs(*cc->saved_ptes, COLLAPSE_SAVED_PTES); - if (!cc->candidates || !cc->saved_ptes) { + cc->retries =3D kmalloc_objs(*cc->retries, COLLAPSE_RETRY_STORE_SIZE); + if (!cc->candidates || !cc->saved_ptes || !cc->retries) { collapse_control_release(cc); return -ENOMEM; } @@ -1855,6 +1884,9 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) cc->select_orders =3D 0; cc->scan_unmapped =3D 0; cc->nr_collapsed =3D 0; + cc->select_result =3D SCAN_FAIL; + cc->smallest_alloc_failed =3D false; + cc->nr_retries =3D 0; } =20 /* @@ -2010,6 +2042,48 @@ static bool collapse_window_eligible(struct collapse= _control *cc, return nr_eligible_ptes >=3D nr_ptes - max_ptes_none; } =20 +/* + * Queue the region [@offset, @end) to re-enter selection at @order. Two + * producers push, both in collapse_classify_result(): a refused region, t= iled at + * the next enabled order down because selection cannot tell which slot re= fused; + * and a region whose in-window allocation missed, at an unchanged order, = asking + * for reclaim next time. + * + * The store is a stack, and the classify loop that feeds it walks the bat= ch by + * ascending address, so entries pop in the order they were refused rather= than + * by address: a round drawn from two of them descends. Nothing may take + * candidates[0] for the lowest -- what a round spans is cc->batch_start a= nd + * cc->batch_end, taken over its candidates by collapse_revalidate(). + * + * Selection terminates because the tiling producer strictly descends, and= the + * unchanged-order one cannot fire twice for a region: its retry arrives w= ith + * reclaim set, so the next miss is a failure that descends. + * + * The store is sized for the most regions a table can hold, so this cannot + * overflow; losing an entry would cost a region its lower-order attempt, = so it + * asserts rather than fails. + */ +static void collapse_push_retry(struct collapse_control *cc, unsigned int = offset, + unsigned int end, unsigned int order, + bool reclaim) +{ + struct collapse_retry *retry; + + if (cc->nr_retries >=3D COLLAPSE_RETRY_STORE_SIZE) { + VM_WARN_ON_ONCE(1); + return; + } + + retry =3D &cc->retries[cc->nr_retries]; + + retry->offset =3D offset; + retry->end =3D end; + retry->order =3D order; + retry->reclaim =3D reclaim; + + cc->nr_retries++; +} + /* * The next window worth attempting, as an (offset, order) pair. False wh= en * selection is exhausted, which is what ends the range. @@ -2019,8 +2093,42 @@ static bool collapse_window_eligible(struct collapse= _control *cc, * particular -- would be stale by construction. */ static bool collapse_next_candidate(struct collapse_control *cc, - unsigned int *offset, unsigned int *order) + unsigned int *offset, unsigned int *order, + bool *reclaim) { + while (cc->nr_retries) { + struct collapse_retry *r =3D &cc->retries[cc->nr_retries - 1]; + unsigned int try, smallest; + + if (r->offset >=3D r->end) { + cc->nr_retries--; + continue; + } + + /* + * The same walk as the table's own: the largest order the + * offset's alignment allows, capped by the region's, descending + * through the enabled orders until one fits. If nothing fits + * here, step over the smallest window tried and carry on -- + * which is what keeps the region's tail in play. + */ + try =3D min(max_order_from_offset(r->offset), r->order); + smallest =3D try; + while (try && !collapse_window_eligible(cc, r->offset, try)) { + smallest =3D try; + try =3D collapse_lower_order(cc, try); + } + + if (try) { + *offset =3D r->offset; + *order =3D try; + *reclaim =3D r->reclaim; + r->offset +=3D 1U << try; + return true; + } + r->offset +=3D 1U << smallest; + } + while (cc->select_offset < cc->select_end) { if (!collapse_window_eligible(cc, cc->select_offset, cc->select_order)) { @@ -2035,6 +2143,7 @@ static bool collapse_next_candidate(struct collapse_c= ontrol *cc, */ *offset =3D cc->select_offset; *order =3D cc->select_order; + *reclaim =3D false; collapse_selection_advance(cc, 1U << cc->select_order); return true; } @@ -2051,7 +2160,68 @@ static bool collapse_classify_result(struct collapse= _control *cc, unsigned int offset, unsigned int order, enum scan_result result) { - return true; + unsigned int lower; + + switch (result) { + /* Done with the region: the cursor moved past it at emission */ + case SCAN_SUCCEED: + cc->nr_collapsed +=3D 1U << order; + fallthrough; + case SCAN_PTE_MAPPED_HUGEPAGE: + return true; + /* Only the light allocation missed: the same order, allowed to reclaim */ + case SCAN_ALLOC_LIGHT_MISS: + collapse_push_retry(cc, offset, offset + (1U << order), order, + /*reclaim=3D*/ true); + return true; + /* A smaller order over the same region might still fit */ + case SCAN_ALLOC_HUGE_PAGE_FAIL: + /* + * Only a failure with nothing left below it says the allocator + * cannot serve this collapse. A failure at a large order says + * nothing about what the region will settle for -- one PMD is + * 512M with 64K pages, so that attempt fails as a matter of + * course -- and the caller answers an allocation failure by + * backing off for a while. + */ + if (!collapse_lower_order(cc, order)) + cc->smallest_alloc_failed =3D true; + fallthrough; + case SCAN_LACK_REFERENCED_PAGE: + case SCAN_EXCEED_NONE_PTE: + case SCAN_EXCEED_SWAP_PTE: + case SCAN_EXCEED_SHARED_PTE: + case SCAN_PAGE_LOCK: + case SCAN_PAGE_COUNT: + case SCAN_PAGE_NOT_EXCLUSIVE: + case SCAN_PAGE_NULL: + case SCAN_DEL_PAGE_LRU: + case SCAN_PTE_NON_PRESENT: + case SCAN_PTE_UFFD: + case SCAN_PAGE_LAZYFREE: + case SCAN_PAGE_DIRTY_OR_WRITEBACK: + cc->select_result =3D result; + lower =3D collapse_lower_order(cc, order); + if (lower) { + /* The whole failed region re-enters, as one entry */ + collapse_push_retry(cc, offset, offset + (1U << order), + lower, /*reclaim=3D*/ false); + } + return true; + /* + * Nothing further is worth attempting in this table. A dropped lock + * belongs here rather than above: it says nothing about any window, so + * lowering the order of every candidate the round was carrying would be + * a verdict nobody reached. The next scan finds the table again. + */ + case SCAN_LOCK_DROPPED: + case SCAN_PMD_MAPPED: + default: + cc->select_result =3D result; + cc->select_offset =3D cc->select_end; + cc->nr_retries =3D 0; + return false; + } } =20 /* @@ -2112,7 +2282,7 @@ static bool collapse_batch_full(struct collapse_contr= ol *cc, unsigned int slots, */ static void collapse_add_candidate(struct collapse_control *cc, unsigned long addr, unsigned int order, - pte_t *saved_ptes) + bool reclaim, pte_t *saved_ptes) { struct collapse_candidate *cand; =20 @@ -2124,7 +2294,7 @@ static void collapse_add_candidate(struct collapse_co= ntrol *cc, cc->nr_candidates++; cand->addr =3D addr; cand->order =3D order; - cand->reclaim =3D false; + cand->reclaim =3D reclaim; cand->state =3D CAND_SELECTED; cand->result =3D SCAN_FAIL; cand->new_folio =3D NULL; @@ -2145,7 +2315,7 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned long= start, unsigned long end, unsigned int offset, order; unsigned long bytes =3D 0; unsigned int slots =3D 0; - bool pending =3D false; + bool pending =3D false, reclaim =3D false; bool cont =3D true; =20 collapse_selection_init(cc, (start - pmd_addr) >> PAGE_SHIFT, @@ -2153,7 +2323,8 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned long= start, unsigned long end, =20 while (cont) { if (!pending) - pending =3D collapse_next_candidate(cc, &offset, &order); + pending =3D collapse_next_candidate(cc, &offset, &order, + &reclaim); =20 if (!pending || collapse_batch_full(cc, slots, bytes, order)) { /* @@ -2177,12 +2348,24 @@ collapse_anon_pmd(struct mm_struct *mm, unsigned lo= ng start, unsigned long end, * full round could not take is kept pending for the next one. */ collapse_add_candidate(cc, pmd_addr + offset * PAGE_SIZE, order, - cc->saved_ptes + slots); + reclaim, cc->saved_ptes + slots); =20 slots +=3D 1U << order; bytes +=3D PAGE_SIZE << order; pending =3D false; } =20 - return cc->nr_collapsed ? SCAN_SUCCEED : SCAN_FAIL; + if (cc->nr_collapsed) + return SCAN_SUCCEED; + /* + * Report an allocation failure over any refusal, the scan's included: it + * is the one outcome the caller acts on, by backing off rather than + * scanning on. + */ + if (cc->smallest_alloc_failed) + return SCAN_ALLOC_HUGE_PAGE_FAIL; + /* Nothing salvaged and nothing to wait for: say what was refused */ + if (cc->scan_refusal !=3D SCAN_SUCCEED) + return cc->scan_refusal; + return cc->select_result; } diff --git a/mm/collapse.h b/mm/collapse.h index 94b796271843..3803f5a89087 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -11,6 +11,7 @@ #define COLLAPSE_MIN_MTHP_ORDER 2 =20 struct collapse_candidate; +struct collapse_retry; =20 enum scan_result { SCAN_FAIL, @@ -135,6 +136,16 @@ struct collapse_control { */ enum scan_result scan_refusal; =20 + /* Why the last window was refused */ + enum scan_result select_result; + + /* A region ran out of orders to try because none could be allocated */ + bool smallest_alloc_failed; + + /* Regions waiting to re-enter selection at a lower order */ + struct collapse_retry *retries; + unsigned int nr_retries; + /* The candidate windows collected for the current round */ struct collapse_candidate *candidates; unsigned int nr_candidates; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 207C63F44E9; Sun, 16 Aug 2026 22:47:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920426; cv=none; b=IXP9J1B1bcEEN56cNRZX9l48BOhVjgAykXLJC46y/yaEFZ1z4gLLUy/aoX77p9RzEssEyREAiGLr4cQyWsC9zbMBH+CiQvHEPNBG7gr2+tOd5DhlPvpqNeyhauGz/wNYYuHne+rC7kbS/OfNAEXEo+1b4x3bUe9VnhHwNlcmq9Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920426; c=relaxed/simple; bh=esfOLY/HwuAXougYMGMGajTc2sL/3wg0eE4TmzYzGaE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ZYu67XzFDdIO5Mj/sW100awIr4fyS1ZTfRlZ0uzNNWQYqOxr0PMWTw7Er85uZtB/FML1a8pEv9qQgxpqj0Egn87YN61cHAXw7VTsMJCX1xcPE5VynaIHpUiRpkQzsh/baEie8YHBBtXMMYB9Qz3+5tyb0O+zluzSvkmj6qmY0og= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=SmT7gsgM; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=HD6CvoxK; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="SmT7gsgM"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="HD6CvoxK" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 7FA5EEC0074; Sun, 16 Aug 2026 18:47:04 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:04 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920424; x= 1787006824; bh=206JYDvQB0ujHXi5pJEt/V4UdyhHvopT44QwBt0napI=; b=S mT7gsgM5hzoFn1e/Zc+g6MXtweaEWB8DtqJUpnR6CPUbPr5gd0rmpDOjnI7B+FJZ txMLHobGB7wjUnh15YuS1D6f9A1qo1XtIhszT6QcE85qy8Ng6FcOXmB+ZP9D8a9U eoLwiTHGvY2J9TaxwDI4LMtRrJOeGgZFniGnCc+dVkjUHq3woiaAyvmBovWPwtXf 7FrR0IgNIW/wbvN7neSBi1ErynrI6HYN5oxFeC19LqBFV5hGrZlaB8W2MdMkbyUn prTIHTooXk1tuKhBRj4FGMd0Ld/CURRaWtTpmQ3M1Sea56np3V90gvgeVj6b5O6A wk9d9RZCe8+y2PrEjeL8Q== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920424; x=1787006824; bh=2 06JYDvQB0ujHXi5pJEt/V4UdyhHvopT44QwBt0napI=; b=HD6CvoxK3cTmYVE2m ReLahvWPkzgEQ4xT4HkS/JQyy1p6N1JC5QRs9cFhD+r6mqagE+WgEjY852OpRFV+ QUbyUv2hgQfdnaxbPeyhLjhSJ7NdwBSUf8CezLenlk2bsJjHJWDk8Yv2zQ+2BnHW NJqMInRKsOVl+590OdISwJAY/+wpCY/KXsjKnc1i0oLKwB5ZuyT78BzdhFleZHZ3 HPgTJd6Tl+RGi0FLD9+ZAqGcvJ8YzLrf1NdGGgAhNz/03iWNX+NEVWRM2Lu0cvfw 69jmBoQG3EJyQP6apVDX+G/Rc4Nm8EZ1FAt3rAaji1H0KQUrfhZqxzrureH9kn8r QQuDA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGeQUP6jehEeI0kOFfHX8Bz1vci7ZS0pF5h1BDhjvrnMAYVx6hGR3OE9yv6W3r0gb n0/FUryAyJMRviPVI/zcKHlW5AI1WzRuWaH5KZl7QKEdlwAk6TndP7sFtmhpcUxoS2T1Ub bJYQdn3ioRn3ttrBnyjoeoqk/8OxdnzTOD/CD/598C7l9KFQmDhpWnCO05OTUOy71ONkYN 1mUBKy2LgCWDUmlELwjFPAnmnv6K5fs53d2O4zI6fbjDWLajXYbzG+y30fmb5KnxmJ+KbM yupMSRc5Aw9K+Bbz/tXrPcir7VA0gcPKiOoeZixTP7fXgAmXhGWtgZIkK9P0VDdjDmjmE/ vWPsFH7ul4cYFqd0Pb083yR7+P0Ep3+LJFuPrtz77cQGVMhStHZpW289eHgvarwxIEFYEM +ALAAqWpg18S5YrkTaLDxuOv8uigPsSWuJ/XKe+OsF7XwjD2hjS1D/a46c+kLVxStz75+O DACM9YC77Y//czQG/9f12QOqbdQHTdxxlslZP+eor7NOYfVbjXDbL6UAhu4Yhl/d8xVJwH n0qOTZ1svyTljhI8LFos0S5R4ZiR/3pFdPEW1V8s/kvfdC+Ux5Mk3EJmhpkcGwhdlTIVdz y3Ba/zg8hH0wHRl0+R61js1Su6FP9MrYAYldXcP+OaLB23y634BaHsTI2LCA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:03 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 24/57] mm/collapse: report each candidate's outcome to tracing Date: Sun, 16 Aug 2026 23:45:36 +0100 Message-ID: <20260816224609.308019-25-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The engine decides per candidate, and every one of those decisions is currently invisible: the mechanism it is about to replace reports through mm_collapse_huge_page_isolate, which the engine never calls. Switching the anonymous path over without something in its place would take existing tracing with it. Add one tracepoint, mm_collapse_candidate: a window's address and order, the pass that reached a verdict on it, and what that verdict was. Every candidate a pass judged produces exactly one -- the pass that refused it, or the install for one that made it. A candidate the round gave up on before any pass judged it produces none. That is enough to follow a round: which windows were attempted, and which ones the batch dropped and where. It is also what a scan of the trace buffer can attribute to an address. It goes in the huge_memory trace system, next to the events it stands in for, so a consumer enabling that system keeps seeing collapses. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 40 ++++++++++++++++++++++++++++++ mm/collapse.c | 32 +++++++++++++++++++++++- mm/collapse.h | 13 ++++++++++ 3 files changed, 84 insertions(+), 1 deletion(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index ff938ac9c43c..86131845b761 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -44,12 +44,21 @@ EM( SCAN_PAGE_NOT_EXCLUSIVE, "page_not_exclusive") \ EMe(SCAN_ALLOC_LIGHT_MISS, "alloc_light_miss") =20 +#define COLLAPSE_PASS_STATUS \ + EM( COLLAPSE_PASS_ALLOC, "alloc") \ + EM( COLLAPSE_PASS_REVALIDATE, "revalidate") \ + EM( COLLAPSE_PASS_FAULTIN, "faultin") \ + EM( COLLAPSE_PASS_FREEZE, "freeze") \ + EM( COLLAPSE_PASS_COPY, "copy") \ + EMe(COLLAPSE_PASS_INSTALL, "install") + #undef EM #undef EMe #define EM(a, b) TRACE_DEFINE_ENUM(a); #define EMe(a, b) TRACE_DEFINE_ENUM(a); =20 SCAN_STATUS +COLLAPSE_PASS_STATUS =20 #undef EM #undef EMe @@ -117,6 +126,37 @@ TRACE_EVENT(mm_collapse_huge_page, __entry->order) ); =20 +TRACE_EVENT(mm_collapse_candidate, + + TP_PROTO(struct mm_struct *mm, unsigned long addr, unsigned int order, + int pass, int result), + + TP_ARGS(mm, addr, order, pass, result), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(unsigned long, addr) + __field(unsigned int, order) + __field(int, pass) + __field(int, result) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->addr =3D addr; + __entry->order =3D order; + __entry->pass =3D pass; + __entry->result =3D result; + ), + + TP_printk("mm=3D%p, addr=3D0x%lx, order=3D%u, pass=3D%s, result=3D%s", + __entry->mm, + __entry->addr, + __entry->order, + __print_symbolic(__entry->pass, COLLAPSE_PASS_STATUS), + __print_symbolic(__entry->result, SCAN_STATUS)) +); + TRACE_EVENT(mm_collapse_huge_page_isolate, =20 TP_PROTO(struct folio *folio, int none_or_zero, diff --git a/mm/collapse.c b/mm/collapse.c index 9b73ebff1103..91ff20138a8e 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -20,6 +20,7 @@ #include =20 #include +#include #include "collapse.h" #include "internal.h" =20 @@ -195,6 +196,14 @@ static unsigned int candidate_nr_pages(const struct co= llapse_candidate *cand) return 1U << cand->order; } =20 +static void collapse_trace_candidate(struct mm_struct *mm, + const struct collapse_candidate *cand, + enum collapse_pass pass) +{ + trace_mm_collapse_candidate(mm, cand->addr, cand->order, pass, + cand->result); +} + /* Where a candidate sits in the table, in the PTE offsets selection count= s in */ static unsigned int candidate_offset(const struct collapse_candidate *cand, unsigned long pmd_addr) @@ -278,6 +287,8 @@ static enum scan_result collapse_revalidate(struct vm_a= rea_struct *vma, BIT(cand->order))) { cand->state =3D CAND_SKIPPED; cand->result =3D SCAN_VMA_CHECK; + collapse_trace_candidate(mm, cand, + COLLAPSE_PASS_REVALIDATE); continue; } =20 @@ -420,6 +431,8 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, if (r =3D=3D SCAN_EXCEED_SWAP_PTE) { cand->state =3D CAND_SKIPPED; cand->result =3D r; + collapse_trace_candidate(vma->vm_mm, cand, + COLLAPSE_PASS_FAULTIN); break; } if (r !=3D SCAN_SUCCEED) { @@ -878,6 +891,7 @@ static void collapse_freeze(struct vm_area_struct *vma, continue; cand->state =3D CAND_SKIPPED; cand->result =3D SCAN_NO_PTE_TABLE; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_FREEZE); } return; } @@ -904,6 +918,7 @@ static void collapse_freeze(struct vm_area_struct *vma, cand->result =3D result; if (result !=3D SCAN_SUCCEED) { cand->state =3D CAND_SKIPPED; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_FREEZE); continue; } =20 @@ -986,6 +1001,7 @@ static void collapse_reserve(struct mm_struct *mm, str= uct collapse_control *cc) =20 cand->state =3D CAND_SKIPPED; cand->result =3D result; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_ALLOC); } } =20 @@ -1016,6 +1032,7 @@ static void collapse_deposit(struct mm_struct *mm, st= ruct collapse_control *cc) if (!cand->deposit) { cand->state =3D CAND_SKIPPED; cand->result =3D SCAN_ALLOC_HUGE_PAGE_FAIL; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_ALLOC); } } =20 @@ -1060,6 +1077,8 @@ static void collapse_provision(struct mm_struct *mm, } cand->result =3D result; } + + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_ALLOC); } } =20 @@ -1106,6 +1125,8 @@ static void collapse_copy(struct vm_area_struct *vma, */ if (copy_mc_user_highpage(dst, src, addr, vma)) { cand->result =3D SCAN_COPY_MC; + collapse_trace_candidate(vma->vm_mm, cand, + COLLAPSE_PASS_COPY); break; } } @@ -1294,6 +1315,7 @@ static void collapse_install_pmd(struct vm_area_struc= t *vma, /* Table gone under us; see collapse_abort_candidate() on @pte */ spin_unlock(pmd_ptl); cand->result =3D SCAN_NO_PTE_TABLE; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_INSTALL); collapse_abort_candidate(vma, cand, NULL); return; } @@ -1315,6 +1337,7 @@ static void collapse_install_pmd(struct vm_area_struc= t *vma, =20 if (!collapse_verify_candidate(cand, pte, &nr_populated)) { cand->result =3D SCAN_PTE_NON_PRESENT; + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_INSTALL); collapse_abort_candidate(vma, cand, pte); goto out_unlock; } @@ -1411,6 +1434,8 @@ static void collapse_install(struct vm_area_struct *v= ma, continue; =20 cand->result =3D SCAN_NO_PTE_TABLE; + collapse_trace_candidate(mm, cand, + COLLAPSE_PASS_INSTALL); collapse_abort_candidate(vma, cand, NULL); } return; @@ -1439,6 +1464,8 @@ static void collapse_install(struct vm_area_struct *v= ma, =20 if (!collapse_verify_candidate(cand, cand_pte, &nr_populated)) { cand->result =3D SCAN_PTE_NON_PRESENT; + collapse_trace_candidate(mm, cand, + COLLAPSE_PASS_INSTALL); collapse_abort_candidate(vma, cand, cand_pte); continue; } @@ -1548,8 +1575,11 @@ static unsigned int collapse_finish(struct mm_struct= *mm, pte_free(mm, cand->deposit); cand->deposit =3D NULL; } - if (cand->state =3D=3D CAND_INSTALLED) + if (cand->state =3D=3D CAND_INSTALLED) { nr_installed++; + collapse_trace_candidate(mm, cand, + COLLAPSE_PASS_INSTALL); + } } =20 return nr_installed; diff --git a/mm/collapse.h b/mm/collapse.h index 3803f5a89087..34de3ebb05e3 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -13,6 +13,19 @@ struct collapse_candidate; struct collapse_retry; =20 +/* + * Which pass of a round reached a verdict on a candidate. Only collapse.c + * produces these; the trace header khugepaged.c builds names them. + */ +enum collapse_pass { + COLLAPSE_PASS_ALLOC, + COLLAPSE_PASS_REVALIDATE, + COLLAPSE_PASS_FAULTIN, + COLLAPSE_PASS_FREEZE, + COLLAPSE_PASS_COPY, + COLLAPSE_PASS_INSTALL, +}; + enum scan_result { SCAN_FAIL, SCAN_SUCCEED, --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D4B6F3F4DD7; Sun, 16 Aug 2026 22:47:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920428; cv=none; b=dsePw5VXDDET76mLJZBnDsvMf4Exl310Rp+8xVr+dpIPyan/BS4ZTOcNUNAANNuzhn1LMG6FxH4fHyc4N8SlNanu/ZQmiRwPSHxesoM1f4sQGgOyew28qPPkhY1lnkc55rwaFl1mXzeXA3tpeCIvLUahr7w3EdLBlk5waKtBv+Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920428; c=relaxed/simple; bh=rXDOP2lazn8hqApcn6UuIp/qgHOhDYu0/+0jCK61xSY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cM63rp+x5KAJ9bHPW2TWD5sKa5GcCDMoWzAMFBfwjIuyMF8+OXgeg23BLM1eNfKoe7j8tua28s1gqwYjR7o1oT0+0IGLK+OYN5TO0v6M07eIMCv8DVBqTOhQVR3l0ASNArweeHc2n7r9dnXNMPJbA1x7+1nVw1gjcOyMfb043kQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ei6Waa0k; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=KctISWAq; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ei6Waa0k"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="KctISWAq" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 33F2514000F8; Sun, 16 Aug 2026 18:47:06 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:06 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920426; x= 1787006826; bh=g/bEmHChGXXUr4LI/M464ypgkDiKtYjukQog+PJvxwc=; b=e i6Waa0kx1JeztwEvkH7Ff200lYEintKhm9n8taz6+bc7eEjLWupgfwUZ7CWItYkM 2Rn43iFJOvlPRnH51Hgg2xJ7+H7866S3YBV7YbWu1LzMtr1v8tS0Ozue/Tf1ZSYQ Fr9Kh7fq0yU1zBXvlJzbNWo6FVL0RQ/i6dy1wYd6jN6FAvBpvDOv19Zp7Ae4VNdK hqIdJA5GeQO18Qm4hGbkOS32fUXEqf512NYKfKFchNpjNRXFGPuuWYL9cK7GUupI E5iOXDGzSn3jQAxqA1rGfPfXYsbPcP5mf8Uu3VdY7Ng2qrsnUURbYEpMw1rXOrH+ AhXf+0FPEOh2Y0pdB3Rqw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920426; x=1787006826; bh=g /bEmHChGXXUr4LI/M464ypgkDiKtYjukQog+PJvxwc=; b=KctISWAqOVGNzO3uI qUUJkvTYfOuq8nQ9ZBuU3O7uZQDSx6Y23srlgaQ2pPIN98XBQHk0crb7X2+82lI9 47S3CEe/ZbBo1+Yd+Q3SszJ5WYGeCCw0Pb+jekaXRk505lB06Ol0XHi2ePnP3p8p JCZSQT2Ct7CIBoci6RVWvambGzUrqgr70a8Nu842IaglBiwaZvCCu2oE2Wb7FteH 44Wb4fDGgn4XCO0TNpb8iMuGW1f7r4/EMOuau7bZqwcquYtnctmwb03QbgRAzCKg tCxf7D/xKNQoW09nKQLobQf+vzzbuWcSTndkfMPifmJnoNcq5rFaii3A7e0mf5dJ jz3Yg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGlsoL4dlGZVhjNGnyng7DWLkzd23+vL0S14uJaq1tszuZ6c/qXpcKjIN3SVavrqW anksLy7isjEKJzVCE9/It/QcezLPAumFeVWWYf/196d5DKxXijJBwwHlwdbhLN8+Q20ZSw VAcGFAGUZDfCfgJltT5NnBIG4Gh7JxdfG7re68r55pgQSoCRLBnoP1i35vKgO+O/kd9+lz qLG43ny3AnvUx0aE1SnxP3FJCiuv2lDE4FRTwfST8/uDABclrMy6sZKUEjfd1WPzb5NlOP S5qghqIOcnVKJHuC+GM/K5ckB8SXpR6Ag+KqE5QSQT/v+DpSzaOOJbUo2XdsQxu9tqIPpv n09TP7n+3O6PIcr4WyyYUg9u90X9r/cVsCqHguDiVmkz5B1ez36ZmXEudDYtVH55+e6bCK KD2uHs8jsljSrOCCQPiL6eMr8S6teHqR+yZdv5bSo99ujysnSUwODNY3DMq2uOuYhc/B1K zQd8gpF7lAHXJCYb5Gf3L/8gFrxBSdbUcXNXwSJdHiMT6VIr0Hbyj5mC/SwERAeAvPFx50 mW3688759fv6PqAubbjBheKUmrOIHeCnR7QwRCZHDUwZPqhY9eJVsFYJ4M8BVHsY7amxmx /+VZW7YyA4oFHyCdj4h7m+U9nJCAvQgbWOMZ55UwAOCzJinUKyhZTVU0sDzA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:05 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 25/57] mm/collapse: collapse anonymous memory with the new engine Date: Sun, 16 Aug 2026 23:45:37 +0100 Message-ID: <20260816224609.308019-26-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Point the anonymous path at the engine. Everything it needs is in place, so this is the whole switch: collapse_single_pmd() calls collapse_scan_anon_pmd() and then collapse_anon_pmd(), where it used to call collapse_scan_pmd(). The scan runs under the mmap_read the caller already holds; the collapse is called after dropping it, and takes the lock itself for each round. Both callers hand the engine a PMD-aligned address with the whole table inside the VMA: khugepaged walks [ALIGN(vm_start), ALIGN_DOWN(vm_end)) a table at a time, and MADV_COLLAPSE aligns its range inwards the same way. So the range passed is always the table. The engine accepts a narrower one, which nothing asks for yet. Two things userspace sees change: - A collapse runs under mmap_read rather than holding mmap_write for its duration, so faults elsewhere in the address space are no longer stopped while it works. - A table that cannot become one huge page still yields the largest windows inside it, where before a single disqualified PTE gave up the whole table. The result the caller gets is the engine's, and it still acts on an allocation failure by backing off. The mechanism this replaces is left in place, now unreferenced. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 15 ++++++--------- mm/collapse.h | 5 +++++ mm/khugepaged.c | 14 ++++++++++++-- 3 files changed, 23 insertions(+), 11 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 91ff20138a8e..df3760e3918b 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -83,9 +83,6 @@ * consecutive pages of one folio -- so partially mapped and compound sour= ces * collapse too: any order below the window's is a source, and a PMD candi= date * takes even a PTE-mapped THP of its own order. - * - * Nothing calls any of this yet: the anon path still uses the mechanism it - * replaces, and is switched over once both halves are complete. */ =20 /* @@ -1926,9 +1923,9 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) * that acts on what it found hands the range to collapse_anon_pmd() after= wards, * without the lock. */ -static enum scan_result __maybe_unused -collapse_scan_anon_pmd(struct vm_area_struct *vma, unsigned long start, - unsigned long end, struct collapse_control *cc) +enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, + unsigned long start, unsigned long end, + struct collapse_control *cc) { const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; struct mm_struct *mm =3D vma->vm_mm; @@ -2337,9 +2334,9 @@ static void collapse_add_candidate(struct collapse_co= ntrol *cc, * largest order downwards. Returns what the table yielded: a collapse, or * the reason it did not. */ -static enum scan_result __maybe_unused -collapse_anon_pmd(struct mm_struct *mm, unsigned long start, unsigned long= end, - struct collapse_control *cc) +enum scan_result collapse_anon_pmd(struct mm_struct *mm, unsigned long sta= rt, + unsigned long end, + struct collapse_control *cc) { const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; unsigned int offset, order; diff --git a/mm/collapse.h b/mm/collapse.h index 34de3ebb05e3..50a9d59bbf03 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -186,6 +186,11 @@ static inline int collapse_test_exit_or_disable(struct= mm_struct *mm) mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); } =20 +enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, + unsigned long start, unsigned long end, + struct collapse_control *cc); +enum scan_result collapse_anon_pmd(struct mm_struct *mm, unsigned long sta= rt, + unsigned long end, struct collapse_control *cc); int collapse_control_init(struct collapse_control *cc); void collapse_control_release(struct collapse_control *cc); =20 diff --git a/mm/khugepaged.c b/mm/khugepaged.c index c7c933e819e2..0662d08f7c60 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -1553,7 +1553,8 @@ static enum scan_result mthp_collapse(struct mm_struc= t *mm, return last_result; } =20 -static enum scan_result collapse_scan_pmd(struct mm_struct *mm, +static enum scan_result __maybe_unused +collapse_scan_pmd(struct mm_struct *mm, struct vm_area_struct *vma, unsigned long start_addr, bool *lock_dropped, struct collapse_control *cc) { @@ -2749,7 +2750,16 @@ static enum scan_result collapse_single_pmd(unsigned= long addr, mmap_assert_locked(mm); =20 if (vma_is_anonymous(vma)) { - result =3D collapse_scan_pmd(mm, vma, addr, lock_dropped, cc); + result =3D collapse_scan_anon_pmd(vma, addr, addr + HPAGE_PMD_SIZE, + cc); + if (!cc->select_orders) + goto end; + + /* collapse_anon_pmd() takes mmap_lock itself, where it needs it */ + mmap_read_unlock(mm); + *lock_dropped =3D true; + + result =3D collapse_anon_pmd(mm, addr, addr + HPAGE_PMD_SIZE, cc); goto end; } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2D34B3F58CC; Sun, 16 Aug 2026 22:47:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920430; cv=none; b=UQCX9Qm7Vws4xmII+2SrxUUXk/4mlVxFVLWpKZ8ee6HSchPo67WBAJ4/r+2JuDRYZItNVcmwr++anPJET/M1M00PsDTgl1FRlCMj+phSf9BCVJ74tnxKeBYr7d5fUYLlmyLVEqNb9kll+EVHBLxQj3sI6ZTJlfqa0SfWhQRhQg8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920430; c=relaxed/simple; bh=OPyny8eAsrINGqeWv//2QHxGxjLp4XvtVK74h/syTDo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kp0uFVHEsFuHUQABWRw0TclLJcl8Kdu+hz7tXqwlMmtrrSkFcIpLn/xLklGqWnPOM0JoHHQBpR7JI0eAuPH/HtPrP96cqc+AyAsRPo8QdYxpPKYclV102p6zbo4GMD6BAQWQq+JlllwX0+RC97cauFgy1rllwMrzPATFC6qZawE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=PfcUTVDQ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=jWnfKufM; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="PfcUTVDQ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="jWnfKufM" Received: from phl-compute-11.internal (phl-compute-11.internal [10.202.2.51]) by mailfout.phl.internal (Postfix) with ESMTP id 4E99CEC0235; Sun, 16 Aug 2026 18:47:08 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-11.internal (MEProxy); Sun, 16 Aug 2026 18:47:08 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920428; x= 1787006828; bh=1c6dedl3WgtUFb9jPAe2ESc5CIX86porar5KZcsrzYc=; b=P fcUTVDQ49jpV/be6XprUQPnuAOC57eVsBvEsEnAEKcY0dsk0NQFcmbVG0sbmdIZl vs8ID/LUhIWEV5y8c6X/pgtS0sR7AGXUaW7g3jR82v35F9itQhU2cyzyFf63iY4z GpW10+8BfZ6Grc84rLt4NcRphubImU6GCxhiyWrahqU0kH/hWzYfvd5yPZgCaDYZ zV1DeXVITKswJ+W1gn5IllQkVfR9EMmDWDKtmZNN58dDhgmIR3B120MvB1qUu+pr d8/SK/JeQCewCxxSWx5sqi+yal7QaWawnuxZi4XIQw8q0a56HC8WNOvNxjtrL4Sh Mnz5o5fcRpSPwEROnfOdA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920428; x=1787006828; bh=1 c6dedl3WgtUFb9jPAe2ESc5CIX86porar5KZcsrzYc=; b=jWnfKufMiTHYpTqhy 8/7x3lxyBFDw/VOiuKQk1ibqBd5vJdZMzpNHMlWhyile4+KrhO5OPk+vSxr1DEol W/LQjamfvEHNoqnOZU/iUFRaAbNz6eH8z/Z1jh+9+sfJZHTTcHGwgrr8AFz4sppl WR6rYXRNFh1pCQbFUVrCOizz0Q4Ij3CFVTWC13w+IJSTqPRMRNWLInF3Bam8Y/Q+ cR+FAuNbMdkwERhB2FY62z2GKSew9MT9dmKrFD1BZwCLKONxWAsshRa8uyR9NNO9 ioKXO+lmu8ZP6GjZjsWkeDRu+x0jROQtHTb0BFWHub8HBgga11qL3AYmcD5ex3VX 6Nf0A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFteH 4aM8NKpVy6af3VJibLjmcthP+/7Eaj5TPpEQ+vEDAUYIkWdK+ywgMPzRIUnik9fpR7cN7f bjkwqvsizU8qA3PCDS7V7krH7+iKXYe5Hg1M4iBEcouwPVEwlyarcK46RRrbYWTslsvg3b 8BxETdBe/8+AXH8UuDZEwRoTeZ6Sk9Nqx/4bSAJ0WAYgoUD2lbzIpPDnAGEFI5TsE7Rnc/ otYH6+OJNrWmpfIFm9pxbzF5fJxzly2qMTts1jVR1CnQc2+liXZQT+8k/3HwwuwlfNmq53 zldbTxNq/mPU+OrhxsdMvQQUI8vVOO98rzAuQrbTusleoWDxxkj+W6cyhFJw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:07 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 26/57] mm/collapse: give collapse_single_pmd() the range to work on Date: Sun, 16 Aug 2026 23:45:38 +0100 Message-ID: <20260816224609.308019-27-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_single_pmd() derives the end of its range from its start: one PMD, always. Both of its callers already know the range they mean. Take the end as an argument and pass it to the scan and the collapse, both of which already accept a partial table. Both callers pass what the function computed for itself. Preparation for scanning a VMA that holds less than a whole table. No functional change intended. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/khugepaged.c | 14 ++++++++------ 1 file changed, 8 insertions(+), 6 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 0662d08f7c60..d1e031ed3e6f 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -2738,8 +2738,8 @@ static enum scan_result collapse_scan_file(struct mm_= struct *mm, * the results. */ static enum scan_result collapse_single_pmd(unsigned long addr, - struct vm_area_struct *vma, bool *lock_dropped, - struct collapse_control *cc) + unsigned long end, struct vm_area_struct *vma, + bool *lock_dropped, struct collapse_control *cc) { struct mm_struct *mm =3D vma->vm_mm; bool triggered_wb =3D false; @@ -2750,8 +2750,7 @@ static enum scan_result collapse_single_pmd(unsigned = long addr, mmap_assert_locked(mm); =20 if (vma_is_anonymous(vma)) { - result =3D collapse_scan_anon_pmd(vma, addr, addr + HPAGE_PMD_SIZE, - cc); + result =3D collapse_scan_anon_pmd(vma, addr, end, cc); if (!cc->select_orders) goto end; =20 @@ -2759,7 +2758,7 @@ static enum scan_result collapse_single_pmd(unsigned = long addr, mmap_read_unlock(mm); *lock_dropped =3D true; =20 - result =3D collapse_anon_pmd(mm, addr, addr + HPAGE_PMD_SIZE, cc); + result =3D collapse_anon_pmd(mm, addr, end, cc); goto end; } =20 @@ -2872,6 +2871,8 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, hend); =20 *result =3D collapse_single_pmd(khugepaged_scan.address, + khugepaged_scan.address + + HPAGE_PMD_SIZE, vma, &lock_dropped, cc); /* move to next address */ khugepaged_scan.address +=3D HPAGE_PMD_SIZE; @@ -3211,7 +3212,8 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, hend =3D min(hend, vma->vm_end & HPAGE_PMD_MASK); } =20 - result =3D collapse_single_pmd(addr, vma, &mmap_unlocked, cc); + result =3D collapse_single_pmd(addr, addr + HPAGE_PMD_SIZE, vma, + &mmap_unlocked, cc); =20 switch (result) { case SCAN_SUCCEED: --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C27613EAC68; Sun, 16 Aug 2026 22:47:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920432; cv=none; b=L8zghN8bTTpVZyutdTgGmDy8W801m+UPQmbfHPrEKeWZ3ToSSppS7aO8ySPH//TaYF2vOGoMKz8Z/UePrF+CzGjlEmpqA4RcjNbkm8hfoFQ0h/xvJd7J4brMpVYh8H5sx8Xrqg1BtU3rUy/B4/2uk9hofrs54ICirhCdxDxrIPk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920432; c=relaxed/simple; bh=qtthvM37H3GLokDgDQmyPduLycS78KAAlnhVAkzzjjs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=u17YoJTAV2g9woqAVoxZJYWtnO1uTTaDnLP50hwFvpCoZEugL6XaXCcVlG2tBiRMR1vFhk+YjlG52/RB6GTSF+nyClTzlUQGEZzy+VPkeIa0v5M3TmliW6yPWlXcTw62SAnDWhUMgoM7/ucv+NQpw1GIyxlD2NdphrHkTBUPWzg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=HW2cTKeV; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=KiLMp3B4; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="HW2cTKeV"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="KiLMp3B4" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id F3212EC0243; Sun, 16 Aug 2026 18:47:09 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:09 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920429; x= 1787006829; bh=I4QTeWOTpJRPuh3OfF+8CEm8xjikxAxj2x56bLCOOSc=; b=H W2cTKeViOUFKg4ugUwcqSC2lzWJ2gzQ7Bl336yHE4NmnTD/qYiXZJsS56PyF3jb2 vRin2AVLS4DXRf50KFmmsvo15XzGj6ZIFMCxW03WGSC/kro0ZFT6IFJAUlgtHdNy hM/fH6yJL6al46oPTa5NXWEtpQQTLqS3moEhY/crVlL3JyNJwfEHpbz1j3kZcrcx uc+i42paLfUjk4TVxpShCfb002SV7PIr7Vvvl/K8Df/Kx3R1Y8cK4Yr/qvdTv+UA VQt7mxI61MBPkXm1Y6/DEljs8O5n3aFOsNl2ZUr5vl8HIK07YW3x6Q/GFZBU2fur VXUI0Zr3/GosIsshFIHOw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920429; x=1787006829; bh=I 4QTeWOTpJRPuh3OfF+8CEm8xjikxAxj2x56bLCOOSc=; b=KiLMp3B4ag400yZym +u+s/kP8DjoiV1jGTUTQh2HrwXArbn4BZkhHRadCrK83d4yfJSLprdZxmKYcteHO 2fer5H+sy+0buosPR8Jmq/oOFdb9PocomY1Io8K3Mo6iOEq6505wWMSEt/j0uJsS 5hMv4hnkWSAXWQURzSUF6TqqU4MjMv2KmIg/QpYD/78uBX3THZSuZWfcbQGr70Ab EcPfrHqlhhLAdYfwzS3Apkpxct6mHHw7ChfflvQq1hghT4TjW3dKGq9D/dK1EvPJ zy8Uv+0mt+J8Ooliq/iNd8BAPNecVWh3XL3zw0M8TFt/CZWDZmMnVXBdtRwKT5CS lSIcg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGlsoL4dlGZVhjNGnyng7DWLkzd23+vL0S14uJaq1tszuZ6c/qXpcKjIN3SVavrqW anksLy7isjEKJzVCE9/It/QcezLPAumFeVWWYf/196d5DKxXijJBwwHlwdbhLN8+Q20ZSw VAcGFAGUZDfCfgJltT5NnBIG4Gh7JxdfG7re68r55pgQSoCRLBnoP1i35vKgO+O/kd9+lz qLG43ny3AnvUx0aE1SnxP3FJCiuv2lDE4FRTwfST8/uDABclrMy6sZKUEjfd1WPzb5NlOP S5qghqIOcnVKJHuC+GM/K5ckB8SXpR6Ag+KqE5QSQT/v+DpSzaOOJbUo2XdsQxu9tqIPVC wtH8RrR+VRBK4rHQP5fUOUd9SDmtkxSJIuqK6X0YhrZ10zdW4WgrIMWQcAlMOIw07y45IF iUQPQjmY18mZD+KeADQUCZgbhRiYVOnmw5psrOcS2zb3A7Ve3TEpym00DdUWWBFRygip4Q Fn+i65Iy6aPDTmuKBbeNhatneA0MOXpV6XTo/iFtVU6PszWR/4qIwju3GvgPVWhmGwEaF8 lOhqjwLWDDCkWCXUSIksZsiHNeLhW050yblRbXZhW9L1t9tojDhndbGZqh5BRIbDI62sIJ 2pCa7pAkdvTJlFFSD2rj0hqoXUfSk4uHGCJ+7xklr79KCQBj6ko9GT3+9ngA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:09 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 27/57] mm/collapse: scan the windows a VMA can hold Date: Sun, 16 Aug 2026 23:45:39 +0100 Message-ID: <20260816224609.308019-28-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" khugepaged covers each VMA in whole PTE tables, from its first PMD-aligned address to its last. A VMA smaller than a table is never scanned at all, and in a larger one everything outside its PMD-aligned span is skipped. That was the right shape when a collapse was always a PMD. It keeps mTHP collapse away from every range that is not PMD-shaped, which is most of what an mTHP is for. The gap is widest where a PMD is largest. On arm64 with 64K base pages a PMD is 512M, so the old walk reached only VMAs big enough and aligned well enough to hold one. A 2M mTHP -- order 5 there -- was unreachable in anything smaller, however many such windows the VMA had room for. Root the coverage at windows of the largest order the VMA allows, and hand the range on one table at a time, clamped to the VMA. The engine already accepts a partial table; this is the first caller that gives it one. The cursor is no longer PMD-aligned, so the assert that said it was goes. The bound beside it goes too: the range is clamped to the VMA where it is computed, leaving nothing for it to catch. For a VMA whose largest allowed order is the PMD order -- every file VMA, and any anonymous VMA with only PMD-order THP enabled -- the walk is exactly what it was. Where smaller orders are enabled, khugepaged now reaches VMAs a table would not fit in, and the edges of VMAs it used to leave out. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/khugepaged.c | 36 ++++++++++++++++++++++++------------ 1 file changed, 24 insertions(+), 12 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index d1e031ed3e6f..895183d92fb8 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -2838,44 +2838,56 @@ static void collapse_scan_mm_slot(unsigned int prog= ress_max, =20 vma_iter_init(&vmi, mm, khugepaged_scan.address); for_each_vma(vmi, vma) { - unsigned long hstart, hend; + unsigned long hstart, hend, window; + unsigned long orders; =20 cond_resched(); if (unlikely(collapse_test_exit_or_disable(mm))) { cc->progress++; break; } - if (!collapse_possible(vma, vma->vm_flags, TVA_KHUGEPAGED)) { + orders =3D collapse_possible_orders(vma, vma->vm_flags, + TVA_KHUGEPAGED); + if (!orders) { cc->progress++; continue; } - hstart =3D ALIGN(vma->vm_start, HPAGE_PMD_SIZE); - hend =3D ALIGN_DOWN(vma->vm_end, HPAGE_PMD_SIZE); + + /* + * Coverage is rooted at windows of the largest order the VMA + * allows: below the PMD order that reaches VMAs a whole table + * would not fit in, and parts of a VMA that a whole table would + * leave out. + */ + window =3D PAGE_SIZE << __fls(orders); + hstart =3D ALIGN(vma->vm_start, window); + hend =3D ALIGN_DOWN(vma->vm_end, window); if (khugepaged_scan.address > hend) { cc->progress++; continue; } if (khugepaged_scan.address < hstart) khugepaged_scan.address =3D hstart; - VM_BUG_ON(khugepaged_scan.address & ~HPAGE_PMD_MASK); =20 while (khugepaged_scan.address < hend) { + unsigned long pmd_addr, range_end; bool lock_dropped =3D false; =20 + /* One table's worth at most, and never past the VMA */ + pmd_addr =3D khugepaged_scan.address & HPAGE_PMD_MASK; + range_end =3D min(hend, pmd_addr + HPAGE_PMD_SIZE); + cond_resched(); if (unlikely(collapse_test_exit_or_disable(mm))) goto breakouterloop; =20 - VM_WARN_ON_ONCE(khugepaged_scan.address < hstart || - khugepaged_scan.address + HPAGE_PMD_SIZE > - hend); + VM_WARN_ON_ONCE(khugepaged_scan.address < hstart); =20 *result =3D collapse_single_pmd(khugepaged_scan.address, - khugepaged_scan.address + - HPAGE_PMD_SIZE, - vma, &lock_dropped, cc); + range_end, vma, + &lock_dropped, cc); /* move to next address */ - khugepaged_scan.address +=3D HPAGE_PMD_SIZE; + khugepaged_scan.address =3D range_end; if (lock_dropped) /* * We released mmap_lock so break loop. Note --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BDD1D3EB108; Sun, 16 Aug 2026 22:47:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920435; cv=none; b=l3PlRmI0gcNP5R9QC8N3OIb/1QLNsxJ0pOaqPy7M783dFuAXZEysVLZLS0ZQErxsQxHX/aZFWhfeTMYScat0ZSMP7BVK6OMcroTWGImoiXP3nTAYMEjJWZKl5ZUqoobeKJgAtjHLtvfMW4wJc2xy3dPpY4o7arvYw2YigE1Iu00= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920435; c=relaxed/simple; bh=uwlqtGAEUSObOWmxKNJ1kgnjj1i4ST8v96cyxT381Lc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eSolyiyGskvo+mFWiwuoqk/QB+M9LHpJBiHOogevAFFeU9DwKXXSb3N600fI5MC7hwbj7PVXDRKFWULYWdIXfaYstAhml2Dil2Cau2U7qZNvDIsGcxtMV8qyGnKQiWOMGsq9QLdr3dgStdcD2XHUOn1sbWzHXG9b+W3DihHw7W4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ej7e+TPd; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=c72X5jXN; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ej7e+TPd"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="c72X5jXN" Received: from phl-compute-08.internal (phl-compute-08.internal [10.202.2.48]) by mailfhigh.phl.internal (Postfix) with ESMTP id DBA9414000EB; Sun, 16 Aug 2026 18:47:11 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-08.internal (MEProxy); Sun, 16 Aug 2026 18:47:11 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920431; x= 1787006831; bh=5vxOkKZOXXy3Z946h/BVNJ6lUnR6XJtenZlUIlDq4Kc=; b=e j7e+TPdpEvJk+CLODVtOPxNeRaoCzZ07RmVma+BC9p5W4liU5WJax2GG8a6uW/Pi Hjeyos4Wx4Bt7aAjpu6D4wQdTusdo98Am1Fyk2O3yNlpRlp4KbL0lqxj0K+suyLo jZouI9aP0r/Od4w86Zmvy4TNQCKar1JICU+WlGHHq5JCDJQ1e+d2s8tFy1Rp1zDw JYOapaeZ/U/nJCfvLTxauQnuwItFgs3evlxrrES8vEvSvdS9xmdFfdJpGjh2npA6 1xmK/OqmO8+AUC5vXojNQ7t/xzjm0XTW5FXxt3XVBc2oPibqkGvmvchn/6uS40Oy LkfGQ+lH0zGHHHEkcfiGw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920431; x=1787006831; bh=5 vxOkKZOXXy3Z946h/BVNJ6lUnR6XJtenZlUIlDq4Kc=; b=c72X5jXN4u6nSNsyR WoVzIes8h0Ty++u5ddkjgZQm7semKWLzEVDAnZllTZfqeCUFiZHi0GX0mZSHR+7h krIjvh/eS+7bzAgopHHnhYnTvpIQqJ/H4dtfiFgIUI+hYIzqwUpS/zGgBE0hCImC WxEiWVNLQIZYKAdxanM6kl7oVvcYubWx0WaD5Qk86iEEXz+YJd9cg+gCot5yjfpC I64sZFy9xDWQfeAuD4KPiK4Cf0hs3K5W+lHKTwCkTPvGqEU2cSoGc2zYJJwBkUwl 4odwSvC+uezGPqU92SyCcFEiCVc/XhUC0NMiVe/7ims/vv+eSzEzrxxRU0Aqm1i+ I0FFA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtIG G5k5WQREbOISbTHWF1AMHMtLYEIvOP7kszcil1gNCk78M3Ta+EVbbPBEJ95pPNjpmGGlVN UK1pXd2EFV2Ai/H2/QkOYFvsvoHN3YWCaJDX2TUcCxivhtPjIx1k4KDwFwEEeqlImpYeZv mJE8eWmYgLKN1rPrVdmcQLlhlCai1L/6WQTF49kw0XwE5e/zmV5JutM4uCFLxO5Bm1qWJs 6IcR3Tz23zpt8wa7bFHRkr00+f03DT+UWDRr53j3CESqVEAiKRRp67A/tpae6m8HuCS+p6 ZUHJGZX4VmVY6to7mPwkjyH8FNQMOB1Po5MJl+t/tB6Y7is9ZoGTDP0vL9dA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:10 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 28/57] mm/collapse: remove the mechanism the engine replaces Date: Sun, 16 Aug 2026 23:45:40 +0100 Message-ID: <20260816224609.308019-29-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Nothing reaches the old anonymous collapse any more: the entry point was rewired to the engine, and every function below it lost its last caller. Delete the chain: the scan, the mTHP order walk, the collapse itself, isolation, swap-in, the copy with its success and failure paths, the PTE release helpers, folio_pte_referenced() and the pmd-still-valid check. What stays is what the file paths and MADV_COLLAPSE still call: alloc_charge_folio() for a file collapse's destination, hugepage_vma_revalidate() for the VMA check after MADV_COLLAPSE drops mmap_lock, and count_collapse_event() and collapse_control_init_scan() for the file scan. Four tracepoints lose their only emitter here: mm_khugepaged_scan_pmd, mm_collapse_huge_page, mm_collapse_huge_page_isolate and mm_collapse_huge_page_swapin. Their definitions stay, now without an emitter, and the engine reports through mm_collapse_candidate. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/khugepaged.c | 984 +----------------------------------------------- 1 file changed, 9 insertions(+), 975 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 895183d92fb8..6203473f4953 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -545,370 +545,6 @@ void __khugepaged_exit(struct mm_struct *mm) } } =20 -static void collapse_control_init_scan(struct collapse_control *cc) -{ - memset(cc->node_load, 0, sizeof(cc->node_load)); - nodes_clear(cc->alloc_nmask); - bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); -} - -static void release_pte_folio(struct folio *folio) -{ - node_stat_mod_folio(folio, - NR_ISOLATED_ANON + folio_is_file_lru(folio), - -folio_nr_pages(folio)); - folio_unlock(folio); - folio_putback_lru(folio); -} - -static void release_pte_pages(pte_t *pte, pte_t *_pte, - struct list_head *compound_pagelist) -{ - struct folio *folio, *tmp; - - while (--_pte >=3D pte) { - pte_t pteval =3D ptep_get(_pte); - unsigned long pfn; - - if (pte_none(pteval)) - continue; - VM_WARN_ON_ONCE(!pte_present(pteval)); - pfn =3D pte_pfn(pteval); - if (is_zero_pfn(pfn)) - continue; - folio =3D pfn_folio(pfn); - if (folio_test_large(folio)) - continue; - release_pte_folio(folio); - } - - list_for_each_entry_safe(folio, tmp, compound_pagelist, lru) { - list_del(&folio->lru); - release_pte_folio(folio); - } -} - -/* - * folio_pte_referenced() - Check if a folio or its PTE mapping was recent= ly used - * - * Return: true if recent access was observed through either the folio sta= te - * or the current PTE mapping. - */ -static inline bool folio_pte_referenced(struct folio *folio, - struct vm_area_struct *vma, unsigned long addr, pte_t pteval) -{ - /* The folio was referenced previously ... */ - if (folio_test_young(folio) || folio_test_referenced(folio)) - return true; - /* ... or the PTE mapping was recently used */ - return pte_young(pteval) || mmu_notifier_test_young(vma->vm_mm, addr); -} - -static void count_collapse_event(unsigned int order, enum vm_event_item vm= _event, - enum mthp_stat_item mthp_event) -{ - if (is_pmd_order(order)) - count_vm_event(vm_event); - count_mthp_stat(order, mthp_event); -} - -static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct= *vma, - unsigned long start_addr, pte_t *pte, struct collapse_control *cc, - unsigned int order, struct list_head *compound_pagelist) -{ - const unsigned int max_ptes_none =3D collapse_max_ptes_none(cc, vma, orde= r); - const unsigned int max_ptes_shared =3D collapse_max_ptes_shared(cc, order= ); - const unsigned long nr_pages =3D 1UL << order; - struct page *page =3D NULL; - struct folio *folio =3D NULL; - unsigned long addr =3D start_addr; - pte_t *_pte; - int none_or_zero =3D 0, shared =3D 0, referenced =3D 0; - enum scan_result result =3D SCAN_FAIL; - - for (_pte =3D pte; _pte < pte + nr_pages; - _pte++, addr +=3D PAGE_SIZE) { - pte_t pteval =3D ptep_get(_pte); - if (pte_none_or_zero(pteval)) { - if (++none_or_zero > max_ptes_none) { - result =3D SCAN_EXCEED_NONE_PTE; - count_collapse_event(order, THP_SCAN_EXCEED_NONE_PTE, - MTHP_STAT_COLLAPSE_EXCEED_NONE); - goto out; - } - continue; - } - if (!pte_present(pteval)) { - result =3D SCAN_PTE_NON_PRESENT; - goto out; - } - if (pte_uffd(pteval)) { - result =3D SCAN_PTE_UFFD; - goto out; - } - page =3D vm_normal_page(vma, addr, pteval); - if (unlikely(!page) || unlikely(is_zone_device_page(page))) { - result =3D SCAN_PAGE_NULL; - goto out; - } - - folio =3D page_folio(page); - VM_BUG_ON_FOLIO(!folio_test_anon(folio), folio); - - /* - * If the vma has the VM_DROPPABLE flag, the collapse will - * preserve the lazyfree property without needing to skip. - */ - if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && - folio_test_lazyfree(folio) && !pte_dirty(pteval)) { - result =3D SCAN_PAGE_LAZYFREE; - goto out; - } - - /* See collapse_scan_pmd(). */ - if (folio_maybe_mapped_shared(folio)) { - /* - * TODO: Support shared pages without leading to further - * mTHP collapses. Currently bringing in new pages via - * shared may cause a future higher order collapse on a - * rescan of the same range. - */ - if (++shared > max_ptes_shared) { - result =3D SCAN_EXCEED_SHARED_PTE; - count_collapse_event(order, THP_SCAN_EXCEED_SHARED_PTE, - MTHP_STAT_COLLAPSE_EXCEED_SHARED); - goto out; - } - } - /* - * TODO: In some cases of partially-mapped folios, we'd actually - * want to collapse. - */ - if (!is_pmd_order(order) && folio_order(folio) >=3D order) { - result =3D SCAN_PTE_MAPPED_HUGEPAGE; - goto out; - } - - if (folio_test_large(folio)) { - struct folio *f; - - /* - * Check if we have dealt with the compound page - * already - */ - list_for_each_entry(f, compound_pagelist, lru) { - if (folio =3D=3D f) - goto next; - } - } - - /* - * We can do it before folio_isolate_lru because the - * folio can't be freed from under us. NOTE: folio lock - * is needed to serialize against split_huge_page() - * when invoked from the VM. - */ - if (!folio_trylock(folio)) { - result =3D SCAN_PAGE_LOCK; - goto out; - } - - /* - * Check if the page has any GUP (or other external) pins. - * - * The page table that maps the page has been already unlinked - * from the page table tree and this process cannot get - * an additional pin on the page. - * - * New pins can come later if the page is shared across fork, - * but not from this process. The other process cannot write to - * the page, only trigger CoW. - */ - if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio)) { - folio_unlock(folio); - result =3D SCAN_PAGE_COUNT; - goto out; - } - - /* - * Isolate the folio to avoid collapsing a hugepage - * currently in use by the VM. - */ - if (!folio_isolate_lru(folio)) { - folio_unlock(folio); - result =3D SCAN_DEL_PAGE_LRU; - goto out; - } - node_stat_mod_folio(folio, - NR_ISOLATED_ANON + folio_is_file_lru(folio), - folio_nr_pages(folio)); - VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); - VM_BUG_ON_FOLIO(folio_test_lru(folio), folio); - - if (folio_test_large(folio)) - list_add_tail(&folio->lru, compound_pagelist); -next: - if (cc->policy.require_referenced && - folio_pte_referenced(folio, vma, addr, pteval)) - referenced++; - } - - if (unlikely(cc->policy.require_referenced && !referenced)) { - result =3D SCAN_LACK_REFERENCED_PAGE; - } else { - result =3D SCAN_SUCCEED; - trace_mm_collapse_huge_page_isolate(folio, none_or_zero, - referenced, result, order); - return result; - } -out: - release_pte_pages(pte, _pte, compound_pagelist); - trace_mm_collapse_huge_page_isolate(folio, none_or_zero, - referenced, result, order); - return result; -} - -static void __collapse_huge_page_copy_succeeded(pte_t *pte, - struct vm_area_struct *vma, unsigned long address, - spinlock_t *ptl, unsigned int order, - struct list_head *compound_pagelist) -{ - const unsigned long nr_pages =3D 1UL << order; - unsigned long end =3D address + (PAGE_SIZE * nr_pages); - struct folio *src, *tmp; - pte_t pteval; - pte_t *_pte; - unsigned int nr_ptes; - - for (_pte =3D pte; _pte < pte + nr_pages; _pte +=3D nr_ptes, - address +=3D nr_ptes * PAGE_SIZE) { - nr_ptes =3D 1; - pteval =3D ptep_get(_pte); - if (pte_none_or_zero(pteval)) { - add_mm_counter(vma->vm_mm, MM_ANONPAGES, 1); - if (pte_none(pteval)) - continue; - /* - * ptl mostly unnecessary. - */ - spin_lock(ptl); - ptep_clear(vma->vm_mm, address, _pte); - spin_unlock(ptl); - ksm_might_unmap_zero_page(vma->vm_mm, pteval); - } else { - struct page *src_page =3D pte_page(pteval); - - src =3D page_folio(src_page); - - if (folio_test_large(src)) { - unsigned int max_nr_ptes =3D (end - address) >> PAGE_SHIFT; - - nr_ptes =3D folio_pte_batch(src, _pte, pteval, max_nr_ptes); - } else { - release_pte_folio(src); - } - - /* - * ptl mostly unnecessary, but preempt has to - * be disabled to update the per-cpu stats - * inside folio_remove_rmap_pte(). - */ - spin_lock(ptl); - clear_ptes(vma->vm_mm, address, _pte, nr_ptes); - folio_remove_rmap_ptes(src, src_page, nr_ptes, vma); - spin_unlock(ptl); - free_swap_cache(src); - folio_put_refs(src, nr_ptes); - } - } - - list_for_each_entry_safe(src, tmp, compound_pagelist, lru) { - list_del(&src->lru); - node_stat_sub_folio(src, NR_ISOLATED_ANON + - folio_is_file_lru(src)); - folio_unlock(src); - free_swap_cache(src); - folio_putback_lru(src); - } -} - -static void __collapse_huge_page_copy_failed(pte_t *pte, - pmd_t *pmd, pmd_t orig_pmd, struct vm_area_struct *vma, - unsigned int order, struct list_head *compound_pagelist) -{ - const unsigned long nr_pages =3D 1UL << order; - spinlock_t *pmd_ptl; - - /* - * Re-establish the PMD to point to the original page table - * entry. Restoring PMD needs to be done prior to releasing - * pages. Since pages are still isolated and locked here, - * acquiring anon_vma_lock_write() is unnecessary. - */ - pmd_ptl =3D pmd_lock(vma->vm_mm, pmd); - pmd_populate(vma->vm_mm, pmd, pmd_pgtable(orig_pmd)); - spin_unlock(pmd_ptl); - /* - * Release both raw and compound pages isolated - * in __collapse_huge_page_isolate. - */ - release_pte_pages(pte, pte + nr_pages, compound_pagelist); -} - -/* - * __collapse_huge_page_copy - attempts to copy memory contents from raw - * pages to a hugepage. Cleans up the raw pages if copying succeeds; - * otherwise restores the original page table and releases isolated raw pa= ges. - * Returns SCAN_SUCCEED if copying succeeds, otherwise returns SCAN_COPY_M= C. - * - * @pte: starting of the PTEs to copy from - * @folio: the new hugepage to copy contents to - * @pmd: pointer to the new hugepage's PMD - * @orig_pmd: the original raw pages' PMD - * @vma: the original raw pages' virtual memory area - * @address: starting address to copy - * @ptl: lock on raw pages' PTEs - * @compound_pagelist: list that stores compound pages - */ -static enum scan_result __collapse_huge_page_copy(pte_t *pte, struct folio= *folio, - pmd_t *pmd, pmd_t orig_pmd, struct vm_area_struct *vma, - unsigned long address, spinlock_t *ptl, unsigned int order, - struct list_head *compound_pagelist) -{ - const unsigned long nr_pages =3D 1UL << order; - unsigned int i; - enum scan_result result =3D SCAN_SUCCEED; - - /* - * Copying pages' contents is subject to memory poison at any iteration. - */ - for (i =3D 0; i < nr_pages; i++) { - pte_t pteval =3D ptep_get(pte + i); - struct page *page =3D folio_page(folio, i); - unsigned long src_addr =3D address + i * PAGE_SIZE; - struct page *src_page; - - if (pte_none_or_zero(pteval)) { - clear_user_highpage(page, src_addr); - continue; - } - src_page =3D pte_page(pteval); - if (copy_mc_user_highpage(page, src_page, src_addr, vma) > 0) { - result =3D SCAN_COPY_MC; - break; - } - } - - if (likely(result =3D=3D SCAN_SUCCEED)) - __collapse_huge_page_copy_succeeded(pte, vma, address, ptl, - order, compound_pagelist); - else - __collapse_huge_page_copy_failed(pte, pmd, orig_pmd, vma, - order, compound_pagelist); - - return result; -} - static void khugepaged_alloc_sleep(void) { DEFINE_WAIT(wait); @@ -1089,119 +725,19 @@ enum scan_result find_pmd_or_thp_or_none(struct mm_= struct *mm, return check_pmd_state(*pmd); } =20 -static enum scan_result check_pmd_still_valid(struct mm_struct *mm, - unsigned long address, pmd_t *pmd) +static void count_collapse_event(unsigned int order, enum vm_event_item vm= _event, + enum mthp_stat_item mthp_event) { - pmd_t *new_pmd; - enum scan_result result =3D find_pmd_or_thp_or_none(mm, address, &new_pmd= ); - - if (result !=3D SCAN_SUCCEED) - return result; - if (new_pmd !=3D pmd) - return SCAN_FAIL; - return SCAN_SUCCEED; + if (is_pmd_order(order)) + count_vm_event(vm_event); + count_mthp_stat(order, mthp_event); } =20 -/* - * Bring missing pages in from swap, to complete THP collapse. - * Only done if collapse_scan_pmd() believes it is worthwhile. - * - * For mTHP orders the function bails on the first swap entry, because - * faulting pages back in during collapse could re-populate PTEs that - * push a later scan over the threshold for a higher-order collapse. - * - * Called and returns without pte mapped or spinlocks held. - * Returns result: if not SCAN_SUCCEED, mmap_lock has been released. - */ -static enum scan_result __collapse_huge_page_swapin(struct mm_struct *mm, - struct vm_area_struct *vma, unsigned long start_addr, - pmd_t *pmd, int referenced, unsigned int order) +static void collapse_control_init_scan(struct collapse_control *cc) { - int swapped_in =3D 0; - vm_fault_t ret =3D 0; - unsigned long addr, end =3D start_addr + (PAGE_SIZE << order); - enum scan_result result; - pte_t *pte =3D NULL; - spinlock_t *ptl; - - for (addr =3D start_addr; addr < end; addr +=3D PAGE_SIZE) { - struct vm_fault vmf =3D { - .vma =3D vma, - .address =3D addr, - .pgoff =3D linear_page_index(vma, addr), - .flags =3D FAULT_FLAG_ALLOW_RETRY, - .pmd =3D pmd, - }; - - if (!pte++) { - /* - * Here the ptl is only used to check pte_same() in - * do_swap_page(), so readonly version is enough. - */ - pte =3D pte_offset_map_ro_nolock(mm, pmd, addr, &ptl); - if (!pte) { - mmap_read_unlock(mm); - result =3D SCAN_NO_PTE_TABLE; - goto out; - } - } - - vmf.orig_pte =3D ptep_get_lockless(pte); - if (pte_none(vmf.orig_pte) || - pte_present(vmf.orig_pte)) - continue; - - /* - * TODO: Support swapin without leading to further mTHP - * collapses. Currently bringing in new pages via swapin may - * cause a future higher order collapse on a rescan of the same - * range. - */ - if (!is_pmd_order(order)) { - count_mthp_stat(order, MTHP_STAT_COLLAPSE_EXCEED_SWAP); - pte_unmap(pte); - mmap_read_unlock(mm); - result =3D SCAN_EXCEED_SWAP_PTE; - goto out; - } - - vmf.pte =3D pte; - vmf.ptl =3D ptl; - ret =3D do_swap_page(&vmf); - /* Which unmaps pte (after perhaps re-checking the entry) */ - pte =3D NULL; - - /* - * do_swap_page() returns VM_FAULT_RETRY with released mmap_lock. - * Note we treat VM_FAULT_RETRY as VM_FAULT_ERROR here because - * we do not retry here and swap entry will remain in pagetable - * resulting in later failure. - */ - if (ret & VM_FAULT_RETRY) { - /* Likely, but not guaranteed, that page lock failed */ - result =3D SCAN_PAGE_LOCK; - goto out; - } - if (ret & VM_FAULT_ERROR) { - mmap_read_unlock(mm); - result =3D SCAN_FAIL; - goto out; - } - swapped_in++; - } - - if (pte) - pte_unmap(pte); - - /* Drain LRU cache to remove extra pin on the swapped in pages */ - if (swapped_in) - lru_add_drain(); - - result =3D SCAN_SUCCEED; -out: - trace_mm_collapse_huge_page_swapin(mm, swapped_in, referenced, result, - order); - return result; + memset(cc->node_load, 0, sizeof(cc->node_load)); + nodes_clear(cc->alloc_nmask); + bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); } =20 static enum scan_result alloc_charge_folio(struct folio **foliop, struct m= m_struct *mm, @@ -1234,197 +770,6 @@ static enum scan_result alloc_charge_folio(struct fo= lio **foliop, struct mm_stru return SCAN_SUCCEED; } =20 -/* - * collapse_huge_page() expects the mmap_lock to be unlocked before enteri= ng and - * will always return with the lock unlocked, to avoid holding the mmap_lo= ck - * while allocating a THP, as that could trigger direct reclaim/compaction. - * Note that the VMA must be rechecked after grabbing the mmap_lock again. - */ -static enum scan_result collapse_huge_page(struct mm_struct *mm, unsigned = long start_addr, - int referenced, int unmapped, struct collapse_control *cc, - unsigned int order) -{ - const unsigned long pmd_addr =3D start_addr & HPAGE_PMD_MASK; - const unsigned long end_addr =3D start_addr + (PAGE_SIZE << order); - LIST_HEAD(compound_pagelist); - pmd_t *pmd, _pmd; - pte_t *pte =3D NULL; - pgtable_t pgtable; - struct folio *folio; - spinlock_t *pmd_ptl, *pte_ptl; - enum scan_result result =3D SCAN_FAIL; - struct vm_area_struct *vma; - struct mmu_notifier_range range; - bool anon_vma_locked =3D false; - - result =3D alloc_charge_folio(&folio, mm, cc, order); - if (result !=3D SCAN_SUCCEED) - goto out_nolock; - - if (folio_memcg_alloc_deferred(folio)) { - result =3D SCAN_ALLOC_HUGE_PAGE_FAIL; - goto out_nolock; - } - - mmap_read_lock(mm); - result =3D hugepage_vma_revalidate(mm, pmd_addr, /*expect_anon=3D*/ true, - &vma, cc, order); - if (result !=3D SCAN_SUCCEED) { - mmap_read_unlock(mm); - goto out_nolock; - } - - result =3D find_pmd_or_thp_or_none(mm, pmd_addr, &pmd); - if (result !=3D SCAN_SUCCEED) { - mmap_read_unlock(mm); - goto out_nolock; - } - - if (unmapped) { - /* - * __collapse_huge_page_swapin() will return with mmap_lock - * released when it fails. So we jump out_nolock directly in - * that case. Continuing to collapse causes inconsistency. - */ - result =3D __collapse_huge_page_swapin(mm, vma, start_addr, pmd, - referenced, order); - if (result !=3D SCAN_SUCCEED) - goto out_nolock; - } - - mmap_read_unlock(mm); - /* - * Prevent all access to pagetables with the exception of - * gup_fast later handled by the pmdp_collapse_flush() and the VM - * handled by the anon_vma lock + folio lock. - * - * UFFDIO_MOVE is prevented to race as well thanks to the - * mmap_lock. - */ - mmap_write_lock(mm); - result =3D hugepage_vma_revalidate(mm, pmd_addr, /*expect_anon=3D*/ true, - &vma, cc, order); - if (result !=3D SCAN_SUCCEED) - goto out_up_write; - /* check if the pmd is still valid */ - vma_start_write(vma); - result =3D check_pmd_still_valid(mm, pmd_addr, pmd); - if (result !=3D SCAN_SUCCEED) - goto out_up_write; - - anon_vma_lock_write(vma->anon_vma); - anon_vma_locked =3D true; - - /* - * Only notify about the PTE range we will actually modify. While we - * temporary unmap the whole PTE table for mTHP collapse, we'll remap - * it later, leaving other PTEs effectively unmodified. The locks we - * hold prevent anybody from stumbling over such temporarily unmapped - * PTE tables. - */ - mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, start_addr, - end_addr); - mmu_notifier_invalidate_range_start(&range); - - pmd_ptl =3D pmd_lock(mm, pmd); /* probably unnecessary */ - /* - * This removes any huge TLB entry from the CPU so we won't allow - * huge and small TLB entries for the same virtual address to - * avoid the risk of CPU bugs in that area. - * - * Parallel GUP-fast is fine since GUP-fast will back off when - * it detects PMD is changed. - */ - _pmd =3D pmdp_collapse_flush(vma, pmd_addr, pmd); - spin_unlock(pmd_ptl); - mmu_notifier_invalidate_range_end(&range); - tlb_remove_table_sync_one(); - - pte =3D pte_offset_map_lock(mm, &_pmd, start_addr, &pte_ptl); - if (pte) { - result =3D __collapse_huge_page_isolate(vma, start_addr, pte, cc, - order, &compound_pagelist); - spin_unlock(pte_ptl); - } else { - result =3D SCAN_NO_PTE_TABLE; - } - - if (unlikely(result !=3D SCAN_SUCCEED)) { - spin_lock(pmd_ptl); - VM_WARN_ON_ONCE(!pmd_none(*pmd)); - /* - * We can only use set_pmd_at() when establishing - * hugepmds and never for establishing regular pmds that - * points to regular pagetables. Use pmd_populate() for that - */ - pmd_populate(mm, pmd, pmd_pgtable(_pmd)); - spin_unlock(pmd_ptl); - goto out_up_write; - } - - /* - * For PMD collapse all pages are isolated and locked so anon_vma - * rmap can't run anymore. For mTHP collapse the PMD entry has been - * removed and not all pages are isolated and locked, so we must hold - * the lock to prevent neighboring folios from attempting to access - * this PMD until its reinstalled. - */ - if (is_pmd_order(order)) { - anon_vma_unlock_write(vma->anon_vma); - anon_vma_locked =3D false; - } - - result =3D __collapse_huge_page_copy(pte, folio, pmd, _pmd, - vma, start_addr, pte_ptl, - order, &compound_pagelist); - if (unlikely(result !=3D SCAN_SUCCEED)) - goto out_up_write; - - /* - * The smp_wmb() inside __folio_mark_uptodate() ensures the - * copy_huge_page writes become visible before the set_pmd_at() - * write. - */ - __folio_mark_uptodate(folio); - spin_lock(pmd_ptl); - VM_WARN_ON_ONCE(!pmd_none(*pmd)); - if (is_pmd_order(order)) { - pgtable =3D pmd_pgtable(_pmd); - pgtable_trans_huge_deposit(mm, pmd, pgtable); - map_anon_folio_pmd_nopf(folio, pmd, vma, pmd_addr); - } else { - /* - * Some architectures (e.g. MIPS) walk the live page table in - * their implementation. update_mmu_cache_range() must be called - * with a valid page table hierarchy and the PTE lock held. - * Acquire it nested inside pmd_ptl when they are distinct locks. - */ - if (pte_ptl !=3D pmd_ptl) - spin_lock_nested(pte_ptl, SINGLE_DEPTH_NESTING); - pmd_populate(mm, pmd, pmd_pgtable(_pmd)); - map_anon_folio_pte_nopf(folio, pte, vma, start_addr, - /*uffd_wp=3D*/ false); - if (pte_ptl !=3D pmd_ptl) - spin_unlock(pte_ptl); - } - spin_unlock(pmd_ptl); - - folio =3D NULL; - - result =3D SCAN_SUCCEED; -out_up_write: - if (pte) - pte_unmap(pte); - if (anon_vma_locked) - anon_vma_unlock_write(vma->anon_vma); - mmap_write_unlock(mm); -out_nolock: - if (folio) - folio_put(folio); - trace_mm_collapse_huge_page(mm, result =3D=3D SCAN_SUCCEED, result, order= ); - return result; -} - /* Return the highest naturally aligned order that fits at @offset within = a PMD. */ unsigned int max_order_from_offset(unsigned int offset) { @@ -1434,317 +779,6 @@ unsigned int max_order_from_offset(unsigned int offs= et) return min_t(unsigned int, __ffs(offset), HPAGE_PMD_ORDER); } =20 -/* - * mthp_collapse() consumes the bitmap that is generated during - * collapse_scan_pmd() to determine what regions and mTHP orders fit best. - * - * Each bit in cc->eligible_ptes marks a PTE the scan accepted as a collap= se - * source. We start at the PMD order and check if it is eligible for colla= pse; - * if not, we check the left and right halves of the PTE page table we are - * examining at a lower order. - * - * For each of these, we determine how many PTE entries are occupied in the - * range of PTE entries we propose to collapse, then we compare this to a - * threshold number of PTE entries which would need to be occupied for a - * collapse to be permitted at that order (accounting for max_ptes_none). - * - * If a collapse is permitted, we attempt to collapse the PTE range into a - * mTHP. - */ -static enum scan_result mthp_collapse(struct mm_struct *mm, - unsigned long address, int referenced, int unmapped, - struct collapse_control *cc, unsigned long enabled_orders) -{ - unsigned int nr_occupied_ptes, nr_ptes, max_ptes_none; - enum scan_result last_result =3D SCAN_FAIL; - int collapsed =3D 0; - bool alloc_failed =3D false; - unsigned long collapse_address; - unsigned int offset =3D 0; - unsigned int order =3D HPAGE_PMD_ORDER; - - while (offset < HPAGE_PMD_NR) { - nr_ptes =3D 1UL << order; - - if (!test_bit(order, &enabled_orders)) - goto next_order; - - max_ptes_none =3D collapse_max_ptes_none(cc, NULL, order); - nr_occupied_ptes =3D bitmap_weight_from(cc->eligible_ptes, offset, - offset + nr_ptes); - - /* - * Swap PTEs accepted during the scan are counted in @unmapped, - * not in the eligible bitmap. Account them for the PMD-order - * candidate. - */ - if (is_pmd_order(order)) - nr_occupied_ptes +=3D unmapped; - - if (nr_occupied_ptes >=3D nr_ptes - max_ptes_none) { - enum scan_result ret; - - collapse_address =3D address + offset * PAGE_SIZE; - ret =3D collapse_huge_page(mm, collapse_address, referenced, - unmapped, cc, order); - - switch (ret) { - /* Cases where we continue to next collapse candidate */ - case SCAN_SUCCEED: - collapsed +=3D nr_ptes; - fallthrough; - case SCAN_PTE_MAPPED_HUGEPAGE: - goto next_offset; - /* Cases where lower orders might still succeed */ - case SCAN_ALLOC_HUGE_PAGE_FAIL: - alloc_failed =3D true; - fallthrough; - case SCAN_LACK_REFERENCED_PAGE: - case SCAN_EXCEED_NONE_PTE: - case SCAN_EXCEED_SWAP_PTE: - case SCAN_EXCEED_SHARED_PTE: - case SCAN_PAGE_LOCK: - case SCAN_PAGE_COUNT: - case SCAN_PAGE_NULL: - case SCAN_DEL_PAGE_LRU: - case SCAN_PTE_NON_PRESENT: - case SCAN_PTE_UFFD: - case SCAN_PAGE_LAZYFREE: - last_result =3D ret; - goto next_order; - /* Cases where no further collapse is possible */ - case SCAN_PMD_MAPPED: - fallthrough; - default: - last_result =3D ret; - goto done; - } - } - -next_order: - /* - * Continue with the next smaller order if there is still - * any smaller order enabled. When at the smallest order - * we must always move to the next offset. - */ - if (order > COLLAPSE_MIN_MTHP_ORDER && - (enabled_orders & GENMASK(order - 1, 0))) { - order--; - continue; - } -next_offset: - /* - * Advance past the region we just processed and determine the - * highest order we can attempt next. Since huge pages must be - * naturally aligned, the max order we can attempt next is - * limited by the alignment of the new offset. - * E.g. if we collapsed a order-2 mTHP at offset 0, offset - * becomes 4 and __ffs(4) =3D=3D 2, so the next attempt starts at - * order 2. - */ - offset +=3D nr_ptes; - order =3D max_order_from_offset(offset); - } -done: - if (collapsed) - return SCAN_SUCCEED; - if (alloc_failed) - return SCAN_ALLOC_HUGE_PAGE_FAIL; - return last_result; -} - -static enum scan_result __maybe_unused -collapse_scan_pmd(struct mm_struct *mm, - struct vm_area_struct *vma, unsigned long start_addr, - bool *lock_dropped, struct collapse_control *cc) -{ - const unsigned int max_ptes_shared =3D collapse_max_ptes_shared(cc, HPAGE= _PMD_ORDER); - const unsigned int max_ptes_swap =3D collapse_max_ptes_swap(cc, HPAGE_PMD= _ORDER); - unsigned int max_ptes_none =3D collapse_max_ptes_none(cc, vma, HPAGE_PMD_= ORDER); - enum tva_type tva_flags =3D cc->policy.tva_type; - pmd_t *pmd; - pte_t *pte, *_pte, pteval; - int i; - int none_or_zero =3D 0, shared =3D 0, referenced =3D 0; - enum scan_result result =3D SCAN_FAIL; - struct page *page =3D NULL; - struct folio *folio =3D NULL; - unsigned long addr; - unsigned long enabled_orders; - spinlock_t *ptl; - int node =3D NUMA_NO_NODE, unmapped =3D 0; - - VM_BUG_ON(start_addr & ~HPAGE_PMD_MASK); - - result =3D find_pmd_or_thp_or_none(mm, start_addr, &pmd); - if (result !=3D SCAN_SUCCEED) { - cc->progress++; - goto out; - } - - collapse_control_init_scan(cc); - - enabled_orders =3D collapse_possible_orders(vma, vma->vm_flags, tva_flags= ); - - /* - * If PMD is the only enabled order, enforce max_ptes_none, otherwise - * scan all pages to populate the bitmap for mTHP collapse. The bitmap - * is then checked again in mthp_collapse() for each attempted order. - */ - if (enabled_orders !=3D BIT(HPAGE_PMD_ORDER)) - max_ptes_none =3D KHUGEPAGED_MAX_PTES_LIMIT; - - pte =3D pte_offset_map_lock(mm, pmd, start_addr, &ptl); - if (!pte) { - cc->progress++; - result =3D SCAN_NO_PTE_TABLE; - goto out; - } - - for (i =3D 0; i < HPAGE_PMD_NR; i++) { - _pte =3D pte + i; - addr =3D start_addr + i * PAGE_SIZE; - pteval =3D ptep_get(_pte); - - cc->progress++; - - if (pte_none_or_zero(pteval)) { - if (++none_or_zero > max_ptes_none) { - result =3D SCAN_EXCEED_NONE_PTE; - count_collapse_event(HPAGE_PMD_ORDER, THP_SCAN_EXCEED_NONE_PTE, - MTHP_STAT_COLLAPSE_EXCEED_NONE); - goto out_unmap; - } - continue; - } - if (!pte_present(pteval)) { - if (++unmapped > max_ptes_swap) { - result =3D SCAN_EXCEED_SWAP_PTE; - count_collapse_event(HPAGE_PMD_ORDER, THP_SCAN_EXCEED_SWAP_PTE, - MTHP_STAT_COLLAPSE_EXCEED_SWAP); - goto out_unmap; - } - /* - * Always be strict with uffd-wp - * enabled swap entries. Please see - * comment below for pte_uffd(). - */ - if (pte_swp_uffd_any(pteval)) { - result =3D SCAN_PTE_UFFD; - goto out_unmap; - } - continue; - } - if (pte_uffd(pteval)) { - /* - * Don't collapse the page if any of the small - * PTEs are armed with uffd write protection. - * Here we can also mark the new huge pmd as - * write protected if any of the small ones is - * marked but that could bring unknown - * userfault messages that falls outside of - * the registered range. So, just be simple. - */ - result =3D SCAN_PTE_UFFD; - goto out_unmap; - } - - page =3D vm_normal_page(vma, addr, pteval); - if (unlikely(!page) || unlikely(is_zone_device_page(page))) { - result =3D SCAN_PAGE_NULL; - goto out_unmap; - } - folio =3D page_folio(page); - - /* - * If the vma has the VM_DROPPABLE flag, the collapse will - * preserve the lazyfree property without needing to skip. - */ - if (cc->policy.skip_lazyfree && !(vma->vm_flags & VM_DROPPABLE) && - folio_test_lazyfree(folio) && !pte_dirty(pteval)) { - result =3D SCAN_PAGE_LAZYFREE; - goto out_unmap; - } - - if (!folio_test_anon(folio)) { - result =3D SCAN_PAGE_ANON; - goto out_unmap; - } - - /* - * We treat a single page as shared if any part of the THP - * is shared. - */ - if (folio_maybe_mapped_shared(folio)) { - if (++shared > max_ptes_shared) { - result =3D SCAN_EXCEED_SHARED_PTE; - count_collapse_event(HPAGE_PMD_ORDER, THP_SCAN_EXCEED_SHARED_PTE, - MTHP_STAT_COLLAPSE_EXCEED_SHARED); - goto out_unmap; - } - } - - __set_bit(i, cc->eligible_ptes); - /* - * Record which node the original page is from and save this - * information to cc->node_load[]. - * Khugepaged will allocate hugepage from the node has the max - * hit record. - */ - node =3D folio_nid(folio); - if (collapse_scan_abort(node, cc)) { - result =3D SCAN_SCAN_ABORT; - goto out_unmap; - } - cc->node_load[node]++; - if (!folio_test_lru(folio)) { - result =3D SCAN_PAGE_LRU; - goto out_unmap; - } - if (folio_test_locked(folio)) { - result =3D SCAN_PAGE_LOCK; - goto out_unmap; - } - - /* - * Check if the page has any GUP (or other external) pins. - * - * Here the check is racy, but such case is ephemeral and - * we could always retry collapse later. Anyway the same - * check will be done again later the risk seems low. - */ - if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio)) { - result =3D SCAN_PAGE_COUNT; - goto out_unmap; - } - - if (cc->policy.require_referenced && - folio_pte_referenced(folio, vma, addr, pteval)) - referenced++; - } - if (cc->policy.require_referenced && - (!referenced || - (unmapped && referenced < HPAGE_PMD_NR / 2))) { - result =3D SCAN_LACK_REFERENCED_PAGE; - } else { - result =3D SCAN_SUCCEED; - } -out_unmap: - pte_unmap_unlock(pte, ptl); - if (result =3D=3D SCAN_SUCCEED) { - /* collapse_huge_page() expects the lock to be dropped before calling */ - mmap_read_unlock(mm); - result =3D mthp_collapse(mm, start_addr, referenced, - unmapped, cc, enabled_orders); - /* mmap_lock was released above, set lock_dropped */ - *lock_dropped =3D true; - } -out: - trace_mm_khugepaged_scan_pmd(mm, folio, referenced, - none_or_zero, result, unmapped); - return result; -} - static void collect_mm_slot(struct mm_slot *slot) { struct mm_struct *mm =3D slot->mm; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C41073EA976; Sun, 16 Aug 2026 22:47:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920436; cv=none; b=N7Mzulwu1i4hqKA7phKRpuJZv71RNlF5NWufn+ywlf8AxS2GZVaOi7A/bo3Xio1q9GlL4ZPLC0TAxIm/Fid09Pjj5xOf+cD3zVt38q/2+47jlm5phVpGjLAcOP2n4UkrZKZV657HMo7ycybzAG0PkihwMzHr5zjvxRSExBIevuA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920436; c=relaxed/simple; bh=QIY6pUWnMRARFJZmXUoPldloXAGvNVEu7FWVprChByM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BA/cF0ijyy07x5rA77EtlG2douy6APukh5tgECFUXC1K+VNCwUt0dcn7ZVWh4N3Efu9HhiUDWyqvzK1cOx02mzdgUfRJpF3p9e/ob3JF4aD/wVxTOm2G2hdVPCKDfrusDdc51zXMC+Zi9TmX0KpUIDQSKL0Nuw55yBtpWdoSiFI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Nzmo2MU9; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=cPw1clLh; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Nzmo2MU9"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="cPw1clLh" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id A1C33EC0241; Sun, 16 Aug 2026 18:47:13 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:13 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920433; x= 1787006833; bh=v4o+UQoQZuN2LWwiMtMQO8gOWSU/QLCCuxZH7iTS8uU=; b=N zmo2MU9TZwrrFRAZ6xGNylCaGbUonvBeA/xtuSMYioyuvFDLHSIkGp9ynrpX/LNk 7cKrZXTFjSkDIQvlM/IKGNuMy42XDNj73hPmv79SntLIxYKwCg2sGRUpHEq1BUSq HaurKhr1NuNDOKleEDWdn4tS+8Lpksv31G0u2S0+rkxPM9TDwRneDQWVxVF8iRna Rq/4eS2CPbjoAaerlyjGFwP4QOUjUvqmhbd5ANkGUC85hZd4PW4ddcInBU9LeAQX LjLMy5fgedqxFG6hWYHZXbV8W/FiM+xQvgUXY17mYC7WD/CLqgnPOuBtNXVveFUU bbBAAhxd9NBZ59MvVDClw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920433; x=1787006833; bh=v 4o+UQoQZuN2LWwiMtMQO8gOWSU/QLCCuxZH7iTS8uU=; b=cPw1clLhWCFCWDCZS eKIgUsyQawxGxJvSYFl3N6kirdJFd0ho8m3Eqa6wrgy+YA6Hp4nGh0FvyoicjECL YpmeWK3qAeMfnCbxdNUgJaBBljXKMVawJztxRg09Sqj/zdVCSRTJ52w8mzfYy2dM 9xjh9Bbne/DfZCp0LPW2uMQtmSR5CXuQZn1yiVI52TywqYBt5KNHM4XBK+4wcDW1 Zt0wV7Xn3DOhcEfanGyDNG3DREcPIuSkU2bSob4PN7ICfYMzmlDIU6bX6QxYjZTT /LWQ6eue81Dn3kfxq1NHvLs2Yu6CrEESW0ezOKfRn2Iu6/w540Zkw6g6RgZdBMX3 KVisQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEDk+RLLWMQ2xmtz+SLrRm4SAAyqR0dcBO/jFiSra6UqVmtacLBiFsIEc4u4O0k4e TprFkQPKvch8mvItdBrQvHbw0cB+L9AXDq7j8E1YHObaRHu/NhSSUGMahPtwIv/D6P/2Ck /ZmAzcEVmv8/4SK/QQlss/OdD6ZYE8aT6ou3Os4qcrsJTku67URPEU5YsnAJoGYWAjn3Up 7SKj7HM1m+Bk9heCO7PBXLylkVDNPLNBQlGmsET8D8iA96BjDQqliz/FRaUJqsbc23hZYK 6FfmpWMzIACF8l0zIVIJGtLfQU4v7CVbie56C+ilpPey1dPs7ANiSUKA9aUYDcXfVHEKt5 QL+mvNaMseb/Ht1UO0WB56E3iZGzyAOebwIkEOYKQzzqtStumThDKCMbnt+pabPBUCi6C4 i1X8mX281nuQJE+kMOftIZFvYGyVcJ4EZfk7iq3IwS7NuqcYuxw+BjUBjNWCnYNXKe3ykr ajmbjgA95dA0JtaZDAeNOemV3+7uqoirAmx5+CghSeFHgxZC/NzlROWzFMZZ9a7SEjVKov CsevLGiDtRPJplVpKu9iUjOIQlaE4smjKTqMP5JqBu7Rj0/yeqEqWcLhYu98GV1a7O039C yStEXlYF+K1jnN4lvNCMC8uN6vMTMzcxa2dkpBw+3kVFpSgx8l0seglGw0jw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:13 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 29/57] mm/collapse: move what a collapse is judged on into collapse.c Date: Sun, 16 Aug 2026 23:45:41 +0100 Message-ID: <20260816224609.308019-30-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse.c had to call back into khugepaged.c for nine things: the PMD state checks, the orders a VMA allows, the per-order limits, the NUMA heuristics and the window alignment. The engine therefore depended on the daemon, which is backwards -- khugepaged is one caller of a collapse, not the place a collapse lives. Move them across. max_order_from_offset() and collapse_max_ptes_shared() have no caller left outside collapse.c and land static; the rest stay exported while khugepaged.c's file paths still ask for them. KHUGEPAGED_MAX_PTES_LIMIT goes to collapse.h with them, being both the ceiling the tunables accept and the value they read as "no limit". Every function body is byte for byte what it was, and the dependency now points one way. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 187 +++++++++++++++++++++++++++++++++++++++++++++++ mm/collapse.h | 11 +-- mm/khugepaged.c | 189 ------------------------------------------------ 3 files changed, 191 insertions(+), 196 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index df3760e3918b..43a4b771bbe1 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -113,6 +113,193 @@ min(COLLAPSE_BATCH_BYTES >> (PAGE_SHIFT + COLLAPSE_MIN_MTHP_ORDER), \ COLLAPSE_TABLE_WINDOWS) =20 +static inline enum scan_result check_pmd_state(pmd_t *pmd) +{ + pmd_t pmde =3D pmdp_get_lockless(pmd); + + if (pmd_none(pmde)) + return SCAN_NO_PTE_TABLE; + + /* + * The folio may be under migration when khugepaged is trying to + * collapse it. Migration success or failure will eventually end + * up with a present PMD mapping a folio again. + */ + if (pmd_is_migration_entry(pmde)) + return SCAN_PMD_MAPPED; + if (!pmd_present(pmde)) + return SCAN_NO_PTE_TABLE; + if (pmd_trans_huge(pmde)) + return SCAN_PMD_MAPPED; + if (pmd_bad(pmde)) + return SCAN_NO_PTE_TABLE; + return SCAN_SUCCEED; +} + +enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, + unsigned long address, pmd_t **pmd) +{ + *pmd =3D mm_find_pmd(mm, address); + if (!*pmd) + return SCAN_NO_PTE_TABLE; + + return check_pmd_state(*pmd); +} + +/* + * Check what orders are possible based on the vma and collapse type. + * This is used to determine if mTHP collapse is a viable option. + */ +unsigned long collapse_possible_orders(struct vm_area_struct *vma, + vm_flags_t vm_flags, enum tva_type tva_flags) +{ + unsigned long orders; + + /* If khugepaged is scanning an anonymous vma, allow mTHP collapse */ + if ((tva_flags =3D=3D TVA_KHUGEPAGED) && vma_is_anonymous(vma)) + orders =3D THP_ORDERS_ALL_ANON; + else + orders =3D BIT(HPAGE_PMD_ORDER); + + return thp_vma_allowable_orders(vma, vm_flags, tva_flags, orders); +} + +/* Return the highest naturally aligned order that fits at @offset within = a PMD. */ +static unsigned int max_order_from_offset(unsigned int offset) +{ + if (offset =3D=3D 0) + return HPAGE_PMD_ORDER; + + return min_t(unsigned int, __ffs(offset), HPAGE_PMD_ORDER); +} + +/** + * collapse_max_ptes_none - Calculate maximum allowed empty PTEs or PTEs m= apping + * the shared zeropage for the given collapse operation. + * @cc: The collapse control struct + * @vma: The vma to check for userfaultfd + * @order: The folio order being collapsed to + * + * Return: Maximum number of empty/shared zeropage PTEs for the collapse o= peration + */ +unsigned int collapse_max_ptes_none(struct collapse_control *cc, + struct vm_area_struct *vma, unsigned int order) +{ + const unsigned int max_ptes_none =3D cc->policy.max_ptes_none; + + if (vma && userfaultfd_armed(vma)) + return 0; + /* The limit as given, at the PMD order and wherever it is not capped */ + if (is_pmd_order(order) || !cc->policy.strict_sub_pmd) + return max_ptes_none; + /* + * for mTHP collapse with the sysctl value set to KHUGEPAGED_MAX_PTES_LIM= IT, + * scale the maximum number of PTEs to the order of the collapse. + */ + if (max_ptes_none =3D=3D KHUGEPAGED_MAX_PTES_LIMIT) + return (1 << order) - 1; + /* + * For mTHP collapse of values other than 0 or KHUGEPAGED_MAX_PTES_LIMIT, + * emit a warning and return 0. + */ + if (max_ptes_none) + pr_warn_once("mTHP collapse does not support max_ptes_none" + " values other than 0 or %u, defaulting to 0.\n", + KHUGEPAGED_MAX_PTES_LIMIT); + return 0; +} + +/** + * collapse_max_ptes_shared - Calculate maximum allowed PTEs that map shar= ed + * anonymous pages for the given collapse operation. + * @cc: The collapse control struct + * @order: The folio order being collapsed to + * + * Return: Maximum number of PTEs that map shared anonymous pages for the + * collapse operation + */ +static unsigned int collapse_max_ptes_shared(struct collapse_control *cc, + unsigned int order) +{ + /* + * A sub-PMD window held to the strict rule takes no shared page at all: + * an mTHP is not worth the CoW-breaking. + */ + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) + return 0; + return cc->policy.max_ptes_shared; +} + +/** + * collapse_max_ptes_swap - Calculate the maximum allowed non-present PTEs= or the + * maximum allowed non-present pagecache entries for the given collapse op= eration. + * @cc: The collapse control struct + * @order: The folio order being collapsed to + * + * Return: Maximum number of non-present PTEs or the maximum allowed non-p= resent + * pagecache entries for the collapse operation. + */ +unsigned int collapse_max_ptes_swap(struct collapse_control *cc, + unsigned int order) +{ + /* + * A sub-PMD window held to the strict rule takes nothing non-present: + * reading pages back to build an mTHP is not worth the latency. + */ + if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) + return 0; + return cc->policy.max_ptes_swap; +} + +bool collapse_scan_abort(int nid, struct collapse_control *cc) +{ + int i; + + /* + * If node_reclaim_mode is disabled, then no extra effort is made to + * allocate memory locally. + */ + if (!node_reclaim_enabled()) + return false; + + /* If there is a count for this node already, it must be acceptable */ + if (cc->node_load[nid]) + return false; + + for (i =3D 0; i < MAX_NUMNODES; i++) { + if (!cc->node_load[i]) + continue; + if (node_distance(nid, i) > node_reclaim_distance) + return true; + } + return false; +} + +#ifdef CONFIG_NUMA +int collapse_find_target_node(struct collapse_control *cc) +{ + int nid, target_node =3D 0, max_value =3D 0; + + /* find first node with max normal pages hit */ + for (nid =3D 0; nid < MAX_NUMNODES; nid++) + if (cc->node_load[nid] > max_value) { + max_value =3D cc->node_load[nid]; + target_node =3D nid; + } + + for_each_online_node(nid) { + if (max_value =3D=3D cc->node_load[nid]) + node_set(nid, cc->alloc_nmask); + } + + return target_node; +} +#else +int collapse_find_target_node(struct collapse_control *cc) +{ + return 0; +} +#endif /* * The saved-PTE pool spans a whole table. The byte cap bounds what a rou= nd * holds, but not what one candidate does: a sub-PMD order goes up to diff --git a/mm/collapse.h b/mm/collapse.h index 50a9d59bbf03..2fdb6418653a 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -7,6 +7,9 @@ #include #include =20 +/* Ceiling the max_ptes_* tunables accept, and the value meaning "no limit= " */ +#define KHUGEPAGED_MAX_PTES_LIMIT (HPAGE_PMD_NR - 1) + /* The smallest order a collapse will build, and so the finest window it c= uts */ #define COLLAPSE_MIN_MTHP_ORDER 2 =20 @@ -194,22 +197,16 @@ enum scan_result collapse_anon_pmd(struct mm_struct *= mm, unsigned long start, int collapse_control_init(struct collapse_control *cc); void collapse_control_release(struct collapse_control *cc); =20 -/* - * Defined in khugepaged.c, which still uses them itself. - * TODO: move each into collapse.c once its last khugepaged.c user is gone. - */ unsigned long collapse_possible_orders(struct vm_area_struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags); +enum scan_result check_pmd_state(pmd_t *pmd); enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, unsigned long address, pmd_t **pmd); int collapse_find_target_node(struct collapse_control *cc); bool collapse_scan_abort(int nid, struct collapse_control *cc); -unsigned int max_order_from_offset(unsigned int offset); unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order); unsigned int collapse_max_ptes_swap(struct collapse_control *cc, unsigned int order); -unsigned int collapse_max_ptes_shared(struct collapse_control *cc, - unsigned int order); =20 #endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 6203473f4953..26c0e961ac9f 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -57,7 +57,6 @@ static DECLARE_WAIT_QUEUE_HEAD(khugepaged_wait); * * Note that these are only respected if collapse was initiated by khugepa= ged. */ -#define KHUGEPAGED_MAX_PTES_LIMIT (HPAGE_PMD_NR - 1) unsigned int khugepaged_max_ptes_none __read_mostly; static unsigned int khugepaged_max_ptes_swap __read_mostly; static unsigned int khugepaged_max_ptes_shared __read_mostly; @@ -296,84 +295,6 @@ struct attribute_group khugepaged_attr_group =3D { }; #endif /* CONFIG_SYSFS */ =20 -/** - * collapse_max_ptes_none - Calculate maximum allowed empty PTEs or PTEs m= apping - * the shared zeropage for the given collapse operation. - * @cc: The collapse control struct - * @vma: The vma to check for userfaultfd - * @order: The folio order being collapsed to - * - * Return: Maximum number of empty/shared zeropage PTEs for the collapse o= peration - */ -unsigned int collapse_max_ptes_none(struct collapse_control *cc, - struct vm_area_struct *vma, unsigned int order) -{ - const unsigned int max_ptes_none =3D cc->policy.max_ptes_none; - - if (vma && userfaultfd_armed(vma)) - return 0; - /* The limit as given, at the PMD order and wherever it is not capped */ - if (is_pmd_order(order) || !cc->policy.strict_sub_pmd) - return max_ptes_none; - /* - * for mTHP collapse with the sysctl value set to KHUGEPAGED_MAX_PTES_LIM= IT, - * scale the maximum number of PTEs to the order of the collapse. - */ - if (max_ptes_none =3D=3D KHUGEPAGED_MAX_PTES_LIMIT) - return (1 << order) - 1; - /* - * For mTHP collapse of values other than 0 or KHUGEPAGED_MAX_PTES_LIMIT, - * emit a warning and return 0. - */ - if (max_ptes_none) - pr_warn_once("mTHP collapse does not support max_ptes_none" - " values other than 0 or %u, defaulting to 0.\n", - KHUGEPAGED_MAX_PTES_LIMIT); - return 0; -} - -/** - * collapse_max_ptes_shared - Calculate maximum allowed PTEs that map shar= ed - * anonymous pages for the given collapse operation. - * @cc: The collapse control struct - * @order: The folio order being collapsed to - * - * Return: Maximum number of PTEs that map shared anonymous pages for the - * collapse operation - */ -unsigned int collapse_max_ptes_shared(struct collapse_control *cc, - unsigned int order) -{ - /* - * A sub-PMD window held to the strict rule takes no shared page at all: - * an mTHP is not worth the CoW-breaking. - */ - if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) - return 0; - return cc->policy.max_ptes_shared; -} - -/** - * collapse_max_ptes_swap - Calculate the maximum allowed non-present PTEs= or the - * maximum allowed non-present pagecache entries for the given collapse op= eration. - * @cc: The collapse control struct - * @order: The folio order being collapsed to - * - * Return: Maximum number of non-present PTEs or the maximum allowed non-p= resent - * pagecache entries for the collapse operation. - */ -unsigned int collapse_max_ptes_swap(struct collapse_control *cc, - unsigned int order) -{ - /* - * A sub-PMD window held to the strict rule takes nothing non-present: - * reading pages back to build an mTHP is not worth the latency. - */ - if (!is_pmd_order(order) && cc->policy.strict_sub_pmd) - return 0; - return cc->policy.max_ptes_swap; -} - int hugepage_madvise(struct vm_area_struct *vma, vm_flags_t *vm_flags, int advice) { @@ -483,24 +404,6 @@ void __khugepaged_enter(struct mm_struct *mm) wake_up_interruptible(&khugepaged_wait); } =20 -/* - * Check what orders are possible based on the vma and collapse type. - * This is used to determine if mTHP collapse is a viable option. - */ -unsigned long collapse_possible_orders(struct vm_area_struct *vma, - vm_flags_t vm_flags, enum tva_type tva_flags) -{ - unsigned long orders; - - /* If khugepaged is scanning an anonymous vma, allow mTHP collapse */ - if ((tva_flags =3D=3D TVA_KHUGEPAGED) && vma_is_anonymous(vma)) - orders =3D THP_ORDERS_ALL_ANON; - else - orders =3D BIT(HPAGE_PMD_ORDER); - - return thp_vma_allowable_orders(vma, vm_flags, tva_flags, orders); -} - static bool collapse_possible(struct vm_area_struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags) { @@ -559,30 +462,6 @@ static struct collapse_control khugepaged_collapse_con= trol =3D { .is_khugepaged =3D true, }; =20 -bool collapse_scan_abort(int nid, struct collapse_control *cc) -{ - int i; - - /* - * If node_reclaim_mode is disabled, then no extra effort is made to - * allocate memory locally. - */ - if (!node_reclaim_enabled()) - return false; - - /* If there is a count for this node already, it must be acceptable */ - if (cc->node_load[nid]) - return false; - - for (i =3D 0; i < MAX_NUMNODES; i++) { - if (!cc->node_load[i]) - continue; - if (node_distance(nid, i) > node_reclaim_distance) - return true; - } - return false; -} - #define khugepaged_defrag() \ (transparent_hugepage_flags & \ (1<tva_type =3D TVA_FORCED_COLLAPSE; } =20 -#ifdef CONFIG_NUMA -int collapse_find_target_node(struct collapse_control *cc) -{ - int nid, target_node =3D 0, max_value =3D 0; - - /* find first node with max normal pages hit */ - for (nid =3D 0; nid < MAX_NUMNODES; nid++) - if (cc->node_load[nid] > max_value) { - max_value =3D cc->node_load[nid]; - target_node =3D nid; - } - - for_each_online_node(nid) { - if (max_value =3D=3D cc->node_load[nid]) - node_set(nid, cc->alloc_nmask); - } - - return target_node; -} -#else -int collapse_find_target_node(struct collapse_control *cc) -{ - return 0; -} -#endif - /* * If mmap_lock temporarily dropped, revalidate vma * after taking the mmap_lock again. @@ -692,39 +545,6 @@ static enum scan_result hugepage_vma_revalidate(struct= mm_struct *mm, unsigned l return SCAN_SUCCEED; } =20 -static inline enum scan_result check_pmd_state(pmd_t *pmd) -{ - pmd_t pmde =3D pmdp_get_lockless(pmd); - - if (pmd_none(pmde)) - return SCAN_NO_PTE_TABLE; - - /* - * The folio may be under migration when khugepaged is trying to - * collapse it. Migration success or failure will eventually end - * up with a present PMD mapping a folio again. - */ - if (pmd_is_migration_entry(pmde)) - return SCAN_PMD_MAPPED; - if (!pmd_present(pmde)) - return SCAN_NO_PTE_TABLE; - if (pmd_trans_huge(pmde)) - return SCAN_PMD_MAPPED; - if (pmd_bad(pmde)) - return SCAN_NO_PTE_TABLE; - return SCAN_SUCCEED; -} - -enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, - unsigned long address, pmd_t **pmd) -{ - *pmd =3D mm_find_pmd(mm, address); - if (!*pmd) - return SCAN_NO_PTE_TABLE; - - return check_pmd_state(*pmd); -} - static void count_collapse_event(unsigned int order, enum vm_event_item vm= _event, enum mthp_stat_item mthp_event) { @@ -770,15 +590,6 @@ static enum scan_result alloc_charge_folio(struct foli= o **foliop, struct mm_stru return SCAN_SUCCEED; } =20 -/* Return the highest naturally aligned order that fits at @offset within = a PMD. */ -unsigned int max_order_from_offset(unsigned int offset) -{ - if (offset =3D=3D 0) - return HPAGE_PMD_ORDER; - - return min_t(unsigned int, __ffs(offset), HPAGE_PMD_ORDER); -} - static void collect_mm_slot(struct mm_slot *slot) { struct mm_struct *mm =3D slot->mm; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CDC343F6C2A; Sun, 16 Aug 2026 22:47:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920437; cv=none; b=jZ2V8cBqhDYv7QRMfkC5ii1CayFhABhtLuzr3WZ1G1cdQ2FcvZPAeI0QdWSEI0WodQqMLL4IteaXq8koXErLFOUDE9iAGJi2OUho0FF97x9qgnaVSmSS2xX+EvMWbhd/0QDgn1zSO1YR51RPdaHUDT6+pUV5A5DPGDA9paIMruY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920437; c=relaxed/simple; bh=KD4sXY94Vuz2+nX0EfaUT6taLHsHeW/DxE8KoKSqk/k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ihTyO7kKNb6lfWOKch1L5OW0QnTENe4Vkday4Rug31cbW7IrfiGDOg357Tv/4oaT88PxXN0fTYNjTCnFhwYZ4wpmcwOltVzGTA3pIo9Szy3NqpUmvaDY8pcX6EZtg+5QYYznfDetySychiF7ckzgjRKHvrTRmu6w5vNufqh2N84= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=V5giir+F; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=GApl5JHf; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="V5giir+F"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="GApl5JHf" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 38940EC0242; Sun, 16 Aug 2026 18:47:15 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:15 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920435; x= 1787006835; bh=0Xh3qJFy3fUbrlN8Fy5RgHt7peoEA3/8bbzab7kXN/s=; b=V 5giir+FskG5CJ7dclo4ePNvd6SAwJqQ4yiFRkHkuQgTdwtX43TgBGCcujN18hpPC RtjvbXARIqnkKrr2hpRHpBYlBOeLlSJ0nyu/Al0vs+246WUCf9WlXUarCqwF2kFE omjK4nbjMDwDvNu6N7pPXuHBG7/Ey3yVmLU6Zv7W1ayBFV4aSyKwfQ5N04Xxtyjc u9glyesF+QjUFHPyqBM+qiOeyYSuAQTTlXgUCAUuC0mSsYGk9W9AxivavcATJvxf Gr42crPrhx4gE+UqaTqeT7txuIL8NINf/LYmAUQ8Yr9ahMsFyPw+ilLbnPy9ht8X RgDOoqXtnvB7pw/3FLNMQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920435; x=1787006835; bh=0 Xh3qJFy3fUbrlN8Fy5RgHt7peoEA3/8bbzab7kXN/s=; b=GApl5JHf4d23ZEorQ EMstbMXZUb1M4UgrkGwCFZMAFdXcPxOT6ROYUSU/cH8dAgE9NiO09jAP+z9H5Xhv VEABoGBQvAtGiWNxDNEfau5NaGo5dfHGNdJv6fzrB6LNxrLVw55XJQ8RwVn8UiNR haN8wb0Ssgxm8r44TK+Kw9fwY7ZWiUvDWxZA4ElLSbfgNDrymH6sIj2P10BWjvVv IADR1Ntgw2akgUPcVEsyf1QHZDh+sFA5OpHEr8RISRW8FjyO7cQR1zV2NE6vPcJc Fpi39Ghn8vQCcTQM+xAHBOLTOWKjWchajHE8H0BYbYn6GXdfYf862gBGJ+iaZDDI +BTtA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEDk+RLLWMQ2xmtz+SLrRm4SAAyqR0dcBO/jFiSra6UqVmtacLBiFsIEc4u4O0k4e TprFkQPKvch8mvItdBrQvHbw0cB+L9AXDq7j8E1YHObaRHu/NhSSUGMahPtwIv/D6P/2Ck /ZmAzcEVmv8/4SK/QQlss/OdD6ZYE8aT6ou3Os4qcrsJTku67URPEU5YsnAJoGYWAjn3Up 7SKj7HM1m+Bk9heCO7PBXLylkVDNPLNBQlGmsET8D8iA96BjDQqliz/FRaUJqsbc23hZYK 6FfmpWMzIACF8l0zIVIJGtLfQU4v7CVbie56C+ilpPey1dPs7ANiSUKA9aUYDcXfVHEKIB CBopL77cQG0VLaTu2WtaEqdiaZ5xCAD/3YiiT0Lq44Ym9vdIpYdhzK+jQXietixt7epVkB if3RNtS6b9w3JJ41Es85ighZoOawaKfc83t0IoYu4uoikAtuRK+JbqxDq/gaHo9YYFOBBq B65m56HfcLycTC3OwAjrSioOFCBDb40VhLH64PtRBYsZjA0+1n3CrhGCunvVAvf9rb6dEp 0jfbxFptvtcfTYM69fB3mqD0o8Q3g3qeOYT+NVvJ+WlCKBClxubhwGBnyR/4vMG7vwnYjS uhnzb4yQEXQ/H5n2RPqlQYRsWTuxmq3VmvuT/9rB2mqs0kuK37xxe+LH8GWw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:14 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 30/57] mm/collapse: name the max_ptes ceiling after collapse Date: Sun, 16 Aug 2026 23:45:42 +0100 Message-ID: <20260816224609.308019-31-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The constant lives in collapse.h and is read by the limit helpers there, so KHUGEPAGED_MAX_PTES_LIMIT names it after one of its callers rather than after what it is. Rename it to COLLAPSE_MAX_PTES_LIMIT, matching the other constant the header defines. khugepaged keeps using it to bound what its tunables accept. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 8 ++++---- mm/collapse.h | 2 +- mm/khugepaged.c | 8 ++++---- 3 files changed, 9 insertions(+), 9 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 43a4b771bbe1..ae7c2777b279 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -193,19 +193,19 @@ unsigned int collapse_max_ptes_none(struct collapse_c= ontrol *cc, if (is_pmd_order(order) || !cc->policy.strict_sub_pmd) return max_ptes_none; /* - * for mTHP collapse with the sysctl value set to KHUGEPAGED_MAX_PTES_LIM= IT, + * for mTHP collapse with the sysctl value set to COLLAPSE_MAX_PTES_LIMIT, * scale the maximum number of PTEs to the order of the collapse. */ - if (max_ptes_none =3D=3D KHUGEPAGED_MAX_PTES_LIMIT) + if (max_ptes_none =3D=3D COLLAPSE_MAX_PTES_LIMIT) return (1 << order) - 1; /* - * For mTHP collapse of values other than 0 or KHUGEPAGED_MAX_PTES_LIMIT, + * For mTHP collapse of values other than 0 or COLLAPSE_MAX_PTES_LIMIT, * emit a warning and return 0. */ if (max_ptes_none) pr_warn_once("mTHP collapse does not support max_ptes_none" " values other than 0 or %u, defaulting to 0.\n", - KHUGEPAGED_MAX_PTES_LIMIT); + COLLAPSE_MAX_PTES_LIMIT); return 0; } =20 diff --git a/mm/collapse.h b/mm/collapse.h index 2fdb6418653a..94c11051f06a 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -8,7 +8,7 @@ #include =20 /* Ceiling the max_ptes_* tunables accept, and the value meaning "no limit= " */ -#define KHUGEPAGED_MAX_PTES_LIMIT (HPAGE_PMD_NR - 1) +#define COLLAPSE_MAX_PTES_LIMIT (HPAGE_PMD_NR - 1) =20 /* The smallest order a collapse will build, and so the finest window it c= uts */ #define COLLAPSE_MIN_MTHP_ORDER 2 diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 26c0e961ac9f..8c770e251c22 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -214,7 +214,7 @@ static ssize_t max_ptes_none_store(struct kobject *kobj, unsigned long max_ptes_none; =20 err =3D kstrtoul(buf, 10, &max_ptes_none); - if (err || max_ptes_none > KHUGEPAGED_MAX_PTES_LIMIT) + if (err || max_ptes_none > COLLAPSE_MAX_PTES_LIMIT) return -EINVAL; =20 khugepaged_max_ptes_none =3D max_ptes_none; @@ -239,7 +239,7 @@ static ssize_t max_ptes_swap_store(struct kobject *kobj, unsigned long max_ptes_swap; =20 err =3D kstrtoul(buf, 10, &max_ptes_swap); - if (err || max_ptes_swap > KHUGEPAGED_MAX_PTES_LIMIT) + if (err || max_ptes_swap > COLLAPSE_MAX_PTES_LIMIT) return -EINVAL; =20 khugepaged_max_ptes_swap =3D max_ptes_swap; @@ -265,7 +265,7 @@ static ssize_t max_ptes_shared_store(struct kobject *ko= bj, unsigned long max_ptes_shared; =20 err =3D kstrtoul(buf, 10, &max_ptes_shared); - if (err || max_ptes_shared > KHUGEPAGED_MAX_PTES_LIMIT) + if (err || max_ptes_shared > COLLAPSE_MAX_PTES_LIMIT) return -EINVAL; =20 khugepaged_max_ptes_shared =3D max_ptes_shared; @@ -330,7 +330,7 @@ int __init khugepaged_init(void) return -ENOMEM; =20 khugepaged_pages_to_scan =3D HPAGE_PMD_NR * 8; - khugepaged_max_ptes_none =3D KHUGEPAGED_MAX_PTES_LIMIT; + khugepaged_max_ptes_none =3D COLLAPSE_MAX_PTES_LIMIT; khugepaged_max_ptes_swap =3D HPAGE_PMD_NR / 8; khugepaged_max_ptes_shared =3D HPAGE_PMD_NR / 2; =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E2A63F823C; Sun, 16 Aug 2026 22:47:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920439; cv=none; b=i8IM+l9UwId3zvT3MjfhHnbqfP71iNYJUNIuo+t3n+pF5rMyGL3TEoTG+3CsybKnc0gYehsfKaA3gO7PYH44r8BpLPBFoPKA8yVwgE3Nbegs3YRDJ0cCMaRj+HKr/mgFL2akDLCAVYIhBnVaSGMarvzuL+Nf22/cTAs8a/CZrc0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920439; c=relaxed/simple; bh=toGrETRjFAsfJJZlY7GOsW89MNqqeWVBezibGn00oKU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gVD382TkiMp01am8bNAvWaklaPqgTmUtEL5DZ6GR75ShEqZadPeaVwnkaSIwZndfeluD1F3AG2pAfsG0fLaGDsqC1X6c8Ga7DUU9JR5WPSNbhY8j9l69nFGCqVnRmjpiW8GKgmt7VK09R1Oep8zNgCYlsRneC/+s5nT+vv1H0xo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=nFwpijY7; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=kzT8VXqp; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="nFwpijY7"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="kzT8VXqp" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.phl.internal (Postfix) with ESMTP id D792414000F8; Sun, 16 Aug 2026 18:47:16 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:47:16 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920436; x= 1787006836; bh=j0njKa/sU97UuU9BC7LS/teFWmvcsdJR3szFnz2Zzu0=; b=n FwpijY7TTDo0MOviKSKnc53xYc9LVC2yNmby+rLfbuoxdBpQG/3VfPzA5ScaZeQk VJBprsmLnSS9QYbRHPPzoMg7QYsT+NyU/2WTHigZ6gqh7N5H0aRa+obNo5lMnz9+ LIUDzFyF/ZjCZErzr6zVwQ1Bqrn2IKjWUutXndjWljG5vv3jDQQftjn0QIoD64xC 2f3eqT0rHMRdKWY2omcSMPpKCKRv08oyTTG89ScCptgCjx1VNjOa5VqeDkOPl38C cUpqL5oI7MPT6jutBz2ks6pLq9EAP+cGTQHvOQC0TApQdjIbK8mUwp1tIg001iOI S0IJE7xr9gVjevZ4YO4uA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920436; x=1787006836; bh=j 0njKa/sU97UuU9BC7LS/teFWmvcsdJR3szFnz2Zzu0=; b=kzT8VXqpML+W8GnWr QIdRiSkVJ/ZiLN2FHS7sFEVJfNIkVJg351UzPogM2LODLj3cfPmCV2O9dlYdokxr umwrk5AyQHqeMU7FJZzpfG6+TZuCyKnYVn3KOU6almFI5JnAvIrL+MV99NFSX5SW BtosKOzNt3bvMBShA6ykTWrhsT0qdk38rzrhAHFBCDBJL/gjhrbFym1ENyCWWliM zZmoMQCjlY9/SJJOUQperbWC6Je6wJBw+2Dfw1mNkNjIR06KkwiROJR3rtRq3F5/ zbmaOWTlngBrg331nR1ETZ1AKbvzjLjbaWWCILyxzYcPsjs4w8vqyype7S48NZtx Y5b1Q== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTD5 ylGlg2zq7m16jfM5uHJgPlIYbMxnPsaFzwmzt7lRLFaMXcTaKz54PlFbIvjMQOW6UMRTdo qJTi5mkLOtGPpWwYdcJg3eekaqRuyFT/TOyP3TJ0aR58vKWg9nB3VYpnFoaCeuEyCZDDvw 1qnFPw7SZrPjoeir7CLHWxY6h3tVea27Z8Lz3Tztdek1L3AyXroPtcmRBi7gaOu/mMWaiG rWJ7/Akg0ggNjjzmQJ916tdv+9CwLGhh6B0ikeVXChonVeoq7RofIRBCPlDV8zcgCVXDWt nLtTsuB2Kqk3/b5bDoi4k+l6PtTz8vhwuFM3WW6Bwr3j7x1JzGPB+AxSWHrg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:16 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 31/57] mm/khugepaged: count collapses where khugepaged makes them Date: Sun, 16 Aug 2026 23:45:43 +0100 Message-ID: <20260816224609.308019-32-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_single_pmd() bumps khugepaged_pages_collapsed for its caller, and tests cc->is_khugepaged to know whether it should: the counter belongs to the daemon, and MADV_COLLAPSE must not touch it. So the one thing the shared path still asks about its caller is bookkeeping, not policy. Count in khugepaged's own walk instead, at the call it already makes. The question goes away, and cc->is_khugepaged with it -- nothing else read it. current_is_khugepaged() is a different test, on the task rather than on the request. Preparation for moving the dispatcher into collapse.c, from where a static in khugepaged.c is out of reach. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 2 -- mm/khugepaged.c | 9 +++------ 2 files changed, 3 insertions(+), 8 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index 94c11051f06a..11f51c6ea444 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -114,8 +114,6 @@ struct collapse_policy { struct collapse_control { struct collapse_policy policy; =20 - bool is_khugepaged; - /* Num pages scanned per node */ u32 node_load[MAX_NUMNODES]; =20 diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 8c770e251c22..907ed1131460 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -458,9 +458,7 @@ static void khugepaged_alloc_sleep(void) remove_wait_queue(&khugepaged_wait, &wait); } =20 -static struct collapse_control khugepaged_collapse_control =3D { - .is_khugepaged =3D true, -}; +static struct collapse_control khugepaged_collapse_control; =20 #define khugepaged_defrag() \ (transparent_hugepage_flags & \ @@ -1639,8 +1637,6 @@ static enum scan_result collapse_single_pmd(unsigned = long addr, mmap_read_unlock(mm); } end: - if (cc->is_khugepaged && result =3D=3D SCAN_SUCCEED) - ++khugepaged_pages_collapsed; return result; } =20 @@ -1731,6 +1727,8 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, *result =3D collapse_single_pmd(khugepaged_scan.address, range_end, vma, &lock_dropped, cc); + if (*result =3D=3D SCAN_SUCCEED) + ++khugepaged_pages_collapsed; /* move to next address */ khugepaged_scan.address =3D range_end; if (lock_dropped) @@ -2039,7 +2037,6 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, cc =3D kmalloc_obj(*cc); if (!cc) return -ENOMEM; - cc->is_khugepaged =3D false; collapse_policy_forced(&cc->policy); cc->progress =3D 0; err =3D collapse_control_init(cc); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8111A3F8882; Sun, 16 Aug 2026 22:47:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920443; cv=none; b=PkjFTllJj4ZAULEgo8L3eF9ERvwe4SF10TlVewBq41PYz64/6TRIaCIBySpgfCOvqgLhr0v8DTgJWfbdZezxglq8YM7wskow3iK+ZCK2fnMgfUAHUuzDdMTVXSEIfPQ6XY71UPMVMBzrAljSBptgP3SacQ5+6cJRiYZHV8cmebw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920443; c=relaxed/simple; bh=LTpf3QiJUzBnBhJbC7mTGRuQ3C28/aalAzECSZJzfUc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=u9dqd0CF1Q1O5MxFpu1AeBq4tk2TSDszsRQL4mqHzi0ewX1e0ctEn7NBtHYN8uOcumVW6ZMirBs/QMPoqLiR4Pp0ATwlKXJmae/q3hEsNMOjwEP634vwYQb7ism0nu2y6CfBNSGZ23NOsTLf0lyT3FVYUjrZt+ZsBqhZZefVmNM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=XSU+L665; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=kxeVJIdK; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="XSU+L665"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="kxeVJIdK" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id CF79FEC0074; Sun, 16 Aug 2026 18:47:18 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:18 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920438; x= 1787006838; bh=KEajPWpWYWlbGhN/CguiGDhcHomV0Q9qRnl8pYmoAV4=; b=X SU+L665jfN6LvWryDqfVXrFM4L1PfaTutPm03RYuGnvG75EY64RNFj1Pz5nhNwC4 YTBBzyAsoWAIthD3pTsb4D9nmjk7yye/igrhfJvan1uYmZm8XwdhDvOo1yq2Pd60 u7wCoPboz7qkRNAj4pL+GLvfE+vAIfTA/ifC7l1XwUu9zQh19JO6o3iKZoMmaKi4 MCyMIZJf52wTHIppKYMSoRVcdj/k10ijuZM8qOMPy1W2C/IczwP6QIdHs72SZ3EW dC5df+q97j8Dw9GIc66+zBp8IG0JP6jMNmAvIw4x62omNKj7ruRjdTZ8hqhpHKCZ /0Aiggb6Hpjsp9UEVPrrg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920438; x=1787006838; bh=K EajPWpWYWlbGhN/CguiGDhcHomV0Q9qRnl8pYmoAV4=; b=kxeVJIdKQXkPD+zDW Hc3s+nxBL2O5P58eG4AxfdxazNX8ZbzU6GfvLdsC03iDk8RwRp6FmwKjuJdl4Ir3 LjFIr2oQMVPl70W9N8oa8Xd/xXGgg8TdxnSjFLogsAQkMTJizWETZI50jF778WAm lomOU1aG4X1eHhmalt4dNrCq37BGkFT/u0y2FCSel13vpMpMYlfkmXkKCZVYW/rH jqDGfOajCuvOrnuMctMNNPr0P+w+RSIU6kKPR9s0H3m+rsfUFGRR3BaWTJETYEaT zTklzGlhO1xW21ihy6WhUHYJhLW5UEi6s+oKo62m/ikw1wJh9Am3dzjdIMriUqZm qlQrQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWPmjqOfRj5F992yj4B4BZj1raD29b+9/uQ9rFyvaV+kr+0iJRDW/7N8sznN5hp3 Jd0ytp1vu59bDzCFT/O8S/efKgYbCfD2rLp5ESvFXHBu+C+T7VPMHXocEWQmqRcTwN39e9 beda2Ri+DmyXoylgECNlOmesAYSvAqjyQbPp39VciQhjDZLAdlalALq+T39tfvcBsAftOx 0pZrMKicu8IddtrFqvBcFdG/5WkfJWvoJzqNqmEsI17lT+SJI8PEV94Md+CC1CjuERkMm2 pW/0xtbQUhoKznxP8UdY68yvPC1FKG1sdD5CGoRhqiz0Y+ik+TmkR1x+Rv8Bp8o0DAMUtZ cE0l5giu9WU8FxtH87akRwOLggtEGz9Dv2wcd/9L9aAdAWMzcaZbfT1wtLjnSx1hT6I9zI 1Y4LA1YLLh2q43e+CQmfwta5t0xx3lsVCiUD7cC2BwIeD1CTfTqiUkc0HojLf7rC4PM4N6 yZ/NnNw9MvpbU4dorT50kf8UOvmLFYyeUvbl7TV6stAyqkFLLKJjVI/ym70PglOCy9e0jG Jzdl+FjX8OQv7exDKeImToZak6fo4n7jWXfEBYv5NrACReeimfTQRLFLcu5aRty+EeRXQj NuAcg5VNidCBlztgm/DhHQAIp0S7HZ535kLsEqSc/5F1AYHi++FW9jf1pWHQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:17 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 32/57] mm/collapse: move the file collapse into collapse.c Date: Sun, 16 Aug 2026 23:45:44 +0100 Message-ID: <20260816224609.308019-33-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapse is split across two files with no boundary to speak of: the engine in collapse.c, the file and shmem half in khugepaged.c, helpers reaching in both directions, and the function that chooses between them sitting with the daemon. Move the rest of the mechanism over -- the file collapse and its scan, the PTE-mapped-THP recollapse and the table retraction it needs, and the allocator they share -- and with them collapse_single_pmd(), which dispatches on the VMA. khugepaged.c keeps what is actually khugepaged: the tunables, the daemon and its scan budget, the mm_slot bookkeeping, and the walk that decides which ranges to offer. Five helpers lose their last caller outside collapse.c and go static, as do the engine's two entries now that the dispatcher reaches them from the same file. check_pmd_state() was already static inline, so its declaration was only ever dead. collapse_single_pmd() takes their place in collapse.h, keeping the interface it has today: a VMA to dispatch on and an out-param saying whether the lock survived. Pure motion: every function arrives in collapse.c exactly as it left khugepaged.c. Beyond that the diff has only the four includes the moved code needs, and the header declarations that changed hands. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 1095 ++++++++++++++++++++++++++++++++++++++++++++++- mm/collapse.h | 17 +- mm/khugepaged.c | 1075 ---------------------------------------------- 3 files changed, 1090 insertions(+), 1097 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index ae7c2777b279..21bfbc038044 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -1,7 +1,9 @@ // SPDX-License-Identifier: GPL-2.0 #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt =20 +#include #include +#include #include #include #include /* x86 flush_tlb_range() uses hstate_vma() */ @@ -11,8 +13,10 @@ #include #include #include +#include #include #include +#include #include #include #include @@ -136,7 +140,7 @@ static inline enum scan_result check_pmd_state(pmd_t *p= md) return SCAN_SUCCEED; } =20 -enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, +static enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, unsigned long address, pmd_t **pmd) { *pmd =3D mm_find_pmd(mm, address); @@ -182,7 +186,7 @@ static unsigned int max_order_from_offset(unsigned int = offset) * * Return: Maximum number of empty/shared zeropage PTEs for the collapse o= peration */ -unsigned int collapse_max_ptes_none(struct collapse_control *cc, +static unsigned int collapse_max_ptes_none(struct collapse_control *cc, struct vm_area_struct *vma, unsigned int order) { const unsigned int max_ptes_none =3D cc->policy.max_ptes_none; @@ -239,7 +243,7 @@ static unsigned int collapse_max_ptes_shared(struct col= lapse_control *cc, * Return: Maximum number of non-present PTEs or the maximum allowed non-p= resent * pagecache entries for the collapse operation. */ -unsigned int collapse_max_ptes_swap(struct collapse_control *cc, +static unsigned int collapse_max_ptes_swap(struct collapse_control *cc, unsigned int order) { /* @@ -251,7 +255,7 @@ unsigned int collapse_max_ptes_swap(struct collapse_con= trol *cc, return cc->policy.max_ptes_swap; } =20 -bool collapse_scan_abort(int nid, struct collapse_control *cc) +static bool collapse_scan_abort(int nid, struct collapse_control *cc) { int i; =20 @@ -276,7 +280,7 @@ bool collapse_scan_abort(int nid, struct collapse_contr= ol *cc) } =20 #ifdef CONFIG_NUMA -int collapse_find_target_node(struct collapse_control *cc) +static int collapse_find_target_node(struct collapse_control *cc) { int nid, target_node =3D 0, max_value =3D 0; =20 @@ -295,7 +299,7 @@ int collapse_find_target_node(struct collapse_control *= cc) return target_node; } #else -int collapse_find_target_node(struct collapse_control *cc) +static int collapse_find_target_node(struct collapse_control *cc) { return 0; } @@ -2110,7 +2114,7 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) * that acts on what it found hands the range to collapse_anon_pmd() after= wards, * without the lock. */ -enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, +static enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, unsigned long start, unsigned long end, struct collapse_control *cc) { @@ -2521,7 +2525,7 @@ static void collapse_add_candidate(struct collapse_co= ntrol *cc, * largest order downwards. Returns what the table yielded: a collapse, or * the reason it did not. */ -enum scan_result collapse_anon_pmd(struct mm_struct *mm, unsigned long sta= rt, +static enum scan_result collapse_anon_pmd(struct mm_struct *mm, unsigned l= ong start, unsigned long end, struct collapse_control *cc) { @@ -2583,3 +2587,1078 @@ enum scan_result collapse_anon_pmd(struct mm_struc= t *mm, unsigned long start, return cc->scan_refusal; return cc->select_result; } + +static void count_collapse_event(unsigned int order, enum vm_event_item vm= _event, + enum mthp_stat_item mthp_event) +{ + if (is_pmd_order(order)) + count_vm_event(vm_event); + count_mthp_stat(order, mthp_event); +} + +static void collapse_control_init_scan(struct collapse_control *cc) +{ + memset(cc->node_load, 0, sizeof(cc->node_load)); + nodes_clear(cc->alloc_nmask); + bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); +} + +static enum scan_result alloc_charge_folio(struct folio **foliop, struct m= m_struct *mm, + struct collapse_control *cc, unsigned int order) +{ + gfp_t gfp =3D cc->policy.gfp; + int node =3D collapse_find_target_node(cc); + struct folio *folio; + + folio =3D __folio_alloc(gfp, order, node, &cc->alloc_nmask); + if (!folio) { + *foliop =3D NULL; + count_collapse_event(order, THP_COLLAPSE_ALLOC_FAILED, + MTHP_STAT_COLLAPSE_ALLOC_FAILED); + return SCAN_ALLOC_HUGE_PAGE_FAIL; + } + + count_collapse_event(order, THP_COLLAPSE_ALLOC, MTHP_STAT_COLLAPSE_ALLOC); + + if (unlikely(mem_cgroup_charge(folio, mm, gfp))) { + folio_put(folio); + *foliop =3D NULL; + return SCAN_CGROUP_CHARGE_FAIL; + } + + if (is_pmd_order(order)) + count_memcg_folio_events(folio, THP_COLLAPSE_ALLOC, 1); + + *foliop =3D folio; + return SCAN_SUCCEED; +} + +/* folio must be locked, and mmap_lock must be held */ +static enum scan_result set_huge_pmd(struct vm_area_struct *vma, unsigned = long addr, + pmd_t *pmdp, struct folio *folio, struct page *page) +{ + struct mm_struct *mm =3D vma->vm_mm; + struct vm_fault vmf =3D { + .vma =3D vma, + .address =3D addr, + .flags =3D 0, + }; + pgd_t *pgdp; + p4d_t *p4dp; + pud_t *pudp; + + mmap_assert_locked(vma->vm_mm); + + if (!pmdp) { + pgdp =3D pgd_offset(mm, addr); + p4dp =3D p4d_alloc(mm, pgdp, addr); + if (!p4dp) + return SCAN_FAIL; + pudp =3D pud_alloc(mm, p4dp, addr); + if (!pudp) + return SCAN_FAIL; + pmdp =3D pmd_alloc(mm, pudp, addr); + if (!pmdp) + return SCAN_FAIL; + } + + vmf.pmd =3D pmdp; + if (do_set_pmd(&vmf, folio, page)) + return SCAN_FAIL; + + folio_get(folio); + return SCAN_SUCCEED; +} + +static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, = unsigned long addr, + bool install_pmd) +{ + enum scan_result result =3D SCAN_FAIL; + int nr_mapped_ptes =3D 0; + unsigned int nr_batch_ptes; + struct mmu_notifier_range range; + bool notified =3D false; + unsigned long haddr =3D addr & HPAGE_PMD_MASK; + unsigned long end =3D haddr + HPAGE_PMD_SIZE; + struct vm_area_struct *vma =3D vma_lookup(mm, haddr); + struct folio *folio; + pte_t *start_pte, *pte; + pmd_t *pmd, pgt_pmd; + spinlock_t *pml =3D NULL, *ptl; + int i; + + mmap_assert_locked(mm); + + /* First check VMA found, in case page tables are being torn down */ + if (!vma || !vma->vm_file || + !range_in_vma(vma, haddr, haddr + HPAGE_PMD_SIZE)) + return SCAN_VMA_CHECK; + + /* Fast check before locking page if already PMD-mapped */ + result =3D find_pmd_or_thp_or_none(mm, haddr, &pmd); + if (result =3D=3D SCAN_PMD_MAPPED) + return result; + + /* + * If we are here, we've succeeded in replacing all the native pages + * in the page cache with a single hugepage. If a mm were to fault-in + * this memory (mapped by a suitably aligned VMA), we'd get the hugepage + * and map it by a PMD, regardless of sysfs THP settings. As such, let's + * analogously elide sysfs THP settings here and force collapse. + */ + if (!thp_vma_allowable_order(vma, vma->vm_flags, TVA_FORCED_COLLAPSE, PMD= _ORDER)) + return SCAN_VMA_CHECK; + + /* + * Keep pmd pgtable while the uffd bit is in use; see comment in + * retract_page_tables(). + */ + if (userfaultfd_protected(vma)) + return SCAN_PTE_UFFD; + + folio =3D filemap_lock_folio(vma->vm_file->f_mapping, + linear_page_index(vma, haddr)); + if (IS_ERR(folio)) + return SCAN_PAGE_NULL; + + if (!is_pmd_order(folio_order(folio))) { + result =3D SCAN_PAGE_COMPOUND; + goto drop_folio; + } + + result =3D find_pmd_or_thp_or_none(mm, haddr, &pmd); + switch (result) { + case SCAN_SUCCEED: + break; + case SCAN_NO_PTE_TABLE: + /* + * All pte entries have been removed and pmd cleared. + * Skip all the pte checks and just update the pmd mapping. + */ + goto maybe_install_pmd; + default: + goto drop_folio; + } + + result =3D SCAN_FAIL; + start_pte =3D pte_offset_map_lock(mm, pmd, haddr, &ptl); + if (!start_pte) /* mmap_lock + page lock should prevent this */ + goto drop_folio; + + /* step 1: check all mapped PTEs are to the right huge page */ + for (i =3D 0, addr =3D haddr, pte =3D start_pte; + i < HPAGE_PMD_NR; i++, addr +=3D PAGE_SIZE, pte++) { + struct page *page; + pte_t ptent =3D ptep_get(pte); + + /* empty pte, skip */ + if (pte_none(ptent)) + continue; + + /* page swapped out, abort */ + if (!pte_present(ptent)) { + result =3D SCAN_PTE_NON_PRESENT; + goto abort; + } + + page =3D vm_normal_page(vma, addr, ptent); + if (WARN_ON_ONCE(page && is_zone_device_page(page))) + page =3D NULL; + /* + * Note that uprobe, debugger, or MAP_PRIVATE may change the + * page table, but the new page will not be a subpage of hpage. + */ + if (folio_page(folio, i) !=3D page) + goto abort; + } + + pte_unmap_unlock(start_pte, ptl); + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, + haddr, haddr + HPAGE_PMD_SIZE); + mmu_notifier_invalidate_range_start(&range); + notified =3D true; + + /* + * pmd_lock covers a wider range than ptl, and (if split from mm's + * page_table_lock) ptl nests inside pml. The less time we hold pml, + * the better; but userfaultfd's mfill_atomic_pte() on a private VMA + * inserts a valid as-if-COWed PTE without even looking up page cache. + * So page lock of folio does not protect from it, so we must not drop + * ptl before pgt_pmd is removed, so uffd private needs pml taken now. + */ + if (userfaultfd_armed(vma) && !(vma->vm_flags & VM_SHARED)) + pml =3D pmd_lock(mm, pmd); + + start_pte =3D pte_offset_map_rw_nolock(mm, pmd, haddr, &pgt_pmd, &ptl); + if (!start_pte) /* mmap_lock + page lock should prevent this */ + goto abort; + if (!pml) + spin_lock(ptl); + else if (ptl !=3D pml) + spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); + + if (unlikely(!pmd_same(pgt_pmd, pmdp_get_lockless(pmd)))) + goto abort; + + /* step 2: clear page table and adjust rmap */ + for (i =3D 0, addr =3D haddr, pte =3D start_pte; i < HPAGE_PMD_NR; + i +=3D nr_batch_ptes, addr +=3D nr_batch_ptes * PAGE_SIZE, + pte +=3D nr_batch_ptes) { + unsigned int max_nr_batch_ptes =3D (end - addr) >> PAGE_SHIFT; + struct page *page; + pte_t ptent =3D ptep_get(pte); + + nr_batch_ptes =3D 1; + + if (pte_none(ptent)) + continue; + /* + * We dropped ptl after the first scan, to do the mmu_notifier: + * page lock stops more PTEs of the folio being faulted in, but + * does not stop write faults COWing anon copies from existing + * PTEs; and does not stop those being swapped out or migrated. + */ + if (!pte_present(ptent)) { + result =3D SCAN_PTE_NON_PRESENT; + goto abort; + } + page =3D vm_normal_page(vma, addr, ptent); + + if (folio_page(folio, i) !=3D page) + goto abort; + + nr_batch_ptes =3D folio_pte_batch(folio, pte, ptent, max_nr_batch_ptes); + + /* + * Must clear entry, or a racing truncate may re-remove it. + * TLB flush can be left until pmdp_collapse_flush() does it. + * PTE dirty? Shmem page is already dirty; file is read-only. + */ + clear_ptes(mm, addr, pte, nr_batch_ptes); + folio_remove_rmap_ptes(folio, page, nr_batch_ptes, vma); + nr_mapped_ptes +=3D nr_batch_ptes; + } + + if (!pml) + spin_unlock(ptl); + + /* step 3: set proper refcount and mm_counters. */ + if (nr_mapped_ptes) { + folio_ref_sub(folio, nr_mapped_ptes); + add_mm_counter(mm, mm_counter_file(folio), -nr_mapped_ptes); + } + + /* step 4: remove empty page table */ + if (!pml) { + pml =3D pmd_lock(mm, pmd); + if (ptl !=3D pml) { + spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); + if (unlikely(!pmd_same(pgt_pmd, pmdp_get_lockless(pmd)))) { + flush_tlb_mm(mm); + goto unlock; + } + } + } + pgt_pmd =3D pmdp_collapse_flush(vma, haddr, pmd); + pmdp_get_lockless_sync(); + pte_unmap_unlock(start_pte, ptl); + if (ptl !=3D pml) + spin_unlock(pml); + + mmu_notifier_invalidate_range_end(&range); + + mm_dec_nr_ptes(mm); + page_table_check_pte_clear_range(mm, haddr, pgt_pmd); + pte_free_defer(mm, pmd_pgtable(pgt_pmd)); + +maybe_install_pmd: + /* step 5: install pmd entry */ + result =3D install_pmd + ? set_huge_pmd(vma, haddr, pmd, folio, &folio->page) + : SCAN_SUCCEED; + goto drop_folio; +abort: + if (nr_mapped_ptes) { + flush_tlb_mm(mm); + folio_ref_sub(folio, nr_mapped_ptes); + add_mm_counter(mm, mm_counter_file(folio), -nr_mapped_ptes); + } +unlock: + if (start_pte) + pte_unmap_unlock(start_pte, ptl); + if (pml && pml !=3D ptl) + spin_unlock(pml); + if (notified) + mmu_notifier_invalidate_range_end(&range); +drop_folio: + folio_unlock(folio); + folio_put(folio); + return result; +} + +/** + * collapse_pte_mapped_thp - Try to collapse a pte-mapped THP for mm at + * address haddr. + * + * @mm: process address space where collapse happens + * @addr: THP collapse address + * @install_pmd: If a huge PMD should be installed + * + * This function checks whether all the PTEs in the PMD are pointing to the + * right THP. If so, retract the page table so the THP can refault in with + * as pmd-mapped. Possibly install a huge PMD mapping the THP. + */ +void collapse_pte_mapped_thp(struct mm_struct *mm, unsigned long addr, + bool install_pmd) +{ + try_collapse_pte_mapped_thp(mm, addr, install_pmd); +} + +/* Can we retract page tables for this file-backed VMA? */ +static bool file_backed_vma_is_retractable(struct vm_area_struct *vma) +{ + /* + * Check vma->anon_vma to exclude MAP_PRIVATE mappings that + * got written to. These VMAs are likely not worth removing + * page tables from, as PMD-mapping is likely to be split later. + */ + if (READ_ONCE(vma->anon_vma)) + return false; + + /* + * When a vma is registered with uffd-wp or RWP, we cannot recycle + * the page table because there may be pte markers installed. + * VM_UFFD_RWP ranges similarly rely on per-PTE uffd state + * and cannot be recycled to a shared PMD. Other vmas can still + * have the same file mapped hugely, but skip this one: it will + * always be mapped in small page size for these registrations. + */ + if (userfaultfd_protected(vma)) + return false; + + /* + * If the VMA contains guard regions then we can't collapse it. + * + * This is set atomically on guard marker installation under mmap/VMA + * read lock, and here we may not hold any VMA or mmap lock at all. + * + * This is therefore serialised on the PTE page table lock, which is + * obtained on guard region installation after the flag is set, so this + * check being performed under this lock excludes races. + */ + if (vma_test_atomic_flag(vma, VMA_MAYBE_GUARD_BIT)) + return false; + + return true; +} + +static void retract_page_tables(struct address_space *mapping, pgoff_t pgo= ff) +{ + struct vm_area_struct *vma; + + i_mmap_lock_read(mapping); + mapping_rmap_tree_foreach(vma, mapping, pgoff, pgoff) { + struct mmu_notifier_range range; + struct mm_struct *mm; + unsigned long addr; + pmd_t *pmd, pgt_pmd; + spinlock_t *pml; + spinlock_t *ptl; + bool success =3D false; + + addr =3D vma->vm_start + + ((pgoff - vma_start_pgoff(vma)) << PAGE_SHIFT); + if (addr & ~HPAGE_PMD_MASK || + vma->vm_end < addr + HPAGE_PMD_SIZE) + continue; + + mm =3D vma->vm_mm; + if (find_pmd_or_thp_or_none(mm, addr, &pmd) !=3D SCAN_SUCCEED) + continue; + + if (collapse_test_exit(mm)) + continue; + + if (!file_backed_vma_is_retractable(vma)) + continue; + + /* PTEs were notified when unmapped; but now for the PMD? */ + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, + addr, addr + HPAGE_PMD_SIZE); + mmu_notifier_invalidate_range_start(&range); + + pml =3D pmd_lock(mm, pmd); + /* + * The lock of new_folio is still held, we will be blocked in + * the page fault path, which prevents the pte entries from + * being set again. So even though the old empty PTE page may be + * concurrently freed and a new PTE page is filled into the pmd + * entry, it is still empty and can be removed. + * + * So here we only need to recheck if the state of pmd entry + * still meets our requirements, rather than checking pmd_same() + * like elsewhere. + */ + if (check_pmd_state(pmd) !=3D SCAN_SUCCEED) + goto drop_pml; + ptl =3D pte_lockptr(mm, pmd); + if (ptl !=3D pml) + spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); + + /* + * Huge page lock is still held, so normally the page table must + * remain empty; and we have already skipped anon_vma and + * userfaultfd_wp() vmas. But since the mmap_lock is not held, + * it is still possible for a racing userfaultfd_ioctl() or + * madvise() to have inserted ptes or markers. Now that we hold + * ptlock, repeating the retractable checks protects us from + * races against the prior checks. + */ + if (likely(file_backed_vma_is_retractable(vma))) { + pgt_pmd =3D pmdp_collapse_flush(vma, addr, pmd); + pmdp_get_lockless_sync(); + success =3D true; + } + + if (ptl !=3D pml) + spin_unlock(ptl); +drop_pml: + spin_unlock(pml); + + mmu_notifier_invalidate_range_end(&range); + + if (success) { + mm_dec_nr_ptes(mm); + page_table_check_pte_clear_range(mm, addr, pgt_pmd); + pte_free_defer(mm, pmd_pgtable(pgt_pmd)); + } + } + i_mmap_unlock_read(mapping); +} + +/** + * collapse_file - collapse filemap/tmpfs/shmem pages into huge one. + * + * @mm: process address space where collapse happens + * @addr: virtual collapse start address + * @file: file that collapse on + * @start: collapse start address + * @cc: collapse context and scratchpad + * + * Basic scheme is simple, details are more complex: + * - allocate and lock a new huge page; + * - scan page cache, locking old pages + * + swap/gup in pages if necessary; + * - copy data to new page + * - handle shmem holes + * + re-validate that holes weren't filled by someone else + * + check for userfaultfd + * - finalize updates to the page cache; + * - if replacing succeeds: + * + unlock huge page; + * + free old pages; + * - if replacing failed; + * + unlock old pages + * + unlock and free huge page; + */ +static enum scan_result collapse_file(struct mm_struct *mm, unsigned long = addr, + struct file *file, pgoff_t start, struct collapse_control *cc) +{ + struct address_space *mapping =3D file->f_mapping; + struct page *dst; + struct folio *folio, *tmp, *new_folio; + pgoff_t index =3D 0, end =3D start + HPAGE_PMD_NR; + LIST_HEAD(pagelist); + XA_STATE_ORDER(xas, &mapping->i_pages, start, HPAGE_PMD_ORDER); + enum scan_result result =3D SCAN_SUCCEED; + int nr_none =3D 0; + bool is_shmem =3D shmem_file(file); + + /* + * MADV_COLLAPSE ignores shmem huge config, so do not check shmem + * + * TODO: once shmem always calls mapping_set_large_folios() on its + * mapping, the shmem check can be removed. + */ + VM_WARN_ON_ONCE(!is_shmem && !mapping_pmd_folio_support(mapping)); + VM_WARN_ON_ONCE(start & (HPAGE_PMD_NR - 1)); + + result =3D alloc_charge_folio(&new_folio, mm, cc, HPAGE_PMD_ORDER); + if (result !=3D SCAN_SUCCEED) + goto out; + + mapping_set_update(&xas, mapping); + + __folio_set_locked(new_folio); + if (is_shmem) + __folio_set_swapbacked(new_folio); + new_folio->index =3D start; + new_folio->mapping =3D mapping; + + /* + * Ensure we have slots for all the pages in the range. This is + * almost certainly a no-op because most of the pages must be present + */ + do { + xas_lock_irq(&xas); + xas_create_range(&xas); + if (!xas_error(&xas)) + break; + xas_unlock_irq(&xas); + if (!xas_nomem(&xas, GFP_KERNEL)) { + result =3D SCAN_FAIL; + goto rollback; + } + } while (1); + + for (index =3D start; index < end;) { + xas_set(&xas, index); + folio =3D xas_load(&xas); + + VM_BUG_ON(index !=3D xas.xa_index); + if (is_shmem) { + if (!folio) { + /* + * Stop if extent has been truncated or + * hole-punched, and is now completely + * empty. + */ + if (index =3D=3D start) { + if (!xas_next_entry(&xas, end - 1)) { + result =3D SCAN_TRUNCATED; + goto xa_locked; + } + } + nr_none++; + index++; + continue; + } + + if (xa_is_value(folio) || !folio_test_uptodate(folio)) { + xas_unlock_irq(&xas); + /* swap in or instantiate fallocated page */ + if (shmem_get_folio(mapping->host, index, 0, + &folio, SGP_NOALLOC)) { + result =3D SCAN_FAIL; + goto xa_unlocked; + } + /* drain lru cache to help folio_isolate_lru() */ + lru_add_drain(); + } else if (folio_trylock(folio)) { + folio_get(folio); + xas_unlock_irq(&xas); + } else { + result =3D SCAN_PAGE_LOCK; + goto xa_locked; + } + } else { /* !is_shmem */ + if (!folio || xa_is_value(folio)) { + xas_unlock_irq(&xas); + page_cache_sync_readahead(mapping, &file->f_ra, + file, index, + end - index); + /* drain lru cache to help folio_isolate_lru() */ + lru_add_drain(); + folio =3D filemap_lock_folio(mapping, index); + if (IS_ERR(folio)) { + result =3D SCAN_FAIL; + goto xa_unlocked; + } + } else if (folio_test_dirty(folio)) { + /* + * This page is dirty because it hasn't + * been flushed since first write. + * + * Trigger async flush for read-only files and + * hope the writeback is done when khugepaged + * revisits this page. Writable files can have + * their folios dirty at any time; blindly + * flushing them would cause undesirable + * system-wide writeback. + * + * This is a one-off situation. We are not + * forcing writeback in loop. + */ + xas_unlock_irq(&xas); + if (!inode_is_open_for_write(mapping->host)) + filemap_flush(mapping); + result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; + goto xa_unlocked; + } else if (folio_test_writeback(folio)) { + xas_unlock_irq(&xas); + result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; + goto xa_unlocked; + } else if (folio_trylock(folio)) { + folio_get(folio); + xas_unlock_irq(&xas); + } else { + result =3D SCAN_PAGE_LOCK; + goto xa_locked; + } + } + + /* + * The folio must be locked, so we can drop the i_pages lock + * without racing with truncate. + */ + VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); + + /* make sure the folio is up to date */ + if (unlikely(!folio_test_uptodate(folio))) { + result =3D SCAN_FAIL; + goto out_unlock; + } + + /* + * If file was truncated then extended, or hole-punched, before + * we locked the first folio, then a THP might be there already. + * This will be discovered on the first iteration. + */ + if (is_pmd_order(folio_order(folio))) { + result =3D SCAN_PTE_MAPPED_HUGEPAGE; + goto out_unlock; + } + + if (folio_mapping(folio) !=3D mapping) { + result =3D SCAN_TRUNCATED; + goto out_unlock; + } + + if (!is_shmem && (folio_test_dirty(folio) || + folio_test_writeback(folio))) { + /* + * khugepaged only works on clean file-backed folios, + * so this folio is dirty because it hasn't been flushed + * since first write. + */ + result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; + goto out_unlock; + } + + if (!folio_isolate_lru(folio)) { + result =3D SCAN_DEL_PAGE_LRU; + goto out_unlock; + } + + if (!filemap_release_folio(folio, GFP_KERNEL)) { + result =3D SCAN_PAGE_HAS_PRIVATE; + folio_putback_lru(folio); + goto out_unlock; + } + + if (folio_mapped(folio)) + try_to_unmap(folio, + TTU_IGNORE_MLOCK | TTU_BATCH_FLUSH); + + xas_lock_irq(&xas); + + VM_BUG_ON_FOLIO(folio !=3D xa_load(xas.xa, index), folio); + + /* + * We control 2 + nr_pages references to the folio: + * - we hold a pin on it; + * - nr_pages reference from page cache; + * - one from lru_isolate_folio; + * If those are the only references, then any new usage + * of the folio will have to fetch it from the page + * cache. That requires locking the folio to handle + * truncate, so any new usage will be blocked until we + * unlock folio after collapse/during rollback. + */ + if (folio_ref_count(folio) !=3D 2 + folio_nr_pages(folio)) { + result =3D SCAN_PAGE_COUNT; + xas_unlock_irq(&xas); + folio_putback_lru(folio); + goto out_unlock; + } + + /* + * At this point, the folio is locked and unmapped. If the PTE + * was dirty, try_to_unmap() has transferred the dirty bit to + * the folio and we must not collapse it into a clean + * file-backed folio. + * + * If the folio is clean here, no one can write it until we + * drop the folio lock. A write through a stale TLB entry came + * from a clean PTE and must fault because the PTE has been + * cleared; the fault path has to take the folio lock before + * installing a writable mapping. Buffered write paths also + * have to take the folio lock before modifying file contents + * without a mapping, typically via write_begin_get_folio(). + */ + if (!is_shmem && folio_test_dirty(folio)) { + result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; + xas_unlock_irq(&xas); + folio_putback_lru(folio); + goto out_unlock; + } + + /* + * Accumulate the folios that are being collapsed. + */ + list_add_tail(&folio->lru, &pagelist); + index +=3D folio_nr_pages(folio); + continue; +out_unlock: + folio_unlock(folio); + folio_put(folio); + goto xa_unlocked; + } + +xa_locked: + xas_unlock_irq(&xas); +xa_unlocked: + + /* + * If collapse is successful, flush must be done now before copying. + * If collapse is unsuccessful, does flush actually need to be done? + * Do it anyway, to clear the state. + */ + try_to_unmap_flush(); + + if (result =3D=3D SCAN_SUCCEED && nr_none && + !shmem_charge(mapping->host, nr_none)) + result =3D SCAN_FAIL; + if (result !=3D SCAN_SUCCEED) { + nr_none =3D 0; + goto rollback; + } + + /* + * The old folios are locked, so they won't change anymore. + */ + index =3D start; + dst =3D folio_page(new_folio, 0); + list_for_each_entry(folio, &pagelist, lru) { + int i, nr_pages =3D folio_nr_pages(folio); + + while (index < folio->index) { + clear_highpage(dst); + index++; + dst++; + } + + for (i =3D 0; i < nr_pages; i++) { + if (copy_mc_highpage(dst, folio_page(folio, i)) > 0) { + result =3D SCAN_COPY_MC; + goto rollback; + } + index++; + dst++; + } + } + while (index < end) { + clear_highpage(dst); + index++; + dst++; + } + + if (nr_none) { + struct vm_area_struct *vma; + int nr_none_check =3D 0; + + i_mmap_lock_read(mapping); + xas_lock_irq(&xas); + + xas_set(&xas, start); + for (index =3D start; index < end; index++) { + if (!xas_next(&xas)) { + xas_store(&xas, XA_RETRY_ENTRY); + if (xas_error(&xas)) { + result =3D SCAN_STORE_FAILED; + goto immap_locked; + } + nr_none_check++; + } + } + + if (nr_none !=3D nr_none_check) { + result =3D SCAN_PAGE_FILLED; + goto immap_locked; + } + + /* + * If userspace observed a missing page in a VMA with + * a MODE_MISSING userfaultfd, then it might expect a + * UFFD_EVENT_PAGEFAULT for that page. If so, we need to + * roll back to avoid suppressing such an event. Since + * wp/minor userfaultfds don't give userspace any + * guarantees that the kernel doesn't fill a missing + * page with a zero page, so they don't matter here. + * + * Any userfaultfds registered after this point will + * not be able to observe any missing pages due to the + * previously inserted retry entries. + */ + mapping_rmap_tree_foreach(vma, mapping, start, end) { + if (userfaultfd_missing(vma)) { + result =3D SCAN_EXCEED_NONE_PTE; + goto immap_locked; + } + } + +immap_locked: + i_mmap_unlock_read(mapping); + if (result !=3D SCAN_SUCCEED) { + xas_set(&xas, start); + for (index =3D start; index < end; index++) { + if (xas_next(&xas) =3D=3D XA_RETRY_ENTRY) + xas_store(&xas, NULL); + } + + xas_unlock_irq(&xas); + goto rollback; + } + } else { + xas_lock_irq(&xas); + } + + if (is_shmem) { + lruvec_stat_mod_folio(new_folio, NR_SHMEM, HPAGE_PMD_NR); + lruvec_stat_mod_folio(new_folio, NR_SHMEM_THPS, HPAGE_PMD_NR); + } else { + lruvec_stat_mod_folio(new_folio, NR_FILE_THPS, HPAGE_PMD_NR); + } + lruvec_stat_mod_folio(new_folio, NR_FILE_PAGES, HPAGE_PMD_NR); + + /* + * Mark new_folio as uptodate before inserting it into the + * page cache so that it isn't mistaken for an fallocated but + * unwritten page. + */ + folio_mark_uptodate(new_folio); + folio_ref_add(new_folio, HPAGE_PMD_NR - 1); + + if (is_shmem) + folio_mark_dirty(new_folio); + folio_add_lru(new_folio); + + /* Join all the small entries into a single multi-index entry. */ + xas_set_order(&xas, start, HPAGE_PMD_ORDER); + xas_store(&xas, new_folio); + WARN_ON_ONCE(xas_error(&xas)); + xas_unlock_irq(&xas); + + /* + * Remove pte page tables, so we can re-fault the page as huge. A caller + * that wants the PMD mapped now is told to go and do that. + */ + retract_page_tables(mapping, start); + if (cc->policy.install_pmd) + result =3D SCAN_PTE_MAPPED_HUGEPAGE; + folio_unlock(new_folio); + + /* + * The collapse has succeeded, so free the old folios. + */ + list_for_each_entry_safe(folio, tmp, &pagelist, lru) { + list_del(&folio->lru); + lruvec_stat_mod_folio(folio, NR_FILE_PAGES, + -folio_nr_pages(folio)); + if (is_shmem) + lruvec_stat_mod_folio(folio, NR_SHMEM, + -folio_nr_pages(folio)); + folio->mapping =3D NULL; + folio_clear_active(folio); + folio_clear_unevictable(folio); + folio_unlock(folio); + folio_put_refs(folio, 2 + folio_nr_pages(folio)); + } + + goto out; + +rollback: + /* Something went wrong: roll back page cache changes */ + if (nr_none) { + xas_lock_irq(&xas); + mapping->nrpages -=3D nr_none; + xas_unlock_irq(&xas); + shmem_uncharge(mapping->host, nr_none); + } + + list_for_each_entry_safe(folio, tmp, &pagelist, lru) { + list_del(&folio->lru); + folio_unlock(folio); + folio_putback_lru(folio); + folio_put(folio); + } + + new_folio->mapping =3D NULL; + + folio_unlock(new_folio); + folio_put(new_folio); +out: + VM_BUG_ON(!list_empty(&pagelist)); + trace_mm_khugepaged_collapse_file(mm, new_folio, index, addr, is_shmem, f= ile, HPAGE_PMD_NR, result); + return result; +} + +static enum scan_result collapse_scan_file(struct mm_struct *mm, + unsigned long addr, struct file *file, pgoff_t start, + struct collapse_control *cc) +{ + const unsigned int max_ptes_none =3D collapse_max_ptes_none(cc, NULL, HPA= GE_PMD_ORDER); + const unsigned int max_ptes_swap =3D collapse_max_ptes_swap(cc, HPAGE_PMD= _ORDER); + struct folio *folio =3D NULL; + struct address_space *mapping =3D file->f_mapping; + XA_STATE(xas, &mapping->i_pages, start); + int present, swap; + int node =3D NUMA_NO_NODE; + enum scan_result result =3D SCAN_SUCCEED; + + present =3D 0; + swap =3D 0; + collapse_control_init_scan(cc); + rcu_read_lock(); + xas_for_each(&xas, folio, start + HPAGE_PMD_NR - 1) { + if (xas_retry(&xas, folio)) + continue; + + if (xa_is_value(folio)) { + swap +=3D 1 << xas_get_order(&xas); + if (swap > max_ptes_swap) { + result =3D SCAN_EXCEED_SWAP_PTE; + count_vm_event(THP_SCAN_EXCEED_SWAP_PTE); + break; + } + continue; + } + + if (!folio_try_get(folio)) { + xas_reset(&xas); + continue; + } + + if (unlikely(folio !=3D xas_reload(&xas))) { + folio_put(folio); + xas_reset(&xas); + continue; + } + + if (is_pmd_order(folio_order(folio))) { + result =3D SCAN_PTE_MAPPED_HUGEPAGE; + /* + * PMD-sized THP implies that we can only try + * retracting the PTE table. + */ + folio_put(folio); + break; + } + + node =3D folio_nid(folio); + if (collapse_scan_abort(node, cc)) { + result =3D SCAN_SCAN_ABORT; + folio_put(folio); + break; + } + cc->node_load[node]++; + + if (!folio_test_lru(folio)) { + result =3D SCAN_PAGE_LRU; + folio_put(folio); + break; + } + + if (folio_expected_ref_count(folio) + 1 !=3D folio_ref_count(folio)) { + result =3D SCAN_PAGE_COUNT; + folio_put(folio); + break; + } + + /* + * We probably should check if the folio is referenced + * here, but nobody would transfer pte_young() to + * folio_test_referenced() for us. And rmap walk here + * is just too costly... + */ + + present +=3D folio_nr_pages(folio); + folio_put(folio); + + if (need_resched()) { + xas_pause(&xas); + cond_resched_rcu(); + } + } + rcu_read_unlock(); + if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) + cc->progress++; + else + cc->progress +=3D HPAGE_PMD_NR; + + if (result =3D=3D SCAN_SUCCEED) { + if (present < HPAGE_PMD_NR - max_ptes_none) { + result =3D SCAN_EXCEED_NONE_PTE; + count_vm_event(THP_SCAN_EXCEED_NONE_PTE); + } else { + result =3D collapse_file(mm, addr, file, start, cc); + } + } + + trace_mm_khugepaged_scan_file(mm, folio, file, present, swap, result); + return result; +} + +/* + * Try to collapse a single PMD starting at a PMD aligned addr, and return + * the results. + */ +enum scan_result collapse_single_pmd(unsigned long addr, + unsigned long end, struct vm_area_struct *vma, + bool *lock_dropped, struct collapse_control *cc) +{ + struct mm_struct *mm =3D vma->vm_mm; + bool triggered_wb =3D false; + enum scan_result result; + struct file *file; + pgoff_t pgoff; + + mmap_assert_locked(mm); + + if (vma_is_anonymous(vma)) { + result =3D collapse_scan_anon_pmd(vma, addr, end, cc); + if (!cc->select_orders) + goto end; + + /* collapse_anon_pmd() takes mmap_lock itself, where it needs it */ + mmap_read_unlock(mm); + *lock_dropped =3D true; + + result =3D collapse_anon_pmd(mm, addr, end, cc); + goto end; + } + + file =3D get_file(vma->vm_file); + pgoff =3D linear_page_index(vma, addr); + + mmap_read_unlock(mm); + *lock_dropped =3D true; +retry: + result =3D collapse_scan_file(mm, addr, file, pgoff, cc); + + /* Dirty pages are worth a writeback and one more try, if asked for */ + if (cc->policy.writeback_dirty && result =3D=3D SCAN_PAGE_DIRTY_OR_WRITEB= ACK && + !triggered_wb && mapping_can_writeback(file->f_mapping)) { + const loff_t lstart =3D (loff_t)pgoff << PAGE_SHIFT; + const loff_t lend =3D lstart + HPAGE_PMD_SIZE - 1; + + filemap_write_and_wait_range(file->f_mapping, lstart, lend); + triggered_wb =3D true; + goto retry; + } + fput(file); + + if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) { + mmap_read_lock(mm); + if (collapse_test_exit_or_disable(mm)) + result =3D SCAN_ANY_PROCESS; + else + result =3D try_collapse_pte_mapped_thp(mm, addr, + cc->policy.install_pmd); + if (result =3D=3D SCAN_PMD_MAPPED) + result =3D SCAN_SUCCEED; + mmap_read_unlock(mm); + } +end: + return result; +} diff --git a/mm/collapse.h b/mm/collapse.h index 11f51c6ea444..dc60806fb81e 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -187,24 +187,13 @@ static inline int collapse_test_exit_or_disable(struc= t mm_struct *mm) mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); } =20 -enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, - unsigned long start, unsigned long end, - struct collapse_control *cc); -enum scan_result collapse_anon_pmd(struct mm_struct *mm, unsigned long sta= rt, - unsigned long end, struct collapse_control *cc); int collapse_control_init(struct collapse_control *cc); void collapse_control_release(struct collapse_control *cc); +enum scan_result collapse_single_pmd(unsigned long addr, unsigned long end, + struct vm_area_struct *vma, bool *lock_dropped, + struct collapse_control *cc); =20 unsigned long collapse_possible_orders(struct vm_area_struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags); -enum scan_result check_pmd_state(pmd_t *pmd); -enum scan_result find_pmd_or_thp_or_none(struct mm_struct *mm, - unsigned long address, pmd_t **pmd); -int collapse_find_target_node(struct collapse_control *cc); -bool collapse_scan_abort(int nid, struct collapse_control *cc); -unsigned int collapse_max_ptes_none(struct collapse_control *cc, - struct vm_area_struct *vma, unsigned int order); -unsigned int collapse_max_ptes_swap(struct collapse_control *cc, - unsigned int order); =20 #endif /* __MM_COLLAPSE_H */ diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 907ed1131460..b7fc93e11d6b 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -543,51 +543,6 @@ static enum scan_result hugepage_vma_revalidate(struct= mm_struct *mm, unsigned l return SCAN_SUCCEED; } =20 -static void count_collapse_event(unsigned int order, enum vm_event_item vm= _event, - enum mthp_stat_item mthp_event) -{ - if (is_pmd_order(order)) - count_vm_event(vm_event); - count_mthp_stat(order, mthp_event); -} - -static void collapse_control_init_scan(struct collapse_control *cc) -{ - memset(cc->node_load, 0, sizeof(cc->node_load)); - nodes_clear(cc->alloc_nmask); - bitmap_zero(cc->eligible_ptes, MAX_PTRS_PER_PTE); -} - -static enum scan_result alloc_charge_folio(struct folio **foliop, struct m= m_struct *mm, - struct collapse_control *cc, unsigned int order) -{ - gfp_t gfp =3D cc->policy.gfp; - int node =3D collapse_find_target_node(cc); - struct folio *folio; - - folio =3D __folio_alloc(gfp, order, node, &cc->alloc_nmask); - if (!folio) { - *foliop =3D NULL; - count_collapse_event(order, THP_COLLAPSE_ALLOC_FAILED, - MTHP_STAT_COLLAPSE_ALLOC_FAILED); - return SCAN_ALLOC_HUGE_PAGE_FAIL; - } - - count_collapse_event(order, THP_COLLAPSE_ALLOC, MTHP_STAT_COLLAPSE_ALLOC); - - if (unlikely(mem_cgroup_charge(folio, mm, gfp))) { - folio_put(folio); - *foliop =3D NULL; - return SCAN_CGROUP_CHARGE_FAIL; - } - - if (is_pmd_order(order)) - count_memcg_folio_events(folio, THP_COLLAPSE_ALLOC, 1); - - *foliop =3D folio; - return SCAN_SUCCEED; -} - static void collect_mm_slot(struct mm_slot *slot) { struct mm_struct *mm =3D slot->mm; @@ -610,1036 +565,6 @@ static void collect_mm_slot(struct mm_slot *slot) } } =20 -/* folio must be locked, and mmap_lock must be held */ -static enum scan_result set_huge_pmd(struct vm_area_struct *vma, unsigned = long addr, - pmd_t *pmdp, struct folio *folio, struct page *page) -{ - struct mm_struct *mm =3D vma->vm_mm; - struct vm_fault vmf =3D { - .vma =3D vma, - .address =3D addr, - .flags =3D 0, - }; - pgd_t *pgdp; - p4d_t *p4dp; - pud_t *pudp; - - mmap_assert_locked(vma->vm_mm); - - if (!pmdp) { - pgdp =3D pgd_offset(mm, addr); - p4dp =3D p4d_alloc(mm, pgdp, addr); - if (!p4dp) - return SCAN_FAIL; - pudp =3D pud_alloc(mm, p4dp, addr); - if (!pudp) - return SCAN_FAIL; - pmdp =3D pmd_alloc(mm, pudp, addr); - if (!pmdp) - return SCAN_FAIL; - } - - vmf.pmd =3D pmdp; - if (do_set_pmd(&vmf, folio, page)) - return SCAN_FAIL; - - folio_get(folio); - return SCAN_SUCCEED; -} - -static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, = unsigned long addr, - bool install_pmd) -{ - enum scan_result result =3D SCAN_FAIL; - int nr_mapped_ptes =3D 0; - unsigned int nr_batch_ptes; - struct mmu_notifier_range range; - bool notified =3D false; - unsigned long haddr =3D addr & HPAGE_PMD_MASK; - unsigned long end =3D haddr + HPAGE_PMD_SIZE; - struct vm_area_struct *vma =3D vma_lookup(mm, haddr); - struct folio *folio; - pte_t *start_pte, *pte; - pmd_t *pmd, pgt_pmd; - spinlock_t *pml =3D NULL, *ptl; - int i; - - mmap_assert_locked(mm); - - /* First check VMA found, in case page tables are being torn down */ - if (!vma || !vma->vm_file || - !range_in_vma(vma, haddr, haddr + HPAGE_PMD_SIZE)) - return SCAN_VMA_CHECK; - - /* Fast check before locking page if already PMD-mapped */ - result =3D find_pmd_or_thp_or_none(mm, haddr, &pmd); - if (result =3D=3D SCAN_PMD_MAPPED) - return result; - - /* - * If we are here, we've succeeded in replacing all the native pages - * in the page cache with a single hugepage. If a mm were to fault-in - * this memory (mapped by a suitably aligned VMA), we'd get the hugepage - * and map it by a PMD, regardless of sysfs THP settings. As such, let's - * analogously elide sysfs THP settings here and force collapse. - */ - if (!thp_vma_allowable_order(vma, vma->vm_flags, TVA_FORCED_COLLAPSE, PMD= _ORDER)) - return SCAN_VMA_CHECK; - - /* - * Keep pmd pgtable while the uffd bit is in use; see comment in - * retract_page_tables(). - */ - if (userfaultfd_protected(vma)) - return SCAN_PTE_UFFD; - - folio =3D filemap_lock_folio(vma->vm_file->f_mapping, - linear_page_index(vma, haddr)); - if (IS_ERR(folio)) - return SCAN_PAGE_NULL; - - if (!is_pmd_order(folio_order(folio))) { - result =3D SCAN_PAGE_COMPOUND; - goto drop_folio; - } - - result =3D find_pmd_or_thp_or_none(mm, haddr, &pmd); - switch (result) { - case SCAN_SUCCEED: - break; - case SCAN_NO_PTE_TABLE: - /* - * All pte entries have been removed and pmd cleared. - * Skip all the pte checks and just update the pmd mapping. - */ - goto maybe_install_pmd; - default: - goto drop_folio; - } - - result =3D SCAN_FAIL; - start_pte =3D pte_offset_map_lock(mm, pmd, haddr, &ptl); - if (!start_pte) /* mmap_lock + page lock should prevent this */ - goto drop_folio; - - /* step 1: check all mapped PTEs are to the right huge page */ - for (i =3D 0, addr =3D haddr, pte =3D start_pte; - i < HPAGE_PMD_NR; i++, addr +=3D PAGE_SIZE, pte++) { - struct page *page; - pte_t ptent =3D ptep_get(pte); - - /* empty pte, skip */ - if (pte_none(ptent)) - continue; - - /* page swapped out, abort */ - if (!pte_present(ptent)) { - result =3D SCAN_PTE_NON_PRESENT; - goto abort; - } - - page =3D vm_normal_page(vma, addr, ptent); - if (WARN_ON_ONCE(page && is_zone_device_page(page))) - page =3D NULL; - /* - * Note that uprobe, debugger, or MAP_PRIVATE may change the - * page table, but the new page will not be a subpage of hpage. - */ - if (folio_page(folio, i) !=3D page) - goto abort; - } - - pte_unmap_unlock(start_pte, ptl); - mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, - haddr, haddr + HPAGE_PMD_SIZE); - mmu_notifier_invalidate_range_start(&range); - notified =3D true; - - /* - * pmd_lock covers a wider range than ptl, and (if split from mm's - * page_table_lock) ptl nests inside pml. The less time we hold pml, - * the better; but userfaultfd's mfill_atomic_pte() on a private VMA - * inserts a valid as-if-COWed PTE without even looking up page cache. - * So page lock of folio does not protect from it, so we must not drop - * ptl before pgt_pmd is removed, so uffd private needs pml taken now. - */ - if (userfaultfd_armed(vma) && !(vma->vm_flags & VM_SHARED)) - pml =3D pmd_lock(mm, pmd); - - start_pte =3D pte_offset_map_rw_nolock(mm, pmd, haddr, &pgt_pmd, &ptl); - if (!start_pte) /* mmap_lock + page lock should prevent this */ - goto abort; - if (!pml) - spin_lock(ptl); - else if (ptl !=3D pml) - spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); - - if (unlikely(!pmd_same(pgt_pmd, pmdp_get_lockless(pmd)))) - goto abort; - - /* step 2: clear page table and adjust rmap */ - for (i =3D 0, addr =3D haddr, pte =3D start_pte; i < HPAGE_PMD_NR; - i +=3D nr_batch_ptes, addr +=3D nr_batch_ptes * PAGE_SIZE, - pte +=3D nr_batch_ptes) { - unsigned int max_nr_batch_ptes =3D (end - addr) >> PAGE_SHIFT; - struct page *page; - pte_t ptent =3D ptep_get(pte); - - nr_batch_ptes =3D 1; - - if (pte_none(ptent)) - continue; - /* - * We dropped ptl after the first scan, to do the mmu_notifier: - * page lock stops more PTEs of the folio being faulted in, but - * does not stop write faults COWing anon copies from existing - * PTEs; and does not stop those being swapped out or migrated. - */ - if (!pte_present(ptent)) { - result =3D SCAN_PTE_NON_PRESENT; - goto abort; - } - page =3D vm_normal_page(vma, addr, ptent); - - if (folio_page(folio, i) !=3D page) - goto abort; - - nr_batch_ptes =3D folio_pte_batch(folio, pte, ptent, max_nr_batch_ptes); - - /* - * Must clear entry, or a racing truncate may re-remove it. - * TLB flush can be left until pmdp_collapse_flush() does it. - * PTE dirty? Shmem page is already dirty; file is read-only. - */ - clear_ptes(mm, addr, pte, nr_batch_ptes); - folio_remove_rmap_ptes(folio, page, nr_batch_ptes, vma); - nr_mapped_ptes +=3D nr_batch_ptes; - } - - if (!pml) - spin_unlock(ptl); - - /* step 3: set proper refcount and mm_counters. */ - if (nr_mapped_ptes) { - folio_ref_sub(folio, nr_mapped_ptes); - add_mm_counter(mm, mm_counter_file(folio), -nr_mapped_ptes); - } - - /* step 4: remove empty page table */ - if (!pml) { - pml =3D pmd_lock(mm, pmd); - if (ptl !=3D pml) { - spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); - if (unlikely(!pmd_same(pgt_pmd, pmdp_get_lockless(pmd)))) { - flush_tlb_mm(mm); - goto unlock; - } - } - } - pgt_pmd =3D pmdp_collapse_flush(vma, haddr, pmd); - pmdp_get_lockless_sync(); - pte_unmap_unlock(start_pte, ptl); - if (ptl !=3D pml) - spin_unlock(pml); - - mmu_notifier_invalidate_range_end(&range); - - mm_dec_nr_ptes(mm); - page_table_check_pte_clear_range(mm, haddr, pgt_pmd); - pte_free_defer(mm, pmd_pgtable(pgt_pmd)); - -maybe_install_pmd: - /* step 5: install pmd entry */ - result =3D install_pmd - ? set_huge_pmd(vma, haddr, pmd, folio, &folio->page) - : SCAN_SUCCEED; - goto drop_folio; -abort: - if (nr_mapped_ptes) { - flush_tlb_mm(mm); - folio_ref_sub(folio, nr_mapped_ptes); - add_mm_counter(mm, mm_counter_file(folio), -nr_mapped_ptes); - } -unlock: - if (start_pte) - pte_unmap_unlock(start_pte, ptl); - if (pml && pml !=3D ptl) - spin_unlock(pml); - if (notified) - mmu_notifier_invalidate_range_end(&range); -drop_folio: - folio_unlock(folio); - folio_put(folio); - return result; -} - -/** - * collapse_pte_mapped_thp - Try to collapse a pte-mapped THP for mm at - * address haddr. - * - * @mm: process address space where collapse happens - * @addr: THP collapse address - * @install_pmd: If a huge PMD should be installed - * - * This function checks whether all the PTEs in the PMD are pointing to the - * right THP. If so, retract the page table so the THP can refault in with - * as pmd-mapped. Possibly install a huge PMD mapping the THP. - */ -void collapse_pte_mapped_thp(struct mm_struct *mm, unsigned long addr, - bool install_pmd) -{ - try_collapse_pte_mapped_thp(mm, addr, install_pmd); -} - -/* Can we retract page tables for this file-backed VMA? */ -static bool file_backed_vma_is_retractable(struct vm_area_struct *vma) -{ - /* - * Check vma->anon_vma to exclude MAP_PRIVATE mappings that - * got written to. These VMAs are likely not worth removing - * page tables from, as PMD-mapping is likely to be split later. - */ - if (READ_ONCE(vma->anon_vma)) - return false; - - /* - * When a vma is registered with uffd-wp or RWP, we cannot recycle - * the page table because there may be pte markers installed. - * VM_UFFD_RWP ranges similarly rely on per-PTE uffd state - * and cannot be recycled to a shared PMD. Other vmas can still - * have the same file mapped hugely, but skip this one: it will - * always be mapped in small page size for these registrations. - */ - if (userfaultfd_protected(vma)) - return false; - - /* - * If the VMA contains guard regions then we can't collapse it. - * - * This is set atomically on guard marker installation under mmap/VMA - * read lock, and here we may not hold any VMA or mmap lock at all. - * - * This is therefore serialised on the PTE page table lock, which is - * obtained on guard region installation after the flag is set, so this - * check being performed under this lock excludes races. - */ - if (vma_test_atomic_flag(vma, VMA_MAYBE_GUARD_BIT)) - return false; - - return true; -} - -static void retract_page_tables(struct address_space *mapping, pgoff_t pgo= ff) -{ - struct vm_area_struct *vma; - - i_mmap_lock_read(mapping); - mapping_rmap_tree_foreach(vma, mapping, pgoff, pgoff) { - struct mmu_notifier_range range; - struct mm_struct *mm; - unsigned long addr; - pmd_t *pmd, pgt_pmd; - spinlock_t *pml; - spinlock_t *ptl; - bool success =3D false; - - addr =3D vma->vm_start + - ((pgoff - vma_start_pgoff(vma)) << PAGE_SHIFT); - if (addr & ~HPAGE_PMD_MASK || - vma->vm_end < addr + HPAGE_PMD_SIZE) - continue; - - mm =3D vma->vm_mm; - if (find_pmd_or_thp_or_none(mm, addr, &pmd) !=3D SCAN_SUCCEED) - continue; - - if (collapse_test_exit(mm)) - continue; - - if (!file_backed_vma_is_retractable(vma)) - continue; - - /* PTEs were notified when unmapped; but now for the PMD? */ - mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, - addr, addr + HPAGE_PMD_SIZE); - mmu_notifier_invalidate_range_start(&range); - - pml =3D pmd_lock(mm, pmd); - /* - * The lock of new_folio is still held, we will be blocked in - * the page fault path, which prevents the pte entries from - * being set again. So even though the old empty PTE page may be - * concurrently freed and a new PTE page is filled into the pmd - * entry, it is still empty and can be removed. - * - * So here we only need to recheck if the state of pmd entry - * still meets our requirements, rather than checking pmd_same() - * like elsewhere. - */ - if (check_pmd_state(pmd) !=3D SCAN_SUCCEED) - goto drop_pml; - ptl =3D pte_lockptr(mm, pmd); - if (ptl !=3D pml) - spin_lock_nested(ptl, SINGLE_DEPTH_NESTING); - - /* - * Huge page lock is still held, so normally the page table must - * remain empty; and we have already skipped anon_vma and - * userfaultfd_wp() vmas. But since the mmap_lock is not held, - * it is still possible for a racing userfaultfd_ioctl() or - * madvise() to have inserted ptes or markers. Now that we hold - * ptlock, repeating the retractable checks protects us from - * races against the prior checks. - */ - if (likely(file_backed_vma_is_retractable(vma))) { - pgt_pmd =3D pmdp_collapse_flush(vma, addr, pmd); - pmdp_get_lockless_sync(); - success =3D true; - } - - if (ptl !=3D pml) - spin_unlock(ptl); -drop_pml: - spin_unlock(pml); - - mmu_notifier_invalidate_range_end(&range); - - if (success) { - mm_dec_nr_ptes(mm); - page_table_check_pte_clear_range(mm, addr, pgt_pmd); - pte_free_defer(mm, pmd_pgtable(pgt_pmd)); - } - } - i_mmap_unlock_read(mapping); -} - -/** - * collapse_file - collapse filemap/tmpfs/shmem pages into huge one. - * - * @mm: process address space where collapse happens - * @addr: virtual collapse start address - * @file: file that collapse on - * @start: collapse start address - * @cc: collapse context and scratchpad - * - * Basic scheme is simple, details are more complex: - * - allocate and lock a new huge page; - * - scan page cache, locking old pages - * + swap/gup in pages if necessary; - * - copy data to new page - * - handle shmem holes - * + re-validate that holes weren't filled by someone else - * + check for userfaultfd - * - finalize updates to the page cache; - * - if replacing succeeds: - * + unlock huge page; - * + free old pages; - * - if replacing failed; - * + unlock old pages - * + unlock and free huge page; - */ -static enum scan_result collapse_file(struct mm_struct *mm, unsigned long = addr, - struct file *file, pgoff_t start, struct collapse_control *cc) -{ - struct address_space *mapping =3D file->f_mapping; - struct page *dst; - struct folio *folio, *tmp, *new_folio; - pgoff_t index =3D 0, end =3D start + HPAGE_PMD_NR; - LIST_HEAD(pagelist); - XA_STATE_ORDER(xas, &mapping->i_pages, start, HPAGE_PMD_ORDER); - enum scan_result result =3D SCAN_SUCCEED; - int nr_none =3D 0; - bool is_shmem =3D shmem_file(file); - - /* - * MADV_COLLAPSE ignores shmem huge config, so do not check shmem - * - * TODO: once shmem always calls mapping_set_large_folios() on its - * mapping, the shmem check can be removed. - */ - VM_WARN_ON_ONCE(!is_shmem && !mapping_pmd_folio_support(mapping)); - VM_WARN_ON_ONCE(start & (HPAGE_PMD_NR - 1)); - - result =3D alloc_charge_folio(&new_folio, mm, cc, HPAGE_PMD_ORDER); - if (result !=3D SCAN_SUCCEED) - goto out; - - mapping_set_update(&xas, mapping); - - __folio_set_locked(new_folio); - if (is_shmem) - __folio_set_swapbacked(new_folio); - new_folio->index =3D start; - new_folio->mapping =3D mapping; - - /* - * Ensure we have slots for all the pages in the range. This is - * almost certainly a no-op because most of the pages must be present - */ - do { - xas_lock_irq(&xas); - xas_create_range(&xas); - if (!xas_error(&xas)) - break; - xas_unlock_irq(&xas); - if (!xas_nomem(&xas, GFP_KERNEL)) { - result =3D SCAN_FAIL; - goto rollback; - } - } while (1); - - for (index =3D start; index < end;) { - xas_set(&xas, index); - folio =3D xas_load(&xas); - - VM_BUG_ON(index !=3D xas.xa_index); - if (is_shmem) { - if (!folio) { - /* - * Stop if extent has been truncated or - * hole-punched, and is now completely - * empty. - */ - if (index =3D=3D start) { - if (!xas_next_entry(&xas, end - 1)) { - result =3D SCAN_TRUNCATED; - goto xa_locked; - } - } - nr_none++; - index++; - continue; - } - - if (xa_is_value(folio) || !folio_test_uptodate(folio)) { - xas_unlock_irq(&xas); - /* swap in or instantiate fallocated page */ - if (shmem_get_folio(mapping->host, index, 0, - &folio, SGP_NOALLOC)) { - result =3D SCAN_FAIL; - goto xa_unlocked; - } - /* drain lru cache to help folio_isolate_lru() */ - lru_add_drain(); - } else if (folio_trylock(folio)) { - folio_get(folio); - xas_unlock_irq(&xas); - } else { - result =3D SCAN_PAGE_LOCK; - goto xa_locked; - } - } else { /* !is_shmem */ - if (!folio || xa_is_value(folio)) { - xas_unlock_irq(&xas); - page_cache_sync_readahead(mapping, &file->f_ra, - file, index, - end - index); - /* drain lru cache to help folio_isolate_lru() */ - lru_add_drain(); - folio =3D filemap_lock_folio(mapping, index); - if (IS_ERR(folio)) { - result =3D SCAN_FAIL; - goto xa_unlocked; - } - } else if (folio_test_dirty(folio)) { - /* - * This page is dirty because it hasn't - * been flushed since first write. - * - * Trigger async flush for read-only files and - * hope the writeback is done when khugepaged - * revisits this page. Writable files can have - * their folios dirty at any time; blindly - * flushing them would cause undesirable - * system-wide writeback. - * - * This is a one-off situation. We are not - * forcing writeback in loop. - */ - xas_unlock_irq(&xas); - if (!inode_is_open_for_write(mapping->host)) - filemap_flush(mapping); - result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; - goto xa_unlocked; - } else if (folio_test_writeback(folio)) { - xas_unlock_irq(&xas); - result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; - goto xa_unlocked; - } else if (folio_trylock(folio)) { - folio_get(folio); - xas_unlock_irq(&xas); - } else { - result =3D SCAN_PAGE_LOCK; - goto xa_locked; - } - } - - /* - * The folio must be locked, so we can drop the i_pages lock - * without racing with truncate. - */ - VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); - - /* make sure the folio is up to date */ - if (unlikely(!folio_test_uptodate(folio))) { - result =3D SCAN_FAIL; - goto out_unlock; - } - - /* - * If file was truncated then extended, or hole-punched, before - * we locked the first folio, then a THP might be there already. - * This will be discovered on the first iteration. - */ - if (is_pmd_order(folio_order(folio))) { - result =3D SCAN_PTE_MAPPED_HUGEPAGE; - goto out_unlock; - } - - if (folio_mapping(folio) !=3D mapping) { - result =3D SCAN_TRUNCATED; - goto out_unlock; - } - - if (!is_shmem && (folio_test_dirty(folio) || - folio_test_writeback(folio))) { - /* - * khugepaged only works on clean file-backed folios, - * so this folio is dirty because it hasn't been flushed - * since first write. - */ - result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; - goto out_unlock; - } - - if (!folio_isolate_lru(folio)) { - result =3D SCAN_DEL_PAGE_LRU; - goto out_unlock; - } - - if (!filemap_release_folio(folio, GFP_KERNEL)) { - result =3D SCAN_PAGE_HAS_PRIVATE; - folio_putback_lru(folio); - goto out_unlock; - } - - if (folio_mapped(folio)) - try_to_unmap(folio, - TTU_IGNORE_MLOCK | TTU_BATCH_FLUSH); - - xas_lock_irq(&xas); - - VM_BUG_ON_FOLIO(folio !=3D xa_load(xas.xa, index), folio); - - /* - * We control 2 + nr_pages references to the folio: - * - we hold a pin on it; - * - nr_pages reference from page cache; - * - one from lru_isolate_folio; - * If those are the only references, then any new usage - * of the folio will have to fetch it from the page - * cache. That requires locking the folio to handle - * truncate, so any new usage will be blocked until we - * unlock folio after collapse/during rollback. - */ - if (folio_ref_count(folio) !=3D 2 + folio_nr_pages(folio)) { - result =3D SCAN_PAGE_COUNT; - xas_unlock_irq(&xas); - folio_putback_lru(folio); - goto out_unlock; - } - - /* - * At this point, the folio is locked and unmapped. If the PTE - * was dirty, try_to_unmap() has transferred the dirty bit to - * the folio and we must not collapse it into a clean - * file-backed folio. - * - * If the folio is clean here, no one can write it until we - * drop the folio lock. A write through a stale TLB entry came - * from a clean PTE and must fault because the PTE has been - * cleared; the fault path has to take the folio lock before - * installing a writable mapping. Buffered write paths also - * have to take the folio lock before modifying file contents - * without a mapping, typically via write_begin_get_folio(). - */ - if (!is_shmem && folio_test_dirty(folio)) { - result =3D SCAN_PAGE_DIRTY_OR_WRITEBACK; - xas_unlock_irq(&xas); - folio_putback_lru(folio); - goto out_unlock; - } - - /* - * Accumulate the folios that are being collapsed. - */ - list_add_tail(&folio->lru, &pagelist); - index +=3D folio_nr_pages(folio); - continue; -out_unlock: - folio_unlock(folio); - folio_put(folio); - goto xa_unlocked; - } - -xa_locked: - xas_unlock_irq(&xas); -xa_unlocked: - - /* - * If collapse is successful, flush must be done now before copying. - * If collapse is unsuccessful, does flush actually need to be done? - * Do it anyway, to clear the state. - */ - try_to_unmap_flush(); - - if (result =3D=3D SCAN_SUCCEED && nr_none && - !shmem_charge(mapping->host, nr_none)) - result =3D SCAN_FAIL; - if (result !=3D SCAN_SUCCEED) { - nr_none =3D 0; - goto rollback; - } - - /* - * The old folios are locked, so they won't change anymore. - */ - index =3D start; - dst =3D folio_page(new_folio, 0); - list_for_each_entry(folio, &pagelist, lru) { - int i, nr_pages =3D folio_nr_pages(folio); - - while (index < folio->index) { - clear_highpage(dst); - index++; - dst++; - } - - for (i =3D 0; i < nr_pages; i++) { - if (copy_mc_highpage(dst, folio_page(folio, i)) > 0) { - result =3D SCAN_COPY_MC; - goto rollback; - } - index++; - dst++; - } - } - while (index < end) { - clear_highpage(dst); - index++; - dst++; - } - - if (nr_none) { - struct vm_area_struct *vma; - int nr_none_check =3D 0; - - i_mmap_lock_read(mapping); - xas_lock_irq(&xas); - - xas_set(&xas, start); - for (index =3D start; index < end; index++) { - if (!xas_next(&xas)) { - xas_store(&xas, XA_RETRY_ENTRY); - if (xas_error(&xas)) { - result =3D SCAN_STORE_FAILED; - goto immap_locked; - } - nr_none_check++; - } - } - - if (nr_none !=3D nr_none_check) { - result =3D SCAN_PAGE_FILLED; - goto immap_locked; - } - - /* - * If userspace observed a missing page in a VMA with - * a MODE_MISSING userfaultfd, then it might expect a - * UFFD_EVENT_PAGEFAULT for that page. If so, we need to - * roll back to avoid suppressing such an event. Since - * wp/minor userfaultfds don't give userspace any - * guarantees that the kernel doesn't fill a missing - * page with a zero page, so they don't matter here. - * - * Any userfaultfds registered after this point will - * not be able to observe any missing pages due to the - * previously inserted retry entries. - */ - mapping_rmap_tree_foreach(vma, mapping, start, end) { - if (userfaultfd_missing(vma)) { - result =3D SCAN_EXCEED_NONE_PTE; - goto immap_locked; - } - } - -immap_locked: - i_mmap_unlock_read(mapping); - if (result !=3D SCAN_SUCCEED) { - xas_set(&xas, start); - for (index =3D start; index < end; index++) { - if (xas_next(&xas) =3D=3D XA_RETRY_ENTRY) - xas_store(&xas, NULL); - } - - xas_unlock_irq(&xas); - goto rollback; - } - } else { - xas_lock_irq(&xas); - } - - if (is_shmem) { - lruvec_stat_mod_folio(new_folio, NR_SHMEM, HPAGE_PMD_NR); - lruvec_stat_mod_folio(new_folio, NR_SHMEM_THPS, HPAGE_PMD_NR); - } else { - lruvec_stat_mod_folio(new_folio, NR_FILE_THPS, HPAGE_PMD_NR); - } - lruvec_stat_mod_folio(new_folio, NR_FILE_PAGES, HPAGE_PMD_NR); - - /* - * Mark new_folio as uptodate before inserting it into the - * page cache so that it isn't mistaken for an fallocated but - * unwritten page. - */ - folio_mark_uptodate(new_folio); - folio_ref_add(new_folio, HPAGE_PMD_NR - 1); - - if (is_shmem) - folio_mark_dirty(new_folio); - folio_add_lru(new_folio); - - /* Join all the small entries into a single multi-index entry. */ - xas_set_order(&xas, start, HPAGE_PMD_ORDER); - xas_store(&xas, new_folio); - WARN_ON_ONCE(xas_error(&xas)); - xas_unlock_irq(&xas); - - /* - * Remove pte page tables, so we can re-fault the page as huge. A caller - * that wants the PMD mapped now is told to go and do that. - */ - retract_page_tables(mapping, start); - if (cc->policy.install_pmd) - result =3D SCAN_PTE_MAPPED_HUGEPAGE; - folio_unlock(new_folio); - - /* - * The collapse has succeeded, so free the old folios. - */ - list_for_each_entry_safe(folio, tmp, &pagelist, lru) { - list_del(&folio->lru); - lruvec_stat_mod_folio(folio, NR_FILE_PAGES, - -folio_nr_pages(folio)); - if (is_shmem) - lruvec_stat_mod_folio(folio, NR_SHMEM, - -folio_nr_pages(folio)); - folio->mapping =3D NULL; - folio_clear_active(folio); - folio_clear_unevictable(folio); - folio_unlock(folio); - folio_put_refs(folio, 2 + folio_nr_pages(folio)); - } - - goto out; - -rollback: - /* Something went wrong: roll back page cache changes */ - if (nr_none) { - xas_lock_irq(&xas); - mapping->nrpages -=3D nr_none; - xas_unlock_irq(&xas); - shmem_uncharge(mapping->host, nr_none); - } - - list_for_each_entry_safe(folio, tmp, &pagelist, lru) { - list_del(&folio->lru); - folio_unlock(folio); - folio_putback_lru(folio); - folio_put(folio); - } - - new_folio->mapping =3D NULL; - - folio_unlock(new_folio); - folio_put(new_folio); -out: - VM_BUG_ON(!list_empty(&pagelist)); - trace_mm_khugepaged_collapse_file(mm, new_folio, index, addr, is_shmem, f= ile, HPAGE_PMD_NR, result); - return result; -} - -static enum scan_result collapse_scan_file(struct mm_struct *mm, - unsigned long addr, struct file *file, pgoff_t start, - struct collapse_control *cc) -{ - const unsigned int max_ptes_none =3D collapse_max_ptes_none(cc, NULL, HPA= GE_PMD_ORDER); - const unsigned int max_ptes_swap =3D collapse_max_ptes_swap(cc, HPAGE_PMD= _ORDER); - struct folio *folio =3D NULL; - struct address_space *mapping =3D file->f_mapping; - XA_STATE(xas, &mapping->i_pages, start); - int present, swap; - int node =3D NUMA_NO_NODE; - enum scan_result result =3D SCAN_SUCCEED; - - present =3D 0; - swap =3D 0; - collapse_control_init_scan(cc); - rcu_read_lock(); - xas_for_each(&xas, folio, start + HPAGE_PMD_NR - 1) { - if (xas_retry(&xas, folio)) - continue; - - if (xa_is_value(folio)) { - swap +=3D 1 << xas_get_order(&xas); - if (swap > max_ptes_swap) { - result =3D SCAN_EXCEED_SWAP_PTE; - count_vm_event(THP_SCAN_EXCEED_SWAP_PTE); - break; - } - continue; - } - - if (!folio_try_get(folio)) { - xas_reset(&xas); - continue; - } - - if (unlikely(folio !=3D xas_reload(&xas))) { - folio_put(folio); - xas_reset(&xas); - continue; - } - - if (is_pmd_order(folio_order(folio))) { - result =3D SCAN_PTE_MAPPED_HUGEPAGE; - /* - * PMD-sized THP implies that we can only try - * retracting the PTE table. - */ - folio_put(folio); - break; - } - - node =3D folio_nid(folio); - if (collapse_scan_abort(node, cc)) { - result =3D SCAN_SCAN_ABORT; - folio_put(folio); - break; - } - cc->node_load[node]++; - - if (!folio_test_lru(folio)) { - result =3D SCAN_PAGE_LRU; - folio_put(folio); - break; - } - - if (folio_expected_ref_count(folio) + 1 !=3D folio_ref_count(folio)) { - result =3D SCAN_PAGE_COUNT; - folio_put(folio); - break; - } - - /* - * We probably should check if the folio is referenced - * here, but nobody would transfer pte_young() to - * folio_test_referenced() for us. And rmap walk here - * is just too costly... - */ - - present +=3D folio_nr_pages(folio); - folio_put(folio); - - if (need_resched()) { - xas_pause(&xas); - cond_resched_rcu(); - } - } - rcu_read_unlock(); - if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) - cc->progress++; - else - cc->progress +=3D HPAGE_PMD_NR; - - if (result =3D=3D SCAN_SUCCEED) { - if (present < HPAGE_PMD_NR - max_ptes_none) { - result =3D SCAN_EXCEED_NONE_PTE; - count_vm_event(THP_SCAN_EXCEED_NONE_PTE); - } else { - result =3D collapse_file(mm, addr, file, start, cc); - } - } - - trace_mm_khugepaged_scan_file(mm, folio, file, present, swap, result); - return result; -} - -/* - * Try to collapse a single PMD starting at a PMD aligned addr, and return - * the results. - */ -static enum scan_result collapse_single_pmd(unsigned long addr, - unsigned long end, struct vm_area_struct *vma, - bool *lock_dropped, struct collapse_control *cc) -{ - struct mm_struct *mm =3D vma->vm_mm; - bool triggered_wb =3D false; - enum scan_result result; - struct file *file; - pgoff_t pgoff; - - mmap_assert_locked(mm); - - if (vma_is_anonymous(vma)) { - result =3D collapse_scan_anon_pmd(vma, addr, end, cc); - if (!cc->select_orders) - goto end; - - /* collapse_anon_pmd() takes mmap_lock itself, where it needs it */ - mmap_read_unlock(mm); - *lock_dropped =3D true; - - result =3D collapse_anon_pmd(mm, addr, end, cc); - goto end; - } - - file =3D get_file(vma->vm_file); - pgoff =3D linear_page_index(vma, addr); - - mmap_read_unlock(mm); - *lock_dropped =3D true; -retry: - result =3D collapse_scan_file(mm, addr, file, pgoff, cc); - - /* Dirty pages are worth a writeback and one more try, if asked for */ - if (cc->policy.writeback_dirty && result =3D=3D SCAN_PAGE_DIRTY_OR_WRITEB= ACK && - !triggered_wb && mapping_can_writeback(file->f_mapping)) { - const loff_t lstart =3D (loff_t)pgoff << PAGE_SHIFT; - const loff_t lend =3D lstart + HPAGE_PMD_SIZE - 1; - - filemap_write_and_wait_range(file->f_mapping, lstart, lend); - triggered_wb =3D true; - goto retry; - } - fput(file); - - if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) { - mmap_read_lock(mm); - if (collapse_test_exit_or_disable(mm)) - result =3D SCAN_ANY_PROCESS; - else - result =3D try_collapse_pte_mapped_thp(mm, addr, - cc->policy.install_pmd); - if (result =3D=3D SCAN_PMD_MAPPED) - result =3D SCAN_SUCCEED; - mmap_read_unlock(mm); - } -end: - return result; -} - static void collapse_scan_mm_slot(unsigned int progress_max, enum scan_result *result, struct collapse_control *cc) __releases(&khugepaged_mm_lock) --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7B2693EC6B0; Sun, 16 Aug 2026 22:47:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920444; cv=none; b=T1dx5he+FK2oDYH92LG1XD+gQyiJPw0RHn4jcO+9twd90EMCZXiufxRN7q2oLo0BY21GMmbtaERBisuc38S46xcwI2VZxmu6ZXnwP7CBPubnob07i53v+EJk39sgyfOL3iD59Q4LgV4vdJN0w2wzGvtia18ewqrx3Wyo89qeIFE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920444; c=relaxed/simple; bh=BRitTdWupruJv/LRxCuJbPg3HjwuvpGClUq2KT0tapg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=V5ThTEUVP/uWhjc9BaNuZ3A7zs2BZ/iw12929XJ/CehPJXKwGl19ss5lpRVhdhgkX3M8AJW13TWm28ikl13iPDJxPykd/HCikrAN5Z6MK7CDcZ4o/zp/NScGDYvKDdrap4ze6OpMSBEVPHsTcZvhWW1syYAhaHiTOAecAoFbePI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=UFFDVb02; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=mGaveY87; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="UFFDVb02"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="mGaveY87" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 9D6A8EC0235; Sun, 16 Aug 2026 18:47:20 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:20 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920440; x= 1787006840; bh=tVba+XTZdHa/dscPQvmcm/U9IhQHpFwuT7yBDkbY0DQ=; b=U FFDVb02XU3S3G1WNkrD19eUsTXs3NXE/7zdrXGCajZ2AlfDj5SMBX+byeh4mxPWJ TKRXOSYfCwXtLLQem3j3pmqKgIjr+sg2eEQr/pJsbRvfzW5rVXTz0AhYQjuN7juN CKAC248x7rsmyYL3l7DDphqHJX7cDiKQHgBP9PHEqgUkX/1pnZTxGVuKfUZoD/7H Thbcgy3SSYVksF1wjmz6frXWWr5x7H8joPfNbSFZM9fO24Unwv1zrETrcWVgy8Cz GaFEEz63U2z5aFBUghZZ51st9lwBI3E8jatPZZjqZ+gKVTecogzsU4CRchpbLUsS mFVX998aI2PSTLzi0BHhA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920440; x=1787006840; bh=t Vba+XTZdHa/dscPQvmcm/U9IhQHpFwuT7yBDkbY0DQ=; b=mGaveY87oB0cm5+aj FjboME9vnpTTfu8BXkJK7CVtLM0qhHtzHBiv2XBjmsu5r9YytlGNkwoekJ8y2ZEH kWudNpiWzyDHp74yr4BMdhhcT2eW3QJ8Y5cRBDjyBIqXv20SHQDghzDxwfKnInjN MeV0pyPNBYEYK8DeNcw1bcKksWHhn8MtKJ/pANsFd/JLBwgIVkrVcbi768+GM5Fc aoRpkWUojerhDgpbIDSPwtLrhTwzuUXTJHA0YtkUOZvoLnbis4vABLunXdjUQAKR m0R6T0NsmqzVcgbtllIFM4cqoBaFLkjnsNJ+H53RYmGXD4YQqXxw63xBy6VUAKcV txqXg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWPmjqOfRj5F992yj4B4BZj1raD29b+9/uQ9rFyvaV+kr+0iJRDW/7N8sznN5hp3 Jd0ytp1vu59bDzCFT/O8S/efKgYbCfD2rLp5ESvFXHBu+C+T7VPMHXocEWQmqRcTwN39e9 beda2Ri+DmyXoylgECNlOmesAYSvAqjyQbPp39VciQhjDZLAdlalALq+T39tfvcBsAftOx 0pZrMKicu8IddtrFqvBcFdG/5WkfJWvoJzqNqmEsI17lT+SJI8PEV94Md+CC1CjuERkMm2 pW/0xtbQUhoKznxP8UdY68yvPC1FKG1sdD5CGoRhqiz0Y+ik+TmkR1x+Rv8Bp8o0DAMUDe xJSVVh9q+2wDLRmgw31VvVHCgFkXWkjkLBfJJme8xEC58rxNNC+IyRv4f9KacmZj6IY27S 0qJosTUoPozktMyQF2/JdGirbhA3Foe35GQVTQBcL2hQC4RgWN+wR3pptB2L1mOR4QgPZV 9Y9aDr4bo2/3/ib4cGB+zVs0l3x4xH9StKATJd518qqststirjaxVvywUp62WSv62liYbN 10KCxs6d5IK3xDagGF7hK8wEZkYw1vdi0HigrV2bOrp6O70pSP6CQz5YssphOA8iCncGAn 8wFh8GOHT2iQyCN4Y+jCXlqjoDjtf8JNd1dIh8jfIInw86kzHOYd5Yy81sqg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:19 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 33/57] mm/collapse: split collapse into a scan and a run Date: Sun, 16 Aug 2026 23:45:45 +0100 Message-ID: <20260816224609.308019-34-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_single_pmd() did both halves of a collapse behind one call, and dropped mmap_lock somewhere in the middle. Which of its paths dropped it was not something a caller could see, so it was handed back a bool and had to keep track. Both callers did that badly: khugepaged broke out of its VMA walk after every table, collapsed or not, and MADV_COLLAPSE carried lock state through its loop and re-took the lock only to hand it back. Split it in two, with the lock as the boundary: - collapse_scan_pmd() judges one table and returns with mmap_lock still held. It only reads, and almost every table it is offered has nothing in it, so a caller walks a whole VMA under the one lock it took to get there. - collapse_run_pmd() is called without the lock, which the caller gives up first, and takes it again per round. What it does is slow enough that a writer would otherwise wait behind all of it. Whether there is anything to run is the scan's return value, so no caller has to ask about the lock. The file side is what makes this more than a rename. A file collapse works on the page cache and never sees a VMA, but the file and the offset have to come from one: the scan takes them while it still has the VMA, in cc->scan_file and cc->scan_pgoff, and the run is what gives the reference back. A scan that found file work therefore has to be run, and collapse_control_release() warns and drops the reference rather than rest on callers getting that right. The orders a VMA allows now come in as an argument. khugepaged's walk already computes that mask once per VMA, where the old shape recomputed it twice for every table. MADV_COLLAPSE keeps its own copy only while it holds the VMA, and clears it beside vma =3D NULL: after the lock is given up, the next lookup may return a different VMA. khugepaged now stays in its VMA walk across every table it refuses, and gives the lock up only for a table it is going to collapse. MADV_COLLAPSE does the same. It gives up the lock the VMA walk left held, so lru_add_drain_all() does not wait on every CPU under it, then takes it once and walks its whole range. It drops that lock only around a collapse, and looks the VMA up again only then, since nothing else can have moved it. Holding on across refusals is not new for anonymous memory -- the old entry already returned with the lock held when a table had nothing in it -- but it was never true of file ranges, and never something a caller could rely on. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 167 +++++++++++++++++++++++++++++++----------- mm/collapse.h | 48 ++++++++++++- mm/khugepaged.c | 188 +++++++++++++++++++++++------------------------- mm/mremap.c | 2 +- 4 files changed, 262 insertions(+), 143 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 21bfbc038044..f0d80204c2bd 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -401,6 +401,10 @@ static unsigned int candidate_offset(const struct coll= apse_candidate *cand, =20 void collapse_control_release(struct collapse_control *cc) { + /* Only a scan that was never run leaves this behind */ + if (WARN_ON_ONCE(cc->scan_file)) + fput(cc->scan_file); + kfree(cc->candidates); kfree(cc->saved_ptes); kfree(cc->retries); @@ -413,6 +417,10 @@ int collapse_control_init(struct collapse_control *cc) { cc->nr_candidates =3D 0; cc->nr_retries =3D 0; + cc->select_orders =3D 0; + cc->scan_refusal =3D SCAN_FAIL; + cc->scan_file =3D NULL; + cc->scan_pgoff =3D 0; cc->candidates =3D kmalloc_objs(*cc->candidates, COLLAPSE_MAX_CANDIDATES); cc->saved_ptes =3D kmalloc_objs(*cc->saved_ptes, COLLAPSE_SAVED_PTES); cc->retries =3D kmalloc_objs(*cc->retries, COLLAPSE_RETRY_STORE_SIZE); @@ -2116,7 +2124,8 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) */ static enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, unsigned long start, unsigned long end, - struct collapse_control *cc) + struct collapse_control *cc, + unsigned long vma_orders) { const unsigned long pmd_addr =3D start & HPAGE_PMD_MASK; struct mm_struct *mm =3D vma->vm_mm; @@ -2135,12 +2144,7 @@ static enum scan_result collapse_scan_anon_pmd(struc= t vm_area_struct *vma, /* Cleared only once a table has turned out to be there */ collapse_anon_scan_init(cc); =20 - cc->select_orders =3D collapse_possible_orders(vma, vma->vm_flags, - cc->policy.tva_type); - if (!cc->select_orders) { - cc->scan_refusal =3D SCAN_VMA_CHECK; - return cc->scan_refusal; - } + cc->select_orders =3D vma_orders; =20 /* The scan narrows select_orders to whatever is left worth trying */ cc->scan_refusal =3D collapse_scan_table(vma, pmd, start, end, cc); @@ -2596,7 +2600,7 @@ static void count_collapse_event(unsigned int order, = enum vm_event_item vm_event count_mthp_stat(order, mthp_event); } =20 -static void collapse_control_init_scan(struct collapse_control *cc) +static void collapse_file_scan_init(struct collapse_control *cc) { memset(cc->node_load, 0, sizeof(cc->node_load)); nodes_clear(cc->alloc_nmask); @@ -3493,7 +3497,7 @@ static enum scan_result collapse_file(struct mm_struc= t *mm, unsigned long addr, return result; } =20 -static enum scan_result collapse_scan_file(struct mm_struct *mm, +static enum scan_result collapse_pagecache_pmd(struct mm_struct *mm, unsigned long addr, struct file *file, pgoff_t start, struct collapse_control *cc) { @@ -3508,7 +3512,7 @@ static enum scan_result collapse_scan_file(struct mm_= struct *mm, =20 present =3D 0; swap =3D 0; - collapse_control_init_scan(cc); + collapse_file_scan_init(cc); rcu_read_lock(); xas_for_each(&xas, folio, start + HPAGE_PMD_NR - 1) { if (xas_retry(&xas, folio)) @@ -3600,46 +3604,65 @@ static enum scan_result collapse_scan_file(struct m= m_struct *mm, } =20 /* - * Try to collapse a single PMD starting at a PMD aligned addr, and return - * the results. + * Judge one table's worth of a file VMA. All it needs of the VMA is the = file and + * the offset, which it takes while it still has both; the collapse works = on the + * page cache and never sees a VMA. */ -enum scan_result collapse_single_pmd(unsigned long addr, - unsigned long end, struct vm_area_struct *vma, - bool *lock_dropped, struct collapse_control *cc) +static enum scan_result collapse_scan_file_pmd(struct vm_area_struct *vma, + unsigned long addr, struct collapse_control *cc) { - struct mm_struct *mm =3D vma->vm_mm; + enum scan_result result; + pmd_t *pmd; + + /* + * A file collapse only ever builds a PMD, so the whole table has to be + * the VMA's -- a PMD shared with another VMA would need all of them + * locked. Not the question collapse_possible_orders() answered, which is + * whether the VMA may use the order at all: this is whether the table at + * @addr is wholly inside it. While a file VMA collapses at PMD order + * alone its callers hand over whole tables and this cannot fire, but the + * anonymous side already hands over parts of one. + */ + if (!thp_vma_suitable_order(vma, addr, HPAGE_PMD_ORDER)) + return SCAN_ADDRESS_RANGE; + + /* + * A PMD that is huge already has nothing left to collapse, and skipping + * it here is what keeps mmap_lock out of a collapse that would find + * nothing. Everything else is worth the page cache scan, pmd_none() + * included: a file range can be collapsed out of the cache without being + * mapped first, which is why this is not the test the anonymous side + * makes. + */ + result =3D find_pmd_or_thp_or_none(vma->vm_mm, addr & HPAGE_PMD_MASK, &pm= d); + if (result =3D=3D SCAN_PMD_MAPPED) + return result; + + cc->scan_file =3D get_file(vma->vm_file); + cc->scan_pgoff =3D linear_page_index(vma, addr); + + return SCAN_SUCCEED; +} + +/* + * Build a PMD over what the page cache holds, and map it over the range i= f a huge + * folio is already there but mapped by PTEs. Runs with no mmap_lock, whi= ch the + * caller gave up, and takes it again only for that last step. + */ +static enum scan_result collapse_file_pmd(struct mm_struct *mm, + unsigned long addr, struct collapse_control *cc) +{ + struct file *file =3D cc->scan_file; bool triggered_wb =3D false; enum scan_result result; - struct file *file; - pgoff_t pgoff; =20 - mmap_assert_locked(mm); - - if (vma_is_anonymous(vma)) { - result =3D collapse_scan_anon_pmd(vma, addr, end, cc); - if (!cc->select_orders) - goto end; - - /* collapse_anon_pmd() takes mmap_lock itself, where it needs it */ - mmap_read_unlock(mm); - *lock_dropped =3D true; - - result =3D collapse_anon_pmd(mm, addr, end, cc); - goto end; - } - - file =3D get_file(vma->vm_file); - pgoff =3D linear_page_index(vma, addr); - - mmap_read_unlock(mm); - *lock_dropped =3D true; retry: - result =3D collapse_scan_file(mm, addr, file, pgoff, cc); + result =3D collapse_pagecache_pmd(mm, addr, file, cc->scan_pgoff, cc); =20 /* Dirty pages are worth a writeback and one more try, if asked for */ if (cc->policy.writeback_dirty && result =3D=3D SCAN_PAGE_DIRTY_OR_WRITEB= ACK && !triggered_wb && mapping_can_writeback(file->f_mapping)) { - const loff_t lstart =3D (loff_t)pgoff << PAGE_SHIFT; + const loff_t lstart =3D (loff_t)cc->scan_pgoff << PAGE_SHIFT; const loff_t lend =3D lstart + HPAGE_PMD_SIZE - 1; =20 filemap_write_and_wait_range(file->f_mapping, lstart, lend); @@ -3647,6 +3670,7 @@ enum scan_result collapse_single_pmd(unsigned long ad= dr, goto retry; } fput(file); + cc->scan_file =3D NULL; =20 if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) { mmap_read_lock(mm); @@ -3659,6 +3683,67 @@ enum scan_result collapse_single_pmd(unsigned long a= ddr, result =3D SCAN_SUCCEED; mmap_read_unlock(mm); } -end: + return result; } + +/* + * Scan one table's worth of @vma and decide whether there is anything to = collapse + * in it. The caller holds mmap_lock for reading and still holds it when = this + * returns: what is looked at is either the VMA or a page table that the l= ock + * keeps in place. + * + * Returns whether collapse_run_pmd() has anything to do, and a scan that = found + * something has to be run: the file side takes a reference on the file wh= ile it + * still has the VMA to take it from, and the run is what gives it back. = What the + * scan turned down is left in cc->scan_refusal either way. + */ +bool collapse_scan_pmd(struct vm_area_struct *vma, unsigned long addr, + unsigned long end, struct collapse_control *cc, + unsigned long vma_orders) +{ + struct mm_struct *mm =3D vma->vm_mm; + + mmap_assert_locked(mm); + + /* + * What the scan answers with, so cleared before it runs. + * collapse_anon_scan_init() clears the orders too, but only once the + * table has turned out to be there. + */ + cc->select_orders =3D 0; + + /* Ours to give back only if the last scan was never run */ + if (WARN_ON_ONCE(cc->scan_file)) { + fput(cc->scan_file); + cc->scan_file =3D NULL; + } + + if (unlikely(collapse_test_exit_or_disable(mm))) + cc->scan_refusal =3D SCAN_ANY_PROCESS; + else if (addr < vma->vm_start || end > vma->vm_end) + cc->scan_refusal =3D SCAN_ADDRESS_RANGE; + else if (!vma_orders) + cc->scan_refusal =3D SCAN_VMA_CHECK; + else if (vma_is_anonymous(vma)) + collapse_scan_anon_pmd(vma, addr, end, cc, vma_orders); + else + cc->scan_refusal =3D collapse_scan_file_pmd(vma, addr, cc); + + return cc->select_orders || cc->scan_file; +} + +/* + * Collapse what the scan selected. Called with no mmap_lock: the caller = gives it + * up first, because a collapse takes it again for each round and revalida= tes + * under it, and holding it across the whole collapse would keep a writer = to the + * address space waiting for it. + */ +enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr, + unsigned long end, struct collapse_control *cc) +{ + if (cc->scan_file) + return collapse_file_pmd(mm, addr, cc); + else + return collapse_anon_pmd(mm, addr, end, cc); +} diff --git a/mm/collapse.h b/mm/collapse.h index dc60806fb81e..4af7bb9c4261 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -150,6 +150,13 @@ struct collapse_control { */ enum scan_result scan_refusal; =20 + /* + * A reference the file side takes while it still has the VMA, since the + * collapse runs without it, and the offset it decided on. + */ + struct file *scan_file; + pgoff_t scan_pgoff; + /* Why the last window was refused */ enum scan_result select_result; =20 @@ -187,12 +194,47 @@ static inline int collapse_test_exit_or_disable(struc= t mm_struct *mm) mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); } =20 +/* + * A caller states what it allows in the policy, takes a control for the a= rrays a + * round needs, and then hands over one PTE table's worth of a VMA at a ti= me: + * + * collapse_control_init(cc); once per control + * fill in cc->policy; what this caller allows + * collapse_scan_pmd(vma, addr, end, cc); per table, as often as wanted + * collapse_run_pmd(mm, addr, end, cc); when the scan found work + * collapse_control_release(cc); + * + * The caller holds mmap_lock for reading and passes a range within one PT= E table + * of @vma. A range the VMA does not cover is refused, which is also how = a caller + * learns that its own range shrank. + * + * A scan returns with that lock still held: it only reads, and almost eve= ry table + * it is offered has nothing to collapse, so a caller walks a whole VMA un= der the + * one lock it took to get there. + * + * A collapse is called without it: the caller gives the lock up first, an= d with it + * @vma and anything derived under it, so a caller carrying on has to look= up + * again. What the collapse does -- allocate, quiesce, copy, flush -- is = slow + * enough that a writer would wait behind it, so it takes the lock again p= er round + * instead, and revalidates rather than trusting what the scan saw. + * + * A scan that found something has to be run: the file side takes a refere= nce on + * the file while it still has the VMA to take it from, and the run is wha= t gives + * it back. What it turned down is left in cc->scan_refusal, for a caller= that has + * to report why a table was not collapsed. + * + * A control is not reentrant: it carries the arrays a round works out of,= so one + * per collapsing thread. + */ int collapse_control_init(struct collapse_control *cc); void collapse_control_release(struct collapse_control *cc); -enum scan_result collapse_single_pmd(unsigned long addr, unsigned long end, - struct vm_area_struct *vma, bool *lock_dropped, - struct collapse_control *cc); +bool collapse_scan_pmd(struct vm_area_struct *vma, unsigned long addr, + unsigned long end, struct collapse_control *cc, + unsigned long vma_orders); +enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr, + unsigned long end, struct collapse_control *cc); =20 +/* Which orders a VMA may collapse to, empty when it may not collapse at a= ll */ unsigned long collapse_possible_orders(struct vm_area_struct *vma, vm_flags_t vm_flags, enum tva_type tva_flags); =20 diff --git a/mm/khugepaged.c b/mm/khugepaged.c index b7fc93e11d6b..47c134cd4129 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -485,64 +485,6 @@ static void collapse_policy_khugepaged(struct collapse= _policy *p) p->tva_type =3D TVA_KHUGEPAGED; } =20 -/* MADV_COLLAPSE was asked for explicitly, so it is not held to those. */ -static void collapse_policy_forced(struct collapse_policy *p) -{ - p->max_ptes_none =3D HPAGE_PMD_NR; - p->max_ptes_swap =3D HPAGE_PMD_NR; - p->max_ptes_shared =3D HPAGE_PMD_NR; - p->strict_sub_pmd =3D false; - p->skip_lazyfree =3D false; - p->require_referenced =3D false; - p->install_pmd =3D true; - p->writeback_dirty =3D true; - p->gfp =3D GFP_TRANSHUGE; - p->tva_type =3D TVA_FORCED_COLLAPSE; -} - -/* - * If mmap_lock temporarily dropped, revalidate vma - * after taking the mmap_lock again. - * Returns enum scan_result value. - */ - -static enum scan_result hugepage_vma_revalidate(struct mm_struct *mm, unsi= gned long address, - bool expect_anon, struct vm_area_struct **vmap, - struct collapse_control *cc, unsigned int order) -{ - struct vm_area_struct *vma; - enum tva_type type =3D cc->policy.tva_type; - - if (unlikely(collapse_test_exit_or_disable(mm))) - return SCAN_ANY_PROCESS; - - *vmap =3D vma =3D find_vma(mm, address); - if (!vma) - return SCAN_VMA_NULL; - - /* - * We cannot collapse VMA regions that do not span the full PMD. This is - * due to the potential of the PMD being shared by another VMA leaving - * us vulnerable to a race condition. Always check the PMD order here to - * ensure its not shared by another VMA. We'd need to lock all VMAs in - * the PMD range to support this. - */ - if (!thp_vma_suitable_order(vma, address, PMD_ORDER)) - return SCAN_ADDRESS_RANGE; - if (!thp_vma_allowable_orders(vma, vma->vm_flags, type, BIT(order))) - return SCAN_VMA_CHECK; - /* - * Anon VMA expected, the address may be unmapped then - * remapped to file after khugepaged reacquired the mmap_lock. - * - * thp_vma_allowable_orders() may return true for qualified file - * vmas. - */ - if (expect_anon && (!(*vmap)->anon_vma || !vma_is_anonymous(*vmap))) - return SCAN_PAGE_ANON; - return SCAN_SUCCEED; -} - static void collect_mm_slot(struct mm_slot *slot) { struct mm_struct *mm =3D slot->mm; @@ -636,8 +578,7 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, khugepaged_scan.address =3D hstart; =20 while (khugepaged_scan.address < hend) { - unsigned long pmd_addr, range_end; - bool lock_dropped =3D false; + unsigned long pmd_addr, range_end, start; =20 /* One table's worth at most, and never past the VMA */ pmd_addr =3D khugepaged_scan.address & HPAGE_PMD_MASK; @@ -649,24 +590,24 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, =20 VM_WARN_ON_ONCE(khugepaged_scan.address < hstart); =20 - *result =3D collapse_single_pmd(khugepaged_scan.address, - range_end, vma, - &lock_dropped, cc); - if (*result =3D=3D SCAN_SUCCEED) - ++khugepaged_pages_collapsed; + start =3D khugepaged_scan.address; /* move to next address */ khugepaged_scan.address =3D range_end; - if (lock_dropped) - /* - * We released mmap_lock so break loop. Note - * that we drop mmap_lock before all hugepage - * allocations, so if allocation fails, we are - * guaranteed to break here and report the - * correct result back to caller. - */ - goto breakouterloop_mmap_lock; - if (cc->progress >=3D progress_max) - goto breakouterloop; + + /* If nothing to collapse, the lock is still ours */ + if (!collapse_scan_pmd(vma, start, range_end, cc, orders)) { + *result =3D cc->scan_refusal; + if (cc->progress >=3D progress_max) + goto breakouterloop; + continue; + } + + /* collapse_run_pmd() takes its own locks, so give this up */ + mmap_read_unlock(mm); + *result =3D collapse_run_pmd(mm, start, range_end, cc); + if (*result =3D=3D SCAN_SUCCEED) + ++khugepaged_pages_collapsed; + goto breakouterloop_mmap_lock; } } breakouterloop: @@ -904,6 +845,21 @@ bool current_is_khugepaged(void) return kthread_func(current) =3D=3D khugepaged; } =20 +/* MADV_COLLAPSE was asked for explicitly, so it is not held to those. */ +static void collapse_policy_forced(struct collapse_policy *p) +{ + p->max_ptes_none =3D HPAGE_PMD_NR; + p->max_ptes_swap =3D HPAGE_PMD_NR; + p->max_ptes_shared =3D HPAGE_PMD_NR; + p->strict_sub_pmd =3D false; + p->skip_lazyfree =3D false; + p->require_referenced =3D false; + p->install_pmd =3D true; + p->writeback_dirty =3D true; + p->gfp =3D GFP_TRANSHUGE; + p->tva_type =3D TVA_FORCED_COLLAPSE; +} + static int madvise_collapse_errno(enum scan_result r) { /* @@ -942,9 +898,10 @@ int madvise_collapse(struct vm_area_struct *vma, unsig= ned long start, struct collapse_control *cc; struct mm_struct *mm =3D vma->vm_mm; unsigned long hstart, hend, addr; + /* What the VMA allows; valid only while its lock is held */ + unsigned long vma_orders; enum scan_result last_fail =3D SCAN_FAIL; int thps =3D 0; - bool mmap_unlocked =3D false; int err; =20 BUG_ON(vma->vm_start > start); @@ -971,28 +928,67 @@ int madvise_collapse(struct vm_area_struct *vma, unsi= gned long start, } =20 mmgrab(mm); + + /* + * Nothing below wants the lock the VMA walk left held, and + * lru_add_drain_all() waits on every CPU, so give it up first. The + * walk carries on under mmap_lock and its own caller is what drops it, + * so reporting this only tells the walk that its VMA is now stale. + */ + mmap_read_unlock(mm); + *lock_dropped =3D true; + vma =3D NULL; + vma_orders =3D 0; lru_add_drain_all(); =20 for (addr =3D hstart; addr < hend; addr +=3D HPAGE_PMD_SIZE) { - enum scan_result result =3D SCAN_FAIL; + enum scan_result result; =20 - if (mmap_unlocked) { + /* + * A collapse gives the lock up, and the VMA has to be found + * again after one: it can shrink while nothing is held. A scan + * that finds nothing to collapse leaves the lock alone, so a + * range that is already collapsed walks it without relocking. + * + * Reschedule only here, where nothing is held: a preemption + * point under a lock is a writer waiting longer. + */ + if (!vma) { cond_resched(); mmap_read_lock(mm); - mmap_unlocked =3D false; - *lock_dropped =3D true; - result =3D hugepage_vma_revalidate(mm, addr, false, &vma, - cc, HPAGE_PMD_ORDER); - if (result !=3D SCAN_SUCCEED) { - last_fail =3D result; - goto out_nolock; + vma =3D vma_lookup(mm, addr); + if (!vma) { + mmap_read_unlock(mm); + hend =3D addr; + break; } - - hend =3D min(hend, vma->vm_end & HPAGE_PMD_MASK); + vma_orders =3D collapse_possible_orders(vma, + vma->vm_flags, TVA_FORCED_COLLAPSE); } =20 - result =3D collapse_single_pmd(addr, addr + HPAGE_PMD_SIZE, vma, - &mmap_unlocked, cc); + /* If nothing to collapse, the lock is still ours */ + if (!collapse_scan_pmd(vma, addr, addr + HPAGE_PMD_SIZE, cc, + vma_orders)) { + result =3D cc->scan_refusal; + } else { + /* collapse_run_pmd() takes its own locks, so give this up */ + mmap_read_unlock(mm); + vma =3D NULL; + /* The mask belonged to that lock, not to this range */ + vma_orders =3D 0; + + result =3D collapse_run_pmd(mm, addr, + addr + HPAGE_PMD_SIZE, cc); + } + + /* + * The VMA shrank under us, so the rest of the range was never + * ours to collapse: stop, and expect only what came before. + */ + if (result =3D=3D SCAN_VMA_NULL || result =3D=3D SCAN_ADDRESS_RANGE) { + hend =3D addr; + break; + } =20 switch (result) { case SCAN_SUCCEED: @@ -1015,18 +1011,14 @@ int madvise_collapse(struct vm_area_struct *vma, un= signed long start, default: last_fail =3D result; /* Other error, exit */ - goto out_maybelock; + goto out; } } =20 -out_maybelock: - /* Caller expects us to hold mmap_lock on return */ - if (mmap_unlocked) { - *lock_dropped =3D true; +out: + /* The VMA walk this returns to expects the lock it was holding */ + if (!vma) mmap_read_lock(mm); - } -out_nolock: - mmap_assert_locked(mm); mmdrop(mm); collapse_control_release(cc); kfree(cc); diff --git a/mm/mremap.c b/mm/mremap.c index e8df5cdb0ac9..a25ac3db787a 100644 --- a/mm/mremap.c +++ b/mm/mremap.c @@ -244,7 +244,7 @@ static int move_ptes(struct pagetable_move_control *pmc, goto out; } /* - * Now new_pte is none, so collapse_scan_file() path can not find + * Now new_pte is none, so collapse_pagecache_pmd() path can not find * this by traversing file->f_mapping, so there is no concurrency with * retract_page_tables(). In addition, we already hold the exclusive * mmap_lock, so this new_pte page is stable, so there is no need to get --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0C1623EC6BE; Sun, 16 Aug 2026 22:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920445; cv=none; b=rJLQ89RsmM7c69cMmC3vsBxtbb/F0REVTw+FgmjzIY35DFDtnAHQ6p4X2MtwQ0/+kiP/KlVKAdpgWHchJc9axe0NB83rQks+wPI/SY/VMjKmN5TyhxszS5qCTp4KIl6T+j8/ptV8jWeEVlj3Ikdf4qdUIe/bXjYaeoAd98Uxnp4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920445; c=relaxed/simple; bh=8AlcU4/uD5ioAfM0LSoQ50JrdK2H5vgan7P2Fzte1tY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uswvTSPcIj5dHPYHoiZY6VDndupLMtjQ/cke2pahSQAQx+CQgO6mrX5WBa0H2mc7Ohn/VadLzrkpGLMqyF+M9SZ4ws81FvSyVMA9TLG/MZwlGTDk2PSlmUSyAS2874equ/XLZ7G94yGe7Kqpovzrzt3thenIPFM4DKjIAgxFLj8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=0TKJLNxH; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Qt1zivLi; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="0TKJLNxH"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Qt1zivLi" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfout.phl.internal (Postfix) with ESMTP id 6ACC9EC0243; Sun, 16 Aug 2026 18:47:22 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:47:22 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920442; x= 1787006842; bh=4iXJOXAQHj2CY1FuoUAjEsT25SgdR4G2WYeHbr672pA=; b=0 TKJLNxHxHvDTi1JO5KJFcfrhZXanz89LFVOf4ilaSlwQgpSTm7QWtQRimntdF5p0 n/EF9qklNjR2uCIUuju2KAWMJconEDDI495F5fd5elc8QcCjl04Xw/HqGjgwwPDy LVmCJPf7Zym5bLWqQoLBElgKw0KaINCDJwYq5PEWclF5fahzo+Al+tVzVZduzc8V iq/kKSfK8pqQSFP26270CsOOjSP14QNSTn3cbn4MiOF0356ySRInnronoVsxOsYo zwiVM3T3eCshiUGrfSZe2bt0w0AU3a8u3ylEZqzIlJhajEFrNxgO/be7y9iknHtj NRgN5LRz2RFsqLwsAnikA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920442; x=1787006842; bh=4 iXJOXAQHj2CY1FuoUAjEsT25SgdR4G2WYeHbr672pA=; b=Qt1zivLiWpZFt3wta bvZjbH7EaUemCPXzSummZwE1ZYZZ+6g95taLyQ5c5hMMMuXVQ/ilzOkJgaxb0puk 0nTZDEf5RW9ym/62Hb3gWGIYzHDgFP8RM/fOS95BZBrzZ89obtp+2N13jNwTVcdI e5913E0SIptTjhlwKKcMAfLybKLnNWDf/8BDWifYF146YNmqh54PiWDeeFELOHK/ Frw4AAZohmormhTAtH4Zhp8BGUAY5zke/tK3SbhhfD558L27tgsKDZbL/bPCbId3 l407ebVIaTf7VKi3FXwmxA1FouJCNGepkRskrPVhE3dZISl/4d0lfWJNApLkm+N0 LAjng== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGQiHMMNOd1xnoN6lXQSIrGAe7MwYo75Qud3VSOh3aIjL1T4IX+wE1i/v23Y8aWnG m9SHT5inuXsr6fNeLF7q+XcMZ+pkZJxjpb2XQQyKrFdYKaLz7QDFEiegvQaT+Y5rGxRgjF euipinLXiFMtCpFgMqQP9Dwn4SHq3YG285QsMU4KB7YfRQHYWKIG6dJTmdP/7bPKv1qUNe 3LkhBXaaROXWrLGafj+UVrogU8Abxo2OvgWTBKIlLl3SY5Eb27sggkfdoaZJgpjY1+W5mK IYgTH3bmg7umgeRjJmx+WAhyhWPmGn2xLzWKHa3nzXVCTdFk2+6Uh9D41+s+1+4ThvFtWC NAyvoHSHbEIY0NAsm0LaLE+G1U7bQRd7xlETqJv5zW95qYCPd700IVMa68R8Moc34GA1i+ gm9TISZCFWIb/mVF32aBjGXoy9/OTcQw1SJLqFnJ6MVH8f+sWdnXGMmNW9GJViGimaynlq Bgx2NQMcbYJRpelOJJtGx1D9s5SAzCfWiUQHyjvO4/TDl+cGEWwR6X5GpK7r/VDxSI0yQM Dekgauzsswp5pbCNuPQ8dAYstyi2KSMy+pRNFz4PJ2Ef5H5Gbge/Q8j2nN2vOVsoIApq4q u2dDI9evWydLU7pmgESPRnN5Pdm1JBhAmbUcJZHRjU0rpmb9GBH+ONBAp33Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:21 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 34/57] mm/collapse: implement MADV_COLLAPSE in madvise.c Date: Sun, 16 Aug 2026 23:45:46 +0100 Message-ID: <20260816224609.308019-35-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" MADV_COLLAPSE is a madvise operation, but its implementation sat in khugepaged.c. The daemon's file therefore also held a syscall's worth of code that has nothing to do with the daemon: the walk over the user's range, the per-PMD loop, and the errno translation that reports back through madvise(2). Move it to madvise.c, among the operations it belongs with, along with the errno map and the policy it states for itself. It takes a struct madvise_behavior like every one of those operations, which is where the range, the VMA and the lock-dropped flag it used to be handed separately already live. It stays a caller of the same interface khugepaged uses, so nothing about the collapse changes. The eligibility test reads collapse_possible_orders() rather than collapse_possible(), a static wrapper around it that madvise.c cannot reach. With the declaration in huge_mm.h no longer needed, the !CONFIG_TRANSPARENT_HUGEPAGE stub moves in with it. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/linux/huge_mm.h | 9 -- mm/khugepaged.c | 182 ------------------------------------- mm/madvise.c | 195 +++++++++++++++++++++++++++++++++++++++- 3 files changed, 193 insertions(+), 193 deletions(-) diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index c745f7ad2298..8ca0fa3be2ac 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -510,8 +510,6 @@ change_huge_pud(struct mmu_gather *tlb, struct vm_area_= struct *vma, =20 int hugepage_madvise(struct vm_area_struct *vma, vm_flags_t *vm_flags, int advice); -int madvise_collapse(struct vm_area_struct *vma, unsigned long start, - unsigned long end, bool *lock_dropped); void vma_adjust_trans_huge(struct vm_area_struct *vma, unsigned long start, unsigned long end, struct vm_area_struct *next); spinlock_t *__pmd_trans_huge_lock(pmd_t *pmd, struct vm_area_struct *vma); @@ -715,13 +713,6 @@ static inline int hugepage_madvise(struct vm_area_stru= ct *vma, return -EINVAL; } =20 -static inline int madvise_collapse(struct vm_area_struct *vma, - unsigned long start, - unsigned long end, bool *lock_dropped) -{ - return -EINVAL; -} - static inline void vma_adjust_trans_huge(struct vm_area_struct *vma, unsigned long start, unsigned long end, diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 47c134cd4129..967cc472b6dc 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -844,185 +844,3 @@ bool current_is_khugepaged(void) { return kthread_func(current) =3D=3D khugepaged; } - -/* MADV_COLLAPSE was asked for explicitly, so it is not held to those. */ -static void collapse_policy_forced(struct collapse_policy *p) -{ - p->max_ptes_none =3D HPAGE_PMD_NR; - p->max_ptes_swap =3D HPAGE_PMD_NR; - p->max_ptes_shared =3D HPAGE_PMD_NR; - p->strict_sub_pmd =3D false; - p->skip_lazyfree =3D false; - p->require_referenced =3D false; - p->install_pmd =3D true; - p->writeback_dirty =3D true; - p->gfp =3D GFP_TRANSHUGE; - p->tva_type =3D TVA_FORCED_COLLAPSE; -} - -static int madvise_collapse_errno(enum scan_result r) -{ - /* - * MADV_COLLAPSE breaks from existing madvise(2) conventions to provide - * actionable feedback to caller, so they may take an appropriate - * fallback measure depending on the nature of the failure. - */ - switch (r) { - case SCAN_ALLOC_HUGE_PAGE_FAIL: - return -ENOMEM; - case SCAN_CGROUP_CHARGE_FAIL: - case SCAN_EXCEED_NONE_PTE: - return -EBUSY; - /* Resource temporary unavailable - trying again might succeed */ - case SCAN_PAGE_COUNT: - case SCAN_PAGE_LOCK: - case SCAN_PAGE_LRU: - case SCAN_DEL_PAGE_LRU: - case SCAN_PAGE_FILLED: - case SCAN_PAGE_HAS_PRIVATE: - case SCAN_PAGE_DIRTY_OR_WRITEBACK: - return -EAGAIN; - /* - * Other: Trying again likely not to succeed / error intrinsic to - * specified memory range. khugepaged likely won't be able to collapse - * either. - */ - default: - return -EINVAL; - } -} - -int madvise_collapse(struct vm_area_struct *vma, unsigned long start, - unsigned long end, bool *lock_dropped) -{ - struct collapse_control *cc; - struct mm_struct *mm =3D vma->vm_mm; - unsigned long hstart, hend, addr; - /* What the VMA allows; valid only while its lock is held */ - unsigned long vma_orders; - enum scan_result last_fail =3D SCAN_FAIL; - int thps =3D 0; - int err; - - BUG_ON(vma->vm_start > start); - BUG_ON(vma->vm_end < end); - - if (!collapse_possible(vma, vma->vm_flags, TVA_FORCED_COLLAPSE)) - return -EINVAL; - - hstart =3D ALIGN(start, HPAGE_PMD_SIZE); - hend =3D ALIGN_DOWN(end, HPAGE_PMD_SIZE); - - if (hstart >=3D hend) - return 0; - - cc =3D kmalloc_obj(*cc); - if (!cc) - return -ENOMEM; - collapse_policy_forced(&cc->policy); - cc->progress =3D 0; - err =3D collapse_control_init(cc); - if (err) { - kfree(cc); - return err; - } - - mmgrab(mm); - - /* - * Nothing below wants the lock the VMA walk left held, and - * lru_add_drain_all() waits on every CPU, so give it up first. The - * walk carries on under mmap_lock and its own caller is what drops it, - * so reporting this only tells the walk that its VMA is now stale. - */ - mmap_read_unlock(mm); - *lock_dropped =3D true; - vma =3D NULL; - vma_orders =3D 0; - lru_add_drain_all(); - - for (addr =3D hstart; addr < hend; addr +=3D HPAGE_PMD_SIZE) { - enum scan_result result; - - /* - * A collapse gives the lock up, and the VMA has to be found - * again after one: it can shrink while nothing is held. A scan - * that finds nothing to collapse leaves the lock alone, so a - * range that is already collapsed walks it without relocking. - * - * Reschedule only here, where nothing is held: a preemption - * point under a lock is a writer waiting longer. - */ - if (!vma) { - cond_resched(); - mmap_read_lock(mm); - vma =3D vma_lookup(mm, addr); - if (!vma) { - mmap_read_unlock(mm); - hend =3D addr; - break; - } - vma_orders =3D collapse_possible_orders(vma, - vma->vm_flags, TVA_FORCED_COLLAPSE); - } - - /* If nothing to collapse, the lock is still ours */ - if (!collapse_scan_pmd(vma, addr, addr + HPAGE_PMD_SIZE, cc, - vma_orders)) { - result =3D cc->scan_refusal; - } else { - /* collapse_run_pmd() takes its own locks, so give this up */ - mmap_read_unlock(mm); - vma =3D NULL; - /* The mask belonged to that lock, not to this range */ - vma_orders =3D 0; - - result =3D collapse_run_pmd(mm, addr, - addr + HPAGE_PMD_SIZE, cc); - } - - /* - * The VMA shrank under us, so the rest of the range was never - * ours to collapse: stop, and expect only what came before. - */ - if (result =3D=3D SCAN_VMA_NULL || result =3D=3D SCAN_ADDRESS_RANGE) { - hend =3D addr; - break; - } - - switch (result) { - case SCAN_SUCCEED: - case SCAN_PMD_MAPPED: - ++thps; - break; - /* Whitelisted set of results where continuing OK */ - case SCAN_NO_PTE_TABLE: - case SCAN_PTE_NON_PRESENT: - case SCAN_PTE_UFFD: - case SCAN_LACK_REFERENCED_PAGE: - case SCAN_PAGE_NULL: - case SCAN_PAGE_COUNT: - case SCAN_PAGE_LOCK: - case SCAN_PAGE_COMPOUND: - case SCAN_PAGE_LRU: - case SCAN_DEL_PAGE_LRU: - last_fail =3D result; - break; - default: - last_fail =3D result; - /* Other error, exit */ - goto out; - } - } - -out: - /* The VMA walk this returns to expects the lock it was holding */ - if (!vma) - mmap_read_lock(mm); - mmdrop(mm); - collapse_control_release(cc); - kfree(cc); - - return thps =3D=3D ((hend - hstart) >> HPAGE_PMD_SHIFT) ? 0 - : madvise_collapse_errno(last_fail); -} diff --git a/mm/madvise.c b/mm/madvise.c index c179938097bf..76ddf61f043f 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -894,6 +894,198 @@ bool madvise_dontneed_free_valid_vma(struct madvise_b= ehavior *madv_behavior) return true; } =20 +#ifdef CONFIG_TRANSPARENT_HUGEPAGE +#include "collapse.h" + +/* MADV_COLLAPSE was asked for explicitly, so it is not held to those. */ +static void collapse_policy_forced(struct collapse_policy *p) +{ + p->max_ptes_none =3D HPAGE_PMD_NR; + p->max_ptes_swap =3D HPAGE_PMD_NR; + p->max_ptes_shared =3D HPAGE_PMD_NR; + p->strict_sub_pmd =3D false; + p->skip_lazyfree =3D false; + p->require_referenced =3D false; + p->install_pmd =3D true; + p->writeback_dirty =3D true; + p->gfp =3D GFP_TRANSHUGE; + p->tva_type =3D TVA_FORCED_COLLAPSE; +} + +static int madvise_collapse_errno(enum scan_result r) +{ + /* + * MADV_COLLAPSE breaks from existing madvise(2) conventions to provide + * actionable feedback to caller, so they may take an appropriate + * fallback measure depending on the nature of the failure. + */ + switch (r) { + case SCAN_ALLOC_HUGE_PAGE_FAIL: + return -ENOMEM; + case SCAN_CGROUP_CHARGE_FAIL: + case SCAN_EXCEED_NONE_PTE: + return -EBUSY; + /* Resource temporary unavailable - trying again might succeed */ + case SCAN_PAGE_COUNT: + case SCAN_PAGE_LOCK: + case SCAN_PAGE_LRU: + case SCAN_DEL_PAGE_LRU: + case SCAN_PAGE_FILLED: + case SCAN_PAGE_HAS_PRIVATE: + case SCAN_PAGE_DIRTY_OR_WRITEBACK: + return -EAGAIN; + /* + * Other: Trying again likely not to succeed / error intrinsic to + * specified memory range. khugepaged likely won't be able to collapse + * either. + */ + default: + return -EINVAL; + } +} + +static int madvise_collapse(struct madvise_behavior *madv_behavior) +{ + struct madvise_behavior_range *range =3D &madv_behavior->range; + struct vm_area_struct *vma =3D madv_behavior->vma; + struct mm_struct *mm =3D madv_behavior->mm; + unsigned long hstart, hend, addr; + struct collapse_control *cc; + unsigned long vma_orders; + enum scan_result last_fail =3D SCAN_FAIL; + int thps =3D 0; + int err; + + BUG_ON(vma->vm_start > range->start); + BUG_ON(vma->vm_end < range->end); + + if (!collapse_possible_orders(vma, vma->vm_flags, TVA_FORCED_COLLAPSE)) + return -EINVAL; + + hstart =3D ALIGN(range->start, HPAGE_PMD_SIZE); + hend =3D ALIGN_DOWN(range->end, HPAGE_PMD_SIZE); + + if (hstart >=3D hend) + return 0; + + cc =3D kmalloc_obj(*cc); + if (!cc) + return -ENOMEM; + collapse_policy_forced(&cc->policy); + cc->progress =3D 0; + err =3D collapse_control_init(cc); + if (err) { + kfree(cc); + return err; + } + + mmgrab(mm); + + /* + * Nothing below wants the lock the VMA walk left held, and + * lru_add_drain_all() waits on every CPU, so give it up first. The + * walk carries on under mmap_lock and its own caller is what drops it, + * so reporting this only tells the walk that its VMA is now stale. + */ + mmap_read_unlock(mm); + mark_mmap_lock_dropped(madv_behavior); + vma =3D NULL; + vma_orders =3D 0; + lru_add_drain_all(); + + for (addr =3D hstart; addr < hend; addr +=3D HPAGE_PMD_SIZE) { + enum scan_result result; + + /* + * A collapse gives the lock up, and the VMA has to be found + * again after one: it can shrink while nothing is held. A scan + * that finds nothing to collapse leaves the lock alone, so a + * range that is already collapsed walks it without relocking. + * + * Reschedule only here, where nothing is held: a preemption + * point under a lock is a writer waiting longer. + */ + if (!vma) { + cond_resched(); + mmap_read_lock(mm); + vma =3D vma_lookup(mm, addr); + if (!vma) { + mmap_read_unlock(mm); + hend =3D addr; + break; + } + vma_orders =3D collapse_possible_orders(vma, + vma->vm_flags, TVA_FORCED_COLLAPSE); + } + + /* If nothing to collapse, the lock is still ours */ + if (!collapse_scan_pmd(vma, addr, addr + HPAGE_PMD_SIZE, cc, + vma_orders)) { + result =3D cc->scan_refusal; + } else { + /* collapse_run_pmd() takes its own locks, so give this up */ + mmap_read_unlock(mm); + vma =3D NULL; + /* The mask belonged to that lock, not to this range */ + vma_orders =3D 0; + + result =3D collapse_run_pmd(mm, addr, + addr + HPAGE_PMD_SIZE, cc); + } + + /* + * The VMA shrank under us, so the rest of the range was never + * ours to collapse: stop, and expect only what came before. + */ + if (result =3D=3D SCAN_VMA_NULL || result =3D=3D SCAN_ADDRESS_RANGE) { + hend =3D addr; + break; + } + + switch (result) { + case SCAN_SUCCEED: + case SCAN_PMD_MAPPED: + ++thps; + break; + /* Whitelisted set of results where continuing OK */ + case SCAN_NO_PTE_TABLE: + case SCAN_PTE_NON_PRESENT: + case SCAN_PTE_UFFD: + case SCAN_LACK_REFERENCED_PAGE: + case SCAN_PAGE_NULL: + case SCAN_PAGE_COUNT: + case SCAN_PAGE_LOCK: + case SCAN_PAGE_COMPOUND: + case SCAN_PAGE_LRU: + case SCAN_DEL_PAGE_LRU: + last_fail =3D result; + break; + default: + last_fail =3D result; + /* Other error, exit */ + goto out; + } + } + +out: + /* The VMA walk this returns to expects the lock it was holding */ + if (!vma) + mmap_read_lock(mm); + mmdrop(mm); + collapse_control_release(cc); + kfree(cc); + + return thps =3D=3D ((hend - hstart) >> HPAGE_PMD_SHIFT) ? 0 + : madvise_collapse_errno(last_fail); +} + +#else +static int madvise_collapse(struct madvise_behavior *madv_behavior) +{ + return -EINVAL; +} +#endif /* CONFIG_TRANSPARENT_HUGEPAGE */ + static long madvise_dontneed_free(struct madvise_behavior *madv_behavior) { struct mm_struct *mm =3D madv_behavior->mm; @@ -1361,8 +1553,7 @@ static int madvise_vma_behavior(struct madvise_behavi= or *madv_behavior) case MADV_DONTNEED_LOCKED: return madvise_dontneed_free(madv_behavior); case MADV_COLLAPSE: - return madvise_collapse(vma, range->start, range->end, - &madv_behavior->lock_dropped); + return madvise_collapse(madv_behavior); case MADV_GUARD_INSTALL: return madvise_guard_install(madv_behavior); case MADV_GUARD_REMOVE: --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D6DC23F9F4D; Sun, 16 Aug 2026 22:47:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920446; cv=none; b=cNFIKDo+GxYfDUME26diFDDx5E+E54Dld+05j2PvDBemungvQLU8HNUJaE8Q8sO9IBiQvbdgkmL5FkLle28knQWRWKo5nLpB78aXAQ2vUs1AUK0v2SIt4OnHVblpGSIf/5Bc47iDFUpAgzFWdVUnSWWAoavHJvc0fVt84yzNEMw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920446; c=relaxed/simple; bh=tH7aBqcMPQp2gHhJGzoYFSPBz3m7t/BA8U185cyR/DI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PMse17rtjava2EoseM0cjwK2PmO+1MizCT6UxYvN9H2NvDitgNAHhyflIq4xLhVW6ZesYaoonSHDxALrSSDgWW6kZ5oEOl3a7ovNyIHomEXxRLd/bF23pfJxgxo44jKeIm6xtRHf5RVrvxhUomn2ApcjyEW0ZpA7NUTDmQh7mhc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=uOcNWDuk; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=FkAIi0+F; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="uOcNWDuk"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="FkAIi0+F" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.phl.internal (Postfix) with ESMTP id 12CC214000FA; Sun, 16 Aug 2026 18:47:24 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:47:24 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920444; x= 1787006844; bh=Z0cmyfIpPsi4r5/pE09ldrrp9VWmAB+T0yJDqsA9a2c=; b=u OcNWDuk6hsisB4OeqPabVopGNy13grQuzvgJq85TLmqA51ucU0aQ37WXnBEH/YNK xRzxr4yYNs1plWx6cVFdGkC/4CHlDzLZwJ5FNrMNvsXwqFRJ12YOdjCsuxnKK5Zm sbm0/BuPHV6sPp9Q/tPr719JszqofCKZtq28lD3HL+MK2V/IsOEMvDOjsyIVvhn7 zbzNCExPuDv65XMmc8G6OlANQB3ADYjpseHsACi/Kpr0QbqAYA29bEEsq5fIOSsL BXsi1bAVxHAiyb5y/sgr6pSFp+BZnpVsRjEOaFLrQxaytlrB2/duubCkOtkx5ZDC B+5g7f0IbuoTH5CGKbZ6Q== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920444; x=1787006844; bh=Z 0cmyfIpPsi4r5/pE09ldrrp9VWmAB+T0yJDqsA9a2c=; b=FkAIi0+FqJANMN1HX dvGYnwXMnZGnQxbrvxE9qJ+ujJP6jYfAta0Qydl/FKwcz8IC3PP1emysQBvt4ION Nz3AcHasVBCmHRxqVSVm5Fyyw84T2J1ss2FDyCNyVmiz7frF1A1rvqywEV0KqPhF H+DfeDI6zOxv8XotiiRhwedQxShlJM5nOYrusFnQ6T+g5cwR6eHP5Qh/liqtUYYp Izjdb3fZgtXiMW+ymZnv/ksh/nGUuXHsbREtOt+2LPnYmnFUqLRZjiLKICrU41gV D+OzmbIrwl5ugKHYdpIa8Z/DkRO3kHJDy3Xhw/ts+xIBTKOGwNgQ00aadEUFRzjC JAlVA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFlVyZDQeKVm/cv0R7Uo2AcdsoROSs+PNN+RMPWNFgEb+eGte/fI2JmPedO/gUmIe OeoZho3fHA3jq9U7FHaKLb7Pve0ST4uqo5iHPPAgdpzTDnnLaz+AJ6xkiozNrpK1meLIn4 2sefU5BkNXUp9QcsQtONjZ0nUOLKoeupvv0JrZV8hAC+7fEjHSt2kKuDAbX3om+BmKbYq3 FepF2Gq8LX713eWlXPmLkfGboRIv9L2g+Y4/FUEyMpnFgJsl0SpF+3p/+THHwbO3LtcR+N 5KYVcWsQpsjE5I4FVrmdf0VHbzDD4H+sgFlZmgmVgtgliSro7zGS30VizlBSv+ooeZSJRL EqJyIKEeMiqcYtpukkHWvUxjLMWEpzHN7YOdlRnSW1LWMwkyRw9mHhg8vb7sytLMPBmOYq gShL0XyER5ei2tmZGVRjzvcxwS73NM45HQRDl8UV+z8O0eThEaY5NVORggJz+kwv5S0Rpt ufdX2pA/xxTUHhdIF1Wdck2LhiU4Kb2rberi3q4k2/E2MOq1pjRebomwn6ZV87is1ziRLa Y5f+B7wVSRzeuFHZnsaKZMrOyZTv3v13m+LvJzRjJZvoYi/mfKk1+4WHcVf6RYvqLO3GgY Sq1f1Yxu3n9naNwyx9jc1/U8r6cROKfXaEKiJDWRNYI8Cax25Ayrhnm8WK0w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:23 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 35/57] mm/madvise: drop MADV_COLLAPSE's redundant mm reference Date: Sun, 16 Aug 2026 23:45:47 +0100 Message-ID: <20260816224609.308019-36-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" madvise_collapse() holds an mmgrab() reference across its work, which nothing needs. mmgrab() pins the mm_struct alone; every caller already holds mm_users, which keeps the address space itself alive and so implies it: - madvise(2) works on current->mm, which lives as long as the task is in the syscall; - process_madvise(2) reaches a remote mm through mm_access(), which takes an mm_users reference and holds it until the syscall returns; - io_uring passes current->mm; - DAMON takes one with get_task_mm() and drops it after the call. Drop the mmgrab()/mmdrop() pair. It has been there since commit 7d8faaf15545 ("mm/madvise: introduce MADV_COLLAPSE sync hugepage col= lapse"). Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/madvise.c | 3 --- 1 file changed, 3 deletions(-) diff --git a/mm/madvise.c b/mm/madvise.c index 76ddf61f043f..c1bb425be3f4 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -979,8 +979,6 @@ static int madvise_collapse(struct madvise_behavior *ma= dv_behavior) return err; } =20 - mmgrab(mm); - /* * Nothing below wants the lock the VMA walk left held, and * lru_add_drain_all() waits on every CPU, so give it up first. The @@ -1071,7 +1069,6 @@ static int madvise_collapse(struct madvise_behavior *= madv_behavior) /* The VMA walk this returns to expects the lock it was holding */ if (!vma) mmap_read_lock(mm); - mmdrop(mm); collapse_control_release(cc); kfree(cc); =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A77503FADFA; Sun, 16 Aug 2026 22:47:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920448; cv=none; b=Jjje3C0IE2X4BqIVEb91NxiTTJ9MO/LFskoeXEAMMg8nExagQk/LYMHzeSu2+OAoxUtADNyKShhlfsryr9W4bK/gjxm0P1Re9lWOxCa+qrhMAn59DypzTsxz3b/swlD10eN6otadFeTLD9YC/d5pk+3WwW1rJr9GsnNK2Waw8BY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920448; c=relaxed/simple; bh=jz4ij5tio1IG99wlK7+3EFva3nFMif1EfCGnf0/5n08=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VS90exYdgQ8qh4HAiVtqlUkldyT1ObRC2FFCHaU/kYE9bj2TM2F43zw5t4p0YEnT7fyPp/hoqlJIoSMES7hk8cbvpQnOPezEGc0MJFqhyw6h/t/ivG1wtR+WW00AJmvD8A7O2K+zZFgFpy6y2rOR2JplIGMxfpAx+9Df08qAIGw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=U8DMpupR; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Ne2vpsJ2; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="U8DMpupR"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Ne2vpsJ2" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfout.phl.internal (Postfix) with ESMTP id BC2ACEC0245; Sun, 16 Aug 2026 18:47:25 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:47:25 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920445; x= 1787006845; bh=jQDVLS5QXQD8PSUIRmPFOeq1RPNNLBEX8FtCDNd5u2M=; b=U 8DMpupRPxqNtjnk8MQ2nr5Tbo5J15VHzqtBgNexbVDLBpQK8kOZV1V9VY+oTuclR FQwN1MvWVLuiLmcoyeMjBjAsCK5jQiMnCByX5sgQcVEV08a7dKIl5WJmD167xKsd yxQ7OnXROVAGW/6T7F5mwovdCEXMiXGeyCbx0Blytwa7iAXZ9fy4r4TDPHRhwZaH vClCvsD2KdWW3DONN2vx+ocz7EhMKeBAZdEmGRcSKeBQrFIA3qDi1+U+CcOSNLE2 5eDbkp2Q8RBM6gtc1s0Xkg2Edc72+x+DGSpowCNEvZ8LT8Ez3oNUoKYnNGCPP2Hm xuahjYF0L026BXVJ9QcSg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920445; x=1787006845; bh=j QDVLS5QXQD8PSUIRmPFOeq1RPNNLBEX8FtCDNd5u2M=; b=Ne2vpsJ2pSqH37gTo Wmx5ryypkj6OLWQeygS4FAdp8lHexFPq9T6hnhqWd3FFoaJ8/u9jg8aewPXuAWAS vfFKx+JtJDAjFXvqvAkkEPEZvhLlVERfDcBvsH4T53nK2mozBL/Cvrt5pq5zfyCl A4+EEFq2kfAuNRTDskQydpsx9uP81j30iaWmjmWzj0Kel4UcWCky1aVAOkzvfuXa Cl+wdwDq8VFhMsWFJNio4nZN4wuuCYjfQ4iX5XsjVIRHlza3dfHgIMbrxIF7iRC3 o5N+ta537koNr/JRGVSGrgGX0lKLqtbooNDWYePfkUiDFtW0Wv6uX1IcV369p2qt GzBww== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFcTQrb7i2Wnr2aEeZt1aoD7asKutBPdJgU3YSD5OfcfuljRo8Y5zG4WAng55KhcU fdYu9dW0n+wtsUDWiYHh8AwjWiiyOeVLRvzra0gUIzgKMFXfu2OsH3wUsDwz3Dpe7vEjB9 JjBmMfD78Db5X6Vl1RDz8Wz4IQF1tjtPvs8yYIcZ4vdtxAZhotrvrKr9o/e/MWe2qw9p1d u/4ngMViLbtWn8zpPvYSIgtjaF8APmxmHCU5Ga5Rd5eIGcQp1MbbGmkchmHrMFMoUd9mrv TpVQ9+TKNOdfLLUK+2CPoHpagx885EUQUSZxhSq0AxxU5QJRQ+JA8yzxyjEYMjW1uFFjmi LqqoNifkXB79/90TSRF1slcFxxwzddkh5G5bZ3d0l80IbLrRpgBlzopY6EpPfytjdWVjfO +DqxD67otuMVUFiKrSCwFY0E3gh/4zA/82p9GG1UXZlq2+64MGfIaNuCFx78ugxlx/Vxd3 zcra6CnnMWps68YaqKx2D1rCx0QlW0Fe4XYEBXJdYnAr+DVTZT5z4gxDDkKwjNDEjaX1X+ 0qqckODmklZPLQLdNdAJbhiPAlLC0CzYGXGbs8lQ4kjAcpKVSed6BeRd67fF25naIaKKH6 DYyWnNDoe85We2Xvv2oHM5trVYOLjMtG281TyArJBr7AEnxa79wrFeGYYOLA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:25 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 36/57] mm/collapse: report what the scan found Date: Sun, 16 Aug 2026 23:45:48 +0100 Message-ID: <20260816224609.308019-37-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The anonymous scan reports nothing, and its verdict decides everything after it: which orders selection may still try, and whether the table is refused outright. The only way to see it was to infer it from what the candidates did afterwards, or from their absence. Add mm_collapse_scan: where the table was scanned, how many slots were holes or the zeropage, how many were swapped out, which orders survived the scan, and the verdict. The first two counts are why a table yields a smaller window than expected. The third is what a PMD candidate will have to read back. The order mask is the scan's whole output to selection in one number: a table refused as a unit shows an empty mask, which distinguishes it at a glance from one that merely lost the PMD order. The file scan already reported through mm_khugepaged_scan_file; this gives the anonymous side the same. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 34 ++++++++++++++++++++++++++++++ mm/collapse.c | 2 ++ 2 files changed, 36 insertions(+) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 86131845b761..573cf5428969 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -126,6 +126,40 @@ TRACE_EVENT(mm_collapse_huge_page, __entry->order) ); =20 +TRACE_EVENT(mm_collapse_scan, + + TP_PROTO(struct mm_struct *mm, unsigned long addr, int none_or_zero, + int unmapped, unsigned long orders, int result), + + TP_ARGS(mm, addr, none_or_zero, unmapped, orders, result), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(unsigned long, addr) + __field(int, none_or_zero) + __field(int, unmapped) + __field(unsigned long, orders) + __field(int, result) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->addr =3D addr; + __entry->none_or_zero =3D none_or_zero; + __entry->unmapped =3D unmapped; + __entry->orders =3D orders; + __entry->result =3D result; + ), + + TP_printk("mm=3D%p, addr=3D0x%lx, none_or_zero=3D%d, unmapped=3D%d, order= s=3D0x%lx, result=3D%s", + __entry->mm, + __entry->addr, + __entry->none_or_zero, + __entry->unmapped, + __entry->orders, + __print_symbolic(__entry->result, SCAN_STATUS)) +); + TRACE_EVENT(mm_collapse_candidate, =20 TP_PROTO(struct mm_struct *mm, unsigned long addr, unsigned int order, diff --git a/mm/collapse.c b/mm/collapse.c index f0d80204c2bd..b750a1fc81a5 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -2097,6 +2097,8 @@ static enum scan_result collapse_scan_table(struct vm= _area_struct *vma, cc->select_orders &=3D ~BIT(HPAGE_PMD_ORDER); =20 cc->scan_unmapped =3D unmapped; + trace_mm_collapse_scan(vma->vm_mm, start, none_or_zero, unmapped, + cc->select_orders, result); return result; } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 46DA73EC837; Sun, 16 Aug 2026 22:47:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920450; cv=none; b=nueKDJDsgvoYQk/QJFqY3OLR+B3tLTrYhQk/sWue4RBjQ0S9/NHOAGOg3se1oh+1TlFv3cWRAylaq7KzvNI3Okl13ZLZPMjEA46eSzaB7qoZhtOIum5S1Qzee/9WA3Hs4d6Px/SRu23spl+gjZ2vZB/PbpmHVvZK9GfKEz5RHR0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920450; c=relaxed/simple; bh=1KeF9AI4FCP0f2B4mcuZTbOO+JQHuKOB3PUuiXmHzlo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WlvSWXxEyPW8IEyIy06jE1oVARBRyyCPJb9rPa12QOIN8o88DADfyf5o40YDT8CFaEXMe44tHFXJds3Px1ciz6euEayf05UmKyrwR7PggyBs/MNwxc8/TbObW0Svf0re36nL3SY9wqmJcm3C4wlnyEurVrzunwnqPJDwegNO8hk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=d6Y2jbrL; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=SnXwxds2; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="d6Y2jbrL"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="SnXwxds2" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.phl.internal (Postfix) with ESMTP id 75E8F14000FB; Sun, 16 Aug 2026 18:47:27 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Sun, 16 Aug 2026 18:47:27 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920447; x= 1787006847; bh=RxOUZtNI4iK6oDUbeznOi5j6qWIw7a+qTCAWQpT1d7E=; b=d 6Y2jbrLxcG+b7UlQQQDlAUCnECDZvaCKHbgYYo96wTHEGmtT61158UIi24EfmWBq 1r0MiJt1qD9EWnwIlRJ1FVDPQSHCuQQk2fnFJeYBzxHGnpFwer1hwqp22UwnDM2i GF665c+8S6CCznhnz7Rxaon2MSmgX0NrQJdycztlwwZuy5MpTU2Teq7AmK9Br+rh izXEyoj2U8owy02g1nD1Qj6i4jC4YyjP5FI2rG4tML9AyVg6sUVTfMzPyT6WUOF9 zAAV5AdKP1cwODl6BUU4DOg8bV+kyvHHxuV+0Y0QmTPawvH2JEVhRfVi3ztWiTCc Vz6hvChnLuvM2EhXEAC4g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920447; x=1787006847; bh=R xOUZtNI4iK6oDUbeznOi5j6qWIw7a+qTCAWQpT1d7E=; b=SnXwxds2phmTEsFOj Q2LaIYTq8f6mE6MN8pVm+gwliOOaKFVpnarOySZ1FjmfvCANGLx6V98d0FLbztWA B9biuWNQ114cLiT0pExzYfTBAjP1lIJeDENfz04uTJKdsENucFxvn8NFnjIzDojQ rN4j4aKvkTKFdK+PWQx9CFZ2bOkedltp9DqnN7IbzuX99k9Y5+1++80cNQmc3arD b8c5YS9+7N68QAUhoUdsSU4gM67AF3m8qbE1G7ElyxXnjn20xzG3lb0IekzfMn5G gxgHPBwT5QWFMdOWoUEnYJE6pNoxkJ3aZJHhY4pVR79n5Yk+ZaFycQjFXiJOknTt X3SxA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTVW SGg4/Aho0z4roUqJEIyLZFKJMrH29ZtJmmfVu0Xk4xI9g5vXJsHBT/1UTWE52f63Iditod QO+tCxiJRIyoY5eyvFqF/AP9BwfZSJJILY0yBU4a+Y77JTTBvzFRPmzoF2NwjJw/SQV0mw 3gBuVTYtebOsIyX9LdjddvoMVP97xwlvjYt5YJaoGGSpdBdloyVecVWIk3bHdHHzpzaiBL dxjak1g91n3VQ8llIqbPfezoTwPoq1nkTX2Ju/5Z66s1JgNHx0dnFQ8yPZHSqhczgAz/jU 7Zani/FxXTb5g+3pyTmIq3aqyP8cdvHFbM+9BhJbzY2xyTHsPI232s/hahjg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:26 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 37/57] mm/collapse: report what the fault-in pass paid Date: Sun, 16 Aug 2026 23:45:49 +0100 Message-ID: <20260816224609.308019-38-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The fault-in pass is the one place a collapse does work on someone else's behalf: a swap read, or a CoW break, for every slot that needs one. How much of that a round pays is invisible, and it is the first thing to look at when collapses are slow, or when a workload notices khugepaged at all. Add mm_collapse_faultin: the faults taken across the round, with the outcome. A round that collapses a full table without faulting anything and one that reads sixty-four pages back from swap are otherwise indistinguishable. The mm is captured before the walk, because the pass returns with mmap_lock dropped on failure and the VMA is then unsafe to touch at the report. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 24 ++++++++++++++++++++++++ mm/collapse.c | 19 ++++++++++++++----- 2 files changed, 38 insertions(+), 5 deletions(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 573cf5428969..c2314e26111c 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -160,6 +160,30 @@ TRACE_EVENT(mm_collapse_scan, __print_symbolic(__entry->result, SCAN_STATUS)) ); =20 +TRACE_EVENT(mm_collapse_faultin, + + TP_PROTO(struct mm_struct *mm, unsigned int nr_faults, int result), + + TP_ARGS(mm, nr_faults, result), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(unsigned int, nr_faults) + __field(int, result) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->nr_faults =3D nr_faults; + __entry->result =3D result; + ), + + TP_printk("mm=3D%p, nr_faults=3D%u, result=3D%s", + __entry->mm, + __entry->nr_faults, + __print_symbolic(__entry->result, SCAN_STATUS)) +); + TRACE_EVENT(mm_collapse_candidate, =20 TP_PROTO(struct mm_struct *mm, unsigned long addr, unsigned int order, diff --git a/mm/collapse.c b/mm/collapse.c index b750a1fc81a5..1b5db42b6991 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -509,8 +509,10 @@ static enum scan_result collapse_revalidate(struct vm_= area_struct *vma, =20 /* * Bring one address to a state the freeze will accept: present, and exclu= sive if - * it is anonymous. Returns with mmap_lock dropped on every failure, beca= use the - * fault path may drop it and the caller cannot tell which case it is in. + * it is anonymous. Every fault it takes to get there counts in *nr_fault= s, each + * one an allocation or a read the round is paying for. Returns with mmap= _lock + * dropped on every failure, because the fault path may drop it and the ca= ller + * cannot tell which case it is in. * * SCAN_EXCEED_SWAP_PTE is the exception: it is a verdict on this candidate * rather than on the round, nothing was faulted to reach it, and it keeps= the @@ -518,7 +520,8 @@ static enum scan_result collapse_revalidate(struct vm_a= rea_struct *vma, */ static enum scan_result collapse_faultin_addr(struct vm_area_struct *vma, struct collapse_candidate *cand, - pmd_t *pmd, unsigned long addr) + pmd_t *pmd, unsigned long addr, + unsigned int *nr_faults) { struct mm_struct *mm =3D vma->vm_mm; const unsigned int flags =3D FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_UNSHARE | @@ -571,6 +574,7 @@ static enum scan_result collapse_faultin_addr(struct vm= _area_struct *vma, =20 /* Only swap or shared PTEs reach here; the rest broke out */ ret =3D handle_mm_fault(vma, addr, flags, NULL); + (*nr_faults)++; /* * Not a verdict on this window: the fault dropped the lock to * wait, which is what a swap-in normally does. Distinct from @@ -600,7 +604,9 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, struct collapse_control *cc, pmd_t *pmd) { + struct mm_struct *mm =3D vma->vm_mm; enum scan_result result =3D SCAN_SUCCEED; + unsigned int nr_faults =3D 0; unsigned int i; =20 for (i =3D 0; i < cc->nr_candidates; i++) { @@ -616,7 +622,8 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, j++, addr +=3D PAGE_SIZE) { enum scan_result r; =20 - r =3D collapse_faultin_addr(vma, cand, pmd, addr); + r =3D collapse_faultin_addr(vma, cand, pmd, addr, + &nr_faults); /* * The one failure that judges this candidate rather * than the round, and so the one that leaves the lock @@ -627,7 +634,7 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, if (r =3D=3D SCAN_EXCEED_SWAP_PTE) { cand->state =3D CAND_SKIPPED; cand->result =3D r; - collapse_trace_candidate(vma->vm_mm, cand, + collapse_trace_candidate(mm, cand, COLLAPSE_PASS_FAULTIN); break; } @@ -638,6 +645,8 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, } } out: + /* @vma is unsafe on the failure path: the callee dropped mmap_lock */ + trace_mm_collapse_faultin(mm, nr_faults, result); return result; } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BCD0D3E9C03; Sun, 16 Aug 2026 22:47:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920452; cv=none; b=R62vmBjU1zvjefZJk6gwUl6dEKQ1h8OFy3UMI27trUAndj64jOMYmcYMFUxW4Jh3AmYXtJmUGMUXdLDxOsOFigTQb+0nF2kFPPPXsn9JACWLzYFY5ZdikIzL3bAdkf6nk5keOa6wQTHrYLyVLbG+KE8T5NXxm7IPNHn5fzA5YVc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920452; c=relaxed/simple; bh=5wEAEo7D2EREMpKRRC37/LWq2WJPnJJyB+LP5SnhqD8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WAEv4CtmNaKvY8K6F95pzW3wi8IVJLDWEDrNk3ljZQSGyT9SZqGdWT90yUXxdsuDn+DTn6XU1IcprLpjBrZOWbw5pO7EIA7Bjkm3IOJDqvz7/i9mLRR+1mp/l7v8vj+cwu5OU1kmOrNKT23U5RjxfTkHAwgMz0PLC4IofxNaL/w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=BhepOWU2; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=k/h55XPP; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="BhepOWU2"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="k/h55XPP" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 1C5C514000FD; Sun, 16 Aug 2026 18:47:29 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:29 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920449; x= 1787006849; bh=aj0XhmlZKMkc4zRSewDPCH6H4jDvf0UGdOKiLur85eQ=; b=B hepOWU2udgD96fUUwtE9ixv2rB8W/u51vG0/xWbPODW9XQc8uFxwwY7kIN25f5tJ RL/ZuQec3DaHCu1usTRQlGM0w7Ub2FJEyf4XCdczxJDeHGJla6Rds+sl25tkSeSy ZZI8wiHH4PMgE60hcAhyY0nLoyUVV/DnA8gYgzm5O4lkUIRpEofVx0fIRrr0bN/1 77WnGWzwEALGiu7r3MQCZlwksyfOB+JFNi71Bv9I07HQPHOc5M+NgIKpxIIu+zxA e+qGCM3DGZ14EcTK3FWOpD8uZO3w0vQOphqYeIJooyT5w9P3DxtLltVC4rQajxCn sfo4e+qWwmTBVxVpOOHOg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920449; x=1787006849; bh=a j0XhmlZKMkc4zRSewDPCH6H4jDvf0UGdOKiLur85eQ=; b=k/h55XPP92QK72IdP n+QMfUYKLUCJQTPtmap+pScLQ+Yp2ae8jj00qGBBg9ob4Hhy3m9T/wMxyXxEdXMr hadvv/jmjBS5vM/lxpv2OmfmUpmllRz237gk/pTbswnncVe4D0V4cm+GA/HACKBG F7XIPXb7Cd0pI1u1TDQ1+60zLMKRvtW/ZzE+Ckozhyovat2U2Patacuew6HTgcQh ch3v4fUul8wDcNmzzL2uRDkRiciRCIM2rip1km2yZIvqiFX6Wf7kmtd5KDA3/2DZ 86/UBUszSBfVSKHnwyps5X0bRgTXKfvaw29WUO27o2CiNV258MIdajXNxy5xKDhN PsgEQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFe0TBuCSyWNWlaTWnsnH38dlt47jt5E6w/XYzg/0RtV5IwtbmYp4d/5OEsZ8sA5Z UrBtcKKYpJCYixx+ecrdsB329OhriQaNA5EMshLJupqSJKXKDlX4l+JTcK5if7dIpIi6UB +BFP7gkTqf7K8f8t8WCOU395N0ODAlGoFtgYkIRXoqS0rIU4lzr7JCi9Af7Sg5tkZu5jni VRAtq63fYeOiOdPNGAwbZJmoZoc8niBcpDBe0ZRzWrmdU+90wGMFCi1p+kJaaIoKSIjuYL B17M0nHwdihDuvGDXIV2beIJ54RuD10DDB0gu/vUckC8ggwfo/FmZIKnlAb1GvBKZWmqU2 UfCGDF4pDeDY2wtd30oVyTlAsy4TVuXAOh0zLmHFyWKfHtL9tC6Bk6daO343oI5hPRxCYt AvxBtZGfE/TipDf//xzVRUb31ZRkOT536+bpGgVZ8Kxgfgtcru0nhyvvmxR1jbSANi4/5p 2evqh5MIBhwBqfmKO5qfjsxs94IIT5rta9PZ+T51i0Pj91x8kBIn94YXuBy1+L0KxPoWkV 3zowkEzOOMvATrbD07q7UbCI9+OSWbkjcmlQGuCGoYv44HFiNF4xr1KtN/wJWrlEAKKVor HMGWpvWt8PaF9v6Rr/7Ww1RdG1DWJQmJnBSNOEqc6Wj+6LY5L57na4cUDNcg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:28 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 38/57] mm/collapse: report the round, and what it made faulters wait Date: Sun, 16 Aug 2026 23:45:50 +0100 Message-ID: <20260816224609.308019-39-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A round is the unit the engine actually works in, and nothing reports one. The per-candidate events say which windows were taken and which were refused. They do not say how large the batch was, how much of it landed, or the number that matters most for whether batching was the right idea: how long a faulter on a source is held up. That wait has a definite span. A thread touching a source sleeps on the folio lock the freeze took, and wakes when the putback drops it. So the interval from the first freeze to the last putback is what the round costs anyone unlucky enough to touch it. Add mm_collapse_round: that interval in microseconds, with the candidates collected, the ones installed, and the outcome. It is the thing to watch if a larger batch is ever proposed. collapse_finish() returns the number of candidates installed, and this is its first consumer. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 31 ++++++++++++++++++++++++++++++ mm/collapse.c | 17 +++++++++++++++- 2 files changed, 47 insertions(+), 1 deletion(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index c2314e26111c..d7c0195ace92 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -160,6 +160,37 @@ TRACE_EVENT(mm_collapse_scan, __print_symbolic(__entry->result, SCAN_STATUS)) ); =20 +TRACE_EVENT(mm_collapse_round, + + TP_PROTO(struct mm_struct *mm, unsigned int nr_candidates, + unsigned int nr_installed, int result, u64 freeze_to_wake_us), + + TP_ARGS(mm, nr_candidates, nr_installed, result, freeze_to_wake_us), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(unsigned int, nr_candidates) + __field(unsigned int, nr_installed) + __field(int, result) + __field(u64, freeze_to_wake_us) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->nr_candidates =3D nr_candidates; + __entry->nr_installed =3D nr_installed; + __entry->result =3D result; + __entry->freeze_to_wake_us =3D freeze_to_wake_us; + ), + + TP_printk("mm=3D%p, nr_candidates=3D%u, nr_installed=3D%u, result=3D%s, f= reeze_to_wake_us=3D%llu", + __entry->mm, + __entry->nr_candidates, + __entry->nr_installed, + __print_symbolic(__entry->result, SCAN_STATUS), + __entry->freeze_to_wake_us) +); + TRACE_EVENT(mm_collapse_faultin, =20 TP_PROTO(struct mm_struct *mm, unsigned int nr_faults, int result), diff --git a/mm/collapse.c b/mm/collapse.c index 1b5db42b6991..d0d28e8dfcea 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -8,6 +8,7 @@ #include #include /* x86 flush_tlb_range() uses hstate_vma() */ #include +#include #include #include #include @@ -20,6 +21,7 @@ #include #include #include +#include #include #include =20 @@ -1803,6 +1805,8 @@ static void collapse_round(struct mm_struct *mm, unsi= gned long pmd_addr, struct mmu_notifier_range range; struct vm_area_struct *vma; enum scan_result result; + unsigned int nr_installed; + u64 latency =3D 0; pmd_t *pmd; =20 collapse_reserve(mm, cc); @@ -1838,6 +1842,13 @@ static void collapse_round(struct mm_struct *mm, uns= igned long pmd_addr, cc->batch_start, cc->batch_end); mmu_notifier_invalidate_range_start(&range); =20 + /* + * What the faulters on this batch's sources are made to wait: they sleep + * from the freeze that took their folio's lock to the putback that drops + * it. Measured per round rather than argued about. + */ + latency =3D ktime_get_ns(); + /* * None of these can fail as a whole: the freeze takes the sources it * can and drops the candidates it cannot, and each pass after it works @@ -1849,12 +1860,16 @@ static void collapse_round(struct mm_struct *mm, un= signed long pmd_addr, collapse_install(vma, cc, pmd); collapse_putback(vma, cc); =20 + latency =3D ktime_get_ns() - latency; + mmu_notifier_invalidate_range_end(&range); =20 out_unlock: mmap_read_unlock(mm); out: - collapse_finish(mm, cc, result); + nr_installed =3D collapse_finish(mm, cc, result); + trace_mm_collapse_round(mm, cc->nr_candidates, nr_installed, result, + div_u64(latency, NSEC_PER_USEC)); } =20 /* --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BC0B43FD152; Sun, 16 Aug 2026 22:47:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920453; cv=none; b=rEy6H0W6E49Ah9NGU+lYvokiWfYyOX9BlDRBNlfGWU9jXtWYrrCinEn1yEI1pc0k3+bkBSFvvGrsqGV7L+Ai9tq5PRl3OnwBbIcK+o2lfU2T+rGyWs27dPGMt3vQbaylgnkKY67fdcMVFjQOhd8SoyPI+E60AeJCwJg7RgXK62E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920453; c=relaxed/simple; bh=escjSTIimt0mq2hDhoROdAmJcfnD1+p08WOxhz6ijas=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Wj4ZnQIOQ/6OMhjSw3MoeRhujSrqH9Gsh8ixpqKqo9NGIAwPJK/0IhV7FGuIzo9QtuotFNUCkgq+P+qRWSlqlzLgjP9QpEd6sBLnL7a0TGlBhzHT7Nb+TxYpBiz6o4pfxapF1bmrQk5Wm7H4Le2pUKq6dWLIqXWAzkClWiQKujQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=fqgBePJx; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=gpO6yFOx; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="fqgBePJx"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="gpO6yFOx" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.phl.internal (Postfix) with ESMTP id D78D914000EB; Sun, 16 Aug 2026 18:47:30 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:47:30 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920450; x= 1787006850; bh=VfEyxC9Im7mJEsMqgxdNAvFiFjPJWbWm6qSLISrElYA=; b=f qgBePJx4Bo3HbFJyeN2dZyhF6W9vnCXweOhotKni8XlBOVASElVU5dDs8OSyDqdK c3ENyv0gVOn0GF/NvBBjd/axEHh/Vu8OexIiFhoZAyggYbtaQDsBhoZ+WmmcvYre rWG8XLy6TdLlRnx+oS3VLWx1tmGctzd4o+aPwimxACFpcLZlVOOHALp/b1osRmjW HHd5fHrMb6H8J6du7ZpmsHf7pvDTzYgS52ZD1dgyZETwlhRC74DhRuOI/eWQNorX HZj3gEeBLBCiC9/17CeZw/zjl8AchF2hQnumgqlX4AT8VXMEYNLd5ZfR2hb3h0bI QYhqjteKd5UVfBPR9QY9A== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920450; x=1787006850; bh=V fEyxC9Im7mJEsMqgxdNAvFiFjPJWbWm6qSLISrElYA=; b=gpO6yFOxNjDM9wgvm 3tsCw8lKEZToheIBY3nKC2J4r3FJzgoxnFqVEsgQQnshctjzWY/pSaLewzUi76q5 3IG+sIUJ92z/e4usee5qInuTP/IcWSBbKXWykSMkyvVJfkpbluuujqyeI9PtAo+s L7W4e5VJvCyWL1xgX9T9IsqPun435x92ISw04DUDOgTZL6alvD4hE6SN6TZGgLqB UiKRUA2q9xuSOfRdxrv4J5brvGf30kjj6McGa82RdHsTOzI6JPFQ56GkeRidnsm7 /HLzOGmhJQBVg3lJKloFnVagBya3S0yGmwutcPJD8kNs0/KvWleWhECOfs0P3Tds oRWTg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7QD //r8SeRFX4Frzvk62JSidcpbkYQUZimo/tP4qXTbike3XiqGAJwN8Xr3kRNfVDMumZRQHN pQp4OueLITAI0UouqFGbnagXu8ExdBZus5wetT1OuQE4tDgzakpVZarjUM7ZWQst2XqwMV fEfmHg1AsnJoug3gmFQcrWEQqoHtEIHh6/2bOBBVF24rcuBgAxfpkyNDVHw+513ohCb20R eUKVVADrXLoZVGdZhVrUf0blIOO6VCtgQS0h6hA9dsHLM66Bq2f2A+gNMrsdllE3L1IMqD Xct95UMn7x9keDJbsXybO6LH2pGdcd4KxecDU1YaEJdlNnE1zEHhod/LDs/w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:30 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 39/57] mm/collapse: name the file collapse's tracepoints after collapse Date: Sun, 16 Aug 2026 23:45:51 +0100 Message-ID: <20260816224609.308019-40-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The two file events are named for khugepaged, from when khugepaged was the only thing that collapsed and the code lived in its file. Neither is true now: MADV_COLLAPSE reaches them through the same entry point, and they are emitted from collapse.c beside the events that do use the collapse name. Rename mm_khugepaged_scan_file to mm_collapse_scan_file, and mm_khugepaged_collapse_file to mm_collapse_file. Every event a collapse emits is then under one prefix: the anonymous and file scans, the per-candidate verdicts, the fault-in, the round, and the file collapse itself. mm_khugepaged_scan keeps its name. That one is the daemon reporting a scan pass of its own, from khugepaged.c, and it is not something a collapse emits. Renaming a tracepoint breaks anything watching the old name. In tree that is raw_tp_null_args[], which tells the BPF verifier that the folio argument of both may be NULL. Without a matching entry the verifier would let a program dereference it unchecked. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 4 ++-- kernel/bpf/btf.c | 4 ++-- mm/collapse.c | 4 ++-- 3 files changed, 6 insertions(+), 6 deletions(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index d7c0195ace92..5d0891e03bb0 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -308,7 +308,7 @@ TRACE_EVENT(mm_collapse_huge_page_swapin, __entry->order) ); =20 -TRACE_EVENT(mm_khugepaged_scan_file, +TRACE_EVENT(mm_collapse_scan_file, =20 TP_PROTO(struct mm_struct *mm, struct folio *folio, struct file *file, int present, int swap, int result), @@ -342,7 +342,7 @@ TRACE_EVENT(mm_khugepaged_scan_file, __print_symbolic(__entry->result, SCAN_STATUS)) ); =20 -TRACE_EVENT(mm_khugepaged_collapse_file, +TRACE_EVENT(mm_collapse_file, TP_PROTO(struct mm_struct *mm, struct folio *new_folio, pgoff_t index, unsigned long addr, bool is_shmem, struct file *file, int nr, int result), diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index c4673a54c4ba..58f78274d989 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -6705,8 +6705,8 @@ static const struct bpf_raw_tp_null_args raw_tp_null_= args[] =3D { /* huge_memory */ { "mm_khugepaged_scan_pmd", 0x10 }, { "mm_collapse_huge_page_isolate", 0x1 }, - { "mm_khugepaged_scan_file", 0x10 }, - { "mm_khugepaged_collapse_file", 0x10 }, + { "mm_collapse_scan_file", 0x10 }, + { "mm_collapse_file", 0x10 }, /* kmem */ { "mm_page_alloc", 0x1 }, { "mm_page_pcpu_drain", 0x1 }, diff --git a/mm/collapse.c b/mm/collapse.c index d0d28e8dfcea..6c17c83a4e21 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -3519,7 +3519,7 @@ static enum scan_result collapse_file(struct mm_struc= t *mm, unsigned long addr, folio_put(new_folio); out: VM_BUG_ON(!list_empty(&pagelist)); - trace_mm_khugepaged_collapse_file(mm, new_folio, index, addr, is_shmem, f= ile, HPAGE_PMD_NR, result); + trace_mm_collapse_file(mm, new_folio, index, addr, is_shmem, file, HPAGE_= PMD_NR, result); return result; } =20 @@ -3625,7 +3625,7 @@ static enum scan_result collapse_pagecache_pmd(struct= mm_struct *mm, } } =20 - trace_mm_khugepaged_scan_file(mm, folio, file, present, swap, result); + trace_mm_collapse_scan_file(mm, folio, file, present, swap, result); return result; } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 455E93FDC0E; Sun, 16 Aug 2026 22:47:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920455; cv=none; b=BvIWNpZj8B5VOeZcgP9OnUmvc0mH7iWcYGD90S0FZh0TOLZShiQG7+BH5EORCDhDknHK0MGp9IH6IjlkjfyFIZ9PJFI/VUJLHYzBZPf00+vp/k95FPPwD5pU2PWK1OmvuOu6rbYXQStrMRvs2BdOtxeOyvxFCMoSEny56a9z5rM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920455; c=relaxed/simple; bh=pcjGGSxxxLz1h9L02X12hc7fivat5KXReUErdn4Truk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dhxzJ8qqOrW4Gc3Jn87fB26hwDI3U+39AzxpcV+COBT6Gbqs7QuKzPTReujFX7raulh6avar6gAb+LcC6xgB4lS397/FxCPf8cOi+kriY5OF84v3R/5hD672vEreIuN9o7TJ/pI1cjdqE2UjcFpeEIeqpybLiUQpXcL+aVRkFPo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=gswbJJkI; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=U1MZ5/c6; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="gswbJJkI"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="U1MZ5/c6" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfout.phl.internal (Postfix) with ESMTP id 9E70FEC0074; Sun, 16 Aug 2026 18:47:32 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:47:32 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920452; x= 1787006852; bh=QzLUHFRii2gq6ACAxatB8F8suZiMKSOoddlYvZzHBZg=; b=g swbJJkID1QsVjOARJRQEemN69kiFpzyr+Mt7D34rSba1wHp/hOzc3cGXLET3iNCS tBqLFgEqokNSzd9iU45z844C2K+jbK5nTP3V+25a+TglzB/VUebojSICLNM1irnG qHqAWBB7Z3PMADn9c+fdCtnltLQ5u7uiJq08DsA6ETTRqtjAF8KJ3l0LTQVxQOF8 +feuumKMO6IpJuTcsTHTvT6lUrwoT8DcaTCBtTpM2eqw/4AubVswrjp4QquzTU8+ qOWKmXU01SRITpbhnsTPr0ArlNrikkjFxxgjxAdZ9agI0FJ0P/krf4dAaYwAP6Fh 08ewTb5Gt02+1ste71YVQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920452; x=1787006852; bh=Q zLUHFRii2gq6ACAxatB8F8suZiMKSOoddlYvZzHBZg=; b=U1MZ5/c6jWxKp9TY2 r+UeUxfnV92NvTdQmc2rYhwT6aaj4KjUx0gmqO3ibqTzkf8IvndbM3Mi34vVJZ8O w0K8qDSl/6fPO5/CUWnLb24ho2vmJv3Kj24OAr/PmySFGM13wChEX5xozXFBGMOJ 5h8jD5Br+4a1eeRHGq4e/Op43jbUeqph2FpNxW8iX6HXA3y/VLpu7UKX+Jd/GZy8 uh7FuYijgc1+pfSNYj1rcEe4F3LkU5JUmtZbBfc50MN4O5/sRUdWrXqieArZDyy0 EogELfhjQGNcolaF+X/BtlVDNUgbVgVzkLwSIWMlBjXx9HWaDXdLVHIl1mfYogcN 4SNcA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFcTQrb7i2Wnr2aEeZt1aoD7asKutBPdJgU3YSD5OfcfuljRo8Y5zG4WAng55KhcU fdYu9dW0n+wtsUDWiYHh8AwjWiiyOeVLRvzra0gUIzgKMFXfu2OsH3wUsDwz3Dpe7vEjB9 JjBmMfD78Db5X6Vl1RDz8Wz4IQF1tjtPvs8yYIcZ4vdtxAZhotrvrKr9o/e/MWe2qw9p1d u/4ngMViLbtWn8zpPvYSIgtjaF8APmxmHCU5Ga5Rd5eIGcQp1MbbGmkchmHrMFMoUd9mrv TpVQ9+TKNOdfLLUK+2CPoHpagx885EUQUSZxhSq0AxxU5QJRQ+JA8yzxyjEYMjW1uFFjH3 KM5L2v8sZoJpHBCvI6Nsa/j6UZDwcdngSl5LZCfOF7f6wAZPSUOnJLz4aLZzXuecwrvcr5 zhoDLChk4xqwGS1DiTGDHhodMYwKhBIzVXpaVBMmQBuijJi+gtMZx9AYlrVSShheB7nXiE D9fJK66yif99eFxt4yYvz4rkygDBGnFSOyJlmhNlLizlPFuj3MnDi4HECMSMT7L57BEJem zxSv2Uc51Suko0811otPJ48mbAVKsa1BGF3GG2NL/bPloFW20n0QACLTfTdRZ/rZ/K/XnH DIZ7dAY4OKAldJMJOl+s79NAIATJ65Dt3g1HhgU92eBcsPMM9Qk9kXneLfeA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:31 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 40/57] mm/collapse: remove the tracepoints of the mechanism that is gone Date: Sun, 16 Aug 2026 23:45:52 +0100 Message-ID: <20260816224609.308019-41-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Four tracepoints lost their only emitter when the anonymous mechanism was deleted: mm_khugepaged_scan_pmd, mm_collapse_huge_page, mm_collapse_huge_page_isolate and mm_collapse_huge_page_swapin. Enabling one now does nothing. Remove them, and the two entries they have in raw_tp_null_args[], which can never match again. A tool that still asks for one fails to attach rather than sitting on an event that never fires. The engine's own events cover the same ground: mm_collapse_scan, mm_collapse_candidate, mm_collapse_faultin and mm_collapse_round. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/huge_memory.h | 123 ----------------------------- kernel/bpf/btf.c | 2 - 2 files changed, 125 deletions(-) diff --git a/include/trace/events/huge_memory.h b/include/trace/events/huge= _memory.h index 5d0891e03bb0..6aabf4235648 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/huge_memory.h @@ -65,67 +65,6 @@ COLLAPSE_PASS_STATUS #define EM(a, b) {a, b}, #define EMe(a, b) {a, b} =20 -TRACE_EVENT(mm_khugepaged_scan_pmd, - - TP_PROTO(struct mm_struct *mm, struct folio *folio, - int referenced, int none_or_zero, int status, int unmapped), - - TP_ARGS(mm, folio, referenced, none_or_zero, status, unmapped), - - TP_STRUCT__entry( - __field(struct mm_struct *, mm) - __field(unsigned long, pfn) - __field(int, referenced) - __field(int, none_or_zero) - __field(int, status) - __field(int, unmapped) - ), - - TP_fast_assign( - __entry->mm =3D mm; - __entry->pfn =3D folio ? folio_pfn(folio) : -1; - __entry->referenced =3D referenced; - __entry->none_or_zero =3D none_or_zero; - __entry->status =3D status; - __entry->unmapped =3D unmapped; - ), - - TP_printk("mm=3D%p, scan_pfn=3D0x%lx, referenced=3D%d, none_or_zero=3D%d,= status=3D%s, unmapped=3D%d", - __entry->mm, - __entry->pfn, - __entry->referenced, - __entry->none_or_zero, - __print_symbolic(__entry->status, SCAN_STATUS), - __entry->unmapped) -); - -TRACE_EVENT(mm_collapse_huge_page, - - TP_PROTO(struct mm_struct *mm, int isolated, int status, unsigned int ord= er), - - TP_ARGS(mm, isolated, status, order), - - TP_STRUCT__entry( - __field(struct mm_struct *, mm) - __field(int, isolated) - __field(int, status) - __field(unsigned int, order) - ), - - TP_fast_assign( - __entry->mm =3D mm; - __entry->isolated =3D isolated; - __entry->status =3D status; - __entry->order =3D order; - ), - - TP_printk("mm=3D%p, isolated=3D%d, status=3D%s, order=3D%u", - __entry->mm, - __entry->isolated, - __print_symbolic(__entry->status, SCAN_STATUS), - __entry->order) -); - TRACE_EVENT(mm_collapse_scan, =20 TP_PROTO(struct mm_struct *mm, unsigned long addr, int none_or_zero, @@ -246,68 +185,6 @@ TRACE_EVENT(mm_collapse_candidate, __print_symbolic(__entry->result, SCAN_STATUS)) ); =20 -TRACE_EVENT(mm_collapse_huge_page_isolate, - - TP_PROTO(struct folio *folio, int none_or_zero, - int referenced, int status, unsigned int order), - - TP_ARGS(folio, none_or_zero, referenced, status, order), - - TP_STRUCT__entry( - __field(unsigned long, pfn) - __field(int, none_or_zero) - __field(int, referenced) - __field(int, status) - __field(unsigned int, order) - ), - - TP_fast_assign( - __entry->pfn =3D folio ? folio_pfn(folio) : -1; - __entry->none_or_zero =3D none_or_zero; - __entry->referenced =3D referenced; - __entry->status =3D status; - __entry->order =3D order; - ), - - TP_printk("scan_pfn=3D0x%lx, none_or_zero=3D%d, referenced=3D%d, status= =3D%s, order=3D%u", - __entry->pfn, - __entry->none_or_zero, - __entry->referenced, - __print_symbolic(__entry->status, SCAN_STATUS), - __entry->order) -); - -TRACE_EVENT(mm_collapse_huge_page_swapin, - - TP_PROTO(struct mm_struct *mm, int swapped_in, int referenced, int ret, - unsigned int order), - - TP_ARGS(mm, swapped_in, referenced, ret, order), - - TP_STRUCT__entry( - __field(struct mm_struct *, mm) - __field(int, swapped_in) - __field(int, referenced) - __field(int, ret) - __field(unsigned int, order) - ), - - TP_fast_assign( - __entry->mm =3D mm; - __entry->swapped_in =3D swapped_in; - __entry->referenced =3D referenced; - __entry->ret =3D ret; - __entry->order =3D order; - ), - - TP_printk("mm=3D%p, swapped_in=3D%d, referenced=3D%d, ret=3D%d, order=3D%= u", - __entry->mm, - __entry->swapped_in, - __entry->referenced, - __entry->ret, - __entry->order) -); - TRACE_EVENT(mm_collapse_scan_file, =20 TP_PROTO(struct mm_struct *mm, struct folio *folio, struct file *file, diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index 58f78274d989..22fc8f974be2 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -6703,8 +6703,6 @@ static const struct bpf_raw_tp_null_args raw_tp_null_= args[] =3D { /* host1x */ { "host1x_cdma_push_gather", 0x10000 }, /* huge_memory */ - { "mm_khugepaged_scan_pmd", 0x10 }, - { "mm_collapse_huge_page_isolate", 0x1 }, { "mm_collapse_scan_file", 0x10 }, { "mm_collapse_file", 0x10 }, /* kmem */ --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 555084052B1; Sun, 16 Aug 2026 22:47:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920457; cv=none; b=eNKU09nk5+HPnILW3q08TDOtA2WQrAPT65JtTXcrtIQ5m62OSPlR5mePqkwMVoYJbNlhgEJ5pw6Ku6tdPtwt4TjZGkiMA5xT2phAVm1FyHMpIQCaxcRGU4FWi0HqZhTtGbwbj0cva06FD4UzjbGuUVyl9db+Dqi6PRf/pPNAGPU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920457; c=relaxed/simple; bh=IO5aq+XYavi19arLYAl9mlbrdb1QIjdbqr57wOBffnE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uFExQT/b+Uok7/Jo19mIs99Jvq/mDd5wXCAFgzEJWkJm94a58u0PKdgV4xqTo001ZnMyvTx2jNYEPLpqAqNDi69+1Gtv6duS5EkmpQjfKWMmmhnrrx6ERGaw04DNfsm6s7BzFBB24u0JUoDMBmQtjeHsQfuxYj206C27Qb95vgg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=aUW9SjDQ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=SV/PYUiA; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="aUW9SjDQ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="SV/PYUiA" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 7D975EC0241; Sun, 16 Aug 2026 18:47:34 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:34 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920454; x= 1787006854; bh=rJV1RXsRkttc2jIpRgeymaLmVF3ajFQ4APp6aG4luFg=; b=a UW9SjDQJ9QSf8yY3eh4kelrOvX7FucErcgTcKgg/JVD6GDv+uNvPb3uo07P/+9sq egf/BaDfcMIUhLC6tMN8CgRdtfFgOtBllGsvlS/Yx3xHNkgHwKgMvLyl5Hrvjzb/ GLGHSuMLCjBsiu1UexoFS7Kv/925N0qBs8p9dvLEnfr7k2QYFWUyTqLaX/7ZGoGM 1uZeVEkYt961YFj8xjfho9ShO5TRQ9mmumxXixs/EZLPnjCrNJmFL5QqznnIv1kC 42rEU5rIbksjeLj6pu+4J3XwvDrUYTTacWANJg9E6E2Wz32Ii9qRZnoBOReFCJ71 XCYUziwLokks8nHCsAWoA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920454; x=1787006854; bh=r JV1RXsRkttc2jIpRgeymaLmVF3ajFQ4APp6aG4luFg=; b=SV/PYUiAhgqOTOdcn B11l+Sm41650eHu3CUJHPTB3l8zPecWbhAyzcwWMv6eTVS+INQ2SdniFfmBWJlOB frudIzxT/3dzI+811LhH7EU0/PsLLjqwDQUVUjiI+aMW/o5V65WIlDM6xYCyNBY6 84DdiEZN9rQs4+7US5dWvD76VYLIcjuWTzyQ4HX+U5Bj4UbP3LmW1NaVVEmX5oJF UGT+zK4SniVIumNnD8dSVRVD2Ak1QcSerB0d+DOaHypeeGMXzlOaQXMaVtP3NYCe pPcAQmTQ6QL9r2Txm1V7+efdLn9XgF5pwHUBbwGks6DF5scm5vBrVkZXYdzstJxB m812A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEs0CqHyYyBrtEDvFsHJsaqDA+PIZ0J3r+Lfgg7JQy+ArVExyWeVwVWITCUEI32DM 4Q07u0N7OV0zkPLdTdS1vdC5LFGL2haYN/6+YadFCAGEdEWDkiBi3rhYdwnM7rVYMLfpOZ IqudTgnad9TAcDUXw0QsBDvc+EbhJwbEI9dJrG3N8zHX8Hs4ScjPRdKN82RY2wyTeNKhqr YTzZzZs4Crypk/DQPLz6+eD8+5uxy24RPSI2P11jXgApnProibFo5ekQZWGTZ43B/F8fQP In8cSdmiaADGyUeORNkJ8RH7e9d6bLJ9AYQCJ2FX0Fia17GnwY/TqnulViI5crsw7X2mUA Nzwc/WiFcwI7tHKK7OGfdBPGyBDExQpnBrt3FnFaR7BTfgwwDb5QfBMizoi7mG8Bw/O8DT U6jSKyk/FWQho1hDU7lvfQNoGZzh+q/vCOLL2rl4DxuxcwdjsLfU8RLC/v49vD1DR6kA5u MGrQJ21Jg8P9hKumkyq+NXql9+UJMcN6A4FIBv73MaWE2m40+JFgdn++B+/HRa6FyG5mln sYZQzNby6qE4y2B+3RFT4/uXBfmE5JYvZ3HoQ3XT7IJPIv0m8T/G3JrCL0ih9WjP3XcsvY mN3JxakrOFQp4y0hkjiOH91KKCirGEtm3mWPzN1H+JaeQEQOdXcISKb3g68g X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:33 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 41/57] mm/collapse: give collapse its own trace header Date: Sun, 16 Aug 2026 23:45:53 +0100 Message-ID: <20260816224609.308019-42-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Every event in include/trace/events/huge_memory.h is a collapse event, and all but one is emitted from mm/collapse.c. The header is named for huge_memory.c, which emits none of them, and the tracepoints are built by khugepaged.c, which emits one. Rename it to trace/events/collapse.h, with TRACE_SYSTEM to match, and build the tracepoints in collapse.c. Two things follow. The pass names a candidate event prints are needed only where the tracepoints are built, so enum collapse_pass moves out of mm/collapse.h into collapse.c. And khugepaged.c becomes an ordinary includer, for the one event it does emit. The tracefs directory moves with the trace system: events/huge_memory becomes events/collapse, so anything enabling those events by system name has to follow. khugepaged's own selftest does that, and is updated here. So are the two other in-tree references to the old name: the MAINTAINERS entry for the header, and the group raw_tp_null_args[] keeps its collapse entries under. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- MAINTAINERS | 2 +- .../trace/events/{huge_memory.h =3D> collapse.h} | 8 ++++---- kernel/bpf/btf.c | 6 +++--- mm/collapse.c | 17 ++++++++++++++++- mm/collapse.h | 13 ------------- mm/khugepaged.c | 3 +-- .../selftests/mm/khugepaged_sync_check.c | 10 +++++----- tools/testing/selftests/mm/vm_util.c | 2 +- 8 files changed, 31 insertions(+), 30 deletions(-) rename include/trace/events/{huge_memory.h =3D> collapse.h} (98%) diff --git a/MAINTAINERS b/MAINTAINERS index 0f513b42bc18..7c179b333e4e 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -17273,7 +17273,7 @@ F: Documentation/ABI/testing/sysfs-kernel-mm-transp= arent-hugepage F: Documentation/admin-guide/mm/transhuge.rst F: include/linux/huge_mm.h F: include/linux/khugepaged.h -F: include/trace/events/huge_memory.h +F: include/trace/events/collapse.h F: mm/huge_memory.c F: mm/khugepaged.c F: mm/mm_slot.h diff --git a/include/trace/events/huge_memory.h b/include/trace/events/coll= apse.h similarity index 98% rename from include/trace/events/huge_memory.h rename to include/trace/events/collapse.h index 6aabf4235648..a3af5d8cc9aa 100644 --- a/include/trace/events/huge_memory.h +++ b/include/trace/events/collapse.h @@ -1,9 +1,9 @@ /* SPDX-License-Identifier: GPL-2.0 */ #undef TRACE_SYSTEM -#define TRACE_SYSTEM huge_memory +#define TRACE_SYSTEM collapse =20 -#if !defined(__HUGE_MEMORY_H) || defined(TRACE_HEADER_MULTI_READ) -#define __HUGE_MEMORY_H +#if !defined(_TRACE_COLLAPSE_H) || defined(TRACE_HEADER_MULTI_READ) +#define _TRACE_COLLAPSE_H =20 #include =20 @@ -282,5 +282,5 @@ TRACE_EVENT(mm_khugepaged_scan, __entry->full_scan_finished) ); =20 -#endif /* __HUGE_MEMORY_H */ +#endif /* _TRACE_COLLAPSE_H */ #include diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index 22fc8f974be2..505bdbb534e4 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -6683,6 +6683,9 @@ static const struct bpf_raw_tp_null_args raw_tp_null_= args[] =3D { { "cachefiles_ondemand_cread", 0x1 }, { "cachefiles_ondemand_fd_write", 0x1 }, { "cachefiles_ondemand_fd_release", 0x1 }, + /* collapse */ + { "mm_collapse_scan_file", 0x10 }, + { "mm_collapse_file", 0x10 }, /* ext4, from ext4__mballoc event class */ { "ext4_mballoc_discard", 0x10 }, { "ext4_mballoc_free", 0x10 }, @@ -6702,9 +6705,6 @@ static const struct bpf_raw_tp_null_args raw_tp_null_= args[] =3D { { "time_out_leases", 0x10 }, /* host1x */ { "host1x_cdma_push_gather", 0x10000 }, - /* huge_memory */ - { "mm_collapse_scan_file", 0x10 }, - { "mm_collapse_file", 0x10 }, /* kmem */ { "mm_page_alloc", 0x1 }, { "mm_page_pcpu_drain", 0x1 }, diff --git a/mm/collapse.c b/mm/collapse.c index 6c17c83a4e21..952e3f62d920 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -26,10 +26,25 @@ #include =20 #include -#include #include "collapse.h" #include "internal.h" =20 +/* + * Which pass of a round reached a verdict on a candidate. Named by the t= race + * header, which only this file builds, so it need not be shared. + */ +enum collapse_pass { + COLLAPSE_PASS_ALLOC, + COLLAPSE_PASS_REVALIDATE, + COLLAPSE_PASS_FAULTIN, + COLLAPSE_PASS_FREEZE, + COLLAPSE_PASS_COPY, + COLLAPSE_PASS_INSTALL, +}; + +#define CREATE_TRACE_POINTS +#include + /* * Anonymous collapse, in rounds. * diff --git a/mm/collapse.h b/mm/collapse.h index 4af7bb9c4261..9e2cec1f250b 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -16,19 +16,6 @@ struct collapse_candidate; struct collapse_retry; =20 -/* - * Which pass of a round reached a verdict on a candidate. Only collapse.c - * produces these; the trace header khugepaged.c builds names them. - */ -enum collapse_pass { - COLLAPSE_PASS_ALLOC, - COLLAPSE_PASS_REVALIDATE, - COLLAPSE_PASS_FAULTIN, - COLLAPSE_PASS_FREEZE, - COLLAPSE_PASS_COPY, - COLLAPSE_PASS_INSTALL, -}; - enum scan_result { SCAN_FAIL, SCAN_SUCCEED, diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 967cc472b6dc..f3a7aad5e8f2 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -31,8 +31,7 @@ #include "page_alloc.h" #include "mm_slot.h" =20 -#define CREATE_TRACE_POINTS -#include +#include =20 static struct task_struct *khugepaged_thread __read_mostly; static DEFINE_MUTEX(khugepaged_mutex); diff --git a/tools/testing/selftests/mm/khugepaged_sync_check.c b/tools/tes= ting/selftests/mm/khugepaged_sync_check.c index 45001996b57a..4c37b697d3dd 100644 --- a/tools/testing/selftests/mm/khugepaged_sync_check.c +++ b/tools/testing/selftests/mm/khugepaged_sync_check.c @@ -41,7 +41,7 @@ static unsigned long hpage_pmd_size; /* * Each step switches the events off again, but a helper can still give up * on us in between (a failing sysfs write ends the test from inside - * thp_write_num()), and huge_memory events left on are the whole machine's + * thp_write_num()), and collapse events left on are the whole machine's * problem, not this test's. */ static void trace_events_off(void) @@ -118,7 +118,7 @@ static void one_step(int iteration) if (tracing_clear_trace()) ksft_exit_fail_msg("Cannot clear the trace buffer\n"); if (tracing_events_enable(trace_events_fd, true)) - ksft_exit_fail_msg("Cannot enable huge_memory events\n"); + ksft_exit_fail_msg("Cannot enable collapse events\n"); =20 if (madvise(p, hpage_pmd_size, MADV_HUGEPAGE)) ksft_exit_fail_perror("madvise(MADV_HUGEPAGE)"); @@ -127,7 +127,7 @@ static void one_step(int iteration) =20 /* Off before anything that can give up: the events are system-wide. */ if (tracing_events_enable(trace_events_fd, false)) - ksft_exit_fail_msg("Cannot disable huge_memory events\n"); + ksft_exit_fail_msg("Cannot disable collapse events\n"); if (!passed) ksft_exit_fail_msg("khugepaged did not complete a full pass\n"); =20 @@ -164,9 +164,9 @@ int main(void) kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); if (kpageflags_fd < 0) ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); - trace_events_fd =3D tracing_events_open("huge_memory"); + trace_events_fd =3D tracing_events_open("collapse"); if (trace_events_fd < 0) - ksft_exit_skip("huge_memory events require tracefs and root\n"); + ksft_exit_skip("collapse events require tracefs and root\n"); atexit(trace_events_off); =20 ksft_set_plan(NR_ITERATIONS); diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index ee1334778391..32cba59ae49c 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -601,7 +601,7 @@ bool is_range_backed_by_folio_orders(char *start, size_= t len, int order, #define TRACEFS_ROOT "/sys/kernel/tracing" =20 /* - * Open the enable file of one ftrace event subsystem (e.g. "huge_memory"). + * Open the enable file of one ftrace event subsystem (e.g. "collapse"). * Returns a descriptor for tracing_events_enable(), or -1 if tracefs or t= he * subsystem is not there. The events are system-wide state: whoever * switches them on owns them until it switches them off, including on the --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3C17F40911F; Sun, 16 Aug 2026 22:47:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920458; cv=none; b=c4KyG5AmXjXuPohwVFla3RU/DEPAHRQTJw7q6aelnmbsFyoMZsA29ahQsIOO6V9FcDRJ0Krq3uE4w3ywCEdGNyvGLqis9UKxTKAK1NHKvp55Mtd3hKfVvfka2BBz99NLpG4HI6ZQ6B3S8OCydjzb0tfKn9LJLQ07+xKfMjpYQ4Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920458; c=relaxed/simple; bh=QED0ivctn/g84zHlwg3dHaWIpmZRCAVoxxfhuvqzEtM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pZ946IOl7s0PPpkoPHQf5odbxN1egVfGcjrUzOLwDlMZdcIM2hq+dTYsuPhMTwpwGEQ0dNHwJm9l5YzJVC+/ZyUBc+llsnFCGUI+mekuNw5B7flbO9X+DT+qJfXntqoLf7SRW5uMVX2C10wwOD2XVSIXtO/ONPjmUZUEBKGNsms= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=nG9tow1d; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=BOnDi6ps; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="nG9tow1d"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="BOnDi6ps" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfhigh.phl.internal (Postfix) with ESMTP id 5785714000F8; Sun, 16 Aug 2026 18:47:36 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-03.internal (MEProxy); Sun, 16 Aug 2026 18:47:36 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920456; x= 1787006856; bh=UyrP/wXPOBJTdhoFE5YGqbURIEa1HSuJzV4RfAXPajU=; b=n G9tow1dkQz77VjVwwF3jcgHCENpHOmT8atxbCLLECcSiWMcmK6RFYtDpUvqxkL9w CDUp/5852TW8lhwrZBw4LnSUM49QS+ubiCBuP/Oc7zR+MnT7KO093rX60nzUsC8k WF259W3rgsUCWMJnwYpO7w3ilcTCjCDYsbN844W1SHCtSPPG2qstrHMDxqNHsLlJ H4Gkc9DevfM9JdR2snF+2PuW8JDFOhRAEBUlmzRJd7C3w7d5o8CPVFsE1fyzWzBQ rHcAlP72IyQuA3Oa18FluUCvGMft6PBoqd3uurqhboVAcG32GFWXZl/UT3Umx44e sLOTkOswIKKnPFy5hb0FA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920456; x=1787006856; bh=U yrP/wXPOBJTdhoFE5YGqbURIEa1HSuJzV4RfAXPajU=; b=BOnDi6psLSSwsXcqy wmRIZYHNoPw/nNbEwui6al4nbqSnxczcCAJHo4343mGOYCf6oK2m18CURKPQZoZs xxZjGtrvx3UEn4n7ABQXcJInQWQkREL/pkyU5g+S7F8sPN341QWSiUWETqSUJ8at iXfU/+DnEsUhOfdIcMKtOJUrlBSisODgaAJdTHLCHvsk9q7M9gqUQvHcTf7waYn7 o++nyGT4JgrTC538YZDVwugzrO0mxcS23JtdrqvkLoBlRIxbWz9SGDhWLHGfAdDQ FDs4Ee5ybxfGyhj4wF+/cg7InCmu16q9iKiW+4yBOQ4G5eYrHMPjDj6dLAXkpJpe fcJuQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFcTQrb7i2Wnr2aEeZt1aoD7asKutBPdJgU3YSD5OfcfuljRo8Y5zG4WAng55KhcU fdYu9dW0n+wtsUDWiYHh8AwjWiiyOeVLRvzra0gUIzgKMFXfu2OsH3wUsDwz3Dpe7vEjB9 JjBmMfD78Db5X6Vl1RDz8Wz4IQF1tjtPvs8yYIcZ4vdtxAZhotrvrKr9o/e/MWe2qw9p1d u/4ngMViLbtWn8zpPvYSIgtjaF8APmxmHCU5Ga5Rd5eIGcQp1MbbGmkchmHrMFMoUd9mrv TpVQ9+TKNOdfLLUK+2CPoHpagx885EUQUSZxhSq0AxxU5QJRQ+JA8yzxyjEYMjW1uFFjcm d7hzQcIDGBAJ9Kg3rplLw7jfe4LEi0F6SofMNmrKPLYc4yjweROnOFdgG3+P+anRpdNgju yFWVmmTGSJ1pkA7ZfQP/zLvd+Qc57RLKd2grwoOvBrCoYfMf2/7JtIEgRKykbeQlJl1t2i o3a7q3cY2oo04ey7M5qY21DAFRN0hvbmP917wCqhCug8/gv9jHuXVVRBq1P4rS/xZjYMWZ +SfkbQhdbebpKUDV9jBzsvGf+1KfqEwOb5sbvgleSn3QvacH+qZWnvUYuTftewxhZ+GZdQ 8ctAXFBvI8gELPNX5p5dipgHyieJ7AirBkuIwt+OgA/k+nC52hzYouNJCAlA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:35 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 42/57] mm/collapse: allow error injection into the freeze Date: Sun, 16 Aug 2026 23:45:54 +0100 Message-ID: <20260816224609.308019-43-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The freeze refuses a candidate for reasons a test cannot arrange: a co-mapper appearing between the two sweeps, a GUP-fast pin, writeback. What the round does then -- the candidate drops out, its neighbours carry on, the region goes back to selection at a lower order -- is worth exercising on purpose. Tag collapse_freeze_candidate(), noinline so fail_function can find the symbol. ERRNO is the only injection type for a function that returns a value rather than a pointer or a bool, so an injected result is a negative errno. It only ever reaches cand->result, where every caller compares it against SCAN_SUCCEED and finds it unequal, which is the refusal the test wants. Tracing prints it as a number, having no symbol for it. A forced return happens at function entry, so this reaches the refusal, not the freeze's own unwind path. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/mm/collapse.c b/mm/collapse.c index 952e3f62d920..76a53616b240 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -4,6 +4,7 @@ #include #include #include +#include #include #include #include /* x86 flush_tlb_range() uses hstate_vma() */ @@ -937,7 +938,7 @@ static enum scan_result collapse_check_candidate(struct= vm_area_struct *vma, * collapse_freeze() issues the ranged TLB flush over everything that froze * before dropping the ptl. No copy may run before it completes. */ -static enum scan_result collapse_freeze_candidate(struct mm_struct *mm, +static noinline enum scan_result collapse_freeze_candidate(struct mm_struc= t *mm, struct collapse_candidate *cand, pte_t *pte) { const unsigned int nr_pages =3D candidate_nr_pages(cand); @@ -1085,6 +1086,7 @@ static enum scan_result collapse_freeze_candidate(str= uct mm_struct *mm, collapse_unfreeze_candidate(mm, cand, pte, nr_saved, nr_frozen); return result; } +ALLOW_ERROR_INJECTION(collapse_freeze_candidate, ERRNO); =20 /* * Raise the two barriers on the sources of every candidate: migration ent= ries in --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A17DF3E6DEC; Sun, 16 Aug 2026 22:47:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920460; cv=none; b=rEhubbYycPXIUfMwaUZoitqqpWoTLz9hJ47z4rYTkTVJrF6FctpAiCnRghaK4vb+EG9d08oyMHxEyCwL52U0HzHIfKVPB6NhRUVPEO5c0aBFhfqQxfkZIiAK0jLCsEqKGg/lmo8vp50lj1/Wz8jfR5BjJ463VuRpRm1XVnRqUH0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920460; c=relaxed/simple; bh=lG+OMVFM3ZFKjzFKQUYMMsSJ6jytNMBH6dhA8jU2SEA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=EeMMbdebbLVCJlVpA+TWPkTKnT/uFT2q6ySQl6SiHR/2pPDu3N5dWu1uoTu+GNtnbQ9J+S8CqQIUH051CdUPUTgjYl4r29lFY8nCZXQXRTnXKrant1f8Uar7L2PfaKXR9EQylhTPhnqArINZuWwYl2yw5kTaXWB7OqIiu8EtjSE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=fu2iY+kc; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=QrkuiCXN; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="fu2iY+kc"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="QrkuiCXN" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 0A975EC0235; Sun, 16 Aug 2026 18:47:38 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:38 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920458; x= 1787006858; bh=hsbtWGmnLS8gNA6v/OmuTl4g7B7VvEJrlwoMT5rjl2I=; b=f u2iY+kcbdqg8XZ7qkFdaWgez1i5hIWQAqWtxE4ir7o14H6QTMAJOsoArRisr8+zF XKgXHadiin9jdcqGoR2OnqoC40c2jyr/XvrAdX2LfRqrDWd3bLnUQTkcfD6EHDVj D4I7m+ZK9rL2kJWu66YtJqYb1oLcF25VQrF9+2dad/xkK71QaojHbSgoznDIOpBW cIkyw02O9OYt0r81b5abH4orCiBcMOCblMYhFVVN8xX2Akp1dVWilQFjS77rI1Mj 66kS42c5eYiXYqmJvRekO+PxrO/rzkJDTG11PqQDWJIswNS0G21VmTRLAStxJZPL PykDV2ghQyRM18uwklarQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920458; x=1787006858; bh=h sbtWGmnLS8gNA6v/OmuTl4g7B7VvEJrlwoMT5rjl2I=; b=QrkuiCXNu9Ag8fKym 7s3tQryl5qs35v3JLJkB8cW2oZM5UeI+B6Y11iiR9Zp5PloR6bNbcTcicmuTnw8J lVP4mCH2TikjEdOjGX1x5Zr5UhMkUfHXOFrFEDpSIO2EZveoX6nS27vjEq+Vy4Px k7DS7Zz/TeZ1h5JsO5TG/3bm5dQdq3v2CTlWHQ3BWl/Jg2ZnVWC+CN1FU3rFPxxq JWnzKxdhJWbGRXRwVewUho2MIIaXdEAY6t7E4nHBipdGWCl5uBesALz3spA3/8Bg vcB6C9sMvd9e7HA0J840icixCdVA0ZoFl6RbinUEjij2zfrZ4Q5cFgIMg4DZkdkC DasBg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFxktHaoHQCGyR7uz2gXQdTgUX0nZzUatWYXQdVxBJb4A/qMuEZ/RMKYvGIBz31Rt TR5Nu17EPBqNyeZSjkVnK7sHElwU57b8x3/cyZ9PHHSfAaSWNWLOdeLuZVBbf5zubb18Bw fy/mf5jpQtxpZDPvMPXwjw+an+TW+7k/Tw/n4sb2tNNa7d9nLnQnX6Rw+vp6wioFQOuuev fpJbCZJSwa4xWiqXHUsCfuVr3o0xcsLHyTasHEYEkAGMtBoAwl7gOHcd9Yu3HcCvuVHzqL dMVburb2rGaTCsZcGe3UWITClIJ34iPVxiKzwGBtfazzXdWK1ezlWroYFXdEl+xlvbcPMN vZtr4rHndVJB562/WoDkew1Y03j21tKV2cnjL43Ck1uOBZ8NRyVlM3puHQQJb8vttWvbCT Vub/4J6dev06TTYjDujIcAGui6NrZT7E0xqFxmiu9BmjvwUAj0YhVXNi3R3McxnqEuozCs Cg+X7m9F7uqX2U6p6Xyjjpfd5s6+mLEaC2r7NYfvnpAv6qt2q0FyIdVJus0Zu4lfXxl41X Ho6g80JFqnXt8OaS01tjl26CBO9wiBFx/ns5VSZG0hA3WGnD4j3/8rFwsQ9O4oX5Shk/Hx K0jzuV1jZV8vH/x8VLuU4iFg4PHzM1/90qd2/8TgEM/6D9aaOPm8UyuE0hvA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:37 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 43/57] mm/khugepaged: check the scan budget before the work, not after Date: Sun, 16 Aug 2026 23:45:55 +0100 Message-ID: <20260816224609.308019-44-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" pages_to_scan is meant to bound what one khugepaged pass does, but collapse_scan_mm_slot() tested it in only one place: after a table had been scanned and turned out to hold nothing. Neither of the other two ways of spending the budget reached that test. A VMA the pass skips is charged for and walked past without asking -- one no order can be collapsed at, or one the cursor is already past the end of. A table that does hold a candidate leaves through the collapse. So a pass over an address space of thousands of VMAs khugepaged cannot use walks every one of them, however low pages_to_scan is set. Ask at the top of both loops instead, where the other reasons to stop a pass are already asked. The outer loop asks before it judges a VMA, the inner one before it scans a table. Stopping the outer loop only works if the cursor moves, and it did not for a skipped VMA. Advance khugepaged_scan.address past one, so a pass that runs out of budget resumes after the VMAs it has already judged. Without that, an address space with more skippable VMAs than the budget would be walked from the same place every pass and never scanned at all. pages_to_scan now bounds a pass that finds nothing to collapse, where before it did not. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/khugepaged.c | 16 +++++++++++++--- 1 file changed, 13 insertions(+), 3 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index f3a7aad5e8f2..cc5ff429d811 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -553,9 +553,19 @@ static void collapse_scan_mm_slot(unsigned int progres= s_max, cc->progress++; break; } + + /* + * Before the VMA is judged, so that a pass over an address space + * of VMAs it skips is bounded by the budget too: each one is + * charged for, and none of them was being asked to be scanned. + */ + if (cc->progress >=3D progress_max) + break; + orders =3D collapse_possible_orders(vma, vma->vm_flags, TVA_KHUGEPAGED); if (!orders) { + khugepaged_scan.address =3D vma->vm_end; cc->progress++; continue; } @@ -570,6 +580,7 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, hstart =3D ALIGN(vma->vm_start, window); hend =3D ALIGN_DOWN(vma->vm_end, window); if (khugepaged_scan.address > hend) { + khugepaged_scan.address =3D vma->vm_end; cc->progress++; continue; } @@ -584,7 +595,8 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, range_end =3D min(hend, pmd_addr + HPAGE_PMD_SIZE); =20 cond_resched(); - if (unlikely(collapse_test_exit_or_disable(mm))) + if (unlikely(collapse_test_exit_or_disable(mm)) || + cc->progress >=3D progress_max) goto breakouterloop; =20 VM_WARN_ON_ONCE(khugepaged_scan.address < hstart); @@ -596,8 +608,6 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, /* If nothing to collapse, the lock is still ours */ if (!collapse_scan_pmd(vma, start, range_end, cc, orders)) { *result =3D cc->scan_refusal; - if (cc->progress >=3D progress_max) - goto breakouterloop; continue; } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4456140B38A; Sun, 16 Aug 2026 22:47:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920462; cv=none; b=k2WSji7Gau60y1N/oDO+/iiCsiaMi+BonerAsv52dazDAapSCsH8LWjPmMW99rTA3otsNdCJBpdR8UzXZEm5CbZUv005FOx7laQPG+jOwF7yh8y0fESi/WHcjQCe5YNS+uAC/fwveBt16tfqNOAHd3m617HP5HMnjZL6bt+GKNo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920462; c=relaxed/simple; bh=z4EXwj9wCEp37cQMY9/xOXa9YkGApZgoarJNA/g3m5o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PHpBVyfpl3aXp0vWztkz4jaA/MW0MOpET8TZ4MEOGhiJWICZ6/3OEFVh4YINStkTrYWqA8uNG7t1nS4yxKMhrp1XH5c9DP2rx6ReN6cayTW/eBMSTeaz8B/P9oVCxSP+dYX947FRKgD/xe5gYgpemsiBCKvCxQdMKLPs61QcOlg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=So8MRp2F; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=WqwC9yGZ; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="So8MRp2F"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="WqwC9yGZ" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id A0F0414000EE; Sun, 16 Aug 2026 18:47:39 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:39 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920459; x= 1787006859; bh=rve110aQ3RGnUPR/SGf577VtfeusB/RC8s8suU08YA0=; b=S o8MRp2F4cOuq2Egj9t90m7zBXPLSOh4s2sJaWpbNCEsDpxDH5niH9wmr309FDA78 1OzaTFbiZFp1h8NrcnJqTBlMadRuMzbQ1kPvkYVQRAml8XfPbMkZ9nktmDOL8ME0 wnn1cQXtTsTsT8knZLP04tQvy1fZdPgYCyP2eDgs/r35BTKg9//jXwYYZXP9waEl rfeI+G3yL7oHE5jvUCYmMNMprAESzri9dKTAaMBByxsyLc5cg5ruWQ1YPU3sHV7s 0ET7jvgag123GHU2HHi72b2NbT4ae26E+ne335R05JPqgaZUy7EsR60SzBZtUgDY UriCvMdhkxPvjwXwFBmZw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920459; x=1787006859; bh=r ve110aQ3RGnUPR/SGf577VtfeusB/RC8s8suU08YA0=; b=WqwC9yGZLz5sX585z qfhZev1SIyDpPNU4cAMc3llRXuREEPAnm0z2vOkFa9aeG2QGBZ2yHpRV/ynjoxLY XDbJuKnzKQXNOEkTHGqOcRTzOau5jyeZj0M8UDQy0rPTIOYC3fT19PG2PmnZtKrh wxvFz4S45pqbmAnfcT5g6r3JOW8KVp+MgM7oe11MUl/Cn0LDYV7YLOLihBhlG6Gb OvUOi2yZEtHoI4VES/ClhjVVLgSt8b6SpFX7nSoEBzwLszdR3AfDkyh14SCcOwDy IQ3gXxc7IxtTISx4Wz9MNFFCRS4IJBPhewm13UZ4w8dGFtUGdMcL8zK36q/iXr+N LZAgQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFxktHaoHQCGyR7uz2gXQdTgUX0nZzUatWYXQdVxBJb4A/qMuEZ/RMKYvGIBz31Rt TR5Nu17EPBqNyeZSjkVnK7sHElwU57b8x3/cyZ9PHHSfAaSWNWLOdeLuZVBbf5zubb18Bw fy/mf5jpQtxpZDPvMPXwjw+an+TW+7k/Tw/n4sb2tNNa7d9nLnQnX6Rw+vp6wioFQOuuev fpJbCZJSwa4xWiqXHUsCfuVr3o0xcsLHyTasHEYEkAGMtBoAwl7gOHcd9Yu3HcCvuVHzqL dMVburb2rGaTCsZcGe3UWITClIJ34iPVxiKzwGBtfazzXdWK1ezlWroYFXdEl+xlvbcPRE EcGSAauGtulKcfHD8ZMEdZcshXjw82HKplkBk4CvuYSebHShQKOtVqcQB5KLfW6zdhaS3a mbuLEIL7G4MY9vEby2aIRAMWWkEF1kzHPkEQkq/lnTs1/5ZDVFuLhKaVNJcmrgM2fUTNkB shFf0eOpWUoU11ZJeeiE5nU9Jt3xDrIEhEMqYNU4bSEQoZ+QrLRVRj7ONEhBcvCVmgucac BfN63HDXEl38M4gvW2+m8SBIw1tIkubTmLQrx+CsHTLU0oP5oubte9XGIc9aAb78AQbkG0 3X8/DFAOIVARw+t/zbAJiZ3xPnE62YxRdeKmqylaoL7xZz3mHhxjUn6SNXdA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:39 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 44/57] mm/khugepaged: hold the address space open across a scan Date: Sun, 16 Aug 2026 23:45:56 +0100 Message-ID: <20260816224609.308019-45-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Preparation for taking a per-VMA read lock instead of mmap_lock. What tells khugepaged an address space is going away is the barrier in __khugepaged_exit(): it runs before exit_mmap() and takes mmap_lock for writing, which waits for a scan holding it for reading. A scan under a per-VMA lock holds no mmap_lock, so nothing waits for it and exit_mmap() frees the page tables it is walking. Take a reference on mm_users for the pass instead. __mmput() cannot start while one is held, so neither can exit_mmap(), whatever lock the pass uses. Drop it with mmput_async(), so the last reference does not tear an address space down inside khugepaged. Drop it before the exiting mm is judged, too: that judgement needs the true count to release the slot. The reference is also what the exiting-mm checks were reading, so an address space whose owner has gone now shows as one reference rather than none. The three checks inside the pass ask collapse_test_exit_mmref() instead; the slot-release judgement keeps the old test, running after the reference is dropped. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.h | 19 +++++++++++++++++-- mm/khugepaged.c | 31 +++++++++++++++++++++++++++---- 2 files changed, 44 insertions(+), 6 deletions(-) diff --git a/mm/collapse.h b/mm/collapse.h index 9e2cec1f250b..74e513c5c76c 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -170,6 +170,11 @@ struct collapse_control { pte_t *saved_ptes; }; =20 +static inline int collapse_disabled(struct mm_struct *mm) +{ + return mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); +} + static inline int collapse_test_exit(struct mm_struct *mm) { return atomic_read(&mm->mm_users) =3D=3D 0; @@ -177,8 +182,18 @@ static inline int collapse_test_exit(struct mm_struct = *mm) =20 static inline int collapse_test_exit_or_disable(struct mm_struct *mm) { - return collapse_test_exit(mm) || - mm_flags_test(MMF_DISABLE_THP_COMPLETELY, mm); + return collapse_test_exit(mm) || collapse_disabled(mm); +} + +/* The owner has gone: the caller's own reference is the only one left */ +static inline int collapse_test_exit_mmref(struct mm_struct *mm) +{ + return atomic_read(&mm->mm_users) =3D=3D 1; +} + +static inline int collapse_test_exit_or_disable_mmref(struct mm_struct *mm) +{ + return collapse_test_exit_mmref(mm) || collapse_disabled(mm); } =20 /* diff --git a/mm/khugepaged.c b/mm/khugepaged.c index cc5ff429d811..f3ea1846990e 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -531,16 +531,31 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, spin_unlock(&khugepaged_mm_lock); =20 mm =3D slot->mm; + vma =3D NULL; + + /* + * A reference on mm_users for as long as the pass works on this address + * space. __mmput() cannot start while one is held, so neither can + * exit_mmap(), and the VMAs and page tables stay where they are. + * + * Once per pass, not once per table: the reference is what makes the + * address space safe to work on, and a pass is how long that is wanted + * for. Nothing else in mm takes it per unit of work -- DAMON takes one + * per target and walks every region under it, swapoff one per mm across + * the whole address space, userfaultfd one per call. + */ + if (!mmget_not_zero(mm)) + goto breakouterloop_no_mmput; + /* * Don't wait for semaphore (to avoid long wait times). Just move to * the next mm on the list. */ - vma =3D NULL; if (unlikely(!mmap_read_trylock(mm))) goto breakouterloop_mmap_lock; =20 cc->progress++; - if (unlikely(collapse_test_exit_or_disable(mm))) + if (unlikely(collapse_test_exit_or_disable_mmref(mm))) goto breakouterloop; =20 vma_iter_init(&vmi, mm, khugepaged_scan.address); @@ -549,7 +564,7 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, unsigned long orders; =20 cond_resched(); - if (unlikely(collapse_test_exit_or_disable(mm))) { + if (unlikely(collapse_test_exit_or_disable_mmref(mm))) { cc->progress++; break; } @@ -595,7 +610,7 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, range_end =3D min(hend, pmd_addr + HPAGE_PMD_SIZE); =20 cond_resched(); - if (unlikely(collapse_test_exit_or_disable(mm)) || + if (unlikely(collapse_test_exit_or_disable_mmref(mm)) || cc->progress >=3D progress_max) goto breakouterloop; =20 @@ -622,6 +637,14 @@ static void collapse_scan_mm_slot(unsigned int progres= s_max, breakouterloop: mmap_read_unlock(mm); /* exit_mmap will destroy ptes after this */ breakouterloop_mmap_lock: + /* + * Not mmput(): the last reference would run exit_mmap() here, and + * khugepaged is not the thread that should tear an address space down. + * Dropped before the exiting mm is judged below, so that judgement still + * sees the true count. + */ + mmput_async(mm); +breakouterloop_no_mmput: =20 spin_lock(&khugepaged_mm_lock); VM_BUG_ON(khugepaged_scan.mm_slot !=3D slot); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 084FA3E63B0; Sun, 16 Aug 2026 22:47:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920464; cv=none; b=WWeskdCDsqWbxXl01EcJH1kBD0zO7bJ/FrNrkNng8Ss4xECRsNNvFNES4oTJk3tw6t42iU+DuDYABZSGndAEu2YdUO90Bfe9Lw2sVJR7svrDspQYhfsjGEW40DITc1p7VE08qRfpy04kfFqrTsK+aCUjOab51wzS3tdd2U9WS3w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920464; c=relaxed/simple; bh=YJNac9oYJhGUaLq234UkKNFFUy+hEfprOURyQoJSfGY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ehIDwTY7UtBGtrY201oamYZWLnlUXAR5toke3lT+jBsaH0y0qZDVEtb6qUaz/H4CYBimQaW7NOBapgX1fNFwiZEs8oAD984+DBiIx9DeJH+ADaIQm0c3T8TmhnxHPInupHJhBFdsjgIfCFy85+jvalzxNW2QZ3sNWh+yo26JE6E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=KxxlfdpL; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=HNhI9M71; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="KxxlfdpL"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="HNhI9M71" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 54C8F14000FA; Sun, 16 Aug 2026 18:47:41 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:41 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920461; x= 1787006861; bh=za2Lpztq5iqcTX4ankoznbvISTw47a74bIKVghH2CR4=; b=K xxlfdpLvgJbGk4rJLZ+Cu+xz0wYL/UBYzRzK0xs5M3gPCKjQvfr1XregUI4dLyaJ JOGJHQSXP207jlhG6Zfe2vHlEWfSGbGP+PzIVNQWZcR5F0DdIoXzZp162HVmS9yL 5yGYd8QXngpAZZAXd+4XT4zuOGBQKze0c3PEfVpiJZ7dmKllV6CBGl4QVR/AWUQu 0KyNVuB8fZ/EI3t5z6UoyYcnTUVPzlKo+dnZUWaV9gyqQahUjvLLKQJGXG5hlg0V 0PqpmucS3m7VrgxF6BsVut/zin+6xNehEM0TlFcqX4sD43RIi6kY2uSsK9CUl/O1 NZW6vo+QPsCSLItTVJB4g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920461; x=1787006861; bh=z a2Lpztq5iqcTX4ankoznbvISTw47a74bIKVghH2CR4=; b=HNhI9M71s+Ygqkqwf tzm19iEIk+cG7lE1f+l5XiTXLFPAK17mb4LAOO3ru8tNF2WL6Ilh1TYgsrFMQ+An qHJ3Qk7U68J+JnBv+4ftfCQnjgK2asvYC2JajkJwtuzgsmsOcIFjA+348tMUpeXO kf2p9jXxVBycdSc42l8RMuOfcWC1HfypcoKK5b8sXGzMiQeqmlpKVRPz7g9Ve7Fm mLoeiKq7eo6ToEzZekyQUN3DxOuYCS6XMFQOGucpdFEDR6Q4eSdZgjALi2c1AF8D xZG+se2qdVYevjKlad9zTjekUReUQsVSFjO4RX8b6if+qXsy+uQcz8lXhFIipXw1 faS6g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFxktHaoHQCGyR7uz2gXQdTgUX0nZzUatWYXQdVxBJb4A/qMuEZ/RMKYvGIBz31Rt TR5Nu17EPBqNyeZSjkVnK7sHElwU57b8x3/cyZ9PHHSfAaSWNWLOdeLuZVBbf5zubb18Bw fy/mf5jpQtxpZDPvMPXwjw+an+TW+7k/Tw/n4sb2tNNa7d9nLnQnX6Rw+vp6wioFQOuuev fpJbCZJSwa4xWiqXHUsCfuVr3o0xcsLHyTasHEYEkAGMtBoAwl7gOHcd9Yu3HcCvuVHzqL dMVburb2rGaTCsZcGe3UWITClIJ34iPVxiKzwGBtfazzXdWK1ezlWroYFXdEl+xlvbcPUF k970nTfEUWVpm0nF0GrdJhNoaD+sHupr3s1dISr8zMy52SPkaDGuLGQgwHwwOn5/OWjE7o iI9kR6MagxLIHHsQAOl+POPwFYCHu912UCVS+gZz4vwQZCuXJllT5y6lqgIpCHJTfZ2IXr 2YYKg6bigb/+ehwZOUN6Q0dVru3sbYzx3M9vxSzXPmrpcxVmT9IDxjuBhNZ7lPGDXel82J aKkrEmSyBN2ASatCWqc3p9ztaVXMSpaLwvia8qeuOuljKZPXCFZq1jVjgkiP92CKnw3SJ2 QrooxykOtBW4d0X/pcHN0dYnFZHWRMMnfBDOeV3njcXtRzkyCJoSSkb3PLfg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:40 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 45/57] mm/collapse: take a per-VMA read lock for the round Date: Sun, 16 Aug 2026 23:45:57 +0100 Message-ID: <20260816224609.308019-46-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A round takes mmap_lock for reading and holds it across the freeze, the copy and the install. It is a read lock, so it blocks no faults -- but every writer to the address space waits behind it, wherever in the mm that writer is working. A round works inside one VMA. Everything it has to be excluded from -- fork, split, merge, unmap, and the free_pgtables() that follows them -- takes vma_start_write() on the VMA it touches first. So a read lock on that VMA excludes exactly what an mmap_read excluded, and an mmap_write elsewhere in the mm stops waiting for a collapse it has nothing to do with. Look the VMA up with lock_vma_under_rcu() per round. The round therefore no longer works on the VMA the scan was given: it may be a different one at that address, or the same one shrunk. collapse_revalidate() already re-asks every question the scan asked, and now also checks that the VMA still covers what the round collected. That lookup also fails on a VMA that is merely being written to, because vma_start_read() fails while a writer holds vma_start_write(). A caller cannot tell that from a VMA that has gone, so report SCAN_VMA_LOCK for both, which MADV_COLLAPSE turns into -EAGAIN. Treating it as a range that shrank would report success for a collapse that never happened. The fault-in pass gives up the same lock it was called under, so its unlocks move with it, and the comments that carry the exclusion argument name the lock they argue from. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/trace/events/collapse.h | 1 + mm/collapse.c | 77 +++++++++++++++++++-------------- mm/collapse.h | 1 + mm/madvise.c | 1 + 4 files changed, 48 insertions(+), 32 deletions(-) diff --git a/include/trace/events/collapse.h b/include/trace/events/collaps= e.h index a3af5d8cc9aa..591030973d66 100644 --- a/include/trace/events/collapse.h +++ b/include/trace/events/collapse.h @@ -30,6 +30,7 @@ EM( SCAN_PAGE_COMPOUND, "page_compound") \ EM( SCAN_ANY_PROCESS, "no_process_for_page") \ EM( SCAN_VMA_NULL, "vma_null") \ + EM( SCAN_VMA_LOCK, "vma_not_lockable") \ EM( SCAN_VMA_CHECK, "vma_check_failed") \ EM( SCAN_ADDRESS_RANGE, "not_suitable_address_range") \ EM( SCAN_DEL_PAGE_LRU, "could_not_delete_page_from_lru")\ diff --git a/mm/collapse.c b/mm/collapse.c index 76a53616b240..b6e92e24a3c5 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -52,7 +52,8 @@ enum collapse_pass { * The folios mapped across a window of PTEs become one folio of that wind= ow's * order, with the sources quiesced by the two barriers migration uses -- * migration entries in their PTEs, then a frozen refcount -- so the copy = itself - * needs no lock. The engine runs under mmap_read throughout. + * needs no lock. A collapse takes a read lock on the VMA for each round;= a + * scan still runs under the mmap_lock its caller holds. * * A round carries a batch of candidate windows through the passes togethe= r, * rather than carrying one window through the whole collapse. [ptl] and @@ -450,11 +451,11 @@ int collapse_control_init(struct collapse_control *cc) } =20 /* - * The scan and the allocation both dropped mmap_lock, so nothing seen bef= ore it - * can be trusted: check the VMA the round just looked up and the PTE table - * again, and that they still allow every provisioned candidate. + * The scan and the allocation both ran unlocked, so nothing seen before c= an be + * trusted: check the VMA the round just locked and the PTE table again, a= nd + * that they still allow every provisioned candidate. * - * The VMA was found by address, so it need not be the one the scan saw, n= or + * The VMA was locked by address, so it need not be the one the scan saw, = nor * still cover everything the round collected -- thp_vma_suitable_order() = asks * that of each candidate, since a window is aligned to its own order. A = VMA * that shrank under a candidate therefore refuses that candidate and no m= ore, @@ -528,9 +529,9 @@ static enum scan_result collapse_revalidate(struct vm_a= rea_struct *vma, /* * Bring one address to a state the freeze will accept: present, and exclu= sive if * it is anonymous. Every fault it takes to get there counts in *nr_fault= s, each - * one an allocation or a read the round is paying for. Returns with mmap= _lock - * dropped on every failure, because the fault path may drop it and the ca= ller - * cannot tell which case it is in. + * one an allocation or a read the round is paying for. Returns with the = VMA + * read lock dropped on every failure, because the fault path may drop it = and + * the caller cannot tell which case it is in. * * SCAN_EXCEED_SWAP_PTE is the exception: it is a verdict on this candidate * rather than on the round, nothing was faulted to reach it, and it keeps= the @@ -543,6 +544,7 @@ static enum scan_result collapse_faultin_addr(struct vm= _area_struct *vma, { struct mm_struct *mm =3D vma->vm_mm; const unsigned int flags =3D FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_UNSHARE | + FAULT_FLAG_VMA_LOCK | (mm !=3D current->mm ? FAULT_FLAG_REMOTE : 0); unsigned int tries; =20 @@ -553,7 +555,7 @@ static enum scan_result collapse_faultin_addr(struct vm= _area_struct *vma, =20 pte =3D pte_offset_map(pmd, addr); if (!pte) { - mmap_read_unlock(mm); + vma_end_read(vma); return SCAN_NO_PTE_TABLE; } ptent =3D ptep_get_lockless(pte); @@ -594,14 +596,14 @@ static enum scan_result collapse_faultin_addr(struct = vm_area_struct *vma, ret =3D handle_mm_fault(vma, addr, flags, NULL); (*nr_faults)++; /* - * Not a verdict on this window: the fault dropped the lock to - * wait, which is what a swap-in normally does. Distinct from + * Not a verdict on this window: the fault dropped the VMA lock + * to wait, which is what a swap-in normally does. Distinct from * SCAN_PAGE_LOCK, a folio someone else holds locked. */ if (ret & VM_FAULT_RETRY) return SCAN_LOCK_DROPPED; if (ret & VM_FAULT_ERROR) { - mmap_read_unlock(mm); + vma_end_read(vma); return SCAN_FAIL; } } @@ -611,7 +613,7 @@ static enum scan_result collapse_faultin_addr(struct vm= _area_struct *vma, =20 /* * Make every source the round needs present and exclusively owned by this= mm, - * by faulting it in as an ordinary access would. Sleeps, and drops mmap_= lock on + * by faulting it in as an ordinary access would. Sleeps, and drops the V= MA read * failure, since a fault may have to be retried with it released. * * Anything faulted in lands on a per-CPU LRU batch, holding a reference t= he @@ -663,7 +665,7 @@ static enum scan_result collapse_faultin(struct vm_area= _struct *vma, } } out: - /* @vma is unsafe on the failure path: the callee dropped mmap_lock */ + /* @vma is unsafe on the failure path: the callee dropped its read lock */ trace_mm_collapse_faultin(mm, nr_faults, result); return result; } @@ -785,7 +787,7 @@ static void collapse_unfreeze_candidate(struct mm_struc= t *mm, * next page. No layout is refused for its shape -- the next slot simply = starts * its own span -- so partially mapped and compound sources collapse too. * - * Caller holds mmap_read and the table's ptl. + * Caller holds the VMA read lock and the table's ptl. */ static enum scan_result collapse_check_candidate(struct vm_area_struct *vm= a, struct collapse_control *cc, @@ -929,7 +931,7 @@ static enum scan_result collapse_check_candidate(struct= vm_area_struct *vma, * * On entry: * - * - mmap_read is held, and the table's ptl for the whole freeze; + * - the VMA read lock is held, and the table's ptl for the whole freeze; * - collapse_check_candidate() has accepted the candidate under that sam= e ptl * hold; * - the round is covered by an mmu_notifier_invalidate_range_start() iss= ued @@ -1382,11 +1384,12 @@ static bool collapse_abort_slot(struct vm_area_stru= ct *vma, struct folio *folio, =20 /* * Abort one frozen candidate at install time: it took a machine check dur= ing the - * copy, or some of its slots no longer hold our migration entries. mmap_= read - * (held freeze..putback) blocks fork, mremap and munmap, and faults wait = on the - * migration entries -- but madvise-class operations run under mmap_read t= oo, so a - * concurrent MADV_DONTNEED may have zapped frozen slots, and a fault may = have - * refilled a zapped one. + * copy, or some of its slots no longer hold our migration entries. The V= MA read + * lock (held freeze..putback) blocks fork, mremap and munmap, each of whi= ch + * takes vma_start_write() on the VMA it touches, and faults wait on the + * migration entries -- but MADV_DONTNEED and MADV_FREE take a VMA read lo= ck of + * their own, which ours does not exclude, so a concurrent zap may have ta= ken + * frozen slots, and a fault may have refilled a zapped one. * * Slots still holding our entries are restored from the saved values (no = TLB * flush: identical translation). Foreign slots are left exactly as found= -- @@ -1500,8 +1503,8 @@ static bool collapse_verify_candidate(struct collapse= _candidate *cand, * The PMD terminal layer: verify, detach the table, deposit a fresh one a= nd * install the leaf, as one atomic section under the pmd lock. A pmd_none= window * never exists -- faults stay held at pte level by the migration entries - * throughout -- which is what lets PMD collapse run under mmap_read like = the rest - * of the engine. + * throughout -- which is what lets PMD collapse run under a VMA read lock= like + * the rest of the engine. */ static void collapse_install_pmd(struct vm_area_struct *vma, struct collapse_control *cc, pmd_t *pmd) @@ -1571,8 +1574,8 @@ static void collapse_install_pmd(struct vm_area_struc= t *vma, * walks on the sources are unreachable -- refcounts frozen, folio locks * held from freeze to putback -- non-rmap pte walkers see migration * entries, pmd-level observers see the old table or the leaf and never an - * intermediate, and fork, mremap and munmap take mmap_write, which our - * mmap_read excludes. + * intermediate, and fork, mremap and munmap take vma_start_write() on the + * VMA they touch, which our VMA read lock excludes. * * The flush inside pmdp_collapse_flush() is the round's second over this * range: the freeze displaced every leaf here and flushed before dropping @@ -1829,13 +1832,23 @@ static void collapse_round(struct mm_struct *mm, un= signed long pmd_addr, collapse_reserve(mm, cc); collapse_deposit(mm, cc); =20 + /* + * A read lock on the VMA rather than on the mm. A round works inside one + * VMA, and everything else it has to be excluded from -- fork, split, + * merge, unmap, and the free_pgtables() that follows them -- takes + * vma_start_write() on the VMA it touches first. An mmap_write elsewhere + * in the mm no longer waits behind a collapse. + */ retry: - mmap_read_lock(mm); - - vma =3D find_vma(mm, pmd_addr); + vma =3D lock_vma_under_rcu(mm, pmd_addr); if (!vma) { - result =3D SCAN_VMA_NULL; - goto out_unlock; + /* + * Not only a VMA that has gone: lock_vma_under_rcu() also fails + * on one being written to right now. A caller cannot tell the + * two apart, so say what is true of both -- try again. + */ + result =3D SCAN_VMA_LOCK; + goto out; } =20 result =3D collapse_revalidate(vma, pmd_addr, cc, &pmd); @@ -1852,7 +1865,7 @@ static void collapse_round(struct mm_struct *mm, unsi= gned long pmd_addr, if (result =3D=3D SCAN_LOCK_DROPPED && --passes) goto retry; if (result !=3D SCAN_SUCCEED) - goto out; /* the callee released mmap_lock */ + goto out; /* the callee released the VMA read lock */ =20 /* One invalidate window spans the batch, as collapse_revalidate() left i= t */ mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, @@ -1882,7 +1895,7 @@ static void collapse_round(struct mm_struct *mm, unsi= gned long pmd_addr, mmu_notifier_invalidate_range_end(&range); =20 out_unlock: - mmap_read_unlock(mm); + vma_end_read(vma); out: nr_installed =3D collapse_finish(mm, cc, result); trace_mm_collapse_round(mm, cc->nr_candidates, nr_installed, result, diff --git a/mm/collapse.h b/mm/collapse.h index 74e513c5c76c..e5a0ffab049c 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -39,6 +39,7 @@ enum scan_result { SCAN_PAGE_COMPOUND, SCAN_ANY_PROCESS, SCAN_VMA_NULL, + SCAN_VMA_LOCK, SCAN_VMA_CHECK, SCAN_ADDRESS_RANGE, SCAN_DEL_PAGE_LRU, diff --git a/mm/madvise.c b/mm/madvise.c index c1bb425be3f4..bd9123ee3cb1 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -933,6 +933,7 @@ static int madvise_collapse_errno(enum scan_result r) case SCAN_PAGE_FILLED: case SCAN_PAGE_HAS_PRIVATE: case SCAN_PAGE_DIRTY_OR_WRITEBACK: + case SCAN_VMA_LOCK: return -EAGAIN; /* * Other: Trying again likely not to succeed / error intrinsic to --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A70CE40F72C; Sun, 16 Aug 2026 22:47:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920465; cv=none; b=QC15jGkjVGW9ImzjMrFYyn7CvCd4eaCJJRFE17xWQumAQR2jReXmJgE67CLaKwLmuLmYZm9mHZ8WOoX2FMnTiHI36KT0Sb0vLKX4fHeZ+xDCj5DJ2Qgespsrh0EkEAv/+VpyVSIEEDsEvsgwOLMwBi/qzIZOlvzkdCVA3f4L3uE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920465; c=relaxed/simple; bh=8b4LkbLMpL4QuaHUG434StKbuBjTZ2oy1sPW1STQicI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=f5NGQbQ2mHwIt34bR51hyqFcbsmIiHkjBGv7+kzwi3Kp0uMU+xNsnEhPZ6eF4rD1oquSgXbf7di4W+sNCVjRO/Dm4Rt1tATrRmv9ZbDh30DS6pek4vOqUF4XFvZSF4YSJ7bmn4WIpy39EN934cMXx/Dky8pX9pR7fwJWUv4FWl0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=OSA76d6n; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=WoWhU4ye; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="OSA76d6n"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="WoWhU4ye" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.phl.internal (Postfix) with ESMTP id 1197514000EB; Sun, 16 Aug 2026 18:47:43 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Sun, 16 Aug 2026 18:47:43 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920463; x= 1787006863; bh=ftB1HJ8Y15YVxr5O+cxOqWPRClS84lYkjsQI5VetnJA=; b=O SA76d6nMTniIxLQ415JPsCD5DMqyGlEennXopHdxKTGu8qymvPyWzNMyvAZDGpAh hJf8Q74k9BzfvmJSAjmYJQCNGMutYbLSPUVm8dhMKEgi6U/lgX2ErtYzuZXEIXvI GxBu39nmjYPX4XycBzI/f1grk/PMJwaT0YdVsz6dXCDdrqmhrqZ5d0meTTOaJlxa JfMCD/yl+P5VEjyHyQai3w42hESA0DEUAwKGw62UAcSBVVTRgKHk0X602Oc3ZbNR Nvn4lar9sCGySd6c7QFVeN7zKbdtMm88FXHM1+ZuosVQ6nJSHTM/xeR+KB7zYZ3y kF/HP0BU4WseZZiNC/Eiw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920463; x=1787006863; bh=f tB1HJ8Y15YVxr5O+cxOqWPRClS84lYkjsQI5VetnJA=; b=WoWhU4yeuVWdXErGT FSsCSrQXGPmIxaMq8DPs7WSJw09tPvWM2cER/r/TNCJ9BYH8ZWlJISEpI4oUWF5A kmPb0AM4PSOX5z/iG6YXq5PvCAuXCa+H8Ns73K+etHAEVSoxSm4THccwyKdSKjy1 BDA8bmo4cPNyiwtWDEUhWSVS/84wWVa+jCp1yG4RTGzDo60VssqMr/EG4wLQlvZU k0zP072hYNpL5H/sLBbS+byxeWL2PWlppyGcLF3vlgBlPeveGj/JBQvEipfOnYd4 79IuMK9evW5eRERezTRxUxcgnpUj26TDiQRnkfrZJxHJPWyBe8nUxRj6Nox3/pAZ WyBUQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkhNC7YI+Ei0MrF6zu90W9yR7z714JmmtstcBEn9fnZ0JEx8KxNru7XjOIPijUAd VrNDumTIvQCtLJHAbrPKMRdbJsCYKY1AWxa28xNEDdkQaDqy3PIOJVA9R+OjsTN2M4dnB4 WQk97ejO2lmWwXR5mw6146Dl5x2vSRCf1uehpA4Z3DNP/uvzS8DmGwE+/LsFmt9BQO6ZtJ eWd0dPZ2F0WrWW1tiYo0zV8d4Xgh8ztMbcUaDgwCK69ZeQi0ebesOWu50OSj19NNGHzec3 w5R1etyOsXXy4NxsI/YxbKwNbzdBv+aWoLN5/PFLP+n740pr735uP0KJxEPqjezkxGUEae VbV3rMngkaPxQKuJpJI13Iuy9BW4VdoLWVjHaMCTYZzuReZcweII6OV6N7/eDhyeqRnM3U ZaRP0X5mK4wstAamKJjGnw1uyl2I1fZsgwkT91DxkdavZuHiuRFeu9WzlEiV6bYzz8TpAA dLfkb9PE5A3JzRKardx2vbuSZ+Ht+1ceDrUA3aw09PbBDyLGrX2AOrLz43qEIUSlbyaGxC tSayGSOXDkXXT0rR32+d3CFxt9XK340pkIj6m0BjxG9p46OD0dxAkc1BBoLJPIJT6NlKs5 LWq6FThsDA/9lLgOC1LEnx6RW9EGDOFGdJ7sV1UKB0t2r9/y4E3uyLBgb/cg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:42 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 46/57] mm/khugepaged: scan under a per-VMA read lock Date: Sun, 16 Aug 2026 23:45:58 +0100 Message-ID: <20260816224609.308019-47-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A pass took mmap_lock for the whole walk: every VMA of the address space judged, and every table of every VMA scanned, under one lock. A writer anywhere in the mm waits for all of it. And a pass that finds nothing to collapse -- which is what a pass over an already-collapsed address space is -- holds the lock for the whole sweep to say so. Take a read lock on one VMA at a time instead. lock_next_vma() finds the next VMA at or after the cursor and locks it, falling back to mmap_lock only where it cannot. A scan holds that lock across every table of that VMA and no longer. What a round has to be excluded from already takes vma_start_write() on the VMA it touches, so the exclusion is the same. What changes is that it is scoped to the VMA being scanned. Two things follow from the iterator no longer being carried by mmap_lock. The cursor moves by hand, because nothing else advances it now: past a VMA that was walked, past one skipped without being looked at, and past each table a scan was offered. And the end of the address space has to be recognised rather than fallen out of. scan_complete says whether lock_next_vma() ran out of VMAs: an error is not the end, and treating it as one would release the slot with the address space half scanned. The scan asserted mmap_assert_locked() on the way in. Its two callers no longer agree on what they hold -- khugepaged a VMA read lock from here, MADV_COLLAPSE still mmap_lock -- so there is no single lock to assert, and the assert goes. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 25 +++++----- mm/khugepaged.c | 128 +++++++++++++++++++++++++++++++----------------- 2 files changed, 96 insertions(+), 57 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index b6e92e24a3c5..b3343595bdf2 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -52,8 +52,8 @@ enum collapse_pass { * The folios mapped across a window of PTEs become one folio of that wind= ow's * order, with the sources quiesced by the two barriers migration uses -- * migration entries in their PTEs, then a frozen refcount -- so the copy = itself - * needs no lock. A collapse takes a read lock on the VMA for each round;= a - * scan still runs under the mmap_lock its caller holds. + * needs no lock. A scan runs under a read lock on the VMA it was handed;= a + * collapse is given none and takes its own, for one round at a time. * * A round carries a batch of candidate windows through the passes togethe= r, * rather than carrying one window through the whole collapse. [ptl] and @@ -1949,9 +1949,10 @@ static enum scan_result collapse_scan_table(struct v= m_area_struct *vma, * wait for a scan of the whole table. * * pte_offset_map() holds rcu_read_lock() until pte_unmap(), which is - * what keeps the table itself from being freed underneath the walk; - * mmap_lock keeps the VMA attached, without which free_pgtables() could - * free it without waiting for RCU at all. Nothing below here sleeps. + * what keeps the table itself from being freed underneath the walk; the + * VMA read lock keeps the VMA attached, without which free_pgtables() + * could free it without waiting for RCU at all. Nothing below here + * sleeps. */ pte =3D pte_offset_map(pmd, start); if (!pte) { @@ -2174,9 +2175,9 @@ static void collapse_anon_scan_init(struct collapse_c= ontrol *cc) /* * Judge one table's worth of @vma, leaving in @cc what a collapse could u= se: * which orders are still worth attempting, and why the table was turned d= own if - * some order was. Holds mmap_lock throughout -- it only reads -- and a c= aller - * that acts on what it found hands the range to collapse_anon_pmd() after= wards, - * without the lock. + * some order was. Holds the read lock it was called under throughout -- = it only + * reads -- and a caller that acts on what it found hands the range to + * collapse_anon_pmd() afterwards, without any lock. */ static enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma, unsigned long start, unsigned long end, @@ -3745,9 +3746,9 @@ static enum scan_result collapse_file_pmd(struct mm_s= truct *mm, =20 /* * Scan one table's worth of @vma and decide whether there is anything to = collapse - * in it. The caller holds mmap_lock for reading and still holds it when = this - * returns: what is looked at is either the VMA or a page table that the l= ock - * keeps in place. + * in it. The caller holds a read lock and still holds it when this retur= ns: + * what is looked at is either the VMA or a page table that the lock keeps= in + * place. * * Returns whether collapse_run_pmd() has anything to do, and a scan that = found * something has to be run: the file side takes a reference on the file wh= ile it @@ -3760,8 +3761,6 @@ bool collapse_scan_pmd(struct vm_area_struct *vma, un= signed long addr, { struct mm_struct *mm =3D vma->vm_mm; =20 - mmap_assert_locked(mm); - /* * What the scan answers with, so cleared before it runs. * collapse_anon_scan_init() clears the orders too, but only once the diff --git a/mm/khugepaged.c b/mm/khugepaged.c index f3ea1846990e..1d77d9a8046d 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -511,10 +511,10 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, __releases(&khugepaged_mm_lock) __acquires(&khugepaged_mm_lock) { - struct vma_iterator vmi; struct mm_slot *slot; struct mm_struct *mm; struct vm_area_struct *vma; + bool scan_complete =3D false; unsigned int progress_prev =3D cc->progress; =20 lockdep_assert_held(&khugepaged_mm_lock); @@ -534,55 +534,82 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, vma =3D NULL; =20 /* - * A reference on mm_users for as long as the pass works on this address - * space. __mmput() cannot start while one is held, so neither can - * exit_mmap(), and the VMAs and page tables stay where they are. + * Hold the address space open for the pass. A collapse works under a + * per-VMA read lock, and the barrier __khugepaged_exit() puts in front + * of exit_mmap() -- mmap_write_lock() -- waits for a reader of + * mmap_lock, not for a reader of one VMA. A reference on mm_users + * stops __mmput(), and so both of those, from starting at all. * - * Once per pass, not once per table: the reference is what makes the - * address space safe to work on, and a pass is how long that is wanted - * for. Nothing else in mm takes it per unit of work -- DAMON takes one - * per target and walks every region under it, swapoff one per mm across - * the whole address space, userfaultfd one per call. + * Once per pass rather than once per table: the reference is what makes + * the address space safe to work on, and the pass is how long that is + * wanted for. Nothing else in mm takes it per unit of work -- DAMON + * takes one per target and walks every region under it, swapoff one per + * mm across the whole address space, userfaultfd one per call. It is + * dropped below before the exiting mm is judged, so that judgement still + * sees the true count. */ if (!mmget_not_zero(mm)) goto breakouterloop_no_mmput; =20 - /* - * Don't wait for semaphore (to avoid long wait times). Just move to - * the next mm on the list. - */ - if (unlikely(!mmap_read_trylock(mm))) - goto breakouterloop_mmap_lock; - cc->progress++; - if (unlikely(collapse_test_exit_or_disable_mmref(mm))) - goto breakouterloop; =20 - vma_iter_init(&vmi, mm, khugepaged_scan.address); - for_each_vma(vmi, vma) { + /* + * One VMA at a time, each held by its own read lock rather than by + * mmap_lock over the whole address space. lock_next_vma() locks what it + * finds, falling back to mmap_lock only where it cannot. + * + * Whether this mm still wants collapsing is asked once, at the top of + * each round of the loop. Asking again before entering it only repeats + * the same question: nothing between the two can answer it differently. + */ + for (;;) { unsigned long hstart, hend, window; + struct vma_iterator vmi; unsigned long orders; =20 cond_resched(); + /* + * Our reference is the reason the count cannot fall to zero, so + * it is also what an address space whose owner has gone looks + * like. Stopping is what frees it: nothing else here would. + */ if (unlikely(collapse_test_exit_or_disable_mmref(mm))) { cc->progress++; - break; + goto breakouterloop; } =20 /* - * Before the VMA is judged, so that a pass over an address space - * of VMAs it skips is bounded by the budget too: each one is - * charged for, and none of them was being asked to be scanned. + * Before a VMA is locked, so that a pass over an address space + * of VMAs it skips is bounded by the budget too, and so that a + * collapse returning here does not lock one to be told it is + * out of budget. */ if (cc->progress >=3D progress_max) - break; + goto breakouterloop; + + /* The first VMA at or after the cursor, which often sits in a gap */ + rcu_read_lock(); + vma_iter_init(&vmi, mm, khugepaged_scan.address); + vma =3D lock_next_vma(mm, &vmi, khugepaged_scan.address); + rcu_read_unlock(); + + /* + * NULL is the end of the address space, and the only thing that + * finishes this mm. An error is a fatal signal or the unlikely + * reference count overflow: leave the mm for the next pass + * rather than treat it as walked. + */ + if (IS_ERR_OR_NULL(vma)) { + scan_complete =3D !IS_ERR(vma); + vma =3D NULL; + goto breakouterloop; + } =20 orders =3D collapse_possible_orders(vma, vma->vm_flags, TVA_KHUGEPAGED); if (!orders) { - khugepaged_scan.address =3D vma->vm_end; cc->progress++; - continue; + goto next_vma; } =20 /* @@ -595,9 +622,8 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, hstart =3D ALIGN(vma->vm_start, window); hend =3D ALIGN_DOWN(vma->vm_end, window); if (khugepaged_scan.address > hend) { - khugepaged_scan.address =3D vma->vm_end; cc->progress++; - continue; + goto next_vma; } if (khugepaged_scan.address < hstart) khugepaged_scan.address =3D hstart; @@ -605,19 +631,24 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, while (khugepaged_scan.address < hend) { unsigned long pmd_addr, range_end, start; =20 + cond_resched(); + + if (unlikely(collapse_test_exit_or_disable_mmref(mm)) || + cc->progress >=3D progress_max) { + vma_end_read(vma); + vma =3D NULL; + goto breakouterloop; + } + /* One table's worth at most, and never past the VMA */ pmd_addr =3D khugepaged_scan.address & HPAGE_PMD_MASK; range_end =3D min(hend, pmd_addr + HPAGE_PMD_SIZE); - - cond_resched(); - if (unlikely(collapse_test_exit_or_disable_mmref(mm)) || - cc->progress >=3D progress_max) - goto breakouterloop; + start =3D khugepaged_scan.address; =20 VM_WARN_ON_ONCE(khugepaged_scan.address < hstart); + VM_WARN_ON_ONCE(range_end > hend); =20 - start =3D khugepaged_scan.address; - /* move to next address */ + /* Move the cursor on regardless of what the scan says */ khugepaged_scan.address =3D range_end; =20 /* If nothing to collapse, the lock is still ours */ @@ -627,21 +658,30 @@ static void collapse_scan_mm_slot(unsigned int progre= ss_max, } =20 /* collapse_run_pmd() takes its own locks, so give this up */ - mmap_read_unlock(mm); + vma_end_read(vma); + vma =3D NULL; + *result =3D collapse_run_pmd(mm, start, range_end, cc); if (*result =3D=3D SCAN_SUCCEED) - ++khugepaged_pages_collapsed; - goto breakouterloop_mmap_lock; + khugepaged_pages_collapsed++; + goto breakouterloop; } +next_vma: + /* + * Past this VMA: the cursor has to move by hand, where the + * mmap_lock iterator used to carry it. A VMA that was walked + * is charged by the scan itself, one table at a time; only one + * passed over without being looked at is charged here. + */ + khugepaged_scan.address =3D vma->vm_end; + vma_end_read(vma); + vma =3D NULL; } + breakouterloop: - mmap_read_unlock(mm); /* exit_mmap will destroy ptes after this */ -breakouterloop_mmap_lock: /* * Not mmput(): the last reference would run exit_mmap() here, and * khugepaged is not the thread that should tear an address space down. - * Dropped before the exiting mm is judged below, so that judgement still - * sees the true count. */ mmput_async(mm); breakouterloop_no_mmput: @@ -652,7 +692,7 @@ static void collapse_scan_mm_slot(unsigned int progress= _max, * Release the current mm_slot if this mm is about to die, or * if we scanned all vmas of this mm, or THP got disabled. */ - if (collapse_test_exit_or_disable(mm) || !vma) { + if (collapse_test_exit_or_disable(mm) || scan_complete) { /* * Make sure that if mm_users is reaching zero while * khugepaged runs here, khugepaged_exit will find --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A5723411FA6; Sun, 16 Aug 2026 22:47:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920467; cv=none; b=l59IG+uCBwVLKlr0fC0MvSBZzuEzdstzwld0/kczYhIujUKTdZ/dDcQq74yhDvrK3JeclpirYcbwnBbQLVRafjebzh6wgZ0taUi16/9mQamTIhANxUfmOWe6gC2AnvadaDeSDTNVsgGBtAqL0ilBdu4IawGEz8UU2iYuDsLQ9hI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920467; c=relaxed/simple; bh=FbrM1RxMkPH21ZpwLArdqcOIaBLPz31ZSnxyLXwVwic=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VZoUG8AHB1SsXWu7wixpb01HjAEv55cghJEZqNCqo8ryWB03iD8WLNFLFSBUL2AkS7vmMHsAlsbfYwlwWDmS4iYbfAMMZ325nm6hFqnCZgmYAncR2sN3tytwYVknU96melBe0z6/X9nTHluWBx9MxNRDQTMhQi93MVj9FWnAn3E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=OJIjznKj; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DN6VmWJH; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="OJIjznKj"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DN6VmWJH" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.phl.internal (Postfix) with ESMTP id E8A1714000FD; Sun, 16 Aug 2026 18:47:44 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Sun, 16 Aug 2026 18:47:44 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920464; x= 1787006864; bh=Rc7LbnNgXQF/PbZDks+SvKPcsvsUstGaFXUd5oKeSUM=; b=O JIjznKjsNnYrZJwbG2uEUZFbMrw0iEtyB2PYr2C07fvU0o43n6DmSy+/NXaQqF9l XdM+SNyedXBr9JJFdltwuRAcVT970OFe049Y916sAT65U5i3DLW77lm5Ya6dLewv dzhODT/0Q5OI9i++R5/39ITM5AGoaPMukebE/394olpt8relbwvJ67POozAGYyOP rcKHcz1vOFUJkhm0A2WsN6XLCm9dXzPTXdb8fCxzyqiqmeDqhS1c7RSm/mKmYOZv wed7fmFQ4UgXp8CgGb/oi5qhQTohqYA7bL3MV/J/s6RNzLBnw6YzNObXC1lO6iO1 wuhB4gKZAQlPQrGpollkA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920464; x=1787006864; bh=R c7LbnNgXQF/PbZDks+SvKPcsvsUstGaFXUd5oKeSUM=; b=DN6VmWJHWJ/BUnpT4 1QmAA+Hm4z/fUd6jjpmq0qmMwJisK58EFfAtlxb5nLYKnFshztmqTU0fvX15yXmx M6n/Ou93dZx7EbnVyCK8QO4Q3CAOVRZPaKpHg4aeYy65LCExZ00uh5N3zCuczfFZ y2lpXB02FA9hVNN4p/zaMl4AUZv9vQz2GR9N/VQKHx2nre+8O8sxNV21Tx7WXYuv 7gzkAxxIWyEjLP4ZJQUDKuK9XRouGPe+nC2UgtjiBzHVMFsQhZEVr1zeM4aIV70J t8JpmOTJISP6JC8Vf5euZRtBCazQdZsrPPOVVt8cDgPyrKCqvKiA9EpB96xCSzW7 hs9Ug== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkhNC7YI+Ei0MrF6zu90W9yR7z714JmmtstcBEn9fnZ0JEx8KxNru7XjOIPijUAd VrNDumTIvQCtLJHAbrPKMRdbJsCYKY1AWxa28xNEDdkQaDqy3PIOJVA9R+OjsTN2M4dnB4 WQk97ejO2lmWwXR5mw6146Dl5x2vSRCf1uehpA4Z3DNP/uvzS8DmGwE+/LsFmt9BQO6ZtJ eWd0dPZ2F0WrWW1tiYo0zV8d4Xgh8ztMbcUaDgwCK69ZeQi0ebesOWu50OSj19NNGHzec3 w5R1etyOsXXy4NxsI/YxbKwNbzdBv+aWoLN5/PFLP+n740pr735uP0KJxEPqjezkxGUEAo z0No3HJZUJpep6tsc5V41NtAo/QmKAX1kCmR1GI4As8yw+j+1WrkBGvJZfJG6IqsJdVFZs 5FIr/7VYJuR8ul6w0vGbtWCITm2mD0nL5V0Cih+VpPedlzBfi4FBcgFxNxgdcHQuReNJxB oByGinK4f+XZOIoDvDzy5ihCsDXKlG1e/qbXGsJ3ayhAWA7YZiHa6xXbrMd3RjWrrsUWwo t1nTkzHa8RrFwbbWD2Zsomhb6fsaDA+a67RbUMyEDFtzlbvMh63VAPk4XH5d3OvGnh31cm KlVuA+iMrvppmUkIPdXNd5uceV7/4WuFPUQtG1l2sUw95+q7De/imXhZpvCA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:44 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 47/57] mm/madvise: collapse under a per-VMA read lock Date: Sun, 16 Aug 2026 23:45:59 +0100 Message-ID: <20260816224609.308019-48-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" MADV_COLLAPSE arrives under mmap_lock, gives it up, and takes it again per PMD to scan and hand a table to the collapse. Every writer to the address space waits behind each of those, and the range can be as large as the caller asked for. Take a read lock on the VMA instead: MADVISE_VMA_READ_LOCK in the lock mode, lock_vma_under_rcu() per PMD, and the walk told not to release what the behaviour already let go. A scan that finds nothing keeps the lock, so a range that is already collapsed walks it without relocking. A collapse gives the lock up and looks the VMA up again afterwards: it can shrink while nothing is held, which the scan reports as a refused range like any other. Remote madvise is the exception. process_madvise() has to untag the range with untagged_addr_remote() before any VMA is looked at. That reads mm state mmap_lock protects, so remote MADV_COLLAPSE keeps the mmap_read it has. A VMA that cannot be locked is reported as SCAN_VMA_LOCK, which reaches the caller as -EAGAIN. lock_vma_under_rcu() also fails on a VMA being written to, and reporting that as a range which shrank would tell the caller a collapse succeeded where none was attempted. With that, nothing produces SCAN_VMA_NULL any more, and the test for it goes. Both callers hold a read lock on the VMA now. That is what the changes outside madvise.c are for: the engine's interface documented mmap_lock as its precondition, and only here does that stop being true of every caller. The scan asserts the lock again too, which was not possible while the callers disagreed. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 24 +++++++------ mm/collapse.h | 8 ++--- mm/madvise.c | 98 +++++++++++++++++++++++++++++++++++++++------------ 3 files changed, 93 insertions(+), 37 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index b3343595bdf2..1e3b2d202ffe 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -3685,8 +3685,8 @@ static enum scan_result collapse_scan_file_pmd(struct= vm_area_struct *vma, =20 /* * A PMD that is huge already has nothing left to collapse, and skipping - * it here is what keeps mmap_lock out of a collapse that would find - * nothing. Everything else is worth the page cache scan, pmd_none() + * it here is what keeps a collapse that would find nothing from being + * run at all. Everything else is worth the page cache scan, pmd_none() * included: a file range can be collapsed out of the cache without being * mapped first, which is why this is not the test the anonymous side * makes. @@ -3703,8 +3703,8 @@ static enum scan_result collapse_scan_file_pmd(struct= vm_area_struct *vma, =20 /* * Build a PMD over what the page cache holds, and map it over the range i= f a huge - * folio is already there but mapped by PTEs. Runs with no mmap_lock, whi= ch the - * caller gave up, and takes it again only for that last step. + * folio is already there but mapped by PTEs. Runs with no lock on the VM= A, + * which the caller gave up, and takes mmap_lock only for that last step. */ static enum scan_result collapse_file_pmd(struct mm_struct *mm, unsigned long addr, struct collapse_control *cc) @@ -3746,9 +3746,9 @@ static enum scan_result collapse_file_pmd(struct mm_s= truct *mm, =20 /* * Scan one table's worth of @vma and decide whether there is anything to = collapse - * in it. The caller holds a read lock and still holds it when this retur= ns: - * what is looked at is either the VMA or a page table that the lock keeps= in - * place. + * in it. The caller holds a read lock on @vma and still holds it when th= is + * returns: what is looked at is either the VMA or a page table that the l= ock + * keeps in place. * * Returns whether collapse_run_pmd() has anything to do, and a scan that = found * something has to be run: the file side takes a reference on the file wh= ile it @@ -3761,6 +3761,8 @@ bool collapse_scan_pmd(struct vm_area_struct *vma, un= signed long addr, { struct mm_struct *mm =3D vma->vm_mm; =20 + vma_assert_locked(vma); + /* * What the scan answers with, so cleared before it runs. * collapse_anon_scan_init() clears the orders too, but only once the @@ -3789,10 +3791,10 @@ bool collapse_scan_pmd(struct vm_area_struct *vma, = unsigned long addr, } =20 /* - * Collapse what the scan selected. Called with no mmap_lock: the caller = gives it - * up first, because a collapse takes it again for each round and revalida= tes - * under it, and holding it across the whole collapse would keep a writer = to the - * address space waiting for it. + * Collapse what the scan selected. Called with no lock on the VMA the sc= an + * looked at: the caller gives that up first, because a collapse takes its= own + * for each round and revalidates under it, and holding one across the who= le + * collapse would keep a writer to the VMA waiting for it. */ enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr, unsigned long end, struct collapse_control *cc) diff --git a/mm/collapse.h b/mm/collapse.h index e5a0ffab049c..aeaca305e71a 100644 --- a/mm/collapse.h +++ b/mm/collapse.h @@ -207,8 +207,8 @@ static inline int collapse_test_exit_or_disable_mmref(s= truct mm_struct *mm) * collapse_run_pmd(mm, addr, end, cc); when the scan found work * collapse_control_release(cc); * - * The caller holds mmap_lock for reading and passes a range within one PT= E table - * of @vma. A range the VMA does not cover is refused, which is also how = a caller + * The caller holds a read lock on @vma and passes a range within one PTE = table + * of it. A range the VMA does not cover is refused, which is also how a = caller * learns that its own range shrank. * * A scan returns with that lock still held: it only reads, and almost eve= ry table @@ -218,8 +218,8 @@ static inline int collapse_test_exit_or_disable_mmref(s= truct mm_struct *mm) * A collapse is called without it: the caller gives the lock up first, an= d with it * @vma and anything derived under it, so a caller carrying on has to look= up * again. What the collapse does -- allocate, quiesce, copy, flush -- is = slow - * enough that a writer would wait behind it, so it takes the lock again p= er round - * instead, and revalidates rather than trusting what the scan saw. + * enough that a writer to the VMA would wait behind it, so it takes its o= wn lock + * per round instead, and revalidates rather than trusting what the scan s= aw. * * A scan that found something has to be run: the file side takes a refere= nce on * the file while it still has the VMA to take it from, and the run is wha= t gives diff --git a/mm/madvise.c b/mm/madvise.c index bd9123ee3cb1..5f6d815d70ad 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -276,6 +276,18 @@ static void mark_mmap_lock_dropped(struct madvise_beha= vior *madv_behavior) madv_behavior->lock_dropped =3D true; } =20 +/* + * The VMA-lock counterpart, for a behaviour that releases the VMA it was = handed + * and locks what it needs for itself. The walk has nothing left to relea= se, + * and unlike the mmap_lock case it has nothing to carry on with either: t= he VMA + * fast path applies to one VMA and returns. + */ +static void mark_vma_lock_dropped(struct madvise_behavior *madv_behavior) +{ + VM_WARN_ON_ONCE(madv_behavior->lock_mode !=3D MADVISE_VMA_READ_LOCK); + madv_behavior->lock_dropped =3D true; +} + /* * Schedule all required I/O operations. Do not wait for completion. */ @@ -948,6 +960,8 @@ static int madvise_collapse_errno(enum scan_result r) static int madvise_collapse(struct madvise_behavior *madv_behavior) { struct madvise_behavior_range *range =3D &madv_behavior->range; + const bool vma_locked =3D + madv_behavior->lock_mode =3D=3D MADVISE_VMA_READ_LOCK; struct vm_area_struct *vma =3D madv_behavior->vma; struct mm_struct *mm =3D madv_behavior->mm; unsigned long hstart, hend, addr; @@ -981,13 +995,22 @@ static int madvise_collapse(struct madvise_behavior *= madv_behavior) } =20 /* - * Nothing below wants the lock the VMA walk left held, and - * lru_add_drain_all() waits on every CPU, so give it up first. The - * walk carries on under mmap_lock and its own caller is what drops it, - * so reporting this only tells the walk that its VMA is now stale. + * Give up whatever the caller locked for us. lru_add_drain_all() below + * must not run under a lock, and the loop locks what it works on for + * itself, one VMA at a time, so the caller's VMA is of no use past here. + * + * Which lock that is depends on how we were reached. A range inside one + * VMA arrives with that VMA read-locked and nothing else; a range that + * spans VMAs arrives under mmap_lock, because try_vma_read_lock() took + * it and turned the walk generic. */ - mmap_read_unlock(mm); - mark_mmap_lock_dropped(madv_behavior); + if (vma_locked) { + vma_end_read(vma); + mark_vma_lock_dropped(madv_behavior); + } else { + mmap_read_unlock(mm); + mark_mmap_lock_dropped(madv_behavior); + } vma =3D NULL; vma_orders =3D 0; lru_add_drain_all(); @@ -996,22 +1019,36 @@ static int madvise_collapse(struct madvise_behavior = *madv_behavior) enum scan_result result; =20 /* - * A collapse gives the lock up, and the VMA has to be found - * again after one: it can shrink while nothing is held. A scan - * that finds nothing to collapse leaves the lock alone, so a - * range that is already collapsed walks it without relocking. + * On another process, the reference this call holds is what + * keeps the address space from being torn down -- so if it is + * the only one left, the owner has gone and every page of it is + * waiting on us to stop. Nothing else here would notice: the + * range is the caller's, and it can be enormous. + */ + if (mm !=3D current->mm && collapse_test_exit_mmref(mm)) { + hend =3D addr; + break; + } + + /* + * A collapse gives the VMA read lock up, and the VMA has to be + * found again after one: it can shrink while nothing is held. * * Reschedule only here, where nothing is held: a preemption * point under a lock is a writer waiting longer. */ if (!vma) { cond_resched(); - mmap_read_lock(mm); - vma =3D vma_lookup(mm, addr); - if (!vma) { - mmap_read_unlock(mm); - hend =3D addr; - break; + vma =3D lock_vma_under_rcu(mm, addr); + if (IS_ERR_OR_NULL(vma)) { + /* + * Not only a VMA that has gone: this also fails + * on one being written to right now. Say what + * is true of both -- try again. + */ + vma =3D NULL; + last_fail =3D SCAN_VMA_LOCK; + goto out; } vma_orders =3D collapse_possible_orders(vma, vma->vm_flags, TVA_FORCED_COLLAPSE); @@ -1023,7 +1060,7 @@ static int madvise_collapse(struct madvise_behavior *= madv_behavior) result =3D cc->scan_refusal; } else { /* collapse_run_pmd() takes its own locks, so give this up */ - mmap_read_unlock(mm); + vma_end_read(vma); vma =3D NULL; /* The mask belonged to that lock, not to this range */ vma_orders =3D 0; @@ -1036,7 +1073,7 @@ static int madvise_collapse(struct madvise_behavior *= madv_behavior) * The VMA shrank under us, so the rest of the range was never * ours to collapse: stop, and expect only what came before. */ - if (result =3D=3D SCAN_VMA_NULL || result =3D=3D SCAN_ADDRESS_RANGE) { + if (result =3D=3D SCAN_ADDRESS_RANGE) { hend =3D addr; break; } @@ -1067,8 +1104,14 @@ static int madvise_collapse(struct madvise_behavior = *madv_behavior) } =20 out: - /* The VMA walk this returns to expects the lock it was holding */ - if (!vma) + if (vma) + vma_end_read(vma); + /* + * Hand mmap_lock back only to a caller that is going to carry on with + * it: the generic walk finds the next VMA under it. The VMA fast path + * applies to one VMA and returns, so it wants nothing back. + */ + if (!vma_locked) mmap_read_lock(mm); collapse_control_release(cc); kfree(cc); @@ -1866,7 +1909,10 @@ int madvise_walk_vmas(struct madvise_behavior *madv_= behavior) if (madv_behavior->lock_mode =3D=3D MADVISE_VMA_READ_LOCK && try_vma_read_lock(madv_behavior)) { error =3D madvise_vma_behavior(madv_behavior); - vma_end_read(madv_behavior->vma); + /* A behaviour that let the VMA go has nothing left to release */ + if (!madv_behavior->lock_dropped) + vma_end_read(madv_behavior->vma); + madv_behavior->lock_dropped =3D false; return error; } =20 @@ -1941,8 +1987,16 @@ static enum madvise_lock_mode get_lock_mode(struct m= advise_behavior *madv_behavi case MADV_PAGEOUT: case MADV_POPULATE_READ: case MADV_POPULATE_WRITE: - case MADV_COLLAPSE: return MADVISE_MMAP_READ_LOCK; + case MADV_COLLAPSE: + /* + * Only for this process. On another one the range has to be + * untagged with untagged_addr_remote(), which reads mm state + * that mmap_lock protects, before any VMA is looked at. + */ + if (madv_behavior->mm !=3D current->mm) + return MADVISE_MMAP_READ_LOCK; + fallthrough; case MADV_GUARD_INSTALL: case MADV_GUARD_REMOVE: case MADV_DONTNEED: --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6F1F5412272; Sun, 16 Aug 2026 22:47:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920468; cv=none; b=TD9nbbzowzfSj1jMbFVb9gANrsGOBtDnC2wVOX9MJeiS6MKEU6C5WQ/b6lQfOtG2bhi/iyQ8YGFAngGZGGLvvEMD/OwCJ7mR9Llg8gngSQ1V7BtzCrG6EwiUXRm2J2fh081NBQgB86hQzRvKiO+yHvzZ1EyvdSTg7Su+uqba7h4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920468; c=relaxed/simple; bh=Z9R28D2/2GGvJxVw8L0LsZFpJElQsJEGRr2o3etmFwo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ICgpJDjtqvWS87korpp/I4Kjy/3XTCcLB69XeJm0shSy5Kj27IiV5db8+zssvwVRi77xe9/QVi/JAGs5AJcC3dmWOXstAy+6fJholjVhNxG58agkSkRwP37jABzSB9yIjYj0isCZM6HrmZ3N+qhi8KPO0PwfU4YN1PCFx5A6GGE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=eLcO8lyA; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=QVdq39AQ; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="eLcO8lyA"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="QVdq39AQ" Received: from phl-compute-07.internal (phl-compute-07.internal [10.202.2.47]) by mailfhigh.phl.internal (Postfix) with ESMTP id B4E2814000F8; Sun, 16 Aug 2026 18:47:46 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-07.internal (MEProxy); Sun, 16 Aug 2026 18:47:46 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920466; x= 1787006866; bh=kXFeu3sLpTizBLZ9/Sqsc9AjC3ALB9bfMUETq02vuTs=; b=e LcO8lyAfOxw6vsbVSfPPyUPeqcUILzjVePX85zdk+4Qa/6P2XtB0oD6EX0GBe78y SbRWyJTI8wg5nuia8lPlnJdCs9jQtnPEEf7DN2tIrE87uzvWtoTX11UnOzY1V2Wg pn8AW0p+dQKPTmNzhl3ShlZJjkfHNYIRvQklrrI21hmjoY1w2Ui/7pw8jjDSLjCK 8xyE5palO5N0A7kf2LJ2htSwjseT2dNljxr9AVM4VPHjpGDv5YonQC/5UAZyoVsO 3fAcia4nQAwlOptaPlX+P9Vkhi9DFkHzJj5Iyn/auiZZzlnL8pgwLfelj0CaRU2O 4X37RwoyrG+IWPcER01lQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920466; x=1787006866; bh=k XFeu3sLpTizBLZ9/Sqsc9AjC3ALB9bfMUETq02vuTs=; b=QVdq39AQH+YK6d1qn 1wRVHgcj9LE+6Mkh8KyLkGFlc8HTgZ/WJynfrsg9F2dxi6jKEfLWtVOjY6qhqALg iyleh6gleGIDi8nh1HrpMT4+nFB31rAe40Sl9BATfK4rF7GG926621n+bkl4P0EG kbjaMc0aseXlqQzJrXL15CC98sNuEqLlDh+sAOSGgg2DBg+O9NScSikFT7dpSFvR FK0+7DuAr9HGf+U5YDaciPGUHDvWB2C43vQzUuPxgc8Uv2lcnbSwJqSsvrurXt1f yggRh8v1cbDQTWa3AsyItOQmX4nQruKyH1ppeWxQuaSowmdqMuR5zZjhH8LNiq4m Ersvw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTVv qOtBqEf6P9b+C81V84k4WZcgjMziTkY36tS6PSJ/ihWMzxssTaILKX1x0S+EnTMrbWn4tx /WsAGYeCiIsEPje2mEGeHo1a7HuyO9eOPf3nZ82C8aboYTb/FTNw5UJkzusnordR7v9U0S 8QOawJGWzEPx+57GKu5OD7HLlhjh4MjCz7vLnL1lFlqsZN4BRmFISTdNaYhQjG1CUBsVPL PM5z9JW4FnIDlhdJUtMfpaAk0cPHGSL3oVddW6H3HoL/aB4BMDSxUrghJHNG/0+Ign85r1 +HsdiW0bYOGDytlMz9aQRKxqsDuwuKKJh2Kd2LSQZV0Kcr7fx8iysiO4aKjQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:45 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 48/57] mm/collapse: assert the mm reference the engine relies on Date: Sun, 16 Aug 2026 23:46:00 +0100 Message-ID: <20260816224609.308019-49-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Every caller of the engine holds a reference on mm_users for as long as it works: khugepaged for a pass, MADV_COLLAPSE for a call. So mm_users cannot reach zero underneath one, yet collapse.c still asked whether it had, in three places where the answer can only be no. Ask only whether collapsing was turned off, which prctl() can do at any point, and assert the rest where the interface begins. A caller that arrives without a reference is a bug in the caller. A debug build says so there, rather than leaving it to be found when an address space is freed under a collapse. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/collapse.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/mm/collapse.c b/mm/collapse.c index 1e3b2d202ffe..28631c734fcb 100644 --- a/mm/collapse.c +++ b/mm/collapse.c @@ -473,7 +473,7 @@ static enum scan_result collapse_revalidate(struct vm_a= rea_struct *vma, enum scan_result result; unsigned int i, nr_live =3D 0; =20 - if (unlikely(collapse_test_exit_or_disable(mm))) + if (unlikely(collapse_disabled(mm))) return SCAN_ANY_PROCESS; =20 if (!vma->anon_vma || !vma_is_anonymous(vma)) @@ -3731,7 +3731,7 @@ static enum scan_result collapse_file_pmd(struct mm_s= truct *mm, =20 if (result =3D=3D SCAN_PTE_MAPPED_HUGEPAGE) { mmap_read_lock(mm); - if (collapse_test_exit_or_disable(mm)) + if (collapse_disabled(mm)) result =3D SCAN_ANY_PROCESS; else result =3D try_collapse_pte_mapped_thp(mm, addr, @@ -3762,6 +3762,8 @@ bool collapse_scan_pmd(struct vm_area_struct *vma, un= signed long addr, struct mm_struct *mm =3D vma->vm_mm; =20 vma_assert_locked(vma); + /* The caller holds a reference on it, so it cannot have gone away */ + VM_WARN_ON_ONCE(collapse_test_exit(mm)); =20 /* * What the scan answers with, so cleared before it runs. @@ -3776,7 +3778,7 @@ bool collapse_scan_pmd(struct vm_area_struct *vma, un= signed long addr, cc->scan_file =3D NULL; } =20 - if (unlikely(collapse_test_exit_or_disable(mm))) + if (unlikely(collapse_disabled(mm))) cc->scan_refusal =3D SCAN_ANY_PROCESS; else if (addr < vma->vm_start || end > vma->vm_end) cc->scan_refusal =3D SCAN_ADDRESS_RANGE; --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 50A8E4156DA; Sun, 16 Aug 2026 22:47:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920470; cv=none; b=sxXUd+jJ0gNvu2hHQ6c6t+kAX0RH6k2oPuJXt40d8Q5+SKQQLF5+BZY1dg5yRco08UMfzUcl9l8uSRWLqFD/0Q/zHmvkIOUxXYIk85UVV1iA5E9GxD1Mkd9JkTRxWQWeIC7J9Esql5Hs0a2Jm2J0AmIglx9ff+8RwK+eNJTU5+A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920470; c=relaxed/simple; bh=5z8qphyv0VlAuEc2ngiQhIwupgoOuFydmJLL53f7IOc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Cdc/Mcf+wDwzL2ADaf1FZ7Wmo4eGlkKnwvB8NMouRT5E2xF906aofjSCJGmrY3K/3LwBs7G+KlT4eVd0uuyTDORhbRl3/z5cPgYe3uPT2TcJiozdapbXCCNE7ra75hudpHqJ+a5YT//ABRaGq6uM73t+HDUu2fBK6r4gGEYi9yA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=R/D9hrJ/; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=cMcHegD+; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="R/D9hrJ/"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="cMcHegD+" Received: from phl-compute-08.internal (phl-compute-08.internal [10.202.2.48]) by mailfhigh.phl.internal (Postfix) with ESMTP id 918E714000FE; Sun, 16 Aug 2026 18:47:48 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-08.internal (MEProxy); Sun, 16 Aug 2026 18:47:48 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920468; x= 1787006868; bh=RDFiXs9UFefjDWAntFW6mYoilxTkA+zz17UfIDflfSQ=; b=R /D9hrJ/A6K3alQkubIQewlCsyxKk99E71DYl22d+XE0UTS8oqhOmPEXALqrmtPT+ fhuQgCJdw2MUES2F9umkeZgwgY7Q+WO2UprI3k8gGgT9QVoKAYH0w2Yrx3UOpMRK T0EH6CMchU5fZMIz+NZ7j8tTZ0ons48Jar8nhm4n4gHQPLkgbj6fAW6o25vFywcN f7ZtVUKrem5+kj6u9Jezo+JKRRKvfGWl1AtB9D598iwPDoetjHIJBonwWcHev36V fpQ2h5896OHtH4/c0B1nLJ+qk+2/ePN4qBAyPZqaVs1L0VusKect+tsF/9Topjf5 cYlwFEB4mSG/xDuea4zLQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920468; x=1787006868; bh=R DFiXs9UFefjDWAntFW6mYoilxTkA+zz17UfIDflfSQ=; b=cMcHegD+siSteXceA eORchYYJpRg5mkjZpTJ4wtDQd1qhM4icwlRxaOng9Siw3K5cRYYAp7ay9pM9UbSi Crhn6tlhUd9aPwxLjlsga0sRqVLTa65D6RmWWVm0Z1/+jSLfVsxsGUFqZDtsD3da VwPhwjZqGp4gyFSzD4nBVo39qToJWxBQlF52iv/0BRsc4m+s5VJF9th88pwnxahQ zmSTYQb6PCLMAAiX1c1BLtjdxHH943hWM85tvUwHXmJNmBX3SXis9thMwQiemlzn VFWqoNl1GjEteU5c4kuGSgUd1IXh7wV2yHsqdNjtc6VrqiO92KGGjP8LSgVjdJT3 3cy8A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTlc EbUeQzhKxJEmsRwlRPWupCudTuURb990558e34xXePZqhKsdittFqFVB13AvKNBPn+zTIC vmM4mIxGpPzx/nK894dSAKdZpvpkCU2McTpnHaRUIk0aWa95Ge9rfXUXULEztRAndnCYbt 4qD7P3eAyIA5OFPCP2oHhYFpAkBT0mo15kPYR7H5/b3GrpRM2b7kXKyktWAQ4dA2U9USmE fM3n0gonMJM6KinITZmpApKtfVndLpnhEgoZeb0eJJWbRvyWRI94p3CA2ZvaB1HejNnPzR Unmc2v7FvenSNzXmIHyeuDNm+B0nonswhiJxC0bKiq5ZtPtGImr9bjx9A2Bw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:47 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 49/57] mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit() Date: Sun, 16 Aug 2026 23:46:01 +0100 Message-ID: <20260816224609.308019-50-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" __khugepaged_exit() took mmap_lock for writing and dropped it again, to wait for a scan holding it for reading before exit_mmap() destroyed the page tables underneath it. A scan now holds a reference on mm_users for as long as it runs, so it cannot be in flight here at all: __mmput() runs only once that count has reached zero, and asserts so on entry. The barrier is redundant. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/khugepaged.c | 10 ---------- 1 file changed, 10 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 1d77d9a8046d..2b3364bdb100 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -434,16 +434,6 @@ void __khugepaged_exit(struct mm_struct *mm) mm_flags_clear(MMF_VM_HUGEPAGE, mm); mm_slot_free(mm_slot_cache, slot); mmdrop(mm); - } else if (slot) { - /* - * This is required to serialize against - * collapse_test_exit() (which is guaranteed to run - * under mmap_lock read mode). Stop here (after we return all - * pagetables will be destroyed) until khugepaged has finished - * working on the pagetables under the mmap_lock. - */ - mmap_write_lock(mm); - mmap_write_unlock(mm); } } =20 --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0B512417D64; Sun, 16 Aug 2026 22:47:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920472; cv=none; b=IDPHbOG2P4gtKuMNj9GByoe8TpQM0bPSTH1N3IxSmcPKwz8+pwRKCN/couub7tX+IGwducQNYITmNQefnybrVwz4AB9ochuggiMEm1P+VBlJl0zJvqu2ogOVzlHWXpWv8X8sK3aIbc4ETcYRwkEEMt9W0zt9/YlAYrzgjdbd7e8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920472; c=relaxed/simple; bh=xfE9665ZBQHh3APTq8yboXBPsMEMe1PYQzkF4sy2Gu8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=X8Sy5/Er2TDj4iFm8F7qlp4HxT2rtJrBtvqlsnYdC9hsNuWTQCYu/ZautuY7b13fdtFzB2ztK+r0WuzbaWPapsDm6jK0OJ40DGUEy3odrsIRiPLs6Q4LW4DtnHJ2oOyUUvkUkcJerNdk+xBFJrHcX5l8OT5+61zJJ4W1/ykBlko= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=0EkC16hO; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=dghigcQF; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="0EkC16hO"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="dghigcQF" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 4ED26EC0235; Sun, 16 Aug 2026 18:47:50 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:50 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920470; x= 1787006870; bh=PyXPtjPVYE8Z0aMRtUKX037WUeKGrSifkaLdS3sKWUk=; b=0 EkC16hOhVsqh30c3Myj5XL1qP+K3wPn+xQsGWQHAIoL0MoTMTirS5zrh9jwW26t2 fBAEm/dy2wTAwkB/h3ZE3kSUvy9x/4nvh1q27BabBNosXddL+HpAY6fRpJu3BipR 8iqy3eHg6CeNMcQeurXur8Xo7z2x5a/ggI37tXw+K9PtbYpcZmV78hM9IvvkRE2R sOeKu+Tgb9UbXSGTfTBeKA659U+9g9mzfmac24WxtEmGHmzOq0bDGXyXmdcu1z5l cl7cPp1RdfPSNbG2FPTMz9N/DF8rPTqKVf+YY96FVzgS3ay4WBA77eGAzC0g05Vd 832iL1c6w9+8FFo0LmIVg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920470; x=1787006870; bh=P yXPtjPVYE8Z0aMRtUKX037WUeKGrSifkaLdS3sKWUk=; b=dghigcQFSuackbjpf i8ccZXUzsilX+jOQyQEmxsRAaw5KlmG+++n/q8Szwyp9fj5VwwLmt6JWIVLJtQPx 9Z6z5+Rf/ctbSVUqrCVcu9kPxz3jn0iASTx1yvIRYNQGmvMMo4kC+KCy/MZyJasx 564WfCpnJnO754Ya6z0dxyDhnjiSYP8aYao3MoBpoIf1W26oSdF9j3YrbHzlyZza JL4TYl/Rj1x3MEd+h/95I+G2hy9ObHkYRC3fw37KisBD9c7OLokNEIDSjvv0AGHF bhsQO+QwiedX376TuXwmQ9L+MjMyO/QDazWFYKgr/eisS7+NypQvIdmQbbH9uSDG z/IGg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGLjSC+Sq1mG4OxzCqJWZCyq6WBdNxixaUzytYe9+wGJyabxJMeXSGoSuAsbaseo8 j4DSEFeK04hDwZtES01hwyaQU7EBbvB2uUyGWGXfBIvFkincgAhOfMPuQIYJEtCVnAaSQR A4Ax8HD1ONlE2jDjL1mVvD+0hWY7OblS1nu5BjF2+enir154N1oLAu20vju0M/bNfdRnaJ kR8ssngVd0jx9kVZmvlOu6xrNJDHPkWuaI7duTsmkJWDtl1nRnZmYTu08hiimZvSmFoDki 6dltoqAKCLj0c9eKP1Vuuwv4Z4pwJTjTf79O9CtQykYTwcA2QZ7OKyC8oLqq7vFkYshLYo e77wWWIENfWWfKsbKC1kP1m+OxwAAiqZx0bEjZKDuPeV22XOXIGcwfKRFV0FN1rHmveHL0 Y6ejQe4oFZtmmgdQW5dIGSQZXfyUU5nLR+TZExZg0Bcb/lQn5/6SQDajFLx0I0bg5ey1S5 VVBBbzfOElkCyYT+2y9mCdYY69R2UkeyosAzehBiW5C//ulhzKnXpLOIAglVsKGDLTD3qg KWjxi/4iYNGoMuWvB6ZUWuS8mA+ypg9x4kPFB/Msx2X7Cx2DUJIV3qpyFE9KM1niesKiwc ddPOQNpAyzTHzdRzd6gdnR6t2t+NbuyxnuFQ8diZkqJAQ06NMSk0cglrCNug X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:49 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 50/57] selftests/mm: attribute collapses by candidate event alone Date: Sun, 16 Aug 2026 23:46:02 +0100 Message-ID: <20260816224609.308019-51-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The check counts collapses of its window by matching isolate events against the source PFNs it recorded beforehand. That tracepoint went with the mechanism that emitted it, so nothing would match. Count the engine's per-candidate events instead: an install that succeeded, at the window's address and order, is one collapse of that window. Preparing the window still checks that the sources are present, which is the other thing those PFN lookups were doing. The recorded PFNs themselves are no longer needed, so the array goes. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- .../selftests/mm/khugepaged_sync_check.c | 55 +++++++++---------- 1 file changed, 26 insertions(+), 29 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_sync_check.c b/tools/tes= ting/selftests/mm/khugepaged_sync_check.c index 4c37b697d3dd..2a0c247aeec5 100644 --- a/tools/testing/selftests/mm/khugepaged_sync_check.c +++ b/tools/testing/selftests/mm/khugepaged_sync_check.c @@ -7,10 +7,9 @@ * advancing by two is a completion barrier for one full pass that * started after setup (khugepaged_full_pass()). Verify the pair gives * deterministic, attributable results: one barrier step over one - * prepared window produces exactly one collapse attempt on that - * window's source pages (mm_collapse_huge_page_isolate events filtered - * by source PFN and order) and the window is collapsed - * afterwards, repeatably. + * prepared window produces exactly one collapse of that window -- + * mm_collapse_candidate install events at its address and order -- and + * the window is collapsed afterwards, repeatably. * * scan_sleep_millisecs is set to 60s to prove the wake path: without * the wake, one barrier step would sleep multiples of that and blow @@ -51,11 +50,10 @@ static void trace_events_off(void) } =20 /* - * Count collapse attempts attributable to our window: isolate events whose - * scan_pfn is one of the window's source PFNs, reported once per attempt. + * Count the collapses attributable to our window: per-candidate install + * events at the window's address and order, one per collapse. */ -static int count_attributed(unsigned long *pfns, int nr_pfns, - unsigned int order) +static int count_attributed(unsigned long addr, unsigned int order) { char line[1024]; int count =3D 0; @@ -66,25 +64,25 @@ static int count_attributed(unsigned long *pfns, int nr= _pfns, ksft_exit_fail_msg("Cannot open trace buffer\n"); =20 while (fgets(line, sizeof(line), fp)) { + char *s; unsigned long val; unsigned int ord; - char *s, *o; - int i; + char *o; =20 - s =3D strstr(line, "mm_collapse_huge_page_isolate:"); - if (!s) - continue; - if (sscanf(s, "mm_collapse_huge_page_isolate: scan_pfn=3D0x%lx", - &val) !=3D 1) - continue; - o =3D strstr(s, "order=3D"); - if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || ord !=3D order) - continue; - for (i =3D 0; i < nr_pfns; i++) { - if (val =3D=3D pfns[i]) { - count++; - break; - } + s =3D strstr(line, "mm_collapse_candidate:"); + if (s) { + if (!strstr(s, "pass=3Dinstall") || + !strstr(s, "result=3Dsucceeded")) + continue; + o =3D strstr(s, "addr=3D"); + if (!o || sscanf(o, "addr=3D0x%lx", &val) !=3D 1 || + val !=3D addr) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || + ord !=3D order) + continue; + count++; } } fclose(fp); @@ -95,7 +93,6 @@ static void one_step(int iteration) { const size_t window =3D getpagesize() << TARGET_ORDER; const int nr_pages =3D 1 << TARGET_ORDER; - unsigned long pfns[1 << TARGET_ORDER]; bool collapsed, passed; int attributed; char *p; @@ -106,11 +103,11 @@ static void one_step(int iteration) if (p !=3D BASE_ADDR) ksft_exit_fail_perror("mmap() window"); =20 - /* Prepare one window; record its source PFNs. */ + /* Prepare one window, and check the sources really are present. */ for (i =3D 0; i < nr_pages; i++) { p[i * getpagesize()] =3D i + 1; - pfns[i] =3D pagemap_get_pfn(pagemap_fd, p + i * getpagesize()); - if (pfns[i] =3D=3D -1UL) + if (pagemap_get_pfn(pagemap_fd, + p + i * getpagesize()) =3D=3D -1UL) ksft_exit_fail_msg("Source page not present\n"); } =20 @@ -133,7 +130,7 @@ static void one_step(int iteration) =20 collapsed =3D is_range_backed_by_folio_orders(p, window, TARGET_ORDER, pagemap_fd, kpageflags_fd); - attributed =3D count_attributed(pfns, nr_pages, TARGET_ORDER); + attributed =3D count_attributed((unsigned long)p, TARGET_ORDER); =20 ksft_test_result(collapsed && attributed =3D=3D 1, "step %d: window collapsed, %d attributed result(s)\n", --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 05D8E418374; Sun, 16 Aug 2026 22:47:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920474; cv=none; b=fOm56kjEcwcSywrTrDSZiAF2CAgwzVrrtFQER8BaaMVLYWKguF3WdhIT65MD+q4NBxf8THh8T/xf3pnnU6DcyNelp+ns9pKM/eVW2HI4PcjczoIBp3pJv6O3jyekQvVcKc6mSZTcjjLPcyZbY45g6ycbFZc58GZtJxTQgBoHllE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920474; c=relaxed/simple; bh=DCU0OBGX8PtBGbjVBDq2QbdRsHPkXfxxwZ0E3flceEI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=r3+tdw6E0y1pLX6yXrxk3ThSTEX93959Zp69QDbks8aRo4KipolFzJ0RN99Ne8hNEEnZIum2sL0PawMBJRFsR+oCCjShxKyAhdbpylDktwBCLXoznkObQ+VHryxgqKa9aPYyNp4L2PD7a0zHXrLkTDn9htMAB6mqvisyN3ILYcU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=xA337mvy; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=QfjNs5uP; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="xA337mvy"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="QfjNs5uP" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfhigh.phl.internal (Postfix) with ESMTP id 3103814000EE; Sun, 16 Aug 2026 18:47:52 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Sun, 16 Aug 2026 18:47:52 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920472; x= 1787006872; bh=wHV+00xH9B3bJYx8y5slE3mzGKqkEhVlu1wRFthp3Pk=; b=x A337mvyUFlpppzBdiF3Cu5k9EWiW/mPwjKlFQ/EXtIodTzvnUTAOUv3VnRXvYjcP BvsI78mXPc70LuodC4qj7bM30YcVCzUF6bWD25/7lIJKsC/SxkFTfrbAyhixXe/a Ot2cgLPAViVXFX20f1ru0+uCEaCqC272zU0ifxB51M0YvHmHo6xZiuHUCJgaHFop sNiaEmjnxlTqfqW1yUJj6/pkllSKYYaoyyXsWDIaZ6i7iNKsQrvVATmoqc9Pqwsg nfy1mTw7sQCJFJS+XFYlzKSlJqzQofUcHkgoKkMpwRYgrpDu20KHTAbDPusrYQOD WLY8l/lRq1E+tcEixhsuA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920472; x=1787006872; bh=w HV+00xH9B3bJYx8y5slE3mzGKqkEhVlu1wRFthp3Pk=; b=QfjNs5uPIoaSLCjY4 UNvENcFHjTNOENXnBPqTi+RNc22z55HCrvzacXE7SByYPz/otrYquY5kNhc9W/3g LJKZRjgZmIW3rhS35U/WAqh93TYt/66eol9QOdQmjq9k275DoFipa+J8BMQSSkmp uYPcyX7FoE3dxI715BbSY+NfHhll5cGsXZF4Baj2Eu7idCmmIGSdXfquxh3juO5Y u10RldDwrSLco3n7O23TY902qashNBNAUBQi18VDEWR3xNFy+jWdJHukqh+svl2R 0Q+F820OiBKJJYrzAvITwog2ofMeMp2mGGznk4AB5awQjF9JUJ9I7O0pSSD433cS r2H9g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7uwKlsd4UyzlmYX1bnlzO4+qp5TmvoTS/qzkFKfMuj25xYssUPl6UU5Ie9xRba Y0bWuaC/AePVMhqrkCnU3j3fo4+kQ4lrNr0qcDw4ZM2xDmZI+m2r8oZbGDZU/DcLlRMJvk pxa8BemzxDXpf0p5hqW+oOweEwIQBvnr/hc1O3uCrLqxLkUCX+QMvT/AmKStBz4XcOBl17 ynvfoFkCi5Qu6ZndF8QgBKMwimBrechIaqu56LTRaFtWn2N2FSoSPFDDGPAClh8o3Wii3h WmXtU4UIpdlwyanZt6DS7CEKxM1G2yBGNUvAI/9KSggaAyLfyuWlS9a0iep84iGUQaZXfq knD0n9YA/kWmvfl0Wjb8E/Beu3ua4NshP/Nio/tZ6O8p5iMb+wpx4YBhk175j5jhIoNquf +gH+mFtvZJEmhVHsfKOwvOI/Zqd9dLt3FKSCAdmKnioYMErInfqEm+HllRAv8ksF/n6S5S h2uR4Bt/+ZJoTA+GRWlNvAH0+HzPvPGFqXohUpm2Y5V9P1mXIEgx/+CO6eALiZ20+Mpnd1 O0fqm0rJ2cvqo0NujMKgCoRTnm8xRyygGCtnH8y5EQi/fGzr2AjM3K+VHQ7HUZbO016+tv bw5ebY3O4xa6hlnqtSatzZA1NE7UEOHzGO6lgunGBISbTEBpHkcnlEiy4S+w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:51 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 51/57] selftests/mm: cover collapse inside a sub-PMD VMA Date: Sun, 16 Aug 2026 23:46:03 +0100 Message-ID: <20260816224609.308019-52-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The new engine can collapse an mTHP inside a VMA smaller than a PMD. The old mechanism could not: its coverage was rooted at whole PMD-aligned spans, so a VMA that could not hold one was passed over however many mTHP-sized windows it held. On arm64 with 64K pages, where a PMD is 512M, that was every VMA below 512M. Nothing covers that, so add two cases: - a VMA of exactly one window, the smallest thing that can be collapsed at all; - a VMA of several windows, which also walks from one window to the next inside a single VMA. The second takes as many windows as fit in half a table. At the order just below the PMD order two windows are already a whole table, so there is no room for several and the case skips. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 65 +++++++++++++++++++++++++ 1 file changed, 65 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 172e7307eeee..85f138cfb9b2 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1536,6 +1536,69 @@ static void collapse_order_mixed_sources(struct coll= apse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * A VMA smaller than a PMD is a valid collapse target, so long as it hold= s a + * naturally aligned window of the target order. This is the case khugepa= ged + * used to pass over entirely, its coverage being rooted at whole PMD-alig= ned + * spans, which on arm64 with 64K pages meant every VMA below 512M. + */ +static void __collapse_order_sub_pmd_vma(struct collapse_context *c, + struct mem_ops *ops, int nr_windows, + const char *name) +{ + size_t size =3D nr_windows * mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D mmap(BASE_ADDR, size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (p !=3D BASE_ADDR) + ksft_exit_fail_msg("Failed to allocate VMA at %p\n", BASE_ADDR); + + fill_memory(p, 0, size); + if (!window_not_collapsed(p, size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, size, MADV_HUGEPAGE); + ksft_print_msg("Collapse inside a sub-PMD VMA (%d windows)...", + nr_windows); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, size)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, size); + munmap(p, size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", name); +} + +static void collapse_order_sub_pmd_vma(struct collapse_context *c, + struct mem_ops *ops) +{ + __collapse_order_sub_pmd_vma(c, ops, 1, __func__); +} + +static void collapse_order_sub_pmd_range(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + int nr_windows =3D 16; + + while (nr_windows > 1 && nr_windows * window > hpage_pmd_size / 2) + nr_windows /=3D 2; + + if (nr_windows =3D=3D 1) { + ksft_test_result_skip("%s: no room for multiple windows below the PMD\n", + __func__); + return; + } + __collapse_order_sub_pmd_vma(c, ops, nr_windows, __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1823,6 +1886,8 @@ int main(int argc, char **argv) TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_sub_pmd_vma, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_sub_pmd_range, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B5214418A51; Sun, 16 Aug 2026 22:47:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920476; cv=none; b=SwWI5xXbiAiTYJA14K7rOfDauaoVxYkltoccz6EhSJPVSFLkU7DbGmeEBzsQk7et6PoRcSQBJZJ7iSapjKX16tFlEQjo80do/SM6baG9fUAKF07NCCCe+Nme4K9w4bUwlEhrU14ZpIM2agzz9NOSdmzYUKYMiqfAqV4eJmICD5E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920476; c=relaxed/simple; bh=0khGCNAZ79erovRTN27euIO+9Af+Ty/ga1jXRTOxdDQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=MOIVFaXveH245XRB1C2NgQqFYbx9oYw1Sj9zVpqo06xajBEOI3QCz9hKXInbmsbAAjMkelBW4mjHj3+O/nVt8YKeT5nieTLhIF+IZAJ0a5fW3XU13xppSTlCW63SCt+H0/0oQI6Gnu9bcx7YTj/iAwuGGwDOTgOU4r03prT+Z+s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=J2MIBzRg; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=J/P8NbYc; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="J2MIBzRg"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="J/P8NbYc" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id ED57E14000FA; Sun, 16 Aug 2026 18:47:53 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:53 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920473; x= 1787006873; bh=6nKA+x05SWCXqc9tAtZ9fIuQuNqLhmkMozez9tYtOg8=; b=J 2MIBzRgxUP1C4fIjfMGKaTm5D0/70cOQRHKAbA2pFShrgKmGO1NRT0pDfPKhBD15 HXSqf3NsaCIleEj/q4hDHVHX9JGeoP9UY13JE55foNll6OzHp2g5Vk0uFAJIdY4w GoO1Yg6tsF2g+kWdbyG7BWiB5m28OxHZoMTpJ/e+M9RAaV2uFSRL40wNl1JkXpKF H0a9hOB0HY8+RNI0kZp1Bn1uVhNnk0D9LTLf0sAdhXUtaXB/Tg+lxrUj4edzFeL/ njv4WlY3IIA3p8BkcnDozs2Bf6Usp5OKhLCy9kW+QzrUena3ltGTKavDK/nUVo1+ IMeSMRBK+DmhTjOnIrvBg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920473; x=1787006873; bh=6 nKA+x05SWCXqc9tAtZ9fIuQuNqLhmkMozez9tYtOg8=; b=J/P8NbYcP6xCTt5pt yUGsWYlcL5jXSEBSpHo5972USkYsQ4clKoiQmAxZfiapt3giYptu6UTfctr+Ux+Z 3RhxGo/0KNda0UZVE2m3Yruac+La6uxGyjkj6L5QP82mf8mjZK1DODa+rEi5LyAu cTaf8ctK+B1NMVYNCbtVDXY5yhmgL4LRnrsHm85REZi0il4ySqfrtaPo2gkjN/gS NzcbaykTcpMTHuwT1q9a/oStrP2GFiGnKqwUf49YL/jxsoX3XZdYDzyUwCCDAeS6 vMUXpAuCeF89aQ8/g2ihCWhko3penp8Csz0Q401gDWY0aBssIjsDYMRm878T1f1y c3e7w== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGLjSC+Sq1mG4OxzCqJWZCyq6WBdNxixaUzytYe9+wGJyabxJMeXSGoSuAsbaseo8 j4DSEFeK04hDwZtES01hwyaQU7EBbvB2uUyGWGXfBIvFkincgAhOfMPuQIYJEtCVnAaSQR A4Ax8HD1ONlE2jDjL1mVvD+0hWY7OblS1nu5BjF2+enir154N1oLAu20vju0M/bNfdRnaJ kR8ssngVd0jx9kVZmvlOu6xrNJDHPkWuaI7duTsmkJWDtl1nRnZmYTu08hiimZvSmFoDki 6dltoqAKCLj0c9eKP1Vuuwv4Z4pwJTjTf79O9CtQykYTwcA2QZ7OKyC8oLqq7vFkYshLpw LbZeDrUOk8rhM5rSxocjaWFP41pGzRhRHl0flKSZneFSf+mPmKgFKw4251gunztkdKVNyH QnRYsGz5L/GXiKiy1v3S/SrN7GzTBcXp0MYjJsgYyrkQ0uihcwm0LmFwzy8ExXjXwl/DcJ Bqw+WfxPjvkY6ZNwBmZbpr0mZaQ4NNtCOskJtyCw9HD9a7mzzSdUorDoJlDWU180C8lZ54 HZEFuOjndqiCWeC0hQcqgUiK5PiUEeazn7abJ+/umTML33AYxBH0wTQylP6xID9xXksZyN k1UDRxZM6cfUp4CdSMNeruahniVs/40VjqqWeFlJs8fuMuUCsfGpnTv8yfQA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:53 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 52/57] selftests/mm: cover a hole-y window in a sub-PMD VMA Date: Sun, 16 Aug 2026 23:46:04 +0100 Message-ID: <20260816224609.308019-53-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A collapse over a partially populated window has to zero the slots it found empty. The destination comes from the allocator holding whatever was last written to it, and a slot the process never touched must read back zero. Nothing in the suite checks that. validate_memory() only re-reads the pattern fill_memory() wrote, and every caller passes it the faulted extent alone, so a collapse that left stale bytes behind the holes would pass every case here. Check it where population and sub-PMD eligibility meet, which no existing case covers either. One page is faulted in a window-sized VMA; the case expects the window collapsed, the faulted page unchanged, and every byte behind the holes zero. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 50 +++++++++++++++++++++++++ 1 file changed, 50 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 85f138cfb9b2..b61e32566d47 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1599,6 +1599,55 @@ static void collapse_order_sub_pmd_range(struct coll= apse_context *c, __collapse_order_sub_pmd_vma(c, ops, nr_windows, __func__); } =20 +/* + * A partially populated window in a sub-PMD VMA: population and + * sub-PMD eligibility at once. The unfaulted slots must come back + * zero-filled in the collapsed folio, and the faulted ones unchanged. + */ +static void collapse_order_sub_pmd_holes(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t size =3D mthp_window_size(); + char *bytes; + void *p; + size_t i; + + mthp_push_target_order(); + + p =3D mmap(BASE_ADDR, size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + bytes =3D p; + if (p !=3D BASE_ADDR) + ksft_exit_fail_msg("Failed to allocate VMA at %p\n", BASE_ADDR); + + fill_memory(p, 0, page_size); + if (!window_not_collapsed(p, size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, size, MADV_HUGEPAGE); + ksft_print_msg("Collapse hole-y window inside a sub-PMD VMA..."); + if (!khugepaged_wait_full_pass()) { + fail("Timeout"); + } else if (window_collapsed(p, size)) { + /* The unfaulted tail must be zero-filled. */ + for (i =3D page_size; i < size; i++) { + if (bytes[i]) + break; + } + if (i =3D=3D size) + success("OK"); + else + fail("Fail"); + } else { + fail("Fail"); + } + + validate_memory(p, 0, page_size); + munmap(p, size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1888,6 +1937,7 @@ int main(int argc, char **argv) TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); TEST(collapse_order_sub_pmd_vma, mthp_khugepaged_context, anon_ops); TEST(collapse_order_sub_pmd_range, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_sub_pmd_holes, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 73D6641A549; Sun, 16 Aug 2026 22:47:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920477; cv=none; b=evGCdNJh9rmwZXrWyKSs5ew5X6ZMsxRuRQso1MszQppQ1hhViW7YytxHbgYAMOGL3j5L5Ne5sbxip8t+duZtfctPidWOnN3hmUqloWNmdU1y7Py4gGRJ9sEyaGAHxVGUdfVH008ZHEZ8jaB0o+Goey4KX6OBBbOzrcksIIm8BEE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920477; c=relaxed/simple; bh=RDTsAz4Oy7wezsLnEkXHTNcA0qQ8+nrFDvVHPQR253A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=diGR0CuuiO4sf8Jy8gzGUEhjUwip82IN0tVbMg4bEKu0NddkMh+d+Gat+t2QfPwajepSM5qyly/Z3NGwXxiSJ3cDWdfjQ7P6rtVVSS/juE0QThQSWRK9ItJ68vlQXPoP6L7uFYdzTD7c0OYR6fPshjYu+zsAFj26HjRlyK5CYQ4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=HgHGN44b; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=G8U1xcnF; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="HgHGN44b"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="G8U1xcnF" Received: from phl-compute-11.internal (phl-compute-11.internal [10.202.2.51]) by mailfhigh.phl.internal (Postfix) with ESMTP id C871A14000FB; Sun, 16 Aug 2026 18:47:55 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-11.internal (MEProxy); Sun, 16 Aug 2026 18:47:55 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920475; x= 1787006875; bh=YVoKODTNuHH9tsvnWaADZQ9VZLR8etvBsMTlfNVNhEw=; b=H gHGN44biqF8ui1hgPrOaWlVx9x42CW2TrTQ1MBpnq9uZohHvhs9TFO8COYmiBtCQ x092f/nn0bFsT4jxcfIqh4yGHRauPzg3rM62tQd+rliAmvov/yjuGh8SL0B1rWY+ hbGt5HaX9ZWi+CPjnEF7Gpu06s9ZzxKtL25W+JLXAm5m95PwqT58EM1QRCpYQ8Y8 A9SznxPCoxfH4y21kjuZpwwI58y6krWsOdIdjexXUSbQAu/m2ylSEVTAiMpyulSj Mf89HJ7jwuv3gx0AfmaaIxd4QsMTQ3GFYKUXKqbtO6l7ASEFDRt1X9lgRkWYJoMx qHj4B8Rlg2SfjUyO/jiNw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920475; x=1787006875; bh=Y VoKODTNuHH9tsvnWaADZQ9VZLR8etvBsMTlfNVNhEw=; b=G8U1xcnFalpJlNLkp cMJJl5lDt/0JdgftmpIqkcaVJ5istmKoLWtdXShZo572ZMnXxNWs8lUmlTfJuUVZ sJ0oKdAlnJeAEi8RfwsAi7Zow0AhAJpDGHKGFzUYuC9vKVZZ3PpKIcX1iGaU62fD ZbuD7OgEZo5Gu2GQntxsilcY+newcgsOQAAVWUddpN6x81U7xrB9zkrs+XTN8Iau jK2PJm6aPa/liiZhSlJ0HCJGn0lbtkeauX7LWN/kL/puId4wJ7pGTwhBDeEaN/Yb mXxkUN6NBnt2KB3cpwfp7jM+S6qVsdwcameuR29CIdoVXBb0UioyoXaPv5xBveZP C9zig== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTnT D5+QXiNA7V0+//1GAWfYdWnFUDbJ4PKAYPebwd2vIw25LxzaUn6QnwsCjtV0vOBDOG68oD Oqk6cCWO4XKcdDE7GZMuUoxBQVzvFGinAnfQRKXDAj//k/5a2GjRPi+pYfH35b9i57TE0h XMcQaJjSBsGRFRkjnhDLw4f5+2D4nvb6hkAfT4rb9iy+9cz89AU2qog1b4xomiUSaQuLKo YugSHkz9PMS2Kwoq88m1eRQ5YhN8GUeGt4kbhSAlFnxsssMN8NRLY8TPq77Lu3JPp+uR9g JtdJwGX3/UFlAhzoIkkuyY2NCBLzeHHV6BWLrw7FTIxQWDTMxU/KKrZWRpRg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:55 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 53/57] selftests/mm: cover collapse of mlocked ranges Date: Sun, 16 Aug 2026 23:46:05 +0100 Message-ID: <20260816224609.308019-54-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapsing an mlocked range makes the teardown do something it does nowhere else: the sources have to be munlocked while the destination arrives already mlocked, and munlocking takes a reference. So a teardown that reaches a source before it is unfrozen fails on a refcount that is not allowed to move. That is the one ordering constraint in the putback with no other way to be caught. An mlocked range was collapsible before, as long as the whole VMA was locked. This case is the other shape. mlock() over part of a VMA splits it, leaving the locked part smaller than a PMD, which khugepaged passed over for as long as its coverage was rooted at PMD-aligned spans. A partially mlocked region therefore went uncollapsed however long it lived. Cover it deterministically: mlock a window, collapse it, check the contents survive. Drive it under contention too, with a thread mlocking and munlocking random spans across the race harness's region, since the ordering only breaks when a teardown and an mlock overlap. The plain racers never touch VM_LOCKED at all. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 38 ++++++++++++++++++++ tools/testing/selftests/mm/khugepaged_race.c | 29 +++++++++++++-- 2 files changed, 64 insertions(+), 3 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index b61e32566d47..208300ecb344 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1599,6 +1599,43 @@ static void collapse_order_sub_pmd_range(struct coll= apse_context *c, __collapse_order_sub_pmd_vma(c, ops, nr_windows, __func__); } =20 +/* + * Collapse of an mlocked window: source teardown munlocks the old + * pages while the new folio arrives mlocked via folio_add_lru_vma(). + * A teardown that touches the sources while they are still frozen + * blows up exactly here (munlock_folio() takes a reference). + */ +static void collapse_order_mlocked(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, window); + if (mlock(p, window)) + ksft_exit_fail_perror("mlock()"); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse fully populated mlocked window..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, window); + munlock(p, window); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + /* * A partially populated window in a sub-PMD VMA: population and * sub-PMD eligibility at once. The unfaulted slots must come back @@ -1938,6 +1975,7 @@ int main(int argc, char **argv) TEST(collapse_order_sub_pmd_vma, mthp_khugepaged_context, anon_ops); TEST(collapse_order_sub_pmd_range, mthp_khugepaged_context, anon_ops); TEST(collapse_order_sub_pmd_holes, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mlocked, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index 6682bbae0a8f..a4710130aabf 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -219,6 +219,29 @@ static void *forker_fn(void *arg) return NULL; } =20 +/* + * mlock/munlock cycling over the shared areas: collapse of an mlocked + * range munlocks the sources at teardown and mlocks the new folio -- + * the interaction the fuzzer caught (munlock on a frozen source) and + * the plain racers never drove. + */ +static void *mlocker_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + unsigned long nr =3D 1UL << (rand_r(&seed) % 8); /* 1..128 pages */ + + if (rand_r(&seed) & 1) + mlock(region + page_idx * page_size, nr * page_size); + else + munlock(region + page_idx * page_size, nr * page_size); + usleep(rand_r(&seed) % 1000); + } + return NULL; +} + static void *mremapper_fn(void *arg) { unsigned int seed =3D (unsigned long)arg; @@ -328,14 +351,14 @@ int main(int argc, char **argv) { static const char * const thread_names[] =3D { "faulter", "faulter2", "dontneed", "pinner", "forker", - "mremapper", "pageout", "compactor", + "mremapper", "mlocker", "pageout", "compactor", }; void *(*const thread_fns[])(void *) =3D { faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, - mremapper_fn, pageout_fn, compactor_fn, + mremapper_fn, mlocker_fn, pageout_fn, compactor_fn, }; enum { T_FAULTER, T_FAULTER2, T_DONTNEED, T_PINNER, T_FORKER, - T_MREMAPPER, T_PAGEOUT, T_COMPACTOR }; + T_MREMAPPER, T_MLOCKER, T_PAGEOUT, T_COMPACTOR }; const unsigned long pageout_bit =3D 1UL << T_PAGEOUT; const unsigned long compactor_bit =3D 1UL << T_COMPACTOR; const int nr_threads =3D ARRAY_SIZE(thread_names); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 658EF3EE1E5; Sun, 16 Aug 2026 22:47:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920480; cv=none; b=L0O055x+N5SkVjhlknujJ98xoqeEADb2Khe/93MCIgLmmgLft+pHDv9jZ1X9hc5qBleGbkIkhODzJwTH+0nfOXdKLEoUAu2Xv2PhFBUSIVpa9waHT+zKrcFNHJw5qh1Cfluazhvj5rwywNFjb8cwh6kCI7GIGQPwLW3vnB/kV/s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920480; c=relaxed/simple; bh=V2amMiMl7R+h4dDAD6Zy7w5nkocePltmqZVhOP7JMNs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bNASkxTpqSo5VGJ0WWKQRDAtbc9y4L+k4FU1eLRqs0b2tPjpicjBPaQsqIlkwIL3z7D0xOZUanVkVOAJFb1Ezp91F+h+9XE8G/cstHAxkQPXn5pJeE+eZKXXlXnodp5JNZp0ucc73lhCqq0Lqovo9zjqSUz/9rQjzLd6MSmHQF4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=0KrLNBOp; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=fwlDkpgk; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="0KrLNBOp"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="fwlDkpgk" Received: from phl-compute-11.internal (phl-compute-11.internal [10.202.2.51]) by mailfout.phl.internal (Postfix) with ESMTP id 692D5EC0074; Sun, 16 Aug 2026 18:47:57 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-11.internal (MEProxy); Sun, 16 Aug 2026 18:47:57 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920477; x= 1787006877; bh=pI4lY+AJfKOOjXvPiR0jbn4HyML/QwSMTUaiVyOT0As=; b=0 KrLNBOpYOPbnXg5hm7ina1g0ZHCTyNqP/VTFRHVM6sdgshZ2wVN2j3ej2iBzttnh NAFDGpntTkBz66SHNLaWbBOblPFQCR+VRnv2V9DPWGeCv5mV9baYtrFc1BbgMSev HrZSAsUrzTpuDBvQAfhsbtKU1sVtWv77Hpc3bl/55w+FTA5CBFcsQZwa8qNvjS9y UA/9nHSkwNijJDlcDkAjGQMwh4ZuG/UIgy1abfmtDFrbBaRzBdjvSB1iLO7s0ccE iL85a+IykYR7qjP+C2ouXneVPivWKeNXjgoL4dE3WJb4TboXzECfi7IQzvjvYPiC 87hASIGf73zI5ZtJaTqEg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920477; x=1787006877; bh=p I4lY+AJfKOOjXvPiR0jbn4HyML/QwSMTUaiVyOT0As=; b=fwlDkpgkeMZAR/C/C ZHIYGibmR+6BagltMa4AF34xr2PMsy5NRw+2XrmedxKjk3KU+BKe3dGZZ2a/c44F Ten7qAoPkTuUp9RW43eRSwxioiaEELXeJ/LPc9gwYawfdl/BJZbACsP9JDuEf6oB Wtvdge7jFJW7ZBqHTQM0Wyvvq9Ee8bkKb0DQr1cQy6ymOJ6H7+CX+twmnrLXi5Lp nfIWtZ9yeTCawpc4qzNfqq3/aRLrxfoOYUnpz7yK2pHtLUYiLCBuyHR/po/ALpbW yS20oqx9rvGQqNSCztj9W2NjE/igqtmYUoTwny+hAuIYNNMGe7F8zve174pc0AVm HCUnw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFX8QZ1Abcp//ElmtP0wYUsUMnb+KjPA544aOlJXhXcQAudkXa81C9Rj4L11ediC/ SYpDnD1LvcyFEAv2YFQNL3ReIx6t28G7cFOXvsWQqNdQRv24cIelnnbcWrHqur16sY2q+b PHK84xbhjrx0R6ZFbdj3lDcQr+wbYteZ9PHsSvZgjQoovnxRkTlpJXxI+K+aK0DY9MqaWr nM0jSjvJAOtmNeFdTIuI9VxKMOFqeJI9DUUNBafqnTlsaU2K5w1xy/oDJpRmozGKRfcEpJ T4FDpTcqBKNo58tTW6I4nVl91YkYNyFfziYdH6IV5aj62MGBqsSUtSYZnaDODTOndFzTg4 3gj7NdlC8U/37zfIU9cOpYnuRdXqkl9pVUWc3zwBpKJzvj+ZfRQlz7GS0Ejlp3UR3S3OSR Yd0A+lV4lpsWSXukhEFEBp0eGov5qfF8JzT+PPjVG2KIhEnHHbHIxwPmWbz8OxyU++EsAh RsLI86qe1gASncpfU8bRy6B8cY8saqXrUVybVt5nqWgtzw4IJAa+6Td5Hbmb1QRUXlJbJx GkafdYrgnhVsK1vDbEOoEqpzGn0iwNtltfuGQDQQEbiiDEAB8frR5GXA+RVgD6DnVcpjQT 3C3sZmLkZkqv6ZVIOyP1kW5sYHRBNe9OT/zSKApbEQuHaXzBQYMRjgclBtCA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:56 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 54/57] selftests/mm: cover collapse beside a MADV_FREE'd page Date: Sun, 16 Aug 2026 23:46:06 +0100 Message-ID: <20260816224609.308019-55-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A clean lazyfree page must not be collapsed. Copying it into a folio that is not lazyfree would quietly take back memory the process offered to the kernel, and reclaim would no longer be free to drop it. khugepaged refuses the window that holds one. What it should not do is give up on the rest of the table. One page a process no longer needs is a poor reason to leave a whole PMD's worth of memory without large folios. Yet refusing per table is what khugepaged did: any single disqualified PTE ended the scan. So collapse a table with one MADV_FREE'd page in it, and expect three things: the windows beside it collapsed, the window holding it not, and the page itself still backed by an order-0 folio. That last one matters because selection descends orders on a refusal, so checking the target order alone would not notice a smaller window swallowing it. The freed page's contents are not checked -- reclaim is entitled to have dropped them -- but everything else must still read back. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 44 +++++++++++++++++++++++++ 1 file changed, 44 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 208300ecb344..58cb7364652e 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1685,6 +1685,49 @@ static void collapse_order_sub_pmd_holes(struct coll= apse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * One MADV_FREE'd page must not stop the windows beside it from collapsin= g. + * khugepaged still refuses the window holding it: collapsing would copy t= he + * page into a folio that is not lazyfree, quietly making memory the proce= ss + * offered up undroppable again. + */ +static void collapse_order_lazyfree_window(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, hpage_pmd_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + /* Clean and lazyfree: do not touch this page again. */ + if (madvise(p, page_size, MADV_FREE)) + ksft_exit_fail_perror("MADV_FREE"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse the windows beside a MADV_FREE'd page..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p + window, hpage_pmd_size - window) && + window_not_collapsed(p, window) && + /* Left alone at every order, not just the target one */ + is_range_backed_by_folio_orders(p, page_size, 0, + pagemap_fd, kpageflags_fd)) + success("OK"); + else + fail("Fail"); + + /* Everything but the freed page, whose contents may be gone. */ + validate_memory(p, page_size, hpage_pmd_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1976,6 +2019,7 @@ int main(int argc, char **argv) TEST(collapse_order_sub_pmd_range, mthp_khugepaged_context, anon_ops); TEST(collapse_order_sub_pmd_holes, mthp_khugepaged_context, anon_ops); TEST(collapse_order_mlocked, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_lazyfree_window, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C760C3EFFB9; Sun, 16 Aug 2026 22:47:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920481; cv=none; b=MQotroKZdihiPuokIzTQdKRT/AxpQbeFa9k9DtIEnrDkCuCH29KrrGsqPIXhpD8/uSSDJmTfsZYhsopM+rEXK+BFyDK7UGsMzclOJr6JmOGO6GtZalhegVR/oTBu2OKgxUoJiaQRunueLVPYTTBPrbr7YyEBFdRiaWPi/q7cha4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920481; c=relaxed/simple; bh=DrYyq6k0YTeD7AW469vi5eSN+TBdQBREionMKZ3o9aM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CxkIX0DDPXSiLJBTU2xukcGRrbcqr7BLtYr8d6v6PiyKvgGsZyoclyWz/0n393MqtlO73uXW4HXDdaUv5eLapuRZ1c2bbC6YPcmai2IkHBFhR62NEQk9fbeJ2QXTWXUnBR6OtR1j4tvFoxD5+lt+DFPiX8/PqiXlItb6IObjOas= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=RxwFjfkg; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=CUyI9ies; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="RxwFjfkg"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="CUyI9ies" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 3095014000EB; Sun, 16 Aug 2026 18:47:59 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:47:59 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920479; x= 1787006879; bh=VGSfdznQJ8iv193S46pZAZAzBKSWT9DlgvTIxMoYJH0=; b=R xwFjfkgyL4cFAowgG/eRTs7SRD2Hk/lfU+igVpiLcbevYOnK6QpUDMSD6n5iu6ZM zuA72uBY8qQt/aXDlQ5itWOvuicsiloRedf0VOGIk3mJfxalsDJ5N8gxwM/ydKuf GJPr3I5IaQojSIhxRjuFn/gDLC7N2Uj+YOzi4andpPBZkkyOHPN+JjXkMht7voog Xtmqwbd0D1tE6wgjTL+814XNv8cbtKxMgYB5ldwonHdOm7fhBTjpHYg+dUGtjWVN Dpgmhp28+G0lQ77eWwELqxLVOTBMpQAjlBJN7Vumg6XzcLNq199lQHyyHgBf0coh MSwwf3GWO64nJwVo3gqTw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920479; x=1787006879; bh=V GSfdznQJ8iv193S46pZAZAzBKSWT9DlgvTIxMoYJH0=; b=CUyI9iesxRmQIJgIH ZchdKM3r/qX6ZQsnBbNAiee8xozAyBRfIHAyuaTwQADy1uOTpELJY0d0rJMGghuX Vg51RrWYlOewtag1HzkpI3hh962E6G7J7MecOytWOTZ9xsjrNtn6/C6De10EqmPB nXxEBZbWim70I4NSgDzqOk6TpCQhbY+sNOBzThSPqwQMmZwKXlIf1kr1ytHQ1nRy dJwRzDh/Ac+R0Nlw3jSdrkL/Tcc/JtMQtFFa6oAF6I83AgRXGyggSzzpIyak3MFY rw85QW7kfxP9MViQVHIJ5loVxX8lWLFItJo5F7LI0YNFHQMfjoFRB6IKD+cyka8S Iu8lQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFmVpV/hJjZHr0In5rst86Yjj8UzaETlwDLpAB00i6iblejFCjXgd4Yu6W0V/gth1 QQk5g5peqevhLXv5oAa6A4wgFqksr02imAlY52J4/X33aBN3Akp7RIr6ZCjxvFrfIU4xSQ 9AheWiTkiHfncTH5fU2zdbKh7twfazXu6MLOYizqj9hOx9PWjvQu87w255iL0QOT3f2Y/i HNIQn9sEDNwBaHw57xp44AudhIMQHXhs3AVBoLAek/6nuxTCRhfFb4YEz1RsUSKGdN99wZ k2GQ9I2a7t9qkq8S1OweCHxSw7aYXSWFm/gAj0eB9X5rOMd4bXVGhGMj8yPothnhUJpprg zkISQdrFK3cnz1l1Kx1HDHsdctZnnvA7Eni2VQr/C/DD3oP+T4LhVxqhuDiLc9MniGzZA5 2ZAzxrFzhLFV4yStAzNOLiigtaOra4QSZduguqacbRIjtObYgtDf9UQKf+XKXOblOYjTIB su8VPGWXVs5MMz/wTe4Agk4jqScxC/DUeqhsEmIhr25STr6PTaupy6pZdh9Rvf8PdX5Pe+ HQT6QCpQ8lQRKf5ktiR2ZePBoSzbiWb+LUWCLc9KU3LXrUaSQ+BlRQJhjxqDGBC4Og9JNj iTLwf+gO+G6lrH03ovxQ7tQImYZ2cdLwvb56qO61EPdPO8ovPG4mjErqmGkQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:47:58 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 55/57] selftests/mm: cover collapse beside a pinned page Date: Sun, 16 Aug 2026 23:46:07 +0100 Message-ID: <20260816224609.308019-56-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" A GUP pin holds a reference nobody else can account for, so a pinned page cannot be collapsed. The freeze needs the folio's reference count to match what its mappings and cache membership explain, and a pin makes it not. The window holding one has to be refused. The windows beside it should not be, as with a lazyfree page. Pin one page for the duration of a khugepaged pass through gup_test, which makes the refusal deterministic instead of a race. Check the same three things the lazyfree case does: the rest of the table collapsed, that window not, and the pinned page still backed by an order-0 folio, since selection descends orders on a refusal. Skipped without CONFIG_GUP_TEST or the privilege to use it. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 64 +++++++++++++++++++++++++ 1 file changed, 64 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 58cb7364652e..bd684cd25fed 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -21,7 +21,9 @@ =20 #include "linux/magic.h" =20 +#include #include "vm_util.h" +#include "../../../../mm/gup_test.h" #include "hugepage_settings.h" =20 #define BASE_ADDR ((void *)(1UL << 30)) @@ -1728,6 +1730,67 @@ static void collapse_order_lazyfree_window(struct co= llapse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * A GUP pin on one page holds a reference the collapse cannot account for= , so + * khugepaged must leave the window holding it alone -- and collapse the r= est. + * The pin is held across the whole khugepaged pass, which is what makes t= he + * refusal deterministic rather than a race. + */ +static void collapse_order_pinned_window(struct collapse_context *c, + struct mem_ops *ops) +{ + struct pin_longterm_test pin =3D {}; + size_t window =3D mthp_window_size(); + void *p; + int fd; + + fd =3D open("/sys/kernel/debug/gup_test", O_RDWR); + if (fd < 0) { + ksft_test_result_skip("%s: gup_test needs CONFIG_GUP_TEST and root\n", + __func__); + return; + } + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, hpage_pmd_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + pin.addr =3D (__u64)(unsigned long)p; + pin.size =3D page_size; + pin.flags =3D PIN_LONGTERM_TEST_FLAG_USE_WRITE; + if (ioctl(fd, PIN_LONGTERM_TEST_START, &pin)) { + ops->cleanup_area(p, hpage_pmd_size); + close(fd); + thp_pop_settings(); + ksft_test_result_skip("%s: cannot pin\n", __func__); + return; + } + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse the windows beside a pinned page..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p + window, hpage_pmd_size - window) && + window_not_collapsed(p, window) && + /* Left alone at every order, not just the target one */ + is_range_backed_by_folio_orders(p, page_size, 0, + pagemap_fd, kpageflags_fd)) + success("OK"); + else + fail("Fail"); + + ioctl(fd, PIN_LONGTERM_TEST_STOP); + close(fd); + + validate_memory(p, 0, hpage_pmd_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -2020,6 +2083,7 @@ int main(int argc, char **argv) TEST(collapse_order_sub_pmd_holes, mthp_khugepaged_context, anon_ops); TEST(collapse_order_mlocked, mthp_khugepaged_context, anon_ops); TEST(collapse_order_lazyfree_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_pinned_window, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fout-a1-smtp.messagingengine.com (fout-a1-smtp.messagingengine.com [103.168.172.144]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A4FCE41CB39; Sun, 16 Aug 2026 22:48:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.144 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920483; cv=none; b=XE9kGvic00UGjPwMWF9TfdXcYCpJ+D1GKojoXQmItIWjCsKDvGZKEkWKnnB7TQ3XHxVnUE/s3lZTOChSqKWxYxAwBoCtRQr1/hiudUtavjW2bRmhpzy2NrhOXB5mOk/B5zi0fXK4g9A/wHmvzKuXmZolj1jPZq3sSgxpZEwVlPs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920483; c=relaxed/simple; bh=fx4uMDLnQu7YiIWADrslnkPRqCeym9piQn5W0aKziC4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=d6AgVunN4oTbUrBAOpH65D0jIsFXZjwoFKCPuNUMdwzTZmHd5rs5y0Vh1Qt+CdGDZudD0tOOqc32CVoFQL5rrv5Gln5rncj1NE5C4xMhaz3vEyaffrsCaVBAol5yRE9/UhTOaAyeo6OyhUXoZf6JmsXNdYwmYgPhB9XDUf3eyUw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Q1nl18jy; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=HBZrNw5X; arc=none smtp.client-ip=103.168.172.144 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Q1nl18jy"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="HBZrNw5X" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.phl.internal (Postfix) with ESMTP id 080F7EC0242; Sun, 16 Aug 2026 18:48:01 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Sun, 16 Aug 2026 18:48:01 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920481; x= 1787006881; bh=GxeIcugRbXVZfwL3UIFb4N2+J7ON7LziwuBbRMyDD8k=; b=Q 1nl18jyTLYzI3yx4j+jwVY/KO10w2lM3QYtsyjfOG4YOYBwTTnXPm4LRNOpGJ7DC dmKxYGXrkLCioU5z597fuaWe1KK1SoBHdOqe8jVVxx8MR5EmpfiZNAoZN4ptk6Do AKXYSDxDlff8PElQJPNGo9Gwn5a5iZywwlOUcAyZ4qnWz5NmpDoicDuH9UgfEJKu oiDVD5rHzcOmcKWYCWbFK1KMpMgUEKeMjzMZzKOtq5veaoK8uWxWeGfaDkyJ99Bt eGtKMBK+kMBbdqMD4Gydwilbq39tES6G7BP86WARBy1JZkAP/jnsRdituYma7+ac lvecsLUJu69i55sDnPcPA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920481; x=1787006881; bh=G xeIcugRbXVZfwL3UIFb4N2+J7ON7LziwuBbRMyDD8k=; b=HBZrNw5XGXWsoSjY/ l2Ty9l7Y2HQdxrIGH/GmBuDYsahOR7E3ZhJ7jrSBaXi6PaevxI/Qf6bG5rxg+k3K 3UvOZ5FlVYhUKt8uwxUk/RSSKVXcVEz01RWIOyi1e2oSV9DwG6jdLF4CFvyFXKGf fiOCfCaV0iuAw9DbMeMee4rZHJ79Do9HuXLHLTBAzscPu8eXBw+dHBbSq+8GQj8q 5PqkmoLrGd3iwwReLQvYbgkfL1D5f+nq/be1sa9A8Fwi/pdDNVC5hAkCmeOVJ2wl iC4dGym4EQV4Gf/TD9kygnf/BmERNW97vZnmtxLZufRwGB+cbK2zMGe+dliAX3uP bo9ug== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGU6CsvhMdBtF1sno4coJRY9LqZz27w5X62/KcvcmaHPoOW0E1BwYMGQvImxADkeF Vrd/NWvQclh95CfVK4mxB+x5IwibSBkynFlxvp5GF5/z4jepOu57IyLHUvB8DzXymwpMun cyoWwWLP4hjAaD7bSTsoDmYKe03ryBERSg0IuR8ZqJolHrsaFnqWf5mL2byVsL9qLL9zPI 9yLFlvwwZlCiJbT4IVFlRVV/9A2lZNuDFrGWvpQWw3cjqX9EcFrpnMQcn/0MdQygQWXSYC yXHZy2pRvgBssh5qObsL1GM1XSEE6m2uHX2m5CwJlvTSIfOodU1appS1XFI/CBVlhtQ7bu WKduOy5KF3rkObzFMiEUvs0Uj+LfiCh6TivtFsv9ZCNMbNZlGDBoD70ARSLp0XqawKgcWO 4RiWTFPOoz+vIHDroCYLvSEQINMB4dtDTXJ4bPtSoOC1A7Uuzl+tv2HVHx+vABiXweeL5Y B3sGlUM6WH7yS7YE4BSgR6uka3P965uA8/oqj2x+5VFy/Ho1o+YZlJ6AzrIDQ2AOtpBe1f oON1he5PaB+Gj11+twxr/dYpMSHdLv7yT7mnf/dci59b0UPnqF9iMScbflIoYgvrHP20+G HvUS/ei21AytaZzWxcEuOCYs0kXcHbXn/msGaoTEwH1OUh7FsC6dL8NQXoJg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:48:00 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 56/57] selftests/mm: cover the scaled max_ptes_shared limit Date: Sun, 16 Aug 2026 23:46:08 +0100 Message-ID: <20260816224609.308019-57-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" max_ptes_shared is written in PTEs of a whole PMD, but a range smaller than a PMD is scanned in full and judged by the same setting. Compared raw the limit is then unreachable: on arm64 with 64K base pages a 2M range holds 32 PTEs and can never exceed a 4096-PTE budget. So it is scaled to what was actually scanned, and what decides is the shared fraction rather than the count. Cover both directions in one sub-PMD range, with a forked child holding the sources shared. One PTE past the scaled budget: nothing in the range collapses. Break CoW on one more page, bringing it inside: it collapses. The counts come from the current max_ptes_shared rather than being hardcoded, so the case follows the setting and the order it runs at. Without the scaling the first half passes wrongly: a few dozen shared PTEs never reach a limit expressed in PMD units, so the range collapses when it should not. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 85 +++++++++++++++++++++++++ 1 file changed, 85 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index bd684cd25fed..dde24756a3a2 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1791,6 +1791,90 @@ static void collapse_order_pinned_window(struct coll= apse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * max_ptes_shared counts PTEs of a whole PMD, but a VMA smaller than one = is + * scanned in full and judged against the same setting, so the limit is sc= aled + * to the range actually scanned: what decides is the shared *fraction*, n= ot + * the raw count. An unscaled comparison against HPAGE_PMD_NR/2 could never + * refuse a range this small, so both directions are checked here. + */ +static void collapse_order_sub_pmd_shared(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + size_t size =3D 4 * window; + unsigned long nr_ptes =3D size / page_size; + unsigned long budget, cow; + int max_shared, wstatus; + void *p; + + /* + * The range has to stay strictly below a PMD to say anything about the + * scaling: at exactly one PMD the scaled limit is the raw one, and the + * case would pass without testing what it is here for. + */ + if (size >=3D hpage_pmd_size) { + ksft_test_result_skip("%s: four windows do not fit below the PMD\n", + __func__); + return; + } + + max_shared =3D thp_read_num("khugepaged/max_ptes_shared"); + /* The same fraction of this range as max_shared is of a PMD. */ + budget =3D (unsigned long)max_shared * nr_ptes / hpage_pmd_nr; + if (budget + 1 > nr_ptes) { + ksft_test_result_skip("%s: max_ptes_shared leaves nothing to exceed\n", + __func__); + return; + } + + mthp_push_target_order(); + + p =3D mmap(BASE_ADDR, size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (p !=3D BASE_ADDR) + ksft_exit_fail_msg("Failed to allocate VMA at %p\n", BASE_ADDR); + fill_memory(p, 0, size); + madvise(p, size, MADV_HUGEPAGE); + + if (!fork()) { + /* + * Everything is shared with the parent now. Break CoW on all + * but budget + 1 PTEs: one PTE over the scaled limit, and far + * below the unscaled one. + */ + cow =3D nr_ptes - budget - 1; + fill_memory(p, 0, cow * page_size); + ksft_print_msg("Refuse a sub-PMD range over the scaled max_ptes_shared..= ."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_not_collapsed(p, size)) + success("OK"); + else + fail("Fail"); + + /* One fewer shared PTE brings it back within the limit. */ + fill_memory(p, cow * page_size, (cow + 1) * page_size); + ksft_print_msg("Collapse once inside it..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, size)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, size); + _exit(exit_status); + } + wait(&wstatus); + if (WEXITSTATUS(wstatus)) + exit_status =3D WEXITSTATUS(wstatus); + + munmap(p, size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -2084,6 +2168,7 @@ int main(int argc, char **argv) TEST(collapse_order_mlocked, mthp_khugepaged_context, anon_ops); TEST(collapse_order_lazyfree_window, mthp_khugepaged_context, anon_ops); TEST(collapse_order_pinned_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_sub_pmd_shared, mthp_khugepaged_context, anon_ops); } =20 TEST(collapse_full, madvise_context, anon_ops); --=20 2.54.0 From nobody Mon Sep 28 22:32:16 2026 Received: from fhigh-a2-smtp.messagingengine.com (fhigh-a2-smtp.messagingengine.com [103.168.172.153]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC4573E6DEB; Sun, 16 Aug 2026 22:48:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.153 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920485; cv=none; b=PFV+xjoP+H6fMAgZQY4lBsarBgEXiHvlb2paM5iiqpq/eQ2hlmgWCrWyV+O/LOG9XCqcnMIhEL7wmXf0W1YIyl6Rsd7yExtv9d8B16Ml7TPvsk9H5fD74GGYRGT9pwGm5VZfbUxmmf4IVK/ce/MsgPH0KQzB5eBqN9Ifv+xssAs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786920485; c=relaxed/simple; bh=y4Oe55bBTpGkU/aNbfP6J95JKmmpSt53Q4zvyPHYMcs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=qLHvBIoLMfIXCyIktMEZIaiWm6gza8pvPT1LvchtzqpbaBqi9XwY+NLvkUnT+C/KteOM0SL7zvn7bchBYd5u7YSCE+yDZ0RBPZMGSfvLwjbkSjprIBka/kaFggFcYLtxOaobznac/vmjXD3gMbLNA9duyHU5wYuqX/grXmLp9OA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=GFVDPNvE; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Hz4HNAaV; arc=none smtp.client-ip=103.168.172.153 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="GFVDPNvE"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Hz4HNAaV" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id E8B7A14000FD; Sun, 16 Aug 2026 18:48:02 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Sun, 16 Aug 2026 18:48:02 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786920482; x= 1787006882; bh=Iai5LJEqyfwXNhMcgzy+VoGlTqvQD8cplGNN5rO3Gn4=; b=G FVDPNvEdF/lL/bDhKjUQATKlA6pTWX5pDq2Wh5U6aMTM9BzGaSDO688Im2i2YsZr RgZLwS2n+5inMq7Sob+xaV5+TCHHdQG1qDsXXUMDes8OSNoX72MFONQN46UfiGaT gmUPI3kN38+KKJmS39xN5epeDUOgbe34UYQSTXxEscLFjf44HpGbu46bsYtgURrZ hz6oW55iwh5iz4ibfqIA0SlbLWZgVWwx5fgQBGNHTH0sDswgjEjq2xeCLlqIwx7k OjZXiy7IskzyAQXUmhglwRWDYBI7GyO6jyOInK7bkJYi7godKmrDUlhG0svHBspK BKCCz6S1+MQLpB6oJhZ+w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786920482; x=1787006882; bh=I ai5LJEqyfwXNhMcgzy+VoGlTqvQD8cplGNN5rO3Gn4=; b=Hz4HNAaVKEcF2CWJo Om6xPuGHGt2tE6CzXz/iroFOtOXu9jVJ8U1KDKTNufChodRcuF6N/9cR+S45Gc7g U2qTpO/u8LNLszJ+pDJwpQfivAb1vmdzkS/riL3/qXvhsLcKxH0fNGh4G6ViFzFC tlhWV4G6yKXiImzcb3QnFXD/nY+PbuOc5yUw15F6eAa5vnCNtOnYN9ZbnxZuL//0 qgPvJyeIyqZPA3gpeNlSE77SZhJx9bF3g3hQp2HA7hU3CDoVLugXJIMJZSUwUysW d7g98jMRabHLZTxC5QOF9JigdxfvevOF9qGHTsmEueBRNVNgfHonEANULwlpslQ/ iq3cQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEdkpHxwPvghylIRiDh2DXj/W+5rHHxiNSRAV/FnUvF+pTHTSRogpF51K0JMFOoNq FXR3xXOJwqbxi3xnpOR1gVnxkOx0+3lmAqj6HIZkTE7HCsj4rfJi12LdCUf4noPIB2xzxm G7I/52H/Ds2kSEPPWl68t4mZYEFuU9+tOum6r1AuUpSuXeb/EuVzDQV2mgKoQzhJqWy9WS TVItjwfT9Ds7erymUHvJC6I4UFiVodQ+tEdI5cAgcJk8Vajp+hbWnlKpegW2smcmqEpO9y vf97AV5jaic/zQ1Bw2Y9IMD/RUnDCUvC3h/3BgvkY0X8lg0WF0YKtDhM/4K5gFEWm8G2hC o7UOq1ARSQoAFN87IS8L7LmQD0NqIPY7B3lg1mFqX/Ux9m5lhs/bBn3ORbVqzOiq5xkBQH eMvoSuKMYDFCUz5886MTOlInYpH1NMOThGQVJpNdwJoccgzj6LI2NAzcCnEuDStOxdS3Ob t7IJbEzdcGjLWgq+oRueBNy4DT1TJGHM6PcqkdDdDDoeUE9kieAliwATf6dSYT+wGv4zS8 PNNAWvqZZQE8EQod+tB+wj3MtC6rQT+AzOTnzegSlTSw16FRcdIH8OOwF9BXYH+8WCobx4 N+1qrWjy3Npgs0clVyYZ0uus3IX3Vf/faSpYjpWgaZtEl4vpp9FuSha/Ga2Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Sun, 16 Aug 2026 18:48:02 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com, willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: [RFC PATCH 57/57] MAINTAINERS: add an entry for collapse Date: Sun, 16 Aug 2026 23:46:09 +0100 Message-ID: <20260816224609.308019-58-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name> References: <20260816224609.308019-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapse is now its own body of code: mm/collapse.c, a header stating what it offers a caller, and its own tracepoint family. It has two callers, the khugepaged daemon and MADV_COLLAPSE, rather than living inside one of them. Add a section for it, and move what belongs to it out of the THP entry: khugepaged.c and its header, the trace header, the selftests, and mm_slot.h. The selftests are matched as khugepaged*.c, which also covers the race and sync-check harnesses the THP entry did not list. mm_slot.h stays listed under KSM as well, which is the other user of it. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- MAINTAINERS | 19 ++++++++++++++----- 1 file changed, 14 insertions(+), 5 deletions(-) diff --git a/MAINTAINERS b/MAINTAINERS index 7c179b333e4e..4e4e5030b979 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -16955,6 +16955,20 @@ T: git git://git.kernel.org/pub/scm/linux/kernel/g= it/akpm/mm F: include/linux/balloon.h F: mm/balloon.c =20 +MEMORY MANAGEMENT - COLLAPSE +M: Kiryl Shutsemau +L: linux-mm@kvack.org +S: Maintained +W: http://www.linux-mm.org +T: git git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm +F: include/linux/khugepaged.h +F: include/trace/events/collapse.h +F: mm/collapse.c +F: mm/collapse.h +F: mm/khugepaged.c +F: mm/mm_slot.h +F: tools/testing/selftests/mm/khugepaged*.c + MEMORY MANAGEMENT - CORE M: Andrew Morton M: David Hildenbrand @@ -17272,12 +17286,7 @@ T: git git://git.kernel.org/pub/scm/linux/kernel/g= it/akpm/mm F: Documentation/ABI/testing/sysfs-kernel-mm-transparent-hugepage F: Documentation/admin-guide/mm/transhuge.rst F: include/linux/huge_mm.h -F: include/linux/khugepaged.h -F: include/trace/events/collapse.h F: mm/huge_memory.c -F: mm/khugepaged.c -F: mm/mm_slot.h -F: tools/testing/selftests/mm/khugepaged.c F: tools/testing/selftests/mm/split_huge_page_test.c F: tools/testing/selftests/mm/transhuge-stress.c =20 --=20 2.54.0