From nobody Fri Sep 25 14:32:51 2026 Received: from mxhk.zte.com.cn (mxhk.zte.com.cn [160.30.148.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A75CA31282F for ; Fri, 11 Sep 2026 08:07:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=160.30.148.35 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114057; cv=none; b=NMpuvqScu7D71ZBk53l74LEdeDQ/M9I3zzcpSKvwd022aoki8demlJHV1ud0k78o8xiy2SrXE9GWJDFFIyzPeYoLRs6B7JgyppV2GYn1MYJKz1FhTuokm6IxRHlt5TSpRNUE3RCkXPRdlsOIeQrWuFIk8NkK1X1tPsRpxDgihMk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114057; c=relaxed/simple; bh=fbUQbqmMa/yKOEdmf/O6VB/1EdcFnMXyJDrT6AKyp0U=; h=Message-ID:In-Reply-To:References:Date:Mime-Version:From:To:Cc: Subject:Content-Type; b=LpLjI/HEKgUq4dinM5qOToykoaJvVptgLrPTMCD8f9JQY/ud0L8aAer3jPFAcLtjTtFdARPH2n/r6udAbFhisbQxoOaVCEYmIEutqNz0iUgKTbuvtzIziPkXfcPgex4vGmILa1pieyQDOiePOaNBdRMQIIlkz2qG+yXmyEQsQSA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn; spf=pass smtp.mailfrom=zte.com.cn; arc=none smtp.client-ip=160.30.148.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zte.com.cn Received: from mse-fl2.zte.com.cn (unknown [10.5.228.133]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mxhk.zte.com.cn (FangMail) with ESMTPS id 4hh6bT5KV0z8Xrrd; Fri, 11 Sep 2026 16:07:33 +0800 (CST) Received: from xaxapp04.zte.com.cn ([10.99.98.157]) by mse-fl2.zte.com.cn with SMTP id 68B87Ng4028463; Fri, 11 Sep 2026 16:07:23 +0800 (+08) (envelope-from xu.xin16@zte.com.cn) Received: from mapi (xaxapp01[null]) by mapi (Zmail) with MAPI id mid32; Fri, 11 Sep 2026 16:07:25 +0800 (CST) X-Zmail-TransId: 2af96aa3b6bd609-b95f5 X-Mailer: Zmail v1.0 Message-ID: <20260911160725440VlRXkblhz6f7GJgQfbVv3@zte.com.cn> In-Reply-To: <20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn> References: 20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn Date: Fri, 11 Sep 2026 16:07:25 +0800 (CST) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 From: To: , , , Cc: , Subject: =?UTF-8?B?W1BBVENIIDEvNF0gbW0vcGFnZXdhbGs6IGRlbGV0ZSB0aGUgdW51c2VkIG1lbWJlcg==?= X-MAIL: mse-fl2.zte.com.cn 68B87Ng4028463 X-TLS: YES X-ENVELOPE-SENDER: xu.xin16@zte.com.cn X-SOURCE-IP: 10.5.228.133 unknown Fri, 11 Sep 2026 16:07:33 +0800 X-CLEAN: YES X-Fangmail-Anti-Spam-Filtered: true X-Fangmail-MID-QID: 6AA3B6C5.001/4hh6bT5KV0z8Xrrd Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Xu Xin (ZTE) The member vma has never been used, we should delete it Fixes: aa39ca6940f1a ("mm/pagewalk: introduce folio_walk_start() + folio_wa= lk_end()") Signed-off-by: Xu Xin (ZTE) --- include/linux/pagewalk.h | 1 - 1 file changed, 1 deletion(-) diff --git a/include/linux/pagewalk.h b/include/linux/pagewalk.h index b41d7265c01b..1c397be0d092 100644 --- a/include/linux/pagewalk.h +++ b/include/linux/pagewalk.h @@ -183,7 +183,6 @@ struct folio_walk { pmd_t pmd; }; /* private */ - struct vm_area_struct *vma; spinlock_t *ptl; }; --=20 2.25.1 From nobody Fri Sep 25 14:32:51 2026 Received: from mxct.zte.com.cn (mxct.zte.com.cn [183.62.165.209]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0FEB838B7D1 for ; Fri, 11 Sep 2026 08:09:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=183.62.165.209 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114200; cv=none; b=YGedII/PgJ+hAv8HXPMbTr95UUrrvaocu9ZaRRLCpU0L3EMMwv4/dbfAGygrHEpQRe91AnFxuIn5faRkIZr+Tm/IiG9VYTcMVrTVSKNELxCS9RS67pXUCrH0CMU01d1mBe3poMkMd0U3r7CqQ6snrWBqLZRfd74WrNvjNh8lsG0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114200; c=relaxed/simple; bh=W6HeEpQNY03uVRRnRjbY7FH/Z/LeZzURBhcyKL8iNrk=; h=Message-ID:In-Reply-To:References:Date:Mime-Version:From:To:Cc: Subject:Content-Type; b=m/h4k3NRX5mR+AAkgO6Zo52ZNO8ri4zC5mA8LI04xkF2tb8qKt0p9Q9CVjPdF+hHI9CTMLg6fsuru953W0PemCEjy1iSSIUbQaXLkDGUJRsAOGxBNs5BhfcyZhv3eGs1QFXF2flUMR0VVL4q0NzOkfZvt5hMdwhc/em1ryOG56g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn; spf=pass smtp.mailfrom=zte.com.cn; arc=none smtp.client-ip=183.62.165.209 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zte.com.cn Received: from mse-fl1.zte.com.cn (unknown [10.5.228.132]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mxct.zte.com.cn (FangMail) with ESMTPS id 4hh6f24N3gz58fdY; Fri, 11 Sep 2026 16:09:46 +0800 (CST) Received: from xaxapp04.zte.com.cn ([10.99.98.157]) by mse-fl1.zte.com.cn with SMTP id 68B89dvq051060; Fri, 11 Sep 2026 16:09:39 +0800 (+08) (envelope-from xu.xin16@zte.com.cn) Received: from mapi (xaxapp04[null]) by mapi (Zmail) with MAPI id mid32; Fri, 11 Sep 2026 16:09:42 +0800 (CST) X-Zmail-TransId: 2afb6aa3b746af2-9faf8 X-Mailer: Zmail v1.0 Message-ID: <20260911160942109wy2iDTGJslK6rHl8E56si@zte.com.cn> In-Reply-To: <20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn> References: 20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn Date: Fri, 11 Sep 2026 16:09:42 +0800 (CST) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 From: To: , , , , , , , , , , , , , , , , , Cc: , Subject: =?UTF-8?B?W1BBVENIIDIvNF0gbW06IG1ha2UgZm9saW9fd2Fsa19zdGFydCgpJ3MgbG9ja2luZyBhc3NlcnRzIHNjYWxhYmxl?= X-MAIL: mse-fl1.zte.com.cn 68B89dvq051060 X-TLS: YES X-ENVELOPE-SENDER: xu.xin16@zte.com.cn X-SOURCE-IP: 10.5.228.132 unknown Fri, 11 Sep 2026 16:09:46 +0800 X-CLEAN: YES X-Fangmail-Anti-Spam-Filtered: true X-Fangmail-MID-QID: 6AA3B74A.000/4hh6f24N3gz58fdY Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Xu Xin (ZTE) Add an additional member 'walk_lock' to folio_walk to indicate locking requirements for the walk. Similar to commit 49b0638502da0 ("mm: enable page walking API to lock vmas during the walk"). But no change is made on any existing locking behavior, all existing folio_walk_start() callers are currently still under mmap_read_lock() protection. This change is prepared for the latter patch to enable VMA locking asserts. folio_walk_start now operate under write-locked mmap_lock. With introduction of vma locks at the next patch, the vmas have to be locked as well during such walks to prevent concurrent page faults in these areas. No functional change intended. Signed-off-by: Xu Xin (ZTE) --- arch/s390/mm/fault.c | 4 +++- include/linux/pagewalk.h | 1 + kernel/events/uprobes.c | 4 +++- mm/huge_memory.c | 4 +++- mm/ksm.c | 4 +++- mm/migrate.c | 8 ++++++-- mm/pagewalk.c | 10 +++++++++- mm/rmap.c | 4 +++- 8 files changed, 31 insertions(+), 8 deletions(-) diff --git a/arch/s390/mm/fault.c b/arch/s390/mm/fault.c index 46d828926009..968f128a5232 100644 --- a/arch/s390/mm/fault.c +++ b/arch/s390/mm/fault.c @@ -412,7 +412,9 @@ __context_unsafe(/* folio_walk_end() not instrumented *= /) unsigned long addr =3D get_fault_address(regs); struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct folio *folio; int rc; diff --git a/include/linux/pagewalk.h b/include/linux/pagewalk.h index 1c397be0d092..455428eaf1fc 100644 --- a/include/linux/pagewalk.h +++ b/include/linux/pagewalk.h @@ -184,6 +184,7 @@ struct folio_walk { }; /* private */ spinlock_t *ptl; + enum page_walk_lock walk_lock; }; struct folio *folio_walk_start(struct folio_walk *fw, diff --git a/kernel/events/uprobes.c b/kernel/events/uprobes.c index 7709ea882477..054fdcb52136 100644 --- a/kernel/events/uprobes.c +++ b/kernel/events/uprobes.c @@ -507,7 +507,9 @@ int uprobe_write(struct arch_uprobe *auprobe, struct vm= _area_struct *vma, int ret, ref_ctr_updated =3D 0; unsigned int gup_flags =3D FOLL_FORCE; struct mmu_notifier_range range; - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct folio *folio; struct page *page; diff --git a/mm/huge_memory.c b/mm/huge_memory.c index dd66c6ad5af1..b714677e2f20 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -4795,7 +4795,9 @@ static int split_huge_pages_pid(int pid, unsigned lon= g vaddr_start, */ for (addr =3D vaddr_start; addr < vaddr_end; addr +=3D PAGE_SIZE) { struct vm_area_struct *vma =3D vma_lookup(mm, addr); - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct folio *folio; struct address_space *mapping; unsigned int target_order =3D new_order; diff --git a/mm/ksm.c b/mm/ksm.c index 624f37975e12..8df66b4e5de0 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -817,8 +817,10 @@ static struct page *get_mergeable_page(struct ksm_rmap= _item *rmap_item) unsigned long addr =3D rmap_item->address; struct vm_area_struct *vma; struct page *page =3D NULL; - struct folio_walk fw; struct folio *folio; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; mmap_read_lock(mm); vma =3D find_mergeable_vma(mm, addr); diff --git a/mm/migrate.c b/mm/migrate.c index 15b45832bcfa..fa638adfb0de 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -2302,7 +2302,9 @@ static int add_folio_for_migration(struct mm_struct *= mm, const void __user *p, int node, struct list_head *pagelist, bool migrate_all) { struct vm_area_struct *vma; - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct folio *folio; unsigned long addr; int err =3D -EFAULT; @@ -2464,7 +2466,9 @@ static void do_pages_stat_array(struct mm_struct *mm,= unsigned long nr_pages, for (i =3D 0; i < nr_pages; i++) { unsigned long addr =3D (unsigned long)(*pages); struct vm_area_struct *vma; - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct folio *folio; int err =3D -EFAULT; diff --git a/mm/pagewalk.c b/mm/pagewalk.c index 7411702a37f5..8eb29fba20ad 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -910,7 +910,15 @@ struct folio *folio_walk_start(struct folio_walk *fw, pgd_t *pgdp; p4d_t *p4dp; - mmap_assert_locked(vma->vm_mm); + /* + * Other locking modes except for mmap or vma read locking are not + * expected. + */ + if (fw->walk_lock !=3D PGWALK_RDLOCK && fw->walk_lock !=3D PGWALK_VMA_RDL= OCK_VERIFY) + WARN_ONCE(1, "walk_lock is not expected!\n"); + process_mm_walk_lock(vma->vm_mm, fw->walk_lock); + process_vma_walk_lock(vma, fw->walk_lock); + vma_pgtable_walk_begin(vma); if (WARN_ON_ONCE(addr < vma->vm_start || addr >=3D vma->vm_end)) diff --git a/mm/rmap.c b/mm/rmap.c index 5fefe5b060b1..fd8e1c9e4424 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -2871,7 +2871,9 @@ struct page *make_device_exclusive(struct mm_struct *= mm, unsigned long addr, struct mmu_notifier_range range; struct folio *folio, *fw_folio; struct vm_area_struct *vma; - struct folio_walk fw; + struct folio_walk fw =3D { + .walk_lock =3D PGWALK_RDLOCK, + }; struct page *page; swp_entry_t entry; pte_t swp_pte; --=20 2.25.1 From nobody Fri Sep 25 14:32:51 2026 Received: from mxhk.zte.com.cn (mxhk.zte.com.cn [160.30.148.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8D4B23537F9 for ; Fri, 11 Sep 2026 08:12:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=160.30.148.35 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114341; cv=none; b=jueirFBKnG+45iHwYQo88eU10Pn+W3pTbkfSZCYC9MO9qEo3bmWJUPHvBTnklLPZGrhe7lR8IRuvcR/qg8BkvStDDf7wkfh0mb2iP4yiqu/iz5cPeEN7KSCMBxkff/85+Ujt9Md31UbzNZdsbQTbjxN90hUuWtdFnG3Da5Jt9Cc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114341; c=relaxed/simple; bh=xyiHQ4GfGPTe8remfAx9trjKNYwZbt2KB++GbUbhKmc=; h=Message-ID:In-Reply-To:References:Date:Mime-Version:From:To:Cc: Subject:Content-Type; b=g1NCw2Z2XXJZWy0f4JoDEoYtwEz8nZRH9mimXOuPpvohUPv23+3j2PIQMqsdVVoNQZjl6yP/HH8XGsmr8PedXzDQEbxS/X3cfKjr9QNfoJxXa2e71EXm5NmHF8m+qVKCyDZIHAsaHtPA9UmHTseK9ZIR0Fni3AV/ejZtFZuyqg0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn; spf=pass smtp.mailfrom=zte.com.cn; arc=none smtp.client-ip=160.30.148.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zte.com.cn Received: from mse-fl2.zte.com.cn (unknown [10.5.228.133]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mxhk.zte.com.cn (FangMail) with ESMTPS id 4hh6hx5Fzxz8XrrX; Fri, 11 Sep 2026 16:12:17 +0800 (CST) Received: from xaxapp04.zte.com.cn ([10.99.98.157]) by mse-fl2.zte.com.cn with SMTP id 68B8C7Bo036013; Fri, 11 Sep 2026 16:12:08 +0800 (+08) (envelope-from xu.xin16@zte.com.cn) Received: from mapi (xaxapp01[null]) by mapi (Zmail) with MAPI id mid32; Fri, 11 Sep 2026 16:12:10 +0800 (CST) X-Zmail-TransId: 2af96aa3b7da071-c2054 X-Mailer: Zmail v1.0 Message-ID: <20260911161210082WKOqK1dumByDY7jeEOdPF@zte.com.cn> In-Reply-To: <20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn> References: 20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn Date: Fri, 11 Sep 2026 16:12:10 +0800 (CST) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 From: To: , , , Cc: , , Subject: =?UTF-8?B?W1BBVENIIDMvNF0gbW0va3NtOiBtYWtlIGJyZWFrX2tzbSgpIG1vcmUgc2NhbGFibGU=?= X-MAIL: mse-fl2.zte.com.cn 68B8C7Bo036013 X-TLS: YES X-ENVELOPE-SENDER: xu.xin16@zte.com.cn X-SOURCE-IP: 10.5.228.133 unknown Fri, 11 Sep 2026 16:12:17 +0800 X-CLEAN: YES X-Fangmail-Anti-Spam-Filtered: true X-Fangmail-MID-QID: 6AA3B7E1.000/4hh6hx5Fzxz8XrrX Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Xu Xin (ZTE) Currently the last argument 'walk_lock' of break_ksm() is used to indicate whether the page_walk is protected by mmap_read_lock or mmap_write_lock. If 'walk_lock' is true, we suppose its context to be under mmap_write_lock() protection, then mark it PGWALK_WRLOCK and make its vma be write-locked during the walk; If 'walk_lock' is false, we suppose its context to be mmap_read_lock(), then mark it PGWALK_RDLOCK. This change is prepared for the latter patch to enable VMA read-locking where break_ksm() might be under the third new proctecion way: VMA read-locking, so we have to replace the boolean variable to the enum 'page_walk_lock', but without any function changed. No functional change intended. Signed-off-by: Xu Xin (ZTE) --- mm/ksm.c | 21 ++++++++------------- 1 file changed, 8 insertions(+), 13 deletions(-) diff --git a/mm/ksm.c b/mm/ksm.c index 8df66b4e5de0..dda105681d7f 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -660,16 +660,11 @@ static int break_ksm_pmd_entry(pmd_t *pmdp, unsigned = long addr, unsigned long en return found; } -static const struct mm_walk_ops break_ksm_ops =3D { +static struct mm_walk_ops break_ksm_ops =3D { .pmd_entry =3D break_ksm_pmd_entry, .walk_lock =3D PGWALK_RDLOCK, }; -static const struct mm_walk_ops break_ksm_lock_vma_ops =3D { - .pmd_entry =3D break_ksm_pmd_entry, - .walk_lock =3D PGWALK_WRLOCK, -}; - /* * Though it's very tempting to unmerge rmap_items from stable tree rather * than check every pte of a given vma, the locking doesn't quite work for @@ -696,11 +691,11 @@ static const struct mm_walk_ops break_ksm_lock_vma_op= s =3D { * protection keys here anyway. */ static int break_ksm(struct vm_area_struct *vma, unsigned long addr, - unsigned long end, bool lock_vma) + unsigned long end, enum page_walk_lock walk_lock) { vm_fault_t ret =3D 0; - const struct mm_walk_ops *ops =3D lock_vma ? - &break_ksm_lock_vma_ops : &break_ksm_ops; + struct mm_walk_ops *ops =3D &break_ksm_ops; + ops->walk_lock =3D walk_lock; do { int ksm_page; @@ -807,7 +802,7 @@ static void break_cow(struct ksm_rmap_item *rmap_item) mmap_read_lock(mm); vma =3D find_mergeable_vma(mm, addr); if (vma) - break_ksm(vma, addr, addr + PAGE_SIZE, false); + break_ksm(vma, addr, addr + PAGE_SIZE, PGWALK_RDLOCK); mmap_read_unlock(mm); } @@ -1245,7 +1240,7 @@ static int unmerge_and_remove_all_rmap_items(void) for_each_vma(vmi, vma) { if (!(vma->vm_flags & VM_MERGEABLE) || !vma->anon_vma) continue; - err =3D break_ksm(vma, vma->vm_start, vma->vm_end, false); + err =3D break_ksm(vma, vma->vm_start, vma->vm_end, PGWALK_RDLOCK); if (err) goto error; } @@ -2885,7 +2880,7 @@ static int __ksm_del_vma(struct vm_area_struct *vma) return 0; if (vma->anon_vma) { - err =3D break_ksm(vma, vma->vm_start, vma->vm_end, true); + err =3D break_ksm(vma, vma->vm_start, vma->vm_end, PGWALK_WRLOCK); if (err) return err; } @@ -3037,7 +3032,7 @@ int ksm_madvise(struct vm_area_struct *vma, unsigned = long start, return 0; /* just ignore the advice */ if (vma->anon_vma) { - err =3D break_ksm(vma, start, end, true); + err =3D break_ksm(vma, start, end, PGWALK_WRLOCK); if (err) return err; } --=20 2.25.1 From nobody Fri Sep 25 14:32:51 2026 Received: from mxhk.zte.com.cn (mxhk.zte.com.cn [160.30.148.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 357C83E451B for ; Fri, 11 Sep 2026 08:13:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=160.30.148.35 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114426; cv=none; b=ok4e6rpiwVyf0lhi1LQSBhBqNf4JIGRfhXqkeFMcKKmY/9SfpeJEvLt9dXkh0mQZGwebCbbS5BLJdrrONPm9cUgRa3T6ju2rgLLysaRUy/Sh0ww33Hhiy3lwVZ1niWoQgw267ajVTwifSAv0Yd8LuVUbl1tMToZZEIumFQ56gZs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789114426; c=relaxed/simple; bh=1gbI2S+X9LJrLiMsBaj+/HyfW8nwdO9zQHB+Lxf+gu4=; h=Message-ID:In-Reply-To:References:Date:Mime-Version:From:To: Subject:Content-Type; b=f2+bunWT8tX+u2IVp1ldEZY5zxOODa3ov6oCrMMp3SUC5eDoPAxVV1Z6gixvWejwCHX9rPLHY/cHXvc+oP+RF31VbGMayklnJTZ8sf2TYqt8GDSDy07shX5HQKeCq/xDI9p0nltfY665ZFxRksR2/VybqW7A8JQuMoolszLOXV8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn; spf=pass smtp.mailfrom=zte.com.cn; arc=none smtp.client-ip=160.30.148.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=zte.com.cn Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zte.com.cn Received: from mse-fl2.zte.com.cn (unknown [10.5.228.133]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange x25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mxhk.zte.com.cn (FangMail) with ESMTPS id 4hh6kZ2hg5z8Xrrb; Fri, 11 Sep 2026 16:13:42 +0800 (CST) Received: from xaxapp01.zte.com.cn ([10.88.99.176]) by mse-fl2.zte.com.cn with SMTP id 68B8DVsW037964; Fri, 11 Sep 2026 16:13:31 +0800 (+08) (envelope-from xu.xin16@zte.com.cn) Received: from mapi (xaxapp02[null]) by mapi (Zmail) with MAPI id mid32; Fri, 11 Sep 2026 16:13:33 +0800 (CST) X-Zmail-TransId: 2afa6aa3b82db48-b5b38 X-Mailer: Zmail v1.0 Message-ID: <20260911161333316ISAQaj9yJhmglKT5A9OnK@zte.com.cn> In-Reply-To: <20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn> References: 20260911160421076_KNXun8Mpp9Xj7fxHG0i7@zte.com.cn Date: Fri, 11 Sep 2026 16:13:33 +0800 (CST) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 From: To: , , , , , , , , , , , , , , , , , , , Subject: =?UTF-8?B?W1BBVENIIDQvNF0gbW0va3NtOiBhZGQgZmluZF9tZXJnZWFibGVfdm1hX2xvY2tlZCgpIHRvIHVzZSBwZXItVk1BIGxvY2tpbmc=?= X-MAIL: mse-fl2.zte.com.cn 68B8DVsW037964 X-TLS: YES X-ENVELOPE-SENDER: xu.xin16@zte.com.cn X-SOURCE-IP: 10.5.228.133 unknown Fri, 11 Sep 2026 16:13:42 +0800 X-CLEAN: YES X-Fangmail-Anti-Spam-Filtered: true X-Fangmail-MID-QID: 6AA3B836.000/4hh6kZ2hg5z8Xrrb Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Xu Xin (ZTE) Purpose =3D=3D=3D=3D=3D=3D=3D Let's add find_mergeable_vma_locked(), which is similar to find_tcp_vma(), using the universal per-VMA locking helper, so that we can avoid mmap_read_lock() to reduce contention. To be used in KSM code to replace find_mergeable_vma() with mmap_read_lock(), the helper find_mergeable_vma_locked() uses the universal per-VMA locking allowing us to lock a struct vm_area_struct without taking the process-wide mmap lock in read mode. Performance =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D As a test, I construct a testcase which follows the approach: Create one victim and several churner threads sharing one mm_struct; The victim registers a 32 MiB anonymous VM_MERGEABLE region containing 8192 pages totally:churners hammer mmap_lock via mmap/munmap repeatedly; ksmd merges the victim's pages; Compare latency baseline VS this per-VMA patch. Before patched After Patched 0 churner: 1.627 seconds 1.426 seconds 4 churners: 72.45 seconds 36.61 seconds In conclusion, when no mmap_lock contention (0 churner), there is little difference between the baseline kernel and the per-VMA optimized kernel; But under interference from 4 churner threads, the merge time of the per-VMA KSM-optimized kernel is significantly reduced by 50%. Signed-off-by: Xu Xin (ZTE) Reported-by: Longlong Xia Suggested-by: David Hildenbrand --- mm/ksm.c | 64 +++++++++++++++++++++++++++++++++++++------------------- 1 file changed, 42 insertions(+), 22 deletions(-) diff --git a/mm/ksm.c b/mm/ksm.c index dda105681d7f..1d85769ec7db 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -765,15 +765,36 @@ static bool vma_ksm_compatible(struct vm_area_struct = *vma) return ksm_compatible(vma->vm_file, vma->flags); } -static struct vm_area_struct *find_mergeable_vma(struct mm_struct *mm, - unsigned long addr) +/** + * find_mergeable_vma_locked() - Find the VMA covering 'address' which is + * VM_MERGEABLE and read-lock it by per-VMA locks. Please use vma_end_read= () + * to unlock vma after finishing reading the VMA (non-NULL). + * + * Return: If a VMA exists which spans @address, return that VMA, read-loc= ked. + * If no VMA is mapped there or, very unlikely, a reference count overflow + * occurred, return NULL, and no read-locked. + * + * IMPORTANT: If a VMA exists but is not VM_MERGEABLE or has no anon_vma, + * this function releases the per-VMA read lock before returning NULL. + * Callers must NOT call vma_end_read() on a NULL return value. + */ +static struct vm_area_struct *find_mergeable_vma_locked(struct mm_struct *= mm, + unsigned long address) { struct vm_area_struct *vma; + if (ksm_test_exit(mm)) return NULL; - vma =3D vma_lookup(mm, addr); - if (!vma || !(vma->vm_flags & VM_MERGEABLE) || !vma->anon_vma) + + vma =3D vma_start_read_unlocked(mm, address); + if (!vma) + return NULL; + + if (!(vma->vm_flags & VM_MERGEABLE) || !vma->anon_vma) { + vma_end_read(vma); return NULL; + } + return vma; } @@ -799,11 +820,12 @@ static void break_cow(struct ksm_rmap_item *rmap_item) */ rmap_item->linear_page_index =3D 0; - mmap_read_lock(mm); - vma =3D find_mergeable_vma(mm, addr); - if (vma) - break_ksm(vma, addr, addr + PAGE_SIZE, PGWALK_RDLOCK); - mmap_read_unlock(mm); + vma =3D find_mergeable_vma_locked(mm, addr); + if (!vma) + return; + + break_ksm(vma, addr, addr + PAGE_SIZE, PGWALK_VMA_RDLOCK_VERIFY); + vma_end_read(vma); } static struct page *get_mergeable_page(struct ksm_rmap_item *rmap_item) @@ -814,13 +836,12 @@ static struct page *get_mergeable_page(struct ksm_rma= p_item *rmap_item) struct page *page =3D NULL; struct folio *folio; struct folio_walk fw =3D { - .walk_lock =3D PGWALK_RDLOCK, + .walk_lock =3D PGWALK_VMA_RDLOCK_VERIFY, }; - mmap_read_lock(mm); - vma =3D find_mergeable_vma(mm, addr); + vma =3D find_mergeable_vma_locked(mm, addr); if (!vma) - goto out; + return NULL; folio =3D folio_walk_start(&fw, vma, addr, 0); if (folio) { @@ -831,12 +852,12 @@ static struct page *get_mergeable_page(struct ksm_rma= p_item *rmap_item) } folio_walk_end(&fw, vma); } -out: + if (page) { flush_anon_page(vma, page, addr); flush_dcache_page(page); } - mmap_read_unlock(mm); + vma_end_read(vma); return page; } @@ -1568,14 +1589,14 @@ static int try_to_merge_with_zero_page(struct ksm_r= map_item *rmap_item, if (ksm_use_zero_pages && (rmap_item->oldchecksum =3D=3D zero_checksum)) { struct vm_area_struct *vma; - mmap_read_lock(mm); - vma =3D find_mergeable_vma(mm, rmap_item->address); + vma =3D find_mergeable_vma_locked(mm, rmap_item->address); if (vma) { err =3D try_to_merge_one_page(vma, page, ZERO_PAGE(rmap_item->address)); trace_ksm_merge_one_page( page_to_pfn(ZERO_PAGE(rmap_item->address)), rmap_item, mm, err); + vma_end_read(vma); } else { /* * If the vma is out of date, we do not need to @@ -1583,7 +1604,6 @@ static int try_to_merge_with_zero_page(struct ksm_rma= p_item *rmap_item, */ err =3D 0; } - mmap_read_unlock(mm); } return err; @@ -1602,10 +1622,9 @@ static int try_to_merge_with_ksm_page(struct ksm_rma= p_item *rmap_item, struct vm_area_struct *vma; int err =3D -EFAULT; - mmap_read_lock(mm); - vma =3D find_mergeable_vma(mm, rmap_item->address); + vma =3D find_mergeable_vma_locked(mm, rmap_item->address); if (!vma) - goto out; + goto out_trace; err =3D try_to_merge_one_page(vma, page, kpage); if (err) @@ -1625,7 +1644,8 @@ static int try_to_merge_with_ksm_page(struct ksm_rmap= _item *rmap_item, rmap_item->linear_page_index =3D linear_anon_page_index(vma, rmap_item->a= ddress); get_anon_vma(vma->anon_vma); out: - mmap_read_unlock(mm); + vma_end_read(vma); +out_trace: trace_ksm_merge_with_ksm_page(kpage, page_to_pfn(kpage ? kpage : page), rmap_item, mm, err); return err; --=20 2.25.1