From nobody Sat Sep 26 07:16:28 2026 Received: from flow-a3-smtp.messagingengine.com (flow-a3-smtp.messagingengine.com [103.168.172.138]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 245DB389114; Thu, 3 Sep 2026 18:29:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.138 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460203; cv=none; b=X5iPkZDrwsTUFmVVEr9ox/GN3xJTgligk1aVqkxmD4MTazD+hHzh0X0FFDYDKfmE7PayZKD07FQ7+1gFKC3EXhSxcJj/46CnJLhhXq+BqwWz7bbzdmrOsOuTN/Qd6K/SvUuI7Ad363Avv2PMEKmeORg0EvmvSjuarrTqSshtTMM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460203; c=relaxed/simple; bh=scTVfOEUItjdosptvl+093FHTl2lw3ZCIX81xPIg3T4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DbE/bngPP7AoKdSrCUAh8u1Gw6LuGJEYkiKFLT+PVqF6UvBvciaDX3Bpf2pvW0QDzIrfP5iC4SA+m/YtFVDuO01qWflG6xawHUHvX6rI2WlK9ByGKe2/kLnk3emjV51EzsLBsWN9vNyhJIA3+EjRvV9CwuIAxe/SED5WDTugze8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=bJdZVe6I; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=kbCRJOxh; arc=none smtp.client-ip=103.168.172.138 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="bJdZVe6I"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="kbCRJOxh" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailflow.phl.internal (Postfix) with ESMTP id 6F8C61380127; Thu, 3 Sep 2026 14:29:52 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Thu, 03 Sep 2026 14:29:52 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1788460192; x= 1788467392; bh=IXAbT3RhlQhLzfRcazVV9LtEcGIthgKmvu2c4KjzvuY=; b=b JdZVe6I4o5MKZYcnL79xg8tI+zVR1JDmMZgYVFSYtR8SqW35SnCD7g/95joxvX5X aXeiCt5fmfyFzCFRue3W5UV8XRQiKuUdwWOZwCtQqvOZT3DWv3JVv9JcG0K0dy6z +1IVzKjiCAydrk/jYoZ1N3nIx389GaoHaCK2qtKzAVusk0BCT4q9ut3ribRdqaFG 3HOpf7GrmP3YSG5BASFZXnNUluqBvZ5WDJcF3TCq0hn7D6N+dW0IClv4L2Igw1yN 0SRbbPhLWrY43keGKQ6H09PipuQM+A3hWHhhodd3Kn23f1D/Bpqli04sb3NZY55g qRsG7UEihGe/b4uDG3AVw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1788460192; x=1788467392; bh=I XAbT3RhlQhLzfRcazVV9LtEcGIthgKmvu2c4KjzvuY=; b=kbCRJOxh0BQ9soMyK QxcADm17Z9hOm4OlolGPi0SrnBWReUD2ckX/nmzuoJlifUXT6nVBhFveGt7kp86I BnaLnr/WTe6DmCjOeP6bgSptQ3TrHOtcweAgVyJsMM4FEVVowAsHxo9HZh2GYxJP uhkXHeZIiWJoRdjnor/KDYrMnLyDXy8aPN9AaPhwQlrA9WGNax8Z6gR43tQYVFt7 SI8w6ywcrUIaoGlEkDxAbqrzOTz/u+xfjzqrsLxoJU2QUR5Na6pquTBGZUvF+Svo GjArvPzD9Y1qBtyLUEtcEeooiKmLA9xiMHAKAYthIN241qIz4BAMOTGEXD1cJLv/ JZoHg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7tpQ23d6ma5ZKwjaUeV/XmiCXSwfUJTD/UMTaTx+tQ6P98fG4/H3PFA7hJFxaU SY7i6N6X86D7HaHpdbLHvMivEhulOmq3NoSfdg1pBrJA6v+lfq2ALheALUZjz6o7Nf8BIz DWIHLTWCbkC3w9bMwM7vsWb978omVdOmMplVzO2fLHjMRy1E/CxFkmf+2uWv7frAXfMOWG H/vdAXIPTaxMrzbmklMnL5PmRnZSJQIp5M1ihNDwszNqcLBuKJXCEkMPvIWqkfKnePof3W BlQ363pN/8dtwFHD474IAw0kAsBZ9JYvF7V9lEJnX7ro/M4XAHwxRodQrXyCYUn1NLN+PM 0vfdcaBQjLc+HudZy9eIbs9FNURWggrBI09gzzjkvwtrJdKDbJanKkcjjxbA2wXLW5vmPV JHY7vjUWVS+tCavA8PcCB9/d4IboR/rmYFx7WSmkm5HX3rO99EzxgghTPold4UY7N8MLio ckoVvpjoVQYAuSOMa6WiILvVCXgNSpWb0YmrSVSnYN92LolbIS1ZBUCtfZZgzD+yHUUE+T 3FBZrWGSTxV07/NgqDrk6AcNrJhpdb5ML5pWJ3Z4RwgqVY2AvQVvFXk+au178of9NJo1VX Jwuxz1BTVNt4FUyEe/8GbpTWtyYMn6RCqGxKo1ht9sgk2Nnrd0PTSdbQ5Fwg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 3 Sep 2026 14:29:51 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, "Matthew Wilcox (Oracle)" , David Hildenbrand Cc: Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jan Kara , Rik van Riel , Harry Yoo , Lance Yang , Jann Horn , Alexander Viro , Christian Brauner , "Darrick J. Wong" , Carlos Maiolino , Usama Arif , Pedro Falcato , linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Kiryl Shutsemau (Meta)" Subject: [RFC PATCH 1/5] mm: let folio_mkclean() report which pages had dirty PTEs Date: Thu, 3 Sep 2026 19:29:39 +0100 Message-ID: <20260903182943.662461-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name> References: <20260903182943.662461-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" folio_mkclean() walks every mapping of a folio and clears the dirty and write bits of each page table entry. It has the per-entry dirty bit in hand while doing that, but only counts how many entries it cleaned. For a large folio those bits are the only record of which parts of the folio were written through a mapping. Everything downstream has to assume the whole folio changed because that information is dropped here. Add folio_mkclean_dirtymap(), which takes a bitmap and sets a bit for every page of the folio whose entry was dirty. folio_mkclean() becomes a wrapper that passes no bitmap, so there is no change in behaviour yet. A PMD entry has one dirty bit for the whole folio, so a PMD-mapped folio reports all of its pages as dirty. That is the best that can be done: the hardware does not track anything finer for a PMD. Assisted-by: Claude:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/linux/rmap.h | 7 ++++++ mm/rmap.c | 56 +++++++++++++++++++++++++++++++++++--------- 2 files changed, 52 insertions(+), 11 deletions(-) diff --git a/include/linux/rmap.h b/include/linux/rmap.h index 8dc0871e5f00..6fc0a6020252 100644 --- a/include/linux/rmap.h +++ b/include/linux/rmap.h @@ -927,6 +927,7 @@ unsigned long page_address_in_vma(const struct folio *f= olio, * returns the number of cleaned PTEs. */ int folio_mkclean(struct folio *); +int folio_mkclean_dirtymap(struct folio *folio, unsigned long *dirty_map); =20 int mapping_wrprotect_range(struct address_space *mapping, pgoff_t pgoff, unsigned long pfn, unsigned long nr_pages); @@ -990,6 +991,12 @@ static inline int folio_mkclean(struct folio *folio) { return 0; } + +static inline int folio_mkclean_dirtymap(struct folio *folio, + unsigned long *dirty_map) +{ + return 0; +} #endif /* CONFIG_MMU */ =20 #endif /* _LINUX_RMAP_H */ diff --git a/mm/rmap.c b/mm/rmap.c index 1f72d279ba68..aaf45ac79fa8 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1100,12 +1100,18 @@ int folio_referenced(struct folio *folio, int is_lo= cked, return rwc.contended ? -1 : pra.referenced; } =20 -static int page_vma_mkclean_one(struct page_vma_mapped_walk *pvmw) +struct mkclean_state { + unsigned long *dirty_map; + int cleaned; +}; + +static int page_vma_mkclean_one(struct page_vma_mapped_walk *pvmw, + unsigned long *dirty_map) { - int cleaned =3D 0; struct vm_area_struct *vma =3D pvmw->vma; - struct mmu_notifier_range range; unsigned long address =3D pvmw->address; + struct mmu_notifier_range range; + int cleaned =3D 0; =20 /* * We have to assume the worse case ie pmd for invalidation. Note that @@ -1134,6 +1140,15 @@ static int page_vma_mkclean_one(struct page_vma_mapp= ed_walk *pvmw) if (!pte_dirty(entry) && !pte_write(entry)) continue; =20 + if (dirty_map && pte_dirty(entry)) { + pgoff_t idx =3D linear_page_index(vma, address) - + pvmw->pgoff; + + /* The walk only visits pages of this folio */ + VM_WARN_ON_ONCE(idx >=3D pvmw->nr_pages); + __set_bit(idx, dirty_map); + } + flush_cache_page(vma, address, pte_pfn(entry)); entry =3D ptep_clear_flush(vma, address, pte); entry =3D pte_wrprotect(entry); @@ -1155,6 +1170,9 @@ static int page_vma_mkclean_one(struct page_vma_mappe= d_walk *pvmw) if (!pmd_dirty(entry) && !pmd_write(entry)) continue; =20 + if (dirty_map && pmd_dirty(entry)) + bitmap_set(dirty_map, 0, pvmw->nr_pages); + flush_cache_range(vma, address, address + HPAGE_PMD_SIZE); entry =3D pmdp_invalidate(vma, address, pmd); @@ -1181,9 +1199,9 @@ static bool page_mkclean_one(struct folio *folio, str= uct vm_area_struct *vma, unsigned long address, void *arg) { DEFINE_FOLIO_VMA_WALK(pvmw, folio, vma, address, PVMW_SYNC); - int *cleaned =3D arg; + struct mkclean_state *state =3D arg; =20 - *cleaned +=3D page_vma_mkclean_one(&pvmw); + state->cleaned +=3D page_vma_mkclean_one(&pvmw, state->dirty_map); =20 return true; } @@ -1196,12 +1214,23 @@ static bool invalid_mkclean_vma(struct vm_area_stru= ct *vma, void *arg) return true; } =20 -int folio_mkclean(struct folio *folio) +/** + * folio_mkclean_dirtymap - Write-protect a folio and report what was dirt= y. + * @folio: The folio to clean. + * @dirty_map: Bitmap of at least folio_nr_pages(@folio) bits, or NULL. + * + * Write-protects and cleans every mapping of @folio. With @dirty_map, se= ts a + * bit for each page whose entry was dirty; a PMD-mapped folio has one dir= ty + * bit for all of it, so every page is reported. + * + * Return: the number of page table entries cleaned. + */ +int folio_mkclean_dirtymap(struct folio *folio, unsigned long *dirty_map) { - int cleaned =3D 0; + struct mkclean_state state =3D { .dirty_map =3D dirty_map }; struct address_space *mapping; struct rmap_walk_control rwc =3D { - .arg =3D (void *)&cleaned, + .arg =3D (void *)&state, .rmap_one =3D page_mkclean_one, .invalid_vma =3D invalid_mkclean_vma, }; @@ -1217,7 +1246,12 @@ int folio_mkclean(struct folio *folio) =20 rmap_walk(folio, &rwc); =20 - return cleaned; + return state.cleaned; +} + +int folio_mkclean(struct folio *folio) +{ + return folio_mkclean_dirtymap(folio, NULL); } EXPORT_SYMBOL_GPL(folio_mkclean); =20 @@ -1241,7 +1275,7 @@ static bool mapping_wrprotect_range_one(struct folio = *folio, .flags =3D PVMW_SYNC, }; =20 - state->cleaned +=3D page_vma_mkclean_one(&pvmw); + state->cleaned +=3D page_vma_mkclean_one(&pvmw, NULL); =20 return true; } @@ -1324,7 +1358,7 @@ int pfn_mkclean_range(unsigned long pfn, unsigned lon= g nr_pages, pgoff_t pgoff, pvmw.address =3D vma_address(vma, pgoff, nr_pages); VM_BUG_ON_VMA(pvmw.address =3D=3D -EFAULT, vma); =20 - return page_vma_mkclean_one(&pvmw); + return page_vma_mkclean_one(&pvmw, NULL); } =20 static void __folio_mod_stat(struct folio *folio, int nr, int nr_pmdmapped) --=20 2.54.0 From nobody Sat Sep 26 07:16:28 2026 Received: from flow-a3-smtp.messagingengine.com (flow-a3-smtp.messagingengine.com [103.168.172.138]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2262244E652; Thu, 3 Sep 2026 18:29:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.138 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460203; cv=none; b=jk9fUcUaaP/JdyStxYp/WhfGB1xLmVNvXw7L0fA8EKYDf6+nnqBrqg2YTKeCbinMvXr7BxYvQVsfBJqN92ZSk5GGrm9Bvcn0c+Xi4YmVC/A2LMxm3wkRxvGvdKX5Hn0gebKh1V+K8X1rTfmJKsR9mQg6N15kj7anAIQRu0S9Sgs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460203; c=relaxed/simple; bh=s6YmXMhsMpKaQ7k8heM3VQkaEqiVhflX2lUvQHY4Pvc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Htztfji8MkNQQoys6+H8JlASakozSoOgr5yFtJXFhr045qs3dXjdccjVJ1fM5H0uJ0TSe4h+SGAzQW+aCsNfkyDqBDKqPxlCaqFDnt04z7cX5WR578N4hAB8meKW8FsYInhTNBwUnaOYFNyGZytBtvWWb/dO2PPBIG6H71ieMak= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=nuyJvhmf; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=ICPXCBfE; arc=none smtp.client-ip=103.168.172.138 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="nuyJvhmf"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="ICPXCBfE" Received: from phl-compute-07.internal (phl-compute-07.internal [10.202.2.47]) by mailflow.phl.internal (Postfix) with ESMTP id 7103013801B4; Thu, 3 Sep 2026 14:29:54 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-07.internal (MEProxy); Thu, 03 Sep 2026 14:29:54 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1788460194; x= 1788467394; bh=G9VB0sdKC1XKz21nLpi2URn5ZyQ5pGM/85eGOTZORQs=; b=n uyJvhmfYQRHlIHNWbptwLQ+0wR/mTXr50Ep4IVzvnBN9VJAGlJfwTEEGbOSmD0sj 8U2wKGhLv9SdB8zvnkG1GrKXEoegcSc+yGRCvoU6QlwwoNX0JAYKZqUevNj2kVtI G+6hWtTgtgr0ZXHhfFigDUYk2bccnYjGouRAjiRntPappxQmXaLRAqHktmdqEwGk 82QA4MLyofmVulSGwYjzl9+i8FoCN+6TsusGfW+AfcJKRjfH3bFdWDD1sxUm/PMR f3t9i8yR/aq1IBkipjapx+gFYvWVRlH0Ou4YkXutcTEN/usEJiWvt5vM41gG4oNd beF9ced3G57SHnblDGYoA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1788460194; x=1788467394; bh=G 9VB0sdKC1XKz21nLpi2URn5ZyQ5pGM/85eGOTZORQs=; b=ICPXCBfErnIq545ty MYKmeRgwuDeQC/flzKJo3829hY3s8E5hZ4FX7tuxaW0UFig94Rj2yI1NQBVXq5vh 7z1NLgeM3bikWeZqc6wVqhLIJPpvVRZ343ESwgBIlvXN6rWuJKvWieKEBnQR2R8J DdJhtRA81wb+lpJ+rK3zPgjpSoIqZ7eNNTEk9kif/pH82rKSCgblKNbfgQ0C9prX LBayyAyHKW1RNd5MOYwr/Yo30POrvtSI+pQNz9HkKGf5kCQABlHCFhe7kLHoPikP JJhMEAKRJjUn2Ug+cbIr7RAcBb6xgrQ2D8X8lWiB5BTzgGiBvM0TFzBHYsHBOCjK FQYHg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7tpQ23d6ma5ZKwjaUeV/XmiCXSwfUJTD/UMTaTx+tQ6P98fG4/H3PFA7hJFxaU SY7i6N6X86D7HaHpdbLHvMivEhulOmq3NoSfdg1pBrJA6v+lfq2ALheALUZjz6o7Nf8BIz DWIHLTWCbkC3w9bMwM7vsWb978omVdOmMplVzO2fLHjMRy1E/CxFkmf+2uWv7frAXfMOWG H/vdAXIPTaxMrzbmklMnL5PmRnZSJQIp5M1ihNDwszNqcLBuKJXCEkMPvIWqkfKnePof3W BlQ363pN/8dtwFHD474IAw0kAsBZ9JYvF7V9lEJnX7ro/M4XAHwxRodQrXyCYUn1NLN+Uj NdPYldr8Mq2IRZx0YyuQhIoiitm1QZtl4Syt/IqGnJWW5uRScoV3lGNcaWi5MYWySny51c TwILw9eVYuCxGIrYrqLo3zB0wXMnvyhES/hTXL4WkawgjP16PklJlBhJZkkqHSu5ZSeipJ vNQsPjbp1knfB3s+6DOtSIcyjNQSo+fvk74i8TsMuVFro5h5T4RQZDS/9sWdiKOQTAS3KJ Q+KPaFI3YoSacqM82ZxFZTjYhHPwfNu0Ug3d4Hx5T5s+kXfQFEYWmgo290U3VCybc+7Eq5 +rvAeQmDrD+p82sLjT/fTXuwlbtObV0rjSi9zthMlSSAgvNSR/aZPKslN9wg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 3 Sep 2026 14:29:53 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, "Matthew Wilcox (Oracle)" , David Hildenbrand Cc: Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jan Kara , Rik van Riel , Harry Yoo , Lance Yang , Jann Horn , Alexander Viro , Christian Brauner , "Darrick J. Wong" , Carlos Maiolino , Usama Arif , Pedro Falcato , linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Kiryl Shutsemau (Meta)" Subject: [RFC PATCH 2/5] mm: add a_ops->dirty_folio_range() and use the mkclean dirty harvest Date: Thu, 3 Sep 2026 19:29:40 +0100 Message-ID: <20260903182943.662461-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name> References: <20260903182943.662461-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Every way of dirtying part of a folio through a mapping ends up at folio_mark_dirty(), which has no way to say which part changed, so a_ops->dirty_folio() dirties all of it. A filesystem that tracks dirty state per block then writes back the whole folio for a single stored byte. Add a_ops->dirty_folio_range() and folio_mark_dirty_range() to pass the range on. The new operation can express everything a_ops->dirty_folio() can, so folio_mark_dirty() goes through it with a range covering the folio and a filesystem needs only one of the two. Filesystems without it dirty the whole folio. Use it in folio_clear_dirty_for_io(), where the page table dirty bits were being turned into a whole-folio dirty. It now collects them with folio_mkclean_dirtymap() and hands the filesystem the runs that were dirty. A folio that is not already dirty still dirties whole, because the clean to dirty transition needs the accounting in folio_mark_dirty(). Assisted-by: Claude:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- include/linux/fs.h | 3 ++ include/linux/mm.h | 1 + mm/page-writeback.c | 87 +++++++++++++++++++++++++++++++++++++++++---- 3 files changed, 85 insertions(+), 6 deletions(-) diff --git a/include/linux/fs.h b/include/linux/fs.h index 072d8cd09a0b..1d98b6c0b880 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -406,6 +406,9 @@ struct address_space_operations { =20 /* Mark a folio dirty. Return true if this dirtied it */ bool (*dirty_folio)(struct address_space *, struct folio *); + /* Mark [off, off + len) of a folio dirty */ + bool (*dirty_folio_range)(struct address_space *mapping, + struct folio *folio, size_t off, size_t len); =20 void (*readahead)(struct readahead_control *); =20 diff --git a/include/linux/mm.h b/include/linux/mm.h index 87feaa5a2b78..7628262c17e1 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -3337,6 +3337,7 @@ struct kvec; struct page *get_dump_page(unsigned long addr, int *locked); =20 bool folio_mark_dirty(struct folio *folio); +bool folio_mark_dirty_range(struct folio *folio, size_t off, size_t len); bool folio_mark_dirty_lock(struct folio *folio); bool set_page_dirty(struct page *page); int set_page_dirty_lock(struct page *page); diff --git a/mm/page-writeback.c b/mm/page-writeback.c index 6c9c7ba89b8a..39b54c25a9fa 100644 --- a/mm/page-writeback.c +++ b/mm/page-writeback.c @@ -2751,6 +2751,21 @@ bool folio_redirty_for_writepage(struct writeback_co= ntrol *wbc, } EXPORT_SYMBOL(folio_redirty_for_writepage); =20 +/* + * Hand a dirtied range of @folio to the filesystem. ->dirty_folio_range(= ) can + * express everything ->dirty_folio() can, so a filesystem that implements= it + * does not need both, and a whole-folio dirty comes through here as a ran= ge + * covering the folio. + */ +static bool mapping_dirty_range(struct address_space *mapping, + struct folio *folio, size_t off, size_t len) +{ + if (!mapping->a_ops->dirty_folio_range) + return mapping->a_ops->dirty_folio(mapping, folio); + + return mapping->a_ops->dirty_folio_range(mapping, folio, off, len); +} + /** * folio_mark_dirty - Mark a folio as being modified. * @folio: The folio. @@ -2782,13 +2797,39 @@ bool folio_mark_dirty(struct folio *folio) */ if (folio_test_reclaim(folio)) folio_clear_reclaim(folio); - return mapping->a_ops->dirty_folio(mapping, folio); + return mapping_dirty_range(mapping, folio, 0, + folio_size(folio)); } =20 return noop_dirty_folio(mapping, folio); } EXPORT_SYMBOL(folio_mark_dirty); =20 +/** + * folio_mark_dirty_range - Mark part of a folio as being modified. + * @folio: The folio. + * @off: Offset of the modified range within the folio. + * @len: Length of the modified range. + * + * Like folio_mark_dirty(), but tells a filesystem that tracks dirty state= per + * block that only [@off, @off + @len) changed, so writeback can skip the = rest + * of the folio. Filesystems without that tracking dirty the whole folio. + * + * Return: True if the folio was newly dirtied, false if it was already di= rty. + */ +bool folio_mark_dirty_range(struct folio *folio, size_t off, size_t len) +{ + struct address_space *mapping =3D folio_mapping(folio); + + if (likely(mapping)) { + if (folio_test_reclaim(folio)) + folio_clear_reclaim(folio); + return mapping_dirty_range(mapping, folio, off, len); + } + + return noop_dirty_folio(mapping, folio); +} + /* * folio_mark_dirty() is racy if the caller has no reference against * folio->mapping->host, and if the folio is unlocked. This is because an= other @@ -2844,6 +2885,41 @@ void __folio_cancel_dirty(struct folio *folio) } EXPORT_SYMBOL(__folio_cancel_dirty); =20 +/* + * Write-protect every mapping of @folio and hand the filesystem the parts= that + * were dirty in a page table. + * + * Without ->dirty_folio_range() there is nowhere to put per-block state, = so + * any PTE dirty bit dirties the whole folio. Same when the folio is not + * already dirty, because then the dirty transition needs the full account= ing + * in folio_mark_dirty(), and for a folio too large for the bitmap, which = the + * page cache does not make. + */ +static void folio_mkclean_for_io(struct folio *folio, + struct address_space *mapping) +{ + DECLARE_BITMAP(map, 1UL << MAX_PAGECACHE_ORDER); + unsigned int nr =3D folio_nr_pages(folio); + unsigned int start, end; + + if (!mapping->a_ops->dirty_folio_range || !folio_test_dirty(folio) || + WARN_ON_ONCE(nr > (1UL << MAX_PAGECACHE_ORDER))) { + if (folio_mkclean(folio)) + folio_mark_dirty(folio); + return; + } + + bitmap_zero(map, nr); + if (!folio_mkclean_dirtymap(folio, map)) + return; + + for_each_set_bitrange(start, end, map, nr) { + mapping_dirty_range(mapping, folio, + (size_t)start << PAGE_SHIFT, + (size_t)(end - start) << PAGE_SHIFT); + } +} + /* * Clear a folio's dirty flag, while caring for dirty memory accounting. * Returns true if the folio was previously dirty. @@ -2875,9 +2951,9 @@ bool folio_clear_dirty_for_io(struct folio *folio) * * We use this sequence to make sure that * (a) we account for dirty stats properly - * (b) we tell the low-level filesystem to - * mark the whole folio dirty if it was - * dirty in a pagetable. Only to then + * (b) we tell the low-level filesystem which + * parts of the folio were dirty in a + * pagetable. Only to then * (c) clean the folio again and return 1 to * cause the writeback. * @@ -2895,8 +2971,7 @@ bool folio_clear_dirty_for_io(struct folio *folio) * as a serialization point for all the different * threads doing their things. */ - if (folio_mkclean(folio)) - folio_mark_dirty(folio); + folio_mkclean_for_io(folio, mapping); /* * We carefully synchronise fault handlers against * installing a dirty pte and marking the folio dirty --=20 2.54.0 From nobody Sat Sep 26 07:16:28 2026 Received: from flow-a3-smtp.messagingengine.com (flow-a3-smtp.messagingengine.com [103.168.172.138]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7E9AE4D90D4; Thu, 3 Sep 2026 18:30:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.138 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460207; cv=none; b=DdydLNKZhCVqO9ur0Zk2G9GQiCLALwHtxcXoKIWq4lRGwxd/kyQ6b2TN6MKN53MjZKvA9L5Bn9ecrAZFGNhelDi4IdYPlyBI7PNruqCnl4g0j7zPC3goeFkdid3xFldiyZw1Ewa9EYlmWlOx19+ApWimZskQSSy/bIv324NhNLk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460207; c=relaxed/simple; bh=AGl0yqzbyZIZ3+w9gehbrVCg4atEJD8pd5d/L+GIIrk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lEiZUTEL3TDLyEzwo45rwWx13ub5hDDSRgUOFtCaUgnZDjUx7UNP7aQ6SQ3Vxq2IM/EXMxPufvviAqdP4SLfr0z381PAM/PuE0CzW8XeyEU4NzS0zWAr9KySXwb6z0hQ7cvi3fSSxcTmC6CyxpJKgYM7uNu/vN+w61/joT3scG8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=MsuyO2Pz; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DbzgCr/L; arc=none smtp.client-ip=103.168.172.138 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="MsuyO2Pz"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DbzgCr/L" Received: from phl-compute-12.internal (phl-compute-12.internal [10.202.2.52]) by mailflow.phl.internal (Postfix) with ESMTP id 6205213801BF; Thu, 3 Sep 2026 14:29:56 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-12.internal (MEProxy); Thu, 03 Sep 2026 14:29:56 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1788460196; x= 1788467396; bh=YaDn1Mr2KlX00kBdrNWD4+vAoqqgrXa7uJO/7qm/H4k=; b=M suyO2PzQqzmatJVuanDQjvKAb46heIMEBntMOnbccW2xwfQ7TL4F5/zuk7rQyUJS rd5/xkqFWmEWenrjDe/V4e/M7dcmR4+TQ8E0SRvNJ6IAkv+n4IKK4sS53vNWigOP nWhIeMS0/gRlzHUx+MIx+6a28K9TfOqLMu2BW+X4LWTCYD+ikvyDBKR3Uo1/Wo/T jI9QP9z+sZJLxVQ769Y+WayAEtFibHUfqPrS8gSf/zCmbLmZF2dTxYjrbpw8y+rI N31IBypoX7BjgTu2pCRXTtyUxqsiwsJnYdM+2P2mTMUUE+Khy+Y/mg3enbD66MGI 7Dd+hHVuD2my0zHurTXEQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1788460196; x=1788467396; bh=Y aDn1Mr2KlX00kBdrNWD4+vAoqqgrXa7uJO/7qm/H4k=; b=DbzgCr/Le1sX11vVP 7GmorVE4FqVpz6jCtJaCoTYDUUaOSIPXO8DJ7CKC/LbnppB56MXf0otj2AskYUxp YdFxPXYda3pL/fKW9u6CU8/4UUGEx6373kLAjc0vvyCtbGu0qkkCDFdknBtgmY52 /R7oem3dvIsTrcmtZoDfLxlRpJXEfAzgU9ngP+dmS4hgSQmRViv5rkyq9XCCCMIi yrCuYj/FuL4CVvhNtDxg1GB10I8ZJ1g15nTKV9ye++E2i19Ow6jwUC8ldnkZFmKg 5Pj5z+Cc9nxtSMZXOD/950oU7spRLNEJZ1Nn55a7F6HByJPQmo+IU3M92QhyltiP RXdeg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7tpQ23d6ma5ZKwjaUeV/XmiCXSwfUJTD/UMTaTx+tQ6P98fG4/H3PFA7hJFxaU SY7i6N6X86D7HaHpdbLHvMivEhulOmq3NoSfdg1pBrJA6v+lfq2ALheALUZjz6o7Nf8BIz DWIHLTWCbkC3w9bMwM7vsWb978omVdOmMplVzO2fLHjMRy1E/CxFkmf+2uWv7frAXfMOWG H/vdAXIPTaxMrzbmklMnL5PmRnZSJQIp5M1ihNDwszNqcLBuKJXCEkMPvIWqkfKnePof3W BlQ363pN/8dtwFHD474IAw0kAsBZ9JYvF7V9lEJnX7ro/M4XAHwxRodQrXyCYUn1NLN+V2 WlXbIyuCNTMYS9ZmYmi7Wi/90XXuwuQ/NA/yJbaBk7n4YmF/Fa7zOdJg1WdYIjDOJdp6hh ptgs59nDPCUW9de3B9V4/fpiGW/B1vB3mzc/YQWuOYT279QwjVDtwX05wgL7hFPCYvgkEC dRxtjhl5ExCKf1Ug6ah2j8izXWrb0RpAwbstfBWXHLZvPuUhLVmo9ZXao281T/269HgiaM BF2KNnURXUG9XwSdzbaOR26vJd7tKUTSlkp/0SwBs4DAf3or5HjA/lXjm8s22buqgPuj4W neilQdGiiklz/Q2Bt/PY2bT3OTIIocN02kwpbSSn+Pc+3VDCbU+3378snzgg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 3 Sep 2026 14:29:55 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, "Matthew Wilcox (Oracle)" , David Hildenbrand Cc: Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jan Kara , Rik van Riel , Harry Yoo , Lance Yang , Jann Horn , Alexander Viro , Christian Brauner , "Darrick J. Wong" , Carlos Maiolino , Usama Arif , Pedro Falcato , linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Kiryl Shutsemau (Meta)" Subject: [RFC PATCH 3/5] mm: keep the mmap dirty range down to the faulting page Date: Thu, 3 Sep 2026 19:29:41 +0100 Message-ID: <20260903182943.662461-4-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name> References: <20260903182943.662461-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Two places in the fault path dirty a whole folio when a single page was written. fault_dirty_shared_page() marks the folio dirty after a write fault on a shared mapping. Only the page the fault was for has been stored to, so pass that range instead. A PMD-mapped folio is still dirtied whole, since a PMD has one dirty bit and there is nothing finer to report. set_pte_range() marks a whole batch of entries dirty on a write fault, which destroys the per-page dirty information before the pages have been written to. Leave the batch clean for a shared mapping of a filesystem that implements a_ops->dirty_folio_range(), and let the hardware set the bit on the first store to each page. Anywhere else nothing reads those bits, and on architectures without a hardware dirty bit a clean batch costs a fault per page for nothing. That covers shmem, which has no dirty state below the folio. The page the fault was for is dirtied right away. It is about to be stored to, so leaving it clean would only move the work to a second page table walk or fault. Assisted-by: Claude:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- mm/memory.c | 58 +++++++++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 54 insertions(+), 4 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index 8da0f945141b..27a059e0c016 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -3778,10 +3778,19 @@ static vm_fault_t fault_dirty_shared_page(struct vm= _fault *vmf) struct vm_area_struct *vma =3D vmf->vma; struct address_space *mapping; struct folio *folio =3D page_folio(vmf->page); + size_t off =3D folio_page_idx(folio, vmf->page) << PAGE_SHIFT; bool dirtied; bool page_mkwrite =3D vma->vm_ops && vma->vm_ops->page_mkwrite; =20 - dirtied =3D folio_mark_dirty(folio); + /* + * A PMD entry has one dirty bit for the whole folio, so there is no + * finer information to pass on. A PTE-mapped folio only has the page + * the fault was for dirtied so far. + */ + if (pmd_trans_huge(pmdp_get_lockless(vmf->pmd))) + dirtied =3D folio_mark_dirty(folio); + else + dirtied =3D folio_mark_dirty_range(folio, off, PAGE_SIZE); VM_BUG_ON_FOLIO(folio_test_anon(folio), folio); /* * Take a local copy of the address_space - folio.mapping may be zeroed @@ -5625,6 +5634,31 @@ vm_fault_t do_set_pmd(struct vm_fault *vmf, struct f= olio *folio, struct page *pa } #endif =20 +/* + * May the whole batch be marked dirty? + * + * Only a shared mapping of a filesystem that tracks dirty state per block= says + * no. There the page table dirty bits are the only record of which parts = of a + * large folio were written through the mapping, so pages nobody has store= d to + * have to stay clean and let the hardware set the bit on the first store. + * + * Everywhere else nothing ever reads those bits, and on hardware without a + * dirty bit leaving them clean costs a fault per page for nothing. That c= overs + * shmem, which has no dirty state below the folio. + */ +static bool can_dirty_whole_batch(struct vm_fault *vmf, unsigned int nr, + bool prefault) +{ + struct vm_area_struct *vma =3D vmf->vma; + + if (!(vmf->flags & FAULT_FLAG_WRITE) || prefault || nr =3D=3D 1) + return true; + if (!(vma->vm_flags & VM_SHARED)) + return true; + + return !vma->vm_file->f_mapping->a_ops->dirty_folio_range; +} + /** * set_pte_range - Set a range of PTEs to point to pages in a folio. * @vmf: Fault description. @@ -5639,6 +5673,7 @@ void set_pte_range(struct vm_fault *vmf, struct folio= *folio, struct vm_area_struct *vma =3D vmf->vma; bool write =3D vmf->flags & FAULT_FLAG_WRITE; bool prefault =3D !in_range(vmf->address, addr, nr * PAGE_SIZE); + bool dirty_batch =3D can_dirty_whole_batch(vmf, nr, prefault); pte_t entry; =20 flush_icache_pages(vma, page, nr); @@ -5649,10 +5684,13 @@ void set_pte_range(struct vm_fault *vmf, struct fol= io *folio, else entry =3D pte_sw_mkyoung(entry); =20 - if (write) - entry =3D maybe_mkwrite(pte_mkdirty(entry), vma); - else if (pte_write(entry) && folio_test_dirty(folio)) + if (write) { + if (dirty_batch) + entry =3D pte_mkdirty(entry); + entry =3D maybe_mkwrite(entry, vma); + } else if (pte_write(entry) && folio_test_dirty(folio)) { entry =3D pte_mkdirty(entry); + } if (unlikely(vmf_orig_pte_uffd_wp(vmf))) entry =3D pte_mkuffd(entry); /* copy-on-write page */ @@ -5665,6 +5703,18 @@ void set_pte_range(struct vm_fault *vmf, struct foli= o *folio, } set_ptes(vma->vm_mm, addr, vmf->pte, entry, nr); =20 + /* + * The page the fault was for is about to be stored to, so dirty it + * here rather than leave the store to a second page table walk, or to + * a second fault where the dirty bit is maintained in software. + */ + if (!dirty_batch) { + pte_t *ptep =3D vmf->pte + ((vmf->address - addr) >> PAGE_SHIFT); + + ptep_set_access_flags(vma, vmf->address, ptep, + pte_mkdirty(ptep_get(ptep)), 1); + } + /* no need to invalidate: a not-present page won't be cached */ update_mmu_cache_range(vmf, vma, addr, vmf->pte, nr); } --=20 2.54.0 From nobody Sat Sep 26 07:16:28 2026 Received: from flow-a3-smtp.messagingengine.com (flow-a3-smtp.messagingengine.com [103.168.172.138]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A0ABA37F300; Thu, 3 Sep 2026 18:30:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.138 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460206; cv=none; b=egFy4RWlZABhcHlZ+WikGk4JcCLCu7Y76aWAoWfNRhvpkikb+NRn6lnjjTr6WOP701nt7Pktn1Q52Jx6OTnQ8LoAjigHnkNQ679g0Z2azH0KBF/qGtI2SZjLZCv2A5/DPi4XP+vNEv51T6xMu4RRb2dzNRAVdmA7RiFoyZboZwY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460206; c=relaxed/simple; bh=zRIFoaHBJ/MC2atUK1+Tog7TQjnA75nOcSUuGToSjjM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=F8sOHvPDlCyGvThiXbuMfoam59c74agcos7YJwY66js8ZUZQ403IzxbjgLfKVxMVS23g9fDtr5nihFvfm+3ox1pam8GRvLCf5f9svn6XUuFo4hmYHEWJsP78rFETuBEdotteFrsZP29jd3SxKQ7VibxPR0Mn6BCnicMKRXpem2g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=PCGah2gW; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=mqQeKYgE; arc=none smtp.client-ip=103.168.172.138 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="PCGah2gW"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="mqQeKYgE" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailflow.phl.internal (Postfix) with ESMTP id 47A8613801C9; Thu, 3 Sep 2026 14:29:58 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Thu, 03 Sep 2026 14:29:58 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1788460198; x= 1788467398; bh=l6U5JT6Kk2/xs1EJ8XhRce8llmfDGx2KTryruEjZgvk=; b=P CGah2gWTAOY3yv9onzlevaf6PtRIBJ9KvNO1H4h/R0CAu5V6o573OecJUZW1hu50 OcmQbEEBiUm3A95KwPzs+qFC+qc0sLn9LyR/Cf8f+pvXz1W5077bFvLSFYz0iHes h3VeE+9hOMYTTukruZdGnajEF1rEAnh2I9G3mNOzCXhbZVEnP8dx0VRV18hTi457 tSTXD693EWCrTXyZG8w3a+YJ4Oks4aVp4y2h/7DqGcLCHWEBoh88TTzikvGkRoP5 htQGU7HY0nCHLtpnfcfcjv+7tdKKmaZB+puK5Hm7Xw42kXoV5iU+toStgRhIsvNj xZCfPWW4Ub5+o4KhpHoVg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1788460198; x=1788467398; bh=l 6U5JT6Kk2/xs1EJ8XhRce8llmfDGx2KTryruEjZgvk=; b=mqQeKYgEAME01DRGP c/4LZYMnGGlrYHdQXWflNww59PC4u8jo0l8zi85PR/DAwDBpzoT1UaxU1Xv5M5F9 vGfOPVdGmRY0Zpn4ZAAnw+p/iSs8DUSejpEgl9RfarDJZboK2qoPGz1L7rA2W9Yk r/PceP3oNQWKpy1uKFhsZau2Q3ZPpvEMGsvdEYdbHbldnirD5VKCfYHN8yyPWyA3 i5h9AewDwpdj/vYmcZwQfp6QOSW6rdaooaOKKiBJybJZIYe45AK4m7Mi/V6Rg44l ubJ1xF4tMXkJ6pVf/7Q04ynsSjc7DWDvTCE3CBVvYYN5z0hyRuYZRTCvxgoI9YZh e6e/A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFt7tpQ23d6ma5ZKwjaUeV/XmiCXSwfUJTD/UMTaTx+tQ6P98fG4/H3PFA7hJFxaU SY7i6N6X86D7HaHpdbLHvMivEhulOmq3NoSfdg1pBrJA6v+lfq2ALheALUZjz6o7Nf8BIz DWIHLTWCbkC3w9bMwM7vsWb978omVdOmMplVzO2fLHjMRy1E/CxFkmf+2uWv7frAXfMOWG H/vdAXIPTaxMrzbmklMnL5PmRnZSJQIp5M1ihNDwszNqcLBuKJXCEkMPvIWqkfKnePof3W BlQ363pN/8dtwFHD474IAw0kAsBZ9JYvF7V9lEJnX7ro/M4XAHwxRodQrXyCYUn1NLN+Ds vTSbwwlQXD/Xc/IRcU3516zJNmer5W7iJlak/JBkZGlLVXh5Jz1RDWZJv/HmQhCZSFe6/+ LFT/aJzbopf1p7NE1DGDPW43exOhHSL2cH4V3IRUOv3StBhh/IZwpAlWaVz/iAgnjLIo3w ASK19DmhYX/RmFv3Osm2M4Ssd32+t4yegh9IdZHWic5Dxo0QmUYnp7ikHtJMkkfzVZTIW4 QEexuTLPtB9MjxXyNFFVeKtVwKxsm7kXcZEh75wncogm4njYAvO9anhmQGa6ywTuDI5cTg /2HiBPo8JwB+W+Rw+qVIlUFyQdOeeenoeStd8WAmWkEgqh6whFzfNtMu6ruw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 3 Sep 2026 14:29:57 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, "Matthew Wilcox (Oracle)" , David Hildenbrand Cc: Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jan Kara , Rik van Riel , Harry Yoo , Lance Yang , Jann Horn , Alexander Viro , Christian Brauner , "Darrick J. Wong" , Carlos Maiolino , Usama Arif , Pedro Falcato , linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Kiryl Shutsemau (Meta)" Subject: [RFC PATCH 4/5] iomap: narrow page_mkwrite() dirtying to the faulting page Date: Thu, 3 Sep 2026 19:29:42 +0100 Message-ID: <20260903182943.662461-5-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name> References: <20260903182943.662461-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" iomap already tracks dirty state per block and iomap_writeback_folio() already walks the bitmap and submits only the dirty ranges. The buffered write path sets just the range it copied. Only the mmap path sets everything, because iomap_dirty_folio() covers the whole folio. Add iomap_dirty_folio_range() for the new a_ops->dirty_folio_range(), and have iomap_page_mkwrite() dirty just the page the fault was for. Blocks are still allocated for the whole folio; only the dirty range narrows. iomap_dirty_folio() becomes a call to it covering the whole folio. Assisted-by: Claude:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- fs/iomap/buffered-io.c | 37 ++++++++++++++++++++++++++++--------- include/linux/iomap.h | 2 ++ 2 files changed, 30 insertions(+), 9 deletions(-) diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 0a5ebfda90f1..8d8a8acab01a 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -847,14 +847,18 @@ void iomap_invalidate_folio(struct folio *folio, size= _t offset, size_t len) } EXPORT_SYMBOL_GPL(iomap_invalidate_folio); =20 +bool iomap_dirty_folio_range(struct address_space *mapping, struct folio *= folio, + size_t off, size_t len) +{ + ifs_alloc(mapping->host, folio, 0); + iomap_set_range_dirty(folio, off, len); + return filemap_dirty_folio(mapping, folio); +} +EXPORT_SYMBOL_GPL(iomap_dirty_folio_range); + bool iomap_dirty_folio(struct address_space *mapping, struct folio *folio) { - struct inode *inode =3D mapping->host; - size_t len =3D folio_size(folio); - - ifs_alloc(inode, folio, 0); - iomap_set_range_dirty(folio, 0, len); - return filemap_dirty_folio(mapping, folio); + return iomap_dirty_folio_range(mapping, folio, 0, folio_size(folio)); } EXPORT_SYMBOL_GPL(iomap_dirty_folio); =20 @@ -1804,7 +1808,7 @@ iomap_truncate_page(struct inode *inode, loff_t pos, = bool *did_zero, EXPORT_SYMBOL_GPL(iomap_truncate_page); =20 static int iomap_folio_mkwrite_iter(struct iomap_iter *iter, - struct folio *folio) + struct folio *folio, size_t dirty_off, size_t dirty_len) { loff_t length =3D iomap_length(iter); int ret; @@ -1817,7 +1821,7 @@ static int iomap_folio_mkwrite_iter(struct iomap_iter= *iter, block_commit_write(folio, 0, length); } else { WARN_ON_ONCE(!folio_test_uptodate(folio)); - folio_mark_dirty(folio); + folio_mark_dirty_range(folio, dirty_off, dirty_len); } =20 return iomap_iter_advance(iter, length); @@ -1832,6 +1836,7 @@ vm_fault_t iomap_page_mkwrite(struct vm_fault *vmf, c= onst struct iomap_ops *ops, .private =3D private, }; struct folio *folio =3D page_folio(vmf->page); + size_t dirty_off, dirty_len; ssize_t ret; =20 folio_lock(folio); @@ -1840,8 +1845,22 @@ vm_fault_t iomap_page_mkwrite(struct vm_fault *vmf, = const struct iomap_ops *ops, goto out_unlock; iter.pos =3D folio_pos(folio); iter.len =3D ret; + + /* + * Blocks are still allocated for the whole folio, but only the page + * the fault was for has been written, so that is all that has to be + * written back. A racing truncate can leave that page beyond i_size, + * and then there is nothing to dirty. + */ + dirty_off =3D folio_page_idx(folio, vmf->page) << PAGE_SHIFT; + if (dirty_off < ret) + dirty_len =3D min_t(size_t, PAGE_SIZE, ret - dirty_off); + else + dirty_len =3D 0; + while ((ret =3D iomap_iter(&iter, ops)) > 0) - iter.status =3D iomap_folio_mkwrite_iter(&iter, folio); + iter.status =3D iomap_folio_mkwrite_iter(&iter, folio, + dirty_off, dirty_len); =20 if (ret < 0) goto out_unlock; diff --git a/include/linux/iomap.h b/include/linux/iomap.h index bc7ae6327dbf..3cf2ac751409 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -443,6 +443,8 @@ struct folio *iomap_get_folio(struct iomap_iter *iter, = loff_t pos, size_t len); bool iomap_release_folio(struct folio *folio, gfp_t gfp_flags); void iomap_invalidate_folio(struct folio *folio, size_t offset, size_t len= ); bool iomap_dirty_folio(struct address_space *mapping, struct folio *folio); +bool iomap_dirty_folio_range(struct address_space *mapping, struct folio *= folio, + size_t off, size_t len); void iomap_folio_mark_uptodate(struct folio *folio); int iomap_file_unshare(struct inode *inode, loff_t pos, loff_t len, const struct iomap_ops *ops, --=20 2.54.0 From nobody Sat Sep 26 07:16:28 2026 Received: from flow-a3-smtp.messagingengine.com (flow-a3-smtp.messagingengine.com [103.168.172.138]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0D69C3921CE; Thu, 3 Sep 2026 18:30:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.138 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460207; cv=none; b=e+f5goL19csa7Ps0FGhhqmZYrvsniuGGzgQde+YWg16q5fL8KvwnCFbq+VY2rLAXQVIEbZWPghEs0voskxij3RbI0hW1xyoEYf4yfvmyac5K1avZ424r1xW+WreocySV7XUthtVMtYPNITvQ2f7fvimYShS56Zm3I7Gd4MtQA6Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788460207; c=relaxed/simple; bh=s9grWpA1MfOtv515ZQpTkDMTmGOhDbSQM3D43txdNWE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GSSNkruQsN5NzezWVlcSz05+DFEnFAxv2V90Yf7VxWgNVWD08doANVnwy4BimMjuVQDGHhgvKkcG1VycOn1AV/ebQpeb8LQ90IvpB+MuUa0FLFpX2dVWfY/1SylacCQJycTx4+zk/mFzEjdD98EbCJU8z6cqXI1p0MuafMul9Yk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ZfVUbLvA; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=dP9rtL8a; arc=none smtp.client-ip=103.168.172.138 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ZfVUbLvA"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="dP9rtL8a" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailflow.phl.internal (Postfix) with ESMTP id 9E44E1380124; Thu, 3 Sep 2026 14:30:00 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Thu, 03 Sep 2026 14:30:00 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm2; t=1788460200; x= 1788467400; bh=vEML2BGHrlYhrMvp5AB1SPVVJb0CVEuxfoeolIQgojE=; b=Z fVUbLvAibQBespoCa9TCmGkV1BnOYkv1M/xJyjDG1gyY6rHvdR22a+kfxPAs9MB+ XF8T/bbG9xC9y88qQ7bqtuGOLmN6GyKlA04aHOkj6hwVYPfX5KvBE9i4zrntsoBc tBOg6sEb36G1eyQGKX139+1gpWGm14daiHOTVoCZKG5VCbGm2Mx0xkUx8T0iPfcd OLUWSUDJe9K4rQ2LTTeG7Ivm/ELObsZvHg8cl9w3j+mk9kQ0l3d6i2Z48RHFI/s4 jj9fCAIrR+St/752PEaYXkx3nnn/FJRII1xpI3iZ+oDRuS6ItlzppRL6SDlA1WSi D3s4QERuEPkjBJ7cEnO4w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1788460200; x=1788467400; bh=v EML2BGHrlYhrMvp5AB1SPVVJb0CVEuxfoeolIQgojE=; b=dP9rtL8aWaMErlaKB Nkd+BmzduYONw4oxo8AFh5+75pLVElq5D7eNb7mpykySo9GBZJ+jVX+i19xGyLFb zM3egKiSN/yx5SzY3GDUm09bVmhK3GpOwZLiGioDnFBlJr0c/oU+1AAcef0TjgFK S3QnZIWeslmK3MSf8KJIbLZ/w6PFdckq2cyhO/UEQTwxQqMB7BqQga5sku0v5NHU mcG472xrqnHIjIaKZhnk29PWFH1EfImKodpDQqHzcc3twt5NjIoBxidJHhtS7o40 tRdiyQq/7NNjATGbFmydEp1mDz7JngxCkXZ0wX/JN5IRKraMy9Fp948Dhg9pgydG Mvj8w== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGHxwC/Lj0Bpod+4rdZ9spd4mDaW4xhMdggMrFSzQCeA9RP9FzsfDxO1aVEQuln3V KlIkun03rqLJS4PF97Rrmwc8zvF+vhZeY8/AccbFhjx8xCTK9U2xEB/cnF/j4CrJu3iI+j BDg6h6fOcXVaNyEdDW/UAbX4yL7zrcocoaRK+UcMxdgx6bHH3qnz957Bwbxgg+ymAK7JCa v7gyu76p9eB8jE5O7T8JvzBRYgOMnbmzr1XCS8fmgHILKnNDK2ydQ7V9ZxlyJmWme3VYwV QfwByNxS35mMCJZpf38kt5LBBSLWPp97jeu/UaoDB9vrULq5ydrUj8GkH6/ySug8LGaySq coUY0N3OP70olSepNjMeWj5mqOhaxOuR5jUiYT9qIhV5gEVXIDgSEmlIq8NPzpG5TPztZT WubPaJmIDKi/lX6eG+dHBjR6VX0xMvJQxZll/fBCukIBGXC1ExxQGK8KqqT5cX/KXlkwXM 4FYqT0LZXJPXcs6JPOJlwZm9nBcXjg5gjUkk4fNwkkMwx1NTEiwp2NyuHl07d80tuzCaLf F/Phi0zWnhMgDDVo4rrYJn7Ffca/6UsJYfLqR2BjyJWxCoYQuqWqF79I5bSSMubeuF+mi9 LY22vDexfWFkE0EbRwniR67TbvH9OSLcNCYrDD+BG4uJRdtDQghYe18qvNFw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Thu, 3 Sep 2026 14:29:59 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, "Matthew Wilcox (Oracle)" , David Hildenbrand Cc: Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jan Kara , Rik van Riel , Harry Yoo , Lance Yang , Jann Horn , Alexander Viro , Christian Brauner , "Darrick J. Wong" , Carlos Maiolino , Usama Arif , Pedro Falcato , linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-xfs@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, "Kiryl Shutsemau (Meta)" Subject: [RFC PATCH 5/5] xfs: track mmap dirty state per block Date: Thu, 3 Sep 2026 19:29:43 +0100 Message-ID: <20260903182943.662461-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903182943.662461-1-kirill@shutemov.name> References: <20260903182943.662461-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" With a_ops->dirty_folio_range() wired up, a store through a shared mapping writes back only the block it touched instead of the whole folio. a_ops->dirty_folio() goes away with it. folio_mark_dirty() reaches a_ops->dirty_folio_range() with a range covering the folio, which is what iomap_dirty_folio() does. On a 512M file held in 2M page cache folios, storing one byte per folio and then calling msync() wrote 512M of the file before this series and writes 1M after it. Minor fault counts are the same either way, in both the PTE-mapped and the batch-mapped case. A PMD-mapped folio still writes back whole. A PMD carries one dirty bit for the 2M it maps, so there is no per-block state to recover. Assisted-by: Claude:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- fs/xfs/xfs_aops.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/fs/xfs/xfs_aops.c b/fs/xfs/xfs_aops.c index 74a6089abadf..600e7de2e599 100644 --- a/fs/xfs/xfs_aops.c +++ b/fs/xfs/xfs_aops.c @@ -863,7 +863,7 @@ const struct address_space_operations xfs_address_space= _operations =3D { .read_folio =3D xfs_vm_read_folio, .readahead =3D xfs_vm_readahead, .writepages =3D xfs_vm_writepages, - .dirty_folio =3D iomap_dirty_folio, + .dirty_folio_range =3D iomap_dirty_folio_range, .release_folio =3D iomap_release_folio, .invalidate_folio =3D iomap_invalidate_folio, .bmap =3D xfs_vm_bmap, --=20 2.54.0