From nobody Mon Sep 28 06:35:53 2026 Received: from relay2-d.mail.gandi.net (relay2-d.mail.gandi.net [217.70.183.194]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 07C872E06E4; Tue, 25 Aug 2026 13:53:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.70.183.194 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666018; cv=none; b=WU8Fy2iUQXSIQ5b8ckGbpUiXqkoMZDKETOKv3AT24+8xTWlDNOLsJzfLpKiY8deL+ZFjXi5LIe/SFCnYpQYIDlgn44x8hbriLAbGHgwKgQII5jlVvAQfJ1YY2HCd0hwKjsImVtzt7KD1qV4xEvSFXN1c6AUesaGc31tXRngymBg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666018; c=relaxed/simple; bh=OU4ZB+Nl8uAWKesx38KT9bX4tXsl6DkXi9UId5ms1ds=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XaPdIScu9Zy0G0tUUT1S7sy4FX3kIrE2AuBvagjcazV9b1EOtPa1Zm2G20XWQHYvid3ejW/SswW5PCaLREOYhmxsDZmvNNx68iw3hmdtTrfK+4rkP0kfrpno5OG+TdwNNMMPOv7A3sokgkzktiYr7Fr/jh0fU3onZ+gf5t/YpjU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr; spf=pass smtp.mailfrom=ghiti.fr; arc=none smtp.client-ip=217.70.183.194 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=ghiti.fr Received: by mail.gandi.net (Postfix) with ESMTPSA id C4D6D3EC40; Tue, 25 Aug 2026 13:53:20 +0000 (UTC) From: Alexandre Ghiti To: Johannes Weiner , Yosry Ahmed , Nhat Pham , Andrew Morton , Chris Li , Kairui Song Cc: Kairui Song , Chengming Zhou , "Matthew Wilcox (Oracle)" , Jan Kara , Kemeng Shi , Baoquan He , Barry Song , Youngjun Park , Alexander Viro , Christian Brauner , David Hildenbrand , Lorenzo Stoakes , Michal Hocko , Axel Rasmussen , Qi Zheng , Shakeel Butt , Wei Xu , Yuanchu Xie , Kunwu Chan , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, Alexandre Ghiti Subject: [PATCH v4 1/3] mm: swap: move LRU insertion out of the swap cache allocator Date: Tue, 25 Aug 2026 15:52:05 +0200 Message-ID: <20260825135209.3135169-2-alex@ghiti.fr> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260825135209.3135169-1-alex@ghiti.fr> References: <20260825135209.3135169-1-alex@ghiti.fr> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-GND-Sasl: alex@ghiti.fr X-GND-Score: -100 X-GND-Cause: dmFkZTGiHDiWlfQ5ccxBZewsSasowgl6gtfYlDY1PWykUS63s55gMKk5eyoNcDBuK/u7C54COj+K24nRLBNKdyIK2RQTz0Gd7w2H9vLC2370zE+Ddl7i6a/AXnrEXG/Z0BlNEt/8xyMeSW6zBiRZZwvjEQff7MVGCDEPhCRkLbIwEbRTVuxjo/fEYI5R/+1hgmHojmsu8cW0sWOxX9UVZyxkZFjKfn8+jU4byiXlitSGcdBMjxKaqXvJ5iTQAfdm2MFjYyiM9uE53CRdiBRxdil8ElrlmgLmGYOBnqlq0fAMw7kI3FGQeaxVI5JUrPkWhh338+VpDFe89mQ4tcVzXfaSQ8hzDc0daAEIRGFZmYT0D7iWGD5qbwVNcKHLyfesFsfK/xpXyoXvm+s3SJlZnm5pqMnFKTlBYmZ/tPlFbha8ECk/iGOxwFzxS2k18DAYhofQFOQl7PhNmSRxtmRghS0HkfJLLVDqgtGFpiMZbNZ4qwuldG/IuC7XTINHXKU9Ox7Mafs95pYy44z5lAsZ7z42M81z8bzR4mgmtFTPz38QShROj60PllvcOBhS8190kONhlBFG2XZV0gWssJzlzOi1h80WAjRrvSYVIuVUxkntGmHdcDl03QGFX9/nPC7Gtw2vwvjDlR4c8x8D/6W1uz8PuK3jKeibQsG+9yx1D3u9ztRX5Q X-GND-State: clean Content-Type: text/plain; charset="utf-8" zswap writeback wants a swap cache folio it can free directly once writeback completes, i.e. one that is not on the LRU (folio_add_lru() stages the folio in a per-CPU batch that holds a reference until it is drained, which keeps remove_mapping() from freeing the folio on synchronous devices, and likely on asynchronous ones too). So defer the LRU addition to the callers of __swap_cache_alloc_folio(), no functional change intended. Suggested-by: Kairui Song Signed-off-by: Alexandre Ghiti Reviewed-by: Kunwu Chan Reviewed-by: Nhat Pham --- mm/swap.h | 6 +++--- mm/swap_state.c | 20 ++++++++++++-------- mm/zswap.c | 5 +++-- 3 files changed, 18 insertions(+), 13 deletions(-) diff --git a/mm/swap.h b/mm/swap.h index 77d2d14eda42..fc44daae1de1 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -304,9 +304,9 @@ bool swap_cache_has_folio(swp_entry_t entry); struct folio *swap_cache_get_folio(swp_entry_t entry); void *swap_cache_get_shadow(swp_entry_t entry); void swap_cache_del_folio(struct folio *folio); -struct folio *swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp_m= ask, - unsigned long orders, struct vm_fault *vmf, - struct mempolicy *mpol, pgoff_t ilx); +struct folio *__swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp= _mask, + unsigned long orders, struct vm_fault *vmf, + struct mempolicy *mpol, pgoff_t ilx); /* Below helpers require the caller to lock and pass in the swap cluster. = */ void __swap_cache_add_folio(struct swap_cluster_info *ci, struct folio *folio, swp_entry_t entry); diff --git a/mm/swap_state.c b/mm/swap_state.c index 727a17ee7821..07418fc94f00 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -483,13 +483,11 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, node_stat_mod_folio(folio, NR_FILE_PAGES, nr_pages); lruvec_stat_mod_folio(folio, NR_SWAPCACHE, nr_pages); =20 - /* Caller will initiate read into locked new_folio */ - folio_add_lru(folio); return folio; } =20 /** - * swap_cache_alloc_folio - Allocate folio for swapped out slot in swap ca= che. + * __swap_cache_alloc_folio - Allocate folio for swapped out slot in swap = cache. * @targ_entry: swap entry indicating the target slot * @gfp: memory allocation flags * @orders: allocation orders, must be non zero @@ -501,13 +499,17 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, * doing IO (e.g. swap in or zswap writeback). The swap slot indicated by * @targ_entry must have a non-zero swap count (swapped out). * + * The returned folio is locked and is NOT on the LRU. The caller must eit= her + * add it to the LRU with folio_add_lru() so page reclaim can find it, or = free + * it directly once done; a folio left off the LRU is unreclaimable and le= aks. + * * Context: Caller must protect the swap device with reference count or lo= cks. * Return: Returns the folio if allocation succeeded and folio is in the s= wap * cache. Returns error code if failed due to race, OOM or invalid argumen= ts. */ -struct folio *swap_cache_alloc_folio(swp_entry_t targ_entry, gfp_t gfp, - unsigned long orders, struct vm_fault *vmf, - struct mempolicy *mpol, pgoff_t ilx) +struct folio *__swap_cache_alloc_folio(swp_entry_t targ_entry, gfp_t gfp, + unsigned long orders, struct vm_fault *vmf, + struct mempolicy *mpol, pgoff_t ilx) { int order, err; struct folio *ret; @@ -643,12 +645,13 @@ static struct folio *swap_cache_read_folio(swp_entry_= t entry, gfp_t gfp, folio =3D swap_cache_get_folio(entry); if (folio) return folio; - folio =3D swap_cache_alloc_folio(entry, gfp, BIT(0), NULL, mpol, ilx); + folio =3D __swap_cache_alloc_folio(entry, gfp, BIT(0), NULL, mpol, ilx); } while (PTR_ERR(folio) =3D=3D -EEXIST); =20 if (IS_ERR_OR_NULL(folio)) return NULL; =20 + folio_add_lru(folio); swap_read_folio(folio, plug); if (readahead) { folio_set_readahead(folio); @@ -683,12 +686,13 @@ struct folio *swapin_sync(swp_entry_t entry, gfp_t gf= p, unsigned long orders, folio =3D swap_cache_get_folio(entry); if (folio) return folio; - folio =3D swap_cache_alloc_folio(entry, gfp, orders, vmf, mpol, ilx); + folio =3D __swap_cache_alloc_folio(entry, gfp, orders, vmf, mpol, ilx); } while (PTR_ERR(folio) =3D=3D -EEXIST); =20 if (IS_ERR(folio)) return folio; =20 + folio_add_lru(folio); swap_read_folio(folio, NULL); return folio; } diff --git a/mm/zswap.c b/mm/zswap.c index 761cd699e0a3..8163e6c5f76c 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1000,8 +1000,8 @@ static int zswap_writeback_entry(struct zswap_entry *= entry, return -EEXIST; =20 mpol =3D get_task_policy(current); - folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, - NO_INTERLEAVE_INDEX); + folio =3D __swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mp= ol, + NO_INTERLEAVE_INDEX); put_swap_device(si); =20 /* @@ -1013,6 +1013,7 @@ static int zswap_writeback_entry(struct zswap_entry *= entry, */ if (IS_ERR(folio)) return PTR_ERR(folio); + folio_add_lru(folio); =20 /* * folio is locked, and the swapcache is now secured against --=20 2.53.0-Meta From nobody Mon Sep 28 06:35:53 2026 Received: from relay2-d.mail.gandi.net (relay2-d.mail.gandi.net [217.70.183.194]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E2AAA423761; Tue, 25 Aug 2026 13:54:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.70.183.194 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666076; cv=none; b=Abd6JJctfGdqQRNoH4eyTOtBNUVThyv+S3wSBiGJ58BImHtOF+BRrO320SSWGSCHsGbIv6Lu3LgIlV/44V9hrP5WqY8ehZtNRbh5ONxBKR/A5R40ePCq6oegunOuUozMOBLIzB+9Qj3pTVKyqWcOmL5ernBx71xcd9G518vbXd4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666076; c=relaxed/simple; bh=gImpJjZ7DEhNHeWoO0HYhy9bzhZqlLwDOLMAAGgGcic=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Xe7mFG8Ax7/yTL27zWYFZagZxU7jalo9mjbBbyRFi8fTV+tv1b5KIpfFa/0Fei30jyV20+sU5+Rg5Rf+9qRrF55hrFR7tAcN/qIt0Qc4CnbhTQwr9P7YZxrZu/HQfeW2obpCPcc3a0TTJAZLG8Cq8jEKXOKT3RHaYUuy9jjmu+g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr; spf=pass smtp.mailfrom=ghiti.fr; arc=none smtp.client-ip=217.70.183.194 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=ghiti.fr Received: by mail.gandi.net (Postfix) with ESMTPSA id 5E0683EDDC; Tue, 25 Aug 2026 13:54:28 +0000 (UTC) From: Alexandre Ghiti To: Johannes Weiner , Yosry Ahmed , Nhat Pham , Andrew Morton , Chris Li , Kairui Song Cc: Kairui Song , Chengming Zhou , "Matthew Wilcox (Oracle)" , Jan Kara , Kemeng Shi , Baoquan He , Barry Song , Youngjun Park , Alexander Viro , Christian Brauner , David Hildenbrand , Lorenzo Stoakes , Michal Hocko , Axel Rasmussen , Qi Zheng , Shakeel Butt , Wei Xu , Yuanchu Xie , Kunwu Chan , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, Alexandre Ghiti Subject: [PATCH v4 2/3] mm: swap: drop dropbehind swap cache folios on writeback completion Date: Tue, 25 Aug 2026 15:52:06 +0200 Message-ID: <20260825135209.3135169-3-alex@ghiti.fr> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260825135209.3135169-1-alex@ghiti.fr> References: <20260825135209.3135169-1-alex@ghiti.fr> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-GND-Sasl: alex@ghiti.fr X-GND-Score: -100 X-GND-Cause: dmFkZTE5yRoPOtXaeVjz9nQYgYdgB2auKAdjCUIr5fKYD4tdDaE1rs3LHZQc4Y+i/VGPnn893m6ZXKNvOHjIpP0WvxT1xRkSxTAmSD74k5h/QvYA5e/Qd5Qp+C+1l/DaTi8AhfimOpo9Hb/M/tYAzEgQ8mJpPSqgcAOHSEeUKAFK6+Jfc5ACHK4WH5T526qhUWIUJwtcTcKZAD6bevELJuJmneh86u1cgECUmL27fSiv7RBksWh3FxjVFmXOYZkitH61oBWQZOvCWZXxb3FwYRES+GIVDfBZysSnq4f/n5MyAXRWt/BFEaqWacVR1vTB2r/AICAplFmbCabDE651DW+YI5R0MiadWc/QraQBB74Gpzn/mui0X/sfjIUwnOZiGiD6R9fjMzia+THvV3eU/pqFJcpa5AiOFBG7jHRqE7m2M0nde/lddf0N/IO9jaAllKtNcpNxHET+eJ2HLgOeurGb3ZJymrfTO5WLDVNy996zV/T9LQt67lWKEZKmb70WvRn5+mEAuJO6NVCNGSofgxM8MpPpSm0bu9a/u8pkn/UPn/Lu7q33O7Kfz64jozefevC9CxyGbgFBTI7CkFDYcLh9/ht9LiQzTzQoSslVYP4S6WdSWnmp2sF/0X6lOPGUqp/GMYUlDpILRmIDABSzGBb8aitWwdtbuaS+/p7SjMyYFD4ZUw X-GND-State: clean Content-Type: text/plain; charset="utf-8" A PG_dropbehind folio is dropped from its cache once writeback completes rather than left for reclaim to find later; this is implemented for file folios in folio_end_dropbehind(). Extend it to swap cache folios. The drop takes the folio and swap cluster locks and may sleep, so it cannot run in interrupt context. Set BIO_COMPLETE_IN_TASK on the write, as the file dropbehind paths do, and drop the folio directly from folio_end_writeback(). Suggested-by: Yosry Ahmed Suggested-by: Johannes Weiner Suggested-by: Nhat Pham Signed-off-by: Alexandre Ghiti Reviewed-by: Kunwu Chan Reviewed-by: Nhat Pham --- include/linux/swap.h | 5 +++++ mm/filemap.c | 19 ++++++++++++++++++ mm/page_io.c | 7 +++++++ mm/swap_state.c | 41 +++++++++++++++++++++++++++++++++++++ mm/vmscan.c | 48 +++++++++++++++++++++++++++++++++++--------- 5 files changed, 110 insertions(+), 10 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 8f0f68e245ba..29ec60dcae21 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -374,6 +374,8 @@ extern unsigned long mem_cgroup_shrink_node(struct mem_= cgroup *mem, extern unsigned long shrink_all_memory(unsigned long nr_pages); extern int vm_swappiness; long remove_mapping(struct address_space *mapping, struct folio *folio); +long remove_mapping_reclaim(struct address_space *mapping, struct folio *f= olio, + struct mem_cgroup *target_memcg); =20 #if defined(CONFIG_SYSFS) && defined(CONFIG_NUMA) extern int reclaim_register_node(struct node *node); @@ -465,6 +467,8 @@ void swap_put_entries_direct(swp_entry_t entry, int nr); */ bool folio_free_swap(struct folio *folio); =20 +void swap_writeback_dropbehind_folio(struct folio *folio); + /* Allocate / free (hibernation) exclusive entries */ swp_entry_t swap_alloc_hibernation_slot(int type); void swap_free_hibernation_slot(swp_entry_t entry); @@ -475,6 +479,7 @@ static inline void put_swap_device(struct swap_info_str= uct *si) } =20 #else /* CONFIG_SWAP */ +static inline void swap_writeback_dropbehind_folio(struct folio *folio) {} static inline struct swap_info_struct *get_swap_device(swp_entry_t entry) { return NULL; diff --git a/mm/filemap.c b/mm/filemap.c index d721986d5f46..040c97a121de 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -1686,6 +1686,8 @@ EXPORT_SYMBOL_GPL(folio_end_writeback_no_dropbehind); */ void folio_end_writeback(struct folio *folio) { + bool swap_dropbehind; + VM_BUG_ON_FOLIO(!folio_test_writeback(folio), folio); =20 /* @@ -1695,7 +1697,24 @@ void folio_end_writeback(struct folio *folio) * reused before the folio_wake_bit(). */ folio_get(folio); + + /* + * Sample this before folio_end_writeback_no_dropbehind() clears + * PG_writeback: until then a racing swapin cannot remove the folio from + * the swap cache. Afterwards it can, and the drop below then finds a + * non-swapcache folio and puts it back on the LRU instead. The + * reference taken above keeps the folio alive across that window. + */ + swap_dropbehind =3D folio_test_swapcache(folio) && + folio_test_dropbehind(folio); + folio_end_writeback_no_dropbehind(folio); + + if (swap_dropbehind) { + swap_writeback_dropbehind_folio(folio); + return; + } + folio_end_dropbehind(folio); folio_put(folio); } diff --git a/mm/page_io.c b/mm/page_io.c index b23f494fcc83..586c79c3bb3d 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -456,6 +456,13 @@ static void swap_writepage_bdev_async(struct folio *fo= lio, bio->bi_end_io =3D end_swap_bio_write; bio_add_folio_nofail(bio, folio, folio_size(folio), 0); =20 + /* + * Dropping the folio from the swap cache takes sleeping locks, so the + * completion must not run in interrupt context. + */ + if (folio_test_dropbehind(folio)) + bio_set_flag(bio, BIO_COMPLETE_IN_TASK); + bio_associate_blkg_from_page(bio, folio); count_swpout_vm_event(folio); folio_start_writeback(folio); diff --git a/mm/swap_state.c b/mm/swap_state.c index 07418fc94f00..231fa87cbbe0 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -537,6 +537,47 @@ struct folio *__swap_cache_alloc_folio(swp_entry_t tar= g_entry, gfp_t gfp, return ret; } =20 +/** + * swap_writeback_dropbehind_folio - drop a dropbehind swap cache folio + * @folio: the off-LRU folio whose writeback has completed + * + * Context: task context, with the reference taken by folio_end_writeback() + * donated to us. + */ +void swap_writeback_dropbehind_folio(struct folio *folio) +{ + struct mem_cgroup *memcg; + + folio_lock(folio); + + /* The folio was allocated off the LRU and nothing re-adds it here. */ + VM_WARN_ON_ONCE_FOLIO(folio_test_lru(folio), folio); + + rcu_read_lock(); + memcg =3D folio_memcg(folio); + if (!mem_cgroup_tryget(memcg)) + memcg =3D NULL; + rcu_read_unlock(); + + /* + * Gate remove_mapping_reclaim() on folio_test_swapcache(): a racing + * swapin may have freed the swap slot (folio_free_swap()) and dropped the + * folio from the cache, and it must not run on a non-swapcache folio (it + * would trip __remove_mapping()'s mapping =3D=3D folio_mapping() check). + */ + if (!folio_test_swapcache(folio) || folio_test_writeback(folio) || + !remove_mapping_reclaim(swap_address_space(folio->swap), folio, memcg= )) { + /* Raced: the folio is now owned by the swapin; put it back. */ + folio_clear_dropbehind(folio); + folio_add_lru(folio); + } + + mem_cgroup_put(memcg); + + folio_unlock(folio); + folio_put(folio); +} + /* * If we are the only user, then try to free up the swap cache. * diff --git a/mm/vmscan.c b/mm/vmscan.c index 848bd3e5eee2..4cc3a3ed6db6 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -780,6 +780,22 @@ static int __remove_mapping(struct address_space *mapp= ing, struct folio *folio, return 0; } =20 +static long __remove_mapping_unfreeze(struct address_space *mapping, + struct folio *folio, bool reclaimed, + struct mem_cgroup *target_memcg) +{ + if (__remove_mapping(mapping, folio, reclaimed, target_memcg)) { + /* + * Unfreezing the refcount with 1 effectively + * drops the pagecache ref for us without requiring another + * atomic operation. + */ + folio_ref_unfreeze(folio, 1); + return folio_nr_pages(folio); + } + return 0; +} + /** * remove_mapping() - Attempt to remove a folio from its mapping. * @mapping: The address space. @@ -794,16 +810,28 @@ static int __remove_mapping(struct address_space *map= ping, struct folio *folio, */ long remove_mapping(struct address_space *mapping, struct folio *folio) { - if (__remove_mapping(mapping, folio, false, NULL)) { - /* - * Unfreezing the refcount with 1 effectively - * drops the pagecache ref for us without requiring another - * atomic operation. - */ - folio_ref_unfreeze(folio, 1); - return folio_nr_pages(folio); - } - return 0; + return __remove_mapping_unfreeze(mapping, folio, false, NULL); +} + +/** + * remove_mapping_reclaim() - Remove a folio from its mapping, as reclaim = does. + * @mapping: The address space. + * @folio: The folio to remove. + * @target_memcg: The memcg to charge the eviction shadow to; the caller m= ust + * keep it alive across the call. + * + * Like remove_mapping(), but stores a workingset eviction shadow the way = page + * reclaim does, so that a later refault can be detected and the folio + * re-activated. + * Return: The number of pages removed from the mapping. 0 if the folio + * could not be removed. + * Context: The caller should have a single refcount on the folio and + * hold its lock. + */ +long remove_mapping_reclaim(struct address_space *mapping, struct folio *f= olio, + struct mem_cgroup *target_memcg) +{ + return __remove_mapping_unfreeze(mapping, folio, true, target_memcg); } =20 /** --=20 2.53.0-Meta From nobody Mon Sep 28 06:35:53 2026 Received: from relay7-d.mail.gandi.net (relay7-d.mail.gandi.net [217.70.183.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C0CD4481256; Tue, 25 Aug 2026 13:55:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.70.183.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666147; cv=none; b=bWg5ILejZNpg4IQEHvEfoygqjHrdg3XtEtAe22ztJkklrLYOFQEookNOhzm9lluPA9pFudsnTaTxV4vyKQyVYVBxTGFC2lv3g596LtkHoD5WbWaDZ1U3CTLuj2gs2lY/fqgevejN3EVMGjGpTDP5trMapp6n8xeF1o4Qmkh0dRM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787666147; c=relaxed/simple; bh=0/WeIqBfPKwCqAW+ElrQnm7JI/vBCWEn6IMLIr1sMYU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KcwqLhswbuAJElkP/OViKvjMxYyUcj07N0h+0gd35LHG9OhCaCcYfVv6bK0exHX1gmk1ZgV4a6VL6u8UQXGLjlOkMmK7IN6pnvq+aOK+7/RAkn0aWVqBXqv/N7ubaWDBI2eDXRgC3EHFd1TeJ/3+JI+np+ACB6TPK15tjZEoNRY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr; spf=pass smtp.mailfrom=ghiti.fr; arc=none smtp.client-ip=217.70.183.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ghiti.fr Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=ghiti.fr Received: by mail.gandi.net (Postfix) with ESMTPSA id 4CF8C3EC48; Tue, 25 Aug 2026 13:55:33 +0000 (UTC) From: Alexandre Ghiti To: Johannes Weiner , Yosry Ahmed , Nhat Pham , Andrew Morton , Chris Li , Kairui Song Cc: Kairui Song , Chengming Zhou , "Matthew Wilcox (Oracle)" , Jan Kara , Kemeng Shi , Baoquan He , Barry Song , Youngjun Park , Alexander Viro , Christian Brauner , David Hildenbrand , Lorenzo Stoakes , Michal Hocko , Axel Rasmussen , Qi Zheng , Shakeel Butt , Wei Xu , Yuanchu Xie , Kunwu Chan , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, Alexandre Ghiti Subject: [PATCH v4 3/3] mm: zswap: drop cold writeback folios via swap dropbehind Date: Tue, 25 Aug 2026 15:52:07 +0200 Message-ID: <20260825135209.3135169-4-alex@ghiti.fr> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260825135209.3135169-1-alex@ghiti.fr> References: <20260825135209.3135169-1-alex@ghiti.fr> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-GND-Sasl: alex@ghiti.fr X-GND-Score: -100 X-GND-Cause: dmFkZTGiHDiWlfQ5ccxBZewsSasowgl6gtfYlDY1PWykUS63s55gMKk5eyoNcDBuK/u7C54COj+K24nRLBNKdyIK2RQTz0Gd7w2H9vLC2370zE+Ddl7i6a/AXnrEXG/Z0BlNEt/8xyMeSW6zBiRZZwvjEQff7MVGCDEPhCRkLbIwEbRTVuxjo/fEYI5R/+1hgmHojmsu8cW0sWOxX9UVZyxkZFjKfn8+jU4byiXlitSGcdBMjxKaqXvJ5iTQAfdm2MFjYyiM9uE53CRdiBRxdil8ElrlmgLmGYOBnqlq0fAMw7kI3FGQeaxVI5JUrPkWhh338+VpDFe89mQ4tcVzXfaSQ8hzO/SV91Cb7MrhV04BuJgiTggGt7Zdy76v+v7di5V0F2GWJCgToIeVaHv/Py1bQ3X7vOAU0B8xrylrXT48j2SQPSTKdpB92XGLcGraHgdEQU91N+ZEkWE0EcFcwWBVjG4UlvtEJ4g3dezAkjFWFHIeOOTWEYYxW6p9quP8zM5YnvsxNCuS9As6YAZcXNMc7qGf3JBHypUwMfNT4exbp8J3xkukcZHG/XNc5XQ7kT1Japjehyh81GsvvoqgbA+JKIPd8DFZWoCTLWUfe7Xjox8gz9Pj2o95Ze/TDPs3/QoqVyZjPFEADf8hZRx/ISlGVw/W3F88XplLHYxgWA1CMdOnDQ X-GND-State: clean Content-Type: text/plain; charset="utf-8" zswap writeback decompresses an entry into a fresh swap cache folio and writes it back. The folio is cold by construction, yet it is left on the LRU for reclaim to find and free later, wasting a reclaim scan and keeping cold memory resident longer than necessary. Allocate the folio off the LRU and mark it PG_dropbehind so the swap dropbehind path frees it from the swap cache once writeback completes. Suggested-by: Johannes Weiner Suggested-by: Nhat Pham Signed-off-by: Alexandre Ghiti Reviewed-by: Kunwu Chan --- mm/zswap.c | 19 ++++++++++++++++--- 1 file changed, 16 insertions(+), 3 deletions(-) diff --git a/mm/zswap.c b/mm/zswap.c index 8163e6c5f76c..d16822a516e8 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1013,7 +1013,6 @@ static int zswap_writeback_entry(struct zswap_entry *= entry, */ if (IS_ERR(folio)) return PTR_ERR(folio); - folio_add_lru(folio); =20 /* * folio is locked, and the swapcache is now secured against @@ -1046,12 +1045,26 @@ static int zswap_writeback_entry(struct zswap_entry= *entry, /* folio is up to date */ folio_mark_uptodate(folio); =20 - /* move it to the tail of the inactive list after end_writeback */ - folio_set_reclaim(folio); + folio_set_dropbehind(folio); + + /* + * Drop our reference before starting writeback so the swap cache holds + * the only one: the drop in folio_end_writeback() needs that for + * remove_mapping_reclaim() to succeed, otherwise the folio is handed + * back to reclaim instead. + * + * Nothing can free the folio in the meantime: we hold the folio lock + * until writeback starts, PG_writeback then blocks swap cache removal, + * and folio_end_writeback() takes its own reference before clearing + * PG_writeback and donates it to the drop. + */ + folio_put(folio); =20 /* start writeback */ __swap_writepage(folio, NULL); =20 + return 0; + out: if (ret) { swap_cache_del_folio(folio); --=20 2.53.0-Meta