From nobody Fri Oct 2 03:41:04 2026 Received: from out-181.mta1.migadu.com (mta1.migadu.com [37.59.57.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 642D938E100 for ; Wed, 5 Aug 2026 07:54:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=37.59.57.117 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916452; cv=none; b=kHAUbYycGkGE8naPe3lCDyQB9oyBfO5Z29VeiN49px/BaRTzwagXMzUMSyDL134U1YUWAN0NCM6xPgwMBgkVvY8E98xG3iGweEMW9nrkB5rX6Poe5zuYE7ahOqS7KUkBFDXPPDGUV2wX6bemH48mt5HcEfESNl1bBhS+XVySfMw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916452; c=relaxed/simple; bh=9XnfkdMqTmO2JyAJ7Ad0CGdZS5toXEqOplv2CvH59Fw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=Zku74+uH+YXDsmy61oGDweDDPzAvje0IAZxIogCEK8j+l+QbpJZQ4MSGeTRfSuuwAmujXdS1MSyPg+zpLfoaO12IdmYLDWM8JI2HQ9jWD9+Bkc0IBo6FsP2arxvRkHEUlQBA6f3N9J6ddrMtJRw+6/RViiwkJmvoJkoWI7eE0uw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=QWiaLdA6; arc=none smtp.client-ip=37.59.57.117 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="QWiaLdA6" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916447; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Uz9E19ZcjXqKY3kPMLAOLyjnTdmqsl8/hR42BfB8vtA=; b=QWiaLdA6YkLRcMr8zbXYeEKDYhkzprRdyodqBuEmPCAtX+mzD2p+a1SQJLaj9X9K3/iY5m 9BZzFhJ4hpWNRH+zYX4qNoFuqVqQ+ltdwBUWTF/QoEcds3iZ4S6QHgp9EJKLf7vwzdgEge qotuMxWFp8impk1IVlzNc3PhPh4OQQQ= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 01/10] mm: xswap support for zswap Date: Wed, 5 Aug 2026 15:53:24 +0800 Message-ID: <20260805075336.3579395-2-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT From: Chris Li Introduce extendable (virtual) swap device support =E2=80=94 xswap. The current zswap requires a backing swapfile. The swap slot used by zswap is not able to be used by the swapfile, wasting swapfile space. An xswap device is a swapfile that only contains the swap header, with the header indicating the size of the virtual swap space. There is no swap data section, therefore no waste of swapfile space. Any write to an xswap device will fail. To prevent accidental read or write, bdev of swap_info_struct is set to NULL. Xswap devices set the SSD flag because there is no rotational disk access when using zswap. Zswap writeback is disabled if all swapfiles in the system are xswap devices (tracked via nr_real_swapfiles). How to create an xswap device: touch swap.1G truncate -s 1G swap.1G mkswap swap.1G dd if=3Dswap.1G of=3Dxswap.1G bs=3D4096 count=3D1 # xswap.1G is 4K on disk but reports 1G capacity swapon xswap.1G Signed-off-by: Chris Li Signed-off-by: Baoquan He --- include/linux/swap.h | 2 ++ mm/page_io.c | 16 +++++++++++++++ mm/swap_state.c | 7 +++++++ mm/swapfile.c | 49 ++++++++++++++++++++++++++++++++++++++++---- mm/zswap.c | 9 ++++++-- 5 files changed, 77 insertions(+), 6 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 45f301d73e2a..78f4302b92e1 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -207,6 +207,7 @@ enum { SWP_STABLE_WRITES =3D (1 << 11), /* no overwrite PG_writeback pages */ SWP_SYNCHRONOUS_IO =3D (1 << 12), /* synchronous IO is efficient */ SWP_HIBERNATION =3D (1 << 13), /* pinned for hibernation */ + SWP_XSWAP =3D (1 << 14), /* extendable swap device */ /* add others here before... */ }; =20 @@ -350,6 +351,7 @@ void free_folio_and_swap_cache(struct folio *folio); void free_pages_and_swap_cache(struct encoded_page **, int); /* linux/mm/swapfile.c */ extern atomic_long_t nr_swap_pages; +extern atomic_t nr_real_swapfiles; extern long total_swap_pages; extern atomic_t nr_rotate_swap; =20 diff --git a/mm/page_io.c b/mm/page_io.c index e4fa7ffffe8b..1f3fa52131ab 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -247,6 +247,17 @@ int swap_writeout(struct swap_io_ctx *ctx, struct foli= o *folio) } rcu_read_unlock(); =20 + /* + * ctx->sis is set by swap_add_folio() which is called from + * __swap_writepage() below. Since we must avoid the writepage + * path for xswap devices, use the swap_info from the folio's + * swap entry directly instead of going through ctx. + */ + if (unlikely(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP)) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writepage(ctx, folio); return 0; out_unlock: @@ -479,6 +490,11 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct f= olio *folio) if (zswap_load(folio) !=3D -ENOENT) goto finish; =20 + if (unlikely(sis->flags & SWP_XSWAP)) { + folio_unlock(folio); + goto finish; + } + /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); swap_add_folio(ctx, folio, READ); diff --git a/mm/swap_state.c b/mm/swap_state.c index 5be825911e64..ebb2d2ac356f 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -829,6 +829,13 @@ struct folio *swap_cluster_readahead(swp_entry_t entry= , gfp_t gfp_mask, struct blk_plug plug; swp_entry_t ra_entry; =20 + /* + * The entry may have been freed by another task. Avoid swap_info_get() + * which will print error message if the race happens. + */ + if (si->flags & SWP_XSWAP) + goto skip; + mask =3D swapin_nr_pages(offset) - 1; if (!mask) goto skip; diff --git a/mm/swapfile.c b/mm/swapfile.c index 4d4e3e3059f6..08c49d5bea84 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -66,6 +66,7 @@ static void move_cluster(struct swap_info_struct *si, static DEFINE_SPINLOCK(swap_lock); static unsigned int nr_swapfiles; atomic_long_t nr_swap_pages; +atomic_t nr_real_swapfiles; /* * Some modules use swappable objects and may try to swap them out under * memory pressure (via the shrinker). Before doing so, they may wish to @@ -1223,6 +1224,8 @@ static void del_from_avail_list(struct swap_info_stru= ct *si, bool swapoff) goto skip; } =20 + if (!(si->flags & SWP_XSWAP)) + atomic_sub(1, &nr_real_swapfiles); plist_del(&si->avail_list, &swap_avail_head); =20 skip: @@ -1265,6 +1268,8 @@ static void add_to_avail_list(struct swap_info_struct= *si, bool swapon) } =20 plist_add(&si->avail_list, &swap_avail_head); + if (!(si->flags & SWP_XSWAP)) + atomic_add(1, &nr_real_swapfiles); =20 skip: spin_unlock(&swap_avail_lock); @@ -2952,6 +2957,19 @@ static int setup_swap_extents(struct swap_info_struc= t *sis, struct inode *inode =3D mapping->host; int ret; =20 + if (sis->flags & SWP_XSWAP) { + *span =3D 0; + /* + * xswap devices have no backing block device and + * physical writeout is skipped in swap_writeout(), + * but sis->ops must still be set so that callers + * like shrink_folio_list() can safely dereference + * ops->flags. + */ + sis->ops =3D &swap_bdev_ops; + return 0; + } + ret =3D sio_pool_init(); if (ret) return ret; @@ -3160,7 +3178,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) =20 destroy_swap_extents(p, p->swap_file); =20 - if (!(p->flags & SWP_SOLIDSTATE)) + if (!(p->flags & SWP_XSWAP) && + !(p->flags & SWP_SOLIDSTATE)) atomic_dec(&nr_rotate_swap); =20 mutex_lock(&swapon_mutex); @@ -3270,6 +3289,19 @@ static void swap_stop(struct seq_file *swap, void *v) mutex_unlock(&swapon_mutex); } =20 +static const char *swap_type_str(struct swap_info_struct *si) +{ + struct file *file =3D si->swap_file; + + if (si->flags & SWP_XSWAP) + return "xswap\t"; + + if (S_ISBLK(file_inode(file)->i_mode)) + return "partition"; + + return "file\t"; +} + static int swap_show(struct seq_file *swap, void *v) { struct swap_info_struct *si =3D v; @@ -3289,8 +3321,7 @@ static int swap_show(struct seq_file *swap, void *v) len =3D seq_file_path(swap, file, " \t\n\\"); seq_printf(swap, "%*s%s\t%lu\t%s%lu\t%s%d\n", len < 40 ? 40 - len : 1, " ", - S_ISBLK(file_inode(file)->i_mode) ? - "partition" : "file\t", + swap_type_str(si), bytes, bytes < 10000000 ? "\t" : "", inuse, inuse < 10000000 ? "\t" : "", si->prio); @@ -3468,6 +3499,7 @@ static unsigned long read_swap_header(struct swap_inf= o_struct *si, unsigned long maxpages; unsigned long swapfilepages; unsigned long last_page; + loff_t size; =20 if (memcmp("SWAPSPACE2", swap_header->magic.magic, 10)) { pr_err("Unable to find swap-space signature\n"); @@ -3510,7 +3542,16 @@ static unsigned long read_swap_header(struct swap_in= fo_struct *si, =20 if (!maxpages) return 0; - swapfilepages =3D i_size_read(inode) >> PAGE_SHIFT; + + size =3D i_size_read(inode); + if (size =3D=3D PAGE_SIZE) { + /* xswap: a swap device with no backing storage */ + si->bdev =3D NULL; + si->flags |=3D SWP_XSWAP | SWP_SOLIDSTATE; + return maxpages; + } + + swapfilepages =3D size >> PAGE_SHIFT; if (swapfilepages && maxpages > swapfilepages) { pr_warn("Swap area shorter than signature indicates\n"); return 0; diff --git a/mm/zswap.c b/mm/zswap.c index f7c9c89f6449..6178465d0583 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -998,7 +998,12 @@ static int zswap_writeback_entry(struct zswap_entry *e= ntry, /* try to allocate swap cache folio */ si =3D get_swap_device(swpentry); if (!si) - return -EEXIST; + return -ENOENT; + + if (si->flags & SWP_XSWAP) { + put_swap_device(si); + return -EINVAL; + } =20 mpol =3D get_task_policy(current); folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, @@ -1535,7 +1540,7 @@ bool zswap_store(struct folio *folio) zswap_pool_put(pool); put_objcg: obj_cgroup_put(objcg); - if (!ret && zswap_pool_reached_full) + if (!ret && zswap_pool_reached_full && atomic_read(&nr_real_swapfiles)) queue_work(shrink_wq, &zswap_shrink_work); check_old: /* --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-170.mta0.migadu.com (out-170.mta0.migadu.com [91.218.175.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0BA5A23BD05 for ; Wed, 5 Aug 2026 07:54:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916457; cv=none; b=hFaxc6ChqPEfqDVAlNxqrfm/fVzJSN6VtMrJG0w6lJhX2rona5A1nSQBQ7aSI0b2PqiB7e7gfHaYtiaTnbRtk89cm0HzRWC7NPiYiNRyFk0trqNpE9tWJcuxXP7XwbKbMIJ0JRN5duEhO5g62a6MZqLtDM3sjx+lgLNcoy1lFWQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916457; c=relaxed/simple; bh=rniLTxIjlWcNpR517leeSpr0oIpPLO2dqGIPhD2euWk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-type; b=fhylA/JNSUE5CVdC79E7RAM2WXHRWNfCWN41EFsCoKcSOUh/GQk8c1J63QPvtxisTgzcPX1ef+qwhyjqmpmnG+WPrPl6wO69l08WlXvHxi5sKXRXrZ9nj0YVbxZMh6FF2AeDH84zb1TS9vTg4TIrOwnStR4L7qj9DnUR+Vhhdz0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=MSUvCXlN; arc=none smtp.client-ip=91.218.175.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="MSUvCXlN" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916453; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Vyg1JuKKGQsuDU1bKZ0uj2oDV98R+6DJe0buLRSbuTU=; b=MSUvCXlNL2jJBQKWRW6MwT/fPVE0vP3PeCSPlPH7WhqywWqLqS1Xh114SiWC01cikNT09/ vfXD2m/xesXSYNU3Uzfu6TaN8oGfy6pMdMom1JUDyeE9I7rEKk2LCr72CVqA+BGcdQoSY6 xBeD/z12DtBqmFC/8ZqDvGdWQfznWFU= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 02/10] mm, swap: add CONFIG_XSWAP and xswap fields to swap_info_struct Date: Wed, 5 Aug 2026 15:53:25 +0800 Message-ID: <20260805075336.3579395-3-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Add CONFIG_XSWAP Kconfig option (depends on SWAP && 64BIT) for extendable (virtual) swap device support. Add three fields to struct swap_info_struct under CONFIG_XSWAP: - cluster_vm: the VM_SPARSE vm_struct backing the cluster_info array - nr_clusters: total number of clusters in the xswap address space - nr_clusters_mapped: number of clusters currently mapped (lazy grow) These fields enable lazy vmalloc-based dynamic cluster management where the cluster_info array is backed by a sparse vmalloc area that grows on demand and shrinks when clusters are freed. Signed-off-by: Baoquan He --- include/linux/swap.h | 5 +++++ mm/Kconfig | 9 +++++++++ 2 files changed, 14 insertions(+) diff --git a/include/linux/swap.h b/include/linux/swap.h index 78f4302b92e1..970232f6359d 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -248,6 +248,11 @@ struct swap_info_struct { signed char type; /* strange name for an index */ unsigned int max; /* size of this swap device */ struct swap_cluster_info *cluster_info; /* cluster info. Only for SSD */ +#ifdef CONFIG_XSWAP + struct vm_struct *cluster_vm; /* VM_SPARSE area for xswap dynamic cluster= _info */ + unsigned long nr_clusters; /* total cluster count for xswap */ + unsigned long nr_clusters_mapped; /* currently mapped cluster count */ +#endif struct list_head free_clusters; /* free clusters list */ struct list_head full_clusters; /* full clusters list */ struct list_head nonfull_clusters[SWAP_NR_ORDERS]; diff --git a/mm/Kconfig b/mm/Kconfig index 060190e12bce..82f62f4c5b23 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -122,6 +122,15 @@ config ZSWAP_COMPRESSOR_DEFAULT default "zstd" if ZSWAP_COMPRESSOR_DEFAULT_ZSTD default "" =20 +config XSWAP + bool "Extendable (virtual) swap device" + depends on SWAP && 64BIT + help + Adds support for extendable swap devices (xswap) that decouple + PTE swap entries from physical backing storage. The cluster_info + array is backed by a sparse vmalloc area that grows and shrinks + on demand, avoiding static pre-allocation overhead. + config ZSMALLOC tristate =20 --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-182.mta0.migadu.com (out-182.mta0.migadu.com [91.218.175.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 292003D9680 for ; Wed, 5 Aug 2026 07:54:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916464; cv=none; b=amfB4pWRaJPoSkPvo4wxreDVfiWpFhljf2AGPzxlZa7ILEOOVuX8MyDIuDYDebQg3vs7E57oYtTuDLZ+Sb8gAa9e8z6NNeyE3MOBZ28kt8VO3f331yfmAYaOpNuH17Oxvssh0d/+68GCNrESXIqfLa0l4R1yVunFXoc0koE/TyQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916464; c=relaxed/simple; bh=g8q6gd1fOjNUMzl/RP/E6LLiEX8sL9Y1id9HTGMLPcI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-type; b=LpDnmEkzMCBBpP//1O0Tke+r+GOW3smYFIzfYGODIwXolXstW6PzMEd6Qns9scox8UwkW9xP+WqXN1Jr9bcxehMkGyx93SkNlQvhR69BA+hWNbmk98cMWodqe/xLcbhvluNn0QqDQVuI2vwLeFmxsgk3Tgw+mySWjPzx3U+dNSk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=AoC3OVWl; arc=none smtp.client-ip=91.218.175.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="AoC3OVWl" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916459; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=8Dy0TRGL3n4s1aM02P+TwJL4y7iAJ4YkyqwwdxQaTLI=; b=AoC3OVWlDqOuR0z8wqrCJCfGHI4NpcOpWGnvGnMNmLw6tDLDqWM5X1qByg7ZBZDk7Eg7Nr sgbL9dsduN4zSGhnGnbysHLo0lCfSIyCuwW4hL5tjdf47c4oAoR6wOpeH5UR9sBiCgjXSI qPR8I3N5cWLQIDuWEKpTAMQYIY2biMM= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 03/10] mm, swap: add xswap cluster grow via VM_SPARSE vmalloc Date: Wed, 5 Aug 2026 15:53:26 +0800 Message-ID: <20260805075336.3579395-4-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Implement dynamic cluster_info array growth for xswap devices using a VM_SPARSE vmalloc area: 1. xswap_map_clusters(): Allocate physical pages and map them into the pre-reserved VM_SPARSE KVA region via vm_area_map_pages(). 2. xswap_unmap_clusters(): Unmap pages from the VM_SPARSE area via vm_area_unmap_pages() (used by the error/teardown paths, shrink comes later). 3. setup_swap_clusters_info() xswap path: Use get_vm_area(VM_SPARSE) for the cluster_info array, lazily mapping only the initial chunk. 4. free_swap_cluster_info(): Refactor to take swap_info_struct*. For xswap, unmap all clusters and free_vm_area(). 5. swapoff: Remove snapshot locals; move p->max/p->cluster_info clearing after free_swap_cluster_info(). The grow path runs in the swap allocation context which may have PF_MEMALLOC set during reclaim, so avoid consuming emergency reserves by using __GFP_HIGH | __GFP_NOMEMALLOC on alloc_page() and kmalloc_array(), switching to vm_area_map_pages_gfp(), and wrapping the entire allocation block with memalloc_noreclaim_save() to prevent recursive reclaim from internal page table allocations. Concurrent grow operations race on vm_area_map_pages(), triggering WARN_ON(!pte_none) in the vmap page table walk. Add a per-device mutex (xswap_lock) held across xswap_map_clusters and xswap_unmap_clusters to serialize page table modifications. The shrink path is already deferred to a workqueue so it does not contend with itself. Signed-off-by: Baoquan He --- include/linux/swap.h | 1 + mm/swapfile.c | 256 +++++++++++++++++++++++++++++++++++++++++-- 2 files changed, 245 insertions(+), 12 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 970232f6359d..536b0e989c48 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -252,6 +252,7 @@ struct swap_info_struct { struct vm_struct *cluster_vm; /* VM_SPARSE area for xswap dynamic cluster= _info */ unsigned long nr_clusters; /* total cluster count for xswap */ unsigned long nr_clusters_mapped; /* currently mapped cluster count */ + struct mutex xswap_lock; /* serialize map/unmap operations */ #endif struct list_head free_clusters; /* free clusters list */ struct list_head full_clusters; /* full clusters list */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 08c49d5bea84..37c5dca153bc 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -49,6 +49,24 @@ #include "internal.h" #include "swap.h" =20 +#ifdef CONFIG_XSWAP +/* + * xswap: dynamically grow the cluster_info array via a VM_SPARSE area. + * + * XSWAP_GROW_CLUSTERS is the number of clusters to map in one grow + * operation. It is set to the number of cluster_info structs that + * fit in a single page (at least 16), so that the vmalloc page table + * overhead is proportional to the number of clusters mapped. + */ +#define XSWAP_GROW_CLUSTERS \ + max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16) + +static int xswap_map_clusters(struct swap_info_struct *si, + unsigned long start_idx, unsigned long nr); +static void xswap_unmap_clusters(struct swap_info_struct *si, + unsigned long start_idx, unsigned long nr); +#endif + static void swap_range_alloc(struct swap_info_struct *si, unsigned int nr_entries); static bool folio_swapcache_freeable(struct folio *folio); @@ -3041,20 +3059,47 @@ static void wait_for_allocation(struct swap_info_st= ruct *si) =20 BUG_ON(si->flags & SWP_WRITEOK); =20 +#ifdef CONFIG_XSWAP + /* + * xswap clusters beyond nr_clusters_mapped have been unmapped + * by the shrinker and their vmalloc pages are no longer + * accessible. Only iterate over currently mapped clusters. + */ + if (si->flags & SWP_XSWAP) + end =3D min(end, READ_ONCE(si->nr_clusters_mapped) * + SWAPFILE_CLUSTER); +#endif + for (offset =3D 0; offset < end; offset +=3D SWAPFILE_CLUSTER) { ci =3D swap_cluster_lock(si, offset); swap_cluster_unlock(ci); } } =20 -static void free_swap_cluster_info(struct swap_cluster_info *cluster_info, - unsigned long maxpages) +static void free_swap_cluster_info(struct swap_info_struct *si) { + struct swap_cluster_info *cluster_info =3D si->cluster_info; + unsigned long maxpages =3D si->max; struct swap_cluster_info *ci; - int i, nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); + int i, nr_clusters; =20 if (!cluster_info) return; + +#ifdef CONFIG_XSWAP + if (si->flags & SWP_XSWAP) { + /* Unmap all mapped clusters and free the VM_SPARSE area */ + if (si->nr_clusters_mapped > 0) + xswap_unmap_clusters(si, 0, si->nr_clusters_mapped); + free_vm_area(si->cluster_vm); + si->cluster_vm =3D NULL; + si->nr_clusters =3D 0; + si->nr_clusters_mapped =3D 0; + return; + } +#endif + + nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); for (i =3D 0; i < nr_clusters; i++) { ci =3D cluster_info + i; /* Cluster with bad marks count will have a remaining table */ @@ -3093,11 +3138,9 @@ static void flush_percpu_swap_cluster(struct swap_in= fo_struct *si) SYSCALL_DEFINE1(swapoff, const char __user *, specialfile) { struct swap_info_struct *p =3D NULL; - struct swap_cluster_info *cluster_info; struct file *swap_file, *victim; struct address_space *mapping; struct inode *inode; - unsigned int maxpages; int err, found =3D 0; =20 if (!capable(CAP_SYS_ADMIN)) @@ -3189,10 +3232,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specia= lfile) =20 swap_file =3D p->swap_file; p->swap_file =3D NULL; - maxpages =3D p->max; - cluster_info =3D p->cluster_info; - p->max =3D 0; - p->cluster_info =3D NULL; spin_unlock(&p->lock); spin_unlock(&swap_lock); arch_swap_invalidate_area(p->type); @@ -3200,7 +3239,9 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) mutex_unlock(&swapon_mutex); kfree(p->global_cluster); p->global_cluster =3D NULL; - free_swap_cluster_info(cluster_info, maxpages); + free_swap_cluster_info(p); + p->max =3D 0; + p->cluster_info =3D NULL; =20 inode =3D mapping->host; =20 @@ -3564,6 +3605,139 @@ static unsigned long read_swap_header(struct swap_i= nfo_struct *si, return maxpages; } =20 +#ifdef CONFIG_XSWAP +static int xswap_map_clusters(struct swap_info_struct *si, + unsigned long start_idx, unsigned long nr) +{ + unsigned long start_addr =3D (unsigned long)si->cluster_info + + (size_t)start_idx * sizeof(struct swap_cluster_info); + unsigned long end_addr =3D start_addr + (size_t)nr * sizeof(struct swap_c= luster_info); + /* + * vm_area_map_pages() requires that start and end be page-aligned. + * If start_addr falls within a page that was already mapped by a + * previous batch (grow path), round it up to skip the already-mapped + * partial page. Always round end_addr up so the vmap page table walk + * terminates correctly (the walk loop exits when addr =3D=3D end, and ad= dr + * advances by PAGE_SIZE each iteration). + */ + unsigned long vm_start =3D PAGE_ALIGN(start_addr); + unsigned long vm_end =3D PAGE_ALIGN(end_addr); + unsigned int noreclaim_flags; + unsigned long npages; + struct page **pages; + unsigned long i; + + mutex_lock(&si->xswap_lock); + + if (vm_start >=3D vm_end) { + /* All requested clusters fall within already-mapped pages. */ + for (i =3D start_idx; i < start_idx + nr; i++) + spin_lock_init(&si->cluster_info[i].lock); + WRITE_ONCE(si->nr_clusters_mapped, start_idx + nr); + mutex_unlock(&si->xswap_lock); + return 0; + } + + npages =3D (vm_end - vm_start) >> PAGE_SHIFT; + + /* + * Prevent recursive reclaim: vm_area_map_pages() internally + * allocates page tables with GFP_PGTABLE_KERNEL, which lacks + * __GFP_NOMEMALLOC. memalloc_noreclaim_save() ensures those + * allocations cannot recurse into swap by disabling __GFP_FS/IO. + */ + noreclaim_flags =3D memalloc_noreclaim_save(); + + pages =3D kmalloc_array(npages, sizeof(*pages), + __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL); + if (!pages) { + memalloc_noreclaim_restore(noreclaim_flags); + mutex_unlock(&si->xswap_lock); + return -ENOMEM; + } + + for (i =3D 0; i < npages; i++) { + /* + * __GFP_ZERO is critical: cluster_info structs contain pointer + * fields (extend_table, zero_bitmap, memcg_table, table) that + * must start as NULL. Without zeroing, stale data from a + * previous user of the page would look like valid pointers. + */ + pages[i] =3D alloc_page(__GFP_HIGH | __GFP_NOMEMALLOC | + GFP_KERNEL | __GFP_ZERO); + if (!pages[i]) + goto fail; + } + + if (vm_area_map_pages(si->cluster_vm, vm_start, vm_end, pages)) { + i =3D npages; /* free all pages on failure */ + goto fail; + } + + kfree(pages); + memalloc_noreclaim_restore(noreclaim_flags); + + /* Initialize spinlocks for newly mapped clusters */ + for (i =3D start_idx; i < start_idx + nr; i++) + spin_lock_init(&si->cluster_info[i].lock); + + /* + * Pairs with READ_ONCE() in shrink/grow paths. + */ + WRITE_ONCE(si->nr_clusters_mapped, start_idx + nr); + mutex_unlock(&si->xswap_lock); + return 0; + +fail: + while (i > 0) { + i--; + if (pages[i]) + __free_page(pages[i]); + } + memalloc_noreclaim_restore(noreclaim_flags); + kfree(pages); + mutex_unlock(&si->xswap_lock); + return -ENOMEM; +} + +static void xswap_unmap_clusters(struct swap_info_struct *si, + unsigned long start_idx, unsigned long nr) +{ + unsigned long start_addr =3D (unsigned long)si->cluster_info + + (size_t)start_idx * sizeof(struct swap_cluster_info); + unsigned long end_addr =3D start_addr + (size_t)nr * sizeof(struct swap_c= luster_info); + /* + * Round to page boundaries: start up (skip partial page that may + * contain clusters still in use before start_idx), end up so the + * entire range is covered. vm_area_unmap_pages() operates on + * whole pages. + */ + unsigned long vm_start =3D PAGE_ALIGN(start_addr); + unsigned long vm_end =3D PAGE_ALIGN(end_addr); + + mutex_lock(&si->xswap_lock); + + if (vm_start >=3D vm_end) { + mutex_unlock(&si->xswap_lock); + goto out; + } + + vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end); + /* + * vm_area_unmap_pages() only clears PTEs; it does not free the + * physical pages. Walk the page table to find and free them. + */ + /* TODO: free backing pages via page table walk or tracking bitmap */ + mutex_unlock(&si->xswap_lock); + +out: + /* + * Pairs with READ_ONCE() in shrink/grow paths. + */ + WRITE_ONCE(si->nr_clusters_mapped, start_idx); +} +#endif /* CONFIG_XSWAP */ + static int setup_swap_clusters_info(struct swap_info_struct *si, union swap_header *swap_header, unsigned long maxpages) @@ -3573,6 +3747,64 @@ static int setup_swap_clusters_info(struct swap_info= _struct *si, int err =3D -ENOMEM; unsigned long i; =20 +#ifdef CONFIG_XSWAP + if (si->flags & SWP_XSWAP) { + unsigned long size =3D PAGE_ALIGN(nr_clusters * sizeof(*cluster_info)); + struct vm_struct *vm; + + vm =3D get_vm_area(size, VM_SPARSE); + if (!vm) + goto err; + + cluster_info =3D vm->addr; + si->cluster_vm =3D vm; + si->nr_clusters =3D nr_clusters; + si->cluster_info =3D cluster_info; + + /* Map the initial chunk (at least cluster 0) */ + if (xswap_map_clusters(si, 0, min_t(unsigned long, + XSWAP_GROW_CLUSTERS, nr_clusters))) + goto err_free_vm; + + /* xswap: only cluster 0 slot 0 is bad */ + err =3D swap_cluster_setup_bad_slot(si, cluster_info, 0, false); + if (err) + goto err_unmap; + + INIT_LIST_HEAD(&si->free_clusters); + INIT_LIST_HEAD(&si->full_clusters); + INIT_LIST_HEAD(&si->discard_clusters); + for (i =3D 0; i < SWAP_NR_ORDERS; i++) { + INIT_LIST_HEAD(&si->nonfull_clusters[i]); + INIT_LIST_HEAD(&si->frag_clusters[i]); + } + + /* Mark mapped clusters: cluster 0 has 1 bad slot, rest free */ + for (i =3D 0; i < si->nr_clusters_mapped; i++) { + struct swap_cluster_info *ci =3D &cluster_info[i]; + + if (i =3D=3D 0) { + ci->flags =3D CLUSTER_FLAG_NONFULL; + list_add_tail(&ci->list, &si->nonfull_clusters[0]); + } else { + ci->flags =3D CLUSTER_FLAG_FREE; + list_add_tail(&ci->list, &si->free_clusters); + } + } + + mutex_init(&si->xswap_lock); + return 0; + +err_unmap: + xswap_unmap_clusters(si, 0, si->nr_clusters_mapped); +err_free_vm: + free_vm_area(si->cluster_vm); + si->cluster_vm =3D NULL; + si->cluster_info =3D NULL; + return err; + } +#endif /* CONFIG_XSWAP */ + cluster_info =3D kvzalloc_objs(*cluster_info, nr_clusters); if (!cluster_info) goto err; @@ -3640,7 +3872,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, si->cluster_info =3D cluster_info; return 0; err: - free_swap_cluster_info(cluster_info, maxpages); + free_swap_cluster_info(si); return err; } =20 @@ -3859,7 +4091,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) si->global_cluster =3D NULL; inode =3D NULL; destroy_swap_extents(si, swap_file); - free_swap_cluster_info(si->cluster_info, si->max); + free_swap_cluster_info(si); si->cluster_info =3D NULL; /* * Clear the SWP_USED flag after all resources are freed so --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-184.mta0.migadu.com (out-184.mta0.migadu.com [91.218.175.184]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3E78C3E6DC9 for ; Wed, 5 Aug 2026 07:54:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.184 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916474; cv=none; b=f9iZsG6czhAnmKxPX3UcnFrXp7FDtRYjDhUreq0A6+xhy3T+yN8Vyx2Ycb/o1bguVCmOj4n29tnhh0s1sgz1qxPH7juF6TZ2yTK+M9VHY1jpopR0CWeGpQN+ATZ0iowp18FSm39nq9onWUV1atdLwkNlZ1ahT+EwZORmbiscluk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916474; c=relaxed/simple; bh=5Zh6OHNqDCaWpKsIRkbDjfjKu/IbHR3clOGt78hFvV8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=EbUjQTnogZZxoQikbI69rQqlpSQaNLAgb4WhMpa+OZdgnV8rMbuORqdvOSwt1daYvr4LIteOPPuhBANr7ez/n5yYYDzXO/OQk6AigqiWD0m/Vlee2kYpjpJvASpCLITbJh2f4rf6ZSq2f5Z4DP0+xsn3RdP6LYZ09KisN+sCU4Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=fP8B4IaP; arc=none smtp.client-ip=91.218.175.184 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="fP8B4IaP" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916466; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=XpLCIEmlJUyWy6CaqTAtX3ScpMnuNBCh2QJ/2LrzeSU=; b=fP8B4IaPHFP1GjMdG9SHR1R6wSCWK42Uq8FnfQxxybz5rUrSwg6TyTomvz/tmRjvsXGbWp AZ+0rZurfV8VxOZs2EI0+H8SAjovddBweoba2G2ZTeF9y6AzYyJb4nO3KcHiUQA0j/nO+s Hoy3fXQiZaEADvfHxCHfIKuB9vq4hlc= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 04/10] mm, swap: add xswap grow trigger on cluster allocation Date: Wed, 5 Aug 2026 15:53:27 +0800 Message-ID: <20260805075336.3579395-5-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT When cluster_alloc_swap_entry() fails to find a free cluster and the xswap device still has room to grow, expand the mapped range by XSWAP_GROW_CLUSTERS clusters. Since xswap is always SWP_SOLIDSTATE, no locks need to be dropped before calling xswap_map_clusters() =E2=80=94 global_cluster_lock is never held on this path. The grow sequence: 1. Check nr_clusters_mapped < nr_clusters and free list empty 2. Call xswap_map_clusters() to allocate and map more physical pages 3. Add newly mapped clusters to si->free_clusters under si->lock 4. Retry allocation from the fresh free clusters This makes the xswap cluster space grow transparently as swap usage increases, without any userspace intervention. Signed-off-by: Baoquan He --- mm/swapfile.c | 101 +++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 100 insertions(+), 1 deletion(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 37c5dca153bc..5d8e10be0159 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -65,6 +65,7 @@ static int xswap_map_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr); static void xswap_unmap_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr); +static int xswap_check_mapped(pte_t *pte, unsigned long addr, void *data); #endif =20 static void swap_range_alloc(struct swap_info_struct *si, @@ -1204,6 +1205,48 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, if (found) goto done; } + +#ifdef CONFIG_XSWAP + /* + * For xswap: if no free cluster was found and more clusters + * can be mapped, grow the cluster_info array and retry. + */ + if (!found && (si->flags & SWP_XSWAP) && + READ_ONCE(si->nr_clusters_mapped) < READ_ONCE(si->nr_clusters) && + list_empty(&si->free_clusters)) { + unsigned long nr_new =3D min(READ_ONCE(si->nr_clusters) - + READ_ONCE(si->nr_clusters_mapped), + XSWAP_GROW_CLUSTERS); + unsigned long start =3D READ_ONCE(si->nr_clusters_mapped); + unsigned long i; + + if (!xswap_map_clusters(si, start, nr_new)) { + unsigned long added =3D 0; + + spin_lock(&si->lock); + for (i =3D start; i < start + nr_new; i++) { + struct swap_cluster_info *ci =3D &si->cluster_info[i]; + spin_lock(&ci->lock); + /* + * A concurrent grower may have already added + * these clusters to the free list. Only add + * clusters that are still off-list (NONE). + */ + if (ci->flags =3D=3D CLUSTER_FLAG_NONE) { + ci->flags =3D CLUSTER_FLAG_FREE; + list_add_tail(&ci->list, &si->free_clusters); + added++; + } + spin_unlock(&ci->lock); + } + spin_unlock(&si->lock); + + /* Retry allocation from the free list */ + found =3D alloc_swap_scan_list(si, &si->free_clusters, + folio, false); + } + } +#endif done: if (!(si->flags & SWP_SOLIDSTATE)) spin_unlock(&si->global_cluster_lock); @@ -3669,7 +3712,32 @@ static int xswap_map_clusters(struct swap_info_struc= t *si, goto fail; } =20 - if (vm_area_map_pages(si->cluster_vm, vm_start, vm_end, pages)) { + /* + * Check if the target pages are already mapped by a concurrent + * grower. We must do this after page allocation because + * alloc_page(GFP_KERNEL) can sleep, opening a race window. + * If someone already mapped these pages, free ours and continue. + */ + if (apply_to_existing_page_range(&init_mm, vm_start, + vm_end - vm_start, + xswap_check_mapped, NULL)) { + i =3D npages; + goto fail_nounmap; + } + + int err =3D vm_area_map_pages(si->cluster_vm, vm_start, vm_end, pages); + if (err) { + /* + * -EBUSY means the PTEs are already present: + * another thread raced with us and mapped the + * same pages between our check above and this + * call. Treat as success =E2=80=94 free our unused + * pages and continue. + */ + if (err =3D=3D -EBUSY) { + i =3D npages; + goto fail_nounmap; + } i =3D npages; /* free all pages on failure */ goto fail; } @@ -3688,6 +3756,28 @@ static int xswap_map_clusters(struct swap_info_struc= t *si, mutex_unlock(&si->xswap_lock); return 0; =20 +fail_nounmap: + /* + * Pages already mapped by a concurrent grower. Free our unused + * pages, then fall through to initialize spinlocks. The vmalloc + * PTEs now point to the concurrent grower's pages. + */ + while (i > 0) { + i--; + if (pages[i]) + __free_page(pages[i]); + } + kfree(pages); + memalloc_noreclaim_restore(noreclaim_flags); + + /* Initialize spinlocks for newly mapped clusters */ + for (i =3D start_idx; i < start_idx + nr; i++) + spin_lock_init(&si->cluster_info[i].lock); + + WRITE_ONCE(si->nr_clusters_mapped, start_idx + nr); + mutex_unlock(&si->xswap_lock); + return 0; + fail: while (i > 0) { i--; @@ -3736,6 +3826,15 @@ static void xswap_unmap_clusters(struct swap_info_st= ruct *si, */ WRITE_ONCE(si->nr_clusters_mapped, start_idx); } + +/* + * Callback for apply_to_existing_page_range(): return 1 to stop at the + * first present PTE, signalling that the range is already mapped. + */ +static int xswap_check_mapped(pte_t *pte, unsigned long addr, void *data) +{ + return 1; +} #endif /* CONFIG_XSWAP */ =20 static int setup_swap_clusters_info(struct swap_info_struct *si, --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-189.mta0.migadu.com (out-189.mta0.migadu.com [91.218.175.189]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8A5423D9680 for ; Wed, 5 Aug 2026 07:54:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.189 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916477; cv=none; b=SI8pt6N2VyoqDw4HBGUHSufTMfS9v+FeHGj+8oq+ydJLdTgRHTlWGUJOftiqGCfEIBkzcaNiH1utLWIu/+aL4QHwneJvI+Se/VinUkqa0GM5blZS0QwKAUpiU2bBT3w6iXUX9NDbU3DzOteMVSbwLUYM5LKAKmsHNjDbvVjbDMQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916477; c=relaxed/simple; bh=lnKOQpE6NFp+c7+J8jsraye3pE5Ig570bqpTZNlFz7g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=jMhrvs3oXSpHO98yznRYqffY5nswGXzDw4EjKpTUG8xU9KkgyiMDDw6JqyYzuldJEoWzldzndgQdvoFGew2i0K0wDHxwU6E13qpAAU4Vam44uEM4LZp/GSxURcFdsXYsv18xUnveMrjF5kLHXHqacmUrLdp3LcGO6FomMrFUNEY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=lwo+LBGq; arc=none smtp.client-ip=91.218.175.189 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="lwo+LBGq" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916472; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=A1v4kpNGwyglXBeEENPVH37QuycHMC781xF2cHBH8qU=; b=lwo+LBGqUGfl4wpOaO9E8CN/JxJ1ABjebdqeWt+2Jb5ttg/9qvrmXbXUodRe8Xw/8+/rLO wU1QcRvzLdgnMeilGnm/Ax9BmV2ojOea+t9HHZxWqTts1wDL8wHzA5CppS+bAcT7WRzqE7 cPXbnT7eE4cmO4RWb9v193K8nF5AYlI= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 05/10] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Date: Wed, 5 Aug 2026 15:53:28 +0800 Message-ID: <20260805075336.3579395-6-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Add xswap_try_shrink() =E2=80=94 the shrink logic that scans backwards from the tail to find contiguous free clusters, then unmaps full pages when >=3D XSWAP_GROW_CLUSTERS free clusters accumulate. Wire the trigger in __free_cluster(): after a cluster is released to the free list, call xswap_try_shrink() to attempt tail shrinking. Also update the XSWAP_GROW_CLUSTERS comment to reflect both grow and shrink semantics. Signed-off-by: Baoquan He --- mm/swapfile.c | 48 ++++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 44 insertions(+), 4 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 5d8e10be0159..0d1d0c5e24c6 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -51,11 +51,13 @@ =20 #ifdef CONFIG_XSWAP /* - * xswap: dynamically grow the cluster_info array via a VM_SPARSE area. + * xswap: dynamically grow and shrink the cluster_info array via a + * VM_SPARSE area. * - * XSWAP_GROW_CLUSTERS is the number of clusters to map in one grow - * operation. It is set to the number of cluster_info structs that - * fit in a single page (at least 16), so that the vmalloc page table + * XSWAP_GROW_CLUSTERS is the number of clusters to map/unmap in one + * grow/shrink operation. It is set to the number of cluster_info + * structs that fit in a single page (at least 16), so that the vmalloc + * page table * overhead is proportional to the number of clusters mapped. */ #define XSWAP_GROW_CLUSTERS \ @@ -66,6 +68,7 @@ static int xswap_map_clusters(struct swap_info_struct *si, static void xswap_unmap_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr); static int xswap_check_mapped(pte_t *pte, unsigned long addr, void *data); +static void xswap_try_shrink(struct swap_info_struct *si); #endif =20 static void swap_range_alloc(struct swap_info_struct *si, @@ -629,6 +632,9 @@ static void __free_cluster(struct swap_info_struct *si,= struct swap_cluster_info swap_cluster_free_table(ci); move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE); ci->order =3D 0; +#ifdef CONFIG_XSWAP + xswap_try_shrink(si); +#endif } =20 /* @@ -3835,6 +3841,40 @@ static int xswap_check_mapped(pte_t *pte, unsigned l= ong addr, void *data) { return 1; } + +/* + * Try to shrink the cluster_info tail: unmap contiguous free clusters + * at the end of the mapped range. + */ +static void xswap_try_shrink(struct swap_info_struct *si) +{ + struct swap_cluster_info *ci; + unsigned long last, idx; + + if (!(si->flags & SWP_XSWAP)) + return; + if (READ_ONCE(si->nr_clusters_mapped) <=3D 1) /* keep cluster 0 */ + return; + + /* Find the last non-free cluster from the tail */ + last =3D READ_ONCE(si->nr_clusters_mapped); + while (last > 1) { + idx =3D last - 1; + ci =3D &si->cluster_info[idx]; + if (ci->count || ci->flags !=3D CLUSTER_FLAG_FREE) + break; + last =3D idx; + } + + if (last =3D=3D si->nr_clusters_mapped) + return; /* nothing to shrink */ + + /* Only unmap if we can free at least one full page of clusters */ + if (si->nr_clusters_mapped - last < XSWAP_GROW_CLUSTERS) + return; + + xswap_unmap_clusters(si, last, si->nr_clusters_mapped - last); +} #endif /* CONFIG_XSWAP */ =20 static int setup_swap_clusters_info(struct swap_info_struct *si, --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-175.mta1.migadu.com (out-175.mta1.migadu.com [95.215.58.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2CF423E5EC0 for ; Wed, 5 Aug 2026 07:54:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916498; cv=none; b=YKp4FlssDzCrvA3aflRCFIiSwGgySjGEAoqsQmNNR2Z98xW/nDhjssfs3zW1ME5HRqXMgD7LdM9hBEGlBDNeJ0EufnoB/R7KIw3mqYGBezc8gAc9IofV2JW8k+Il4qLBtbRMUGuV17w8xNUUfgS6gtVBBebRoNajYg/ZqgxxmuE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916498; c=relaxed/simple; bh=RejxSqib9aBuLsvfbPxFQxtz4EFWAblNB5sG5n0YL+c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=ZqZLYlE4n4Rkiltto8nmwxwErPvAhT31iO1Gd1th/FxCqg2/Xo9gUuOaEhJA4eX/SZIR3nu58DCGIxYkl3qHyYwoCWpUae0l7q7+i/VmQkq6ejMcaxWGqxBd4PChVHo0PBrFoMUiyPXLhfWHP4zCGlh/AnfaLwQg81N5WPxjhw4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=WTf6PIsg; arc=none smtp.client-ip=95.215.58.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="WTf6PIsg" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916493; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=KQPF9NpoXms5JqUk+wWG0b5ssSUK6sTaQPnJLeO/Wgc=; b=WTf6PIsgJKrF3nj0UfO/kXgdoFHWcAZIMMiQtM5TabV1fluZmpgG0uKdbdZKRSg62tLfXT sFrrnvfe5jPh5tVLZVdGvzdHlRyVVhmPUwQeLsbGawGSKzWB+6MVaKMxIW99TXl3MxwIJM ZnAaw97ElGJcefECReEWYaPZmZChrqU= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 06/10] mm, swap: free backing pages in xswap_unmap_clusters Date: Wed, 5 Aug 2026 15:53:29 +0800 Message-ID: <20260805075336.3579395-7-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT vm_area_unmap_pages() only clears PTEs and frees intermediate page table pages =E2=80=94 it does not free the backing physical pages allocated by xswap_map_clusters(). Fix this by walking the page table with apply_to_existing_page_range() before the unmap to collect all struct pages in the range. After vunmap_range() clears the PTEs, free the collected pages via __free_page(). Use a simple xswap_page_data collector callback: for each present PTE, collect pte_page() into a dynamically allocated array. The array is freed after the pages are released. Signed-off-by: Baoquan He --- mm/swapfile.c | 52 +++++++++++++++++++++++++++++++++++++++++++++------ 1 file changed, 46 insertions(+), 6 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 0d1d0c5e24c6..c43c8746378e 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -3796,6 +3796,23 @@ static int xswap_map_clusters(struct swap_info_struc= t *si, return -ENOMEM; } =20 +struct xswap_page_data { + struct page **pages; + int nr; + int max; +}; + +static int xswap_collect_page(pte_t *pte, unsigned long addr, void *data) +{ + struct xswap_page_data *xpd =3D data; + + if (!pte_present(*pte)) + return 0; + if (xpd->nr < xpd->max) + xpd->pages[xpd->nr++] =3D pte_page(*pte); + return 0; +} + static void xswap_unmap_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr) { @@ -3805,11 +3822,15 @@ static void xswap_unmap_clusters(struct swap_info_s= truct *si, /* * Round to page boundaries: start up (skip partial page that may * contain clusters still in use before start_idx), end up so the - * entire range is covered. vm_area_unmap_pages() operates on - * whole pages. + * entire range is covered. vm_area_unmap_pages() and + * apply_to_existing_page_range() operate on whole pages. */ unsigned long vm_start =3D PAGE_ALIGN(start_addr); unsigned long vm_end =3D PAGE_ALIGN(end_addr); + unsigned long size; + unsigned long npages; + struct xswap_page_data xpd; + int i; =20 mutex_lock(&si->xswap_lock); =20 @@ -3818,12 +3839,31 @@ static void xswap_unmap_clusters(struct swap_info_s= truct *si, goto out; } =20 - vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end); + size =3D vm_end - vm_start; + npages =3D size >> PAGE_SHIFT; + /* - * vm_area_unmap_pages() only clears PTEs; it does not free the - * physical pages. Walk the page table to find and free them. + * Walk the page table to collect physical pages before unmapping. + * vm_area_unmap_pages() only clears PTEs and frees intermediate + * page table pages =E2=80=94 it does not free backing pages. */ - /* TODO: free backing pages via page table walk or tracking bitmap */ + xpd.pages =3D kmalloc_array(npages, sizeof(*xpd.pages), GFP_KERNEL); + if (xpd.pages) { + xpd.nr =3D 0; + xpd.max =3D npages; + apply_to_existing_page_range(&init_mm, vm_start, size, + xswap_collect_page, &xpd); + } + + vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end); + + /* Free the collected backing pages */ + if (xpd.pages) { + for (i =3D 0; i < xpd.nr; i++) + __free_page(xpd.pages[i]); + kfree(xpd.pages); + } + mutex_unlock(&si->xswap_lock); =20 out: --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-182.mta0.migadu.com (out-182.mta0.migadu.com [91.218.175.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 144053E4508 for ; Wed, 5 Aug 2026 07:55:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916514; cv=none; b=QysU5YGg5g/76oRtFCa115rDypqy5R8AmOIxXz2mFbsbx581qZmS581Up2Ifx272J2BmvBDstRprLXmm3ZF9I85Lt1ihTIILzlCRjyosfYFV8t8gR/FLveTZZmWYKGoul9NchSark5pv2do54FNMmrqZs3CNmPkozByObUS/2Xg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916514; c=relaxed/simple; bh=fntQdSEMbh/KmHgZw8c7vJR0PXNZ6tMLvHrYqumv8lg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=UizJ2OBX+9/s2zT9NXZp7/1y7sF+zmgBCpgOM8B/7QFJ8rjIoYRhbjdmlVuw2ZDQZS/sB6GEWKPZTwBL9ydXQLf2OiXf37JqXR/whw4mfdDLAzon+auqVngGXsypwlvYzb2byQsHpgN8pPN4li2mCPpevkKwk84MAju7Y2JJtdk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=Lx78EHCd; arc=none smtp.client-ip=91.218.175.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="Lx78EHCd" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916510; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=QdHk8BKCu5NpvOOk3FzW3mOqj3l0dtnoNHZ0kMxVMyE=; b=Lx78EHCdPBRTQG/pdg90gUbgWcqQGF1jBtSYPeOk3wTe4nmTceDcudcIlIobgcmt+Ry7kx l08+xVJ6/nAMcNaJSQ4FKeud6Ym7IKkNj7NSuvaqm+/ufcCqNbR3386ENQDFeIfYpBZBZt CokuTCBgeKebSIFE6F/avpEFpWnsLHM= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 07/10] mm, swap: add nr_free_tail for O(1) xswap shrink detection Date: Wed, 5 Aug 2026 15:53:30 +0800 Message-ID: <20260805075336.3579395-8-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Track contiguous free clusters at the tail of the mapped range in si->nr_free_tail, maintained across alloc/free/grow paths. This eliminates the backwards scan on every shrink check. Three paths maintain the counter: 1. xswap_update_free_tail(): called on cluster free. If the freed cluster is adjacent to the existing tail boundary, increment and extend backwards to include already-free clusters now connected. 2. xswap_trim_free_tail(): called on cluster allocation. If the allocated cluster lies within the tail free region, truncate the count to end just before it. 3. Grow path: nr_free_tail +=3D nr_new =E2=80=94 all newly mapped clusters are immediately free. xswap_try_shrink() simplifies to a threshold check: if nr_free_tail >=3D XSWAP_GROW_CLUSTERS =E2=86=92 unmap Setup initializes nr_free_tail =3D nr_clusters_mapped - 1 (all but cluster 0 are free at the tail). Signed-off-by: Baoquan He --- include/linux/swap.h | 1 + mm/swapfile.c | 121 +++++++++++++++++++++++++++++++++++++------ 2 files changed, 105 insertions(+), 17 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 536b0e989c48..72b28116ed0f 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -252,6 +252,7 @@ struct swap_info_struct { struct vm_struct *cluster_vm; /* VM_SPARSE area for xswap dynamic cluster= _info */ unsigned long nr_clusters; /* total cluster count for xswap */ unsigned long nr_clusters_mapped; /* currently mapped cluster count */ + unsigned long nr_free_tail; /* contiguous free clusters at tail */ struct mutex xswap_lock; /* serialize map/unmap operations */ #endif struct list_head free_clusters; /* free clusters list */ diff --git a/mm/swapfile.c b/mm/swapfile.c index c43c8746378e..3f536495b8cf 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -68,6 +68,9 @@ static int xswap_map_clusters(struct swap_info_struct *si, static void xswap_unmap_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr); static int xswap_check_mapped(pte_t *pte, unsigned long addr, void *data); +static void xswap_trim_free_tail(struct swap_info_struct *si, unsigned lon= g idx); +static void xswap_update_free_tail(struct swap_info_struct *si, + unsigned long freed_idx); static void xswap_try_shrink(struct swap_info_struct *si); #endif =20 @@ -633,6 +636,7 @@ static void __free_cluster(struct swap_info_struct *si,= struct swap_cluster_info move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE); ci->order =3D 0; #ifdef CONFIG_XSWAP + xswap_update_free_tail(si, ci - si->cluster_info); xswap_try_shrink(si); #endif } @@ -981,6 +985,9 @@ static bool __swap_cluster_alloc_entries(struct swap_in= fo_struct *si, if (cluster_is_empty(ci)) ci->order =3D order; ci->count +=3D nr_pages; +#ifdef CONFIG_XSWAP + xswap_trim_free_tail(si, cluster_index(si, ci)); +#endif swap_range_alloc(si, nr_pages); =20 return true; @@ -1246,6 +1253,8 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, spin_unlock(&ci->lock); } spin_unlock(&si->lock); + WRITE_ONCE(si->nr_free_tail, + READ_ONCE(si->nr_free_tail) + nr_new); =20 /* Retry allocation from the free list */ found =3D alloc_swap_scan_list(si, &si->free_clusters, @@ -3883,37 +3892,115 @@ static int xswap_check_mapped(pte_t *pte, unsigned= long addr, void *data) } =20 /* - * Try to shrink the cluster_info tail: unmap contiguous free clusters - * at the end of the mapped range. + * Maintain si->nr_free_tail, the number of contiguous free clusters at + * the tail of the mapped range. Called when a cluster at @freed_idx is + * freed. Provides O(1) shrink detection: if nr_free_tail is non-zero, + * the tail can be unmapped without scanning cluster_info[]. + * + * Only increments when @freed_idx is the cluster immediately before the + * existing tail region. Then scans backwards for already-free clusters + * now connected to the tail, bounded by XSWAP_GROW_CLUSTERS at a time. */ -static void xswap_try_shrink(struct swap_info_struct *si) +static void xswap_update_free_tail(struct swap_info_struct *si, + unsigned long freed_idx) { + unsigned long nr_mapped, nr_tail, tid, i; struct swap_cluster_info *ci; - unsigned long last, idx; =20 if (!(si->flags & SWP_XSWAP)) return; - if (READ_ONCE(si->nr_clusters_mapped) <=3D 1) /* keep cluster 0 */ + + nr_mapped =3D READ_ONCE(si->nr_clusters_mapped); + nr_tail =3D READ_ONCE(si->nr_free_tail); + + /* Protect against concurrent shrink that races past us */ + if (nr_tail >=3D nr_mapped) return; =20 - /* Find the last non-free cluster from the tail */ - last =3D READ_ONCE(si->nr_clusters_mapped); - while (last > 1) { - idx =3D last - 1; - ci =3D &si->cluster_info[idx]; - if (ci->count || ci->flags !=3D CLUSTER_FLAG_FREE) + tid =3D nr_mapped - nr_tail - 1; + + /* Only the cluster immediately before the tail region counts */ + if (freed_idx !=3D tid) + return; + + nr_tail++; + WRITE_ONCE(si->nr_free_tail, nr_tail); + + /* Extend: include already-free clusters now connected to the tail */ + for (i =3D 1; i < XSWAP_GROW_CLUSTERS; i++) { + nr_mapped =3D READ_ONCE(si->nr_clusters_mapped); + nr_tail =3D READ_ONCE(si->nr_free_tail); + if (nr_tail >=3D nr_mapped - 1) + break; /* reached cluster 0 */ + tid =3D nr_mapped - nr_tail - 1; + ci =3D &si->cluster_info[tid]; + + if (READ_ONCE(ci->count) || + READ_ONCE(ci->flags) !=3D CLUSTER_FLAG_FREE) break; - last =3D idx; + nr_tail++; + WRITE_ONCE(si->nr_free_tail, nr_tail); } +} + +/* + * Trim si->nr_free_tail when a cluster in the tail region is allocated. + * @idx: index of the cluster being allocated. + */ +static void xswap_trim_free_tail(struct swap_info_struct *si, unsigned lon= g idx) +{ + unsigned long nr_mapped, nr_tail, tail_start; =20 - if (last =3D=3D si->nr_clusters_mapped) - return; /* nothing to shrink */ + if (!(si->flags & SWP_XSWAP)) + return; + + /* + * nr_clusters_mapped and nr_free_tail are read locklessly; + * concurrent updates may cause nr_free_tail to be trimmed + * slightly less than ideally, which is harmless. + */ + nr_mapped =3D READ_ONCE(si->nr_clusters_mapped); + nr_tail =3D READ_ONCE(si->nr_free_tail); + tail_start =3D nr_mapped - nr_tail; + if (idx >=3D tail_start) + WRITE_ONCE(si->nr_free_tail, nr_mapped - idx - 1); +} + +/* + * Try to shrink the cluster_info tail. Uses si->nr_free_tail which + * is maintained incrementally during alloc/free =E2=80=94 no scanning nee= ded. + */ +static void xswap_try_shrink(struct swap_info_struct *si) +{ + unsigned long start_idx, nr_unmap, i; + struct swap_cluster_info *ci; + + if (!(si->flags & SWP_XSWAP)) + return; + if (si->nr_free_tail < XSWAP_GROW_CLUSTERS) + return; + + nr_unmap =3D round_down(si->nr_free_tail, XSWAP_GROW_CLUSTERS); + start_idx =3D si->nr_clusters_mapped - nr_unmap; + + /* Verify the tail clusters are still free before unmapping */ + spin_lock(&si->lock); + for (i =3D start_idx; i < si->nr_clusters_mapped; i++) { + ci =3D &si->cluster_info[i]; + if (ci->flags !=3D CLUSTER_FLAG_FREE) { + nr_unmap =3D i - start_idx; + break; + } + list_del(&ci->list); + ci->flags =3D CLUSTER_FLAG_NONE; + } + spin_unlock(&si->lock); =20 - /* Only unmap if we can free at least one full page of clusters */ - if (si->nr_clusters_mapped - last < XSWAP_GROW_CLUSTERS) + if (nr_unmap < XSWAP_GROW_CLUSTERS) return; =20 - xswap_unmap_clusters(si, last, si->nr_clusters_mapped - last); + xswap_unmap_clusters(si, start_idx, nr_unmap); + si->nr_free_tail -=3D nr_unmap; } #endif /* CONFIG_XSWAP */ =20 --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-178.mta0.migadu.com (out-178.mta0.migadu.com [91.218.175.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 42AD53E5A2B for ; Wed, 5 Aug 2026 07:55:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916523; cv=none; b=p34vaRVK4F/aqWvpAuVX92Ku/bXWm/vnZKsjLyTsO/5J/y5XSvJWmF0p+NVctGxZaD+tCGWiLdKu55jq2v+NHLfM6Sf+buDzKjpLfP33jUYTyv0/OWrGHt5EMddrlYFA46QBEV6QSCrbIP6JpRj5P2EBp0erP+thBUrwNiqyPVU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916523; c=relaxed/simple; bh=8vxDYPJT7zxVP55sojXBKnEYKhnuZozdf0JkOyXR3EY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=jKx8Me5mhpszgAZc3kYZ9wdBNck04o/1wwYgbT8EqNp2FAGZW4Np5b6wwmKnjgvyudarEbEyOsoIlfFgEEtHDI3JmdPVTojCAXTiJn85N6lR2KnSChwrLd+buObT6/bzQ9VcixssSmnnAUVUHAFF8fhLkrklz3/6EI2BGZJeFVY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=PlSv69tq; arc=none smtp.client-ip=91.218.175.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="PlSv69tq" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916519; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=U88ICcPSN2njwbDn4B5fxEExcCgx6VDyaCcBrToBcqI=; b=PlSv69tqP76D2gH+DzpHn9NJAtmrWNmjrsjXO2gTTvQO8sSuDIk1Uv3fYkWE4mxeIfCDix ecWzoHmDJlrdX+zGS+47GIEeOxLMRuEWuV6UrA0qYE/sezgXnuhjRSlZ142pLut59SZ9+u EkiXDQpBJeKa0PfL3qo/aaNaCzKbiHI= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 08/10] mm, swap: add adjustable runtime ceiling (nr_clusters) for xswap Date: Wed, 5 Aug 2026 15:53:31 +0800 Message-ID: <20260805075336.3579395-9-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Split the xswap cluster limit into two fields: - nr_clusters_max: immutable hard limit set at swapon from swap header - nr_clusters: current growth ceiling, adjustable at runtime (=E2=89=A4 nr_= clusters_max) The grow path already uses nr_clusters as the ceiling. Shrink now also respects it: when nr_clusters drops below nr_clusters_mapped, shrinking fires on free until the mapped count reaches the ceiling. When nr_clusters =3D=3D nr_clusters_max (default), shrink is effectively disabled =E2=80=94 all growth and no shrink. At swapon, nr_clusters starts at nr_clusters_max (full size). Signed-off-by: Baoquan He --- include/linux/swap.h | 3 ++- mm/swapfile.c | 31 ++++++++++++++++++++++--------- 2 files changed, 24 insertions(+), 10 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 72b28116ed0f..1159153459a1 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -250,7 +250,8 @@ struct swap_info_struct { struct swap_cluster_info *cluster_info; /* cluster info. Only for SSD */ #ifdef CONFIG_XSWAP struct vm_struct *cluster_vm; /* VM_SPARSE area for xswap dynamic cluster= _info */ - unsigned long nr_clusters; /* total cluster count for xswap */ + unsigned long nr_clusters_max;/* upper limit from swap header */ + unsigned long nr_clusters; /* current growth ceiling (=E2=89=A4 nr_clust= ers_max) */ unsigned long nr_clusters_mapped; /* currently mapped cluster count */ unsigned long nr_free_tail; /* contiguous free clusters at tail */ struct mutex xswap_lock; /* serialize map/unmap operations */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 3f536495b8cf..3037f428f217 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -72,6 +72,7 @@ static void xswap_trim_free_tail(struct swap_info_struct = *si, unsigned long idx) static void xswap_update_free_tail(struct swap_info_struct *si, unsigned long freed_idx); static void xswap_try_shrink(struct swap_info_struct *si); + #endif =20 static void swap_range_alloc(struct swap_info_struct *si, @@ -3151,6 +3152,7 @@ static void free_swap_cluster_info(struct swap_info_s= truct *si) xswap_unmap_clusters(si, 0, si->nr_clusters_mapped); free_vm_area(si->cluster_vm); si->cluster_vm =3D NULL; + si->nr_clusters_max =3D 0; si->nr_clusters =3D 0; si->nr_clusters_mapped =3D 0; return; @@ -3972,20 +3974,31 @@ static void xswap_trim_free_tail(struct swap_info_s= truct *si, unsigned long idx) */ static void xswap_try_shrink(struct swap_info_struct *si) { - unsigned long start_idx, nr_unmap, i; + unsigned long nr_mapped, nr_ceiling, nr_tail, nr_unmap; + unsigned long start_idx, i; struct swap_cluster_info *ci; =20 if (!(si->flags & SWP_XSWAP)) return; - if (si->nr_free_tail < XSWAP_GROW_CLUSTERS) + + nr_mapped =3D READ_ONCE(si->nr_clusters_mapped); + nr_ceiling =3D READ_ONCE(si->nr_clusters); + nr_tail =3D READ_ONCE(si->nr_free_tail); + + if (nr_mapped <=3D nr_ceiling) + return; + if (nr_tail < XSWAP_GROW_CLUSTERS) return; =20 - nr_unmap =3D round_down(si->nr_free_tail, XSWAP_GROW_CLUSTERS); - start_idx =3D si->nr_clusters_mapped - nr_unmap; + nr_unmap =3D min(round_down(nr_tail, XSWAP_GROW_CLUSTERS), + nr_mapped - nr_ceiling); + if (nr_unmap < XSWAP_GROW_CLUSTERS) + return; + start_idx =3D nr_mapped - nr_unmap; =20 /* Verify the tail clusters are still free before unmapping */ spin_lock(&si->lock); - for (i =3D start_idx; i < si->nr_clusters_mapped; i++) { + for (i =3D start_idx; i < nr_mapped; i++) { ci =3D &si->cluster_info[i]; if (ci->flags !=3D CLUSTER_FLAG_FREE) { nr_unmap =3D i - start_idx; @@ -4000,7 +4013,7 @@ static void xswap_try_shrink(struct swap_info_struct = *si) return; =20 xswap_unmap_clusters(si, start_idx, nr_unmap); - si->nr_free_tail -=3D nr_unmap; + WRITE_ONCE(si->nr_free_tail, nr_tail - nr_unmap); } #endif /* CONFIG_XSWAP */ =20 @@ -4024,6 +4037,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, =20 cluster_info =3D vm->addr; si->cluster_vm =3D vm; + si->nr_clusters_max =3D nr_clusters; si->nr_clusters =3D nr_clusters; si->cluster_info =3D cluster_info; =20 @@ -4058,6 +4072,8 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, } } =20 + /* All mapped clusters except cluster 0 are free at the tail */ + si->nr_free_tail =3D si->nr_clusters_mapped - 1; mutex_init(&si->xswap_lock); return 0; =20 @@ -4469,7 +4485,6 @@ void __folio_throttle_swaprate(struct folio *folio, g= fp_t gfp) static int __init swapfile_init(void) { swapfile_maximum_size =3D arch_max_swapfile_size(); - /* * Once a cluster is freed, it's swap table content is read * only, and all swap cache readers (swap_cache_*) verifies @@ -4479,12 +4494,10 @@ static int __init swapfile_init(void) swap_table_cachep =3D kmem_cache_create("swap_table", sizeof(struct swap_table), 0, SLAB_PANIC | SLAB_TYPESAFE_BY_RCU, NULL); - #ifdef CONFIG_MIGRATION if (swapfile_maximum_size >=3D (1UL << SWP_MIG_TOTAL_BITS)) swap_migration_ad_supported =3D true; #endif /* CONFIG_MIGRATION */ - return 0; } subsys_initcall(swapfile_init); --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-186.mta0.migadu.com (out-186.mta0.migadu.com [91.218.175.186]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E28F923BD05 for ; Wed, 5 Aug 2026 07:55:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.186 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916529; cv=none; b=SSJRbuwDReXWpWQHEmU17OQuamjySDQBfrNu5+L3E7DxBBCyxKTpwdBtxGa4x9hef0iRK4dAQTdPg1uGHjQJ8I0SW6NSvxrjxHaxEjqeWJNLsRKdsNTBJhwZLijNCb0LELMq9haerbERRg0iEaIW5ydV4UwaegfIQ3WKU9GDn6A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916529; c=relaxed/simple; bh=0we31HKYZr34O6Xt3px9AkXvZRqjTZ5SxzVApMods9w=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-type; b=P4oNXGW68PcRWilHKrex+hySLEkSD+AO4Zr4Sh45D3v0V+t1/RG0AotySupz3FqfionD/HaNeCa2fvDmG6tKWAMNSQbko/ELzSoP/K6yaPr/z/E+Wc/ophK8z1K+u4U1ez7zCiYp5YdeprpUuAPBle/IXc/voX+kTMgX3/t0/Oo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=pyc4i7e3; arc=none smtp.client-ip=91.218.175.186 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="pyc4i7e3" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916525; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=3lEj0de+6RdDBzrWHGVYXCvS0Mqk6pZo6dba7WT6t2E=; b=pyc4i7e3V7GAwOHRbjgP022+4yb0wxUFzkJor0j4yxH7X3QZ6m6S92NaTu7tPaeMdETqz4 Kxp+0xEtMUKaOHZyFsEclJU/kMdGQbRZo/bfns17uVIaW92yeQhcoCGqsYSbDK8JfkH6aW o1v7Fm/ifbDVSk4DoPKF5mdxhc4gLg4= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 09/10] mm, swap: add debugfs knob for xswap per-device cluster limit Date: Wed, 5 Aug 2026 15:53:32 +0800 Message-ID: <20260805075336.3579395-10-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Add a per-device debugfs file for runtime adjustment of xswap cluster limit: /sys/kernel/debug/xswap/type_cluster_limit Reading shows the current ceiling (in clusters); writing sets it (clamped to [0, nr_clusters_max]). Setting below nr_clusters_mapped triggers an immediate shrink check via xswap_try_shrink(). The debugfs entry is created at swapon and removed at swapoff. Signed-off-by: Baoquan He --- include/linux/swap.h | 1 + mm/swapfile.c | 97 +++++++++++++++++++++++++++++++++++++++++++- 2 files changed, 97 insertions(+), 1 deletion(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 1159153459a1..4f4583f9a4e5 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -254,6 +254,7 @@ struct swap_info_struct { unsigned long nr_clusters; /* current growth ceiling (\ufffd\ufffd\ufffd= nr_clusters_max) */ unsigned long nr_clusters_mapped; /* currently mapped cluster count */ unsigned long nr_free_tail; /* contiguous free clusters at tail */ + struct dentry *debugfs_entry; /* debugfs: type_max_clusters */ struct mutex xswap_lock; /* serialize map/unmap operations */ #endif struct list_head free_clusters; /* free clusters list */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 3037f428f217..265f2bac1a12 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -48,6 +48,9 @@ #include "swap_table.h" #include "internal.h" #include "swap.h" +#include + +static DEFINE_SPINLOCK(swap_lock); =20 #ifdef CONFIG_XSWAP /* @@ -63,6 +66,8 @@ #define XSWAP_GROW_CLUSTERS \ max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16) =20 +static struct dentry *xswap_debugfs_root; + static int xswap_map_clusters(struct swap_info_struct *si, unsigned long start_idx, unsigned long nr); static void xswap_unmap_clusters(struct swap_info_struct *si, @@ -73,6 +78,90 @@ static void xswap_update_free_tail(struct swap_info_stru= ct *si, unsigned long freed_idx); static void xswap_try_shrink(struct swap_info_struct *si); =20 +/* + * debugfs read/write for per-device max cluster count. + * Shows/sets si->nr_clusters (current growth ceiling), clamped to + * [0, si->nr_clusters_max]. + */ +static ssize_t xswap_max_clusters_read(struct file *file, char __user *buf, + size_t count, loff_t *ppos) +{ + struct swap_info_struct *si =3D file->private_data; + char tmp[32]; + int len; + + len =3D snprintf(tmp, sizeof(tmp), "%lu\n", READ_ONCE(si->nr_clusters)); + return simple_read_from_buffer(buf, count, ppos, tmp, len); +} + +static ssize_t xswap_max_clusters_write(struct file *file, + const char __user *buf, + size_t count, loff_t *ppos) +{ + struct swap_info_struct *si =3D file->private_data; + unsigned long val, new_pages; + int err; + + err =3D kstrtoul_from_user(buf, count, 0, &val); + if (err) + return err; + + if (val > si->nr_clusters_max) + val =3D si->nr_clusters_max; + + spin_lock(&si->lock); + si->nr_clusters =3D val; + spin_unlock(&si->lock); + + /* Keep the visible swap size in sync with the new ceiling. */ + new_pages =3D min_t(unsigned long, val * SWAPFILE_CLUSTER, si->max); + if (new_pages) + new_pages--; + if (new_pages !=3D si->pages) { + long delta =3D (long)new_pages - (long)si->pages; + + spin_lock(&swap_lock); + si->pages =3D new_pages; + atomic_long_add(delta, &nr_swap_pages); + total_swap_pages +=3D delta; + spin_unlock(&swap_lock); + } + + /* + * Lowering the ceiling may make tail clusters eligible for + * shrinking. Trigger an immediate check. + */ + xswap_try_shrink(si); + + return count; +} + +static const struct file_operations xswap_debugfs_fops =3D { + .read =3D xswap_max_clusters_read, + .write =3D xswap_max_clusters_write, + .open =3D simple_open, + .llseek =3D default_llseek, +}; + +static void xswap_debugfs_add(struct swap_info_struct *si) +{ + char name[32]; + + if (!xswap_debugfs_root) + return; + + snprintf(name, sizeof(name), "type%d_cluster_limit", si->type); + si->debugfs_entry =3D debugfs_create_file(name, 0644, xswap_debugfs_root, + si, &xswap_debugfs_fops); +} + +static void xswap_debugfs_del(struct swap_info_struct *si) +{ + debugfs_remove(si->debugfs_entry); + si->debugfs_entry =3D NULL; +} + + #endif =20 static void swap_range_alloc(struct swap_info_struct *si, @@ -89,7 +178,6 @@ static void move_cluster(struct swap_info_struct *si, * * Also protects swap_active_head total_swap_pages, and the SWP_WRITEOK fl= ag. */ -static DEFINE_SPINLOCK(swap_lock); static unsigned int nr_swapfiles; atomic_long_t nr_swap_pages; atomic_t nr_real_swapfiles; @@ -3147,6 +3235,7 @@ static void free_swap_cluster_info(struct swap_info_s= truct *si) =20 #ifdef CONFIG_XSWAP if (si->flags & SWP_XSWAP) { + xswap_debugfs_del(si); /* Unmap all mapped clusters and free the VM_SPARSE area */ if (si->nr_clusters_mapped > 0) xswap_unmap_clusters(si, 0, si->nr_clusters_mapped); @@ -4075,6 +4164,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, /* All mapped clusters except cluster 0 are free at the tail */ si->nr_free_tail =3D si->nr_clusters_mapped - 1; mutex_init(&si->xswap_lock); + xswap_debugfs_add(si); return 0; =20 err_unmap: @@ -4498,6 +4588,11 @@ static int __init swapfile_init(void) if (swapfile_maximum_size >=3D (1UL << SWP_MIG_TOTAL_BITS)) swap_migration_ad_supported =3D true; #endif /* CONFIG_MIGRATION */ + +#ifdef CONFIG_XSWAP + xswap_debugfs_root =3D debugfs_create_dir("xswap", NULL); +#endif + return 0; } subsys_initcall(swapfile_init); --=20 2.54.0 From nobody Fri Oct 2 03:41:04 2026 Received: from out-185.mta0.migadu.com (out-185.mta0.migadu.com [91.218.175.185]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0C8E13E6DD5 for ; Wed, 5 Aug 2026 07:55:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.185 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916538; cv=none; b=J37qPbhDSTCKw4bk87ZcBcptwhZcmUGz6BX8MxhLZo53A112yVEm8jvy8k/TfCsIm0YhLNf1OU6FmZmMfrXf+GWeBm2UsIQEikwMZdendkGGLKcoOnVUW6TqA7R7OLkZXQsHkBQeaj4v6vlBtVxiwiI5yN4HOKoUkDhaeUeJPMU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785916538; c=relaxed/simple; bh=UNs7LmYi40SjWRCbHikpB3XszHaogZTJOoVSP3fg/D0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type:Content-type; b=Qx2o6UQ5TrClLZ/O2rV7ZmauOD169n0O3QRf8a+PCVkAZriwZmeN+9/YmL2emkxOQn+Yw5vxY8h+PALSnmcGnIxFXoDBAyWKGkH82Xi8X0R6VK0992AzbiyINNI0RcR45aWwndNEjeyJt1LQX1V4WaHB8l2AF3Ro02IPx16mBY0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=GLgVwNRU; arc=none smtp.client-ip=91.218.175.185 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="GLgVwNRU" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785916532; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=EbfFddbqSLBXoL9Z6NQe5+gvAdUEfreYVifyxE81zg0=; b=GLgVwNRUwiBcSuY7U3D9LcolFHX5qp94+NHSDFNCWLyiNOSj/IDuA2OtulWLjCHMnDp0FP GBcfFr+2W4RrKIkXygjMqYPSkVmDYXphGEoyWSvWVXWVH/H8YIKDRPjIo6WiWBaP2iyTrV FpdB/Tm+iQCJEZO94Gpyn5ghqq5yloA= From: Baoquan He To: linux-mm@kvack.org Cc: chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com, hannes@cmpxchg.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org, Baoquan He Subject: [RFC PATCH v2 10/10] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Date: Wed, 5 Aug 2026 15:53:33 +0800 Message-ID: <20260805075336.3579395-11-baoquan.he@linux.dev> In-Reply-To: <20260805075336.3579395-1-baoquan.he@linux.dev> References: <20260805075336.3579395-1-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-type: text/plain Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT xswap_try_shrink() was called directly from __free_cluster() while holding ci->lock. The shrink path calls xswap_unmap_clusters() which unmaps vmalloc pages backing cluster_info, and on return swap_cache_del_folio() tries swap_cluster_unlock(ci) on the now- unmapped address =E2=80=94 crashing on a not-present page. Replace direct calls with schedule_work() so shrink runs in an independent workqueue context where no cluster locks are held. Use cancel_work_sync() during swapoff to ensure no pending shrink work races with the VM area teardown. Signed-off-by: Baoquan He --- include/linux/swap.h | 1 + mm/swapfile.c | 19 +++++++++++++------ 2 files changed, 14 insertions(+), 6 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 4f4583f9a4e5..a5b323374769 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -255,6 +255,7 @@ struct swap_info_struct { unsigned long nr_clusters_mapped; /* currently mapped cluster count */ unsigned long nr_free_tail; /* contiguous free clusters at tail */ struct dentry *debugfs_entry; /* debugfs: type_max_clusters */ + struct work_struct xswap_shrink_work; /* deferred shrink trigger */ struct mutex xswap_lock; /* serialize map/unmap operations */ #endif struct list_head free_clusters; /* free clusters list */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 265f2bac1a12..ca406556626e 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -127,11 +127,8 @@ static ssize_t xswap_max_clusters_write(struct file *f= ile, spin_unlock(&swap_lock); } =20 - /* - * Lowering the ceiling may make tail clusters eligible for - * shrinking. Trigger an immediate check. - */ - xswap_try_shrink(si); + /* Shrink trigger: lowering the ceiling may free tail clusters. */ + schedule_work(&si->xswap_shrink_work); =20 return count; } @@ -726,7 +723,7 @@ static void __free_cluster(struct swap_info_struct *si,= struct swap_cluster_info ci->order =3D 0; #ifdef CONFIG_XSWAP xswap_update_free_tail(si, ci - si->cluster_info); - xswap_try_shrink(si); + schedule_work(&si->xswap_shrink_work); #endif } =20 @@ -3236,6 +3233,7 @@ static void free_swap_cluster_info(struct swap_info_s= truct *si) #ifdef CONFIG_XSWAP if (si->flags & SWP_XSWAP) { xswap_debugfs_del(si); + cancel_work_sync(&si->xswap_shrink_work); /* Unmap all mapped clusters and free the VM_SPARSE area */ if (si->nr_clusters_mapped > 0) xswap_unmap_clusters(si, 0, si->nr_clusters_mapped); @@ -4057,6 +4055,13 @@ static void xswap_trim_free_tail(struct swap_info_st= ruct *si, unsigned long idx) WRITE_ONCE(si->nr_free_tail, nr_mapped - idx - 1); } =20 +static void xswap_shrink_work_fn(struct work_struct *work) +{ + struct swap_info_struct *si =3D container_of(work, + struct swap_info_struct, xswap_shrink_work); + xswap_try_shrink(si); +} + /* * Try to shrink the cluster_info tail. Uses si->nr_free_tail which * is maintained incrementally during alloc/free =E2=80=94 no scanning nee= ded. @@ -4163,7 +4168,9 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, =20 /* All mapped clusters except cluster 0 are free at the tail */ si->nr_free_tail =3D si->nr_clusters_mapped - 1; + mutex_init(&si->xswap_lock); + INIT_WORK(&si->xswap_shrink_work, xswap_shrink_work_fn); xswap_debugfs_add(si); return 0; =20 --=20 2.54.0