From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f12.google.com (mail-oi2-f12.google.com [74.125.231.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D40CF47ACDF for ; Fri, 18 Sep 2026 18:02:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754569; cv=none; b=fJ939zOzBaLxxW91pfUhQ0mTUC6p/BtedVUbVxWv6vwFuJ4kgrLDREZRg34oYFjRRXoNpkb3n7N8ycrHOXnA51kqTo9SgTMG2Lg+PpeEJB3HGpJFnZxVIq1/Ywf401TiZKOSItlEGx9/lokmSFMpum1rMEt1fWZb3xztoUHGiCk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754569; c=relaxed/simple; bh=zQyC5c4E6kvoG3yUOvItxLrDtTg6uY3jvdEVYUQ7bAM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=mpfQ4xKHDdneBeVJ89AX6IKxTgwHzXzv2QsI3EfAHN6UbC6i3UTbHu1JFyk9kPu/Hxg0eaLQ2ycuRs7QUqdK27qH3RjdmpEdM4xHPAHS2/CnI5EMcrP+YmyZuoJRx+JkyXi7/AqWjf5++gLK4aFnz1D5tfH8RkfWqSplfabENis= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=AJyw9pVF; arc=none smtp.client-ip=74.125.231.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="AJyw9pVF" Received: by mail-oi2-f12.google.com with SMTP id 46e09a7af769-7f4f0d37f9dso419785a34.1 for ; Fri, 18 Sep 2026 11:02:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754565; x=1790359365; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Y8Lc0jqpGY0qUf0efcamB1alDhLPge61ilYZq9DXRxo=; b=AJyw9pVF6gbVCKSdCWZ8qvFEJ1821f7BR4Ckiw4uoOjpKkiQ3dm6Lcqpnt+oA7L3KY GxLEo79FNEjKB6d5lli3637VibfxLvqxE1glg1Q1lljCE6w0VnqhkC9c7onOoPZbGIMc aekgwTbNSH9gzmege0foCEY8sDEy2RKilktJxreN576w+hgTyq7gGIb2ZeMvpHKgWfIq EM5MA7JJzO/oezT42UhnnOmvqBOU2J6p49qGikgP3hbL4cAmWedD1KhvEL5oulcmSdFi Xa3rWAWArk/IyMEEcVqSzOUA6vHqEbA8qFrqYSdRzD+4lop6wJQadfgc8+i4YceYDiPS n6Ug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754565; x=1790359365; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Y8Lc0jqpGY0qUf0efcamB1alDhLPge61ilYZq9DXRxo=; b=UGsZXUvWmqigOXlU3HARS/O/OsS1WKkGRqCzQCXOCpbJ+ZymNJ3lNUacqcZrNEtHxo me7CassMKly6cn9MbF8PSPWwAGUEwoFM8KNA8K4rM42LOjY5Xk/Hdd4OrcIymOPIqsiF bk0NqWRLlgwcXcWyxfglgTDiDQ6KpaONnYw7fjx5yM1zAtvwAywB4uwK85ohGUfey2v3 wn5PVh4/YpgdwdBj8gKjVEDkDXdrAQb/H7Si+Q0kVY35jMXmJotBsgF5Vdqj7Fbh5h5l 4xdpw+cCICuTXL03ZVnHWWbZIP8F4ypB8RutGCuhVz1OD+Es4zuX2bV5aC0gVx6N4rcE pftA== X-Forwarded-Encrypted: i=1; AKwUvBwiAnIW3AoClAWrRpSGTJxEKOXB4OGEOEwrfPEqgxUR2lh6iLPWAXs+Ax1Pkua5IKk+pU8vCSSQ8zfEGF0=@vger.kernel.org X-Gm-Message-State: AFuF++lVq8gplMzotu4d+x4k24sPXeCFA/jzMh7VW6NIDrmc0wZQUX++ 4NDeV8SzCNXJvOsU9NQ7WGnPPtLakBymqZKAkRucn44KFPq1js2kM1Sh X-Gm-Gg: AYBFou0ZQ1H0rgCjfLKwdvDBSFcMVvR/IseC9xypTOD8Em4mntQWaTO3B/y12MxywGh YYWWNzC3HC8oX3fRkRA2RwzPXnQ/N/AzVPVL6pRXo83z0BDSCLHHTUWLX+4Ed4oHktTRv0WKbcg hb+4o/DpWPQwdhVlpdP7TCOjXziPXG735812UC9fMcy1PJixj8JgZ9hWPMQsqxIPJtXL8WjumOZ kWMm9F09xVtWs1hhcZXJYmylSnirz1Lw6ovhEs/GpqNXuKYqJ/mCEWp3RD6hmLehGoI3Ht4kbn4 dbKeCumsbE1gwp+MDLJjaMVzlbzLVpPsFP2zyIIKXeKBHv9+qruxJeYHjLqR6ZBvObYosP/IuI2 ul/cS3/rtYjcYbQTa9SvSwmj622O49cmkwY8SftnQyVtIKrQQ7L8RI3oYEco/WKHOkcsR9JXBvc rAPtfZlNhA7GXam3AkKT3Fmb4O5PKBLlo3n2ghVu1gPvQbGbGBJO12hHjP7bJL89VCvEyF98nU0 z/0CkCHjRCc3QOTH+m9tg== X-Received: by 2002:a05:6830:6819:b0:7fa:5c49:524b with SMTP id 46e09a7af769-80de2bd367fmr3986990a34.18.1789754564497; Fri, 18 Sep 2026 11:02:44 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4a::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-8107e389850sm138395a34.10.2026.09.18.11.02.43 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:43 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 01/11] mm, swap: add virtual swap device infrastructure Date: Fri, 18 Sep 2026 11:02:31 -0700 Message-ID: <20260918180241.3424851-2-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Create a virtual swap device (just under 8 TiB with 4 KiB pages), along with the dynamic cluster infrastructure that the rest of the vswap layer is built on. swap_cluster_info_dynamic keeps per-cluster info in an xarray, so a device can be sized without a static cluster_info[] array. For now, vswap requires a 64-bit architecture. swap_table_lookup() resolves a swap entry's cluster and reads its slot, both under RCU. A vswap cluster is allocated on demand and freed by kfree_rcu(), so an unpinned entry cannot have its cluster resolved beforehand, and the lookup can return NULL. Callers holding a pinned entry keep resolving it themselves: the pin holds ci->count above zero, so the cluster cannot be destroyed. mm/workingset.c includes swap_table.h with CONFIG_SWAP off, so the helpers swap_table_lookup() calls gain !CONFIG_SWAP stubs. The dynamic-cluster allocator is wired in, but nothing reaches it yet. vswap_si is kept off the swap device lists, and no allocation path can select it. Backends (zswap, zero, physical disk) and the vswap-aware swap-out / swap-in / writeback paths arrive in subsequent patches. Routing is controlled by the "vswap=3D" kernel parameter, defaulting to CONFIG_VSWAP_DEFAULT_ON. When off, no device exists, every vswap path is skipped, and swap behavior is unchanged. When on, vswap_init() creates the device and enables a static key only after it is fully published, so callers never observe a half-built device. The device lives for the lifetime of the kernel and cannot be swapon'd or swapoff'd. The device is capped at REFCOUNT_MAX pages, just under 8 TiB with 4 KiB pages. Every swapped page takes a reference on its cgroup's memcg->private_id_ref, a 32-bit refcount_t, so a cgroup that swaps past that saturates it and leaks the memcg until reboot. The limit is not specific to vswap, a physical swapfile that large hits it too, but vswap is the first device that can reach it without provisioning the storage. Suggested-by: Kairui Song Signed-off-by: Nhat Pham --- .../admin-guide/kernel-parameters.txt | 7 + MAINTAINERS | 1 + include/linux/swap.h | 9 + mm/Kconfig | 20 ++ mm/page_io.c | 14 + mm/swap.h | 77 ++++- mm/swap_state.c | 40 +-- mm/swap_table.h | 27 ++ mm/swapfile.c | 275 ++++++++++++++++-- mm/vswap.h | 34 +++ mm/zswap.c | 6 + 11 files changed, 468 insertions(+), 42 deletions(-) create mode 100644 mm/vswap.h diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentatio= n/admin-guide/kernel-parameters.txt index 68647ff4bdd2..646e689d9e80 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -8431,6 +8431,13 @@ Kernel parameters force - force vulnerability detection even on unaffected processors =20 + vswap=3D [MM,EARLY] + Route swapouts through the virtual swap layer, which + allows zswap and zero-filled pages to be used without + a physical swap device. 64-bit only. + Format: { on | off } + Default: on if CONFIG_VSWAP_DEFAULT_ON=3Dy, else off. + vsyscall=3D [X86-64,EARLY] Controls the behavior of vsyscalls (i.e. calls to fixed addresses of 0xffffffffff600x00 from legacy diff --git a/MAINTAINERS b/MAINTAINERS index e4412c3d8d45..e49a3d4e332b 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -17411,6 +17411,7 @@ F: mm/swap.h F: mm/swap_table.h F: mm/swap_state.c F: mm/swapfile.c +F: mm/vswap.h =20 MEMORY MANAGEMENT - THP (TRANSPARENT HUGE PAGE) M: Andrew Morton diff --git a/include/linux/swap.h b/include/linux/swap.h index 61005501888c..0b341301d8fa 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -201,6 +201,7 @@ enum { SWP_STABLE_WRITES =3D (1 << 11), /* no overwrite PG_writeback pages */ SWP_SYNCHRONOUS_IO =3D (1 << 12), /* synchronous IO is efficient */ SWP_HIBERNATION =3D (1 << 13), /* pinned for hibernation */ + SWP_VSWAP =3D (1 << 14), /* virtual swap device */ /* add others here before... */ }; =20 @@ -270,8 +271,14 @@ struct swap_info_struct { struct list_head discard_clusters; /* discard clusters list */ struct plist_node avail_list; /* entry in swap_avail_head */ const struct swap_ops *ops; + struct xarray cluster_info_pool; /* Xarray for vswap dynamic cluster info= */ }; =20 +static inline bool swap_is_vswap(struct swap_info_struct *si) +{ + return si->flags & SWP_VSWAP; +} + /** * folio_swap_entry - Return the swap entry at a page index within a folio. * @folio: The folio. @@ -427,6 +434,8 @@ void swap_free_hibernation_slot(swp_entry_t entry); =20 static inline void put_swap_device(struct swap_info_struct *si) { + if (swap_is_vswap(si)) + return; percpu_ref_put(&si->users); } =20 diff --git a/mm/Kconfig b/mm/Kconfig index 30170a936f1f..beef56e870ce 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -19,6 +19,26 @@ menuconfig SWAP used to provide more virtual memory than the actual RAM present in your computer. If unsure say Y. =20 +config VSWAP_DEFAULT_ON + bool "Route swapouts through virtual swap by default" + depends on SWAP && 64BIT + default n + help + Virtual swap allows zswap and zero-filled pages to be used + without swapping on a physical device first, and lets a page + move between zswap and a swapfile without invalidating the page + table entries that refer to it. + + Swap entries are handed out by a virtual swap device instead of + naming a slot on a real one, so the backing can be chosen and + changed after the entry exists. + + Say Y to make "vswap=3Don" the default, routing swapouts through + the virtual swap layer from boot. + + Say N (default) to leave vswap off unless "vswap=3Don" is passed + on the kernel command line. + config ZSWAP bool "Compressed cache for swap pages" depends on SWAP diff --git a/mm/page_io.c b/mm/page_io.c index 1da4ff484f09..5cd77bd1b12c 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -28,6 +28,7 @@ #include #include "swap.h" #include "swap_table.h" +#include "vswap.h" =20 int generic_swapfile_activate(struct swap_info_struct *sis, struct file *swap_file, @@ -248,6 +249,14 @@ int swap_writeout(struct swap_io_ctx *ctx, struct foli= o *folio) } rcu_read_unlock(); =20 + /* + * A vswap folio has no physical slot to write to, so keep it dirty. + */ + if (is_vswap_entry(folio->swap)) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writeout(ctx, folio); return 0; out_unlock: @@ -482,6 +491,11 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct f= olio *folio) if (zswap_load(folio) !=3D -ENOENT) goto finish; =20 + if (unlikely(swap_is_vswap(sis))) { + folio_unlock(folio); + goto finish; + } + /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); swap_add_folio(ctx, folio, READ); diff --git a/mm/swap.h b/mm/swap.h index b3b54c28929a..528a27afb335 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -67,6 +67,12 @@ struct swap_cluster_info { struct list_head list; }; =20 +struct swap_cluster_info_dynamic { + struct swap_cluster_info ci; + unsigned int index; /* for cluster_index() */ + struct rcu_head rcu; +}; + /* All on-list cluster must have a non-zero flag. */ enum swap_cluster_flags { CLUSTER_FLAG_NONE =3D 0, /* For temporary off-list cluster */ @@ -77,6 +83,7 @@ enum swap_cluster_flags { CLUSTER_FLAG_USABLE =3D CLUSTER_FLAG_FRAG, CLUSTER_FLAG_FULL, CLUSTER_FLAG_DISCARD, + CLUSTER_FLAG_DEAD, /* Vswap dynamic cluster pending kfree_rcu */ CLUSTER_FLAG_MAX, }; =20 @@ -119,12 +126,33 @@ static inline struct swap_info_struct *__swap_entry_t= o_info(swp_entry_t entry) return __swap_type_to_info(swp_type(entry)); } =20 +/** + * __swap_offset_to_cluster - look up the cluster holding a swap offset + * @si: the swap device + * @offset: the swap entry offset + * + * Context: A vswap cluster is freed by kfree_rcu(). Callers must hold the + * RCU read lock, or know the cluster is pinned by an in-use entry. + * + * Return: the cluster, or NULL if @si is a vswap device with no cluster + * allocated at @offset. + */ static inline struct swap_cluster_info *__swap_offset_to_cluster( struct swap_info_struct *si, pgoff_t offset) { + unsigned int cluster_idx =3D offset / SWAPFILE_CLUSTER; + VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ VM_WARN_ON_ONCE(offset >=3D roundup(si->max, SWAPFILE_CLUSTER)); - return &si->cluster_info[offset / SWAPFILE_CLUSTER]; + + if (swap_is_vswap(si)) { + struct swap_cluster_info_dynamic *ci_dyn; + + ci_dyn =3D xa_load(&si->cluster_info_pool, cluster_idx); + return ci_dyn ? &ci_dyn->ci : NULL; + } + + return &si->cluster_info[cluster_idx]; } =20 static inline struct swap_cluster_info *__swap_entry_to_cluster(swp_entry_= t entry) @@ -133,10 +161,36 @@ static inline struct swap_cluster_info *__swap_entry_= to_cluster(swp_entry_t entr swp_offset(entry)); } =20 +static inline struct swap_cluster_info *__vswap_cluster_lock( + struct swap_info_struct *si, unsigned long offset, bool irq) +{ + struct swap_cluster_info *ci; + + rcu_read_lock(); + ci =3D __swap_offset_to_cluster(si, offset); + if (ci) { + if (irq) + spin_lock_irq(&ci->lock); + else + spin_lock(&ci->lock); + + /* The cluster can be torn down while we wait for the lock. */ + if (ci->flags =3D=3D CLUSTER_FLAG_DEAD) { + if (irq) + spin_unlock_irq(&ci->lock); + else + spin_unlock(&ci->lock); + ci =3D NULL; + } + } + rcu_read_unlock(); + return ci; +} + static __always_inline struct swap_cluster_info *__swap_cluster_lock( struct swap_info_struct *si, unsigned long offset, bool irq) { - struct swap_cluster_info *ci =3D __swap_offset_to_cluster(si, offset); + struct swap_cluster_info *ci; =20 /* * Nothing modifies swap cache in an IRQ context. All access to @@ -149,6 +203,11 @@ static __always_inline struct swap_cluster_info *__swa= p_cluster_lock( */ VM_WARN_ON_ONCE(!in_task()); VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ + + if (swap_is_vswap(si)) + return __vswap_cluster_lock(si, offset, irq); + + ci =3D __swap_offset_to_cluster(si, offset); if (irq) spin_lock_irq(&ci->lock); else @@ -159,10 +218,12 @@ static __always_inline struct swap_cluster_info *__sw= ap_cluster_lock( /** * swap_cluster_lock - Lock and return the swap cluster of given offset. * @si: swap device the cluster belongs to. - * @offset: the swap entry offset, pointing to a valid slot. + * @offset: the swap entry offset. * * Context: The caller must ensure the offset is in the valid range and * protect the swap device with reference count or locks. + * Return: the locked cluster, or NULL if it is gone. Only a vswap device + * can return NULL, as its clusters are allocated and freed on demand. */ static inline struct swap_cluster_info *swap_cluster_lock( struct swap_info_struct *si, unsigned long offset) @@ -363,6 +424,16 @@ static inline struct swap_info_struct *__swap_entry_to= _info(swp_entry_t entry) return NULL; } =20 +static inline struct swap_cluster_info *__swap_entry_to_cluster(swp_entry_= t entry) +{ + return NULL; +} + +static inline unsigned int swp_cluster_offset(swp_entry_t entry) +{ + return 0; +} + static inline int folio_alloc_swap(struct folio *folio) { return -EINVAL; diff --git a/mm/swap_state.c b/mm/swap_state.c index cef44aadee61..424a0040b4b8 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -96,8 +96,7 @@ struct folio *swap_cache_get_folio(swp_entry_t entry) struct folio *folio; =20 for (;;) { - swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), - swp_cluster_offset(entry)); + swp_tb =3D swap_table_lookup(entry); if (!swp_tb_is_folio(swp_tb)) return NULL; folio =3D swp_tb_to_folio(swp_tb); @@ -119,8 +118,7 @@ bool swap_cache_has_folio(swp_entry_t entry) { unsigned long swp_tb; =20 - swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), - swp_cluster_offset(entry)); + swp_tb =3D swap_table_lookup(entry); return swp_tb_is_folio(swp_tb); } =20 @@ -136,8 +134,7 @@ void *swap_cache_get_shadow(swp_entry_t entry) { unsigned long swp_tb; =20 - swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), - swp_cluster_offset(entry)); + swp_tb =3D swap_table_lookup(entry); if (swp_tb_is_shadow(swp_tb)) return swp_tb_to_shadow(swp_tb); return NULL; @@ -414,14 +411,16 @@ void __swap_cache_replace_folio(struct swap_cluster_i= nfo *ci, * -ENOENT / -EEXIST: Target swap entry is unavailable or cached, the call= er * should abort or try to use the cached folio instead */ -static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci, - swp_entry_t targ_entry, gfp_t gfp, +static struct folio *__swap_cache_alloc(swp_entry_t targ_entry, gfp_t gfp, unsigned int order, struct vm_fault *vmf, struct mempolicy *mpol, pgoff_t ilx) { int err; swp_entry_t entry; struct folio *folio; + struct swap_cluster_info *ci; + struct swap_info_struct *si =3D __swap_entry_to_info(targ_entry); + unsigned long offset =3D swp_offset(targ_entry); void *shadow =3D NULL; unsigned short memcg_id; unsigned long address, nr_pages =3D 1UL << order; @@ -431,9 +430,12 @@ static struct folio *__swap_cache_alloc(struct swap_cl= uster_info *ci, entry.val =3D round_down(targ_entry.val, nr_pages); =20 /* Check if the slot and range are available, skip allocation if not */ - spin_lock(&ci->lock); - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); - spin_unlock(&ci->lock); + err =3D -ENOENT; + ci =3D swap_cluster_lock(si, offset); + if (ci) { + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); + swap_cluster_unlock(ci); + } if (unlikely(err)) return ERR_PTR(err); =20 @@ -454,10 +456,13 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, return ERR_PTR(-ENOMEM); =20 /* Double check the range is still not in conflict */ - spin_lock(&ci->lock); - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg_= id); + err =3D -ENOENT; + ci =3D swap_cluster_lock(si, offset); + if (ci) + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg= _id); if (unlikely(err)) { - spin_unlock(&ci->lock); + if (ci) + swap_cluster_unlock(ci); folio_put(folio); return ERR_PTR(err); } @@ -465,10 +470,11 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, __folio_set_locked(folio); __folio_set_swapbacked(folio); __swap_cache_do_add_folio(ci, folio, entry); - spin_unlock(&ci->lock); + swap_cluster_unlock(ci); =20 if (mem_cgroup_swapin_charge_folio(folio, memcg_id, vmf ? vmf->vma->vm_mm : NULL, gfp)) { + /* The folio pins the cluster */ spin_lock(&ci->lock); __swap_cache_do_del_folio(ci, folio, entry, shadow); spin_unlock(&ci->lock); @@ -525,9 +531,7 @@ struct folio *swap_cache_alloc_folio(swp_entry_t targ_e= ntry, gfp_t gfp, { int order, err; struct folio *ret; - struct swap_cluster_info *ci; =20 - ci =3D __swap_entry_to_cluster(targ_entry); order =3D highest_order(orders); =20 /* orders must be non-zero, and must not exceed cluster size. */ @@ -535,7 +539,7 @@ struct folio *swap_cache_alloc_folio(swp_entry_t targ_e= ntry, gfp_t gfp, return ERR_PTR(-EINVAL); =20 do { - ret =3D __swap_cache_alloc(ci, targ_entry, gfp, order, + ret =3D __swap_cache_alloc(targ_entry, gfp, order, vmf, mpol, ilx); if (!IS_ERR(ret)) break; diff --git a/mm/swap_table.h b/mm/swap_table.h index e6613e62f8d0..3d64d1629882 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -6,6 +6,8 @@ #include #include "swap.h" =20 +extern struct swap_info_struct *vswap_si; + /* A typical flat array in each cluster as swap table */ struct swap_table { atomic_long_t entries[SWAPFILE_CLUSTER]; @@ -264,6 +266,31 @@ static inline unsigned long swap_table_get(struct swap= _cluster_info *ci, return swp_tb; } =20 +/* + * Resolve @entry's cluster and read its slot, both under RCU. A vswap + * cluster is allocated on demand and freed by kfree_rcu(), so a caller + * starting from an entry cannot resolve it beforehand. + */ +static inline unsigned long swap_table_lookup(swp_entry_t entry) +{ + struct swap_cluster_info *ci; + atomic_long_t *table; + unsigned long swp_tb; + + rcu_read_lock(); + ci =3D __swap_entry_to_cluster(entry); + if (!ci) { + rcu_read_unlock(); + return null_to_swp_tb(); + } + table =3D rcu_dereference(ci->table); + swp_tb =3D table ? atomic_long_read(&table[swp_cluster_offset(entry)]) + : null_to_swp_tb(); + rcu_read_unlock(); + + return swp_tb; +} + static inline void __swap_table_set_zero(struct swap_cluster_info *ci, unsigned int ci_off) { diff --git a/mm/swapfile.c b/mm/swapfile.c index c1c5fbb3c909..edd7bddf7ff9 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -36,6 +36,7 @@ #include #include #include +#include #include #include #include @@ -46,6 +47,7 @@ #include #include #include "swap_table.h" +#include "vswap.h" #include "internal.h" #include "swap.h" =20 @@ -399,6 +401,8 @@ static inline bool cluster_is_usable(struct swap_cluste= r_info *ci, int order) static inline unsigned int cluster_index(struct swap_info_struct *si, struct swap_cluster_info *ci) { + if (swap_is_vswap(si)) + return container_of(ci, struct swap_cluster_info_dynamic, ci)->index; return ci - si->cluster_info; } =20 @@ -594,10 +598,15 @@ static void move_cluster(struct swap_info_struct *si, lockdep_assert_held(&ci->lock); =20 spin_lock(&si->lock); - if (ci->flags =3D=3D CLUSTER_FLAG_NONE) + if (!list) { + /* Going away. An isolated cluster is already off its list. */ + if (ci->flags !=3D CLUSTER_FLAG_NONE) + list_del(&ci->list); + } else if (ci->flags =3D=3D CLUSTER_FLAG_NONE) { list_add_tail(&ci->list, list); - else + } else { list_move_tail(&ci->list, list); + } spin_unlock(&si->lock); ci->flags =3D new_flags; } @@ -615,6 +624,18 @@ static void __free_cluster(struct swap_info_struct *si= , struct swap_cluster_info { swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, false); swap_cluster_free_table(ci); + + if (swap_is_vswap(si)) { + struct swap_cluster_info_dynamic *ci_dyn; + + /* vswap clusters are destroyed, not returned to free_clusters. */ + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + xa_erase(&si->cluster_info_pool, ci_dyn->index); + move_cluster(si, ci, NULL, CLUSTER_FLAG_DEAD); + kfree_rcu(ci_dyn, rcu); + return; + } + move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE); ci->order =3D 0; } @@ -851,6 +872,8 @@ static bool cluster_reclaim_range(struct swap_info_stru= ct *si, unsigned long offset =3D start, end =3D start + nr_pages; unsigned long swp_tb; =20 + VM_WARN_ON_ONCE(swap_is_vswap(si)); + spin_unlock(&ci->lock); do { swp_tb =3D swap_table_get(ci, offset % SWAPFILE_CLUSTER); @@ -1042,6 +1065,44 @@ static unsigned int alloc_swap_scan_list(struct swap= _info_struct *si, return found; } =20 +static unsigned int vswap_alloc_cluster(struct swap_info_struct *si, + struct folio *folio) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; + unsigned long offset; + + VM_WARN_ON(!swap_is_vswap(si)); + + ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_ATOMIC | __GFP_NOWARN); + if (!ci_dyn) + return SWAP_ENTRY_INVALID; + + spin_lock_init(&ci_dyn->ci.lock); + INIT_LIST_HEAD(&ci_dyn->ci.list); + + if (swap_cluster_alloc_table(&ci_dyn->ci, GFP_ATOMIC | __GFP_NOWARN)) { + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + + /* Lock before publishing: xa_alloc makes the cluster findable by offset.= */ + ci =3D &ci_dyn->ci; + spin_lock(&ci->lock); + + if (xa_alloc(&si->cluster_info_pool, &ci_dyn->index, ci_dyn, + XA_LIMIT(1, DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER) - 1), + GFP_ATOMIC | __GFP_NOWARN)) { + spin_unlock(&ci->lock); + swap_cluster_free_table(&ci_dyn->ci); + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + + offset =3D cluster_offset(si, ci); + return alloc_swap_scan_cluster(si, ci, folio, offset); +} + static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool f= orce) { long to_scan =3D 1; @@ -1064,7 +1125,9 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) spin_unlock(&ci->lock); nr_reclaim =3D __try_to_reclaim_swap(si, offset, TTRS_ANYWAY); - spin_lock(&ci->lock); + ci =3D swap_cluster_lock(si, offset); + if (!ci) + goto next; if (nr_reclaim) { offset +=3D abs(nr_reclaim); continue; @@ -1078,6 +1141,7 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) relocate_cluster(si, ci); =20 swap_cluster_unlock(ci); +next: if (to_scan <=3D 0) break; =20 @@ -1154,6 +1218,12 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, goto done; } =20 + if (swap_is_vswap(si)) { + found =3D vswap_alloc_cluster(si, folio); + if (found) + goto done; + } + if (!(si->flags & SWP_PAGE_DISCARD)) { found =3D alloc_swap_scan_list(si, &si->free_clusters, folio, false); if (found) @@ -1290,8 +1360,10 @@ static bool swap_usage_add(struct swap_info_struct *= si, unsigned int nr_entries) /* * If device is full, and SWAP_USAGE_OFFLIST_BIT is not set, * remove it from the plist. + * + * Vswap is never on the avail list, so skip it. */ - if (unlikely(val =3D=3D si->pages)) { + if (unlikely(val =3D=3D si->pages) && !swap_is_vswap(si)) { del_from_avail_list(si, false); return true; } @@ -1306,8 +1378,10 @@ static void swap_usage_sub(struct swap_info_struct *= si, unsigned int nr_entries) /* * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, * add it to the plist. + * + * Vswap is never on the avail list, so skip it. */ - if (unlikely(val & SWAP_USAGE_OFFLIST_BIT)) + if (unlikely(val & SWAP_USAGE_OFFLIST_BIT) && !swap_is_vswap(si)) add_to_avail_list(si, false); } =20 @@ -1352,6 +1426,10 @@ static void swap_range_free(struct swap_info_struct = *si, unsigned long offset, =20 static bool get_swap_device_info(struct swap_info_struct *si) { + /* The vswap device is always alive, so it needs no refcount. */ + if (swap_is_vswap(si)) + return true; + if (!percpu_ref_tryget_live(&si->users)) return false; /* @@ -1387,11 +1465,11 @@ static bool swap_alloc_fast(struct folio *folio) return false; =20 ci =3D swap_cluster_lock(si, offset); - if (cluster_is_usable(ci, order)) { + if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); alloc_swap_scan_cluster(si, ci, folio, offset); - } else { + } else if (ci) { swap_cluster_unlock(ci); } =20 @@ -1513,6 +1591,7 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t g= fp) if (IS_ERR_OR_NULL(si)) return 0; =20 + /* The source PTE pins the entry, so its cluster is alive. */ ci =3D __swap_offset_to_cluster(si, offset); ret =3D swap_extend_table_alloc(si, ci, swp_cluster_offset(entry), gfp); =20 @@ -1923,7 +2002,7 @@ struct swap_info_struct *get_swap_device(swp_entry_t = entry) return NULL; put_out: pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_offset, entry.val); - percpu_ref_put(&si->users); + put_swap_device(si); return ERR_PTR(-EIO); } =20 @@ -2002,7 +2081,7 @@ bool swap_entry_swapped(struct swap_info_struct *si, = swp_entry_t entry) unsigned long swp_tb; =20 ci =3D swap_cluster_lock(si, offset); - swp_tb =3D swap_table_get(ci, offset % SWAPFILE_CLUSTER); + swp_tb =3D __swap_table_get(ci, swp_cluster_offset(entry)); swap_cluster_unlock(ci); =20 return swp_tb_get_count(swp_tb) > 0; @@ -2055,6 +2134,7 @@ static bool folio_maybe_swapped(struct folio *folio) VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio); =20 + /* Folio is locked and in swap cache, so ci->count > 0: cluster is alive.= */ ci =3D __swap_entry_to_cluster(entry); ci_off =3D swp_cluster_offset(entry); ci_end =3D ci_off + folio_nr_pages(folio); @@ -2249,6 +2329,9 @@ static int __find_hibernation_swap_type(dev_t device,= sector_t offset) =20 if (!(sis->flags & SWP_WRITEOK)) continue; + /* vswap has no bdev, so it is never a hibernation target. */ + if (swap_is_vswap(sis)) + continue; =20 if (device =3D=3D sis->bdev->bd_dev) { struct swap_extent *se =3D first_se(sis); @@ -2375,6 +2458,9 @@ int find_first_swap(dev_t *device) =20 if (!(sis->flags & SWP_WRITEOK)) continue; + /* vswap has no bdev, so it is never a hibernation target. */ + if (swap_is_vswap(sis)) + continue; *device =3D sis->bdev->bd_dev; spin_unlock(&swap_lock); return type; @@ -2591,8 +2677,7 @@ static int unuse_pte_range(struct vm_area_struct *vma= , pmd_t *pmd, &vmf); } if (!folio) { - swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), - swp_cluster_offset(entry)); + swp_tb =3D swap_table_lookup(entry); if (swp_tb_get_count(swp_tb) <=3D 0) continue; return -ENOMEM; @@ -3022,15 +3107,24 @@ static int setup_swap_extents(struct swap_info_stru= ct *sis, =20 static void _enable_swap_info(struct swap_info_struct *si) { - atomic_long_add(si->pages, &nr_swap_pages); - total_swap_pages +=3D si->pages; + if (!swap_is_vswap(si)) { + atomic_long_add(si->pages, &nr_swap_pages); + total_swap_pages +=3D si->pages; + } =20 assert_spin_locked(&swap_lock); =20 - plist_add(&si->list, &swap_active_head); + /* + * Vswap has no backing file and no swapoff support, so keep it + * off swap_active_head (used by swapoff filename lookup and + * swap_sync_discard) and swap_avail_head (physical allocator). + */ + if (!swap_is_vswap(si)) { + plist_add(&si->list, &swap_active_head); =20 - /* Add back to available list */ - add_to_avail_list(si, true); + /* Add back to available list */ + add_to_avail_list(si, true); + } } =20 /* @@ -3074,12 +3168,31 @@ static void wait_for_allocation(struct swap_info_st= ruct *si) } } =20 -static void free_swap_cluster_info(struct swap_cluster_info *cluster_info, +static void free_swap_cluster_info(struct swap_info_struct *si, + struct swap_cluster_info *cluster_info, unsigned long maxpages) { + struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *ci; + unsigned long idx; int i, nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); =20 + if (swap_is_vswap(si)) { + xa_for_each(&si->cluster_info_pool, idx, ci_dyn) { + ci =3D &ci_dyn->ci; + spin_lock(&ci->lock); + if (cluster_table_is_alloced(ci)) { + swap_cluster_assert_empty(ci, 0, + SWAPFILE_CLUSTER, true); + swap_cluster_free_table(ci); + } + spin_unlock(&ci->lock); + kfree(ci_dyn); + } + xa_destroy(&si->cluster_info_pool); + return; + } + if (!cluster_info) return; for (i =3D 0; i < nr_clusters; i++) { @@ -3226,7 +3339,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) mutex_unlock(&swapon_mutex); kfree(p->global_cluster); p->global_cluster =3D NULL; - free_swap_cluster_info(cluster_info, maxpages); + free_swap_cluster_info(p, cluster_info, maxpages); =20 inode =3D mapping->host; =20 @@ -3573,10 +3686,39 @@ static int setup_swap_clusters_info(struct swap_inf= o_struct *si, unsigned long maxpages) { unsigned long nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); - struct swap_cluster_info *cluster_info; + struct swap_cluster_info *cluster_info =3D NULL; + struct swap_cluster_info_dynamic *ci_dyn =3D NULL; int err =3D -ENOMEM; unsigned long i; =20 + /* A vswap device uses an xarray pool instead of a static array. */ + if (swap_is_vswap(si)) { + nr_clusters =3D 0; + xa_init_flags(&si->cluster_info_pool, XA_FLAGS_ALLOC); + + /* + * Pre-allocate cluster 0 and mark slot 0 (header page) + * as bad so the allocator never hands out page offset 0. + */ + ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_KERNEL); + if (!ci_dyn) + goto err; + spin_lock_init(&ci_dyn->ci.lock); + INIT_LIST_HEAD(&ci_dyn->ci.list); + + err =3D xa_insert(&si->cluster_info_pool, 0, ci_dyn, GFP_KERNEL); + if (err) { + kfree(ci_dyn); + goto err; + } + + err =3D swap_cluster_setup_bad_slot(si, &ci_dyn->ci, 0, false); + if (err) + goto err; + + goto setup_cluster_info; + } + cluster_info =3D kvzalloc_objs(*cluster_info, nr_clusters); if (!cluster_info) goto err; @@ -3620,6 +3762,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, goto err; } =20 +setup_cluster_info: INIT_LIST_HEAD(&si->free_clusters); INIT_LIST_HEAD(&si->full_clusters); INIT_LIST_HEAD(&si->discard_clusters); @@ -3641,10 +3784,16 @@ static int setup_swap_clusters_info(struct swap_inf= o_struct *si, } } =20 + /* Slot 0 is bad, so cluster 0 never empties. The rest of it is usable. */ + if (swap_is_vswap(si)) { + ci_dyn->ci.flags =3D CLUSTER_FLAG_NONFULL; + list_add_tail(&ci_dyn->ci.list, &si->nonfull_clusters[0]); + } + si->cluster_info =3D cluster_info; return 0; err: - free_swap_cluster_info(cluster_info, maxpages); + free_swap_cluster_info(si, cluster_info, maxpages); return err; } =20 @@ -3866,7 +4015,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) si->global_cluster =3D NULL; inode =3D NULL; destroy_swap_extents(si, swap_file); - free_swap_cluster_info(si->cluster_info, si->max); + free_swap_cluster_info(si, si->cluster_info, si->max); si->cluster_info =3D NULL; /* * Clear the SWP_USED flag after all resources are freed so @@ -3997,3 +4146,87 @@ static int __init swapfile_init(void) return 0; } subsys_initcall(swapfile_init); + +struct swap_info_struct *vswap_si; +DEFINE_STATIC_KEY_FALSE(vswap_key); + +static bool vswap_enabled_early __initdata =3D IS_ENABLED(CONFIG_VSWAP_DEF= AULT_ON); + +static int __init early_vswap(char *buf) +{ + return kstrtobool(buf, &vswap_enabled_early); +} +early_param("vswap", early_vswap); + +/* vswap does no IO on its own. */ +static const struct swap_ops vswap_ops =3D { }; + +static int __init vswap_init(void) +{ + struct swap_info_struct *si; + unsigned long maxpages; + int err; + + if (!IS_ENABLED(CONFIG_64BIT)) { + if (vswap_enabled_early) + pr_warn("vswap: requires 64-bit architecture; vswap disabled, swapout f= alls back to direct physical swap\n"); + return 0; + } + + if (!vswap_enabled_early) + return 0; + + si =3D alloc_swap_info(); + if (IS_ERR(si)) { + pr_warn("vswap: alloc_swap_info failed (%ld); vswap disabled, swapout fa= lls back to direct physical swap\n", + PTR_ERR(si)); + return 0; + } + + /* + * Each swapped page holds a reference on its cgroup's + * memcg->private_id_ref, a 32-bit refcount_t, so a cgroup that swaps + * REFCOUNT_MAX pages saturates it and leaks the memcg until reboot. + * Cap the device there; that also keeps page counts inside si->max's + * unsigned int. Lifting it needs the memcg refcount widened, which a + * physical swapfile of this size needs too. + */ + maxpages =3D min(swapfile_maximum_size, + ALIGN_DOWN((unsigned long)REFCOUNT_MAX, SWAPFILE_CLUSTER)); + /* + * SWP_WRITEOK enables slot allocation. SWP_SOLIDSTATE selects + * per-CPU cluster allocation; vswap has no si->global_cluster. + */ + si->flags |=3D SWP_VSWAP | SWP_SOLIDSTATE | SWP_WRITEOK; + si->ops =3D &vswap_ops; + si->bdev =3D NULL; + si->max =3D maxpages; + si->pages =3D maxpages - 1; + + INIT_WORK(&si->discard_work, swap_discard_work); + INIT_WORK(&si->reclaim_work, swap_reclaim_work); + + err =3D setup_swap_clusters_info(si, NULL, maxpages); + if (err) + goto fail; + + mutex_lock(&swapon_mutex); + enable_swap_info(si); + mutex_unlock(&swapon_mutex); + + vswap_si =3D si; + pr_info("vswap: created virtual swap device (%lu pages)\n", maxpages); + + /* Last: everything above must be visible before routing starts. */ + static_branch_enable(&vswap_key); + return 0; + +fail: + pr_warn("vswap: setup_swap_clusters_info failed (%d); vswap disabled, swa= pout falls back to direct physical swap\n", + err); + spin_lock(&swap_lock); + si->flags =3D 0; + spin_unlock(&swap_lock); + return 0; +} +late_initcall(vswap_init); diff --git a/mm/vswap.h b/mm/vswap.h new file mode 100644 index 000000000000..16395f357955 --- /dev/null +++ b/mm/vswap.h @@ -0,0 +1,34 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Virtual swap space + * + * Copyright (C) 2026 Nhat Pham + */ +#ifndef _MM_VSWAP_H +#define _MM_VSWAP_H + +#include +#include +#include "swap.h" + +#ifdef CONFIG_SWAP + +DECLARE_STATIC_KEY_FALSE(vswap_key); + +/* + * Only true once vswap_init() has published vswap_si, so callers never + * see the device half built. + */ +static inline bool vswap_is_enabled(void) +{ + return static_branch_unlikely(&vswap_key); +} + +static inline bool is_vswap_entry(swp_entry_t entry) +{ + return swap_is_vswap(__swap_entry_to_info(entry)); +} + +#endif /* CONFIG_SWAP */ + +#endif /* _MM_VSWAP_H */ diff --git a/mm/zswap.c b/mm/zswap.c index 507f2d19fd2a..73728b510ff6 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1018,6 +1018,12 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, if (IS_ERR_OR_NULL(si)) return -ENOENT; =20 + /* Vswap entries have no physical backing to write to. */ + if (swap_is_vswap(si)) { + put_swap_device(si); + return -EINVAL; + } + mpol =3D get_task_policy(current); folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, NO_INTERLEAVE_INDEX); --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-ot1-f49.google.com (mail-ot1-f49.google.com [209.85.210.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1595446072 for ; Fri, 18 Sep 2026 18:02:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754571; cv=none; b=iJNhE0wf2/OwGcIX5pRiZqaVOAMzPxOq2R5lG7YQUgd2+Vf4yDZRChWcbXZRCcVDz1NIDq6p2wJKYZocoF89aVvRA0OVGub7fii8UDtKdHG63Zal5EpEIdjom+x2IZppJI3Vlxh6KC3q9+UD5C6yrWnKC2LuhMKSGPhSFLOxTLw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754571; c=relaxed/simple; bh=hTXAM+oFcps6nmIiD1HqMSxqPU9Lt5lnKtD8HOxM4Cw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bPeZKZNM6CIqVrTHfk8OblU62aJ9bUNywGDwate5q9dTwBrQmGiKmxc7gDQgKS3GI8SKlcSRHHW+lbHKH1Ur0Yi0OwGII6p31ma2GFRA1A4lqQm60d29+pz9Lkd3UPfp8q/WNjK7fPZjTHmlI6c3ZdyYYM7uA06mPdfxSx38Dzw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=P2Du8ByJ; arc=none smtp.client-ip=209.85.210.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="P2Du8ByJ" Received: by mail-ot1-f49.google.com with SMTP id 46e09a7af769-80638c24bedso851260a34.0 for ; Fri, 18 Sep 2026 11:02:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754566; x=1790359366; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=szbh+CjqxPqdhW//FC3nMgY8saM/5ZX/IiWyhRbMQJs=; b=P2Du8ByJDDIR4WlTkQttlnKGu2XZu6qcpFoXf366uSg/J7hhiHG4NFhLLZPQLmRQEe s5FOT1VHEE1EHGkrxJoavAYV2B+scCYp0iXh4AEjlD4pK8lWXiwnRjIPa2Y+U57GI24t LwY5ORg7o+taVRU0Kd143+u6YLgkpIt4+fp4iFJsihytsC0GM/9f2Xp97cu5VXZicgRH rSe+5xmsSDSOXUXy4OMNlPWpub1L4AjYmaM495b0eGrHg4I7R2WE0HbBY5TA6jiZQn2e tqSe0LdqqwIcGe5G2OuUX/xU3niRV38I12GnLtOmRLJ88fgTnscHPlh5yBXp4zgsU0eB 3Ixw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754566; x=1790359366; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=szbh+CjqxPqdhW//FC3nMgY8saM/5ZX/IiWyhRbMQJs=; b=GHaXtU39veIUVasABhQTPIfNHXX8Db7a2WIt9O6uY4MNN6+6BYN7BI/i0sUCM7lAPd 7YKfOgHflimAUDkaqShzVumteMBtdzzgXqk0zm11VeGQFNWrPw/IrAyByh+6hfEAq2O1 ddGIPKKKeS44XR483ncPEt/gfNNH6WQsppqaV6vYyAn92Vce42FCcZ/rhITbsDNrAW8T 9UxwO4FqWdcYM6RCpT/yIpPVTMs555TsWHSsOVbRn0DLX3QTFpsus9X6xr4jNkO6SRCn dk0MnPNC0xldVM1jLsBplpdFXtBSFqBEluxds918eBHZ9H0wTR0cqp4RXx+WB2sphY7Z nlBw== X-Forwarded-Encrypted: i=1; AKwUvBwKzvZmmkrYkqZiTKbjE1zstycB1QtlT7jvab/lB0cdm9d2ANoYnbs57uaZYqV6PXY7AvYObHbczH9rgWY=@vger.kernel.org X-Gm-Message-State: AFuF++nBYA+rMvMhNtBp6XBvlCIrw/gdl4Fuzk3bR4wRbfzuisiLBsO9 JNdhDq/CMilS7qt+Hk2AiSJismCPvHAVT/FmpNFVsx0iNxRvSygqmiJm X-Gm-Gg: AYBFou25hRso4HsIFQomDReavnjivYRTmEY5mH/ILpXNoOzJoxtamv7ck+cDsz0tEB2 s8cLx8UEz3cDAWSu2rZ2PH/Vj1JZUAYTmVFBkfhn45lpwJS5XJ/WSrAHpKLKMZgi41TTS9LbKL0 LiuXMEOE7VELBp/W2jepgPw0oOezckOatdvZS2eUqqytuaHj/b4d6Lg8SlEb6V+Y3nKtIOxR1jn DjH8qYInAwSKncOr0HZRr2esTT5mfVNgBkJTtjIhLjwV5q61TLvUV/eRzalvSeP+gsMeORY5Eiu qr1nKE1Fn/BxSlQonPl0kGuZUx+ky3/YM9c34jtAmq69oHKdnwQkVyp/gyRiJJY63wl3PVli2lI sy0l6gLywE2lLuscaRtl/inGUdsPxDLmbNPcherrRBzUkkpKrE7lyU2HhbygC3EJCOupV1JMMAV UgpjyZpdkAmRJQaDVvM5EdeT/YFhzvd/3ajR16odp53+tb+/DEM9kdSgeCPy0VIsXYtyxwjJZUJ S/QnsdasC2naNE19Bovpg== X-Received: by 2002:a9d:7085:0:b0:802:9f53:3a66 with SMTP id 46e09a7af769-80c4bd14aedmr4916798a34.6.1789754566141; Fri, 18 Sep 2026 11:02:46 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4e::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-8107dd8bd19sm170602a34.8.2026.09.18.11.02.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:45 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 02/11] mm, swap: support zswap and zero-filled swap pages as vswap backends Date: Fri, 18 Sep 2026 11:02:32 -0700 Message-ID: <20260918180241.3424851-3-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Build the virtual swap layer on top of the swap-table infrastructure. Virtual swap entries decouple PTE swap entries from physical backing, allowing pages to be compressed by zswap (or detected as zero-filled) without pre-allocating a physical swap slot. This patch only supports zswap and zero-page backends. If zswap_store fails, the page stays dirty in the swap cache. Physical disk backing arrives in later patches. Zswap writeback of vswap-backed entries is also disabled: they have no physical slot to write back to yet, so the zswap shrinker (both the dynamic count path and the pool-full worker path) is skipped while vswap is enabled. Physical backing and real writeback come in later patches. THP swapin is disabled for vswap entries for now. vswap_alloc() only routes a swapout through vswap when vswap is enabled, either by CONFIG_VSWAP_DEFAULT_ON or by "vswap=3Don" on the kernel command line. It also declines the folio when zswap is off, or when the folio's cgroup is already over its zswap limit. In both cases writeout would have to find a physical slot anyway, so the indirection would buy nothing. Anon reclaim gating needs a matching change: a vswap zswap-backed swapout consumes no physical slot, so the physical free count no longer tells reclaim whether anon is reclaimable. Add mem_cgroup_can_swap(), which answers that question and gates on the swap.max headroom instead when vswap and zswap are both on. Its callers only ever compared mem_cgroup_get_nr_swap_pages() against a threshold, so that function keeps its meaning and its physical free count. The OVERCOMMIT_GUESS heuristic gates a single allocation on totalram_pages() + total_swap_pages. Vswap deliberately contributes to neither, so on a host with no swapfile the bound collapses to RAM alone and allocations that vswap could absorb are refused. Compression swap is backed by memory rather than by a device, so reference the bound to RAM: add 2 * totalram_pages() when vswap and zswap are both on, for a 3x RAM ceiling plus whatever physical swap exists. mem_cgroup_can_vswap() is added alongside, so the !CONFIG_MEMCG stub of mem_cgroup_can_swap() gets the same carve-out as the real one. Without it, get_nr_swap_pages() reads 0 under vswap and MGLRU concludes anon can never be swapped. Suggested-by: Kairui Song Signed-off-by: Nhat Pham --- include/linux/swap.h | 13 +++ include/linux/zswap.h | 4 + mm/memcontrol.c | 29 +++++++ mm/memory.c | 11 ++- mm/page_io.c | 12 ++- mm/shmem.c | 4 +- mm/swap.h | 1 + mm/swap_state.c | 8 ++ mm/swapfile.c | 187 ++++++++++++++++++++++++++++++++++++++++-- mm/util.c | 13 ++- mm/vmscan.c | 18 ++-- mm/vswap.h | 185 +++++++++++++++++++++++++++++++++++++++++ mm/workingset.c | 2 +- mm/zswap.c | 68 ++++++++++++--- 14 files changed, 521 insertions(+), 34 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 0b341301d8fa..abf658f7861f 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -367,6 +367,8 @@ extern void __meminit kswapd_stop(int nid); bool current_is_kswapd(void); =20 #ifdef CONFIG_SWAP +bool mem_cgroup_can_vswap(struct mem_cgroup *memcg); + int add_swap_extent(struct swap_info_struct *sis, unsigned long start_page, unsigned long nr_pages, sector_t start_block); int generic_swapfile_activate(struct swap_info_struct *, struct file *, @@ -440,6 +442,11 @@ static inline void put_swap_device(struct swap_info_st= ruct *si) } =20 #else /* CONFIG_SWAP */ +static inline bool mem_cgroup_can_vswap(struct mem_cgroup *memcg) +{ + return false; +} + static inline struct swap_info_struct *get_swap_device(swp_entry_t entry) { return NULL; @@ -538,6 +545,7 @@ static inline void mem_cgroup_uncharge_swap(unsigned sh= ort id, unsigned int nr_p =20 long mem_cgroup_get_folio_swap_margin(struct folio *folio); extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg); +bool mem_cgroup_can_swap(struct mem_cgroup *memcg, long nr_pages); extern bool mem_cgroup_swap_full(struct folio *folio); #else static inline int mem_cgroup_try_charge_swap(struct folio *folio) @@ -560,6 +568,11 @@ static inline long mem_cgroup_get_nr_swap_pages(struct= mem_cgroup *memcg) return get_nr_swap_pages(); } =20 +static inline bool mem_cgroup_can_swap(struct mem_cgroup *memcg, long nr_p= ages) +{ + return mem_cgroup_can_vswap(memcg) || get_nr_swap_pages() >=3D nr_pages; +} + static inline bool mem_cgroup_swap_full(struct folio *folio) { return vm_swap_full(); diff --git a/include/linux/zswap.h b/include/linux/zswap.h index df6cafbe95dc..0cc2cc878566 100644 --- a/include/linux/zswap.h +++ b/include/linux/zswap.h @@ -6,6 +6,7 @@ #include =20 struct lruvec; +struct zswap_entry; =20 extern atomic_long_t zswap_stored_pages; =20 @@ -28,6 +29,7 @@ unsigned long zswap_total_pages(void); bool zswap_store(struct folio *folio); int zswap_load(struct folio *folio); void zswap_invalidate(int type, pgoff_t offset, unsigned long nr_entries); +void zswap_entry_free(struct zswap_entry *entry); int zswap_swapon(int type, unsigned long nr_pages); void zswap_swapoff(int type); void zswap_memcg_offline_cleanup(struct mem_cgroup *memcg); @@ -54,6 +56,8 @@ static inline void zswap_invalidate(int type, pgoff_t off= set, { } =20 +static inline void zswap_entry_free(struct zswap_entry *entry) {} + static inline int zswap_swapon(int type, unsigned long nr_pages) { return 0; diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 791e536efaeb..63d2c9e3dbe1 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -66,6 +66,7 @@ #include "internal.h" #include "swap.h" #include "swap_table.h" +#include "vswap.h" #include #include #include "slab.h" @@ -6049,6 +6050,34 @@ long mem_cgroup_get_folio_swap_margin(struct folio *= folio) return margin; } =20 +/** + * mem_cgroup_can_swap - can @memcg swap out at least @nr_pages more pages? + * @memcg: the memcg to query + * @nr_pages: the number of pages the caller wants to swap out + * + * A vswap zswap-backed swapout needs no physical slot, so gate on the + * swap.max headroom rather than the physical free count. + * + * Return: true if @memcg can swap out at least @nr_pages more pages. + */ +bool mem_cgroup_can_swap(struct mem_cgroup *memcg, long nr_pages) +{ + long avail; + + if (mem_cgroup_can_vswap(memcg)) + return true; + + if (!vswap_is_enabled() || !zswap_is_enabled()) + return mem_cgroup_get_nr_swap_pages(memcg) >=3D nr_pages; + + avail =3D PAGE_COUNTER_MAX; + for (; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) + avail =3D min_t(long, avail, + READ_ONCE(memcg->swap.max) - + page_counter_read(&memcg->swap)); + return avail >=3D nr_pages; +} + bool mem_cgroup_swap_full(struct folio *folio) { struct mem_cgroup *memcg; diff --git a/mm/memory.c b/mm/memory.c index 338fce99e711..73ebc59d2b03 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -89,6 +89,7 @@ #include "pgalloc-track.h" #include "internal.h" #include "swap.h" +#include "vswap.h" =20 #if defined(LAST_CPUPID_NOT_IN_PAGE_FLAGS) && !defined(CONFIG_COMPILE_TEST) #warning Unfortunate NUMA and NUMA Balancing config, growing page-frame fo= r last_cpupid. @@ -4712,6 +4713,9 @@ static inline bool should_try_to_free_swap(struct swa= p_info_struct *si, { if (!folio_test_swapcache(folio)) return false; + /* A vswap entry holds no physical slot, so keeping it saves no IO. */ + if (is_vswap_entry(folio->swap)) + return true; /* * Always try to free swap cache for SWP_SYNCHRONOUS_IO devices. Swap * cache can help save some IO or memory overhead, but these devices @@ -4868,15 +4872,18 @@ static unsigned long thp_swapin_suitable_orders(str= uct vm_fault *vmf) if (unlikely(userfaultfd_armed(vma))) return 0; =20 + entry =3D softleaf_from_pte(vmf->orig_pte); + /* * A large swapped out folio could be partially or fully in zswap. We * lack handling for such cases, so fallback to swapping in order-0 * folio. + * + * THP swapin for vswap is not supported yet either. */ - if (!zswap_never_enabled()) + if (is_vswap_entry(entry) || !zswap_never_enabled()) return 0; =20 - entry =3D softleaf_from_pte(vmf->orig_pte); /* * Get a list of all the (large) orders below PMD_ORDER that are enabled * and suitable for swapping THP. diff --git a/mm/page_io.c b/mm/page_io.c index 5cd77bd1b12c..2e7fb335eb01 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -162,13 +162,18 @@ static void swap_zeromap_folio_set(struct folio *foli= o) int nr_pages =3D folio_nr_pages(folio); struct swap_cluster_info *ci; swp_entry_t entry =3D folio->swap; - unsigned int i; + unsigned int voff, i; =20 VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); =20 ci =3D swap_cluster_get_and_lock(folio); - for (i =3D 0; i < folio_nr_pages(folio); i++) { + if (is_vswap_entry(folio->swap)) { + /* Free any prior backing (e.g. ZSWAP entry from earlier swapout) */ + voff =3D swp_cluster_offset(folio->swap); + __vswap_release_backing(ci, voff, nr_pages); + } + for (i =3D 0; i < nr_pages; i++) { __swap_table_set_zero(ci, swp_cluster_offset(entry)); entry.val++; } @@ -236,6 +241,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio= *folio) */ swap_zeromap_folio_clear(folio); =20 + if (is_vswap_entry(folio->swap)) + folio_release_vswap_backing(folio); + if (zswap_store(folio)) { count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT); goto out_unlock; diff --git a/mm/shmem.c b/mm/shmem.c index b572c60f2af8..2871a1f618ec 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -86,6 +86,7 @@ static struct vfsmount *shm_mnt __ro_after_init; #include =20 #include "internal.h" +#include "vswap.h" =20 #define VM_ACCT(size) (PAGE_ALIGN(size) >> PAGE_SHIFT) =20 @@ -1819,7 +1820,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct fo= lio *folio, if ((info->flags & SHMEM_F_LOCKED) || sbinfo->noswap) goto redirty; =20 - if (!total_swap_pages) + /* vswap doesn't contribute to total_swap_pages */ + if (!total_swap_pages && !(vswap_is_enabled() && zswap_is_enabled())) goto redirty; =20 /* diff --git a/mm/swap.h b/mm/swap.h index 528a27afb335..81e47dc36a02 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -71,6 +71,7 @@ struct swap_cluster_info_dynamic { struct swap_cluster_info ci; unsigned int index; /* for cluster_index() */ struct rcu_head rcu; + atomic_long_t *virtual_table; /* Backing pointers for vswap slots */ }; =20 /* All on-list cluster must have a non-zero flag. */ diff --git a/mm/swap_state.c b/mm/swap_state.c index 424a0040b4b8..54e4fefdab95 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -27,6 +27,7 @@ #include "internal.h" #include "swap_table.h" #include "swap.h" +#include "vswap.h" =20 /* Swap readahead cluster size, as a power of 2 pages. */ static int page_cluster; @@ -188,6 +189,13 @@ static int __swap_cache_add_check(struct swap_cluster_= info *ci, if (nr =3D=3D 1) return 0; =20 + /* + * Reject a vswap batch so swap_cache_alloc_folio falls back to + * order 0. + */ + if (is_vswap_entry(targ_entry)) + return -EBUSY; + is_zero =3D __swap_table_test_zero(ci, ci_off); ci_off =3D round_down(ci_off, nr); ci_end =3D ci_off + nr; diff --git a/mm/swapfile.c b/mm/swapfile.c index edd7bddf7ff9..67a2399cf1ae 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -46,6 +46,7 @@ =20 #include #include +#include "memcontrol-v1.h" #include "swap_table.h" #include "vswap.h" #include "internal.h" @@ -131,6 +132,16 @@ static DEFINE_PER_CPU(struct percpu_swap_cluster, perc= pu_swap_cluster) =3D { .lock =3D INIT_LOCAL_LOCK(), }; =20 +struct percpu_vswap_cluster { + unsigned long offset[SWAP_NR_ORDERS]; + local_lock_t lock; +}; + +static DEFINE_PER_CPU(struct percpu_vswap_cluster, percpu_vswap_cluster) = =3D { + .offset =3D { [0 ... SWAP_NR_ORDERS - 1] =3D SWAP_ENTRY_INVALID }, + .lock =3D INIT_LOCAL_LOCK(), +}; + /* May return NULL on invalid type, caller must check for NULL return */ static struct swap_info_struct *swap_type_to_info(int type) { @@ -236,7 +247,8 @@ static int __try_to_reclaim_swap(struct swap_info_struc= t *si, =20 need_reclaim =3D ((flags & TTRS_ANYWAY) || ((flags & TTRS_UNMAPPED) && !folio_mapped(folio)) || - ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio))); + ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio) && + !is_vswap_entry(folio->swap))); if (!need_reclaim || !folio_swapcache_freeable(folio)) goto out_unlock; =20 @@ -544,7 +556,9 @@ swap_cluster_populate(struct swap_info_struct *si, /* * Only cluster isolation from the allocator does table allocation. * Swap allocator uses percpu clusters and holds the local lock. + * vswap clusters are destroyed rather than freed to si->free_clusters. */ + VM_WARN_ON_ONCE(swap_is_vswap(si)); lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); if (!(si->flags & SWP_SOLIDSTATE)) lockdep_assert_held(&si->global_cluster_lock); @@ -632,6 +646,7 @@ static void __free_cluster(struct swap_info_struct *si,= struct swap_cluster_info ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); xa_erase(&si->cluster_info_pool, ci_dyn->index); move_cluster(si, ci, NULL, CLUSTER_FLAG_DEAD); + vswap_cluster_free_vtable(ci); kfree_rcu(ci_dyn, rcu); return; } @@ -926,7 +941,8 @@ static bool cluster_scan_range(struct swap_info_struct = *si, if (swp_tb_is_null(swp_tb)) continue; if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) { - if (!vm_swap_full()) + /* vswap slots are abundant; never reclaim to reuse one */ + if (swap_is_vswap(si) || !vm_swap_full()) return false; *need_reclaim =3D true; continue; @@ -1034,6 +1050,10 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, out: relocate_cluster(si, ci); swap_cluster_unlock(ci); + if (swap_is_vswap(si)) { + this_cpu_write(percpu_vswap_cluster.offset[order], next); + return found; + } if (si->flags & SWP_SOLIDSTATE) { this_cpu_write(percpu_swap_cluster.offset[order], next); this_cpu_write(percpu_swap_cluster.si[order], si); @@ -1086,6 +1106,12 @@ static unsigned int vswap_alloc_cluster(struct swap_= info_struct *si, return SWAP_ENTRY_INVALID; } =20 + if (vswap_cluster_alloc_vtable(ci_dyn, GFP_ATOMIC | __GFP_NOWARN)) { + swap_cluster_free_table(&ci_dyn->ci); + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + /* Lock before publishing: xa_alloc makes the cluster findable by offset.= */ ci =3D &ci_dyn->ci; spin_lock(&ci->lock); @@ -1095,6 +1121,7 @@ static unsigned int vswap_alloc_cluster(struct swap_i= nfo_struct *si, GFP_ATOMIC | __GFP_NOWARN)) { spin_unlock(&ci->lock); swap_cluster_free_table(&ci_dyn->ci); + vswap_cluster_free_vtable(&ci_dyn->ci); kfree(ci_dyn); return SWAP_ENTRY_INVALID; } @@ -1178,7 +1205,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, * Swapfile is not block device so unable * to allocate large entries. */ - if (order && !(si->flags & SWP_BLKDEV)) + if (order && !(si->flags & SWP_BLKDEV) && !swap_is_vswap(si)) return 0; =20 if (!(si->flags & SWP_SOLIDSTATE)) { @@ -1231,7 +1258,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, } =20 /* Try reclaim full clusters if free and nonfull lists are drained */ - if (vm_swap_full()) + if (!swap_is_vswap(si) && vm_swap_full()) swap_reclaim_full_clusters(si, false); =20 if (order < PMD_ORDER) { @@ -1392,7 +1419,8 @@ static void swap_range_alloc(struct swap_info_struct = *si, if (vm_swap_full()) schedule_work(&si->reclaim_work); } - atomic_long_sub(nr_entries, &nr_swap_pages); + if (!swap_is_vswap(si)) + atomic_long_sub(nr_entries, &nr_swap_pages); } =20 static void swap_range_free(struct swap_info_struct *si, unsigned long off= set, @@ -1401,7 +1429,8 @@ static void swap_range_free(struct swap_info_struct *= si, unsigned long offset, unsigned long end =3D offset + nr_entries - 1; void (*swap_slot_free_notify)(struct block_device *, unsigned long); =20 - zswap_invalidate(si->type, offset, nr_entries); + if (!swap_is_vswap(si)) + zswap_invalidate(si->type, offset, nr_entries); =20 if (si->flags & SWP_BLKDEV) swap_slot_free_notify =3D @@ -1420,7 +1449,8 @@ static void swap_range_free(struct swap_info_struct *= si, unsigned long offset, * only after the above cleanups are done. */ smp_wmb(); - atomic_long_add(nr_entries, &nr_swap_pages); + if (!swap_is_vswap(si)) + atomic_long_add(nr_entries, &nr_swap_pages); swap_usage_sub(si, nr_entries); } =20 @@ -1808,6 +1838,68 @@ static int swap_dup_entries_cluster(struct swap_info= _struct *si, return err; } =20 +/* + * Virtual swap is unbounded for a zswap-capable memcg: vswap charges the + * physical backing, not the allocation, so the swap.max walk would starve + * anon reclaim. swap.max is still enforced when the backing is charged. + */ +bool mem_cgroup_can_vswap(struct mem_cgroup *memcg) +{ + return vswap_is_enabled() && zswap_is_enabled() && + (mem_cgroup_disabled() || do_memsw_account()); +} + +static bool vswap_alloc(struct folio *folio) +{ + unsigned int order =3D folio_order(folio); + struct swap_cluster_info *ci; + struct obj_cgroup *objcg; + unsigned long offset; + bool may_zswap; + + if (!vswap_is_enabled() || !zswap_is_enabled()) + return false; + + /* + * If zswap will not take the folio, writeout has to find a physical + * slot anyway. We are just incurring indirection overhead + * unnecessarily. + */ + objcg =3D get_obj_cgroup_from_folio(folio); + may_zswap =3D !objcg || obj_cgroup_may_zswap(objcg); + if (objcg) + obj_cgroup_put(objcg); + if (!may_zswap) + return false; + + local_lock(&percpu_vswap_cluster.lock); + offset =3D this_cpu_read(percpu_vswap_cluster.offset[order]); + + if (offset !=3D SWAP_ENTRY_INVALID) { + ci =3D swap_cluster_lock(vswap_si, offset); + if (ci && cluster_is_usable(ci, order)) { + if (cluster_is_empty(ci)) + offset =3D cluster_offset(vswap_si, ci); + alloc_swap_scan_cluster(vswap_si, ci, folio, offset); + } else if (ci) { + swap_cluster_unlock(ci); + } + } + + if (!folio_test_swapcache(folio)) + cluster_alloc_swap_entry(vswap_si, folio); + + if (folio_test_swapcache(folio)) { + /* alloc_swap_scan_cluster updated percpu offset already */ + local_unlock(&percpu_vswap_cluster.lock); + return true; + } + + this_cpu_write(percpu_vswap_cluster.offset[order], SWAP_ENTRY_INVALID); + local_unlock(&percpu_vswap_cluster.lock); + return false; +} + /** * folio_alloc_swap - allocate swap space for a folio * @folio: folio we want to move to swap @@ -1846,12 +1938,16 @@ int folio_alloc_swap(struct folio *folio) } } =20 + if (vswap_alloc(folio)) + goto done; + again: local_lock(&percpu_swap_cluster.lock); if (!swap_alloc_fast(folio)) swap_alloc_slow(folio); local_unlock(&percpu_swap_cluster.lock); =20 +done: if (!order && unlikely(!folio_test_swapcache(folio))) { if (swap_sync_discard()) goto again; @@ -1869,7 +1965,7 @@ int folio_alloc_swap(struct folio *folio) return 0; =20 failed: - if (get_nr_swap_pages() <=3D 0) + if (get_nr_swap_pages() <=3D 0 && !vswap_is_enabled()) return -ENOSPC; if (mem_cgroup_get_folio_swap_margin(folio) <=3D 0) return -ENOMEM; @@ -1877,6 +1973,73 @@ int folio_alloc_swap(struct folio *folio) return order ? -E2BIG : -ENOMEM; } =20 +/** + * __vswap_release_backing - release the backing of a range of vtable slots + * @ci: the locked vswap cluster + * @ci_start: first slot offset within @ci + * @nr: number of slots + * + * Releases the backing of each slot in [@ci_start, @ci_start + @nr). + * Clears the zero marks if set. + * + * Context: caller must hold @ci->lock. + */ +void __vswap_release_backing(struct swap_cluster_info *ci, + unsigned int ci_start, unsigned int nr) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int ci_off; + unsigned long vt; + + lockdep_assert_held(&ci->lock); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + + for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { + vt =3D __vtable_get(ci_dyn, ci_off); + + switch (vtable_type(vt)) { + case VSWAP_ZSWAP: + zswap_entry_free(vtable_to_zswap(vt)); + break; + case VSWAP_NONE: + break; + default: + /* VSWAP_ZERO/VSWAP_FOLIO are return-only, not vtable tags */ + break; + } + + __vtable_set(ci_dyn, ci_off, VSWAP_NONE); + /* Zero-backed state lives in swap_table; clear it too. */ + if (__swap_table_test_zero(ci, ci_off)) + __swap_table_clear_zero(ci, ci_off); + } +} + +/** + * folio_release_vswap_backing() - Drop all backing for a folio's vswap en= try. + * @folio: the folio, occupying a virtual swap entry. + * + * Release whatever backing the folio's virtual swap slots currently hold = and + * reset them to empty, so a fresh backing can be installed. Used when a + * folio's swap backend is replaced. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + */ +void folio_release_vswap_backing(struct folio *folio) +{ + struct swap_cluster_info *ci; + int nr =3D folio_nr_pages(folio); + unsigned int voff; + + ci =3D __swap_entry_to_cluster(folio->swap); + voff =3D swp_cluster_offset(folio->swap); + + spin_lock(&ci->lock); + __vswap_release_backing(ci, voff, nr); + spin_unlock(&ci->lock); +} + /** * folio_dup_swap() - Increase swap count of swap entries of a folio. * @folio: folio with swap entries bounded. @@ -2022,6 +2185,9 @@ void __swap_cluster_free_entries(struct swap_info_str= uct *si, =20 VM_WARN_ON(ci->count < nr_pages); =20 + if (swap_is_vswap(si)) + __vswap_release_backing(ci, ci_start, nr_pages); + ci->count -=3D nr_pages; do { old_tb =3D __swap_table_get(ci, ci_off); @@ -3187,6 +3353,7 @@ static void free_swap_cluster_info(struct swap_info_s= truct *si, swap_cluster_free_table(ci); } spin_unlock(&ci->lock); + vswap_cluster_free_vtable(ci); kfree(ci_dyn); } xa_destroy(&si->cluster_info_pool); @@ -3716,6 +3883,10 @@ static int setup_swap_clusters_info(struct swap_info= _struct *si, if (err) goto err; =20 + err =3D vswap_cluster_alloc_vtable(ci_dyn, GFP_KERNEL); + if (err) + goto err; + goto setup_cluster_info; } =20 diff --git a/mm/util.c b/mm/util.c index c5ee52aede1e..bde985f242df 100644 --- a/mm/util.c +++ b/mm/util.c @@ -33,6 +33,7 @@ =20 #include "internal.h" #include "swap.h" +#include "vswap.h" =20 /** * kfree_const - conditionally free memory @@ -970,7 +971,17 @@ int __vm_enough_memory(const struct mm_struct *mm, lon= g pages, int cap_sys_admin return 0; =20 if (sysctl_overcommit_memory =3D=3D OVERCOMMIT_GUESS) { - if (pages > totalram_pages() + total_swap_pages) + allowed =3D totalram_pages() + total_swap_pages; + /* + * Vswap is in neither term, so with no swapfile the bound + * would collapse to RAM alone even though zswap can absorb + * more. Compression swap is backed by memory, so budget a 3x + * compression ratio, typical of most workloads. Anything that + * compresses better can use OVERCOMMIT_ALWAYS. + */ + if (vswap_is_enabled() && zswap_is_enabled()) + allowed +=3D 2 * totalram_pages(); + if (pages > allowed) goto error; return 0; } diff --git a/mm/vmscan.c b/mm/vmscan.c index 836f50814ffa..6cf689817f2e 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -69,6 +69,7 @@ #include "internal.h" #include "page_alloc.h" #include "swap.h" +#include "vswap.h" =20 #define CREATE_TRACE_POINTS #include @@ -415,10 +416,12 @@ static inline bool can_reclaim_anon_pages(struct mem_= cgroup *memcg, if (memcg =3D=3D NULL) { /* * For non-memcg reclaim, is there space in any swap device? - * And under GFP_NOIO, is there enough swapcached anon to make - * scanning anon worthwhile? + * vswap does not contribute to nr_swap_pages. And under + * GFP_NOIO, is there enough swapcached anon to make scanning + * anon worthwhile? */ - if (get_nr_swap_pages() > 0 && + if ((get_nr_swap_pages() > 0 || + (vswap_is_enabled() && zswap_is_enabled())) && !reclaimable_anon_is_low(memcg, nid, sc)) return true; } else { @@ -426,7 +429,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cg= roup *memcg, * Is the memcg below its swap limit, and under GFP_NOIO does * it have enough swapcached anon to make scanning worthwhile? */ - if (mem_cgroup_get_nr_swap_pages(memcg) > 0 && + if (mem_cgroup_can_swap(memcg, 1) && !reclaimable_anon_is_low(memcg, nid, sc)) return true; } @@ -1603,7 +1606,8 @@ static unsigned int shrink_folio_list(struct list_hea= d *folio_list, activate_locked: /* Not a candidate for swapping, so reclaim swap space. */ if (folio_test_swapcache(folio) && - (mem_cgroup_swap_full(folio) || folio_test_mlocked(folio))) + ((mem_cgroup_swap_full(folio) && !is_vswap_entry(folio->swap)) || + folio_test_mlocked(folio))) folio_free_swap(folio); VM_BUG_ON_FOLIO(folio_test_active(folio), folio); if (!folio_test_mlocked(folio)) { @@ -2760,7 +2764,7 @@ static bool can_age_anon_pages(struct lruvec *lruvec, struct scan_control *sc) { /* Aging the anon LRU is valuable if swap is present: */ - if (total_swap_pages > 0) + if (total_swap_pages > 0 || (vswap_is_enabled() && zswap_is_enabled())) return true; =20 /* Also valuable if anon pages can be demoted: */ @@ -2853,7 +2857,7 @@ static int get_swappiness(struct lruvec *lruvec, stru= ct scan_control *sc) return 0; =20 if (!can_demote(pgdat->node_id, sc, memcg) && - mem_cgroup_get_nr_swap_pages(memcg) < MIN_LRU_BATCH) + !mem_cgroup_can_swap(memcg, MIN_LRU_BATCH)) return 0; =20 return swappiness; diff --git a/mm/vswap.h b/mm/vswap.h index 16395f357955..5334c77b6b84 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -11,8 +11,22 @@ #include #include "swap.h" =20 +struct zswap_entry; + +/* + * VSWAP_ZERO and VSWAP_FOLIO are return-only values synthesized from + * swap_table state; the rest are stored in the vtable per slot. + */ +enum vswap_backing_type { + VSWAP_NONE =3D 0, + VSWAP_ZSWAP =3D 1, + VSWAP_ZERO, + VSWAP_FOLIO, +}; + #ifdef CONFIG_SWAP =20 +#include "swap_table.h" DECLARE_STATIC_KEY_FALSE(vswap_key); =20 /* @@ -29,6 +43,177 @@ static inline bool is_vswap_entry(swp_entry_t entry) return swap_is_vswap(__swap_entry_to_info(entry)); } =20 +/* + * Virtual table entry encoding for vswap clusters. + * + * Each entry in ci_dyn->virtual_table stores the backing type and + * pointer for a virtual swap slot. Tag in low 3 bits, payload in + * upper 61 bits. + * + * NONE: |----- 0000 ------|000| - no separate backend pointer + * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low bi= ts) + * + * Pointer payloads (ZSWAP) are stored directly with the tag OR'd into the + * low bits (kernel pointers are >=3D 8-byte aligned, same approach as xar= ray). + * + * vtable[i] =3D NONE does not by itself mean "free". The swap_table entry + * and the per-slot zero flag carry the rest of the state. The full + * per-slot state table is: + * + * vtable[i] | swap_table[i] | zero | meaning + * ----------+---------------+-------+-------------------------------- + * NONE | NULL | clear | truly free / unbacked + * NONE | PFN | clear | folio cached, no backing + * NONE | shadow | clear | evicted, no backing: data lost + * NONE | * | set | zero-backed; cached if PFN set + * ZSWAP | PFN | clear | folio cached + zswap entry + * ZSWAP | shadow / NULL | clear | evicted, only in zswap + * + * Locking: a slot's vtable entry (the vswap entry's backend) is only + * stable while the caller owns and holds the lock on that entry's swap + * cache folio. The cluster lock (ci_dyn->ci.lock) only makes an individual + * vtable read atomic, and by itself does not give the caller the right to + * change the backend. A backend read without the folio lock is + * best-effort and must be re-validated under the folio lock before + * being acted on. + * + * Zero-backed slots use the swap_table per-slot zero flag (same as + * direct-mapped physical swap), via __swap_table_test_zero() and friends, + * which fall back to ci->zero_bitmap where the flag does not fit. Cached + * folios are read out of the swap_table PFN entry; there is no separate F= OLIO + * vtable type because the folio pointer would duplicate that PFN and + * would go stale on folio migration / split. + */ + +#define VTABLE_TAG_BITS 3 +#define VTABLE_TAG_MASK ((1UL << VTABLE_TAG_BITS) - 1) + +static inline enum vswap_backing_type vtable_type(unsigned long vt) +{ + return vt & VTABLE_TAG_MASK; +} + +static inline struct zswap_entry *vtable_to_zswap(unsigned long vt) +{ + VM_WARN_ON(vtable_type(vt) !=3D VSWAP_ZSWAP); + return (struct zswap_entry *)(vt & ~VTABLE_TAG_MASK); +} + +/* Virtual table accessors */ + +static inline unsigned long __vtable_get(struct swap_cluster_info_dynamic = *ci_dyn, + unsigned int off) +{ + VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); + return atomic_long_read(&ci_dyn->virtual_table[off]); +} + +static inline void __vtable_set(struct swap_cluster_info_dynamic *ci_dyn, + unsigned int off, unsigned long vt) +{ + VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); + atomic_long_set(&ci_dyn->virtual_table[off], vt); +} + +/** + * vswap_lock_cluster - look up and lock the vswap cluster for an entry + * @entry: the virtual swap entry + * @voff: out param, receives @entry's slot offset within the cluster + * + * Return: the locked vswap cluster, or NULL if @entry has no live cluster. + */ +static inline struct swap_cluster_info_dynamic * +vswap_lock_cluster(swp_entry_t entry, unsigned int *voff) +{ + struct swap_cluster_info *ci; + + ci =3D swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry)); + if (!ci) + return NULL; + *voff =3D swp_cluster_offset(entry); + return container_of(ci, struct swap_cluster_info_dynamic, ci); +} + +void __vswap_release_backing(struct swap_cluster_info *ci, + unsigned int ci_start, unsigned int nr); + +/** + * vswap_zswap_store - record a zswap entry as the backing for a vswap ent= ry. + * @entry: the vswap entry + * @ze: the zswap entry now holding @entry's compressed data + * + * Releases @entry's previous backing, and sets the zswap entry @ze as the= new + * backing. + * + * Context: takes and drops the vswap cluster lock internally. + */ +static inline void vswap_zswap_store(swp_entry_t entry, + struct zswap_entry *ze) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + __vswap_release_backing(&ci_dyn->ci, voff, 1); + __vtable_set(ci_dyn, voff, (unsigned long)ze | VSWAP_ZSWAP); + swap_cluster_unlock(&ci_dyn->ci); +} + +/** + * vswap_zswap_load - return the zswap entry backing a vswap entry + * @entry: the virtual swap entry + * + * Context: takes and drops the vswap cluster lock internally. + * Return: the backing zswap entry, or NULL if @entry is not zswap-backed. + */ +static inline struct zswap_entry *vswap_zswap_load(swp_entry_t entry) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + unsigned long vt; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) + return NULL; + vt =3D __vtable_get(ci_dyn, voff); + swap_cluster_unlock(&ci_dyn->ci); + + if (vtable_type(vt) !=3D VSWAP_ZSWAP) + return NULL; + return vtable_to_zswap(vt); +} + +void folio_release_vswap_backing(struct folio *folio); + +static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn, + gfp_t gfp) +{ + ci_dyn->virtual_table =3D kcalloc(SWAPFILE_CLUSTER, + sizeof(*ci_dyn->virtual_table), gfp); + return ci_dyn->virtual_table ? 0 : -ENOMEM; +} + +static inline void vswap_cluster_free_vtable(struct swap_cluster_info *ci) +{ + struct swap_cluster_info_dynamic *ci_dyn; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + kfree(ci_dyn->virtual_table); + ci_dyn->virtual_table =3D NULL; +} + +#else /* !CONFIG_SWAP */ + +static inline bool vswap_is_enabled(void) +{ + return false; +} + +static inline bool is_vswap_entry(swp_entry_t entry) +{ + return false; +} + #endif /* CONFIG_SWAP */ =20 #endif /* _MM_VSWAP_H */ diff --git a/mm/workingset.c b/mm/workingset.c index 8412f4840ae3..fe80cd50f3e9 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -523,7 +523,7 @@ bool workingset_test_recent(void *shadow, bool file, bo= ol *workingset, workingset_size +=3D lruvec_page_state(eviction_lruvec, NR_INACTIVE_FILE); } - if (mem_cgroup_get_nr_swap_pages(eviction_memcg) > 0) { + if (mem_cgroup_can_swap(eviction_memcg, 1)) { workingset_size +=3D lruvec_page_state(eviction_lruvec, NR_ACTIVE_ANON); if (file) { diff --git a/mm/zswap.c b/mm/zswap.c index 73728b510ff6..3466c80ac188 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -39,6 +39,7 @@ #include =20 #include "swap.h" +#include "vswap.h" #include "internal.h" =20 /********************************* @@ -257,6 +258,25 @@ static inline struct xarray *swap_zswap_tree(swp_entry= _t swp) return zswap_tree(swp_type(swp), swp_offset(swp)); } =20 +static struct zswap_entry *zswap_entry_load(swp_entry_t swp) +{ + if (is_vswap_entry(swp)) + return vswap_zswap_load(swp); + return xa_load(swap_zswap_tree(swp), swp_offset(swp)); +} + +static struct zswap_entry *zswap_entry_store(swp_entry_t swp, + struct zswap_entry *entry) +{ + if (is_vswap_entry(swp)) { + vswap_zswap_store(swp, entry); + return NULL; + } + + return xa_store(swap_zswap_tree(swp), swp_offset(swp), entry, + GFP_KERNEL); +} + #define zswap_pool_debug(msg, p) \ pr_debug("%s pool %s\n", msg, (p)->tfm_name) =20 @@ -774,7 +794,7 @@ static void zswap_entry_cache_free(struct zswap_entry *= entry) * Carries out the common pattern of freeing an entry's zsmalloc allocatio= n, * freeing the entry itself, and decrementing the number of stored pages. */ -static void zswap_entry_free(struct zswap_entry *entry) +void zswap_entry_free(struct zswap_entry *entry) { struct zswap_pool *pool =3D zswap_entry_pool(entry); =20 @@ -1226,6 +1246,9 @@ static unsigned long zswap_shrinker_count(struct shri= nker *shrinker, if (!zswap_shrinker_enabled || !mem_cgroup_zswap_writeback_enabled(memcg)) return 0; =20 + if (vswap_is_enabled()) + return 0; + /* * The shrinker resumes swap writeback, which will enter block * and may enter fs. XXX: Harmonize with vmscan.c __GFP_FS @@ -1308,6 +1331,8 @@ static struct shrinker *zswap_alloc_shrinker(void) * Return: 0 if at least one entry was written back, -EAGAIN if entries * were scanned but none could be written back, or -ENOENT if @memcg has * writeback disabled, is a zombie cgroup, or has empty zswap LRUs. + * + * Also returns -ENOENT when vswap is enabled. */ static int shrink_memcg(struct mem_cgroup *memcg) { @@ -1316,6 +1341,9 @@ static int shrink_memcg(struct mem_cgroup *memcg) if (!mem_cgroup_zswap_writeback_enabled(memcg)) return -ENOENT; =20 + if (vswap_is_enabled()) + return -ENOENT; + /* * Skip zombies because their LRUs are reparented and we would be * reclaiming from the parent instead of the dead memcg. @@ -1344,6 +1372,13 @@ static void shrink_worker(struct work_struct *w) int ret, failures =3D 0, attempts =3D 0; unsigned long thr; =20 + /* + * When vswap is enabled, zswap entries are almost all vswap backed, + * with no slot to write back to. + */ + if (vswap_is_enabled()) + return; + /* Reclaim down to the accept threshold */ thr =3D zswap_accept_thr_pages(); =20 @@ -1447,15 +1482,13 @@ static bool zswap_store_page(struct folio *folio, l= ong index, goto compress_failed; =20 /* - * Set pool_idx before the xa_store() below publishes the entry, or a + * Set pool_idx before the store below publishes the entry, or a * concurrent reader could resolve a stale pool_idx left by slab reuse * to an unrelated live pool. */ entry->pool_idx =3D pool->idx; =20 - old =3D xa_store(swap_zswap_tree(page_swpentry), - swp_offset(page_swpentry), - entry, GFP_KERNEL); + old =3D zswap_entry_store(page_swpentry, entry); if (xa_is_err(old)) { int err =3D xa_err(old); =20 @@ -1523,7 +1556,7 @@ bool zswap_store(struct folio *folio) struct mem_cgroup *memcg =3D NULL; struct zswap_pool *pool; bool ret =3D false; - long index; + long index =3D 0; =20 VM_WARN_ON_ONCE(!folio_test_locked(folio)); VM_WARN_ON_ONCE(!folio_test_swapcache(folio)); @@ -1576,14 +1609,21 @@ bool zswap_store(struct folio *folio) if (!ret && zswap_pool_reached_full) queue_work(shrink_wq, &zswap_shrink_work); check_old: + if (ret) + return ret; + /* * If the zswap store fails or zswap is disabled, we must invalidate * the possibly stale entries which were previously stored at the * offsets corresponding to each page of the folio. Otherwise, * writeback could overwrite the new data in the swapfile. */ - if (!ret) + if (is_vswap_entry(swp)) { + if (index > 0) + folio_release_vswap_backing(folio); + } else { zswap_invalidate(swp_type(swp), swp_offset(swp), nr_pages); + } =20 return ret; } @@ -1634,8 +1674,7 @@ static bool zswap_is_present(swp_entry_t entry, unsig= ned int nr) int zswap_load(struct folio *folio) { swp_entry_t swp =3D folio->swap; - pgoff_t offset =3D swp_offset(swp); - struct xarray *tree =3D swap_zswap_tree(swp); + struct swap_info_struct *si =3D __swap_entry_to_info(swp); struct zswap_entry *entry; =20 VM_WARN_ON_ONCE(!folio_test_locked(folio)); @@ -1659,7 +1698,7 @@ int zswap_load(struct folio *folio) return -ENOENT; } =20 - entry =3D xa_load(tree, offset); + entry =3D zswap_entry_load(swp); if (!entry) return -ENOENT; =20 @@ -1682,8 +1721,13 @@ int zswap_load(struct folio *folio) * compression work. */ folio_mark_dirty(folio); - xa_erase(tree, offset); - zswap_entry_free(entry); + + if (swap_is_vswap(si)) { + folio_release_vswap_backing(folio); + } else { + xa_erase(swap_zswap_tree(swp), swp_offset(swp)); + zswap_entry_free(entry); + } =20 folio_unlock(folio); return 0; --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f13.google.com (mail-oi2-f13.google.com [74.125.231.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 061D951B178 for ; Fri, 18 Sep 2026 18:02:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754572; cv=none; b=hYva3kCg8HmRMVoK51TedtwkDpBxNspKygKjLXYKgGf3AcqqcPnAskCzXlui/EIuRACad3UjH6NdNWJ0MqFzULtZLj7hA6yLZG/CGz7dB0v9oMrVV+9KQTtEJjNRBoPd5cl8Tc0nHHRZwA2QPSdYUWZcqzZYY1wnwQQjyiVKpk8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754572; c=relaxed/simple; bh=a4iX01xeKFkOjusL7wUAwQPEtM/OuMizB+7wXsVzgZw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LOsuFGGqCvpSNibC3kntnRZA48UJv5dZiVxAvn4+BTkEY9jDpCW7DIXfjW+KQTKZ4i1hN93N+jzaan+bVqS6y2/UCeWTOgCeZEK65muK3OGb8d0psiM2RtFdorFR+Q+Lq1k0zy2XlbgBEU47mvm0EccNzg7rV6Puw9BrfF9iSX4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=fGrzyr7u; arc=none smtp.client-ip=74.125.231.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="fGrzyr7u" Received: by mail-oi2-f13.google.com with SMTP id 5614622812f47-4b37a316adfso897923b6e.2 for ; Fri, 18 Sep 2026 11:02:49 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754569; x=1790359369; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ku1DiWNPMjUzjw3oBYfS28zIvxt+bdOecf2pjvKy8dE=; b=fGrzyr7uJTs3/WNJ0FyvQdvRjCVZdxPuEqm2eRj7MXFqDNYyL0+J5vxPck4QaZUOyo 5YC+Ko3ig5ecqBJRbgXx4JvslbvPVhSjOE0s/R+pIt/OdAK5m4jobwmjjVqyzN19hpuE bQD1GJo5hgYCxtglJIODPhg4cS5j3Il1DeSGm3pPE7PGUcDunK/HQ1ry+2GhMpzohQH6 LCqgLTo6pUmNcCj6sAHlp5Jxckdt00FY3ZEvv4nTtiDF7P7hBsRvMMdn/laqtNdgtVVw WVz7qauqHTF6tuLkqh63IGwTgH3zpZ3ABnHR9pdPmbJSvFTH0iDUvXR3VCNvTy7v2V52 9A2A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754569; x=1790359369; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ku1DiWNPMjUzjw3oBYfS28zIvxt+bdOecf2pjvKy8dE=; b=jJVgM40rQA/qtOx1wKUYqZlmLCLrUosyjEXuMKTXl2UAheXX75w40gkEVVocEhAKHT TIT4PK4nbBvZu278f+XUtwKkg1dMLw61nCWPdsMPPtnGU6ZypluIB3HEt+nHoQKcVVdc nUpVL271sQbIeiHjGoRU/YfxkSK9IuJ8bfgyuiY3L1u3FDoiThc892r8Ckjo4kaQkFXH z92BlB1/sbskKfOWvKm0ywDoJp48zu860YKCcUhusdWgb6AE5Cr6kFn+kr6NHVSnR9c2 CTdOrvsxCMYJYntY9a+fvpBtALUnUujGf0KruhKrscEAQAOt/BD62qoUNfK9jDLaadPk vCtw== X-Forwarded-Encrypted: i=1; AKwUvBw+/0L0NejFhnrr1GwPtoiOxeX5zsFVa7UzoQJpT2F7v4Fe8DzS9pYUl4NQWdBpxYL5lqJ0MpOrVx/x9k4=@vger.kernel.org X-Gm-Message-State: AFuF++msZtywb7Q68h0YNq+sypzlYUTcPN7TpK5HC0M27sZuqSjLYEqb 263dXV78A37WsRa7Pbd1EIiQzH0y5TcsJw9KWDHPucXMUV832PyBq3U3 X-Gm-Gg: AYBFou29ZAlP+uideyF9XRCkeVcK+oY5qOf5Z4m+aHioQxG5wU3jy7ATORVWBliEJor 3IbMIg6i19hCKTR2j8t7as2YNsZg3RxsUR7S2RPW7F7X2+Zvts7R80O3bcvzqGbhwOpwwu7K/1a Q/nT9jTkpRwhyek4YtVncXbwf1g7k39XJm09F05KlPKSnGfjBDUV0XCxu6MsdIvJ5ewRfKYyoX0 1jGDmJs3fLtcXtRgAJ5SMz9sOh3ThtCFFZz6D9xPaWxTumemCrNLxC7T+KuVXrDDjEFKRsJN/MU JUU6MnHgoaHFrycksHsZo44gfvxqRweH83RWm9JmJtTln67P8pR1fPBHBRPIv9QBuuuFsr7SKWG 14a9Rsm8Chun9yR3JsTmzJtjSZboKc8IpXKflSK/KuNxNqs8ygXQa3QXK/WjV6kSSOxGi41MLIi 9IElq1HIhEVCto3irwrb822Pwj+NxmQx+HMzM6iBxyxloCVRyRUEdt/QOYc5QGxyotLgZoFfXYj 1s85xyCoce2z6wyAYG5Xw== X-Received: by 2002:a05:6820:f007:b0:6b7:83d6:292a with SMTP id 006d021491bc7-6ca9cb53b30mr3234729eaf.45.1789754568673; Fri, 18 Sep 2026 11:02:48 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:11::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6cd3591d0b6sm565054eaf.12.2026.09.18.11.02.46 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:47 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 03/11] mm, swap: prepare the swap IO path for vswap Date: Fri, 18 Sep 2026 11:02:33 -0700 Message-ID: <20260918180241.3424851-4-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In preparation for adding a physical swap backend for vswap, make the swap IO path able to submit IO for a swap entry other than folio->swap. The swap IO path derives the target device and sector from folio->swap. For a vswap folio backed by a physical slot that entry is virtual, so it identifies neither the backing device nor the sector to submit IO against. Compute the sector from an explicit entry (swap_folio_sector becomes swap_entry_sector), thread that entry through swap_add_folio, __swap_writeout and ops->can_merge, and stash it in swap_iocb so the submit and completion paths address the IO from it rather than from folio->swap. This lets the batching path serve both vswap entries (backed by a physical slot) and physical entries mapped directly into PTEs. ops->can_merge changes anchor with it: swap_bdev_can_merge() compares sectors and swap_fs_can_merge() byte offsets, and both now measure from the batch head plus the accumulated length instead of from the previous folio plus its size. The two agree at every step, since a batch only grows through merges that passed the same test. The blkg comparison is unchanged and still anchors on the last folio in the batch. All callers pass folio->swap, so there is no functional change. Signed-off-by: Nhat Pham --- include/linux/swap.h | 2 +- include/linux/swap_ops.h | 9 ++++--- mm/page_io.c | 54 +++++++++++++++++++--------------------- mm/swap.h | 3 ++- mm/swapfile.c | 6 ++--- mm/zswap.c | 2 +- 6 files changed, 38 insertions(+), 38 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index abf658f7861f..5ab050b2457c 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -409,7 +409,7 @@ extern int __swap_count(swp_entry_t entry); extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t en= try); extern int swp_swapcount(swp_entry_t entry); extern struct swap_info_struct *get_swap_device(swp_entry_t entry); -sector_t swap_folio_sector(struct folio *folio); +sector_t swap_entry_sector(swp_entry_t entry); =20 /* * If there is an existing swap slot reference (swap entry) and the caller diff --git a/include/linux/swap_ops.h b/include/linux/swap_ops.h index 57ac6c703f68..223c84548bde 100644 --- a/include/linux/swap_ops.h +++ b/include/linux/swap_ops.h @@ -12,6 +12,7 @@ struct swap_iocb { struct bio_vec bvecs[SWAP_CLUSTER_MAX]; int nr_bvecs; int len; + swp_entry_t entry; /* first slot in the batch; addresses the IO */ }; =20 struct swap_io_ctx { @@ -30,15 +31,15 @@ struct swap_io_ctx { struct swap_ops { unsigned int flags; =20 - bool (*can_merge)(struct folio *folio, struct folio *prev_folio, - size_t prev_folio_size, int rw); + bool (*can_merge)(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw); void (*submit_write)(struct swap_io_ctx *ctx); void (*submit_read)(struct swap_io_ctx *ctx); }; =20 void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *= iter); -bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio, - size_t prev_folio_size, int rw); +bool swap_fs_can_merge(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw); int swap_fs_activate(struct swap_info_struct *sis, const struct swap_ops *= ops); =20 #endif /* _MM_SWAP_OPS_H */ diff --git a/mm/page_io.c b/mm/page_io.c index 2e7fb335eb01..632683232622 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -265,7 +265,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio= *folio) return AOP_WRITEPAGE_ACTIVATE; } =20 - __swap_writeout(ctx, folio); + __swap_writeout(ctx, folio, folio->swap); return 0; out_unlock: folio_unlock(folio); @@ -336,24 +336,22 @@ int sio_pool_init(void) } =20 static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio, - int rw) + swp_entry_t phys, int rw) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); - struct bio_vec *last_bv =3D &ctx->sio->bvecs[ctx->sio->nr_bvecs - 1]; - struct folio *prev_folio =3D bvec_folio(last_bv); - size_t prev_folio_size =3D folio_size(prev_folio); + struct swap_info_struct *sis =3D __swap_entry_to_info(phys); =20 if (ctx->sis !=3D sis) return false; - return sis->ops->can_merge(folio, prev_folio, prev_folio_size, rw); + return sis->ops->can_merge(folio, phys, ctx->sio, rw); } =20 -static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, i= nt rw) +static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys, int rw) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); + struct swap_info_struct *sis =3D __swap_entry_to_info(phys); struct swap_iocb *sio =3D ctx->sio; =20 - if (sio && !swap_can_merge(ctx, folio, rw)) { + if (sio && !swap_can_merge(ctx, folio, phys, rw)) { if (rw =3D=3D WRITE) swap_write_submit(ctx); else @@ -366,6 +364,7 @@ static void swap_add_folio(struct swap_io_ctx *ctx, str= uct folio *folio, int rw) ctx->sio =3D sio =3D mempool_alloc(sio_pool, GFP_NOIO); sio->nr_bvecs =3D 0; sio->len =3D 0; + sio->entry =3D phys; } bvec_set_folio(&sio->bvecs[sio->nr_bvecs], folio, folio_size(folio), 0); sio->len +=3D folio_size(folio); @@ -386,7 +385,8 @@ static void swap_add_folio(struct swap_io_ctx *ctx, str= uct folio *folio, int rw) } } =20 -void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio) +void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys) { VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); =20 @@ -402,7 +402,7 @@ void __swap_writeout(struct swap_io_ctx *ctx, struct fo= lio *folio) =20 folio_start_writeback(folio); folio_unlock(folio); - swap_add_folio(ctx, folio, WRITE); + swap_add_folio(ctx, folio, phys, WRITE); } =20 /* @@ -506,7 +506,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio) =20 /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); - swap_add_folio(ctx, folio, READ); + swap_add_folio(ctx, folio, folio->swap, READ); =20 finish: if (workingset) { @@ -538,8 +538,6 @@ static void swap_fs_write_complete(struct kiocb *iocb, = long ret) bool failed =3D ret !=3D sio->len; =20 if (failed) { - struct folio *folio =3D bvec_folio(&sio->bvecs[0]); - /* * In the case of swap-over-nfs, this can be a temporary failure * if the system has limited memory for allocating transmit @@ -547,7 +545,7 @@ static void swap_fs_write_complete(struct kiocb *iocb, = long ret) * folio_rotate_reclaimable but rate-limit the messages. */ pr_err_ratelimited("Write error %ld on dio swapfile (%llu)\n", - ret, swap_dev_pos(folio->swap)); + ret, swap_dev_pos(sio->entry)); } =20 swap_write_end(sio, failed); @@ -619,7 +617,7 @@ static void swap_bdev_submit_write(struct swap_io_ctx *= ctx) bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs), REQ_OP_WRITE | REQ_SWAP); bio->bi_iter.bi_size =3D sio->len; - bio->bi_iter.bi_sector =3D swap_folio_sector(bio_first_folio_all(bio)); + bio->bi_iter.bi_sector =3D swap_entry_sector(sio->entry); bio_associate_blkg_from_folio(bio, bio_first_folio_all(bio)); =20 if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) { @@ -639,7 +637,7 @@ static void swap_bdev_submit_read(struct swap_io_ctx *c= tx) bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs), REQ_OP_READ); bio->bi_iter.bi_size =3D sio->len; - bio->bi_iter.bi_sector =3D swap_folio_sector(bio_first_folio_all(bio)); + bio->bi_iter.bi_sector =3D swap_entry_sector(sio->entry); =20 if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) { /* @@ -657,13 +655,14 @@ static void swap_bdev_submit_read(struct swap_io_ctx = *ctx) } } =20 -static bool swap_bdev_can_merge(struct folio *folio, struct folio *prev_fo= lio, - size_t prev_folio_size, int rw) +static bool swap_bdev_can_merge(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw) { - if (swap_folio_sector(folio) !=3D - swap_folio_sector(prev_folio) + (prev_folio_size >> SECTOR_SHIFT)) + if (swap_entry_sector(phys) !=3D + swap_entry_sector(sio->entry) + (sio->len >> SECTOR_SHIFT)) return false; - if (rw =3D=3D WRITE && !folio_blkg_can_merge(folio, prev_folio)) + if (rw =3D=3D WRITE && !folio_blkg_can_merge(folio, + bvec_folio(&sio->bvecs[sio->nr_bvecs - 1]))) return false; return true; } @@ -679,7 +678,7 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw= , struct iov_iter *iter) struct swap_iocb *sio =3D ctx->sio; =20 init_sync_kiocb(&sio->iocb, ctx->sis->swap_file); - sio->iocb.ki_pos =3D swap_dev_pos(bvec_folio(&sio->bvecs[0])->swap); + sio->iocb.ki_pos =3D swap_dev_pos(sio->entry); if (rw =3D=3D WRITE) sio->iocb.ki_complete =3D swap_fs_write_complete; else @@ -690,11 +689,10 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int = rw, struct iov_iter *iter) } EXPORT_SYMBOL_GPL(swap_fs_prepare_rw); =20 -bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio, - size_t prev_folio_size, int rw) +bool swap_fs_can_merge(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw) { - return swap_dev_pos(folio->swap) =3D=3D - swap_dev_pos(prev_folio->swap) + prev_folio_size; + return swap_dev_pos(phys) =3D=3D swap_dev_pos(sio->entry) + sio->len; } EXPORT_SYMBOL_GPL(swap_fs_can_merge); =20 diff --git a/mm/swap.h b/mm/swap.h index 81e47dc36a02..df323d5e8da8 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -320,7 +320,8 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio); void swap_read_submit(struct swap_io_ctx *ctx); void swap_write_submit(struct swap_io_ctx *ctx); int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio); -void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio); +void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys); =20 /* linux/mm/swap_state.c */ extern struct address_space swap_space __read_mostly; diff --git a/mm/swapfile.c b/mm/swapfile.c index 67a2399cf1ae..af6a162c4450 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -342,14 +342,14 @@ offset_to_swap_extent(struct swap_info_struct *sis, u= nsigned long offset) BUG(); } =20 -sector_t swap_folio_sector(struct folio *folio) +sector_t swap_entry_sector(swp_entry_t entry) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); + struct swap_info_struct *sis =3D __swap_entry_to_info(entry); struct swap_extent *se; sector_t sector; pgoff_t offset; =20 - offset =3D swp_offset(folio->swap); + offset =3D swp_offset(entry); se =3D offset_to_swap_extent(sis, offset); sector =3D se->start_block + (offset - se->start_page); return sector << (PAGE_SHIFT - 9); diff --git a/mm/zswap.c b/mm/zswap.c index 3466c80ac188..c2430dbfc653 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1094,7 +1094,7 @@ static int zswap_writeback_entry(struct zswap_entry *= entry, folio_set_reclaim(folio); =20 /* start writeback */ - __swap_writeout(&ctx, folio); + __swap_writeout(&ctx, folio, folio->swap); swap_write_submit(&ctx); =20 out: --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oa2-f24.google.com (mail-oa2-f24.google.com [74.125.231.88]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 27EE4519905 for ; Fri, 18 Sep 2026 18:02:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.88 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754578; cv=none; b=kOX9xQEhUF3EqMpLeGTMosreANb6VVkR+fixtWzaLzj6GH5AotGqdp3ynHJP8y1jvghdWhropnoOBuFk5C5D16kNPrOpXjUFKrNsPls9xIAZSCcZwaqeYPNpv5r5rCuTBRmHwEkGgXtHdxqyu1b+M8DlsRhoc96jFiRvZq+4pic= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754578; c=relaxed/simple; bh=8AjVfqUA6NMwDbS2TvxNHsnOJ/xgrJLzd/b5QK2PCDk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=F45AkbBq+nM1cSILyyaXWQAMmGoSjA805S5Rht12KY7pVZHchGRQGP7iSkg2h3T0GoQRdqQpijbwAAmK8EO75Q9KfVPkyXHRqasRr2aw9XU6KXnfmqvYpLcrZF8EJWNSyvAZWD0MPMJLi/zZwwzQnESG6Pr0xNQZwX8nPpltsek= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EQf1VSWS; arc=none smtp.client-ip=74.125.231.88 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EQf1VSWS" Received: by mail-oa2-f24.google.com with SMTP id 586e51a60fabf-482620dc91bso786437fac.2 for ; Fri, 18 Sep 2026 11:02:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754571; x=1790359371; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=x3hJ+4xEcO7dPr9U7rXAyf2WO4OSYhm84DLpduOrn64=; b=EQf1VSWSBluP4rDhF3BBtxrpPX4DAY0w5vXXFjmXnSWnI7yH9CdXIuLWww4fIXDq+a aMo72AR4fHgDFqRE7bsvr11p7z4gE/JxZGEGNn3h41NDyN+Pa2GOoY1geEJGaAHss22l Xvsy/FIHbWN9tQTUZ6hP3PJFlbf38ifJDjqgcSZH1G3IdZyiJ2punKeqbDU/+wc4y7S5 3WHtCRxcVjvFzNfi/hU9j2S8WNwQTLirfD1+Y0VhuknJnmERk6rDUP+A2jgCEg7tlPb3 x1IlI+5bkUZhTONDBXxvUQvfCydfcIb6KwwvqSrO6XAqYtIA+O4vYjBCSt9UVQO+sGmv kW/g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754571; x=1790359371; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=x3hJ+4xEcO7dPr9U7rXAyf2WO4OSYhm84DLpduOrn64=; b=ZXq0y4zGFQIfr84rxvvZcZ0F9fq94Dtf2BsCMCWwpQzLu2TPCB+sxDGtEvBs3yZMKT r7syZ6rXSRFwuu21lXk7ppe5S89BjCpkUMxyh2x7LBYaIy7SqlcIbHr8xzqvlQjW7J7C P2iRMFEQncBqjoA3zwh70eOUBFhJbBFuwdN/gQWpYYNYVCepY78Z8G88oHL65rEYE1Yz VRbw5XAZFgLS6ZH1YlsEQkgy1uc9CMXWONzMCvgEl8074JQ6zJTollrYNM3p9ICXzA9D dZYLIroQI/bv2q/hdYEB2LGdriKgZ2oVsVGjcRFYdtM09rEs6cSFzrwLyFflAlNkbANE DVCw== X-Forwarded-Encrypted: i=1; AKwUvBygR/7aRlv8Z2PGaxz2dJhQGN+l87dSWWPVQTXt57FtcmI9YU8GmHzKi6gaStp2itYh97NcemkCVZG13CI=@vger.kernel.org X-Gm-Message-State: AFuF++lxxKtlu/nqH/q/Ju8s+ZtA/+ySP2qfctuvEb4gslwxuzgA2ViG GYjFISBuHpVlj8AdT+qgApkmSd2nANmtJhfP3d4PctiScr0F06jfB6bF X-Gm-Gg: AYBFou23RkuXvqnu7hrTf8b3D55OjpbDjTbTvvmkG97lpKecyFfKY5gY4vi10i5vRFi Cpla76SfymGGNkVEUzQSXOQjV8TG9OtA5hfI8igzi2GLJtsavYurjAKVHPMY+sobltdBY0E/+03 xUbuQKAodHCHrqQh5qCC6RdYz2AXFjVzL9JnlBdwoZQKTJO4bZtRm5OaPf4RRmOTVkag4e18lEx L1SDzP2at3sqzd8PCt6bV7AGf3jkyeJGMI9LwM634bS2DREjq9pB8PK/bsQ5Xh978ahyoiMKZK4 1F9gM9f+y50BrSy6gHG26WsKpd0L1roEgHiXcRKsRCyui+GYNVaMkiGJshb25XyZAK8fcgE67xw Lcc9zGnaei8luieKZfrYpTnWW8k4qy5/R2X360/5XOfI7wPHynC08kDXHzvmOz7y+WxSrTxOHjN BdkbJ5g0Qz3g4yIZSjjeOXvo702OD89QCvSLN8NoSLsepqr/A3ojzSyVLXd4OD+jp/WKSThoqk3 xZgS8Jdbv43yi5hXcpGqoYAxs3q66E= X-Received: by 2002:a05:6870:330d:b0:448:c946:9ae1 with SMTP id 586e51a60fabf-486e66b4448mr3606944fac.18.1789754570354; Fri, 18 Sep 2026 11:02:50 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:6::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-487398590c6sm1626103fac.8.2026.09.18.11.02.49 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:49 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 04/11] mm, swap: support physical swap as a vswap backend Date: Fri, 18 Sep 2026 11:02:34 -0700 Message-ID: <20260918180241.3424851-5-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add physical swap as a backend for the virtual swap layer. When zswap declines a page, the swapout path allocates a physical slot on demand for swap out. Each vswap entry's physical slot is tracked via a pointer-tagged swap_table entry on the physical cluster (an rmap back to the vswap entry). Physical readahead scans a whole offset window and would trip over these rmap slots, so __swap_cache_add_check() now skips swp_tb_is_pointer() entries. Nothing is lost: a backing slot is faulted through its owning vswap entry, never through the physical offset. swapoff reads each vswap entry back through its rmap slot before freeing the physical slot. A failed read leaves the folio not uptodate, so drop it instead of marking it dirty: dirtying would write uninitialised memory out to swap, and the loss is now reported as SIGBUS on the next fault rather than silently returning stale data. If zswap is disabled at the host level or for the folio's cgroup, the folio still gets a physical swap slot, bypassing vswap and mapping the slot directly into the PTEs. In practice the swapfile backend is only reached when zswap declines the folio at swap_writeout() time, most often because the pool is full. The machinery is in place, but its main consumer is not. Zswap writeback to physical swap is added in a following patch. Reclaim of physical slots backing cache-only vswap entries follows it. Suggested-by: Kairui Song Signed-off-by: Nhat Pham --- mm/memory.c | 13 +- mm/page_io.c | 41 ++++-- mm/swap_state.c | 6 +- mm/swap_table.h | 42 +++++- mm/swapfile.c | 352 +++++++++++++++++++++++++++++++++++++++++++----- mm/vmscan.c | 2 +- mm/vswap.h | 198 ++++++++++++++++++++++++++- mm/zswap.c | 2 +- 8 files changed, 593 insertions(+), 63 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index 73ebc59d2b03..e9e05e31c4f8 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -4713,8 +4713,11 @@ static inline bool should_try_to_free_swap(struct sw= ap_info_struct *si, { if (!folio_test_swapcache(folio)) return false; - /* A vswap entry holds no physical slot, so keeping it saves no IO. */ - if (is_vswap_entry(folio->swap)) + /* + * Non-swapfile backends cannot be reused for future swapouts. + * Free the swap slot unless backed by contiguous physical swap. + */ + if (!folio_phys_swap_backed(folio)) return true; /* * Always try to free swap cache for SWP_SYNCHRONOUS_IO devices. Swap @@ -4722,7 +4725,7 @@ static inline bool should_try_to_free_swap(struct swa= p_info_struct *si, * are fast, and meanwhile, swap cache pinning the slot deferring the * release of metadata or fragmentation is a more critical issue. */ - if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) + if (swap_entry_backend_has_flag(si, folio->swap, SWP_SYNCHRONOUS_IO)) return true; if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) || folio_test_mlocked(folio)) @@ -5035,7 +5038,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) swap_update_readahead(folio, vma, vmf->address); if (!folio) { /* Swapin bypasses readahead for SWP_SYNCHRONOUS_IO devices */ - if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) + if (swap_entry_backend_has_flag(si, entry, SWP_SYNCHRONOUS_IO)) folio =3D swapin_sync(entry, GFP_HIGHUSER_MOVABLE, thp_swapin_suitable_orders(vmf) | BIT(0), vmf, NULL, 0); @@ -5200,7 +5203,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) */ exclusive =3D true; } else if (exclusive && folio_test_writeback(folio) && - data_race(si->flags & SWP_STABLE_WRITES)) { + swap_entry_backend_has_flag(si, entry, SWP_STABLE_WRITES)) { /* * This is tricky: not all swap backends support * concurrent page modifications while under writeback. diff --git a/mm/page_io.c b/mm/page_io.c index 632683232622..4858953fc522 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -209,6 +209,7 @@ static void swap_zeromap_folio_clear(struct folio *foli= o) */ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio) { + swp_entry_t phys; int ret =3D 0; =20 if (folio_free_swap(folio)) @@ -241,8 +242,14 @@ int swap_writeout(struct swap_io_ctx *ctx, struct foli= o *folio) */ swap_zeromap_folio_clear(folio); =20 + /* + * For vswap: release stale non-swapfile backings (e.g. ZSWAP from a + * previous swapout cycle) so zswap_store or folio_realloc_swap + * starts on clean slots. Contiguous PHYS backing is preserved for + * reuse by folio_realloc_swap. + */ if (is_vswap_entry(folio->swap)) - folio_release_vswap_backing(folio); + folio_release_non_phys_swap_backing(folio); =20 if (zswap_store(folio)) { count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT); @@ -257,12 +264,15 @@ int swap_writeout(struct swap_io_ctx *ctx, struct fol= io *folio) } rcu_read_unlock(); =20 - /* - * A vswap folio has no physical slot to write to, so keep it dirty. - */ + /* zswap declined it, so it needs a physical slot. */ if (is_vswap_entry(folio->swap)) { - folio_mark_dirty(folio); - return AOP_WRITEPAGE_ACTIVATE; + phys =3D folio_realloc_swap(folio); + if (!phys.val) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writeout(ctx, folio, phys); + return 0; } =20 __swap_writeout(ctx, folio, folio->swap); @@ -475,6 +485,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio) bool workingset =3D folio_test_workingset(folio); unsigned long pflags; bool in_thrashing; + swp_entry_t phys; =20 VM_BUG_ON_FOLIO(!folio_test_swapcache(folio) && !synchronous, folio); VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); @@ -499,14 +510,24 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct = folio *folio) if (zswap_load(folio) !=3D -ENOENT) goto finish; =20 - if (unlikely(swap_is_vswap(sis))) { - folio_unlock(folio); - goto finish; + /* + * Resolve the physical slot to read from. A vswap entry keeps + * folio->swap virtual, so map it to its physical backing; a folio with + * no backing has nothing to read. + */ + if (swap_is_vswap(sis)) { + phys =3D vswap_to_phys(folio->swap); + if (!phys.val) { + folio_unlock(folio); + goto finish; + } + } else { + phys =3D folio->swap; } =20 /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); - swap_add_folio(ctx, folio, folio->swap, READ); + swap_add_folio(ctx, folio, phys, READ); =20 finish: if (workingset) { diff --git a/mm/swap_state.c b/mm/swap_state.c index 54e4fefdab95..657622cfd7f1 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -177,6 +177,9 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, return -ENOENT; ci_off =3D swp_cluster_offset(targ_entry); old_tb =3D __swap_table_get(ci, ci_off); + /* Physical readahead can hit a vswap-backing rmap slot; skip it. */ + if (swp_tb_is_pointer(old_tb)) + return -ENOENT; if (swp_tb_is_folio(old_tb)) return -EEXIST; if (!__swp_tb_get_count(old_tb)) @@ -201,7 +204,8 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, ci_end =3D ci_off + nr; do { old_tb =3D __swap_table_get(ci, ci_off); - if (unlikely(swp_tb_is_folio(old_tb) || + if (unlikely(swp_tb_is_pointer(old_tb) || + swp_tb_is_folio(old_tb) || !__swp_tb_get_count(old_tb) || is_zero !=3D __swap_table_test_zero(ci, ci_off) || (memcg_id && *memcg_id !=3D __swap_cgroup_get(ci, ci_off)))) diff --git a/mm/swap_table.h b/mm/swap_table.h index 3d64d1629882..7d094694ec37 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -4,6 +4,7 @@ =20 #include #include +#include #include "swap.h" =20 extern struct swap_info_struct *vswap_si; @@ -30,7 +31,7 @@ struct swap_memcg_table { * NULL: |---------------- 0 ---------------| - Free slot * Shadow: |SWAP_COUNT|Z|---- SHADOW_VAL ---|1| - Swapped out slot * PFN: |SWAP_COUNT|Z|------ PFN -------|10| - Cached slot - * Pointer: |----------- Pointer ----------|100| - (Unused) + * Pointer: |-------- vswap offset --------|100| - vswap rmap * Bad: |------------- 1 -------------|1000| - Bad slot * * COUNT is `SWP_TB_COUNT_BITS` long, Z is the `SWP_TB_ZERO_FLAG` bit, @@ -51,9 +52,8 @@ struct swap_memcg_table { * - PFN: Swap slot is in use, and cached. Memcg info is recorded on the p= age * struct. * - * - Pointer: Unused yet. `0b100` is reserved for potential pointer usage - * because only the lower three bits can be used as a marker for 8 bytes - * aligned pointers. + * - Pointer: Reverse map from a physical slot to the vswap entry that owns + * it. See the layout below. * * - Bad: Swap slot is reserved, protects swap header or holes on swap dev= ices. */ @@ -393,4 +393,38 @@ static inline unsigned short __swap_cgroup_clear(struc= t swap_cluster_info *ci, } #endif =20 +/* + * Pointer-tagged swap table entry: rmap for vswap-backing physical slots. + * + * On physical clusters, a Pointer-tagged entry stores the offset of the + * vswap entry that owns this physical slot (the reverse map). Only the + * offset is stored; the swap type is implicit (always vswap_si->type, + * since there is exactly one vswap device). + * + * Pointer: |---- vswap offset ----|100| + */ +#define SWP_TB_PTR_MARK_BITS 3 +#define SWP_TB_PTR_MARK 0b100UL +#define SWP_TB_PTR_MARK_MASK ((1UL << SWP_TB_PTR_MARK_BITS) - 1) +#define SWP_RMAP_ENTRY_MASK (~SWP_TB_PTR_MARK_MASK) + +static inline bool swp_tb_is_pointer(unsigned long swp_tb) +{ + return (swp_tb & SWP_TB_PTR_MARK_MASK) =3D=3D SWP_TB_PTR_MARK; +} + +static inline unsigned long swp_entry_to_swp_tb_ptr(swp_entry_t entry) +{ + return (swp_offset(entry) << SWP_TB_PTR_MARK_BITS) | SWP_TB_PTR_MARK; +} + +static inline swp_entry_t swp_tb_ptr_to_swp_entry(unsigned long swp_tb) +{ + unsigned long offset; + + VM_WARN_ON(!swp_tb_is_pointer(swp_tb)); + offset =3D (swp_tb & SWP_RMAP_ENTRY_MASK) >> SWP_TB_PTR_MARK_BITS; + return swp_entry(vswap_si->type, offset); +} + #endif diff --git a/mm/swapfile.c b/mm/swapfile.c index af6a162c4450..13d2ae60fe4c 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -248,7 +248,7 @@ static int __try_to_reclaim_swap(struct swap_info_struc= t *si, need_reclaim =3D ((flags & TTRS_ANYWAY) || ((flags & TTRS_UNMAPPED) && !folio_mapped(folio)) || ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio) && - !is_vswap_entry(folio->swap))); + folio_phys_swap_backed(folio))); if (!need_reclaim || !folio_swapcache_freeable(folio)) goto out_unlock; =20 @@ -962,6 +962,8 @@ static bool __swap_cluster_alloc_entries(struct swap_in= fo_struct *si, { unsigned int order; unsigned long nr_pages; + swp_entry_t vswap_entry, v; + unsigned int i; =20 lockdep_assert_held(&ci->lock); =20 @@ -981,8 +983,26 @@ static bool __swap_cluster_alloc_entries(struct swap_i= nfo_struct *si, order =3D folio_order(folio); nr_pages =3D 1 << order; swap_cluster_assert_empty(ci, ci_off, nr_pages, false); - __swap_cache_add_folio(ci, folio, swp_entry(si->type, - ci_off + cluster_offset(si, ci))); + if (folio_test_swapcache(folio)) { + /* + * Folio already in the swap cache: we are allocating + * physical backing for its vswap entry. Point each + * physical slot back at its own vswap entry + * (Pointer-tagged rmap). + */ + VM_WARN_ON(!is_vswap_entry(folio->swap)); + vswap_entry =3D folio->swap; + for (i =3D 0; i < nr_pages; i++) { + v =3D vswap_entry; + v.val +=3D i; + __swap_table_set(ci, ci_off + i, + swp_entry_to_swp_tb_ptr(v)); + } + } else { + __swap_cache_add_folio(ci, folio, + swp_entry(si->type, + ci_off + cluster_offset(si, ci))); + } } else if (IS_ENABLED(CONFIG_HIBERNATION)) { order =3D 0; nr_pages =3D 1; @@ -1478,12 +1498,14 @@ static bool get_swap_device_info(struct swap_info_s= truct *si) * Fast path try to get swap entries with specified order from current * CPU's swap entry pool (a cluster). */ -static bool swap_alloc_fast(struct folio *folio) +static swp_entry_t swap_alloc_fast(struct folio *folio) { unsigned int order =3D folio_order(folio); struct swap_cluster_info *ci; struct swap_info_struct *si; - unsigned int offset; + unsigned long offset, found =3D 0; + + lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); =20 /* * Once allocated, swap_info_struct will never be completely freed, @@ -1492,25 +1514,28 @@ static bool swap_alloc_fast(struct folio *folio) si =3D this_cpu_read(percpu_swap_cluster.si[order]); offset =3D this_cpu_read(percpu_swap_cluster.offset[order]); if (!si || !offset || !get_swap_device_info(si)) - return false; + return (swp_entry_t){}; =20 ci =3D swap_cluster_lock(si, offset); if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); - alloc_swap_scan_cluster(si, ci, folio, offset); + found =3D alloc_swap_scan_cluster(si, ci, folio, offset); } else if (ci) { swap_cluster_unlock(ci); } =20 put_swap_device(si); - return folio_test_swapcache(folio); + if (found) + return swp_entry(si->type, found); + return (swp_entry_t){}; } =20 /* Rotate the device and switch to a new cluster */ -static void swap_alloc_slow(struct folio *folio) +static swp_entry_t swap_alloc_slow(struct folio *folio) { struct swap_info_struct *si, *next; + unsigned long found; =20 spin_lock(&swap_avail_lock); start_over: @@ -1519,12 +1544,12 @@ static void swap_alloc_slow(struct folio *folio) plist_requeue(&si->avail_list, &swap_avail_head); spin_unlock(&swap_avail_lock); if (get_swap_device_info(si)) { - cluster_alloc_swap_entry(si, folio); + found =3D cluster_alloc_swap_entry(si, folio); put_swap_device(si); - if (folio_test_swapcache(folio)) - return; + if (found) + return swp_entry(si->type, found); if (folio_test_large(folio)) - return; + return (swp_entry_t){}; } =20 spin_lock(&swap_avail_lock); @@ -1542,6 +1567,7 @@ static void swap_alloc_slow(struct folio *folio) goto start_over; } spin_unlock(&swap_avail_lock); + return (swp_entry_t){}; } =20 /* @@ -1900,6 +1926,23 @@ static bool vswap_alloc(struct folio *folio) return false; } =20 +static swp_entry_t folio_alloc_phys_swap(struct folio *folio) +{ + swp_entry_t entry; + +again: + local_lock(&percpu_swap_cluster.lock); + entry =3D swap_alloc_fast(folio); + if (!entry.val) + entry =3D swap_alloc_slow(folio); + local_unlock(&percpu_swap_cluster.lock); + + if (!entry.val && !folio_order(folio) && swap_sync_discard()) + goto again; + + return entry; +} + /** * folio_alloc_swap - allocate swap space for a folio * @folio: folio we want to move to swap @@ -1938,20 +1981,8 @@ int folio_alloc_swap(struct folio *folio) } } =20 - if (vswap_alloc(folio)) - goto done; - -again: - local_lock(&percpu_swap_cluster.lock); - if (!swap_alloc_fast(folio)) - swap_alloc_slow(folio); - local_unlock(&percpu_swap_cluster.lock); - -done: - if (!order && unlikely(!folio_test_swapcache(folio))) { - if (swap_sync_discard()) - goto again; - } + if (!vswap_alloc(folio)) + folio_alloc_phys_swap(folio); =20 /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */ if (unlikely(mem_cgroup_try_charge_swap(folio))) { @@ -1973,6 +2004,11 @@ int folio_alloc_swap(struct folio *folio) return order ? -E2BIG : -ENOMEM; } =20 +static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, + struct swap_cluster_info *pci, + unsigned int ci_start, + unsigned int nr_pages); + /** * __vswap_release_backing - release the backing of a range of vtable slots * @ci: the locked vswap cluster @@ -1988,8 +2024,11 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, unsigned int ci_start, unsigned int nr) { struct swap_cluster_info_dynamic *ci_dyn; + struct swap_info_struct *psi; + unsigned long phys_off_start =3D 0, phys_off_end =3D 0; unsigned int ci_off; unsigned long vt; + swp_entry_t phys_first =3D {}; =20 lockdep_assert_held(&ci->lock); ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); @@ -1997,7 +2036,30 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { vt =3D __vtable_get(ci_dyn, ci_off); =20 + /* The free helper takes one contiguous run within one cluster. */ + if (phys_off_start !=3D phys_off_end && + (vtable_type(vt) !=3D VSWAP_SWAPFILE || + swp_type(vtable_to_phys(vt)) !=3D swp_type(phys_first) || + swp_offset(vtable_to_phys(vt)) !=3D phys_off_end || + phys_off_end % SWAPFILE_CLUSTER =3D=3D 0)) { + psi =3D __swap_entry_to_info(phys_first); + __swap_cluster_free_phys_backing(psi, + __swap_entry_to_cluster(phys_first), + phys_off_start % SWAPFILE_CLUSTER, + phys_off_end - phys_off_start); + phys_off_start =3D phys_off_end =3D 0; + } + switch (vtable_type(vt)) { + case VSWAP_SWAPFILE: + if (phys_off_start =3D=3D phys_off_end) { + phys_first =3D vtable_to_phys(vt); + phys_off_start =3D swp_offset(phys_first); + phys_off_end =3D phys_off_start + 1; + } else { + phys_off_end++; + } + break; case VSWAP_ZSWAP: zswap_entry_free(vtable_to_zswap(vt)); break; @@ -2013,6 +2075,14 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, if (__swap_table_test_zero(ci, ci_off)) __swap_table_clear_zero(ci, ci_off); } + + if (phys_off_start !=3D phys_off_end) { + psi =3D __swap_entry_to_info(phys_first); + __swap_cluster_free_phys_backing(psi, + __swap_entry_to_cluster(phys_first), + phys_off_start % SWAPFILE_CLUSTER, + phys_off_end - phys_off_start); + } } =20 /** @@ -2040,6 +2110,103 @@ void folio_release_vswap_backing(struct folio *foli= o) spin_unlock(&ci->lock); } =20 +/** + * folio_release_non_phys_swap_backing() - Drop a folio's non-physical vsw= ap backing. + * @folio: the folio, occupying a virtual swap entry. + * + * Release the zswap backing recorded for @folio's virtual swap entry, + * leaving the slots empty so the writeout path can install fresh physical + * backing. Does nothing when the entry is already backed by physical + * swapfile slots, which are kept for reuse, or when it has no backing + * beyond the swap cache folio itself. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + */ +void folio_release_non_phys_swap_backing(struct folio *folio) +{ + struct swap_cluster_info *ci; + struct swap_cluster_info_dynamic *ci_dyn; + int nr =3D folio_nr_pages(folio); + unsigned int voff; + unsigned long vt; + enum vswap_backing_type type; + + ci =3D __swap_entry_to_cluster(folio->swap); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + voff =3D swp_cluster_offset(folio->swap); + + spin_lock(&ci->lock); + /* + * A folio's slots cannot mix swapfile with other backends, except + * mid-backend-change, which always starts from slot 0. + */ + vt =3D __vtable_get(ci_dyn, voff); + type =3D vtable_type(vt); + + if (type =3D=3D VSWAP_SWAPFILE || type =3D=3D VSWAP_NONE) { + spin_unlock(&ci->lock); + return; + } + + __vswap_release_backing(ci, voff, nr); + spin_unlock(&ci->lock); +} + +/** + * folio_realloc_swap() - Back a virtual swap folio with a physical swap s= lot. + * @folio: the folio, occupying a virtual swap entry. + * + * Ensure @folio's virtual swap entry has physical (swapfile) backing, + * allocating a physical slot on demand if it has none. If @folio is + * already physically backed, the existing physical entry is returned + * unchanged. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + * Return: The physical swap entry now backing @folio, or an empty entry + * (.val =3D=3D 0) on failure. + */ +swp_entry_t folio_realloc_swap(struct folio *folio) +{ + swp_entry_t vswap_entry =3D folio->swap; + struct swap_cluster_info *ci; + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + swp_entry_t phys_entry =3D {}; + swp_entry_t pe; + int i, nr =3D folio_nr_pages(folio); + + VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); + VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); + VM_WARN_ON(!is_vswap_entry(vswap_entry)); + + phys_entry =3D vswap_to_phys(vswap_entry); + if (phys_entry.val) + return phys_entry; + + phys_entry =3D folio_alloc_phys_swap(folio); + if (!phys_entry.val) + return (swp_entry_t){}; + + voff =3D swp_cluster_offset(vswap_entry); + + ci =3D __swap_entry_to_cluster(vswap_entry); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + spin_lock(&ci->lock); + /* + * Install PHYS backing without freeing any prior contents of the + * vtable. Releasing the old backing is the caller's job. + */ + for (i =3D 0; i < nr; i++) { + pe.val =3D phys_entry.val + i; + __vtable_set(ci_dyn, voff + i, vtable_mk_phys(pe)); + } + spin_unlock(&ci->lock); + + return phys_entry; +} + /** * folio_dup_swap() - Increase swap count of swap entries of a folio. * @folio: folio with swap entries bounded. @@ -2169,6 +2336,47 @@ struct swap_info_struct *get_swap_device(swp_entry_t= entry) return ERR_PTR(-EIO); } =20 +/* + * Common tail for freeing swap slots: device-level accounting + * and cluster list management. + */ +static void __swap_cluster_finish_free(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_start, + unsigned int nr_pages) +{ + lockdep_assert_held(&ci->lock); + swap_range_free(si, cluster_offset(si, ci) + ci_start, nr_pages); + swap_cluster_assert_empty(ci, ci_start, nr_pages, false); + + if (!ci->count) + free_cluster(si, ci); + else + partial_free_cluster(si, ci); +} + +/* + * Free physical swap slots that were backing vswap entries (Pointer-tagge= d). + */ +static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, + struct swap_cluster_info *pci, + unsigned int ci_start, + unsigned int nr_pages) +{ + unsigned int ci_off; + + spin_lock_nested(&pci->lock, SINGLE_DEPTH_NESTING); + VM_WARN_ON(pci->count < nr_pages); + pci->count -=3D nr_pages; + for (ci_off =3D ci_start; ci_off < ci_start + nr_pages; ci_off++) { + __swap_table_set(pci, ci_off, null_to_swp_tb()); + if (!SWAP_TABLE_HAS_ZEROFLAG) + __swap_table_clear_zero(pci, ci_off); + } + __swap_cluster_finish_free(psi, pci, ci_start, nr_pages); + swap_cluster_unlock(pci); +} + /* * Free a set of swap slots after their swap count dropped to zero, or wil= l be * zero after putting the last ref (saves one __swap_cluster_put_entry cal= l). @@ -2180,7 +2388,6 @@ void __swap_cluster_free_entries(struct swap_info_str= uct *si, unsigned long old_tb; unsigned short batch_id =3D 0, id_cur; unsigned int ci_off =3D ci_start, ci_end =3D ci_start + nr_pages; - unsigned long ci_head =3D cluster_offset(si, ci); unsigned int batch_off =3D ci_off; =20 VM_WARN_ON(ci->count < nr_pages); @@ -2218,13 +2425,7 @@ void __swap_cluster_free_entries(struct swap_info_st= ruct *si, if (batch_id) mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); =20 - swap_range_free(si, ci_head + ci_start, nr_pages); - swap_cluster_assert_empty(ci, ci_start, nr_pages, false); - - if (!ci->count) - free_cluster(si, ci); - else - partial_free_cluster(si, ci); + __swap_cluster_finish_free(si, ci, ci_start, nr_pages); } =20 int __swap_count(swp_entry_t entry) @@ -3024,19 +3225,94 @@ static unsigned int find_next_to_unuse(struct swap_= info_struct *si, =20 static int try_to_unuse(unsigned int type) { + struct mempolicy *mpol =3D get_task_policy(current); struct mm_struct *prev_mm; struct mm_struct *mm; struct list_head *p; int retval =3D 0; struct swap_info_struct *si =3D swap_info[type]; struct folio *folio; - swp_entry_t entry; - unsigned int i; + struct swap_io_ctx ctx; + swp_entry_t entry, vswap_entry, phys; + unsigned long swp_tb; + unsigned int i, j; =20 if (!swap_usage_in_pages(si)) goto success; =20 retry: + /* + * Free vswap-backing slots (Pointer-tagged) first. Walk physical + * clusters, read the vswap entry from the rmap, ensure the data + * is in the swap cache, and transition PHYS to FOLIO. Freeing the + * physical backing is enough, so no page table walk is needed. + */ + i =3D 0; + while (vswap_is_enabled() && + swap_usage_in_pages(si) && + !signal_pending(current) && + (i =3D find_next_to_unuse(si, i)) !=3D 0) { + swp_tb =3D swap_table_lookup(swp_entry(si->type, i)); + if (!swp_tb_is_pointer(swp_tb)) + continue; + + vswap_entry =3D swp_tb_ptr_to_swp_entry(swp_tb); + + folio =3D swap_cache_get_folio(vswap_entry); + if (!folio) { + folio =3D swap_cache_alloc_folio(vswap_entry, + GFP_HIGHUSER_MOVABLE, + BIT(0), NULL, mpol, + NO_INTERLEAVE_INDEX); + if (IS_ERR(folio)) { + if (PTR_ERR(folio) =3D=3D -ENOMEM) + return -ENOMEM; + continue; + } + ctx =3D (struct swap_io_ctx){}; + swap_read_folio(&ctx, folio); + swap_read_submit(&ctx); + } + folio_lock(folio); + + if (!folio_matches_swap_entry(folio, vswap_entry)) { + folio_unlock(folio); + folio_put(folio); + continue; + } + + /* + * Re-validate under folio lock: rmap holds folio->swap + j + * for some j in [0, nr_pages). Check folio->swap still maps + * to the contiguous physical run that includes our slot i. + */ + j =3D vswap_entry.val - folio->swap.val; + phys =3D vswap_to_phys(folio->swap); + if (!phys.val || swp_type(phys) !=3D type || + swp_offset(phys) + j !=3D i) { + folio_unlock(folio); + folio_put(folio); + continue; + } + + folio_wait_writeback(folio); + folio_release_vswap_backing(folio); + /* + * Drop a folio whose read failed rather than dirtying + * uninitialised memory; the next fault finds no backing and + * gets SIGBUS. + */ + if (unlikely(!folio_test_uptodate(folio))) + swap_cache_del_folio(folio); + else + folio_mark_dirty(folio); + folio_unlock(folio); + folio_put(folio); + } + + if (!swap_usage_in_pages(si)) + goto success; + retval =3D shmem_unuse(type); if (retval) return retval; @@ -3079,6 +3355,8 @@ static int try_to_unuse(unsigned int type) (i =3D find_next_to_unuse(si, i)) !=3D 0) { =20 entry =3D swp_entry(type, i); + + /* Pointer-tagged rmap slots have no folio; the pre-pass took them. */ folio =3D swap_cache_get_folio(entry); if (!folio) continue; diff --git a/mm/vmscan.c b/mm/vmscan.c index 6cf689817f2e..2c9cc5c20ea9 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1606,7 +1606,7 @@ static unsigned int shrink_folio_list(struct list_hea= d *folio_list, activate_locked: /* Not a candidate for swapping, so reclaim swap space. */ if (folio_test_swapcache(folio) && - ((mem_cgroup_swap_full(folio) && !is_vswap_entry(folio->swap)) || + ((mem_cgroup_swap_full(folio) && folio_phys_swap_backed(folio)) || folio_test_mlocked(folio))) folio_free_swap(folio); VM_BUG_ON_FOLIO(folio_test_active(folio), folio); diff --git a/mm/vswap.h b/mm/vswap.h index 5334c77b6b84..6f952b2591d1 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -20,6 +20,7 @@ struct zswap_entry; enum vswap_backing_type { VSWAP_NONE =3D 0, VSWAP_ZSWAP =3D 1, + VSWAP_SWAPFILE =3D 2, VSWAP_ZERO, VSWAP_FOLIO, }; @@ -50,11 +51,15 @@ static inline bool is_vswap_entry(swp_entry_t entry) * pointer for a virtual swap slot. Tag in low 3 bits, payload in * upper 61 bits. * - * NONE: |----- 0000 ------|000| - no separate backend pointer - * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low bi= ts) + * NONE: |----- 0000 ------|000| - no separate backend pointer + * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low = bits) + * SWAPFILE: |- type:5,off:56 -|010| - on a physical swapfile * - * Pointer payloads (ZSWAP) are stored directly with the tag OR'd into the - * low bits (kernel pointers are >=3D 8-byte aligned, same approach as xar= ray). + * SWAPFILE packs swp_type in the top MAX_SWAPFILES_SHIFT bits and swp_off= set in + * the middle VTABLE_PHYS_OFF_BITS bits, both above the tag, so the type is + * not shifted off the word. Pointer payloads (ZSWAP) are stored directly = with + * the tag OR'd into the low bits (kernel pointers are >=3D 8-byte aligned= , same + * approach as xarray). * * vtable[i] =3D NONE does not by itself mean "free". The swap_table entry * and the per-slot zero flag carry the rest of the state. The full @@ -68,6 +73,8 @@ static inline bool is_vswap_entry(swp_entry_t entry) * NONE | * | set | zero-backed; cached if PFN set * ZSWAP | PFN | clear | folio cached + zswap entry * ZSWAP | shadow / NULL | clear | evicted, only in zswap + * SWAPFILE | PFN | clear | folio cached + physical slot + * SWAPFILE | shadow / NULL | clear | evicted, only on the swapfile * * Locking: a slot's vtable entry (the vswap entry's backend) is only * stable while the caller owns and holds the lock on that entry's swap @@ -93,6 +100,23 @@ static inline enum vswap_backing_type vtable_type(unsig= ned long vt) return vt & VTABLE_TAG_MASK; } =20 +/* swp_offset field width in a physical backend slot; layout described abo= ve. */ +#define VTABLE_PHYS_OFF_BITS (BITS_PER_LONG - VTABLE_TAG_BITS - MAX_SWAPFI= LES_SHIFT) + +static inline unsigned long vtable_mk_phys(swp_entry_t entry) +{ + VM_WARN_ON_ONCE(swp_offset(entry) >> VTABLE_PHYS_OFF_BITS); + return ((unsigned long)swp_type(entry) << (VTABLE_TAG_BITS + VTABLE_PHYS_= OFF_BITS)) | + (swp_offset(entry) << VTABLE_TAG_BITS) | VSWAP_SWAPFILE; +} + +static inline swp_entry_t vtable_to_phys(unsigned long vt) +{ + VM_WARN_ON(vtable_type(vt) !=3D VSWAP_SWAPFILE); + return swp_entry(vt >> (VTABLE_TAG_BITS + VTABLE_PHYS_OFF_BITS), + (vt >> VTABLE_TAG_BITS) & ((1UL << VTABLE_PHYS_OFF_BITS) - 1)); +} + static inline struct zswap_entry *vtable_to_zswap(unsigned long vt) { VM_WARN_ON(vtable_type(vt) !=3D VSWAP_ZSWAP); @@ -134,6 +158,33 @@ vswap_lock_cluster(swp_entry_t entry, unsigned int *vo= ff) return container_of(ci, struct swap_cluster_info_dynamic, ci); } =20 +/** + * vswap_to_phys - resolve a vswap entry's physical swap backing + * @entry: the virtual swap entry + * + * Context: takes and drops the vswap cluster lock internally. + * Return: the backing physical swp_entry_t, or the null entry (.val =3D= =3D 0) + * when @entry has no physical backing (NONE/ZSWAP/ZERO). + */ +static inline swp_entry_t vswap_to_phys(swp_entry_t entry) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + unsigned long vt; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) + return (swp_entry_t){}; + + vt =3D __vtable_get(ci_dyn, voff); + swap_cluster_unlock(&ci_dyn->ci); + + if (vtable_type(vt) !=3D VSWAP_SWAPFILE) + return (swp_entry_t){}; + + return vtable_to_phys(vt); +} + void __vswap_release_backing(struct swap_cluster_info *ci, unsigned int ci_start, unsigned int nr); =20 @@ -184,6 +235,104 @@ static inline struct zswap_entry *vswap_zswap_load(sw= p_entry_t entry) } =20 void folio_release_vswap_backing(struct folio *folio); +swp_entry_t folio_realloc_swap(struct folio *folio); +void folio_release_non_phys_swap_backing(struct folio *folio); + +/* + * Walk nr vtable slots starting at voff in ci_dyn. Returns the prefix + * length of slots sharing one effective backing type. For SWAPFILE, + * the prefix is also restricted to contiguous offsets in the same + * swapfile. + * + * Effective type per slot: + * vtable=3DNONE + zero flag set -> VSWAP_ZERO + * vtable=3DNONE + swap_table PFN tag -> VSWAP_FOLIO + * vtable=3DNONE + neither -> VSWAP_NONE + * vtable=3DSWAPFILE -> VSWAP_SWAPFILE + * vtable=3DZSWAP -> VSWAP_ZSWAP + * + * *typep returns the effective type of slot 0. Caller holds + * ci_dyn->ci.lock. + */ +static inline int __vswap_check_backing(struct swap_cluster_info_dynamic *= ci_dyn, + unsigned int voff, int nr, + enum vswap_backing_type *typep) +{ + enum vswap_backing_type first_type =3D VSWAP_NONE; + enum vswap_backing_type slot_type; + swp_entry_t first_phys =3D {}; + unsigned long vt, swap_tb; + int i; + + lockdep_assert_held(&ci_dyn->ci.lock); + + for (i =3D 0; i < nr; i++) { + vt =3D __vtable_get(ci_dyn, voff + i); + if (vtable_type(vt) =3D=3D VSWAP_NONE) { + swap_tb =3D __swap_table_get(&ci_dyn->ci, voff + i); + if (__swap_table_test_zero(&ci_dyn->ci, voff + i)) + slot_type =3D VSWAP_ZERO; + else if (swp_tb_is_folio(swap_tb)) + slot_type =3D VSWAP_FOLIO; + else + slot_type =3D VSWAP_NONE; + } else { + slot_type =3D vtable_type(vt); + } + + if (!i) { + first_type =3D slot_type; + if (first_type =3D=3D VSWAP_SWAPFILE) + first_phys =3D vtable_to_phys(vt); + } else if (slot_type !=3D first_type) { + break; + } else if (first_type =3D=3D VSWAP_SWAPFILE && + vtable_to_phys(vt).val !=3D first_phys.val + i) { + break; + } + } + + if (typep) + *typep =3D first_type; + return i; +} + +static inline int vswap_check_backing(swp_entry_t entry, int nr, + enum vswap_backing_type *typep) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + int ret; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) { + if (typep) + *typep =3D VSWAP_NONE; + return 0; + } + ret =3D __vswap_check_backing(ci_dyn, voff, nr, typep); + swap_cluster_unlock(&ci_dyn->ci); + return ret; +} + +/** + * folio_phys_swap_backed - test whether a folio is backed by a contiguous + * range of physical swap slots. + * @folio: a swap-cache resident folio + * + * Return: %true if @folio->swap is not a vswap entry, or if these vswap + * entries are backed by a contiguous range of physical slots. + */ +static inline bool folio_phys_swap_backed(struct folio *folio) +{ + swp_entry_t entry =3D folio->swap; + int nr =3D folio_nr_pages(folio); + enum vswap_backing_type type; + + return !is_vswap_entry(entry) || + (vswap_check_backing(entry, nr, &type) =3D=3D nr && + type =3D=3D VSWAP_SWAPFILE); +} =20 static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn, gfp_t gfp) @@ -214,6 +363,47 @@ static inline bool is_vswap_entry(swp_entry_t entry) return false; } =20 +static inline swp_entry_t vswap_to_phys(swp_entry_t entry) +{ + return (swp_entry_t){}; +} + +static inline bool folio_phys_swap_backed(struct folio *folio) +{ + return true; +} + #endif /* CONFIG_SWAP */ =20 +/* + * Test a per-backend swap flag (SWP_SYNCHRONOUS_IO, SWP_STABLE_WRITES, ..= .) + * for @entry. For a vswap entry the property belongs to the current + * physical backing rather than vswap_si itself; resolve to the backing + * and test there. Returns false for zswap/zero/unbacked vswap entries + * as they don't have a backing bdev. + */ +static inline bool swap_entry_backend_has_flag(struct swap_info_struct *si, + swp_entry_t entry, + unsigned long flag) +{ + struct swap_info_struct *phys_si; + swp_entry_t phys; + bool has_flag; + + if (!swap_is_vswap(si)) + return data_race(si->flags & flag); + + phys =3D vswap_to_phys(entry); + if (!phys.val) + return false; + + phys_si =3D get_swap_device(phys); + if (IS_ERR_OR_NULL(phys_si)) + return false; + + has_flag =3D data_race(phys_si->flags & flag); + put_swap_device(phys_si); + return has_flag; +} + #endif /* _MM_VSWAP_H */ diff --git a/mm/zswap.c b/mm/zswap.c index c2430dbfc653..bbfaeac00355 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1620,7 +1620,7 @@ bool zswap_store(struct folio *folio) */ if (is_vswap_entry(swp)) { if (index > 0) - folio_release_vswap_backing(folio); + folio_release_non_phys_swap_backing(folio); } else { zswap_invalidate(swp_type(swp), swp_offset(swp), nr_pages); } --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9D58851C061 for ; Fri, 18 Sep 2026 18:02:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754576; cv=none; b=rlJXX1ctqZhI1mpsFr9WKclT3INBHWV6l5MFAu6fLzRw3dvnQMCmjJegd8mSCFka6DwNc0kNBq+nAveGGJaqTWPPAKAPZdaC0Zla7fkKX6wWQlxTcMcBRP8hnWLNBrm+BKsRK5DISJK1OFCzzTZ/oLOSu7jvv2OfsdGWgXKxjR0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754576; c=relaxed/simple; bh=RO9jvV3V9zXQK+hpLOst1hPa/YVS04E+RkiRewZGdwM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=S8/RPOXrZmmscxjLW4xe84WsXJD8/kJiCBI+Oc4ZncU3qRtkX/LXB0QKqPl+Mefu+W0IAefsU6xJhXmLKUOE9/Pa8ca2ZA23KP2r1lyZ68dp+KGjugQ3FSnJeOZmEdnGkeu7JJWXuMgf5Jh1dt9aZ4DKztwMx2Aq3OaM2egkwxs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=DWdgEdz/; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="DWdgEdz/" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-466ccab774cso659976fac.1 for ; Fri, 18 Sep 2026 11:02:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754572; x=1790359372; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=im6K772PTWgmT32pMTVGaO8Gd5j7dB48NMbCPlbIlg4=; b=DWdgEdz/BUB875vJORxYqGWY3TkyglLZYuuR7Z5GNVQVm5du70FyAfIOVk6GUiW8xe BFm929V/GFhF/9TPAdxkScjQHcDVG4xQx/LDIxljcN0mJnXWz1q5dVQL8/R9SSHAYMJq F2Bq6c14ji3YNKv6ODmBXmCFkSVMTQDhbECMkaWni1TbIkQuDSMWa6FUYX0ttDcgYp5D iAlS/k4VnFq2NhGZwxGd6ikyid176zOe3zI0XA5EmTdCLmlQjtoLU1FRVPYZGIYFhpCn 0khAwgWmtTUU0+gYcIUbSb8kvGLSwVMBLElo+H+SYqFY3vop9/oJpLv2MSmLM3gQ0grO DHFw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754572; x=1790359372; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=im6K772PTWgmT32pMTVGaO8Gd5j7dB48NMbCPlbIlg4=; b=USLythcIn6Kr/82/ucAv5s9Io5sLTFlJNrf8LZL3flbIoe/fWfppG++KjcSXOGLd/p /wolcXDXvErf91phrGH5R0175/0Ibd3wI9uNA+rcWewryqB5RaKFbjzTX9eRqVT9LT9c O10ciKyhd1/xv/rJ2GbKKgYbIqdH6aAlGwiWHWB0JKIhpzmQyXWo60ajLQh+bXbrv2qP 8s78aOT2NB26sbFcDDb+32hng22FXAF4mFiUKURAxLzuqNU92GriozmYb8ZKugfpzK8+ BL/rzCtaBy0ea7wTzAHkuVjKN4nqxarewSKi1kcKfUl1oCCLORTx453YsABf603GIAOr 2l6A== X-Forwarded-Encrypted: i=1; AKwUvBz15du+zY2rAi2mDTNdVCzjsXbxK2EEOMHhabak+k5Tb+m923NkQiRyoA129ioz2mLhmhvmYPK1ETeU4WE=@vger.kernel.org X-Gm-Message-State: AFuF++nW1mpZiylRHzVQkaTojiXw6JWtaWAd5TVAdJNM8riAnspCNckY 0xx9hX2jUJPEAvN51kL5pLHnYUB6jC+FzX5cnbcxbKrNcUrONNi/QcFE X-Gm-Gg: AYBFou32TEzeXNnP74BTNa5hUOeS4OirQZj7qEU3Yb9tN3hh8mC1pWBJVmrLlF7fwt4 e1ujRrx2pC/NhtDsJ7i1WhspZck9nAQdIoKaRoPM4T9zFFGThBmK/XvfVJ7Z9NfC7kNBXSsinmD oSB8eM7jQESGSxQCwtngAXnZexEm0pQlBfguUkOQVE1ijF/g181d4gcCoY1N02wa9aMiDl1AFJP ND5k9TxUrE4wLHk3apqtuBvfZZZUJsq6Ivr/Jur1hc3SP17CXNOUvvd9bUbpTrf/ZQvGC0ZxJlp rNpYYSCAWriAEGNS4qrUa72Biul8Gor/l1WHdKBwzd+l+u/QLET3nM4WctZUmWC/mHCOfpMEiWx t1gkg+ZocM9AajCTxJZ19I+fi/+z2ZTh3EHQYlElk4vWNO6HBO+NrePBQvRnJTWegnFmsNa3A/z hxi9xIme8dQViPaNGaEE+NKVyZ2/PAZp4I8r2TIzyNFVVL3v1Y3ms9Go2YVY06gevbfa5nGvoFo 8F7VC5MMIQ8bJKsPAlsuQ== X-Received: by 2002:a05:6870:b402:b0:47d:a38a:6218 with SMTP id 586e51a60fabf-486e6ac67d0mr3577258fac.25.1789754572094; Fri, 18 Sep 2026 11:02:52 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:40::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-48739852afasm1601298fac.7.2026.09.18.11.02.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:51 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 05/11] mm, swap: enable THP swapin for vswap entries Date: Fri, 18 Sep 2026 11:02:35 -0700 Message-ID: <20260918180241.3424851-6-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Swap a large anon folio back in as a unit when its vswap entries share a contiguous run of physical swap slots on a synchronous IO device, instead of always falling back to order-0 faults. A zswap-backed or mixed-backing batch is still refused, and the fault retries at a smaller order. Signed-off-by: Nhat Pham --- mm/memory.c | 5 +++-- mm/swap_state.c | 17 +++++++++++++---- mm/zswap.c | 6 +++++- 3 files changed, 21 insertions(+), 7 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index e9e05e31c4f8..e052de3b4461 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -4882,9 +4882,10 @@ static unsigned long thp_swapin_suitable_orders(stru= ct vm_fault *vmf) * lack handling for such cases, so fallback to swapping in order-0 * folio. * - * THP swapin for vswap is not supported yet either. + * Vswap entries are checked later, under the cluster lock in + * __swap_cache_add_check(). */ - if (is_vswap_entry(entry) || !zswap_never_enabled()) + if (!is_vswap_entry(entry) && !zswap_never_enabled()) return 0; =20 /* diff --git a/mm/swap_state.c b/mm/swap_state.c index 657622cfd7f1..2107d05ae8d5 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -165,6 +165,9 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, unsigned int ci_off, ci_end; unsigned long old_tb; bool is_zero; + struct swap_cluster_info_dynamic *ci_dyn; + enum vswap_backing_type type; + int ret; =20 lockdep_assert_held(&ci->lock); =20 @@ -193,11 +196,17 @@ static int __swap_cache_add_check(struct swap_cluster= _info *ci, return 0; =20 /* - * Reject a vswap batch so swap_cache_alloc_folio falls back to - * order 0. + * For a vswap entry batch, reject if the backing is not THP-amenable + * (e.g. uniformly ZSWAP, or mixed). The order-fallback loop in + * swap_cache_alloc_folio will retry with a smaller order on -EBUSY. */ - if (is_vswap_entry(targ_entry)) - return -EBUSY; + if (is_vswap_entry(targ_entry)) { + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + ret =3D __vswap_check_backing(ci_dyn, round_down(ci_off, nr), + nr, &type); + if (ret !=3D nr || type =3D=3D VSWAP_ZSWAP) + return -EBUSY; + } =20 is_zero =3D __swap_table_test_zero(ci, ci_off); ci_off =3D round_down(ci_off, nr); diff --git a/mm/zswap.c b/mm/zswap.c index bbfaeac00355..3e1aa295f9dd 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1688,9 +1688,13 @@ int zswap_load(struct folio *folio) * range on the backing device, so scan the range rather than rejecting * it outright. The caller has pinned every slot, so zswap cannot start * a store or a writeback into the range while we look. + * + * A vswap batch is checked when the folio enters the swap cache, and + * its backing cannot change after that. */ if (folio_test_large(folio)) { - if (WARN_ON_ONCE(zswap_is_present(swp, + if (WARN_ON_ONCE(!swap_is_vswap(si) && + zswap_is_present(swp, folio_nr_pages(folio)))) { folio_unlock(folio); return -EIO; --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f13.google.com (mail-oi2-f13.google.com [74.125.231.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B790051C345 for ; Fri, 18 Sep 2026 18:02:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754579; cv=none; b=EUWgOd+igFNUpmKIWE30F7wG4x1wrOLtsXev3XAZM/jhqWK7KtQ8ZRTDNaPSI4BqCeuKdlvbbiHnMsUu7MitWfGwnZrrgooYollOmVcUN3xTgTE1IvaF4U2zoOLVGMkXQkVmGfDixnv9zh3h/rfkjeEILlbtj0+5DzGyplJkZLQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754579; c=relaxed/simple; bh=4gZxRi8s9AvaXJeqTNZNQaZREkrt6u2IyZlzTPcet/4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rvQ9f/FhyendpRvYfWVzRGjZkIw4OA4ZeUr2rPHDed7DFeJMCfv1I/NXNs+2PIiI2ExUT1NDxdSwJ0fQb6/kUb84w6/UGn5FsIglGw1lW5yozX0xFzWh13hGTT1YeSCDGSGF1sBQkjV1NpyU7BQs6zl/zq8X5RMVEFofo/4KAPo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NHRK1fW7; arc=none smtp.client-ip=74.125.231.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NHRK1fW7" Received: by mail-oi2-f13.google.com with SMTP id 5614622812f47-4c2c08ff3f8so783188b6e.1 for ; Fri, 18 Sep 2026 11:02:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754574; x=1790359374; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wHkgIfdkKYj8pIFBCxdLHIsV2Dv3mqE/3ti2RlbB1g0=; b=NHRK1fW7E+laUZ4ltR98byxU3hHbcBIvad1I4rskbHrwoVHM1nH/smR5kY+lATMg86 CINiimukzZAMi7eQSOApt6+wamZBquiwL0Mt9D9LVH9t3a3s6it5whBxH1eBBSUwnlTJ O8inhLuEXzDoR99Ques5yrfCjEyoFGFjIPhYKycxAHuAm8XhdKoMN4hV6Uw09UKBY87n coBFGD7HXErx1Psl4w49G5dSnlhTYSMN0yYaUoxH93AVqV1/WZdQ3i9Co/IP2qchsahY DD0i18zm0iJlXvFPxVZuFLAVrrU8aHyd5Yb9SzTTnWJl9T7HdnQ8ZiPq1O5ySon6dYLL WMvA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754574; x=1790359374; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=wHkgIfdkKYj8pIFBCxdLHIsV2Dv3mqE/3ti2RlbB1g0=; b=jC3gd2kZv8HVaRx1BIPdXzPPBv8N4d0rmKXglbQnoTPj31ckw92f/MunuBL6ZEXJgh jYPGLfpzjk1DfddkeTB2BRNGtcsQqOw2Vj+fizZpdUYSUXooHfLBnxqSVwbpRwZFbZgX EEesB9C2s5WjpIDyCABzjxVtuzepqJ9aJLbOpugg64zhW1neEUv0SIpiuCyqbIPWWsA7 NQ+Qszcm+97VGSNTzFBTIrWxTws+BoFUWeiWaYNU99XT+eqWQyqdoQlIW/6QiYPnRaiQ 3UIzgBFmJEsLxCNEscs30BpDneOhM9QjdlQwHid3eNO0r376iH+/Y0/Vg4WuGr0ICdz6 +YWA== X-Forwarded-Encrypted: i=1; AKwUvBzTrkseyuoQRgKu1vXuRfC+G3/fKO9MzXNlQuPUacxysMtesdguGmDdscHhWkb5jhfSVHagvY83qB2eGb4=@vger.kernel.org X-Gm-Message-State: AFuF++leMITlNFxO3LpNp83EHNYS6ub1Ou01mEfpcp8Rw80bxyZQpQrK VAhsnsTN0fMCpRhIXzVvXxoRXeXq6uQxeS0q+kHClB+mTkHawycu4jgR X-Gm-Gg: AYBFou140SkOujP9xFyF6nswQwz2GjIjsDiWLM0F2jolD8yEs9Mcre8BnjJNfRIFOg4 WrThCaedQCxZNmOqTsjCNmedo4c7cbX1j7s7vZ96w4Y4h0YeXekDGDLsWvuupLLvzVB3Dxihx5+ kaKGwSKYH/29hTcfmNSabbUFH7XzzyZNki5WY+3Yw99dgBTkLVj5HdFVjuVzJctU/dUjm4/9ZHt ZPBo412oOF4CJKZ9sA/3q2MLHC1P5vo8s6CKmowbCJtfYEGJz0wo0iAB1uZgu0Y8R+M4nvuqG6c kc5/MLfWXW+/Eo7OglVkPk5wtsHGHAiKn63CjKumgGMOnPnzimtjmAxgSrr5KLLlJXOtYuPQ09j 1F9ojvcpuADh3yhOu//33R1aSWhxDQ19kdxszrbOG8swROpoO9r54t+VobtBddZfwgquvUxJWbW vj6tjFKAff5IUSR2CWcYIpi+UNneuo5113VqkBWAhNkZISF4BoD65IZ7q7MEJqfHd6qkL7uwQIf c253fWRnZhw1iN55V30 X-Received: by 2002:a05:6808:3c47:b0:4b9:a829:f016 with SMTP id 5614622812f47-4ccf74b986fmr3815644b6e.34.1789754573817; Fri, 18 Sep 2026 11:02:53 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:5::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-8107dd8bd0esm154912a34.5.2026.09.18.11.02.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:53 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 06/11] mm, swap: write back vswap zswap entries to physical swap Date: Fri, 18 Sep 2026 11:02:36 -0700 Message-ID: <20260918180241.3424851-7-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add support for writing back zswap-backed vswap entries to physical swap. The mechanism mirrors the existing zswap writeback path, except the backing physical slot is allocated on demand at writeback time rather than already being pinned by the PTE. The zswap shrinker no longer skips vswap entries, unless we are out of physical swap space. Signed-off-by: Nhat Pham --- mm/zswap.c | 67 +++++++++++++++++++++++++++++++++++++----------------- 1 file changed, 46 insertions(+), 21 deletions(-) diff --git a/mm/zswap.c b/mm/zswap.c index 3e1aa295f9dd..56315298c291 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1025,12 +1025,13 @@ static bool zswap_decompress(struct zswap_entry *en= try, struct folio *folio) static int zswap_writeback_entry(struct zswap_entry *entry, swp_entry_t swpentry) { - struct xarray *tree; pgoff_t offset =3D swp_offset(swpentry); struct folio *folio; struct mempolicy *mpol; struct swap_info_struct *si; struct swap_io_ctx ctx =3D {}; + swp_entry_t phys =3D {}; + bool is_vswap; int ret =3D 0; =20 /* try to allocate swap cache folio */ @@ -1038,12 +1039,7 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, if (IS_ERR_OR_NULL(si)) return -ENOENT; =20 - /* Vswap entries have no physical backing to write to. */ - if (swap_is_vswap(si)) { - put_swap_device(si); - return -EINVAL; - } - + is_vswap =3D swap_is_vswap(si); mpol =3D get_task_policy(current); folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, NO_INTERLEAVE_INDEX); @@ -1062,24 +1058,44 @@ static int zswap_writeback_entry(struct zswap_entry= *entry, /* * folio is locked, and the swapcache is now secured against * concurrent swapping to and from the slot, and concurrent - * swapoff so we can safely dereference the zswap tree here. + * swapoff so we can safely dereference the zswap tree (or vswap + * vtable) here. * Verify that the swap entry hasn't been invalidated and recycled * behind our backs, to avoid overwriting a new swap folio with * old compressed data. Only when this is successful can the entry * be dereferenced. */ - tree =3D swap_zswap_tree(swpentry); - if (entry !=3D xa_load(tree, offset)) { + if (entry !=3D zswap_entry_load(swpentry)) { ret =3D -ENOMEM; goto out; } =20 + if (is_vswap) { + /* + * Allocate physical backing before decompress so a failure + * wastes no work. + */ + phys =3D folio_realloc_swap(folio); + if (!phys.val) { + ret =3D -ENOMEM; + goto out; + } + } + if (!zswap_decompress(entry, folio)) { ret =3D -EIO; + /* + * The phys allocation above took the entry out of the vtable. + * Restore the zswap entry to the vtable, which also frees the + * allocated physical swap space. + */ + if (is_vswap) + vswap_zswap_store(swpentry, entry); goto out; } =20 - xa_erase(tree, offset); + if (!is_vswap) + xa_erase(swap_zswap_tree(swpentry), offset); =20 count_vm_event(ZSWPWB); if (entry->objcg) @@ -1094,7 +1110,10 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, folio_set_reclaim(folio); =20 /* start writeback */ - __swap_writeout(&ctx, folio, folio->swap); + if (is_vswap) + __swap_writeout(&ctx, folio, phys); + else + __swap_writeout(&ctx, folio, folio->swap); swap_write_submit(&ctx); =20 out: @@ -1109,6 +1128,15 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, /********************************* * shrinker functions **********************************/ +/* + * vswap zswap entries get a physical slot allocated on demand at writeback + * time. Skip the shrinker when none is available. + */ +static bool zswap_writeback_possible(void) +{ + return !vswap_is_enabled() || get_nr_swap_pages() > 0; +} + /* * The dynamic shrinker is modulated by the following factors: * @@ -1246,7 +1274,7 @@ static unsigned long zswap_shrinker_count(struct shri= nker *shrinker, if (!zswap_shrinker_enabled || !mem_cgroup_zswap_writeback_enabled(memcg)) return 0; =20 - if (vswap_is_enabled()) + if (!zswap_writeback_possible()) return 0; =20 /* @@ -1332,7 +1360,8 @@ static struct shrinker *zswap_alloc_shrinker(void) * were scanned but none could be written back, or -ENOENT if @memcg has * writeback disabled, is a zombie cgroup, or has empty zswap LRUs. * - * Also returns -ENOENT when vswap is enabled. + * Also returns -ENOENT when vswap is enabled and there is no physical + * swap to write back to. */ static int shrink_memcg(struct mem_cgroup *memcg) { @@ -1341,7 +1370,7 @@ static int shrink_memcg(struct mem_cgroup *memcg) if (!mem_cgroup_zswap_writeback_enabled(memcg)) return -ENOENT; =20 - if (vswap_is_enabled()) + if (!zswap_writeback_possible()) return -ENOENT; =20 /* @@ -1372,11 +1401,7 @@ static void shrink_worker(struct work_struct *w) int ret, failures =3D 0, attempts =3D 0; unsigned long thr; =20 - /* - * When vswap is enabled, zswap entries are almost all vswap backed, - * with no slot to write back to. - */ - if (vswap_is_enabled()) + if (!zswap_writeback_possible()) return; =20 /* Reclaim down to the accept threshold */ @@ -1457,7 +1482,7 @@ static void shrink_worker(struct work_struct *w) break; resched: cond_resched(); - } while (zswap_total_pages() > thr); + } while (zswap_total_pages() > thr && zswap_writeback_possible()); } =20 /********************************* --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oo2-f40.google.com (mail-oo2-f40.google.com [74.125.231.168]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F2CA551CF4A for ; Fri, 18 Sep 2026 18:02:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.168 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754580; cv=none; b=IDM8lyCtvhhq2XAvr6MZTuypoI1FNcf3kRJ/X+mbyTA48c0L9Xh1wVU5sMFhBZzaOAjOV9TCTCkrgH0Fwsx4GEUOlucbLdQYDHYVOdDR2yNPI/VTGbK/AA8Q5mWbaF7jN6f7EXymkylQm5jVVNFM6Ph9eFbZsNytHySbyjIDRfk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754580; c=relaxed/simple; bh=IJvK4Vvy0VCt+V23WX0pSPCnBeAybxb27qXqZn6DYsc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eLb+iNr+/QWxGuA0RqQRr4b0lK+pXWI4Qjo4hzkSkS8SY2PThmHwyvG1yp7rcHvpN4W9X3x9eicwQNgxQ8M08dPEtG4iU6nc/3gEoSJGERyVe7iUAjtu6LOmZCMWFRMP1d+wadLust+3XX6jq0be55faL4K1ZnFp2viZOiFuXKA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Uryn2Xnt; arc=none smtp.client-ip=74.125.231.168 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Uryn2Xnt" Received: by mail-oo2-f40.google.com with SMTP id 006d021491bc7-6c72bd8a017so662163eaf.3 for ; Fri, 18 Sep 2026 11:02:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754575; x=1790359375; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=iKzNgqV0ZA0AZzF6ydZQcV12V6+MZv1Kiypu+ZDENgo=; b=Uryn2XntMQmAIkaYR7Cii24/foDihxsCV5gSYUVWt7q9yJUigTcHjWzqV0eE6zuSlW o/czTtTnPbKY9aZ/22ANz/VJZrFQJod9F3llZambAKuNyXs7uli1FTklM/D8vovyJvqx BD8fbMnRXV9Z7udumeUCOCCgMIbr7fsY6V0muOlvn3TYBzPLrq66PaUpP2iDjyOO4ejI bvQ2/iEEK8useDyRF/VNKhJoYch5TCKzTh+3OI5ew5g2kIndT955Hvw0ybDolbUoW+oU 2Xsjh7RDi88oqELySswrvh4vALVVRBUj54sbwS/dozhRGf7wrZsUWBMmDUcdd9fpmL+e WnRA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754575; x=1790359375; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=iKzNgqV0ZA0AZzF6ydZQcV12V6+MZv1Kiypu+ZDENgo=; b=naEmI/HL1dJnB4tmHtf/b25VjEt7s0joQ1f4TK7efrF2rzN6LJ278y4FR3nz7jZeWA FelvmIu87YOnedO6oMQolByekVU92kPDcbq/Z8y2E8Pa8u1ZCzODn6LQOt+y9yTjj4UU klaYX7fjdGpU9p0vxuSlRncDbtmcnO8olWO1FQUC8DIDI+n7EyfjhBcqGKFcLT4L0xOl 3/WhInHqmLqEo7n2HT2ZnoEU/9RAztcFZ7VOaupU/WzU4Oeqjv0t7z3NpGKg7r/eK3D2 PkbnX7gPUiP4XR6isqsETGF70laIeX5wS3QgBJqxsxmWqYYdifA6fUC0FLLNxrpiKs3i 1++w== X-Forwarded-Encrypted: i=1; AKwUvBx30pHWtalRlEX6P2BEVPGWTdy97FJkqlpDurf0Mr8fantfRinIfihVBCmp+27zlq8wG1q2RnXZQ0jZhC0=@vger.kernel.org X-Gm-Message-State: AFuF++npFE12opzBa7S2p7fD3yX7v1021YSEjSwJC4XH5fenb+YXJlgK 6X5+nG+nZcGQUNy7kxSl+aMkvF1NrVaKUxTjni3/YH86736QN+DLpd38 X-Gm-Gg: AYBFou09l/HDTOQWgwY1Jf7TzO4cTxWWseH3DrNbOvEwl2UZhOmybg9vnDRe5i09/bW 0ICo9W/BAnn4ucAmmqYPYEa8BiCIRmchhT51yuoeDQ1I8IvFs3hQfg3nuXcFBOwPt7NOOV/TRtJ 42LN5FBBqezGYV3Xdth+FN+uCpNT79Ie3eofzsf5iJnomDXTKLeN1ztYdtc9K25RoL4nYjRJsIL +5sR/63dXX3NaeCkbtcqwSnqS5sDHZMsB+1E97VkVTDaHeUlXdgIx9/i4aHRHZs3lr+wHIV/+aM WwjXOGGB2fXCrPCjiXVb4G1NOPqWPOoLM8asBH7OSdnJ53A7DEAJxPA1ZUYFj3BAZwlzzKrVrAy FgwJ73RdbUdgAAs/KMYRWQFVidjRcIEp07H08GxmUwoYi5tiuF+RuhepFdN0kUydoqDbKBJqWc8 KlvEbM7i3Qi3otBbSKXj6ViXQ/JD3Boo2u3uYO4qboVeK2KeE5QRDepUeMpwSbzbbY0kJCX9bkr J85ODf+VJPMkyyDcpCpyw== X-Received: by 2002:a05:6820:190d:b0:6cd:3fbd:1d5f with SMTP id 006d021491bc7-6cd3fbd2951mr340212eaf.73.1789754575416; Fri, 18 Sep 2026 11:02:55 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:12::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6cd34b35d71sm522300eaf.7.2026.09.18.11.02.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:54 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 07/11] mm, swap: reclaim physical slots backing cache-only vswap entries Date: Fri, 18 Sep 2026 11:02:37 -0700 Message-ID: <20260918180241.3424851-8-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When a vswap entry's swap_count drops to 0 while its folio is still in the swap cache, the entry is cache-only and its physical slot is redundant. swap_put_entries_direct() tries to reclaim it at put time, but folio_put_swap() does not, and even the direct route can come back empty: the folio may be under the writeback that allocated the slot, its trylock may be lost, it may still be mapped, or swap_only_has_cache() may be false on a partial run. Nothing retried afterwards, so the slot stayed pinned. Reclaim such slots from the physical reclaim scanner, once swap is more than half used (vm_swap_full()), to free physical capacity for new allocations. Reclaim goes through folio_free_swap(), which ends in folio_set_dirty(), so a reclaimed slot costs a re-swapout if the folio is dropped again. Signed-off-by: Nhat Pham --- mm/swap_table.h | 12 +++-- mm/swapfile.c | 117 ++++++++++++++++++++++++++++++++++++++++++++++++ mm/vswap.h | 31 +++++++++++++ 3 files changed, 156 insertions(+), 4 deletions(-) diff --git a/mm/swap_table.h b/mm/swap_table.h index 7d094694ec37..79f06642a553 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -31,7 +31,7 @@ struct swap_memcg_table { * NULL: |---------------- 0 ---------------| - Free slot * Shadow: |SWAP_COUNT|Z|---- SHADOW_VAL ---|1| - Swapped out slot * PFN: |SWAP_COUNT|Z|------ PFN -------|10| - Cached slot - * Pointer: |-------- vswap offset --------|100| - vswap rmap + * Pointer: |C|------- vswap offset -------|100| - vswap rmap * Bad: |------------- 1 -------------|1000| - Bad slot * * COUNT is `SWP_TB_COUNT_BITS` long, Z is the `SWP_TB_ZERO_FLAG` bit, @@ -399,14 +399,18 @@ static inline unsigned short __swap_cgroup_clear(stru= ct swap_cluster_info *ci, * On physical clusters, a Pointer-tagged entry stores the offset of the * vswap entry that owns this physical slot (the reverse map). Only the * offset is stored; the swap type is implicit (always vswap_si->type, - * since there is exactly one vswap device). + * since there is exactly one vswap device). The top bit is reserved as + * a cache-only flag, set when vswap swap_count drops to 0 but the folio + * is still in swap cache. * - * Pointer: |---- vswap offset ----|100| + * Pointer: |C|---- vswap offset ----|100| + * C =3D SWP_RMAP_CACHE_ONLY (the top bit) */ #define SWP_TB_PTR_MARK_BITS 3 #define SWP_TB_PTR_MARK 0b100UL #define SWP_TB_PTR_MARK_MASK ((1UL << SWP_TB_PTR_MARK_BITS) - 1) -#define SWP_RMAP_ENTRY_MASK (~SWP_TB_PTR_MARK_MASK) +#define SWP_RMAP_CACHE_ONLY (1UL << (BITS_PER_LONG - 1)) +#define SWP_RMAP_ENTRY_MASK (~(SWP_RMAP_CACHE_ONLY | SWP_TB_PTR_MARK_MASK)) =20 static inline bool swp_tb_is_pointer(unsigned long swp_tb) { diff --git a/mm/swapfile.c b/mm/swapfile.c index 13d2ae60fe4c..2efb47b4cc4f 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -142,6 +142,11 @@ static DEFINE_PER_CPU(struct percpu_vswap_cluster, per= cpu_vswap_cluster) =3D { .lock =3D INIT_LOCAL_LOCK(), }; =20 +static void vswap_mark_cache_only(struct swap_cluster_info *ci, + unsigned int ci_off); +static void vswap_clear_cache_only(struct swap_cluster_info *ci, + unsigned int ci_start, int nr); + /* May return NULL on invalid type, caller must check for NULL return */ static struct swap_info_struct *swap_type_to_info(int type) { @@ -874,6 +879,55 @@ static int swap_cluster_setup_bad_slot(struct swap_inf= o_struct *si, return ret; } =20 +/* + * Try to reclaim a Pointer-tagged physical slot backing a vswap entry. + * The physical cluster lock must NOT be held. Returns the backing folio's + * page count, negated if the slots could not be reclaimed, or 0 if the + * folio could not be shown to own @offset (i.e. there is a race). + */ +static int try_to_reclaim_vswap_backing(struct swap_info_struct *si, + unsigned long offset, + swp_entry_t vswap_entry) +{ + swp_entry_t phys_base; + struct folio *folio; + unsigned int i; + int ret; + + folio =3D swap_cache_get_folio(vswap_entry); + if (!folio) + return 0; + + if (!folio_trylock(folio)) { + /* Unvalidated: the folio may not even own @offset. */ + folio_put(folio); + return 0; + } + + if (!folio_matches_swap_entry(folio, vswap_entry)) { + folio_unlock(folio); + folio_put(folio); + return 0; + } + + i =3D vswap_entry.val - folio->swap.val; + phys_base =3D vswap_to_phys(folio->swap); + if (!phys_base.val || swp_type(phys_base) !=3D si->type || + swp_offset(phys_base) + i !=3D offset) { + folio_unlock(folio); + folio_put(folio); + return 0; + } + + /* The run is ours: skip it all, whether or not the free succeeds. */ + ret =3D folio_nr_pages(folio); + if (!folio_free_swap(folio)) + ret =3D -ret; + folio_unlock(folio); + folio_put(folio); + return ret; +} + /* * Reclaim drops the ci lock, so the cluster may become unusable (freed or * stolen by a lower order). @usable will be set to false if that happens. @@ -1155,6 +1209,7 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) long to_scan =3D 1; unsigned long offset, end; struct swap_cluster_info *ci; + swp_entry_t vswap_entry; unsigned long swp_tb; int nr_reclaim; =20 @@ -1179,6 +1234,19 @@ static void swap_reclaim_full_clusters(struct swap_i= nfo_struct *si, bool force) offset +=3D abs(nr_reclaim); continue; } + } else if (swp_tb_is_pointer(swp_tb) && + (swp_tb & SWP_RMAP_CACHE_ONLY)) { + vswap_entry =3D swp_tb_ptr_to_swp_entry(swp_tb); + spin_unlock(&ci->lock); + nr_reclaim =3D try_to_reclaim_vswap_backing(si, offset, + vswap_entry); + ci =3D swap_cluster_lock(si, offset); + if (!ci) + goto next; + if (nr_reclaim) { + offset +=3D abs(nr_reclaim); + continue; + } } offset++; } @@ -1749,6 +1817,8 @@ static void swap_put_entries_cluster(struct swap_info= _struct *si, } /* count will be 0 after put, slot can be reclaimed */ need_reclaim =3D true; + if (swap_is_vswap(si)) + vswap_mark_cache_only(ci, ci_off); } /* * A count !=3D 1 or cached slot can't be freed. Put its swap @@ -1855,6 +1925,8 @@ static int swap_dup_entries_cluster(struct swap_info_= struct *si, goto failed; } } while (++ci_off < ci_end); + if (swap_is_vswap(si)) + vswap_clear_cache_only(ci, ci_start, nr); swap_cluster_unlock(ci); return 0; failed: @@ -2004,6 +2076,51 @@ int folio_alloc_swap(struct folio *folio) return order ? -E2BIG : -ENOMEM; } =20 +static void vswap_mark_cache_only(struct swap_cluster_info *ci, + unsigned int ci_off) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *pci; + swp_entry_t phys; + unsigned long vt; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + vt =3D __vtable_get(ci_dyn, ci_off); + + if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { + phys =3D vtable_to_phys(vt); + pci =3D __swap_entry_to_cluster(phys); + swap_rmap_mark_cache_only(pci, swp_cluster_offset(phys)); + } +} + +/* + * Clear the cache-only rmap hint for entries re-referenced from count 0 t= o 1 + * (no longer reclaimable), so the physical reclaim scanner skips them. + */ +static void vswap_clear_cache_only(struct swap_cluster_info *ci, + unsigned int ci_start, int nr) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *pci; + unsigned long swp_tb, vt; + swp_entry_t phys; + unsigned int off; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + for (off =3D ci_start; off < ci_start + nr; off++) { + swp_tb =3D __swap_table_get(ci, off); + if (!swp_tb_is_folio(swp_tb) || swp_tb_get_count(swp_tb) !=3D 1) + continue; + vt =3D __vtable_get(ci_dyn, off); + if (vtable_type(vt) !=3D VSWAP_SWAPFILE) + continue; + phys =3D vtable_to_phys(vt); + pci =3D __swap_entry_to_cluster(phys); + swap_rmap_clear_cache_only(pci, swp_cluster_offset(phys)); + } +} + static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, struct swap_cluster_info *pci, unsigned int ci_start, diff --git a/mm/vswap.h b/mm/vswap.h index 6f952b2591d1..b79866d5999c 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -44,6 +44,37 @@ static inline bool is_vswap_entry(swp_entry_t entry) return swap_is_vswap(__swap_entry_to_info(entry)); } =20 +/* + * Rmap cache-only helpers for physical cluster Pointer-tagged entries. + * SWP_RMAP_CACHE_ONLY records, inline on the physical swap_table entry, + * that the backing vswap entry has swap_count =3D=3D 0 (swap-cache-only, = so + * reclaimable). The physical reclaim scanner reads it directly instead of + * chasing the rmap into the vswap layer and paying the cluster-lookup + * indirection. + * + * Callers hold the vswap cluster lock, not the physical one. The rmap is + * only touched while the vtable holds the slot as SWAPFILE, and that + * window is opened and closed under the vswap cluster lock, so the + * allocator has finished writing the entry by then. + */ +static inline void swap_rmap_mark_cache_only(struct swap_cluster_info *ci, + unsigned int off) +{ + atomic_long_t *table; + + table =3D rcu_dereference_protected(ci->table, true); + atomic_long_or(SWP_RMAP_CACHE_ONLY, &table[off]); +} + +static inline void swap_rmap_clear_cache_only(struct swap_cluster_info *ci, + unsigned int off) +{ + atomic_long_t *table; + + table =3D rcu_dereference_protected(ci->table, true); + atomic_long_and(~SWP_RMAP_CACHE_ONLY, &table[off]); +} + /* * Virtual table entry encoding for vswap clusters. * --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f12.google.com (mail-oi2-f12.google.com [74.125.231.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B0D751B168 for ; Fri, 18 Sep 2026 18:02:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; cv=none; b=fvmQtLB8WWbthmW9ZUyNaQtM3/XJT+ntavjj22V7IcGjq28K13kFTj8tVwFz+UY21cfIzZ7SNCVJuDKDAaKkzmj43A55i0KYOIFfWiOKb0BMIKtKq7REHNZEzRORzzChApNbEJuLUvcPxwl8Qa2y0tuajcIKhZv80xXxUmeGoqU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; c=relaxed/simple; bh=WKTpSp0xD3gasRqakvuWowl65qpZ319TiioffPYIquc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FiqZmtbbLJcbDeV1/uxpFKXjbFzmBg7I41iKdcoCtyRwVs2SOMHRGlhMrRnZnISydRDJ00ni5rPCXMaibLbvdXt6q7X2GSTtlCQ43cmikkxW62Yz7j4El7fLfz34oYlOzc5D1uI+2UbZ5uc0u24gjKARnMkF/3ponpDvXFxVkjQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EyXvcV5o; arc=none smtp.client-ip=74.125.231.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EyXvcV5o" Received: by mail-oi2-f12.google.com with SMTP id 5614622812f47-4b37a2ffef2so531094b6e.3 for ; Fri, 18 Sep 2026 11:02:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754577; x=1790359377; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Psu72RUYq22I6IhK7xLj3nCdKxN5UQ7y0NifhPwjoWA=; b=EyXvcV5oabkk/fO1oXwQeyzt0XsAQiezJI1iunwokswGDtpVm7Su/aK6ec1/RZ88YQ OjnHtKfxALSbZMb5pHILDLgfW7rs/5GLoORDJr5P9qt/S1QGBT2vTd1o1tS0wb5Z7mkn jMhY7qy5hUcHQR7NbZ0aMUQupclrRjf4ZrZGxjbpYAbWRZ5Ddkm6ap6aib0xTK3zvqt5 NhxLM9ywTnIJIFXk2ge3NDwYdXunRFuig6SVP47w9va1Xis53mm6KDksSXhSDD2NgDQQ 2kIYwPhu1SmAyMvFVnHf3t4Pl8ttF/Opw4tEXh9NLFthRdUm+A+KK6iP4M9oByF+UozQ qAZg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754577; x=1790359377; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Psu72RUYq22I6IhK7xLj3nCdKxN5UQ7y0NifhPwjoWA=; b=Ms3q5nanZl65gsi2GCzQ0YITpOT9hi8s4TD2Kduh0DRC7SV69nmzvNX5pzJMLnrZER rSHaUQhUKWY/9iU5LIAp6pbFZO4dpAlfEhm4/OxWddzS6dwmxfvOtOWPVJEi683Wc0f8 7nOakDNlD3dctpqloOL3e/eDowSk/xa4Pt8dcsMwx0m3f9XYaKWp1tXw64haXNUYrR2j MnypMrFzeSgeJ23PZfeRAR4c4JayzFG+rMjmcWVxLV3h2WrhLx+GCN9FnBg/dGpYbRVt 7cnZBZd30tGaHZ4/Isti6vQgJhy/YAlRanNI7kThgtp1O8CMF1PdSGpjQUyvHmh/4K3I 2D7g== X-Forwarded-Encrypted: i=1; AKwUvBwBcjsg45zzvM5dLp2TEWSbr3inVd+IveAVieE0yQPodLEce1vzBt/oH7kCKVuvt0ACHS+4BaV/n1o5dXM=@vger.kernel.org X-Gm-Message-State: AFuF++nQme6I1FNUoZ4hRL/+ZMtL2h0QmmVOEMKT5k2yGUq7thrpLhV9 VImpbfhk2Q5bDm4DkxIjUrJyhmGEh0Jx87bIUFaUxmpI38PyH92LK5sN X-Gm-Gg: AYBFou2IjTxU5642MPKhyPRweKyzdPElz9qHEoOY0/5oE98a/vK8rq4m5nZ/K2FQNDT X8cl/4B/fBliIUrA6C1PjYSsTxA2/0QU7ve/aLPC9exhXizJFYFJpl0pscLGU2OcufCqvZv3jiL Yf4XjEjckl0xci20kCLXN09t7v9UzzAHeSSFXvKaBw5fSZHlbw+wTmCqSFbRm5EnP0TQv8QpufE YQ4Nb84DrrtEPlhKYbMG5n26uoYeO4v5zSjbjhyikfGQLMd9L6uLc9VdkVQKzdVNbkCC+jIovM9 SbYMMCxBtKILgu8oSXNZeTNJ2ZcuP993Z72hRp+LKvLe4+K0M0RNCPaWLVwE7QuwZesJ5jsiWYb 2f/dTfjt7J9zRI217EueRcBPZC2q0l3w2HIn4eSbzZhBSE4jIJ7u/Xv77QxtJGxs/2XuCr9jIbb 3/nIHjlTuODblr0m6enp4Xw+lMHMcZlvKU3GE77YRsmJMAtc4nrp/Q3r31ZX/tBAbUEDqkPG0Al IoxWEdtQ5mvunPvalyr X-Received: by 2002:a05:6808:1203:b0:4c2:e1d3:757d with SMTP id 5614622812f47-4ccf8b363d5mr3469594b6e.21.1789754577081; Fri, 18 Sep 2026 11:02:57 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:f::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4cd6a148c35sm1936178b6e.13.2026.09.18.11.02.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:56 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 08/11] mm, swap: only charge physical swap entries Date: Fri, 18 Sep 2026 11:02:38 -0700 Message-ID: <20260918180241.3424851-9-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Zswap-backed and zero-filled pages occupy no swap space, but were charged against memcg->swap as though they did. Charge memcg->swap when a vswap entry acquires physical backing rather than when it is allocated. This changes what the counter means and when the charge can fail: * memory.swap.current counts only on-disk swap usage, not zswap-backed or zero-filled pages. * A cgroup can reclaim its anon memory even with memory.swap.max set to 0, provided zswap is allowed for it. * The charge can fail at writeback rather than at allocation. swap_writeout() returns AOP_WRITEPAGE_ACTIVATE and zswap_writeback_entry() returns -ENOMEM for a cgroup at its limit. * The zswap shrinker skips such a cgroup rather than walking its LRUs. Also refactor the swap memcg operations into separate get, record, charge, uncharge and put helpers, since recording the owner and charging it no longer happen at the same time. Direct-mapped physical swap charging is unchanged. So is cgroup v1 memsw accounting: the folio's memsw charge is retained across swapout regardless of backing, and released when the entry is freed. Suggested-by: Johannes Weiner Signed-off-by: Nhat Pham --- .../admin-guide/cgroup-v1/memcg_test.rst | 2 +- Documentation/admin-guide/cgroup-v2.rst | 46 +++--- include/linux/memcontrol.h | 6 + include/linux/swap.h | 61 ++++++- mm/memcontrol-v1.c | 10 +- mm/memcontrol.c | 152 ++++++++++-------- mm/swapfile.c | 128 +++++++++++++-- mm/zswap.c | 29 ++-- 8 files changed, 312 insertions(+), 122 deletions(-) diff --git a/Documentation/admin-guide/cgroup-v1/memcg_test.rst b/Documenta= tion/admin-guide/cgroup-v1/memcg_test.rst index d9951c319ef5..cd565626c435 100644 --- a/Documentation/admin-guide/cgroup-v1/memcg_test.rst +++ b/Documentation/admin-guide/cgroup-v1/memcg_test.rst @@ -43,7 +43,7 @@ Please note that implementation details can be changed. mem_cgroup_uncharge() Called when a page's refcount goes down to 0. =20 - mem_cgroup_uncharge_swap() + mem_cgroup_swap_uncharge() Called when swp_entry's refcnt goes down to 0. A charge against swap disappears. =20 diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-= guide/cgroup-v2.rst index 8d2603751c51..2a1cba7f01ff 100644 --- a/Documentation/admin-guide/cgroup-v2.rst +++ b/Documentation/admin-guide/cgroup-v2.rst @@ -1845,16 +1845,17 @@ The following nested keys are defined. A read-only single value file which exists on non-root cgroups. =20 - The total amount of swap currently being used by the cgroup - and its descendants. + The total amount of physical swap currently being used by the + cgroup and its descendants. =20 memory.swap.high A read-write single value file which exists on non-root cgroups. The default is "max". =20 - Swap usage throttle limit. If a cgroup's swap usage exceeds - this limit, all its further allocations will be throttled to - allow userspace to implement custom out-of-memory procedures. + Physical swap usage throttle limit. If a cgroup's physical + swap usage exceeds this limit, all its further allocations will + be throttled to allow userspace to implement custom + out-of-memory procedures. =20 This limit marks a point of no return for the cgroup. It is NOT designed to manage the amount of swapping a workload does @@ -1867,8 +1868,9 @@ The following nested keys are defined. memory.swap.peak A read-write single value file which exists on non-root cgroups. =20 - The max swap usage recorded for the cgroup and its descendants since - the creation of the cgroup or the most recent reset for that FD. + The max physical swap usage recorded for the cgroup and its + descendants since the creation of the cgroup or the most recent + reset for that FD. =20 A write of any non-empty string to this file resets it to the current memory usage for subsequent reads through the same @@ -1878,8 +1880,9 @@ The following nested keys are defined. A read-write single value file which exists on non-root cgroups. The default is "max". =20 - Swap usage hard limit. If a cgroup's swap usage reaches this - limit, anonymous memory of the cgroup will not be swapped out. + Physical swap usage hard limit. If a cgroup's physical swap + usage reaches this limit, anonymous memory of the cgroup will + not be swapped out to a physical swap device. =20 memory.swap.events A read-only flat-keyed file which exists on non-root cgroups. @@ -1888,22 +1891,23 @@ The following nested keys are defined. modified event. =20 high - The number of times the cgroup's swap usage was over - the high threshold. + The number of times the cgroup's physical swap usage + was over the high threshold. =20 max - The number of times the cgroup's swap usage was about - to go over the max boundary and swap allocation - failed. + The number of times the cgroup's physical swap usage + was about to go over the max boundary and physical + swap allocation failed. =20 fail - The number of times swap allocation failed either - because of running out of swap system-wide or max - limit. - - When reduced under the current usage, the existing swap - entries are reclaimed gradually and the swap usage may stay - higher than the limit for an extended period of time. This + The number of times physical swap allocation failed + either because of running out of physical swap + system-wide or max limit. + + When reduced under the current usage, the existing physical + swap entries are reclaimed gradually and the physical swap + usage may stay higher than the limit for an extended period of + time. This reduces the impact on the workload and memory management. =20 memory.zswap.current diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 46bf724cae7a..4add06affefa 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -1933,6 +1933,7 @@ static inline void mem_cgroup_calculate_protection_pa= th(struct mem_cgroup *root, =20 #if defined(CONFIG_MEMCG) && defined(CONFIG_ZSWAP) bool obj_cgroup_may_zswap(struct obj_cgroup *objcg); +bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush); void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size); void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size); bool mem_cgroup_zswap_writeback_enabled(struct mem_cgroup *memcg); @@ -1941,6 +1942,11 @@ static inline bool obj_cgroup_may_zswap(struct obj_c= group *objcg) { return true; } + +static inline bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may= _flush) +{ + return true; +} static inline void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size) { diff --git a/include/linux/swap.h b/include/linux/swap.h index 5ab050b2457c..cd22db50b44c 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -527,20 +527,49 @@ static inline void folio_throttle_swaprate(struct fol= io *folio, gfp_t gfp) #endif =20 #if defined(CONFIG_MEMCG) && defined(CONFIG_SWAP) -int __mem_cgroup_try_charge_swap(struct folio *folio); -static inline int mem_cgroup_try_charge_swap(struct folio *folio) +struct mem_cgroup *__mem_cgroup_swap_get(struct folio *folio); +static inline struct mem_cgroup *mem_cgroup_swap_get(struct folio *folio) +{ + if (mem_cgroup_disabled()) + return NULL; + return __mem_cgroup_swap_get(folio); +} + +int __mem_cgroup_swap_charge(struct mem_cgroup *memcg, unsigned int nr_pag= es); +static inline int mem_cgroup_swap_charge(struct mem_cgroup *memcg, + unsigned int nr_pages) { if (mem_cgroup_disabled()) return 0; - return __mem_cgroup_try_charge_swap(folio); + return __mem_cgroup_swap_charge(memcg, nr_pages); } =20 -extern void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_= pages); -static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned in= t nr_pages) +void __mem_cgroup_swap_record(struct folio *folio, struct mem_cgroup *memc= g); +static inline void mem_cgroup_swap_record(struct folio *folio, + struct mem_cgroup *memcg) { if (mem_cgroup_disabled()) return; - __mem_cgroup_uncharge_swap(id, nr_pages); + __mem_cgroup_swap_record(folio, memcg); +} + +void __mem_cgroup_swap_uncharge(struct mem_cgroup *memcg, + unsigned int nr_pages); +static inline void mem_cgroup_swap_uncharge(struct mem_cgroup *memcg, + unsigned int nr_pages) +{ + if (mem_cgroup_disabled()) + return; + __mem_cgroup_swap_uncharge(memcg, nr_pages); +} + +void __mem_cgroup_swap_put(struct mem_cgroup *memcg, unsigned int nr_pages= ); +static inline void mem_cgroup_swap_put(struct mem_cgroup *memcg, + unsigned int nr_pages) +{ + if (mem_cgroup_disabled()) + return; + __mem_cgroup_swap_put(memcg, nr_pages); } =20 long mem_cgroup_get_folio_swap_margin(struct folio *folio); @@ -548,16 +577,32 @@ extern long mem_cgroup_get_nr_swap_pages(struct mem_c= group *memcg); bool mem_cgroup_can_swap(struct mem_cgroup *memcg, long nr_pages); extern bool mem_cgroup_swap_full(struct folio *folio); #else -static inline int mem_cgroup_try_charge_swap(struct folio *folio) +static inline struct mem_cgroup *mem_cgroup_swap_get(struct folio *folio) +{ + return NULL; +} + +static inline int mem_cgroup_swap_charge(struct mem_cgroup *memcg, + unsigned int nr_pages) { return 0; } =20 -static inline void mem_cgroup_uncharge_swap(unsigned short id, +static inline void mem_cgroup_swap_record(struct folio *folio, + struct mem_cgroup *memcg) +{ +} + +static inline void mem_cgroup_swap_uncharge(struct mem_cgroup *memcg, unsigned int nr_pages) { } =20 +static inline void mem_cgroup_swap_put(struct mem_cgroup *memcg, + unsigned int nr_pages) +{ +} + static inline long mem_cgroup_get_folio_swap_margin(struct folio *folio) { return PAGE_COUNTER_MAX; diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index bf2c7d53b01b..3e06a8bdf46e 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -341,6 +341,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_= cluster_info *ci) void memcg1_swapin(struct folio *folio) { struct swap_cluster_info *ci; + struct mem_cgroup *memcg; unsigned long nr_pages; unsigned short id; =20 @@ -372,7 +373,14 @@ void memcg1_swapin(struct folio *folio) id =3D __swap_cgroup_clear(ci, swp_cluster_offset(folio->swap), nr_pages); swap_cluster_unlock(ci); - mem_cgroup_uncharge_swap(id, nr_pages); + + rcu_read_lock(); + memcg =3D mem_cgroup_from_private_id(id); + if (memcg) { + mem_cgroup_swap_uncharge(memcg, nr_pages); + mem_cgroup_swap_put(memcg, nr_pages); + } + rcu_read_unlock(); } #endif =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 63d2c9e3dbe1..bba9148b745d 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5941,80 +5941,111 @@ int __init mem_cgroup_init(void) =20 #ifdef CONFIG_SWAP /** - * __mem_cgroup_try_charge_swap - try charging swap space for a folio + * __mem_cgroup_swap_get - pin the memcg to account a folio's swap slots to * @folio: folio being added to swap * - * Try to charge @folio's memcg for the swap space at folio->swap. + * Pins one private ID ref per page of @folio on its memcg, or on its clos= est + * online ancestor if it has been offlined. The caller charges and records + * against whichever memcg is returned, so both land on the same one. * - * Returns 0 on success, -ENOMEM on failure. + * Return: the pinned memcg, or NULL if there is nothing to account. Drop = the + * pins with __mem_cgroup_swap_put(). */ -int __mem_cgroup_try_charge_swap(struct folio *folio) +struct mem_cgroup *__mem_cgroup_swap_get(struct folio *folio) { unsigned int nr_pages =3D folio_nr_pages(folio); - struct swap_cluster_info *ci; - struct page_counter *counter; struct mem_cgroup *memcg; struct obj_cgroup *objcg; =20 if (do_memsw_account()) - return 0; + return NULL; =20 objcg =3D folio_objcg(folio); VM_WARN_ON_ONCE_FOLIO(!objcg, folio); if (!objcg) - return 0; + return NULL; =20 rcu_read_lock(); memcg =3D obj_cgroup_memcg(objcg); if (!folio_test_swapcache(folio)) { memcg_memory_event(memcg, MEMCG_SWAP_FAIL); rcu_read_unlock(); - return 0; + return NULL; } =20 memcg =3D mem_cgroup_private_id_get_online(memcg, nr_pages); /* memcg is pined by memcg ID. */ rcu_read_unlock(); =20 + return memcg; +} + +/** + * __mem_cgroup_swap_charge - charge physical swap space + * @memcg: the mem_cgroup to charge (may be NULL) + * @nr_pages: the amount of swap space to charge + * + * Return: 0 on success, -ENOMEM if memory.swap.max is exceeded. + */ +int __mem_cgroup_swap_charge(struct mem_cgroup *memcg, unsigned int nr_pag= es) +{ + struct page_counter *counter; + + if (do_memsw_account() || !memcg) + return 0; + if (!mem_cgroup_is_root(memcg) && !page_counter_try_charge(&memcg->swap, nr_pages, &counter)) { memcg_memory_event(memcg, MEMCG_SWAP_MAX); memcg_memory_event(memcg, MEMCG_SWAP_FAIL); - mem_cgroup_private_id_put(memcg, nr_pages); return -ENOMEM; } mod_memcg_state(memcg, MEMCG_SWAP, nr_pages); + return 0; +} + +/** + * __mem_cgroup_swap_record - record the owner of a folio's swap slots + * @folio: folio being added to swap + * @memcg: the memcg pinned by __mem_cgroup_swap_get() + */ +void __mem_cgroup_swap_record(struct folio *folio, struct mem_cgroup *memc= g) +{ + struct swap_cluster_info *ci; =20 ci =3D swap_cluster_get_and_lock(folio); - __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages, - mem_cgroup_private_id(memcg)); + __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), + folio_nr_pages(folio), mem_cgroup_private_id(memcg)); swap_cluster_unlock(ci); - - return 0; } =20 /** - * __mem_cgroup_uncharge_swap - uncharge swap space - * @id: cgroup id to uncharge + * __mem_cgroup_swap_uncharge - uncharge physical swap space + * @memcg: the mem_cgroup to uncharge (may be NULL) * @nr_pages: the amount of swap space to uncharge */ -void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages) +void __mem_cgroup_swap_uncharge(struct mem_cgroup *memcg, unsigned int nr_= pages) { - struct mem_cgroup *memcg; + if (!memcg) + return; =20 - rcu_read_lock(); - memcg =3D mem_cgroup_from_private_id(id); - if (memcg) { - if (!mem_cgroup_is_root(memcg)) { - if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, nr_pages); - else - page_counter_uncharge(&memcg->swap, nr_pages); - } - mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages); - mem_cgroup_private_id_put(memcg, nr_pages); + if (!mem_cgroup_is_root(memcg)) { + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, nr_pages); + else + page_counter_uncharge(&memcg->swap, nr_pages); } - rcu_read_unlock(); + mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages); +} + +/** + * __mem_cgroup_swap_put - drop the private ID refs taken for swap slots + * @memcg: the pinned mem_cgroup + * @nr_pages: number of refs to drop + */ +void __mem_cgroup_swap_put(struct mem_cgroup *memcg, unsigned int nr_pages) +{ + mem_cgroup_private_id_put(memcg, nr_pages); } =20 long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg) @@ -6055,27 +6086,12 @@ long mem_cgroup_get_folio_swap_margin(struct folio = *folio) * @memcg: the memcg to query * @nr_pages: the number of pages the caller wants to swap out * - * A vswap zswap-backed swapout needs no physical slot, so gate on the - * swap.max headroom rather than the physical free count. - * * Return: true if @memcg can swap out at least @nr_pages more pages. */ bool mem_cgroup_can_swap(struct mem_cgroup *memcg, long nr_pages) { - long avail; - - if (mem_cgroup_can_vswap(memcg)) - return true; - - if (!vswap_is_enabled() || !zswap_is_enabled()) - return mem_cgroup_get_nr_swap_pages(memcg) >=3D nr_pages; - - avail =3D PAGE_COUNTER_MAX; - for (; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) - avail =3D min_t(long, avail, - READ_ONCE(memcg->swap.max) - - page_counter_read(&memcg->swap)); - return avail >=3D nr_pages; + return mem_cgroup_can_vswap(memcg) || + mem_cgroup_get_nr_swap_pages(memcg) >=3D nr_pages; } =20 bool mem_cgroup_swap_full(struct folio *folio) @@ -6241,8 +6257,10 @@ static struct cftype swap_files[] =3D { =20 #ifdef CONFIG_ZSWAP /** - * obj_cgroup_may_zswap - check if this cgroup can zswap - * @objcg: the object cgroup + * mem_cgroup_may_zswap - check if this cgroup can zswap + * @memcg: the memcg to query + * @may_flush: force-flush stats for an accurate check (sleeps). Pass false + * from atomic contexts; the check is then best-effort. * * Check if the hierarchical zswap limit has been reached. * @@ -6252,36 +6270,38 @@ static struct cftype swap_files[] =3D { * spending cycles on compression when there is already no room left * or zswap is disabled altogether somewhere in the hierarchy. */ -bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) +bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush) { - struct mem_cgroup *memcg, *original_memcg; - bool ret =3D true; - if (!cgroup_subsys_on_dfl(memory_cgrp_subsys)) return true; =20 - original_memcg =3D get_mem_cgroup_from_objcg(objcg); - for (memcg =3D original_memcg; !mem_cgroup_is_root(memcg); - memcg =3D parent_mem_cgroup(memcg)) { + for (; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) { unsigned long max =3D READ_ONCE(memcg->zswap_max); unsigned long pages; =20 if (max =3D=3D PAGE_COUNTER_MAX) continue; - if (max =3D=3D 0) { - ret =3D false; - break; - } + if (max =3D=3D 0) + return false; =20 /* Force flush to get accurate stats for charging */ - __mem_cgroup_flush_stats(memcg, true); + if (may_flush) + __mem_cgroup_flush_stats(memcg, true); pages =3D memcg_page_state(memcg, MEMCG_ZSWAP_B) / PAGE_SIZE; - if (pages < max) - continue; - ret =3D false; - break; + if (pages >=3D max) + return false; } - mem_cgroup_put(original_memcg); + return true; +} + +bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) +{ + struct mem_cgroup *memcg; + bool ret; + + memcg =3D get_mem_cgroup_from_objcg(objcg); + ret =3D mem_cgroup_may_zswap(memcg, true); + mem_cgroup_put(memcg); return ret; } =20 diff --git a/mm/swapfile.c b/mm/swapfile.c index 2efb47b4cc4f..3c3fc3b87b9b 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -1944,7 +1944,8 @@ static int swap_dup_entries_cluster(struct swap_info_= struct *si, bool mem_cgroup_can_vswap(struct mem_cgroup *memcg) { return vswap_is_enabled() && zswap_is_enabled() && - (mem_cgroup_disabled() || do_memsw_account()); + (mem_cgroup_disabled() || do_memsw_account() || + mem_cgroup_may_zswap(memcg, false)); } =20 static bool vswap_alloc(struct folio *folio) @@ -2030,6 +2031,7 @@ static swp_entry_t folio_alloc_phys_swap(struct folio= *folio) int folio_alloc_swap(struct folio *folio) { unsigned int order =3D folio_order(folio); + struct mem_cgroup *memcg; unsigned int size =3D 1 << order; =20 VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); @@ -2056,10 +2058,20 @@ int folio_alloc_swap(struct folio *folio) if (!vswap_alloc(folio)) folio_alloc_phys_swap(folio); =20 - /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */ - if (unlikely(mem_cgroup_try_charge_swap(folio))) { - swap_cache_del_folio(folio); - goto failed; + /* + * Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. + * A vswap entry has no physical swap yet, so only record the memcg. + * folio_realloc_swap() charges it once backing is allocated. + */ + memcg =3D mem_cgroup_swap_get(folio); + if (memcg) { + if (!is_vswap_entry(folio->swap) && + unlikely(mem_cgroup_swap_charge(memcg, size))) { + mem_cgroup_swap_put(memcg, size); + swap_cache_del_folio(folio); + } else { + mem_cgroup_swap_record(folio, memcg); + } } =20 if (unlikely(!folio_test_swapcache(folio))) @@ -2126,6 +2138,36 @@ static void __swap_cluster_free_phys_backing(struct = swap_info_struct *psi, unsigned int ci_start, unsigned int nr_pages); =20 +static void vswap_uncharge_cgroup_batch(unsigned short memcg_id, + unsigned int batch_nr, + unsigned int batch_nr_swapfile) +{ + struct mem_cgroup *memcg; + unsigned int n; + + /* + * v1 (memsw): entries keep their memsw charge across swapout + * regardless of backing, so uncharge all of them. v2: only + * swapfile-backed entries are charged, so uncharge just those. + * + * On v1 the id is written by __memcg1_swapout() as the folio leaves the + * swap cache and cleared by memcg1_swapin() when it comes back, both + * under the cluster lock. Callers still holding a cached folio are + * outside that window and see @memcg_id =3D=3D 0, so only the free path + * uncharges. On v2 the id is set when swap is allocated, so those + * callers do uncharge, which balances the charge folio_realloc_swap() + * took. + */ + n =3D do_memsw_account() ? batch_nr : batch_nr_swapfile; + if (!n) + return; + + rcu_read_lock(); + memcg =3D memcg_id ? mem_cgroup_from_private_id(memcg_id) : NULL; + rcu_read_unlock(); + mem_cgroup_swap_uncharge(memcg, n); +} + /** * __vswap_release_backing - release the backing of a range of vtable slots * @ci: the locked vswap cluster @@ -2146,12 +2188,25 @@ void __vswap_release_backing(struct swap_cluster_in= fo *ci, unsigned int ci_off; unsigned long vt; swp_entry_t phys_first =3D {}; + unsigned short batch_id, cur_id; + unsigned int batch_nr =3D 0, batch_nr_swapfile =3D 0; =20 lockdep_assert_held(&ci->lock); ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + batch_id =3D __swap_cgroup_get(ci, ci_start); =20 for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { vt =3D __vtable_get(ci_dyn, ci_off); + cur_id =3D __swap_cgroup_get(ci, ci_off); + + if (cur_id !=3D batch_id) { + vswap_uncharge_cgroup_batch(batch_id, batch_nr, + batch_nr_swapfile); + batch_id =3D cur_id; + batch_nr =3D 0; + batch_nr_swapfile =3D 0; + } + batch_nr++; =20 /* The free helper takes one contiguous run within one cluster. */ if (phys_off_start !=3D phys_off_end && @@ -2169,6 +2224,7 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, =20 switch (vtable_type(vt)) { case VSWAP_SWAPFILE: + batch_nr_swapfile++; if (phys_off_start =3D=3D phys_off_end) { phys_first =3D vtable_to_phys(vt); phys_off_start =3D swp_offset(phys_first); @@ -2200,6 +2256,8 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, phys_off_start % SWAPFILE_CLUSTER, phys_off_end - phys_off_start); } + + vswap_uncharge_cgroup_batch(batch_id, batch_nr, batch_nr_swapfile); } =20 /** @@ -2289,7 +2347,10 @@ swp_entry_t folio_realloc_swap(struct folio *folio) swp_entry_t vswap_entry =3D folio->swap; struct swap_cluster_info *ci; struct swap_cluster_info_dynamic *ci_dyn; + struct mem_cgroup *memcg; unsigned int voff; + unsigned long vt; + unsigned short memcg_id; swp_entry_t phys_entry =3D {}; swp_entry_t pe; int i, nr =3D folio_nr_pages(folio); @@ -2298,18 +2359,37 @@ swp_entry_t folio_realloc_swap(struct folio *folio) VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); VM_WARN_ON(!is_vswap_entry(vswap_entry)); =20 - phys_entry =3D vswap_to_phys(vswap_entry); - if (phys_entry.val) - return phys_entry; + voff =3D swp_cluster_offset(vswap_entry); + ci =3D __swap_entry_to_cluster(vswap_entry); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + + spin_lock(&ci->lock); + vt =3D __vtable_get(ci_dyn, voff); + if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { + spin_unlock(&ci->lock); + return vtable_to_phys(vt); + } + memcg_id =3D __swap_cgroup_get(ci, voff); + spin_unlock(&ci->lock); =20 phys_entry =3D folio_alloc_phys_swap(folio); if (!phys_entry.val) return (swp_entry_t){}; =20 - voff =3D swp_cluster_offset(vswap_entry); + rcu_read_lock(); + memcg =3D folio_memcg(folio); + if (!memcg || mem_cgroup_private_id(memcg) !=3D memcg_id) + memcg =3D memcg_id ? mem_cgroup_from_private_id(memcg_id) : NULL; + rcu_read_unlock(); + + if (mem_cgroup_swap_charge(memcg, nr)) { + __swap_cluster_free_phys_backing(__swap_entry_to_info(phys_entry), + __swap_entry_to_cluster(phys_entry), + swp_cluster_offset(phys_entry), + nr); + return (swp_entry_t){}; + } =20 - ci =3D __swap_entry_to_cluster(vswap_entry); - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); spin_lock(&ci->lock); /* * Install PHYS backing without freeing any prior contents of the @@ -2494,6 +2574,25 @@ static void __swap_cluster_free_phys_backing(struct = swap_info_struct *psi, swap_cluster_unlock(pci); } =20 +/* + * Release the cgroup accounting of a batch of freed slots. For vswap the + * physical swap was already uncharged by __vswap_release_backing(), so on= ly + * the ID ref is left to drop. + */ +static void memcg_swap_free(unsigned short id, unsigned int nr, bool is_vs= wap) +{ + struct mem_cgroup *memcg; + + rcu_read_lock(); + memcg =3D mem_cgroup_from_private_id(id); + if (memcg) { + if (!is_vswap) + mem_cgroup_swap_uncharge(memcg, nr); + mem_cgroup_swap_put(memcg, nr); + } + rcu_read_unlock(); +} + /* * Free a set of swap slots after their swap count dropped to zero, or wil= l be * zero after putting the last ref (saves one __swap_cluster_put_entry cal= l). @@ -2506,10 +2605,11 @@ void __swap_cluster_free_entries(struct swap_info_s= truct *si, unsigned short batch_id =3D 0, id_cur; unsigned int ci_off =3D ci_start, ci_end =3D ci_start + nr_pages; unsigned int batch_off =3D ci_off; + bool is_vswap =3D swap_is_vswap(si); =20 VM_WARN_ON(ci->count < nr_pages); =20 - if (swap_is_vswap(si)) + if (is_vswap) __vswap_release_backing(ci, ci_start, nr_pages); =20 ci->count -=3D nr_pages; @@ -2533,14 +2633,14 @@ void __swap_cluster_free_entries(struct swap_info_s= truct *si, id_cur =3D __swap_cgroup_clear(ci, ci_off, 1); if (batch_id !=3D id_cur) { if (batch_id) - mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + memcg_swap_free(batch_id, ci_off - batch_off, is_vswap); batch_id =3D id_cur; batch_off =3D ci_off; } } while (++ci_off < ci_end); =20 if (batch_id) - mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + memcg_swap_free(batch_id, ci_off - batch_off, is_vswap); =20 __swap_cluster_finish_free(si, ci, ci_start, nr_pages); } diff --git a/mm/zswap.c b/mm/zswap.c index 56315298c291..aac09970c4d1 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1130,11 +1130,17 @@ static int zswap_writeback_entry(struct zswap_entry= *entry, **********************************/ /* * vswap zswap entries get a physical slot allocated on demand at writeback - * time. Skip the shrinker when none is available. + * time, and that slot is charged to @memcg. Skip the shrinker when either + * the device or the cgroup has no room left. @memcg may be NULL to check + * the device alone. */ -static bool zswap_writeback_possible(void) +static bool zswap_writeback_possible(struct mem_cgroup *memcg) { - return !vswap_is_enabled() || get_nr_swap_pages() > 0; + if (!vswap_is_enabled()) + return true; + if (!memcg) + return get_nr_swap_pages() > 0; + return mem_cgroup_get_nr_swap_pages(memcg) > 0; } =20 /* @@ -1199,9 +1205,10 @@ static enum lru_status shrink_memcg_cb(struct list_h= ead *item, struct list_lru_o * * Temporary failures, where the same entry should be tried * again immediately, almost never happen for this shrinker. - * We don't do any trylocking; -ENOMEM comes closest, - * but that's extremely rare and doesn't happen spuriously - * either. Don't bother distinguishing this case. + * We don't do any trylocking; -ENOMEM comes closest, but + * zswap_writeback_possible() keeps the shrinker off cgroups + * with no physical swap headroom, and it doesn't happen + * spuriously either. Don't bother distinguishing this case. */ list_move_tail(item, &l->list); =20 @@ -1274,7 +1281,7 @@ static unsigned long zswap_shrinker_count(struct shri= nker *shrinker, if (!zswap_shrinker_enabled || !mem_cgroup_zswap_writeback_enabled(memcg)) return 0; =20 - if (!zswap_writeback_possible()) + if (!zswap_writeback_possible(memcg)) return 0; =20 /* @@ -1361,7 +1368,7 @@ static struct shrinker *zswap_alloc_shrinker(void) * writeback disabled, is a zombie cgroup, or has empty zswap LRUs. * * Also returns -ENOENT when vswap is enabled and there is no physical - * swap to write back to. + * swap for @memcg to write back to. */ static int shrink_memcg(struct mem_cgroup *memcg) { @@ -1370,7 +1377,7 @@ static int shrink_memcg(struct mem_cgroup *memcg) if (!mem_cgroup_zswap_writeback_enabled(memcg)) return -ENOENT; =20 - if (!zswap_writeback_possible()) + if (!zswap_writeback_possible(memcg)) return -ENOENT; =20 /* @@ -1401,7 +1408,7 @@ static void shrink_worker(struct work_struct *w) int ret, failures =3D 0, attempts =3D 0; unsigned long thr; =20 - if (!zswap_writeback_possible()) + if (!zswap_writeback_possible(NULL)) return; =20 /* Reclaim down to the accept threshold */ @@ -1482,7 +1489,7 @@ static void shrink_worker(struct work_struct *w) break; resched: cond_resched(); - } while (zswap_total_pages() > thr && zswap_writeback_possible()); + } while (zswap_total_pages() > thr && zswap_writeback_possible(NULL)); } =20 /********************************* --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f23.google.com (mail-oi2-f23.google.com [74.125.231.215]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1512051CF7F for ; Fri, 18 Sep 2026 18:02:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.215 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754584; cv=none; b=ft6TQs4Q2jDyaUBmSSsbo6HEGQsWxYg00+GrtePYqBNNHzwkYaLrPXopQFz+6ol/TtKb9xED1FHeK/xHfJOQSGO9XUXytED4jVsbGC9VSAEkoJxKbyusTFsxnKYed6t5grUyEXBWGbWSdFA4ooueDhXz37WZzwuTCtRvCkJLTIk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754584; c=relaxed/simple; bh=DKrV8Dk2Xf7+0f4DO4cy+Nlg5v572tfl82Ug1Z8WZHA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KbXotB3wo++dmRA5dGfJnXJormoijEtICQ+0PLPFtAaYa3rQk/R8Y4tKMBYHo0mwcR4mZEUYa/gMZvBjbNl/9IqHSJewcmD2pQWUwcZU+WQbU/bVQUXThncl6+6L4Wx68ppaC30MqFnRrZRwOoEjwyW4U2WNg97EKDq56Bgn5ZI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NKrZdAbX; arc=none smtp.client-ip=74.125.231.215 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NKrZdAbX" Received: by mail-oi2-f23.google.com with SMTP id 46e09a7af769-7f4f0d37dccso589422a34.3 for ; Fri, 18 Sep 2026 11:02:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754578; x=1790359378; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WkWDXoev3ue4855J8KR8M6HtT2XHzciSRhLql9R/iaw=; b=NKrZdAbX37tFFps5wJG4gDwuD+7Amk62HrEVxFb/x1qyabY3KErXXE+KjUcrEOyYmL K4OwhcMHOEeh+DxWeefb/D0MzN/Tkp4CezgeyEEXrSRlKXIOCLI2xN9GZ+EIEVVwEa7p GJ/p0t8SrCvbbixIM+pkLNA20Vkh6DiJdCp/xYd2FTdukvboriPIQU1toHND7YOwAvhh kDAabakX9n/RP4yzyoR5YnGsD5tXBCXAlLfk/msencEMtGQxxQwahA3rlzpIlPkU67zt QEczZjrWWPzmHWBS8lWuCRi93i6JJil4KnMTv96lYSbMWkoXeXvBkWz3ID8x7fo46Yl5 FyNg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754578; x=1790359378; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=WkWDXoev3ue4855J8KR8M6HtT2XHzciSRhLql9R/iaw=; b=0YJHF3/yVy1782N52mVIaNQS6LM8Hb48j1EhuhEADqiWiYj9N082gXaCYCA3s59q+Q SKIcgeJGSMaqCfWFsEaw9AmKG2oPky2IIKGvHhBus30b2JPjhbFTMA8MHfiyXru7IZXd /PJ77X9ZyTq9NONE3T0yHSB5O4QSWGGZ0R7o3qeMkEfdrKwnt2gR3cMGIcWK9eIt/En7 FnkR20oN8KUgRYGEhAALsBVw4MCbQ4C/t11rrb/PbJ/o00cy8GKU+l88Vjn5SmWtlKxt mBpjZsFEjyie/F+gXW7ZmU1XgjPMhHlVKRaAIy8iQKx75167Hwiy0vZDJ4uVDABE4Z91 i8mQ== X-Forwarded-Encrypted: i=1; AKwUvBy3sFnxL/rMoHXd+6MzPSfhHN1Th7U9gEurk423Z9zQmKNtOQeGl4zjiagw7R7plpcWQzT8gK+MvgNdPfY=@vger.kernel.org X-Gm-Message-State: AFuF++mKN6wafe60sZYmN57e1Gt5SNsxDvGLbfGqSqpwdmVI6YRl//Gy B5tGI7E2SmaAWY+G+BPPmV+jmYO0kWQlc45kQR25def6T2wagUyPwojN X-Gm-Gg: AYBFou34FFWtSGhZT1H2i7w4xJeOn8LmwmIzcoCpkAGxm+yFU13GZ6wtkZPc8hapZ+E PlKZ0Hvxpas4Opc0yFWaziTguH7dHio/C0ya/VquM5SabxBkZAOD5b4UHSFoOYf7QpTLEx3bxE+ Y/NVeOh9eNKyf19rCLoB7rVZIEhF3XcC0jRazrFOJdHhEsPmhlHKWak7huIU3JPrsLN9KjN9n1e FYSLxCxf7OYXJi5naWmh0jdR1Ygxzs8HVYDUthlfpzJ8deTyS+2+WsPasJ8vvZOqfS1eLEtURnR 7AJm/UKS3ZYUojnlN3AFw78UYwJT0J58Ik5kPeCOe1me6TRqOHaFq+9Hn1T1VCQeSoDqJ/XKvGM F7uT9iHBF1rTmDQ40N7mf01gBoavyQSiPSJa3+48E2bL24dZbKs4AxSxlY/mUg+T7LNpWyk4dcv u1rm3iIVDBe2HScRsgZaa+7ECH//XeRkK9C3snFgY5X/mF3ZK1pfQoiZyluWZS8Ivf03KzlvkB9 4tjjq5kEgfdIIgEatgI8w== X-Received: by 2002:a05:6820:81c2:b0:6c3:38ca:c9c0 with SMTP id 006d021491bc7-6ca9ca59232mr3229844eaf.53.1789754578532; Fri, 18 Sep 2026 11:02:58 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4e::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6cd3591d0b6sm565715eaf.12.2026.09.18.11.02.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:57 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 09/11] mm, swap: add debugfs counters for vswap Date: Fri, 18 Sep 2026 11:02:39 -0700 Message-ID: <20260918180241.3424851-10-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add /sys/kernel/debug/vswap/ with two counters: * used: virtual swap slots (pages) currently allocated * alloc_reject: cumulative pages that failed to get a vswap slot Signed-off-by: Nhat Pham --- mm/swapfile.c | 25 +++++++++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/mm/swapfile.c b/mm/swapfile.c index 3c3fc3b87b9b..cf07d87d2301 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -7,6 +7,7 @@ */ =20 #include +#include #include #include #include @@ -142,6 +143,7 @@ static DEFINE_PER_CPU(struct percpu_vswap_cluster, perc= pu_vswap_cluster) =3D { .lock =3D INIT_LOCAL_LOCK(), }; =20 +static atomic_long_t vswap_alloc_reject =3D ATOMIC_LONG_INIT(0); static void vswap_mark_cache_only(struct swap_cluster_info *ci, unsigned int ci_off); static void vswap_clear_cache_only(struct swap_cluster_info *ci, @@ -1996,6 +1998,7 @@ static bool vswap_alloc(struct folio *folio) =20 this_cpu_write(percpu_vswap_cluster.offset[order], SWAP_ENTRY_INVALID); local_unlock(&percpu_vswap_cluster.lock); + atomic_long_add(folio_nr_pages(folio), &vswap_alloc_reject); return false; } =20 @@ -4827,9 +4830,25 @@ early_param("vswap", early_vswap); /* vswap does no IO on its own. */ static const struct swap_ops vswap_ops =3D { }; =20 +static int vswap_used_get(void *data, u64 *val) +{ + *val =3D swap_usage_in_pages(vswap_si); + return 0; +} +DEFINE_DEBUGFS_ATTRIBUTE(vswap_used_fops, vswap_used_get, NULL, "%llu\n"); + +static int vswap_alloc_reject_get(void *data, u64 *val) +{ + *val =3D atomic_long_read(&vswap_alloc_reject); + return 0; +} +DEFINE_DEBUGFS_ATTRIBUTE(vswap_alloc_reject_fops, vswap_alloc_reject_get, = NULL, + "%llu\n"); + static int __init vswap_init(void) { struct swap_info_struct *si; + struct dentry *root; unsigned long maxpages; int err; =20 @@ -4881,6 +4900,12 @@ static int __init vswap_init(void) mutex_unlock(&swapon_mutex); =20 vswap_si =3D si; + + root =3D debugfs_create_dir("vswap", NULL); + debugfs_create_file("used", 0444, root, NULL, &vswap_used_fops); + debugfs_create_file("alloc_reject", 0444, root, NULL, + &vswap_alloc_reject_fops); + pr_info("vswap: created virtual swap device (%lu pages)\n", maxpages); =20 /* Last: everything above must be visible before routing starts. */ --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oi2-f43.google.com (mail-oi2-f43.google.com [74.125.231.235]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE7C951D536 for ; Fri, 18 Sep 2026 18:03:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.235 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; cv=none; b=FzYgvngeCN4I43MqJGBrSpQXWiwWia3msqbwifiOhSPechu4Alxo9Skgb5Z1Vd2kfL/8Yra5fUW7T0UNuZJwQbEWzfW06Jm6xNBUvsmh46HGpGd42qihSadvgnzoxgOm0ew9biB6cFm5BrUpQAc1zWwxNVyuXR9SfbpofU5iIGk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; c=relaxed/simple; bh=248meyirULeKvPgqKHcpHfv3DScKvY8uep/W5NAxx1s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aFnsxHELmhiWF+CDTs3ZQSBLqWgvw35KJ+0T+D1jrjDicewJTK/5b6g2xNUSoCvY654YpFXGa0e6by+V22089LAojsGu98CQIRo64f9MWUCLORQTr+bWyIyKGQR/5xawueZGM1osdsF347dwDz7mIvY1Rr/XvdTosMJSFvLIVxs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=K3zgXMub; arc=none smtp.client-ip=74.125.231.235 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="K3zgXMub" Received: by mail-oi2-f43.google.com with SMTP id 46e09a7af769-7f4f0d1779dso938539a34.2 for ; Fri, 18 Sep 2026 11:03:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754580; x=1790359380; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=leUxoZNbrIHhJwMkVXew1sBjFqiId1Gr67nxTcdiA8U=; b=K3zgXMubFzyRxcXj6UaURZWLgP8JGNZ9k+S8ptsKVpiH0+quEGRhzVcew29tLaHNGC J961N1sOL1y2qMWgshUIedJzrcgEybXUiUOL/yeQGPkgU6J+RfQ/r2XVFlRwD/I5dYzy Brq7E+TSuirObrPLZ1MVU47LZi9HzeEWFZghRprpW0Y9pbt/6j0cl0mkb3qjjh/OTupb IbMHRsn4QqvFT7jdIc4gSum4hV5EQCMYMjGQE71fPNDgj2m8gDlv0ZXe3ZsFk1Dj//O0 PzhAtsWwZgj97lsHWinTMGD/HwgviefqzlFKnlQn5QHGRDRau69FvItijwTr7yd4MtAg W9qA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754580; x=1790359380; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=leUxoZNbrIHhJwMkVXew1sBjFqiId1Gr67nxTcdiA8U=; b=FyAds++XJLPazFyn+iOoUVq0gjNLgcntto1iPW86oWgkEcOkGtOaNrELrXJeYAP+ZY hR/Lnj0JZv+jWxn9Rs+Yq5k21pW9BceuoGZH7R3AckM/Abv/QQ3ZhVv5JXucvpgCgXkO ++QBzO+ZZkMr46hITMbYVcdj0DpJ5Q4VWu+yZNOBlfEnO9diizVv9EvJLiBiYfNJ8Hvq P5SoATwluLJuk5GMJybLoxfgMviMmradeIUROYchmu65597jS72M51AnbIHgpWqxdCE+ JZxQPS4r5y9mczi8o428O4RbdgVVGuhoiKN2VVjGXPUMyi5tDpsgpj4kh94AOO3vRiPt /GTg== X-Forwarded-Encrypted: i=1; AKwUvBzr+GJ/YNUqu3PLkBYMFMwuQZqABGYmAABkbaoqVLo2U7q4/nFvRjNhvtM3DbZhSHgbLq5XcymqkyEYks0=@vger.kernel.org X-Gm-Message-State: AFuF++nHDbRE58xGv2Y3oj0fLvM6ElokOsiUG5egcfLryIdnyZ5rgBRn /0GWRNiccVtYAGUR4WlJ26wIRSJgLrVB42MngXzEM1BmdUDdM/GaQdeD X-Gm-Gg: AYBFou3LqxcHdms7p+LNkYv0C8CydNtdsUbnm8K3bZ6X6DvCcOE6eSCw6ARNcEXgEI4 z4HyaZ+92rWfuxDEazVXSbTNKvNlaQT1VA5ebqc1b80fGXNku55vT0EoR/JULfJ8IRbf5NuVned DzaMDzxVgE3AO/XPTQWo5CDkizOBnQnrdmWKrE7x73ITWpWUVGYUiDoEzSNTc88OmfZ2FJjTCEV 0h6zLhbIiphbJ5Y+5sQ6UZPuD75phMViqZK/EcKwec2VhJq3ttwBggZWuK0N1myPRWAPwqEaOuk CqYWw/iPIR11LCREaNow205pKCmxAajHLhqOfv7MgeXV557MRDLqVioWHNUenkv+Nup66UXF9IF YzX6YSc/NPffpVZ5N6WYviiXKoJ0tdl8b8Wo0xzuTGcLU3GDkkv+4Q7rlIT/xlC+qO6sithrXMq X6qOfLBz3cuo41lzb0BGm9a7KfmfjJvHkCNmBPZWCmb035jbthzuzcuBuKB6dGrKXJF7y1MPN9R ethYkMTM9hdRJAY8YaNcfDq2cFXx/aY X-Received: by 2002:a05:6830:650a:b0:804:ca33:4aff with SMTP id 46e09a7af769-80de0e50a9cmr3677995a34.9.1789754580097; Fri, 18 Sep 2026 11:03:00 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:43::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-8107e894949sm141880a34.17.2026.09.18.11.02.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:59 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 10/11] mm, swap: defer memcg_table allocation for physical swap clusters Date: Fri, 18 Sep 2026 11:02:40 -0700 Message-ID: <20260918180241.3424851-11-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Stop allocating a memcg table for every physical swap cluster that only ever holds vswap backings. The table costs SWAPFILE_CLUSTER * sizeof(unsigned short) per cluster, 1 KB per 2 MB of swap on a 64-bit kernel with 4 KB pages. On a vswap-heavy workload, where zswap writeback is the only consumer of physical swap, that is the common case. Such clusters never have their memcg_table read or written: vswap-layer charging records on the vswap cluster's table, not the physical one. Allocate eagerly only where the table is known to be needed: every vswap cluster, and, when vswap is off, every physical cluster, since none of its slots is then a vswap backing. A physical cluster otherwise allocates on its first direct-use slot, and skips entirely if it only holds vswap backings. Hibernation slots have no folio and record no cgroup, so they do not trigger it. That deferred allocation is on the allocator's fast path and can fail; the allocation it serves then fails too, and the caller falls back to another cluster. Signed-off-by: Nhat Pham --- mm/swapfile.c | 83 +++++++++++++++++++++++++++++++++++++++------------ 1 file changed, 64 insertions(+), 19 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index cf07d87d2301..39d1840b0d36 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -472,7 +472,8 @@ static void swap_cluster_free_table(struct swap_cluster= _info *ci) swap_cluster_free_count_table(table); } =20 -static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gf= p) +static int swap_cluster_alloc_table(struct swap_info_struct *si, + struct swap_cluster_info *ci, gfp_t gfp) { struct swap_table *table =3D NULL; struct folio *folio; @@ -493,7 +494,14 @@ static int swap_cluster_alloc_table(struct swap_cluste= r_info *ci, gfp_t gfp) return -ENOMEM; =20 #ifdef CONFIG_MEMCG - if (!mem_cgroup_disabled()) { + /* + * A physical cluster under vswap may hold only vswap backings, which + * record their memcg on the vswap cluster's table, not this one. Such + * clusters defer memcg_table allocation until they hand out a slot + * that maps directly into the PTEs. + */ + if ((!vswap_is_enabled() || swap_is_vswap(si)) && + !mem_cgroup_disabled()) { VM_WARN_ON_ONCE(ci->memcg_table); ci->memcg_table =3D kzalloc_obj(*ci->memcg_table, gfp); if (!ci->memcg_table) { @@ -571,8 +579,8 @@ swap_cluster_populate(struct swap_info_struct *si, lockdep_assert_held(&si->global_cluster_lock); lockdep_assert_held(&ci->lock); =20 - if (!swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - __GFP_NOWARN)) + if (!swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + __GFP_NOWARN)) return ci; =20 /* @@ -585,8 +593,8 @@ swap_cluster_populate(struct swap_info_struct *si, spin_unlock(&si->global_cluster_lock); local_unlock(&percpu_swap_cluster.lock); =20 - ret =3D swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - GFP_KERNEL); + ret =3D swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + GFP_KERNEL); =20 /* * Back to atomic context. We might have migrated to a new CPU with a @@ -863,7 +871,7 @@ static int swap_cluster_setup_bad_slot(struct swap_info= _struct *si, =20 ci =3D cluster_info + idx; /* Need to allocate swap table first for initial bad slot marking. */ - if (!ci->count && swap_cluster_alloc_table(ci, GFP_KERNEL)) + if (!ci->count && swap_cluster_alloc_table(si, ci, GFP_KERNEL)) return -ENOMEM; spin_lock(&ci->lock); /* Check for duplicated bad swap slots. */ @@ -1086,7 +1094,9 @@ static bool __swap_cluster_alloc_entries(struct swap_= info_struct *si, /* Try use a new cluster for current CPU and allocate from it. */ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci, - struct folio *folio, unsigned long offset) + struct folio *folio, + unsigned long offset, + bool *nomem) { unsigned int next =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; unsigned long start =3D ALIGN_DOWN(offset, SWAPFILE_CLUSTER); @@ -1115,6 +1125,23 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, if (!ret) continue; } +#ifdef CONFIG_MEMCG + /* + * Lazy-allocate memcg_table on the first direct-use slot of a + * physical cluster. + */ + if (vswap_is_enabled() && folio && + !folio_test_swapcache(folio) && !mem_cgroup_disabled() && + !ci->memcg_table) { + ci->memcg_table =3D kzalloc_obj(*ci->memcg_table, + GFP_ATOMIC | __GFP_NOWARN); + if (!ci->memcg_table) { + if (nomem) + *nomem =3D true; + goto out; + } + } +#endif if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUST= ER)) break; found =3D offset; @@ -1124,7 +1151,15 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, break; } out: - relocate_cluster(si, ci); + /* + * On a discard-capable device, relocating a cluster whose memcg_table + * allocation failed queues a discard for slots that were never used, + * which folio_alloc_phys_swap() reads as progress and retries on. + */ + if (nomem && *nomem && !ci->count) + __free_cluster(si, ci); + else + relocate_cluster(si, ci); swap_cluster_unlock(ci); if (swap_is_vswap(si)) { this_cpu_write(percpu_vswap_cluster.offset[order], next); @@ -1145,7 +1180,13 @@ static unsigned int alloc_swap_scan_list(struct swap= _info_struct *si, bool scan_all) { unsigned int found =3D SWAP_ENTRY_INVALID; + bool nomem =3D false; =20 + /* + * In rare cases alloc_swap_scan_cluster() can fail due to + * memcg_table allocation failure. Short-circuit to avoid looping + * over the list indefinitely. + */ do { struct swap_cluster_info *ci =3D isolate_lock_cluster(si, list); unsigned long offset; @@ -1153,10 +1194,10 @@ static unsigned int alloc_swap_scan_list(struct swa= p_info_struct *si, if (!ci) break; offset =3D cluster_offset(si, ci); - found =3D alloc_swap_scan_cluster(si, ci, folio, offset); + found =3D alloc_swap_scan_cluster(si, ci, folio, offset, &nomem); if (found) break; - } while (scan_all); + } while (scan_all && !nomem); =20 return found; } @@ -1177,7 +1218,8 @@ static unsigned int vswap_alloc_cluster(struct swap_i= nfo_struct *si, spin_lock_init(&ci_dyn->ci.lock); INIT_LIST_HEAD(&ci_dyn->ci.list); =20 - if (swap_cluster_alloc_table(&ci_dyn->ci, GFP_ATOMIC | __GFP_NOWARN)) { + if (swap_cluster_alloc_table(si, &ci_dyn->ci, + GFP_ATOMIC | __GFP_NOWARN)) { kfree(ci_dyn); return SWAP_ENTRY_INVALID; } @@ -1203,7 +1245,7 @@ static unsigned int vswap_alloc_cluster(struct swap_i= nfo_struct *si, } =20 offset =3D cluster_offset(si, ci); - return alloc_swap_scan_cluster(si, ci, folio, offset); + return alloc_swap_scan_cluster(si, ci, folio, offset, NULL); } =20 static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool f= orce) @@ -1310,7 +1352,8 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, if (cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); - found =3D alloc_swap_scan_cluster(si, ci, folio, offset); + found =3D alloc_swap_scan_cluster(si, ci, folio, offset, + NULL); } else { swap_cluster_unlock(ci); } @@ -1354,7 +1397,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, if (order < PMD_ORDER) { /* * Scan only one fragment cluster is good enough. Order 0 - * allocation will surely success, and large allocation + * allocation rarely fails, and large allocation * failure is not critical. Scanning one cluster still * keeps the list rotated and reclaimed (for clean swap cache). */ @@ -1369,7 +1412,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, /* Order 0 stealing from higher order */ for (int o =3D 1; o < SWAP_NR_ORDERS; o++) { /* - * Clusters here have at least one usable slots and can't fail order 0 + * Clusters here have at least one usable slots and rarely fail order 0 * allocation, but reclaim may drop si->lock and race with another user. */ found =3D alloc_swap_scan_list(si, &si->frag_clusters[o], folio, true); @@ -1590,7 +1633,7 @@ static swp_entry_t swap_alloc_fast(struct folio *foli= o) if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); - found =3D alloc_swap_scan_cluster(si, ci, folio, offset); + found =3D alloc_swap_scan_cluster(si, ci, folio, offset, NULL); } else if (ci) { swap_cluster_unlock(ci); } @@ -1981,7 +2024,8 @@ static bool vswap_alloc(struct folio *folio) if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(vswap_si, ci); - alloc_swap_scan_cluster(vswap_si, ci, folio, offset); + alloc_swap_scan_cluster(vswap_si, ci, folio, offset, + NULL); } else if (ci) { swap_cluster_unlock(ci); } @@ -2860,7 +2904,8 @@ swp_entry_t swap_alloc_hibernation_slot(int type) if (pcp_si =3D=3D si && pcp_offset) { ci =3D swap_cluster_lock(si, pcp_offset); if (cluster_is_usable(ci, 0)) - offset =3D alloc_swap_scan_cluster(si, ci, NULL, pcp_offset); + offset =3D alloc_swap_scan_cluster(si, ci, NULL, + pcp_offset, NULL); else swap_cluster_unlock(ci); } --=20 2.53.0-Meta From nobody Wed Sep 23 17:09:52 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2DF5951DB05 for ; Fri, 18 Sep 2026 18:03:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754591; cv=none; b=dcOMTlS9WWSnMxFCBCA00Dh+aQ4uGoC8U2IVCWu7hSabWtft7vSdo4IX0/t7Ys2ExNEK/cw7MR7W85OQMPiC1wIB7XkNttD19XXY9Ows89jGoh7TEqwbyO6aAYJRkvNFzXKAbG3ERN0C70rPHK+C5zU/z4y2udUbnzT+K+ksGuY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754591; c=relaxed/simple; bh=1ru2Zp5ILl53osMtDD38RHdLyI5pG1nL61SzFYE36Sc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=J4wzCtWFzAPTwJO/9ABD//YbPo7xfE0InzC+VCJeuw0qRG5VGBNmrbTqE62AJGx98UtNeZ1RyPUxu+Wfyf2OAD7GU1wws53XTgeYFq8CtNGBwJMSwZvAtoXkk7gwuB4ZFOEqkS7bnRMXuqER99lo8hsKh4ZXIln25npQOACRv28= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pgDk5Rvd; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pgDk5Rvd" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-46accbdfc39so1198235fac.3 for ; Fri, 18 Sep 2026 11:03:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754582; x=1790359382; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=eMvS9UEOG0U7dnKspyTQkECTAwb3ShdMHwYlkFC5G4c=; b=pgDk5Rvdhn2zbZE5GFZHOH+rzBX3UBNEnmEVQuJ5stiMrtL6JGlXn7OSYSwyXpeGTB Vu3c5/nF/rOq5vUe1tTINKCXwYqpyODexEnE3SPddXdKOwN28lF6/mKnbh62QRopNzZV XUQF2JgIF1DGXUf8fwnNzXhoJmzrh2pMPRvc08KR12S5HuwWeYqiEq+zdFIyO7xcw+8R cZ2T4WaXyynPjAdsxcjk7nQzlSPCk2qdzGD9TFTscYFXs90Bob6C56qfCjalxhnRpKHL e3v6cAwqKm2XHLmRQ7qvXWXKmhMRhOV10e94YlhkDiGV9BoKpVaHPYHLUAfrOo4p+53W pY7g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754582; x=1790359382; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=eMvS9UEOG0U7dnKspyTQkECTAwb3ShdMHwYlkFC5G4c=; b=dU3g4YOcMrRqJU1kongJEsRiVUmy6Jjx6er9KOLJ4JS/35E70aSXbyvoYw4hRDjGkL cAmgYMm1YWhXG/uM98IcylibV0KpMj1BBvpYhSPbMn4NsrByZMv1RZRvoqVkTpZYxJfb r0Q46G0e1z45vgBLHGaAgdacwp7wIktrBH2FnyUy3vePF8ulqyCIC9/i0JjXyLeHbUKh En99DtpD6T4h5BGPDahPY3ScJpTFret6ap6SQYoU/6OHQ6e9W3f8eCpVMxJ1QQYohXG9 t31xuqC0wlh6KqDrEx02Rktv1/8n7YP48WDTOKhnRcOv1zwmgVC6An9hQ+9d7Xx1eCrq kn/A== X-Forwarded-Encrypted: i=1; AKwUvBympR1QBJwOMUyECWQmCyU8o2/p4RTelpit6VnezGevnsGAZrMInC7SM7UuBexRx15SOsLgsTHy9k3Y2wQ=@vger.kernel.org X-Gm-Message-State: AFuF++keSklKXeSqF3Ehb9wvGXRLKDMfmFN3rkQH4sspXJXT7ErvAB5e IQbIrdI8rFJ23z8zpYorjNV7mFNGhxa9HCQrYQrNQ+z2qwAkbKNaFYbx X-Gm-Gg: AYBFou0rHXLNVAnupLjW2iX01mRFx2Etp0MfIql4A9g8+dwTqt5OPLB6rHmgNAv7HFY BR0FfvSv5GTbUoqzJ/mkl0KE6FOSDXfu+CM4Acg2tzTbPeWLZ4PzpnYxJhRMOyvN+63cAv04Zro wmdxTbKO3ahT6kxH9gJYAbgqtlvuObPWQfEZJwtq3mGh5snKYAgEZhc5AGo7RXtFIAqDRKNV2l3 q0+TBnv+UqvvCLMCMnaO3NQTLHf+gfKH6MWYYpdoksI8cxMZE5Sl5ezhiw1XiC5u3azJ4To63x6 p35vY2D+k1uYLKYPleZYtIpA7e7BI01P0f0R7eboVKEhNHzTeZwRygn0PqDrYoCuTdUv/o7fd33 5k42LuhMff3lKZWYdV+dhEgAItL4/QrJZthXpqXC7YsoLml5jbK+HQDPWlUBSWr8xSOxhIICb6o DSS0Vfx8j4Kl03llNLZnbRFcjjMsIxhVnY1/bP8eqiV+lYXFgwp/i87dtiwXaq0WhZ7C1Wzs4Qv YpnGr1QvbA2cawk4bIzsQ== X-Received: by 2002:a05:6871:3865:b0:41b:e633:baf4 with SMTP id 586e51a60fabf-486e4bc49dfmr4029647fac.3.1789754581881; Fri, 18 Sep 2026 11:03:01 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:46::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-4873922d1besm1668863fac.6.2026.09.18.11.03.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:03:01 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [RFC PATCH v5 11/11] mm, swap: back vswap clusters with a VM_SPARSE array Date: Fri, 18 Sep 2026 11:02:41 -0700 Message-ID: <20260918180241.3424851-12-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" vswap keeps its cluster_info in an xarray of individually allocated clusters. Replace it with the VM_SPARSE vmalloc array Baoquan He designed for xswap: one reservation at init, mapped a page of clusters at a time as the device grows. That drops the per-cluster allocation and the xarray nodes, and it simplifies access, because the index gives the address. A lookup becomes arithmetic instead of an xa_load() that can return NULL, and no cluster needs an RCU grace period to be freed, so the NULL arm goes away in every caller along with the index and rcu_head fields, the kfree_rcu(), CLUSTER_FLAG_DEAD and __vswap_cluster_lock(). Only the grow side is ported; there is no shrink. That leaves swap_cluster_info_dynamic wrapping swap_cluster_info for a single pointer, so move the virtual table into swap_cluster_info and delete the wrapper. Both cluster arrays then have the same element type and merge into si->cluster_info, which restores cluster_index() to upstream's subtraction and leaves __swap_offset_to_cluster() a plain array index. A vswap cluster ends up smaller, having lost the index and rcu_head; a physical cluster grows by the one pointer it never uses. A cluster is no longer destroyed when it empties. There is nothing left to free, since it is now an element of a fixed array, so it goes onto si->free_clusters like a physical device's cluster and waits to be reused. Its virtual table is still freed, and that is the bulk of it: SWAPFILE_CLUSTER pointers, a full page at the usual layout, against a few dozen bytes for the cluster itself. An emptied cluster gives back almost all of what it held; only the mapping stays. Most of this is Baoquan's code, adapted to vswap's existing cluster layer rather than to a new device type, so I am keeping his attributions from the original posting. Co-developed-by: Baoquan He Signed-off-by: Baoquan He Signed-off-by: Nhat Pham Link: https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos= .cn/ --- include/linux/swap.h | 5 +- mm/swap.h | 61 +------- mm/swap_state.c | 19 +-- mm/swap_table.h | 9 -- mm/swapfile.c | 353 ++++++++++++++++++++++++++++--------------- mm/vswap.h | 96 +++++------- 6 files changed, 280 insertions(+), 263 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index cd22db50b44c..dffdec14c407 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -164,6 +164,7 @@ static inline void mm_account_reclaimed_pages(unsigned = long pages) =20 struct address_space; struct sysinfo; +struct vm_struct; struct zone; =20 /* @@ -271,7 +272,9 @@ struct swap_info_struct { struct list_head discard_clusters; /* discard clusters list */ struct plist_node avail_list; /* entry in swap_avail_head */ const struct swap_ops *ops; - struct xarray cluster_info_pool; /* Xarray for vswap dynamic cluster info= */ + struct vm_struct *cluster_info_area; /* Vswap cluster array reservation */ + unsigned int nr_mapped_clusters; /* Mapped prefix of cluster_info */ + struct mutex cluster_grow_lock; /* Serialize growth of the array */ }; =20 static inline bool swap_is_vswap(struct swap_info_struct *si) diff --git a/mm/swap.h b/mm/swap.h index df323d5e8da8..83015ff5f390 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -64,16 +64,10 @@ struct swap_cluster_info { #if !SWAP_TABLE_HAS_ZEROFLAG unsigned long *zero_bitmap; #endif + atomic_long_t *virtual_table; /* Backing pointers, vswap clusters only */ struct list_head list; }; =20 -struct swap_cluster_info_dynamic { - struct swap_cluster_info ci; - unsigned int index; /* for cluster_index() */ - struct rcu_head rcu; - atomic_long_t *virtual_table; /* Backing pointers for vswap slots */ -}; - /* All on-list cluster must have a non-zero flag. */ enum swap_cluster_flags { CLUSTER_FLAG_NONE =3D 0, /* For temporary off-list cluster */ @@ -84,7 +78,6 @@ enum swap_cluster_flags { CLUSTER_FLAG_USABLE =3D CLUSTER_FLAG_FRAG, CLUSTER_FLAG_FULL, CLUSTER_FLAG_DISCARD, - CLUSTER_FLAG_DEAD, /* Vswap dynamic cluster pending kfree_rcu */ CLUSTER_FLAG_MAX, }; =20 @@ -127,17 +120,6 @@ static inline struct swap_info_struct *__swap_entry_to= _info(swp_entry_t entry) return __swap_type_to_info(swp_type(entry)); } =20 -/** - * __swap_offset_to_cluster - look up the cluster holding a swap offset - * @si: the swap device - * @offset: the swap entry offset - * - * Context: A vswap cluster is freed by kfree_rcu(). Callers must hold the - * RCU read lock, or know the cluster is pinned by an in-use entry. - * - * Return: the cluster, or NULL if @si is a vswap device with no cluster - * allocated at @offset. - */ static inline struct swap_cluster_info *__swap_offset_to_cluster( struct swap_info_struct *si, pgoff_t offset) { @@ -145,13 +127,8 @@ static inline struct swap_cluster_info *__swap_offset_= to_cluster( =20 VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ VM_WARN_ON_ONCE(offset >=3D roundup(si->max, SWAPFILE_CLUSTER)); - - if (swap_is_vswap(si)) { - struct swap_cluster_info_dynamic *ci_dyn; - - ci_dyn =3D xa_load(&si->cluster_info_pool, cluster_idx); - return ci_dyn ? &ci_dyn->ci : NULL; - } + VM_WARN_ON_ONCE(swap_is_vswap(si) && + cluster_idx >=3D READ_ONCE(si->nr_mapped_clusters)); =20 return &si->cluster_info[cluster_idx]; } @@ -162,32 +139,6 @@ static inline struct swap_cluster_info *__swap_entry_t= o_cluster(swp_entry_t entr swp_offset(entry)); } =20 -static inline struct swap_cluster_info *__vswap_cluster_lock( - struct swap_info_struct *si, unsigned long offset, bool irq) -{ - struct swap_cluster_info *ci; - - rcu_read_lock(); - ci =3D __swap_offset_to_cluster(si, offset); - if (ci) { - if (irq) - spin_lock_irq(&ci->lock); - else - spin_lock(&ci->lock); - - /* The cluster can be torn down while we wait for the lock. */ - if (ci->flags =3D=3D CLUSTER_FLAG_DEAD) { - if (irq) - spin_unlock_irq(&ci->lock); - else - spin_unlock(&ci->lock); - ci =3D NULL; - } - } - rcu_read_unlock(); - return ci; -} - static __always_inline struct swap_cluster_info *__swap_cluster_lock( struct swap_info_struct *si, unsigned long offset, bool irq) { @@ -205,9 +156,6 @@ static __always_inline struct swap_cluster_info *__swap= _cluster_lock( VM_WARN_ON_ONCE(!in_task()); VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ =20 - if (swap_is_vswap(si)) - return __vswap_cluster_lock(si, offset, irq); - ci =3D __swap_offset_to_cluster(si, offset); if (irq) spin_lock_irq(&ci->lock); @@ -223,8 +171,7 @@ static __always_inline struct swap_cluster_info *__swap= _cluster_lock( * * Context: The caller must ensure the offset is in the valid range and * protect the swap device with reference count or locks. - * Return: the locked cluster, or NULL if it is gone. Only a vswap device - * can return NULL, as its clusters are allocated and freed on demand. + * Return: The locked cluster. */ static inline struct swap_cluster_info *swap_cluster_lock( struct swap_info_struct *si, unsigned long offset) diff --git a/mm/swap_state.c b/mm/swap_state.c index 2107d05ae8d5..627fee08593c 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -165,7 +165,6 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, unsigned int ci_off, ci_end; unsigned long old_tb; bool is_zero; - struct swap_cluster_info_dynamic *ci_dyn; enum vswap_backing_type type; int ret; =20 @@ -201,8 +200,7 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, * swap_cache_alloc_folio will retry with a smaller order on -EBUSY. */ if (is_vswap_entry(targ_entry)) { - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); - ret =3D __vswap_check_backing(ci_dyn, round_down(ci_off, nr), + ret =3D __vswap_check_backing(ci, round_down(ci_off, nr), nr, &type); if (ret !=3D nr || type =3D=3D VSWAP_ZSWAP) return -EBUSY; @@ -451,12 +449,9 @@ static struct folio *__swap_cache_alloc(swp_entry_t ta= rg_entry, gfp_t gfp, entry.val =3D round_down(targ_entry.val, nr_pages); =20 /* Check if the slot and range are available, skip allocation if not */ - err =3D -ENOENT; ci =3D swap_cluster_lock(si, offset); - if (ci) { - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); - swap_cluster_unlock(ci); - } + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); + swap_cluster_unlock(ci); if (unlikely(err)) return ERR_PTR(err); =20 @@ -477,13 +472,10 @@ static struct folio *__swap_cache_alloc(swp_entry_t t= arg_entry, gfp_t gfp, return ERR_PTR(-ENOMEM); =20 /* Double check the range is still not in conflict */ - err =3D -ENOENT; ci =3D swap_cluster_lock(si, offset); - if (ci) - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg= _id); + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg_= id); if (unlikely(err)) { - if (ci) - swap_cluster_unlock(ci); + swap_cluster_unlock(ci); folio_put(folio); return ERR_PTR(err); } @@ -495,7 +487,6 @@ static struct folio *__swap_cache_alloc(swp_entry_t tar= g_entry, gfp_t gfp, =20 if (mem_cgroup_swapin_charge_folio(folio, memcg_id, vmf ? vmf->vma->vm_mm : NULL, gfp)) { - /* The folio pins the cluster */ spin_lock(&ci->lock); __swap_cache_do_del_folio(ci, folio, entry, shadow); spin_unlock(&ci->lock); diff --git a/mm/swap_table.h b/mm/swap_table.h index 79f06642a553..7d005a943881 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -266,11 +266,6 @@ static inline unsigned long swap_table_get(struct swap= _cluster_info *ci, return swp_tb; } =20 -/* - * Resolve @entry's cluster and read its slot, both under RCU. A vswap - * cluster is allocated on demand and freed by kfree_rcu(), so a caller - * starting from an entry cannot resolve it beforehand. - */ static inline unsigned long swap_table_lookup(swp_entry_t entry) { struct swap_cluster_info *ci; @@ -279,10 +274,6 @@ static inline unsigned long swap_table_lookup(swp_entr= y_t entry) =20 rcu_read_lock(); ci =3D __swap_entry_to_cluster(entry); - if (!ci) { - rcu_read_unlock(); - return null_to_swp_tb(); - } table =3D rcu_dereference(ci->table); swp_tb =3D table ? atomic_long_read(&table[swp_cluster_offset(entry)]) : null_to_swp_tb(); diff --git a/mm/swapfile.c b/mm/swapfile.c index 39d1840b0d36..d529a27fdd89 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -144,6 +144,35 @@ static DEFINE_PER_CPU(struct percpu_vswap_cluster, per= cpu_vswap_cluster) =3D { }; =20 static atomic_long_t vswap_alloc_reject =3D ATOMIC_LONG_INIT(0); + +/* + * Vswap allocates from its own device with a separate percpu cluster cach= e, + * so the allocator has two local locks to pick from. + */ +static void swap_percpu_cluster_lock(struct swap_info_struct *si) +{ + if (swap_is_vswap(si)) + local_lock(&percpu_vswap_cluster.lock); + else + local_lock(&percpu_swap_cluster.lock); +} + +static void swap_percpu_cluster_unlock(struct swap_info_struct *si) +{ + if (swap_is_vswap(si)) + local_unlock(&percpu_vswap_cluster.lock); + else + local_unlock(&percpu_swap_cluster.lock); +} + +static void swap_percpu_cluster_assert_held(struct swap_info_struct *si) +{ + if (swap_is_vswap(si)) + lockdep_assert_held(&this_cpu_ptr(&percpu_vswap_cluster)->lock); + else + lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); +} + static void vswap_mark_cache_only(struct swap_cluster_info *ci, unsigned int ci_off); static void vswap_clear_cache_only(struct swap_cluster_info *ci, @@ -420,8 +449,6 @@ static inline bool cluster_is_usable(struct swap_cluste= r_info *ci, int order) static inline unsigned int cluster_index(struct swap_info_struct *si, struct swap_cluster_info *ci) { - if (swap_is_vswap(si)) - return container_of(ci, struct swap_cluster_info_dynamic, ci)->index; return ci - si->cluster_info; } =20 @@ -450,10 +477,14 @@ static void swap_cluster_free_count_table(struct swap= _table *table) swap_cluster_free_table_folio_rcu_cb); } =20 -static void swap_cluster_free_table(struct swap_cluster_info *ci) +static void swap_cluster_free_table(struct swap_info_struct *si, + struct swap_cluster_info *ci) { struct swap_table *table; =20 + if (swap_is_vswap(si)) + vswap_cluster_free_vtable(ci); + #ifdef CONFIG_MEMCG kfree(ci->memcg_table); ci->memcg_table =3D NULL; @@ -515,12 +546,19 @@ static int swap_cluster_alloc_table(struct swap_info_= struct *si, VM_WARN_ON_ONCE(ci->zero_bitmap); ci->zero_bitmap =3D bitmap_zalloc(SWAPFILE_CLUSTER, gfp); if (!ci->zero_bitmap) { - swap_cluster_free_table(ci); + swap_cluster_free_table(si, ci); swap_cluster_free_count_table(table); return -ENOMEM; } #endif =20 + /* The virtual table shares the swap table's lifetime. */ + if (swap_is_vswap(si) && vswap_cluster_alloc_vtable(ci, gfp)) { + swap_cluster_free_table(si, ci); + swap_cluster_free_count_table(table); + return -ENOMEM; + } + /* * Make tables visible to cluster_is_usable() after everything is * ready. @@ -571,10 +609,8 @@ swap_cluster_populate(struct swap_info_struct *si, /* * Only cluster isolation from the allocator does table allocation. * Swap allocator uses percpu clusters and holds the local lock. - * vswap clusters are destroyed rather than freed to si->free_clusters. */ - VM_WARN_ON_ONCE(swap_is_vswap(si)); - lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); + swap_percpu_cluster_assert_held(si); if (!(si->flags & SWP_SOLIDSTATE)) lockdep_assert_held(&si->global_cluster_lock); lockdep_assert_held(&ci->lock); @@ -591,7 +627,7 @@ swap_cluster_populate(struct swap_info_struct *si, spin_unlock(&ci->lock); if (!(si->flags & SWP_SOLIDSTATE)) spin_unlock(&si->global_cluster_lock); - local_unlock(&percpu_swap_cluster.lock); + swap_percpu_cluster_unlock(si); =20 ret =3D swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL); @@ -604,7 +640,7 @@ swap_cluster_populate(struct swap_info_struct *si, * could happen with ignoring the percpu cluster is fragmentation, * which is acceptable since this fallback and race is rare. */ - local_lock(&percpu_swap_cluster.lock); + swap_percpu_cluster_lock(si); if (!(si->flags & SWP_SOLIDSTATE)) spin_lock(&si->global_cluster_lock); spin_lock(&ci->lock); @@ -652,20 +688,7 @@ static void swap_cluster_schedule_discard(struct swap_= info_struct *si, static void __free_cluster(struct swap_info_struct *si, struct swap_cluste= r_info *ci) { swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, false); - swap_cluster_free_table(ci); - - if (swap_is_vswap(si)) { - struct swap_cluster_info_dynamic *ci_dyn; - - /* vswap clusters are destroyed, not returned to free_clusters. */ - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); - xa_erase(&si->cluster_info_pool, ci_dyn->index); - move_cluster(si, ci, NULL, CLUSTER_FLAG_DEAD); - vswap_cluster_free_vtable(ci); - kfree_rcu(ci_dyn, rcu); - return; - } - + swap_cluster_free_table(si, ci); move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE); ci->order =3D 0; } @@ -1202,50 +1225,147 @@ static unsigned int alloc_swap_scan_list(struct sw= ap_info_struct *si, return found; } =20 -static unsigned int vswap_alloc_cluster(struct swap_info_struct *si, - struct folio *folio) +/* + * Reserve address space for the vswap cluster array. Nothing is mapped ye= t, + * so this costs address space only, plus an eighth of it in shadow under + * CONFIG_KASAN_VMALLOC. + */ +static int vswap_reserve_cluster_array(struct swap_info_struct *si, + unsigned long maxpages) +{ + unsigned long nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); + + mutex_init(&si->cluster_grow_lock); + si->cluster_info_area =3D get_vm_area(nr_clusters * + sizeof(*si->cluster_info), + VM_SPARSE); + if (!si->cluster_info_area) + return -ENOMEM; + + si->cluster_info =3D si->cluster_info_area->addr; + return 0; +} + +static void vswap_free_cluster_array(struct swap_info_struct *si) +{ + unsigned long addr, end; + struct page *page; + + if (!si->cluster_info_area) + return; + + end =3D round_up((unsigned long)&si->cluster_info[si->nr_mapped_clusters], + PAGE_SIZE); + for (addr =3D (unsigned long)si->cluster_info; addr < end; + addr +=3D PAGE_SIZE) { + page =3D vmalloc_to_page((void *)addr); + vm_area_unmap_pages(si->cluster_info_area, addr, + addr + PAGE_SIZE); + __free_page(page); + } + + free_vm_area(si->cluster_info_area); + si->cluster_info_area =3D NULL; + si->cluster_info =3D NULL; + si->nr_mapped_clusters =3D 0; +} + +static bool vswap_can_grow(struct swap_info_struct *si) +{ + return READ_ONCE(si->nr_mapped_clusters) < + DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER); +} + +/* Clusters added per growth of the vswap cluster array, one page worth. */ +#define VSWAP_GROW_CLUSTERS \ + max_t(unsigned long, \ + PAGE_SIZE / sizeof(struct swap_cluster_info), 16) + +/* + * Map one more page of the vswap cluster array and hand the clusters it + * covers to the allocator. The caller must not hold the percpu cluster + * lock: vm_area_map_pages() might sleep. + * + * The mapped prefix only ever grows, so the pages already backing clusters + * [0, si->nr_mapped_clusters) are exactly those below the page boundary + * above the last one. A grow whose clusters all fall inside an already + * mapped page maps nothing. + */ +static int vswap_grow_clusters(struct swap_info_struct *si) { - struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *ci; - unsigned long offset; + unsigned int noreclaim_flags; + unsigned long start, end; + struct page *page; + unsigned int i, first, nr; + int err =3D -ENOSPC; =20 + BUILD_BUG_ON(VSWAP_GROW_CLUSTERS * + sizeof(struct swap_cluster_info) > PAGE_SIZE); VM_WARN_ON(!swap_is_vswap(si)); =20 - ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_ATOMIC | __GFP_NOWARN); - if (!ci_dyn) - return SWAP_ENTRY_INVALID; + /* Rechecked under the mutex, this only keeps a full device cheap. */ + if (!vswap_can_grow(si)) + return -ENOSPC; =20 - spin_lock_init(&ci_dyn->ci.lock); - INIT_LIST_HEAD(&ci_dyn->ci.list); + /* Outside the mutex, so this one may still reclaim. */ + page =3D alloc_page(__GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL | + __GFP_ZERO); =20 - if (swap_cluster_alloc_table(si, &ci_dyn->ci, - GFP_ATOMIC | __GFP_NOWARN)) { - kfree(ci_dyn); - return SWAP_ENTRY_INVALID; - } + mutex_lock(&si->cluster_grow_lock); + first =3D si->nr_mapped_clusters; + nr =3D min_t(unsigned int, VSWAP_GROW_CLUSTERS, + DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER) - first); + if (!nr) + goto out; =20 - if (vswap_cluster_alloc_vtable(ci_dyn, GFP_ATOMIC | __GFP_NOWARN)) { - swap_cluster_free_table(&ci_dyn->ci); - kfree(ci_dyn); - return SWAP_ENTRY_INVALID; - } + start =3D round_up((unsigned long)&si->cluster_info[first], + PAGE_SIZE); + end =3D round_up((unsigned long)&si->cluster_info[first + nr], + PAGE_SIZE); =20 - /* Lock before publishing: xa_alloc makes the cluster findable by offset.= */ - ci =3D &ci_dyn->ci; - spin_lock(&ci->lock); + if (start !=3D end) { + err =3D -ENOMEM; + if (!page) + goto out; + /* + * vm_area_map_pages() allocates page tables with + * GFP_PGTABLE_KERNEL, so they carry __GFP_DIRECT_RECLAIM. + * A non-reclaim caller of folio_alloc_swap() would otherwise + * recurse back here and deadlock on the mutex it already + * holds. Callers already under PF_MEMALLOC do not need this, + * swapon does. It grants the page tables reserve access, at + * most three pages per grow. + */ + noreclaim_flags =3D memalloc_noreclaim_save(); + err =3D vm_area_map_pages(si->cluster_info_area, start, end, + &page); + memalloc_noreclaim_restore(noreclaim_flags); + if (err) + goto out; + page =3D NULL; + } =20 - if (xa_alloc(&si->cluster_info_pool, &ci_dyn->index, ci_dyn, - XA_LIMIT(1, DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER) - 1), - GFP_ATOMIC | __GFP_NOWARN)) { + /* + * Publish the new clusters before they become reachable by offset. + * A zeroed page leaves them off-list with CLUSTER_FLAG_NONE, which + * is what move_cluster() expects. + */ + WRITE_ONCE(si->nr_mapped_clusters, first + nr); + for (i =3D first; i < first + nr; i++) { + ci =3D &si->cluster_info[i]; + spin_lock_init(&ci->lock); + INIT_LIST_HEAD(&ci->list); + spin_lock(&ci->lock); + move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE); spin_unlock(&ci->lock); - swap_cluster_free_table(&ci_dyn->ci); - vswap_cluster_free_vtable(&ci_dyn->ci); - kfree(ci_dyn); - return SWAP_ENTRY_INVALID; } - - offset =3D cluster_offset(si, ci); - return alloc_swap_scan_cluster(si, ci, folio, offset, NULL); + err =3D 0; +out: + mutex_unlock(&si->cluster_grow_lock); + if (page) + __free_page(page); + return err; } =20 static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool f= orce) @@ -1272,8 +1392,6 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) nr_reclaim =3D __try_to_reclaim_swap(si, offset, TTRS_ANYWAY); ci =3D swap_cluster_lock(si, offset); - if (!ci) - goto next; if (nr_reclaim) { offset +=3D abs(nr_reclaim); continue; @@ -1285,8 +1403,6 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) nr_reclaim =3D try_to_reclaim_vswap_backing(si, offset, vswap_entry); ci =3D swap_cluster_lock(si, offset); - if (!ci) - goto next; if (nr_reclaim) { offset +=3D abs(nr_reclaim); continue; @@ -1300,7 +1416,6 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) relocate_cluster(si, ci); =20 swap_cluster_unlock(ci); -next: if (to_scan <=3D 0) break; =20 @@ -1378,10 +1493,19 @@ static unsigned long cluster_alloc_swap_entry(struc= t swap_info_struct *si, goto done; } =20 - if (swap_is_vswap(si)) { - found =3D vswap_alloc_cluster(si, folio); - if (found) - goto done; + /* + * Grow the vswap cluster array and let the free list scan below pick + * up the new clusters. Growth sleeps, so drop the percpu cluster lock + * across it; the scan does not care which CPU it lands back on. The + * list_empty() test is racy either way: a stale empty costs one page, + * a stale non-empty skips the grow and leaves the caller to the + * fragment and stealing scans below. + */ + if (swap_is_vswap(si) && list_empty(&si->free_clusters) && + vswap_can_grow(si)) { + local_unlock(&percpu_vswap_cluster.lock); + vswap_grow_clusters(si); + local_lock(&percpu_vswap_cluster.lock); } =20 if (!(si->flags & SWP_PAGE_DISCARD)) { @@ -1630,11 +1754,11 @@ static swp_entry_t swap_alloc_fast(struct folio *fo= lio) return (swp_entry_t){}; =20 ci =3D swap_cluster_lock(si, offset); - if (ci && cluster_is_usable(ci, order)) { + if (cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); found =3D alloc_swap_scan_cluster(si, ci, folio, offset, NULL); - } else if (ci) { + } else { swap_cluster_unlock(ci); } =20 @@ -1760,7 +1884,6 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t g= fp) if (IS_ERR_OR_NULL(si)) return 0; =20 - /* The source PTE pins the entry, so its cluster is alive. */ ci =3D __swap_offset_to_cluster(si, offset); ret =3D swap_extend_table_alloc(si, ci, swp_cluster_offset(entry), gfp); =20 @@ -2021,12 +2144,12 @@ static bool vswap_alloc(struct folio *folio) =20 if (offset !=3D SWAP_ENTRY_INVALID) { ci =3D swap_cluster_lock(vswap_si, offset); - if (ci && cluster_is_usable(ci, order)) { + if (cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(vswap_si, ci); alloc_swap_scan_cluster(vswap_si, ci, folio, offset, NULL); - } else if (ci) { + } else { swap_cluster_unlock(ci); } } @@ -2138,13 +2261,11 @@ int folio_alloc_swap(struct folio *folio) static void vswap_mark_cache_only(struct swap_cluster_info *ci, unsigned int ci_off) { - struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *pci; swp_entry_t phys; unsigned long vt; =20 - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); - vt =3D __vtable_get(ci_dyn, ci_off); + vt =3D __vtable_get(ci, ci_off); =20 if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { phys =3D vtable_to_phys(vt); @@ -2160,18 +2281,16 @@ static void vswap_mark_cache_only(struct swap_clust= er_info *ci, static void vswap_clear_cache_only(struct swap_cluster_info *ci, unsigned int ci_start, int nr) { - struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *pci; unsigned long swp_tb, vt; swp_entry_t phys; unsigned int off; =20 - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); for (off =3D ci_start; off < ci_start + nr; off++) { swp_tb =3D __swap_table_get(ci, off); if (!swp_tb_is_folio(swp_tb) || swp_tb_get_count(swp_tb) !=3D 1) continue; - vt =3D __vtable_get(ci_dyn, off); + vt =3D __vtable_get(ci, off); if (vtable_type(vt) !=3D VSWAP_SWAPFILE) continue; phys =3D vtable_to_phys(vt); @@ -2229,7 +2348,6 @@ static void vswap_uncharge_cgroup_batch(unsigned shor= t memcg_id, void __vswap_release_backing(struct swap_cluster_info *ci, unsigned int ci_start, unsigned int nr) { - struct swap_cluster_info_dynamic *ci_dyn; struct swap_info_struct *psi; unsigned long phys_off_start =3D 0, phys_off_end =3D 0; unsigned int ci_off; @@ -2239,11 +2357,10 @@ void __vswap_release_backing(struct swap_cluster_in= fo *ci, unsigned int batch_nr =3D 0, batch_nr_swapfile =3D 0; =20 lockdep_assert_held(&ci->lock); - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); batch_id =3D __swap_cgroup_get(ci, ci_start); =20 for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { - vt =3D __vtable_get(ci_dyn, ci_off); + vt =3D __vtable_get(ci, ci_off); cur_id =3D __swap_cgroup_get(ci, ci_off); =20 if (cur_id !=3D batch_id) { @@ -2290,7 +2407,7 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, break; } =20 - __vtable_set(ci_dyn, ci_off, VSWAP_NONE); + __vtable_set(ci, ci_off, VSWAP_NONE); /* Zero-backed state lives in swap_table; clear it too. */ if (__swap_table_test_zero(ci, ci_off)) __swap_table_clear_zero(ci, ci_off); @@ -2348,14 +2465,12 @@ void folio_release_vswap_backing(struct folio *foli= o) void folio_release_non_phys_swap_backing(struct folio *folio) { struct swap_cluster_info *ci; - struct swap_cluster_info_dynamic *ci_dyn; int nr =3D folio_nr_pages(folio); unsigned int voff; unsigned long vt; enum vswap_backing_type type; =20 ci =3D __swap_entry_to_cluster(folio->swap); - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); voff =3D swp_cluster_offset(folio->swap); =20 spin_lock(&ci->lock); @@ -2363,7 +2478,7 @@ void folio_release_non_phys_swap_backing(struct folio= *folio) * A folio's slots cannot mix swapfile with other backends, except * mid-backend-change, which always starts from slot 0. */ - vt =3D __vtable_get(ci_dyn, voff); + vt =3D __vtable_get(ci, voff); type =3D vtable_type(vt); =20 if (type =3D=3D VSWAP_SWAPFILE || type =3D=3D VSWAP_NONE) { @@ -2393,7 +2508,6 @@ swp_entry_t folio_realloc_swap(struct folio *folio) { swp_entry_t vswap_entry =3D folio->swap; struct swap_cluster_info *ci; - struct swap_cluster_info_dynamic *ci_dyn; struct mem_cgroup *memcg; unsigned int voff; unsigned long vt; @@ -2408,10 +2522,9 @@ swp_entry_t folio_realloc_swap(struct folio *folio) =20 voff =3D swp_cluster_offset(vswap_entry); ci =3D __swap_entry_to_cluster(vswap_entry); - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); =20 spin_lock(&ci->lock); - vt =3D __vtable_get(ci_dyn, voff); + vt =3D __vtable_get(ci, voff); if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { spin_unlock(&ci->lock); return vtable_to_phys(vt); @@ -2444,7 +2557,7 @@ swp_entry_t folio_realloc_swap(struct folio *folio) */ for (i =3D 0; i < nr; i++) { pe.val =3D phys_entry.val + i; - __vtable_set(ci_dyn, voff + i, vtable_mk_phys(pe)); + __vtable_set(ci, voff + i, vtable_mk_phys(pe)); } spin_unlock(&ci->lock); =20 @@ -2765,7 +2878,6 @@ static bool folio_maybe_swapped(struct folio *folio) VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio); =20 - /* Folio is locked and in swap cache, so ci->count > 0: cluster is alive.= */ ci =3D __swap_entry_to_cluster(entry); ci_off =3D swp_cluster_offset(entry); ci_end =3D ci_off + folio_nr_pages(folio); @@ -3881,25 +3993,22 @@ static void free_swap_cluster_info(struct swap_info= _struct *si, struct swap_cluster_info *cluster_info, unsigned long maxpages) { - struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *ci; - unsigned long idx; int i, nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); =20 if (swap_is_vswap(si)) { - xa_for_each(&si->cluster_info_pool, idx, ci_dyn) { - ci =3D &ci_dyn->ci; + nr_clusters =3D si->nr_mapped_clusters; + for (i =3D 0; i < nr_clusters; i++) { + ci =3D &si->cluster_info[i]; spin_lock(&ci->lock); if (cluster_table_is_alloced(ci)) { swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, true); - swap_cluster_free_table(ci); + swap_cluster_free_table(si, ci); } spin_unlock(&ci->lock); - vswap_cluster_free_vtable(ci); - kfree(ci_dyn); } - xa_destroy(&si->cluster_info_pool); + vswap_free_cluster_array(si); return; } =20 @@ -3911,7 +4020,7 @@ static void free_swap_cluster_info(struct swap_info_s= truct *si, spin_lock(&ci->lock); if (cluster_table_is_alloced(ci)) { swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, true); - swap_cluster_free_table(ci); + swap_cluster_free_table(si, ci); } spin_unlock(&ci->lock); } @@ -4397,39 +4506,18 @@ static int setup_swap_clusters_info(struct swap_inf= o_struct *si, { unsigned long nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); struct swap_cluster_info *cluster_info =3D NULL; - struct swap_cluster_info_dynamic *ci_dyn =3D NULL; + struct swap_cluster_info *ci; int err =3D -ENOMEM; unsigned long i; =20 - /* A vswap device uses an xarray pool instead of a static array. */ + /* A vswap device grows its cluster array on demand. */ if (swap_is_vswap(si)) { nr_clusters =3D 0; - xa_init_flags(&si->cluster_info_pool, XA_FLAGS_ALLOC); - - /* - * Pre-allocate cluster 0 and mark slot 0 (header page) - * as bad so the allocator never hands out page offset 0. - */ - ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_KERNEL); - if (!ci_dyn) - goto err; - spin_lock_init(&ci_dyn->ci.lock); - INIT_LIST_HEAD(&ci_dyn->ci.list); - - err =3D xa_insert(&si->cluster_info_pool, 0, ci_dyn, GFP_KERNEL); - if (err) { - kfree(ci_dyn); - goto err; - } - - err =3D swap_cluster_setup_bad_slot(si, &ci_dyn->ci, 0, false); + err =3D vswap_reserve_cluster_array(si, maxpages); if (err) goto err; - - err =3D vswap_cluster_alloc_vtable(ci_dyn, GFP_KERNEL); - if (err) - goto err; - + /* Reservation owns the array; keep the tail store idempotent. */ + cluster_info =3D si->cluster_info; goto setup_cluster_info; } =20 @@ -4487,7 +4575,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, } =20 for (i =3D 0; i < nr_clusters; i++) { - struct swap_cluster_info *ci =3D &cluster_info[i]; + ci =3D &cluster_info[i]; =20 if (ci->count) { ci->flags =3D CLUSTER_FLAG_NONFULL; @@ -4500,8 +4588,23 @@ static int setup_swap_clusters_info(struct swap_info= _struct *si, =20 /* Slot 0 is bad, so cluster 0 never empties. The rest of it is usable. */ if (swap_is_vswap(si)) { - ci_dyn->ci.flags =3D CLUSTER_FLAG_NONFULL; - list_add_tail(&ci_dyn->ci.list, &si->nonfull_clusters[0]); + err =3D vswap_grow_clusters(si); + if (err) + goto err; + + ci =3D si->cluster_info; + spin_lock(&ci->lock); + move_cluster(si, ci, NULL, CLUSTER_FLAG_NONE); + spin_unlock(&ci->lock); + + err =3D swap_cluster_setup_bad_slot(si, ci, 0, false); + if (err) + goto err; + + spin_lock(&ci->lock); + move_cluster(si, ci, &si->nonfull_clusters[0], + CLUSTER_FLAG_NONFULL); + spin_unlock(&ci->lock); } =20 si->cluster_info =3D cluster_info; diff --git a/mm/vswap.h b/mm/vswap.h index b79866d5999c..f8235882f3f0 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -78,7 +78,7 @@ static inline void swap_rmap_clear_cache_only(struct swap= _cluster_info *ci, /* * Virtual table entry encoding for vswap clusters. * - * Each entry in ci_dyn->virtual_table stores the backing type and + * Each entry in ci->virtual_table stores the backing type and * pointer for a virtual swap slot. Tag in low 3 bits, payload in * upper 61 bits. * @@ -109,7 +109,7 @@ static inline void swap_rmap_clear_cache_only(struct sw= ap_cluster_info *ci, * * Locking: a slot's vtable entry (the vswap entry's backend) is only * stable while the caller owns and holds the lock on that entry's swap - * cache folio. The cluster lock (ci_dyn->ci.lock) only makes an individual + * cache folio. The cluster lock (ci->lock) only makes an individual * vtable read atomic, and by itself does not give the caller the right to * change the backend. A backend read without the folio lock is * best-effort and must be re-validated under the folio lock before @@ -156,18 +156,18 @@ static inline struct zswap_entry *vtable_to_zswap(uns= igned long vt) =20 /* Virtual table accessors */ =20 -static inline unsigned long __vtable_get(struct swap_cluster_info_dynamic = *ci_dyn, +static inline unsigned long __vtable_get(struct swap_cluster_info *ci, unsigned int off) { VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); - return atomic_long_read(&ci_dyn->virtual_table[off]); + return atomic_long_read(&ci->virtual_table[off]); } =20 -static inline void __vtable_set(struct swap_cluster_info_dynamic *ci_dyn, +static inline void __vtable_set(struct swap_cluster_info *ci, unsigned int off, unsigned long vt) { VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); - atomic_long_set(&ci_dyn->virtual_table[off], vt); + atomic_long_set(&ci->virtual_table[off], vt); } =20 /** @@ -175,18 +175,13 @@ static inline void __vtable_set(struct swap_cluster_i= nfo_dynamic *ci_dyn, * @entry: the virtual swap entry * @voff: out param, receives @entry's slot offset within the cluster * - * Return: the locked vswap cluster, or NULL if @entry has no live cluster. + * Return: the locked vswap cluster. */ -static inline struct swap_cluster_info_dynamic * +static inline struct swap_cluster_info * vswap_lock_cluster(swp_entry_t entry, unsigned int *voff) { - struct swap_cluster_info *ci; - - ci =3D swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry)); - if (!ci) - return NULL; *voff =3D swp_cluster_offset(entry); - return container_of(ci, struct swap_cluster_info_dynamic, ci); + return swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry)); } =20 /** @@ -199,16 +194,13 @@ vswap_lock_cluster(swp_entry_t entry, unsigned int *v= off) */ static inline swp_entry_t vswap_to_phys(swp_entry_t entry) { - struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; unsigned int voff; unsigned long vt; =20 - ci_dyn =3D vswap_lock_cluster(entry, &voff); - if (!ci_dyn) - return (swp_entry_t){}; - - vt =3D __vtable_get(ci_dyn, voff); - swap_cluster_unlock(&ci_dyn->ci); + ci =3D vswap_lock_cluster(entry, &voff); + vt =3D __vtable_get(ci, voff); + swap_cluster_unlock(ci); =20 if (vtable_type(vt) !=3D VSWAP_SWAPFILE) return (swp_entry_t){}; @@ -232,13 +224,13 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, static inline void vswap_zswap_store(swp_entry_t entry, struct zswap_entry *ze) { - struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; unsigned int voff; =20 - ci_dyn =3D vswap_lock_cluster(entry, &voff); - __vswap_release_backing(&ci_dyn->ci, voff, 1); - __vtable_set(ci_dyn, voff, (unsigned long)ze | VSWAP_ZSWAP); - swap_cluster_unlock(&ci_dyn->ci); + ci =3D vswap_lock_cluster(entry, &voff); + __vswap_release_backing(ci, voff, 1); + __vtable_set(ci, voff, (unsigned long)ze | VSWAP_ZSWAP); + swap_cluster_unlock(ci); } =20 /** @@ -250,15 +242,13 @@ static inline void vswap_zswap_store(swp_entry_t entr= y, */ static inline struct zswap_entry *vswap_zswap_load(swp_entry_t entry) { - struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; unsigned int voff; unsigned long vt; =20 - ci_dyn =3D vswap_lock_cluster(entry, &voff); - if (!ci_dyn) - return NULL; - vt =3D __vtable_get(ci_dyn, voff); - swap_cluster_unlock(&ci_dyn->ci); + ci =3D vswap_lock_cluster(entry, &voff); + vt =3D __vtable_get(ci, voff); + swap_cluster_unlock(ci); =20 if (vtable_type(vt) !=3D VSWAP_ZSWAP) return NULL; @@ -270,7 +260,7 @@ swp_entry_t folio_realloc_swap(struct folio *folio); void folio_release_non_phys_swap_backing(struct folio *folio); =20 /* - * Walk nr vtable slots starting at voff in ci_dyn. Returns the prefix + * Walk nr vtable slots starting at voff in ci. Returns the prefix * length of slots sharing one effective backing type. For SWAPFILE, * the prefix is also restricted to contiguous offsets in the same * swapfile. @@ -283,9 +273,9 @@ void folio_release_non_phys_swap_backing(struct folio *= folio); * vtable=3DZSWAP -> VSWAP_ZSWAP * * *typep returns the effective type of slot 0. Caller holds - * ci_dyn->ci.lock. + * ci->lock. */ -static inline int __vswap_check_backing(struct swap_cluster_info_dynamic *= ci_dyn, +static inline int __vswap_check_backing(struct swap_cluster_info *ci, unsigned int voff, int nr, enum vswap_backing_type *typep) { @@ -295,13 +285,13 @@ static inline int __vswap_check_backing(struct swap_c= luster_info_dynamic *ci_dyn unsigned long vt, swap_tb; int i; =20 - lockdep_assert_held(&ci_dyn->ci.lock); + lockdep_assert_held(&ci->lock); =20 for (i =3D 0; i < nr; i++) { - vt =3D __vtable_get(ci_dyn, voff + i); + vt =3D __vtable_get(ci, voff + i); if (vtable_type(vt) =3D=3D VSWAP_NONE) { - swap_tb =3D __swap_table_get(&ci_dyn->ci, voff + i); - if (__swap_table_test_zero(&ci_dyn->ci, voff + i)) + swap_tb =3D __swap_table_get(ci, voff + i); + if (__swap_table_test_zero(ci, voff + i)) slot_type =3D VSWAP_ZERO; else if (swp_tb_is_folio(swap_tb)) slot_type =3D VSWAP_FOLIO; @@ -331,18 +321,13 @@ static inline int __vswap_check_backing(struct swap_c= luster_info_dynamic *ci_dyn static inline int vswap_check_backing(swp_entry_t entry, int nr, enum vswap_backing_type *typep) { - struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; unsigned int voff; int ret; =20 - ci_dyn =3D vswap_lock_cluster(entry, &voff); - if (!ci_dyn) { - if (typep) - *typep =3D VSWAP_NONE; - return 0; - } - ret =3D __vswap_check_backing(ci_dyn, voff, nr, typep); - swap_cluster_unlock(&ci_dyn->ci); + ci =3D vswap_lock_cluster(entry, &voff); + ret =3D __vswap_check_backing(ci, voff, nr, typep); + swap_cluster_unlock(ci); return ret; } =20 @@ -365,21 +350,18 @@ static inline bool folio_phys_swap_backed(struct foli= o *folio) type =3D=3D VSWAP_SWAPFILE); } =20 -static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn, +static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info *ci, gfp_t gfp) { - ci_dyn->virtual_table =3D kcalloc(SWAPFILE_CLUSTER, - sizeof(*ci_dyn->virtual_table), gfp); - return ci_dyn->virtual_table ? 0 : -ENOMEM; + ci->virtual_table =3D kcalloc(SWAPFILE_CLUSTER, + sizeof(*ci->virtual_table), gfp); + return ci->virtual_table ? 0 : -ENOMEM; } =20 static inline void vswap_cluster_free_vtable(struct swap_cluster_info *ci) { - struct swap_cluster_info_dynamic *ci_dyn; - - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); - kfree(ci_dyn->virtual_table); - ci_dyn->virtual_table =3D NULL; + kfree(ci->virtual_table); + ci->virtual_table =3D NULL; } =20 #else /* !CONFIG_SWAP */ --=20 2.53.0-Meta