From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oo1-f53.google.com (mail-oo1-f53.google.com [209.85.161.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE743395AD4 for ; Thu, 6 Aug 2026 18:42:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041782; cv=none; b=mNAgZ29Ss/oZqJA1Gf8eTVd1c6HKY6Rok0c+jnOyAp1e6lK7wmEJediASydKZjPkyGmwpK2ETpL6QGMIF6F0JnXSykEcp4RYFSRc/BhDlE8yGyKzjPNIqpUusOMWvlB25byhpPvAD51fnqOW4Hp7cHCSX1IEUgPMDhiyYpFvNAI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041782; c=relaxed/simple; bh=7WjEpUVtPDvbIdTqUuJtxvkGmoaHNOoPwe8qjG6fx5k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o+/7bRmujbP5F/hLecD85Lsaa9enJHYbqY8784RcsdL9puuPszqTiCR3ncHy42eAPYm4j762o9vAKcLiWSk1yl+xgSreByEyEGLDBV9p9t6hXE1M8MuL5r1fLSTKpxEpIQJJTgZSfdAdmBbgVMFqo1rxlcmLlZiQOoB3FIttEIE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Kt3y3RLb; arc=none smtp.client-ip=209.85.161.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Kt3y3RLb" Received: by mail-oo1-f53.google.com with SMTP id 006d021491bc7-6b01abe5d03so478962eaf.1 for ; Thu, 06 Aug 2026 11:42:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041777; x=1786646577; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7Z8Rr4fXv8HnfyepjCtSg1DMwHI70e7GbioRfg7aAs4=; b=Kt3y3RLbvBvNmovlTqVSsbimUzMjZrzkxC4QFiPgJjBC/6ipMsW8jJvySmZaWgr1RW 2P+gnh0ngizmEMYzXN29/83srXslR3ZnlJAZJ5GrA9Ehpf2NDlK87pd1f4ScLENa33Ci MkYHnOMjFSPBy1V6xgxO5OTDzPoSm+httiYNJtmgoowfibZq24fy8z93LkHlDt5Xz6gU hjeoAFLiJOprFQ98vMQfwd+rAu6zph+jF/2/0dNh5gHjez2zb0KL1w1I2RClVVQVlfQP rau6w4RO5mKzq0I0OHyktcnpEAFDRTaRZ0pW2AZD+7HnsOTGiKfTHQFz0npB0etdEsdU qzew== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041777; x=1786646577; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7Z8Rr4fXv8HnfyepjCtSg1DMwHI70e7GbioRfg7aAs4=; b=gA4+QCS9FDz27XQhaLvf6ymhms2cVgpvGTLqW/S7QnvOVvqiZBVCz6CIGxYmIOuxJg lU0Wj7jNGYYEimYgCfK0qI9GAFOzF6rISDMVwdPJXvCsyXxY+6lnW1G+ex0P9BfwDUJ4 sx+eRGKGEeav6DRRRte/ElAZXCYeTALmGgEQADOUB/md+utBScrURHsWV/Aj6ZyvBJSZ EC3VxD6VPfJwsLVFHrjhpuS8CghJLNkRfr9kuCZUKOcXj7XviGhUQmyPktY2xkxqfC5Q kb9RcT5LBQbGifqC7G13l9FVguvbd1CtSd/V2SmEZbQI5tWBpatnVeVDgajnkDHbv3fW jhbA== X-Forwarded-Encrypted: i=1; AHgh+RppwldqozbLt0Rz+Mem3UX67KTkewtcCq7SYa2L/v7SFUKZrr6uZUijY3IuOXIB9yeFCPfe4Q0lLVR/4wk=@vger.kernel.org X-Gm-Message-State: AOJu0Yx9bD/4TXy+EarqOpznP/6V91QGqLRFjNSqf6nQNUbL7X6K23YD CGyT4Zqq4umTKMwHUAVOCvwKL/veQMqUDKjHseehsJ6aSIGww7rf+5xE X-Gm-Gg: AR+sD13Ms+9rBK2R0kTDqEc9gbT9YIIed/29J6LIzZG+xaqK4ek7SV2ZIxSKEd/5h6y vqjxltiqFpKFLSxFz4IHL8z1pItCbS7vS8cXOduz1yxhw1hkRnYlNuzjEe0tGCpfdEu304LDTUI lxH3DFDAqcjiY7/Ck+u/q0+Er9o6wXefspYi8xBbCxUQmxhHDsim5dWJ7qUgNoY0gh7HuS+thEM WCxAUNsHP+sFcWGeOner5bbBs10PB1rtj2Y00c67VA24Ye+cf/+2OhIaniCvUwX7XuDKmdb3SVI x7OsG+btZortBv4m1LDqsxKikUcc13IuyffnBf4GSSS91Hx0UHZPMZDtDTc+VU+NDCV0dyvdnG+ D3HFz0A2mjDXVnh2ySuQ/pelRKKUdlMiKsYF10JHaGUQHLHqbf5kPIs/pnWBgOFpzFm/br/3nuE alUOpHXh9mhscyESz0zxX8czH1/R2aR0QfcVsogonqEhpISpxppw5RYNlr7+YB5S8G2a9nix9DH /m+qHqjWHrlycaHmK5d0g== X-Received: by 2002:a05:6820:4b8e:b0:6a1:4a27:cd22 with SMTP id 006d021491bc7-6ae968d0ddbmr9679694eaf.0.1786041777343; Thu, 06 Aug 2026 11:42:57 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:21::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02be475b6sm166741eaf.11.2026.08.06.11.42.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:42:56 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 01/11] mm, swap: add virtual swap device infrastructure Date: Thu, 6 Aug 2026 11:42:44 -0700 Message-ID: <20260806184254.3790858-2-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Create a 16 TB virtual swap device at boot, along with the dynamic cluster infrastructure that the rest of the vswap layer is built on. swap_cluster_info_dynamic keeps per-cluster info in an xarray, so a device can be sized without a static cluster_info[] array. Gated by a new CONFIG_VSWAP (depends on SWAP && 64BIT). For now the vswap device cannot be swapon'd or swapoff'd. It is created unconditionally at boot when CONFIG_VSWAP=3Dy and lives for the lifetime of the kernel. The SWP_VSWAP flag and swap_is_vswap() helper let hot paths skip per-device bookkeeping that doesn't apply (avail-list management, percpu_ref get/put, hibernation target lookup, etc.). This patch is pure scaffolding. It wires the dynamic-cluster allocator into cluster_alloc_swap_entry (via an SWP_VSWAP branch that dispatches to alloc_swap_scan_dynamic), but the branch is not yet reachable because vswap_si is kept off swap_avail_head and swap_active_head and folio_alloc_swap has no path that calls into vswap_si directly. Backends (zswap, zero, physical disk) and the vswap-aware swap-out / swap-in / writeback paths arrive in subsequent patches. Suggested-by: Kairui Song Co-developed-by: Kairui Song Signed-off-by: Kairui Song Signed-off-by: Nhat Pham --- MAINTAINERS | 1 + include/linux/swap.h | 16 +++ mm/Kconfig | 10 ++ mm/page_io.c | 15 +++ mm/swap.h | 47 ++++++-- mm/swap_state.c | 41 ++++--- mm/swap_table.h | 2 + mm/swapfile.c | 270 +++++++++++++++++++++++++++++++++++++++---- mm/vswap.h | 31 +++++ mm/zswap.c | 6 + 10 files changed, 396 insertions(+), 43 deletions(-) create mode 100644 mm/vswap.h diff --git a/MAINTAINERS b/MAINTAINERS index e9c8567308a7..d0da9a29a910 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -17248,6 +17248,7 @@ F: mm/swap.h F: mm/swap_table.h F: mm/swap_state.c F: mm/swapfile.c +F: mm/vswap.h =20 MEMORY MANAGEMENT - THP (TRANSPARENT HUGE PAGE) M: Andrew Morton diff --git a/include/linux/swap.h b/include/linux/swap.h index 45f301d73e2a..a955bd60dd58 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -207,6 +207,7 @@ enum { SWP_STABLE_WRITES =3D (1 << 11), /* no overwrite PG_writeback pages */ SWP_SYNCHRONOUS_IO =3D (1 << 12), /* synchronous IO is efficient */ SWP_HIBERNATION =3D (1 << 13), /* pinned for hibernation */ + SWP_VSWAP =3D (1 << 14), /* virtual swap device */ /* add others here before... */ }; =20 @@ -276,8 +277,21 @@ struct swap_info_struct { struct list_head discard_clusters; /* discard clusters list */ struct plist_node avail_list; /* entry in swap_avail_head */ const struct swap_ops *ops; + struct xarray cluster_info_pool; /* Xarray for vswap dynamic cluster info= */ }; =20 +#ifdef CONFIG_VSWAP +static inline bool swap_is_vswap(struct swap_info_struct *si) +{ + return si->flags & SWP_VSWAP; +} +#else +static inline bool swap_is_vswap(struct swap_info_struct *si) +{ + return false; +} +#endif + static inline swp_entry_t page_swap_entry(struct page *page) { struct folio *folio =3D page_folio(page); @@ -402,6 +416,8 @@ void swap_free_hibernation_slot(swp_entry_t entry); =20 static inline void put_swap_device(struct swap_info_struct *si) { + if (swap_is_vswap(si)) + return; percpu_ref_put(&si->users); } =20 diff --git a/mm/Kconfig b/mm/Kconfig index 331daf7fcfab..32d38b552845 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -19,6 +19,16 @@ menuconfig SWAP used to provide more virtual memory than the actual RAM present in your computer. If unsure say Y. =20 +config VSWAP + bool "Virtual swap device" + depends on SWAP && 64BIT + help + Adds a virtual swap layer that decouples swap entries in page + tables from physical backing storage. Swap entries are allocated + from a virtual swap device and can be backed by zswap, a physical + swapfile, or kept in memory - with the backing changeable at + runtime without invalidating page table entries. + config ZSWAP bool "Compressed cache for swap pages" depends on SWAP diff --git a/mm/page_io.c b/mm/page_io.c index e4fa7ffffe8b..fca1718056af 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -27,6 +27,7 @@ #include #include "swap.h" #include "swap_table.h" +#include "vswap.h" =20 int generic_swapfile_activate(struct swap_info_struct *sis, struct file *swap_file, @@ -247,6 +248,15 @@ int swap_writeout(struct swap_io_ctx *ctx, struct foli= o *folio) } rcu_read_unlock(); =20 + /* + * A vswap folio that reaches here could not be stored to a backend + * (zswap) and has no physical slot to write to, so keep it dirty. + */ + if (is_vswap_entry(folio->swap)) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writepage(ctx, folio); return 0; out_unlock: @@ -479,6 +489,11 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct f= olio *folio) if (zswap_load(folio) !=3D -ENOENT) goto finish; =20 + if (unlikely(swap_is_vswap(sis))) { + folio_unlock(folio); + goto finish; + } + /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); swap_add_folio(ctx, folio, READ); diff --git a/mm/swap.h b/mm/swap.h index ec580c713204..b593ad3214ef 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -66,6 +66,12 @@ struct swap_cluster_info { struct list_head list; }; =20 +struct swap_cluster_info_dynamic { + struct swap_cluster_info ci; + unsigned int index; /* for cluster_index() */ + struct rcu_head rcu; +}; + /* All on-list cluster must have a non-zero flag. */ enum swap_cluster_flags { CLUSTER_FLAG_NONE =3D 0, /* For temporary off-list cluster */ @@ -76,6 +82,7 @@ enum swap_cluster_flags { CLUSTER_FLAG_USABLE =3D CLUSTER_FLAG_FRAG, CLUSTER_FLAG_FULL, CLUSTER_FLAG_DISCARD, + CLUSTER_FLAG_DEAD, /* Vswap dynamic cluster pending kfree_rcu */ CLUSTER_FLAG_MAX, }; =20 @@ -143,9 +150,19 @@ static inline struct swap_info_struct *__swap_entry_to= _info(swp_entry_t entry) static inline struct swap_cluster_info *__swap_offset_to_cluster( struct swap_info_struct *si, pgoff_t offset) { + unsigned int cluster_idx =3D offset / SWAPFILE_CLUSTER; + VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ VM_WARN_ON_ONCE(offset >=3D roundup(si->max, SWAPFILE_CLUSTER)); - return &si->cluster_info[offset / SWAPFILE_CLUSTER]; + + if (swap_is_vswap(si)) { + struct swap_cluster_info_dynamic *ci_dyn; + + ci_dyn =3D xa_load(&si->cluster_info_pool, cluster_idx); + return ci_dyn ? &ci_dyn->ci : NULL; + } + + return &si->cluster_info[cluster_idx]; } =20 static inline struct swap_cluster_info *__swap_entry_to_cluster(swp_entry_= t entry) @@ -157,7 +174,7 @@ static inline struct swap_cluster_info *__swap_entry_to= _cluster(swp_entry_t entr static __always_inline struct swap_cluster_info *__swap_cluster_lock( struct swap_info_struct *si, unsigned long offset, bool irq) { - struct swap_cluster_info *ci =3D __swap_offset_to_cluster(si, offset); + struct swap_cluster_info *ci; =20 /* * Nothing modifies swap cache in an IRQ context. All access to @@ -170,20 +187,36 @@ static __always_inline struct swap_cluster_info *__sw= ap_cluster_lock( */ VM_WARN_ON_ONCE(!in_task()); VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ - if (irq) - spin_lock_irq(&ci->lock); - else - spin_lock(&ci->lock); + + rcu_read_lock(); + ci =3D __swap_offset_to_cluster(si, offset); + if (ci) { + if (irq) + spin_lock_irq(&ci->lock); + else + spin_lock(&ci->lock); + + if (ci->flags =3D=3D CLUSTER_FLAG_DEAD) { + if (irq) + spin_unlock_irq(&ci->lock); + else + spin_unlock(&ci->lock); + ci =3D NULL; + } + } + rcu_read_unlock(); return ci; } =20 /** * swap_cluster_lock - Lock and return the swap cluster of given offset. * @si: swap device the cluster belongs to. - * @offset: the swap entry offset, pointing to a valid slot. + * @offset: the swap entry offset. * * Context: The caller must ensure the offset is in the valid range and * protect the swap device with reference count or locks. + * Return: the locked cluster, or NULL if it is gone. Only a vswap device + * can return NULL, as its clusters are allocated and freed on demand. */ static inline struct swap_cluster_info *swap_cluster_lock( struct swap_info_struct *si, unsigned long offset) diff --git a/mm/swap_state.c b/mm/swap_state.c index 5be825911e64..9e0d71fcdc24 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -95,8 +95,10 @@ struct folio *swap_cache_get_folio(swp_entry_t entry) struct folio *folio; =20 for (;;) { + rcu_read_lock(); swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), swp_cluster_offset(entry)); + rcu_read_unlock(); if (!swp_tb_is_folio(swp_tb)) return NULL; folio =3D swp_tb_to_folio(swp_tb); @@ -118,8 +120,10 @@ bool swap_cache_has_folio(swp_entry_t entry) { unsigned long swp_tb; =20 + rcu_read_lock(); swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), swp_cluster_offset(entry)); + rcu_read_unlock(); return swp_tb_is_folio(swp_tb); } =20 @@ -135,8 +139,10 @@ void *swap_cache_get_shadow(swp_entry_t entry) { unsigned long swp_tb; =20 + rcu_read_lock(); swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), swp_cluster_offset(entry)); + rcu_read_unlock(); if (swp_tb_is_shadow(swp_tb)) return swp_tb_to_shadow(swp_tb); return NULL; @@ -405,14 +411,16 @@ void __swap_cache_replace_folio(struct swap_cluster_i= nfo *ci, * -ENOENT / -EEXIST: Target swap entry is unavailable or cached, the call= er * should abort or try to use the cached folio instead */ -static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci, - swp_entry_t targ_entry, gfp_t gfp, +static struct folio *__swap_cache_alloc(swp_entry_t targ_entry, gfp_t gfp, unsigned int order, struct vm_fault *vmf, struct mempolicy *mpol, pgoff_t ilx) { int err; swp_entry_t entry; struct folio *folio; + struct swap_cluster_info *ci; + struct swap_info_struct *si =3D __swap_entry_to_info(targ_entry); + unsigned long offset =3D swp_offset(targ_entry); void *shadow =3D NULL; unsigned short memcg_id; unsigned long address, nr_pages =3D 1UL << order; @@ -422,9 +430,12 @@ static struct folio *__swap_cache_alloc(struct swap_cl= uster_info *ci, entry.val =3D round_down(targ_entry.val, nr_pages); =20 /* Check if the slot and range are available, skip allocation if not */ - spin_lock(&ci->lock); - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); - spin_unlock(&ci->lock); + err =3D -ENOENT; + ci =3D swap_cluster_lock(si, offset); + if (ci) { + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, NULL, NULL); + swap_cluster_unlock(ci); + } if (unlikely(err)) return ERR_PTR(err); =20 @@ -445,10 +456,13 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, return ERR_PTR(-ENOMEM); =20 /* Double check the range is still not in conflict */ - spin_lock(&ci->lock); - err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg_= id); + err =3D -ENOENT; + ci =3D swap_cluster_lock(si, offset); + if (ci) + err =3D __swap_cache_add_check(ci, targ_entry, nr_pages, &shadow, &memcg= _id); if (unlikely(err)) { - spin_unlock(&ci->lock); + if (ci) + swap_cluster_unlock(ci); folio_put(folio); return ERR_PTR(err); } @@ -456,13 +470,14 @@ static struct folio *__swap_cache_alloc(struct swap_c= luster_info *ci, __folio_set_locked(folio); __folio_set_swapbacked(folio); __swap_cache_do_add_folio(ci, folio, entry); - spin_unlock(&ci->lock); + swap_cluster_unlock(ci); =20 if (mem_cgroup_swapin_charge_folio(folio, memcg_id, vmf ? vmf->vma->vm_mm : NULL, gfp)) { - spin_lock(&ci->lock); + /* The folio pins the cluster */ + ci =3D swap_cluster_lock(si, offset); __swap_cache_do_del_folio(ci, folio, entry, shadow); - spin_unlock(&ci->lock); + swap_cluster_unlock(ci); folio_unlock(folio); /* nr_pages refs from swap cache, 1 from allocation */ folio_put_refs(folio, nr_pages + 1); @@ -516,9 +531,7 @@ struct folio *swap_cache_alloc_folio(swp_entry_t targ_e= ntry, gfp_t gfp, { int order, err; struct folio *ret; - struct swap_cluster_info *ci; =20 - ci =3D __swap_entry_to_cluster(targ_entry); order =3D highest_order(orders); =20 /* orders must be non-zero, and must not exceed cluster size. */ @@ -526,7 +539,7 @@ struct folio *swap_cache_alloc_folio(swp_entry_t targ_e= ntry, gfp_t gfp, return ERR_PTR(-EINVAL); =20 do { - ret =3D __swap_cache_alloc(ci, targ_entry, gfp, order, + ret =3D __swap_cache_alloc(targ_entry, gfp, order, vmf, mpol, ilx); if (!IS_ERR(ret)) break; diff --git a/mm/swap_table.h b/mm/swap_table.h index e6613e62f8d0..fd7f0fb9836a 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -255,6 +255,8 @@ static inline unsigned long swap_table_get(struct swap_= cluster_info *ci, unsigned long swp_tb; =20 VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); + if (!ci) + return SWP_TB_NULL; =20 rcu_read_lock(); table =3D rcu_dereference(ci->table); diff --git a/mm/swapfile.c b/mm/swapfile.c index 4d4e3e3059f6..fea3a8eccbc1 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -42,10 +42,12 @@ #include #include #include +#include =20 #include #include #include "swap_table.h" +#include "vswap.h" #include "internal.h" #include "swap.h" =20 @@ -401,6 +403,8 @@ static inline bool cluster_is_usable(struct swap_cluste= r_info *ci, int order) static inline unsigned int cluster_index(struct swap_info_struct *si, struct swap_cluster_info *ci) { + if (swap_is_vswap(si)) + return container_of(ci, struct swap_cluster_info_dynamic, ci)->index; return ci - si->cluster_info; } =20 @@ -712,6 +716,34 @@ static void swap_users_ref_free(struct percpu_ref *ref) complete(&si->comp); } =20 +#ifdef CONFIG_VSWAP +static void vswap_free_cluster(struct swap_info_struct *si, + struct swap_cluster_info *ci) +{ + struct swap_cluster_info_dynamic *ci_dyn; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + if (ci->flags !=3D CLUSTER_FLAG_NONE) { + spin_lock(&si->lock); + list_del(&ci->list); + spin_unlock(&si->lock); + } + swap_cluster_free_table(ci); + /* + * Ordering vs the RCU cluster lookup: erase from the xarray first + * (new lookups miss it), mark DEAD under the held ci->lock (a lookup + * that already has ci sees DEAD on relock and bails), then kfree_rcu + * so the cluster outlives any reader still in its RCU section. + */ + xa_erase(&si->cluster_info_pool, ci_dyn->index); + ci->flags =3D CLUSTER_FLAG_DEAD; + kfree_rcu(ci_dyn, rcu); +} +#else +static inline void vswap_free_cluster(struct swap_info_struct *si, + struct swap_cluster_info *ci) {} +#endif + /* * Must be called after freeing if ci->count =3D=3D 0, moves the cluster t= o free * or discard list. @@ -733,6 +765,11 @@ static void free_cluster(struct swap_info_struct *si, = struct swap_cluster_info * return; } =20 + if (swap_is_vswap(si)) { + vswap_free_cluster(si, ci); + return; + } + __free_cluster(si, ci); } =20 @@ -835,14 +872,21 @@ static int swap_cluster_setup_bad_slot(struct swap_in= fo_struct *si, * stolen by a lower order). @usable will be set to false if that happens. */ static bool cluster_reclaim_range(struct swap_info_struct *si, - struct swap_cluster_info *ci, + struct swap_cluster_info **pcip, unsigned long start, unsigned int order, bool *usable) { + struct swap_cluster_info *ci =3D *pcip; unsigned int nr_pages =3D 1 << order; unsigned long offset =3D start, end =3D start + nr_pages; unsigned long swp_tb; =20 + /* + * Take RCU read lock before releasing the cluster lock to keep ci + * alive - for vswap dynamic clusters, ci is freed via kfree_rcu + * and the grace period could otherwise elapse in the window. + */ + rcu_read_lock(); spin_unlock(&ci->lock); do { swp_tb =3D swap_table_get(ci, offset % SWAPFILE_CLUSTER); @@ -852,7 +896,15 @@ static bool cluster_reclaim_range(struct swap_info_str= uct *si, if (__try_to_reclaim_swap(si, offset, TTRS_ANYWAY) < 0) break; } while (++offset < end); - spin_lock(&ci->lock); + rcu_read_unlock(); + + /* Re-lookup: dynamic cluster may have been freed while lock was dropped = */ + ci =3D swap_cluster_lock(si, start); + *pcip =3D ci; + if (!ci) { + *usable =3D false; + return false; + } =20 /* * We just dropped ci->lock so cluster could be used by another @@ -983,7 +1035,8 @@ static unsigned int alloc_swap_scan_cluster(struct swa= p_info_struct *si, if (!cluster_scan_range(si, ci, offset, nr_pages, &need_reclaim)) continue; if (need_reclaim) { - ret =3D cluster_reclaim_range(si, ci, offset, order, &usable); + ret =3D cluster_reclaim_range(si, &ci, offset, order, + &usable); if (!usable) goto out; if (cluster_is_empty(ci)) @@ -1001,8 +1054,10 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, break; } out: - relocate_cluster(si, ci); - swap_cluster_unlock(ci); + if (ci) { + relocate_cluster(si, ci); + swap_cluster_unlock(ci); + } if (si->flags & SWP_SOLIDSTATE) { this_cpu_write(percpu_swap_cluster.offset[order], next); this_cpu_write(percpu_swap_cluster.si[order], si); @@ -1034,6 +1089,41 @@ static unsigned int alloc_swap_scan_list(struct swap= _info_struct *si, return found; } =20 +static unsigned int alloc_swap_scan_dynamic(struct swap_info_struct *si, + struct folio *folio) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *ci; + unsigned long offset; + + VM_WARN_ON(!swap_is_vswap(si)); + + ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_ATOMIC); + if (!ci_dyn) + return SWAP_ENTRY_INVALID; + + spin_lock_init(&ci_dyn->ci.lock); + INIT_LIST_HEAD(&ci_dyn->ci.list); + + if (swap_cluster_alloc_table(&ci_dyn->ci, GFP_ATOMIC)) { + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + + if (xa_alloc(&si->cluster_info_pool, &ci_dyn->index, ci_dyn, + XA_LIMIT(1, DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER) - 1), + GFP_ATOMIC)) { + swap_cluster_free_table(&ci_dyn->ci); + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + + ci =3D &ci_dyn->ci; + spin_lock(&ci->lock); + offset =3D cluster_offset(si, ci); + return alloc_swap_scan_cluster(si, ci, folio, offset); +} + static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool f= orce) { long to_scan =3D 1; @@ -1056,7 +1146,9 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) spin_unlock(&ci->lock); nr_reclaim =3D __try_to_reclaim_swap(si, offset, TTRS_ANYWAY); - spin_lock(&ci->lock); + ci =3D swap_cluster_lock(si, offset); + if (!ci) + goto next; if (nr_reclaim) { offset +=3D abs(nr_reclaim); continue; @@ -1070,6 +1162,7 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) relocate_cluster(si, ci); =20 swap_cluster_unlock(ci); +next: if (to_scan <=3D 0) break; =20 @@ -1146,6 +1239,12 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, goto done; } =20 + if (swap_is_vswap(si)) { + found =3D alloc_swap_scan_dynamic(si, folio); + if (found) + goto done; + } + if (!(si->flags & SWP_PAGE_DISCARD)) { found =3D alloc_swap_scan_list(si, &si->free_clusters, folio, false); if (found) @@ -1264,6 +1363,13 @@ static void add_to_avail_list(struct swap_info_struc= t *si, bool swapon) goto skip; } =20 + /* + * Keep vswap off the avail list - it is not allocated from by + * the physical swap allocator (swap_alloc_fast/slow). + */ + if (swap_is_vswap(si)) + goto skip; + plist_add(&si->avail_list, &swap_avail_head); =20 skip: @@ -1280,10 +1386,10 @@ static bool swap_usage_add(struct swap_info_struct = *si, unsigned int nr_entries) long val =3D atomic_long_add_return_relaxed(nr_entries, &si->inuse_pages); =20 /* - * If device is full, and SWAP_USAGE_OFFLIST_BIT is not set, - * remove it from the plist. + * If device is full, and SWAP_USAGE_OFFLIST_BIT is not set, remove it + * from the plist. Vswap is never on the avail list, so skip it. */ - if (unlikely(val =3D=3D si->pages)) { + if (unlikely(val =3D=3D si->pages) && !swap_is_vswap(si)) { del_from_avail_list(si, false); return true; } @@ -1296,10 +1402,10 @@ static void swap_usage_sub(struct swap_info_struct = *si, unsigned int nr_entries) long val =3D atomic_long_sub_return_relaxed(nr_entries, &si->inuse_pages); =20 /* - * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, - * add it to the plist. + * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, add it to + * the plist. Vswap is never on the avail list, so skip it. */ - if (unlikely(val & SWAP_USAGE_OFFLIST_BIT)) + if (unlikely(val & SWAP_USAGE_OFFLIST_BIT) && !swap_is_vswap(si)) add_to_avail_list(si, false); } =20 @@ -1346,6 +1452,10 @@ static void swap_range_free(struct swap_info_struct = *si, unsigned long offset, =20 static bool get_swap_device_info(struct swap_info_struct *si) { + /* vswap device is always alive - no ref counting needed */ + if (swap_is_vswap(si)) + return true; + if (!percpu_ref_tryget_live(&si->users)) return false; /* @@ -1381,11 +1491,11 @@ static bool swap_alloc_fast(struct folio *folio) return false; =20 ci =3D swap_cluster_lock(si, offset); - if (cluster_is_usable(ci, order)) { + if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); alloc_swap_scan_cluster(si, ci, folio, offset); - } else { + } else if (ci) { swap_cluster_unlock(ci); } =20 @@ -1507,6 +1617,7 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t g= fp) if (!si) return 0; =20 + /* Entry is in use (being faulted in), so its cluster is alive. */ ci =3D __swap_offset_to_cluster(si, offset); ret =3D swap_extend_table_alloc(si, ci, swp_cluster_offset(entry), gfp); =20 @@ -1742,6 +1853,7 @@ int folio_alloc_swap(struct folio *folio) unsigned int order =3D folio_order(folio); unsigned int size =3D 1 << order; =20 + VM_WARN_ON_FOLIO(folio_test_swapcache(folio), folio); VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); VM_BUG_ON_FOLIO(!folio_test_uptodate(folio), folio); =20 @@ -1904,7 +2016,8 @@ struct swap_info_struct *get_swap_device(swp_entry_t = entry) return NULL; put_out: pr_err("%s: %s%08lx\n", __func__, Bad_offset, entry.val); - percpu_ref_put(&si->users); + if (!swap_is_vswap(si)) + percpu_ref_put(&si->users); return NULL; } =20 @@ -2036,6 +2149,7 @@ static bool folio_maybe_swapped(struct folio *folio) VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio); =20 + /* Folio is locked and in swap cache, so ci->count > 0: cluster is alive.= */ ci =3D __swap_entry_to_cluster(entry); ci_off =3D swp_cluster_offset(entry); ci_end =3D ci_off + folio_nr_pages(folio); @@ -2223,6 +2337,9 @@ static int __find_hibernation_swap_type(dev_t device,= sector_t offset) =20 if (!(sis->flags & SWP_WRITEOK)) continue; + /* vswap has no bdev - never a hibernation target */ + if (swap_is_vswap(sis)) + continue; =20 if (device =3D=3D sis->bdev->bd_dev) { struct swap_extent *se =3D first_se(sis); @@ -2349,6 +2466,9 @@ int find_first_swap(dev_t *device) =20 if (!(sis->flags & SWP_WRITEOK)) continue; + /* vswap has no bdev - never a hibernation target */ + if (swap_is_vswap(sis)) + continue; *device =3D sis->bdev->bd_dev; spin_unlock(&swap_lock); return type; @@ -2565,8 +2685,10 @@ static int unuse_pte_range(struct vm_area_struct *vm= a, pmd_t *pmd, &vmf); } if (!folio) { + rcu_read_lock(); swp_tb =3D swap_table_get(__swap_entry_to_cluster(entry), swp_cluster_offset(entry)); + rcu_read_unlock(); if (swp_tb_get_count(swp_tb) <=3D 0) continue; return -ENOMEM; @@ -2712,8 +2834,10 @@ static unsigned int find_next_to_unuse(struct swap_i= nfo_struct *si, * allocations from this area (while holding swap_lock). */ for (i =3D prev + 1; i < si->max; i++) { + rcu_read_lock(); swp_tb =3D swap_table_get(__swap_offset_to_cluster(si, i), i % SWAPFILE_CLUSTER); + rcu_read_unlock(); if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb)) break; if ((i % LATENCY_LIMIT) =3D=3D 0) @@ -2952,6 +3076,11 @@ static int setup_swap_extents(struct swap_info_struc= t *sis, struct inode *inode =3D mapping->host; int ret; =20 + if (swap_is_vswap(sis)) { + *span =3D 0; + return 0; + } + ret =3D sio_pool_init(); if (ret) return ret; @@ -2977,15 +3106,24 @@ static int setup_swap_extents(struct swap_info_stru= ct *sis, =20 static void _enable_swap_info(struct swap_info_struct *si) { - atomic_long_add(si->pages, &nr_swap_pages); - total_swap_pages +=3D si->pages; + if (!swap_is_vswap(si)) { + atomic_long_add(si->pages, &nr_swap_pages); + total_swap_pages +=3D si->pages; + } =20 assert_spin_locked(&swap_lock); =20 - plist_add(&si->list, &swap_active_head); + /* + * Vswap has no backing file and no swapoff support - keep it + * off swap_active_head (used by swapoff filename lookup and + * swap_sync_discard) and swap_avail_head (physical allocator). + */ + if (!swap_is_vswap(si)) { + plist_add(&si->list, &swap_active_head); =20 - /* Add back to available list */ - add_to_avail_list(si, true); + /* Add back to available list */ + add_to_avail_list(si, true); + } } =20 /* @@ -3022,6 +3160,8 @@ static void wait_for_allocation(struct swap_info_stru= ct *si) struct swap_cluster_info *ci; =20 BUG_ON(si->flags & SWP_WRITEOK); + if (swap_is_vswap(si)) + return; =20 for (offset =3D 0; offset < end; offset +=3D SWAPFILE_CLUSTER) { ci =3D swap_cluster_lock(si, offset); @@ -3528,10 +3668,43 @@ static int setup_swap_clusters_info(struct swap_inf= o_struct *si, unsigned long maxpages) { unsigned long nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); - struct swap_cluster_info *cluster_info; + struct swap_cluster_info *cluster_info =3D NULL; + struct swap_cluster_info_dynamic *ci_dyn; int err =3D -ENOMEM; unsigned long i; =20 + /* For SWP_VSWAP files, initialize Xarray pool instead of static array */ + if (swap_is_vswap(si)) { + /* + * Pre-allocate cluster 0 and mark slot 0 (header page) + * as bad so the allocator never hands out page offset 0. + */ + ci_dyn =3D kzalloc_obj(*ci_dyn, GFP_KERNEL); + if (!ci_dyn) + goto err; + spin_lock_init(&ci_dyn->ci.lock); + INIT_LIST_HEAD(&ci_dyn->ci.list); + + nr_clusters =3D 0; + xa_init_flags(&si->cluster_info_pool, XA_FLAGS_ALLOC); + err =3D xa_insert(&si->cluster_info_pool, 0, ci_dyn, GFP_KERNEL); + if (err) { + kfree(ci_dyn); + goto err; + } + + err =3D swap_cluster_setup_bad_slot(si, &ci_dyn->ci, 0, false); + if (err) { + xa_erase(&si->cluster_info_pool, 0); + swap_cluster_free_table(&ci_dyn->ci); + kfree(ci_dyn); + xa_destroy(&si->cluster_info_pool); + goto err; + } + + goto setup_cluster_info; + } + cluster_info =3D kvzalloc_objs(*cluster_info, nr_clusters); if (!cluster_info) goto err; @@ -3556,6 +3729,10 @@ static int setup_swap_clusters_info(struct swap_info= _struct *si, err =3D swap_cluster_setup_bad_slot(si, cluster_info, 0, false); if (err) goto err; + + if (!swap_header) + goto setup_cluster_info; + for (i =3D 0; i < swap_header->info.nr_badpages; i++) { unsigned int page_nr =3D swap_header->info.badpages[i]; =20 @@ -3575,6 +3752,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, goto err; } =20 +setup_cluster_info: INIT_LIST_HEAD(&si->free_clusters); INIT_LIST_HEAD(&si->full_clusters); INIT_LIST_HEAD(&si->discard_clusters); @@ -3611,7 +3789,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) struct dentry *dentry; int prio; int error; - union swap_header *swap_header; + union swap_header *swap_header =3D NULL; int nr_extents; sector_t span; unsigned long maxpages; @@ -3949,3 +4127,51 @@ static int __init swapfile_init(void) return 0; } subsys_initcall(swapfile_init); + +#ifdef CONFIG_VSWAP +struct swap_info_struct *vswap_si; + +/* vswap does no IO on its own. */ +static const struct swap_ops vswap_ops =3D { }; + +static int __init vswap_init(void) +{ + struct swap_info_struct *si; + unsigned long maxpages; + int err; + + si =3D alloc_swap_info(); + if (IS_ERR(si)) + return PTR_ERR(si); + + maxpages =3D min(swapfile_maximum_size, + ALIGN_DOWN((unsigned long)UINT_MAX, SWAPFILE_CLUSTER)); + si->flags |=3D SWP_VSWAP | SWP_SOLIDSTATE | SWP_WRITEOK; + si->ops =3D &vswap_ops; + si->bdev =3D NULL; + si->max =3D maxpages; + si->pages =3D maxpages - 1; + si->prio =3D SHRT_MAX; + si->list.prio =3D -si->prio; + si->avail_list.prio =3D -si->prio; + + err =3D setup_swap_clusters_info(si, NULL, maxpages); + if (err) + goto fail; + + mutex_lock(&swapon_mutex); + enable_swap_info(si); + mutex_unlock(&swapon_mutex); + + vswap_si =3D si; + pr_info("vswap: created virtual swap device (%lu pages)\n", maxpages); + return 0; + +fail: + spin_lock(&swap_lock); + si->flags =3D 0; + spin_unlock(&swap_lock); + return err; +} +late_initcall(vswap_init); +#endif diff --git a/mm/vswap.h b/mm/vswap.h new file mode 100644 index 000000000000..5641692f5be3 --- /dev/null +++ b/mm/vswap.h @@ -0,0 +1,31 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Virtual swap space + * + * Copyright (C) 2026 Nhat Pham + */ +#ifndef _MM_VSWAP_H +#define _MM_VSWAP_H + +#include +#include "swap.h" + +#ifdef CONFIG_VSWAP + +extern struct swap_info_struct *vswap_si; + +static inline bool is_vswap_entry(swp_entry_t entry) +{ + return swap_is_vswap(__swap_entry_to_info(entry)); +} + +#else + +static inline bool is_vswap_entry(swp_entry_t entry) +{ + return false; +} + +#endif /* CONFIG_VSWAP */ + +#endif /* _MM_VSWAP_H */ diff --git a/mm/zswap.c b/mm/zswap.c index f7c9c89f6449..354bf8bd7482 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1000,6 +1000,12 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, if (!si) return -EEXIST; =20 + /* Vswap entries have no physical backing to write to. */ + if (swap_is_vswap(si)) { + put_swap_device(si); + return -EINVAL; + } + mpol =3D get_task_policy(current); folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, NO_INTERLEAVE_INDEX); --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oo1-f50.google.com (mail-oo1-f50.google.com [209.85.161.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 19FF93ADBA5 for ; Thu, 6 Aug 2026 18:43:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041784; cv=none; b=IZ/6d6U1EhXNP039faKFqE7x9S/3UJME0JmYFL6Spz47J4355W8Y4vhNB/liDV5UOyFFoFl5ubXsD0r2dUIJtioHYtK8Bya+Jbk5I5GXoeXZP+hEFGiJvr6x0Vq6ByYvxwZWsZwM3hLo/8aMik6AX+0DuHsEVx2xl7bB/ccinoI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041784; c=relaxed/simple; bh=Xlrj4d4m4FjdXOJ7jfpG6cNLD9R9wjDtkY64UqMTSks=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KKsbK/Sl17ENVhWpF1dlXBpNw4SLQIsf4x6Xih4849sMj9ifT3rTtD+KvNUopLlQ8cg4R/nR7zV2ICUPTrEEB0NQjd4KPUlozCuyXXuezgLa+cOpmAc9rSULZF1pHdn02nKwtpGRIi85tAbefUhFVXE1mM/EDp1orHwQH7ahcds= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=lzaALNys; arc=none smtp.client-ip=209.85.161.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="lzaALNys" Received: by mail-oo1-f50.google.com with SMTP id 006d021491bc7-6ae534c2aadso1782763eaf.3 for ; Thu, 06 Aug 2026 11:43:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041780; x=1786646580; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Lo0KJy4ET0eLH/vHkWIDi1pEtJHPb1YP4byd4zqxg8s=; b=lzaALNysMar9hC4CZFkQWflgk8UAdMvUc1+eWqLOaYVWdOhCD7u7tq/o61/+YiYkrz VBbSbe2wkXHPrwyBiAQWFVRrT62zRu9WgEClgXhbiPeIPVrcQNvlMniyQuKcgZWpKq6e emyOmUArdoQxcJjVQYVgGAZ/wCAcxV67gAqWIRc+D/RWO6LIREQ8zsuCYs/at0wgw4jb eEXoK0u2L23E9f4iW9U84ePR53SiCKqVogfuh085/8VjDH9Rv7xVz1YMazIevabljr+I bBJvIMZwVSDz0t5oSwuQjuiTmbTw6yqFTatYD0+yiKwGQgqdCKYVWljooqWMY1vPS9uw a23Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041780; x=1786646580; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Lo0KJy4ET0eLH/vHkWIDi1pEtJHPb1YP4byd4zqxg8s=; b=igBusg/x8KZOtz+klAFz1rZ9OcDsigVIQwJrvFrLLLf2MJL1ihpaidw052AzgJV65h ljwvxPV6KjzeqILgd6yDO36DnIe6vFzDLR6XTGk4jxv5/R+M9ep2kP+A7ZFvo+VanMGt U3Uo24P8Zj0t//9oIhZwX1Qtd5Lo1SrXdKD09FuOqA98K7uYedtA15/dnSqY2XaaCiUf 5kbI5uH6ZkeXvw9psX5stoZC/dUB42SDX7cukUBxJM+rp2XATuauhrSAEkVd6PsPupCO C90+K6RaJEXnQ74Gx+e6+vtqZbt3/dGGtd6N4r4rDqmwUkgs+fdKrHlazbucWIqRlkzS gpNw== X-Forwarded-Encrypted: i=1; AHgh+RorAbOCE8DAUkst0CS+N8CE/T8CfkDWc01UoDdZvcYP2r5BD426GqoyJV832M3tTh8uNlT01I6ySga0tb0=@vger.kernel.org X-Gm-Message-State: AOJu0YzkmvXmz/pXfPWrJexGOAI2UAypcSA22lr/d9jyXjR435iRlXum G7dHjhVqJM/8cs6v08zrEb9VXtKqa4GbomkbD8GIaCn8P1ABoUxwZC96 X-Gm-Gg: AR+sD12dcwGPSSX4zHEFaFJZgMWr3de8fTdUc0i2RjAYovtxJapchVdEg7naz5RhK+2 vtbb4Vn8oBQQFXCiuA+rw07mDy2BcM1uluxl0OGOCQkCtC93ECWXA7ulNqq5nXD6um+8DiPshtu RXqPN05nEdEA048hfnlg813LdC8cz/SJ7OrY3FBSgpuN9LCOL+wKzPXUO7o1RSPn6xHE1zPSL0q bSElS1OVtVOAZilHucGAkPsn377LO/eWtgC9UtHCV8JMcD2smwjuZA9hNiArFrMlFii4Gv1faIs bJ5tuCQF2K7PisZB1aQI5h4PyDyW6V12EMHAB845mVT21Jx/NY0xBlf72SNOblSW7KmUNKE/rCF ldFJ8sWzVHEStXAhNUP9aUagfx7qSgMjTY0iDTYjU2Vr7/2lBVc/iD8VgMnzy1A3MoSCbk2hsq1 kXU4XMVm7B4n5VcmwkbS6V0aZZDr4JYWLrHL+vb1Euc22ltVgQkLAKOZ5KK18Grdb4L13lbK2LG rB3z5JxxrSNJBlrutVtVg== X-Received: by 2002:a05:6820:290d:b0:6a1:80a7:2c8d with SMTP id 006d021491bc7-6ae97013777mr8599984eaf.32.1786041779594; Thu, 06 Aug 2026 11:42:59 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:53::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bfa6379sm143949eaf.14.2026.08.06.11.42.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:42:58 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 02/11] mm, swap: support zswap and zeroswap as vswap backends Date: Thu, 6 Aug 2026 11:42:45 -0700 Message-ID: <20260806184254.3790858-3-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Build the virtual swap layer on top of the swap-table infrastructure. Virtual swap entries decouple PTE swap entries from physical backing, allowing pages to be compressed by zswap (or detected as zero-filled) without pre-allocating a physical swap slot. This patch only supports zswap and zero-page backends. If zswap_store fails, the page stays dirty in the swap cache. Physical disk backing arrives in the next patch. Zswap writeback of vswap-backed entries is also disabled: they have no physical slot to write back to yet, so the zswap shrinker (both the dynamic count path and the pool-full worker path) is skipped while vswap is enabled. Physical backing and real writeback come in later patches. THP swapin is disabled for vswap entries for now. Add a /proc/sys/vm/vswap_enabled sysctl and a CONFIG_VSWAP_DEFAULT_ON build option so vswap allocation can be enabled and disabled at runtime, defaulting off unless CONFIG_VSWAP_DEFAULT_ON=3Dy. The knob only gates vswap_alloc(), so existing virtual entries keep resolving their backend and drain naturally when it is turned off. Suggested-by: Kairui Song Signed-off-by: Nhat Pham --- Documentation/admin-guide/sysctl/vm.rst | 16 ++ include/linux/zswap.h | 3 + mm/Kconfig | 11 ++ mm/memcontrol.c | 8 + mm/memory.c | 18 +- mm/page_io.c | 12 +- mm/shmem.c | 4 +- mm/swap.h | 1 + mm/swap_state.c | 8 + mm/swapfile.c | 242 ++++++++++++++++++++++-- mm/vmscan.c | 14 +- mm/vswap.h | 206 +++++++++++++++++++- mm/zswap.c | 56 ++++-- 13 files changed, 561 insertions(+), 38 deletions(-) diff --git a/Documentation/admin-guide/sysctl/vm.rst b/Documentation/admin-= guide/sysctl/vm.rst index 5b318d17aa4b..50b41f292631 100644 --- a/Documentation/admin-guide/sysctl/vm.rst +++ b/Documentation/admin-guide/sysctl/vm.rst @@ -74,6 +74,7 @@ Currently, these files are in /proc/sys/vm: - user_reserve_kbytes - vfs_cache_pressure - vfs_cache_pressure_denom +- vswap_enabled - watermark_boost_factor - watermark_scale_factor - zone_reclaim_mode @@ -1152,6 +1153,21 @@ vfs_cache_pressure_denom Defaults to 100 (minimum allowed value). Requires corresponding vfs_cache_pressure setting to take effect. =20 +vswap_enabled +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Controls whether new swapouts are routed through the virtual swap layer +(only present when the kernel is built with CONFIG_VSWAP). Set to 1 to +route swapouts through vswap, 0 to send them straight to the physical +swap device. + +The default is 0 unless the kernel was built with +CONFIG_VSWAP_DEFAULT_ON=3Dy. + +Disabling is allocation-only: it only stops new swapouts from using +vswap. Swap entries already backed by vswap keep being served and drain +naturally as they are faulted back in or freed. + watermark_boost_factor =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =20 diff --git a/include/linux/zswap.h b/include/linux/zswap.h index 30c193a1207e..4b4f211f3301 100644 --- a/include/linux/zswap.h +++ b/include/linux/zswap.h @@ -6,6 +6,7 @@ #include =20 struct lruvec; +struct zswap_entry; =20 extern atomic_long_t zswap_stored_pages; =20 @@ -28,6 +29,7 @@ unsigned long zswap_total_pages(void); bool zswap_store(struct folio *folio); int zswap_load(struct folio *folio); void zswap_invalidate(swp_entry_t swp); +void zswap_entry_free(struct zswap_entry *entry); int zswap_swapon(int type, unsigned long nr_pages); void zswap_swapoff(int type); void zswap_memcg_offline_cleanup(struct mem_cgroup *memcg); @@ -50,6 +52,7 @@ static inline int zswap_load(struct folio *folio) } =20 static inline void zswap_invalidate(swp_entry_t swp) {} +static inline void zswap_entry_free(struct zswap_entry *entry) {} static inline int zswap_swapon(int type, unsigned long nr_pages) { return 0; diff --git a/mm/Kconfig b/mm/Kconfig index 32d38b552845..8d147c0483ef 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -29,6 +29,17 @@ config VSWAP swapfile, or kept in memory - with the backing changeable at runtime without invalidating page table entries. =20 +config VSWAP_DEFAULT_ON + bool "Route swapouts through virtual swap by default" + depends on VSWAP + default n + help + Say Y to route swapouts through the virtual swap layer from + boot. + + Say N (default) to leave vswap off until it is enabled at + runtime via /proc/sys/vm/vswap_enabled. + config ZSWAP bool "Compressed cache for swap pages" depends on SWAP diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 77582acd8ee5..7a426db06222 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -65,6 +65,7 @@ #include "internal.h" #include "swap.h" #include "swap_table.h" +#include "vswap.h" #include #include #include "slab.h" @@ -5728,6 +5729,13 @@ long mem_cgroup_get_nr_swap_pages(struct mem_cgroup = *memcg) { long nr_swap_pages =3D get_nr_swap_pages(); =20 + /* + * vswap zswap-backed swapout needs no physical slot, so gate anon + * reclaim on the swap.max headroom instead of the physical free count. + */ + if (vswap_is_enabled() && zswap_is_enabled()) + nr_swap_pages =3D PAGE_COUNTER_MAX; + if (mem_cgroup_disabled() || do_memsw_account()) return nr_swap_pages; for (; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) diff --git a/mm/memory.c b/mm/memory.c index 6ae52e3869b1..de3573b7c6b1 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -89,6 +89,7 @@ #include "pgalloc-track.h" #include "internal.h" #include "swap.h" +#include "vswap.h" =20 #if defined(LAST_CPUPID_NOT_IN_PAGE_FLAGS) && !defined(CONFIG_COMPILE_TEST) #warning Unfortunate NUMA and NUMA Balancing config, growing page-frame fo= r last_cpupid. @@ -4657,6 +4658,12 @@ static inline bool should_try_to_free_swap(struct sw= ap_info_struct *si, */ if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) return true; + /* + * Non-swapfile backends cannot be reused for future swapouts. + * Free the swap slot unless backed by contiguous physical swap. + */ + if (is_vswap_entry(folio->swap)) + return true; if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) || folio_test_mlocked(folio)) return true; @@ -4805,15 +4812,16 @@ static unsigned long thp_swapin_suitable_orders(str= uct vm_fault *vmf) if (unlikely(userfaultfd_armed(vma))) return 0; =20 + entry =3D softleaf_from_pte(vmf->orig_pte); + /* - * A large swapped out folio could be partially or fully in zswap. We - * lack handling for such cases, so fallback to swapping in order-0 - * folio. + * THP swapin for vswap is not supported yet. Also, a large swapped + * out folio could be partially or fully in zswap, which we lack + * handling for. In both cases, fall back to order-0 swapin. */ - if (!zswap_never_enabled()) + if (is_vswap_entry(entry) || !zswap_never_enabled()) return 0; =20 - entry =3D softleaf_from_pte(vmf->orig_pte); /* * Get a list of all the (large) orders below PMD_ORDER that are enabled * and suitable for swapping THP. diff --git a/mm/page_io.c b/mm/page_io.c index fca1718056af..b1894cd014b3 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -160,14 +160,19 @@ static void swap_zeromap_folio_set(struct folio *foli= o) struct obj_cgroup *objcg =3D get_obj_cgroup_from_folio(folio); int nr_pages =3D folio_nr_pages(folio); struct swap_cluster_info *ci; + unsigned int voff, i; swp_entry_t entry; - unsigned int i; =20 VM_WARN_ON_ONCE_FOLIO(!folio_test_swapcache(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); =20 ci =3D swap_cluster_get_and_lock(folio); - for (i =3D 0; i < folio_nr_pages(folio); i++) { + if (is_vswap_entry(folio->swap)) { + /* Free any prior backing (e.g. ZSWAP entry from earlier swapout) */ + voff =3D swp_cluster_offset(folio->swap); + __vswap_release_backing(ci, voff, nr_pages); + } + for (i =3D 0; i < nr_pages; i++) { entry =3D page_swap_entry(folio_page(folio, i)); __swap_table_set_zero(ci, swp_cluster_offset(entry)); } @@ -235,6 +240,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio= *folio) */ swap_zeromap_folio_clear(folio); =20 + if (is_vswap_entry(folio->swap)) + folio_release_vswap_backing(folio); + if (zswap_store(folio)) { count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT); goto out_unlock; diff --git a/mm/shmem.c b/mm/shmem.c index 2e4dacdcce11..153ac7433fb0 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -85,6 +85,7 @@ static struct vfsmount *shm_mnt __ro_after_init; #include =20 #include "internal.h" +#include "vswap.h" =20 #define VM_ACCT(size) (PAGE_ALIGN(size) >> PAGE_SHIFT) =20 @@ -1617,7 +1618,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct fo= lio *folio, if ((info->flags & SHMEM_F_LOCKED) || sbinfo->noswap) goto redirty; =20 - if (!total_swap_pages) + /* vswap doesn't contribute to total_swap_pages */ + if (!total_swap_pages && !(vswap_is_enabled() && zswap_is_enabled())) goto redirty; =20 /* diff --git a/mm/swap.h b/mm/swap.h index b593ad3214ef..d241e81e967e 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -70,6 +70,7 @@ struct swap_cluster_info_dynamic { struct swap_cluster_info ci; unsigned int index; /* for cluster_index() */ struct rcu_head rcu; + atomic_long_t *virtual_table; /* Backing pointers for vswap slots */ }; =20 /* All on-list cluster must have a non-zero flag. */ diff --git a/mm/swap_state.c b/mm/swap_state.c index 9e0d71fcdc24..9f6377b32911 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -26,6 +26,7 @@ #include "internal.h" #include "swap_table.h" #include "swap.h" +#include "vswap.h" =20 /* Swap readahead cluster size, as a power of 2 pages. */ static int page_cluster; @@ -196,6 +197,13 @@ static int __swap_cache_add_check(struct swap_cluster_= info *ci, if (nr =3D=3D 1) return 0; =20 + /* + * THP swapin for vswap is not supported yet; reject the batch so + * swap_cache_alloc_folio falls back to order 0. + */ + if (is_vswap_entry(targ_entry)) + return -EBUSY; + is_zero =3D __swap_table_test_zero(ci, ci_off); ci_off =3D round_down(ci_off, nr); ci_end =3D ci_off + nr; diff --git a/mm/swapfile.c b/mm/swapfile.c index fea3a8eccbc1..a26cfe5751c8 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -23,6 +23,7 @@ #include #include #include +#include #include #include #include @@ -131,6 +132,29 @@ static DEFINE_PER_CPU(struct percpu_swap_cluster, perc= pu_swap_cluster) =3D { .lock =3D INIT_LOCAL_LOCK(), }; =20 +#ifdef CONFIG_VSWAP +static int sysctl_vswap_enabled =3D IS_ENABLED(CONFIG_VSWAP_DEFAULT_ON); + +bool vswap_is_enabled(void) +{ + return sysctl_vswap_enabled; +} + +struct percpu_vswap_cluster { + unsigned long offset[SWAP_NR_ORDERS]; + local_lock_t lock; +}; + +static DEFINE_PER_CPU(struct percpu_vswap_cluster, percpu_vswap_cluster) = =3D { + .offset =3D { [0 ... SWAP_NR_ORDERS - 1] =3D SWAP_ENTRY_INVALID }, + .lock =3D INIT_LOCAL_LOCK(), +}; + +static bool vswap_alloc(struct folio *folio); +#else +static inline bool vswap_alloc(struct folio *folio) { return false; } +#endif + /* May return NULL on invalid type, caller must check for NULL return */ static struct swap_info_struct *swap_type_to_info(int type) { @@ -236,7 +260,8 @@ static int __try_to_reclaim_swap(struct swap_info_struc= t *si, =20 need_reclaim =3D ((flags & TTRS_ANYWAY) || ((flags & TTRS_UNMAPPED) && !folio_mapped(folio)) || - ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio))); + ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio) && + !is_vswap_entry(folio->swap))); if (!need_reclaim || !folio_swapcache_freeable(folio)) goto out_unlock; =20 @@ -537,7 +562,12 @@ swap_cluster_populate(struct swap_info_struct *si, * Only cluster isolation from the allocator does table allocation. * Swap allocator uses percpu clusters and holds the local lock. */ - lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); +#ifdef CONFIG_VSWAP + if (swap_is_vswap(si)) + lockdep_assert_held(&this_cpu_ptr(&percpu_vswap_cluster)->lock); +#endif + if (!swap_is_vswap(si)) + lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); if (!(si->flags & SWP_SOLIDSTATE)) lockdep_assert_held(&si->global_cluster_lock); lockdep_assert_held(&ci->lock); @@ -554,7 +584,12 @@ swap_cluster_populate(struct swap_info_struct *si, spin_unlock(&ci->lock); if (!(si->flags & SWP_SOLIDSTATE)) spin_unlock(&si->global_cluster_lock); - local_unlock(&percpu_swap_cluster.lock); +#ifdef CONFIG_VSWAP + if (swap_is_vswap(si)) + local_unlock(&percpu_vswap_cluster.lock); +#endif + if (!swap_is_vswap(si)) + local_unlock(&percpu_swap_cluster.lock); =20 ret =3D swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL); @@ -567,7 +602,12 @@ swap_cluster_populate(struct swap_info_struct *si, * could happen with ignoring the percpu cluster is fragmentation, * which is acceptable since this fallback and race is rare. */ - local_lock(&percpu_swap_cluster.lock); +#ifdef CONFIG_VSWAP + if (swap_is_vswap(si)) + local_lock(&percpu_vswap_cluster.lock); +#endif + if (!swap_is_vswap(si)) + local_lock(&percpu_swap_cluster.lock); if (!(si->flags & SWP_SOLIDSTATE)) spin_lock(&si->global_cluster_lock); spin_lock(&ci->lock); @@ -729,6 +769,7 @@ static void vswap_free_cluster(struct swap_info_struct = *si, spin_unlock(&si->lock); } swap_cluster_free_table(ci); + vswap_cluster_free_vtable(ci); /* * Ordering vs the RCU cluster lookup: erase from the xarray first * (new lookups miss it), mark DEAD under the held ci->lock (a lookup @@ -765,6 +806,10 @@ static void free_cluster(struct swap_info_struct *si, = struct swap_cluster_info * return; } =20 + /* + * Vswap dynamic clusters need explicit cleanup (xarray erase, + * kfree_rcu, virtual_table free if allocated). + */ if (swap_is_vswap(si)) { vswap_free_cluster(si, ci); return; @@ -947,7 +992,8 @@ static bool cluster_scan_range(struct swap_info_struct = *si, if (swp_tb_is_null(swp_tb)) continue; if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) { - if (!vm_swap_full()) + /* vswap slots are unlimited; never reclaim to reuse one */ + if (swap_is_vswap(si) || !vm_swap_full()) return false; *need_reclaim =3D true; continue; @@ -1015,7 +1061,8 @@ static bool __swap_cluster_alloc_entries(struct swap_= info_struct *si, /* Try use a new cluster for current CPU and allocate from it. */ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci, - struct folio *folio, unsigned long offset) + struct folio *folio, + unsigned long offset) { unsigned int next =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; unsigned long start =3D ALIGN_DOWN(offset, SWAPFILE_CLUSTER); @@ -1058,6 +1105,12 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, relocate_cluster(si, ci); swap_cluster_unlock(ci); } +#ifdef CONFIG_VSWAP + if (swap_is_vswap(si)) { + this_cpu_write(percpu_vswap_cluster.offset[order], next); + return found; + } +#endif if (si->flags & SWP_SOLIDSTATE) { this_cpu_write(percpu_swap_cluster.offset[order], next); this_cpu_write(percpu_swap_cluster.si[order], si); @@ -1110,10 +1163,17 @@ static unsigned int alloc_swap_scan_dynamic(struct = swap_info_struct *si, return SWAP_ENTRY_INVALID; } =20 + if (vswap_cluster_alloc_vtable(ci_dyn)) { + swap_cluster_free_table(&ci_dyn->ci); + kfree(ci_dyn); + return SWAP_ENTRY_INVALID; + } + if (xa_alloc(&si->cluster_info_pool, &ci_dyn->index, ci_dyn, XA_LIMIT(1, DIV_ROUND_UP(si->max, SWAPFILE_CLUSTER) - 1), GFP_ATOMIC)) { swap_cluster_free_table(&ci_dyn->ci); + vswap_cluster_free_vtable(&ci_dyn->ci); kfree(ci_dyn); return SWAP_ENTRY_INVALID; } @@ -1199,7 +1259,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, * Swapfile is not block device so unable * to allocate large entries. */ - if (order && !(si->flags & SWP_BLKDEV)) + if (order && !(si->flags & SWP_BLKDEV) && !swap_is_vswap(si)) return 0; =20 if (!(si->flags & SWP_SOLIDSTATE)) { @@ -1252,7 +1312,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, } =20 /* Try reclaim full clusters if free and nonfull lists are drained */ - if (vm_swap_full()) + if (!swap_is_vswap(si) && vm_swap_full()) swap_reclaim_full_clusters(si, false); =20 if (order < PMD_ORDER) { @@ -1416,7 +1476,8 @@ static void swap_range_alloc(struct swap_info_struct = *si, if (vm_swap_full()) schedule_work(&si->reclaim_work); } - atomic_long_sub(nr_entries, &nr_swap_pages); + if (!swap_is_vswap(si)) + atomic_long_sub(nr_entries, &nr_swap_pages); } =20 static void swap_range_free(struct swap_info_struct *si, unsigned long off= set, @@ -1426,8 +1487,10 @@ static void swap_range_free(struct swap_info_struct = *si, unsigned long offset, void (*swap_slot_free_notify)(struct block_device *, unsigned long); unsigned int i; =20 - for (i =3D 0; i < nr_entries; i++) - zswap_invalidate(swp_entry(si->type, offset + i)); + if (!swap_is_vswap(si)) { + for (i =3D 0; i < nr_entries; i++) + zswap_invalidate(swp_entry(si->type, offset + i)); + } =20 if (si->flags & SWP_BLKDEV) swap_slot_free_notify =3D @@ -1446,7 +1509,8 @@ static void swap_range_free(struct swap_info_struct *= si, unsigned long offset, * only after the above cleanups are done. */ smp_wmb(); - atomic_long_add(nr_entries, &nr_swap_pages); + if (!swap_is_vswap(si)) + atomic_long_add(nr_entries, &nr_swap_pages); swap_usage_sub(si, nr_entries); } =20 @@ -1838,6 +1902,49 @@ static int swap_dup_entries_cluster(struct swap_info= _struct *si, return err; } =20 +#ifdef CONFIG_VSWAP +static bool vswap_alloc(struct folio *folio) +{ + unsigned int order =3D folio_order(folio); + struct swap_cluster_info *ci; + unsigned long offset; + + if (!sysctl_vswap_enabled) + return false; + + /* vswap_init failed: fall back to direct physical swap */ + if (!vswap_si) + return false; + + local_lock(&percpu_vswap_cluster.lock); + offset =3D this_cpu_read(percpu_vswap_cluster.offset[order]); + + if (offset !=3D SWAP_ENTRY_INVALID) { + ci =3D swap_cluster_lock(vswap_si, offset); + if (ci && cluster_is_usable(ci, order)) { + if (cluster_is_empty(ci)) + offset =3D cluster_offset(vswap_si, ci); + alloc_swap_scan_cluster(vswap_si, ci, folio, offset); + } else if (ci) { + swap_cluster_unlock(ci); + } + } + + if (!folio_test_swapcache(folio)) + cluster_alloc_swap_entry(vswap_si, folio); + + if (folio_test_swapcache(folio)) { + /* alloc_swap_scan_cluster updated percpu offset already */ + local_unlock(&percpu_vswap_cluster.lock); + return true; + } + + this_cpu_write(percpu_vswap_cluster.offset[order], SWAP_ENTRY_INVALID); + local_unlock(&percpu_vswap_cluster.lock); + return false; +} +#endif + /** * folio_alloc_swap - allocate swap space for a folio * @folio: folio we want to move to swap @@ -1875,12 +1982,17 @@ int folio_alloc_swap(struct folio *folio) } } =20 + /* Without zswap a vswap entry has nowhere to go on writeout. */ + if (zswap_is_enabled() && vswap_alloc(folio)) + goto done; + again: local_lock(&percpu_swap_cluster.lock); if (!swap_alloc_fast(folio)) swap_alloc_slow(folio); local_unlock(&percpu_swap_cluster.lock); =20 +done: if (!order && unlikely(!folio_test_swapcache(folio))) { if (swap_sync_discard()) goto again; @@ -1896,6 +2008,80 @@ int folio_alloc_swap(struct folio *folio) return 0; } =20 +#ifdef CONFIG_VSWAP + +/** + * __vswap_release_backing - release the backing of a range of vtable slots + * @ci: the locked vswap cluster + * @ci_start: first slot offset within @ci + * @nr: number of slots + * + * Releases each slot in [@ci_start, @ci_start + @nr): physical swap slots, + * zswap entries, etc. Clears the zero marks if set. + * + * Context: caller must hold @ci->lock. The entire range must belong to the + * same memcg. + */ +void __vswap_release_backing(struct swap_cluster_info *ci, + unsigned int ci_start, unsigned int nr) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int ci_off; + unsigned long vt; + + lockdep_assert_held(&ci->lock); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + + for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { + vt =3D __vtable_get(ci_dyn, ci_off); + + switch (vtable_type(vt)) { + case VSWAP_ZSWAP: + zswap_entry_free(vtable_to_zswap(vt)); + break; + case VSWAP_NONE: + break; + default: + /* VSWAP_ZERO/VSWAP_FOLIO are return-only, not vtable tags */ + break; + } + + __vtable_set(ci_dyn, ci_off, VSWAP_NONE); + /* Zero-backed state lives in swap_table; clear it too. */ + if (__swap_table_test_zero(ci, ci_off)) + __swap_table_clear_zero(ci, ci_off); + } +} + +/** + * folio_release_vswap_backing() - Drop all backing for a folio's vswap en= try. + * @folio: the folio, occupying a virtual swap entry. + * + * Release whatever backing the folio's virtual swap slots currently hold = and + * reset them to empty, so a fresh backing can be installed. Used when a + * folio's swap backend is replaced. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + */ +void folio_release_vswap_backing(struct folio *folio) +{ + struct swap_cluster_info *ci; + int nr =3D folio_nr_pages(folio); + unsigned int voff; + + ci =3D __swap_entry_to_cluster(folio->swap); + if (!ci) + return; + voff =3D swp_cluster_offset(folio->swap); + + spin_lock(&ci->lock); + __vswap_release_backing(ci, voff, nr); + spin_unlock(&ci->lock); +} + +#endif /* CONFIG_VSWAP */ + /** * folio_dup_swap() - Increase swap count of swap entries of a folio. * @folio: folio with swap entries bounded. @@ -2037,6 +2223,9 @@ void __swap_cluster_free_entries(struct swap_info_str= uct *si, =20 VM_WARN_ON(ci->count < nr_pages); =20 + if (swap_is_vswap(si)) + __vswap_release_backing(ci, ci_start, nr_pages); + ci->count -=3D nr_pages; do { old_tb =3D __swap_table_get(ci, ci_off); @@ -2907,6 +3096,7 @@ static int try_to_unuse(unsigned int type) (i =3D find_next_to_unuse(si, i)) !=3D 0) { =20 entry =3D swp_entry(type, i); + folio =3D swap_cache_get_folio(entry); if (!folio) continue; @@ -4134,6 +4324,18 @@ struct swap_info_struct *vswap_si; /* vswap does no IO on its own. */ static const struct swap_ops vswap_ops =3D { }; =20 +static const struct ctl_table vswap_sysctls[] =3D { + { + .procname =3D "vswap_enabled", + .data =3D &sysctl_vswap_enabled, + .maxlen =3D sizeof(sysctl_vswap_enabled), + .mode =3D 0644, + .proc_handler =3D proc_dointvec_minmax, + .extra1 =3D SYSCTL_ZERO, + .extra2 =3D SYSCTL_ONE, + }, +}; + static int __init vswap_init(void) { struct swap_info_struct *si; @@ -4141,8 +4343,12 @@ static int __init vswap_init(void) int err; =20 si =3D alloc_swap_info(); - if (IS_ERR(si)) - return PTR_ERR(si); + if (IS_ERR(si)) { + pr_warn("vswap: alloc_swap_info failed (%ld); vswap disabled, swapout fa= lls back to direct physical swap\n", + PTR_ERR(si)); + sysctl_vswap_enabled =3D 0; + return 0; + } =20 maxpages =3D min(swapfile_maximum_size, ALIGN_DOWN((unsigned long)UINT_MAX, SWAPFILE_CLUSTER)); @@ -4164,14 +4370,20 @@ static int __init vswap_init(void) mutex_unlock(&swapon_mutex); =20 vswap_si =3D si; + + register_sysctl_init("vm", vswap_sysctls); + pr_info("vswap: created virtual swap device (%lu pages)\n", maxpages); return 0; =20 fail: + pr_warn("vswap: setup_swap_clusters_info failed (%d); vswap disabled, swa= pout falls back to direct physical swap\n", + err); + sysctl_vswap_enabled =3D 0; spin_lock(&swap_lock); si->flags =3D 0; spin_unlock(&swap_lock); - return err; + return 0; } late_initcall(vswap_init); #endif diff --git a/mm/vmscan.c b/mm/vmscan.c index 17d2b793cbfc..78ec51f53757 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -68,6 +68,7 @@ #include "internal.h" #include "page_alloc.h" #include "swap.h" +#include "vswap.h" =20 #define CREATE_TRACE_POINTS #include @@ -352,6 +353,9 @@ static inline bool can_reclaim_anon_pages(struct mem_cg= roup *memcg, */ if (get_nr_swap_pages() > 0) return true; + /* vswap doesn't contribute to nr_swap_pages */ + if (vswap_is_enabled() && zswap_is_enabled()) + return true; } else { /* Is the memcg below its swap limit? */ if (mem_cgroup_get_nr_swap_pages(memcg) > 0) @@ -1521,9 +1525,13 @@ static unsigned int shrink_folio_list(struct list_he= ad *folio_list, nr_pages =3D 1; } activate_locked: - /* Not a candidate for swapping, so reclaim swap space. */ + /* + * Not a candidate for swapping, so reclaim physical swap + * space if we are running out. + */ if (folio_test_swapcache(folio) && - (mem_cgroup_swap_full(folio) || folio_test_mlocked(folio))) + ((mem_cgroup_swap_full(folio) && !is_vswap_entry(folio->swap)) || + folio_test_mlocked(folio))) folio_free_swap(folio); VM_BUG_ON_FOLIO(folio_test_active(folio), folio); if (!folio_test_mlocked(folio)) { @@ -2680,7 +2688,7 @@ static bool can_age_anon_pages(struct lruvec *lruvec, struct scan_control *sc) { /* Aging the anon LRU is valuable if swap is present: */ - if (total_swap_pages > 0) + if (total_swap_pages > 0 || (vswap_is_enabled() && zswap_is_enabled())) return true; =20 /* Also valuable if anon pages can be demoted: */ diff --git a/mm/vswap.h b/mm/vswap.h index 5641692f5be3..6d25e0911fa9 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -10,8 +10,23 @@ #include #include "swap.h" =20 +struct zswap_entry; + +/* + * VSWAP_ZERO and VSWAP_FOLIO are return-only values synthesized from + * swap_table state; the rest are stored in the vtable per slot. + */ +enum vswap_backing_type { + VSWAP_NONE =3D 0, + VSWAP_ZSWAP =3D 1, + VSWAP_ZERO, + VSWAP_FOLIO, +}; + #ifdef CONFIG_VSWAP =20 +#include "swap_table.h" + extern struct swap_info_struct *vswap_si; =20 static inline bool is_vswap_entry(swp_entry_t entry) @@ -19,13 +34,202 @@ static inline bool is_vswap_entry(swp_entry_t entry) return swap_is_vswap(__swap_entry_to_info(entry)); } =20 -#else +bool vswap_is_enabled(void); + +/* + * Virtual table entry encoding for vswap clusters. + * + * Each entry in ci_dyn->virtual_table stores the backing type and + * pointer for a virtual swap slot. Tag in low 3 bits, payload in + * upper 61 bits. + * + * NONE: |----- 0000 ------|000| - no separate backend pointer + * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low bi= ts) + * + * Pointer payloads (ZSWAP) are stored directly with the tag OR'd into the + * low bits (kernel pointers are >=3D 8-byte aligned, same approach as xar= ray). + * + * vtable[i] =3D NONE does not by itself mean "free". The swap_table entry + * and the per-slot zero flag carry the rest of the state. The full + * per-slot state table is: + * + * vtable[i] | swap_table[i] | zero | meaning + * ----------+---------------+-------+-------------------------------- + * NONE | NULL | clear | truly free / unbacked + * NONE | PFN | clear | folio cached, no backing + * NONE | shadow | clear | folio evicted, no backing (bug) + * NONE | * | set | zero-backed; cached if PFN set + * ZSWAP | PFN | clear | folio cached + zswap entry + * ZSWAP | shadow / NULL | clear | evicted, only in zswap + * + * Locking: a slot's vtable entry (the vswap entry's backend) is only + * stable while the caller owns and holds the lock on that entry's swap + * cache folio. The cluster lock (ci_dyn->ci.lock) only makes an individual + * vtable read atomic, and by itself does not give the caller the right to + * change the backend. A backend read without the folio lock is + * best-effort and must be re-validated under the folio lock before + * being acted on. + * + * Zero-backed slots use the swap_table per-slot zero flag (same as + * direct-mapped physical swap), since CONFIG_VSWAP requires 64BIT and + * SWAP_TABLE_HAS_ZEROFLAG is always true on 64-bit. Cached folios are + * read out of the swap_table PFN entry; there is no separate FOLIO + * vtable type because the folio pointer would duplicate that PFN and + * would go stale on folio migration / split. + */ + +#define VTABLE_TAG_BITS 3 +#define VTABLE_TAG_MASK ((1UL << VTABLE_TAG_BITS) - 1) + +static inline enum vswap_backing_type vtable_type(unsigned long vt) +{ + return vt & VTABLE_TAG_MASK; +} + +static inline struct zswap_entry *vtable_to_zswap(unsigned long vt) +{ + VM_WARN_ON(vtable_type(vt) !=3D VSWAP_ZSWAP); + return (struct zswap_entry *)(vt & ~VTABLE_TAG_MASK); +} + +/* Virtual table accessors */ + +static inline unsigned long __vtable_get(struct swap_cluster_info_dynamic = *ci_dyn, + unsigned int off) +{ + VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); + return atomic_long_read(&ci_dyn->virtual_table[off]); +} + +static inline void __vtable_set(struct swap_cluster_info_dynamic *ci_dyn, + unsigned int off, unsigned long vt) +{ + VM_WARN_ON_ONCE(off >=3D SWAPFILE_CLUSTER); + atomic_long_set(&ci_dyn->virtual_table[off], vt); +} + +/** + * vswap_lock_cluster - look up and lock the vswap cluster for an entry + * @entry: the virtual swap entry + * @voff: out param, receives @entry's slot offset within the cluster + * + * Return: the locked vswap cluster, or NULL if no cluster is found for @e= ntry. + */ +static inline struct swap_cluster_info_dynamic * +vswap_lock_cluster(swp_entry_t entry, unsigned int *voff) +{ + struct swap_cluster_info *ci; + struct swap_cluster_info_dynamic *ci_dyn; + + ci =3D __swap_entry_to_cluster(entry); + if (!ci) + return NULL; + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + *voff =3D swp_cluster_offset(entry); + spin_lock(&ci->lock); + return ci_dyn; +} + +void __vswap_release_backing(struct swap_cluster_info *ci, + unsigned int ci_start, unsigned int nr); + +/** + * vswap_zswap_store - record a zswap entry as the backing for a vswap ent= ry. + * @entry: the vswap entry + * @ze: the zswap entry now holding @entry's compressed data + * + * Releases @entry's previous backing, and sets the zswap entry @ze as the= new + * backing. + * + * Context: takes and drops the vswap cluster lock internally. + */ +static inline void vswap_zswap_store(swp_entry_t entry, + struct zswap_entry *ze) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) + return; + __vswap_release_backing(&ci_dyn->ci, voff, 1); + __vtable_set(ci_dyn, voff, (unsigned long)ze | VSWAP_ZSWAP); + spin_unlock(&ci_dyn->ci.lock); +} + +/** + * vswap_zswap_load - return the zswap entry backing a vswap entry + * @entry: the virtual swap entry + * + * Context: takes and drops the vswap cluster lock internally. + * Return: the backing zswap entry, or NULL if @entry is not zswap-backed. + */ +static inline struct zswap_entry *vswap_zswap_load(swp_entry_t entry) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + unsigned long vt; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) + return NULL; + vt =3D __vtable_get(ci_dyn, voff); + spin_unlock(&ci_dyn->ci.lock); + + if (vtable_type(vt) !=3D VSWAP_ZSWAP) + return NULL; + return vtable_to_zswap(vt); +} + +void folio_release_vswap_backing(struct folio *folio); + +static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn) +{ + ci_dyn->virtual_table =3D kcalloc(SWAPFILE_CLUSTER, + sizeof(*ci_dyn->virtual_table), + GFP_ATOMIC); + return ci_dyn->virtual_table ? 0 : -ENOMEM; +} + +static inline void vswap_cluster_free_vtable(struct swap_cluster_info *ci) +{ + struct swap_cluster_info_dynamic *ci_dyn; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + kfree(ci_dyn->virtual_table); + ci_dyn->virtual_table =3D NULL; +} + +#else /* !CONFIG_VSWAP */ =20 static inline bool is_vswap_entry(swp_entry_t entry) { return false; } =20 +static inline bool vswap_is_enabled(void) { return false; } + +static inline void __vswap_release_backing(struct swap_cluster_info *ci, + unsigned int ci_start, + unsigned int nr) {} + +static inline void vswap_zswap_store(swp_entry_t entry, + struct zswap_entry *ze) {} + +static inline struct zswap_entry *vswap_zswap_load(swp_entry_t entry) +{ + return NULL; +} + +static inline void folio_release_vswap_backing(struct folio *folio) {} + +static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn) +{ + return 0; +} + +static inline void vswap_cluster_free_vtable(struct swap_cluster_info *ci)= {} + #endif /* CONFIG_VSWAP */ =20 #endif /* _MM_VSWAP_H */ diff --git a/mm/zswap.c b/mm/zswap.c index 354bf8bd7482..e19bde9df722 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -38,6 +38,7 @@ #include =20 #include "swap.h" +#include "vswap.h" #include "internal.h" =20 /********************************* @@ -234,6 +235,25 @@ static inline struct xarray *swap_zswap_tree(swp_entry= _t swp) >> ZSWAP_ADDRESS_SPACE_SHIFT]; } =20 +static struct zswap_entry *zswap_entry_load(swp_entry_t swp) +{ + if (is_vswap_entry(swp)) + return vswap_zswap_load(swp); + return xa_load(swap_zswap_tree(swp), swp_offset(swp)); +} + +static struct zswap_entry *zswap_entry_store(swp_entry_t swp, + struct zswap_entry *entry) +{ + if (is_vswap_entry(swp)) { + vswap_zswap_store(swp, entry); + return NULL; + } + + return xa_store(swap_zswap_tree(swp), swp_offset(swp), entry, + GFP_KERNEL); +} + #define zswap_pool_debug(msg, p) \ pr_debug("%s pool %s\n", msg, (p)->tfm_name) =20 @@ -762,7 +782,7 @@ static void zswap_entry_cache_free(struct zswap_entry *= entry) * Carries out the common pattern of freeing an entry's zsmalloc allocatio= n, * freeing the entry itself, and decrementing the number of stored pages. */ -static void zswap_entry_free(struct zswap_entry *entry) +void zswap_entry_free(struct zswap_entry *entry) { zswap_lru_del(entry); zs_free(entry->pool->zs_pool, entry->handle); @@ -1208,6 +1228,9 @@ static unsigned long zswap_shrinker_count(struct shri= nker *shrinker, if (!zswap_shrinker_enabled || !mem_cgroup_zswap_writeback_enabled(memcg)) return 0; =20 + if (vswap_is_enabled()) + return 0; + /* * The shrinker resumes swap writeback, which will enter block * and may enter fs. XXX: Harmonize with vmscan.c __GFP_FS @@ -1290,6 +1313,9 @@ static int shrink_memcg(struct mem_cgroup *memcg) if (!mem_cgroup_zswap_writeback_enabled(memcg)) return -ENOENT; =20 + if (vswap_is_enabled()) + return -ENOENT; + /* * Skip zombies because their LRUs are reparented and we would be * reclaiming from the parent instead of the dead memcg. @@ -1418,9 +1444,7 @@ static bool zswap_store_page(struct page *page, if (!zswap_compress(page, entry, pool)) goto compress_failed; =20 - old =3D xa_store(swap_zswap_tree(page_swpentry), - swp_offset(page_swpentry), - entry, GFP_KERNEL); + old =3D zswap_entry_store(page_swpentry, entry); if (xa_is_err(old)) { int err =3D xa_err(old); =20 @@ -1489,7 +1513,7 @@ bool zswap_store(struct folio *folio) struct mem_cgroup *memcg =3D NULL; struct zswap_pool *pool; bool ret =3D false; - long index; + long index =3D 0; =20 VM_WARN_ON_ONCE(!folio_test_locked(folio)); VM_WARN_ON_ONCE(!folio_test_swapcache(folio)); @@ -1544,13 +1568,19 @@ bool zswap_store(struct folio *folio) if (!ret && zswap_pool_reached_full) queue_work(shrink_wq, &zswap_shrink_work); check_old: + if (ret) + return ret; + /* * If the zswap store fails or zswap is disabled, we must invalidate * the possibly stale entries which were previously stored at the * offsets corresponding to each page of the folio. Otherwise, * writeback could overwrite the new data in the swapfile. */ - if (!ret) { + if (is_vswap_entry(swp)) { + if (index > 0) + folio_release_vswap_backing(folio); + } else { unsigned type =3D swp_type(swp); pgoff_t offset =3D swp_offset(swp); struct zswap_entry *entry; @@ -1590,8 +1620,7 @@ bool zswap_store(struct folio *folio) int zswap_load(struct folio *folio) { swp_entry_t swp =3D folio->swap; - pgoff_t offset =3D swp_offset(swp); - struct xarray *tree =3D swap_zswap_tree(swp); + struct swap_info_struct *si =3D __swap_entry_to_info(swp); struct zswap_entry *entry; =20 VM_WARN_ON_ONCE(!folio_test_locked(folio)); @@ -1610,7 +1639,7 @@ int zswap_load(struct folio *folio) return -EINVAL; } =20 - entry =3D xa_load(tree, offset); + entry =3D zswap_entry_load(swp); if (!entry) return -ENOENT; =20 @@ -1633,8 +1662,13 @@ int zswap_load(struct folio *folio) * compression work. */ folio_mark_dirty(folio); - xa_erase(tree, offset); - zswap_entry_free(entry); + + if (swap_is_vswap(si)) { + folio_release_vswap_backing(folio); + } else { + xa_erase(swap_zswap_tree(swp), swp_offset(swp)); + zswap_entry_free(entry); + } =20 folio_unlock(folio); return 0; --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oo1-f47.google.com (mail-oo1-f47.google.com [209.85.161.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3FCED3AE1A2 for ; Thu, 6 Aug 2026 18:43:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041784; cv=none; b=nqmAejw+INyZVNvUnlbd2VYToKFRU8s3athHCWhx0KMc00/bBX89Rvlyh+Z9+TjiLLOdXWlOV6Fnh+TEXQjAI97lXg2jcul6joZVei3HyUpRVOWPZdMtg10NirkHEr2l03lGTC04inGsNv8o7o/z8AYVYGxHiygwjdeOdOWNtjg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041784; c=relaxed/simple; bh=Z/V1Am93T/YF1wpxfRBBd7cgF7WxgTBIYTqrA/LI+Cg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NZEWnyhU7Q9LJ2nj1HW2Ja1Xf7Y1dYqerKFlDK1ycS+g7B6FLRnqoCxXxRz45a+zOIG7c+x6tz4giTSAylVWTWtpfY0XYDYKuuaqgSogmOOwdJLkEYacd3zsI5bpll6dAIXKQAy7cgyBB0ThNKg3NjP8jcT0l+RS9BMiMFVbfXU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=JrdpIn3h; arc=none smtp.client-ip=209.85.161.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="JrdpIn3h" Received: by mail-oo1-f47.google.com with SMTP id 006d021491bc7-6b01981b541so562411eaf.0 for ; Thu, 06 Aug 2026 11:43:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041781; x=1786646581; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7uMnAqaoMp1MP/2YeQ3uTyWm7ZJYM3hk0so38h+N6NA=; b=JrdpIn3hpQoTAWxShBprh3ajI796WtrN2YjCgo/qcVvwZreM8+VBzRa4SDYVELNdM2 TcIvKsFe6YpbYZnuXenO4mjC6SXItCaXQorvpJYy/d1QC0h1GeeikP18Rllk+paGzd0o xFR/3uxRPK9aPck2Sx6ojkct2x+p6Wwv2nJX/5JIWbUe33+i9oJ06QKWHcujpQ2gvtFQ 5uh4qJw0ENPiCjQFK0/o/Ex3nu27JW/OMpAcxruYhbIq+QsuzHqLDRMRU9spWTiLmYQI h7fB0qq+ei698MMzNv6t6qYu6aWLzv/iF5CAthFFiY6Zn1c7CfuYhS8lZePpFesQ7qAq FTPw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041781; x=1786646581; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7uMnAqaoMp1MP/2YeQ3uTyWm7ZJYM3hk0so38h+N6NA=; b=S+0ArJOvT09BFg1JHt5rriYdki9DPfgSfbMUpyaqnOuS1LrSNoXKp4W9avoTZpw/G2 ELsL9U/QpB04GeNu4SwXsQF0eNncKHsJwGI3S4T3Rdt8X7wxFk5r7TUgSXrlwz3HIhxH 7MoXR0/F7eTRimGJ4i7K8TqTBI4q2su9NVz7QGEHuc7PjHzIhbXm2uSYa53NJTguOvTR R+phzm2huNEbXs0PhtRS/vMbKy9JH4llTyu0KX15J57Y29CmYP2vzJkUdnwXCwl6nSFq W49kOYdocseKgjc7gDgHApsfCpnOTd+3vorzpj1kCawHy4NhL8E9qrEcsOkocHnfRb/S u5oA== X-Forwarded-Encrypted: i=1; AHgh+RpqlFDbKZjT8OjveBcuOMXPXZmLd0zfHaxrQU5fLzAEhs0d9vhGBo5r0f+1Kcg6b5ZvuM1W0nEYL6JpsSw=@vger.kernel.org X-Gm-Message-State: AOJu0YxYIvji1T0ccPz832VMjjVzryKTeHqmjR4juGx/YRS9Vo6aclds VuvCTDEsUQwr0pKi49XS4uwXCj5vdce1fqft7kImQdh5RXvgjfPdYGe5 X-Gm-Gg: AR+sD1027rP0OXIdbNwvdTnmuPx/KmG3QzUq/1SHHBvOduFIkaSWkH7zWjLmIPEhto7 ZaEOwbBa2j2zKkxK+7tuhbvBgH0t5efh1QWFP4lbt5IZTxvoLfxC0bk5I/ZYfyLZGJw96lfowQA FRLtz2p2gZ5qWireP8ugw4TyBnKxJzPHpWg6hYIYbj6LrAOo2lO7x08a+4cyCv1cSwgn2nI8TTw +WgRHu5nipD2gqVuvDduH/hEdLnRHvLcemxiZ4Xb1ltT10lk2Yh6ZKqCqjjg5n/6fjg3hoSvE8P 5BEBaXSJJQgMUAUShnK8dsH88M0EqcaheLjaRH3cxEEkxofDF+w5yDf+LGjeLdtacLn6lrTF/nu KlOI7hIACpgpkOC1h/TWa09sNwTzTYFbEBmJBDaDve9RWUODptDVzLrgiqOuKLEh33ZD4Qoh93R RVhoKwWzHGYfOIw1da+f5D7Graftbq5u0gGRTW1WNzyii2IeM48D2R8QA77z3qKSwxvxey2T7b6 1PVJSgM6ic= X-Received: by 2002:a05:6820:4b8b:b0:6ac:b3b1:a5a4 with SMTP id 006d021491bc7-6ae968c1761mr9319789eaf.0.1786041781003; Thu, 06 Aug 2026 11:43:01 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4d::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bc7f3d1sm191834eaf.8.2026.08.06.11.43.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:00 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 03/11] mm, swap: prepare the swap IO path for vswap Date: Thu, 6 Aug 2026 11:42:46 -0700 Message-ID: <20260806184254.3790858-4-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In preparation for adding a physical swap backend for vswap, make the swap IO path able to submit IO for a swap entry other than folio->swap. swap_add_folio() and __swap_writepage() derive the target device and sector from folio->swap. For a vswap folio backed by a physical slot that entry is virtual, so it identifies neither the backing device nor the sector to submit IO against. Compute the sector from an explicit entry (swap_folio_sector becomes swap_entry_sector), thread that entry through swap_add_folio, __swap_writepage and ops->can_merge, and stash it in swap_iocb so the submit path addresses the IO from it rather than from folio->swap. This lets the batching path serve both vswap entries (backed by a physical slot) and physical entries mapped directly into PTEs. All callers pass folio->swap for now, so there is no functional change. Signed-off-by: Nhat Pham --- include/linux/swap.h | 2 +- mm/page_io.c | 51 ++++++++++++++++++++++---------------------- mm/swap.h | 7 +++--- mm/swapfile.c | 6 +++--- mm/zswap.c | 2 +- 5 files changed, 35 insertions(+), 33 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index a955bd60dd58..8359134d08cb 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -389,7 +389,7 @@ extern int __swap_count(swp_entry_t entry); extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t en= try); extern int swp_swapcount(swp_entry_t entry); extern struct swap_info_struct *get_swap_device(swp_entry_t entry); -sector_t swap_folio_sector(struct folio *folio); +sector_t swap_entry_sector(swp_entry_t entry); =20 /* * If there is an existing swap slot reference (swap entry) and the caller diff --git a/mm/page_io.c b/mm/page_io.c index b1894cd014b3..5c780bda92bb 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -265,7 +265,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio= *folio) return AOP_WRITEPAGE_ACTIVATE; } =20 - __swap_writepage(ctx, folio); + __swap_writepage(ctx, folio, folio->swap); return 0; out_unlock: folio_unlock(folio); @@ -326,6 +326,7 @@ struct swap_iocb { struct bio_vec bvecs[SWAP_CLUSTER_MAX]; int nr_bvecs; int len; + swp_entry_t entry; /* first slot in the batch; addresses the IO */ }; static mempool_t *sio_pool; =20 @@ -343,24 +344,22 @@ int sio_pool_init(void) } =20 static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio, - int rw) + swp_entry_t phys, int rw) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); - struct bio_vec *last_bv =3D &ctx->sio->bvecs[ctx->sio->nr_bvecs - 1]; - struct folio *prev_folio =3D bvec_folio(last_bv); - size_t prev_folio_size =3D folio_size(prev_folio); + struct swap_info_struct *sis =3D __swap_entry_to_info(phys); =20 if (ctx->sis !=3D sis) return false; - return sis->ops->can_merge(folio, prev_folio, prev_folio_size, rw); + return sis->ops->can_merge(folio, phys, ctx->sio, rw); } =20 -static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, i= nt rw) +static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys, int rw) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); + struct swap_info_struct *sis =3D __swap_entry_to_info(phys); struct swap_iocb *sio =3D ctx->sio; =20 - if (sio && !swap_can_merge(ctx, folio, rw)) { + if (sio && !swap_can_merge(ctx, folio, phys, rw)) { if (rw =3D=3D WRITE) swap_write_submit(ctx); else @@ -373,6 +372,7 @@ static void swap_add_folio(struct swap_io_ctx *ctx, str= uct folio *folio, int rw) ctx->sio =3D sio =3D mempool_alloc(sio_pool, GFP_NOIO); sio->nr_bvecs =3D 0; sio->len =3D 0; + sio->entry =3D phys; } bvec_set_folio(&sio->bvecs[sio->nr_bvecs], folio, folio_size(folio), 0); sio->len +=3D folio_size(folio); @@ -384,7 +384,8 @@ static void swap_add_folio(struct swap_io_ctx *ctx, str= uct folio *folio, int rw) } } =20 -void __swap_writepage(struct swap_io_ctx *ctx, struct folio *folio) +void __swap_writepage(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys) { VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); =20 @@ -400,7 +401,7 @@ void __swap_writepage(struct swap_io_ctx *ctx, struct f= olio *folio) =20 folio_start_writeback(folio); folio_unlock(folio); - swap_add_folio(ctx, folio, WRITE); + swap_add_folio(ctx, folio, phys, WRITE); } =20 /* @@ -504,7 +505,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio) =20 /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); - swap_add_folio(ctx, folio, READ); + swap_add_folio(ctx, folio, folio->swap, READ); =20 finish: if (workingset) { @@ -617,7 +618,7 @@ static void swap_bdev_submit_write(struct swap_io_ctx *= ctx) bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs), REQ_OP_WRITE | REQ_SWAP); bio->bi_iter.bi_size =3D sio->len; - bio->bi_iter.bi_sector =3D swap_folio_sector(bio_first_folio_all(bio)); + bio->bi_iter.bi_sector =3D swap_entry_sector(sio->entry); bio_associate_blkg_from_page(bio, bio_first_folio_all(bio)); =20 if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) { @@ -637,7 +638,7 @@ static void swap_bdev_submit_read(struct swap_io_ctx *c= tx) bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs), REQ_OP_READ); bio->bi_iter.bi_size =3D sio->len; - bio->bi_iter.bi_sector =3D swap_folio_sector(bio_first_folio_all(bio)); + bio->bi_iter.bi_sector =3D swap_entry_sector(sio->entry); =20 if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) { /* @@ -655,13 +656,14 @@ static void swap_bdev_submit_read(struct swap_io_ctx = *ctx) } } =20 -static bool swap_bdev_can_merge(struct folio *folio, struct folio *prev_fo= lio, - size_t prev_folio_size, int rw) +static bool swap_bdev_can_merge(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw) { - if (swap_folio_sector(folio) !=3D - swap_folio_sector(prev_folio) + (prev_folio_size >> SECTOR_SHIFT)) + if (swap_entry_sector(phys) !=3D + swap_entry_sector(sio->entry) + (sio->len >> SECTOR_SHIFT)) return false; - if (rw =3D=3D WRITE && !folio_blkg_can_merge(folio, prev_folio)) + if (rw =3D=3D WRITE && !folio_blkg_can_merge(folio, + bvec_folio(&sio->bvecs[sio->nr_bvecs - 1]))) return false; return true; } @@ -679,7 +681,7 @@ static void swap_fs_submit(struct swap_io_ctx *ctx, int= rw) int ret; =20 init_sync_kiocb(&sio->iocb, ctx->sis->swap_file); - sio->iocb.ki_pos =3D swap_dev_pos(bvec_folio(&sio->bvecs[0])->swap); + sio->iocb.ki_pos =3D swap_dev_pos(sio->entry); if (rw =3D=3D WRITE) sio->iocb.ki_complete =3D swap_fs_write_complete; else @@ -702,11 +704,10 @@ static void swap_fs_submit_read(struct swap_io_ctx *c= tx) swap_fs_submit(ctx, READ); } =20 -static bool swap_fs_can_merge(struct folio *folio, struct folio *prev_foli= o, - size_t prev_folio_size, int rw) +static bool swap_fs_can_merge(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw) { - return swap_dev_pos(folio->swap) =3D=3D - swap_dev_pos(prev_folio->swap) + prev_folio_size; + return swap_dev_pos(phys) =3D=3D swap_dev_pos(sio->entry) + sio->len; } =20 static const struct swap_ops swap_fs_ops =3D { diff --git a/mm/swap.h b/mm/swap.h index d241e81e967e..88ca9be71b7e 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -103,8 +103,8 @@ struct swap_io_ctx { struct swap_ops { unsigned int flags; =20 - bool (*can_merge)(struct folio *folio, struct folio *prev_folio, - size_t prev_folio_size, int rw); + bool (*can_merge)(struct folio *folio, swp_entry_t phys, + struct swap_iocb *sio, int rw); void (*submit_write)(struct swap_io_ctx *ctx); void (*submit_read)(struct swap_io_ctx *ctx); }; @@ -313,7 +313,8 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio); void swap_read_submit(struct swap_io_ctx *ctx); void swap_write_submit(struct swap_io_ctx *ctx); int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio); -void __swap_writepage(struct swap_io_ctx *ctx, struct folio *folio); +void __swap_writepage(struct swap_io_ctx *ctx, struct folio *folio, + swp_entry_t phys); =20 /* linux/mm/swap_state.c */ extern struct address_space swap_space __read_mostly; diff --git a/mm/swapfile.c b/mm/swapfile.c index a26cfe5751c8..b8fdb426514c 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -355,14 +355,14 @@ offset_to_swap_extent(struct swap_info_struct *sis, u= nsigned long offset) BUG(); } =20 -sector_t swap_folio_sector(struct folio *folio) +sector_t swap_entry_sector(swp_entry_t entry) { - struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); + struct swap_info_struct *sis =3D __swap_entry_to_info(entry); struct swap_extent *se; sector_t sector; pgoff_t offset; =20 - offset =3D swp_offset(folio->swap); + offset =3D swp_offset(entry); se =3D offset_to_swap_extent(sis, offset); sector =3D se->start_block + (offset - se->start_page); return sector << (PAGE_SHIFT - 9); diff --git a/mm/zswap.c b/mm/zswap.c index e19bde9df722..789079c3945b 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1076,7 +1076,7 @@ static int zswap_writeback_entry(struct zswap_entry *= entry, folio_set_reclaim(folio); =20 /* start writeback */ - __swap_writepage(&ctx, folio); + __swap_writepage(&ctx, folio, folio->swap); swap_write_submit(&ctx); =20 out: --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-ot1-f52.google.com (mail-ot1-f52.google.com [209.85.210.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E782C3AE197 for ; Thu, 6 Aug 2026 18:43:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041787; cv=none; b=dAa9tYK0rhWZRjhkZV7Y49OnJ4N/LOTRCdWLRVRVdw9ylSCwGWfm4untcuPChY0lwg4bShlWapvSTKB7TFBEIII5Q7h1i7wGYWKdNaNs5jw+JFcC9kNZoGkYeVofIe4pXWAwKbvSnf70KU7Uk0Ai5lZm+sH/DaJuT8eo+zkQFV4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041787; c=relaxed/simple; bh=KV+u7lcKkOTUwfrTC0yeM7I9pnsd3IZNHNNfy4poLcg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PQk9VFcxdq4gQu23btBh/oE29RsHSiWlTJg7sza50uclZEvKv131WAurE0eni5u7MdgPahLmE62KpzeD2tzzdvVOykWpJHGwf4qAbH7CMlzTb5gJLO2JWGO3ozJmWWegdY0uX9LT8snuZ8A0rZSRSLaVP3VC4AUnbjk36NF+M7M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=e6quyqs4; arc=none smtp.client-ip=209.85.210.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="e6quyqs4" Received: by mail-ot1-f52.google.com with SMTP id 46e09a7af769-7eb68bdf53aso1079768a34.3 for ; Thu, 06 Aug 2026 11:43:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041783; x=1786646583; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=q8wSea+R53zhaHFEktUy45t+1cs96tBlwwCzw1xF8Mw=; b=e6quyqs4g5/t7t4D2Y6KcAT25M2wKRUdVv1aD4lhKVCdI4EDgEb/qN2C6BVXXzns82 dCF0N/OqLm8y30IADrPETDMaoHC+ZZaVmviytqiPn0scPkDgCIqTlSE+ZlbHwU1I7NY5 rZ3jJtOyTjZalI+08obnKNj50R6AJssJcaJ2d6nmUHG4YYOwlsGYdTBkGDapQ8ffozcZ t7Zuqv9kXMJOkwwpP2aC62rYfV6xGZ4THoJOMBySHGvymmOUOBhbTdm0lOHvhWwKBbLr 78+uar7lGcZHmo+yyrdI4pQFC5XeOMzRf8yy1FqLCn1KKUz7NsQV8Q4q7rx3oEnBOUYi NBHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041783; x=1786646583; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=q8wSea+R53zhaHFEktUy45t+1cs96tBlwwCzw1xF8Mw=; b=QTMODGNcuGkxXTe3JELh91jMIZgkIOofD3BRf3UBmFtKDs8TfftIqhgbfUGGK2j2k9 wiAfTRYEGgyFjBxfU3Ch5OfCS3Aml7ZoImGgm+u3ILlfxNQMWClmapzwv20n5vlcnxXH 2jtlXPX77jHmHgAIFw0a0eGfYkLCQWqWkzwnM8BR1MwNFs0qDIcVYXFpKIJCK72Fwwpz IW6rf8/9UMX3MgvWWWQtmCH75/ncIDntObR4LC6qw3XDFqDf3t28G+dYyu/9LaIHrS8e NGJYDQ0L5FrMdNZMVR3xRxJAfm4/71Xo2CXVGujHTn2gyukNQYjRMJjtV0CpBaV/hirh eMRg== X-Forwarded-Encrypted: i=1; AHgh+Ro40dX18E2IBtDtXyCZa36eq7L8AL7EJ9QrILONj3fV/zx1ndcVe6gZZ1+OYwKDfNo2BSm+TdU7B9OhQ/o=@vger.kernel.org X-Gm-Message-State: AOJu0Yzdj4noNiqoS5y5UtxNoZUu9F5sMvVjcJ+6ZCaQExunc0kl1J8U VAfXktCMSNPWP1fJbmyUvCBoG7M8YJMznO2WcUmxsCFq49eM0G2dNcvy X-Gm-Gg: AR+sD113svk1UZfccVNWN89stxgfPFmuyBCTXzILvqUfPqSKa6CPSseQNvHR/I7iW5g l2lqee8DgWVssVW2AGzyjDcyllrdo+M9e+u8QM7yNRFSplHb6D/BEchWN101YAgXl5QLF4SErwH RgprIoXDxsYqHcdsi/DmBL3fIbXWFC3Tpw6VvXI44aHxM1VnIfDASe4UaJWNTq+GQW834zDXiGX sgUTn9ixzq7TFB6NI12Fj1MQshx64uJSj+sod98ydE5VxWDYPsW7lwaftwIPropQqJ/k0rSJfdY 9ble5m/hr7+CoGeX2LZWz4XQ/nWB+k2yagBW7xQPoVuuKj5gWjtoubgxA81gBJ+c4wc6Rq+Dwr0 4k+sWKS6vrLNzav1ee4lXNPODOHUA5PrBJOMdznyJSGs962605Af0aS/Pq1OQ+asEWfXgpgB8XX d+d8IsDUxtLI5g0nDcVAQfh9L3+ajf8GXSoROO2l0s6AX964vQFYYhE1/y5wnKcJykaExe1xQe+ 2MIeHHfyEU= X-Received: by 2002:a4a:edce:0:b0:6aa:d860:afad with SMTP id 006d021491bc7-6ae96ce17c0mr9058138eaf.15.1786041782548; Thu, 06 Aug 2026 11:43:02 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:1a::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bc76279sm210527eaf.7.2026.08.06.11.43.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:02 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 04/11] mm, swap: support physical swap as a vswap backend Date: Thu, 6 Aug 2026 11:42:47 -0700 Message-ID: <20260806184254.3790858-5-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add physical swap as a backend for the virtual swap layer. When zswap declines a page, the swapout path allocates a physical slot on demand for swap out. Each vswap entry's physical slot is tracked via a pointer-tagged swap_table entry on the physical cluster (an rmap back to the vswap entry). Physical readahead scans a whole offset window and would trip over these rmap slots, so __swap_cache_add_check() now skips swp_tb_is_pointer() entries. Nothing is lost: a backing slot is faulted through its owning vswap entry, never through the physical offset. Writeback of zswap-backed vswap entries to physical swap, and reclaim of physical slots backing cache-only vswap entries, are added in the following patches. Suggested-by: Kairui Song Signed-off-by: Nhat Pham --- include/linux/swap.h | 9 ++ mm/memory.c | 8 +- mm/page_io.c | 43 ++++-- mm/swap_state.c | 6 +- mm/swap_table.h | 55 +++++++ mm/swapfile.c | 352 +++++++++++++++++++++++++++++++++++++++---- mm/vmscan.c | 2 +- mm/vswap.h | 198 +++++++++++++++++++++++- mm/zswap.c | 2 +- 9 files changed, 627 insertions(+), 48 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 8359134d08cb..2b2bd56afffa 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -391,6 +391,15 @@ extern int swp_swapcount(swp_entry_t entry); extern struct swap_info_struct *get_swap_device(swp_entry_t entry); sector_t swap_entry_sector(swp_entry_t entry); =20 +#ifdef CONFIG_VSWAP +swp_entry_t folio_realloc_swap(struct folio *folio); +#else +static inline swp_entry_t folio_realloc_swap(struct folio *folio) +{ + return (swp_entry_t){}; +} +#endif + /* * If there is an existing swap slot reference (swap entry) and the caller * guarantees that there is no race modification of it (e.g., PTL diff --git a/mm/memory.c b/mm/memory.c index de3573b7c6b1..ba84565605a1 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -4656,13 +4656,13 @@ static inline bool should_try_to_free_swap(struct s= wap_info_struct *si, * are fast, and meanwhile, swap cache pinning the slot deferring the * release of metadata or fragmentation is a more critical issue. */ - if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) + if (swap_entry_backend_has_flag(si, folio->swap, SWP_SYNCHRONOUS_IO)) return true; /* * Non-swapfile backends cannot be reused for future swapouts. * Free the swap slot unless backed by contiguous physical swap. */ - if (is_vswap_entry(folio->swap)) + if (!folio_phys_swap_backed(folio)) return true; if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) || folio_test_mlocked(folio)) @@ -4968,7 +4968,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) swap_update_readahead(folio, vma, vmf->address); if (!folio) { /* Swapin bypasses readahead for SWP_SYNCHRONOUS_IO devices */ - if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) + if (swap_entry_backend_has_flag(si, entry, SWP_SYNCHRONOUS_IO)) folio =3D swapin_sync(entry, GFP_HIGHUSER_MOVABLE, thp_swapin_suitable_orders(vmf) | BIT(0), vmf, NULL, 0); @@ -5133,7 +5133,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) */ exclusive =3D true; } else if (exclusive && folio_test_writeback(folio) && - data_race(si->flags & SWP_STABLE_WRITES)) { + swap_entry_backend_has_flag(si, entry, SWP_STABLE_WRITES)) { /* * This is tricky: not all swap backends support * concurrent page modifications while under writeback. diff --git a/mm/page_io.c b/mm/page_io.c index 5c780bda92bb..605a66a32604 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -208,6 +208,7 @@ static void swap_zeromap_folio_clear(struct folio *foli= o) */ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio) { + swp_entry_t phys; int ret =3D 0; =20 if (folio_free_swap(folio)) @@ -240,8 +241,14 @@ int swap_writeout(struct swap_io_ctx *ctx, struct foli= o *folio) */ swap_zeromap_folio_clear(folio); =20 + /* + * For vswap: release stale non-swapfile backings (e.g. ZSWAP from a + * previous swapout cycle) so zswap_store or folio_realloc_swap + * starts on clean slots. Contiguous PHYS backing is preserved for + * reuse by folio_realloc_swap. + */ if (is_vswap_entry(folio->swap)) - folio_release_vswap_backing(folio); + folio_release_non_phys_swap_backing(folio); =20 if (zswap_store(folio)) { count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT); @@ -257,12 +264,19 @@ int swap_writeout(struct swap_io_ctx *ctx, struct fol= io *folio) rcu_read_unlock(); =20 /* - * A vswap folio that reaches here could not be stored to a backend - * (zswap) and has no physical slot to write to, so keep it dirty. + * A vswap folio with no backend needs a physical slot to write to. + * zswap_store rolled back any partial vtable state on failure, so + * PHYS backing from a prior cycle is still there to reuse. If none + * is free, keep it dirty. */ if (is_vswap_entry(folio->swap)) { - folio_mark_dirty(folio); - return AOP_WRITEPAGE_ACTIVATE; + phys =3D folio_realloc_swap(folio); + if (!phys.val) { + folio_mark_dirty(folio); + return AOP_WRITEPAGE_ACTIVATE; + } + __swap_writepage(ctx, folio, phys); + return 0; } =20 __swap_writepage(ctx, folio, folio->swap); @@ -474,6 +488,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio) bool workingset =3D folio_test_workingset(folio); unsigned long pflags; bool in_thrashing; + swp_entry_t phys; =20 VM_BUG_ON_FOLIO(!folio_test_swapcache(folio) && !synchronous, folio); VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); @@ -498,14 +513,24 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct = folio *folio) if (zswap_load(folio) !=3D -ENOENT) goto finish; =20 - if (unlikely(swap_is_vswap(sis))) { - folio_unlock(folio); - goto finish; + /* + * Resolve the physical slot to read from. A vswap entry keeps + * folio->swap virtual, so map it to its physical backing; a folio with + * no backing has nothing to read. + */ + if (swap_is_vswap(sis)) { + phys =3D vswap_to_phys(folio->swap); + if (!phys.val) { + folio_unlock(folio); + goto finish; + } + } else { + phys =3D folio->swap; } =20 /* We have to read from slower devices. Increase zswap protection. */ zswap_folio_swapin(folio); - swap_add_folio(ctx, folio, folio->swap, READ); + swap_add_folio(ctx, folio, phys, READ); =20 finish: if (workingset) { diff --git a/mm/swap_state.c b/mm/swap_state.c index 9f6377b32911..c61bb3eef62a 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -185,6 +185,9 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, return -ENOENT; ci_off =3D swp_cluster_offset(targ_entry); old_tb =3D __swap_table_get(ci, ci_off); + /* Physical readahead can hit a vswap-backing rmap slot; skip it. */ + if (swp_tb_is_pointer(old_tb)) + return -ENOENT; if (swp_tb_is_folio(old_tb)) return -EEXIST; if (!__swp_tb_get_count(old_tb)) @@ -209,7 +212,8 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, ci_end =3D ci_off + nr; do { old_tb =3D __swap_table_get(ci, ci_off); - if (unlikely(swp_tb_is_folio(old_tb) || + if (unlikely(swp_tb_is_pointer(old_tb) || + swp_tb_is_folio(old_tb) || !__swp_tb_get_count(old_tb) || is_zero !=3D __swap_table_test_zero(ci, ci_off) || (memcg_id && *memcg_id !=3D __swap_cgroup_get(ci, ci_off)))) diff --git a/mm/swap_table.h b/mm/swap_table.h index fd7f0fb9836a..5b0eca07a821 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -4,8 +4,11 @@ =20 #include #include +#include #include "swap.h" =20 +struct zswap_entry; + /* A typical flat array in each cluster as swap table */ struct swap_table { atomic_long_t entries[SWAPFILE_CLUSTER]; @@ -368,4 +371,56 @@ static inline unsigned short __swap_cgroup_clear(struc= t swap_cluster_info *ci, } #endif =20 +/* + * Pointer-tagged swap table entry: rmap for vswap-backing physical slots. + * + * On physical clusters, a Pointer-tagged entry stores the offset of the + * vswap entry that owns this physical slot (the reverse map). Only the + * offset is stored; the swap type is implicit (always vswap_si->type, + * since there is exactly one vswap device). + * + * Pointer: |---- vswap offset ----|100| + */ +#ifdef CONFIG_VSWAP +extern struct swap_info_struct *vswap_si; + +#define SWP_TB_PTR_MARK_BITS 3 +#define SWP_TB_PTR_MARK 0b100UL +#define SWP_TB_PTR_MARK_MASK ((1UL << SWP_TB_PTR_MARK_BITS) - 1) +#define SWP_RMAP_ENTRY_MASK (~SWP_TB_PTR_MARK_MASK) + +static inline bool swp_tb_is_pointer(unsigned long swp_tb) +{ + return (swp_tb & SWP_TB_PTR_MARK_MASK) =3D=3D SWP_TB_PTR_MARK; +} + +static inline unsigned long swp_entry_to_swp_tb_ptr(swp_entry_t entry) +{ + return (swp_offset(entry) << SWP_TB_PTR_MARK_BITS) | SWP_TB_PTR_MARK; +} + +static inline swp_entry_t swp_tb_ptr_to_swp_entry(unsigned long swp_tb) +{ + unsigned long offset; + + VM_WARN_ON(!swp_tb_is_pointer(swp_tb)); + offset =3D (swp_tb & SWP_RMAP_ENTRY_MASK) >> SWP_TB_PTR_MARK_BITS; + return swp_entry(vswap_si->type, offset); +} +#else +static inline bool swp_tb_is_pointer(unsigned long swp_tb) +{ + return false; +} +static inline unsigned long swp_entry_to_swp_tb_ptr(swp_entry_t entry) +{ + return 0; +} +static inline swp_entry_t swp_tb_ptr_to_swp_entry(unsigned long swp_tb) +{ + return (swp_entry_t){}; +} + +#endif /* CONFIG_VSWAP */ + #endif diff --git a/mm/swapfile.c b/mm/swapfile.c index b8fdb426514c..0874f57d3124 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -261,7 +261,7 @@ static int __try_to_reclaim_swap(struct swap_info_struc= t *si, need_reclaim =3D ((flags & TTRS_ANYWAY) || ((flags & TTRS_UNMAPPED) && !folio_mapped(folio)) || ((flags & TTRS_FULL) && mem_cgroup_swap_full(folio) && - !is_vswap_entry(folio->swap))); + folio_phys_swap_backed(folio))); if (!need_reclaim || !folio_swapcache_freeable(folio)) goto out_unlock; =20 @@ -1013,6 +1013,8 @@ static bool __swap_cluster_alloc_entries(struct swap_= info_struct *si, { unsigned int order; unsigned long nr_pages; + swp_entry_t vswap_entry, v; + unsigned int i; =20 lockdep_assert_held(&ci->lock); =20 @@ -1032,8 +1034,26 @@ static bool __swap_cluster_alloc_entries(struct swap= _info_struct *si, order =3D folio_order(folio); nr_pages =3D 1 << order; swap_cluster_assert_empty(ci, ci_off, nr_pages, false); - __swap_cache_add_folio(ci, folio, swp_entry(si->type, - ci_off + cluster_offset(si, ci))); + if (folio_test_swapcache(folio)) { + /* + * Folio already in the swap cache: we are allocating + * physical backing for its vswap entry. Point each + * physical slot back at its own vswap entry + * (Pointer-tagged rmap). + */ + VM_WARN_ON(!is_vswap_entry(folio->swap)); + vswap_entry =3D folio->swap; + for (i =3D 0; i < nr_pages; i++) { + v =3D vswap_entry; + v.val +=3D i; + __swap_table_set(ci, ci_off + i, + swp_entry_to_swp_tb_ptr(v)); + } + } else { + __swap_cache_add_folio(ci, folio, + swp_entry(si->type, + ci_off + cluster_offset(si, ci))); + } } else if (IS_ENABLED(CONFIG_HIBERNATION)) { order =3D 0; nr_pages =3D 1; @@ -1538,12 +1558,14 @@ static bool get_swap_device_info(struct swap_info_s= truct *si) * Fast path try to get swap entries with specified order from current * CPU's swap entry pool (a cluster). */ -static bool swap_alloc_fast(struct folio *folio) +static swp_entry_t swap_alloc_fast(struct folio *folio) { unsigned int order =3D folio_order(folio); struct swap_cluster_info *ci; struct swap_info_struct *si; - unsigned int offset; + unsigned long offset, found =3D 0; + + lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); =20 /* * Once allocated, swap_info_struct will never be completely freed, @@ -1552,25 +1574,28 @@ static bool swap_alloc_fast(struct folio *folio) si =3D this_cpu_read(percpu_swap_cluster.si[order]); offset =3D this_cpu_read(percpu_swap_cluster.offset[order]); if (!si || !offset || !get_swap_device_info(si)) - return false; + return (swp_entry_t){}; =20 ci =3D swap_cluster_lock(si, offset); if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset =3D cluster_offset(si, ci); - alloc_swap_scan_cluster(si, ci, folio, offset); + found =3D alloc_swap_scan_cluster(si, ci, folio, offset); } else if (ci) { swap_cluster_unlock(ci); } =20 put_swap_device(si); - return folio_test_swapcache(folio); + if (found) + return swp_entry(si->type, found); + return (swp_entry_t){}; } =20 /* Rotate the device and switch to a new cluster */ -static void swap_alloc_slow(struct folio *folio) +static swp_entry_t swap_alloc_slow(struct folio *folio) { struct swap_info_struct *si, *next; + unsigned long found; =20 spin_lock(&swap_avail_lock); start_over: @@ -1579,12 +1604,12 @@ static void swap_alloc_slow(struct folio *folio) plist_requeue(&si->avail_list, &swap_avail_head); spin_unlock(&swap_avail_lock); if (get_swap_device_info(si)) { - cluster_alloc_swap_entry(si, folio); + found =3D cluster_alloc_swap_entry(si, folio); put_swap_device(si); - if (folio_test_swapcache(folio)) - return; + if (found) + return swp_entry(si->type, found); if (folio_test_large(folio)) - return; + return (swp_entry_t){}; } =20 spin_lock(&swap_avail_lock); @@ -1602,6 +1627,7 @@ static void swap_alloc_slow(struct folio *folio) goto start_over; } spin_unlock(&swap_avail_lock); + return (swp_entry_t){}; } =20 /* @@ -1982,13 +2008,12 @@ int folio_alloc_swap(struct folio *folio) } } =20 - /* Without zswap a vswap entry has nowhere to go on writeout. */ - if (zswap_is_enabled() && vswap_alloc(folio)) + if (vswap_alloc(folio)) goto done; =20 again: local_lock(&percpu_swap_cluster.lock); - if (!swap_alloc_fast(folio)) + if (!swap_alloc_fast(folio).val) swap_alloc_slow(folio); local_unlock(&percpu_swap_cluster.lock); =20 @@ -2010,6 +2035,11 @@ int folio_alloc_swap(struct folio *folio) =20 #ifdef CONFIG_VSWAP =20 +static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, + struct swap_cluster_info *pci, + unsigned int ci_start, + unsigned int nr_pages); + /** * __vswap_release_backing - release the backing of a range of vtable slots * @ci: the locked vswap cluster @@ -2026,8 +2056,12 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, unsigned int ci_start, unsigned int nr) { struct swap_cluster_info_dynamic *ci_dyn; + struct swap_info_struct *psi; + unsigned long phys_start =3D 0, phys_end =3D 0; + unsigned int phys_type =3D 0; unsigned int ci_off; unsigned long vt; + swp_entry_t phys; =20 lockdep_assert_held(&ci->lock); ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); @@ -2035,7 +2069,37 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { vt =3D __vtable_get(ci_dyn, ci_off); =20 + /* + * Flush batched physical slots when the next entry + * breaks contiguity, changes type/device, or would + * cross a SWAPFILE_CLUSTER boundary (the free helper + * operates on a single cluster). + */ + if (phys_start !=3D phys_end && + (vtable_type(vt) !=3D VSWAP_SWAPFILE || + swp_type(vtable_to_phys(vt)) !=3D phys_type || + swp_offset(vtable_to_phys(vt)) !=3D phys_end || + phys_end % SWAPFILE_CLUSTER =3D=3D 0)) { + psi =3D __swap_type_to_info(phys_type); + __swap_cluster_free_phys_backing(psi, + __swap_entry_to_cluster( + swp_entry(phys_type, phys_start)), + phys_start % SWAPFILE_CLUSTER, + phys_end - phys_start); + phys_start =3D phys_end =3D 0; + } + switch (vtable_type(vt)) { + case VSWAP_SWAPFILE: + if (phys_start =3D=3D phys_end) { + phys =3D vtable_to_phys(vt); + phys_start =3D swp_offset(phys); + phys_end =3D phys_start + 1; + phys_type =3D swp_type(phys); + } else { + phys_end++; + } + break; case VSWAP_ZSWAP: zswap_entry_free(vtable_to_zswap(vt)); break; @@ -2051,6 +2115,15 @@ void __vswap_release_backing(struct swap_cluster_inf= o *ci, if (__swap_table_test_zero(ci, ci_off)) __swap_table_clear_zero(ci, ci_off); } + + if (phys_start !=3D phys_end) { + psi =3D __swap_type_to_info(phys_type); + __swap_cluster_free_phys_backing(psi, + __swap_entry_to_cluster( + swp_entry(phys_type, phys_start)), + phys_start % SWAPFILE_CLUSTER, + phys_end - phys_start); + } } =20 /** @@ -2080,6 +2153,106 @@ void folio_release_vswap_backing(struct folio *foli= o) spin_unlock(&ci->lock); } =20 +/** + * folio_release_non_phys_swap_backing() - Drop a folio's non-physical vsw= ap backing. + * @folio: the folio, occupying a virtual swap entry. + * + * Release any ZSWAP or zero-filled backing recorded for @folio's virtual + * swap entry, leaving the slots empty so the writeout path can install fr= esh + * physical backing. If the first slot is already VSWAP_SWAPFILE or + * VSWAP_NONE, nothing is released: physical backing is kept for reuse. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + */ +void folio_release_non_phys_swap_backing(struct folio *folio) +{ + struct swap_cluster_info *ci; + struct swap_cluster_info_dynamic *ci_dyn; + int nr =3D folio_nr_pages(folio); + unsigned int voff; + unsigned long vt; + enum vswap_backing_type type; + + ci =3D __swap_entry_to_cluster(folio->swap); + if (!ci) + return; + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + voff =3D swp_cluster_offset(folio->swap); + + spin_lock(&ci->lock); + vt =3D __vtable_get(ci_dyn, voff); + type =3D vtable_type(vt); + + if (type =3D=3D VSWAP_SWAPFILE || type =3D=3D VSWAP_NONE) { + spin_unlock(&ci->lock); + return; + } + + __vswap_release_backing(ci, voff, nr); + spin_unlock(&ci->lock); +} + +/** + * folio_realloc_swap() - Back a virtual swap folio with a physical swap s= lot. + * @folio: the folio, occupying a virtual swap entry. + * + * Ensure @folio's virtual swap entry has physical (swapfile) backing, + * allocating a physical slot on demand if it has none. Called from the + * writeout path and from zswap writeback to move a vswap entry onto a real + * swapfile slot. If @folio is already physically backed, the existing + * physical entry is returned unchanged. + * + * Context: Caller must hold the folio lock; @folio must be in the swap ca= che + * and occupy a virtual swap entry. + * Return: The physical swap entry now backing @folio, or an empty entry + * (.val =3D=3D 0) on failure. + */ +swp_entry_t folio_realloc_swap(struct folio *folio) +{ + swp_entry_t vswap_entry =3D folio->swap; + struct swap_cluster_info *ci; + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + swp_entry_t phys_entry =3D {}; + swp_entry_t pe; + int i, nr =3D folio_nr_pages(folio); + + VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); + VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); + VM_WARN_ON(!is_vswap_entry(vswap_entry)); + + phys_entry =3D vswap_to_phys(vswap_entry); + if (phys_entry.val) + return phys_entry; + + local_lock(&percpu_swap_cluster.lock); + phys_entry =3D swap_alloc_fast(folio); + if (!phys_entry.val) + phys_entry =3D swap_alloc_slow(folio); + local_unlock(&percpu_swap_cluster.lock); + + if (!phys_entry.val) + return (swp_entry_t){}; + + voff =3D swp_cluster_offset(vswap_entry); + + ci =3D __swap_entry_to_cluster(vswap_entry); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + spin_lock(&ci->lock); + /* + * Install PHYS backing without freeing any prior contents of the + * vtable. Releasing the old backing is the caller's job: it may + * still need the slot, or may have released it already. + */ + for (i =3D 0; i < nr; i++) { + pe.val =3D phys_entry.val + i; + __vtable_set(ci_dyn, voff + i, vtable_mk_phys(pe)); + } + spin_unlock(&ci->lock); + + return phys_entry; +} #endif /* CONFIG_VSWAP */ =20 /** @@ -2207,6 +2380,63 @@ struct swap_info_struct *get_swap_device(swp_entry_t= entry) return NULL; } =20 +#ifdef CONFIG_VSWAP +/* + * Clear swap table entries to NULL and reset zero flags. + * Does not touch memcg or count - caller handles those. + */ +static void __swap_cluster_clear_table(struct swap_cluster_info *ci, + unsigned int ci_start, + unsigned int nr_pages) +{ + unsigned int ci_off; + + lockdep_assert_held(&ci->lock); + for (ci_off =3D ci_start; ci_off < ci_start + nr_pages; ci_off++) { + __swap_table_set(ci, ci_off, null_to_swp_tb()); + if (!SWAP_TABLE_HAS_ZEROFLAG) + __swap_table_clear_zero(ci, ci_off); + } +} +#endif + +/* + * Common tail for freeing swap slots: device-level accounting + * and cluster list management. + */ +static void __swap_cluster_finish_free(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_start, + unsigned int nr_pages) +{ + lockdep_assert_held(&ci->lock); + swap_range_free(si, cluster_offset(si, ci) + ci_start, nr_pages); + swap_cluster_assert_empty(ci, ci_start, nr_pages, false); + + if (!ci->count) + free_cluster(si, ci); + else + partial_free_cluster(si, ci); +} + +#ifdef CONFIG_VSWAP +/* + * Free physical swap slots that were backing vswap entries (Pointer-tagge= d). + */ +static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, + struct swap_cluster_info *pci, + unsigned int ci_start, + unsigned int nr_pages) +{ + spin_lock_nested(&pci->lock, SINGLE_DEPTH_NESTING); + VM_WARN_ON(pci->count < nr_pages); + pci->count -=3D nr_pages; + __swap_cluster_clear_table(pci, ci_start, nr_pages); + __swap_cluster_finish_free(psi, pci, ci_start, nr_pages); + swap_cluster_unlock(pci); +} +#endif + /* * Free a set of swap slots after their swap count dropped to zero, or wil= l be * zero after putting the last ref (saves one __swap_cluster_put_entry cal= l). @@ -2218,7 +2448,6 @@ void __swap_cluster_free_entries(struct swap_info_str= uct *si, unsigned long old_tb; unsigned short batch_id =3D 0, id_cur; unsigned int ci_off =3D ci_start, ci_end =3D ci_start + nr_pages; - unsigned long ci_head =3D cluster_offset(si, ci); unsigned int batch_off =3D ci_off; =20 VM_WARN_ON(ci->count < nr_pages); @@ -2256,13 +2485,7 @@ void __swap_cluster_free_entries(struct swap_info_st= ruct *si, if (batch_id) mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); =20 - swap_range_free(si, ci_head + ci_start, nr_pages); - swap_cluster_assert_empty(ci, ci_start, nr_pages, false); - - if (!ci->count) - free_cluster(si, ci); - else - partial_free_cluster(si, ci); + __swap_cluster_finish_free(si, ci, ci_start, nr_pages); } =20 int __swap_count(swp_entry_t entry) @@ -3041,19 +3264,88 @@ static unsigned int find_next_to_unuse(struct swap_= info_struct *si, =20 static int try_to_unuse(unsigned int type) { + struct mempolicy *mpol =3D get_task_policy(current); struct mm_struct *prev_mm; struct mm_struct *mm; struct list_head *p; int retval =3D 0; struct swap_info_struct *si =3D swap_info[type]; struct folio *folio; - swp_entry_t entry; - unsigned int i; + struct swap_io_ctx ctx; + swp_entry_t entry, vswap_entry; + unsigned long swp_tb; + unsigned int i, j; =20 if (!swap_usage_in_pages(si)) goto success; =20 retry: + /* + * Free vswap-backing slots (Pointer-tagged) first. Walk physical + * clusters, read the vswap entry from the rmap, ensure the data + * is in the swap cache, and transition PHYS to FOLIO. No page table + * walk needed - just free the physical backing. + */ + i =3D 0; + while (IS_ENABLED(CONFIG_VSWAP) && + swap_usage_in_pages(si) && + !signal_pending(current) && + (i =3D find_next_to_unuse(si, i)) !=3D 0) { + swp_entry_t phys; + + swp_tb =3D swap_table_get(__swap_offset_to_cluster(si, i), + i % SWAPFILE_CLUSTER); + if (!swp_tb_is_pointer(swp_tb)) + continue; + + vswap_entry =3D swp_tb_ptr_to_swp_entry(swp_tb); + + folio =3D swap_cache_get_folio(vswap_entry); + if (!folio) { + folio =3D swap_cache_alloc_folio(vswap_entry, + GFP_KERNEL, BIT(0), NULL, + mpol, NO_INTERLEAVE_INDEX); + if (IS_ERR(folio)) + continue; + ctx =3D (struct swap_io_ctx){}; + swap_read_folio(&ctx, folio); + swap_read_submit(&ctx); + folio_lock(folio); + } else { + folio_lock(folio); + } + + if (!folio_matches_swap_entry(folio, vswap_entry)) { + folio_unlock(folio); + folio_put(folio); + continue; + } + + /* + * Re-validate under folio lock: rmap holds folio->swap + j + * for some j in [0, nr_pages). Check folio->swap still maps + * to the contiguous physical run that includes our slot i. + */ + j =3D vswap_entry.val - folio->swap.val; + phys =3D vswap_to_phys(folio->swap); + if (!phys.val || swp_type(phys) !=3D type || + swp_offset(phys) + j !=3D i || + j >=3D folio_nr_pages(folio)) { + folio_unlock(folio); + folio_put(folio); + continue; + } + + folio_wait_writeback(folio); + folio_release_vswap_backing(folio); + folio_mark_dirty(folio); + folio_unlock(folio); + folio_put(folio); + } + + if (!swap_usage_in_pages(si)) + goto success; + retval =3D shmem_unuse(type); if (retval) return retval; @@ -3097,6 +3389,14 @@ static int try_to_unuse(unsigned int type) =20 entry =3D swp_entry(type, i); =20 + if (IS_ENABLED(CONFIG_VSWAP)) { + swp_tb =3D swap_table_get( + __swap_offset_to_cluster(si, i), + i % SWAPFILE_CLUSTER); + if (swp_tb_is_pointer(swp_tb)) + continue; + } + folio =3D swap_cache_get_folio(entry); if (!folio) continue; diff --git a/mm/vmscan.c b/mm/vmscan.c index 78ec51f53757..f3f9e3993215 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1530,7 +1530,7 @@ static unsigned int shrink_folio_list(struct list_hea= d *folio_list, * space if we are running out. */ if (folio_test_swapcache(folio) && - ((mem_cgroup_swap_full(folio) && !is_vswap_entry(folio->swap)) || + ((mem_cgroup_swap_full(folio) && folio_phys_swap_backed(folio)) || folio_test_mlocked(folio))) folio_free_swap(folio); VM_BUG_ON_FOLIO(folio_test_active(folio), folio); diff --git a/mm/vswap.h b/mm/vswap.h index 6d25e0911fa9..239b47b577d5 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -19,6 +19,7 @@ struct zswap_entry; enum vswap_backing_type { VSWAP_NONE =3D 0, VSWAP_ZSWAP =3D 1, + VSWAP_SWAPFILE =3D 2, VSWAP_ZERO, VSWAP_FOLIO, }; @@ -27,8 +28,6 @@ enum vswap_backing_type { =20 #include "swap_table.h" =20 -extern struct swap_info_struct *vswap_si; - static inline bool is_vswap_entry(swp_entry_t entry) { return swap_is_vswap(__swap_entry_to_info(entry)); @@ -43,11 +42,15 @@ bool vswap_is_enabled(void); * pointer for a virtual swap slot. Tag in low 3 bits, payload in * upper 61 bits. * - * NONE: |----- 0000 ------|000| - no separate backend pointer - * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low bi= ts) + * NONE: |----- 0000 ------|000| - no separate backend pointer + * ZSWAP: |--- zswap_entry* |001| - compressed in zswap (tag in low = bits) + * SWAPFILE: |- type:5,off:56 -|010| - on a physical swapfile * - * Pointer payloads (ZSWAP) are stored directly with the tag OR'd into the - * low bits (kernel pointers are >=3D 8-byte aligned, same approach as xar= ray). + * SWAPFILE packs swp_type in the top MAX_SWAPFILES_SHIFT bits and swp_off= set in + * the middle VTABLE_PHYS_OFF_BITS bits, both above the tag, so the type is + * not shifted off the word. Pointer payloads (ZSWAP) are stored directly = with + * the tag OR'd into the low bits (kernel pointers are >=3D 8-byte aligned= , same + * approach as xarray). * * vtable[i] =3D NONE does not by itself mean "free". The swap_table entry * and the per-slot zero flag carry the rest of the state. The full @@ -86,6 +89,23 @@ static inline enum vswap_backing_type vtable_type(unsign= ed long vt) return vt & VTABLE_TAG_MASK; } =20 +/* swp_offset field width in a physical backend slot; layout described abo= ve. */ +#define VTABLE_PHYS_OFF_BITS (BITS_PER_LONG - VTABLE_TAG_BITS - MAX_SWAPFI= LES_SHIFT) + +static inline unsigned long vtable_mk_phys(swp_entry_t entry) +{ + VM_WARN_ON_ONCE(swp_offset(entry) >> VTABLE_PHYS_OFF_BITS); + return ((unsigned long)swp_type(entry) << (VTABLE_TAG_BITS + VTABLE_PHYS_= OFF_BITS)) | + (swp_offset(entry) << VTABLE_TAG_BITS) | VSWAP_SWAPFILE; +} + +static inline swp_entry_t vtable_to_phys(unsigned long vt) +{ + VM_WARN_ON(vtable_type(vt) !=3D VSWAP_SWAPFILE); + return swp_entry(vt >> (VTABLE_TAG_BITS + VTABLE_PHYS_OFF_BITS), + (vt >> VTABLE_TAG_BITS) & ((1UL << VTABLE_PHYS_OFF_BITS) - 1)); +} + static inline struct zswap_entry *vtable_to_zswap(unsigned long vt) { VM_WARN_ON(vtable_type(vt) !=3D VSWAP_ZSWAP); @@ -130,6 +150,33 @@ vswap_lock_cluster(swp_entry_t entry, unsigned int *vo= ff) return ci_dyn; } =20 +/** + * vswap_to_phys - resolve a vswap entry's physical swap backing + * @entry: the virtual swap entry + * + * Context: takes and drops the vswap cluster lock internally. + * Return: the backing physical swp_entry_t, or the null entry (.val =3D= =3D 0) + * when @entry has no physical backing (NONE/ZSWAP/ZERO). + */ +static inline swp_entry_t vswap_to_phys(swp_entry_t entry) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + unsigned long vt; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) + return (swp_entry_t){}; + + vt =3D __vtable_get(ci_dyn, voff); + spin_unlock(&ci_dyn->ci.lock); + + if (vtable_type(vt) !=3D VSWAP_SWAPFILE) + return (swp_entry_t){}; + + return vtable_to_phys(vt); +} + void __vswap_release_backing(struct swap_cluster_info *ci, unsigned int ci_start, unsigned int nr); =20 @@ -182,6 +229,103 @@ static inline struct zswap_entry *vswap_zswap_load(sw= p_entry_t entry) } =20 void folio_release_vswap_backing(struct folio *folio); +void folio_release_non_phys_swap_backing(struct folio *folio); + +/* + * Walk nr vtable slots starting at voff in ci_dyn. Returns the prefix + * length of slots sharing one effective backing type. For SWAPFILE, + * the prefix is also restricted to contiguous offsets in the same + * swapfile. + * + * Effective type per slot: + * vtable=3DNONE + zero flag set -> VSWAP_ZERO + * vtable=3DNONE + swap_table PFN tag -> VSWAP_FOLIO + * vtable=3DNONE + neither -> VSWAP_NONE + * vtable=3DSWAPFILE -> VSWAP_SWAPFILE + * vtable=3DZSWAP -> VSWAP_ZSWAP + * + * *typep returns the effective type of slot 0. Caller holds + * ci_dyn->ci.lock. + */ +static inline int __vswap_check_backing(struct swap_cluster_info_dynamic *= ci_dyn, + unsigned int voff, int nr, + enum vswap_backing_type *typep) +{ + enum vswap_backing_type first_type =3D VSWAP_NONE; + enum vswap_backing_type slot_type; + swp_entry_t first_phys =3D {}; + unsigned long vt, swap_tb; + int i; + + lockdep_assert_held(&ci_dyn->ci.lock); + + for (i =3D 0; i < nr; i++) { + vt =3D __vtable_get(ci_dyn, voff + i); + if (vtable_type(vt) =3D=3D VSWAP_NONE) { + swap_tb =3D __swap_table_get(&ci_dyn->ci, voff + i); + if (__swap_table_test_zero(&ci_dyn->ci, voff + i)) + slot_type =3D VSWAP_ZERO; + else if (swp_tb_is_folio(swap_tb)) + slot_type =3D VSWAP_FOLIO; + else + slot_type =3D VSWAP_NONE; + } else { + slot_type =3D vtable_type(vt); + } + + if (!i) { + first_type =3D slot_type; + if (first_type =3D=3D VSWAP_SWAPFILE) + first_phys =3D vtable_to_phys(vt); + } else if (slot_type !=3D first_type) { + break; + } else if (first_type =3D=3D VSWAP_SWAPFILE && + vtable_to_phys(vt).val !=3D first_phys.val + i) { + break; + } + } + + if (typep) + *typep =3D first_type; + return i; +} + +static inline int vswap_check_backing(swp_entry_t entry, int nr, + enum vswap_backing_type *typep) +{ + struct swap_cluster_info_dynamic *ci_dyn; + unsigned int voff; + int ret; + + ci_dyn =3D vswap_lock_cluster(entry, &voff); + if (!ci_dyn) { + if (typep) + *typep =3D VSWAP_NONE; + return 0; + } + ret =3D __vswap_check_backing(ci_dyn, voff, nr, typep); + spin_unlock(&ci_dyn->ci.lock); + return ret; +} + +/** + * folio_phys_swap_backed - test whether a folio is backed by a contiguous + * range of physical swap slots. + * @folio: a swap-cache resident folio + * + * Return: %true if @folio->swap is not a vswap entry, or if these vswap + * entries are backed by a contiguous range of physical slots. + */ +static inline bool folio_phys_swap_backed(struct folio *folio) +{ + swp_entry_t entry =3D folio->swap; + int nr =3D folio_nr_pages(folio); + enum vswap_backing_type type; + + return !is_vswap_entry(entry) || + (vswap_check_backing(entry, nr, &type) =3D=3D nr && + type =3D=3D VSWAP_SWAPFILE); +} =20 static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn) { @@ -209,6 +353,16 @@ static inline bool is_vswap_entry(swp_entry_t entry) =20 static inline bool vswap_is_enabled(void) { return false; } =20 +static inline swp_entry_t vswap_to_phys(swp_entry_t entry) +{ + return (swp_entry_t){}; +} + +static inline bool folio_phys_swap_backed(struct folio *folio) +{ + return true; +} + static inline void __vswap_release_backing(struct swap_cluster_info *ci, unsigned int ci_start, unsigned int nr) {} @@ -222,6 +376,7 @@ static inline struct zswap_entry *vswap_zswap_load(swp_= entry_t entry) } =20 static inline void folio_release_vswap_backing(struct folio *folio) {} +static inline void folio_release_non_phys_swap_backing(struct folio *folio= ) {} =20 static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn) { @@ -232,4 +387,35 @@ static inline void vswap_cluster_free_vtable(struct sw= ap_cluster_info *ci) {} =20 #endif /* CONFIG_VSWAP */ =20 +/* + * Test a per-backend swap flag (SWP_SYNCHRONOUS_IO, SWP_STABLE_WRITES, ..= .) + * for @entry. For a vswap entry the property belongs to the current + * physical backing rather than vswap_si itself; resolve to the backing + * and test there. Returns false for zswap/zero/unbacked vswap entries + * as they don't have a backing bdev. + */ +static inline bool swap_entry_backend_has_flag(struct swap_info_struct *si, + swp_entry_t entry, + unsigned long flag) +{ + struct swap_info_struct *phys_si; + swp_entry_t phys; + bool has_flag; + + if (!swap_is_vswap(si)) + return data_race(si->flags & flag); + + phys =3D vswap_to_phys(entry); + if (!phys.val) + return false; + + phys_si =3D get_swap_device(phys); + if (!phys_si) + return false; + + has_flag =3D data_race(phys_si->flags & flag); + put_swap_device(phys_si); + return has_flag; +} + #endif /* _MM_VSWAP_H */ diff --git a/mm/zswap.c b/mm/zswap.c index 789079c3945b..d0c6ce2aa092 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1579,7 +1579,7 @@ bool zswap_store(struct folio *folio) */ if (is_vswap_entry(swp)) { if (index > 0) - folio_release_vswap_backing(folio); + folio_release_non_phys_swap_backing(folio); } else { unsigned type =3D swp_type(swp); pgoff_t offset =3D swp_offset(swp); --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-ot1-f43.google.com (mail-ot1-f43.google.com [209.85.210.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B43F03B388B for ; Thu, 6 Aug 2026 18:43:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041792; cv=none; b=luSZcy82z+Zl1iZ/oUWERCiIKgMAEQ9eX0/zhoYmgKDzm8URhg+TXjkpDJFy/FN9He6unU+LBw9ScRE9k1W9E4axyqCjJYicomg+dSJPVhmkD16Zk6Blg0F4DuVVH3sqSgzImNlWJt0lyd8F43CUUkLo0yvciJNsGT8SH+Wlhds= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041792; c=relaxed/simple; bh=k1/KSPa5L2DOT9UD7QQB8WB9l25sWbdPKEqNWKJbyWk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Tas9jcPZGg3KLzHqyOMBAdSIR7hZrra2Sf2wJ2Da+S10kDJhKvNp90B8LDDcNp8CPpD3aS3i8vMS0V9quU5xUonRlLfLcrsPCvQPQU3R+LWMJ5pPgVC22F4zjSbEGgxpj0zHDNNGEtOIvXYOVL6RMb/J/SSMfFZqdtwvh81y1ec= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=DJrAUnj/; arc=none smtp.client-ip=209.85.210.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="DJrAUnj/" Received: by mail-ot1-f43.google.com with SMTP id 46e09a7af769-7eb63dbd229so958315a34.1 for ; Thu, 06 Aug 2026 11:43:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041788; x=1786646588; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JkErdUV7HA1hqz0H4I2ABEL6H9gZP+ZkKpUGTLLTXSY=; b=DJrAUnj/RSFNrEFyG+jPHuwV9T2LlnbHdhs27v8bfpEZgj6aU5vxmHQkMGMY22mNfu UOhDkqe1Se7Ll7UvEKzHNCMhqQG38ta7IFbT1s4Y4/gZQTMerfhW0pZi1pCKTh+eLeDL DAhHEB+EYejIFd/BwDt6ByxAkUdTD30zBKFRQfp/ewG4fMcYIxaUcoFtxVB7seJtEOFe 8916TftEGV1K8xmcU9wt223IOScQ7mXGFcAhgL535PWdar6a//Peoj8uf0qmtsB0Mkjk GgcpVzYqSpSVwGNWroSgIaAlbYlQivHCWY+8VYRQyAiZleoa+xbUkmHBoRPepJC/1Q6y +yzA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041788; x=1786646588; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=JkErdUV7HA1hqz0H4I2ABEL6H9gZP+ZkKpUGTLLTXSY=; b=Qo+k+vrLKa62GAlQQrqRuVS0kusdZ5AiDDdMnIzwcYJbR29I2xZG1SCaayg8TQWGUQ bWugqiEmyW+trgtVPz0URaRXJa3VERWKgXXs0ZL+PaP5V+Ch1K84gu5lo3Cz2QxbGz8A 3H1HbWU6UsUOXy+RYtsZqwgWh54xl3uP1rHVn/3J2Rew+a31+eZjaZVGYgObuiBWvJct 4baV9ezKU84KdmrA3FUWeZq1pamSzdDs7NtfULpuROGkwCGulSZgL4dklYC1ZnCR1oPz R91D0Gp8D3/Ia3OhiW1W1x8EPjFF3imZOfKfjWiooYE3v9w5qqtcWQ6wq2CYXWYh/oZZ kYgw== X-Forwarded-Encrypted: i=1; AHgh+RqYlo6c/lkUBuwJ2EhuvwZ4qCeDX/MF8RSmGIkoFrKPjQ5kgx7vykvqbl8jgus8UJXbpZ2nYfkIBuqu9qo=@vger.kernel.org X-Gm-Message-State: AOJu0YyItRWfLpkHT4AvsCHV0StqlwnEvR4JoFC3xm5subgbCUkE9VDH VdF518LR5gciE0JG70iWlTjnQwxTrXI/1bPKIcxPVwBGz6UzDnU8NOcE X-Gm-Gg: AR+sD12/7/jPUMP6N9vj/XxVKFa/0N6fYKesgb1Ka4hdH5N2o6u/m7dJN6JuqLpDSBV +KPamwmdsGLluOkcisN6eNpyPDFnfyPvFmDzyi6AkM8wrAPrB581fG0LZ2FL50xI2FR76ibyMGN UjR4YyatX7U9r7F9+xibcnbipKM44kdIelrg6zyxcRIe200T3nXLOzbbVa4X2+Myf85eSw3GWTA sT4D5kSIulWJ6kCkGZhCHr/fqp0X+QB9dB2YEK5ILjgifEE5+Ynqu2YmrMoCIEIQA22+9sZ8kuz TxRd4i+tcf/oRl0jwroCckDWoXvg1AbLw2Lak6OeOtP+oM3uNNodl6o1PQcURbMOOazFsmpVAlr H/95ze/c1ge+0Wq7iW/En0owL30is2o6PkrY1NYf17DOTnyRutWc/Yq3tDWDCK58AQsj1fzqNFQ MMjVOVANQeedfuv8IIPK8Nx3QSs/YlQB7VWyphkZB0nOhvIxsJyfda9JXFMXoL0pcXz58ln+f5X OHg2AToRg== X-Received: by 2002:a05:6830:6d19:b0:7e9:dabf:fba0 with SMTP id 46e09a7af769-7f3444f2529mr1882168a34.15.1786041783901; Thu, 06 Aug 2026 11:43:03 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:b::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f1df34556asm5010219a34.9.2026.08.06.11.43.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:03 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 05/11] mm, swap: enable THP swapin for vswap entries Date: Thu, 6 Aug 2026 11:42:48 -0700 Message-ID: <20260806184254.3790858-6-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Swap a large folio back in as a unit when its vswap entries share a THP-amenable backing (a contiguous physical run, or all zero-filled), instead of always falling back to order-0 faults. A zswap-backed or mixed-backing batch is still refused, and the fault retries at a smaller order. Signed-off-by: Nhat Pham --- mm/memory.c | 12 ++++++++---- mm/swap_state.c | 17 +++++++++++++---- mm/vswap.h | 7 +++++++ mm/zswap.c | 18 ++++++++++++------ 4 files changed, 40 insertions(+), 14 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index ba84565605a1..a1e106a5c5e6 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -4815,11 +4815,15 @@ static unsigned long thp_swapin_suitable_orders(str= uct vm_fault *vmf) entry =3D softleaf_from_pte(vmf->orig_pte); =20 /* - * THP swapin for vswap is not supported yet. Also, a large swapped - * out folio could be partially or fully in zswap, which we lack - * handling for. In both cases, fall back to order-0 swapin. + * A large swapped out folio could be partially or fully in zswap. + * For vswap entries the THP-amenability of the backing is checked + * later under the cluster lock in __swap_cache_add_check, which + * rejects ZSWAP and mixed batches via -EBUSY and triggers + * order-fallback. For non-vswap entries we still need the + * zswap_never_enabled() bail: zswap_load rejects large folios with + * -EINVAL, which would SIGBUS the fault. */ - if (is_vswap_entry(entry) || !zswap_never_enabled()) + if (!is_vswap_entry(entry) && !zswap_never_enabled()) return 0; =20 /* diff --git a/mm/swap_state.c b/mm/swap_state.c index c61bb3eef62a..479814d19f50 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -173,6 +173,9 @@ static int __swap_cache_add_check(struct swap_cluster_i= nfo *ci, unsigned int ci_off, ci_end; unsigned long old_tb; bool is_zero; + struct swap_cluster_info_dynamic *ci_dyn; + enum vswap_backing_type type; + int ret; =20 lockdep_assert_held(&ci->lock); =20 @@ -201,11 +204,17 @@ static int __swap_cache_add_check(struct swap_cluster= _info *ci, return 0; =20 /* - * THP swapin for vswap is not supported yet; reject the batch so - * swap_cache_alloc_folio falls back to order 0. + * For a vswap entry batch, reject if the backing is not THP-amenable + * (e.g. uniformly ZSWAP, or mixed). The order-fallback loop in + * swap_cache_alloc_folio will retry with a smaller order on -EBUSY. */ - if (is_vswap_entry(targ_entry)) - return -EBUSY; + if (is_vswap_entry(targ_entry)) { + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + ret =3D __vswap_check_backing(ci_dyn, round_down(ci_off, nr), + nr, &type); + if (ret !=3D nr || type =3D=3D VSWAP_ZSWAP) + return -EBUSY; + } =20 is_zero =3D __swap_table_test_zero(ci, ci_off); ci_off =3D round_down(ci_off, nr); diff --git a/mm/vswap.h b/mm/vswap.h index 239b47b577d5..a921620f08be 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -378,6 +378,13 @@ static inline struct zswap_entry *vswap_zswap_load(swp= _entry_t entry) static inline void folio_release_vswap_backing(struct folio *folio) {} static inline void folio_release_non_phys_swap_backing(struct folio *folio= ) {} =20 +static inline int __vswap_check_backing(struct swap_cluster_info_dynamic *= ci_dyn, + unsigned int voff, int nr, + enum vswap_backing_type *typep) +{ + return 0; +} + static inline int vswap_cluster_alloc_vtable(struct swap_cluster_info_dyna= mic *ci_dyn) { return 0; diff --git a/mm/zswap.c b/mm/zswap.c index d0c6ce2aa092..5dc338188a29 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1630,13 +1630,19 @@ int zswap_load(struct folio *folio) return -ENOENT; =20 /* - * Large folios should not be swapped in while zswap is being used, as - * they are not properly handled. Zswap does not properly load large - * folios, and a large folio may only be partially in zswap. + * zswap_load() does not support large folios. For non-vswap + * entries this is unexpected on the swapin path: WARN and + * sigbus. For vswap entries __swap_cache_add_check() has already + * filtered out ZSWAP-backed THPs under the cluster lock, so the + * large folio here is zero- or phys-backed; return -ENOENT so the + * phys/zero IO path handles it. */ - if (WARN_ON_ONCE(folio_test_large(folio))) { - folio_unlock(folio); - return -EINVAL; + if (folio_test_large(folio)) { + if (WARN_ON_ONCE(!swap_is_vswap(si))) { + folio_unlock(folio); + return -EINVAL; + } + return -ENOENT; } =20 entry =3D zswap_entry_load(swp); --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-ot1-f45.google.com (mail-ot1-f45.google.com [209.85.210.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEE5F3093DB for ; Thu, 6 Aug 2026 18:43:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041789; cv=none; b=HextDBso7WY66NEeVm0GiBBA2sp+N/9lh40kjUq9X9m+HcEigpapmMA9jLCyEJ9JA7OHsWmuf5QCHjR8KKPv8LhC/q9nQtnJ1C3ekAj83cGHaugGEz2ne5lSaevzRc0YSY5X6MyaSsYvHtDyo8dDjS3ZQ1UM6wx6jqF6TrNFKao= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041789; c=relaxed/simple; bh=SsQaumNHv0ocAFn7305gZZiU249imNn6N7fAB+LJizE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=i5znYW1/thaDzBoI9dpCJM4AmdK+pRZIWkGH6d/SOx+vkneqD7vqvhoAlzIKljm3snMYzkqW9hwZSzYmaZ/fN9UB4vmUBRXILP6RVc2abB+m+Ew3AWwMBt6NWOQSlUk+BFEJrx54xe0F7lY/U7CN1hhA0PmoDQe0cmDaw0aD13I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=L4kAkju9; arc=none smtp.client-ip=209.85.210.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="L4kAkju9" Received: by mail-ot1-f45.google.com with SMTP id 46e09a7af769-7ec49608332so1342496a34.3 for ; Thu, 06 Aug 2026 11:43:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041785; x=1786646585; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/z0Ae7xOJ//F3kxqKAMic+EYDc2ZHsdvQ58o+7Z+X/w=; b=L4kAkju98PBwATYR9Y3B1XqbUIS9ld8hLep2g35y9eWYnPda+V/UaRphr1JlRI9ixs Si3xkaLXrmJHrN/h2YgTDbldF7+PGRbTWeduqQzAWbCC/446omlNWvmAH/8mxnKSORZw 5DZ0GKJY96kLPlBkif9UyfSmTSqfg45BHBBnrzHGr4U8Wnw446MeIA+xVCSKJebAJcS1 nQ2Na4fQf/7loBEAydPCMEKSQ9udJsOzCHGOQxSD3jNU8DbSnJgOEDBj1S11JOBlgQKJ 8mZ4Imdy5MoP8DMVwRqaq0BdoW6mfGirmstWNIJB/Qt2FBYq2zYLRyYxrjCNGTZeHwJc VPHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041785; x=1786646585; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=/z0Ae7xOJ//F3kxqKAMic+EYDc2ZHsdvQ58o+7Z+X/w=; b=d5GC0F5kgFC43sbKww5ZrIFDh8/HfVubZZU3CKr0eiNmiJjD7rkT4vleJxzcVIRjOH MwuGmGYxAiJrWgYdMIUO+HnXwLIrV7i3oPBrQKjOXzbCxcxKMsfxFIXnpC7yR1XwUy4B p0MLYhF/Y+yQYjltKLiCRtMSEHp2JVMBhKZWkAa9Xt7IPS+YvhKCzPwbiwsQB0Gt6QMM rKWp/1V9i9ZSjYjNLycns1liHzBuks4Yxq9zrozx2K6AFr5xzTfbMi19/HxQ88H0rLxH /pdurLchbJ8Jqzfx5g49hkd6hJNrLd9LfevHr89OdQLe9n/hALSYCCJniYwpv4lCyZxT wK7w== X-Forwarded-Encrypted: i=1; AHgh+RoavMyECwxiw0cZCNX8sG62gIv1PidI243vpZZgPsVAOxNbLzfPPXWBKm6th1uBFL9hHtpPeHvVcG27TfU=@vger.kernel.org X-Gm-Message-State: AOJu0YzaJnCnlff7c/RNk0Y5PHDfh8AabMpt1FhooYmDglCAZhTfZOOv exPDnWwIqBoFfYsDe1iVxCho0DsSuLUtuc4SY1Iw4qr9LGALGWKiwASz X-Gm-Gg: AR+sD10l5M7iwjxxaS4kceru5/9Pmg++5xUeQDurQK8Y226kc00f+MgcUKw55nXPqOL Gn4IG1ANfbZ0aAmcH+qJC2Ms4CLnkB540+ZuPSOwso373uBKbwXUhO2FTDRjwNjOekw2Yrdh6Pp UgYoP/L+408xcwmznp5O1dz6XFgQaO2scF7/7iPcfNHGgpkdWeZbxA5iO02tr+d7LK+hU+P7FZv lx6VITZe0HiWBaz1kvAa3R/t8rq1rzwqbR7m/Kpw8kKleE2URN1woiSP6y3PN2rI6exHADUsMmZ 0qMH44FViyNtT26FPf7H9NH/mP2kLI9Ph0gsMT1MAMgBKi4QKEUqYlwVq7bV8x5c/Z9CE8Ki2OX ufLOxevHjhEjc5DHTKz1/eb6pmb3ZwOYMWo5Hjd3efN1eCr1GbNjUOIXuv/alltDXx7tdahVVGT xZpfv440XCBQyYVH25bCccN2bpk1hm9YsF+jxMxgmffx8zCPhO4QV3Nluh0lpsKbA/l38PWFubm JtgJx7QE98= X-Received: by 2002:a05:6830:828e:b0:7dc:c4ae:a689 with SMTP id 46e09a7af769-7f1e5c13d47mr11440177a34.2.1786041785156; Thu, 06 Aug 2026 11:43:05 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:59::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f1df5deeeasm5000722a34.26.2026.08.06.11.43.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:04 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 06/11] mm, swap: write back vswap zswap entries to physical swap Date: Thu, 6 Aug 2026 11:42:49 -0700 Message-ID: <20260806184254.3790858-7-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add support for writing back zswap-backed vswap entries to physical swap. The mechanism mirrors the existing zswap writeback path, except the backing physical slot is allocated on demand at writeback time rather than already being pinned by the PTE. Now that vswap entries can be written back, relax the zswap shrinker gate: replace the blanket "skip while vswap is enabled" check with can_zswap_writeback(), which only skips when no physical slot is free to allocate on demand. This keeps the shrinker off futile writeback when no physical swap is available while letting it drain vswap entries otherwise. Signed-off-by: Nhat Pham --- mm/zswap.c | 76 +++++++++++++++++++++++++++++++++++++++--------------- 1 file changed, 55 insertions(+), 21 deletions(-) diff --git a/mm/zswap.c b/mm/zswap.c index 5dc338188a29..128309d5063b 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1007,12 +1007,12 @@ static bool zswap_decompress(struct zswap_entry *en= try, struct folio *folio) static int zswap_writeback_entry(struct zswap_entry *entry, swp_entry_t swpentry) { - struct xarray *tree; pgoff_t offset =3D swp_offset(swpentry); struct folio *folio; struct mempolicy *mpol; struct swap_info_struct *si; struct swap_io_ctx ctx =3D {}; + swp_entry_t phys =3D {}; int ret =3D 0; =20 /* try to allocate swap cache folio */ @@ -1020,12 +1020,6 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, if (!si) return -EEXIST; =20 - /* Vswap entries have no physical backing to write to. */ - if (swap_is_vswap(si)) { - put_swap_device(si); - return -EINVAL; - } - mpol =3D get_task_policy(current); folio =3D swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol, NO_INTERLEAVE_INDEX); @@ -1044,41 +1038,71 @@ static int zswap_writeback_entry(struct zswap_entry= *entry, /* * folio is locked, and the swapcache is now secured against * concurrent swapping to and from the slot, and concurrent - * swapoff so we can safely dereference the zswap tree here. - * Verify that the swap entry hasn't been invalidated and recycled - * behind our backs, to avoid overwriting a new swap folio with - * old compressed data. Only when this is successful can the entry - * be dereferenced. + * swapoff so we can safely dereference the zswap tree (or vswap + * vtable) here. Verify that the swap entry hasn't been + * invalidated and recycled behind our backs, to avoid overwriting + * a new swap folio with old compressed data. Only when this is + * successful can the entry be dereferenced. */ - tree =3D swap_zswap_tree(swpentry); - if (entry !=3D xa_load(tree, offset)) { + if (entry !=3D zswap_entry_load(swpentry)) { ret =3D -ENOMEM; goto out; } =20 + if (swap_is_vswap(si)) { + /* + * Allocate physical backing before decompress so a failure + * wastes no work. folio_realloc_swap retags the vtable to + * PHYS, leaving the entry pointer held only by the caller. + */ + phys =3D folio_realloc_swap(folio); + if (!phys.val) { + ret =3D -ENOMEM; + goto out; + } + } + if (!zswap_decompress(entry, folio)) { ret =3D -EIO; + /* + * For vswap: folio_realloc_swap already moved the entry + * out of the vtable. Restore it via vswap_zswap_store so + * the entry stays tracked (and the just-allocated PHYS + * slot is freed). For non-vswap: entry is still in the + * zswap tree. + */ + if (swap_is_vswap(si) && phys.val) + vswap_zswap_store(swpentry, entry); goto out; } =20 - xa_erase(tree, offset); + if (!swap_is_vswap(si)) + xa_erase(swap_zswap_tree(swpentry), offset); =20 count_vm_event(ZSWPWB); if (entry->objcg) count_objcg_events(entry->objcg, ZSWPWB, 1); =20 - zswap_entry_free(entry); - /* folio is up to date */ folio_mark_uptodate(folio); =20 /* move it to the tail of the inactive list after end_writeback */ folio_set_reclaim(folio); =20 - /* start writeback */ - __swap_writepage(&ctx, folio, folio->swap); + /* + * Start writeback. The entry has been moved out of its prior location + * (vtable PHYS for vswap, removed from the tree otherwise), so we own + * the free. vswap writes to the on-demand physical slot; others write + * to the folio's own entry. + */ + if (swap_is_vswap(si)) + __swap_writepage(&ctx, folio, phys); + else + __swap_writepage(&ctx, folio, folio->swap); swap_write_submit(&ctx); =20 + zswap_entry_free(entry); + out: if (ret) { swap_cache_del_folio(folio); @@ -1091,6 +1115,16 @@ static int zswap_writeback_entry(struct zswap_entry = *entry, /********************************* * shrinker functions **********************************/ +/* + * vswap zswap entries need a physical slot allocated on demand (via + * folio_realloc_swap) for writeback; if none is free, writeback fails, so + * skip the shrinker to avoid spinning on entries we cannot drain. + */ +static bool can_zswap_writeback(void) +{ + return !vswap_is_enabled() || get_nr_swap_pages(); +} + /* * The dynamic shrinker is modulated by the following factors: * @@ -1228,7 +1262,7 @@ static unsigned long zswap_shrinker_count(struct shri= nker *shrinker, if (!zswap_shrinker_enabled || !mem_cgroup_zswap_writeback_enabled(memcg)) return 0; =20 - if (vswap_is_enabled()) + if (!can_zswap_writeback()) return 0; =20 /* @@ -1313,7 +1347,7 @@ static int shrink_memcg(struct mem_cgroup *memcg) if (!mem_cgroup_zswap_writeback_enabled(memcg)) return -ENOENT; =20 - if (vswap_is_enabled()) + if (!can_zswap_writeback()) return -ENOENT; =20 /* --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oo1-f54.google.com (mail-oo1-f54.google.com [209.85.161.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DE97D3AD526 for ; Thu, 6 Aug 2026 18:43:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041790; cv=none; b=QsT1UzY7xiv3LczreAh10l7n00cLvEY+Arf9cFkTGPA1JxUrp5N9Stv4BxJtr+KbEQoh5soSEObEv51+sfF+6bD6z55cK5QIeKvzABATJwqWbkF0MNpDN1NoTLA/CiU1mbgWWLdzb1us4sCPj1VRkH3TL8yARo247+Qo8A+Yrl8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041790; c=relaxed/simple; bh=nwodC1vLGvhsbb6BdnopoPES0nNfkVSIPRTTmoidO1A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jIWzpExWm3GLWti85hDFatkXJLW+CgU35Msr0hXdhs9DLSjAFEM/AtzlEYxbH2fy9LG0+5KurtfS8ijB8PK+APbo4FEi1E9+hJLHwGabdVCHXwbm0931q49P0BBTU+0AhDJvq9CTbZ6gfXcrfo+H6JcfaCDxcJEXQV2wMYebJO8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=nTtBhumL; arc=none smtp.client-ip=209.85.161.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="nTtBhumL" Received: by mail-oo1-f54.google.com with SMTP id 006d021491bc7-6acc2a10023so938583eaf.0 for ; Thu, 06 Aug 2026 11:43:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041786; x=1786646586; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=2jkqf0E4cj9fQT5qd3L4w2YuapP3ELsmfSVpofHj8EA=; b=nTtBhumLC00+UmchPylMz1nXnNQePLZBmfJ+i5rKmGt9w/hOxLxDByF9rFK2l9mDxc gsIg7IOZ64jcu6osqWnuAt65RgLoMvMd3C0jTeintObP78j8tDvb7kBdZMUR9P6LQPAe QHlU/8/4EjDsEFnvGmGFXyEcO1d2XOyOagfEcK6pSq6+L88/5JBuqswtLU1QF6gSXAdT CeZLF7YqrHJNWl/MT/KhzjF6dxOA6S8Hz/Yzcl8zuRV1pnh/KPCY1cW3csI7m4SVS9T4 nLTcqEcS2p2YVEsT5BuU+tInTvw3R4JOmBchEhg4WOR4YEqjElSsyTnrJ99zQ3ZbhvpB j9VQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041786; x=1786646586; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=2jkqf0E4cj9fQT5qd3L4w2YuapP3ELsmfSVpofHj8EA=; b=hqv8NrAjF3x0gIocEWf0x+xuWLXzH1XCCJZDCrlpj+tNNVcubsIngu734KxLleRiWV 4Dn+cY8MjZZXqQkn/DzjIALBp9Y+rIZtvlbz2B9Un6THlDQ0fTADL1Nhy7pYTIctNANA k3YJohQBmdGNcuQUIN5gCgAYLFKn/7GHN4LRqkHL+n+s8nCNwdtc+Jfz9JkxRBABj4Or 3S6bnO/cldkh2hhiuwPAEmnZiG4KXOKOR4SIimAw+s7eRqcglaOTF1tPkjdRwH4RoqZK lMb7nZEsij1czYbuTzcVc/7ONjQ/doh4Vz5+LisdREsq7t32dp6eXrtk+EQ3ft+Yy4Xo 7Edw== X-Forwarded-Encrypted: i=1; AHgh+RqlEgGzKQEbl3+2YsgPQz1FAeUo8Vz8PLVJrGCdeoeT/ju3Olhn8+HE17HXyZ7ZvdqRcMYc+QWvVqf4+W0=@vger.kernel.org X-Gm-Message-State: AOJu0YzvvZruW/rce35luRhHwpMeSqpuki0FPWdE+HFndQNwym9nurGI KNmErOlLPd2GcyGoUFb5PSPvqlpSharjDdWFwTQwBBAPuXeAB5P6i+IB X-Gm-Gg: AR+sD125uiOZuAB8Y7FlNN17Aae0TWZ6CyiVgesAkwQ7+yUkCxKB5zEoPy2aru9pO4Z c75EHI/XyvZs4xLPklk3AafEKQB2ABi4bdZ49YybUYcJK9xlIX8fQm7Fz5hFyCc8WkVFsMrYih/ sNhz8eLjV78IDcV/wRgokLEadl3FD9ceP+HzPgPTyvXyfF1vApxg4VhxXgF2xpSBE43+rpSraJ0 aybqjQ8yfAhCzuTdmknTOttWtqkSMcUPmAM1G0QMS+mhPtVvWYc7fyw6J9r9OfMZAp4qpAai4iL 6iMGQDti9Qo74OJ5PNXp0jssRNnT0uOrIaCIZowMmMKoxwhKuQ6EXwqn1cmuzYZAqygH3hFoorb RNaD2b/5iO95OC7GaNIyi2LJ8DIFecEgd4LEkScf9yGWUDo6KsbgLW8f2C4LiMZejZF5tjJdvqB WZ2K6o7R2QPhni1eS3Kg65AXmMR3Kf5cc0pilDNqGhhaD1M/4ZDplpxh3AxZFaPAlY0fqMRemQO KethN2NVyo= X-Received: by 2002:a05:6820:4c14:b0:6ac:a9c6:92c with SMTP id 006d021491bc7-6ae96c8868bmr8744138eaf.10.1786041786572; Thu, 06 Aug 2026 11:43:06 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:51::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bfa6379sm144245eaf.14.2026.08.06.11.43.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:06 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 07/11] mm, swap: reclaim physical slots backing cache-only vswap entries Date: Thu, 6 Aug 2026 11:42:50 -0700 Message-ID: <20260806184254.3790858-8-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A vswap entry backed by a physical slot can become cache-only: its swap_count drops to 0 while the folio is still in the swap cache, so the physical slot is redundant and reclaimable. Until now such a slot was only freed when the vswap entry itself was freed, pinning otherwise reclaimable physical capacity. Reclaim such slots from the physical reclaim scanner, once swap is more than half used (vm_swap_full()), to free physical capacity for new allocations. Signed-off-by: Nhat Pham --- mm/swap_table.h | 11 +++- mm/swapfile.c | 142 ++++++++++++++++++++++++++++++++++++++++++++++++ mm/vswap.h | 26 +++++++++ 3 files changed, 176 insertions(+), 3 deletions(-) diff --git a/mm/swap_table.h b/mm/swap_table.h index 5b0eca07a821..b50ebcd9e4de 100644 --- a/mm/swap_table.h +++ b/mm/swap_table.h @@ -377,9 +377,12 @@ static inline unsigned short __swap_cgroup_clear(struc= t swap_cluster_info *ci, * On physical clusters, a Pointer-tagged entry stores the offset of the * vswap entry that owns this physical slot (the reverse map). Only the * offset is stored; the swap type is implicit (always vswap_si->type, - * since there is exactly one vswap device). + * since there is exactly one vswap device). The top bit is reserved as + * a cache-only flag, set when vswap swap_count drops to 0 but the folio + * is still in swap cache. * - * Pointer: |---- vswap offset ----|100| + * Pointer: |C|---- vswap offset ----|100| + * C =3D SWP_RMAP_CACHE_ONLY (bit 63) */ #ifdef CONFIG_VSWAP extern struct swap_info_struct *vswap_si; @@ -387,7 +390,8 @@ extern struct swap_info_struct *vswap_si; #define SWP_TB_PTR_MARK_BITS 3 #define SWP_TB_PTR_MARK 0b100UL #define SWP_TB_PTR_MARK_MASK ((1UL << SWP_TB_PTR_MARK_BITS) - 1) -#define SWP_RMAP_ENTRY_MASK (~SWP_TB_PTR_MARK_MASK) +#define SWP_RMAP_CACHE_ONLY (1UL << (BITS_PER_LONG - 1)) +#define SWP_RMAP_ENTRY_MASK (~(SWP_RMAP_CACHE_ONLY | SWP_TB_PTR_MARK_MASK)) =20 static inline bool swp_tb_is_pointer(unsigned long swp_tb) { @@ -408,6 +412,7 @@ static inline swp_entry_t swp_tb_ptr_to_swp_entry(unsig= ned long swp_tb) return swp_entry(vswap_si->type, offset); } #else +#define SWP_RMAP_CACHE_ONLY 0UL static inline bool swp_tb_is_pointer(unsigned long swp_tb) { return false; diff --git a/mm/swapfile.c b/mm/swapfile.c index 0874f57d3124..ab4bb57707e6 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -151,8 +151,20 @@ static DEFINE_PER_CPU(struct percpu_vswap_cluster, per= cpu_vswap_cluster) =3D { }; =20 static bool vswap_alloc(struct folio *folio); +static void vswap_mark_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_off); +static void vswap_clear_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_start, int nr); #else static inline bool vswap_alloc(struct folio *folio) { return false; } +static inline void vswap_mark_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_off) {} +static inline void vswap_clear_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_start, int nr) {} #endif =20 /* May return NULL on invalid type, caller must check for NULL return */ @@ -912,6 +924,59 @@ static int swap_cluster_setup_bad_slot(struct swap_inf= o_struct *si, return ret; } =20 +/* + * Try to reclaim a Pointer-tagged physical slot backing a vswap entry. + * The physical cluster lock must NOT be held. Returns the number of physi= cal + * slots reclaimed (the backing folio's page count), or < 0 on failure. + */ +static int try_to_reclaim_vswap_backing(struct swap_info_struct *si, + unsigned long offset, + swp_entry_t vswap_entry) +{ + swp_entry_t phys_base; + struct folio *folio; + unsigned int i; + int ret; + + folio =3D swap_cache_get_folio(vswap_entry); + if (!folio) + return -1; + + if (!folio_trylock(folio)) { + folio_put(folio); + return -1; + } + + if (!folio_matches_swap_entry(folio, vswap_entry)) { + folio_unlock(folio); + folio_put(folio); + return -1; + } + + /* + * Re-validate under folio lock. The folio's first vswap entry is + * folio->swap; the rmap value we just read is folio->swap + i for + * some i in [0, nr_pages). Check the folio's first entry still maps + * to the contiguous physical run that includes our target offset. + */ + i =3D vswap_entry.val - folio->swap.val; + phys_base =3D vswap_to_phys(folio->swap); + if (!phys_base.val || swp_type(phys_base) !=3D si->type || + swp_offset(phys_base) + i !=3D offset || + i >=3D folio_nr_pages(folio)) { + folio_unlock(folio); + folio_put(folio); + return -1; + } + + ret =3D folio_nr_pages(folio); + if (!folio_free_swap(folio)) + ret =3D -1; + folio_unlock(folio); + folio_put(folio); + return ret; +} + /* * Reclaim drops the ci lock, so the cluster may become unusable (freed or * stolen by a lower order). @usable will be set to false if that happens. @@ -935,6 +1000,16 @@ static bool cluster_reclaim_range(struct swap_info_st= ruct *si, spin_unlock(&ci->lock); do { swp_tb =3D swap_table_get(ci, offset % SWAPFILE_CLUSTER); + if (swp_tb_is_pointer(swp_tb)) { + rcu_read_unlock(); + if (!(swp_tb & SWP_RMAP_CACHE_ONLY)) + goto relock; + if (try_to_reclaim_vswap_backing(si, offset, + swp_tb_ptr_to_swp_entry(swp_tb)) < 0) + goto relock; + rcu_read_lock(); + continue; + } if (swp_tb_get_count(swp_tb)) break; if (swp_tb_is_folio(swp_tb)) @@ -942,6 +1017,7 @@ static bool cluster_reclaim_range(struct swap_info_str= uct *si, break; } while (++offset < end); rcu_read_unlock(); +relock: =20 /* Re-lookup: dynamic cluster may have been freed while lock was dropped = */ ci =3D swap_cluster_lock(si, start); @@ -1209,6 +1285,7 @@ static void swap_reclaim_full_clusters(struct swap_in= fo_struct *si, bool force) long to_scan =3D 1; unsigned long offset, end; struct swap_cluster_info *ci; + swp_entry_t vswap_entry; unsigned long swp_tb; int nr_reclaim; =20 @@ -1233,6 +1310,19 @@ static void swap_reclaim_full_clusters(struct swap_i= nfo_struct *si, bool force) offset +=3D abs(nr_reclaim); continue; } + } else if (swp_tb_is_pointer(swp_tb) && + (swp_tb & SWP_RMAP_CACHE_ONLY)) { + vswap_entry =3D swp_tb_ptr_to_swp_entry(swp_tb); + spin_unlock(&ci->lock); + nr_reclaim =3D try_to_reclaim_vswap_backing(si, offset, + vswap_entry); + ci =3D swap_cluster_lock(si, offset); + if (!ci) + goto next; + if (nr_reclaim > 0) { + offset +=3D nr_reclaim; + continue; + } } offset++; } @@ -1812,6 +1902,8 @@ static void swap_put_entries_cluster(struct swap_info= _struct *si, } /* count will be 0 after put, slot can be reclaimed */ need_reclaim =3D true; + if (swap_is_vswap(si)) + vswap_mark_cache_only(si, ci, ci_off); } /* * A count !=3D 1 or cached slot can't be freed. Put its swap @@ -1918,6 +2010,7 @@ static int swap_dup_entries_cluster(struct swap_info_= struct *si, goto failed; } } while (++ci_off < ci_end); + vswap_clear_cache_only(si, ci, ci_start, nr); swap_cluster_unlock(ci); return 0; failed: @@ -2034,6 +2127,55 @@ int folio_alloc_swap(struct folio *folio) } =20 #ifdef CONFIG_VSWAP +static void vswap_mark_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_off) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *pci; + swp_entry_t phys; + unsigned long vt; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + vt =3D __vtable_get(ci_dyn, ci_off); + + if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { + phys =3D vtable_to_phys(vt); + pci =3D __swap_entry_to_cluster(phys); + swap_rmap_mark_cache_only(pci, swp_cluster_offset(phys)); + } +} + +/* + * Clear the cache-only rmap hint for entries re-referenced from count 0 t= o 1 + * (no longer reclaimable), so the physical reclaim scanner skips them. + */ +static void vswap_clear_cache_only(struct swap_info_struct *si, + struct swap_cluster_info *ci, + unsigned int ci_start, int nr) +{ + struct swap_cluster_info_dynamic *ci_dyn; + struct swap_cluster_info *pci; + unsigned long swp_tb, vt; + swp_entry_t phys; + unsigned int off; + + if (!swap_is_vswap(si)) + return; + + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + for (off =3D ci_start; off < ci_start + nr; off++) { + swp_tb =3D __swap_table_get(ci, off); + if (!swp_tb_is_folio(swp_tb) || swp_tb_get_count(swp_tb) !=3D 1) + continue; + vt =3D __vtable_get(ci_dyn, off); + if (vtable_type(vt) !=3D VSWAP_SWAPFILE) + continue; + phys =3D vtable_to_phys(vt); + pci =3D __swap_entry_to_cluster(phys); + swap_rmap_clear_cache_only(pci, swp_cluster_offset(phys)); + } +} =20 static void __swap_cluster_free_phys_backing(struct swap_info_struct *psi, struct swap_cluster_info *pci, diff --git a/mm/vswap.h b/mm/vswap.h index a921620f08be..803e9a3271fe 100644 --- a/mm/vswap.h +++ b/mm/vswap.h @@ -35,6 +35,32 @@ static inline bool is_vswap_entry(swp_entry_t entry) =20 bool vswap_is_enabled(void); =20 +/* + * Rmap cache-only helpers for physical cluster Pointer-tagged entries. + * SWP_RMAP_CACHE_ONLY records, inline on the physical swap_table entry, + * that the backing vswap entry has swap_count =3D=3D 0 (swap-cache-only, = so + * reclaimable). The physical reclaim scanner reads it directly instead of + * chasing the rmap into the vswap layer and paying the cluster-lookup + * indirection. + */ +static inline void swap_rmap_mark_cache_only(struct swap_cluster_info *ci, + unsigned int off) +{ + atomic_long_t *table; + + table =3D rcu_dereference_check(ci->table, true); + atomic_long_or(SWP_RMAP_CACHE_ONLY, &table[off]); +} + +static inline void swap_rmap_clear_cache_only(struct swap_cluster_info *ci, + unsigned int off) +{ + atomic_long_t *table; + + table =3D rcu_dereference_check(ci->table, true); + atomic_long_and(~SWP_RMAP_CACHE_ONLY, &table[off]); +} + /* * Virtual table entry encoding for vswap clusters. * --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oi1-f170.google.com (mail-oi1-f170.google.com [209.85.167.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E37723B42EE for ; Thu, 6 Aug 2026 18:43:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041794; cv=none; b=IwdHdWekvRdxpiNNLniS0Nqs4mOAM3SZwivcx1OgF6CwGbTWOFr8WUoe6Nd0eiMQdJ1rBDmKuEaztSF4FMTY7fAPbWyAPwzQL1T6RZOjKGZ7/ngU1J/PfgmWijd0FlaaVG+HLt+HD/mJ7T6llX+qPsxFjV0/5nhxC9Z50Dgwk8I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041794; c=relaxed/simple; bh=ma1FTOVudEGvtBVXKUv0bZv+jA1FRxkPDymfsUFFypA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=t2nhgekraLe5N5ZJH9y3DP8IlF3gPt5GzU55oLbnSOussZ2gx1Bxqvv8Qb9ojE3tkDiY8w+zPbwfdYZvsFh3MXCeiMr3oiIVyYq3AyRFaTl8uhY0x1g1D9etDZ6C16RaJmEKBrLVgqnpjGeTJB2VFxBNuqeeKCPqBc+EdQfmG9c= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pZibNla9; arc=none smtp.client-ip=209.85.167.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pZibNla9" Received: by mail-oi1-f170.google.com with SMTP id 5614622812f47-4a46a53abc9so1470298b6e.3 for ; Thu, 06 Aug 2026 11:43:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041790; x=1786646590; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nVsuALmW4jG8uLihjQKBr5T9l1kxnVLHlQECAc8aSrQ=; b=pZibNla9RSj0XlYHSSuUSrD2Pfe2ZHWlZpqRsEpu2MML+Z/ihVwE2UKCS0koQthNBq 9lpE2Qxh+vii0JyeCV4xr08TXfEcDLXPEjD0C5SnYQhha4qP1pdLwifmh5GMj8qjz00N MHSA+bCqi785ud1/Cic9S1OjjTqEkCwdBRwznjJMhkY6KjCpIPQVQFb9PVDo3Feyuphl 8byOkkWQP3HwWr1jtH74veKUF0UltQNok2EtBhpsXwK3yRWlk/xKyGbCAIrdNMsAZPCg jbVihozwSOIcGQ2xoZcD+D14sinyXYSacSRBFIeLUrk5upvzM+AixC2L1xkW7KqMnIgo Jjuw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041790; x=1786646590; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=nVsuALmW4jG8uLihjQKBr5T9l1kxnVLHlQECAc8aSrQ=; b=j/lMzrqSLahOaZXJrMyCFfE2qC1YlQuFvCof6nWQdH/aa6XTRw4LNijyZbeyTb1fpP mrALrfiKLm8sac6YrRTIF0DLHcsdZHuVW4p6emubn3Ylp39oOX69fwsc8kdyifb9hbOw 4gIhrMUfVvIyzuLhQippnszM5rqcWv+28VaNW6inxeaI1geDFeBZwNS6g9Hbtwoyv2qI bTP2WuYl9NOgwLvxROA/T3k3MKPbBqxE65UxV+8lNn6bD/NMwJaZDfxcM+/QGwN6u8mj m1jCQ4Sh9Eiq1Nu0PJES6sRBJ0qdbQqMWYM0J06+YZRmmuGxljtKVHA+VHH7OaKavdBK YDTw== X-Forwarded-Encrypted: i=1; AHgh+Rr6RRXtFc8koyDMrT5sgB5OLUGrjCESfoBN/z+5AixaWfeKXFAl5w8hu5x8ho4hyt+4PO9r22cR4tnOPZQ=@vger.kernel.org X-Gm-Message-State: AOJu0YxcfYqkCCQ7YPd7/46Jw7P9L/c2hxcuJfcyyTH6chsNPAYnfyzd pKNxaA76NfQyLQtFpfp8naD51GcN3LU+oCNO4IBD2W2pHok1QUT/8AVn X-Gm-Gg: AR+sD13tbiW8cHZpPHBkXLOb5hdXgDjWakj/ZZqvV2djNm8lABFvI1ASNyR+tKRTqDt DaDnnxpOFtQ3nd8TME3uGrRGefWlQCUdF9ls2IFDN/qtC2OQL2U/Fav2f5MbjdBnuYn1MA+MbqK u2ccb5CBZHK4BNZJ0xXi/5Ex+0lQXbU7ai72tFn7LeEdYMOgQP/nkKaZAN573zyLz3FkM5oTRwX nHMYiVlczgd50/77JOAX2wP1rISV3mLWsSTsiIqITzg1iFX3X7E6MrY1KUmPBrX+6soPAtJ+rt3 /tb4mFmqro+YruV0bIMg7j2OcfVpgeyG1KQDH0KA0OYAOQy64A2yg0G3A1+k19EVLXRgsNjhkRM WfPWAV6sWVhNGwgv71YT+BHyd7rfUwyAFaA2ItHBwp9y/JPwn+oO0v+it/nDR9Zmks54618Qt2Z rtLeJNgerVnTA9i/nBU6f1Q7x4tHdJc45nvPwpvNJgTOTZO9VUyYZlYGOXpsckejOzEhkhkTT/U Y8YYSwjF4U= X-Received: by 2002:a05:6808:bc6:b0:496:9b3:486 with SMTP id 5614622812f47-4b13edea938mr2290585b6e.14.1786041789495; Thu, 06 Aug 2026 11:43:09 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:24::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4afae75ce1bsm4972147b6e.15.2026.08.06.11.43.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:08 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 08/11] mm, swap: only charge physical swap entries Date: Thu, 6 Aug 2026 11:42:51 -0700 Message-ID: <20260806184254.3790858-9-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Charge memcg->swap when a vswap entry acquires physical backing rather than when it is allocated, so memory.swap.current tracks on-disk swap usage. Zswap-backed and zero-filled pages occupy no swap space but were charged as though they did. memory.swap.current therefore no longer counts them, and a cgroup whose pages all land in zswap can now reclaim anon memory with memory.swap.max set to 0. Direct-mapped physical swap charging is unchanged. Signed-off-by: Nhat Pham --- include/linux/memcontrol.h | 5 ++ include/linux/swap.h | 57 +++++++++++++ mm/memcontrol.c | 166 ++++++++++++++++++++++++++++++++----- mm/swapfile.c | 111 +++++++++++++++++++++---- 4 files changed, 303 insertions(+), 36 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index e78bc98ab229..0a2f85ac7b6a 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -1900,6 +1900,7 @@ static inline bool memcg_is_dying(struct mem_cgroup *= memcg) =20 #if defined(CONFIG_MEMCG) && defined(CONFIG_ZSWAP) bool obj_cgroup_may_zswap(struct obj_cgroup *objcg); +bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush); void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size); void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size); bool mem_cgroup_zswap_writeback_enabled(struct mem_cgroup *memcg); @@ -1908,6 +1909,10 @@ static inline bool obj_cgroup_may_zswap(struct obj_c= group *objcg) { return true; } +static inline bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may= _flush) +{ + return true; +} static inline void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size) { diff --git a/include/linux/swap.h b/include/linux/swap.h index 2b2bd56afffa..19a703510675 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -523,6 +523,43 @@ static inline int mem_cgroup_try_charge_swap(struct fo= lio *folio) return __mem_cgroup_try_charge_swap(folio); } =20 +extern void __mem_cgroup_record_swap(struct folio *folio); +static inline void mem_cgroup_record_swap(struct folio *folio) +{ + if (mem_cgroup_disabled()) + return; + __mem_cgroup_record_swap(folio); +} + +extern int __mem_cgroup_charge_backing_phys_swap(struct mem_cgroup *memcg, + unsigned int nr_pages); +static inline int mem_cgroup_charge_backing_phys_swap(struct mem_cgroup *m= emcg, + unsigned int nr_pages) +{ + if (mem_cgroup_disabled()) + return 0; + return __mem_cgroup_charge_backing_phys_swap(memcg, nr_pages); +} + +extern void __mem_cgroup_uncharge_backing_phys_swap(struct mem_cgroup *mem= cg, + unsigned int nr_pages); +static inline void mem_cgroup_uncharge_backing_phys_swap(struct mem_cgroup= *memcg, + unsigned int nr_pages) +{ + if (mem_cgroup_disabled()) + return; + __mem_cgroup_uncharge_backing_phys_swap(memcg, nr_pages); +} + +extern void __mem_cgroup_id_put_swap(unsigned short id, unsigned int nr_pa= ges); +static inline void mem_cgroup_id_put_swap(unsigned short id, + unsigned int nr_pages) +{ + if (mem_cgroup_disabled()) + return; + __mem_cgroup_id_put_swap(id, nr_pages); +} + extern void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_= pages); static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned in= t nr_pages) { @@ -539,6 +576,26 @@ static inline int mem_cgroup_try_charge_swap(struct fo= lio *folio) return 0; } =20 +static inline void mem_cgroup_record_swap(struct folio *folio) +{ +} + +static inline int mem_cgroup_charge_backing_phys_swap(struct mem_cgroup *m= emcg, + unsigned int nr_pages) +{ + return 0; +} + +static inline void mem_cgroup_uncharge_backing_phys_swap(struct mem_cgroup= *memcg, + unsigned int nr_pages) +{ +} + +static inline void mem_cgroup_id_put_swap(unsigned short id, + unsigned int nr_pages) +{ +} + static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 7a426db06222..f6aee32ef542 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -48,6 +48,7 @@ #include #include #include +#include #include #include #include @@ -5701,6 +5702,116 @@ int __mem_cgroup_try_charge_swap(struct folio *foli= o) return 0; } =20 +/** + * __mem_cgroup_record_swap - record memcg for swap without charging + * @folio: folio being added to swap + * + * Pin the memcg private ID ref and record it in the swap cgroup table + * without charging memcg->swap; the charge is deferred to physical-backing + * allocation (vswap). + */ +void __mem_cgroup_record_swap(struct folio *folio) +{ + unsigned int nr_pages =3D folio_nr_pages(folio); + struct swap_cluster_info *ci; + struct mem_cgroup *memcg; + struct obj_cgroup *objcg; + + if (do_memsw_account()) + return; + + objcg =3D folio_objcg(folio); + VM_WARN_ON_ONCE_FOLIO(!objcg, folio); + if (!objcg) + return; + + rcu_read_lock(); + memcg =3D obj_cgroup_memcg(objcg); + if (!folio_test_swapcache(folio)) { + rcu_read_unlock(); + return; + } + + memcg =3D mem_cgroup_private_id_get_online(memcg, nr_pages); + rcu_read_unlock(); + + ci =3D swap_cluster_get_and_lock(folio); + __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages, + mem_cgroup_private_id(memcg)); + swap_cluster_unlock(ci); +} + +/** + * __mem_cgroup_charge_backing_phys_swap - charge memcg->swap + * @memcg: the mem_cgroup to charge (may be NULL) + * @nr_pages: number of physical swap pages to charge + * + * Charge the swap counter when a vswap entry gains physical backing. The + * private ID ref is already held (pinned by __mem_cgroup_record_swap() at + * vswap allocation), so this only moves the counter. + * + * Return: 0 on success, -ENOMEM on failure. + */ +int __mem_cgroup_charge_backing_phys_swap(struct mem_cgroup *memcg, + unsigned int nr_pages) +{ + struct page_counter *counter; + + if (do_memsw_account()) + return 0; + if (!memcg) + return 0; + + if (!mem_cgroup_is_root(memcg) && + !page_counter_try_charge(&memcg->swap, nr_pages, &counter)) { + memcg_memory_event(memcg, MEMCG_SWAP_MAX); + memcg_memory_event(memcg, MEMCG_SWAP_FAIL); + return -ENOMEM; + } + mod_memcg_state(memcg, MEMCG_SWAP, nr_pages); + return 0; +} + +/** + * __mem_cgroup_uncharge_backing_phys_swap - uncharge memcg->swap counter + * @memcg: the mem_cgroup to uncharge (may be NULL) + * @nr_pages: number of physical swap pages to uncharge + * + * Uncharge the swap counter on physical backing release for a vswap entry. + * The private ID ref is dropped separately via __mem_cgroup_id_put_swap()= when + * the vswap entry is freed. + */ +void __mem_cgroup_uncharge_backing_phys_swap(struct mem_cgroup *memcg, + unsigned int nr_pages) +{ + if (!memcg) + return; + + if (!mem_cgroup_is_root(memcg)) { + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, nr_pages); + else + page_counter_uncharge(&memcg->swap, nr_pages); + } + mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages); +} + +/** + * __mem_cgroup_id_put_swap - drop memcg private ID ref without uncharging + * @id: cgroup private id + * @nr_pages: number of refs to drop + */ +void __mem_cgroup_id_put_swap(unsigned short id, unsigned int nr_pages) +{ + struct mem_cgroup *memcg; + + rcu_read_lock(); + memcg =3D mem_cgroup_from_private_id(id); + if (memcg) + mem_cgroup_private_id_put(memcg, nr_pages); + rcu_read_unlock(); +} + /** * __mem_cgroup_uncharge_swap - uncharge swap space * @id: cgroup id to uncharge @@ -5727,15 +5838,21 @@ void __mem_cgroup_uncharge_swap(unsigned short id, = unsigned int nr_pages) =20 long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg) { - long nr_swap_pages =3D get_nr_swap_pages(); + long nr_swap_pages; =20 /* - * vswap zswap-backed swapout needs no physical slot, so gate anon - * reclaim on the swap.max headroom instead of the physical free count. + * vswap charges only physical backing (folio_realloc_swap), not + * allocation. For a zswap-capable memcg virtual swap is unbounded, so + * the swap.max walk below would underestimate it and starve anon + * reclaim; report unbounded. swap.max is still enforced at + * phys-backing charge time. */ - if (vswap_is_enabled() && zswap_is_enabled()) - nr_swap_pages =3D PAGE_COUNTER_MAX; + if (vswap_is_enabled() && zswap_is_enabled() && + (mem_cgroup_disabled() || do_memsw_account() || + mem_cgroup_may_zswap(memcg, false))) + return PAGE_COUNTER_MAX; =20 + nr_swap_pages =3D get_nr_swap_pages(); if (mem_cgroup_disabled() || do_memsw_account()) return nr_swap_pages; for (; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) @@ -5907,8 +6024,10 @@ static struct cftype swap_files[] =3D { =20 #ifdef CONFIG_ZSWAP /** - * obj_cgroup_may_zswap - check if this cgroup can zswap - * @objcg: the object cgroup + * mem_cgroup_may_zswap - check if this cgroup hierarchy can zswap + * @original_memcg: the memcg to query + * @may_flush: force-flush stats for an accurate check (sleeps). Pass false + * from atomic contexts; the check is then best-effort. * * Check if the hierarchical zswap limit has been reached. * @@ -5918,15 +6037,13 @@ static struct cftype swap_files[] =3D { * spending cycles on compression when there is already no room left * or zswap is disabled altogether somewhere in the hierarchy. */ -bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) +bool mem_cgroup_may_zswap(struct mem_cgroup *original_memcg, bool may_flus= h) { - struct mem_cgroup *memcg, *original_memcg; - bool ret =3D true; + struct mem_cgroup *memcg; =20 if (!cgroup_subsys_on_dfl(memory_cgrp_subsys)) return true; =20 - original_memcg =3D get_mem_cgroup_from_objcg(objcg); for (memcg =3D original_memcg; !mem_cgroup_is_root(memcg); memcg =3D parent_mem_cgroup(memcg)) { unsigned long max =3D READ_ONCE(memcg->zswap_max); @@ -5934,20 +6051,27 @@ bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) =20 if (max =3D=3D PAGE_COUNTER_MAX) continue; - if (max =3D=3D 0) { - ret =3D false; - break; - } + if (max =3D=3D 0) + return false; =20 /* Force flush to get accurate stats for charging */ - __mem_cgroup_flush_stats(memcg, true); + if (may_flush) + __mem_cgroup_flush_stats(memcg, true); pages =3D memcg_page_state(memcg, MEMCG_ZSWAP_B) / PAGE_SIZE; - if (pages < max) - continue; - ret =3D false; - break; + if (pages >=3D max) + return false; } - mem_cgroup_put(original_memcg); + return true; +} + +bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) +{ + struct mem_cgroup *memcg; + bool ret; + + memcg =3D get_mem_cgroup_from_objcg(objcg); + ret =3D mem_cgroup_may_zswap(memcg, true); + mem_cgroup_put(memcg); return ret; } =20 diff --git a/mm/swapfile.c b/mm/swapfile.c index ab4bb57707e6..1a2d9d9625fc 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -47,6 +47,7 @@ =20 #include #include +#include "memcontrol-v1.h" #include "swap_table.h" #include "vswap.h" #include "internal.h" @@ -2116,8 +2117,16 @@ int folio_alloc_swap(struct folio *folio) goto again; } =20 - /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */ - if (unlikely(mem_cgroup_try_charge_swap(folio))) + /* + * A vswap entry has no physical swap yet, so only record the memcg; + * folio_realloc_swap() charges once backing is allocated. + * + * Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. + */ + if (folio_test_swapcache(folio) && + is_vswap_entry(folio->swap)) + mem_cgroup_record_swap(folio); + else if (unlikely(mem_cgroup_try_charge_swap(folio))) swap_cache_del_folio(folio); =20 if (unlikely(!folio_test_swapcache(folio))) @@ -2182,6 +2191,28 @@ static void __swap_cluster_free_phys_backing(struct = swap_info_struct *psi, unsigned int ci_start, unsigned int nr_pages); =20 +static void vswap_uncharge_cgroup_batch(unsigned short memcg_id, + unsigned int batch_nr, + unsigned int batch_nr_swapfile) +{ + struct mem_cgroup *memcg; + unsigned int n; + + /* + * v1 (memsw): __memcg1_swapout() charges memsw for every swapped-out + * entry regardless of backing, so uncharge all of them. v2: only + * swapfile-backed entries are charged, so uncharge just those. + */ + n =3D do_memsw_account() ? batch_nr : batch_nr_swapfile; + if (!n) + return; + + rcu_read_lock(); + memcg =3D memcg_id ? mem_cgroup_from_private_id(memcg_id) : NULL; + rcu_read_unlock(); + mem_cgroup_uncharge_backing_phys_swap(memcg, n); +} + /** * __vswap_release_backing - release the backing of a range of vtable slots * @ci: the locked vswap cluster @@ -2191,8 +2222,7 @@ static void __swap_cluster_free_phys_backing(struct s= wap_info_struct *psi, * Releases each slot in [@ci_start, @ci_start + @nr): physical swap slots, * zswap entries, etc. Clears the zero marks if set. * - * Context: caller must hold @ci->lock. The entire range must belong to the - * same memcg. + * Context: caller must hold @ci->lock. */ void __vswap_release_backing(struct swap_cluster_info *ci, unsigned int ci_start, unsigned int nr) @@ -2204,12 +2234,27 @@ void __vswap_release_backing(struct swap_cluster_in= fo *ci, unsigned int ci_off; unsigned long vt; swp_entry_t phys; + unsigned short batch_id; + unsigned int batch_nr =3D 0, batch_nr_swapfile =3D 0; =20 lockdep_assert_held(&ci->lock); ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + batch_id =3D __swap_cgroup_get(ci, ci_start); =20 for (ci_off =3D ci_start; ci_off < ci_start + nr; ci_off++) { + unsigned short cur_id; + vt =3D __vtable_get(ci_dyn, ci_off); + cur_id =3D __swap_cgroup_get(ci, ci_off); + + if (cur_id !=3D batch_id) { + vswap_uncharge_cgroup_batch(batch_id, batch_nr, + batch_nr_swapfile); + batch_id =3D cur_id; + batch_nr =3D 0; + batch_nr_swapfile =3D 0; + } + batch_nr++; =20 /* * Flush batched physical slots when the next entry @@ -2233,6 +2278,7 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, =20 switch (vtable_type(vt)) { case VSWAP_SWAPFILE: + batch_nr_swapfile++; if (phys_start =3D=3D phys_end) { phys =3D vtable_to_phys(vt); phys_start =3D swp_offset(phys); @@ -2266,6 +2312,8 @@ void __vswap_release_backing(struct swap_cluster_info= *ci, phys_start % SWAPFILE_CLUSTER, phys_end - phys_start); } + + vswap_uncharge_cgroup_batch(batch_id, batch_nr, batch_nr_swapfile); } =20 /** @@ -2355,7 +2403,10 @@ swp_entry_t folio_realloc_swap(struct folio *folio) swp_entry_t vswap_entry =3D folio->swap; struct swap_cluster_info *ci; struct swap_cluster_info_dynamic *ci_dyn; + struct mem_cgroup *memcg; unsigned int voff; + unsigned long vt; + unsigned short memcg_id; swp_entry_t phys_entry =3D {}; swp_entry_t pe; int i, nr =3D folio_nr_pages(folio); @@ -2364,9 +2415,18 @@ swp_entry_t folio_realloc_swap(struct folio *folio) VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); VM_WARN_ON(!is_vswap_entry(vswap_entry)); =20 - phys_entry =3D vswap_to_phys(vswap_entry); - if (phys_entry.val) - return phys_entry; + voff =3D swp_cluster_offset(vswap_entry); + ci =3D __swap_entry_to_cluster(vswap_entry); + ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); + + spin_lock(&ci->lock); + vt =3D __vtable_get(ci_dyn, voff); + if (vtable_type(vt) =3D=3D VSWAP_SWAPFILE) { + spin_unlock(&ci->lock); + return vtable_to_phys(vt); + } + memcg_id =3D __swap_cgroup_get(ci, voff); + spin_unlock(&ci->lock); =20 local_lock(&percpu_swap_cluster.lock); phys_entry =3D swap_alloc_fast(folio); @@ -2377,10 +2437,20 @@ swp_entry_t folio_realloc_swap(struct folio *folio) if (!phys_entry.val) return (swp_entry_t){}; =20 - voff =3D swp_cluster_offset(vswap_entry); + rcu_read_lock(); + memcg =3D folio_memcg(folio); + if (!memcg || mem_cgroup_private_id(memcg) !=3D memcg_id) + memcg =3D memcg_id ? mem_cgroup_from_private_id(memcg_id) : NULL; + rcu_read_unlock(); + + if (mem_cgroup_charge_backing_phys_swap(memcg, nr)) { + __swap_cluster_free_phys_backing( + __swap_entry_to_info(phys_entry), + __swap_entry_to_cluster(phys_entry), + swp_cluster_offset(phys_entry), nr); + return (swp_entry_t){}; + } =20 - ci =3D __swap_entry_to_cluster(vswap_entry); - ci_dyn =3D container_of(ci, struct swap_cluster_info_dynamic, ci); spin_lock(&ci->lock); /* * Install PHYS backing without freeing any prior contents of the @@ -2591,10 +2661,11 @@ void __swap_cluster_free_entries(struct swap_info_s= truct *si, unsigned short batch_id =3D 0, id_cur; unsigned int ci_off =3D ci_start, ci_end =3D ci_start + nr_pages; unsigned int batch_off =3D ci_off; + bool is_vswap =3D swap_is_vswap(si); =20 VM_WARN_ON(ci->count < nr_pages); =20 - if (swap_is_vswap(si)) + if (is_vswap) __vswap_release_backing(ci, ci_start, nr_pages); =20 ci->count -=3D nr_pages; @@ -2614,18 +2685,28 @@ void __swap_cluster_free_entries(struct swap_info_s= truct *si, /* * Uncharge swap slots by memcg in batches. Consecutive * slots with the same cgroup id are uncharged together. + * For vswap, only drop the ID ref - physical swap was + * already uncharged in __vswap_release_backing above. */ id_cur =3D __swap_cgroup_clear(ci, ci_off, 1); if (batch_id !=3D id_cur) { - if (batch_id) - mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + if (batch_id) { + if (is_vswap) + mem_cgroup_id_put_swap(batch_id, ci_off - batch_off); + else + mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + } batch_id =3D id_cur; batch_off =3D ci_off; } } while (++ci_off < ci_end); =20 - if (batch_id) - mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + if (batch_id) { + if (is_vswap) + mem_cgroup_id_put_swap(batch_id, ci_off - batch_off); + else + mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); + } =20 __swap_cluster_finish_free(si, ci, ci_start, nr_pages); } --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-ot1-f52.google.com (mail-ot1-f52.google.com [209.85.210.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 90FAA3BED7D for ; Thu, 6 Aug 2026 18:43:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041798; cv=none; b=dpkuK1+IfTiN1xWRnwaDrZulNaYnAg/3N870DOX0QGGxQYTgX9GDgzhvIvdqvH6o4u3c273f9EBgM4UO3P7wSnC7HbJDOIPVP8U2lRNg6MA2rAGek/CVNjcrryCVB5NiPf4j+FyOovq6XOONbLGe4306D+cl5SpcDc3wm4YVCZQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041798; c=relaxed/simple; bh=SNocFtXa8C4f89zudX1FgBzt9ndc5ukWPrs95OuM/mQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s4YsxUGZ2fF/d9S5nA5ikgcYBMVHi1gFfZCjsReIFUGwAqgCHa4spHZUfoTJ8TmfRjlzOT2JXb+A7N9EqzckDszWzi0UVUbGD8GGbHK5qlIif+EqG4TwCjD6toZYof8trY7lmRKGroakTHTNMtYDnlh3DMA927St4gdarcUY3Ow= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=MuH1RETk; arc=none smtp.client-ip=209.85.210.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="MuH1RETk" Received: by mail-ot1-f52.google.com with SMTP id 46e09a7af769-7e9ecd7216cso1520125a34.3 for ; Thu, 06 Aug 2026 11:43:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041795; x=1786646595; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XJDCmTvRD2vNfT2YGrgT7EosXDcYWyreOqvvzwLRIg0=; b=MuH1RETknsQ10GggJCh/GIl5KRUeTQoyvhiynmsTumjN5IJUffOqFSkKZnIpjYj0cQ u1RDvCB2bOgy3Z9tskxkNhVwnL1MSHwVdzHTw/K3ERQwCOSCrIv9AodwhbFauOmomqnB hMhBGx4SW3pxoWHr/LDgu+ZxFS68ZZZLU/rRYW3/XE98N0EQabPCzLMzMBX/pUsvEvzm 92lQEoyrxpZu/+STMBMWEY9tI7LHKKc3n6m9hf1Wdw7CmXdFqXRjjoHLkR2zpKIEVj82 PiqRFN0vqTZ81SVYMmtiYOeV5SGt66blbMjzQuWsKQ0R7Om4Dpggzs7IDRDgO3hmaGZ2 uy5w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041795; x=1786646595; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=XJDCmTvRD2vNfT2YGrgT7EosXDcYWyreOqvvzwLRIg0=; b=dQ245yO+sfwDFa8GfEk7UosVVN1Qc30ungVjPXlfBd63Zzq27ABOBvLf652uBUF1pp HoZCHcZo9U7ZeeNhR6Or2BfjQNUssaF8aSt1XYrPVFnk3EMy4N8CdJPLPYdsVRPRS/dE TEdHC/EfNlGE5oJUKxj562PKrquLtCIr3V68GDrObpXca/qJGhahzuaPNT0Hp2gkQvfI PX06xl0TqA/0ygeWReUQu0IIlSOEUg17z3d5yIdt5TbunPnEru/pCyW5/7EsLLkbg1xi /gNlSvTS2yx6lQBDCxDBY2bcooevSYu7IZJtWtWWfUTQ1zYc8a1LjVCki1qHxRFNfkhI KBfg== X-Forwarded-Encrypted: i=1; AHgh+Rr0IAcszDCv7H+kqWCP8Tp59cZwVbc1fxzuSwx3SAwADL519ntZbw6EpvACTNd7GTi9uTToU0ZPlqV/Es8=@vger.kernel.org X-Gm-Message-State: AOJu0YyJEOcG7b39aGN4xqnd7UH0DBW8jUFwtbexvibLQaoyuuf0y9mW rkmRIhiE3lh4BBAWV/NTnaQl1YfZ5yzYEI6yqNlcDnETsFx3ie1cBsYI X-Gm-Gg: AR+sD114p4lQN1C7G9LKtekBGlWnhh/29ZyTuDlDn/My4OkkZPrgTtHcT3c91/WVM6u yHqk2hVBG657BWw/hUUkZfoJHzPAwLjrHPPYOILUaEacwUiHBC8I1Qalo2GW4z63lNXp3GkRnXv ZPf25105pKaPN0ZilpJfnwUePl+Q7f/9y47PnotiRfrheANAS9RuHFENGkT1ov53tvIy/xPDfA4 FJEr65BAkYh1PhamqnQtiMFmPGnh9m6s7yyHANECi35+hRNA1duqWs7BqT9ofRF3U8w4PjdfEwZ 9iI9PLYHRtqnkalB0nqSD35ifUOKkOyv90M7QCfu+29rl8rX4dBw6FpL/VrqEwmZz4mv/2J+7s8 cqoRn1cCIY8agKEVhjYd3M5q1R0US4x7DAlaqJIv0+7ReoHZriNpKzS7bCAA1BPGfKDLuBZObuF 5sTBwz+FDrBFGXiI6LMkKMOTf5mFc6nKlxEhWS1uqWLEDo1lscqGqzBrio2gwySLAeACxyv1yAH SOFKqZbGLME/t5RFVRjKyQ= X-Received: by 2002:a05:6830:6404:b0:7e9:ead3:4449 with SMTP id 46e09a7af769-7f1e5d0f423mr10084761a34.6.1786041790716; Thu, 06 Aug 2026 11:43:10 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:52::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f1df5a4f9bsm5051028a34.23.2026.08.06.11.43.10 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:10 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 09/11] mm, swap: add debugfs counters for vswap Date: Thu, 6 Aug 2026 11:42:52 -0700 Message-ID: <20260806184254.3790858-10-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add /sys/kernel/debug/vswap/ with two counters: * used: virtual swap slots (pages) currently allocated * alloc_reject: cumulative pages that failed to get a vswap slot Signed-off-by: Nhat Pham --- mm/swapfile.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 1a2d9d9625fc..b4d7af21ca1c 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -7,6 +7,7 @@ */ =20 #include +#include #include #include #include @@ -133,6 +134,9 @@ static DEFINE_PER_CPU(struct percpu_swap_cluster, percp= u_swap_cluster) =3D { .lock =3D INIT_LOCAL_LOCK(), }; =20 +static atomic_t __maybe_unused vswap_used =3D ATOMIC_INIT(0); +static atomic_t __maybe_unused vswap_alloc_reject =3D ATOMIC_INIT(0); + #ifdef CONFIG_VSWAP static int sysctl_vswap_enabled =3D IS_ENABLED(CONFIG_VSWAP_DEFAULT_ON); =20 @@ -2056,11 +2060,13 @@ static bool vswap_alloc(struct folio *folio) if (folio_test_swapcache(folio)) { /* alloc_swap_scan_cluster updated percpu offset already */ local_unlock(&percpu_vswap_cluster.lock); + atomic_add(folio_nr_pages(folio), &vswap_used); return true; } =20 this_cpu_write(percpu_vswap_cluster.offset[order], SWAP_ENTRY_INVALID); local_unlock(&percpu_vswap_cluster.lock); + atomic_add(folio_nr_pages(folio), &vswap_alloc_reject); return false; } #endif @@ -2665,8 +2671,10 @@ void __swap_cluster_free_entries(struct swap_info_st= ruct *si, =20 VM_WARN_ON(ci->count < nr_pages); =20 - if (is_vswap) + if (is_vswap) { __vswap_release_backing(ci, ci_start, nr_pages); + atomic_sub(nr_pages, &vswap_used); + } =20 ci->count -=3D nr_pages; do { @@ -4862,6 +4870,7 @@ static const struct ctl_table vswap_sysctls[] =3D { static int __init vswap_init(void) { struct swap_info_struct *si; + struct dentry *root; unsigned long maxpages; int err; =20 @@ -4896,6 +4905,10 @@ static int __init vswap_init(void) =20 register_sysctl_init("vm", vswap_sysctls); =20 + root =3D debugfs_create_dir("vswap", NULL); + debugfs_create_atomic_t("used", 0444, root, &vswap_used); + debugfs_create_atomic_t("alloc_reject", 0444, root, &vswap_alloc_reject); + pr_info("vswap: created virtual swap device (%lu pages)\n", maxpages); return 0; =20 --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oo1-f44.google.com (mail-oo1-f44.google.com [209.85.161.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E1A83BA22E for ; Thu, 6 Aug 2026 18:43:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041795; cv=none; b=i9anppgCZU4vBvL7FwBSSvNltUQgJ+cxbfFCp2Pd2S8xDrFsROOdKSCsNrjQ9ZAxcZMcMH7yd0JAK1IyTM3HP9qpbkyJxYXAfZi0MsXOj5SW6s4G+ztO1N3SJoZ0Y6UzamcnqB+MTsiCWNKsoJIloBEGJfhFRqtY/4MR/vjiQLg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041795; c=relaxed/simple; bh=kUgIWKYFgahkHMUnPHN6PokGq9uBRZ5QpQiL27RcoVo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UMXhli9eZI5wmDhhMfXIzXfIprJihT7JwqIpEfNreiEV4ZbR09Tw2lOuNUEO3dtaT9y5CnWq37U9uSYClAGBSkePtKgnhQsmFI1EqrTYUVbGOkFRK43FE/qF48yxd+SevVlJhf8nlV3mbFWzMH+3m1BiMCx0fIz2+ARxWBYz8Ac= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=hywnn4mg; arc=none smtp.client-ip=209.85.161.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="hywnn4mg" Received: by mail-oo1-f44.google.com with SMTP id 006d021491bc7-6aae36ea5c4so1684496eaf.1 for ; Thu, 06 Aug 2026 11:43:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041792; x=1786646592; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6q7yoTOhS8xVdFoqcdm5Nyx0rOXnYxJmiHBoOvthnIM=; b=hywnn4mgf3I5l8Rr1u+ZDavWqsSeQbYLOkXATsuEurMB66Ye8aIKLNj5hBxrt2ZTQp 69Uvh7CGgZQQMgQOSJOQFxOUyXlWupnl6IhwDKgRblrneRWjgxhLMBNGf8h6FPWBQ0ro VGKZ7XQs6SuJit/gRscc/4nF/tIHr/jWL/nugBgOrTycGUN7fmvMJEACzn7+K8UFyxN4 N+B/bWUvqtdk8gCXRPmp3n84+FB1J3MDEhb0VuA7QvqLIbzTwMczcSyiX+EEXuUIbRjn UmRgS90YEikSE6Bf4Vnt4Vx7relOmMMqhk815P04ppX1CQLOo9oMh0EE3dvUy/gXGsbA zDUw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041792; x=1786646592; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=6q7yoTOhS8xVdFoqcdm5Nyx0rOXnYxJmiHBoOvthnIM=; b=EBptkrbQZnEwuabQ+ibSO015wn+hgsjCPROCQO+VSUKgLqHOqhflve+zClHBsTgG7e Ym3LlSOjcH1ds8aTGoogCbt2/IrsRh1g7xI71+Itqf5df09jx+YW7JCMwXrRKlgbcdIc fYqTsfCf2YeSfJzRvTAIo/eEGm/VyWUiO6Yg5JysFRrT6OahmIiW5s6Cir0VHw6FECbw jxJlFyRDoS5xa5hDktTf3tMFRhBdtMBdy6LSj73kp84I9ooW1IlV4+yJjFa4LxBHNqaA H6GPXhenDUiBSqztJzMwAHAjrl3JQeqUyCaGJiHBnlnX41h64eH96jCmWPIHtJ/yHot1 4NSg== X-Forwarded-Encrypted: i=1; AHgh+Rr3wAYjSR8evJSyMbh+9jCSkKSYBkErKlmK7C1DGEbPeMg+YnSLKI5mJYVTMgx4Luh62rqF+eRRfEqH9mQ=@vger.kernel.org X-Gm-Message-State: AOJu0YyO+MhuQjvHF+Gb4AO2h8VfvAZsOPvGmO8OQHTduVNxkr1YiHUe 35pLE5CVfvqpK16jasVNC3137hrhRjT3okCIIlCk4Et3GGqTh6qiHCHj X-Gm-Gg: AR+sD11hZ5QU2hmYxFbNRJi2XaQEOtkv8KcInD2Zn7bl8608xAMqPY6eAcV+g02iFZ8 ePKcJimOOmLkVNMZ8Rt1Tz2UZO4O3QCtIRZb3tWHQA7aL5U4c5K+6Ua3gJv3Hv7TwHDkEYPJMpt TsYfLEpRFx8mQC9RJwXMwYUkSYRAe1pvdd+IPgHOctCVJH0PMPQ+R5K1526lmF0Kr1dxq/Ahl5S 4uByhSLkjO+CgjXWUh9m30jbqgo+TGW/9j/Y8ISuRztq7IfekJGO7SmJTP8VWqniAPHIH521xut 3Fl1SFXjv38A0Ob3JoXgqLNIbrf1+5gnlJgbXdrlNjL5MJm+mNMQseOV4fj5kXjmMc1t87Hw8Xe lBCkEdOmVwPeSN3WTpS3SD+N6w6s/dKKxh/Anej0z3U8GMAB5+FdGlmuCPSGzp58KXAuqaUvdyD bJNYGEKwbwelPhV8RDm5R41gQ8DKdYwhO3pfHwd3EfPtzfR97TR0Ydsubl+ZimAvq14HzOKePV4 QLFXPVIqdw8XUbY5g== X-Received: by 2002:a05:6820:c8b:b0:6ac:8e23:3078 with SMTP id 006d021491bc7-6ae96c10021mr8748486eaf.5.1786041791947; Thu, 06 Aug 2026 11:43:11 -0700 (PDT) Received: from localhost ([2a03:2880:10ff::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bc2631esm203436eaf.4.2026.08.06.11.43.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:11 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 10/11] mm, swap: defer memcg_table allocation for physical swap clusters Date: Thu, 6 Aug 2026 11:42:53 -0700 Message-ID: <20260806184254.3790858-11-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Stop allocating a memcg table for every physical swap cluster that only ever holds vswap backings. The table costs SWAPFILE_CLUSTER * sizeof(unsigned short) per cluster, 1 KB per 2 MB of swap on a 64-bit kernel with 4 KB pages. On a vswap-heavy workload, where zswap writeback is the only consumer of physical swap, that is the common case. Such clusters never have their memcg_table read or written: vswap-layer charging records on the vswap cluster's table, not the physical one. Allocate eagerly only when the cluster is known to need a table: any cluster in a !CONFIG_VSWAP build, or any vswap cluster. For physical clusters in CONFIG_VSWAP builds, defer to alloc_swap_scan_cluster(), which allocates on the first direct-use slot and skips entirely when the cluster only holds pointer-tagged vswap backings. Signed-off-by: Nhat Pham --- mm/swapfile.c | 40 ++++++++++++++++++++++++++++++++-------- 1 file changed, 32 insertions(+), 8 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index b4d7af21ca1c..65559647eeb4 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -492,7 +492,8 @@ static void swap_cluster_free_table(struct swap_cluster= _info *ci) swap_cluster_free_table_folio_rcu_cb); } =20 -static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gf= p) +static int swap_cluster_alloc_table(struct swap_info_struct *si, + struct swap_cluster_info *ci, gfp_t gfp) { struct swap_table *table =3D NULL; struct folio *folio; @@ -515,7 +516,16 @@ static int swap_cluster_alloc_table(struct swap_cluste= r_info *ci, gfp_t gfp) rcu_assign_pointer(ci->table, table); =20 #ifdef CONFIG_MEMCG - if (!mem_cgroup_disabled()) { + /* + * Allocate memcg_table eagerly only when we know it will be used: + * any cluster in a !CONFIG_VSWAP build (all slots are direct use), + * or any vswap cluster (every vswap alloc records memcg). Physical + * clusters in a CONFIG_VSWAP build defer to alloc_swap_scan_cluster, + * which allocates on the first direct-use slot and skips entirely + * when the cluster only holds Pointer-tagged vswap backings. + */ + if ((!IS_ENABLED(CONFIG_VSWAP) || swap_is_vswap(si)) && + !mem_cgroup_disabled()) { VM_WARN_ON_ONCE(ci->memcg_table); ci->memcg_table =3D kzalloc_obj(*ci->memcg_table, gfp); if (!ci->memcg_table) { @@ -589,8 +599,8 @@ swap_cluster_populate(struct swap_info_struct *si, lockdep_assert_held(&si->global_cluster_lock); lockdep_assert_held(&ci->lock); =20 - if (!swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - __GFP_NOWARN)) + if (!swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + __GFP_NOWARN)) return ci; =20 /* @@ -608,8 +618,8 @@ swap_cluster_populate(struct swap_info_struct *si, if (!swap_is_vswap(si)) local_unlock(&percpu_swap_cluster.lock); =20 - ret =3D swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - GFP_KERNEL); + ret =3D swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + GFP_KERNEL); =20 /* * Back to atomic context. We might have migrated to a new CPU with a @@ -911,7 +921,7 @@ static int swap_cluster_setup_bad_slot(struct swap_info= _struct *si, =20 ci =3D cluster_info + idx; /* Need to allocate swap table first for initial bad slot marking. */ - if (!ci->count && swap_cluster_alloc_table(ci, GFP_KERNEL)) + if (!ci->count && swap_cluster_alloc_table(si, ci, GFP_KERNEL)) return -ENOMEM; spin_lock(&ci->lock); /* Check for duplicated bad swap slots. */ @@ -1193,6 +1203,20 @@ static unsigned int alloc_swap_scan_cluster(struct s= wap_info_struct *si, if (!ret) continue; } +#ifdef CONFIG_MEMCG + /* + * Lazy-allocate memcg_table on the first direct-use slot of a + * physical cluster. + */ + if (IS_ENABLED(CONFIG_VSWAP) && folio && + !folio_test_swapcache(folio) && !mem_cgroup_disabled() && + !ci->memcg_table) { + ci->memcg_table =3D kzalloc_obj(*ci->memcg_table, + GFP_ATOMIC | __GFP_NOWARN); + if (!ci->memcg_table) + goto out; + } +#endif if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUST= ER)) break; found =3D offset; @@ -1259,7 +1283,7 @@ static unsigned int alloc_swap_scan_dynamic(struct sw= ap_info_struct *si, spin_lock_init(&ci_dyn->ci.lock); INIT_LIST_HEAD(&ci_dyn->ci.list); =20 - if (swap_cluster_alloc_table(&ci_dyn->ci, GFP_ATOMIC)) { + if (swap_cluster_alloc_table(si, &ci_dyn->ci, GFP_ATOMIC)) { kfree(ci_dyn); return SWAP_ENTRY_INVALID; } --=20 2.53.0-Meta From nobody Tue Sep 29 14:53:50 2026 Received: from mail-oa1-f41.google.com (mail-oa1-f41.google.com [209.85.160.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8B4203BBFAE for ; Thu, 6 Aug 2026 18:43:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041798; cv=none; b=g52IlYd+6z5kBk1pPR55AuGc5raOQoMXFIxVcmRZ8jVK1SvNMnuBXd/I3h4e+/GLAeRPw+PQPcWiM6WQWlRuv1+v4B3IYbIu3Gq/wVUxRQj3WIw5bmyaQZJDScD8N+p3w75dtOogNt43jJlter8h7+TUcbbbxO1mYo+PKbrRM+0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786041798; c=relaxed/simple; bh=eumqLR4Sq/NYD9JMwB+6UlrMy9SA9TwY89K9+nUggY8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FuDAdx/kBJ2BjLz+qjM46mVksgMH054NyOWqkc4+OFN54bDxU4MLLpVQVdy1UcO4tgvvvZ8cn6slAnqjuGQFI3hqkaI58qdd9nnFO0Up/j1otQ1gHMTkx/947LYP7RP/v3ySvVhAGL3kOoy1gTcU4lBBY/xpD5b4muESiZPzKXs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=n3Ljuaoa; arc=none smtp.client-ip=209.85.160.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="n3Ljuaoa" Received: by mail-oa1-f41.google.com with SMTP id 586e51a60fabf-4560d6f82edso1677231fac.3 for ; Thu, 06 Aug 2026 11:43:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786041793; x=1786646593; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Mfv2zxxfrg+kFLvg5TbxKt0slusQCG6TQNqV0DBz5Rs=; b=n3Ljuaoa8nr72IKOltrNsr/At03BXwE/qrtHmVD4FnI83dQdrpNr+/ooSe6yVy+ukh 0mokHADi001Es0WbLhgifq7N9KIpWnn0/r4udzilzKmfWzT73TdI54x8SeaHNLk3o6/d 6Us6Tkha2Q75xl51ywksdLmntYigywFy0ztBQWcvN8bcyUWWq09NWl1ACQBKjTdj2RRj rkpN+XR4+2kbMDOGWpTpq2waqlXvNvqgI4CviABrxATH+fLT71gT+fzUkQwAlpaoPgnc iZc6ivvZ5q40dAt5qgDKiGoBBvtsh3h190ULgWsE9Y5QFD0n5G9G6ODasFM4+PNJVioD gevw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786041793; x=1786646593; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Mfv2zxxfrg+kFLvg5TbxKt0slusQCG6TQNqV0DBz5Rs=; b=TM7IgFaP6qiLLIOFBA81zOdNBab8ssN7T6izHU/MeRkzKu6qWxntm5DVFcAm1Xh9YG F/RPRZ3fHO2oQQdgnXNhG35bDeXewhKZ55aCEMq+dnpACA5R/Et9iMNoSWad4Y+wVYv8 U4Pfgv34S0q4SYbTyzVOZgDUu+1FcjwZeKFBGXgqxVlBXY+b0iTJw0CmuCpZlSXP1Kiu mmgENTl3zJRCCPq/0trnpo+GTFaPmbUAuR7PBBDVm8xRuK79nnTaIpnQsCfk0FMOrgcw c4tLDmU4saWMdHU9s6spdOV0yP7hPg/8Cylxi8W+rH5m9WC42OEipnkvwk4ub5AYGXY+ ieiw== X-Forwarded-Encrypted: i=1; AHgh+Rq830vJ8e09S0AwuV8vDMY2NrT4ke+dnvr1ODBzo17ApwNvC5RubXpn9adIWP6cmlbQPIQusM8hYfJLih8=@vger.kernel.org X-Gm-Message-State: AOJu0YxheXZzq939XW5SeWuzu1IJ21SmIAM8OtWF50BjqaXj+5Su3Iqp MhX6HNEWMdwpWF6b8a6Zdm2iSzbeTSbAbO4XuEsRuLFeDBMnOuWOuENi X-Gm-Gg: AR+sD12SXGAobKW/oRJMgiuvc3Wd+Fjo4ghWH5OJijnGcy2eISHYjDkQ94qyCmqhSPF vYh4djha7HV6T02C4ssTGWN+u7WmKvNOZ91xryLZRoRcvlE3ZGQGzegz5mo2GEXyvrNPEyErmRL ievytOnXDZ+WPwsFtcX/IaSJAN0G4A1JoVenkog5iRVRqCWHUinABHlmYT5QYbvnU/w2johjeZg Inh95uWdwrR7grg+Y0GpUQw8KUtSpsY/VgdubTB9sGlSBI36OPeAeNtS9dvtsOcKKeCgxe60Z+g CwtG8QrTJoxq5aqF4056rIZd+/px8cKTVd9BNAZh70aY5Mqx/vKr1F8AjTzv5QqIdVfOGRwptv2 77sapxpRG5wLuRD7jzrLGpzjMgIpm6gEnna23Rim4WTUYZxgCX1VCMCW7bBjuC6LoOWarZswxng QbpN+oublG/ujpwdUcm4XWw+sxe70xsqeswnoHU6rHuYk3lWtKOph1r9a/G2BMNkvGjHVXu8hrW 64e6lkXI96tGqTmqOJuYA== X-Received: by 2002:a05:6870:8929:b0:441:f3f6:622b with SMTP id 586e51a60fabf-4599eb1f492mr9226089fac.5.1786041793178; Thu, 06 Aug 2026 11:43:13 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:15::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1a0586asm197246fac.2.2026.08.06.11.43.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 11:43:12 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v3 11/11] mm, swap: widen swap_info_struct max/pages to unsigned long Date: Thu, 6 Aug 2026 11:42:54 -0700 Message-ID: <20260806184254.3790858-12-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260806184254.3790858-1-nphamcs@gmail.com> References: <20260806184254.3790858-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Widen swap_info_struct->max and ->pages from unsigned int to unsigned long so the vswap device can exceed the current 16 TB cap (ALIGN_DOWN(UINT_MAX, SWAPFILE_CLUSTER) pages). Physical swap is unaffected; backing files/bdevs continue to bound it independently of the field width. The new vswap cap is the cluster_info_pool xarray's allocator limit. XA_FLAGS_ALLOC stores allocated IDs in u32, so max_pages =3D UINT_MAX * SWAPFILE_CLUSTER (~8 PB at the typical SWAPFILE_CLUSTER=3D512 layout). Signed-off-by: Nhat Pham --- include/linux/swap.h | 4 +-- mm/swapfile.c | 62 +++++++++++++++++++++++--------------------- 2 files changed, 34 insertions(+), 32 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 19a703510675..44353eb554da 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -246,7 +246,7 @@ struct swap_info_struct { signed short prio; /* swap priority of this type */ struct plist_node list; /* entry in swap_active_head */ signed char type; /* strange name for an index */ - unsigned int max; /* size of this swap device */ + unsigned long max; /* size of this swap device */ struct swap_cluster_info *cluster_info; /* cluster info. Only for SSD */ struct list_head free_clusters; /* free clusters list */ struct list_head full_clusters; /* full clusters list */ @@ -254,7 +254,7 @@ struct swap_info_struct { /* list of cluster that contains at least one free slot */ struct list_head frag_clusters[SWAP_NR_ORDERS]; /* list of cluster that are fragmented or contented */ - unsigned int pages; /* total of usable pages of swap */ + unsigned long pages; /* total of usable pages of swap */ atomic_long_t inuse_pages; /* number of those currently in use */ struct swap_sequential_cluster *global_cluster; /* Use one global cluster= for rotating device */ spinlock_t global_cluster_lock; /* Serialize usage of global cluster */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 65559647eeb4..6db605594d48 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -450,10 +450,10 @@ static inline unsigned int cluster_index(struct swap_= info_struct *si, return ci - si->cluster_info; } =20 -static inline unsigned int cluster_offset(struct swap_info_struct *si, - struct swap_cluster_info *ci) +static inline unsigned long cluster_offset(struct swap_info_struct *si, + struct swap_cluster_info *ci) { - return cluster_index(si, ci) * SWAPFILE_CLUSTER; + return (unsigned long)cluster_index(si, ci) * SWAPFILE_CLUSTER; } =20 static void swap_cluster_free_table_folio_rcu_cb(struct rcu_head *head) @@ -904,7 +904,7 @@ static int swap_cluster_setup_bad_slot(struct swap_info= _struct *si, =20 /* si->max may got shrunk by swap swap_activate() */ if (offset >=3D si->max && !mask) { - pr_debug("Ignoring bad slot %u (max: %u)\n", offset, si->max); + pr_debug("Ignoring bad slot %u (max: %lu)\n", offset, si->max); return 0; } /* @@ -1170,12 +1170,12 @@ static bool __swap_cluster_alloc_entries(struct swa= p_info_struct *si, } =20 /* Try use a new cluster for current CPU and allocate from it. */ -static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, - struct swap_cluster_info *ci, - struct folio *folio, - unsigned long offset) +static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si, + struct swap_cluster_info *ci, + struct folio *folio, + unsigned long offset) { - unsigned int next =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; + unsigned long next =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; unsigned long start =3D ALIGN_DOWN(offset, SWAPFILE_CLUSTER); unsigned int order =3D likely(folio) ? folio_order(folio) : 0; unsigned long end =3D start + SWAPFILE_CLUSTER; @@ -1245,12 +1245,12 @@ static unsigned int alloc_swap_scan_cluster(struct = swap_info_struct *si, return found; } =20 -static unsigned int alloc_swap_scan_list(struct swap_info_struct *si, - struct list_head *list, - struct folio *folio, - bool scan_all) +static unsigned long alloc_swap_scan_list(struct swap_info_struct *si, + struct list_head *list, + struct folio *folio, + bool scan_all) { - unsigned int found =3D SWAP_ENTRY_INVALID; + unsigned long found =3D SWAP_ENTRY_INVALID; =20 do { struct swap_cluster_info *ci =3D isolate_lock_cluster(si, list); @@ -1267,8 +1267,8 @@ static unsigned int alloc_swap_scan_list(struct swap_= info_struct *si, return found; } =20 -static unsigned int alloc_swap_scan_dynamic(struct swap_info_struct *si, - struct folio *folio) +static unsigned long alloc_swap_scan_dynamic(struct swap_info_struct *si, + struct folio *folio) { struct swap_cluster_info_dynamic *ci_dyn; struct swap_cluster_info *ci; @@ -1392,7 +1392,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, { struct swap_cluster_info *ci; unsigned int order =3D likely(folio) ? folio_order(folio) : 0; - unsigned int offset =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; + unsigned long offset =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; =20 /* * Swapfile is not block device so unable @@ -3488,10 +3488,10 @@ static int unuse_mm(struct mm_struct *mm, unsigned = int type) * Return 0 if there are no inuse entries after prev till end of * the map. */ -static unsigned int find_next_to_unuse(struct swap_info_struct *si, - unsigned int prev) +static unsigned long find_next_to_unuse(struct swap_info_struct *si, + unsigned long prev) { - unsigned int i; + unsigned long i; unsigned long swp_tb; =20 /* @@ -3529,7 +3529,8 @@ static int try_to_unuse(unsigned int type) struct swap_io_ctx ctx; swp_entry_t entry, vswap_entry; unsigned long swp_tb; - unsigned int i, j; + unsigned long i; + unsigned int j; =20 if (!swap_usage_in_pages(si)) goto success; @@ -3964,7 +3965,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) struct file *swap_file, *victim; struct address_space *mapping; struct inode *inode; - unsigned int maxpages; + unsigned long maxpages; int err, found =3D 0; =20 if (!capable(CAP_SYS_ADMIN)) @@ -4386,12 +4387,8 @@ static unsigned long read_swap_header(struct swap_in= fo_struct *si, pr_warn("Truncating oversized swap area, only using %luk out of %luk\n", K(maxpages), K(last_page)); } - if (maxpages > last_page) { + if (maxpages > last_page) maxpages =3D last_page + 1; - /* p->max is an unsigned int: don't overflow it */ - if ((unsigned int)maxpages =3D=3D 0) - maxpages =3D UINT_MAX; - } =20 if (!maxpages) return 0; @@ -4630,7 +4627,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) goto bad_swap_unlock_inode; } if (si->pages !=3D si->max - 1) { - pr_err("swap:%u !=3D (max:%u - 1)\n", si->pages, si->max); + pr_err("swap:%lu !=3D (max:%lu - 1)\n", si->pages, si->max); error =3D -EINVAL; goto bad_swap_unlock_inode; } @@ -4718,7 +4715,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) /* Sets SWP_WRITEOK, resurrect the percpu ref, expose the swap device */ enable_swap_info(si); =20 - pr_info("Adding %uk swap on %s. Priority:%d extents:%d across:%lluk %s%s= %s%s\n", + pr_info("Adding %luk swap on %s. Priority:%d extents:%d across:%lluk %s%= s%s%s\n", K(si->pages), name->name, si->prio, nr_extents, K((unsigned long long)span), (si->flags & SWP_SOLIDSTATE) ? "SS" : "", @@ -4906,8 +4903,13 @@ static int __init vswap_init(void) return 0; } =20 + /* + * Cap at the cluster_info_pool xarray's allocator limit + * (XA_FLAGS_ALLOC stores IDs in u32, tops out at UINT_MAX). + */ maxpages =3D min(swapfile_maximum_size, - ALIGN_DOWN((unsigned long)UINT_MAX, SWAPFILE_CLUSTER)); + ALIGN_DOWN((unsigned long)UINT_MAX * SWAPFILE_CLUSTER, + SWAPFILE_CLUSTER)); si->flags |=3D SWP_VSWAP | SWP_SOLIDSTATE | SWP_WRITEOK; si->ops =3D &vswap_ops; si->bdev =3D NULL; --=20 2.53.0-Meta