From nobody Fri Sep 25 17:45:55 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 263343955F9; Thu, 10 Sep 2026 03:46:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789012020; cv=none; b=lSlqMXuzC4JRYdauafGpxg5SPDvUK7ZNLM0F5SUqsGRm13epY8gv1SZT6p1uAV/APeGyHDAipxy0CBpRiqgcUfJPaAgqqEwQuh30nIqbcooKjEgPy4+2DV2gdXX54MdU1XdgSXPTmGLJL8mWlyzGhBzx3hORByGPGTD9zrN28ZI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789012020; c=relaxed/simple; bh=rnmFLhQ+W54V6NRlDXw9yxyyyS6OM1jdtMbAVv5lvkA=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:To:Cc; b=bogoBU8h2CZZhDjiQSR9ekFMyczNAKEuyoEpicZta0udwdbuDz6S3MyyulhNQUsjVg/heefsffZQAwYHb5sL9wg2u+0eTH2e/Nb7uJWurKGL68bDXZbZVFPsOnyQnWCnc/J/eU7HBdHhidKxYurAZPLzlU+4zDN/AskvDEbTgQ0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=kuPgYMI7; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="kuPgYMI7" Received: by smtp.kernel.org (Postfix) with ESMTPS id ACFD3C2BCC7; Thu, 10 Sep 2026 03:46:59 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1789012019; bh=rnmFLhQ+W54V6NRlDXw9yxyyyS6OM1jdtMbAVv5lvkA=; h=From:Date:Subject:To:Cc:Reply-To:From; b=kuPgYMI7y3r9FhNTNN2j/mds3devLqSBkcPBny7+VF5pns+ARBL8ri3iLO6nEE774 sXr9CUzqodA+dc43taN3q1N6wpM1wdTBJtqWNhl/8+LBjM2j7w4FiqKhgflq6bJXXl dGKVzj2ehrCCX1jsyEFfkWlQrZO1o9xtTDcJhPd82jwmpkquzQFfLIl+Qb+Nbm3Z43 O3AqCX4SagUVLx73QV5FU89J+GNL6qC/8FGVVOPyq0WU4nM4Pc1YPQRluj2D1c5lXl e4XSUXx/O/LncjiwykCjg6dkIwpxwMLETU9Jm+clFl2yTyBGLB+hDOb73iroriH5b0 TQ60i2eFH7T+g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 972D5C79FB9; Thu, 10 Sep 2026 03:46:59 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 10 Sep 2026 11:46:58 +0800 Subject: [PATCH v5] mm/memcg: clear folio memcg after changing per memcg stats Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260910-memcg-swapcache-stats-fix-v5-1-033f510ba748@tencent.com> X-B4-Tracking: v=1; b=H4sIADEoomoC/33NQWrDMBAF0KsErasgaSzHyqr3KFlIo1Gshe1gC bch+O6dhEJSWsys/of/5iYKzZmKOO5uYqYllzyNHOzbTmDvxzPJHDkLo0yrOtPJgQY8y/LpL+i xJ1mqr0Wm/CV1IEoQEXwKgveXmbh+2B8nzn0udZqvj1eLvrc/KugNddGSD1WnHYUYuva90og01 j1Og7i7i3laTm1ahi2H3kcA1VgX/lrwapktC9g6OJtsbKxNh3+s5tWyW1bDVoTWtBEUe+m3ta7 rN2415qqoAQAA To: Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Qi Zheng , Shakeel Butt , Axel Rasmussen , Yuanchu Xie , Wei Xu , Johannes Weiner , David Hildenbrand , Michal Hocko , Lorenzo Stoakes , Bingfang Guo Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org, Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1789012018; l=6933; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=8n9tFaBEaZV6w7VSVd9OL8zcVOGtD1Os/VxlJVMxjmA=; b=u7kryou6X6EpuZcoj1fEwBBnTXHEqJB/8nmrzB3FB3H7zV3Cx6QjLZQANBvbnSkqGtXN7g5l3 447z0fdva94CF1x0/qEmjP84l6qrAuEi4T7xL+1moF3EQxzrS3V3DuQ X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo I notice extremely high swapcached count in the per memcg level memory.stat when running tests with cgroupv1 setup by swapping pages in and out. It seems that the counter never gets decreased so the value is rather useless and confusing to users reading it. So I think fixing it so that the value can reflect the actual swapcache usage correctly could be helpful. __memcg1_swapout() transfers the memsw charge of a folio to its swap entry and clears folio->memcg_data as part of that. In the vmscan swapout path it runs before __swap_cache_del_folio(), which then decrements the swapcache stats through lruvec_stat_mod_folio(). Since folio->memcg_data has already been cleared, folio_memcg() returns NULL and the NR_SWAPCACHE decrement only updates the node-level counter instead of the memcg's lruvec, leaking the per-memcg swapcache count. Move the __memcg1_swapout() call into __swap_cache_del_folio(), after the NR_FILE_PAGES and NR_SWAPCACHE updates but before __swap_cache_do_del_folio() removes the folio from the swap cache. This keeps the stats attributed to the folio's memcg while still recording the swap cgroup with a valid folio->swap. Add a swapout parameter so the plain swap_cache_del_folio() path is left unchanged. Fixes: 2732acda82c9 ("mm, swap: use swap cache as the swap in synchronize l= ayer") Signed-off-by: Bingfang Guo Acked-by: Kairui Song --- ** Testing ** The problem can be reproduced by populating 256MB of shmem pages in a memcg and then swapping them out using pageout. After the operation, we can still see those pages in the swapcached counter by reading memory.stats. After applying the given patch, the value is correctly decreased to 0. The test script and source can be found from v4 (below). ** Test Results ** before the patch ``` =3D=3D before reclaim =3D=3D swapcached 0 allocated and dirtied 256 MiB, paging out... paged out; sleeping so memory.stat can be read. pid=3D4778 =3D=3D after reclaim (swap cache should drain to ~0) =3D=3D swapcached 268435456 LEAK DETECTED: swapcached =3D 268435456 bytes (expected ~0) [BUGGY kernel] ``` after the patch ``` =3D=3D before reclaim =3D=3D swapcached 0 allocated and dirtied 256 MiB, paging out... paged out; sleeping so memory.stat can be read. pid=3D2601 =3D=3D after reclaim (swap cache should drain to ~0) =3D=3D swapcached 0 OK: swapcached =3D 0 bytes [FIXED kernel] ``` --- Changes in v5: - Refine the param description. (Kairui) - Collect Acked-by tags. - Link to v4: https://lore.kernel.org/r/20260905-memcg-swapcache-stats-fix-= v4-1-d3626d305d4f@tencent.com Changes in v4: - Update and use the correct fixes tag. (Kairui) - Briefly describe the context for @swapout param in the kdoc. (Kairui) - Cc the stable list. (Kairui) - Link to v3: https://lore.kernel.org/r/20260902-memcg-swapcache-stats-fix-= v3-1-795f5d455f7b@tencent.com Changes in v3: - Add doc for the new parameter. - Link to v2: https://lore.kernel.org/r/20260901-memcg-swapcache-stats-fix-= v2-1-9caad330459b@tencent.com Changes in v2: - Update the commit message to describe the problem in the beginnning. - Change function declaration for !CONFIG_SWAP as well. - Link to v1: https://lore.kernel.org/r/20260831-memcg-swapcache-stats-fix-= v1-1-1c0819ebdb86@tencent.com --- mm/swap.h | 6 ++++-- mm/swap_state.c | 13 ++++++++++--- mm/vmscan.c | 3 +-- 3 files changed, 15 insertions(+), 7 deletions(-) diff --git a/mm/swap.h b/mm/swap.h index 0b5d507739bcb..b3b54c28929a1 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -319,7 +319,8 @@ struct folio *swap_cache_alloc_folio(swp_entry_t target= _entry, gfp_t gfp_mask, void __swap_cache_add_folio(struct swap_cluster_info *ci, struct folio *folio, swp_entry_t entry); void __swap_cache_del_folio(struct swap_cluster_info *ci, - struct folio *folio, swp_entry_t entry, void *shadow); + struct folio *folio, swp_entry_t entry, void *shadow, + bool swapout); void __swap_cache_replace_folio(struct swap_cluster_info *ci, struct folio *old, struct folio *new); =20 @@ -452,7 +453,8 @@ static inline void swap_cache_del_folio(struct folio *f= olio) } =20 static inline void __swap_cache_del_folio(struct swap_cluster_info *ci, - struct folio *folio, swp_entry_t entry, void *shadow) + struct folio *folio, swp_entry_t entry, void *shadow, + bool swapout) { } =20 diff --git a/mm/swap_state.c b/mm/swap_state.c index 305877e1f4d7b..625c185a1ca4d 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -306,21 +306,28 @@ static void __swap_cache_do_del_folio(struct swap_clu= ster_info *ci, * @folio: The folio. * @entry: The first swap entry that the folio corresponds to. * @shadow: shadow value to be filled in the swap cache. + * @swapout: whether this folio is being reclaimed after swapout. * * Removes a folio from the swap cache and fills a shadow in place. * This won't put the folio's refcount. The caller has to do that. * * Context: Caller must ensure the folio is locked and in the swap cache * using the index of @entry, and lock the cluster that holds the entries. + * If @swapout is set, the folio should be in reclaim path and IRQs + * should be disabled. */ void __swap_cache_del_folio(struct swap_cluster_info *ci, struct folio *fo= lio, - swp_entry_t entry, void *shadow) + swp_entry_t entry, void *shadow, bool swapout) { unsigned long nr_pages =3D folio_nr_pages(folio); =20 - __swap_cache_do_del_folio(ci, folio, entry, shadow); node_stat_mod_folio(folio, NR_FILE_PAGES, -nr_pages); lruvec_stat_mod_folio(folio, NR_SWAPCACHE, -nr_pages); + + if (swapout) + __memcg1_swapout(folio, ci); + + __swap_cache_do_del_folio(ci, folio, entry, shadow); } =20 /** @@ -339,7 +346,7 @@ void swap_cache_del_folio(struct folio *folio) swp_entry_t entry =3D folio->swap; =20 ci =3D swap_cluster_lock(__swap_entry_to_info(entry), swp_offset(entry)); - __swap_cache_del_folio(ci, folio, entry, NULL); + __swap_cache_del_folio(ci, folio, entry, NULL, false); swap_cluster_unlock(ci); =20 folio_ref_sub(folio, folio_nr_pages(folio)); diff --git a/mm/vmscan.c b/mm/vmscan.c index 245f68c75b289..29bced43b0ce6 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -755,8 +755,7 @@ static int __remove_mapping(struct address_space *mappi= ng, struct folio *folio, =20 if (reclaimed && !mapping_exiting(mapping)) shadow =3D workingset_eviction(folio, target_memcg); - __memcg1_swapout(folio, ci); - __swap_cache_del_folio(ci, folio, swap, shadow); + __swap_cache_del_folio(ci, folio, swap, shadow, true); swap_cluster_unlock_irq(ci); } else { void (*free_folio)(struct folio *); --- base-commit: b5529903123d9535bcf74386c1f4e185e8632dd9 change-id: 20260828-memcg-swapcache-stats-fix-1beef3dc3afb Best regards, --=20 Bingfang Guo