From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BE0EB3A544A; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; cv=none; b=CCR4P3wL/M5dbzXH7Fp31SDAocksv1KMhxCgGVXYu5VJEbHRRNCjDv478OAAGZF/DIJYgrD70yxgi7qh5ph4G76z1+DtMn0cYsUGtiYpwTcL3rripQQcDGARvTuvkTRiXxmUcfDhpQ1RbOBacohgwVeTIGsfrVoa8blgO0ZsoP8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; c=relaxed/simple; bh=NO0EAYWsY2ldlhO3iQN1yFZTxIcpW5QnD5tDVs0IMKg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WA8vZtW+BCeO+ZUNsiTVsxIZGSBxXsni/LSzjB00RIQyskJgN2kS6RTqvs3ul3bhtM979UdfiApTdh6TPWywXBrJ/1WuidOV7PUd/bIcq6Y1+fErDk8ooPKJlOxCO2+cX4KPIRB1CJsSgQ3Kgk1A53Q4TvRwG4TIWjDUaFN5rhs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=udvo/HNH; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="udvo/HNH" Received: by smtp.kernel.org (Postfix) with ESMTPS id 4D230C2BCB9; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=NO0EAYWsY2ldlhO3iQN1yFZTxIcpW5QnD5tDVs0IMKg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=udvo/HNHAcQodDJe4VeBYpgCYG+6cyZ5dGXMZLKK+zBmGUlId2GrM2uqxC+mbR9r1 AwH4n2l0kIGVLI20cOsgLkpCTcDG2z95fCSZ590MWFVy1CzySezuJUfqRAdajLDeEu NLxbxiU4FrYvG7+oJpExh6UMxBJ71f/yrX0D+iy8G4BtTzcS6FodOdS7EpvXhDfFXU QW6a3qa2hMO4pVpti9hfkuEdXTnaohOuFXf3y8IWI6Z/iZ+ZDk3GcLg2HB2WORhanZ 5GLUQNFbE14/9PoYuAKrFgq96MleabxwAzqCB4cnYHooFWsdqM+SGFWqhomOaJlIJl y39iVeyyRnbRQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 2B5DDC55182; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:46:57 +0800 Subject: [PATCH RFC 01/15] mm/memcontrol: make lru_zone_size atomic and simplify sanity check Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-1-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=4666; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=VmGMj3oQ5WlKjZNi9BvygZeEmYIa6a/WBZg5ksVdh2w=; b=a0JTNWuPi48o3Fm1tmk0HF76NfE1tnogbipupV7MsafExoMnGyTXKGtgtIoml53Gk31D5uRJs EGfJ/FlOQdeCuvTUuafmuiYvO38wQEj4NOhVPlp6aOqx5ZKeCI4/sef X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song commit ca707239e8a7 ("mm: update_lru_size warn and reset bad lru_size") introduced a sanity check to catch memcg counter underflow, which is more like a workaround for another bug: lru_zone_size is unsigned, so underflow will wrap it around and return an enormously large number, then the memcg shrinker will loop almost forever as the calculated number of folios to shrink is huge. That commit also checks if a zero value matches the empty LRU list, so we have to hold the LRU lock, and do the counter adding differently depending on whether the nr_pages is negative. But later commit b4536f0c829c ("mm, memcg: fix the active list aging for lowmem requests when memcg is enabled") already removed the LRU emptiness check, doing the adding differently is meaningless now. And if we just turn it into an atomic long, underflow isn't a big issue either, and can be checked at the reader side. The reader size is much less frequently called than the updater. So let's turn the counter into an atomic long and check at the reader side instead, which has a smaller overhead. Use atomic to avoid potential locking issue. The underflow correction is removed, which should be fine as if there is a mass leaking of the LRU size counter, something else may also have gone very wrong, and one should fix that leaking site instead. Besides, doing the sanity check in updater is unlikely to catch the leaking site, e.g. a folio was removed minutes ago without updating the counter, while there are still other folios on the LRU, the WARN won't be triggered until other folios are removed from a likely correct callsite. For now still keep the LRU lock context, in theory that can be removed too since the update is atomic, if we can tolerate a temporary inaccurate reading, but currently there is no benefit doing so yet. Signed-off-by: Kairui Song --- include/linux/memcontrol.h | 9 +++++++-- mm/memcontrol.c | 18 +----------------- mm/vmscan.c | 5 ----- 3 files changed, 8 insertions(+), 24 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index e78bc98ab229..b13e3f056319 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -113,7 +113,7 @@ struct mem_cgroup_per_node { /* Fields which get updated often at the end. */ struct lruvec lruvec; CACHELINE_PADDING(_pad2_); - unsigned long lru_zone_size[MAX_NR_ZONES][NR_LRU_LISTS]; + atomic_long_t lru_zone_size[MAX_NR_ZONES][NR_LRU_LISTS]; struct mem_cgroup_reclaim_iter iter; =20 /* @@ -897,10 +897,15 @@ static inline unsigned long mem_cgroup_get_zone_lru_size(struct lruvec *lruvec, enum lru_list lru, int zone_idx) { + long val; struct mem_cgroup_per_node *mz; =20 mz =3D container_of(lruvec, struct mem_cgroup_per_node, lruvec); - return READ_ONCE(mz->lru_zone_size[zone_idx][lru]); + val =3D atomic_long_read(&mz->lru_zone_size[zone_idx][lru]); + if (WARN_ON_ONCE(val < 0)) + return 0; + + return val; } =20 void __mem_cgroup_handle_over_high(gfp_t gfp_mask); diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 1bde9d5af88a..7511773b3491 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -1529,28 +1529,12 @@ void mem_cgroup_update_lru_size(struct lruvec *lruv= ec, enum lru_list lru, int zid, long nr_pages) { struct mem_cgroup_per_node *mz; - unsigned long *lru_size; - long size; =20 if (mem_cgroup_disabled()) return; =20 mz =3D container_of(lruvec, struct mem_cgroup_per_node, lruvec); - lru_size =3D &mz->lru_zone_size[zid][lru]; - - if (nr_pages < 0) - *lru_size +=3D nr_pages; - - size =3D *lru_size; - if (WARN_ONCE(size < 0, - "%s(%p, %d, %ld): lru_size %ld\n", - __func__, lruvec, lru, nr_pages, size)) { - VM_BUG_ON(1); - *lru_size =3D 0; - } - - if (nr_pages > 0) - *lru_size +=3D nr_pages; + atomic_long_add(nr_pages, &mz->lru_zone_size[zid][lru]); } =20 /** diff --git a/mm/vmscan.c b/mm/vmscan.c index 17d2b793cbfc..fe7cf5d42e0c 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1636,10 +1636,6 @@ unsigned int reclaim_clean_pages_from_list(struct zo= ne *zone, return nr_reclaimed; } =20 -/* - * Update LRU sizes after isolating pages. The LRU size updates must - * be complete before mem_cgroup_update_lru_size due to a sanity check. - */ static __always_inline void update_lru_sizes(struct lruvec *lruvec, enum lru_list lru, unsigned long *nr_zone_taken) { @@ -1651,7 +1647,6 @@ static __always_inline void update_lru_sizes(struct l= ruvec *lruvec, =20 update_lru_size(lruvec, lru, zid, -nr_zone_taken[zid]); } - } =20 /* --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CB9BC3A5452; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; cv=none; b=Oh/fO5e5HSWIC7/Bv3VLBFcy1MQBu1fv2AZv7JouYwRTzkRSvYenqsJPoi/AuLotx5J6h2k2ayjWXjNr/CrdmP6jN7LU0YC/5P93S+ymAGndDQBnJvK6g6ggJKxm+L56dF5lxksbXNE+0LO8QmaodCLoyGFDnayvewH9dTKgXrE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; c=relaxed/simple; bh=TjoR3Q+A/75X5CmlKSstBL+V/RoGE3aGMjE+kfk6FR0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=M8DZq+TXAdlkNTuPRjFlZ/4EXlsncDOx+h6PLkoO2HrCaLVp12Pg/+++2VNFXok84S6M9AtollNrCOfXRzA/o6RJC3+d1werCsSDJmFeVr932fouJ76uWy49wjz2NUyH2T7KYy0M+w2H9nr8WP0uCDa/5x55qPNqlcFUjF4vDEg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=R7P1axca; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="R7P1axca" Received: by smtp.kernel.org (Postfix) with ESMTPS id 60311C2BCFA; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=TjoR3Q+A/75X5CmlKSstBL+V/RoGE3aGMjE+kfk6FR0=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=R7P1axcakJNma1d6DuaGGoEgfK2OzXAzNi+Mv/cg47QiL1r1DdF7o44v9HsRZtfSj GcAjs7OZPa+Dyp7aWJVjrQrwO/lD0mKmSNLayy+yI5cdGOfeL8H7nKWhSWR+qYFesJ nCd94i7xeq3HnPzNX6ystfYRy2auErIRpTW5YebkMSlmQZgmVxeWVP5GOxnwJHzLwT Nx7gdGo6citbgBfA04y+d4jcCKf9uaKOQS5u2ql5cshdJHeRrF1tZA+bdq7hBfw1FT QPSsF68Z7AZ9nYAZbu6b5EnXX8NrVQWkhXpaMV7w1uDLoSV712syYsX6r4YcmJBu4f tRvI3jtpwT10g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 3EDDEC55822; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:46:58 +0800 Subject: [PATCH RFC 02/15] mm/memcontrol: allow update of LRU statistic without holding LRU lock Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-2-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=2443; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=5ef1JMmzTBWQNcW2bqr39zuDrTRlv1xJUE0rVPZJ1C4=; b=UvqFhP7b4fq/TB0tz3nC3HBptpJwoGpzzJy32pvAzSQ2v9hQM5SwhY+4eceXnqcf0tEbwa++W fjNuHNjeJheAat9Cj1uzPEjUbiHDcOtF8tc1fq2VEat9TX/fYD3goID X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song To enable moving file pages in folio_mark_accessed directly and lazily for MGLRU, allow updating the LRU statistic atomically without holding a lock. It may cause temporary counter underflow, which should be fine as we still follow final consistency of the counter, and it only serves as a factor for calculating the reclaim budget in vmscan. A little inaccuracy has no visible effect. Signed-off-by: Kairui Song --- include/linux/memcontrol.h | 2 +- include/linux/mm_inline.h | 3 +-- mm/memcontrol.c | 4 ++-- 3 files changed, 4 insertions(+), 5 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index b13e3f056319..68f363000d7f 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -902,7 +902,7 @@ unsigned long mem_cgroup_get_zone_lru_size(struct lruve= c *lruvec, =20 mz =3D container_of(lruvec, struct mem_cgroup_per_node, lruvec); val =3D atomic_long_read(&mz->lru_zone_size[zone_idx][lru]); - if (WARN_ON_ONCE(val < 0)) + if (val < 0) return 0; =20 return val; diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 621c8653d8f7..3ad1619ce8ca 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -36,11 +36,10 @@ static __always_inline void __update_lru_size(struct lr= uvec *lruvec, { struct pglist_data *pgdat =3D lruvec_pgdat(lruvec); =20 - lockdep_assert_held(&lruvec->lru_lock); WARN_ON_ONCE(nr_pages !=3D (int)nr_pages); =20 mod_lruvec_state(lruvec, NR_LRU_BASE + lru, nr_pages); - __mod_zone_page_state(&pgdat->node_zones[zid], + mod_zone_page_state(&pgdat->node_zones[zid], NR_ZONE_LRU_BASE + lru, nr_pages); } =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 7511773b3491..3fbd7a6f6650 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -1522,8 +1522,8 @@ struct lruvec *folio_lruvec_lock_irqsave(struct folio= *folio, * @zid: zone id of the accounted pages * @nr_pages: positive when adding or negative when removing * - * This function must be called under lru_lock, just before a page is added - * to or just after a page is removed from an lru list. + * This function must be called when a page is added to or removed from + * an lru list. Caller need to protect the lruvec from being freed. */ void mem_cgroup_update_lru_size(struct lruvec *lruvec, enum lru_list lru, int zid, long nr_pages) --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DE5D93A5E6F; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; cv=none; b=a217Ub28yKnZ+mSceY7Y5swTluzZn0aoovTrLk5x33a621sxbNUQ/ABFa4Sx7EOoVg/RCCGT5zjJx3Ky6/VRwS+8FgXmHqPTtAwhb0hwH9B3EcHinmEPKuOQa85oZCbY5JwOkIe8+nLBOLibFgbrCWdMrBxIZvhLO4k+lUY/4NQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; c=relaxed/simple; bh=FNUTYJ+a+7ODLOGzOoW6j4FqgHRc+4ZX4B7EJpjWRyU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=XxGz13MbzWIiKW5RBqvFmQHkZds53n3Su4+cQixSG8PkuCbN1PoI1J49lyPTF/iX177HMcyU3Lw1GP36MoNit2QRlkdUdlurGwWumFtd3frzWZ9dnRYdfSMZHeJFJEjNabUYayHXK4eid4Ec38yMtMJfrML3T9HxUz+E8lERfic= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OLCl9gKu; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OLCl9gKu" Received: by smtp.kernel.org (Postfix) with ESMTPS id 70D35C2BD01; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=FNUTYJ+a+7ODLOGzOoW6j4FqgHRc+4ZX4B7EJpjWRyU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=OLCl9gKuH/4x/H5dhI5afXs7bdIYFItM6PKprk3rTJJbTj86fp8bcSKrbUewdDdUM xhPEYivmGbsGCKHIFeZVEPEH+0TK+XTuDOzkB2EwEvhi3YPnnmOk9sjQTm2PcNnYcP jMjppwii10+nNyETaX9qTgiRptPJe57z5PEMAmzqVFNU2DllIis3WDHkJF+sFjCTEs hdgw9yvlmcXxZJ3UBZksT/iA+c+ofNjtZ5U0g2U2OFGwPnxjkoP+cFDucQ2bh3jQdo PgPXUqDR4SbMcQnS8MNtWTp01TqOe96skJg1hj7tPEzCoPXxtefHRumMWODtDA7l7O AH7n4sLMAxQfA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 52579C55196; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:46:59 +0800 Subject: [PATCH RFC 03/15] mm/mglru: introduce and always use helpers for manipulating page flags Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-3-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=13228; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=j32su+iOIBqsw3MNs5eaPY0pZ0cjNkd2GlGBsSoBXBs=; b=23sX+FLqku/fL+mFJbJV6+sZUakmQeoahPbxBx/yHOVrlQz7z5P1g1FK6gtGQRbjVH4Ml8WCq 1A2/BR9mDJMCAceZpoBxrrfiazVFmdXq+GG59v4qO356e8pH6hCQcVw X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Instead of doing bit ops on folio->flags.f, introduce helpers for adjusting folio's refs and gen info, make the code easier to debug and understand. Signed-off-by: Kairui Song --- include/linux/mm_inline.h | 83 ++++++++++++++++++++++++++++++++++++++++---= ---- include/linux/mmzone.h | 2 ++ mm/folio.c | 19 ++++++----- mm/migrate.c | 2 -- mm/vmscan.c | 68 +++++++++++++++++++++----------------- 5 files changed, 122 insertions(+), 52 deletions(-) diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 3ad1619ce8ca..4076e3f7dcc8 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -141,10 +141,42 @@ static inline int lru_tier_from_refs(int refs, bool w= orkingset) return workingset ? MAX_NR_TIERS - 1 : order_base_2(refs); } =20 -static inline int folio_lru_refs(const struct folio *folio) +/** + * lru_gen_from_flags - Return the LRU generation number from folio flags. + * @flags: folio flags + * + * Returns: A number between 0 and LRU_GEN_MAX, inclusive. Returns -1 if t= he + * flags indicate the folio is off the list (e.g., isolated). + */ +static inline int lru_gen_from_flags(unsigned long flags) { - unsigned long flags =3D READ_ONCE(folio->flags.f); + int gen =3D ((flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF); =20 + gen -=3D 1; + VM_WARN_ON_ONCE(gen !=3D -1 && gen > LRU_GEN_MAX); + return gen; +} + +/** + * lru_gen_set_flags - Set the LRU generation number to specified folio fl= ags. + * @flags: pointer to the folio flags + * @gen: generation number, between 0 and LRU_GEN_MAX, inclusive. + */ +static inline void lru_gen_set_flags(unsigned long *flags, int gen) +{ + VM_WARN_ON_ONCE(gen > LRU_GEN_MAX || gen < 0); + BUILD_BUG_ON((LRU_GEN_MAX + 1) !=3D MAX_NR_GENS); + + *flags &=3D ~LRU_GEN_MASK; + *flags |=3D (gen + 1UL) << LRU_GEN_PGOFF; +} + +/** + * lru_refs_from_flags - Return LRU referenced / access count from folio f= lags. + * @flags: folio flags + */ +static inline int lru_refs_from_flags(unsigned long flags) +{ if (!(flags & BIT(PG_referenced))) return 0; /* @@ -154,18 +186,46 @@ static inline int folio_lru_refs(const struct folio *= folio) return ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) + 1; } =20 -static inline int folio_lru_gen(const struct folio *folio) +/** + * lru_refs_set_flags - Set the LRU referenced / access count to specified= folio flags. + * @flags: pointer to the folio flags + * @refs: referenced / access count number, between 0 and LRU_REFS_MAX, in= clusive. + */ +static inline void lru_refs_set_flags(unsigned long *flags, unsigned int r= efs) +{ + VM_WARN_ON_ONCE(refs > LRU_REFS_MAX); + + *flags &=3D ~LRU_REFS_FLAGS; + if (!refs) + return; + *flags |=3D (BIT(PG_referenced) | ((refs - 1UL) << LRU_REFS_PGOFF)); +} + +static inline int folio_lru_refs(const struct folio *folio) +{ + return lru_refs_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); +} + +static inline void folio_set_lru_refs(struct folio *folio, unsigned int re= fs) { - unsigned long flags =3D READ_ONCE(folio->flags.f); + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + + do { + new_flags =3D old_flags; + lru_refs_set_flags(&new_flags, refs); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); +} =20 - return ((flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1; +static inline int folio_lru_gen(const struct folio *folio) +{ + return lru_gen_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); } =20 static inline bool lru_gen_is_active(const struct lruvec *lruvec, int gen) { unsigned long max_seq =3D lruvec->lrugen.max_seq; =20 - VM_WARN_ON_ONCE(gen >=3D MAX_NR_GENS); + VM_WARN_ON_ONCE(gen > LRU_GEN_MAX); =20 /* see the comment on MIN_NR_GENS */ return gen =3D=3D lru_gen_from_seq(max_seq) || gen =3D=3D lru_gen_from_se= q(max_seq - 1); @@ -269,7 +329,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lru= vec, struct folio *folio, gen =3D lru_gen_from_seq(seq); flags =3D (gen + 1UL) << LRU_GEN_PGOFF; /* see the comment on MIN_NR_GENS about PG_active */ - set_mask_bits(&folio->flags.f, LRU_GEN_MASK | BIT(PG_active), flags); + set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK | BIT(PG_active), flags= ); =20 lru_gen_update_size(lruvec, folio, -1, gen); /* for folio_rotate_reclaimable() */ @@ -294,7 +354,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lru= vec, struct folio *folio, =20 /* for folio_migrate_flags() */ flags =3D !reclaiming && lru_gen_is_active(lruvec, gen) ? BIT(PG_active) = : 0; - flags =3D set_mask_bits(&folio->flags.f, LRU_GEN_MASK, flags); + flags =3D set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK, flags); gen =3D ((flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1; =20 lru_gen_update_size(lruvec, folio, gen, -1); @@ -305,9 +365,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lru= vec, struct folio *folio, =20 static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) { - unsigned long refs =3D READ_ONCE(old->flags.f) & LRU_REFS_MASK; - - set_mask_bits(&new->flags.f, LRU_REFS_MASK, refs); + folio_set_lru_refs(new, folio_lru_refs(old)); } #else /* !CONFIG_LRU_GEN */ =20 @@ -338,7 +396,8 @@ static inline bool lru_gen_del_folio(struct lruvec *lru= vec, struct folio *folio, =20 static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) { - + if (folio_test_referenced(old)) + folio_set_referenced(new); } #endif /* CONFIG_LRU_GEN */ =20 diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index a26c8b855222..8048c6b0544d 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -496,7 +496,9 @@ enum lruvec_flags { #ifndef __GENERATING_BOUNDS_H =20 #define LRU_GEN_MASK ((BIT(LRU_GEN_WIDTH) - 1) << LRU_GEN_PGOFF) +#define LRU_GEN_MAX (BIT(LRU_GEN_WIDTH - 1) - 1) #define LRU_REFS_MASK ((BIT(LRU_REFS_WIDTH) - 1) << LRU_REFS_PGOFF) +#define LRU_REFS_MAX BIT(LRU_REFS_WIDTH) =20 /* * For folios accessed multiple times through file descriptors, diff --git a/mm/folio.c b/mm/folio.c index a9e328c3f21b..f90b7f86dbe3 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -353,26 +353,28 @@ static void __lru_cache_activate_folio(struct folio *= folio) =20 static void lru_gen_inc_refs(struct folio *folio) { - unsigned long new_flags, old_flags =3D READ_ONCE(folio->flags.f); + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + int refs; =20 if (folio_test_unevictable(folio)) return; =20 /* see the comment on LRU_REFS_FLAGS */ - if (!folio_test_referenced(folio)) { - set_mask_bits(&folio->flags.f, LRU_REFS_MASK, BIT(PG_referenced)); + if (!folio_lru_refs(folio)) { + folio_set_lru_refs(folio, 1); return; } =20 do { - if ((old_flags & LRU_REFS_MASK) =3D=3D LRU_REFS_MASK) { + new_flags =3D old_flags; + refs =3D lru_refs_from_flags(old_flags); + if (refs =3D=3D LRU_REFS_MAX) { if (!folio_test_workingset(folio)) folio_set_workingset(folio); return; } - - new_flags =3D old_flags + BIT(LRU_REFS_PGOFF); - } while (!try_cmpxchg(&folio->flags.f, &old_flags, new_flags)); + lru_refs_set_flags(&new_flags, refs + 1); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); } =20 static bool lru_gen_clear_refs(struct folio *folio) @@ -384,7 +386,8 @@ static bool lru_gen_clear_refs(struct folio *folio) if (gen < 0) return true; =20 - set_mask_bits(&folio->flags.f, LRU_REFS_FLAGS | BIT(PG_workingset), 0); + folio_set_lru_refs(folio, 0); + folio_clear_workingset(folio); =20 rcu_read_lock(); seq =3D READ_ONCE(folio_lruvec(folio)->lrugen.min_seq[type]); diff --git a/mm/migrate.c b/mm/migrate.c index b937cbd76480..c737d0682fa4 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -779,8 +779,6 @@ void folio_migrate_flags(struct folio *newfolio, struct= folio *folio) { int cpupid; =20 - if (folio_test_referenced(folio)) - folio_set_referenced(newfolio); if (folio_test_uptodate(folio)) folio_mark_uptodate(newfolio); if (folio_test_clear_active(folio)) { diff --git a/mm/vmscan.c b/mm/vmscan.c index fe7cf5d42e0c..f5b0a7c63a3a 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -842,19 +842,22 @@ static bool lru_gen_set_refs(struct folio *folio, con= st vma_flags_t *vma_flags) if (!folio_test_referenced(folio) && !folio_test_workingset(folio)) { /* Activate file-backed executable folios after first usage. */ if (is_exec_file_folio(folio, vma_flags)) { - set_mask_bits(&folio->flags.f, LRU_REFS_FLAGS, BIT(PG_workingset)); + folio_set_lru_refs(folio, 0); + folio_set_workingset(folio); return true; } =20 - set_mask_bits(&folio->flags.f, LRU_REFS_MASK, BIT(PG_referenced)); + folio_set_lru_refs(folio, 1); return false; } =20 /* Promote on second access */ - if (folio_lru_refs(folio) > 1) - set_mask_bits(&folio->flags.f, LRU_REFS_FLAGS, BIT(PG_workingset)); - else + if (folio_lru_refs(folio) > 1) { + folio_set_lru_refs(folio, 0); + folio_set_workingset(folio); + } else { folio_mark_accessed(folio); + } return true; } #else @@ -3260,11 +3263,10 @@ static bool positive_ctrl_err(struct ctrl_pos *sp, = struct ctrl_pos *pv) *************************************************************************= *****/ =20 /* promote pages accessed through page tables */ -static int folio_update_gen(struct folio *folio, int gen, const vma_flags_= t *vma_flags) +static int folio_update_gen(struct folio *folio, int new_gen, const vma_fl= ags_t *vma_flags) { - unsigned long new_flags, old_flags =3D READ_ONCE(folio->flags.f); - - VM_WARN_ON_ONCE(gen >=3D MAX_NR_GENS); + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + int old_gen; =20 /* * See the comment on LRU_REFS_FLAGS, and activate file-backed @@ -3273,20 +3275,24 @@ static int folio_update_gen(struct folio *folio, in= t gen, const vma_flags_t *vma */ if (!folio_test_referenced(folio) && !folio_test_workingset(folio) && !is_exec_file_folio(folio, vma_flags)) { - set_mask_bits(&folio->flags.f, LRU_REFS_MASK, BIT(PG_referenced)); + folio_set_lru_refs(folio, 1); return -1; } =20 do { + old_gen =3D lru_gen_from_flags(old_flags); + new_flags =3D old_flags; + /* lru_gen_del_folio() has isolated this page? */ - if (!(old_flags & LRU_GEN_MASK)) - return -1; + if (old_gen < 0) + break; =20 - new_flags =3D old_flags & ~(LRU_GEN_MASK | LRU_REFS_FLAGS); - new_flags |=3D ((gen + 1UL) << LRU_GEN_PGOFF) | BIT(PG_workingset); - } while (!try_cmpxchg(&folio->flags.f, &old_flags, new_flags)); + lru_gen_set_flags(&new_flags, new_gen); + lru_refs_set_flags(&new_flags, 0); + new_flags |=3D BIT(PG_workingset); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); =20 - return ((old_flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1; + return old_gen; } =20 /* protect pages accessed multiple times through file descriptors */ @@ -3294,22 +3300,22 @@ static int folio_inc_gen(struct lruvec *lruvec, str= uct folio *folio) { int type =3D folio_is_file_lru(folio); struct lru_gen_folio *lrugen =3D &lruvec->lrugen; - int new_gen, old_gen =3D lru_gen_from_seq(lrugen->min_seq[type]); - unsigned long new_flags, old_flags =3D READ_ONCE(folio->flags.f); - - VM_WARN_ON_ONCE_FOLIO(!(old_flags & LRU_GEN_MASK), folio); + int old_gen, new_gen, min_gen =3D lru_gen_from_seq(lrugen->min_seq[type]); + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); =20 do { - new_gen =3D ((old_flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1; + old_gen =3D lru_gen_from_flags(old_flags); + VM_WARN_ON_ONCE_FOLIO(old_gen < 0, folio); + /* folio_update_gen() has promoted this page? */ - if (new_gen >=3D 0 && new_gen !=3D old_gen) - return new_gen; + if (old_gen >=3D 0 && old_gen !=3D min_gen) + return old_gen; =20 + new_flags =3D old_flags; new_gen =3D (old_gen + 1) % MAX_NR_GENS; - - new_flags =3D old_flags & ~(LRU_GEN_MASK | LRU_REFS_FLAGS); - new_flags |=3D (new_gen + 1UL) << LRU_GEN_PGOFF; - } while (!try_cmpxchg(&folio->flags.f, &old_flags, new_flags)); + lru_gen_set_flags(&new_flags, new_gen); + lru_refs_set_flags(&new_flags, 0); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); =20 lru_gen_update_size(lruvec, folio, old_gen, new_gen); =20 @@ -4712,7 +4718,7 @@ static bool isolate_folio(struct lruvec *lruvec, stru= ct folio *folio, struct sca =20 /* see the comment on LRU_REFS_FLAGS */ if (!folio_test_referenced(folio)) - set_mask_bits(&folio->flags.f, LRU_REFS_MASK, 0); + folio_set_lru_refs(folio, 0); =20 success =3D lru_gen_del_folio(lruvec, folio, true); VM_WARN_ON_ONCE_FOLIO(!success, folio); @@ -4930,8 +4936,10 @@ static int evict_folios(unsigned long nr_to_scan, st= ruct lruvec *lruvec, } =20 /* don't add rejected folios to the oldest generation */ - if (lru_gen_folio_seq(lruvec, folio, false) =3D=3D min_seq[type]) - set_mask_bits(&folio->flags.f, LRU_REFS_FLAGS, BIT(PG_active)); + if (lru_gen_folio_seq(lruvec, folio, false) =3D=3D min_seq[type]) { + folio_set_lru_refs(folio, 0); + folio_set_active(folio); + } } =20 move_folios_to_lru(&list); --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CF35C3A1682; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; cv=none; b=DvEy6BiB9DtZbfaR1/QqGaVkTjV7jG+H0mdO8XGy49ygzAJYHyCA02DHtkdb2WvIrmi1Lw1UWRbC6f0MFHQ/ixXMGamItd3rp9EYHkMhfFv3XsXK6CtpmmG+gTjnuomEtMeQRfA1wBxjtFM3U7TCk02nYwfzPhD7XH4VOG/A9yg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786442; c=relaxed/simple; bh=SDJd3vW3qkKBS3CdNbHJyIAmbUP3QRPpQi1oINDrVZE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=a3nNcah55niK6+ZQkO60XTjfh/JV3lS/BSXZmjxMNH58F6oUfJnqh2UVsOT1lQvgtMDYp73y+CeTrq/CEugnK4GDQMXhbLpmbsoAekztdH3kUYP2OmGvVlHYJPANZP2KUct0k4uHcxhzQiU95O4Y5pZjUTb20A6N3HVhu4JH6ag= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=nsHo+j7j; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="nsHo+j7j" Received: by smtp.kernel.org (Postfix) with ESMTPS id 7AC70C4AF11; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=SDJd3vW3qkKBS3CdNbHJyIAmbUP3QRPpQi1oINDrVZE=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=nsHo+j7jd2vx98UxP4tyzOOtNN+FVFsQsoXSJloYLREMymCHSA4L03zFuvjoTcnDQ 33yc8gdxMpla5D1wIL9JJj1tGx4BE8+q0pkzWpZRerE7yBVMaIQg7NXXTleRtDb43a nmTDNhEHccHHoHnGkTj87GK9aj3Ck1oy0rshvCCcY6lXgLYwVXgtq7624h7kni5KQC YmLNgVlrQfdJG4M3Rjr9i+MCZ/C0KgR97AuK+6qNwkzY0oCYHYLEe5RXzUr5mXSQxi sqhTSikM1KJDctKIRJDentqEpzDlGei0oKesll+47UrpMKKFqGs/PU2oL+uSOtt5ms IJn0+ERzT+lYg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 64A88C55AB9; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:00 +0800 Subject: [PATCH RFC 04/15] mm/mglru: make generation page counters atomic Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-4-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=5062; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=t4exwehi8Ei4/lRorOvsWBcZjkV0WY4vPQIntyWBAJ0=; b=XbPr3kZEY7eYz5dQbFhpzZo36GfAY5LKWawkIw/5lYUPRWbkiMxoYdg7f0L3E4rW2cRtlhkF2 sNRk59Xa+OyCI1aHCtTsfGqGIryV/wtrtu+tPl7iFjY4mYLkqpEcW95 X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song No feature change, convert them to atomic so we can update them without holding the LRU lock. There is no risk of overflow. The reader always compares and uses zero instead if the counter values are negative. It follows final consistency. Signed-off-by: Kairui Song --- include/linux/mm_inline.h | 6 ++---- include/linux/mmzone.h | 2 +- mm/vmscan.c | 20 ++++++++++---------- 3 files changed, 13 insertions(+), 15 deletions(-) diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 4076e3f7dcc8..018a2f54a5c9 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -245,11 +245,9 @@ static inline void lru_gen_update_size(struct lruvec *= lruvec, struct folio *foli VM_WARN_ON_ONCE(old_gen =3D=3D -1 && new_gen =3D=3D -1); =20 if (old_gen >=3D 0) - WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone], - lrugen->nr_pages[old_gen][type][zone] - delta); + atomic_long_sub(delta, &lrugen->nr_pages[old_gen][type][zone]); if (new_gen >=3D 0) - WRITE_ONCE(lrugen->nr_pages[new_gen][type][zone], - lrugen->nr_pages[new_gen][type][zone] + delta); + atomic_long_add(delta, &lrugen->nr_pages[new_gen][type][zone]); =20 /* addition */ if (old_gen < 0) { diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 8048c6b0544d..4225dab760ba 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -572,7 +572,7 @@ struct lru_gen_folio { /* the multi-gen LRU lists, lazily sorted on eviction */ struct list_head folios[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES]; /* the multi-gen LRU sizes, eventually consistent */ - long nr_pages[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES]; + atomic_long_t nr_pages[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES]; /* the exponential moving average of refaulted */ unsigned long avg_refaulted[ANON_AND_FILE][MAX_NR_TIERS]; /* the exponential moving average of evicted+protected */ diff --git a/mm/vmscan.c b/mm/vmscan.c index f5b0a7c63a3a..b02d2ec8ff4b 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -3354,8 +3354,7 @@ static void reset_batch_size(struct lru_gen_mm_walk *= walk) continue; =20 walk->nr_pages[gen][type][zone] =3D 0; - WRITE_ONCE(lrugen->nr_pages[gen][type][zone], - lrugen->nr_pages[gen][type][zone] + delta); + atomic_long_add(delta, &lrugen->nr_pages[gen][type][zone]); =20 if (lru_gen_is_active(lruvec, gen)) lru +=3D LRU_ACTIVE; @@ -4044,8 +4043,8 @@ static bool inc_max_seq(struct lruvec *lruvec, unsign= ed long seq, int swappiness for (type =3D 0; type < ANON_AND_FILE; type++) { for (zone =3D 0; zone < MAX_NR_ZONES; zone++) { enum lru_list lru =3D type * LRU_INACTIVE_FILE; - long delta =3D lrugen->nr_pages[prev][type][zone] - - lrugen->nr_pages[next][type][zone]; + long delta =3D atomic_long_read(&lrugen->nr_pages[prev][type][zone]) - + atomic_long_read(&lrugen->nr_pages[next][type][zone]); =20 if (!delta) continue; @@ -4163,7 +4162,8 @@ static unsigned long lruvec_evictable_size(struct lru= vec *lruvec, int swappiness for (seq =3D min_seq[type]; seq <=3D max_seq; seq++) { gen =3D lru_gen_from_seq(seq); for (zone =3D 0; zone < MAX_NR_ZONES; zone++) - total +=3D max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L); + total +=3D max(atomic_long_read(&lrugen->nr_pages[gen][type][zone]), + 0L); } } =20 @@ -4598,7 +4598,7 @@ static void __lru_gen_reparent_memcg(struct lruvec *c= hild_lruvec, struct lruvec =20 for (i =3D 0; i < get_nr_gens(child_lruvec, type); i++) { int gen =3D lru_gen_from_seq(child_lrugen->max_seq - i); - long nr_pages =3D child_lrugen->nr_pages[gen][type][zone]; + long nr_pages =3D atomic_long_read(&child_lrugen->nr_pages[gen][type][zo= ne]); int child_lru_active =3D lru_gen_is_active(child_lruvec, gen) ? LRU_ACTI= VE : 0; int parent_lru_active =3D lru_gen_is_active(parent_lruvec, gen) ? LRU_AC= TIVE : 0; =20 @@ -4606,9 +4606,8 @@ static void __lru_gen_reparent_memcg(struct lruvec *c= hild_lruvec, struct lruvec list_splice_tail_init(&child_lrugen->folios[gen][type][zone], &parent_lrugen->folios[gen][type][zone]); =20 - WRITE_ONCE(child_lrugen->nr_pages[gen][type][zone], 0); - WRITE_ONCE(parent_lrugen->nr_pages[gen][type][zone], - parent_lrugen->nr_pages[gen][type][zone] + nr_pages); + atomic_long_set(&child_lrugen->nr_pages[gen][type][zone], 0); + atomic_long_add(nr_pages, &parent_lrugen->nr_pages[gen][type][zone]); =20 if (lru_gen_is_active(child_lruvec, gen) !=3D lru_gen_is_active(parent_l= ruvec, gen)) { __update_lru_size(child_lruvec, lru + child_lru_active, zone, -nr_pages= ); @@ -5650,7 +5649,8 @@ static int lru_gen_seq_show(struct seq_file *m, void = *v) char mark =3D full && seq < min_seq[type] ? 'x' : ' '; =20 for (zone =3D 0; zone < MAX_NR_ZONES; zone++) - size +=3D max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L); + size +=3D max(atomic_long_read(&lrugen->nr_pages[gen][type][zone]), + 0L); =20 seq_printf(m, " %10lu%c", size, mark); } --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E2923A718D; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=lQW0VFO+Hl5jXsxQuEdD6GLQ9sZMIQPh8gQBMCUBSE/3oEm5Bl64lxMPNQQ1hPSKdgMDzBroHkHajd7H6I4rIdlz3+bldySI/WB9xsrWjSYMPY3zy2VWKFnVtBKrrjUVChe4ikIjRoCmNjiLGVqTKMCTNC26qLMZGpddZHT6eeA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=7oMjq8JOU3eHb6zNoW7+5UIf8RcJuFNkmV1SUXaIj68=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=H8yr2zOy4PzHVerwYFitfB3NvFzc5n+TOy/vwRN05x+rlE4c60+SLyK2deYPCeyI0j5f0GClJmkQxb+93AvJ7B/8SgrFu0xcgTqgiphEMN38LNxx6FtUBX3MyiBBgvS8m42QlzdlsISVFAPtkEFGFhulfl8XygDee0qv2oFSMwQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=kjJMqvJc; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="kjJMqvJc" Received: by smtp.kernel.org (Postfix) with ESMTPS id 8E6A4C4AF1B; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=7oMjq8JOU3eHb6zNoW7+5UIf8RcJuFNkmV1SUXaIj68=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=kjJMqvJcoxW6UFujRhq8Ekv2nNzw0LvezVEjFKCRIKjsCT+5GXzs7u6gy1yiHWEP3 SkrKom0mDiWfg+AjcTMAAa0OWGbVcSwpCjs1qTfCY3Qk2iymoPrRnzznVK2TtjpBjU wF6uIqiCkPWTGdBPDFL7xUF1nqQ3qW1xawfTG1/Yvb55+xZRRvv0B+Ecl4d3lnf6oa 3tesC7OUTMVURDnDNcFE8N2+WJNvw3naNwI64uu2B1ZM0GhTRNhS5/Xu8N8rYyH1dw zZ37M6qeq8ihBeiDpF0iGa+BQMx2ykFPljHIqpW9IDGWZ1L+AoFB5GItcGaRoK2waI 7srdIPifotxzQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 79284C55182; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:01 +0800 Subject: [PATCH RFC 05/15] mm/mglru: move max_seq read into walk_update_folio Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-5-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=4925; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=wSChZul1NmIryDYTKVLvOQpbv0SaiqhBZY/xmEybbCE=; b=1D47c3BXTihaaxXk5hhFvD/DLD/AcvdEQ6h35D1WEGFowZbRbvMRS5rfkKZOmT8jVlgf7xrkl UmLoLumuOA5DlyahcOpO0Xo67vojMx6kuNn57rXr4wHXutlCf/PRsKM X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song walk_pte_range(), walk_pmd_range_locked(), and lru_gen_look_around() each read lrugen->max_seq to compute the target generation for folio_update_gen(), then pass it as a parameter to walk_update_folio(). Move the read into walk_update_folio() itself so the callers no longer need to compute or pass the value. This is a pure refactoring: no functional change. Signed-off-by: Kairui Song Reviewed-by: Baoquan He --- mm/vmscan.c | 29 ++++++++++++----------------- 1 file changed, 12 insertions(+), 17 deletions(-) diff --git a/mm/vmscan.c b/mm/vmscan.c index b02d2ec8ff4b..c2ea92c2b69e 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -3511,13 +3511,15 @@ static bool suitable_to_scan(int total, int young) } =20 static void walk_update_folio(struct lru_gen_mm_walk *walk, struct vm_area= _struct *vma, - struct folio *folio, int new_gen, bool dirty) + struct lruvec *lruvec, struct folio *folio, bool dirty) { - int old_gen; + int new_gen, old_gen; =20 if (!folio) return; =20 + new_gen =3D lru_gen_from_seq(READ_ONCE(lruvec->lrugen.max_seq)); + if (dirty && !folio_test_dirty(folio) && !(folio_test_anon(folio) && folio_test_swapbacked(folio) && !folio_test_swapcache(folio))) @@ -3548,8 +3550,6 @@ static bool walk_pte_range(pmd_t *pmd, unsigned long = start, unsigned long end, struct lru_gen_mm_walk *walk =3D args->private; struct mem_cgroup *memcg =3D lruvec_memcg(walk->lruvec); struct pglist_data *pgdat =3D lruvec_pgdat(walk->lruvec); - DEFINE_MAX_SEQ(walk->lruvec); - int gen =3D lru_gen_from_seq(max_seq); unsigned int nr; pmd_t pmdval; =20 @@ -3600,7 +3600,7 @@ static bool walk_pte_range(pmd_t *pmd, unsigned long = start, unsigned long end, continue; =20 if (last !=3D folio) { - walk_update_folio(walk, args->vma, last, gen, dirty); + walk_update_folio(walk, args->vma, walk->lruvec, last, dirty); =20 last =3D folio; dirty =3D false; @@ -3613,7 +3613,7 @@ static bool walk_pte_range(pmd_t *pmd, unsigned long = start, unsigned long end, walk->mm_stats[MM_LEAF_YOUNG] +=3D nr; } =20 - walk_update_folio(walk, args->vma, last, gen, dirty); + walk_update_folio(walk, args->vma, walk->lruvec, last, dirty); last =3D NULL; =20 if (i < PTRS_PER_PTE && get_next_vma(PMD_MASK, PAGE_SIZE, args, &start, &= end)) @@ -3636,8 +3636,6 @@ static void walk_pmd_range_locked(pud_t *pud, unsigne= d long addr, struct vm_area struct lru_gen_mm_walk *walk =3D args->private; struct mem_cgroup *memcg =3D lruvec_memcg(walk->lruvec); struct pglist_data *pgdat =3D lruvec_pgdat(walk->lruvec); - DEFINE_MAX_SEQ(walk->lruvec); - int gen =3D lru_gen_from_seq(max_seq); =20 VM_WARN_ON_ONCE(pud_leaf(*pud)); =20 @@ -3691,7 +3689,7 @@ static void walk_pmd_range_locked(pud_t *pud, unsigne= d long addr, struct vm_area goto next; =20 if (last !=3D folio) { - walk_update_folio(walk, vma, last, gen, dirty); + walk_update_folio(walk, vma, walk->lruvec, last, dirty); =20 last =3D folio; dirty =3D false; @@ -3705,7 +3703,7 @@ static void walk_pmd_range_locked(pud_t *pud, unsigne= d long addr, struct vm_area i =3D i > MIN_LRU_BATCH ? 0 : find_next_bit(bitmap, MIN_LRU_BATCH, i) + = 1; } while (i <=3D MIN_LRU_BATCH); =20 - walk_update_folio(walk, vma, last, gen, dirty); + walk_update_folio(walk, vma, walk->lruvec, last, dirty); =20 lazy_mmu_mode_disable(); spin_unlock(ptl); @@ -4270,8 +4268,6 @@ bool lru_gen_look_around(struct page_vma_mapped_walk = *pvmw, unsigned int nr) struct pglist_data *pgdat =3D folio_pgdat(folio); struct lruvec *lruvec; struct lru_gen_mm_state *mm_state; - unsigned long max_seq; - int gen; =20 lockdep_assert_held(pvmw->ptl); VM_WARN_ON_ONCE_FOLIO(folio_test_lru(folio), folio); @@ -4308,8 +4304,6 @@ bool lru_gen_look_around(struct page_vma_mapped_walk = *pvmw, unsigned int nr) =20 memcg =3D get_mem_cgroup_from_folio(folio); lruvec =3D mem_cgroup_lruvec(memcg, pgdat); - max_seq =3D READ_ONCE((lruvec)->lrugen.max_seq); - gen =3D lru_gen_from_seq(max_seq); mm_state =3D get_mm_state(lruvec); =20 lazy_mmu_mode_enable(); @@ -4341,7 +4335,7 @@ bool lru_gen_look_around(struct page_vma_mapped_walk = *pvmw, unsigned int nr) continue; =20 if (last !=3D folio) { - walk_update_folio(walk, vma, last, gen, dirty); + walk_update_folio(walk, vma, lruvec, last, dirty); =20 last =3D folio; dirty =3D false; @@ -4353,13 +4347,14 @@ bool lru_gen_look_around(struct page_vma_mapped_wal= k *pvmw, unsigned int nr) young +=3D nr; } =20 - walk_update_folio(walk, vma, last, gen, dirty); + walk_update_folio(walk, vma, lruvec, last, dirty); =20 lazy_mmu_mode_disable(); =20 /* feedback from rmap walkers to page table walkers */ if (mm_state && suitable_to_scan(i, young)) - update_bloom_filter(mm_state, max_seq, pvmw->pmd); + update_bloom_filter(mm_state, READ_ONCE(lruvec->lrugen.max_seq), + pvmw->pmd); =20 mem_cgroup_put(memcg); =20 --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E1FF3A718C; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=nsmSxaZ78AB9Gg4gI154dJAPbipGTS0V8ZwbMgxf93VS1x3kivcrp979Yd9tGSeZ4shr8K+YG6cYWzBbsu41/m6uCBKgZuSJ3bz0bkiCFTIE3iULaZ0e1yfD0KHlvrKmwmBigc/L9bFYB5W0k65l0/ftO5BWU8BGDvcErOHas20= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=HQaTSb9gKEiXAW8L884mD+Mrb/OnWgT7I1M4oD8pA9I=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=PW6PwrBk8lb0rNxju+8fXxVcqc7zUKYbBY8ALrd5dhgOylxaTfMFea6JNCtUYC2Q5dhgAZetvIXJ+zQduUmH964ji0fMCrgyz5NSIFhev+10F2cFbg2kRjeMGknid1wTSbVBEleNU61ERR2rgxlwKMaayH3HXzagiOJnR5I8s3M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=jbTvdM4k; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="jbTvdM4k" Received: by smtp.kernel.org (Postfix) with ESMTPS id A4EAAC4AF50; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=HQaTSb9gKEiXAW8L884mD+Mrb/OnWgT7I1M4oD8pA9I=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=jbTvdM4kU319OryzxazU0+wim1SrBjWHRGzrjsZ/dE/8RKZ/X/pWewA2SgMACUbRU 768yYi9K8pPMEYD+lgTURjA56h4lGz3rhufrxlusn1dEWblp0AJlCmyqqa0jEJLtVz yCipkzFcNfSrEETg7dAvwy0VMAUtAayuIa+9OaV0rGWUQm/P+uTUuMRZSlYR/egXhM oocuydPT8RyO4FtJTpZYTYFHOXpf7m1MloNKOGK9cn07MQuIgPetlA1jku8E7YOQ/Q b+QXd5oV1eZxq79+1X5IAB8QGFoieiVHIvp1OYXaBw+Nz3vKJ/Db5hgSTtQxBfpu15 9Pl9ZwmygdQtQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8F13EC55843; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:02 +0800 Subject: [PATCH RFC 06/15] mm/mglru: use explicit tier range in read_ctrl_pos() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-6-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=2788; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=sdC9n1WO6u/L2tFXmmCkir7R9Dni9ihu1pJnxCifFTc=; b=xv52kVfAPZlxWefP3f8wGwS9plTRIu50H/DgY4be7zffsSSa+ZutP0PLQZylfBQsByG3RFGPe vxKVGP4iz+OCGKU2vFXYDZzUJzjoXQqLB4c3nywiA7e41WfE2Dqkv4H X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song read_ctrl_pos() encodes the tier range in a single "tier" parameter via "tier % MAX_NR_TIERS" as the start and "min(tier, MAX_NR_TIERS-1)" as the end. This is hard to follow or maintain or extend. Tier values 0..3 select a single tier, while tier =3D=3D MAX_NR_TIERS selects the full range. Replace it with explicit (tier_min, tier_max) parameters using a half-open [tier_min, tier_max) interval, which is the conventional C idiom. The call sites become self-documenting: - get_tier_idx: (0, 1) for tier 0, (tier, tier+1) for each tier - get_type_to_scan: (0, MAX_NR_TIERS) for the full range No functional change. Signed-off-by: Kairui Song --- mm/vmscan.c | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/mm/vmscan.c b/mm/vmscan.c index c2ea92c2b69e..a359d5a1ff41 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -3192,8 +3192,8 @@ struct ctrl_pos { int gain; }; =20 -static void read_ctrl_pos(struct lruvec *lruvec, int type, int tier, int g= ain, - struct ctrl_pos *pos) +static void read_ctrl_pos(struct lruvec *lruvec, int type, int tier_min, + int tier_max, int gain, struct ctrl_pos *pos) { int i; struct lru_gen_folio *lrugen =3D &lruvec->lrugen; @@ -3202,7 +3202,7 @@ static void read_ctrl_pos(struct lruvec *lruvec, int = type, int tier, int gain, pos->gain =3D gain; pos->refaulted =3D pos->total =3D 0; =20 - for (i =3D tier % MAX_NR_TIERS; i <=3D min(tier, MAX_NR_TIERS - 1); i++) { + for (i =3D tier_min; i < tier_max; i++) { pos->refaulted +=3D lrugen->avg_refaulted[type][i] + atomic_long_read(&lrugen->refaulted[hist][type][i]); pos->total +=3D lrugen->avg_total[type][i] + @@ -4805,9 +4805,9 @@ static int get_tier_idx(struct lruvec *lruvec, int ty= pe) * This value is chosen because any other tier would have at least twice * as many refaults as the first tier. */ - read_ctrl_pos(lruvec, type, 0, 2, &sp); + read_ctrl_pos(lruvec, type, 0, 1, 2, &sp); for (tier =3D 1; tier < MAX_NR_TIERS; tier++) { - read_ctrl_pos(lruvec, type, tier, 3, &pv); + read_ctrl_pos(lruvec, type, tier, tier + 1, 3, &pv); if (!positive_ctrl_err(&sp, &pv)) break; } @@ -4828,8 +4828,8 @@ static int get_type_to_scan(struct lruvec *lruvec, in= t swappiness) * Compare the sum of all tiers of anon with that of file to determine * which type to scan. */ - read_ctrl_pos(lruvec, LRU_GEN_ANON, MAX_NR_TIERS, swappiness, &sp); - read_ctrl_pos(lruvec, LRU_GEN_FILE, MAX_NR_TIERS, MAX_SWAPPINESS - swappi= ness, &pv); + read_ctrl_pos(lruvec, LRU_GEN_ANON, 0, MAX_NR_TIERS, swappiness, &sp); + read_ctrl_pos(lruvec, LRU_GEN_FILE, 0, MAX_NR_TIERS, MAX_SWAPPINESS - swa= ppiness, &pv); =20 return positive_ctrl_err(&sp, &pv); } --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F4083A71AD; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=WIsjWki3cDafqIU31VvvJVjMzymmldNwG9gJk26D5zjp2AY+Fu6Nnaal0HWC8L7nSDAoGod/eSHtzrIEvyxFtRXCyJnbQo+kFkSS/ZWq4RBu+p8Rw0aGeQXe/F7nz3W8kx885U0yfg/0Dc1gHoE9tU0dYkVqXWMMKvuXM8UhI60= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=cYbOLfejrKyQ9BPd8CMH99ZNpXZ1w9drr+TtG7Ncs1I=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=CfMSG2O9WrPe4YoXKgQNaE6m5MUkLiaD/AsXz+amp2WHsNw+Nqm8TqXzy/R2f9IOVg1PMgXNxdQCbVc0x4ydbps/OI24dceSdoVkafvGnOR4ITYw0nX4r7A0ipPhJ8UrQ9jkWFz8lzhuZbcm8vnJ7+1CwpyOY/RZ0HgU0DjEXaE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=DR6NlXu1; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="DR6NlXu1" Received: by smtp.kernel.org (Postfix) with ESMTPS id BBDDAC32782; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=cYbOLfejrKyQ9BPd8CMH99ZNpXZ1w9drr+TtG7Ncs1I=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=DR6NlXu1rq5Z+SXhDN9ulerai7GEJwrzpL9Q56W/L9ph2wNEmAl6eYzYKPIeLnBsn NgZAi1VRVCaYS0vFiXM+6nfNaChvZoDUTxS7se3JaKl7wkHXinNHVOBblqePpDMDKF 9iz/D1V3MZJeehVkPHTeffNs+iRIP/Zur2ckBO+qxQhZCDHgnXolj2iJ1VES1sIrAp EHRKdgFz1n993debHt7uHf1+Cb/4LSsNfpSSLksax6/8KgV8/mMg3pw+BirsymwJqY qii9NSRFISffUDeKF6BuX/lzEKGvUyejypMd/sxebeJSJXVh4949yokP/lf4MG1NLW PUWum8Y5zcOJg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id A426DC55184; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:03 +0800 Subject: [PATCH RFC 07/15] mm/mglru: move refault workingset activation into lru_gen_refault Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-7-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=2070; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=1vC+xi9jNiCocZFLKmWvRW3U2STkCUDTFpNlV8n9RbA=; b=52vn6IfvHhVTl+P2eITAdJpvk846TJfwLfY88yxmTc88qbpfSMc32svILrESnyC2fHKo7ZD0j ZaK57q/2igSCfTqDJ7iXTgDStAVyQJT+CccWDakvux35J5eCy4eOMqP X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Move the folio_set_active() for refaulted workingset folios from folio_add_lru() into lru_gen_refault(), where the refault detection already happens. No functional change: the ordering and logic are preserved in all cases, and no other paths reach the removed branch. This is a preparatory cleanup for MGLRU-FG. --- mm/folio.c | 6 +----- mm/workingset.c | 9 ++++----- 2 files changed, 5 insertions(+), 10 deletions(-) diff --git a/mm/folio.c b/mm/folio.c index f90b7f86dbe3..fab00cb02970 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -474,17 +474,13 @@ void folio_add_lru(struct folio *folio) VM_BUG_ON_FOLIO(folio_test_lru(folio), folio); =20 /* - * For refaulted workingset folios, set PG_active so they - * can be added to active generations. * For prefaulted file folios, folio_mark_accessed() sets * PG_referenced so lru_gen_folio_seq() places them into * the second oldest generation. */ if (lru_gen_enabled() && !folio_test_unevictable(folio) && lru_gen_in_fault() && !(current->flags & PF_MEMALLOC)) { - if (folio_test_workingset(folio)) - folio_set_active(folio); - else if (!folio_test_referenced(folio)) + if (!folio_test_referenced(folio) && !folio_test_workingset(folio)) folio_mark_accessed(folio); } =20 diff --git a/mm/workingset.c b/mm/workingset.c index 7ac2b88c80ae..5438e9390011 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -320,12 +320,11 @@ static void lru_gen_refault(struct folio *folio, void= *shadow) atomic_long_add(delta, &lrugen->refaulted[hist][type][tier]); =20 if (workingset) { - /* - * see folio_add_lru(), where folio_set_active() is - * called for workingset folios - */ - if (lru_gen_in_fault()) + /* Send refaulted workingset folios to active generations. */ + if (lru_gen_in_fault()) { + folio_set_active(folio); mod_lruvec_state(lruvec, WORKINGSET_ACTIVATE_BASE + type, delta); + } folio_set_workingset(folio); mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + type, delta); } else --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3458B3A785C; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=hxF0XY1+jjAzfToJwxLISrwp+Tz5RgWS37DfhwzbS+953zHEvM5N2NNuF6s/3tbAwxSYHwmDOmmDkA2Ceg4NVUySWRjJ+dYTQ3KW8IZ2Q1HjWldkuruNnZzGc8dDolp8ZGykzJx0Le+bTIokajKtoD297k4VPezKej+8ov2pCgo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=sFUdOmE7PPq4dTTxszgE5i0xfhfb5mMlbqX0i1zr9aU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Nd2kTDoSF4uATPyWDtsCeCOmJLyJKWomh4sLuhb0t4B1DJi8uwmhjfJqT6yKuDKiNLdOHGtj78YAD0xVg3bxbqiPdOuLJxqjGwwdBw40EfgdutLNuC2vMK4g4tATvsRItUmWEMzL+yaXaEJmJgdLGRSQgMAaKeC3iiooRlNcHac= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ZUq52jCl; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ZUq52jCl" Received: by smtp.kernel.org (Postfix) with ESMTPS id D0011C4AF14; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786442; bh=sFUdOmE7PPq4dTTxszgE5i0xfhfb5mMlbqX0i1zr9aU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ZUq52jClfqCojXBb7kZZeNU1BU0+C14rPWs5LSmrSLot7mRQPqxYH7HZN8pT4lyIv XJKm+VV9jWxwZOIbxBfBm0drKR8pQX8YVsh/ZTV7tAlGDznW98wxaWmAaLpBgDjvfK O0pBgmRo60yQhS2T85Sx9edHG+54OcMp+2x6Dnrcp0wiFSfNRGmqsGK94o9gVtEJIE 5kbKn+diWyCivlu/yL9VSo+WVPKBm8M6OfxAJ1tsuQU8dFF4C4EuB+6EM0S7LKLmcQ dcnc8fB/eMTj8GYwv8Bn4uQxqZ4ctz7QjaILV1dnDsh+XL8a7DF+0n+mU1y//xtwz4 ZGZOaLcmkMzug== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id B9AA8C55822; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:04 +0800 Subject: [PATCH RFC 08/15] mm/memcg: add folio-based lruvec live helper Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-8-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=1942; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=CYaQx+ZxB9rM3tJj/qRBO+O5Fxa1QWpmGb5iZwXshX8=; b=++Am5L1YzqiSnxDr9o1VMbCVEwEukS5UN1/zmmuMnpfuljIk/qOzAXVOhUiwyoRYhYQDYt45R xgE/vtt5N5wAukyhcBnL6BpQE79+Z9dNxi2hwZQC227lqiolRkVLMAv X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Add a helper that resolves a stable lruvec for a folio under RCU without taking the lruvec lock. It takes a folio directly so the lruvec lookup happens inside the RCU read-side critical section, which a lruvec-based interface cannot guarantee. The lock-taking variant now inlines the ancestor walk instead of calling a separate helper. No functional change. Signed-off-by: Kairui Song --- include/linux/memcontrol.h | 38 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 38 insertions(+) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 68f363000d7f..ea0111392b9b 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -1506,6 +1506,44 @@ static inline void lruvec_lock_irq(struct lruvec *lr= uvec) spin_lock_irq(&lruvec->lru_lock); } =20 +/** + * folio_lruvec_live_get - get a live lruvec for a folio under RCU + * @folio: the folio + * + * Computes @folio's lruvec and walks up to the nearest live ancestor + * if the folio's memcg is dying. Must be paired with + * folio_lruvec_live_put(). + * + * Return: the live lruvec, with rcu_read_lock held. + */ +static inline struct lruvec *folio_lruvec_live_get(struct folio *folio) +{ +#ifdef CONFIG_MEMCG + struct lruvec *lruvec; + struct pglist_data *pgdat; + struct mem_cgroup *memcg; + + rcu_read_lock(); + lruvec =3D folio_lruvec(folio); + pgdat =3D lruvec_pgdat(lruvec); + memcg =3D lruvec_memcg(lruvec); + while (unlikely(memcg && css_is_dying(&memcg->css))) { + memcg =3D parent_mem_cgroup(memcg); + lruvec =3D mem_cgroup_lruvec(memcg, pgdat); + } + return lruvec; +#else + return folio_lruvec(folio); +#endif +} + +static inline void folio_lruvec_live_put(struct lruvec *lruvec) +{ +#ifdef CONFIG_MEMCG + rcu_read_unlock(); +#endif +} + static inline struct lruvec *lruvec_live_lock_irq(struct lruvec *lruvec) { #ifdef CONFIG_MEMCG --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 40E233A7D7A; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=pCDIybkYb75g3qFfMLaQxhHgAu7qlVX6uSRexVCpxwM7BY2Bj5gbE8Lc4hzZ7mBs2aNXmWlEvhlEeh+a3Fu7jbXe43xmpAdWPzfH0pu/O1JPxVvfwQgTDIs2QnPMH0Ndolr0paAdLBCeBSyRwLNODKEtbSARXBb+wNqA3hoJrOQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=k8Ezd9Qdeh2DboRyTdsd3mjI2la+SQMFmr9yieWzrhU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Py0P7iTpq6KNk7/KNY6PPlhmI90t7Sd+jU/PSw0MKM1PqOdy3WrCgI1xezmA6hb1WzABA/fRBYwL0EmSMwIEBRvcduLe3lTjJdPIYxkA75/fM9hXF430QkAhuiqtRQ1MLsViG89EdU1Y1JgWI9gmJvJ9NkRCHwzORQnc3oJFuHs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=kcChvQYm; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="kcChvQYm" Received: by smtp.kernel.org (Postfix) with ESMTPS id EC9EEC2BCF6; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=k8Ezd9Qdeh2DboRyTdsd3mjI2la+SQMFmr9yieWzrhU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=kcChvQYmZDRdOo6FZcEONIrnfrBUpZlU+GRo+mxEdkFrgWGV2dT++VL0yjwkTdd7d efGiLZT27UOAzvP0WiUwafIH/BhaQzaULsldXBzc/NZr0OlVKuf8N7bUXci9eWBWqB 5HKUfRbJNjDW/a+FZDWkBRH1E/tauhoZbG7Y+0RVAUBc9u+3cTrHuR1GKdJ8W0UVkE qN+DCAOlQFAR/owpYTGtLONGw2Bno+V3euHp0SQcpibrOiwlwl/yFmpyVO80Tscm7y w2q82DATazXhoDUNlTPJrR1ZaGDb/NAz4Wl3y9T5MGeASEUPlXvC5rh5NLlqRuU5z5 uTBrP0w9fCVUQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id D481DC55182; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:05 +0800 Subject: [PATCH RFC 09/15] mm/mglru: frequency guided workingset promotion (MGLRU-FG) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-9-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=43085; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=PZVvdYTo29B+se/YGXvix1IVIJ4PBPQU7sKD9u9WNsE=; b=eOyEQ66pEutii2irfmeNMyZIjccwM8pMYrHNSfCqJ1S6VhxQd0nrq/COXHmYPXRLnXFB46nAU M30N7LZ64ZIBKNQ7Gm3WRRjwzWJSxFZHSVhCfy4vsDWiKs1Ns+DaRSf X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Complement MGLRU's eviction-time tier-PID protection with access-time frequency-guided promotion. Introduce a unified set of helpers built based on referenced (access) count of a folio. Each access increments a folio's referenced count stored in folio flags (refs), refs still mappes to a logarithmic tier just like before, but with more formal bit definitions, a few special thresholds are introduced: LRU_REFS_REFERENCED (1), LRU_REFS_WORKINGSET (2), LRU_REFS_PROTECTED (3), and LRU_REFS_MAX(7). When it reaches certain threshold, the folio is promoted proactively instead of wait for the PID controller to kick in. Also simplify MGLRU's usage of PG_workingset and PG_referenced, now these 2 flags are purely used as the lower 2 bit of refs for MGLRU. This doesn't effect classical LRU in any way. This will actually simplify and make MGLRU's certain metric reading more accurate, and reduced MGLRU's original tier / referenced count bit by one since only one extra bit is now needed to record a max referenced count of 7 (previously 2 extra bits are needed). This changes make sense because MGLRU doesn't have demotion so these 2 flags are never separately useful for MGLRU. This addresses several shortcomings of the old model: - Long feedback loop: protection only activated after enough re-faults, by which time the folio is often no longer hot. - Limited tier resolution: once referenced count exceeded the bits limit (8 previously), MGLRU could no longer distinguish hotter folios as they are capped by the tier. And what's worse, PG_workingset forces a folio to stay on tier 3. - Eviction-time bias: because PID protection activates upon eviction and always targets the LRU tail, it tends to protect cold tail folios at the expense of hotter head folios. Once the tail folios consume the PID protection budget, head folios lose their protection. Additionally, the PID cannot distinguish the access time of folios that share the same reference count. Besides reworking the LRU_REFS related helpers and definitions, most of the work is done by the helpers below; the implementation details are described in their inline comments. - folio_inc_lru_refs(): Used by both cache access (folio_mark_accessed) and page table access. The folio could be off-list (isolated), unlocked, or unmapped. This helper uses PG_lru to stabilize the folio and performs a speculative and lazy promotion. - folio_inc_lru_refs_walk(): Used by the PTE walk path during aging, where generations are stable; performs lazy promotion. - folio_inc_lru_refs_isolated(): Used by the rmap check before eviction. The folio is isolated and hence this doesn't perform promotion by itself; the folio will be added back to the right gen upon return. The eviction-time folio_inc_gen() still handles PID protection, but the protection ratio is softer than before, and it caps refs at WORKINGSET so the folio retains enough history to stay above the cold tier. The PID controller gain factors in get_tier_idx() are also relaxed from (2:3) to (1:2). Since the new folio gen bump paths already proactively protect hot folios, PID protection can afford to be more permissive without increasing the refault rate. PG_workingset and PG_referenced are repurposed as the low two bits of the unified LRU reference count. LRU_REFS_MASK provides the higher bits. This eliminates the old restriction where LRU_REFS_MASK was only valid when PG_referenced was set, and allows all paths to use the same encoding consistently. Hence, a workingset folio is now defined as refs >=3D LRU_REFS_WORKINGSET (2), matching the active/inactive LRU's definition and giving in-kernel consumers (PSI, readahead) consistent behavior on MGLRU, which will be done in later commits. Note that PG_workingset and PG_referenced are no longer independent flags under MGLRU. Adjusting existing raw folio_test_*() callers to the new semantics is left as follow-ups. Signed-off-by: Kairui Song --- include/linux/mm_inline.h | 83 ++++++++----- include/linux/mmzone.h | 135 ++++++++++++++------- kernel/bounds.c | 2 +- mm/folio.c | 46 +------- mm/vmscan.c | 290 ++++++++++++++++++++++++++++++------------= ---- mm/workingset.c | 55 ++++++--- 6 files changed, 385 insertions(+), 226 deletions(-) diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 018a2f54a5c9..944baa91bf18 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -133,12 +133,13 @@ static inline int lru_hist_from_seq(unsigned long seq) return seq % NR_HIST_GENS; } =20 -static inline int lru_tier_from_refs(int refs, bool workingset) +static inline int lru_tier_from_refs(unsigned int refs) { - VM_WARN_ON_ONCE(refs > BIT(LRU_REFS_WIDTH)); - - /* see the comment on MAX_NR_TIERS */ - return workingset ? MAX_NR_TIERS - 1 : order_base_2(refs); + BUILD_BUG_ON(fls(LRU_REFS_MAX - 1) > MAX_NR_TIERS - 1); + VM_WARN_ON_ONCE(refs > LRU_REFS_MAX); + if (refs < LRU_REFS_WORKINGSET) + return 0; + return fls(refs - 1); } =20 /** @@ -164,9 +165,8 @@ static inline int lru_gen_from_flags(unsigned long flag= s) */ static inline void lru_gen_set_flags(unsigned long *flags, int gen) { - VM_WARN_ON_ONCE(gen > LRU_GEN_MAX || gen < 0); BUILD_BUG_ON((LRU_GEN_MAX + 1) !=3D MAX_NR_GENS); - + VM_WARN_ON_ONCE(gen > LRU_GEN_MAX || gen < 0); *flags &=3D ~LRU_GEN_MASK; *flags |=3D (gen + 1UL) << LRU_GEN_PGOFF; } @@ -177,13 +177,16 @@ static inline void lru_gen_set_flags(unsigned long *f= lags, int gen) */ static inline int lru_refs_from_flags(unsigned long flags) { - if (!(flags & BIT(PG_referenced))) - return 0; + int refs; + /* - * Return the total number of accesses including PG_referenced. Also see - * the comment on LRU_REFS_FLAGS. + * Return the total number of accesses. Also see the comment on + * LRU_REFS_FLAGS. */ - return ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) + 1; + refs =3D (flags & BIT(PG_referenced)) ? BIT(0) : 0; + refs +=3D (flags & BIT(PG_workingset)) ? BIT(1) : 0; + refs +=3D ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) << 2; + return refs; } =20 /** @@ -194,11 +197,13 @@ static inline int lru_refs_from_flags(unsigned long f= lags) static inline void lru_refs_set_flags(unsigned long *flags, unsigned int r= efs) { VM_WARN_ON_ONCE(refs > LRU_REFS_MAX); - + BUILD_BUG_ON((LRU_REFS_MAX >> 2) > (BIT(LRU_REFS_WIDTH) - 1)); *flags &=3D ~LRU_REFS_FLAGS; - if (!refs) - return; - *flags |=3D (BIT(PG_referenced) | ((refs - 1UL) << LRU_REFS_PGOFF)); + if (refs & BIT(0)) + *flags |=3D BIT(PG_referenced); + if (refs & BIT(1)) + *flags |=3D BIT(PG_workingset); + *flags |=3D (((unsigned long)refs) >> 2) << LRU_REFS_PGOFF; } =20 static inline int folio_lru_refs(const struct folio *folio) @@ -216,6 +221,8 @@ static inline void folio_set_lru_refs(struct folio *fol= io, unsigned int refs) } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); } =20 +int folio_inc_lru_refs(struct folio *folio, bool is_fault, bool is_exec); + static inline int folio_lru_gen(const struct folio *folio) { return lru_gen_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); @@ -223,7 +230,7 @@ static inline int folio_lru_gen(const struct folio *fol= io) =20 static inline bool lru_gen_is_active(const struct lruvec *lruvec, int gen) { - unsigned long max_seq =3D lruvec->lrugen.max_seq; + unsigned long max_seq =3D READ_ONCE(lruvec->lrugen.max_seq); =20 VM_WARN_ON_ONCE(gen > LRU_GEN_MAX); =20 @@ -280,23 +287,24 @@ static inline unsigned long lru_gen_folio_seq(const s= truct lruvec *lruvec, bool reclaiming) { int gen; + int refs =3D folio_lru_refs(folio); int type =3D folio_is_file_lru(folio); const struct lru_gen_folio *lrugen =3D &lruvec->lrugen; =20 /* - * +-----------------------------------+---------------------------------= --+ - * | Accessed through page tables and | Accessed through file descriptor= s | - * | promoted by folio_update_gen() | and protected by folio_inc_gen()= | - * +-----------------------------------+---------------------------------= --+ - * | PG_active (set while isolated) | = | - * +-----------------+-----------------+-----------------+---------------= --+ - * | PG_workingset | PG_referenced | PG_workingset | LRU_REFS_FLAG= S | - * +-----------------------------------+---------------------------------= --+ - * |<---------- MIN_NR_GENS ---------->| = | - * |<---------------------------- MAX_NR_GENS ---------------------------= ->| + * +------------------------------------------+--------------------------= ----------------+ + * | Accessed through page tables and | Accessed through file= descriptors | + * | promoted by folio_inc_lru_refs_walk() | protected by folio_inc_lr= u_refs/inc_gen | + * +------------------------------------------+--------------------------= ----------------+ + * | PG_active (set at isolation or refault) | = | + * +--------------------+---------------------+--------------------+-----= ----------------+ + * | LRU_REFS_MAX | LRU_REFS_WORKINGSET | LRU_REFS_MAX | LRU_= REFS_WORKINGSET | + * +------------------------------------------+--------------------------= ----------------+ + * |<-------------- MIN_NR_GENS ------------->| = | + * |<----------------------------------- MAX_NR_GENS --------------------= --------------->| */ if (folio_test_active(folio)) - gen =3D MIN_NR_GENS - folio_test_workingset(folio); + gen =3D MIN_NR_GENS - (refs >=3D LRU_REFS_WORKINGSET); else if (reclaiming) gen =3D MAX_NR_GENS; else if ((!folio_is_file_lru(folio) && !folio_test_swapcache(folio)) || @@ -304,7 +312,7 @@ static inline unsigned long lru_gen_folio_seq(const str= uct lruvec *lruvec, (folio_test_dirty(folio) || folio_test_writeback(folio)))) gen =3D MIN_NR_GENS; else - gen =3D MAX_NR_GENS - (folio_test_workingset(folio) || folio_test_refere= nced(folio)); + gen =3D MAX_NR_GENS - (refs >=3D LRU_REFS_WORKINGSET); =20 return max(READ_ONCE(lrugen->max_seq) - gen + 1, READ_ONCE(lrugen->min_se= q[type])); } @@ -365,6 +373,7 @@ static inline void folio_migrate_refs(struct folio *new= , const struct folio *old { folio_set_lru_refs(new, folio_lru_refs(old)); } + #else /* !CONFIG_LRU_GEN */ =20 static inline bool lru_gen_enabled(void) @@ -392,10 +401,26 @@ static inline bool lru_gen_del_folio(struct lruvec *l= ruvec, struct folio *folio, return false; } =20 +static inline int folio_lru_refs(const struct folio *folio) +{ + return 0; +} + +static inline void folio_set_lru_refs(struct folio *folio, unsigned int re= fs) +{ +} + +static inline int folio_inc_lru_refs(struct folio *folio, bool promote, bo= ol is_exec) +{ + return 0; +} + static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) { if (folio_test_referenced(old)) folio_set_referenced(new); + if (folio_test_workingset(old)) + folio_set_workingset(new); } #endif /* CONFIG_LRU_GEN */ =20 diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 4225dab760ba..e4f7efc02e50 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -472,56 +472,111 @@ enum lruvec_flags { #define MAX_NR_GENS 4U =20 /* - * Each generation is divided into multiple tiers. A folio accessed N times - * through file descriptors is in tier order_base_2(N). A folio in the fir= st - * tier (N=3D0,1) is marked by PG_referenced unless it was faulted in thro= ugh page - * tables or read ahead. A folio in the last tier (MAX_NR_TIERS-1) is mark= ed by - * PG_workingset. A folio in any other tier (1flags. + * Each generation is divided into multiple tiers. A folio's referenced + * count maps to a tier as shown below: * - * In contrast to moving across generations which requires the LRU lock, m= oving - * across tiers only involves atomic operations on folio->flags and theref= ore - * has a negligible cost in the buffered access path. In the eviction path, - * comparisons of refaulted/(evicted+protected) from the first tier and th= e rest - * infer whether folios accessed multiple times through file descriptors a= re - * statistically hot and thus worth protecting. + * MGLRU (frequency guidance) + * Refs Tier |- Refs: how many times (at least) a folio has been refere= nced. + * 0 0 |- Mostly cold pages, readahead, etc. [1] + * 1 0 |=3D LRU_REFS_REFERENCED: Used at least once. [2] + * -WORKINGSET-+|- Pages beyond are workingset and never fall below this f= loor. [3] + * 2 1<-+|=3D LRU_REFS_WORKINGSET: Classical workingset, accessed t= wice, protected. [4] + * 3 2 |- LRU_REFS_PROTECTED: Protected workingset, promoted page= s capped at here. [5] + * 4* 2 | + * 5* 3 |- The tier here is MAX_NR_TIERS - 1 + * 6* 3 | + * 7* 3 |=3D LRU_REFS_MAX: Promotion candidate. [6] + * -PROMOTION->-/ * - * MAX_NR_TIERS is set to 4 so that the multi-gen LRU can support twice the - * number of categories of the active/inactive LRU when keeping track of - * accesses through file descriptors. This uses MAX_NR_TIERS-2 spare bits = in - * folio->flags, masked by LRU_REFS_MASK. + * Ideally each tier holds folios of similar access patterns: lower tiers + * are less important and evicted faster. A page's reference count and + * tier are capped when it changes generation, preventing it from + * dominating the new generation based on old-generation access history. + * Generation ordering already ensures a newer-gen page is hotter than an + * older-gen one regardless of tier. + * + * Refs tracks accesses from two sources: page table (lazily collected by + * the page table aging walk or rmap eviction lookup) and file descriptors + * (by folio_mark_accessed). Page table accesses are weighted heavier + * because the accessed bit is sticky (undercounts repeated accesses), + * passively collected, and page faults are generally more important as + * userspace does not expect a memory access to block on reclaim. Both + * access types increment refs by one; the result is capped at + * LRU_REFS_PROTECTED on promotion or deferral, or LRU_REFS_MAX otherwise. + * + * 1. Tier is fls(N-1) for N > 1, 0 otherwise. Folios with zero + * accesses (refs =3D=3D 0) are generally cold, e.g. readahead folios. + * + * Page table access advances a folio by one generation even at the + * lowest refs or tier. Freshly allocated folios start with refs =3D= =3D 0; + * faulted and mapped folios have their page table access bit set, so + * the first page table access check always sets LRU_REFS_REFERENCED and + * moves them one generation forward, driving aging and workingset shif= t. + * + * 2. Folios accessed once stay on tier 0: one-time usage does not + * qualify for protection. A second access advances the folio, + * aligning with classical LRU's use-twice threshold. A second page + * table access promotes to the latest gen; file access only defers + * eviction from the oldest gen. + * + * 3. Folios accessed at least twice are considered workingset. This + * mostly aligns with classical LRU: at least one I/O is saved by + * keeping them in memory. Folios at or above this level never fall + * below tier 1 (the workingset floor), so tier 0 stays a clean tier + * for cold cache while tier 1 serves as the fallback line for + * actually reused or historically hot folios. + * + * Folios refaulted through a page fault at refs 1 will enter the second + * newest gen, so faulting will be protected better. + * + * 4. Starting from tier 1, PID protection sacrifices lower tiers to + * protect higher tiers by comparing refault rates for long-term + * accuracy, and caps higher refs to this value. Since PID protection + * bypasses page table lookup and clearing, when a further eviction + * attempt occurs after PID loosens, the folio's page table access is + * rechecked and the folio is sent back to LRU_REFS_PROTECTED. This + * also gives folios a fair opportunity to be promoted by file access + * again. + * + * Folios refaulted through a page fault at tier 1 or above are activat= ed + * and enter the newest gen. Non fault page will enter second oldest ge= n, + * driving aging and workingset shifting. + * + * 5. Pages beyond the ordinary workingset tier form new tiers for the + * PID controller to protect differently. Folios at or above this + * level are capped at LRU_REFS_PROTECTED on promotion or deferral, + * and at LRU_REFS_WORKINGSET under PID protection in the oldest + * generation, where they represent a historical workingset. + * + * 6. Folios that reach LRU_REFS_MAX are advanced to the next generation + * on further access, with refs capped to LRU_REFS_PROTECTED. This + * gives them a fair start for advancement to an even newer generation + * while keeping hot folios distinguishable. + * + * Tiering uses PG_referenced and PG_workingset as the lower two bits, + * and the bits masked by LRU_REFS_MASK as the higher bits. + * + * A folio's referenced count never goes backwards except upon gen + * increase as described above. Refault of a reclaimed folio restores + * its referenced count, capped at LRU_REFS_PROTECTED, which aligns with + * promotion. Page table refaults of previous workingset folios send + * them to the latest gen, driving aging faster. + * + * MAX_NR_TIERS is set to 4 so that the multi-gen LRU can support twice + * the number of categories of the active/inactive LRU. */ #define MAX_NR_TIERS 4U +#define LRU_REFS_REFERENCED 0x1 +#define LRU_REFS_WORKINGSET 0x2 +#define LRU_REFS_PROTECTED 0x3 =20 #ifndef __GENERATING_BOUNDS_H =20 #define LRU_GEN_MASK ((BIT(LRU_GEN_WIDTH) - 1) << LRU_GEN_PGOFF) #define LRU_GEN_MAX (BIT(LRU_GEN_WIDTH - 1) - 1) #define LRU_REFS_MASK ((BIT(LRU_REFS_WIDTH) - 1) << LRU_REFS_PGOFF) -#define LRU_REFS_MAX BIT(LRU_REFS_WIDTH) - -/* - * For folios accessed multiple times through file descriptors, - * lru_gen_inc_refs() sets additional bits of LRU_REFS_WIDTH in folio->fla= gs - * after PG_referenced, then PG_workingset after LRU_REFS_WIDTH. After all= its - * bits are set, i.e., LRU_REFS_FLAGS|BIT(PG_workingset), a folio is lazily - * promoted into the second oldest generation in the eviction path. And wh= en - * folio_inc_gen() does that, it clears LRU_REFS_FLAGS so that - * lru_gen_inc_refs() can start over. Note that for this case, LRU_REFS_MA= SK is - * only valid when PG_referenced is set. - * - * For folios accessed multiple times through page tables, folio_update_ge= n() - * from a page table walk or lru_gen_set_refs() from a rmap walk sets - * PG_referenced after the accessed bit is cleared for the first time. - * Thereafter, those two paths set PG_workingset and promote folios to the - * youngest generation. Like folio_inc_gen(), folio_update_gen() also clea= rs - * PG_referenced. Note that for this case, LRU_REFS_MASK is not used. - * - * For both cases above, after PG_workingset is set on a folio, it remains= until - * this folio is either reclaimed, or "deactivated" by lru_gen_clear_refs(= ). It - * can be set again if lru_gen_test_recent() returns true upon a refault. - */ -#define LRU_REFS_FLAGS (LRU_REFS_MASK | BIT(PG_referenced)) +#define LRU_REFS_FLAGS (LRU_REFS_MASK | BIT(PG_referenced) | BIT(PG_worki= ngset)) +#define LRU_REFS_MAX (BIT(LRU_REFS_WIDTH + 2) - 1) =20 struct lruvec; struct page_vma_mapped_walk; diff --git a/kernel/bounds.c b/kernel/bounds.c index 02b619eb6106..06a034713b5d 100644 --- a/kernel/bounds.c +++ b/kernel/bounds.c @@ -25,7 +25,7 @@ int main(void) DEFINE(SPINLOCK_SIZE, sizeof(spinlock_t)); #ifdef CONFIG_LRU_GEN DEFINE(LRU_GEN_WIDTH, order_base_2(MAX_NR_GENS + 1)); - DEFINE(__LRU_REFS_WIDTH, MAX_NR_TIERS - 2); + DEFINE(__LRU_REFS_WIDTH, MAX_NR_TIERS - 3); #else DEFINE(LRU_GEN_WIDTH, 0); DEFINE(__LRU_REFS_WIDTH, 0); diff --git a/mm/folio.c b/mm/folio.c index fab00cb02970..a326602a59fe 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -272,7 +272,6 @@ static void lru_activate(struct lruvec *lruvec, struct = folio *folio) if (folio_test_active(folio) || folio_test_unevictable(folio)) return; =20 - lruvec_del_folio(lruvec, folio); folio_set_active(folio); lruvec_add_folio(lruvec, folio); @@ -351,32 +350,6 @@ static void __lru_cache_activate_folio(struct folio *f= olio) =20 #ifdef CONFIG_LRU_GEN =20 -static void lru_gen_inc_refs(struct folio *folio) -{ - unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); - int refs; - - if (folio_test_unevictable(folio)) - return; - - /* see the comment on LRU_REFS_FLAGS */ - if (!folio_lru_refs(folio)) { - folio_set_lru_refs(folio, 1); - return; - } - - do { - new_flags =3D old_flags; - refs =3D lru_refs_from_flags(old_flags); - if (refs =3D=3D LRU_REFS_MAX) { - if (!folio_test_workingset(folio)) - folio_set_workingset(folio); - return; - } - lru_refs_set_flags(&new_flags, refs + 1); - } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); -} - static bool lru_gen_clear_refs(struct folio *folio) { int gen =3D folio_lru_gen(folio); @@ -387,7 +360,6 @@ static bool lru_gen_clear_refs(struct folio *folio) return true; =20 folio_set_lru_refs(folio, 0); - folio_clear_workingset(folio); =20 rcu_read_lock(); seq =3D READ_ONCE(folio_lruvec(folio)->lrugen.min_seq[type]); @@ -398,10 +370,6 @@ static bool lru_gen_clear_refs(struct folio *folio) =20 #else /* !CONFIG_LRU_GEN */ =20 -static void lru_gen_inc_refs(struct folio *folio) -{ -} - static bool lru_gen_clear_refs(struct folio *folio) { return false; @@ -427,7 +395,8 @@ void folio_mark_accessed(struct folio *folio) if (folio_test_dropbehind(folio)) return; if (lru_gen_enabled()) { - lru_gen_inc_refs(folio); + if (!folio_test_unevictable(folio)) + folio_inc_lru_refs(folio, false, false); return; } =20 @@ -473,17 +442,6 @@ void folio_add_lru(struct folio *folio) folio_test_unevictable(folio), folio); VM_BUG_ON_FOLIO(folio_test_lru(folio), folio); =20 - /* - * For prefaulted file folios, folio_mark_accessed() sets - * PG_referenced so lru_gen_folio_seq() places them into - * the second oldest generation. - */ - if (lru_gen_enabled() && !folio_test_unevictable(folio) && - lru_gen_in_fault() && !(current->flags & PF_MEMALLOC)) { - if (!folio_test_referenced(folio) && !folio_test_workingset(folio)) - folio_mark_accessed(folio); - } - folio_batch_add_and_move(folio, lru_add); } EXPORT_SYMBOL(folio_add_lru); diff --git a/mm/vmscan.c b/mm/vmscan.c index a359d5a1ff41..9c8d9e3af375 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -830,38 +830,177 @@ enum folio_references { }; =20 #ifdef CONFIG_LRU_GEN +/*************************************************************************= ***** + * Referenced count feedback + *************************************************************************= *****/ + /* - * Only used on a mapped folio in the eviction (rmap walk) path, where pro= motion - * needs to be done by taking the folio off the LRU list and then adding i= t back - * with PG_active set. In contrast, the aging (page table walk) path uses - * folio_update_gen(). + * The folio_inc_lru_refs{_*} helpers below collect the referenced info + * (hotness) from other parts, including the page table walker, the rmap w= alk + * upon eviction, the rmap lookaround, and file descriptors + * (folio_mark_accessed). + * + * Page table accesses escalate a folio in two steps. The first access + * advances it one generation; a second access sends it to the newest + * generation. Executable file folios skip the first step and are promoted + * immediately, as reclaiming them causes IO thrashing. + * + * File descriptor accesses do not promote. They only defer eviction from + * the oldest generation, and only once the folio is a workingset folio + * (LRU_REFS_WORKINGSET), leaving the rest to PID protection. Page table + * accesses are treated more generously because the accessed bit is sticky + * (it under-counts repeated accesses) and because a page fault is more + * costly than file descriptor I/O. + * + * PID protection operates on tier > 0 folios. The one proactive promotion + * outside of it and the page table path is the overflow case where the + * referenced count exceeds LRU_REFS_MAX, which means the folio is hotter + * than everything else in its generation. + * + * Whenever a folio changes generation here its referenced count is capped= at + * LRU_REFS_PROTECTED, so it starts at or below the protected tier regardl= ess + * of its old-generation access history. PID protection (folio_inc_gen) c= aps + * at LRU_REFS_WORKINGSET independently. */ -static bool lru_gen_set_refs(struct folio *folio, const vma_flags_t *vma_f= lags) -{ - /* see the comment on LRU_REFS_FLAGS */ - if (!folio_test_referenced(folio) && !folio_test_workingset(folio)) { - /* Activate file-backed executable folios after first usage. */ - if (is_exec_file_folio(folio, vma_flags)) { - folio_set_lru_refs(folio, 0); - folio_set_workingset(folio); - return true; + +/* + * Update the folio's lru refs indicator without taking the folio lock, + * isolation, or lruvec lock. Used by both page table access (@is_fault=3D= true) + * and by file access (@is_fault=3Dfalse). + */ +int folio_inc_lru_refs(struct folio *folio, bool is_fault, bool is_exec) +{ + int max_gen, min_gen; + int type, refs, gen, new_gen; + unsigned long new_flags, old_flags, max_seq; + struct lru_gen_folio *lrugen; + struct lruvec *lruvec; + + type =3D folio_is_file_lru(folio); + lruvec =3D folio_lruvec_live_get(folio); + lrugen =3D &lruvec->lrugen; + + old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + do { + new_flags =3D old_flags; + gen =3D lru_gen_from_flags(old_flags); + refs =3D lru_refs_from_flags(old_flags) + 1; + new_gen =3D gen; + if (!(old_flags & BIT(PG_lru)) || gen < 0) + goto out; + + max_seq =3D READ_ONCE(lrugen->max_seq); + max_gen =3D lru_gen_from_seq(max_seq); + min_gen =3D lru_gen_from_seq(READ_ONCE(lrugen->min_seq[type])); + if (gen =3D=3D max_gen) + goto out; + + if (is_fault || is_exec) { + /* Promote second page table access or executable */ + if (refs > LRU_REFS_REFERENCED || is_exec) + new_gen =3D max_gen; + else + new_gen =3D (gen + 1UL) % MAX_NR_GENS; + refs =3D min(refs, LRU_REFS_PROTECTED); + } else if (refs > LRU_REFS_MAX) { + /* LRU refs counting overflow, bump the gen */ + new_gen =3D (gen + 1UL) % MAX_NR_GENS; + refs =3D LRU_REFS_PROTECTED; + } else if (gen =3D=3D min_gen && refs >=3D LRU_REFS_WORKINGSET) { + /* Defer eviction of just accessed workingset */ + new_gen =3D (gen + 1UL) % MAX_NR_GENS; + refs =3D min(refs, LRU_REFS_PROTECTED); } +out: + refs =3D min(refs, LRU_REFS_MAX); + lru_refs_set_flags(&new_flags, refs); + if (new_gen >=3D 0) + lru_gen_set_flags(&new_flags, new_gen); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); =20 - folio_set_lru_refs(folio, 1); - return false; + if (new_gen !=3D gen) { + /* + * Gen can only go forward, so concurrent aging is + * usually fine, except when multiple aging increase + * max_seq multiple times, new_gen may have go beyond + * the new max_seq's current gen border and causes + * hotness inversion. In that very unlikely case, + * just activate the folio. + */ + lru_gen_update_size(lruvec, folio, gen, new_gen); + if (unlikely(READ_ONCE(lrugen->max_seq) - max_seq > MIN_NR_GENS)) + folio_activate(folio); } =20 - /* Promote on second access */ - if (folio_lru_refs(folio) > 1) { - folio_set_lru_refs(folio, 0); - folio_set_workingset(folio); - } else { - folio_mark_accessed(folio); - } - return true; + folio_lruvec_live_put(lruvec); + return refs; +} + +/* + * Update the folio's lru refs indicator during a page table walk. + * max_seq is stable since this runs inside the aging process. + * + * Returns the old generation and stores the new generation in @new_gen wh= en + * the folio is promoted (to max_gen) or advanced by one generation. + * Returns -1 if no gen change occurred. + */ +static int folio_inc_lru_refs_walk(struct folio *folio, struct lruvec *lru= vec, + const vma_flags_t *vma_flags, int *new_gen) +{ + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + unsigned long max_seq =3D READ_ONCE(lruvec->lrugen.max_seq); + int refs, gen, max_gen, ret; + + max_gen =3D lru_gen_from_seq(max_seq); + + do { + gen =3D lru_gen_from_flags(old_flags); + refs =3D lru_refs_from_flags(old_flags) + 1; + new_flags =3D old_flags; + + if (gen >=3D 0 && gen !=3D max_gen) { + ret =3D gen; + /* Promote second page table access or executable */ + if (refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags)) + *new_gen =3D max_gen; + else + *new_gen =3D (gen + 1) % MAX_NR_GENS; + lru_gen_set_flags(&new_flags, *new_gen); + lru_refs_set_flags(&new_flags, min(refs, LRU_REFS_PROTECTED)); + } else { + ret =3D -1; + lru_refs_set_flags(&new_flags, min(refs, LRU_REFS_MAX)); + } + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); + + return ret; +} + +/* + * Update the folio's lru refs indicator while the folio is isolated. + * Only used on mapped folios upon the final eviction, when the folio is + * off the LRU list (isolated). + * + * Increments the refs count (capped at LRU_REFS_PROTECTED). Returns true + * if the caller should activate the folio (second access or + * executable), false to keep it in the eviction list. + */ +static bool folio_inc_lru_refs_isolated(struct folio *folio, const vma_fla= gs_t *vma_flags) +{ + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + int refs; + + do { + new_flags =3D old_flags; + refs =3D lru_refs_from_flags(old_flags) + 1; + lru_refs_set_flags(&new_flags, min(refs, LRU_REFS_PROTECTED)); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); + + /* Promote second page table access or executable */ + return refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags); } #else -static bool lru_gen_set_refs(struct folio *folio, const vma_flags_t *vma_f= lags) +static bool folio_inc_lru_refs_isolated(struct folio *folio, const vma_fla= gs_t *vma_flags) { return false; } @@ -896,7 +1035,8 @@ static enum folio_references folio_check_references(st= ruct folio *folio, if (!referenced_ptes) return FOLIOREF_RECLAIM; =20 - return lru_gen_set_refs(folio, &vma_flags) ? FOLIOREF_ACTIVATE : FOLIORE= F_KEEP; + return folio_inc_lru_refs_isolated(folio, &vma_flags) ? + FOLIOREF_ACTIVATE : FOLIOREF_KEEP; } =20 referenced_folio =3D folio_test_clear_referenced(folio); @@ -3262,59 +3402,31 @@ static bool positive_ctrl_err(struct ctrl_pos *sp, = struct ctrl_pos *pv) * the aging *************************************************************************= *****/ =20 -/* promote pages accessed through page tables */ -static int folio_update_gen(struct folio *folio, int new_gen, const vma_fl= ags_t *vma_flags) -{ - unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); - int old_gen; - - /* - * See the comment on LRU_REFS_FLAGS, and activate file-backed - * executable folios after first usage to avoid typical IO - * thrashing from reclaiming. - */ - if (!folio_test_referenced(folio) && !folio_test_workingset(folio) && - !is_exec_file_folio(folio, vma_flags)) { - folio_set_lru_refs(folio, 1); - return -1; - } - - do { - old_gen =3D lru_gen_from_flags(old_flags); - new_flags =3D old_flags; - - /* lru_gen_del_folio() has isolated this page? */ - if (old_gen < 0) - break; - - lru_gen_set_flags(&new_flags, new_gen); - lru_refs_set_flags(&new_flags, 0); - new_flags |=3D BIT(PG_workingset); - } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); - - return old_gen; -} - -/* protect pages accessed multiple times through file descriptors */ +/* + * Force bump a folio's generation. Used for PID protection or defer the + * eviction of temporarily unevictable folio. + */ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio) { + int refs; int type =3D folio_is_file_lru(folio); struct lru_gen_folio *lrugen =3D &lruvec->lrugen; int old_gen, new_gen, min_gen =3D lru_gen_from_seq(lrugen->min_seq[type]); unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); =20 do { + new_flags =3D old_flags; + refs =3D lru_refs_from_flags(old_flags); old_gen =3D lru_gen_from_flags(old_flags); VM_WARN_ON_ONCE_FOLIO(old_gen < 0, folio); =20 - /* folio_update_gen() has promoted this page? */ + /* folio has been promoted? */ if (old_gen >=3D 0 && old_gen !=3D min_gen) return old_gen; =20 - new_flags =3D old_flags; new_gen =3D (old_gen + 1) % MAX_NR_GENS; lru_gen_set_flags(&new_flags, new_gen); - lru_refs_set_flags(&new_flags, 0); + lru_refs_set_flags(&new_flags, min(refs, LRU_REFS_WORKINGSET)); } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); =20 lru_gen_update_size(lruvec, folio, old_gen, new_gen); @@ -3518,21 +3630,17 @@ static void walk_update_folio(struct lru_gen_mm_wal= k *walk, struct vm_area_struc if (!folio) return; =20 - new_gen =3D lru_gen_from_seq(READ_ONCE(lruvec->lrugen.max_seq)); - if (dirty && !folio_test_dirty(folio) && !(folio_test_anon(folio) && folio_test_swapbacked(folio) && !folio_test_swapcache(folio))) folio_mark_dirty(folio); =20 if (walk) { - old_gen =3D folio_update_gen(folio, new_gen, &vma->flags); - if (old_gen >=3D 0 && old_gen !=3D new_gen) + old_gen =3D folio_inc_lru_refs_walk(folio, lruvec, &vma->flags, &new_gen= ); + if (old_gen >=3D 0) update_batch_size(walk, folio, old_gen, new_gen); - } else if (lru_gen_set_refs(folio, &vma->flags)) { - old_gen =3D folio_lru_gen(folio); - if (old_gen >=3D 0 && old_gen !=3D new_gen) - folio_activate(folio); + } else { + folio_inc_lru_refs(folio, true, is_exec_file_folio(folio, &vma->flags)); } } =20 @@ -3917,7 +4025,8 @@ static bool inc_min_seq(struct lruvec *lruvec, int ty= pe, int swappiness) while (!list_empty(head)) { struct folio *folio =3D lru_to_folio(head); int refs =3D folio_lru_refs(folio); - bool workingset =3D folio_test_workingset(folio); + int delta =3D folio_nr_pages(folio); + int tier =3D lru_tier_from_refs(refs); =20 VM_WARN_ON_ONCE_FOLIO(folio_test_unevictable(folio), folio); VM_WARN_ON_ONCE_FOLIO(folio_test_active(folio), folio); @@ -3927,14 +4036,8 @@ static bool inc_min_seq(struct lruvec *lruvec, int t= ype, int swappiness) new_gen =3D folio_inc_gen(lruvec, folio); list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]); =20 - /* don't count the workingset being lazily promoted */ - if (refs + workingset !=3D BIT(LRU_REFS_WIDTH) + 1) { - int tier =3D lru_tier_from_refs(refs, workingset); - int delta =3D folio_nr_pages(folio); - - WRITE_ONCE(lrugen->protected[hist][type][tier], - lrugen->protected[hist][type][tier] + delta); - } + WRITE_ONCE(lrugen->protected[hist][type][tier], + lrugen->protected[hist][type][tier] + delta); =20 if (!--remaining) return false; @@ -4649,8 +4752,7 @@ static bool sort_folio(struct lruvec *lruvec, struct = folio *folio, struct scan_c int zone =3D folio_zonenum(folio); int delta =3D folio_nr_pages(folio); int refs =3D folio_lru_refs(folio); - bool workingset =3D folio_test_workingset(folio); - int tier =3D lru_tier_from_refs(refs, workingset); + int tier =3D lru_tier_from_refs(refs); struct lru_gen_folio *lrugen =3D &lruvec->lrugen; =20 VM_WARN_ON_ONCE_FOLIO(gen >=3D MAX_NR_GENS, folio); @@ -4672,17 +4774,15 @@ static bool sort_folio(struct lruvec *lruvec, struc= t folio *folio, struct scan_c } =20 /* protected */ - if (tier > tier_idx || refs + workingset =3D=3D BIT(LRU_REFS_WIDTH) + 1) { + if (tier > tier_idx) { + int hist =3D lru_hist_from_seq(lrugen->min_seq[type]); + gen =3D folio_inc_gen(lruvec, folio); list_move(&folio->lru, &lrugen->folios[gen][type][zone]); =20 - /* don't count the workingset being lazily promoted */ - if (refs + workingset !=3D BIT(LRU_REFS_WIDTH) + 1) { - int hist =3D lru_hist_from_seq(lrugen->min_seq[type]); + WRITE_ONCE(lrugen->protected[hist][type][tier], + lrugen->protected[hist][type][tier] + delta); =20 - WRITE_ONCE(lrugen->protected[hist][type][tier], - lrugen->protected[hist][type][tier] + delta); - } return true; } =20 @@ -4710,10 +4810,6 @@ static bool isolate_folio(struct lruvec *lruvec, str= uct folio *folio, struct sca return false; } =20 - /* see the comment on LRU_REFS_FLAGS */ - if (!folio_test_referenced(folio)) - folio_set_lru_refs(folio, 0); - success =3D lru_gen_del_folio(lruvec, folio, true); VM_WARN_ON_ONCE_FOLIO(!success, folio); =20 @@ -4801,13 +4897,13 @@ static int get_tier_idx(struct lruvec *lruvec, int = type) struct ctrl_pos sp, pv =3D {}; =20 /* - * To leave a margin for fluctuations, use a larger gain factor (2:3). + * To leave a margin for fluctuations, use a larger gain factor (1:2). * This value is chosen because any other tier would have at least twice * as many refaults as the first tier. */ - read_ctrl_pos(lruvec, type, 0, 1, 2, &sp); for (tier =3D 1; tier < MAX_NR_TIERS; tier++) { - read_ctrl_pos(lruvec, type, tier, tier + 1, 3, &pv); + read_ctrl_pos(lruvec, type, 0, tier, 1, &sp); + read_ctrl_pos(lruvec, type, tier, tier + 1, 2, &pv); if (!positive_ctrl_err(&sp, &pv)) break; } @@ -4930,10 +5026,8 @@ static int evict_folios(unsigned long nr_to_scan, st= ruct lruvec *lruvec, } =20 /* don't add rejected folios to the oldest generation */ - if (lru_gen_folio_seq(lruvec, folio, false) =3D=3D min_seq[type]) { - folio_set_lru_refs(folio, 0); + if (lru_gen_folio_seq(lruvec, folio, false) =3D=3D min_seq[type]) folio_set_active(folio); - } } =20 move_folios_to_lru(&list); diff --git a/mm/workingset.c b/mm/workingset.c index 5438e9390011..452fe8554990 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -189,6 +189,13 @@ #define EVICTION_MASK (~0UL >> EVICTION_SHIFT) #define EVICTION_MASK_ANON (~0UL >> EVICTION_SHIFT_ANON) =20 +/* + * LRU refs uses LRU_REFS_WIDTH + 2 bits, the 2 bits are PG_workingset and + * PG_referenced. But here we record PG_workingset separately (to reuse + * pack_shadow). + */ +#define LRU_REFS_BITS ((LRU_REFS_WIDTH + 2) - 1) + /* * Eviction timestamps need to be able to cover the full range of * actionable refaults. However, bits are tight in the xarray @@ -242,13 +249,12 @@ static void *lru_gen_eviction(struct folio *folio) int type =3D folio_is_file_lru(folio); int delta =3D folio_nr_pages(folio); int refs =3D folio_lru_refs(folio); - bool workingset =3D folio_test_workingset(folio); - int tier =3D lru_tier_from_refs(refs, workingset); + int tier =3D lru_tier_from_refs(refs); struct mem_cgroup *memcg; struct pglist_data *pgdat =3D folio_pgdat(folio); unsigned short memcg_id; =20 - BUILD_BUG_ON(LRU_GEN_WIDTH + LRU_REFS_WIDTH > + BUILD_BUG_ON(LRU_GEN_WIDTH + LRU_REFS_BITS > BITS_PER_LONG - max(EVICTION_SHIFT, EVICTION_SHIFT_ANON)); =20 rcu_read_lock(); @@ -256,14 +262,14 @@ static void *lru_gen_eviction(struct folio *folio) lruvec =3D mem_cgroup_lruvec(memcg, pgdat); lrugen =3D &lruvec->lrugen; min_seq =3D READ_ONCE(lrugen->min_seq[type]); - token =3D (min_seq << LRU_REFS_WIDTH) | max(refs - 1, 0); + token =3D (min_seq << LRU_REFS_BITS) | refs >> 1; =20 hist =3D lru_hist_from_seq(min_seq); atomic_long_add(delta, &lrugen->evicted[hist][type][tier]); memcg_id =3D mem_cgroup_private_id(memcg); rcu_read_unlock(); =20 - return pack_shadow(memcg_id, pgdat, token, workingset, type); + return pack_shadow(memcg_id, pgdat, token, refs & 1, type); } =20 /* @@ -284,11 +290,24 @@ static bool lru_gen_test_recent(void *shadow, struct = lruvec **lruvec, *lruvec =3D mem_cgroup_lruvec(memcg, pgdat); =20 max_seq =3D READ_ONCE((*lruvec)->lrugen.max_seq); - max_seq &=3D (file ? EVICTION_MASK : EVICTION_MASK_ANON) >> LRU_REFS_WIDT= H; + max_seq &=3D (file ? EVICTION_MASK : EVICTION_MASK_ANON) >> LRU_REFS_BITS; =20 - return abs_diff(max_seq, *token >> LRU_REFS_WIDTH) < MAX_NR_GENS; + return abs_diff(max_seq, *token >> LRU_REFS_BITS) < MAX_NR_GENS; } =20 +/* + * Restore the refs of a refaulted folio from its shadow entry. + * + * Any folio that was accessed at least once before eviction (refs >=3D + * LRU_REFS_REFERENCED) is activated on a fault-driven refault, giving it a + * strong gen placement. Non-fault refaults (e.g. readahead) are not + * activated regardless of refs. + * + * The restored refs is capped at LRU_REFS_PROTECTED to prevent stale + * high-tier history from carrying over across eviction cycles. The + * WORKINGSET_RESTORE stat is bumped only for refs >=3D LRU_REFS_WORKINGSET + * to track genuine workingset restoration. + */ static void lru_gen_refault(struct folio *folio, void *shadow) { bool recent; @@ -314,21 +333,29 @@ static void lru_gen_refault(struct folio *folio, void= *shadow) lrugen =3D &lruvec->lrugen; =20 hist =3D lru_hist_from_seq(READ_ONCE(lrugen->min_seq[type])); - refs =3D (token & (BIT(LRU_REFS_WIDTH) - 1)) + 1; - tier =3D lru_tier_from_refs(refs, workingset); + refs =3D ((token & (BIT(LRU_REFS_BITS) - 1)) << 1) + workingset; + tier =3D lru_tier_from_refs(refs); =20 atomic_long_add(delta, &lrugen->refaulted[hist][type][tier]); =20 - if (workingset) { - /* Send refaulted workingset folios to active generations. */ + /* + * Activate a fault-driven refault: the folio was accessed at + * least once before eviction and would have been promoted had + * it stayed in memory. + */ + if (refs >=3D LRU_REFS_REFERENCED) { if (lru_gen_in_fault()) { folio_set_active(folio); mod_lruvec_state(lruvec, WORKINGSET_ACTIVATE_BASE + type, delta); } - folio_set_workingset(folio); + /* Cap restored refs to prevent stale high-tier carry-over */ + folio_set_lru_refs(folio, min(refs, LRU_REFS_PROTECTED)); + } + + /* WORKINGSET_RESTORE tracks genuine workingset-level refaults */ + if (refs >=3D LRU_REFS_WORKINGSET) mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + type, delta); - } else - set_mask_bits(&folio->flags.f, LRU_REFS_MASK, (refs - 1UL) << LRU_REFS_P= GOFF); + unlock: rcu_read_unlock(); } --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5310C3A7F48; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=DtfDY3/Pjlkzvff2y0f6qM1u//7xsLmQOjuj5/elb2G2F9rz8hZbPYMzALxta+9ekLc3WhOwoluLcVmv0ZQCWyXGYxnuagCGl0DevZ3ejXNsJfNNcQU9v/2BpWLdLzSWnkWTM9kSR2B3hGMN1aTuj4OEPL1rYU3+/r5KkQCN2RA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=5Spxi6AE7pqNpQLt2qla4NEl4WlMj9DJOAfIyYr9o9A=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=qqf9dfreYlcn6eloYg8iyGMsgM9qPdeLf6TfcHFHSEWqOxAchM1MDyVC20CPTmmyld8nWqQW7zq2tlSY6IYPcUQVUK2jdD2cHC8EdXIr+c9dYvRbVeHV0R6Nq5hD8+T+4KYO23GiQKkT9UFP+GXleRL0z8cT343BVusiDN3H4oU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=pJliByqC; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="pJliByqC" Received: by smtp.kernel.org (Postfix) with ESMTPS id 16D94C2BCFA; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=5Spxi6AE7pqNpQLt2qla4NEl4WlMj9DJOAfIyYr9o9A=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=pJliByqCy5ody7eU79GVRSPSk+eexK73f7T9s2SBA49hSB0o4C8OvJINOuOy/0Jst Zf5hdfZyM3BmdZVlhGa9zwvPZ1qDpxZajvv36UNiDD9+TQyCF1PF3iKIS1IRw58cEH S1jWRFCCHZPkg9acN8GoaJYbQsmHXUvnFfAvmigjHZ72t6HDzavzGxtQJdRuRutMuI rit3kvALB2TShQJNLsrks/iT+tefbNu+yb5c9/D8oiZCntRW+LBOYSasZJZWzBXP0R 2Wwz0shzNISUK1tDmHAqLUyuj0SooD4iiZRQSuup57dcIK3OXX/Bcpz2Hvd6o03mc+ f8U0t7cC5Gh4g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id F0079C55196; Mon, 3 Aug 2026 19:47:22 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:06 +0800 Subject: [PATCH RFC 10/15] mm/mglru: make folio lru referenced times count a generic API Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-10-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=9394; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=pavYdhuAU5OwgfCxwjbABkIxiu5hT6fbcrjm/x7ACsw=; b=d3wg6Yy52t36UmlNDlDQoCaeid1WRHhSSSH1k9fIA07ImY/PrYXUV4Ynfyh44gvYQPVityi4D GpazfrxHVV0APh8Tm1fpf0blBikgvUg/l5TmvuUwYEJSj4j9ENJVudp X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song To prepare for unifying the API for checking folio referenced status, expose the referenced times counting as a generic API. For MGLRU this helps to adapt other subsystem based on the referenced times counting, for non-MGLRU this is still bitwise compatible and there won't be major behavior change. Signed-off-by: Kairui Song --- include/linux/mm_inline.h | 233 ++++++++++++++++++++++++++++++------------= ---- mm/migrate.c | 2 - 2 files changed, 155 insertions(+), 80 deletions(-) diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 944baa91bf18..a13b7d3c033a 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -94,6 +94,161 @@ static __always_inline enum lru_list folio_lru_list(con= st struct folio *folio) return lru; } =20 +/** + * lru_refs_from_flags - Return LRU referenced / access count from folio f= lags. + * @flags: folio flags + */ +static inline int lru_refs_from_flags(unsigned long flags) +{ + int refs; + + /* + * Return the total number of accesses. Also see the comment on + * LRU_REFS_FLAGS. + */ + refs =3D (flags & BIT(PG_referenced)) ? BIT(0) : 0; + refs +=3D (flags & BIT(PG_workingset)) ? BIT(1) : 0; + refs +=3D ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) << 2; + return refs; +} + +/** + * lru_refs_set_flags - Set the LRU referenced / access count to specified= folio flags. + * @flags: pointer to the folio flags + * @refs: referenced / access count number, between 0 and LRU_REFS_MAX, in= clusive. + */ +static inline void lru_refs_set_flags(unsigned long *flags, unsigned int r= efs) +{ + VM_WARN_ON_ONCE(refs > LRU_REFS_MAX); + BUILD_BUG_ON((LRU_REFS_MAX >> 2) > (BIT(LRU_REFS_WIDTH) - 1)); + *flags &=3D ~LRU_REFS_FLAGS; + if (refs & BIT(0)) + *flags |=3D BIT(PG_referenced); + if (refs & BIT(1)) + *flags |=3D BIT(PG_workingset); + *flags |=3D (((unsigned long)refs) >> 2) << LRU_REFS_PGOFF; +} + +static inline int folio_lru_refs(const struct folio *folio) +{ + return lru_refs_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); +} + +static inline void folio_set_lru_refs(struct folio *folio, unsigned int re= fs) +{ + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + + do { + new_flags =3D old_flags; + lru_refs_set_flags(&new_flags, refs); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); +} + +int folio_inc_lru_refs(struct folio *folio, bool is_fault, bool is_exec); + +/** + * folio_is_referenced - Tell if a folio was accessed before. + * @folio: the folio. + * + * This helper currently only works as intended for MGLRU, as it checks + * all LRU_REFS_FLAGS. It might be fine for non-MGLRU to replace + * folio_test_referenced in some cases but the user should be careful. + * + * Returns: true if the folio's LRU referenced / accessed count > 0. + */ +static inline bool folio_is_referenced(const struct folio *folio) +{ + return folio_lru_refs(folio) >=3D LRU_REFS_REFERENCED; +} + +/** + * folio_mark_referenced - Mark a folio as referenced. + * @folio: the folio. + * + * Ensures the folio's LRU referenced count is at least + * LRU_REFS_REFERENCED. Won't do anything if the count is already larger + * than that. This helper currently only works as intended for MGLRU. + * Not a drop-in replacement, but should be fine for non-MGLRU to replace + * folio_set_referenced with this after audit. + */ +static inline void folio_mark_referenced(struct folio *folio) +{ + unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); + + do { + new_flags =3D old_flags; + if (lru_refs_from_flags(new_flags) >=3D LRU_REFS_REFERENCED) + return; + lru_refs_set_flags(&new_flags, LRU_REFS_REFERENCED); + } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); +} + +/** + * folio_mark_referenced_by_bit - Mark a folio as referenced by bit. + * @folio: the folio. + * + * non-MGLRU may want to make use of the lowest LRU referenced count bit + * explicitly as a referenced mark. + */ +static inline void folio_mark_referenced_by_bit(struct folio *folio) +{ + set_bit(PG_referenced, folio_flags(folio, 0)); +} + +/** + * folio_clear_referenced_by_bit - Clear the referenced bit of a folio. + * @folio: the folio. + */ +static inline void folio_clear_referenced_by_bit(struct folio *folio) +{ + clear_bit(PG_referenced, folio_flags(folio, 0)); +} + +/** + * folio_test_clear_referenced_by_bit - Test and clear the referenced bit + * @folio: the folio. + */ +static inline bool folio_test_clear_referenced_by_bit(struct folio *folio) +{ + return test_and_clear_bit(PG_referenced, folio_flags(folio, 0)); +} + +/** + * folio_is_referenced_by_bit - Test if the referenced bit of a folio is s= et. + * @folio: the folio. + */ +static inline bool folio_is_referenced_by_bit(const struct folio *folio) +{ + return test_bit(PG_referenced, const_folio_flags(folio, 0)); +} + +/** + * folio_is_workingset - Tell if a folio is part of the workingset. + * @folio: the folio. + * + * Can be used to replace folio_test_workingset safely. For MGLRU the LRU + * referenced count tells if a folio is a workingset as intended. For non-= MGLRU, + * the check below only holds true if the PG_workingset bit is set. + */ +static inline bool folio_is_workingset(const struct folio *folio) +{ + return folio_lru_refs(folio) >=3D LRU_REFS_WORKINGSET; +} + +/** + * folio_mark_workingset_by_bit - Set the workingset bit of a folio. + * @folio: the folio. + */ +static inline void folio_mark_workingset_by_bit(struct folio *folio) +{ + set_bit(PG_workingset, folio_flags(folio, 0)); +} + +static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) +{ + folio_set_lru_refs(new, folio_lru_refs(old)); +} + #ifdef CONFIG_LRU_GEN =20 static inline bool lru_gen_switching(void) @@ -171,58 +326,6 @@ static inline void lru_gen_set_flags(unsigned long *fl= ags, int gen) *flags |=3D (gen + 1UL) << LRU_GEN_PGOFF; } =20 -/** - * lru_refs_from_flags - Return LRU referenced / access count from folio f= lags. - * @flags: folio flags - */ -static inline int lru_refs_from_flags(unsigned long flags) -{ - int refs; - - /* - * Return the total number of accesses. Also see the comment on - * LRU_REFS_FLAGS. - */ - refs =3D (flags & BIT(PG_referenced)) ? BIT(0) : 0; - refs +=3D (flags & BIT(PG_workingset)) ? BIT(1) : 0; - refs +=3D ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) << 2; - return refs; -} - -/** - * lru_refs_set_flags - Set the LRU referenced / access count to specified= folio flags. - * @flags: pointer to the folio flags - * @refs: referenced / access count number, between 0 and LRU_REFS_MAX, in= clusive. - */ -static inline void lru_refs_set_flags(unsigned long *flags, unsigned int r= efs) -{ - VM_WARN_ON_ONCE(refs > LRU_REFS_MAX); - BUILD_BUG_ON((LRU_REFS_MAX >> 2) > (BIT(LRU_REFS_WIDTH) - 1)); - *flags &=3D ~LRU_REFS_FLAGS; - if (refs & BIT(0)) - *flags |=3D BIT(PG_referenced); - if (refs & BIT(1)) - *flags |=3D BIT(PG_workingset); - *flags |=3D (((unsigned long)refs) >> 2) << LRU_REFS_PGOFF; -} - -static inline int folio_lru_refs(const struct folio *folio) -{ - return lru_refs_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); -} - -static inline void folio_set_lru_refs(struct folio *folio, unsigned int re= fs) -{ - unsigned long new_flags, old_flags =3D READ_ONCE(*folio_flags(folio, 0)); - - do { - new_flags =3D old_flags; - lru_refs_set_flags(&new_flags, refs); - } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags)); -} - -int folio_inc_lru_refs(struct folio *folio, bool is_fault, bool is_exec); - static inline int folio_lru_gen(const struct folio *folio) { return lru_gen_from_flags(READ_ONCE(*const_folio_flags(folio, 0))); @@ -369,11 +472,6 @@ static inline bool lru_gen_del_folio(struct lruvec *lr= uvec, struct folio *folio, return true; } =20 -static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) -{ - folio_set_lru_refs(new, folio_lru_refs(old)); -} - #else /* !CONFIG_LRU_GEN */ =20 static inline bool lru_gen_enabled(void) @@ -401,27 +499,6 @@ static inline bool lru_gen_del_folio(struct lruvec *lr= uvec, struct folio *folio, return false; } =20 -static inline int folio_lru_refs(const struct folio *folio) -{ - return 0; -} - -static inline void folio_set_lru_refs(struct folio *folio, unsigned int re= fs) -{ -} - -static inline int folio_inc_lru_refs(struct folio *folio, bool promote, bo= ol is_exec) -{ - return 0; -} - -static inline void folio_migrate_refs(struct folio *new, const struct foli= o *old) -{ - if (folio_test_referenced(old)) - folio_set_referenced(new); - if (folio_test_workingset(old)) - folio_set_workingset(new); -} #endif /* CONFIG_LRU_GEN */ =20 static __always_inline diff --git a/mm/migrate.c b/mm/migrate.c index c737d0682fa4..806f1e913a38 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -786,8 +786,6 @@ void folio_migrate_flags(struct folio *newfolio, struct= folio *folio) folio_set_active(newfolio); } else if (folio_test_clear_unevictable(folio)) folio_set_unevictable(newfolio); - if (folio_test_workingset(folio)) - folio_set_workingset(newfolio); if (folio_test_checked(folio)) folio_set_checked(newfolio); /* --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4C66E3A7F45; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=ma/X/JB+NOmXS6LYGvDWDEjPznF2Hgr0RreLNFVm5bcOwsJivbNos2rHMw5rbL8mchh/DAjZ/TFPgFsyJA3dEw2P3r8g0NxPLYcSSRbMkkXGfRNz1haLWtMwWY02OHHaOvlxJWX3o6z0Gzz4s1W+8KVWsp6coF0swgzwcZhqwDs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=9OIzYG9v5ayIbesiCljJXQsRCZ3b0ppcAUlzxbinN9A=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=jpNHQcSnz0uT936F0FOcbX7n4VMwI6pP4EVsel+BktH9N8KPh36/8l2uiRoODNVzqCuEmYLCxTouXKoabE8pwU4WFuYmZPd/k6zlfs7JVfiNXU863A5yxx8RkmXOfbDuhTww58MULW4ogTczWYFcrjUTGFGZrplJ7LvVqRA8mhQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=uo+RiEy4; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="uo+RiEy4" Received: by smtp.kernel.org (Postfix) with ESMTPS id 271B3C2BD05; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=9OIzYG9v5ayIbesiCljJXQsRCZ3b0ppcAUlzxbinN9A=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=uo+RiEy4b/s/f7OV8igrBPb+6kmSRLVoYXNapHhJR9A0cHDGv9Mu1MZ+AAvGMPwN+ zWbMdBQHs7Wg/yLxiDO1btNE6L9Kp6ZTESKJEfMoYeLw9wbNhk5208pQBV4FZrbAIZ Um1R58QjeGm3VFfflOatSM+BZI/mcz8vw95b3SCCixa3s3YXilW8eNZmgpyuudRaWj j8ux8M0UoDDVtzvLKzywZeH4DvMyxfoPCXx01eiiWLeajQiR8/1vxA9Vk8tmkZ4azd AjxEK50uddf3Xfhalntmw6DTdC6pTw/kL4sbZ8GvDOx4l5MzL1W1upf7RIBPsCUUZM uWPecPaa2MXHA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 121B6C55843; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:07 +0800 Subject: [PATCH RFC 11/15] mm/mglru: replace folio workinset check and update with new helper Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-11-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=7823; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=ps3GPccJmNKKFg75e7bKrntkjE3H1Gau9WECQyN6J8c=; b=SsYlfhIoI9qTSRrpCosoUsyZvOU0ul0m78d3Qr/DFKMMXzkq44g4dt3l0QbyqaioklkvikIIo XQr9zDOideJDFLJAUTIr+IvQonmZ1XL5DQ4gdvvRSPbhot5/9ELDycb X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song With the new folio LRU refs tracking API, when MGLRU enabled, a folio is considered a workingset folio if its referenced count > 1. This is compatible with classical LRU and reasonable in many ways: The PG_referenced and PG_workingset are used as the lower bits of the LRU refs counter, and MGLRU will make use of extra bits as higher bits. So when MGLRU is disabled, all higher bits are always 0, making the check bit-wise equal to the old behavior. Active/inactive LRU sets PG_workingset explicitly for folios moved from active list to inactive list, and that makes the LRU refs tracking API (folio_is_workingset) report a referenced number > 1. Clearing PG_workingset will always return a value <=3D 1. When MGLRU is enabled, a folio referenced twice is considered a workingset folio, which is basically the same as how active/inactive LRU used to promote a file page to the active list. Note for active/inactive, the folio has to be marked inactive and PG_workingset before eviction, but MGLRU doesn't have a demotion process, so this simplified check ensures we have a stable definition and accurate readings in PSI and readaheads just like before. Signed-off-by: Kairui Song --- fs/btrfs/compression.c | 3 ++- mm/filemap.c | 8 ++++---- mm/madvise.c | 4 ++-- mm/page_io.c | 3 ++- mm/readahead.c | 8 ++++---- mm/vmscan.c | 2 +- mm/workingset.c | 4 ++-- 7 files changed, 17 insertions(+), 15 deletions(-) diff --git a/fs/btrfs/compression.c b/fs/btrfs/compression.c index ffb6b52863a7..e756403e8bd5 100644 --- a/fs/btrfs/compression.c +++ b/fs/btrfs/compression.c @@ -21,6 +21,7 @@ #include #include #include +#include #include "misc.h" #include "ctree.h" #include "fs.h" @@ -448,7 +449,7 @@ static noinline int add_ra_bio_folios(struct inode *ino= de, u64 compressed_end, continue; } =20 - if (!*memstall && folio_test_workingset(folio)) { + if (!*memstall && folio_is_workingset(folio)) { psi_memstall_enter(pflags); *memstall =3D 1; } diff --git a/mm/filemap.c b/mm/filemap.c index 6afec636881f..a88a6140ed09 100644 --- a/mm/filemap.c +++ b/mm/filemap.c @@ -1259,7 +1259,7 @@ static inline int folio_wait_bit_common(struct folio = *folio, int bit_nr, bool in_thrashing; =20 if (bit_nr =3D=3D PG_locked && - !folio_test_uptodate(folio) && folio_test_workingset(folio)) { + !folio_test_uptodate(folio) && folio_is_workingset(folio)) { delayacct_thrashing_start(&in_thrashing); psi_memstall_enter(&pflags); thrashing =3D true; @@ -1414,7 +1414,7 @@ void softleaf_entry_wait_on_locked(softleaf_t entry, = spinlock_t *ptl) struct folio *folio =3D softleaf_to_folio(entry); =20 q =3D folio_waitqueue(folio); - if (!folio_test_uptodate(folio) && folio_test_workingset(folio)) { + if (!folio_test_uptodate(folio) && folio_is_workingset(folio)) { delayacct_thrashing_start(&in_thrashing); psi_memstall_enter(&pflags); thrashing =3D true; @@ -2510,7 +2510,7 @@ static void filemap_get_read_batch(struct address_spa= ce *mapping, static int filemap_read_folio(struct file *file, filler_t filler, struct folio *folio) { - bool workingset =3D folio_test_workingset(folio); + bool workingset =3D folio_is_workingset(folio); unsigned long pflags; int error; =20 @@ -3981,7 +3981,7 @@ vm_fault_t filemap_map_pages(struct vm_fault *vmf, */ if ((map_ret & VM_FAULT_NOPAGE) && !(vmf->flags & FAULT_FLAG_TRIED) && - !folio_test_workingset(folio) && + !folio_is_workingset(folio) && !(vma->vm_flags & (VM_SEQ_READ | VM_EXEC))) { unsigned short mmap_miss; =20 diff --git a/mm/madvise.c b/mm/madvise.c index 07a21ca31bad..abb17760b8b5 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -427,7 +427,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd, folio_clear_referenced(folio); folio_test_clear_young(folio); if (folio_test_active(folio)) - folio_set_workingset(folio); + folio_mark_workingset_by_bit(folio); if (pageout) { if (folio_isolate_lru(folio)) { if (folio_test_unevictable(folio)) @@ -542,7 +542,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd, folio_clear_referenced(folio); folio_test_clear_young(folio); if (folio_test_active(folio)) - folio_set_workingset(folio); + folio_mark_workingset_by_bit(folio); if (pageout) { if (folio_isolate_lru(folio)) { if (folio_test_unevictable(folio)) diff --git a/mm/page_io.c b/mm/page_io.c index e4fa7ffffe8b..e7efc5bff668 100644 --- a/mm/page_io.c +++ b/mm/page_io.c @@ -25,6 +25,7 @@ #include #include #include +#include #include "swap.h" #include "swap_table.h" =20 @@ -452,7 +453,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct fo= lio *folio) { struct swap_info_struct *sis =3D __swap_entry_to_info(folio->swap); bool synchronous =3D sis->flags & SWP_SYNCHRONOUS_IO; - bool workingset =3D folio_test_workingset(folio); + bool workingset =3D folio_is_workingset(folio); unsigned long pflags; bool in_thrashing; =20 diff --git a/mm/readahead.c b/mm/readahead.c index 558c92957518..3ab796af6490 100644 --- a/mm/readahead.c +++ b/mm/readahead.c @@ -302,7 +302,7 @@ void page_cache_ra_unbounded(struct readahead_control *= ractl, } if (i =3D=3D mark) folio_set_readahead(folio); - ractl->_workingset |=3D folio_test_workingset(folio); + ractl->_workingset |=3D folio_is_workingset(folio); ractl->_nr_pages +=3D min_nrpages; i +=3D min_nrpages; } @@ -474,7 +474,7 @@ static inline int ra_alloc_folio(struct readahead_contr= ol *ractl, pgoff_t index, } =20 ractl->_nr_pages +=3D 1UL << order; - ractl->_workingset |=3D folio_test_workingset(folio); + ractl->_workingset |=3D folio_is_workingset(folio); return 0; } =20 @@ -817,7 +817,7 @@ void readahead_expand(struct readahead_control *ractl, folio_put(folio); return; } - if (unlikely(folio_test_workingset(folio)) && + if (unlikely(folio_is_workingset(folio)) && !ractl->_workingset) { ractl->_workingset =3D true; psi_memstall_enter(&ractl->_pflags); @@ -846,7 +846,7 @@ void readahead_expand(struct readahead_control *ractl, folio_put(folio); return; } - if (unlikely(folio_test_workingset(folio)) && + if (unlikely(folio_is_workingset(folio)) && !ractl->_workingset) { ractl->_workingset =3D true; psi_memstall_enter(&ractl->_pflags); diff --git a/mm/vmscan.c b/mm/vmscan.c index 9c8d9e3af375..913e69eae534 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -2268,7 +2268,7 @@ static void shrink_active_list(unsigned long nr_to_sc= an, } =20 folio_clear_active(folio); /* we are de-activating */ - folio_set_workingset(folio); + folio_mark_workingset_by_bit(folio); list_add(&folio->lru, &l_inactive); } =20 diff --git a/mm/workingset.c b/mm/workingset.c index 452fe8554990..568a3e44cd5e 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -438,7 +438,7 @@ void *workingset_eviction(struct folio *folio, struct m= em_cgroup *target_memcg) eviction >>=3D bucket_order[file]; workingset_age_nonresident(lruvec, folio_nr_pages(folio)); return pack_shadow(memcgid, pgdat, eviction, - folio_test_workingset(folio), file); + folio_is_workingset(folio), file); } =20 /** @@ -609,7 +609,7 @@ void workingset_refault(struct folio *folio, void *shad= ow) =20 /* Folio was active prior to eviction */ if (workingset) { - folio_set_workingset(folio); + folio_mark_workingset_by_bit(folio); mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + file, nr); } out: --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5DA693A7F73; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=GQdqTlol6nRWc90obL2bobvz6YCZmRZ6I+AVGOenCp6YawMaxvAly6chHTa/SzpZc16nH204H6gCbO9sA3nXfbkJcHH+PPpnJUDppbQZ3dGgfWqmRn5PWMUdY8yFkomL0GzT8bcfxupv1ypsoOdpi0t0ws31IiQzlL2+i6f381Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=ozlP16v/HxlVKrcEOAOjQCQyb4sgtwD0X3QHzHntDkc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=sbWvVTNLwQOuM3F8BjB9WbMM5eap4/3QSZ+1fXtVlEBV1v16hneD0jU6289q8ua6XGu5HMKqSmAvHrtfaUlZ81xxf00n3uqP4No2FdJ2i4STo7QYIQ7YqazOCcS/jurMyXNaGm2SB9VGZrodli7SVw5mpbaqHEkpD1Rxnvd7S6s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=GlzGBOqy; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="GlzGBOqy" Received: by smtp.kernel.org (Postfix) with ESMTPS id 3A2B7C2BD01; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=ozlP16v/HxlVKrcEOAOjQCQyb4sgtwD0X3QHzHntDkc=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=GlzGBOqyDrbpUYrwIhtK2XJ3zvl+rkTFDs9PAIpQWtKeWSEXbz4Syzhc+hM9k4SHc 3MHiEprVgCIkXc5yiqHKgW0yb8YXDKugvRJ8hsVm1hwiI3hDjYGdKdC23zYhSK1uPV y55wgrUR3NM8Q/Ue8nH3GE1QyRBdmUJddSBV9GZsmcdklwrwLoVO0mWmrsAqu+FF5a vRpK7fMS/rOrRjB8vm0mnj0xZIuTvXlbCydOLJ2lm9bEj9L3v+mD5as1JeYgL3Xmhv 6XKgsb8jwSErBUj/R41/6gaiAR0OenavrSznfQ6Ep4b4gMcfwekSUvgRrL/koQOhNG hQ47cJDYQnGkg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 25F34C55184; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:08 +0800 Subject: [PATCH RFC 12/15] mm/smap: report workingset folios as referenced Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-12-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=2528; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=w6RlPQJYy1xwfLtM1LXlBT4HMO6wCueMStGvS21qkQ8=; b=cic1gW3I0SoP78woxinTEezcXi3rjJJlFplB1hZ3sbyv9AiX0gCSAui2nXMP3uCgUxCTP1R2d 5cvp+hpla5+BDyG2oQdBDsnw2Dj9Zc4FvGYKcB/svhLvwI4iYWV+la3 X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song For MGLRU, switch smap to use the folio refs count API so smap will report all folio with referenced count >=3D 1 as "Referenced". Current smap checking PG_referenced is causing folios to flick between referenced and not-reference status, because for both MGLRU and active/inactive LRU, PG_referenced may got cleared on second access. (Increase of LRU referenced times count for MGLRU, and movig to active list active/inactive all clears that bit). After this, we will have a more reliable and useful reading for MGLRU. Signed-off-by: Kairui Song --- fs/proc/task_mmu.c | 22 +++++++++++++++++++--- 1 file changed, 19 insertions(+), 3 deletions(-) diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c index 817e3e0f9194..c5c96c523291 100644 --- a/fs/proc/task_mmu.c +++ b/fs/proc/task_mmu.c @@ -944,6 +944,22 @@ static void smaps_page_accumulate(struct mem_size_stat= s *mss, } } =20 +static bool smap_check_folio_referenced(struct folio *folio) +{ + if (lru_gen_enabled()) + return folio_is_referenced(folio); + else + return folio_is_referenced_by_bit(folio); +} + +static void smap_clear_folio_referenced(struct folio *folio) +{ + if (lru_gen_enabled()) + folio_set_lru_refs(folio, 0); + else + folio_clear_referenced_by_bit(folio); +} + static void smaps_account(struct mem_size_stats *mss, struct page *page, bool compound, bool young, bool dirty, bool locked, bool present) @@ -970,7 +986,7 @@ static void smaps_account(struct mem_size_stats *mss, s= truct page *page, =20 mss->resident +=3D size; /* Accumulate the size in pages that have been accessed. */ - if (young || folio_test_young(folio) || folio_test_referenced(folio)) + if (young || folio_test_young(folio) || smap_check_folio_referenced(folio= )) mss->referenced +=3D size; =20 /* @@ -1791,7 +1807,7 @@ static int clear_refs_pte_range(pmd_t *pmd, unsigned = long addr, /* Clear accessed and referenced bits. */ pmdp_test_and_clear_young(vma, addr, pmd); folio_test_clear_young(folio); - folio_clear_referenced(folio); + smap_clear_folio_referenced(folio); out: spin_unlock(ptl); return 0; @@ -1820,7 +1836,7 @@ static int clear_refs_pte_range(pmd_t *pmd, unsigned = long addr, /* Clear accessed and referenced bits. */ ptep_test_and_clear_young(vma, addr, pte); folio_test_clear_young(folio); - folio_clear_referenced(folio); + smap_clear_folio_referenced(folio); } pte_unmap_unlock(pte - 1, ptl); cond_resched(); --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6C1873A873B; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=mMcat7syzUVeaBpkD/BD0u4RTQncw3Jloqd/hzDmATjsg4kD6wbizoJYtSu7tsv51FWdKS3Ik4VvQWL2gTYVSdHNb0byd0e78D2E8LhE7BunapIFTh2/4XrzdkS/pHIw2DR2WbncfkQ4Jk7RyXqm55+YLLXbOcj88gmUBLA+efA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=S5TdssiSZAp76+zKDFhJcytSk6RO5zBrHmYRHuO0F6E=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ZP3tJTFMSqb4b7xbKi/8uoJizNt9OR4ab3sb2hNN7Vxiuh2+9i0rE9UcBMed4Gm57vGfMs76n8CYJMccPH5MzVNQjsjasfhS/JcLkPCDXGGw9j2lMCvv6dZCo0LxIJMDb2WqrHY+NfbafUQdlENiRCQ8XChr6vNYxZ9uEfmyUmI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Af+zpRdy; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Af+zpRdy" Received: by smtp.kernel.org (Postfix) with ESMTPS id 4C583C2BCFC; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=S5TdssiSZAp76+zKDFhJcytSk6RO5zBrHmYRHuO0F6E=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Af+zpRdydIutQ6SAx43Q3inKaMriMXd5MoFrCXgchJohl7fa4v3MGhkt426MP4s0B aBQe3XG1uMO90L4AN6ite34KW+M8twVLkYW8/PbaLgDGs8OgAEMKZ+jiZz0Cww+3HB KBOkMBc4plqjFAfjNVtfZofSlj8RpyhR1nErY1+QTGRRMdvCR+91CsjcB0S6eerdIb 06tImMnTumDIpFJz1AW7oYqmRmY6K0RnADpsiUSMDdzuBua1JjiwSBZBSVXpHTUvt8 kHWGffKCuAs4CV3fMPesksrzTepX8fdMVmBlXJARn76QLo2jxpQ67U3YSNZ5o48rYC fNOzrxASBVfNQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 39035C55182; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:09 +0800 Subject: [PATCH RFC 13/15] mm/huge_memory: mark file folio as accessed more accurately on split Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-13-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=1826; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=CJ94TqcOoIUBLTvAe1Ubz/gZmv43vqcSsxqrYElZJYw=; b=VAOrM0iKC69AZ7qWggVe+gnW+wnbWJ/uxpCHJjI4Gjx+GlW4HXfap1IE5/CsinA+y3hXAUa+7 vT7U8ic/6W3CGO/AEQt2ELX54Q2EI94FR/sfhb7iinRxP5+wthKZiST X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song The behavior of updating the folio's access info isn't consistent for huge mapping splitting or ordinary unmapping. The page table's young flag has to be translated into folio's access info. Right now it only check and set folio's referenced flag, which isn't enough since folio flags update on access have its rules. Ordinary unmapping (zapping) calls folio_mark_accessed(), and it also checks if the VMA has recency to avoid false updates. So first just use the right helper here to be more consistent. Signed-off-by: Kairui Song --- mm/huge_memory.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 21c92ee48e46..043c9ac963b4 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3058,8 +3058,8 @@ static void __split_huge_pud_locked(struct vm_area_st= ruct *vma, pud_t *pud, =20 if (!folio_test_dirty(folio) && pud_dirty(old_pud)) folio_mark_dirty(folio); - if (!folio_test_referenced(folio) && pud_young(old_pud)) - folio_set_referenced(folio); + if (pud_young(old_pud) && vma_has_recency(vma)) + folio_mark_accessed(folio); folio_remove_rmap_pud(folio, page, vma); add_mm_counter(vma->vm_mm, mm_counter_file(folio), -HPAGE_PUD_NR); @@ -3181,8 +3181,8 @@ static void __split_huge_pmd_locked(struct vm_area_st= ruct *vma, pmd_t *pmd, folio =3D page_folio(page); if (!folio_test_dirty(folio) && pmd_dirty(old_pmd)) folio_mark_dirty(folio); - if (!folio_test_referenced(folio) && pmd_young(old_pmd)) - folio_set_referenced(folio); + if (pmd_young(old_pmd) && vma_has_recency(vma)) + folio_mark_accessed(folio); folio_remove_rmap_pmd(folio, page, vma); add_mm_counter(mm, mm_counter_file(folio), -HPAGE_PMD_NR); folio_put(folio); --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 840373A9605; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=DzrzklzSx5UGROq7XsqTuwZv0p1d28eM/f6kc7txgjpfSy7E0J2ipFxWg3pLFPNFDw6joru5b3ZHS4s6md1L7J0GHdi+ghLx4GARN6We0ojmNXdt+kXIT+fjInE0W4+pdHIyh2a5E4pSfJI7SvcFWhQiZAyK9l9vzdWhyIKqD8g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=KnWmpT9oaTCGf6Iw2Wd4BLWFdHkuktfOvH/1O23SoQ4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Iu8Pw/1cPENUBGbtYWQieCKY3kvm3qEjokVzbKjPeo3YhGMMitcVRHOivEj00XGWSOqjIu+d3g0hN/sAYBIWtrShQXQXIrlMpnFpFWmDjhPv2kWup5KTjlgoE3oIIIFNJX8PbKn3SYnQHjvNbNjXsCn4b4+zg3kC7sP5MGyk/U4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Vzi4z2uQ; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Vzi4z2uQ" Received: by smtp.kernel.org (Postfix) with ESMTPS id 61217C2BD00; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=KnWmpT9oaTCGf6Iw2Wd4BLWFdHkuktfOvH/1O23SoQ4=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Vzi4z2uQBatDLSnAke9C7dGrFDH32+Nw8VbM64cqPyRHnSJVf4sNLTmTo+byjHoq5 g12QYA+zrgxjU0zoYbN6Qb6Fi4sPXGLMOEbUq+5jhHxsJUcnSkwyYCuQL5SdqC5Hhe 5Z4rKW3CuEoLFUq16XFgTOAB2zR8P4qPV8qnRAeItPxDh6DQj2lvnuWCvdjU/PM3DR +JKumL6fI8LU6o4U8A+sSv9+7GxY8MxBNCewE3p/ojT8q+3CBL2i0+lE3706up+xdC 91j1xPPybG2pmacfz4OUkPm/F1jKGvpuKNkCljTxfMgADdcCecFuEFoE7Dnvm02qlJ aBwDsOXWAn9cw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 4D1A9C55822; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:10 +0800 Subject: [PATCH RFC 14/15] mm/khugepaged: consider workingset folios as referenced Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-14-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=1957; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=HMlEqaXAqytxjGsOSQ1y0vS58P8dpXTZRRV+ITOk0K8=; b=hnHXcfL7XYuZF3ifhfyr250b99EXPZETH4Gh62F4nWF3Hz6MzoB9rwDrPS18A+Ns4mOMfBCXI K4dGqE5c7HxA0+8DZ2H8QLMV297MA/7uO9Suo8aKeXAAYU/OtiXsZ+E X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song The folio_test_referenced check here is clearly trying to test if the folio was ever referenced. It was first introduced by commit 8ee53820edfd ("thp: mmu_notifier_test_young") as an supplement of the young bit check. Folios are marked as PG_referenced on first access, but following access will clear their PG_referenced on second access. So checking only the referenced flag is not accurate enough. Switch to use the new helper, so we can cover the secondary and following access from MGLRU side. For non-MGLRU, this will make it return positve for workingset folios too though, which should be OK. Signed-off-by: Kairui Song --- mm/khugepaged.c | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/mm/khugepaged.c b/mm/khugepaged.c index b237f6e7662a..86c9b07dece6 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -809,7 +809,7 @@ static enum scan_result __collapse_huge_page_isolate(st= ruct vm_area_struct *vma, */ if (cc->is_khugepaged && (pte_young(pteval) || folio_test_young(folio) || - folio_test_referenced(folio) || + folio_is_referenced(folio) || mmu_notifier_test_young(vma->vm_mm, addr))) referenced++; } @@ -1767,7 +1767,7 @@ static enum scan_result collapse_scan_pmd(struct mm_s= truct *mm, */ if (cc->is_khugepaged && (pte_young(pteval) || folio_test_young(folio) || - folio_test_referenced(folio) || + folio_is_referenced(folio) || mmu_notifier_test_young(vma->vm_mm, addr))) referenced++; } @@ -2752,7 +2752,7 @@ static enum scan_result collapse_scan_file(struct mm_= struct *mm, /* * We probably should check if the folio is referenced * here, but nobody would transfer pte_young() to - * folio_test_referenced() for us. And rmap walk here + * folio_is_referenced() for us. And rmap walk here * is just too costly... */ =20 --=20 2.55.0 From nobody Fri Oct 2 07:46:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 931B63A963B; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; cv=none; b=dJUN7nuSdQ2AB2/+x8KRVwgIiM84qJUL9yfSExE0h8UrXH6NEMEp8eZYulHBTGxfAuvsYtRbNxlM8UODyWW5Uv8waKGgWJRQx5fDY9TecchFP+bWrJL4SkhAqNYazfTY8ZV9nLr3RrI98lpDjAHyHdAX+uiwU2qX0nf9SlquDtg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785786443; c=relaxed/simple; bh=1n+hcBYJ2VLCPzu8Xg2KTxM1AQOc1aeNpaYZxBPUvng=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Rrzeq3aKDJeB33M8gF5Rac/YCbzl7Hdmnb0UdE0l8w8JdXkQzblnCcg49H+aY2jmC7JLOkShiqvJLCiGb/9l/AHu8OX1SC4EdpFMxaZi9gwtkCeBMZ0uXk4thsjp7b+NU/L8Qme0E313l9fyFHsRYpZ3xZfpNDHwo/L1bS9Rz94= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=nRCymV0U; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="nRCymV0U" Received: by smtp.kernel.org (Postfix) with ESMTPS id 7248EC2BCF6; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785786443; bh=1n+hcBYJ2VLCPzu8Xg2KTxM1AQOc1aeNpaYZxBPUvng=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=nRCymV0UwDUe8D10YefBzAs8HYuwJOYHOY/54Em4EiZScroS/J3oWbwC7w3BVWN9w QmcSHQm009697qseGapcMPFvqMZLb1jWkBMLrJ6klAp/KUMSrSM+XYMmtMLRnRv+tN D8g0NABxxd7WwkHevK8tpSJG7hiWcBVSvOAHOaUzUTQvkf/yKBHs9VH40ro0kXnNpK w8oGXyVwoIlbrmMMg0DFz6kJeGqDBGbAlpUeUUnDtq2NI1IkIZZdCI1XaceUzOH5cX q4ZtF1eMgPLoyAsMXMBiuvavWyYgsuGymFOU9PuNjLrYrslmV4hKpJU2PUXf03kpGk RreVEaECiQ2Og== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 5F2BFC55196; Mon, 3 Aug 2026 19:47:23 +0000 (UTC) From: Kairui Song via B4 Relay Date: Tue, 04 Aug 2026 03:47:11 +0800 Subject: [PATCH RFC 15/15] mm/madvise: convert to new lru refs API and better support for MGLRU Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260804-mglru-fg-v1-15-4d8dad39dad6@tencent.com> References: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> In-Reply-To: <20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com> To: linux-mm@kvack.org Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Ying Huang , Chris Li , Baoquan He , Nico Pache , Usama Arif , Michal Hocko , Roman Gushchin , Shakeel Butt , David Hildenbrand , Lorenzo Stoakes , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Vlastimil Babka , Suren Baghdasaryan , Kemeng Shi , Nhat Pham , Youngjun Park , Zi Yan , Gregory Price , "Matthew Wilcox (Oracle)" , Baolin Wang , Ryan Roberts , Dev Jain , Lance Yang , Hugh Dickins , SeongJae Park , David Rientjes , Yu Zhao , Vernon Yang , Zicheng Wang , Chen Ridong , Tal Zussman , Kairui Song , linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Kairui Song , Baoquan He , Nico Pache X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785786439; l=2598; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=IZ9piP3HZQ3l6Pb7dyuhGVIV9UajVhIyv5fRhBDYZ6o=; b=Ilfa9kf1jZoKpnpzjRcdb54/uo49AYoBym1Mxkn0Z4Zm9OWDSdCMbgnMvIATbnuQPI2/EEr0d PVsDj5Gc3FuCRO92Q5+ThtGtYa277T/0VLQwzHhJJM+3TojEEjijN16 X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song For active/inactive LRU, madvise wants evicted folios from active LRU to be considered for PSI too, so some special handling are added. But MGLRU doesn't really need this, as it has a different activation logic. Switch to new helpers and improve the support for MGLRU here. Signed-off-by: Kairui Song --- mm/madvise.c | 37 +++++++++++++++++++++++-------------- 1 file changed, 23 insertions(+), 14 deletions(-) diff --git a/mm/madvise.c b/mm/madvise.c index abb17760b8b5..f132dd7418f5 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -350,6 +350,27 @@ static inline int madvise_folio_pte_batch(unsigned lon= g addr, unsigned long end, FPB_MERGE_YOUNG_DIRTY); } =20 +/* + * We are deactivating a folio for accelerating reclaiming. + * VM couldn't reclaim the folio unless we clear PG_young. + * As a side effect, it makes confuse idle-page tracking + * because they will miss recent referenced history. + */ +static void madvise_cold_or_pageout_prep_folio(struct folio *folio) +{ + folio_test_clear_young(folio); + + /* + * MGLRU clears all reference flags in folio_deactivate, + * no need to touch it here. + */ + if (!lru_gen_enabled()) { + folio_clear_referenced_by_bit(folio); + if (folio_test_active(folio)) + folio_mark_workingset_by_bit(folio); + } +} + static int madvise_cold_or_pageout_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end, struct mm_walk *walk) @@ -424,10 +445,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pm= d, tlb_remove_pmd_tlb_entry(tlb, pmd, addr); } =20 - folio_clear_referenced(folio); - folio_test_clear_young(folio); - if (folio_test_active(folio)) - folio_mark_workingset_by_bit(folio); + madvise_cold_or_pageout_prep_folio(folio); if (pageout) { if (folio_isolate_lru(folio)) { if (folio_test_unevictable(folio)) @@ -533,16 +551,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pm= d, tlb_remove_tlb_entries(tlb, pte, nr, addr); } =20 - /* - * We are deactivating a folio for accelerating reclaiming. - * VM couldn't reclaim the folio unless we clear PG_young. - * As a side effect, it makes confuse idle-page tracking - * because they will miss recent referenced history. - */ - folio_clear_referenced(folio); - folio_test_clear_young(folio); - if (folio_test_active(folio)) - folio_mark_workingset_by_bit(folio); + madvise_cold_or_pageout_prep_folio(folio); if (pageout) { if (folio_isolate_lru(folio)) { if (folio_test_unevictable(folio)) --=20 2.55.0