From nobody Tue Sep 29 02:34:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E0054374E5B; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; cv=none; b=WS6AMPh+ey8lmaAyp7OHheYb5Ptve1NV6Abqt9ggmEtgtUyHsTQizzpIo+emQ9oqRIuvfDXR4r5yvUYVs5Vcg7Mf9FvMJa18Mbf3ztI9mUDw3FYKCJ+37vpZQEDnRowRFWaaJziqk6HMxmwyOD+OtK9YAvBy0wdUg5dxBXPzM5A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; c=relaxed/simple; bh=/jY3lCvj/S6s/mtmnK+py0Ta55+PVDMUn3Xj23a2mm0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FIOPPmE/1tmLQyQszeshINuPQYovPpz/kV+eDu9P3I2BqQ7lV6kVoO4QkrklAm8asMIyZZUkH9b36IucQ1iWD6uYWBnhrTP6J+q5iKqGpGBtIkorBdTHerCluUz4TJGNs8AWDxNFUnEw93dWMChPUcLn7/T4OKnTERGJkWslRpo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=MG30tG0j; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="MG30tG0j" Received: by smtp.kernel.org (Postfix) with ESMTPS id 9EBC9C2BCF6; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786611132; bh=/jY3lCvj/S6s/mtmnK+py0Ta55+PVDMUn3Xj23a2mm0=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=MG30tG0jIAAQO2Ms23J181a/6qTNlDEYvNG5iUsZ7Qz0g/wihw24Na1LKrGju3ORc BEaLfFTEPlqQVhqZJLgm6kPMqXp7/rmT5K2fn86gfLsLSXsFR9bBDRHPM2Ne1RnfxV ImdPv5IN2lK2NY/lBp2mLFp1nBjEYNA1MYESY6Cb1GEvNr7Kags3r7XGOfv2Kw+IzV lbsFmj/ix9YtbOFDTVi9QPty6BwWEQbSXUnRN8/XSyi3FFj4p8z4aP+JZ6ustXBlOp 0qJYLaajY9wtQMyDcvKppGMwiD7SXh5w5XnrlCvlYwKCjt7yZDrS3B6rMuW6BfWTy7 3g7fhapR9qMSA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7C3DAC5AC67; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 13 Aug 2026 16:52:02 +0800 Subject: [PATCH RFC 1/5] memcg: move memcg private ID refcount to objcg Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-memcgid-objcg-v1-1-83d21c685b77@tencent.com> References: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> In-Reply-To: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , Dave Chinner , Qi Zheng , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu Cc: cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Bingfang Guo , Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786611130; l=8568; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=OoA5ZvY6U2Ek43M1DA077LOuBUDVK9W5rd1MBAq9D6g=; b=NwGka+2hgzQ2Y53VCMURN987ApMGmBKou7N9EagAAGcK2Q2VVI1BH6fu3bLRHKx9tvNOph6Lt eXCDryxHaP5CdWdcA+uqx7j5Po2BLAQ14+bFpf6rPcRc1GuJt1MoIGm X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo In the previous series by Muchun Song and Qi Zheng, folios are charged to the objcg and reparented as the memcg offlines. Same can be done to memcg private ID and its main user: swap entries. Make the memcgid xarray hold a pointer and a reference to an objcg of the memcg, which is used to find the memcg (or its parent) later on. The online state now pins the objcg instead of the css, so swapped out pages no longer pin the dying memcg. The id reference held by the online state is released in css_released() after reparenting instead of in css_offline(). This is the key invariant the rest of the series builds on: css_offline() runs while other css references may still be held, but css_released() only runs once the last reference is gone, so a caller holding a memcg reference can always count on the id refcount being alive. To prevent races between memcgid put in css offline and memcgid get, the release is put off till css_released, which could delay the release of the memcgid and the objcg it pins, but overall it should be fine. After reparenting, the objcg points to a live ancestor, so mem_cgroup_from_private_id() now returns that ancestor instead of the memcg the ID originally belonged to. Callers that need the exact memcg are fixed in patch 5. Signed-off-by: Bingfang Guo --- include/linux/memcontrol.h | 10 +++-- mm/memcontrol.c | 99 ++++++++++++++++++++++++++++++++----------= ---- 2 files changed, 76 insertions(+), 33 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 8170bb8066a22..c33ec7efad50b 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -191,6 +191,7 @@ struct obj_cgroup { struct rcu_head rcu; }; bool is_root; + refcount_t id_ref; }; =20 /* @@ -202,8 +203,8 @@ struct obj_cgroup { struct mem_cgroup { struct cgroup_subsys_state css; =20 - /* Private memcg ID. Used to ID objects that outlive the cgroup */ - struct mem_cgroup_private_id id; + /* The objcg holding private memcg ID. */ + struct obj_cgroup *id_objcg; =20 /* Accounted resources */ struct page_counter memory; /* Both v1 & v2 */ @@ -270,6 +271,9 @@ struct mem_cgroup { #endif int kmemcg_id; =20 + /* Private memcg ID. Used to ID objects that outlive the cgroup */ + int id; + struct memcg_vmstats_percpu __percpu *vmstats_percpu; =20 #ifdef CONFIG_CGROUP_WRITEBACK @@ -820,7 +824,7 @@ static inline unsigned short mem_cgroup_private_id(stru= ct mem_cgroup *memcg) if (mem_cgroup_disabled()) return 0; =20 - return memcg->id.id; + return memcg->id; } struct mem_cgroup *mem_cgroup_from_private_id(unsigned short id); =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 8319ad8c5c23a..5f30e76ee93d7 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -3697,7 +3697,7 @@ static void memcg_online_kmem(struct mem_cgroup *memc= g) =20 static_branch_enable(&memcg_kmem_online_key); =20 - memcg->kmemcg_id =3D memcg->id.id; + memcg->kmemcg_id =3D memcg->id; } =20 static void memcg_offline_kmem(struct mem_cgroup *memcg) @@ -3956,25 +3956,40 @@ static DEFINE_XARRAY_ALLOC1(mem_cgroup_private_ids); =20 static void mem_cgroup_private_id_remove(struct mem_cgroup *memcg) { - if (memcg->id.id > 0) { - xa_erase(&mem_cgroup_private_ids, memcg->id.id); - memcg->id.id =3D 0; + if (memcg->id > 0) { + xa_erase(&mem_cgroup_private_ids, memcg->id); + memcg->id =3D 0; } } =20 -static inline void mem_cgroup_private_id_put(struct mem_cgroup *memcg, uns= igned int n) +/** + * @objcg: the objcg returned by mem_cgroup_private_id_objcg + * @id: the corresponding memcg private id + */ +static void __mem_cgroup_private_id_put(struct obj_cgroup *objcg, + unsigned short id, unsigned int n) { - if (refcount_sub_and_test(n, &memcg->id.ref)) { - mem_cgroup_private_id_remove(memcg); + struct obj_cgroup *objcg_free; =20 - /* Memcg ID pins CSS */ - css_put(&memcg->css); + if (refcount_sub_and_test(n, &objcg->id_ref)) { + objcg_free =3D xa_erase(&mem_cgroup_private_ids, id); + VM_WARN_ON(objcg_free !=3D objcg); + + /* Memcg ID pins the objcg */ + obj_cgroup_put(objcg); } } =20 +static inline void mem_cgroup_private_id_put(struct mem_cgroup *memcg, uns= igned int n) +{ + __mem_cgroup_private_id_put(memcg->id_objcg, memcg->id, n); +} + struct mem_cgroup *mem_cgroup_private_id_get_online(struct mem_cgroup *mem= cg, unsigned int n) { - while (!refcount_add_not_zero(n, &memcg->id.ref)) { + struct obj_cgroup *objcg =3D memcg->id_objcg; + + while (!refcount_add_not_zero(n, &objcg->id_ref)) { /* * The root cgroup cannot be destroyed, so it's refcount must * always be >=3D 1. @@ -3984,6 +3999,7 @@ struct mem_cgroup *mem_cgroup_private_id_get_online(s= truct mem_cgroup *memcg, un break; } memcg =3D parent_mem_cgroup(memcg); + objcg =3D memcg->id_objcg; } return memcg; } @@ -3996,8 +4012,29 @@ struct mem_cgroup *mem_cgroup_private_id_get_online(= struct mem_cgroup *memcg, un */ struct mem_cgroup *mem_cgroup_from_private_id(unsigned short id) { + struct obj_cgroup *objcg; WARN_ON_ONCE(!rcu_read_lock_held()); - return xa_load(&mem_cgroup_private_ids, id); + + objcg =3D xa_load(&mem_cgroup_private_ids, id); + if (!objcg) + return NULL; + + return obj_cgroup_memcg(objcg); +} + +static struct mem_cgroup *mem_cgroup_take_from_private_id(unsigned short i= d, unsigned int n) +{ + struct obj_cgroup *objcg; + struct mem_cgroup *memcg; + + objcg =3D xa_load(&mem_cgroup_private_ids, id); + if (!objcg) + return NULL; + + memcg =3D get_mem_cgroup_from_objcg(objcg); + + __mem_cgroup_private_id_put(objcg, id, n); + return memcg; } =20 struct mem_cgroup *mem_cgroup_get_from_id(u64 id) @@ -4098,7 +4135,7 @@ static struct mem_cgroup *mem_cgroup_alloc(struct mem= _cgroup *parent) if (!memcg) return ERR_PTR(-ENOMEM); =20 - error =3D xa_alloc(&mem_cgroup_private_ids, &memcg->id.id, NULL, + error =3D xa_alloc(&mem_cgroup_private_ids, &memcg->id, NULL, XA_LIMIT(1, MEM_CGROUP_ID_MAX), GFP_KERNEL); if (error) goto fail; @@ -4243,9 +4280,10 @@ static int mem_cgroup_css_online(struct cgroup_subsy= s_state *css) FLUSH_TIME); lru_gen_online_memcg(memcg); =20 - /* Online state pins memcg ID, memcg ID pins CSS */ - refcount_set(&memcg->id.ref, 1); - css_get(css); + /* CSS pins memcg ID, memcg ID pins obj cgroup */ + memcg->id_objcg =3D memcg->nodeinfo[0]->objcg; + refcount_set(&memcg->id_objcg->id_ref, 1); + obj_cgroup_get(memcg->id_objcg); =20 /* * Ensure mem_cgroup_from_private_id() works once we're fully online. @@ -4257,7 +4295,7 @@ static int mem_cgroup_css_online(struct cgroup_subsys= _state *css) * publish it here at the end of onlining. This matches the * regular ID destruction during offlining. */ - xa_store(&mem_cgroup_private_ids, memcg->id.id, memcg, GFP_KERNEL); + xa_store(&mem_cgroup_private_ids, memcg->id, objcg, GFP_KERNEL); =20 return 0; free_objcg: @@ -4308,8 +4346,6 @@ static void mem_cgroup_css_offline(struct cgroup_subs= ys_state *css) lru_gen_offline_memcg(memcg); =20 drain_all_stock(memcg); - - mem_cgroup_private_id_put(memcg, 1); } =20 static void mem_cgroup_css_released(struct cgroup_subsys_state *css) @@ -4318,6 +4354,9 @@ static void mem_cgroup_css_released(struct cgroup_sub= sys_state *css) =20 invalidate_reclaim_iterators(memcg); lru_gen_release_memcg(memcg); + + mem_cgroup_private_id_put(memcg, 1); + memcg->id_objcg =3D NULL; } =20 static void mem_cgroup_css_free(struct cgroup_subsys_state *css) @@ -5651,19 +5690,19 @@ void __mem_cgroup_uncharge_swap(unsigned short id, = unsigned int nr_pages) { struct mem_cgroup *memcg; =20 - rcu_read_lock(); - memcg =3D mem_cgroup_from_private_id(id); - if (memcg) { - if (!mem_cgroup_is_root(memcg)) { - if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, nr_pages); - else - page_counter_uncharge(&memcg->swap, nr_pages); - } - mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages); - mem_cgroup_private_id_put(memcg, nr_pages); + memcg =3D mem_cgroup_take_from_private_id(id, nr_pages); + if (!memcg) + return; + + if (!mem_cgroup_is_root(memcg)) { + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, nr_pages); + else + page_counter_uncharge(&memcg->swap, nr_pages); } - rcu_read_unlock(); + mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages); + + mem_cgroup_put(memcg); } =20 long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg) --=20 2.43.7 From nobody Tue Sep 29 02:34:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DFF4B32B114; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; cv=none; b=drv8tCF5kpd2+4PwH4rINv+RwKMcUTB96KKvenjgjjtmYp9wMhyHJYAoQzxnV0IIyyaVrgftvUO1fygaOfqrbQ/kOCv7ZXW4TK65WRBeZ4+7trTMsysrXejyMspP8o32lnS7mHe+GhWnwdJDajO4iydKuMrDDmv3mXvzChBszS8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; c=relaxed/simple; bh=kR5zbkSvv/RlKYuZXiHgRuu0UFnNidkDfN2007OyReY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=JlZiHeRQq/6rBrwg29ENcG77LN5mRxIkvGvJ0T3yYl07DUNcpP5Q2NUQx2EYriqrhltsOD1nFJL3kGP7QScv8FBVcEM6lN2Edw7EXsDSgM4gVOvYL4wK4CPt5ltInaZ3+Vm39Sz/O29KZIuQ+UdEYB+rqYNBC9xTc2zIfQwNZKo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=RfP6jZ1g; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="RfP6jZ1g" Received: by smtp.kernel.org (Postfix) with ESMTPS id AB260C2BCF7; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786611132; bh=kR5zbkSvv/RlKYuZXiHgRuu0UFnNidkDfN2007OyReY=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=RfP6jZ1ggMa6trUcUGQtCnAYEDpNKuCrnzBMVI5PbQIOu9Xzi6sfL9dnF6vFsVEFu +Xy9acYBGaWqIXo9XIYkcJqrZK1Rn0USRIjDu3py131+38MIq/SpGi5Ob4ltsBRN4n MEg4iKW99OT+2Ne7qVFsqWsJOGwZdB0Jlfv5EN0oQHhk4Bl/o014To5HR3fKHRgNqT ELFNsRYpztYG+EdSGTs4Z1JPWzDErROInTyz4uOS7Yz0bIMjmwE2SYK0PAg46gSP1U M0ok0XuDrdfFvdgndo6P2fWpDEO39+cE5PfaQmmjxtl3eLFy9aZ9+6zEO2Src/KbbJ 3b5XjEjRsXSRw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8F204C5CFDB; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 13 Aug 2026 16:52:03 +0800 Subject: [PATCH RFC 2/5] memcg: get stable memcg first before getting memcgid reference Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-memcgid-objcg-v1-2-83d21c685b77@tencent.com> References: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> In-Reply-To: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , Dave Chinner , Qi Zheng , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu Cc: cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Bingfang Guo , Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786611130; l=4263; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=jO1EJ+YiSvB3LpF8HaBSpgO3Wgmvg81030gxkA7owJQ=; b=IuoSLuzL/85amN4CY0wJlbVtvGU0mcWEE5s5ZLz/r+pG0gu1bd9lArAGvkUhpLoBQdNeOaZQn /8mlXbtlBnPBY0MdmBu+DAQ0AO/qUgV3bOyPNMKrDjv8rLi+hCRWcs4 X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo Storing the memcg private ID in a swap entry used to take the ID reference under the RCU read lock and rely on mem_cgroup_private_id_get_online() to hand back a usable (possibly parent) memcg. Now that the ID refcount lives on the objcg and stays alive until css_released(), holding a memcg reference is enough to pin the ID. Both __memcg1_swapout() and __mem_cgroup_try_charge_swap() take a stable memcg reference first via get_mem_cgroup_from_objcg() and pin the ID afterwards, dropping the rcu_read_lock() usage and the get-error-put handling, and recording exactly the memcg the folio belongs to in the swap entry. Signed-off-by: Bingfang Guo --- mm/memcontrol-v1.c | 23 +++++++++-------------- mm/memcontrol.c | 9 ++++++--- 2 files changed, 15 insertions(+), 17 deletions(-) diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 2dc599484d006..a913d32ad1e17 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -618,7 +618,7 @@ void memcg1_commit_charge(struct folio *folio, struct m= em_cgroup *memcg) */ void __memcg1_swapout(struct folio *folio, struct swap_cluster_info *ci) { - struct mem_cgroup *memcg, *swap_memcg; + struct mem_cgroup *memcg; struct obj_cgroup *objcg; unsigned int nr_entries; =20 @@ -638,19 +638,20 @@ void __memcg1_swapout(struct folio *folio, struct swa= p_cluster_info *ci) if (!objcg) return; =20 - rcu_read_lock(); - memcg =3D obj_cgroup_memcg(objcg); /* * In case the memcg owning these pages has been offlined and doesn't * have an ID allocated to it anymore, charge the closest online - * ancestor for the swap instead and transfer the memory+swap charge. + * ancestor for the swap instead. */ + memcg =3D get_mem_cgroup_from_objcg(objcg); nr_entries =3D folio_nr_pages(folio); - swap_memcg =3D mem_cgroup_private_id_get_online(memcg, nr_entries); - mod_memcg_state(swap_memcg, MEMCG_SWAP, nr_entries); + mod_memcg_state(memcg, MEMCG_SWAP, nr_entries); + + /* we have a reference to it, so we should get exact memcg itself */ + mem_cgroup_private_id_get_online(memcg, nr_entries); =20 __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_entries, - mem_cgroup_private_id(swap_memcg)); + mem_cgroup_private_id(memcg)); =20 folio_unqueue_deferred_split(folio); folio->memcg_data =3D 0; @@ -658,12 +659,6 @@ void __memcg1_swapout(struct folio *folio, struct swap= _cluster_info *ci) if (!obj_cgroup_is_root(objcg)) page_counter_uncharge(&memcg->memory, nr_entries); =20 - if (memcg !=3D swap_memcg) { - if (!mem_cgroup_is_root(swap_memcg)) - page_counter_charge(&swap_memcg->memsw, nr_entries); - page_counter_uncharge(&memcg->memsw, nr_entries); - } - /* * The caller must hold the swap cluster lock with IRQ off. It is * important here to have the interrupts disabled because it is the @@ -675,7 +670,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_= cluster_info *ci) preempt_enable_nested(); memcg1_check_events(memcg, folio_nid(folio)); =20 - rcu_read_unlock(); + mem_cgroup_put(memcg); obj_cgroup_put(objcg); } =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 5f30e76ee93d7..a210fe2501219 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5660,24 +5660,27 @@ int __mem_cgroup_try_charge_swap(struct folio *foli= o) return 0; } =20 - memcg =3D mem_cgroup_private_id_get_online(memcg, nr_pages); - /* memcg is pined by memcg ID. */ + memcg =3D get_mem_cgroup_from_objcg(objcg); rcu_read_unlock(); =20 if (!mem_cgroup_is_root(memcg) && !page_counter_try_charge(&memcg->swap, nr_pages, &counter)) { memcg_memory_event(memcg, MEMCG_SWAP_MAX); memcg_memory_event(memcg, MEMCG_SWAP_FAIL); - mem_cgroup_private_id_put(memcg, nr_pages); + mem_cgroup_put(memcg); return -ENOMEM; } mod_memcg_state(memcg, MEMCG_SWAP, nr_pages); =20 + /* we have a reference to it, so we should get exact memcg itself */ + mem_cgroup_private_id_get_online(memcg, nr_pages); + ci =3D swap_cluster_get_and_lock(folio); __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages, mem_cgroup_private_id(memcg)); swap_cluster_unlock(ci); =20 + mem_cgroup_put(memcg); return 0; } =20 --=20 2.43.7 From nobody Tue Sep 29 02:34:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DFEA330E83F; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; cv=none; b=c026n7H/Jz3e4210dNK8SpDPHLsYB22a2Hd9D2xCPsJp+0u/NbtJtXn8vWwr2mFREskWKdQYDe+efn6uvjb8OpE1G8Bjeu39/8gF6XQWnXaBmY8R6KKDY78Bq+X1eBSZVJSj6jgtU/fr49X6UtHdpo69gP0uMIQu87MdJrQfy7M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; c=relaxed/simple; bh=YuyPCd/TXGvgbAvhPChGeYw4WxnJWp82+drwUmZOF3s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=GY5wQn7zmayXQbgrhffSlnjh1qs+WN7iwQcjZGR68NykcKyJfE4vAa3XeSdY4CwoIRuvHwASlVj+Ck5ELzHos0YgmcSVStw3Q6L8BWC+XzndhXU+7gp92xmJR+C/XeKMe3wny5VOFf7nJJ6b70AfO/uNJ3EPkim3wAyA9VANXQU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=QEPWjCtG; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="QEPWjCtG" Received: by smtp.kernel.org (Postfix) with ESMTPS id B5D4DC2BCF5; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786611132; bh=YuyPCd/TXGvgbAvhPChGeYw4WxnJWp82+drwUmZOF3s=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=QEPWjCtGG8O34+4NMiSIXqSwp4U5bnOEDqslYKlj0rvgdylZpAcndqddK2iVtTx6u zmgM0y2ji/Zd8ehQPJGGnCmGFWuD/4Psth4bG9a8TeJovlALBwehTzCvCUANWoizMh e1VAHRWscZw11V8E3A/rk6ciwW1Tgy9w8YBH7LTxA3qDQPw1hFy91XyGGHVhRNkSVy lpBFDJMkMJfksU54MNE7ok6EAEBHbFXb08Tm2iSssx0wSApAfcD8XUpj3IPDWS2cij lJ/5814zGkTbUKkOjnqN0/bmMGQhO/HgBfUMJ0B7k1t8+FUxPIZKc9+r6mR53inH3r QIyzUIdctyb2g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 9ECC6C5DF66; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 13 Aug 2026 16:52:04 +0800 Subject: [PATCH RFC 3/5] memcg: remove retry logic in mem_cgroup_private_id_get_online Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-memcgid-objcg-v1-3-83d21c685b77@tencent.com> References: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> In-Reply-To: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , Dave Chinner , Qi Zheng , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu Cc: cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Bingfang Guo , Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786611130; l=3214; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=5i2b3adAqDemDU/NAVoq8sK/uF2LIbB9NIx+Y5tqrZ8=; b=R9G5tW/eOsF1fEjMg6cHo6RdBDAD5yzOljC1iwmE4ddG37Cg5xTgcu5n+bSnG4To3/L8nTiny eDLxTl9DEDGAgfNU01C5tFnLzDpgWesTl3zgcaES2tabu+sgDAb5DeH X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo With the ID released only in css_released() (patch 1) and every caller holding a stable reference obtained from get_mem_cgroup_from_objcg() (patch 2), the id refcount is guaranteed to be non-zero whenever the ID is taken, so the retry loop that walked up the parent chain can never trigger. Remove the retry logic, rename the function to mem_cgroup_private_id_get() and turn the fallible return value into a VM_WARN_ON() that documents the invariant. Signed-off-by: Bingfang Guo --- mm/memcontrol-v1.c | 2 +- mm/memcontrol-v1.h | 3 +-- mm/memcontrol.c | 20 +++++--------------- 3 files changed, 7 insertions(+), 18 deletions(-) diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index a913d32ad1e17..e5161e061bd11 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -648,7 +648,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_= cluster_info *ci) mod_memcg_state(memcg, MEMCG_SWAP, nr_entries); =20 /* we have a reference to it, so we should get exact memcg itself */ - mem_cgroup_private_id_get_online(memcg, nr_entries); + mem_cgroup_private_id_get(memcg, nr_entries); =20 __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_entries, mem_cgroup_private_id(memcg)); diff --git a/mm/memcontrol-v1.h b/mm/memcontrol-v1.h index 0f703f239c80f..9c74400aa7ddb 100644 --- a/mm/memcontrol-v1.h +++ b/mm/memcontrol-v1.h @@ -21,8 +21,7 @@ void drain_all_stock(struct mem_cgroup *root_memcg); =20 int memory_stat_show(struct seq_file *m, void *v); =20 -struct mem_cgroup *mem_cgroup_private_id_get_online(struct mem_cgroup *mem= cg, - unsigned int n); +void mem_cgroup_private_id_get(struct mem_cgroup *memcg, unsigned int n); =20 /* Cgroup v1-specific declarations */ #ifdef CONFIG_MEMCG_V1 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index a210fe2501219..12545ca48194d 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -3985,23 +3985,13 @@ static inline void mem_cgroup_private_id_put(struct= mem_cgroup *memcg, unsigned __mem_cgroup_private_id_put(memcg->id_objcg, memcg->id, n); } =20 -struct mem_cgroup *mem_cgroup_private_id_get_online(struct mem_cgroup *mem= cg, unsigned int n) +void mem_cgroup_private_id_get(struct mem_cgroup *memcg, unsigned int n) { + bool success; struct obj_cgroup *objcg =3D memcg->id_objcg; =20 - while (!refcount_add_not_zero(n, &objcg->id_ref)) { - /* - * The root cgroup cannot be destroyed, so it's refcount must - * always be >=3D 1. - */ - if (WARN_ON_ONCE(mem_cgroup_is_root(memcg))) { - VM_BUG_ON(1); - break; - } - memcg =3D parent_mem_cgroup(memcg); - objcg =3D memcg->id_objcg; - } - return memcg; + success =3D refcount_add_not_zero(n, &objcg->id_ref); + VM_WARN_ON(!success); } =20 /** @@ -5673,7 +5663,7 @@ int __mem_cgroup_try_charge_swap(struct folio *folio) mod_memcg_state(memcg, MEMCG_SWAP, nr_pages); =20 /* we have a reference to it, so we should get exact memcg itself */ - mem_cgroup_private_id_get_online(memcg, nr_pages); + mem_cgroup_private_id_get(memcg, nr_pages); =20 ci =3D swap_cluster_get_and_lock(folio); __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages, --=20 2.43.7 From nobody Tue Sep 29 02:34:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B34E442FDB; Thu, 13 Aug 2026 08:52:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; cv=none; b=T6kSYc+k8AcKOIj3HfZaSABqGLknVxggzzEy8dCHVpMRdnwAuybbEugIXv1ovfSVGtNU1eIRxirxorq/nZETu+/4E2otjBn666kyzkDObB4XNXTJlCtcIgBZD2p0+dQsJ0O9+H/ILy5K5NJnAlHuTzEqf1Qu4nxhAk9P4G+Ze44= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; c=relaxed/simple; bh=TVNyWPxC3vugi2lXQarBpbfKCo9Zrfp4Eb3m0alZ6VQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=KaZdknFvpN12/9laoo5X6inikdGRDd/jC+F8hkACY0n1CTyhSCHBTAm2gXCvj+AUw7TjPgAV8rYWOX+9/FEv9TJZbmUcMxJEM9G5DoN8865YZIBDK+uYcl5S9W39wkzDwyhen6AQetT5xGDV9EXrduW2z/zDI9VvG/10HqZb4ns= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=HNeVQYsl; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="HNeVQYsl" Received: by smtp.kernel.org (Postfix) with ESMTPS id C9275C2BCFD; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786611132; bh=TVNyWPxC3vugi2lXQarBpbfKCo9Zrfp4Eb3m0alZ6VQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=HNeVQYslTmxKPiegC7pSF3twWJOh8xZnlXJxQYbxOcP6ULvEYLjy+F+lWXsuif0fy E/oxHp5u5J0SmXTxILeT+sS9kexpAn+dCXtRzDuxgMmCwJt1tXXdAuTD8S5/P4cQiQ hRgXABdXtGXw4KbbQrc6pWhdmcMzzjHqTk9oPiTSb41O5zCG3OLkuvPRrpXvev0VRF uSnm6o02cp8JItit158JHyThWQMdg0MQEk0NTrMvs1aDZg4xq5odOT2u/MTD7XU+YT bg4O6wbHWlQVDcyZfVW5SPJFygUXLpG5yAaqAgTYcwfZJ3rDUmz+G7kqZu+wjyk5tl v79EsUGSP4kmg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id B098DC5CFEB; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 13 Aug 2026 16:52:05 +0800 Subject: [PATCH RFC 4/5] memcg: add a helper to get online memcg from memcgid Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-memcgid-objcg-v1-4-83d21c685b77@tencent.com> References: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> In-Reply-To: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , Dave Chinner , Qi Zheng , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu Cc: cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Bingfang Guo , Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786611130; l=2588; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=9eOtm2zMK/wYtzZgGx4xX57blghdDku36wwXT3ij+M4=; b=WM11z9bUgck0UICbTCfnt20vuTyNOqdecwoNJvQu5KCqITzAtCi1u5O7CuPml+516VwM/0RIf vroA/tx5h04A0FUcWL3DZ/ltvYeJU8pD2+eJ3eqxKUBOjxHI2+PZ/95 X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo When swapping in, the folio is charged back to the memcg that swapped it out, or to one of its ancestors if that memcg is gone. mem_cgroup_swapin_charge_folio() currently does the id lookup and the css_tryget_online() check by hand under the RCU read lock. The objcg behind the id is reparented to an online memcg when its own memcg is destroyed, so looking the id up and taking a reference through the objcg is enough to guarantee an online memcg. Add mem_cgroup_from_private_id_online() for that purpose and use it in mem_cgroup_swapin_charge_folio(), dropping the RCU read lock usage. Signed-off-by: Bingfang Guo --- include/linux/memcontrol.h | 1 + mm/memcontrol.c | 22 ++++++++++++++++++---- 2 files changed, 19 insertions(+), 4 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index c33ec7efad50b..fef8a1c4191b1 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -827,6 +827,7 @@ static inline unsigned short mem_cgroup_private_id(stru= ct mem_cgroup *memcg) return memcg->id; } struct mem_cgroup *mem_cgroup_from_private_id(unsigned short id); +struct mem_cgroup *mem_cgroup_from_private_id_online(unsigned short id); =20 static inline u64 mem_cgroup_id(struct mem_cgroup *memcg) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 12545ca48194d..fdf2e0d1f17e5 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4012,6 +4012,22 @@ struct mem_cgroup *mem_cgroup_from_private_id(unsign= ed short id) return obj_cgroup_memcg(objcg); } =20 +/** + * mem_cgroup_from_private_id - look up an online memcg from a memcg id + * and get a reference. + * @id: the memcg id to look up + */ +struct mem_cgroup *mem_cgroup_from_private_id_online(unsigned short id) +{ + struct obj_cgroup *objcg; + + objcg =3D xa_load(&mem_cgroup_private_ids, id); + if (!objcg) + return NULL; + + return get_mem_cgroup_from_objcg(objcg); +} + static struct mem_cgroup *mem_cgroup_take_from_private_id(unsigned short i= d, unsigned int n) { struct obj_cgroup *objcg; @@ -5248,11 +5264,9 @@ int mem_cgroup_swapin_charge_folio(struct folio *fol= io, unsigned short id, if (mem_cgroup_disabled()) return 0; =20 - rcu_read_lock(); - memcg =3D mem_cgroup_from_private_id(id); - if (!memcg || !css_tryget_online(&memcg->css)) + memcg =3D mem_cgroup_from_private_id_online(id); + if (!memcg) memcg =3D get_mem_cgroup_from_mm(mm); - rcu_read_unlock(); =20 ret =3D charge_memcg(folio, memcg, gfp); =20 --=20 2.43.7 From nobody Tue Sep 29 02:34:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B274442FB2; Thu, 13 Aug 2026 08:52:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; cv=none; b=Kqi2PrApvZpUxnbOc9Kp+czal5FZpVQ5AGLGorYOvlPysjHG3a+0PwGRSkcXT6eimJ/0mDhKK5LoKif6cAM6W0nKAykuqac80v9MeODiPZuzUIoY7J1+EgNpqT8N++vksEZqz5HghiKgZShN46npeC7x+N0SnQPssgVXe2gNTjo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611133; c=relaxed/simple; bh=meapVXEGFdDb6NF7sii95w8iasCv07pp1/gAqS5rawg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=d+x4+msaA1Gw+iqUGJoqguhdxFFN3azKOiM2OT9yrVQpKyuHEEo4lGKI5sdaCm0Damz425UN9h+weCkt7kJeYywwgmNfk5ynaTwvMoINCqq6G9i7RaWwwpc0OdlqT7TPTLAQ1xvu0mHh0Pm8aDZB0ZC3YLP6gwZqPD+jCK7wmZ0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=fF3ITC1F; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="fF3ITC1F" Received: by smtp.kernel.org (Postfix) with ESMTPS id D2653C2BD04; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786611132; bh=meapVXEGFdDb6NF7sii95w8iasCv07pp1/gAqS5rawg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=fF3ITC1FGfkSq+//ZPnwjAjYDzwqhnID02bRoPebzrOkwCsZkcIoaLMNULErC4KM/ J4HjoSjFkzSgDho8I3gN+xfJGQtvFyq/UuIwnHVUeXUJLroRG0jg0cZV9gKBw+1NmM LH59MDg4B8EiF1uFv4ZCuqOyP5a6IVByqZO59CIinw1Ti9wy46GK5YWHmcrKO0Rzun xDW/71H7UFy8nFxLvVDPhn3txw2pPMdlOBCIbHMzLpHqgkK7oELkFSDPJneKoOLB85 1v1WCRmMPEHpAyZv9pikwgrcFqx07I6Ikzefyg4+wrVawxS8Egz8SDX4dGTnEgUEoH 4K6ab0qMorjpQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id C0F93C5DF67; Thu, 13 Aug 2026 08:52:12 +0000 (UTC) From: Bingfang Guo via B4 Relay Date: Thu, 13 Aug 2026 16:52:06 +0800 Subject: [PATCH RFC 5/5] memcg: filter out reparented memcgs got using memcgid Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-memcgid-objcg-v1-5-83d21c685b77@tencent.com> References: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> In-Reply-To: <20260813-memcgid-objcg-v1-0-83d21c685b77@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , Dave Chinner , Qi Zheng , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu Cc: cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Bingfang Guo , Bingfang Guo X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786611130; l=2805; i=bingfangguo@tencent.com; s=20260812; h=from:subject:message-id; bh=zY1JXeipRRE4x+8HShYlKEllHTJG9Rn8M3UwjRUwSJk=; b=CusELUTT7HqgDWh2KGQvd8fGGqtxEXggEBLMydLTG5KzIJxwbuc5HVMPhaTVT/r1MjFsRGMGg Wp/8xlAWmeCC0QRq+pJq/R8DkaLIDm1Gd2BHzegRT2mnbwkdEmg/i4o X-Developer-Key: i=bingfangguo@tencent.com; a=ed25519; pk=q7u9S7d5e7VijWrYF9O+vv9/XvqWduVGtwGtuCObAHE= X-Endpoint-Received: by B4 Relay for bingfangguo@tencent.com/20260812 with auth_id=942 X-Original-From: Bingfang Guo Reply-To: bingfangguo@tencent.com From: Bingfang Guo mem_cgroup_from_private_id() looks up the objcg that owns the id and returns the objcg's current memcg. After reparenting, that memcg can differ from the one the id originally belonged to. Callers such as the list lru and workingset refault code expect to get back exactly the memcg referred to by the memcgid, so check that the returned memcg still owns the id and return NULL otherwise, letting the callers skip the entry. Signed-off-by: Bingfang Guo --- mm/list_lru.c | 2 +- mm/memcontrol.c | 9 ++++++++- mm/workingset.c | 6 +++++- 3 files changed, 14 insertions(+), 3 deletions(-) diff --git a/mm/list_lru.c b/mm/list_lru.c index 36662d02ff963..bc956267f6835 100644 --- a/mm/list_lru.c +++ b/mm/list_lru.c @@ -428,7 +428,7 @@ unsigned long list_lru_walk_node(struct list_lru *lru, = int nid, xa_for_each(&lru->xa, index, mlru) { rcu_read_lock(); memcg =3D mem_cgroup_from_private_id(index); - if (!mem_cgroup_tryget(memcg)) { + if (!memcg || !mem_cgroup_tryget(memcg)) { rcu_read_unlock(); continue; } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index fdf2e0d1f17e5..e7555eca77019 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -3999,17 +3999,24 @@ void mem_cgroup_private_id_get(struct mem_cgroup *m= emcg, unsigned int n) * @id: the memcg id to look up * * Caller must hold rcu_read_lock(). + * + * @return: the memcg, or NULL if the memcg is already reparented. */ struct mem_cgroup *mem_cgroup_from_private_id(unsigned short id) { struct obj_cgroup *objcg; + struct mem_cgroup *memcg; WARN_ON_ONCE(!rcu_read_lock_held()); =20 objcg =3D xa_load(&mem_cgroup_private_ids, id); if (!objcg) return NULL; =20 - return obj_cgroup_memcg(objcg); + memcg =3D obj_cgroup_memcg(objcg); + if (mem_cgroup_private_id(memcg) !=3D id) + return NULL; + + return memcg; } =20 /** diff --git a/mm/workingset.c b/mm/workingset.c index f351798e723ac..b6e22536a5240 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -283,6 +283,10 @@ static bool lru_gen_test_recent(void *shadow, struct l= ruvec **lruvec, memcg =3D mem_cgroup_from_private_id(memcg_id); *lruvec =3D mem_cgroup_lruvec(memcg, pgdat); =20 + /* reparented memcg loses its max_seq */ + if (!memcg) + return false; + max_seq =3D READ_ONCE((*lruvec)->lrugen.max_seq); max_seq &=3D (file ? EVICTION_MASK : EVICTION_MASK_ANON) >> LRU_REFS_WIDT= H; =20 @@ -470,7 +474,7 @@ bool workingset_test_recent(void *shadow, bool file, bo= ol *workingset, * configurations instead. */ eviction_memcg =3D mem_cgroup_from_private_id(memcgid); - if (!mem_cgroup_tryget(eviction_memcg)) + if (!eviction_memcg || !mem_cgroup_tryget(eviction_memcg)) eviction_memcg =3D NULL; rcu_read_unlock(); =20 --=20 2.43.7