From nobody Fri Sep 25 19:16:06 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED8663B6C14; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; cv=none; b=luj7mye8ifgKCmwQPx48PosKQfe7WO6t8Ec9jU15V9LU+Uahh0N7CyhTlB2r0MojyZqiHAY9WhEs342vEQPHuZQXK3zf/uIRNheGVvyqT2KRcoqzs3x+goba/8Ha/5A1i6upRKavkN4o8sHubzXehTuo+By+1lk0zsOsUCXg+5U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; c=relaxed/simple; bh=DzE6a8ucUIKezdLyiT87i1fHKPmzHSKWoTm6ztJGBrE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dkIxKChxLRveRJ4EYOxpkOCnolJBVNjQ96Kv8sdPao5vhxR6rKK7+zcod3VnIJyeS334pyySaz3fYgzmqhu3r+0BAmHiHk07Yix9LPJeyoo/iHd2VF1YgD7CjVgjudjWutWcXPzdx+Bf2R1HKhottOG1WKkT8SbGVZwrr8WWYaM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Hj8TJM05; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Hj8TJM05" Received: by smtp.kernel.org (Postfix) with ESMTPS id 967D2C2BCFC; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788947062; bh=DzE6a8ucUIKezdLyiT87i1fHKPmzHSKWoTm6ztJGBrE=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Hj8TJM05RMbUz9l8edeAbgRMFTc2BxpJxnU49GGVlssA5yqJhzTXQWVFa75FQQdPO J4YNpewm/4ro6Y07XRGNqQL9cyOySQMVgEjUxIke3w8Oo3Fb4DdRSEuUgrbsnrzlNb 1PHUoFfpWtXrr4dtkw8qbvWyPKKVq2GmCyMVn9YQouAr1HiKc8i9a7oSWwU1y7WLvJ htEEwBsrMUDiJiPqabaUxG/VxfLFjo8rmBkvgao1rD7D105RjhTVaetCJDL26KiXap xXAGno8zLj2EJFQvj9BbhsxFkOa/Wyda2kvb6t8qeqAekqRjR0ruPXiwwYxelzU18l zSNvHb8QC44ww== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 74424C79FAD; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) From: linuszeng via B4 Relay Date: Wed, 09 Sep 2026 17:44:19 +0800 Subject: [PATCH v2 1/3] mm: page_counter: add page_counter_protection struct and init API Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-descriptive-name-v2-1-d7dd7c099049@tencent.com> References: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> In-Reply-To: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Maarten Lankhorst , Maxime Ripard , Natalie Vock , Tejun Heo , =?utf-8?q?Michal_Koutn=C3=BD?= , Oscar Salvador , Jingxiang Zeng Cc: Michal Hocko , cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, linuszeng X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788947061; l=8820; i=linuszeng@tencent.com; s=20260909; h=from:subject:message-id; bh=qRo7579G0THOdatUrKxYMO5fc5xqKHZM+hDjreNI4ZA=; b=sm9aZHKSL3UgtbwsCx3EepH0eABzrIRsPuV/OQa6YtjA4YYenyUMrOOwjC/BuCeTXIw+F2a4w E94ynhOvlhRA+SfoBNns+mQRxSsyDGgong2OQVembX7Fcbr5Hz7AA6T X-Developer-Key: i=linuszeng@tencent.com; a=ed25519; pk=6K54xRzYIRWqatrAPy86M4E0MsI92BVJBhXwz5NdC74= X-Endpoint-Received: by B4 Relay for linuszeng@tencent.com/20260909 with auth_id=1017 X-Original-From: linuszeng Reply-To: linuszeng@tencent.com From: linuszeng This commit extracts the hierarchical protection state (memory.min and memory.low) from struct page_counter into a new page_counter_protection structure. It introduces page_counter_init_protection() to attach this context, saving space for counters that don't support protection. The dmem pool allocator now points its counter at the embedded protection context, and the pool fix-up path in get_cg_pool_locked() links the new prot->parent the same way it links cnt.parent, so pools created bottom-up do not lose hierarchical protection. No functional change. --- include/linux/memcontrol.h | 7 ++++++ include/linux/page_counter.h | 59 +++++++++++++++++++++++++++++++++++++++-= ---- kernel/cgroup/dmem.c | 9 ++++--- mm/hugetlb_cgroup.c | 4 +-- mm/memcontrol.c | 21 ++++++++++------ mm/page_counter.c | 2 +- 6 files changed, 82 insertions(+), 20 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 058ebd73ff16..ed863f4ed233 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -195,6 +195,13 @@ struct mem_cgroup { /* Accounted resources */ struct page_counter memory; /* Both v1 & v2 */ =20 + /* + * Hierarchical memory.min/memory.low protection tracking for the + * memory page counter. swap/memsw, kmem and tcpmem counters do not + * support protection and have no such context. + */ + struct page_counter_protection memory_prot; + union { struct page_counter swap; /* v2 only */ struct page_counter memsw; /* v1 only */ diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index 07b7cb12249c..b81f16702764 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -7,6 +7,32 @@ #include #include =20 +/* + * Hierarchical protection (memory.min / memory.low) tracking. + * + * Only the memory page counter (and dmem pools) participate in protection. + * swap/memsw, kmem and tcpmem page counters never do, so the protection + * fields are kept out of struct page_counter in this separate structure to + * save space in the common case. struct page_counter links to it via ->pr= ot, + * which is NULL for counters without protection support. + */ +struct page_counter_protection { + struct page_counter_protection *parent; + + /* effective memory.min and memory.min usage tracking */ + unsigned long emin; + atomic_long_t min_usage; + atomic_long_t children_min_usage; + + /* effective memory.low and memory.low usage tracking */ + unsigned long elow; + atomic_long_t low_usage; + atomic_long_t children_low_usage; + + unsigned long min; + unsigned long low; +}; + struct page_counter { /* * Make sure 'usage' does not share cacheline with any other field in @@ -41,6 +67,12 @@ struct page_counter { unsigned long high; unsigned long max; struct page_counter *parent; + + /* + * Hierarchical protection context, NULL for counters that do not + * support memory.min/memory.low (swap, memsw, kmem, tcpmem, ...). + */ + struct page_counter_protection *prot; } ____cacheline_internodealigned_in_smp; =20 #if BITS_PER_LONG =3D=3D 32 @@ -49,18 +81,33 @@ struct page_counter { #define PAGE_COUNTER_MAX (LONG_MAX / PAGE_SIZE) #endif =20 -/* - * Protection is supported only for the first counter (with id 0). - */ static inline void page_counter_init(struct page_counter *counter, - struct page_counter *parent, - bool protection_support) + struct page_counter *parent) { counter->usage =3D (atomic_long_t)ATOMIC_LONG_INIT(0); counter->max =3D PAGE_COUNTER_MAX; counter->parent =3D parent; - counter->protection_support =3D protection_support; counter->track_failcnt =3D false; + counter->prot =3D NULL; +} + +/* + * Enable hierarchical protection (memory.min/memory.low) on @counter. + * @prot and @parent are the protection contexts of @counter and its + * parent page counter respectively. Only the memory page counter (and + * dmem pools) call this. + * + * The remaining members of @prot (emin, elow and the usage counters) are + * expected to be zero already, so @prot must come from zeroed memory. + */ +static inline void page_counter_init_protection(struct page_counter *count= er, + struct page_counter_protection *prot, + struct page_counter_protection *parent) +{ + counter->prot =3D prot; + prot->parent =3D parent; + prot->min =3D 0; + prot->low =3D 0; } =20 static inline unsigned long page_counter_read(struct page_counter *counter) diff --git a/kernel/cgroup/dmem.c b/kernel/cgroup/dmem.c index 4683f3d68022..a4bac0d5ac3b 100644 --- a/kernel/cgroup/dmem.c +++ b/kernel/cgroup/dmem.c @@ -88,6 +88,7 @@ struct dmem_cgroup_pool_state { struct rcu_head rcu; =20 struct page_counter cnt; + struct page_counter_protection prot; struct dmem_cgroup_pool_state *parent; =20 refcount_t ref; @@ -426,8 +427,9 @@ alloc_pool_single(struct dmemcg_state *dmemcs, struct d= mem_cgroup_region *region if (parent) ppool =3D find_cg_pool_locked(parent, region); =20 - page_counter_init(&pool->cnt, - ppool ? &ppool->cnt : NULL, true); + page_counter_init(&pool->cnt, ppool ? &ppool->cnt : NULL); + page_counter_init_protection(&pool->cnt, &pool->prot, + ppool ? &ppool->prot : NULL); reset_all_resource_limits(pool); refcount_set(&pool->ref, 1); kref_get(®ion->ref); @@ -480,8 +482,9 @@ get_cg_pool_locked(struct dmemcg_state *dmemcs, struct = dmem_cgroup_region *regio /* ppool was created if it didn't exist by above loop. */ ppool =3D find_cg_pool_locked(pp, region); =20 - /* Fix up parent links, mark as inited. */ + /* Fix up parent links (counter and protection), mark as inited. */ pool->cnt.parent =3D &ppool->cnt; + pool->prot.parent =3D &ppool->prot; if (ppool && !pool->parent) { pool->parent =3D ppool; dmemcg_pool_get(ppool); diff --git a/mm/hugetlb_cgroup.c b/mm/hugetlb_cgroup.c index ecb6e0b7819a..7fdae504cfc6 100644 --- a/mm/hugetlb_cgroup.c +++ b/mm/hugetlb_cgroup.c @@ -108,8 +108,8 @@ static void hugetlb_cgroup_init(struct hugetlb_cgroup *= h_cgroup, fault =3D hugetlb_cgroup_counter_from_cgroup(h_cgroup, idx); rsvd =3D hugetlb_cgroup_counter_from_cgroup_rsvd(h_cgroup, idx); =20 - page_counter_init(fault, fault_parent, false); - page_counter_init(rsvd, rsvd_parent, false); + page_counter_init(fault, fault_parent); + page_counter_init(rsvd, rsvd_parent); =20 if (!cgroup_subsys_on_dfl(hugetlb_cgrp_subsys)) { fault->track_failcnt =3D true; diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 86ff580c7018..ffa1ced3baae 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4267,25 +4267,30 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *pa= rent_css) #endif page_counter_set_high(&memcg->swap, PAGE_COUNTER_MAX); if (parent) { - page_counter_init(&memcg->memory, &parent->memory, memcg_on_dfl); - page_counter_init(&memcg->swap, &parent->swap, false); + page_counter_init(&memcg->memory, &parent->memory); + if (memcg_on_dfl) + page_counter_init_protection(&memcg->memory, &memcg->memory_prot, + &parent->memory_prot); + page_counter_init(&memcg->swap, &parent->swap); #ifdef CONFIG_MEMCG_V1 WRITE_ONCE(memcg->swappiness, mem_cgroup_swappiness(parent)); memcg->memory.track_failcnt =3D !memcg_on_dfl; memcg->memsw.track_failcnt =3D !memcg_on_dfl; WRITE_ONCE(memcg->oom_kill_disable, READ_ONCE(parent->oom_kill_disable)); - page_counter_init(&memcg->kmem, &parent->kmem, false); - page_counter_init(&memcg->tcpmem, &parent->tcpmem, false); + page_counter_init(&memcg->kmem, &parent->kmem); + page_counter_init(&memcg->tcpmem, &parent->tcpmem); memcg->tcpmem.track_failcnt =3D !memcg_on_dfl; #endif } else { init_memcg_stats(); init_memcg_events(); - page_counter_init(&memcg->memory, NULL, true); - page_counter_init(&memcg->swap, NULL, false); + page_counter_init(&memcg->memory, NULL); + page_counter_init_protection(&memcg->memory, &memcg->memory_prot, + NULL); + page_counter_init(&memcg->swap, NULL); #ifdef CONFIG_MEMCG_V1 - page_counter_init(&memcg->kmem, NULL, false); - page_counter_init(&memcg->tcpmem, NULL, false); + page_counter_init(&memcg->kmem, NULL); + page_counter_init(&memcg->tcpmem, NULL); #endif root_mem_cgroup =3D memcg; return &memcg->css; diff --git a/mm/page_counter.c b/mm/page_counter.c index 450543f4b318..38cb99f5f50e 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -15,7 +15,7 @@ =20 static bool track_protection(struct page_counter *c) { - return c->protection_support; + return c->prot !=3D NULL; } =20 static void propagate_protected_usage(struct page_counter *c, --=20 2.43.7 From nobody Fri Sep 25 19:16:06 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED5DB3A9856; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; cv=none; b=aavwAdEOhCpKtUonjds9lckIaphMbHYOgUxzMjBf5mgYtKN+kU7MSZZVcn2Cf+1n1gr3FIj0R9WA/AES4ERMWVAxClTqOg1PmF6jjf/DKYSmfeaNmp0jG5aYnW/Ew3js5iSsEMH2bjJL4c2yrC06hLBmQmMlqRB89aC2Sz6nyL4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; c=relaxed/simple; bh=CRNCNFOv1d/GN5B5RDOOI3W6Az4pcT/yOZ7OHd2PmTU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=BBVxST5h7z4SaXihnNlvtIQChSlsEd7U6t+61qnqCZeDAYSj4nAMutaFnA4AiOaE7t244srODWKZAguKN714l++30jT0nvFrTUrYjH83m5W3amQyNzksV5zGr/WXaL3VOm5rmxrQkauQMqSiguFJrvfXnHJMS05tUhSVBdW3jTw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Ca+kykha; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Ca+kykha" Received: by smtp.kernel.org (Postfix) with ESMTPS id A6EB0C4AF0B; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788947062; bh=CRNCNFOv1d/GN5B5RDOOI3W6Az4pcT/yOZ7OHd2PmTU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Ca+kykhajnNckAaTgWwCtYDMwntf2X0wkKrXlvNUTTINw9pfmajZwXe2mCtVHDXkS eSzJjxzbPFw+aTWc9xCtU02u6fR0Fduew3A08cvSMVtz5i9fAxyb8eoiApeB+XL1dl uvP2eAAOTxcPT2MLZXToPLVcJDF+jwUjhY0Vtuqtv7prm0zkdH87TjRBrtX19ZT/2e fGyIhNRN4Bfb55rxaR5Sa8pvR2jYydZmc5J4NRllPJvBCJFzVTwUvmmRffLXGe4BfQ f3BnW+tiHI/RNfGA5ewq3n9IwTS5keBbKWxfKx75PoNf+xLoyVyBRkbRHBM/7fVO5g GG9JeGilqKlgA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 885CCC79F8C; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) From: linuszeng via B4 Relay Date: Wed, 09 Sep 2026 17:44:20 +0800 Subject: [PATCH v2 2/3] mm: page_counter: track protection state in page_counter_protection Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-descriptive-name-v2-2-d7dd7c099049@tencent.com> References: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> In-Reply-To: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Maarten Lankhorst , Maxime Ripard , Natalie Vock , Tejun Heo , =?utf-8?q?Michal_Koutn=C3=BD?= , Oscar Salvador , Jingxiang Zeng Cc: Michal Hocko , cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, linuszeng X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788947061; l=9367; i=linuszeng@tencent.com; s=20260909; h=from:subject:message-id; bh=XnbhI6+XRiaUlggWoaB82VSvLpQ+T9RolGkd/4opeEI=; b=5VgMCSoPCaXiCqb5wg+bntHG8d6Ak9j8BQa82fZVBGqZccF2FPppQ4a8KIUtrqj7tSklZX58w J1XUmGOlDWkCMxeVIF2N7KV/UqP5LXntzd+BqFlzw8g7w4RgPHAeSVX X-Developer-Key: i=linuszeng@tencent.com; a=ed25519; pk=6K54xRzYIRWqatrAPy86M4E0MsI92BVJBhXwz5NdC74= X-Endpoint-Received: by B4 Relay for linuszeng@tencent.com/20260909 with auth_id=1017 X-Original-From: linuszeng Reply-To: linuszeng@tencent.com From: linuszeng Move the read/write side of hierarchical protection from struct page_counter to struct page_counter_protection: propagate_protected_usage() updates the protection context of the parent, page_counter_set_min()/low() and page_counter_calculate_protection() operate on it, and memcg and dmem accessors (including dmem_cgroup_below_min()/below_low()) read min/low/emin/elow and children_*_usage from it. struct page_counter keeps its now-unused protection fields for now; they are removed in a follow-up commit. No functional change. --- include/linux/memcontrol.h | 8 +++---- kernel/cgroup/dmem.c | 12 +++++----- mm/memcontrol.c | 8 +++---- mm/page_counter.c | 59 +++++++++++++++++++++++++++++-------------= ---- 4 files changed, 52 insertions(+), 35 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index ed863f4ed233..44065001a66a 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -591,8 +591,8 @@ static inline void mem_cgroup_protection(struct mem_cgr= oup *root, if (root =3D=3D memcg) return; =20 - *min =3D READ_ONCE(memcg->memory.emin); - *low =3D READ_ONCE(memcg->memory.elow); + *min =3D READ_ONCE(memcg->memory_prot.emin); + *low =3D READ_ONCE(memcg->memory_prot.elow); } =20 void mem_cgroup_calculate_protection(struct mem_cgroup *root, @@ -616,7 +616,7 @@ static inline bool mem_cgroup_below_low(struct mem_cgro= up *target, if (mem_cgroup_unprotected(target, memcg)) return false; =20 - return READ_ONCE(memcg->memory.elow) >=3D + return READ_ONCE(memcg->memory_prot.elow) >=3D page_counter_read(&memcg->memory); } =20 @@ -626,7 +626,7 @@ static inline bool mem_cgroup_below_min(struct mem_cgro= up *target, if (mem_cgroup_unprotected(target, memcg)) return false; =20 - return READ_ONCE(memcg->memory.emin) >=3D + return READ_ONCE(memcg->memory_prot.emin) >=3D page_counter_read(&memcg->memory); } =20 diff --git a/kernel/cgroup/dmem.c b/kernel/cgroup/dmem.c index a4bac0d5ac3b..4027d3d309c8 100644 --- a/kernel/cgroup/dmem.c +++ b/kernel/cgroup/dmem.c @@ -212,12 +212,12 @@ set_resource_max(struct dmem_cgroup_pool_state *pool,= u64 val, bool nonblock) =20 static u64 get_resource_low(struct dmem_cgroup_pool_state *pool) { - return pool ? READ_ONCE(pool->cnt.low) : 0; + return pool ? READ_ONCE(pool->cnt.prot->low) : 0; } =20 static u64 get_resource_min(struct dmem_cgroup_pool_state *pool) { - return pool ? READ_ONCE(pool->cnt.min) : 0; + return pool ? READ_ONCE(pool->cnt.prot->min) : 0; } =20 static u64 get_resource_max(struct dmem_cgroup_pool_state *pool) @@ -388,13 +388,13 @@ bool dmem_cgroup_state_evict_valuable(struct dmem_cgr= oup_pool_state *limit_pool, dmem_cgroup_calculate_protection(limit_pool, test_pool); =20 used =3D page_counter_read(ctest); - min =3D READ_ONCE(ctest->emin); + min =3D READ_ONCE(ctest->prot->emin); =20 if (used <=3D min) return false; =20 if (!ignore_low) { - low =3D READ_ONCE(ctest->elow); + low =3D READ_ONCE(ctest->prot->elow); if (used > low) return true; =20 @@ -787,7 +787,7 @@ bool dmem_cgroup_below_min(struct dmem_cgroup_pool_stat= e *root, * here. */ dmem_cgroup_calculate_protection(root, test); - return page_counter_read(&test->cnt) <=3D READ_ONCE(test->cnt.emin); + return page_counter_read(&test->cnt) <=3D READ_ONCE(test->cnt.prot->emin); } EXPORT_SYMBOL_GPL(dmem_cgroup_below_min); =20 @@ -818,7 +818,7 @@ bool dmem_cgroup_below_low(struct dmem_cgroup_pool_stat= e *root, * here. */ dmem_cgroup_calculate_protection(root, test); - return page_counter_read(&test->cnt) <=3D READ_ONCE(test->cnt.elow); + return page_counter_read(&test->cnt) <=3D READ_ONCE(test->cnt.prot->elow); } EXPORT_SYMBOL_GPL(dmem_cgroup_below_low); =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index ffa1ced3baae..b4c01a0dfd4f 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4823,7 +4823,7 @@ static ssize_t memory_peak_write(struct kernfs_open_f= ile *of, char *buf, static int memory_min_show(struct seq_file *m, void *v) { return seq_puts_memcg_tunable(m, - READ_ONCE(mem_cgroup_from_seq(m)->memory.min)); + READ_ONCE(mem_cgroup_from_seq(m)->memory_prot.min)); } =20 static ssize_t memory_min_write(struct kernfs_open_file *of, @@ -4846,7 +4846,7 @@ static ssize_t memory_min_write(struct kernfs_open_fi= le *of, static int memory_low_show(struct seq_file *m, void *v) { return seq_puts_memcg_tunable(m, - READ_ONCE(mem_cgroup_from_seq(m)->memory.low)); + READ_ONCE(mem_cgroup_from_seq(m)->memory_prot.low)); } =20 static ssize_t memory_low_write(struct kernfs_open_file *of, @@ -6271,6 +6271,6 @@ void mem_cgroup_show_protected_memory(struct mem_cgro= up *memcg) memcg =3D root_mem_cgroup; =20 pr_warn("Memory cgroup min protection %lukB -- low protection %lukB", - K(atomic_long_read(&memcg->memory.children_min_usage)), - K(atomic_long_read(&memcg->memory.children_low_usage))); + K(atomic_long_read(&memcg->memory_prot.children_min_usage)), + K(atomic_long_read(&memcg->memory_prot.children_low_usage))); } diff --git a/mm/page_counter.c b/mm/page_counter.c index 38cb99f5f50e..401201c8e390 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -21,28 +21,29 @@ static bool track_protection(struct page_counter *c) static void propagate_protected_usage(struct page_counter *c, unsigned long usage) { + struct page_counter_protection *prot =3D c->prot; unsigned long protected, old_protected; long delta; =20 - if (!c->parent) + if (!prot || !prot->parent) return; =20 - protected =3D min(usage, READ_ONCE(c->min)); - old_protected =3D atomic_long_read(&c->min_usage); + protected =3D min(usage, READ_ONCE(prot->min)); + old_protected =3D atomic_long_read(&prot->min_usage); if (protected !=3D old_protected) { - old_protected =3D atomic_long_xchg(&c->min_usage, protected); + old_protected =3D atomic_long_xchg(&prot->min_usage, protected); delta =3D protected - old_protected; if (delta) - atomic_long_add(delta, &c->parent->children_min_usage); + atomic_long_add(delta, &prot->parent->children_min_usage); } =20 - protected =3D min(usage, READ_ONCE(c->low)); - old_protected =3D atomic_long_read(&c->low_usage); + protected =3D min(usage, READ_ONCE(prot->low)); + old_protected =3D atomic_long_read(&prot->low_usage); if (protected !=3D old_protected) { - old_protected =3D atomic_long_xchg(&c->low_usage, protected); + old_protected =3D atomic_long_xchg(&prot->low_usage, protected); delta =3D protected - old_protected; if (delta) - atomic_long_add(delta, &c->parent->children_low_usage); + atomic_long_add(delta, &prot->parent->children_low_usage); } } =20 @@ -257,7 +258,10 @@ void page_counter_set_min(struct page_counter *counter= , unsigned long nr_pages) { struct page_counter *c; =20 - WRITE_ONCE(counter->min, nr_pages); + if (!counter->prot) + return; + + WRITE_ONCE(counter->prot->min, nr_pages); =20 for (c =3D counter; c; c =3D c->parent) propagate_protected_usage(c, atomic_long_read(&c->usage)); @@ -274,7 +278,10 @@ void page_counter_set_low(struct page_counter *counter= , unsigned long nr_pages) { struct page_counter *c; =20 - WRITE_ONCE(counter->low, nr_pages); + if (!counter->prot) + return; + + WRITE_ONCE(counter->prot->low, nr_pages); =20 for (c =3D counter; c; c =3D c->parent) propagate_protected_usage(c, atomic_long_read(&c->usage)); @@ -445,9 +452,18 @@ void page_counter_calculate_protection(struct page_cou= nter *root, struct page_counter *counter, bool recursive_protection) { + struct page_counter_protection *prot =3D counter->prot; + struct page_counter_protection *parent_prot; unsigned long usage, parent_usage; struct page_counter *parent =3D counter->parent; =20 + /* + * Only counters with protection support (memory, dmem pools) are + * ever passed here, but guard anyway. + */ + if (!prot) + return; + /* * Effective values of the reclaim targets are ignored so they * can be stale. Have a look at mem_cgroup_protection for more @@ -463,23 +479,24 @@ void page_counter_calculate_protection(struct page_co= unter *root, return; =20 if (parent =3D=3D root) { - counter->emin =3D READ_ONCE(counter->min); - counter->elow =3D READ_ONCE(counter->low); + prot->emin =3D READ_ONCE(prot->min); + prot->elow =3D READ_ONCE(prot->low); return; } =20 + parent_prot =3D parent->prot; parent_usage =3D page_counter_read(parent); =20 - WRITE_ONCE(counter->emin, effective_protection(usage, parent_usage, - READ_ONCE(counter->min), - READ_ONCE(parent->emin), - atomic_long_read(&parent->children_min_usage), + WRITE_ONCE(prot->emin, effective_protection(usage, parent_usage, + READ_ONCE(prot->min), + READ_ONCE(parent_prot->emin), + atomic_long_read(&parent_prot->children_min_usage), recursive_protection)); =20 - WRITE_ONCE(counter->elow, effective_protection(usage, parent_usage, - READ_ONCE(counter->low), - READ_ONCE(parent->elow), - atomic_long_read(&parent->children_low_usage), + WRITE_ONCE(prot->elow, effective_protection(usage, parent_usage, + READ_ONCE(prot->low), + READ_ONCE(parent_prot->elow), + atomic_long_read(&parent_prot->children_low_usage), recursive_protection)); } #endif /* CONFIG_MEMCG || CONFIG_CGROUP_DMEM */ --=20 2.43.7 From nobody Fri Sep 25 19:16:06 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED7853AD522; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; cv=none; b=Eg+7a3Pd2h1ZL8cBgeuPDi9EHI4FdfVzBM4izRxp8ulPz9E07nrmXZrqpcDV2iOvoIVFsZZ/nbw/gNhDDGgPFY6tF0EYhYkP9m6tJsSEkMerhV7Qybzz3GBAJSTdRSmrp6KhlNSAlCKxv8Vjv4uzLrF+FlFPhh/+S+Dp03/wvO8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947063; c=relaxed/simple; bh=Ad5PxfNFH0KjUpTuLE+nETx4q7cGEdziN7xPAPqWwOw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=c3qUccrUb/E8lPj+WLkf7nyLudwbTZGDzAuv1WRyXk0VbAmMw0gO3LDCcUKrJCrDM2awyoidnLkqqyofZv7t7RQFg8wBMp84Uj0uGDu/7vv6N9LFF6OIM8rdPmb8f2IDcGFn1WlcojgLGBt6bpX5UMeCFmBSxMksAxhUR3gHhIw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=IXM1iDLR; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="IXM1iDLR" Received: by smtp.kernel.org (Postfix) with ESMTPS id B63EBC2BD04; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788947062; bh=Ad5PxfNFH0KjUpTuLE+nETx4q7cGEdziN7xPAPqWwOw=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=IXM1iDLR+2q7Qj4UZE1d1OdKRnk1SWc6Th/SW9s7EtMl9ukUkzwFJuXlchIh0ZG7V kXMbB4b1c6/O/PP2klrteeuRTIuJqi2PoU0vgOhuvXMAyCemST3VLF0BvP5z9c6fVZ y+4TMX5lsYxOPQiaNVERUL+X5SPpA0gat5nSgRwKCHry0y9mZ8ujjEc6eHMpkVHgfs zACqQPA6A4gpxAlCKvKDS9dB/TaADmeL+QK6k9Y6aZYyHi7S1FqPaO9yGFT/phq03M LW3lmhU0IHCcxGH6bpffE+uzcLe0vH9Ypl9fvpONY/EC+o6Ol3W1/sKLT9X0l1669x pvSSIYNs+XIjA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 9AF87C79FB9; Wed, 9 Sep 2026 09:44:22 +0000 (UTC) From: linuszeng via B4 Relay Date: Wed, 09 Sep 2026 17:44:21 +0800 Subject: [PATCH v2 3/3] mm: page_counter: drop protection fields from struct page_counter Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-descriptive-name-v2-3-d7dd7c099049@tencent.com> References: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> In-Reply-To: <20260909-descriptive-name-v2-0-d7dd7c099049@tencent.com> To: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Maarten Lankhorst , Maxime Ripard , Natalie Vock , Tejun Heo , =?utf-8?q?Michal_Koutn=C3=BD?= , Oscar Salvador , Jingxiang Zeng Cc: Michal Hocko , cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, linuszeng X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788947061; l=1664; i=linuszeng@tencent.com; s=20260909; h=from:subject:message-id; bh=Ty2hOT3RbwUGz3En7B5H9g1JCCLU1A+cfCkKn52iXOE=; b=DabIe6gyrzCaIlUMZ+sK5yF+k8V1AXEAlJEyfeNhNSOm4taa8h7olgRc3rZRs1NDNlh3wZvec 6uuVi36nj8mDzWloXIy3VsJaBD8D3LXRVtrsGSlLIRQ9LQUXjqdvkPm X-Developer-Key: i=linuszeng@tencent.com; a=ed25519; pk=6K54xRzYIRWqatrAPy86M4E0MsI92BVJBhXwz5NdC74= X-Endpoint-Received: by B4 Relay for linuszeng@tencent.com/20260909 with auth_id=1017 X-Original-From: linuszeng Reply-To: linuszeng@tencent.com From: linuszeng The protection state now lives in struct page_counter_protection, so remove the emin/min_usage/children_min_usage, elow/low_usage/ children_low_usage, min, low and protection_support fields from struct page_counter. Also drop the now-orphaned _pad2_ padding and its comment: with the protection fields gone it no longer separates the read-mostly fields from anything, and the structure's cacheline alignment already pads the tail out. swap/memsw, kmem, tcpmem and hugetlb counters no longer carry this unused state: on 64-bit the structure shrinks from three cache lines to two, saving one cache line. --- include/linux/page_counter.h | 16 ---------------- 1 file changed, 16 deletions(-) diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index b81f16702764..0007960ba5a4 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -43,27 +43,11 @@ struct page_counter { =20 CACHELINE_PADDING(_pad1_); =20 - /* effective memory.min and memory.min usage tracking */ - unsigned long emin; - atomic_long_t min_usage; - atomic_long_t children_min_usage; - - /* effective memory.low and memory.low usage tracking */ - unsigned long elow; - atomic_long_t low_usage; - atomic_long_t children_low_usage; - unsigned long watermark; /* Latest cg2 reset watermark */ unsigned long local_watermark; =20 - /* Keep all the read most fields in a separete cacheline. */ - CACHELINE_PADDING(_pad2_); - - bool protection_support; bool track_failcnt; - unsigned long min; - unsigned long low; unsigned long high; unsigned long max; struct page_counter *parent; --=20 2.43.7