From nobody Tue Sep 29 06:05:24 2026 Received: from out-172.mta0.migadu.com (out-172.mta0.migadu.com [91.218.175.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 12A53364053 for ; Tue, 11 Aug 2026 20:32:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480339; cv=none; b=gSOw25WjnvkmvBHj25ONaXAiF+huTeIBnhAlg2Mnf2N5rRFoQ+zd8GHS4kncoU7JjDAWvk45RxMUyUqGsUpl4EdaeeQos78I6snEiXXV0xrw435FNfCtMHGl0Yh+htVCpmEaXHceEDZbQOwgfVUoJnFBB9jaIYqTKARkfRm7Kk8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480339; c=relaxed/simple; bh=InC3oa/piHRn7TfFr/YWrB0pvUFS081ccjqwsa/EqP4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s2Y9gn7XZv3FTlhA7FqSxKYfNa1MQztQEij7ifZFfYpbWIUiYswpzsNM236nUh7skNbeU8Diw7wYwZOXaxtTCYbyQSA15OdKoovLtlMRqzQlwuaISzUPPXubv8Qf4yKfH93zSl0WOcEUei2lnlCbd+8SCWKM6KvwY9O6udZrE0o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=GHFUcx52; arc=none smtp.client-ip=91.218.175.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="GHFUcx52" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480334; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=L1wlHWvlNYK5X1K+L0iWyXB0vamRy/Y38tCPlEOhHAM=; b=GHFUcx52I5JenvTVMdsP5ma5oWDaCPHj2bexFUFhqqbihu1co2+uFmNkT8mKZHxMr0yi7i bcdbNyVplWAsJ89GZkLTLksflSlJn217QSyr0h1zPUpfP84aKJLREgJ3RT0lhhNZkQ/dRO fnnMMvFdTpu2vvtFkIF0itKpByVkWvY= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, syzbot+12ee2725d5fde63a9c96@syzkaller.appspotmail.com Subject: [PATCH 1/9] memcg: make the v1 soft limit knob inert Date: Tue, 11 Aug 2026 13:31:55 -0700 Message-ID: <20260811203203.3456029-2-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" The v1 soft limit has been deprecated since v6.12 and nobody has reported depending on it. Start the removal by decoupling the interface from the implementation: keep memory.soft_limit_in_bytes, but ignore writes to it and always report the maximum value on read similar to what memory.kmem.limit_in_bytes already does. Writes are still parsed, so malformed input keeps returning -EINVAL. The knob now also behaves the same everywhere: it used to return -EOPNOTSUPP on PREEMPT_RT, where soft limit reclaim has always been disabled. This also fixes the syzbot report linked below. Soft limit reclaim is the only caller that runs shrink_lruvec() from kswapd against a specific memcg, so it is the only way to reach lru_gen_shrink_lruvec() and in turn set_mm_walk(), which warns when called from kswapd. Reported-by: syzbot+12ee2725d5fde63a9c96@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a7a6929.b50370da.49fe0.005e.GAE@google= .com/ Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- .../admin-guide/cgroup-v1/memory.rst | 49 +++---------------- mm/memcontrol-v1.c | 43 +++++++++------- 2 files changed, 32 insertions(+), 60 deletions(-) diff --git a/Documentation/admin-guide/cgroup-v1/memory.rst b/Documentation= /admin-guide/cgroup-v1/memory.rst index 7db63c002922..7d2a44af52c9 100644 --- a/Documentation/admin-guide/cgroup-v1/memory.rst +++ b/Documentation/admin-guide/cgroup-v1/memory.rst @@ -47,7 +47,6 @@ Features: - pages are linked to per-memcg LRU exclusively, and there is no global L= RU. - optionally, memory+swap usage can be accounted and limited. - hierarchical accounting - - soft limit - moving (recharging) account at moving a task is selectable. - usage threshold notifier - memory pressure notifier @@ -76,10 +75,9 @@ Brief summary of control files. memory.memsw.failcnt show the number of memory+Swap hits limits memory.max_usage_in_bytes show max memory usage recorded memory.memsw.max_usage_in_bytes show max memory+Swap usage recorded - memory.soft_limit_in_bytes set/show soft limit of memory usage - This knob is not available on CONFIG_PREEMPT_RT systems. - This knob is deprecated and shouldn't= be - used. + memory.soft_limit_in_bytes This knob is deprecated and has no effect. + Writes are ignored and reads always + return the maximum value. memory.stat show various statistics memory.use_hierarchy set/show hierarchical account enabled This knob is deprecated and shouldn't= be @@ -340,9 +338,6 @@ memory.kmem.usage_in_bytes, or in a separate counter wh= en it makes sense. The main "kmem" counter is fed into the main counter, so kmem charges will also be visible from the user counter. =20 -Currently no soft limit is implemented for kernel memory. It is future work -to trigger slab reclaim when those limits are reached. - 2.7.1 Current Kernel Memory resources accounted ----------------------------------------------- =20 @@ -710,42 +705,10 @@ For compatibility reasons writing 1 to memory.use_hie= rarchy will always pass:: =20 THIS IS DEPRECATED! =20 -Soft limits allow for greater sharing of memory. The idea behind soft limi= ts -is to allow control groups to use as much of the memory as needed, provided - -a. There is no memory contention -b. They do not exceed their hard limit - -When the system detects memory contention or low memory, control groups -are pushed back to their soft limits. If the soft limit of each control -group is very high, they are pushed back as much as possible to make -sure that one control group does not starve the others of memory. - -Please note that soft limits is a best-effort feature; it comes with -no guarantees, but it does its best to make sure that when memory is -heavily contended for, memory is allocated based on the soft limit -hints/setup. Currently soft limit based reclaim is set up such that -it gets invoked from balance_pgdat (kswapd). - -7.1 Interface -------------- - -Soft limits can be setup by using the following commands (in this example = we -assume a soft limit of 256 MiB):: - - # echo 256M > memory.soft_limit_in_bytes - -If we want to change this to 1G, we can at any time use:: +Writing to memory.soft_limit_in_bytes has no effect and reading it will +always return the maximum value. =20 - # echo 1G > memory.soft_limit_in_bytes - -.. note:: - Soft limits take effect over a long period of time, since they invo= lve - reclaiming memory for balancing between memory cgroups - -.. note:: - It is recommended to set the soft limit always below the hard limit, - otherwise the hard limit will take precedence. +Use memory.low and memory.min in cgroup v2 instead. =20 .. _cgroup-v1-memory-move-charges: =20 diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 835fc8e51184..05ef55cae4dc 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -96,7 +96,6 @@ enum { RES_LIMIT, RES_MAX_USAGE, RES_FAILCNT, - RES_SOFT_LIMIT, }; =20 #ifdef CONFIG_LOCKDEP @@ -1888,6 +1887,30 @@ static int mem_cgroup_hierarchy_write(struct cgroup_= subsys_state *css, return -EINVAL; } =20 +static u64 mem_cgroup_soft_limit_read(struct cgroup_subsys_state *css, + struct cftype *cft) +{ + return (u64)PAGE_COUNTER_MAX * PAGE_SIZE; +} + +static ssize_t mem_cgroup_soft_limit_write(struct kernfs_open_file *of, + char *buf, size_t nbytes, loff_t off) +{ + unsigned long nr_pages; + int ret; + + ret =3D page_counter_memparse(strstrip(buf), "-1", &nr_pages); + if (ret) + return ret; + + pr_warn_once("soft_limit_in_bytes is deprecated and will be removed. " + "Writing any value to this file has no effect. " + "Please report your usecase to linux-mm@kvack.org if you " + "depend on this functionality.\n"); + + return nbytes; +} + static u64 mem_cgroup_read_u64(struct cgroup_subsys_state *css, struct cftype *cft) { @@ -1924,8 +1947,6 @@ static u64 mem_cgroup_read_u64(struct cgroup_subsys_s= tate *css, return (u64)counter->watermark * PAGE_SIZE; case RES_FAILCNT: return counter->failcnt; - case RES_SOFT_LIMIT: - return (u64)READ_ONCE(memcg->soft_limit) * PAGE_SIZE; default: BUG(); } @@ -2020,17 +2041,6 @@ static ssize_t mem_cgroup_write(struct kernfs_open_f= ile *of, break; } break; - case RES_SOFT_LIMIT: - if (IS_ENABLED(CONFIG_PREEMPT_RT)) { - ret =3D -EOPNOTSUPP; - } else { - pr_warn_once("soft_limit_in_bytes is deprecated and will be removed. " - "Please report your usecase to linux-mm@kvack.org if you " - "depend on this functionality.\n"); - WRITE_ONCE(memcg->soft_limit, nr_pages); - ret =3D 0; - } - break; } return ret ?: nbytes; } @@ -2384,9 +2394,8 @@ struct cftype mem_cgroup_legacy_files[] =3D { }, { .name =3D "soft_limit_in_bytes", - .private =3D MEMFILE_PRIVATE(_MEM, RES_SOFT_LIMIT), - .write =3D mem_cgroup_write, - .read_u64 =3D mem_cgroup_read_u64, + .write =3D mem_cgroup_soft_limit_write, + .read_u64 =3D mem_cgroup_soft_limit_read, }, { .name =3D "failcnt", --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-179.mta0.migadu.com (out-179.mta0.migadu.com [91.218.175.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C45A9403B17 for ; Tue, 11 Aug 2026 20:32:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480343; cv=none; b=XMICYL2z/UwVyoBczQlqQlyFZqHWoh3AyFpjYoVX9GKoT9+ydPvWhKO/K1nGGv5Ke/RgfEGq+1L7Ms1b6iWbeUbKxdGU8UXCoJ2bU9sNfj2adFpjYCvjARGcMpIM/rWlgxaU9zu2vZml1n02FNkwU89l2AqzNXDKfKCVITcKYDU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480343; c=relaxed/simple; bh=49CMDiE/D5hIF/qR0Hif5omINvteXSpiYpKk1JPfRFw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jVXI1Sd9f+p2FRagqS46Kym11jNnu6MOZSCsvTmO4IdHd7Ho2nwbqWUFMJ7ddHRkL3MAKWsmNFderwG0l64rpJwLzNqUgcc0YoP9Cxr6v2RFPddFuJOYI6ZPxe0T/UqOb7fIrZJB6rBTGKiFjOuUJOPGYcNlbhC5CdbbTI2UCbo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=nHgA5+SY; arc=none smtp.client-ip=91.218.175.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="nHgA5+SY" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480339; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=zDWP20iIxrMfICDaLl7H7nQinULjLOkDjQsbQNGYncg=; b=nHgA5+SYM16j2OtMexfEx2mAgozqWPUc+BgYMXWhfCjZY39sfhcIjgaKO+J/fP46aGFVXU DmF2PXcl/gvCk5FaZzVA9eJ61sIkBaC3lrVf3vC5TZI/PjBLfe5AosdjVe3wMuz2ONIb1w 9LUndI7ySNQThqSIfNKz7oL9HOxc/1U= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 2/9] memcg: remove v1 soft limit reclaim Date: Tue, 11 Aug 2026 13:31:56 -0700 Message-ID: <20260811203203.3456029-3-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Nothing can put a cgroup on the soft limit rbtree anymore, so the tree is always empty and both callers of memcg1_soft_limit_reclaim() are guaranteed no-ops. Remove the reclaim pass from direct reclaim and from kswapd, along with its implementation. In shrink_zones() this leaves the global reclaim branch with a last_pgdat check that is now redundant with the identical check right below it, so drop it and move the explaining comment down to the check that remains. That check could only ever fire once last_pgdat was set, which implies first_pgdat had already been assigned, so skipping it does not change which node consider_reclaim_throttle() gets. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- include/linux/memcontrol.h | 12 --- mm/memcontrol-v1.c | 175 ------------------------------------- mm/vmscan.c | 39 ++------- 3 files changed, 6 insertions(+), 220 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index e78bc98ab229..7b02f1b3bb88 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -1927,10 +1927,6 @@ static inline bool mem_cgroup_zswap_writeback_enable= d(struct mem_cgroup *memcg) /* Cgroup v1-related declarations */ =20 #ifdef CONFIG_MEMCG_V1 -unsigned long memcg1_soft_limit_reclaim(pg_data_t *pgdat, int order, - gfp_t gfp_mask, - unsigned long *total_scanned); - bool mem_cgroup_oom_synchronize(bool wait); =20 static inline bool task_in_memcg_oom(struct task_struct *p) @@ -1951,14 +1947,6 @@ static inline void mem_cgroup_exit_user_fault(void) } =20 #else /* CONFIG_MEMCG_V1 */ -static inline -unsigned long memcg1_soft_limit_reclaim(pg_data_t *pgdat, int order, - gfp_t gfp_mask, - unsigned long *total_scanned) -{ - return 0; -} - static inline bool task_in_memcg_oom(struct task_struct *p) { return false; diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 05ef55cae4dc..b38b8d0f7f51 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -34,13 +34,6 @@ struct mem_cgroup_tree { =20 static struct mem_cgroup_tree soft_limit_tree __read_mostly; =20 -/* - * Maximum loops in mem_cgroup_soft_reclaim(), used for soft - * limit reclaim to prevent infinite loops, if they ever occur. - */ -#define MEM_CGROUP_MAX_RECLAIM_LOOPS 100 -#define MEM_CGROUP_MAX_SOFT_LIMIT_RECLAIM_LOOPS 2 - /* for OOM */ struct mem_cgroup_eventfd_list { struct list_head list; @@ -233,174 +226,6 @@ void memcg1_remove_from_trees(struct mem_cgroup *memc= g) } } =20 -static struct mem_cgroup_per_node * -__mem_cgroup_largest_soft_limit_node(struct mem_cgroup_tree_per_node *mctz) -{ - struct mem_cgroup_per_node *mz; - -retry: - mz =3D NULL; - if (!mctz->rb_rightmost) - goto done; /* Nothing to reclaim from */ - - mz =3D rb_entry(mctz->rb_rightmost, - struct mem_cgroup_per_node, tree_node); - /* - * Remove the node now but someone else can add it back, - * we will to add it back at the end of reclaim to its correct - * position in the tree. - */ - __mem_cgroup_remove_exceeded(mz, mctz); - if (!soft_limit_excess(mz->memcg) || - !css_tryget(&mz->memcg->css)) - goto retry; -done: - return mz; -} - -static struct mem_cgroup_per_node * -mem_cgroup_largest_soft_limit_node(struct mem_cgroup_tree_per_node *mctz) -{ - struct mem_cgroup_per_node *mz; - - spin_lock_irq(&mctz->lock); - mz =3D __mem_cgroup_largest_soft_limit_node(mctz); - spin_unlock_irq(&mctz->lock); - return mz; -} - -static int mem_cgroup_soft_reclaim(struct mem_cgroup *root_memcg, - pg_data_t *pgdat, - gfp_t gfp_mask, - unsigned long *total_scanned) -{ - struct mem_cgroup *victim =3D NULL; - int total =3D 0; - int loop =3D 0; - unsigned long excess; - unsigned long nr_scanned; - struct mem_cgroup_reclaim_cookie reclaim =3D { - .pgdat =3D pgdat, - }; - - excess =3D soft_limit_excess(root_memcg); - - while (1) { - victim =3D mem_cgroup_iter(root_memcg, victim, &reclaim); - if (!victim) { - loop++; - if (loop >=3D 2) { - /* - * If we have not been able to reclaim - * anything, it might because there are - * no reclaimable pages under this hierarchy - */ - if (!total) - break; - /* - * We want to do more targeted reclaim. - * excess >> 2 is not to excessive so as to - * reclaim too much, nor too less that we keep - * coming back to reclaim from this cgroup - */ - if (total >=3D (excess >> 2) || - (loop > MEM_CGROUP_MAX_RECLAIM_LOOPS)) - break; - } - continue; - } - total +=3D mem_cgroup_shrink_node(victim, gfp_mask, false, - pgdat, &nr_scanned); - *total_scanned +=3D nr_scanned; - if (!soft_limit_excess(root_memcg)) - break; - } - mem_cgroup_iter_break(root_memcg, victim); - return total; -} - -unsigned long memcg1_soft_limit_reclaim(pg_data_t *pgdat, int order, - gfp_t gfp_mask, - unsigned long *total_scanned) -{ - unsigned long nr_reclaimed =3D 0; - struct mem_cgroup_per_node *mz, *next_mz =3D NULL; - unsigned long reclaimed; - int loop =3D 0; - struct mem_cgroup_tree_per_node *mctz; - unsigned long excess; - - if (lru_gen_enabled()) - return 0; - - if (order > 0) - return 0; - - mctz =3D soft_limit_tree.rb_tree_per_node[pgdat->node_id]; - - /* - * Do not even bother to check the largest node if the root - * is empty. Do it lockless to prevent lock bouncing. Races - * are acceptable as soft limit is best effort anyway. - */ - if (!mctz || RB_EMPTY_ROOT(&mctz->rb_root)) - return 0; - - /* - * This loop can run a while, specially if mem_cgroup's continuously - * keep exceeding their soft limit and putting the system under - * pressure - */ - do { - if (next_mz) - mz =3D next_mz; - else - mz =3D mem_cgroup_largest_soft_limit_node(mctz); - if (!mz) - break; - - reclaimed =3D mem_cgroup_soft_reclaim(mz->memcg, pgdat, - gfp_mask, total_scanned); - nr_reclaimed +=3D reclaimed; - spin_lock_irq(&mctz->lock); - - /* - * If we failed to reclaim anything from this memory cgroup - * it is time to move on to the next cgroup - */ - next_mz =3D NULL; - if (!reclaimed) - next_mz =3D __mem_cgroup_largest_soft_limit_node(mctz); - - excess =3D soft_limit_excess(mz->memcg); - /* - * One school of thought says that we should not add - * back the node to the tree if reclaim returns 0. - * But our reclaim could return 0, simply because due - * to priority we are exposing a smaller subset of - * memory to reclaim from. Consider this as a longer - * term TODO. - */ - /* If excess =3D=3D 0, no tree ops */ - __mem_cgroup_insert_exceeded(mz, mctz, excess); - spin_unlock_irq(&mctz->lock); - css_put(&mz->memcg->css); - loop++; - /* - * Could not reclaim anything and there are no more - * mem cgroups to try or we seem to be looping without - * reclaiming anything. - */ - if (!nr_reclaimed && - (next_mz =3D=3D NULL || - loop > MEM_CGROUP_MAX_SOFT_LIMIT_RECLAIM_LOOPS)) - break; - } while (!nr_reclaimed); - if (next_mz) - css_put(&next_mz->memcg->css); - return nr_reclaimed; -} - static u64 mem_cgroup_move_charge_read(struct cgroup_subsys_state *css, struct cftype *cft) { diff --git a/mm/vmscan.c b/mm/vmscan.c index be6bd26e8c57..032b14793d91 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -6429,8 +6429,6 @@ static void shrink_zones(struct zonelist *zonelist, s= truct scan_control *sc) { struct zoneref *z; struct zone *zone; - unsigned long nr_soft_reclaimed; - unsigned long nr_soft_scanned; gfp_t orig_mask; pg_data_t *last_pgdat =3D NULL; pg_data_t *first_pgdat =3D NULL; @@ -6472,35 +6470,17 @@ static void shrink_zones(struct zonelist *zonelist,= struct scan_control *sc) sc->compaction_ready =3D true; continue; } - - /* - * Shrink each node in the zonelist once. If the - * zonelist is ordered by zone (not the default) then a - * node may be shrunk multiple times but in that case - * the user prefers lower zones being preserved. - */ - if (zone->zone_pgdat =3D=3D last_pgdat) - continue; - - /* - * This steals pages from memory cgroups over softlimit - * and returns the number of reclaimed pages and - * scanned pages. This works for global memory pressure - * and balancing, not for a memcg's limit. - */ - nr_soft_scanned =3D 0; - nr_soft_reclaimed =3D memcg1_soft_limit_reclaim(zone->zone_pgdat, - sc->order, sc->gfp_mask, - &nr_soft_scanned); - sc->nr_reclaimed +=3D nr_soft_reclaimed; - sc->nr_scanned +=3D nr_soft_scanned; - /* need some check for avoid more shrink_zone() */ } =20 if (!first_pgdat) first_pgdat =3D zone->zone_pgdat; =20 - /* See comment about same check for global reclaim above */ + /* + * Shrink each node in the zonelist once. If the zonelist is + * ordered by zone (not the default) then a node may be shrunk + * multiple times but in that case the user prefers lower zones + * being preserved. + */ if (zone->zone_pgdat =3D=3D last_pgdat) continue; last_pgdat =3D zone->zone_pgdat; @@ -7161,8 +7141,6 @@ clear_reclaim_active(pg_data_t *pgdat, int highest_zo= neidx) static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx) { int i; - unsigned long nr_soft_reclaimed; - unsigned long nr_soft_scanned; unsigned long pflags; unsigned long nr_boost_reclaim; unsigned long zone_boosts[MAX_NR_ZONES] =3D { 0, }; @@ -7268,12 +7246,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order= , int highest_zoneidx) */ kswapd_age_node(pgdat, &sc); =20 - /* Call soft limit reclaim before calling shrink_node. */ sc.nr_scanned =3D 0; - nr_soft_scanned =3D 0; - nr_soft_reclaimed =3D memcg1_soft_limit_reclaim(pgdat, sc.order, - sc.gfp_mask, &nr_soft_scanned); - sc.nr_reclaimed +=3D nr_soft_reclaimed; =20 /* * There should be no need to raise the scanning priority if --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-183.mta0.migadu.com (out-183.mta0.migadu.com [91.218.175.183]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E12A38C2BF; Tue, 11 Aug 2026 20:32:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.183 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480351; cv=none; b=DC8b3JRL2i9kED0WIDU2uFcAZ8cAdIlb52S64afR0twUeCv8zUMEHpx9MZaDwLgnXSAWA3Y6S3TrDc7EBUKA/CBkpkZHyuuFRtcz6rWGkdD/dbrPNXVrfwdba1E1eBu4RzdE2A/1IsTb60zK/YtI1eT1vDB+wwqfCbVoPTievB8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480351; c=relaxed/simple; bh=Qu4TLEITm3F7E4hKsZQWmQFX+kD67VWRsH9jTq8HbYY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JmmqpJwzs5ugxWcRVu4fOsyD5NE6DFhWa+D+Ixftmz/JoK8g8KM2E0Nt5e0aE9i9QSGRuUMNa1MPSRHgLG7Q/66CCDcWfyy4tpQNHlM/WG8MYwN0cNXhrGRfyR+QWFQtTeFO+tA1P9pVEdzNsoLEQfe2C/+h8yiC94jR2gGotjw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=jnixA1Lb; arc=none smtp.client-ip=91.218.175.183 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="jnixA1Lb" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480348; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Iz9pq9nN+pdRG2ifYD8weZHP3Mu9osKE7mH+dgefD54=; b=jnixA1LbmTaxkzgar3JKuaJ4k1vUWWyf7VEGK8rae8kCoq1f3oexACKpQ7r+bDe6+dKL3K RqBk78yFC/7Y/qhBF1y28UcuSmr0hRfRkmOknSvcbVpGSio38zp8SRfhMHhp6BWZSAcSu7 1Am0QQuHpJuo03Fodg5Ph06odqaUi/c= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 3/9] memcg: remove mem_cgroup_shrink_node() Date: Tue, 11 Aug 2026 13:31:57 -0700 Message-ID: <20260811203203.3456029-4-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Its only caller was soft limit reclaim, which is gone. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- mm/internal.h | 4 ---- mm/vmscan.c | 41 ----------------------------------------- 2 files changed, 45 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 678ce8d03515..b2315bdb7350 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -85,10 +85,6 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cg= roup *memcg, gfp_t gfp_mask, unsigned int reclaim_options, int *swappiness); -unsigned long mem_cgroup_shrink_node(struct mem_cgroup *memcg, - gfp_t gfp_mask, bool noswap, - pg_data_t *pgdat, - unsigned long *nr_scanned); =20 #ifdef CONFIG_NUMA extern int sysctl_min_unmapped_ratio; diff --git a/mm/vmscan.c b/mm/vmscan.c index 032b14793d91..790b50c78a2e 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -6795,47 +6795,6 @@ unsigned long try_to_free_pages(struct zonelist *zon= elist, int order, =20 #ifdef CONFIG_MEMCG =20 -/* Only used by soft limit reclaim. Do not reuse for anything else. */ -unsigned long mem_cgroup_shrink_node(struct mem_cgroup *memcg, - gfp_t gfp_mask, bool noswap, - pg_data_t *pgdat, - unsigned long *nr_scanned) -{ - struct lruvec *lruvec =3D mem_cgroup_lruvec(memcg, pgdat); - struct scan_control sc =3D { - .nr_to_reclaim =3D SWAP_CLUSTER_MAX, - .target_mem_cgroup =3D memcg, - .may_writepage =3D 1, - .may_unmap =3D 1, - .reclaim_idx =3D MAX_NR_ZONES - 1, - .may_swap =3D !noswap, - }; - - WARN_ON_ONCE(!current->reclaim_state); - - sc.gfp_mask =3D (gfp_mask & GFP_RECLAIM_MASK) | - (GFP_HIGHUSER_MOVABLE & ~GFP_RECLAIM_MASK); - - trace_mm_vmscan_memcg_softlimit_reclaim_begin(sc.gfp_mask, - sc.order, - memcg); - - /* - * NOTE: Although we can get the priority field, using it - * here is not a good idea, since it limits the pages we can scan. - * if we don't reclaim here, the shrink_node from balance_pgdat - * will pick up pages from other mem cgroup's as well. We hack - * the priority and make it zero. - */ - shrink_lruvec(lruvec, &sc); - - trace_mm_vmscan_memcg_softlimit_reclaim_end(sc.nr_reclaimed, memcg); - - *nr_scanned =3D sc.nr_scanned; - - return sc.nr_reclaimed; -} - unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-180.mta0.migadu.com (out-180.mta0.migadu.com [91.218.175.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EDBE404BC5 for ; Tue, 11 Aug 2026 20:32:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480356; cv=none; b=PBF8QZ/iAXy0Clnit9bU85+XNSpqRCs1uZx5tiSzcz0Lz3kZMox87zWvHk60XBXhq5DwO520fzp6MkriKoClP+pT3KJ5dHvpGG54DFMxcGcyRuF0inZKPNWAdz8TTQpyKm1H/AwZWNsRQa2KjzcjIyfadW+TJPchL9g1HnIdxqU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480356; c=relaxed/simple; bh=U9K7f0QJFSHjM3kZX+nRyD7KRwwpOhhPRYAfnJjDSoo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=RKzeLB89hv4jL7DHnodRzmO3O4eOTcsJgG4jX/nPvGGk8HjB1+rshfjbI4NfDai26/YcsaTdVIuItA/iWH7QJ6AHkJ67i9xiIh/CAHdWNgDO2nnQydfIVHModUM7jcS9aN1AmV0cy9U165AytfiI9VTLVo8MGd4D42DrMeFl598= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=q4KemcBI; arc=none smtp.client-ip=91.218.175.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="q4KemcBI" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480353; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=kMaHE95s1N2ro/hCTjeuWjiZbKrQjnsN0zD3VNnyWTQ=; b=q4KemcBI3JXk3fWT/q8wUMtoNfqZJlPmQG1Me6eNlskfW+iCFIzGJNVvpEBAagPzCazkqR 7ouS09BROAyYGe5iZK9V1kMaN3WFVRXDwVzGHJdmOnVN3t8huWXWZkz6iU6N6bskDEjHJb rATBlfgk6qxpHPSnQq6bmoh3y+Lo5WQ= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 4/9] memcg: remove the soft limit reclaim tracepoints Date: Tue, 11 Aug 2026 13:31:58 -0700 Message-ID: <20260811203203.3456029-5-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" mm_vmscan_memcg_softlimit_reclaim_begin and mm_vmscan_memcg_softlimit_reclaim_end were only emitted by mem_cgroup_shrink_node(), which is gone, so they can never fire again. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- include/trace/events/vmscan.h | 14 -------------- 1 file changed, 14 deletions(-) diff --git a/include/trace/events/vmscan.h b/include/trace/events/vmscan.h index b4bf7b8def1f..8a872990b4be 100644 --- a/include/trace/events/vmscan.h +++ b/include/trace/events/vmscan.h @@ -214,13 +214,6 @@ DEFINE_EVENT(mm_vmscan_direct_reclaim_begin_template, = mm_vmscan_memcg_reclaim_be =20 TP_ARGS(gfp_flags, order, memcg) ); - -DEFINE_EVENT(mm_vmscan_direct_reclaim_begin_template, mm_vmscan_memcg_soft= limit_reclaim_begin, - - TP_PROTO(gfp_t gfp_flags, int order, struct mem_cgroup *memcg), - - TP_ARGS(gfp_flags, order, memcg) -); #endif /* CONFIG_MEMCG */ =20 DECLARE_EVENT_CLASS(mm_vmscan_direct_reclaim_end_template, @@ -260,13 +253,6 @@ DEFINE_EVENT(mm_vmscan_direct_reclaim_end_template, mm= _vmscan_memcg_reclaim_end, =20 TP_ARGS(nr_reclaimed, memcg) ); - -DEFINE_EVENT(mm_vmscan_direct_reclaim_end_template, mm_vmscan_memcg_softli= mit_reclaim_end, - - TP_PROTO(unsigned long nr_reclaimed, struct mem_cgroup *memcg), - - TP_ARGS(nr_reclaimed, memcg) -); #endif /* CONFIG_MEMCG */ =20 TRACE_EVENT(mm_shrink_slab_start, --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-189.mta0.migadu.com (out-189.mta0.migadu.com [91.218.175.189]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 549714582CB for ; Tue, 11 Aug 2026 20:32:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.189 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480363; cv=none; b=MJz3FPa73n4Ptn9k4paK/y4A/nbT1S6DKM/Hvh8pH/y57B9eRkVFQtxxGRBqeLnUJVlhZAtQ5SbD/FpbbB4btcsHzZtP/VVtPLIZ9HzZGHrybaUcyFPMMxzs/eoOMXC8HbZAD1DBo9QrN7ZxsrGkdaH27SRYBOrGQj69MR1yTs4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480363; c=relaxed/simple; bh=YqMU9RNZUNzFsU+YziNRY+pkIuQVvc6LRzjzYDGTFVA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=i6EA/ld5Vb3tcCpegNfZv4yqPW8rj4ndiSqSiowcmG/dDjMNYdkaKjtLLKJ1oJYDTdbMRmkK8hSu6WpZTQKjKfz5ZD9f3mF7EmeNcoVMrUbwYoWwGJgTuuHK78EFvax+a12UMpobEeB+lz+Qpr9/NNCRzuvfu0+agRy9odr8MO8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=uN9AUmp6; arc=none smtp.client-ip=91.218.175.189 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="uN9AUmp6" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480358; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=mfoSEEh1wmZsAndftJOLrD9wrMxUb9iE4ebdJd4337w=; b=uN9AUmp6oVwQ9Y1xltyPWBsey4c26TsO6lpQJe3mzugLda498cXA4sDAqL0xIPmFpYHhrF baUc0Yc7yojMKFA+kP4oF9WjOnBtnr74qoSCWjFglB9iVM+arTO+dlCqlzityWDxbV9BFX Hc+5lXXRt8MPFJZqpySkmpMUr20fD6U= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 5/9] memcg: remove the soft limit rbtree Date: Tue, 11 Aug 2026 13:31:59 -0700 Message-ID: <20260811203203.3456029-6-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" With soft limit reclaim gone, the per-node rbtree of cgroups in excess has no readers left. Remove the tree, the helpers maintaining it, and the subsys_initcall that existed only to allocate it. memcg1_check_events() no longer needs to feed it, which also drops the last caller of lru_gen_soft_reclaim(). Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- mm/memcontrol-v1.c | 176 +-------------------------------------------- mm/memcontrol-v1.h | 2 - mm/memcontrol.c | 1 - 3 files changed, 2 insertions(+), 177 deletions(-) diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index b38b8d0f7f51..475f998b7643 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -17,23 +17,6 @@ #include "swap_table.h" #include "memcontrol-v1.h" =20 -/* - * Cgroups above their limits are maintained in a RB-Tree, independent of - * their hierarchy representation - */ - -struct mem_cgroup_tree_per_node { - struct rb_root rb_root; - struct rb_node *rb_rightmost; - spinlock_t lock; -}; - -struct mem_cgroup_tree { - struct mem_cgroup_tree_per_node *rb_tree_per_node[MAX_NUMNODES]; -}; - -static struct mem_cgroup_tree soft_limit_tree __read_mostly; - /* for OOM */ struct mem_cgroup_eventfd_list { struct list_head list; @@ -99,133 +82,6 @@ static struct lockdep_map memcg_oom_lock_dep_map =3D { =20 DEFINE_SPINLOCK(memcg_oom_lock); =20 -static void __mem_cgroup_insert_exceeded(struct mem_cgroup_per_node *mz, - struct mem_cgroup_tree_per_node *mctz, - unsigned long new_usage_in_excess) -{ - struct rb_node **p =3D &mctz->rb_root.rb_node; - struct rb_node *parent =3D NULL; - struct mem_cgroup_per_node *mz_node; - bool rightmost =3D true; - - if (mz->on_tree) - return; - - mz->usage_in_excess =3D new_usage_in_excess; - if (!mz->usage_in_excess) - return; - while (*p) { - parent =3D *p; - mz_node =3D rb_entry(parent, struct mem_cgroup_per_node, - tree_node); - if (mz->usage_in_excess < mz_node->usage_in_excess) { - p =3D &(*p)->rb_left; - rightmost =3D false; - } else { - p =3D &(*p)->rb_right; - } - } - - if (rightmost) - mctz->rb_rightmost =3D &mz->tree_node; - - rb_link_node(&mz->tree_node, parent, p); - rb_insert_color(&mz->tree_node, &mctz->rb_root); - mz->on_tree =3D true; -} - -static void __mem_cgroup_remove_exceeded(struct mem_cgroup_per_node *mz, - struct mem_cgroup_tree_per_node *mctz) -{ - if (!mz->on_tree) - return; - - if (&mz->tree_node =3D=3D mctz->rb_rightmost) - mctz->rb_rightmost =3D rb_prev(&mz->tree_node); - - rb_erase(&mz->tree_node, &mctz->rb_root); - mz->on_tree =3D false; -} - -static void mem_cgroup_remove_exceeded(struct mem_cgroup_per_node *mz, - struct mem_cgroup_tree_per_node *mctz) -{ - unsigned long flags; - - spin_lock_irqsave(&mctz->lock, flags); - __mem_cgroup_remove_exceeded(mz, mctz); - spin_unlock_irqrestore(&mctz->lock, flags); -} - -static unsigned long soft_limit_excess(struct mem_cgroup *memcg) -{ - unsigned long nr_pages =3D page_counter_read(&memcg->memory); - unsigned long soft_limit =3D READ_ONCE(memcg->soft_limit); - unsigned long excess =3D 0; - - if (nr_pages > soft_limit) - excess =3D nr_pages - soft_limit; - - return excess; -} - -static void memcg1_update_tree(struct mem_cgroup *memcg, int nid) -{ - unsigned long excess; - struct mem_cgroup_per_node *mz; - struct mem_cgroup_tree_per_node *mctz; - - if (lru_gen_enabled()) { - if (soft_limit_excess(memcg)) - lru_gen_soft_reclaim(memcg, nid); - return; - } - - mctz =3D soft_limit_tree.rb_tree_per_node[nid]; - if (!mctz) - return; - /* - * Necessary to update all ancestors when hierarchy is used. - * because their event counter is not touched. - */ - for (; memcg; memcg =3D parent_mem_cgroup(memcg)) { - mz =3D memcg->nodeinfo[nid]; - excess =3D soft_limit_excess(memcg); - /* - * We have to update the tree if mz is on RB-tree or - * mem is over its softlimit. - */ - if (excess || mz->on_tree) { - unsigned long flags; - - spin_lock_irqsave(&mctz->lock, flags); - /* if on-tree, remove it */ - if (mz->on_tree) - __mem_cgroup_remove_exceeded(mz, mctz); - /* - * Insert again. mz->usage_in_excess will be updated. - * If excess is 0, no tree ops. - */ - __mem_cgroup_insert_exceeded(mz, mctz, excess); - spin_unlock_irqrestore(&mctz->lock, flags); - } - } -} - -void memcg1_remove_from_trees(struct mem_cgroup *memcg) -{ - struct mem_cgroup_tree_per_node *mctz; - struct mem_cgroup_per_node *mz; - int nid; - - for_each_node(nid) { - mz =3D memcg->nodeinfo[nid]; - mctz =3D soft_limit_tree.rb_tree_per_node[nid]; - if (mctz) - mem_cgroup_remove_exceeded(mz, mctz); - } -} - static u64 mem_cgroup_move_charge_read(struct cgroup_subsys_state *css, struct cftype *cft) { @@ -336,7 +192,7 @@ static void mem_cgroup_threshold(struct mem_cgroup *mem= cg) } } =20 -/* Cgroup1: threshold notifications & softlimit tree updates */ +/* Cgroup1: threshold notifications */ =20 /* * Per memcg event counter is incremented at every pagein/pageout. With TH= P, @@ -405,17 +261,8 @@ static void memcg1_check_events(struct mem_cgroup *mem= cg, int nid) if (IS_ENABLED(CONFIG_PREEMPT_RT)) return; =20 - /* threshold event is triggered in finer grain than soft limit */ - if (unlikely(memcg1_event_ratelimit(memcg, - MEM_CGROUP_TARGET_THRESH))) { - bool do_softlimit; - - do_softlimit =3D memcg1_event_ratelimit(memcg, - MEM_CGROUP_TARGET_SOFTLIMIT); + if (unlikely(memcg1_event_ratelimit(memcg, MEM_CGROUP_TARGET_THRESH))) mem_cgroup_threshold(memcg); - if (unlikely(do_softlimit)) - memcg1_update_tree(memcg, nid); - } } =20 void memcg1_commit_charge(struct folio *folio, struct mem_cgroup *memcg) @@ -2391,22 +2238,3 @@ void memcg1_free_events(struct mem_cgroup *memcg) { free_percpu(memcg->events_percpu); } - -static int __init memcg1_init(void) -{ - int node; - - for_each_node(node) { - struct mem_cgroup_tree_per_node *rtpn; - - rtpn =3D kzalloc_node(sizeof(*rtpn), GFP_KERNEL, node); - - rtpn->rb_root =3D RB_ROOT; - rtpn->rb_rightmost =3D NULL; - spin_lock_init(&rtpn->lock); - soft_limit_tree.rb_tree_per_node[node] =3D rtpn; - } - - return 0; -} -subsys_initcall(memcg1_init); diff --git a/mm/memcontrol-v1.h b/mm/memcontrol-v1.h index 1e394269c613..fd611e66859a 100644 --- a/mm/memcontrol-v1.h +++ b/mm/memcontrol-v1.h @@ -41,7 +41,6 @@ bool memcg1_alloc_events(struct mem_cgroup *memcg); void memcg1_free_events(struct mem_cgroup *memcg); =20 void memcg1_memcg_init(struct mem_cgroup *memcg); -void memcg1_remove_from_trees(struct mem_cgroup *memcg); =20 static inline void memcg1_soft_limit_reset(struct mem_cgroup *memcg) { @@ -98,7 +97,6 @@ static inline bool memcg1_alloc_events(struct mem_cgroup = *memcg) { return true; static inline void memcg1_free_events(struct mem_cgroup *memcg) {} =20 static inline void memcg1_memcg_init(struct mem_cgroup *memcg) {} -static inline void memcg1_remove_from_trees(struct mem_cgroup *memcg) {} static inline void memcg1_soft_limit_reset(struct mem_cgroup *memcg) {} static inline void memcg1_css_offline(struct mem_cgroup *memcg) {} =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 1d3339520809..b68f1f16ae54 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4394,7 +4394,6 @@ static void mem_cgroup_css_free(struct cgroup_subsys_= state *css) =20 vmpressure_cleanup(&memcg->vmpressure); cancel_work_sync(&memcg->high_work); - memcg1_remove_from_trees(memcg); free_shrinker_info(memcg); mem_cgroup_free(memcg); } --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-172.mta0.migadu.com (out-172.mta0.migadu.com [91.218.175.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5FF0F40B10D for ; Tue, 11 Aug 2026 20:32:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480366; cv=none; b=B5AfQNM+N5wF8FmKMe65xcNNRP4bMwZ5ptxYAWLbvpvPcrLW8lbWALJhF2zKhqdsHl/4T644e+J/FVO/X6P+KdQEsizTOhvtI1UdO10ntmPJQqiMwezQT3BF7FxsXfkIUdw6rgFvR8qMWezGiTFguxRp3vC2xnogMdAtM+7ktPM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480366; c=relaxed/simple; bh=sX2noSKb34ahg6g5tpkNnlYlb8jDcEdQTjM16JJvcKE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s27x3GWGlE9+/HlghHRe2iApVDVwha0AkPfFaEi+Wsyb8CyTI8+bNEaLWvLN3oE1UiljIkNZGNKik8BI8KDc4RJcYD+J/vcYxKUSjx7w0Jkikts9FJk5lzMBPs+px0+szCVYGPopvDXAZtFNegETn6xVUfaisHaIeWg/isLA87o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=rqcdEhl9; arc=none smtp.client-ip=91.218.175.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="rqcdEhl9" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480362; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=RBoE71wnBwMOfVKAQwpkXO2zfBhsgRR5WTCUnnYK0XI=; b=rqcdEhl97KNlZGT6AQVoZldHYtPVXRxa52iXODLrzeenxnlhF05ds+4XLpENW7ZnPwYmcu bdVbTk10D1q5qehMWWCSl3vwMyvuL85qTD7qEebW4amHgeE0yZhA3S5VqgwC8gYWBHWkf4 cDLA1UgoMGl3/SBAjSBG+MP6mV/ajA4= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 6/9] memcg: remove lru_gen_soft_reclaim() Date: Tue, 11 Aug 2026 13:32:00 -0700 Message-ID: <20260811203203.3456029-7-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" The soft limit rbtree was the only caller. Dropping it leaves MEMCG_LRU_HEAD unreachable, since nothing else ever rotates a memcg with that op, so remove the op too and update the memcg LRU comment. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Reviewed-by: T.J. Mercier --- include/linux/mmzone.h | 30 +++++++++++------------------- mm/vmscan.c | 16 ++-------------- 2 files changed, 13 insertions(+), 33 deletions(-) diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 94f9c3ff5416..01fabd0ece0d 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -635,35 +635,32 @@ struct lru_gen_mm_walk { * For each node, memcgs are divided into two generations: the old and the * young. For each generation, memcgs are randomly sharded into multiple b= ins * to improve scalability. For each bin, the hlist_nulls is virtually divi= ded - * into three segments: the head, the tail and the default. + * into two segments: the tail and the default. * * An onlining memcg is added to the tail of a random bin in the old gener= ation. * The eviction starts at the head of a random bin in the old generation. = The * per-node memcg generation counter, whose reminder (mod MEMCG_NR_GENS) i= ndexes * the old generation, is incremented when all its bins become empty. * - * There are four operations: - * 1. MEMCG_LRU_HEAD, which moves a memcg to the head of a random bin in i= ts - * current generation (old or young) and updates its "seg" to "head"; - * 2. MEMCG_LRU_TAIL, which moves a memcg to the tail of a random bin in i= ts + * There are three operations: + * 1. MEMCG_LRU_TAIL, which moves a memcg to the tail of a random bin in i= ts * current generation (old or young) and updates its "seg" to "tail"; - * 3. MEMCG_LRU_OLD, which moves a memcg to the head of a random bin in th= e old + * 2. MEMCG_LRU_OLD, which moves a memcg to the head of a random bin in th= e old * generation, updates its "gen" to "old" and resets its "seg" to "defa= ult"; - * 4. MEMCG_LRU_YOUNG, which moves a memcg to the tail of a random bin in = the + * 3. MEMCG_LRU_YOUNG, which moves a memcg to the tail of a random bin in = the * young generation, updates its "gen" to "young" and resets its "seg" = to * "default". * * The events that trigger the above operations are: - * 1. Exceeding the soft limit, which triggers MEMCG_LRU_HEAD; - * 2. The first attempt to reclaim a memcg below low, which triggers + * 1. The first attempt to reclaim a memcg below low, which triggers * MEMCG_LRU_TAIL; - * 3. The first attempt to reclaim a memcg offlined or below reclaimable s= ize + * 2. The first attempt to reclaim a memcg offlined or below reclaimable s= ize * threshold, which triggers MEMCG_LRU_TAIL; - * 4. The second attempt to reclaim a memcg offlined or below reclaimable = size + * 3. The second attempt to reclaim a memcg offlined or below reclaimable = size * threshold, which triggers MEMCG_LRU_YOUNG; - * 5. Attempting to reclaim a memcg below min, which triggers MEMCG_LRU_YO= UNG; - * 6. Finishing the aging on the eviction path, which triggers MEMCG_LRU_Y= OUNG; - * 7. Offlining a memcg, which triggers MEMCG_LRU_OLD. + * 4. Attempting to reclaim a memcg below min, which triggers MEMCG_LRU_YO= UNG; + * 5. Finishing the aging on the eviction path, which triggers MEMCG_LRU_Y= OUNG; + * 6. Offlining a memcg, which triggers MEMCG_LRU_OLD. * * Notes: * 1. Memcg LRU only applies to global reclaim, and the round-robin increm= enting @@ -696,7 +693,6 @@ void lru_gen_exit_memcg(struct mem_cgroup *memcg); void lru_gen_online_memcg(struct mem_cgroup *memcg); void lru_gen_offline_memcg(struct mem_cgroup *memcg); void lru_gen_release_memcg(struct mem_cgroup *memcg); -void lru_gen_soft_reclaim(struct mem_cgroup *memcg, int nid); void max_lru_gen_memcg(struct mem_cgroup *memcg, int nid); bool recheck_lru_gen_max_memcg(struct mem_cgroup *memcg, int nid); void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *p= arent, int nid); @@ -737,10 +733,6 @@ static inline void lru_gen_release_memcg(struct mem_cg= roup *memcg) { } =20 -static inline void lru_gen_soft_reclaim(struct mem_cgroup *memcg, int nid) -{ -} - static inline void max_lru_gen_memcg(struct mem_cgroup *memcg, int nid) { } diff --git a/mm/vmscan.c b/mm/vmscan.c index 790b50c78a2e..71244cf33d59 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -4373,7 +4373,6 @@ bool lru_gen_look_around(struct page_vma_mapped_walk = *pvmw, unsigned int nr) /* see the comment on MEMCG_NR_GENS */ enum { MEMCG_LRU_NOP, - MEMCG_LRU_HEAD, MEMCG_LRU_TAIL, MEMCG_LRU_OLD, MEMCG_LRU_YOUNG, @@ -4395,9 +4394,7 @@ static void lru_gen_rotate_memcg(struct lruvec *lruve= c, int op) new =3D old =3D lruvec->lrugen.gen; =20 /* see the comment on MEMCG_NR_GENS */ - if (op =3D=3D MEMCG_LRU_HEAD) - seg =3D MEMCG_LRU_HEAD; - else if (op =3D=3D MEMCG_LRU_TAIL) + if (op =3D=3D MEMCG_LRU_TAIL) seg =3D MEMCG_LRU_TAIL; else if (op =3D=3D MEMCG_LRU_OLD) new =3D get_memcg_gen(pgdat->memcg_lru.seq); @@ -4411,7 +4408,7 @@ static void lru_gen_rotate_memcg(struct lruvec *lruve= c, int op) =20 hlist_nulls_del_rcu(&lruvec->lrugen.list); =20 - if (op =3D=3D MEMCG_LRU_HEAD || op =3D=3D MEMCG_LRU_OLD) + if (op =3D=3D MEMCG_LRU_OLD) hlist_nulls_add_head_rcu(&lruvec->lrugen.list, &pgdat->memcg_lru.fifo[ne= w][bin]); else hlist_nulls_add_tail_rcu(&lruvec->lrugen.list, &pgdat->memcg_lru.fifo[ne= w][bin]); @@ -4489,15 +4486,6 @@ void lru_gen_release_memcg(struct mem_cgroup *memcg) } } =20 -void lru_gen_soft_reclaim(struct mem_cgroup *memcg, int nid) -{ - struct lruvec *lruvec =3D get_lruvec(memcg, nid); - - /* see the comment on MEMCG_NR_GENS */ - if (READ_ONCE(lruvec->lrugen.seg) !=3D MEMCG_LRU_HEAD) - lru_gen_rotate_memcg(lruvec, MEMCG_LRU_HEAD); -} - bool recheck_lru_gen_max_memcg(struct mem_cgroup *memcg, int nid) { struct lruvec *lruvec =3D get_lruvec(memcg, nid); --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-185.mta0.migadu.com (out-185.mta0.migadu.com [91.218.175.185]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0FCFC403EA5 for ; Tue, 11 Aug 2026 20:32:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.185 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480369; cv=none; b=gM3nqFr5KLuFaMRmtyVnZK5AeJubZxhhmhTrkualTG6TOgj4gi6wMW+gEhedn9ioS7RNq7Mz9Ceuf2p+cAb1pgZIcDcVV02Q6FwbhK4KPDlximwjzPON9KIiAbNzGiFVrlMbPhbBSFBvhvo1koUOJo4/wGR+3QOkPn/WKhcqt4s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480369; c=relaxed/simple; bh=01J9kJaMy8MmdGZHRnNMI7PSucR8E7CkO327BXjw33I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gXU+0btXrNEowxZfxgBOBX96hQvCtDHvQRdGHlXCeYDHxqs6fkD0hW24SqpZ0aotncKWnsd/JfogGfafYQn23ouEVK83PYsmqWGSxu5j9cVlkatjVXF74evo19OlDffUg18yOSSWvKa4B1cGvTDRU7aSQoE7GVHN+sUIdT6Bkk8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=tjWLYDoX; arc=none smtp.client-ip=91.218.175.185 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="tjWLYDoX" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480366; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=5Q4GswXsYa23HjspIYsiXDi4Yy+VY6y3TioX/Zf5Hfc=; b=tjWLYDoXNvQDIAuW73cOM4BIp4z0OlAyBGCwh4bxtqlJ2fpxssixILMpQuX6ySxZ9dsGWm t0/Kmbh9XwTL5oDehABSwwMvkVmGStIVxebzieS00SwnrQ5eUz5rJet0XjbwZp90GrwcOb 9/zoKxC2UKg06CoyX/oIWVH1BJFijFg= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 7/9] memcg: remove the per-node soft limit tree fields Date: Tue, 11 Aug 2026 13:32:01 -0700 Message-ID: <20260811203203.3456029-8-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" tree_node, usage_in_excess and on_tree only existed for the soft limit rbtree. They also doubled as the buffer between the read-mostly head of struct mem_cgroup_per_node and its update-often tail, so replace them with the explicit padding that CONFIG_MEMCG_V1=3Dn already used. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- include/linux/memcontrol.h | 13 ------------- 1 file changed, 13 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 7b02f1b3bb88..ce24e04967d8 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -95,20 +95,7 @@ struct mem_cgroup_per_node { struct lruvec_stats *lruvec_stats; struct shrinker_info __rcu *shrinker_info; =20 -#ifdef CONFIG_MEMCG_V1 - /* - * Memcg-v1 only stuff in middle as buffer between read mostly fields - * and update often fields to avoid false sharing. If v1 stuff is - * not present, an explicit padding is needed. - */ - - struct rb_node tree_node; /* RB tree node */ - unsigned long usage_in_excess;/* Set to the value by which */ - /* the soft limit is exceeded*/ - bool on_tree; -#else CACHELINE_PADDING(_pad1_); -#endif =20 /* Fields which get updated often at the end. */ struct lruvec lruvec; --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-175.mta1.migadu.com (out-175.mta1.migadu.com [95.215.58.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 98C2F403E8A for ; Tue, 11 Aug 2026 20:32:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480379; cv=none; b=pagKQkpi4gamehbbS87SMbwqvXl6LWCKSnazK/ay1Svx+rjmE4SGrAE/epaL00uiKzz5bdhvYOdvg8BE5WJ05d4iQKYU6Rcb36SVs2rWEmnH8KYOqAwUjb3YeYu1oXEYP1aonZVRN9jf9AX5AhYApsA6nm2+Mq9D28Fl7Skp5yM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480379; c=relaxed/simple; bh=7gVWnR3Zrj0TkNt8a0EdnW9fGTbUaafAV3U3q/Me5tE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=f27GaAuWGsG0TDHMYz7znbB96tNR/kAI0K0kD6sOUZ1I1BBLMNEANvBNfkWm9od/DKoR5QaVLrM6SOuOnDOwfvLT9kTRMTbj6Z5jTcUPMeCJ02ig4vKi2V9eE3xoPjqKRL5Tu7HNcDcJPAPfksQozoQTJEPlq5AX1lzlKzRSi20= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=af9Yzsml; arc=none smtp.client-ip=95.215.58.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="af9Yzsml" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480375; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=CN2l9vWwoY/I0+/xUhc13X8fJjrgmTbCFVIi09J5CeY=; b=af9YzsmlUM8QyJokSTb4jl52AGvGhzDmHQV45zQAnc+V3AZGcRml/gBTj66J2r2qxPKzv8 eN/7vvJORDm58oraWpUHvex7ga5Jyw3DfjmD9HzwJMXBlnKuHfY767ETMMCySch/fkEOPa AMO+TEodrg0AxugIsaFVYFIbmwzFnZA= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 8/9] memcg: remove mem_cgroup->soft_limit Date: Tue, 11 Aug 2026 13:32:02 -0700 Message-ID: <20260811203203.3456029-9-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Nothing reads it anymore, so the field and the helper that reset it on css alloc and css reset can go. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- include/linux/memcontrol.h | 2 -- mm/memcontrol-v1.h | 6 ------ mm/memcontrol.c | 2 -- 3 files changed, 10 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index ce24e04967d8..526da1d869ed 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -275,8 +275,6 @@ struct mem_cgroup { =20 struct memcg1_events_percpu __percpu *events_percpu; =20 - unsigned long soft_limit; - /* protected by memcg_oom_lock */ bool oom_lock; int under_oom; diff --git a/mm/memcontrol-v1.h b/mm/memcontrol-v1.h index fd611e66859a..f48d0e22e615 100644 --- a/mm/memcontrol-v1.h +++ b/mm/memcontrol-v1.h @@ -42,11 +42,6 @@ void memcg1_free_events(struct mem_cgroup *memcg); =20 void memcg1_memcg_init(struct mem_cgroup *memcg); =20 -static inline void memcg1_soft_limit_reset(struct mem_cgroup *memcg) -{ - WRITE_ONCE(memcg->soft_limit, PAGE_COUNTER_MAX); -} - struct cgroup_taskset; void memcg1_css_offline(struct mem_cgroup *memcg); =20 @@ -97,7 +92,6 @@ static inline bool memcg1_alloc_events(struct mem_cgroup = *memcg) { return true; static inline void memcg1_free_events(struct mem_cgroup *memcg) {} =20 static inline void memcg1_memcg_init(struct mem_cgroup *memcg) {} -static inline void memcg1_soft_limit_reset(struct mem_cgroup *memcg) {} static inline void memcg1_css_offline(struct mem_cgroup *memcg) {} =20 static inline bool memcg1_oom_prepare(struct mem_cgroup *memcg, bool *lock= ed) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index b68f1f16ae54..ba3ef821553d 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4222,7 +4222,6 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *pare= nt_css) return ERR_CAST(memcg); =20 page_counter_set_high(&memcg->memory, PAGE_COUNTER_MAX); - memcg1_soft_limit_reset(memcg); #ifdef CONFIG_ZSWAP memcg->zswap_max =3D PAGE_COUNTER_MAX; WRITE_ONCE(memcg->zswap_writeback, true); @@ -4429,7 +4428,6 @@ static void mem_cgroup_css_reset(struct cgroup_subsys= _state *css) page_counter_set_min(&memcg->memory, 0); page_counter_set_low(&memcg->memory, 0); page_counter_set_high(&memcg->memory, PAGE_COUNTER_MAX); - memcg1_soft_limit_reset(memcg); page_counter_set_high(&memcg->swap, PAGE_COUNTER_MAX); memcg_wb_domain_size_changed(memcg); } --=20 2.53.0-Meta From nobody Tue Sep 29 06:05:24 2026 Received: from out-184.mta1.migadu.com (out-184.mta1.migadu.com [95.215.58.184]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 98C6A386C17 for ; Tue, 11 Aug 2026 20:33:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.184 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480396; cv=none; b=GwWEEuNJHOB5P9rtFHg0Imgi2dq3S+w+zO1ADmws6LyDqM9buRbxCN7ZLz5wyhtRb7+XqhsqHyk0O7gok3kgSdubREFUb9bne3x3rqEXgXGi/IKvAGnA5Yui+8JQO9UKRQWICWO9Wq5jFzAi0B1oXHylThmcif+GtV77uuuqrns= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786480396; c=relaxed/simple; bh=fEyOzM6T+7pQ0yygBxA49pbQ0bITR05mOXP8WTqab3I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PtM2k5xMfiwhtaRf+Td/OFWvFIuwWvt37L4tzDDmzIkgB7rvhZxOrlLbRKPgpvnISmkuJctTMBNEn0O4Zk1U1AUcTfSjWefuF2UDTYEHQHc37fpoaibNhxhAxZLxVv9cJTZ00D4mZq/PYV8Rmitc2C2B00onF8b8ecpb/4bw9pQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=UQ+63i7M; arc=none smtp.client-ip=95.215.58.184 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="UQ+63i7M" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1786480392; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=lhLQTHM31HiibodbQJGZPDVgE2DmcphLRRks8z+Fmcw=; b=UQ+63i7MGIYBEySTOISWQSlyGmVRykIdbn6QYgoo951mXXsO1jO4SD0VHvfVVsyZivVyLG toSZNLG94NUFaJ6zxz2k1cYdeHR2HgeMPTytvvBQeOOK8GquY0fOgUf0MaHDuKkebOfMxr heUlhJ7ebcAnlddFji+a+M5Kj+irWMw= From: Shakeel Butt To: Andrew Morton Cc: Michal Hocko , Johannes Weiner , Roman Gushchin , Muchun Song , David Hildenbrand , Lorenzo Stoakes , Kairui Song , Qi Zheng , Barry Song , Axel Rasmussen , Meta kernel team , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 9/9] memcg: simplify v1 event ratelimiting Date: Tue, 11 Aug 2026 13:32:03 -0700 Message-ID: <20260811203203.3456029-10-shakeel.butt@linux.dev> In-Reply-To: <20260811203203.3456029-1-shakeel.butt@linux.dev> References: <20260811203203.3456029-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" Thresholds are the only periodic v1 event left, so the target enum, the per-cpu target array and the switch in memcg1_event_ratelimit() all collapse to a single counter. memcg1_check_events() no longer needs a node id either, which lets memcg1_uncharge_batch() drop its nid argument and struct uncharge_gather drop the field feeding it. Signed-off-by: Shakeel Butt Acked-by: Lorenzo Stoakes (ARM) Acked-by: Michal Hocko --- mm/memcontrol-v1.c | 43 +++++++++++-------------------------------- mm/memcontrol-v1.h | 4 ++-- mm/memcontrol.c | 4 +--- 3 files changed, 14 insertions(+), 37 deletions(-) diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 475f998b7643..bf2c7d53b01b 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -200,15 +200,9 @@ static void mem_cgroup_threshold(struct mem_cgroup *me= mcg) * to trigger some periodic events. This is straightforward and better * than using jiffies etc. to handle periodic memcg event. */ -enum mem_cgroup_events_target { - MEM_CGROUP_TARGET_THRESH, - MEM_CGROUP_TARGET_SOFTLIMIT, - MEM_CGROUP_NTARGETS, -}; - struct memcg1_events_percpu { unsigned long nr_page_events; - unsigned long targets[MEM_CGROUP_NTARGETS]; + unsigned long threshold_target; }; =20 static void memcg1_charge_statistics(struct mem_cgroup *memcg, int nr_page= s) @@ -225,43 +219,28 @@ static void memcg1_charge_statistics(struct mem_cgrou= p *memcg, int nr_pages) } =20 #define THRESHOLDS_EVENTS_TARGET 128 -#define SOFTLIMIT_EVENTS_TARGET 1024 =20 -static bool memcg1_event_ratelimit(struct mem_cgroup *memcg, - enum mem_cgroup_events_target target) +static bool memcg1_event_ratelimit(struct mem_cgroup *memcg) { unsigned long val, next; =20 val =3D __this_cpu_read(memcg->events_percpu->nr_page_events); - next =3D __this_cpu_read(memcg->events_percpu->targets[target]); + next =3D __this_cpu_read(memcg->events_percpu->threshold_target); /* from time_after() in jiffies.h */ if ((long)(next - val) < 0) { - switch (target) { - case MEM_CGROUP_TARGET_THRESH: - next =3D val + THRESHOLDS_EVENTS_TARGET; - break; - case MEM_CGROUP_TARGET_SOFTLIMIT: - next =3D val + SOFTLIMIT_EVENTS_TARGET; - break; - default: - break; - } - __this_cpu_write(memcg->events_percpu->targets[target], next); + __this_cpu_write(memcg->events_percpu->threshold_target, + val + THRESHOLDS_EVENTS_TARGET); return true; } return false; } =20 -/* - * Check events in order. - * - */ -static void memcg1_check_events(struct mem_cgroup *memcg, int nid) +static void memcg1_check_events(struct mem_cgroup *memcg) { if (IS_ENABLED(CONFIG_PREEMPT_RT)) return; =20 - if (unlikely(memcg1_event_ratelimit(memcg, MEM_CGROUP_TARGET_THRESH))) + if (unlikely(memcg1_event_ratelimit(memcg))) mem_cgroup_threshold(memcg); } =20 @@ -271,7 +250,7 @@ void memcg1_commit_charge(struct folio *folio, struct m= em_cgroup *memcg) =20 local_irq_save(flags); memcg1_charge_statistics(memcg, folio_nr_pages(folio)); - memcg1_check_events(memcg, folio_nid(folio)); + memcg1_check_events(memcg); local_irq_restore(flags); } =20 @@ -344,7 +323,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_= cluster_info *ci) VM_WARN_ON_IRQS_ENABLED(); memcg1_charge_statistics(memcg, -folio_nr_pages(folio)); preempt_enable_nested(); - memcg1_check_events(memcg, folio_nid(folio)); + memcg1_check_events(memcg); =20 rcu_read_unlock(); obj_cgroup_put(objcg); @@ -398,14 +377,14 @@ void memcg1_swapin(struct folio *folio) #endif =20 void memcg1_uncharge_batch(struct mem_cgroup *memcg, unsigned long pgpgout, - unsigned long nr_memory, int nid) + unsigned long nr_memory) { unsigned long flags; =20 local_irq_save(flags); count_memcg_events(memcg, PGPGOUT, pgpgout); __this_cpu_add(memcg->events_percpu->nr_page_events, nr_memory); - memcg1_check_events(memcg, nid); + memcg1_check_events(memcg); local_irq_restore(flags); } =20 diff --git a/mm/memcontrol-v1.h b/mm/memcontrol-v1.h index f48d0e22e615..b9a21f0fd2c3 100644 --- a/mm/memcontrol-v1.h +++ b/mm/memcontrol-v1.h @@ -59,7 +59,7 @@ void memcg1_oom_recover(struct mem_cgroup *memcg); =20 void memcg1_commit_charge(struct folio *folio, struct mem_cgroup *memcg); void memcg1_uncharge_batch(struct mem_cgroup *memcg, unsigned long pgpgout, - unsigned long nr_memory, int nid); + unsigned long nr_memory); =20 void memcg1_stat_format(struct mem_cgroup *memcg, struct seq_buf *s); void reparent_memcg1_state_local(struct mem_cgroup *memcg, struct mem_cgro= up *parent); @@ -107,7 +107,7 @@ static inline void memcg1_commit_charge(struct folio *f= olio, =20 static inline void memcg1_uncharge_batch(struct mem_cgroup *memcg, unsigned long pgpgout, - unsigned long nr_memory, int nid) {} + unsigned long nr_memory) {} =20 static inline void memcg1_stat_format(struct mem_cgroup *memcg, struct seq= _buf *s) {} =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index ba3ef821553d..44ef376d657b 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5293,7 +5293,6 @@ struct uncharge_gather { unsigned long nr_memory; unsigned long pgpgout; unsigned long nr_kmem; - int nid; }; =20 static inline void uncharge_gather_clear(struct uncharge_gather *ug) @@ -5316,7 +5315,7 @@ static void uncharge_batch(const struct uncharge_gath= er *ug) memcg1_oom_recover(memcg); } =20 - memcg1_uncharge_batch(memcg, ug->pgpgout, ug->nr_memory, ug->nid); + memcg1_uncharge_batch(memcg, ug->pgpgout, ug->nr_memory); rcu_read_unlock(); =20 /* drop reference from uncharge_folio */ @@ -5345,7 +5344,6 @@ static void uncharge_folio(struct folio *folio, struc= t uncharge_gather *ug) uncharge_gather_clear(ug); } ug->objcg =3D objcg; - ug->nid =3D folio_nid(folio); =20 /* pairs with obj_cgroup_put in uncharge_batch */ obj_cgroup_get(objcg); --=20 2.53.0-Meta