From nobody Tue Sep 29 13:19:51 2026 Received: from mail-ot1-f49.google.com (mail-ot1-f49.google.com [209.85.210.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE85C43F4C2 for ; Fri, 7 Aug 2026 20:21:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134064; cv=none; b=RnaDjo08IxbO+DEwwndKJ0CrIFswvfXK4EPfVnVBdS+D6Wo2UjTWnKdQ9WPCumr94aA0liRxAKKk4YbxO48TRFZHsd5HdlAB7CXOFTk06y8bUEDmpTfI+4GlICokGoSRbhthATa2DPKP+qAgKGECDavV0Tr0wWV648Nyd8Tadlk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134064; c=relaxed/simple; bh=x6+zULnjVSGwKnGUnc1/UzJDwmZxjXg8KnI3HpUc3JI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=m3ZllitqKG7Sm5Ysnz0EOGjEvzckHLjsYjLfvp5smJeFYfpLZEYc2faxHFRSX1Pq57Mw6jJa3DIh/ZxXCVFHyoRHk29SOfx1z9Vfm25LudowXgobn3MEyCDWjKAu2U6qfWorKED2jYJ1O7lIjJOc97K/UJl0/sU+BaiSJEMUvJE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=D0Y2G/E+; arc=none smtp.client-ip=209.85.210.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="D0Y2G/E+" Received: by mail-ot1-f49.google.com with SMTP id 46e09a7af769-7eb64085c45so2957146a34.2 for ; Fri, 07 Aug 2026 13:21:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134061; x=1786738861; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HwNwt7qn0Dj2QjnweBxCilOvMKex3+n+ntJRoPVYrE4=; b=D0Y2G/E+peOkW9difLI0B2hdOxQhXBum0CQfAU5NerFA7QrCfSs4+JUHybGD8PI3hx 8veDigyVuMn1Uu3q4VN6zCu72zVR0USFLNecNETtes/YY2sbWz9SN8ZlURH4fYeCPhmn KrnrEkFA80zfW5sQreM5DW26cG4/xs8W1YC659sq6fyMGR57apg70E8K7fS61ck+x80y j3qlPzekZj4NgSCKdLJY4nQ07fF53YoiFvdOJbZPZHQsampYJTZQObl7y3tbBjewqb9M Fz1dnzLal2Vja4UGC7xc+LWKi52qJ4QIPgN5PiFB6M8eiar2b5BCUb5QFa9VTZmo0Jhb b/Ng== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134061; x=1786738861; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=HwNwt7qn0Dj2QjnweBxCilOvMKex3+n+ntJRoPVYrE4=; b=TjZQ1jb3gUFS0YVuzrOpuhf50GXuxo2A9/tksno5Vwo75iEJHGapCqOM8/DCaAax/b GFUeIHub6/dFnA12FGaax4XRq8h/yDuKValLj43sJlC7QRZOzj3IfnC9hDit63HAQilD Hrm0I40GTi8vejUG82005L2g2NyWfoi0W7LiMLRKrUrdyCL6t+iHFFBOH+CGSwF4btDy QFw/WzXPggiQqVX3h44ysswYHWZmQfbVIA6J8qj47vZz3BgdlM8Cgetxe4D8rnvM6Gz1 uZu+Dqs6dxgtI7NGZuImJVC6mwARi8syWXMMeznPYGUVWgNZaclCrjffzQ5Rg6B5Z2Vr BESA== X-Forwarded-Encrypted: i=1; AHgh+RrqKn8n7gSNxW7g4N8TIcpXDik0CMwvDwZk/+kEdvWYvbWnAca68ef7Oh6JUc3+scvI76ZHDic6o9uMceI=@vger.kernel.org X-Gm-Message-State: AOJu0Yz3XXA/eurQJf6lT8JUtpNRDrq0w/tW2LOBW7qyumMKUJh2kOvP UCvjBw+4u6NoMN1rRwR44dlSCYKNa6NOHnI39T5b1u6pGCf3wF8H0jUx X-Gm-Gg: AR+sD118mnLGY+cRwFEoM5hzFVACGPLXxO223Kvz/OMAe8vviZtAJsPyW0zFXL/EGvS GdEcpPPL14mLuXxUvLHXMwON6ooV3EcRZH2pgU46kP3vJNBdsEdR05iNuQfzMfia4T/Bl8UR9gZ 0LJswhUlPnw3gudQ6TD9Op4Zeta+QqMvWC7LJ4o2ldzd2dpbay2k8/4KCJDVDi1cMA8Yaun3R2D MTIYyvjh3L+FI2de5lwVVxrmMXD6CAtEvbKnSsItdKjrqr7D6qvIVKOju9qtDlZGTI2+KftVC8a fPkoT6Ho2V1AqFGjtQ5qeIvvOUFbVxGJH9Ik7ELRVk+pRmHEzcKYUCvvbi45dLXr5XKDpBTRLMP 5UYQJUS+mhmjgZIIC3/80C8WLiFuoOa0KceSrcCukrJgidpJe9um2Vlj8f0TRlN/9cE/9endtFe CunPmmMLLSKSLFLkIdDzIk5/frDkCj4wzpdl3ON/bWD/LBnt8YjL8e/dc3BW57kOg6j+NQRp6zQ pCD2oGCIOBLG4n0zzVT0Ut4tQM/ug== X-Received: by 2002:a05:6820:1791:b0:6ae:ab01:9199 with SMTP id 006d021491bc7-6b042266f4fmr1570536eaf.34.1786134061552; Fri, 07 Aug 2026 13:21:01 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:25::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02bc2631esm3229086eaf.4.2026.08.07.13.21.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:00 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 01/14] mm/memcontrol: Introduce cgroup.memory=tiered_limits boot parameter Date: Fri, 7 Aug 2026 13:20:44 -0700 Message-ID: <20260807202059.2620949-2-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce a "tiered_limits" option for the cgroup.memory=3D kernel commandline parameter to enable tier-proportional scaling and enforcing of the memory cgroup controller limits memory.{min, low, high}. Since mem_cgroup_tiered_limits() will become a hotpath in the later commits to gate charging, demotion, and promotion decisions, use a static key so that cgroups not using tier-aware-memcg limits has minimal overhead. Enable it by adding to the kernel command line: cgroup.memory=3Dtiered_limits The option is boot-time only, since flipping the bit at runtime could leave charges uncharged in the future, or uncharges for folios that were never charged. This feature is incompatible with cgroup v1, and wil raise a single warning statement if a system booted with tiered limits mounts a legacy cgroup: [XXX] cgroup.memory=3Dtiered_limits should not be enabled with cgroupv1 Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 18 ++++++++++++++++++ mm/memcontrol.c | 19 +++++++++++++++++++ 2 files changed, 37 insertions(+) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 2118d5b33d051..dce03df7eae05 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -530,6 +530,19 @@ static inline bool mem_cgroup_disabled(void) return !cgroup_subsys_enabled(memory_cgrp_subsys); } =20 +#ifdef CONFIG_NUMA +DECLARE_STATIC_KEY_FALSE(memcg_tiered_limits_key); +static inline bool mem_cgroup_tiered_limits(void) +{ + return static_branch_unlikely(&memcg_tiered_limits_key); +} +#else +static inline bool mem_cgroup_tiered_limits(void) +{ + return false; +} +#endif + static inline void mem_cgroup_protection(struct mem_cgroup *root, struct mem_cgroup *memcg, unsigned long *min, @@ -1083,6 +1096,11 @@ static inline bool mem_cgroup_disabled(void) return true; } =20 +static inline bool mem_cgroup_tiered_limits(void) +{ + return false; +} + static inline void memcg_memory_event(struct mem_cgroup *memcg, enum memcg_memory_event event) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 29330f5f9d4eb..cefe33b5fd285 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -320,6 +320,13 @@ EXPORT_SYMBOL(memcg_kmem_online_key); DEFINE_STATIC_KEY_FALSE(memcg_bpf_enabled_key); EXPORT_SYMBOL(memcg_bpf_enabled_key); =20 +#ifdef CONFIG_NUMA +DEFINE_STATIC_KEY_FALSE(memcg_tiered_limits_key); + +/* Tier-proportional scaling of memory controller limits enabled? */ +static bool cgroup_memory_tiered_limits __ro_after_init; +#endif + /** * get_mem_cgroup_css_from_folio - acquire a css of the memcg associated w= ith a folio * @folio: folio of interest @@ -4202,6 +4209,9 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *pare= nt_css) struct mem_cgroup *memcg, *old_memcg; bool memcg_on_dfl =3D cgroup_subsys_on_dfl(memory_cgrp_subsys); =20 + if (mem_cgroup_tiered_limits() && !memcg_on_dfl) + pr_warn_once("cgroup.memory=3Dtiered_limits should not be enabled with c= groupv1\n"); + old_memcg =3D set_active_memcg(parent); memcg =3D mem_cgroup_alloc(parent); set_active_memcg(old_memcg); @@ -5584,6 +5594,10 @@ static int __init cgroup_memory(char *s) cgroup_memory_nokmem =3D true; if (!strcmp(token, "nobpf")) cgroup_memory_nobpf =3D true; +#ifdef CONFIG_NUMA + if (!strcmp(token, "tiered_limits")) + cgroup_memory_tiered_limits =3D true; +#endif } return 1; } @@ -5630,6 +5644,11 @@ int __init mem_cgroup_init(void) memcg_pn_cachep =3D KMEM_CACHE(mem_cgroup_per_node, SLAB_PANIC | SLAB_HWCACHE_ALIGN); =20 +#ifdef CONFIG_NUMA + if (cgroup_memory_tiered_limits) + static_branch_enable(&memcg_tiered_limits_key); +#endif + return 0; } =20 --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-ot1-f48.google.com (mail-ot1-f48.google.com [209.85.210.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A0BFD42640D for ; Fri, 7 Aug 2026 20:21:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.48 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134065; cv=none; b=pNvO+AKmlRRcIEYP83/IyZHp7zYbrQyQj476lav69Go2H28pQlzWfkTu98zM4tQh4z/KTJJxbpwF1rDawH8OYiaVFl/dKOYuv2DrO6qH+LKdaw5Dn+f2+lUep507NpYh1E/ug95APoaYeY2ps79aRAL9CRGQ7AV9rOaD32x8CiA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134065; c=relaxed/simple; bh=lwHpAgFN6fA3DPa6IS11A3vLJx+TAT5kwwP52j6vCDQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uWEv+ozlOBGwcQAGKVZaxnsKDdSDqxPG3ekQSdF/txcyi1S8Lb1KGLCiUPVsdpYrXeABrtHcfsaN6zqLQK/K2QE29LNiGnvmr37WhKwlocIy9WAq9sEPdp2Sxf/wGZFfG4YcYEVFSAU5Zbw1I43Ifef1ux+67LRk4cSALboLMtk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=mw1qNUEP; arc=none smtp.client-ip=209.85.210.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="mw1qNUEP" Received: by mail-ot1-f48.google.com with SMTP id 46e09a7af769-7eb68bdf53aso1531200a34.3 for ; Fri, 07 Aug 2026 13:21:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134063; x=1786738863; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=uIJ1lfVFAjreFSGxK42rkml7ixTeVT5CBMzYSJnOvv4=; b=mw1qNUEPm7EcjyCBraD6oiaQghMYIqN8PN3ximGJj5onVdu5N8b4U6GGane6XRAIvh TuMWBmmQXV7AJrNViqHUl4mbhV2WGsWHwF2AOC7Q+Duon+cVVf+Jr2I1IfYqTH3B4QL4 PG2HrO8c9IwE/Wf4L06at7kdGbyfsuE8sfryS1OG1rq90gQL+AK1j8uu7PFh9tA1qYKi oeuBzYS6AmhOd8ulRCfU7+9c2zmYadCvIngHfjHZSsCBkJlfV8YStvgS5SqIWnWKD3Mo Ci7eXCFisZzJRZKeHUBlxwh/ttWU91wIINJ9MvunZfZ9a/PIO1QOGFWjj1o9OFOIUeZ+ CHGQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134063; x=1786738863; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=uIJ1lfVFAjreFSGxK42rkml7ixTeVT5CBMzYSJnOvv4=; b=nE+3IahuUGGsW7zP5VIuCGdmOVgcuKyWvIajLs6i2rb5h0zfA1ncBK4bYRmJToyzoc /dgrcUDJ2KCrD5RqTxexQ8N2wqHYf0mOVxpKjNbOOS8hUHf92uTGy1YYKgz/SFTWVRYA crKNOlfOLvikq5SIvLNl+vvnE49A9ntwZSsc1aehdBYHJQv8ZCH/7gOvr3eIollTAaix 2TuNjMm7BHmn0RuPCLsjD3/pb6fzG7enhMaN+mtHYZxJiSyBKx+sCBiDuFaTKWbd+EQQ rdOw+lYgstCm1qw2Yd67AFHh0dwv6dGRqPENq1lBizsoA7xLpFWqhGj+inuShY9nU/7X 37HA== X-Forwarded-Encrypted: i=1; AHgh+RoI5WcNk2asc3mjrJDqpAh3+dAXt+6a6PptS9angMD+2VP5/dcF1CIf0AaCYLwL81U2ROCLHmXe6FRwL9Q=@vger.kernel.org X-Gm-Message-State: AOJu0YxPNToHHrqnwCglGJ1mX9TSAHC1P5w8WhgARCZrYoH7UE3cSrip 2LAot8UhHzjTHNnJ/K+kH72TLcShuUvZI1fh3zMhshKxRfCsiStwqyvK X-Gm-Gg: AR+sD11DgXiJ7Wd6xAw7VqaHLt9pt9M0ZEfrKpEkNg5AzB9Qb1Iu1l675geLHmi35qt WRkY/fvOLVeOevSt/ReljI2Z761CwZHoT6D4Drpi0YEf7ZUv6JmqeGXht1rDR+lIuIhSBnP6i5j 3HP0vtnlbSbTlqKwkmvayuKZSz80DSuf0oH90ALO4KsX7F0M6jbpg30L51+Td4nZQbcZ6Hj/p4A bcyqQzS4vcJV0qqZciPMJP85eOpJzKVci9F6HjRWw2qMVxDOTk2JdVZs88ZkzsqRIlPBmqQeADA 36W9Lz3ARVgEXywowy0OntMJ1fkWydbvIMQHx3Scps33LdTWYcD5dPsNbo4g19M8QVb4x3dZLWE eYzwVrM6J8xa2wmbxsnJ1/Kp4+RfwKJ/iA+dPEBuwaxvtub3idSRlfVdMsgGU6BkvdEKKTEygMk 0wmC7KV/42OJpnQqENXhktdFsWGFOs+8OhVySq08qSVDcnMsA5vtusqiWCXiJbtZP9hYpmGGmUu Oo5Oo6h295wc5QVkRE= X-Received: by 2002:a05:6820:8118:b0:6ae:55f6:2cdb with SMTP id 006d021491bc7-6b041ee0996mr1744079eaf.1.1786134063536; Fri, 07 Aug 2026 13:21:03 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:46::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1a76701sm2665817fac.5.2026.08.07.13.21.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:03 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 02/14] mm/memcontrol: Refactor page_counter charging in try_charge_memcg Date: Fri, 7 Aug 2026 13:20:45 -0700 Message-ID: <20260807202059.2620949-3-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In preparation for adding charging and uncharging of a new page_counter toptier to try_charge_memcg, refactor the code so that it is easier to distinguish between the memcg v1/v2 cases. No functional changes intended. Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 20 ++++++++++++-------- 1 file changed, 12 insertions(+), 8 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index cefe33b5fd285..ec28512de6a23 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2668,18 +2668,22 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, batch =3D nr_pages; =20 reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP; - if (!do_memsw_account() || - page_counter_try_charge(&memcg->memsw, batch, &counter)) { - if (page_counter_try_charge(&memcg->memory, batch, &counter)) - goto done_restock; - if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, batch); - mem_over_limit =3D mem_cgroup_from_counter(counter, memory); - } else { + + if (do_memsw_account() && + !page_counter_try_charge(&memcg->memsw, batch, &counter)) { mem_over_limit =3D mem_cgroup_from_counter(counter, memsw); reclaim_options &=3D ~MEMCG_RECLAIM_MAY_SWAP; + goto reclaim; } =20 + if (page_counter_try_charge(&memcg->memory, batch, &counter)) + goto done_restock; + + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, batch); + mem_over_limit =3D mem_cgroup_from_counter(counter, memory); + +reclaim: if (batch > nr_pages) { batch =3D nr_pages; goto retry; --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-oi1-f169.google.com (mail-oi1-f169.google.com [209.85.167.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4D105443303 for ; Fri, 7 Aug 2026 20:21:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134067; cv=none; b=PC47RhWLdG8jTRRCz4DotcAd2hjSl/U829l7g9Pud0YmhrUAQEDIqYcUKcObuKSdK+CFknmS1cdsT2RgQOoG+mvFEbG12fjo/ofWk/W1R8FkWT3pOnuS81w2klVcc2xe9gCQQQPrZsw0hYS0pjk7lJR6mpT+idFSVFl2+87m7wQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134067; c=relaxed/simple; bh=w0aiWvuXLSFHMY0mV8BwfkcBTVFPsPvUsMB0BenoBF8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=U6eec3ip4EewaNYA4CEhV1pFXPK4DKtBODCPItin5q0oNzJbqkpNBu6NrmKB4AY+7lKOdZAxVK9tnrixEd0Bzfo0sX3G2PA/CoJiDwscGkXn9RbEVneJDnINlPkep9kEiUMDuq5QPaZ9M+PfHgUut1960oX07PmKka6+uhrcnWk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=MuyuqUWP; arc=none smtp.client-ip=209.85.167.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="MuyuqUWP" Received: by mail-oi1-f169.google.com with SMTP id 5614622812f47-4af81963f35so1491088b6e.0 for ; Fri, 07 Aug 2026 13:21:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134065; x=1786738865; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Veey7tCntVTCDMITmGRXBJciLEPyIkZpuk6PEQ2O088=; b=MuyuqUWP4oGSy3+zAazXy+q9CKwIC8LyAcpu0V/19RRT7ZAO+a9DOcZdhKYJoMsJER E6c5I4wixJe9Qt3DN5tFVUbUsOKVK6lrMZZuJVC0y9jSFF10a67O4q++e0/PjSJRc68U gOl4zF3W2TP2/Hez2yhA7b4ONvH7z86Nq2YWICmD58S07rcHePfL+eKExDcJacXnKNag ESYAFNjaJ9wS0XCap1lcckAsKCcS7cXeDx7mVZoy3h8G3DWyw0yktS8acie9rWR3U7Y2 uTwTNBnvg+jdSNWhhq7+H5eOtBGFrn/7x3omHIoGh8XiB7frBr9JAEx0wZpCXEHDYEGL 1zaw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134065; x=1786738865; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Veey7tCntVTCDMITmGRXBJciLEPyIkZpuk6PEQ2O088=; b=PdBQbScj2ZGMZllPcpFEM6OA5fpipEvdHKlduAaubOjnQEdZCitsCdEe0+JiojImj7 tyUhg/Nd5zj33b0xyvfmCuE7Tz1iNDfpdY4Ml2YDsaCrMNTfqLFZcZLh6nx2CnCt66W5 +vc6deYB2pzSLeb1Z9693HRDHTULA1GkHNIk3nK24wDMi0bTRPHN0+8gmC0sxZL3Dz6z bCbAeTogp9XTS1Gn+/I0BZ3rKFOS4eizMtNDddtzXKsDcZ78fkm5CGAutLszYpABc4L2 0pYpusWNARh7wjY4p4Gpt/PLoM+YX9uLrmeowvVP8vZGN+k3Bs8Hgw3dqStaxnKCn0J1 5tLA== X-Forwarded-Encrypted: i=1; AHgh+Roi5kPnmPBeMHtfTSDZLpkUWyzgbmeeAfitGGULVNSjKvAHWAPWJQybxzsO35d3YyJCZOfn+BAyULoI9wY=@vger.kernel.org X-Gm-Message-State: AOJu0YyRKUbqDuToDFvQg/8EbhHnkzFkBAwpZ6p55o/a0IMS2XbQWTWR 70Kod4F6JULWN6AgQ2TEOvPWvTEG6BuGH8yNGzaMRSDaNxVVgY8P0jYw X-Gm-Gg: AR+sD112Ss4AlOPXs/RyKkIBcroLw7IzjGeSgz9MnZimZFO49w3EkOHi1MlD4BqX8Jv XWc5rINNjd896thswMC1Tgz5WAsrb4/zovXHt4ciM+keQ+LXY0z/TkyBjm3kpXatMd9hJGVSjUB LioZeYJFb0/luKuKF2hnmjLAmfMArueg20NNBTZnmULmlhoU1xT49Bv3qSz1CgJZD/Wq55wGmoG /2qbblbHUKfxgwMIlevzYaVeV1O3yv6cNjY6qoXVXvvs4lBiK633KwlITtQlywEJoHFxXbTJGof QmM1miKwFGq8ZLA2lpSOcfG0V5yA9VamM6f/iyYpmHAgvtGnvRG0kfuwbH38y+1PgxUqrmi/K6S JBSBit2By/+8evO7zmoHP5pPIV73eN7Iz5bsINRlsXN1UllNbYTviAK4vbZOMmS77Of3NhQwWwb YJrauf2I0MaHIKwanLxHjg3qdPxBwW4skk+/rO+CnyON7hLY01gMzBcSu/i9bklOdnNwemx2ZIO 7NlUa5RrRIQp4Kr1g== X-Received: by 2002:a05:6808:150b:b0:496:b7c:274b with SMTP id 5614622812f47-4afae16a91fmr12782189b6e.19.1786134065034; Fri, 07 Aug 2026 13:21:05 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:1::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4b1af5e4fc0sm422662b6e.11.2026.08.07.13.21.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:04 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 03/14] mm/memory-tiers: Introduce a mapping from nid to tier_slot Date: Fri, 7 Aug 2026 13:20:46 -0700 Message-ID: <20260807202059.2620949-4-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Establishing tiered memcg limits will require an ordering of tiers, as well as a way to account how much memory is present in each tier. This will need to be done starting at boot, so that all memory becomes properly accounted. However, tiers can come online and offline at runtime due to DAX memory whose nodes can be hotplugged / hot-unplugged, and these nodes' tiers are not available at boot. Therefore, to establish a fixed mapping from nid to tier that isn't sparse like the tier_ids, introduce a new "tier_slot" which is a dense index that does not change once a tier comes online. Also introduce a helper to retrieve the nodemask associated with a tier. Signed-off-by: Joshua Hahn --- include/linux/memory-tiers.h | 18 ++++++++ mm/memory-tiers.c | 86 +++++++++++++++++++++++++++++++++++- 2 files changed, 102 insertions(+), 2 deletions(-) diff --git a/include/linux/memory-tiers.h b/include/linux/memory-tiers.h index 7999c58629eeb..0e49645cdd1a9 100644 --- a/include/linux/memory-tiers.h +++ b/include/linux/memory-tiers.h @@ -41,6 +41,8 @@ extern struct memory_dev_type *default_dram_type; extern nodemask_t default_dram_nodes; struct memory_dev_type *alloc_memory_type(int adistance); void put_memory_type(struct memory_dev_type *memtype); +int mt_nr_tier_slots(void); +int nid_tier_slot(int nid); void init_node_memory_type(int node, struct memory_dev_type *default_type); void clear_node_memory_type(int node, struct memory_dev_type *memtype); int register_mt_adistance_algorithm(struct notifier_block *nb); @@ -52,6 +54,7 @@ int mt_perf_to_adistance(struct access_coordinate *perf, = int *adist); struct memory_dev_type *mt_find_alloc_memory_type(int adist, struct list_head *memory_types); void mt_put_memory_types(struct list_head *memory_types); +const nodemask_t *mt_tier_nodes(int slot); #ifdef CONFIG_NUMA_MIGRATION int next_demotion_node(int node, const nodemask_t *allowed_mask); void node_get_allowed_targets(pg_data_t *pgdat, nodemask_t *targets); @@ -151,5 +154,20 @@ static inline struct memory_dev_type *mt_find_alloc_me= mory_type(int adist, static inline void mt_put_memory_types(struct list_head *memory_types) { } + +static inline int mt_nr_tier_slots(void) +{ + return 0; +} + +static inline int nid_tier_slot(int nid) +{ + return -1; +} + +static inline const nodemask_t *mt_tier_nodes(int slot) +{ + return NULL; +} #endif /* CONFIG_NUMA */ #endif /* _LINUX_MEMORY_TIERS_H */ diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c index 54851d8a195b0..36187c0ea9ded 100644 --- a/mm/memory-tiers.c +++ b/mm/memory-tiers.c @@ -43,6 +43,16 @@ static LIST_HEAD(memory_tiers); */ static LIST_HEAD(default_memory_types); static struct node_memory_type_map node_memory_types[MAX_NUMNODES]; + +/* + * nr_tier_slots and tier_slot_ids are written with memory_tier_lock and + * read locklessly. nr_tier_slots is monotonically increasing. + */ +static int nr_tier_slots; +static int tier_slot_ids[MAX_NUMNODES] =3D {[0 ... MAX_NUMNODES - 1] =3D= -1,}; +static int node_tier_slots[MAX_NUMNODES] =3D {[0 ... MAX_NUMNODES - 1] =3D= -1,}; +static nodemask_t tier_nodemasks[MAX_NUMNODES]; + struct memory_dev_type *default_dram_type; nodemask_t default_dram_nodes __initdata =3D NODE_MASK_NONE; =20 @@ -273,6 +283,65 @@ static struct memory_tier *__node_get_memory_tier(int = node) lockdep_is_held(&memory_tier_lock)); } =20 +/* Caller must hold memory_tier_lock */ +static int tier_id_slot(int tier_id) +{ + int slot, free_slot =3D -1; + + for (slot =3D 0; slot < nr_node_ids; slot++) { + if (tier_slot_ids[slot] =3D=3D tier_id) + return slot; + if (tier_slot_ids[slot] =3D=3D -1 && free_slot =3D=3D -1) { + free_slot =3D slot; + tier_slot_ids[slot] =3D tier_id; + } + } + + return free_slot; +} + +static void establish_tier_slots(void) +{ + int old_nr_tier_slots =3D mt_nr_tier_slots(); + int highest_slot =3D old_nr_tier_slots; + + lockdep_assert_held_once(&memory_tier_lock); + + for (int slot =3D 0; slot < old_nr_tier_slots; slot++) + nodes_clear(tier_nodemasks[slot]); + + for (int nid =3D 0; nid < nr_node_ids; nid++) { + struct memory_tier *memtier =3D NULL; + int slot =3D -1; + + if (node_state(nid, N_MEMORY)) + memtier =3D __node_get_memory_tier(nid); + if (memtier) { + slot =3D tier_id_slot(memtier->dev.id); + highest_slot =3D max(highest_slot, slot + 1); + } + + WRITE_ONCE(node_tier_slots[nid], slot); + + if (slot !=3D -1) + node_set(nid, tier_nodemasks[slot]); + } + WRITE_ONCE(nr_tier_slots, highest_slot); +} + +int mt_nr_tier_slots(void) +{ + return READ_ONCE(nr_tier_slots); +} + +int nid_tier_slot(int nid) +{ + if (nid < 0 || nid >=3D MAX_NUMNODES) + return -1; + + return READ_ONCE(node_tier_slots[nid]); +} + #ifdef CONFIG_NUMA_MIGRATION bool node_is_toptier(int node) { @@ -729,6 +798,7 @@ static int __init memory_tier_late_init(void) } =20 establish_demotion_targets(); + establish_tier_slots(); put_online_mems(); =20 return 0; @@ -878,6 +948,14 @@ int mt_calc_adistance(int node, int *adist) } EXPORT_SYMBOL_GPL(mt_calc_adistance); =20 +const nodemask_t *mt_tier_nodes(int slot) +{ + if (slot < 0) + return NULL; + + return &tier_nodemasks[slot]; +} + static int __meminit memtier_hotplug_callback(struct notifier_block *self, unsigned long action, void *_arg) { @@ -887,15 +965,19 @@ static int __meminit memtier_hotplug_callback(struct = notifier_block *self, switch (action) { case NODE_REMOVED_LAST_MEMORY: mutex_lock(&memory_tier_lock); - if (clear_node_memory_tier(nn->nid)) + if (clear_node_memory_tier(nn->nid)) { establish_demotion_targets(); + establish_tier_slots(); + } mutex_unlock(&memory_tier_lock); break; case NODE_ADDED_FIRST_MEMORY: mutex_lock(&memory_tier_lock); memtier =3D set_node_memory_tier(nn->nid); - if (!IS_ERR(memtier)) + if (!IS_ERR(memtier)) { establish_demotion_targets(); + establish_tier_slots(); + } mutex_unlock(&memory_tier_lock); break; } --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-oi1-f169.google.com (mail-oi1-f169.google.com [209.85.167.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 05A74441618 for ; Fri, 7 Aug 2026 20:21:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134069; cv=none; b=UyCcwduobM30rgRX6FLVwEKzrCJeA+skEjLdrbpVk2s9iwfVthnESW5Vsz5drQr6z6ENAizc8e7JsSrASilBS1FTYftbso0HllL+6qlSk4N43FQ9uw1WwZVYmuD5/iWmMzCb4QvPTH8Kjfz5PO1Ih1pfH25PVzK6erEBBrUh1FQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134069; c=relaxed/simple; bh=HOogZ2bJZOmmIkJdY9y2MvQwya6iwHgtaiK7/Q90pew=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=F7pyzRviEd1On7irgfajDmjElOLdyigZPzPEjrTFh2bt1wi16rje6EhFBcTiIRqDEjiw4MyTDV7JhiwDZ2/xy4G+1rXqEvdZ+EcGa6E6ZXrdS52j2tp0632LrrmqnScMV928FC4nFbcS6oV0jTruyNRgPQQwKPQRqcD9UEUFVy0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=M9hMxhB9; arc=none smtp.client-ip=209.85.167.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="M9hMxhB9" Received: by mail-oi1-f169.google.com with SMTP id 5614622812f47-4a4c6081f9fso1170055b6e.3 for ; Fri, 07 Aug 2026 13:21:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134067; x=1786738867; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=k/MfpY/lSQ8qNOiw/HDxNskiAGRLvzb2J5gjoLFAHE4=; b=M9hMxhB9Mc0kTXstxl+cRR9/7LncgF70NavK4eY8cZqzmEnowNmrOIJXooDSIn3X1j WhPqb/FVSXTJoWMRNIiB9aB5YI67OpqwzwNeXEP+EvzM3KICEGWAKHkkHCdJxu5nAv2f 1PMINmLpHOA50SfoEQa3qIwAr5EKSUE7F6IpOf7xDghkureoKdZVOhw+defdahYELp3U UF1I4BDH82TnNKUOM9raSOLr3ec5ZceCH8D5jSNKgnLOTcekclZlcWgq5X11VUBOqW4m dDl4ST52QN9Q107+haUtnsfR+TTlCOjAalYgKbMb4ShmanJJtE+vCV+uiMr0UO2EHHCI rgxQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134067; x=1786738867; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=k/MfpY/lSQ8qNOiw/HDxNskiAGRLvzb2J5gjoLFAHE4=; b=LksMQlFF895X0H54N5BagTBqqCrpYMtHpJ3vtV4fPPEpvvSvrzohUJx+K86y1No+Pj BzgEEUgOMvSt6X6ZdzMb1DH1P55E/IEVpzJXnJjGQH+Dlh15JaXp7k+pWRCJdVVVEloD TmIPmx3mQdJJT+RgsPRM331YqyS/+bx5nBZUWQCzOoQlce+hyEAIYtZUxdiJbCYMtGOo RGHOKAnRKXKCotqmdeP9XGLaGfx+u66+AesdZ0Tp0gQ2AsZxbeUpHj9nYECqS0mX4v4f ONIqy35VVOB7PdGb+wJk7ugmgUzC5lw1KnvFBtILdT3CMdGQQ0tMLZx556ir9N1xAAiy GI3Q== X-Forwarded-Encrypted: i=1; AHgh+Rq1qvBe3XUEXi+H2DXLPR34RL1JryImvMu9hDgF/tVJD5sEMMi/8tMLhR1ul95PQ+A3ivl2ftSSTN594iU=@vger.kernel.org X-Gm-Message-State: AOJu0Yz7D8A6CjRIpkmn7KaU/d0nJm7dZB1w14ux1autl8pLsmS9zJSJ zzLI8YXrVQ3oE9J28nNWgIjmNHir+S3iBkdlUGIZjCWdL41LgXR+cO5H X-Gm-Gg: AR+sD13oSof9tdmvek3Fqpg0RBTGpHkQMt1o/LPNnDGzQ8mdx+PHYZ1cqPxxcb9eNNi a010Ar4Rs0uaep7SK1Wm8y7zjGEEQvhSGTkBJ2llIuJVlotjVQhmsGNoQJ3l1Xzi5MBReqsQuzK jO/9chuMF024Lj4poim1Sl5myly/fn4em4QoD2tke6t5nOBiN067xfATuLhx4noQbWHDbuMgCvY i4ZtfbK3Od8xwYqNpKpmp7O0kw/m69bqk+lnTcPHzvCm2iJOx07ASGOC4S7cyb1Uv/Dh2g63QPX JNfxMoL63H2teB+ZK5nfgsjvMuIdOg0D6uk+7mrg7Yi+j/GK2jgR9CKIX5gXtFXWJjkzwH1vyp+ Tle9rd4vifunatnrO5CRPMhomIvMzmgyXb+HsxCrdxoxQaJ54D32zlZdjh5G12Dgia3sppUU0vC vHW8PdwBizB1F16fXF7JCf7OVtfASJ0XW1fwobIZ5ysTBfjgAUPvy49Y0+tzuq1L5zwhYJpeOUe hREvyPidRP+9OSt87E= X-Received: by 2002:a05:6808:15a6:b0:49b:c387:36c5 with SMTP id 5614622812f47-4afae16d6cfmr14165498b6e.17.1786134066779; Fri, 07 Aug 2026 13:21:06 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:42::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4b1af0b57f6sm442378b6e.5.2026.08.07.13.21.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:06 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 04/14] mm/memcontrol: Allocate per-tier page_counters Date: Fri, 7 Aug 2026 13:20:47 -0700 Message-ID: <20260807202059.2620949-5-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Tier-aware limits need one page_counter per memory tier. Add a page_counter array (memcg->tier) to struct mem_cgroup, sized to nr_node_ids (an upper bound on the number of tiers) and allocated only when tiered limits are enabled, so memcgs pay nothing when the feature is off. Initialise and parent-link every slot in mem_cgroup_css_alloc(), and free the array in __mem_cgroup_free(). Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 1 + mm/memcontrol.c | 23 +++++++++++++++++++++++ mm/memory-tiers.c | 17 ++++++++++++++--- 3 files changed, 38 insertions(+), 3 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index dce03df7eae05..bb5bde87ac85a 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -207,6 +207,7 @@ struct mem_cgroup { =20 /* Accounted resources */ struct page_counter memory; /* Both v1 & v2 */ + struct page_counter *tier; /* v2 only */ =20 union { struct page_counter swap; /* v2 only */ diff --git a/mm/memcontrol.c b/mm/memcontrol.c index ec28512de6a23..d096010366515 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -53,6 +53,7 @@ #include #include #include +#include #include #include #include @@ -4126,6 +4127,7 @@ static void __mem_cgroup_free(struct mem_cgroup *memc= g) memcg1_free_events(memcg); kfree(memcg->vmstats); free_percpu(memcg->vmstats_percpu); + kfree(memcg->tier); kfree(memcg); } =20 @@ -4178,6 +4180,13 @@ static struct mem_cgroup *mem_cgroup_alloc(struct me= m_cgroup *parent) if (!alloc_mem_cgroup_per_node_info(memcg, node)) goto fail; =20 + if (mem_cgroup_tiered_limits()) { + memcg->tier =3D kcalloc(nr_node_ids, sizeof(*memcg->tier), + GFP_KERNEL); + if (!memcg->tier) + goto fail; + } + if (memcg_wb_domain_init(memcg, GFP_KERNEL)) goto fail; =20 @@ -4206,6 +4215,16 @@ static struct mem_cgroup *mem_cgroup_alloc(struct me= m_cgroup *parent) return ERR_PTR(error); } =20 +static void memcg_init_tier_counters(struct mem_cgroup *memcg, + struct mem_cgroup *parent, bool protection) +{ + for (int i =3D 0; i < nr_node_ids; i++) { + page_counter_init(&memcg->tier[i], + parent ? &parent->tier[i] : NULL, protection); + page_counter_set_high(&memcg->tier[i], PAGE_COUNTER_MAX); + } +} + static struct cgroup_subsys_state * __ref mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css) { @@ -4231,6 +4250,8 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *pare= nt_css) page_counter_set_high(&memcg->swap, PAGE_COUNTER_MAX); if (parent) { page_counter_init(&memcg->memory, &parent->memory, memcg_on_dfl); + if (mem_cgroup_tiered_limits()) + memcg_init_tier_counters(memcg, parent, memcg_on_dfl); page_counter_init(&memcg->swap, &parent->swap, false); #ifdef CONFIG_MEMCG_V1 WRITE_ONCE(memcg->swappiness, mem_cgroup_swappiness(parent)); @@ -4243,6 +4264,8 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *pare= nt_css) init_memcg_stats(); init_memcg_events(); page_counter_init(&memcg->memory, NULL, true); + if (mem_cgroup_tiered_limits()) + memcg_init_tier_counters(memcg, NULL, true); page_counter_init(&memcg->swap, NULL, false); #ifdef CONFIG_MEMCG_V1 page_counter_init(&memcg->kmem, NULL, false); diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c index 36187c0ea9ded..bd5c78cd26ec4 100644 --- a/mm/memory-tiers.c +++ b/mm/memory-tiers.c @@ -52,6 +52,7 @@ static int nr_tier_slots; static int tier_slot_ids[MAX_NUMNODES] =3D {[0 ... MAX_NUMNODES - 1] =3D= -1,}; static int node_tier_slots[MAX_NUMNODES] =3D {[0 ... MAX_NUMNODES - 1] =3D= -1,}; static nodemask_t tier_nodemasks[MAX_NUMNODES]; +static unsigned long tier_capacity[MAX_NUMNODES]; =20 struct memory_dev_type *default_dram_type; nodemask_t default_dram_nodes __initdata =3D NODE_MASK_NONE; @@ -307,8 +308,10 @@ static void establish_tier_slots(void) =20 lockdep_assert_held_once(&memory_tier_lock); =20 - for (int slot =3D 0; slot < old_nr_tier_slots; slot++) + for (int slot =3D 0; slot < old_nr_tier_slots; slot++) { nodes_clear(tier_nodemasks[slot]); + WRITE_ONCE(tier_capacity[slot], 0); + } =20 for (int nid =3D 0; nid < nr_node_ids; nid++) { struct memory_tier *memtier =3D NULL; @@ -323,8 +326,16 @@ static void establish_tier_slots(void) =20 WRITE_ONCE(node_tier_slots[nid], slot); =20 - if (slot !=3D -1) - node_set(nid, tier_nodemasks[slot]); + if (slot < 0) + continue; + + node_set(nid, tier_nodemasks[slot]); + for (int i =3D 0; i < MAX_NR_ZONES; i++) { + struct zone *zone =3D &NODE_DATA(nid)->node_zones[i]; + + WRITE_ONCE(tier_capacity[slot], + tier_capacity[slot] + zone_managed_pages(zone)); + } } WRITE_ONCE(nr_tier_slots, highest_slot); } --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-ot1-f42.google.com (mail-ot1-f42.google.com [209.85.210.42]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 571EE444704 for ; Fri, 7 Aug 2026 20:21:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.42 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134071; cv=none; b=iYr71rgGkk32ZYnnQ5Gn4kPaxQqFYDvXueDo8hVT8rNhg21XENAcRvQDWZblXXpaFNqYqBsNesQddlooiHkYvSSRmM0chL18mIcYA72GiIuGX5tOUeG/tyEUwVyPwXe5zhcMGAcD5HRzD1b4e0xfWlpuHGPcjH7rX5DZoZMwQsk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134071; c=relaxed/simple; bh=cvy8xs5odfKdlZvQWmVQYRS0OJ7R2YTEIC6VhLVlmYM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gxw7NVAja0DpFDiXSCVWMeM/BJswlgs4VrwHHoscpViCM9d84OKeHMPcSmqE8Lv+49bYRS74eFfq8TK8nGo0eDLNDp6wahJhzAVyEYCNdK2MLZcjY2wAWmXzHtu3m85R4dyuJNILmTSXEN4UT5FKeb0Bv6li3pk3uajguHq9IpE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=m/2p5+DC; arc=none smtp.client-ip=209.85.210.42 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="m/2p5+DC" Received: by mail-ot1-f42.google.com with SMTP id 46e09a7af769-7eb545db3afso5768a34.0 for ; Fri, 07 Aug 2026 13:21:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134069; x=1786738869; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=5vEVg85rEvrzvIwoN3Uc+WskLG7iLQIioETffuSYNDk=; b=m/2p5+DCfwh579mm/HZea72TYbGe8crCuFyyofCi0a3O8phGn1dASa0vl4MbFUDaVc 6RQdJKkPxcgB5yfLJLJ6a4opsxS1I8EtaeNslV6UhTBFPwVKII+WVxkStDfGNCt9uHXG 5ugJp+ORVYI8I8kqYQ/zKw17tuWDZE2sAIcnWMgujVfoj0Gon7z0JxfErBJHHz1GlhE5 dPNHRpseCBQktp0JMI1sJt7VOo5zaeRAjp9qNCxga4XcPZsJE1DfdtQ+32Sb7ipdopYv VocqQNa3aptoknWpVUw6lBILejtklpuIjrmEgFRibQTrEQEAcdsjQZuluLb4gipw+DVj uhbw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134069; x=1786738869; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=5vEVg85rEvrzvIwoN3Uc+WskLG7iLQIioETffuSYNDk=; b=chcVt+vCjfb6hV8FvZhVGRsQP4d3/3AVg71anwBAlRWu1l1z8kN+XKYblKQRcuBcYU pGC2K2pNwUVTEIjDyWHjnhk+cJm7LSOheosEY0vB2T2aNQlhNl2tLFuVZdksTx2R72AN NQVX6EOfqCJnEXahD21T3Z3ZoTKCWnaZkDg2D0Nse8VDW3+jA9JlGVRSGEuUyZ+ZaGGm ebyaidZy0h3zhriJKF359y2LCc+EkFciRlh/jAydbalvplJRnWkJfVHZ8M3eCaol6lxw sPUUHiB7Dyxidx7MqLqdUjlng7Of623sr2bXZEQbwhJq91azF5VUrDlntyeShA7tzb9u gWXw== X-Forwarded-Encrypted: i=1; AHgh+RrabHfaLXSzotBCSZc9c2kKhVMhgL1rJMipNXBczaMvVbYasdqb4tIcXlnDxRL+AymsDv0LJlsXRVPF7IY=@vger.kernel.org X-Gm-Message-State: AOJu0Yy/jgTrjUhZiGEnZr4k3pGAR1dDrZ3KtfGEu3yTCBOYBNYfUPID 9wDvN71c2+ONw509dGc2J0eSUmhjomGDTOv9qjUy3hyGN+ahIMR2b/Cm X-Gm-Gg: AR+sD10vr1PnyEbF3eNNZMp32NoWMb3xD1UJN04BJ4GNjdfN71qPHi6zDAs0JHCf0q0 T+VG9gC1auf/ONbCwgkpTTXds/caE6nUzgnixoSq4nq/sBPVxce9tRxVJO9URInzfCwlHwFd02K oXlnoGiW1R05jQtMYR16NWlHzN9qx4IM4gDBWm7kiagsAwoyEalUsOJtiFd6l/KwZ0Er0ilYH0o yXXh+Xf62o+r5sL9ZX+5zVgYF+y+Ns3lPg9ebI0fym+lcCw3pwMuP13U5cXll4py3DLZDpx3HMs HFiWTRkjfrfaZFki+HhIFvVvoQKYRofEpAWwiJ2mZCh4rI6FMjeczQQo9Wm0IbtVo13x70S70cf Qp/kBgnMmAPfRjTM7lsLfmjJaIYF0gzbt+Wahxbda1baTnM4HM0etmLW10PWYuzSMhurznm42ML kYeC3V8nhq86V7u7eYFW++n83wHxrUXnzzLtxvnPkM+CtyquICirAARrGXgxUjmSK4ye2bH3ooa ppgU1sOvoKDbWqPXA== X-Received: by 2002:a05:6820:168f:b0:6ae:9b39:68d1 with SMTP id 006d021491bc7-6b0420d5753mr1429497eaf.23.1786134068227; Fri, 07 Aug 2026 13:21:08 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02be475f7sm2836904eaf.9.2026.08.07.13.21.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:07 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 05/14] mm/memcontrol: Set tier limits proportional to memory limits Date: Fri, 7 Aug 2026 13:20:48 -0700 Message-ID: <20260807202059.2620949-6-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Compute proportional per-tier limits based on memory limits when users write to memory limit sysfs files, or when memory hotplug causes tier proportions to be shifted. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 10 +++++++ include/linux/memory-tiers.h | 6 ++++ mm/memcontrol.c | 54 ++++++++++++++++++++++++++++++++++++ mm/memory-tiers.c | 13 +++++++++ 4 files changed, 83 insertions(+) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index bb5bde87ac85a..f7a92b66330ec 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -537,11 +537,17 @@ static inline bool mem_cgroup_tiered_limits(void) { return static_branch_unlikely(&memcg_tiered_limits_key); } + +void establish_memcg_tier_limits(void); #else static inline bool mem_cgroup_tiered_limits(void) { return false; } + +static inline void establish_memcg_tier_limits(void) +{ +} #endif =20 static inline void mem_cgroup_protection(struct mem_cgroup *root, @@ -1102,6 +1108,10 @@ static inline bool mem_cgroup_tiered_limits(void) return false; } =20 +static inline void establish_memcg_tier_limits(void) +{ +} + static inline void memcg_memory_event(struct mem_cgroup *memcg, enum memcg_memory_event event) { diff --git a/include/linux/memory-tiers.h b/include/linux/memory-tiers.h index 0e49645cdd1a9..04b396f60b457 100644 --- a/include/linux/memory-tiers.h +++ b/include/linux/memory-tiers.h @@ -55,6 +55,7 @@ struct memory_dev_type *mt_find_alloc_memory_type(int adi= st, struct list_head *memory_types); void mt_put_memory_types(struct list_head *memory_types); const nodemask_t *mt_tier_nodes(int slot); +unsigned long mt_scale_by_tier(unsigned long val, int slot); #ifdef CONFIG_NUMA_MIGRATION int next_demotion_node(int node, const nodemask_t *allowed_mask); void node_get_allowed_targets(pg_data_t *pgdat, nodemask_t *targets); @@ -169,5 +170,10 @@ static inline const nodemask_t *mt_tier_nodes(int slot) { return NULL; } + +static inline unsigned long mt_scale_by_tier(unsigned long val, int slot) +{ + return val; +} #endif /* CONFIG_NUMA */ #endif /* _LINUX_MEMORY_TIERS_H */ diff --git a/mm/memcontrol.c b/mm/memcontrol.c index d096010366515..defd04acfb3fd 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -4421,6 +4421,35 @@ static void mem_cgroup_css_free(struct cgroup_subsys= _state *css) mem_cgroup_free(memcg); } =20 +static inline unsigned long page_counter_max_or_scale(unsigned long val, + int slot) +{ + return val =3D=3D PAGE_COUNTER_MAX ? PAGE_COUNTER_MAX : + mt_scale_by_tier(val, slot); +} + +static void memcg_scale_tier_limits(struct mem_cgroup *memcg) +{ + unsigned long min =3D READ_ONCE(memcg->memory.min); + unsigned long low =3D READ_ONCE(memcg->memory.low); + unsigned long high =3D READ_ONCE(memcg->memory.high); + unsigned long max =3D READ_ONCE(memcg->memory.max); + int nr_tier_slots =3D mt_nr_tier_slots(); + + for (int slot =3D 0; slot < nr_tier_slots; slot++) { + unsigned long new_min =3D page_counter_max_or_scale(min, slot); + unsigned long new_low =3D page_counter_max_or_scale(low, slot); + unsigned long new_high =3D page_counter_max_or_scale(high, slot); + unsigned long new_max =3D page_counter_max_or_scale(max, slot); + struct page_counter *tier =3D &memcg->tier[slot]; + + page_counter_set_min(tier, new_min); + page_counter_set_low(tier, new_low); + page_counter_set_high(tier, new_high); + xchg(&tier->max, new_max); + } +} + /** * mem_cgroup_css_reset - reset the states of a mem_cgroup * @css: the target css @@ -4454,6 +4483,8 @@ static void mem_cgroup_css_reset(struct cgroup_subsys= _state *css) page_counter_set_high(&memcg->memory, PAGE_COUNTER_MAX); memcg1_soft_limit_reset(memcg); page_counter_set_high(&memcg->swap, PAGE_COUNTER_MAX); + if (mem_cgroup_tiered_limits()) + memcg_scale_tier_limits(memcg); memcg_wb_domain_size_changed(memcg); } =20 @@ -4797,6 +4828,21 @@ static ssize_t memory_peak_write(struct kernfs_open_= file *of, char *buf, &memcg->memory_peaks); } =20 +#ifdef CONFIG_NUMA +void establish_memcg_tier_limits(void) +{ + struct mem_cgroup *memcg; + + if (!mem_cgroup_tiered_limits()) + return; + + for_each_mem_cgroup_tree(memcg, NULL) { + if (memcg !=3D root_mem_cgroup) + memcg_scale_tier_limits(memcg); + } +} +#endif + #undef OFP_PEAK_UNSET =20 static int memory_min_show(struct seq_file *m, void *v) @@ -4818,6 +4864,8 @@ static ssize_t memory_min_write(struct kernfs_open_fi= le *of, return err; =20 page_counter_set_min(&memcg->memory, min); + if (mem_cgroup_tiered_limits()) + memcg_scale_tier_limits(memcg); =20 return nbytes; } @@ -4841,6 +4889,8 @@ static ssize_t memory_low_write(struct kernfs_open_fi= le *of, return err; =20 page_counter_set_low(&memcg->memory, low); + if (mem_cgroup_tiered_limits()) + memcg_scale_tier_limits(memcg); =20 return nbytes; } @@ -4866,6 +4916,8 @@ static ssize_t memory_high_write(struct kernfs_open_f= ile *of, return err; =20 page_counter_set_high(&memcg->memory, high); + if (mem_cgroup_tiered_limits()) + memcg_scale_tier_limits(memcg); =20 if (of->file->f_flags & O_NONBLOCK) goto out; @@ -4925,6 +4977,8 @@ static ssize_t memory_max_write(struct kernfs_open_fi= le *of, return err; =20 xchg(&memcg->memory.max, max); + if (mem_cgroup_tiered_limits()) + memcg_scale_tier_limits(memcg); =20 if (of->file->f_flags & O_NONBLOCK) goto out; diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c index bd5c78cd26ec4..e2c99f51c36d1 100644 --- a/mm/memory-tiers.c +++ b/mm/memory-tiers.c @@ -810,6 +810,7 @@ static int __init memory_tier_late_init(void) =20 establish_demotion_targets(); establish_tier_slots(); + establish_memcg_tier_limits(); put_online_mems(); =20 return 0; @@ -967,6 +968,16 @@ const nodemask_t *mt_tier_nodes(int slot) return &tier_nodemasks[slot]; } =20 +unsigned long mt_scale_by_tier(unsigned long val, int slot) +{ + unsigned long total_capacity =3D totalram_pages(); + + if (slot < 0 || !total_capacity) + return 0; + + return mult_frac(val, READ_ONCE(tier_capacity[slot]), total_capacity); +} + static int __meminit memtier_hotplug_callback(struct notifier_block *self, unsigned long action, void *_arg) { @@ -979,6 +990,7 @@ static int __meminit memtier_hotplug_callback(struct no= tifier_block *self, if (clear_node_memory_tier(nn->nid)) { establish_demotion_targets(); establish_tier_slots(); + establish_memcg_tier_limits(); } mutex_unlock(&memory_tier_lock); break; @@ -988,6 +1000,7 @@ static int __meminit memtier_hotplug_callback(struct n= otifier_block *self, if (!IS_ERR(memtier)) { establish_demotion_targets(); establish_tier_slots(); + establish_memcg_tier_limits(); } mutex_unlock(&memory_tier_lock); break; --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-oa1-f43.google.com (mail-oa1-f43.google.com [209.85.160.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 904D744470C for ; Fri, 7 Aug 2026 20:21:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134073; cv=none; b=RwsPvqpqjmOMCNZFZU9eJeQFiM5dpTexjRiVBa6byckm/cOciOYZC2QdPtb4NOvsnqXbCHeXNxLtXY7pLgiPxVBiGVnCpAT/AJ+YRsCgtIn3m1s9zuUVPO9MVXp00qXegW9OP1x2bF6Pe9JoSzSd+fqb4SThKyXXuKdidLcCXls= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134073; c=relaxed/simple; bh=ZG8mhq/oRDYPChzRStWZmZ7lpPUxxEq9apfKos3UvZY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rs708VHUK4FvixZqq4IM5I0djTq4jxvXE4WfExMbgNnrlUAlnN4yFZFDa80iRNSbLMkWzbDLNDly5FKgzr92LtZJ10u9gy/tX2FFlW7XIfedY/ryRy+Oa80WvmmCAdyjYj1DT6xaZZZRpNmAl9q1jQJ/YU1grmBJPtRDNfQiwu0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Qm7J3Ovx; arc=none smtp.client-ip=209.85.160.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Qm7J3Ovx" Received: by mail-oa1-f43.google.com with SMTP id 586e51a60fabf-451fd21113cso2130725fac.1 for ; Fri, 07 Aug 2026 13:21:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134069; x=1786738869; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VEYEX02IKWE90JSQ/XOJ17bp5Gj7dbgYbm77S6gqa2M=; b=Qm7J3OvxDalvZ2pKHv3CBSQQbttnQmHcCJ94rfl7H29NISOsC3XyDumI/CbhJtqy1j RomUEomOOgBIusPKaxLA+9pncuzliPNnv9BKCGuCchoK5DWDR2kY8CPonezWqHJSBo5/ r5mIyY65ci7GSc6hAACm5rdHtxLiDdDd7YFobHQ+f2zUAjBI//boLdEcjT80YNU9WHxq q92+OV7YNsGvN1h33vBANDRvIhA0Vcdwv5OmTOM6RzTineVEDxnG046jE9S+Y08UrH5n 68tyxY/lHCXe3YtwguRZpD19TxXjmCDjXS5ILz4ZLvvD4MmEaMTlTyNFl6R1bXQLM1n+ 6oqA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134069; x=1786738869; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VEYEX02IKWE90JSQ/XOJ17bp5Gj7dbgYbm77S6gqa2M=; b=kAbKiyiTpy7h36zKG/eASfBmnIARgOsaUw0gQ+kONvbhMOwTuo9rESoG2Fj4r4+0gx TRKoejbWSqTTwXOmchUyLXxOOh4809Ejn4BNon/i0xPn50JTY9VTTR1mH9rzTtc8WoAF eat6/jAZMIRFXEYLcnEeUGj+2rVPXmfaXhgbZ4dgsN1hlPfV0z1qgI3L+jh7ZvIUb8Wv WenlvOM7At+td6SrYVwEC7OnKxKHMDcURJFFy+RSlGw2MuceGVEf0kYrSDcJlrD3XbXr 3N5VBvXvyIynARmhg5+FF+hUf6o6TgbJT+wA9DjeLVcGcm02OFv72vr1F1gOcMYLu8tc /2ag== X-Forwarded-Encrypted: i=1; AHgh+RqNh6PQB+m/xeMXhOoinstecC+Y+TJvbGG158ZC2NzzhYHgVZJOQpq/mT1VWFeiY+t6wm/+gb3GMl3hRSI=@vger.kernel.org X-Gm-Message-State: AOJu0Yy2xxJTkMN1TC5r/5HNNLajG5UjNCeD1iXW1Hp0iyn4/0SgQGV3 XKZ7JjEI7amfgaK6L8fs53V67Js6UOYxkZNPhpJ/kEdM/FDCc3mUtju5 X-Gm-Gg: AR+sD12kuWRK87hMPXuak2E71iwHEJph5NEBxRR1Po0b3NrF5m7X7d+Qjl8cjNGThD1 74NiW29EO83HGagli2Mgi+aXdM0DlDaHLvs8c+Z83SPDQ4iXyrBsbei1dC5q2BS6zQC+yKKLyWo sTEIeyN2KE0vlSoTYyoHfSihbV0H7oy0o9+ksEmPTZr7ayFlMW3MdV7/WU/d6AWcPjjZcFBiUsA +3RM4xJhE+VdjqHRYKPCip+xbjQ/06jAlhsae6gKaqv77QEEIrVCt2qm8J7vLZqiRt0QmatCCcm NX/n1ivnjMwd4n7P0LujpaDMA6XEwRiQWyiN3XX7iOEA+JGuEhLbj6qdn370k3+C0Ms6npYp+ZG kRiWDgVu4nwoiGhtOPtOnqqTAn5ePA5mj/1YFwJR9VLcy0i5b8NSSGKW7DL7EGXlnhbjmsT/SPC k/U4OmVSeHUA80FnJSKUtiv/Bs0AYlL5p5OCSkeWVTe70JWn1CksZkYP+QLV9tUQO2JpAvEjX9l 7EUQt4rVnHipNOzrhg/39CEPW+U75so4Rqu3FQ= X-Received: by 2002:a05:6870:478d:b0:424:801:49fe with SMTP id 586e51a60fabf-45a15e2f25dmr397157fac.8.1786134069430; Fri, 07 Aug 2026 13:21:09 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:2::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1a9e431sm2643069fac.6.2026.08.07.13.21.08 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:09 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 06/14] mm/vmscan, memcontrol: Add nodemask to try_to_free_mem_cgroup_pages Date: Fri, 7 Aug 2026 13:20:49 -0700 Message-ID: <20260807202059.2620949-7-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a new nodemask parameter to try_to_free_mem_cgroup_pages to allow selective reclaim on certain nodes. This new function signature can be used in future patches to selectively perform reclaim on toptier and place downward pressure when toptier limits are breached but memcg-wide limits are not yet breached. Signed-off-by: Joshua Hahn --- mm/internal.h | 3 ++- mm/memcontrol-v1.c | 5 +++-- mm/memcontrol.c | 11 +++++++---- mm/vmscan.c | 10 +++++++--- 4 files changed, 19 insertions(+), 10 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index f47f06c555481..5f2f6a1757bd9 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -84,7 +84,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgr= oup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness); + int *swappiness, + const nodemask_t *allowed); unsigned long mem_cgroup_shrink_node(struct mem_cgroup *memcg, gfp_t gfp_mask, bool noswap, pg_data_t *pgdat, diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 835fc8e511844..688896b59efe4 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -1815,7 +1815,8 @@ static int mem_cgroup_resize_max(struct mem_cgroup *m= emcg, } =20 if (!try_to_free_mem_cgroup_pages(memcg, 1, GFP_KERNEL, - memsw ? 0 : MEMCG_RECLAIM_MAY_SWAP, NULL)) { + memsw ? 0 : MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL)) { ret =3D -EBUSY; break; } @@ -1851,7 +1852,7 @@ static int mem_cgroup_force_empty(struct mem_cgroup *= memcg) break; =20 if (!try_to_free_mem_cgroup_pages(memcg, 1, GFP_KERNEL, - MEMCG_RECLAIM_MAY_SWAP, NULL)) + MEMCG_RECLAIM_MAY_SWAP, NULL, NULL)) nr_retries--; } =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index defd04acfb3fd..6c67d9d2c9ac7 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2397,7 +2397,7 @@ static unsigned long reclaim_high(struct mem_cgroup *= memcg, nr_reclaimed +=3D try_to_free_mem_cgroup_pages(memcg, nr_pages, gfp_mask, MEMCG_RECLAIM_MAY_SWAP, - NULL); + NULL, NULL); psi_memstall_leave(&pflags); } while ((memcg =3D parent_mem_cgroup(memcg)) && !mem_cgroup_is_root(memcg)); @@ -2710,7 +2710,8 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, =20 psi_memstall_enter(&pflags); nr_reclaimed =3D try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages, - gfp_mask, reclaim_options, NULL); + gfp_mask, reclaim_options, + NULL, NULL); psi_memstall_leave(&pflags); =20 if (mem_cgroup_margin(mem_over_limit) >=3D nr_pages) @@ -4946,7 +4947,8 @@ static ssize_t memory_high_write(struct kernfs_open_f= ile *of, } =20 reclaimed =3D try_to_free_mem_cgroup_pages(memcg, nr_pages - high, - GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, NULL); + GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL); =20 if (!reclaimed && !nr_retries--) break; @@ -5007,7 +5009,8 @@ static ssize_t memory_max_write(struct kernfs_open_fi= le *of, =20 if (nr_reclaims) { if (!try_to_free_mem_cgroup_pages(memcg, nr_pages - max, - GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, NULL)) + GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL)) nr_reclaims--; continue; } diff --git a/mm/vmscan.c b/mm/vmscan.c index 17d2b793cbfc4..ffe7ea3c5aff6 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -6859,7 +6859,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness) + int *swappiness, + const nodemask_t *allowed) { unsigned long nr_reclaimed; unsigned int noreclaim_flag; @@ -6875,6 +6876,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, .may_unmap =3D 1, .may_swap =3D !!(reclaim_options & MEMCG_RECLAIM_MAY_SWAP), .proactive =3D !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE), + .nodemask =3D allowed, }; /* * Traverse the ZONELIST_FALLBACK zonelist of the current node to put @@ -6900,7 +6902,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness) + int *swappiness, + const nodemask_t *allowed) { return 0; } @@ -8033,7 +8036,8 @@ int user_proactive_reclaim(char *buf, reclaimed =3D try_to_free_mem_cgroup_pages(memcg, batch_size, gfp_mask, reclaim_options, - swappiness =3D=3D -1 ? NULL : &swappiness); + swappiness =3D=3D -1 ? NULL : &swappiness, + NULL); } else { struct scan_control sc =3D { .gfp_mask =3D current_gfp_context(gfp_mask), --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:51 2026 Received: from mail-oa1-f47.google.com (mail-oa1-f47.google.com [209.85.160.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D7E3E442FAD for ; Fri, 7 Aug 2026 20:21:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134074; cv=none; b=CPwr7ZGi3HwUKTm1YcFrLW5JJ4BYBaXcIZBrcKY5cyReXG7lcbQMqXm7j2ofSfJ9/bHDZ+TqQVLtqL+DmzG1y1nYtIhwruCi47d40MbGAm/O+yLCCcb10bBb2gEH2nfBsmd7rL2eKnSw6YKo9jX37fM20GUN99eXJ/gn/VkWIUY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134074; c=relaxed/simple; bh=+IzWGgaNbj6Y02X6WE/c6I/buYlv0yEEBkp/xShOgGQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=H/Rg6EMLVrwXVW8pHSuRpjsBSJ0snvRWWENCKYNmE6PiMNRPppxcq4mT8tphtFhAkmVHEefaIBz0sr62RSv5eweU/txTCEyuCPmKi/td7pbkeKvPaiuoFduh1PUBgXfJWmxT2cuImpwfNrZN/eRbk+yrBSkxl2oR+njQZbIzQc0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=sc4SzZ0h; arc=none smtp.client-ip=209.85.160.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="sc4SzZ0h" Received: by mail-oa1-f47.google.com with SMTP id 586e51a60fabf-455ca262ccbso2163211fac.1 for ; Fri, 07 Aug 2026 13:21:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134072; x=1786738872; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=lxxJcfOvuS2DiwqIjrqP14riK8k/np3VBj64PdkOyLA=; b=sc4SzZ0hWlHxRydlGqtK6Iwsz0JMmjpThjdRZ//Bw7660Vd9ueHOacG4d984f4V1Fy /xx1LRTgLX0y1k5+OhZ7NAmRdAGPrwIKKdhA80gNiur6xeXH937tSeZDwgqUy1Ijn4hE DQjncRVO64HYESqynNfoIB3PoXsp96zl5vgwP/uGWUlZRR77WWYh0ttoKG2tb3xwIqr2 kECE8zjQC7S8U3Dm/AgOeJQKklfPc/GzH1lGD84+/oPfF59iBVRnsyKu9lBTRYGMC7ih TV0T2hJgmH5J6Qe6dVX0Yltkxp+3/Ld2LN0Q9shKqDZHc4g0ySfoSIuadpaf6tLvUxlu gzVg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134072; x=1786738872; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=lxxJcfOvuS2DiwqIjrqP14riK8k/np3VBj64PdkOyLA=; b=KBUOTTUuQZHb1zohHIdN2CyiGOWMmVtwtOQiumPfDmUJWQ7JcMoflIeQiYD6VWrI5h lZ0TSYnP9ThOSEoMsDaWoy7Qrr1rGlskYC0k6xQIAf7ntub6e/QWElB8hqfaweLn8jrN MSeSiFBdLbEgtt/lo80sllZoC2a1/XnnaR2mJr7OES8NyYg6WrCH+1x8CIYo4uZWJTFR VtYk6kFlmlV9WlZXanIAKO1fgJlTPokqT8r4IV6XtoO1I/GylQaRk62Z6yKqamLIW0fX suhC6jh+d6jPa9jbpbctQ7rJOeAtUk0SjydPUjmav7SL4fvFWfvSka9QD/jMiifrc22z LW7w== X-Forwarded-Encrypted: i=1; AHgh+RpacsSBJGbRP83IeMqh8jNmyPFts70uMuijBPmUgULDjciL02DOv4c6y+/S+O2ksVdAOsm4BqA6vWmRqrI=@vger.kernel.org X-Gm-Message-State: AOJu0YyJc8N1oztJu5ZwkrnpZQaTKZHEYmbhi7fJjDr7bJUxTADi/ixq MssLv97EfdVh41s4EQLOCbjNCNCBiiISsUNkb5ocBXggZf1aNTTbIGe7 X-Gm-Gg: AR+sD12/ZK2B7gaGycDi0f675rGe7DwHjSCBAWsWK9t7RIItIZbRWUn2YQldZxgCyGF NbrXbBkI1jA+/F2l2GQkUjpeKme1nmo7JMPrXMU0jtIcA6Hh5G0a9ZJdFFqu+5AoZM/uyMbNnlV OKnBj/irkyOKet3nPIEmdX1+LgerK/LrEBCZgBkN7WHMqPQBVoaz4GD9Sg382ZcY7xG7PMpBMQf ybhblg38gH4u/U9Z9+VsPYkZE9KJtMngZW+ijt+BKAdArhM4WFoNgjTNqfQXFbcLwR2cn9+dHwE Bq59fqzbW2s7rQwEZCyKgf76lzWAPo0C1mSB/szSMnYwtGNqrwOkBVNKGO3fhtwzmJj6DRvqBeU luoYRmTpdhe4Ce0H77v1JQtw90kAwWKTuWetC+338oG3DlBwOhMgR2CMwtRp9frM1lK7UXWgvX7 KrA4P1VmVTX/rdQjaNS7P2cwCz3tNudQ3mO4B0A4bN6i+J8+iRYEBRyrfvmzR6CDrPRy9XfYX9F HHnczxw649PduGgdvj5xoVbha+2TQ== X-Received: by 2002:a05:6808:1a0b:b0:490:af61:8830 with SMTP id 5614622812f47-4afadf217ccmr13020013b6e.3.1786134070676; Fri, 07 Aug 2026 13:21:10 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:73::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4b1af63b441sm422380b6e.14.2026.08.07.13.21.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:10 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 07/14] mm/memcontrol: Charge/uncharge tiered memory to mem_cgroup Date: Fri, 7 Aug 2026 13:20:50 -0700 Message-ID: <20260807202059.2620949-8-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Memory cgroup limits isolate memory as a resource, but treat all memory as equally valuable regardless of which memory tier it resides in. Account tiered memory usage in parallel with existing memory accounting. Add an nid parameter to try_charge_memcg(); callers resolve it to a tier slot with nid_tier_slot() and charge memcg->tier[slot] alongside memcg->memory. Uncharging reuses uncharge_gather to batch. Because the high-volume free path reclaims per-node, a batch is normally single-tier; if a folio in a different tier appears mid-batch, flush the accumulated uncharge and start accumulating for the new tier. Folios on different nodes within the same tier do not force a flush. Also, mem_cgroup_migrate() and mem_cgroup_replace_folio() now move the tier charge when the folio changes tier. Currently this only tracks LRU folios (try_charge_memcg() callers from charge_memcg()). The other two sites, obj_cgroup_charge_pages() and mem_cgroup_sk_charge(), will be handled by a future series that transitions enum memcg_stat_item to a per-lruvec counter (enum node_stat_item). The per-tier limits computed in the previous patch are not consulted yet, this patch only acounts the memory. Enforcement will come in the following patches. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 116 ++++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 108 insertions(+), 8 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 6c67d9d2c9ac7..f3714dfd85aa0 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -1561,6 +1561,15 @@ void mem_cgroup_update_lru_size(struct lruvec *lruve= c, enum lru_list lru, *lru_size +=3D nr_pages; } =20 +static struct page_counter *mem_cgroup_tier_counter(struct mem_cgroup *mem= cg, + int slot) +{ + if (slot < 0) + return NULL; + + return &memcg->tier[slot]; +} + /** * mem_cgroup_margin - calculate chargeable space of a memory cgroup * @memcg: the memory cgroup @@ -2645,13 +2654,32 @@ void __mem_cgroup_handle_over_high(gfp_t gfp_mask) css_put(&memcg->css); } =20 +static void mem_cgroup_uncharge_tier(struct mem_cgroup *memcg, + int slot, unsigned int nr_pages) +{ + struct page_counter *tier_counter =3D mem_cgroup_tier_counter(memcg, slot= ); + + if (tier_counter) + page_counter_uncharge(tier_counter, nr_pages); +} + +static void mem_cgroup_charge_tier(struct mem_cgroup *memcg, + int slot, unsigned int nr_pages) +{ + struct page_counter *tier_counter =3D mem_cgroup_tier_counter(memcg, slot= ); + + if (tier_counter) + page_counter_charge(tier_counter, nr_pages); +} + static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask, - unsigned int nr_pages) + unsigned int nr_pages, int nid) { unsigned int batch =3D max(MEMCG_CHARGE_BATCH, nr_pages); int nr_retries =3D MAX_RECLAIM_RETRIES; struct mem_cgroup *mem_over_limit; struct page_counter *counter; + struct page_counter *tier_counter =3D NULL; unsigned long nr_reclaimed; bool passed_oom =3D false; unsigned int reclaim_options; @@ -2659,10 +2687,19 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, bool raised_max_event =3D false; unsigned long pflags; bool allow_spinning =3D gfpflags_allow_spinning(gfp_mask); + int slot =3D -1; + + if (mem_cgroup_tiered_limits()) { + slot =3D nid_tier_slot(nid); + tier_counter =3D mem_cgroup_tier_counter(memcg, slot); + } =20 retry: - if (consume_stock(memcg, nr_pages)) + if (consume_stock(memcg, nr_pages)) { + if (tier_counter) + page_counter_charge(tier_counter, nr_pages); return 0; + } =20 if (!allow_spinning) /* Avoid the refill and flush of the older stock */ @@ -2677,8 +2714,11 @@ static int try_charge_memcg(struct mem_cgroup *memcg= , gfp_t gfp_mask, goto reclaim; } =20 - if (page_counter_try_charge(&memcg->memory, batch, &counter)) + if (page_counter_try_charge(&memcg->memory, batch, &counter)) { + if (tier_counter) + page_counter_charge(tier_counter, nr_pages); goto done_restock; + } =20 if (do_memsw_account()) page_counter_uncharge(&memcg->memsw, batch); @@ -2781,6 +2821,8 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, * temporarily by force charging it. */ page_counter_charge(&memcg->memory, nr_pages); + if (tier_counter) + page_counter_charge(tier_counter, nr_pages); if (do_memsw_account()) page_counter_charge(&memcg->memsw, nr_pages); =20 @@ -2852,7 +2894,7 @@ static inline int try_charge(struct mem_cgroup *memcg= , gfp_t gfp_mask, if (mem_cgroup_is_root(memcg)) return 0; =20 - return try_charge_memcg(memcg, gfp_mask, nr_pages); + return try_charge_memcg(memcg, gfp_mask, nr_pages, NUMA_NO_NODE); } =20 static void commit_charge(struct folio *folio, struct obj_cgroup *objcg) @@ -3152,7 +3194,7 @@ static int obj_cgroup_charge_pages(struct obj_cgroup = *objcg, gfp_t gfp, =20 memcg =3D get_mem_cgroup_from_objcg(objcg); =20 - ret =3D try_charge_memcg(memcg, gfp, nr_pages); + ret =3D try_charge_memcg(memcg, gfp, nr_pages, NUMA_NO_NODE); if (ret) goto out; =20 @@ -5280,7 +5322,8 @@ static int charge_memcg(struct folio *folio, struct m= em_cgroup *memcg, objcg =3D get_obj_cgroup_from_memcg(memcg); /* Do not account at the root objcg level. */ if (!obj_cgroup_is_root(objcg)) - ret =3D try_charge_memcg(memcg, gfp, folio_nr_pages(folio)); + ret =3D try_charge_memcg(memcg, gfp, folio_nr_pages(folio), + folio_nid(folio)); if (ret) { obj_cgroup_put(objcg); return ret; @@ -5376,6 +5419,8 @@ struct uncharge_gather { unsigned long pgpgout; unsigned long nr_kmem; int nid; + int tier_slot; + unsigned long tier_nr; }; =20 static inline void uncharge_gather_clear(struct uncharge_gather *ug) @@ -5383,6 +5428,34 @@ static inline void uncharge_gather_clear(struct unch= arge_gather *ug) memset(ug, 0, sizeof(*ug)); } =20 +static void flush_tier_charge(struct mem_cgroup *memcg, + const struct uncharge_gather *ug) +{ + struct page_counter *tier_counter; + + tier_counter =3D &memcg->tier[ug->tier_slot]; + page_counter_uncharge(tier_counter, ug->tier_nr); +} + +static void gather_tier_charge(struct uncharge_gather *ug, struct folio *f= olio, + unsigned long nr_pages) +{ + int slot =3D nid_tier_slot(folio_nid(folio)); + + if (slot < 0) + return; + + if (ug->tier_nr && slot !=3D ug->tier_slot) { + rcu_read_lock(); + flush_tier_charge(obj_cgroup_memcg(ug->objcg), ug); + rcu_read_unlock(); + ug->tier_nr =3D 0; + } + + ug->tier_slot =3D slot; + ug->tier_nr +=3D nr_pages; +} + static void uncharge_batch(const struct uncharge_gather *ug) { struct mem_cgroup *memcg; @@ -5395,6 +5468,8 @@ static void uncharge_batch(const struct uncharge_gath= er *ug) mod_memcg_state(memcg, MEMCG_KMEM, -ug->nr_kmem); memcg1_account_kmem(memcg, -ug->nr_kmem); } + if (ug->tier_nr) + flush_tier_charge(memcg, ug); memcg1_oom_recover(memcg); } =20 @@ -5440,8 +5515,11 @@ static void uncharge_folio(struct folio *folio, stru= ct uncharge_gather *ug) ug->nr_kmem +=3D nr_pages; } else { /* LRU pages aren't accounted at the root level */ - if (!obj_cgroup_is_root(objcg)) + if (!obj_cgroup_is_root(objcg)) { ug->nr_memory +=3D nr_pages; + if (mem_cgroup_tiered_limits()) + gather_tier_charge(ug, folio, nr_pages); + } ug->pgpgout++; =20 WARN_ON_ONCE(folio_unqueue_deferred_split(folio)); @@ -5514,6 +5592,11 @@ void mem_cgroup_replace_folio(struct folio *old, str= uct folio *new) /* Force-charge the new page. The old one will be freed soon */ if (!obj_cgroup_is_root(objcg)) { page_counter_charge(&memcg->memory, nr_pages); + if (mem_cgroup_tiered_limits()) { + int slot =3D nid_tier_slot(folio_nid(new)); + + mem_cgroup_charge_tier(memcg, slot, nr_pages); + } if (do_memsw_account()) page_counter_charge(&memcg->memsw, nr_pages); } @@ -5558,6 +5641,23 @@ void mem_cgroup_migrate(struct folio *old, struct fo= lio *new) if (!objcg) return; =20 + if (!obj_cgroup_is_root(objcg) && mem_cgroup_tiered_limits()) { + struct mem_cgroup *memcg; + unsigned long nr_pages =3D folio_nr_pages(old); + int old_slot, new_slot; + + rcu_read_lock(); + memcg =3D obj_cgroup_memcg(objcg); + old_slot =3D nid_tier_slot(folio_nid(old)); + new_slot =3D nid_tier_slot(folio_nid(new)); + + if (old_slot !=3D new_slot) { + mem_cgroup_uncharge_tier(memcg, old_slot, nr_pages); + mem_cgroup_charge_tier(memcg, new_slot, nr_pages); + } + rcu_read_unlock(); + } + /* Transfer the charge and the objcg ref */ commit_charge(new, objcg); =20 @@ -5633,7 +5733,7 @@ bool mem_cgroup_sk_charge(const struct sock *sk, unsi= gned int nr_pages, if (!cgroup_subsys_on_dfl(memory_cgrp_subsys)) return memcg1_charge_skmem(memcg, nr_pages, gfp_mask); =20 - if (try_charge_memcg(memcg, gfp_mask, nr_pages) =3D=3D 0) { + if (try_charge_memcg(memcg, gfp_mask, nr_pages, NUMA_NO_NODE) =3D=3D 0) { mod_memcg_state(memcg, MEMCG_SOCK, nr_pages); return true; } --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oa1-f54.google.com (mail-oa1-f54.google.com [209.85.160.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6926C446072 for ; Fri, 7 Aug 2026 20:21:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134075; cv=none; b=d2EEUrn4ZWtNjV6qmHtmPnY8l+SxpIhgl7Yst5CplIwTYpuM2eDwcWK2RStwqymYfElnEwq+N1lFCi5/1DJ8ffw494V8ckP0GbdN15k2l1Zs2HfaZgq3cYuXKTXh24TbM90zIUe26mM+9JAlevhEsSqeSfOCfNZNKDiA38Q+sjQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134075; c=relaxed/simple; bh=REawthIeH+7C0lmHL+f/QK7bxBDTbcKr3jAoxiOaBOs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kG636eZe2Tf/CGl9ZaChZAgOd/XadkLYCTMHCs9wVol02NzKVKqzQBAGnbl8TXEwR4FA34tt4EgRUW3/kp4tgs6vt2zzceT0b/fUm93o4yXSu6Y+oKT3iPSpRKZcFqhCpaKKc3H9cgPy7VCXOnWfvlmuTSAFB7+qzxOHpg2KAZw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=phSnllFw; arc=none smtp.client-ip=209.85.160.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="phSnllFw" Received: by mail-oa1-f54.google.com with SMTP id 586e51a60fabf-455ca262ccbso2163217fac.1 for ; Fri, 07 Aug 2026 13:21:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134072; x=1786738872; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PmLOPzI7usjs7Qmnspd34ZMOrobza2FnhZDO80Oyo7g=; b=phSnllFwKWfjOo/wzy5MD2OXQvxwhFGBU6Sib/Cx5jJJeSswL5h720is5v7ual5QfK WX5C+aFU2zbA6co05x9t6ZdK1OmfAWcQpiqMEXzQBCUNgobMVYTQLNHxxCGs/+c6o3C3 YQnZ/SxavZ4dFXJqGrAyc9gnkN+DjXK3bMCj/hOt+2FGjZCCeRmBybWyS7M6hXfWkIOg UI47jYN/McIMlcZtztDEU+fZO4wekpZWZPh2sPsOcuCkBkOTAayUP3EwBz7HGWLhIgir dH+ZFhQQnPc8dYJ2JqBpEI1+LgMCSSmtb1V4OGuqMwtVArfcixaeMeKa22FiYUeGr/30 D+eQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134072; x=1786738872; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PmLOPzI7usjs7Qmnspd34ZMOrobza2FnhZDO80Oyo7g=; b=rnQkist14EY3g++UEa/wJhflx3CAFnv3aO+pgXiRX5EQya0NG4g+vZh/RIwiJ2wlWb 44XpmQ7BuFDGO2d2ynBAjjOKQy/vcEI9OOiokgMjYWEH8zufRs5OwnnK3A1ARiw1o149 7PE2X8NyT4ISuULS/EZkF/ZQfAj/rVDCXvKsf8CMV8WiBKd0CWtvAbemIaZUhrEPfOwn XeDuJUtf7X9Omch7iEFPXT4y3MuCpscjSxF7agB3ybIciTzQ4KsHWINuXTwWkKlQ3o1R oy6USletdCxlthKCV+zL6EWTNYBfShpf4Qbpx5VzoRLlIhHJrTUlV7WunNhTZ3yxN5cS 8+Dg== X-Forwarded-Encrypted: i=1; AHgh+Rq3s7FOncQ5u4HGqvZC9FUI0TXKUf7FksKe2FF9p9zKb2/WY3lV5SWQSu645ZNtf2P+TGx3iZB9D1nnFXA=@vger.kernel.org X-Gm-Message-State: AOJu0YyTM1nVxcgHLNAMbVWHdxwAlPcjiM2jr4usYUQ1irssrPKX3VAT mLmwVa/8R5R2GsereUHLoWD89m0xjE8773F46HmxEH0GfHvezbRAA1ik X-Gm-Gg: AR+sD10w0i738zF9GOmweliPwBSqnM4HrBHFacFmUSagojCJB4hwxlBAt/mkdaY8fAc VtEuA6vpBbuUo1Tb0oNxd9Pe4uRQFrBkPNBvHSmHbv04UICMM3rV8tNoz02/mQqU3ZoZGsUFFyR fmeiQ1SwoG8l3DlHl+svBwLw6n1H6R+5UMspbIXiMlqUsABhDv2cXn403qNg8EMcqGEr4TrjiuD 3GMsisZ/u+20olhQLDgEASXsh7RoxLQKaUfnlHik41vjaikzEvY9ZhW87t4IvF3JRUjlWkEXNkM /hSrnu/SwhIU6r1l4z3F99UN6fxzlY/NkdyUC9EsvO9teUFQB4Y7a3jPcyRO8WwWNgIh6uQsFy+ L7t3y1cdZ1TE0jcyFY0kaYqr15gQRgdcHRvQo4qibYjjhAB7foM00hhxBNzowZEcUlamuY06Kku BHTOqMsg5iCIdHOMbqgTCMqi3oQxbNzhWH3/SUq8mK9KgAqZhBG6uR+tYHL2rEwcSqbKKQ5Z/Da vE++lC+IR0YYUSDFGw= X-Received: by 2002:a05:6871:79a9:b0:455:d6a6:6fdf with SMTP id 586e51a60fabf-4599f0ef42dmr14021135fac.21.1786134072221; Fri, 07 Aug 2026 13:21:12 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:42::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1e348d0sm2692743fac.16.2026.08.07.13.21.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:11 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 08/14] mm/memcontrol: Make memory.low and memory.min tier-aware Date: Fri, 7 Aug 2026 13:20:51 -0700 Message-ID: <20260807202059.2620949-9-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On machines serving multiple workloads whose memory is isolated via the memory cgroup controller, it is currently impossible to enforce a fair distribution of tiered memory among the workloads, as the only enforceable limits have to do with total memory footprint, but not where that memory resides. This makes ensuring a consistent and baseline performance difficult, as each workload's performance is heavily impacted by workload-external factors such as which other workloads are co-located in the same host, and the order at which different workloads are started. Extend the existing memory.{low, min} protection to be tier-aware in order to enforce proportional best-effort and guaranteed memory protection of higher-tier memory. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 25 +++++++++++++++++++++---- mm/memcontrol.c | 11 ++++++++++- mm/vmscan.c | 15 +++++++++------ 3 files changed, 40 insertions(+), 11 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index f7a92b66330ec..ceba0fd6de184 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -18,6 +18,7 @@ #include #include #include +#include #include #include #include @@ -618,21 +619,37 @@ static inline bool mem_cgroup_unprotected(struct mem_= cgroup *target, } =20 static inline bool mem_cgroup_below_low(struct mem_cgroup *target, - struct mem_cgroup *memcg) + struct mem_cgroup *memcg, int nid) { if (mem_cgroup_unprotected(target, memcg)) return false; =20 + if (mem_cgroup_tiered_limits()) { + int slot =3D nid_tier_slot(nid); + + if (slot >=3D 0) + return READ_ONCE(memcg->tier[slot].elow) >=3D + page_counter_read(&memcg->tier[slot]); + } + return READ_ONCE(memcg->memory.elow) >=3D page_counter_read(&memcg->memory); } =20 static inline bool mem_cgroup_below_min(struct mem_cgroup *target, - struct mem_cgroup *memcg) + struct mem_cgroup *memcg, int nid) { if (mem_cgroup_unprotected(target, memcg)) return false; =20 + if (mem_cgroup_tiered_limits()) { + int slot =3D nid_tier_slot(nid); + + if (slot >=3D 0) + return READ_ONCE(memcg->tier[slot].emin) >=3D + page_counter_read(&memcg->tier[slot]); + } + return READ_ONCE(memcg->memory.emin) >=3D page_counter_read(&memcg->memory); } @@ -1142,13 +1159,13 @@ static inline bool mem_cgroup_unprotected(struct me= m_cgroup *target, return true; } static inline bool mem_cgroup_below_low(struct mem_cgroup *target, - struct mem_cgroup *memcg) + struct mem_cgroup *memcg, int nid) { return false; } =20 static inline bool mem_cgroup_below_min(struct mem_cgroup *target, - struct mem_cgroup *memcg) + struct mem_cgroup *memcg, int nid) { return false; } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index f3714dfd85aa0..025496794cb91 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5310,7 +5310,16 @@ void mem_cgroup_calculate_protection(struct mem_cgro= up *root, if (!root) root =3D root_mem_cgroup; =20 - page_counter_calculate_protection(&root->memory, &memcg->memory, recursiv= e_protection); + page_counter_calculate_protection(&root->memory, &memcg->memory, + recursive_protection); + + if (mem_cgroup_tiered_limits()) { + int nr_tier_slots =3D mt_nr_tier_slots(); + + for (int slot =3D 0; slot < nr_tier_slots; slot++) + page_counter_calculate_protection(&root->tier[slot], + &memcg->tier[slot], recursive_protection); + } } =20 static int charge_memcg(struct folio *folio, struct mem_cgroup *memcg, diff --git a/mm/vmscan.c b/mm/vmscan.c index ffe7ea3c5aff6..29f3f12042650 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -4190,7 +4190,7 @@ static bool lruvec_is_reclaimable(struct lruvec *lruv= ec, struct scan_control *sc struct mem_cgroup *memcg =3D lruvec_memcg(lruvec); DEFINE_MIN_SEQ(lruvec); =20 - if (mem_cgroup_below_min(NULL, memcg)) + if (mem_cgroup_below_min(NULL, memcg, lruvec_pgdat(lruvec)->node_id)) return false; =20 if (!lruvec_is_sizable(lruvec, sc)) @@ -5057,6 +5057,7 @@ static bool try_to_shrink_lruvec(struct lruvec *lruve= c, struct scan_control *sc) bool need_rotate =3D false, should_age =3D false; long nr_batch, nr_to_scan; int swappiness =3D get_swappiness(lruvec, sc); + int nid =3D lruvec_pgdat(lruvec)->node_id; struct mem_cgroup *memcg =3D lruvec_memcg(lruvec); =20 nr_to_scan =3D get_nr_to_scan(lruvec, sc, memcg, swappiness); @@ -5064,7 +5065,7 @@ static bool try_to_shrink_lruvec(struct lruvec *lruve= c, struct scan_control *sc) int delta; DEFINE_MAX_SEQ(lruvec); =20 - if (mem_cgroup_below_min(sc->target_mem_cgroup, memcg)) { + if (mem_cgroup_below_min(sc->target_mem_cgroup, memcg, nid)) { need_rotate =3D true; break; } @@ -5104,12 +5105,13 @@ static int shrink_one(struct lruvec *lruvec, struct= scan_control *sc) unsigned long reclaimed =3D sc->nr_reclaimed; struct mem_cgroup *memcg =3D lruvec_memcg(lruvec); struct pglist_data *pgdat =3D lruvec_pgdat(lruvec); + int nid =3D pgdat->node_id; =20 /* lru_gen_age_node() called mem_cgroup_calculate_protection() */ - if (mem_cgroup_below_min(NULL, memcg)) + if (mem_cgroup_below_min(NULL, memcg, nid)) return MEMCG_LRU_YOUNG; =20 - if (mem_cgroup_below_low(NULL, memcg)) { + if (mem_cgroup_below_low(NULL, memcg, nid)) { /* see the comment on MEMCG_NR_GENS */ if (READ_ONCE(lruvec->lrugen.seg) !=3D MEMCG_LRU_TAIL) return MEMCG_LRU_TAIL; @@ -6168,6 +6170,7 @@ static void shrink_node_memcgs(pg_data_t *pgdat, stru= ct scan_control *sc) }; struct mem_cgroup_reclaim_cookie *partial =3D &reclaim; struct mem_cgroup *memcg; + int nid =3D pgdat->node_id; =20 /* * In most cases, direct reclaimers can do partial walks @@ -6197,13 +6200,13 @@ static void shrink_node_memcgs(pg_data_t *pgdat, st= ruct scan_control *sc) =20 mem_cgroup_calculate_protection(target_memcg, memcg); =20 - if (mem_cgroup_below_min(target_memcg, memcg)) { + if (mem_cgroup_below_min(target_memcg, memcg, nid)) { /* * Hard protection. * If there is no reclaimable memory, OOM. */ continue; - } else if (mem_cgroup_below_low(target_memcg, memcg)) { + } else if (mem_cgroup_below_low(target_memcg, memcg, nid)) { /* * Soft protection. * Respect the protection only as long as --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-ot1-f41.google.com (mail-ot1-f41.google.com [209.85.210.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F1FC644780B for ; Fri, 7 Aug 2026 20:21:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134076; cv=none; b=RWP2+F0ghaOwOO0jssjeR4m1RvCKRKuofzynkz3p3IuKm+36VYKLNi4M1+V2lwDwXyBZMWlxNY5abmD/B2TIdA2UK2WiJul33AVkRu95RGH0Bjs8gJP5AT5Wbs1TxZSqUkp0W0IftOAGhH0VMxDhc9v4OrrCK/IQwJESua4NjYI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134076; c=relaxed/simple; bh=U4D2bds6vIRZf7Z7iayGAVoiAt9LSNRErOps55OMStg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WMqPhN9AnTrD/6HsnChAeSkRDBJ1SMPr4aJKNVRd2EQYOdpszpQtS6AyphcLnWM027d+rEb/LaiJw9uZIsPft+f2Xx8U22WzG/WdHdxPth9ZTtBOTI3+1cmCr4WLurcXyUYvaZPmxVfV5YC78GlD/1xjo5LI/W1fg/HpzOiGYfI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Jm2mHeYu; arc=none smtp.client-ip=209.85.210.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Jm2mHeYu" Received: by mail-ot1-f41.google.com with SMTP id 46e09a7af769-7eb34c17b96so2639682a34.2 for ; Fri, 07 Aug 2026 13:21:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134074; x=1786738874; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=qFD6EfRm/L4Zg2p1Y1cVpL/1kzGCx7Oqw9Sl1lvWZlg=; b=Jm2mHeYuVe1C9oWOfxJU+hn/H+3QP6Jmt88fiTN3I+2i9EwwDoJ+ULai5XxsbqdGdQ gWgKagcnMTa1TVQCz4QWA9PPWgN3WZo/X2APni0M6W7kb+PKf34iLxSN5DpIheKMjjOO JdOMvRIl3WdxXPVR3yeGZBgZknRAi0dj6x/gRSznp9N6k+GNiDM5PtfVNaHhAkazDkge ucya7o9Xrem5hzfQZY4YLKZUxjdnBdbhV82qsEftKMVsQpJvOB097B+D9fHsproSeyBD cIYYpqZTdhVnRnRH5IejK0HCpQPTrOUEmPAABdMTpgK/VePrtKHxlAfhJPOqbnyUACOB xemg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134074; x=1786738874; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=qFD6EfRm/L4Zg2p1Y1cVpL/1kzGCx7Oqw9Sl1lvWZlg=; b=Uvjfu5YxXsYtKpW6pYCG/e4FsDOsYk9Gs44Z0joX2pJt/hQw/lm6q3VotWNbNyq/XA HH6/mgqZshQoCeWsYlgOtsPJa5jZ30LYv1vywQ5DZfYqYnuXpEHyRbW82ugdyK0T5ntg ArRKqYv72yaKZINd9VmHRJwkzhc85Ndp4P1qljVQ5GeJLh0O63P7tNx5caCjxhRNKehz pYZMjCfHQozsBzN8FkszklYPCZwQI4pW09jmLZBOElx9RnfhOCz7wXujJs+UBchFP9Ne GQ8R5kVl63brI7Xa5SAHOKUneyVRH0lSsunnIlvds1lR38TWr94ygB7MSNOG1G17B5XY QS5Q== X-Forwarded-Encrypted: i=1; AHgh+Rpc9DnlieBIBq5MrHD/pKipF4nmDN7ReVziff0v/97zv5wChydQT61RzPQTHeoE5+DuRE5cv/u6nm0/PT0=@vger.kernel.org X-Gm-Message-State: AOJu0YxcSaRv7RgD0c/3TZsE3QNIfqPp3/QBpUc+sD0WBL+2sbkcYdlu EZwbKfZAF4/5xiltq7cn/lxY5hnncahU7P4govSpnJ6S+GFc/R81esW7 X-Gm-Gg: AR+sD13PyO8iDYFtrMHyhBczecREYq4T7UfVwXE2bMCsjVe+p6rv06Y5KCbRhWgtBuG TwUJ+x5ABJS08ioC9oFGYnueg6H6QpWvdavs6yDCmET+ITMdnXXuFB7VOw3l0sMlC6iI1eM8LVM x2BX43y8/YoPuGTlRbbCWgmGP794lPEOTtylQ2gUVJlr9IOOpfpfgPQ7B+bEIXKZGizsrujgqS/ IFYPR8XfD4V3NO5QOx8k1vF05Fatw3cBJeCTQJ2C5mXdVUWkSdzp3d3g3AnWzqVjDLAUSk+a13a kr+n7dZ9czvs6+MxIvqU7hz8Tn5XukJUknT8CSZLsQEce4flDYLXYfFnivyt+s5DwNzeAdvdz1U q/02CVqbSZRv3YNwG4PJ3Z8MzuIp2dS7sZABq8xnHzOH9/oRQv+8xf4gpV5xA8jTHRHV2G27Q7f 4nfBtX40aWlAvE9ygaeGZtPcTmiO3P56XjCVT8bu9dsmyKBMpLq3JeF9GylRLLdkYa9cOARZHt9 N03b8oi5Z9q1tYBn1Y= X-Received: by 2002:a05:6830:82fb:b0:7dc:c7aa:22bd with SMTP id 46e09a7af769-7f1e5ce763dmr15477492a34.6.1786134073875; Fri, 07 Aug 2026 13:21:13 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:58::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f35b56391dsm1914394a34.1.2026.08.07.13.21.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:13 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 09/14] mm/memcontrol: Make memory.high tier-aware Date: Fri, 7 Aug 2026 13:20:52 -0700 Message-ID: <20260807202059.2620949-10-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On machines serving multiple workloads whose memory is isolated via the memory cgroup controller, it is currently impossible to enforce a fair distribution of tiered memory among the workloads, as the only enforceable limits have to do with total memory footprint, but not where that memory resides. This makes ensuring consistent baseline performance difficult, as each workload's performance is heavily impacted by workload-external factors such as which other workloads are co-located in the same host, and the order in which the workloads are started. Extend the existing memory.high protection to be tier-aware. Depending on the combination of limit breaches, selectively reclaim on tiers: when memory.high is breached, perform reclaim on all tiers. When memory.high is safe but individual tier limits are breached, perform targeted reclaim on those tiers only. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 66 ++++++++++++++++++++++++++++++++++++++++--------- 1 file changed, 55 insertions(+), 11 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 025496794cb91..44ea465b2005d 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2387,6 +2387,28 @@ static int memcg_hotplug_cpu_dead(unsigned int cpu) return 0; } =20 +static bool memcg_tier_over_limit(struct mem_cgroup *memcg, + unsigned long *overage, int *breached_slot) +{ + int nr_tier_slots =3D mt_nr_tier_slots(); + + for (int slot =3D 0; slot < nr_tier_slots; slot++) { + unsigned long usage =3D page_counter_read(&memcg->tier[slot]); + unsigned long limit =3D READ_ONCE(memcg->tier[slot].high); + + if (usage <=3D limit) + continue; + + if (overage) + *overage =3D usage - limit; + if (breached_slot) + *breached_slot =3D slot; + return true; + } + + return false; +} + static unsigned long reclaim_high(struct mem_cgroup *memcg, unsigned int nr_pages, gfp_t gfp_mask) @@ -2395,10 +2417,19 @@ static unsigned long reclaim_high(struct mem_cgroup= *memcg, =20 do { unsigned long pflags; + const nodemask_t *reclaim_nodes =3D NULL; =20 if (page_counter_read(&memcg->memory) <=3D - READ_ONCE(memcg->memory.high)) - continue; + READ_ONCE(memcg->memory.high)) { + int slot; + + if (!mem_cgroup_tiered_limits()) + continue; + if (!memcg_tier_over_limit(memcg, NULL, &slot)) + continue; + + reclaim_nodes =3D mt_tier_nodes(slot); + } =20 memcg_memory_event(memcg, MEMCG_HIGH); =20 @@ -2406,7 +2437,7 @@ static unsigned long reclaim_high(struct mem_cgroup *= memcg, nr_reclaimed +=3D try_to_free_mem_cgroup_pages(memcg, nr_pages, gfp_mask, MEMCG_RECLAIM_MAY_SWAP, - NULL, NULL); + NULL, reclaim_nodes); psi_memstall_leave(&pflags); } while ((memcg =3D parent_mem_cgroup(memcg)) && !mem_cgroup_is_root(memcg)); @@ -2842,23 +2873,25 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, * reclaim, the cost of mismatch is negligible. */ do { - bool mem_high, swap_high; + bool mem_high, swap_high, tier_high; =20 mem_high =3D page_counter_read(&memcg->memory) > READ_ONCE(memcg->memory.high); swap_high =3D page_counter_read(&memcg->swap) > READ_ONCE(memcg->swap.high); + tier_high =3D mem_cgroup_tiered_limits() && + memcg_tier_over_limit(memcg, NULL, NULL); =20 /* Don't bother a random interrupted task */ if (!in_task()) { - if (mem_high) { + if (mem_high || tier_high) { schedule_work(&memcg->high_work); break; } continue; } =20 - if (mem_high || swap_high) { + if (mem_high || swap_high || tier_high) { /* * The allocating tasks in this cgroup will need to do * reclaim or be throttled to prevent further growth @@ -4967,13 +5000,24 @@ static ssize_t memory_high_write(struct kernfs_open= _file *of, =20 for (;;) { unsigned long nr_pages =3D page_counter_read(&memcg->memory); - unsigned long reclaimed; + unsigned long reclaimed, charge; + const nodemask_t *reclaim_nodes =3D NULL; =20 if (high !=3D READ_ONCE(memcg->memory.high)) break; =20 - if (nr_pages <=3D high) - break; + if (nr_pages <=3D high) { + int slot; + + if (!mem_cgroup_tiered_limits()) + break; + if (!memcg_tier_over_limit(memcg, &charge, &slot)) + break; + + reclaim_nodes =3D mt_tier_nodes(slot); + } else { + charge =3D nr_pages - high; + } =20 if (signal_pending(current)) break; @@ -4988,9 +5032,9 @@ static ssize_t memory_high_write(struct kernfs_open_f= ile *of, continue; } =20 - reclaimed =3D try_to_free_mem_cgroup_pages(memcg, nr_pages - high, + reclaimed =3D try_to_free_mem_cgroup_pages(memcg, charge, GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, - NULL, NULL); + NULL, reclaim_nodes); =20 if (!reclaimed && !nr_retries--) break; --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oo1-f54.google.com (mail-oo1-f54.google.com [209.85.161.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CF40544BCB5 for ; Fri, 7 Aug 2026 20:21:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134079; cv=none; b=W1EnFG1iWWfA6BSbTfo0hakm0nScy4I8KUkLQjfVwmmDNaw1hYbM2j0HR20QGx5CJag24bmQsGNMRVs3GYwlUizICqYOCGs+6WEVgsn3BrhLs6EbWf06nIuYP5xsILnH7a19H+p8Ps2aozi5ZMqz9VYtrmheCOkTF80t+BV6c3Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134079; c=relaxed/simple; bh=WtW+xUg5a0joD1Hd5vyHXE9LR+eLMwzjl5LF9juBCBA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=C1s+wY/2gB7in7mN2JBeDdgp5Ki/z5IRH5ysrZ4+TeKIOsLJChdUwZTehK4vGKlFUTUuoXFWg6BZTtVmSSgRnur0hylP5E+eubUUb4tzdNwHzX5brzpFTNghH5mA5SplBuRl4WYLXWTsEJjSqz4iIZc0fOJQnM3brExi7WJNDw8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=sJdBQJ02; arc=none smtp.client-ip=209.85.161.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="sJdBQJ02" Received: by mail-oo1-f54.google.com with SMTP id 006d021491bc7-6acc15016f1so2315821eaf.3 for ; Fri, 07 Aug 2026 13:21:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134076; x=1786738876; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UreOU1bs1UnYmmJ7pTF9AE47T3l1T/OGD/BO7qMZ90c=; b=sJdBQJ02FjELftf2PnRoZy3Gak1O7QnghngDYWblm1KT4ebWqv1U24Zf9qiW6A2+jB 0BD4fp1Rzr/UhHN/jNbc4Yy1BqYsZmOXQ6Jw61PmyFUqfDBaZ8Yu5rDqYGB6NFzy8P98 EwO2s295AERRPL8OK2VI1ocjoZ7l7GD/LksAMT+w2iZAhRoN9uwYlPMiSC+ClX/OjkpE oi7ZXfG4Gwn9+uQsF0q4a9Z9xrcCsc3l0qnvwwCAg4VK+j/biv62AAC0F1LbTqqryuZD YSRbqdDQXFXczAqt8WkdeDAajhJHcnxqk/MLY/66fhZVWaggbb5kWmCHLgLpsdDKXLmk ja3A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134076; x=1786738876; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=UreOU1bs1UnYmmJ7pTF9AE47T3l1T/OGD/BO7qMZ90c=; b=M9adlv2Rhe/dKw+RamXQpI14Dq89WCnxwWO0P0PL8M0UaEHINAPyY4V9T2NjAbQ4Zg Z98rLCu6rIRx9I3/5bicXXBQ2uXS15H+RwDs50LtVuO8xU3r6Gh9Gtp1Nb8OPBtt2uaf tJIBgXsDrz9ykVcLdg54KTpHHXrN5/6Oi0KnIBh2skH8XumDw6HVOtPc/vehuKGmux/W QgDeMxfqlNLYI9AHgj2pXdagyxA5llSa/BJdjJSpm5XhyvCU6RR3gorz6Le4vAGTPwOU yIVsC6qorKJxTN4nOBiZK4S/IbreA/rz46LzQtbSgaQSdhEFlZnRhrwR2RmTWmSTCa6y lzrg== X-Forwarded-Encrypted: i=1; AHgh+RpUfIEGN1FItKxl7bL+T3iL/u0f2JNNPQb4FQVEvSeVQcDSAG4r7j9Ku7L+voB0RVq3JjChImSMtZzEK4E=@vger.kernel.org X-Gm-Message-State: AOJu0Yxm35hBAr6kP5Tvtudq0dEJDPOswNfA6qJ7vb+wZv98ITUkZtfm 09hiUj4kieUdms4cvoppMBIa90yN9lyrM9mLc7L3WQuyYTtVdOhuOQqk X-Gm-Gg: AR+sD12L0gQX1WhnLtl7tfE6NsgdS+IlFVp5LG8bCeXl7BtjIyqvhrshVczlEobFWX9 JJqodLUhn5IJAHRcVhcoy+QI9Xgve4PH7tPeHVaKG/UaAbhFg3prkwfX3g59BZNSIlxErbLaYFp 9lOJi4m8o9MAadpcykeBMyU/HOSqbfRb/DwE1graaL6pm3kxFwSlVLvf6vDi3kdy0NEXkT+qnlE pca0GI9xqsyimGfkbMvxXHxsBhRnYOzaQdPMIaU/fX/ntXAQ24p5HcSxREG5Jud8r6X1JrhfmEz MthvKVtTwZH96wPSrEzdeGGl+dJVSJwcDN5j4u5kvje1oBU0819dUEwHwfbs8Wd9g0QGaBSzEe5 MeLt4WL9+O3O0q7GXIBmRVIGindlIBJ0BNyyM/Ikulr94jWCAJG2hq//YuCxRzNAELky1ssSkkD Y1bbcb1tN9IGIsc+msugPaGUMKQKnjy+446RJmNRVPZ22FVTtdr8K2dOMr50pgzqoi7sYT0DMyK 3vgNS48/Zd6p0p+kf4= X-Received: by 2002:a05:6820:4dfc:b0:6a1:7790:258e with SMTP id 006d021491bc7-6ae96ec782bmr14629606eaf.18.1786134075612; Fri, 07 Aug 2026 13:21:15 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:58::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02be8ede0sm3225544eaf.12.2026.08.07.13.21.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:14 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 10/14] mm/memcontrol: Make memory.max tier-aware Date: Fri, 7 Aug 2026 13:20:53 -0700 Message-ID: <20260807202059.2620949-11-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On machines serving multiple workloads whose memory is isolated via the memory cgroup controller, it is currently impossible to enforce a fair distribution of tiered memory among the workloads, as the only enforceable limits have to do with total memory footprint, but not where that memory resides. Extend the existing memory.max limit to be tier-aware. A folio charge is now attempted against the page_counter of the tier the folio was allocated on, in addition to memory and memsw. When the tier counter is over its limit, reclaim is targeted at that tier's nodes only. Tier counters are parent-linked index-for-index with the memcg hierarchy, so the counter reported by page_counter_try_charge() belongs to an ancestor of the charging memcg. memcg->tier is a separate allocation, so container_of() cannot recover the owner; tier_counter_mem_cgroup() walks the chain instead. This is only done on the charge failure path. mem_cgroup_margin() takes the tier slot so that the retry check and the OOM re-check under oom_lock consider the tier that actually failed. Taking the minimum across all tiers would report no headroom whenever any tier is full, which would disable the pile-on guard in mem_cgroup_out_of_memory() for plain memory.max breaches as well. Note that stock is currently a per-memcg resource and does not distinguish between tiers. There is ongoing work to change this, however. Until then, tiered usage may transiently breach the max limit. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 112 ++++++++++++++++++++++++++++++++++++------------ 1 file changed, 84 insertions(+), 28 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 44ea465b2005d..de6762520f475 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -1570,14 +1570,32 @@ static struct page_counter *mem_cgroup_tier_counter= (struct mem_cgroup *memcg, return &memcg->tier[slot]; } =20 +static struct mem_cgroup *tier_counter_mem_cgroup(struct mem_cgroup *memcg, + struct page_counter *counter, + int slot) +{ + struct mem_cgroup *iter; + + for (iter =3D memcg; iter; iter =3D parent_mem_cgroup(iter)) { + if (&iter->tier[slot] =3D=3D counter) + return iter; + } + + /* the failing counter is always an ancestor of the given memcg */ + VM_WARN_ON_ONCE(1); + return memcg; +} + /** * mem_cgroup_margin - calculate chargeable space of a memory cgroup * @memcg: the memory cgroup + * @slot: the memory tier slot * - * Returns the maximum amount of memory @mem can be charged with, in - * pages. + * Returns the maximum amount of memory @mem can be charged with, in pages. + * If the system has tiered memcg limits, then it returns the minimum of t= he + * tiered margin and the memcg margin. */ -static unsigned long mem_cgroup_margin(struct mem_cgroup *memcg) +static unsigned long mem_cgroup_margin(struct mem_cgroup *memcg, int slot) { unsigned long margin =3D 0; unsigned long count; @@ -1597,6 +1615,21 @@ static unsigned long mem_cgroup_margin(struct mem_cg= roup *memcg) margin =3D 0; } =20 + if (mem_cgroup_tiered_limits()) { + struct page_counter *tier_counter; + + tier_counter =3D mem_cgroup_tier_counter(memcg, slot); + if (!tier_counter) + return margin; + + count =3D page_counter_read(tier_counter); + limit =3D READ_ONCE(tier_counter->max); + if (count < limit) + margin =3D min(margin, limit - count); + else + margin =3D 0; + } + return margin; } =20 @@ -1945,7 +1978,7 @@ void __memcg_memory_event(struct mem_cgroup *memcg, EXPORT_SYMBOL_GPL(__memcg_memory_event); =20 static bool mem_cgroup_out_of_memory(struct mem_cgroup *memcg, gfp_t gfp_m= ask, - int order) + int order, int slot) { struct oom_control oc =3D { .zonelist =3D NULL, @@ -1959,7 +1992,7 @@ static bool mem_cgroup_out_of_memory(struct mem_cgrou= p *memcg, gfp_t gfp_mask, if (mutex_lock_killable(&oom_lock)) return true; =20 - if (mem_cgroup_margin(memcg) >=3D (1 << order)) + if (mem_cgroup_margin(memcg, slot) >=3D (1 << order)) goto unlock; =20 /* @@ -1977,7 +2010,8 @@ static bool mem_cgroup_out_of_memory(struct mem_cgrou= p *memcg, gfp_t gfp_mask, * Returns true if successfully killed one or more processes. Though in so= me * corner cases it can return true even without killing any process. */ -static bool mem_cgroup_oom(struct mem_cgroup *memcg, gfp_t mask, int order) +static bool mem_cgroup_oom(struct mem_cgroup *memcg, gfp_t mask, int order, + int slot) { bool locked, ret; =20 @@ -1989,7 +2023,7 @@ static bool mem_cgroup_oom(struct mem_cgroup *memcg, = gfp_t mask, int order) if (!memcg1_oom_prepare(memcg, &locked)) return false; =20 - ret =3D mem_cgroup_out_of_memory(memcg, mask, order); + ret =3D mem_cgroup_out_of_memory(memcg, mask, order, slot); =20 memcg1_oom_finish(memcg, locked); =20 @@ -2388,13 +2422,15 @@ static int memcg_hotplug_cpu_dead(unsigned int cpu) } =20 static bool memcg_tier_over_limit(struct mem_cgroup *memcg, - unsigned long *overage, int *breached_slot) + unsigned long *overage, int *breached_slot, + bool high) { int nr_tier_slots =3D mt_nr_tier_slots(); =20 for (int slot =3D 0; slot < nr_tier_slots; slot++) { unsigned long usage =3D page_counter_read(&memcg->tier[slot]); - unsigned long limit =3D READ_ONCE(memcg->tier[slot].high); + unsigned long limit =3D high ? READ_ONCE(memcg->tier[slot].high) : + READ_ONCE(memcg->tier[slot].max); =20 if (usage <=3D limit) continue; @@ -2425,7 +2461,7 @@ static unsigned long reclaim_high(struct mem_cgroup *= memcg, =20 if (!mem_cgroup_tiered_limits()) continue; - if (!memcg_tier_over_limit(memcg, NULL, &slot)) + if (!memcg_tier_over_limit(memcg, NULL, &slot, true)) continue; =20 reclaim_nodes =3D mt_tier_nodes(slot); @@ -2719,6 +2755,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, unsigned long pflags; bool allow_spinning =3D gfpflags_allow_spinning(gfp_mask); int slot =3D -1; + const nodemask_t *reclaim_nodes; =20 if (mem_cgroup_tiered_limits()) { slot =3D nid_tier_slot(nid); @@ -2737,6 +2774,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, batch =3D nr_pages; =20 reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP; + reclaim_nodes =3D NULL; =20 if (do_memsw_account() && !page_counter_try_charge(&memcg->memsw, batch, &counter)) { @@ -2745,15 +2783,23 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, goto reclaim; } =20 - if (page_counter_try_charge(&memcg->memory, batch, &counter)) { + if (tier_counter && + !page_counter_try_charge(tier_counter, nr_pages, &counter)) { + mem_over_limit =3D tier_counter_mem_cgroup(memcg, counter, slot); + reclaim_nodes =3D mt_tier_nodes(slot); + goto reclaim; + } + + if (!page_counter_try_charge(&memcg->memory, batch, &counter)) { + mem_over_limit =3D mem_cgroup_from_counter(counter, memory); + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, batch); if (tier_counter) - page_counter_charge(tier_counter, nr_pages); - goto done_restock; + page_counter_uncharge(tier_counter, nr_pages); + goto reclaim; } =20 - if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, batch); - mem_over_limit =3D mem_cgroup_from_counter(counter, memory); + goto done_restock; =20 reclaim: if (batch > nr_pages) { @@ -2782,13 +2828,13 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, psi_memstall_enter(&pflags); nr_reclaimed =3D try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages, gfp_mask, reclaim_options, - NULL, NULL); + NULL, reclaim_nodes); psi_memstall_leave(&pflags); =20 - if (mem_cgroup_margin(mem_over_limit) >=3D nr_pages) + if (mem_cgroup_margin(mem_over_limit, slot) >=3D nr_pages) goto retry; =20 - if (!drained) { + if (!drained && !reclaim_nodes) { drain_all_stock(mem_over_limit); drained =3D true; goto retry; @@ -2824,7 +2870,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, * couldn't make any progress. */ if (mem_cgroup_oom(mem_over_limit, gfp_mask, - get_order(nr_pages * PAGE_SIZE))) { + get_order(nr_pages * PAGE_SIZE), slot)) { passed_oom =3D true; nr_retries =3D MAX_RECLAIM_RETRIES; goto retry; @@ -2880,7 +2926,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, swap_high =3D page_counter_read(&memcg->swap) > READ_ONCE(memcg->swap.high); tier_high =3D mem_cgroup_tiered_limits() && - memcg_tier_over_limit(memcg, NULL, NULL); + memcg_tier_over_limit(memcg, NULL, NULL, true); =20 /* Don't bother a random interrupted task */ if (!in_task()) { @@ -5011,7 +5057,7 @@ static ssize_t memory_high_write(struct kernfs_open_f= ile *of, =20 if (!mem_cgroup_tiered_limits()) break; - if (!memcg_tier_over_limit(memcg, &charge, &slot)) + if (!memcg_tier_over_limit(memcg, &charge, &slot, true)) break; =20 reclaim_nodes =3D mt_tier_nodes(slot); @@ -5073,12 +5119,22 @@ static ssize_t memory_max_write(struct kernfs_open_= file *of, =20 for (;;) { unsigned long nr_pages =3D page_counter_read(&memcg->memory); + unsigned long charge; + const nodemask_t *reclaim_nodes =3D NULL; + int slot =3D -1; =20 if (max !=3D READ_ONCE(memcg->memory.max)) break; =20 - if (nr_pages <=3D max) - break; + if (nr_pages <=3D max) { + if (!mem_cgroup_tiered_limits()) + break; + if (!memcg_tier_over_limit(memcg, &charge, &slot, false)) + break; + reclaim_nodes =3D mt_tier_nodes(slot); + } else { + charge =3D nr_pages - max; + } =20 if (signal_pending(current)) break; @@ -5087,22 +5143,22 @@ static ssize_t memory_max_write(struct kernfs_open_= file *of, if (memcg_is_dying(memcg)) break; =20 - if (!drained) { + if (!drained && !reclaim_nodes) { drain_all_stock(memcg); drained =3D true; continue; } =20 if (nr_reclaims) { - if (!try_to_free_mem_cgroup_pages(memcg, nr_pages - max, + if (!try_to_free_mem_cgroup_pages(memcg, charge, GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, - NULL, NULL)) + NULL, reclaim_nodes)) nr_reclaims--; continue; } =20 memcg_memory_event(memcg, MEMCG_OOM); - if (!mem_cgroup_out_of_memory(memcg, GFP_KERNEL, 0)) + if (!mem_cgroup_out_of_memory(memcg, GFP_KERNEL, 0, slot)) break; cond_resched(); } --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oa1-f46.google.com (mail-oa1-f46.google.com [209.85.160.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B226844A400 for ; Fri, 7 Aug 2026 20:21:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134080; cv=none; b=fzjg8UYdZ9RdsC2ip1cnVyOpaVhkpWDKC/L351G0xo0Ku28vnMi1/FHWBpcq72a8v1MwQYAi1N8WGXkiYhxETNchJQxX6TgDOlerk1iFIShmRQiHcQptkMI3RmAcMkd8wV/Ygo0GDK5CQGy38homy2ze+LEV/1FA3MFF5CIcGkk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134080; c=relaxed/simple; bh=GZJxuSitrZmCJ6vWZysBfmLLWrvpgKrXs+OpthlLIUk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=AaUZxhJMEnZtILCsGi1Zb3VM1RrymnuTe/riJKIu5UYDMjsoqDhG0hvh67xS9LETudr/4WHz0nja60vs0iXa9UJeTnhXq2cF44NJBCi11JYZIgM8E7hl9vzCe2Dmi33isMVCfIiUK36AQkN966SYf0ITC2U3Usb34+g740b8dVI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Grn9cZH1; arc=none smtp.client-ip=209.85.160.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Grn9cZH1" Received: by mail-oa1-f46.google.com with SMTP id 586e51a60fabf-451a49abd8aso2062439fac.3 for ; Fri, 07 Aug 2026 13:21:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134077; x=1786738877; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wktIFByfmvNWwU5aKwDRUJNG+vu9pXnbr6sGB9f+FbU=; b=Grn9cZH1nmBb7d8y0HdAuqVm+AehKqoQAHhCe482dRInzIloDi+T+PRn4de+RWo88K Evlv+u3oLZjpNRknXnIYhFdexyNPAmcgQAkDEiBgfLHSYBf7wuCiqQH/CFqOYKieb91C SuiGF8sTtRoS3v4VImaGezQMpRbFaAwhpkuLOli2kA+IfovGmZ8D5sJpGXuzjNsab1Go 3GT+6/VULX8P0GIlXpbSnNR5EbAKlLWFT+QcoFekRcl/MA1XyU468SBa9+KSxDqVTig1 KYggbGsu33JXsLhBXTjW40oLGG5RrVt5zAoSmQjESPMpDfICFEXb2apPnjm9yzEz6NwP MB2A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134077; x=1786738877; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=wktIFByfmvNWwU5aKwDRUJNG+vu9pXnbr6sGB9f+FbU=; b=dnL4NwMWMMaoqALYsaCqOgCjmaw06EfHwAE5bCksYeBPoqViNhqC/5mR6gtgZilWnB XGWwWDCNuYaVZsDMcUHROyOeuru7diaP2LyJHuXFnhzI+64BhyiuIwEa0+eIjKf2vMMg 9plWnEvTG0kwK5qpmZl9+P4LuXbYqbodUdaAu3mi2Iparb1ttZr5udiPlGIX2G1+Dtv3 HvOrA6A/7laWe3M5S7oO4DG7D2VqRA7OpS1b4EJUjbdokhVj0ymPM1Hm+7pH/zpB+eCr thcJUov2GhyS5aqPZJeqLW1VhEWDzT8hkXhu8vnQiAFZ1LX4yeVBYcdPZgqv7A3gMNLL dEgg== X-Forwarded-Encrypted: i=1; AHgh+RoRop20K+jCb92OLdy+KBNn6I1vcb4QbiW00F+UE5jAiJDiEh/8DDeek1ELemfO5pS2VXhdr9OERFGtOaI=@vger.kernel.org X-Gm-Message-State: AOJu0YxhfxDHdGEN0eloN5vn6p62L03V7TLjUv895jcM9zluRGDg8ign j6tbP8VIuZ4VvHJ4xPfHFtxnwPg7LbYGenk1AWCocKnIkNE730Ju7veR X-Gm-Gg: AR+sD11/6BnuvMsYoqqRQQFzzea4FlEXCULUd/fjZptw7ThWJ3HuzBwwpOEgQ1Oro4m 5vFZtYAPwd4rJxKg6M8uI7EB1rNFhR/s3agjIdra4RsJba+V1lq+C/L0EYrWuclC7vS/oBnYaRn tbDZp+7n8AwPHPUMxQ+tuQD+alZuJGpM+deHZzgAH9Ij0xZHhQsIvoiYHG0FJkxOoB4yaBOCqxI vBaUSVv4SzWr4U7SdzzEuy7eMxoOl7eQ+EHnrg7WlEsRlHWjj6bqBMkIXMvjBQRnnrh8up9351X 5VEtBJnUIGxbN7PyG9BbA70Sb0WgqqlIRWbHurMaceDp4z/G+C4Jn/LvPK7V/EnZFTMdSiGCWrE 73Gj4OKfq+zopsiOmT+CMaYonMdWJOtVsevh63o85w5sc6Wsdc0LXqZKYnoDqxOWzjiW1qcd3Vt Uc02FXJPe7ucHaDqsWQclehlbZDNsPJlrct1MzgqyQGP+Nzdd35UABsyJ0jvfAi+OMJ0crfaF/P AJploAsmTZE1G5OZHLa1Duw9+IYcMInn/UE0Ylk4g== X-Received: by 2002:a05:6870:c250:b0:453:9b2b:f3f1 with SMTP id 586e51a60fabf-459ff6b0630mr3346442fac.5.1786134077083; Fri, 07 Aug 2026 13:21:17 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:55::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1e28910sm2664792fac.15.2026.08.07.13.21.16 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:16 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 11/14] mm/memcontrol, migrate: Transfer tier charge on migration Date: Fri, 7 Aug 2026 13:20:54 -0700 Message-ID: <20260807202059.2620949-12-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Folio migration is charge-neutral, and existing migration paths take advantage of this fact to simply force destination folio charges or just transfer memcg data across folios. Per-tier memcg limits break this assumption. A migration across tiers (i.e. promotion or demotion) keeps the memcg-level charge neutral, but the per-memcg tier charges change. As a result, the destination tier may go over the limit. Charge the destination separately instead, from migrate_folio_unmap where the destination folio has just been allocated but can still be rolled back. This charge attempts a single pass at reclaim if it goes over the hard limit, and fails the migration if not enough headroom is created on the destination memcg tier. Note that this source of migration failure returns -EBUSY and not -ENOMEM since -ENOMEM will attempt the migration again by splitting the folio and aborting the batch, which both do nothing to reduce the memory usage of the memcg tier. We also don't try too hard to reclaim here (__GFP_NORETRY) since failing migrations is cheap, and we don't want to OOM kill because of a promotion attempt. One side effect is that cross-tier migrations now hold both folios' charges until the source is freed, the same way mem_cgroup_replace_folio temporarily holds a duplicate charge. No-op unless the system has tiered memcg limits enabled. Suggested-by: Johannes Weiner Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 8 ++++++++ mm/memcontrol.c | 35 +++++++++++++++++++++++++++++++++++ mm/migrate.c | 30 +++++++++++++++++++++++++++--- 3 files changed, 70 insertions(+), 3 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index ceba0fd6de184..9c2f11191a499 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -708,6 +708,8 @@ static inline void mem_cgroup_uncharge_folios(struct fo= lio_batch *folios) =20 void mem_cgroup_replace_folio(struct folio *old, struct folio *new); void mem_cgroup_migrate(struct folio *old, struct folio *new); +int mem_cgroup_migrate_charge(struct folio *src, struct folio *dst, + bool force); =20 /** * mem_cgroup_lruvec - get the lru list vector for a memcg & node @@ -1204,6 +1206,12 @@ static inline void mem_cgroup_migrate(struct folio *= old, struct folio *new) { } =20 +static inline int mem_cgroup_migrate_charge(struct folio *src, + struct folio *dst, bool force) +{ + return 0; +} + static inline struct lruvec *mem_cgroup_lruvec(struct mem_cgroup *memcg, struct pglist_data *pgdat) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index de6762520f475..1161934e81380 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5716,6 +5716,41 @@ void mem_cgroup_replace_folio(struct folio *old, str= uct folio *new) rcu_read_unlock(); } =20 +/** + * mem_cgroup_migrate_charge - Charge a migration destination up front. + * @src: Folio being migrated away from. + * @dst: Folio being migrated to. + * @force: Charge even if the destination tier is at its limit. + * + * Folio migrations result in a net 0 memcg charge, but the node location = of the + * charge may change during promotions or demotions. When this happens, ch= arge + * @dst in its own right instead of inheriting @src's charge. + * + * Return: 0, or -ENOMEM if @dst could not be charged. + */ +int mem_cgroup_migrate_charge(struct folio *src, struct folio *dst, bool f= orce) +{ + struct mem_cgroup *memcg; + gfp_t gfp =3D GFP_KERNEL; + int ret; + + if (mem_cgroup_disabled() || !folio_memcg_charged(src)) + return 0; + + if (!mem_cgroup_tiered_limits() || + nid_tier_slot(folio_nid(src)) =3D=3D nid_tier_slot(folio_nid(dst))) + return 0; + + /* Refuse the migration if the first reclaim round fails */ + gfp |=3D force ? __GFP_NOFAIL : __GFP_NORETRY; + + memcg =3D get_mem_cgroup_from_folio(src); + ret =3D charge_memcg(dst, memcg, gfp); + mem_cgroup_put(memcg); + + return ret; +} + /** * mem_cgroup_migrate - Transfer the memcg data from the old to the new fo= lio. * @old: Currently circulating folio. diff --git a/mm/migrate.c b/mm/migrate.c index ab15a4dddd047..45d6d23d53859 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -862,7 +862,13 @@ void folio_migrate_flags(struct folio *newfolio, struc= t folio *folio) folio_copy_owner(newfolio, folio); pgalloc_tag_swap(newfolio, folio); =20 - mem_cgroup_migrate(folio, newfolio); + /* + * For failable memcg charge transfers (demotion / promotion) the charge + * has already been transferred at this point. For everyone else simply + * transfer the charge here, where it can no longer fail. + */ + if (!folio_memcg_charged(newfolio)) + mem_cgroup_migrate(folio, newfolio); } EXPORT_SYMBOL(folio_migrate_flags); =20 @@ -1216,7 +1222,7 @@ static void migrate_folio_done(struct folio *src, static int migrate_folio_unmap(new_folio_t get_new_folio, free_folio_t put_new_folio, unsigned long private, struct folio *src, struct folio **dstp, enum migrate_mode mode, - struct list_head *ret) + bool force_charge, struct list_head *ret) { struct folio *dst; int rc =3D -EAGAIN; @@ -1228,6 +1234,17 @@ static int migrate_folio_unmap(new_folio_t get_new_f= olio, dst =3D get_new_folio(src, private); if (!dst) return -ENOMEM; + + if (mem_cgroup_migrate_charge(src, dst, force_charge)) { + if (put_new_folio) + put_new_folio(dst, private); + else + folio_put(dst); + if (ret) + list_move_tail(&src->lru, ret); + return -EBUSY; + } + *dstp =3D dst; =20 dst->migrate_info =3D 0; @@ -1918,8 +1935,15 @@ static int migrate_pages_batch(struct list_head *fro= m, continue; } =20 + /* + * Hotplug must not be refused: offline_pages() retries + * indefinitely and ignores migration failures, so a + * refusal would hang it rather than fail it. + */ rc =3D migrate_folio_unmap(get_new_folio, put_new_folio, - private, folio, &dst, mode, ret_folios); + private, folio, &dst, mode, + reason =3D=3D MR_MEMORY_HOTPLUG, + ret_folios); /* * The rules are: * 0: folio will be put on unmap_folios list, --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oa1-f46.google.com (mail-oa1-f46.google.com [209.85.160.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 11F9F45349A for ; Fri, 7 Aug 2026 20:21:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134081; cv=none; b=trZMthQVXzO88VKBUsEPaexlYt6wBb+/bwtY+LFwTciCRAdJTiPFaDJxt9FIFSIvGOX687tKXPhefmlbm+T1xRdHkc2BHQiFTSQra6rscNiBT5KJTVZceEIuoHqjAmNgkhQ3p2nnIAgxq/e9w/KCcqZQyecA9hYPA8bhAtIFtvk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134081; c=relaxed/simple; bh=ZV2VwtllGfxcjLjy+Ic7pvWCqB5qypDbPzE1Afw8BhU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=QeEimYvxRgoUdWfb6x+KbyFDTcuBA2xNVAIkEYKRFe6wjEH2SRyT6c00Y+0GGz47NVzasY+fF8ITwHemOE1Kwn8J+shvjXfkY8D947XoOncXnmjIGAGJgq3CQK7QlKtRCUcDzM6JXUTr6TA05tkXZko3HaDDwU/LUaXD4CyDCrw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=k4uwRzY3; arc=none smtp.client-ip=209.85.160.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="k4uwRzY3" Received: by mail-oa1-f46.google.com with SMTP id 586e51a60fabf-43b7e186a0cso1334282fac.0 for ; Fri, 07 Aug 2026 13:21:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134079; x=1786738879; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=8XkMxJlHKp/V+Gx87DIQUe3nc4ANw9n5bBTpUGWIjt4=; b=k4uwRzY39CqdgQlzoByv3mAnaL/d0rpYEV4EHO/JzzWMkbwakeJqvOJi3PWdtqyVDA RY47niC6f07RohHD8DYWIW//ImootfUXQwEUoClV2sUurLWQDzi9yuDdR4BecBaXPR5l 2kg6Vgj1ufN9Y9aC9KZ33YPMys/WQVVzkcBod+IqPGQF6EK8/jI5jwj6UVXDL8YKZUny nduMSV7S64naM+iwfq9dKSL31+/2730FM1UnhFkxyMbbK/LzyI201ZJ+/chiflop/RqS WTTIICeD/BZTCXoq4rHLAbLNsR92DRVCFG4o107jjjIdQJdWvs2PJNnghcHXGrhCBWWr BpIA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134079; x=1786738879; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=8XkMxJlHKp/V+Gx87DIQUe3nc4ANw9n5bBTpUGWIjt4=; b=V5yQK0zyvvEgoMZHQnjGQV5iUwJFa/pYgPQch9pEHlSvI/MneGgHohWtIuLVYFpoUG O4OuSI2p8OhpeM70t5Ye/0Kd6hhc6uUNEeug7o5drZ5P+Usu1kS+42e4nfAyEmZVIuMV hrCC7oHDmvS46sCCzMfCEVbHmJNoxJFt3ulYoG5+AqZPLbqzMCKKEhRWFiyVvfi6IK/V O/rCLGMrjqamRlxIIgDh4QcM4eog7+liq/CrfrycFZb8So+VOb1YZyQlcrrSHpOf0gRZ do/wDtS/EOPbRt1U+aZY7iHOdk6yM7BvRtuV6zAXmR38cL4mCsQKCMpID9YX1fbGcccB O4nA== X-Forwarded-Encrypted: i=1; AHgh+Ro2ZQm/IPx0J8O716obcSUssJg8BaC4AeaW332MnyYTLm8gM9QG5WbJIOBARemO8AACjCp7oggcJvSeutY=@vger.kernel.org X-Gm-Message-State: AOJu0YzBuNvhSzK+1T82SQa/n0TM3NeAFHMFJKf+xcfJ8OAFN8xAEM2Y mBVtq6Hkygjdt6GPjjsTPsMbpS1XYq2s5tHYKcsDAkw+i6Fm4ZgkB+eI X-Gm-Gg: AR+sD10pakELvJzXvfzgjcFyMaVW3FZGiP9MRToFpDPHENYl+i2FyqOb+BlIZ320eDc TyrB7glMcymduZ/1jof/IHrnq6mqAaJI1o9XOTPQRL7ske4U/0nPgVQgQIkIsO/kdhLYLj0nKTc rcQc9JyTiU2iLDPfvP2hDiWJHZdoM6FngH1q+QQdeYBgwOX/Hag2bw/7nxDbvDWe8QH3pJjis6w A6lr56yQTeDavW2ywjcL6xgmsGMwYsA2Kp+mrcBxAcymKGyWfKaLIqGyy55l9q9lX9i31E8iAFT FhzBxF2gq6R8WFfdM35BvLPmzyHUl+ysFVGrmwuvcf2RtUTRpOF8UbXQPYDyoie/YqlNGG2jIlW f0kyEzGRW/zBDwX+bfJ68dvK59ABIhQ4xktCMaQ2I+yxgesDHT5+Uonlb00Sa2YUIkSUxwW5BZC HMfW+fyLgZlceVPvyE6yNQDEb+vla5ugSA9dOcNIJyKArgMfp7T3yT81LzvR9UwzPoqGhHvT+IR RduANSrxrKJXdecrA== X-Received: by 2002:a05:6871:b0f:b0:456:9b53:c6b5 with SMTP id 586e51a60fabf-4599ef8fbcfmr13854626fac.13.1786134078812; Fri, 07 Aug 2026 13:21:18 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:4::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1ece1fbsm2770188fac.17.2026.08.07.13.21.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:17 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 12/14] mm/memcontrol: Kick async reclaim on migration and folio replacement Date: Fri, 7 Aug 2026 13:20:55 -0700 Message-ID: <20260807202059.2620949-13-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Memcg tier charges are now transferred during folio migrations, but this still leaves forced charges and unenforced paths like folio replacement, hugetlb migration, and memory hotplug. Forced charges can and should not be enforced by synchronous reclaim, but if soft limits are set, we can kick async reclaimers to try and bring the usage below the high memcg tier limit. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 1161934e81380..4dce7c6fefd98 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5663,6 +5663,16 @@ void __mem_cgroup_uncharge_folios(struct folio_batch= *folios) uncharge_batch(&ug); } =20 +static void tier_kick_high(struct mem_cgroup *memcg, int slot) +{ + struct page_counter *tier_counter; + + tier_counter =3D mem_cgroup_tier_counter(memcg, slot); + if (tier_counter && page_counter_read(tier_counter) > + READ_ONCE(tier_counter->high)) + schedule_work(&memcg->high_work); +} + /** * mem_cgroup_replace_folio - Charge a folio's replacement. * @old: Currently circulating folio. @@ -5705,6 +5715,7 @@ void mem_cgroup_replace_folio(struct folio *old, stru= ct folio *new) int slot =3D nid_tier_slot(folio_nid(new)); =20 mem_cgroup_charge_tier(memcg, slot, nr_pages); + tier_kick_high(memcg, slot); } if (do_memsw_account()) page_counter_charge(&memcg->memsw, nr_pages); @@ -5798,6 +5809,7 @@ void mem_cgroup_migrate(struct folio *old, struct fol= io *new) if (old_slot !=3D new_slot) { mem_cgroup_uncharge_tier(memcg, old_slot, nr_pages); mem_cgroup_charge_tier(memcg, new_slot, nr_pages); + tier_kick_high(memcg, new_slot); } rcu_read_unlock(); } --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oo1-f50.google.com (mail-oo1-f50.google.com [209.85.161.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 50FBF45516D for ; Fri, 7 Aug 2026 20:21:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134083; cv=none; b=lsbJklmUoE92HYE1+ddmP07TMg4xjNGfSL+BEtxs6e5KiSQUNQNF25nlnwpYyqDhDiLF2+VOQII53gEs9PbYvut/6XTJPeEXh6Y7kqDcDUaPAjICJhjS9unzclKTQOFFk6s01RdK9uHqH3QvIs4Iq+RswBHmfoFIEcKge5swKfw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134083; c=relaxed/simple; bh=4YHmAs4imI5gKBDc4Wehbn5cIF5tjkqbuwyC0X3KRSI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iPp87rsZuIzA2A8aBPj55wZ2XYahxPedsFoG70XhXTZ5DctpMUy03Z0UemQazNwQZRxrlklhjOh4XeLavt/brwKLYQLyotAwdYiwufj2rbNbC9McVed9HmiiwGaP0cv6lOgCAq3la9EFPPGgmiOWksNk9nVEBARH6ejKtCCiah8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=aJaHmT5o; arc=none smtp.client-ip=209.85.161.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="aJaHmT5o" Received: by mail-oo1-f50.google.com with SMTP id 006d021491bc7-6ae5baaef5dso1914978eaf.1 for ; Fri, 07 Aug 2026 13:21:21 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134080; x=1786738880; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7/o7PyJxYr/auJmws9zryGpCICrELhYgOPRDbTKuSBc=; b=aJaHmT5oyPYDad1NKHaiW0BNFW2HugGkOZeSpMyYtJcyrz8nLE1vCoKTznIdZxMOvB IQQ09H8e404BOub7VB+I13fSbKSpz/AZ1APhY+eYPW5ElLF6PYoC3O1jaKcnA8D9ffs/ lpYGA6WXdRgwoBRgQxa/vYKbPUcPTOiXtmle4UssWrHguNwpJyeqPqFFWPnAMqpbgXP8 u+wabKckWHK1lj1MZ2LsULNxenxpZp7G6yJ41nke79Z9qOVITz7uPZOhSKHZIAB8eZ/X hpQSNPlGD5Xw3azQ+oueJfiAh6EEol4mWa5cG3RABNOXWqYwcJpWymBCnmhat76uPqGw UISQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134080; x=1786738880; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7/o7PyJxYr/auJmws9zryGpCICrELhYgOPRDbTKuSBc=; b=EcvTu3aLwhCnf6+hEwLriXtXkPybBMUErOW0plDmXAgvrDMf/WbHgpncQKJoQzWuQf wNAGL1VqLM4MMSxIvn3tKOptu7X5R96aSmjWrncmL/ZZMWGmVuYIz7nY1UGEBuanklPu lZ/cP8/mmrJ3QtaGL/dc+pEsOgupBPBbplkYWV95kMfhcrigmpdo32UCuZoMfAi9v+BH KqwxnmHHyysLwhvwud0Dk7dob19cpjniGQJwmdbbxgIZRclqgCjAC++2Ch1+NbT3XQUz PV0oghbu6sh0ZdVGsGgAS96/sf6LAseACi/aK8Dt9/2VBfLr3jyYsOxULtqeKadyeT+1 pE2Q== X-Forwarded-Encrypted: i=1; AHgh+Rpy31JNh+xo69UPM7XWD+Idn5o2iTalm7ZTyLXT497BnedZWUsvi7WiN+PXOHvtdMveJzfh0mOsk9RqebU=@vger.kernel.org X-Gm-Message-State: AOJu0Yz0idhzpgMOXOwPu36xdkBwNemR8fsrlZ7Z+oppLZSLoLAI88j2 PzTsoYmCimlfsDCmcn227C6EatZJrdKcYr6usw+EUfTW5m90YDBN9swg X-Gm-Gg: AR+sD106IPMW+1UsDL9MVKQ0qIkxNwJrhvA6FDK7jlFTid4NjgtgXhbNvQjmPtaiWpp RPUkMIRiRqhtxX0yl84/wGN8pmCErVNrd/J8M7ZhTZ8MrQMAgQSB99Oz+EJL/IRRxXdl2YvR/He O8VP5G7ufOstvZUO99dALxa7lufSaewpqmKIVjTueHzmVXB+GYbYcYsI+BsWnj7+oNbFAnthWBZ 1t1H7Dz65Mkd4SnR2gJE4PoT/GXq9nchBrj76jUjLiOGtuGGWdA4fehAKuyOzbjpazrXZK6WdTv +hzyvEo8Yo+TvGR3TybHmJ4vZ8s4bVvAjeHn/4Gz7X2L9Bj9a8nxlPnjehm7R/ZPMZhRRhvNn9m H1YwjxVqEYUrQPcYtUaM81m3buwDWwJJX5DVIzUpmaKlIxAhM3YXHQvVcKOYYojN9HpnC854334 hOpycYDvjOQAGY0BmVtSLViRG0PHAbC+0ZpMNL1xBNf5k3jy43QsG4AaNRwAP2aaDopKtYcnUkv MgqdqdsxEnaKjtGGJg= X-Received: by 2002:a05:6820:8c7:b0:6a3:6f5d:4d6c with SMTP id 006d021491bc7-6b044e7bbdfmr825692eaf.6.1786134080128; Fri, 07 Aug 2026 13:21:20 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:12::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-459f1a9d9b9sm2606796fac.7.2026.08.07.13.21.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:19 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 13/14] mm/memcontrol, sched/numa: Gate NUMA promotions into memcg tiers Date: Fri, 7 Aug 2026 13:20:56 -0700 Message-ID: <20260807202059.2620949-14-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Memory promotions that go through should_numa_migrate_memory determine if a promotion should be ratelimited / throttled by checking how much headroom there is in the destination node. If there is enough headroom, there is no reason to be throttling promotions. On tiered systems, however, a promotion may trigger reclaim on a node that has plenty of promotion headroom since the memcg tier may be at the limit. For these allocations, we should make sure that memcg tier fullness is also considered when determining whether a promotion should be able to go through without getting limited. Add an additional condition to check before letting a promotion candidate go through un-ratelimited, by checking if the memcg tier is already at its limit. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 7 +++++++ kernel/sched/fair.c | 3 ++- mm/memcontrol.c | 35 +++++++++++++++++++++++++++++++++++ 3 files changed, 44 insertions(+), 1 deletion(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 9c2f11191a499..a7c366b431a0e 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -654,6 +654,8 @@ static inline bool mem_cgroup_below_min(struct mem_cgro= up *target, page_counter_read(&memcg->memory); } =20 +bool mem_cgroup_tier_over_limit(struct folio *folio, int dst_nid); + int __mem_cgroup_charge(struct folio *folio, struct mm_struct *mm, gfp_t g= fp); =20 /** @@ -1172,6 +1174,11 @@ static inline bool mem_cgroup_below_min(struct mem_c= group *target, return false; } =20 +static inline bool mem_cgroup_tier_over_limit(struct folio *folio, int dst= _nid) +{ + return false; +} + static inline int mem_cgroup_charge(struct folio *folio, struct mm_struct *mm, gfp_t gfp) { diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index d78467ec6ee13..397b0f3e67f5f 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -2697,7 +2697,8 @@ bool should_numa_migrate_memory(struct task_struct *p= , struct folio *folio, long nr =3D folio_nr_pages(folio); =20 pgdat =3D NODE_DATA(dst_nid); - if (pgdat_free_space_enough(pgdat)) { + if (pgdat_free_space_enough(pgdat) && + !mem_cgroup_tier_over_limit(folio, dst_nid)) { /* workload changed, reset hot threshold */ pgdat->nbp_threshold =3D 0; mod_node_page_state(pgdat, PGPROMOTE_CANDIDATE_NRL, nr); diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 4dce7c6fefd98..05611a01aa082 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2590,6 +2590,41 @@ static u64 swap_find_max_overage(struct mem_cgroup *= memcg) return max_overage; } =20 +bool mem_cgroup_tier_over_limit(struct folio *folio, int dst_nid) +{ + struct mem_cgroup *memcg; + int dst_slot; + + if (!mem_cgroup_tiered_limits()) + return false; + + dst_slot =3D nid_tier_slot(dst_nid); + if (nid_tier_slot(folio_nid(folio)) =3D=3D dst_slot) + return false; + + guard(rcu)(); + memcg =3D folio_memcg(folio); + if (!memcg || mem_cgroup_is_root(memcg)) + return false; + + do { + struct page_counter *tier_counter; + unsigned long limit; + + tier_counter =3D mem_cgroup_tier_counter(memcg, dst_slot); + if (!tier_counter) + continue; + + limit =3D min(READ_ONCE(tier_counter->max), + READ_ONCE(tier_counter->high)); + if (page_counter_read(tier_counter) > limit) + return true; + } while ((memcg =3D parent_mem_cgroup(memcg)) && + !mem_cgroup_is_root(memcg)); + + return false; +} + /* * Get the number of jiffies that we should penalise a mischievous cgroup = which * is exceeding its memory.high by checking both it and its ancestors. --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:52 2026 Received: from mail-oo1-f43.google.com (mail-oo1-f43.google.com [209.85.161.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C0F0F456284 for ; Fri, 7 Aug 2026 20:21:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134084; cv=none; b=e86n4op7vnWAFjD0hRUQyMHffT+J8cNDsHYX0j0ZVSO5WOvJjY+jP7xAJh0egSiReO1szk7dF8bEBAgxrhIVnxXQuIhNywKfiSsI79nEhv0MVQDQbPxzSR6VgnwWH3Y21PTOPftXUmVsd/Iyl6Sw66vtUNtGBdI8QeKqG2UIxOo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786134084; c=relaxed/simple; bh=m1j3khm0CttzgXKBo/Rg7T/3ykhCkrLJzu6QmUaFGDY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YVgtSVCFYkg9XdE/5WVsUppsPEvbearfoIsz6H0qzV7IYs+alo/qyZuui9W6ClGEC9Ef8/9X51x8UlfUZp1hR/Zxt8gKZPC1IIjN8TkcV7Weqs0TcI/BPSNJ/k6156bMNtPGsm0auCXxMtL2hS6Ui7sZ02gesD2PXAwHDrLYeU4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jaALfnBw; arc=none smtp.client-ip=209.85.161.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jaALfnBw" Received: by mail-oo1-f43.google.com with SMTP id 006d021491bc7-6afcf32fb65so1025669eaf.2 for ; Fri, 07 Aug 2026 13:21:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786134081; x=1786738881; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=gNurNSPYyJZesKOZSqzZdKIcfvAei88zp/AV04IWMMU=; b=jaALfnBw5s6tdC8C0YBmMa/B4zG27HBNqxv0gNLThRJZrMFL2qrNJ1z1rpRyfTmtbb N55DHFRC9E7L02IskqpOv589uUkfDrWq4ctWIVZp+i3pYmTJ1+ZpxDAb/yHdveLku9Yb H1xktTDLEricz9APSFVSMBumSasz/1LRpgE9x+73kR61FWum2U42QbOsLXny++QeT1zg +1X1eR+wrDkv+gPP5suuPNujALcEGvjnhG/Zz4NWgASmDp3Tx4FRhK9tldj1vyMlcPRq rxugOymphuuiU4Ro5+ZDMDj2c4ussDFp4K9sPBWBYyNb9FMOfVMVkIZFt9G67CPvEZAK HE9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786134081; x=1786738881; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=gNurNSPYyJZesKOZSqzZdKIcfvAei88zp/AV04IWMMU=; b=eRTyydDrL/stAFsy5ii11hHr7eM/6KowZZPbGAfabpE72b1oEHlv/FnboFU+8MuSuO E1+uFb3g6KxoP5HE2tCOJZXROTT3TjUuxDKsL+BNj4YHx/TGNkOQLgIX+7cHi1i6afvi b89qb0XowkuNtPrE6X78pPxB/DW+GkVBXSNfQfixEc7OBuC49N4b0KarSFYOt1HI3wcW d7qQxH+dx7EywvRBjdzzOaO1x9GanTSyEFQUAdpHPtCKG1CPhrVlXEaBQfnT5VJc4z5v /EvpRjje6M/ac1w5K3eA93Km/UkN5lhsI+aFMn62shvfTmWRBaJFkXq3C6nI8ouysSbA LsIA== X-Forwarded-Encrypted: i=1; AHgh+RpOUAXa8TOJK/pwAYW7qkJhmsUhVgokF7XTEqkpxQrN2rOx9lKzgpg/4+LaTk8ewOs8jaCPM4uAWJantqY=@vger.kernel.org X-Gm-Message-State: AOJu0YzxzprDwBL0xipK69Ycbhh2uhPthpe1mDsbOfMVXrUgEBykWI+s P8tBAuZ9SMvkPsz57WDq1VaVFy1LDIt0jYwmrqtEAKdF+ZEsl2CUpSmt X-Gm-Gg: AR+sD12AS0QVezPpgPpvBQg4fmGBGH4aWJwnOi8cpFhwVqQHDLur6SmF1oFrm31dZM7 SSBJru6Y4CZnegwkm0gmEvgRnaVUC1bOLoGi6xLoQM3FpwF+al0U+7nY42gH9/qeY7D87tL4FqT zVIxTafOF7dHi+ka0FEWFAyMS2yVKV+8D5SwnXq7w5KlVAE5y8i8qW23afU+PE3X6X5cXPpmxjr 4FkulijuLDFLChahsasUaZOG8d5Fl5kj3jUiVd+eONGd6CDw02x2nCgFTQQjSayKKg085hxwcaR mxJDZpE5WtFsr+i2wy+/w8UX/GztxMQljrUqyR6xrEHyFCxYDG2GpDdeEUM28TuM1PMltXWvKqk lbN1/24Ribt26BHL0e47R2L+Y0I3q2NFprAbtMaX72DQm+2QJhZ8xK1+OwZ4HZAELN6XMEbica+ Doesv9AF5CqyjreNKLIRa1SIjp4g2YQA8BK/FUq8xr9uXJXGiDJvtaU6jTCVCUs8+PNSgHBChAV 2/eYGIEZErWaDeiVZQ= X-Received: by 2002:a05:6820:150f:b0:6ab:3c7:da56 with SMTP id 006d021491bc7-6ae96f6cfdbmr12464583eaf.24.1786134081584; Fri, 07 Aug 2026 13:21:21 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:1c::]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6b02be475b6sm3130491eaf.11.2026.08.07.13.21.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:21:21 -0700 (PDT) From: Joshua Hahn To: Johannes Weiner , Gregory Price Cc: Alistair Popple , Andrew Morton , Axel Rasmussen , Barry Song , Ben Segall , Brendan Jackman , Byungchul Park , David Hildenbrand , David Rientjes , Dietmar Eggemann , "Harry Yoo (Oracle)" , Ingo Molnar , Juri Lelli , K Prateek Nayak , Kairui Song , "Liam R. Howlett" , Lorenzo Stoakes , Matthew Brost , Mel Gorman , Michal Hocko , Michal Hocko , Mike Rapoport , Muchun Song , Peter Zijlstra , Qi Zheng , Rakie Kim , Roman Gushchin , Shakeel Butt , Steven Rostedt , Suren Baghdasaryan , "T.J. Mercier" , Valentin Schneider , Vincent Guittot , Vlastimil Babka , Wei Xu , Ying Huang , Yosry Ahmed , Yuanchu Xie , Zi Yan , cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: [RFC PATCH v3 14/14] mm/page_alloc: steer allocations away from exhausted memory tiers Date: Fri, 7 Aug 2026 13:20:57 -0700 Message-ID: <20260807202059.2620949-15-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> References: <20260807202059.2620949-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A memcg only finds out that it is over a tier's limit at try_charge_memcg time, after the page has already been allocated on the full tier's node. This triggers reclaim to push memory to lower tiers or to swap, much like how zone_reclaim_mode triggers reclaim when a node is full. This causes a lot of unnecessary churn. Instead of allocating a page on a full tier only to immediately reclaim, make the page allocator aware of tier limits at allocation time and steer the first allocation attempt to nodes belonging to tiers with tiered limit headroom. This only affects the ac->nodemask for the fastpath, meaning if no node can satisfy this allocation, it falls back to the caller's nodemask and places a page on a full tier, triggering reclaim. Moreover, if it turns out that the intersection of the allocation context nodemask and the under-limit nodemask is empty, skip the steering. This steering only affects MIGRATE_MOVABLE allocations. We try not to steer kernel allocations using NUMA_NO_NODE which would prefer to remain on higher tiers, even if they are full. This steering also does not affect alloc_pages_bulk_noprof, since its callers do not charge their memory allocations to a tier. No-op unless the system has tiered memcg limits enabled. Signed-off-by: Joshua Hahn --- include/linux/memcontrol.h | 6 +++++ mm/memcontrol.c | 46 ++++++++++++++++++++++++++++++++++++++ mm/page_alloc.c | 19 +++++++++++++--- 3 files changed, 68 insertions(+), 3 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index a7c366b431a0e..b700f3224cbaa 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -655,6 +655,7 @@ static inline bool mem_cgroup_below_min(struct mem_cgro= up *target, } =20 bool mem_cgroup_tier_over_limit(struct folio *folio, int dst_nid); +bool mem_cgroup_tier_allowed_nodemask(nodemask_t *mask); =20 int __mem_cgroup_charge(struct folio *folio, struct mm_struct *mm, gfp_t g= fp); =20 @@ -1179,6 +1180,11 @@ static inline bool mem_cgroup_tier_over_limit(struct= folio *folio, int dst_nid) return false; } =20 +static inline bool mem_cgroup_tier_allowed_nodemask(nodemask_t *mask) +{ + return false; +} + static inline int mem_cgroup_charge(struct folio *folio, struct mm_struct *mm, gfp_t gfp) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 05611a01aa082..e499e58baa287 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2625,6 +2625,52 @@ bool mem_cgroup_tier_over_limit(struct folio *folio,= int dst_nid) return false; } =20 +/* + * Returns whether a mem_cgroup is above a tier's limit. + * The nodemask becomes populated with nodes that are under their limits. + */ +bool mem_cgroup_tier_allowed_nodemask(nodemask_t *mask) +{ + struct mem_cgroup *memcg; + int nr_tier_slots; + bool restricted =3D false; + + if (!mem_cgroup_tiered_limits()) + return false; + + nr_tier_slots =3D mt_nr_tier_slots(); + *mask =3D node_states[N_MEMORY]; + + rcu_read_lock(); + memcg =3D active_memcg(); + if (!memcg && in_task() && current->mm) + memcg =3D mem_cgroup_from_task(rcu_dereference(current->mm->owner)); + + for (; memcg && !mem_cgroup_is_root(memcg); + memcg =3D parent_mem_cgroup(memcg)) { + if (READ_ONCE(memcg->memory.max) =3D=3D PAGE_COUNTER_MAX && + READ_ONCE(memcg->memory.high) =3D=3D PAGE_COUNTER_MAX) + continue; + + for (int slot =3D 0; slot < nr_tier_slots; slot++) { + struct page_counter *tier =3D &memcg->tier[slot]; + unsigned long limit; + + limit =3D min(READ_ONCE(tier->high), + READ_ONCE(tier->max)); + + if (page_counter_read(tier) <=3D limit) + continue; + + restricted =3D true; + nodes_andnot(*mask, *mask, *mt_tier_nodes(slot)); + } + } + rcu_read_unlock(); + + return restricted; +} + /* * Get the number of jiffies that we should penalise a mischievous cgroup = which * is exceeding its memory.high by checking both it and its ancestors. diff --git a/mm/page_alloc.c b/mm/page_alloc.c index 0aeca106a4fde..7ca0acc37d2af 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -5153,7 +5153,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int o= rder, static inline bool prepare_alloc_pages(gfp_t gfp_mask, unsigned int order, int preferred_nid, nodemask_t *nodemask, struct alloc_context *ac, gfp_t *alloc_gfp, - unsigned int *alloc_flags) + unsigned int *alloc_flags, nodemask_t *tier_nodes) { ac->highest_zoneidx =3D gfp_zone(gfp_mask); ac->zonelist =3D node_zonelist(preferred_nid, gfp_mask); @@ -5187,6 +5187,17 @@ static inline bool prepare_alloc_pages(gfp_t gfp_mas= k, unsigned int order, /* Dirty zone balancing only done in the fast path */ ac->spread_dirty_pages =3D (gfp_mask & __GFP_WRITE); =20 + if (mem_cgroup_tiered_limits() && tier_nodes && + !(gfp_mask & __GFP_THISNODE) && + ac->migratetype =3D=3D MIGRATE_MOVABLE && + mem_cgroup_tier_allowed_nodemask(tier_nodes)) { + if (ac->nodemask) + nodes_and(*tier_nodes, *tier_nodes, *ac->nodemask); + /* tier_nodes must outlive the function call since ac uses it */ + if (!nodes_empty(*tier_nodes)) + ac->nodemask =3D tier_nodes; + } + /* * The preferred zone is used for statistics but crucially it is * also used as the starting point for the zonelist iterator. It @@ -5269,7 +5280,8 @@ unsigned long alloc_pages_bulk_noprof(gfp_t gfp, int = preferred_nid, =20 /* May set ALLOC_NOFRAGMENT, fragmentation will return 1 page. */ gfp &=3D gfp_allowed_mask; - if (!prepare_alloc_pages(gfp, 0, preferred_nid, nodemask, &ac, &gfp, &all= oc_flags)) + if (!prepare_alloc_pages(gfp, 0, preferred_nid, nodemask, &ac, &gfp, + &alloc_flags, NULL)) goto out; =20 /* Find an allowed local zone that meets the low watermark. */ @@ -5451,6 +5463,7 @@ struct page *__alloc_frozen_pages_noprof(gfp_t gfp, u= nsigned int order, .alloc_flags =3D alloc_flags, }; unsigned int fastpath_alloc_flags =3D alloc_flags; + nodemask_t tier_nodes; =20 /* Other flags could be supported later if needed. */ if (WARN_ON(alloc_flags & ~(ALLOC_NOLOCK | ALLOC_NO_CODETAG))) @@ -5481,7 +5494,7 @@ struct page *__alloc_frozen_pages_noprof(gfp_t gfp, u= nsigned int order, gfp =3D current_gfp_context(gfp); alloc_gfp =3D gfp; if (!prepare_alloc_pages(gfp, order, preferred_nid, nodemask, &ac, - &alloc_gfp, &fastpath_alloc_flags)) + &alloc_gfp, &fastpath_alloc_flags, &tier_nodes)) return NULL; =20 if (!(alloc_flags & ALLOC_NOLOCK)) { --=20 2.53.0-Meta