From nobody Sun Apr 5 21:35:14 2026 Received: from mail-ot1-f46.google.com (mail-ot1-f46.google.com [209.85.210.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9BE0D37B41A for ; Mon, 23 Feb 2026 22:38:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1771886326; cv=none; b=lOHWMhcQB+CUKSzl6Gb0feLLG961cu4VnbanI2YptlU9kCY2BmrFavrK5YVM6q2+Hrwki5RqArbmW6s1ZWh8QzXc/3xyb4yUWtOJSJz6Ri3kvmCtY4G8am+SipCw00KE3N8kw/eUQzTNdEp5l45pms3gvPD7LgffEIHJPLgfHLo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1771886326; c=relaxed/simple; bh=sBvHCRLOSKHT83gdccnLmheOXBIOB5m3ZeZY9+qfNQA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B6XRR8v+w1fJzFzFxLxa4sTHc6UzEceFnXGBAti3Oa3G0U/TgAh+A8GPVoDRoxAeFBVUVvewrPstHkW6sj7j8YwVnuRRkI6NJvmZJH25HXiPRuynjculIY4cpi8Atgg/ydz8wU1ERd9Pk2qvrAurSw5E093wKKgesG755uVNkH4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=MJwHVpTS; arc=none smtp.client-ip=209.85.210.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="MJwHVpTS" Received: by mail-ot1-f46.google.com with SMTP id 46e09a7af769-7d4c68f0e47so2859270a34.1 for ; Mon, 23 Feb 2026 14:38:44 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1771886323; x=1772491123; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=1/pi4GMyqfUzxedG8y+xzMj5Pky+ci5h56Uzs52BODo=; b=MJwHVpTSrjWbYibbWrJ0R73PEr0g9TRcUEdz7li+J4GMN5D6xVnlFthNRBVtg7SuQ0 mwP5xmsaVwTlAJozaLGVhPo+uC+nlB3tlRrMi0YFTQkP0ErFhKwAQj59WYxJz03QsS2V bPU300jPsoJdEhMBklIEPs6xcK7YA4fr4seeRv45Bm5jtK4KA2TnamjTg440ItepjBPI j1/usqUOeMorV8qDDDTZ0HuvkZAj2rlHUQEkcHkCh+0BUfEoSSHvOFITQPaILHSoAREn PYAnIl2yViMCvKmDVPJuXN3DJM9DcC3Cdt0GLCxyYsvzvxz+xBNpDW73poU/oergbY+e RyuA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1771886323; x=1772491123; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=1/pi4GMyqfUzxedG8y+xzMj5Pky+ci5h56Uzs52BODo=; b=G3PStY3YUi+k6PJ+NcEMmmbf5h8O2mKqI2Y0tNfR0eNFydBZXsGebBRR5BFzsd7TAW h9BYpVeQ8uZh1LRSqg7YpCWsVsVLPhtw4CqSRvx1ZdYtuMnKLeatzsitTqJ+Iyu0IwZP 1RHE2yHUjRp4ZoIsleUUiWfOvZeNKYlSbWF3+hPOWjLT6e02gI/1Lfcp8TVY6rBffeC/ kZMqbTipXSujaBT/e/Fq5PmLcvOJo8V6zw14S3CB4G60aEbWqItZS3nA6QJ0IMfzKv8E xmSSovY5t7XuGnoTfBovXM+V9o3tb5QRmNiNdi7hW/bJhGfVAet52ZdPjgnCkbUWypdx l1PQ== X-Forwarded-Encrypted: i=1; AJvYcCV5U50Ts97hBYWx+GMUCjGuxgEuF9lsU0gGTaOtd5nQxSEHheafMRx0VfBujyMpYKNT+UN+I6MLo0X1vMk=@vger.kernel.org X-Gm-Message-State: AOJu0YykgrRXnFTYz9s+pQMmbm/I1qAQ2SQ53/MjN9hJK5CErmkPHDPO V7I9lnJULPX5uo5PrgloF0QAgpHtaZlSMfZWX8nKX0PX97jMDPfF5uSB X-Gm-Gg: AZuq6aJyUvdA5vnG0bU29Hv/p9WllhCFbTuGx3QeP1WOzNWY1VH3GqLy+8a+CHn1I/C QcZz3zyxwTGVF01QhuaUBDfYc4+Nt7s0hkLCjFczIDoQIuBPa7tlIthjL+Z7X63LHJRchvVvADE iB6WtlfE6WL13GqBuoXIu0mZYKgABnnT6IDcVvFPQFsl5HMosx+i6ViMDrLbx9T6RJQEW0pLs5d p7uS00FkfKlvotblfWKSyYCpFEg4PQCIur3M4Fb2a15QO1N4NmHZ5ic5KbhwQgMRMiOnGOnMhN0 fXW+/4YM3Li2MihOShadlLw+SeFCgLgretDBi78oCfnRuCza99xFBGQVzN5ND4FmE3IDk8qTZTu ncqXEzu9hIQ8d8wQgbXGhwOOWVUQ8tqjDiiGD3x3AfCTakHrGim+M58TdiAsbLlI+FDLFCZYwyW jj1doiJSf6jUTfXEFlTZeUJb4jeUr6PTb6 X-Received: by 2002:a05:6830:6610:b0:7d1:9da9:c6e with SMTP id 46e09a7af769-7d52bf6bb20mr5168975a34.25.1771886323539; Mon, 23 Feb 2026 14:38:43 -0800 (PST) Received: from localhost ([2a03:2880:10ff:71::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7d52ce63069sm8952812a34.0.2026.02.23.14.38.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 23 Feb 2026 14:38:42 -0800 (PST) From: Joshua Hahn To: Joshua Hahn Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Qi Zheng , Axel Rasmussen , Yuanchu Xie , Wei Xu , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com Subject: [RFC PATCH 6/6] mm/memcontrol: Make memory.high tier-aware Date: Mon, 23 Feb 2026 14:38:29 -0800 Message-ID: <20260223223830.586018-7-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260223223830.586018-1-joshua.hahnjy@gmail.com> References: <20260223223830.586018-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On machines serving multiple workloads whose memory is isolated via the memory cgroup controller, it is currently impossible to enforce a fair distribution of toptier memory among the workloads, as the only enforcable limits have to do with total memory footprint, but not where that memory resides. This makes ensuring a consistent and baseline performance difficult, as each workload's performance is heavily impacted by workload-external factors wuch as which other workloads are co-located in the same host, and the order at which different workloads are started. Extend the existing memory.high protection to be tier-aware in the charging and enforcement to limit toptier-hogging for workloads. Also, add a new nodemask parameter to try_to_free_mem_cgroup_pages, which can be used to selectively reclaim from memory at the memcg-tier interection of a cgroup. Signed-off-by: Joshua Hahn --- include/linux/swap.h | 3 +- mm/memcontrol-v1.c | 6 ++-- mm/memcontrol.c | 85 +++++++++++++++++++++++++++++++++++++------- mm/vmscan.c | 11 +++--- 4 files changed, 84 insertions(+), 21 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 0effe3cc50f5..c6037ac7bf6e 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -368,7 +368,8 @@ extern unsigned long try_to_free_mem_cgroup_pages(struc= t mem_cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness); + int *swappiness, + nodemask_t *allowed); extern unsigned long mem_cgroup_shrink_node(struct mem_cgroup *mem, gfp_t gfp_mask, bool noswap, pg_data_t *pgdat, diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 0b39ba608109..29630c7f3567 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -1497,7 +1497,8 @@ static int mem_cgroup_resize_max(struct mem_cgroup *m= emcg, } =20 if (!try_to_free_mem_cgroup_pages(memcg, 1, GFP_KERNEL, - memsw ? 0 : MEMCG_RECLAIM_MAY_SWAP, NULL)) { + memsw ? 0 : MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL)) { ret =3D -EBUSY; break; } @@ -1529,7 +1530,8 @@ static int mem_cgroup_force_empty(struct mem_cgroup *= memcg) return -EINTR; =20 if (!try_to_free_mem_cgroup_pages(memcg, 1, GFP_KERNEL, - MEMCG_RECLAIM_MAY_SWAP, NULL)) + MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL)) nr_retries--; } =20 diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 8aa7ae361a73..ebd4a1b73c51 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2184,18 +2184,30 @@ static unsigned long reclaim_high(struct mem_cgroup= *memcg, =20 do { unsigned long pflags; - - if (page_counter_read(&memcg->memory) <=3D - READ_ONCE(memcg->memory.high)) + nodemask_t toptier_nodes, *reclaim_nodes; + bool mem_high_ok, toptier_high_ok; + + mt_get_toptier_nodemask(&toptier_nodes, NULL); + mem_high_ok =3D page_counter_read(&memcg->memory) <=3D + READ_ONCE(memcg->memory.high); + toptier_high_ok =3D !(tier_aware_memcg_limits && + mem_cgroup_toptier_usage(memcg) > + page_counter_toptier_high(&memcg->memory)); + if (mem_high_ok && toptier_high_ok) continue; =20 + if (mem_high_ok && !toptier_high_ok) + reclaim_nodes =3D &toptier_nodes; + else + reclaim_nodes =3D NULL; + memcg_memory_event(memcg, MEMCG_HIGH); =20 psi_memstall_enter(&pflags); nr_reclaimed +=3D try_to_free_mem_cgroup_pages(memcg, nr_pages, gfp_mask, MEMCG_RECLAIM_MAY_SWAP, - NULL); + NULL, reclaim_nodes); psi_memstall_leave(&pflags); } while ((memcg =3D parent_mem_cgroup(memcg)) && !mem_cgroup_is_root(memcg)); @@ -2296,6 +2308,24 @@ static u64 mem_find_max_overage(struct mem_cgroup *m= emcg) return max_overage; } =20 +static u64 toptier_find_max_overage(struct mem_cgroup *memcg) +{ + u64 overage, max_overage =3D 0; + + if (!tier_aware_memcg_limits) + return 0; + + do { + unsigned long usage =3D mem_cgroup_toptier_usage(memcg); + unsigned long high =3D page_counter_toptier_high(&memcg->memory); + + overage =3D calculate_overage(usage, high); + max_overage =3D max(overage, max_overage); + } while ((memcg =3D parent_mem_cgroup(memcg)) && + !mem_cgroup_is_root(memcg)); + + return max_overage; +} static u64 swap_find_max_overage(struct mem_cgroup *memcg) { u64 overage, max_overage =3D 0; @@ -2401,6 +2431,14 @@ void __mem_cgroup_handle_over_high(gfp_t gfp_mask) penalty_jiffies +=3D calculate_high_delay(memcg, nr_pages, swap_find_max_overage(memcg)); =20 + /* + * Don't double-penalize for toptier high overage if system-wide + * memory.high has already been breached. + */ + if (!penalty_jiffies) + penalty_jiffies +=3D calculate_high_delay(memcg, nr_pages, + toptier_find_max_overage(memcg)); + /* * Clamp the max delay per usermode return so as to still keep the * application moving forwards and also permit diagnostics, albeit @@ -2503,7 +2541,8 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, =20 psi_memstall_enter(&pflags); nr_reclaimed =3D try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages, - gfp_mask, reclaim_options, NULL); + gfp_mask, reclaim_options, + NULL, NULL); psi_memstall_leave(&pflags); =20 if (mem_cgroup_margin(mem_over_limit) >=3D nr_pages) @@ -2592,23 +2631,26 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, * reclaim, the cost of mismatch is negligible. */ do { - bool mem_high, swap_high; + bool mem_high, swap_high, toptier_high =3D false; =20 mem_high =3D page_counter_read(&memcg->memory) > READ_ONCE(memcg->memory.high); swap_high =3D page_counter_read(&memcg->swap) > READ_ONCE(memcg->swap.high); + toptier_high =3D tier_aware_memcg_limits && + (mem_cgroup_toptier_usage(memcg) > + page_counter_toptier_high(&memcg->memory)); =20 /* Don't bother a random interrupted task */ if (!in_task()) { - if (mem_high) { + if (mem_high || toptier_high) { schedule_work(&memcg->high_work); break; } continue; } =20 - if (mem_high || swap_high) { + if (mem_high || swap_high || toptier_high) { /* * The allocating tasks in this cgroup will need to do * reclaim or be throttled to prevent further growth @@ -4476,7 +4518,7 @@ static ssize_t memory_high_write(struct kernfs_open_f= ile *of, struct mem_cgroup *memcg =3D mem_cgroup_from_css(of_css(of)); unsigned int nr_retries =3D MAX_RECLAIM_RETRIES; bool drained =3D false; - unsigned long high; + unsigned long high, toptier_high; int err; =20 buf =3D strstrip(buf); @@ -4485,15 +4527,22 @@ static ssize_t memory_high_write(struct kernfs_open= _file *of, return err; =20 page_counter_set_high(&memcg->memory, high); + toptier_high =3D page_counter_toptier_high(&memcg->memory); =20 if (of->file->f_flags & O_NONBLOCK) goto out; =20 for (;;) { unsigned long nr_pages =3D page_counter_read(&memcg->memory); + unsigned long toptier_pages =3D mem_cgroup_toptier_usage(memcg); unsigned long reclaimed; + unsigned long to_free; + nodemask_t toptier_nodes, *reclaim_nodes; + bool mem_high_ok =3D nr_pages <=3D high; + bool toptier_high_ok =3D !(tier_aware_memcg_limits && + toptier_pages > toptier_high); =20 - if (nr_pages <=3D high) + if (mem_high_ok && toptier_high_ok) break; =20 if (signal_pending(current)) @@ -4505,8 +4554,17 @@ static ssize_t memory_high_write(struct kernfs_open_= file *of, continue; } =20 - reclaimed =3D try_to_free_mem_cgroup_pages(memcg, nr_pages - high, - GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, NULL); + mt_get_toptier_nodemask(&toptier_nodes, NULL); + if (mem_high_ok && !toptier_high_ok) { + reclaim_nodes =3D &toptier_nodes; + to_free =3D toptier_pages - toptier_high; + } else { + reclaim_nodes =3D NULL; + to_free =3D nr_pages - high; + } + reclaimed =3D try_to_free_mem_cgroup_pages(memcg, to_free, + GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, + NULL, reclaim_nodes); =20 if (!reclaimed && !nr_retries--) break; @@ -4558,7 +4616,8 @@ static ssize_t memory_max_write(struct kernfs_open_fi= le *of, =20 if (nr_reclaims) { if (!try_to_free_mem_cgroup_pages(memcg, nr_pages - max, - GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, NULL)) + GFP_KERNEL, MEMCG_RECLAIM_MAY_SWAP, + NULL, NULL)) nr_reclaims--; continue; } diff --git a/mm/vmscan.c b/mm/vmscan.c index 5b4cb030a477..94498734b4f5 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -6652,7 +6652,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness) + int *swappiness, nodemask_t *allowed) { unsigned long nr_reclaimed; unsigned int noreclaim_flag; @@ -6668,6 +6668,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, .may_unmap =3D 1, .may_swap =3D !!(reclaim_options & MEMCG_RECLAIM_MAY_SWAP), .proactive =3D !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE), + .nodemask =3D allowed, }; /* * Traverse the ZONELIST_FALLBACK zonelist of the current node to put @@ -6693,7 +6694,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem= _cgroup *memcg, unsigned long nr_pages, gfp_t gfp_mask, unsigned int reclaim_options, - int *swappiness) + int *swappiness, nodemask_t *allowed) { return 0; } @@ -7806,9 +7807,9 @@ int user_proactive_reclaim(char *buf, reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP | MEMCG_RECLAIM_PROACTIVE; reclaimed =3D try_to_free_mem_cgroup_pages(memcg, - batch_size, gfp_mask, - reclaim_options, - swappiness =3D=3D -1 ? NULL : &swappiness); + batch_size, gfp_mask, reclaim_options, + swappiness =3D=3D -1 ? NULL : &swappiness, + NULL); } else { struct scan_control sc =3D { .gfp_mask =3D current_gfp_context(gfp_mask), --=20 2.47.3