From nobody Sat Sep 26 14:39:01 2026 Received: from mail-oa1-f52.google.com (mail-oa1-f52.google.com [209.85.160.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3D43A4A2A67 for ; Mon, 31 Aug 2026 16:37:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194277; cv=none; b=CUgkbNbUxvWEeaIKAhiStFcuZeroU2WlhuP0vDJqBhNZvFTdd/uOsmCRv3gdUDa0KO0LV+4Oq/WNOzMoiWMOy1GS+p/GT3Vyl8Hgp3CJN8A+gj2dPaK4nNkpJm8H9e9m4R1GLHK+psThElljjx6jlVQk27O80S8b9d7zgkVh84k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194277; c=relaxed/simple; bh=GYbb3qr5zDuvMPL4+OMnJHI4rslRuEhrqPm0smt3Vuk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eCkRfm4xq5Yyx8V6m9MWaAG88WAs2gISmKj5LLSmLJBQVK/F5BXZctQWO8xzwm/IolUfJmXiEkVHhlrt8tmjKxZ6stfQYztvaLFj36nC4dnXnKnhKSRZd5LfIUJZ7+ypxCaoon3NooUZU0Do8s/toEx4Zttgt0jHjPKCc1X4cSg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BY1bOPRO; arc=none smtp.client-ip=209.85.160.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BY1bOPRO" Received: by mail-oa1-f52.google.com with SMTP id 586e51a60fabf-46ac24346b1so683089fac.2 for ; Mon, 31 Aug 2026 09:37:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194275; x=1788799075; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=q9afrDjj3OiNFQNFlcWPaAJKUnSQO06bnnIAXSq7zMg=; b=BY1bOPRO72g7bDo/cyTQmRavUOzPDHdR3Om4uMaxzchkjXXfGZSh303BWPAHG96Szk 7s2jisFglfTJY7EKhcs3qlTjtzjS+ImJ54HEod4HHG7oPCofSpfwfUyq+8Pmg0gT9Lf6 3xd50Ivu+m9qbpJLvXGLoxXziYvqPRvJ0V1dYABL0C74Xr2w93dxj7Qz/yYCjSrvi4eJ SGQUh7rnaYZX43a+VIRKpJV2HX/H7QMRAWKqFrsnrUoZmpglSKwzRFDJTepnZGa4Fkkk DE1twcwhIjF8t3nU+HaOeIX1+LzYpAYxI4dLqvx1r//YVDLL79tWS5eaDXZyO9z8b2PB hhGg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194275; x=1788799075; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=q9afrDjj3OiNFQNFlcWPaAJKUnSQO06bnnIAXSq7zMg=; b=auT2l9PdZCRYdWLkKINUSOFP+Gl8730Ase8o9br/fl9oMONAV3f+YU9WGtHUfGem8n pfloGhiuoFEuoPkoVfRIzQIvXj6gmBwXad8XIE4DlT98PUW+PwaxIjFjVBR3EVHxqNqS 4pH/qqQaDpSeAtBmTgv2hxureVkhvyWV6ouKK5nwLS+gjlEL/aZkVzQJKaQyCmtPNL7W cWezCHmO7ggdrI/DVmAdjwp7609UHw/SdanEbVLWrATjqjb8pXsVkRQzVZTYwzgHEWy2 vOEVWtZ528ExmVNm8VT6kklZJ7nj2i/YDFj3hFfOEMGljpj0Wie/RYEBQLT39E4xS42h LpYg== X-Forwarded-Encrypted: i=1; AHgh+Rry2In/LqBwrBKBY4XfSTnegK8r6p9u95T+X+8C4UkvVDvD7cKdepGEOi2TmnvUkYhvLlTjE9BClXwVioU=@vger.kernel.org X-Gm-Message-State: AFuF++n9I/VEu3BkNjyStRrehRgeDLk6ui9D6Y5d2IVCUmFzJSe57Au1 MWOwdhKXpXLNYLuvocuo9doFx1mArCLv6B4BIgLukGZV70a0brxCGJ25 X-Gm-Gg: AR+sD1123m7PGi0JGi60cnAEDGF/+6pPq8QkHyjSqSwdeQDjWNEjUHHT4XVTvNsS/Ad kcYAtQrp2HmSosmNArMHEKi2u8Qs4HxcgBlfTpFSO4hwxrT8XG5EYNGYosa4/thDnTW4QtFKOy7 ZnycyjRGWKrHhVngyF7J/Qif1jLqT4f3Z8ctpOWW3AisSwiCEbRbYXg5ALSd7k23bav7+HEEf3E HptEp/K7Zy5JMmP6trWlfqwpMN8e5gDn4nt9xS1OXDPqJ8ITIDhwpfrK7Hy3c1XbCQXck3q+JiC AiFITKG+RQNErRvOhk//oQLMKStKoW8LPs7AO0FNN94u4AhWP+P9j3OW4pfodEdF4fnzCJkIaHi g5nRyKebRumBaEWmQCkEdqaPCX5iO0p5RFHtmVbr3x2X4gn8wFmmbamTSs85mdPrUHhNjvwu0RT fOOMsAunHfhkrhlZ77wJvq5sdqwJUzc1dPPhY4UXQOU1x+Mr91PQtuWh4EaieSLgTcjxRd5kzgT wI3SdaRqZSeuf6GNA== X-Received: by 2002:a05:6808:1921:b0:4b1:b83a:5878 with SMTP id 5614622812f47-4b5b81e2f8emr2412176b6e.14.1788194274813; Mon, 31 Aug 2026 09:37:54 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:8::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-4b3a1995f2fsm8537880b6e.11.2026.08.31.09.37.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:37:54 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 1/7] mm/memcontrol: flatten try_charge_memcg control flow Date: Mon, 31 Aug 2026 09:37:45 -0700 Message-ID: <20260831163752.2193337-2-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Refactor try_charge_memcg by flattening the nested memsw/memory page_counter operations to separate the logic between the two. When page_counter_try_charge is made stock-aware, this flattening makes the control flow easier to follow since each page counter now has its own success/failure paths. No functional changes intended. Signed-off-by: Joshua Hahn Acked-by: Shakeel Butt --- mm/memcontrol.c | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index bf829638524b5..93c2fa04da4fd 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2679,18 +2679,21 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, batch =3D nr_pages; =20 reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP; - if (!do_memsw_account() || - page_counter_try_charge(&memcg->memsw, batch, &counter)) { - if (page_counter_try_charge(&memcg->memory, batch, &counter)) - goto done_restock; - if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, batch); - mem_over_limit =3D mem_cgroup_from_counter(counter, memory); - } else { + if (do_memsw_account() && + !page_counter_try_charge(&memcg->memsw, batch, &counter)) { mem_over_limit =3D mem_cgroup_from_counter(counter, memsw); reclaim_options &=3D ~MEMCG_RECLAIM_MAY_SWAP; + goto reclaim; } =20 + if (page_counter_try_charge(&memcg->memory, batch, &counter)) + goto done_restock; + + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, batch); + mem_over_limit =3D mem_cgroup_from_counter(counter, memory); + +reclaim: if (batch > nr_pages) { batch =3D nr_pages; goto retry; --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-oa1-f47.google.com (mail-oa1-f47.google.com [209.85.160.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 03F2E4A2A78 for ; Mon, 31 Aug 2026 16:37:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194280; cv=none; b=liG/7TlKZ1wpbAjvVx/jCyyhE560FifIg4AEzlN0qPrLzasKLmjbjHwi7LdVmg7PIU1l6yRuSmwGF1M40S2vR3ekmWWbxoiGiFZ/wYWU3uGGwT0gvUrJedIsisyumVExVyJwJ2JyuT1wArEbyVUSSfihLMrGO17GYy8o3PWnrhU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194280; c=relaxed/simple; bh=PKwARw2KbFSmGIOyQgOWgTdCsHCypuENVQPYBlHabig=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NLKXoTJ1LJXmfPCOjawfeYeEVWBA0BLMxnLACVtXvYHfRkYpVR/8DFsveBGG/sG4dHtk++D2Q1W6ouOfWYPWkaOZdaFIcDWI9dGwmX7fSecYhIsSl0kmkrT6VPUKvDoD4Jq+oGxOci5Lr8v/pZ6rg5evSB+H4JSo6n5lKVOERdo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NhMMScEC; arc=none smtp.client-ip=209.85.160.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NhMMScEC" Received: by mail-oa1-f47.google.com with SMTP id 586e51a60fabf-44856d185bcso3328643fac.3 for ; Mon, 31 Aug 2026 09:37:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194277; x=1788799077; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=x77j6vdZ8LqsWXipiJe0J8DOL7X4B8I6QMSB+YfLeSc=; b=NhMMScEC4aVNhUD8AbEq8e717vgYClbGN/HCQ/5dnyw17zcWAAcx5ibhqJty84/41y Q7MI2/eiOxQ+R08yEncHkjW/W2x3pW0RGFy2WdmeD7ub+sNBa8WyKlj+F68Ry2hARN1u 1/YwCOlg5wxKBkxudafrCXtegn6tOMV+e95BSdAvW4pm34AbmTUd70Cm9ONieOFmvxrq DbuZrhvUyQEzD+q4Gi0H0QhJEtqR9gRseLgjhCs8fFgDCFM/akiezrx5WC1ecYMcCXJg JS10T1pHOaGVXZU4N2rLC0rEr/W6Dz+3JNI45A008I03xl9sLU7mYgyxpqlEGxi2Ovne EGIQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194277; x=1788799077; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=x77j6vdZ8LqsWXipiJe0J8DOL7X4B8I6QMSB+YfLeSc=; b=bIgfwfl2m4YHTmScj5nULU5x9uQB5NkzPfrqLtbgZj/vd4skTA8CB4G7JI8NWFCawP RMH5oPjLlWL0LRcjMUCw6us0xAdOPnz7U7IRBQmxaJ+5e7LDWOKREyW4hDmwBADrEV7D LXXHZZw4bIdYQn9TV1+AOqMwUxlxdYgtq6ImD+2gBR39OlZq0B70vDxcQPggWI/8us4Y VPE+4gby+wP3GbC2TpA4NdyUp0WqGcymE9GVARbejnCE6I6hU/IphcqY26xznOdZEEPl hdy0e4zYmPQppX0BsSqacdMC4asr4qKNJAr6q44+5yvJKrGMB8lxlO6aNHYe9/CF7Wc6 wPFw== X-Forwarded-Encrypted: i=1; AHgh+Rp2ALpnMjCxE25e8t7HRThuSgGZnNrTrffKHBq7CUvAf0yXsUpDVhwIMvgb2PVWHHXhtPRVpEp5vVes3aE=@vger.kernel.org X-Gm-Message-State: AFuF++mqMKhpRBRuuLqXEuZjDj9h6ZKUbBOZg1KlZGbY1WkkZ6uqovl2 gqMoQwAfTSmBwofFOODKkOBv/7Q0DWIkHy3I+dteqCMiyPzAxA91HMK2 X-Gm-Gg: AYBFou3PkdEPOGcRvFypT5gpP881HwwS8ySohYIlb4e5xsQtGHFS0cNi5vMW9/BWHJz bGe5H201s0OjnIxXh2XoKAHnL7m//T9Gx4q4qbB9zQgZUhiRsEe8b3OESjl9V9YyGEsm2ZfUEIP 9H4be5mK/AWX7X2gkpRR2DyJ44/yQMIGn0KfnLc2OtTvqCU5tPDUuA1UNY1v4oe/BOoIW1HppwU 9RuEKJtbjVCWHaV47W9bBXJW3mtkR4yw1zy9MAkLAFb5KxQ18ByivQ1DEtfI0XByYOSj5uhSWCy YXzP5tFLb/0s5xuwbuwC8IBUUw4Xd7MxZ4+vHaFQYuoohpMbqOAQ+UUqEliZ+U+QnzLn68HoQiR WtFZWx2xCCuhfr3bQpT99gSADgxK7BVqw7CCf4470AGKRq8HRlKQpZtSdOs6m141GAz3en6SXtV Lv/+ODF0eloDzrhytJhCKuo2I1zGHkxEyIp7QX4At/4nZym040LXzJ4vBZ6WLDidsqEJUUGI02S f34i1EXBOPVxfShroyZ X-Received: by 2002:a05:6870:3313:b0:448:3eb2:8c8c with SMTP id 586e51a60fabf-46835d7d79fmr26983191fac.6.1788194276047; Mon, 31 Aug 2026 09:37:56 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:2c::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-468a545dfe6sm11149786fac.14.2026.08.31.09.37.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:37:55 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 2/7] mm/page_counter: report the number of pages charged Date: Mon, 31 Aug 2026 09:37:46 -0700 Message-ID: <20260831163752.2193337-3-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add an optional @nr_charged parameter to page_counter_try_charge. On success, it will be set to the number of pages actually charged to the hierarchy. Today this number is always @nr_pages, so there is no functional change. Of the 6 callsites, only one user (try_charge_memcg) uses that information. The number of charged pages is added to current->memcg_nr_pages_over_high to indicate how many pages it charged to the hierarchy while over high. Today, try_charge_memcg requests "batch" from page_counter_try_charge and adds that same amount to memcg_nr_pages_over_high on success, since page_counter_try_charge's only source of charges is the hierarchy. However, this invariant changes later in the series when stock is pushed down from the memcg level to the page_counter level, and a page_counter charge can be successful without growing the hierarchy size. Plumb the new parameter to all callsites, passing NULL where the source of charge does not matter to the caller, and passing &nr_charged in try_charge_memcg to account the hierarchy size growth. No functional change intended. Signed-off-by: Joshua Hahn --- include/linux/page_counter.h | 4 ++-- kernel/cgroup/dmem.c | 2 +- mm/hugetlb_cgroup.c | 2 +- mm/memcontrol-v1.c | 2 +- mm/memcontrol.c | 10 ++++++---- mm/page_counter.c | 10 ++++++++-- 6 files changed, 19 insertions(+), 11 deletions(-) diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index d649b6bbbc871..89a083f16fbf7 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -71,8 +71,8 @@ static inline unsigned long page_counter_read(struct page= _counter *counter) void page_counter_cancel(struct page_counter *counter, unsigned long nr_pa= ges); void page_counter_charge(struct page_counter *counter, unsigned long nr_pa= ges); bool page_counter_try_charge(struct page_counter *counter, - unsigned long nr_pages, - struct page_counter **fail); + unsigned long nr_pages, struct page_counter **fail, + unsigned long *nr_charged); void page_counter_uncharge(struct page_counter *counter, unsigned long nr_= pages); void page_counter_set_min(struct page_counter *counter, unsigned long nr_p= ages); void page_counter_set_low(struct page_counter *counter, unsigned long nr_p= ages); diff --git a/kernel/cgroup/dmem.c b/kernel/cgroup/dmem.c index 4683f3d680226..fbbbd0b09d290 100644 --- a/kernel/cgroup/dmem.c +++ b/kernel/cgroup/dmem.c @@ -736,7 +736,7 @@ int dmem_cgroup_try_charge(struct dmem_cgroup_region *r= egion, u64 size, goto err; } =20 - if (!page_counter_try_charge(&pool->cnt, size, &fail)) { + if (!page_counter_try_charge(&pool->cnt, size, &fail, NULL)) { if (ret_limit_pool) { *ret_limit_pool =3D container_of(fail, struct dmem_cgroup_pool_state, c= nt); css_get(&(*ret_limit_pool)->cs->css); diff --git a/mm/hugetlb_cgroup.c b/mm/hugetlb_cgroup.c index ecb6e0b7819a0..6df4a69b0d529 100644 --- a/mm/hugetlb_cgroup.c +++ b/mm/hugetlb_cgroup.c @@ -274,7 +274,7 @@ static int __hugetlb_cgroup_charge_cgroup(int idx, unsi= gned long nr_pages, =20 if (!page_counter_try_charge( __hugetlb_cgroup_counter_from_cgroup(h_cg, idx, rsvd), - nr_pages, &counter)) { + nr_pages, &counter, NULL)) { ret =3D -ENOMEM; hugetlb_event(h_cg, idx, HUGETLB_MAX); css_put(&h_cg->css); diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index bf2c7d53b01b1..cf514d1bd7c38 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -2194,7 +2194,7 @@ bool memcg1_charge_skmem(struct mem_cgroup *memcg, un= signed int nr_pages, { struct page_counter *fail; =20 - if (page_counter_try_charge(&memcg->tcpmem, nr_pages, &fail)) { + if (page_counter_try_charge(&memcg->tcpmem, nr_pages, &fail, NULL)) { memcg->tcpmem_pressure =3D 0; return true; } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 93c2fa04da4fd..71410084fa7fc 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2663,6 +2663,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, struct mem_cgroup *mem_over_limit; struct page_counter *counter; unsigned long nr_reclaimed; + unsigned long nr_charged =3D 0; bool passed_oom =3D false; unsigned int reclaim_options; bool drained =3D false; @@ -2680,13 +2681,14 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, =20 reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP; if (do_memsw_account() && - !page_counter_try_charge(&memcg->memsw, batch, &counter)) { + !page_counter_try_charge(&memcg->memsw, batch, &counter, NULL)) { mem_over_limit =3D mem_cgroup_from_counter(counter, memsw); reclaim_options &=3D ~MEMCG_RECLAIM_MAY_SWAP; goto reclaim; } =20 - if (page_counter_try_charge(&memcg->memory, batch, &counter)) + if (page_counter_try_charge(&memcg->memory, batch, &counter, + &nr_charged)) goto done_restock; =20 if (do_memsw_account()) @@ -2847,7 +2849,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, * and distribute reclaim work and delay penalties * based on how much each task is actually allocating. */ - current->memcg_nr_pages_over_high +=3D batch; + current->memcg_nr_pages_over_high +=3D nr_charged; set_notify_resume(current); break; } @@ -5771,7 +5773,7 @@ int __mem_cgroup_try_charge_swap(struct folio *folio) rcu_read_unlock(); =20 if (!mem_cgroup_is_root(memcg) && - !page_counter_try_charge(&memcg->swap, nr_pages, &counter)) { + !page_counter_try_charge(&memcg->swap, nr_pages, &counter, NULL)) { memcg_memory_event(memcg, MEMCG_SWAP_MAX); memcg_memory_event(memcg, MEMCG_SWAP_FAIL); mem_cgroup_private_id_put(memcg, nr_pages); diff --git a/mm/page_counter.c b/mm/page_counter.c index 661e0f2a5127a..a934619cc7bf7 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -111,13 +111,15 @@ void page_counter_charge(struct page_counter *counter= , unsigned long nr_pages) * @counter: counter * @nr_pages: number of pages to charge * @fail: points first counter to hit its limit, if any + * @nr_charged: optional; on success, set to the number of pages actually + * charged to the hierarchy * * Returns %true on success, or %false and @fail if the counter or one * of its ancestors has hit its configured limit. */ bool page_counter_try_charge(struct page_counter *counter, - unsigned long nr_pages, - struct page_counter **fail) + unsigned long nr_pages, struct page_counter **fail, + unsigned long *nr_charged) { struct page_counter *c; bool protection =3D track_protection(counter); @@ -162,6 +164,10 @@ bool page_counter_try_charge(struct page_counter *coun= ter, WRITE_ONCE(c->watermark, new); } } + + if (nr_charged) + *nr_charged =3D nr_pages; + return true; =20 failed: --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-ot1-f47.google.com (mail-ot1-f47.google.com [209.85.210.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B0F524A2E1B for ; Mon, 31 Aug 2026 16:37:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194280; cv=none; b=mloBC08qxAETdbwiOtEXNUtf9odQODF8kOHKTNFF+U1Gf6OzGb4qp3rp7firUuR5HZ083RbEhWwtNiQKydlq5zyOfQISbrauVAdKslPmsAm2x4UqbqWUmh2QVQ4AQsxCVFkUA89RxWYyp/3MOPUSK/rBL1S5+A18k9FiBnLSfaM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194280; c=relaxed/simple; bh=rJFpJjImstzEz9X/weAONaIfNQqNwroPzDQP73j8CWs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DeEzY/ogVoozCPpYSnBSeY6T7bnAK3ngiNUB6OE7SB8u5EWBkS0yBxOtgVwNfeg3vIlUe1P3mUtTVXQLvrfSQ49rhcurPX+XmpBt0aTqDG0fywjfPSVFsJPVzEHymLGR6iNMGT12Fh+h6EI94pWBBk9RKcrS6I86sGqI3gbXJc4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=N6GXDAXH; arc=none smtp.client-ip=209.85.210.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="N6GXDAXH" Received: by mail-ot1-f47.google.com with SMTP id 46e09a7af769-7f4e729368fso3010911a34.0 for ; Mon, 31 Aug 2026 09:37:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194277; x=1788799077; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UZzK6JqUVQtEF4ncuPC417FBn/U4i41SWhLzyFT/8Js=; b=N6GXDAXHbUOQxq0ydFPAAvDt7iR+d1ZmdUByPJOY/l2X6jK1oDCkMPXVfVF0/VAP4u HxBQFDHDixYpSJNz6c1qEm5RMsPtgtzTxAfFlDFRRXOKdxRlRQNfzZORlgyvRDHJb/zj KDVgLkHNwASS7gIA26C5uL1EOENAms6s8voa9q0TiwlurerJLTVLZjmdDqt9d7+pGi9l gJn+i3qnOpxg2z9+nN0ApdVh5zFGIH+VL3xLqQ6sKRg2Kkh7HiUVL7venX4xqhzQwfH3 i+PEonuzCBdz4f/lXiL1bPS27EGH0vfJZnJxor/iiokojuR51FbPLEo2aZgcDPXVXe00 X2Rw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194277; x=1788799077; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=UZzK6JqUVQtEF4ncuPC417FBn/U4i41SWhLzyFT/8Js=; b=Pu3K+tEOwOTH95iU7t9EGDg+e+o4LbrseCSm5G/APp5dGRdJ+bB+PMhd7ttdw8D95k 9zthtrWxelIjbLMEDDu6nEwa96mjAxAOKQK+1o+b2yyKtCbQLxtyx0xTF/OW0VCCdscQ n2NKSxM6G079ZRHB25UhCHznabZ0zk8abBs8qR6SO8U5+Nr5+RMkp3CNL2aI1JdZdr7D lUTKz9fzejXoocGjyoeRr6ng8LzlUmr8UwCRwpArAWvZgrBWTO93mrBAblYMsYlPxfWf gKETdoaVtwcA4+OjwT4lVP0PsjJ6uYKPbYFKRuBKSSM6kjlm8DNTdz61yrUomPwdCfyb 7umw== X-Forwarded-Encrypted: i=1; AHgh+Rp0/GOmEIlie+uEhUe+wrZR8rNxf20GUDe5j/mt4FqyIzlVFSkYYWgujRimEO7Jqqq+MfW0a1fibH+uN84=@vger.kernel.org X-Gm-Message-State: AFuF++kj40c9kcjA7LRbeJfjgYQOQHeFvy1Le7t8oN1bY/bKMuNAtjb5 D3HvvGfmeVbJ+Q2zwZc2/wL2ePgOTqvTfogoqiJyrShaCH6rPTyd2lH5 X-Gm-Gg: AR+sD11kTgNR5BicMOnjNqJFyZRYmpVzNX5gyTVlxat5rBJCOD+Vj37yw6bbCO8wqpF 05IcjU2FuNzE7mh7h285+E4CPgllaNbzy07f4Rjqlp7c+ZkZ7oV+0tpdd2/VpH4zNPPerrRMOjU IFM35YJsHORi0tRK93kqV2mrZW6bwU3UbMa0kRiYEPjm6YRaj9RsJpISmY7gtIU3muTygXFCzj6 9FmRH2kc74teSEw0I0cRjQuDOVPH9E9ju9jDiCJS2S0qKJSHnOmJcepg1APhlfpk4xzrL/J33sg FzMmIYHhIyYSRn+sBBtCzvb805+dT16KVtxVdtmhEpiroGc98g5nYKYFHsdLVQPHX2dewNVeDrh U+uUs3+wntUNpXODZOfHFvwUM/liKuXGpa+9+Nf5cfb4d2wDCHwFFyGCxjBLwHouMjjjIQQig7x io3I9rZZc660pkrgzwTMClVbj4O5TdQFUOCE9gsx/wJzIw0qMQdN0oP6l6Jry0+ZAhOC/rI6yXF 6pQrcVXQzXDFmSCgA== X-Received: by 2002:a05:6830:349a:b0:7e9:e288:2b60 with SMTP id 46e09a7af769-7f4f21cb626mr27446752a34.1.1788194277265; Mon, 31 Aug 2026 09:37:57 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:3::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f4fa96cf17sm8719481a34.15.2026.08.31.09.37.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:37:56 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 3/7] mm/page_counter: introduce per-page_counter stock Date: Mon, 31 Aug 2026 09:37:47 -0700 Message-ID: <20260831163752.2193337-4-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In order to avoid expensive hierarchy walks on every memcg charge and limit check, memcontrol uses per-cpu stocks (memcg_stock_pcp) to cache pre-charged pages and introduce a fast path to try_charge_memcg. However, there are a few quirks with the current implementation that can be improved upon. First, each memcg_stock_pcp can only cache the charges of 7 memcgs (NR_MEMCG_STOCK). When an 8th memcg wants to cache its charge on a CPU, a victim memcg is chosen among the 7 cached memcgs and is evicted, losing all cached charges. Second, stock draining is per-CPU rather than per-memcg. That is, when a memcg is under pressure and must retrieve all cached charges, it iterates through every CPU and drains the stock charges of all present memcgs. This means that one under-pressure memcg evicts the caches of all co-cpu-resident memcg stock caches. Finally, stock is tightly coupled with memcg, so adding new page_counters to memcg is an unscalable operation where only one counter gets to use the fastpath. We can address all of these concerns by pushing stock caches down to the page_counter level, and making each counter responsible for its own charge. Introduce struct page_counter_stock along with its allocation, free, and per-CPU drain helpers. No functional change intended. Suggested-by: Johannes Weiner Signed-off-by: Joshua Hahn --- include/linux/page_counter.h | 16 +++++++ mm/page_counter.c | 90 ++++++++++++++++++++++++++++++++++++ 2 files changed, 106 insertions(+) diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index 89a083f16fbf7..c1fe331f34e7e 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -5,8 +5,11 @@ #include #include #include +#include #include =20 +struct page_counter_stock; + struct page_counter { /* * Make sure 'usage' does not share cacheline with any other field in @@ -41,6 +44,13 @@ struct page_counter { unsigned long high; unsigned long max; struct page_counter *parent; + struct page_counter_stock __percpu *stock; + unsigned long batch; + + /* make sure the work_struct is separate from the read most fields */ + CACHELINE_PADDING(_pad3_); + + struct work_struct drain_work; } ____cacheline_internodealigned_in_smp; =20 #if BITS_PER_LONG =3D=3D 32 @@ -61,6 +71,8 @@ static inline void page_counter_init(struct page_counter = *counter, counter->parent =3D parent; counter->protection_support =3D protection_support; counter->track_failcnt =3D false; + counter->stock =3D NULL; + counter->batch =3D 0; } =20 static inline unsigned long page_counter_read(struct page_counter *counter) @@ -99,6 +111,10 @@ static inline void page_counter_reset_watermark(struct = page_counter *counter) counter->watermark =3D usage; } =20 +void page_counter_drain_cpu_stock(struct page_counter *counter, int cpu); +void page_counter_alloc_stock(struct page_counter *counter, unsigned long = batch); +void page_counter_free_stock(struct page_counter *counter); + #if IS_ENABLED(CONFIG_MEMCG) || IS_ENABLED(CONFIG_CGROUP_DMEM) void page_counter_calculate_protection(struct page_counter *root, struct page_counter *counter, diff --git a/mm/page_counter.c b/mm/page_counter.c index a934619cc7bf7..3f61eba695518 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -8,11 +8,18 @@ #include #include #include +#include #include #include +#include #include #include =20 +struct page_counter_stock { + raw_spinlock_t lock; + unsigned long nr_pages; +}; + static bool track_protection(struct page_counter *c) { return c->protection_support; @@ -295,6 +302,89 @@ int page_counter_memparse(const char *buf, const char = *max, return 0; } =20 +/** + * page_counter_drain_cpu_stock - release @cpu's cached charges + * @counter: counter whose stock to drain + * @cpu: CPU whose stock is drained + */ +void page_counter_drain_cpu_stock(struct page_counter *counter, int cpu) +{ + struct page_counter_stock __percpu *stock =3D READ_ONCE(counter->stock); + struct page_counter_stock *pcp_stock; + unsigned long nr_pages; + unsigned long flags; + + if (!stock) + return; + + pcp_stock =3D per_cpu_ptr(stock, cpu); + raw_spin_lock_irqsave(&pcp_stock->lock, flags); + nr_pages =3D pcp_stock->nr_pages; + pcp_stock->nr_pages =3D 0; + raw_spin_unlock_irqrestore(&pcp_stock->lock, flags); + + if (nr_pages) + page_counter_uncharge(counter, nr_pages); +} + +/** + * page_counter_alloc_stock - allocate the percpu stock for a page_counter + * @counter: counter to allocate percpu stock for + * @batch: maximum number of pages a CPU may cache + * + * Failure to allocate is not fatal; @counter falls back to hierarchy char= ges. + * The caller must not (un)charge @counter concurrently with this call, an= d this + * must not be called twice on the same counter. A concurrent drain is fine + * since the stock is published with a release store the drain paths pair = with. + * + * Context: Process context. May sleep, the percpu alloc uses GFP_KERNEL. + */ +void page_counter_alloc_stock(struct page_counter *counter, unsigned long = batch) +{ + struct page_counter_stock __percpu *stock; + int cpu; + + if (WARN_ON_ONCE(counter->stock)) + return; + + stock =3D alloc_percpu_gfp(struct page_counter_stock, GFP_KERNEL_ACCOUNT); + if (!stock) + return; + + for_each_possible_cpu(cpu) { + struct page_counter_stock *pcp_stock =3D per_cpu_ptr(stock, cpu); + + raw_spin_lock_init(&pcp_stock->lock); + } + + counter->batch =3D batch; + /* Publish stock only after percpu allocs / inits are finished */ + smp_store_release(&counter->stock, stock); +} + +/** + * page_counter_free_stock - free @counter's percpu cached charge + * @counter: page_counter whose stock to free + * + * Caller must guarantee no (un)charge or drain of @counter is in flight o= r can + * start. memcg only calls this once the cgroup is dead and unreachable. + */ +void page_counter_free_stock(struct page_counter *counter) +{ + struct page_counter_stock __percpu *stock =3D counter->stock; + int cpu; + + if (!stock) + return; + + /* Stop greedy over-charging before the stock goes away */ + counter->batch =3D 0; + for_each_possible_cpu(cpu) + page_counter_drain_cpu_stock(counter, cpu); + + WRITE_ONCE(counter->stock, NULL); + free_percpu(stock); +} =20 #if IS_ENABLED(CONFIG_MEMCG) || IS_ENABLED(CONFIG_CGROUP_DMEM) /* --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-oa1-f51.google.com (mail-oa1-f51.google.com [209.85.160.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B6C8D4A33E1 for ; Mon, 31 Aug 2026 16:37:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194282; cv=none; b=RwM2FdFKEwMGsPzo8KVR0Y//B9zyXr+TWoLo/4s9+/HmZrdGHBnx14MnnhGLTN/Ru5y9dxcC/IEzcyzmB392oPcuUyendEN1haYSXCAhPt7JGu/qFnav2WzWMC8QkpvhMauU4+nOtE3dn8deF5b2QZLRGSjGXC6CvKxJ3swUOGc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194282; c=relaxed/simple; bh=dZraJ7RTBatWlxTVjR14xRqPdJ/jnRpH20RMutxDuWU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aERM1QiZiFSxlblko3j+/7/PpptEX4EHzN6B3OSedwroLgxfFtAReNoSYMhrq4ZjZEF7vUvMN1AzLviTOlFK6HLe4Iuk/D4c4ts1sWlGRb6y3cg7EhxPfBCGQQLG12ThE4ZmWQrNPoyi+fnNEssHZpgnguqi0rz4gk3xEC0XvgI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NeHe404P; arc=none smtp.client-ip=209.85.160.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NeHe404P" Received: by mail-oa1-f51.google.com with SMTP id 586e51a60fabf-46ac530b9ebso74674fac.0 for ; Mon, 31 Aug 2026 09:37:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194278; x=1788799078; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=vXFNKbTdOrOMYaajm/HLCYSs/GSBNBi/73jyNLJmP+Y=; b=NeHe404Ph+4dXuR07BEA3+0AbDRA5pVVSNy6dyeXGmpqXVAilXlE2iceT9/J2hsX4D nw+hl1IELiqpe8dUx3tfaTaM6RiHyK6mrTImr1c4tG4TSKh0+p/dNU0r4MWWpev/UjtN mR1xyiLxt1+Lk50O2ibMz/r3OWX7tAL7zMOsj30hvwAVfjlwY6rMcfsUMaUmXUFSuSfQ ED6UYYR75xCWHOJCqK9o+m365DgABZ8+c6Qa49OGqyrrL8EnIdtJmGkTwfNerXnffwSI KERozlAfHRzSiT/0r7v44O/1YzP0F+WDz1D+fsoydAXPlWeplbnF5XIaxZgi8V5bQQ60 ypMw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194278; x=1788799078; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=vXFNKbTdOrOMYaajm/HLCYSs/GSBNBi/73jyNLJmP+Y=; b=Mdg4ux+6bZzZfYuvcov8WFN3aY7gG2r+EOsN/1ZCn7mrD49BsZoSKrf7w+cLeB8dvh 3vZENelCV8FVgKIUHYZcm1t6aEoJItl3j8oSXw2V2lG+Pxon9EYkkVNm1vcn43aZd12g TJnRUn2Bp4ReJCgC2SfWxml5/eudW/36EpvVlKyv+NJ6Fs9PzElyiL0qvwY0PVhnxlBR G3h8v7wD9+bcRRheVzDJ7yANn2d/KhqNtxTW5D81iM+tyenfpsCopbs4y5ZIs+zhEbDK S1gNaq6a5dD6bFB6Fz75xFJTjc3FZtWXbN+OPsWadyEjFtA7yoP1mHmmgOiCsLwNI9qO aBiw== X-Forwarded-Encrypted: i=1; AHgh+RoGzrlD3RvwjsM0s/LHBko4i8HZ9qK5k/qsXBbvi1kQJv7aUShb+4JubdwUstXLs3nLFlQfYbP2EI39gbE=@vger.kernel.org X-Gm-Message-State: AFuF++m/z58AJ3H/p+LbSV0pduBe9CVxqE4Zg4Eo2MH8djoBuZb0agNr mdrM3gtL7p9HJJ5953Y6K2mnEFygmgxst59JwO4yRoigYtzIRwkCIFKg X-Gm-Gg: AYBFou3Wh83wh6Jp5f1Vh+YU5Si1IN7nVCIOHrXamZznMkoDjVrj6d8LZV7qU8yH2Ed 75xjtSJIk7Jd6hL7qnE1ZktOmmlD8SrIO5nHxRlZ35MINLmdoKksj5wzg0ZCQi749kBj15TO7DB jE5qd8k981D7rY3vyojtatd2mJa/EsZFI7B1VN1pGfEOksxLv5G8HRBvM4X6k69DFLFv6YkXTQG HUi/G1LNOmhgL1cTDtoaNqlMhJLyu31HoWt0zvBZGmMzw6kebKhWDObHNlWcmlQBHXRs7IkALa0 bGYU2Ie8mKuZ/DGkHEeqdjyKmQRYW2aKLgY4VhGoAN0B545UIspfdr1jnHD8pC23+m3AozivBMU UxWWgvGvANTemqYp12jDwlNZvUBd8io9XuTRMsLmFsjQ8fef+P7ODdWGfE4TKkC1CDGHAVChEvC yziGZe1fhl38M6nk6/Chu2HVH/7gThVnLD3AMfZRsZyX3Kbf/eTpV3Qv3yJ2RxWMjgudTjAcGYi zRvLNE48wLNEUY+gQM= X-Received: by 2002:a05:6870:2b11:b0:456:4c6b:773a with SMTP id 586e51a60fabf-46aca5a5f8amr7723915fac.9.1788194278410; Mon, 31 Aug 2026 09:37:58 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:30::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-468a4effb4dsm10497204fac.10.2026.08.31.09.37.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:37:58 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 4/7] mm/page_counter: use stock in page_counter_try_charge Date: Mon, 31 Aug 2026 09:37:48 -0700 Message-ID: <20260831163752.2193337-5-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Transparently make page_counter_try_charge attempt to service the charge from its stock. We preserve the same semantics as the existing stock management in try_charge_memcg: 1. Limit-check against the stock. If there is enough, then skip the hierarchy walk and charge to the stock. 2. Greedily attempt to fulfill the charge request and refill the stock simultaneously to the hierarchy. 3. If this fails, retry the stock and charge without trying to refill the stock, i.e. with the number of pages requested. 4. If the greedy attempt succeeds, return excess pages to the stock. page_counter_refill_stock() falls back to a hierarchical uncharge when there is no stock, in NMI contexts, on lock contention, or for a refill larger than the batch. The greedy charge is also skipped in NMI where both stock helpers bail out since the batch charge would be undone again. No functional change intended, since no page_counter enables stock yet and counter->batch is left at 0. Suggested-by: Johannes Weiner Signed-off-by: Joshua Hahn --- include/linux/page_counter.h | 2 + mm/page_counter.c | 135 +++++++++++++++++++++++++++++++---- 2 files changed, 125 insertions(+), 12 deletions(-) diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index c1fe331f34e7e..428ca8e7b2da5 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -82,6 +82,8 @@ static inline unsigned long page_counter_read(struct page= _counter *counter) =20 void page_counter_cancel(struct page_counter *counter, unsigned long nr_pa= ges); void page_counter_charge(struct page_counter *counter, unsigned long nr_pa= ges); +unsigned long page_counter_refill_stock(struct page_counter *counter, + unsigned long overage); bool page_counter_try_charge(struct page_counter *counter, unsigned long nr_pages, struct page_counter **fail, unsigned long *nr_charged); diff --git a/mm/page_counter.c b/mm/page_counter.c index 3f61eba695518..a76949abf04e7 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -113,25 +113,126 @@ void page_counter_charge(struct page_counter *counte= r, unsigned long nr_pages) } } =20 +static bool page_counter_consume_stock(struct page_counter *counter, + unsigned long nr_pages) +{ + struct page_counter_stock __percpu *stock =3D READ_ONCE(counter->stock); + struct page_counter_stock *pcp_stock; + unsigned long flags; + bool charged =3D false; + + if (!stock || nr_pages > counter->batch) + return false; + + /* raw_spin_trylock isn't enough to protect against nested NMI in UP */ + if (in_nmi()) + return false; + + /* It's OK to migrate here, since stock is fungible within a counter. */ + pcp_stock =3D raw_cpu_ptr(stock); + + if (!raw_spin_trylock_irqsave(&pcp_stock->lock, flags)) + return false; + + if (pcp_stock->nr_pages >=3D nr_pages) { + pcp_stock->nr_pages -=3D nr_pages; + charged =3D true; + } + + raw_spin_unlock_irqrestore(&pcp_stock->lock, flags); + return charged; +} + +/** + * page_counter_refill_stock - return pages to a page_counter's stock + * @counter: counter to return the pages to + * @overage: number of pages to return + * + * Return: how many of @overage went to the hierarchy rather than the stoc= k. + * The flush itself can be larger, since it also returns what earlier call= ers + * stocked. + */ +unsigned long page_counter_refill_stock(struct page_counter *counter, + unsigned long overage) +{ + struct page_counter_stock __percpu *stock =3D READ_ONCE(counter->stock); + struct page_counter_stock *pcp_stock; + unsigned long high =3D counter->batch; + unsigned long low =3D high / 2; + unsigned long to_flush =3D overage; + unsigned long stocked; + unsigned long flags; + + if (!stock || overage > high) + goto uncharge_counter; + + /* See page_counter_consume_stock() for why NMI skips the stock. */ + if (in_nmi()) + goto uncharge_counter; + + /* It's OK to migrate here, since stock is fungible within a counter. */ + pcp_stock =3D raw_cpu_ptr(stock); + if (!raw_spin_trylock_irqsave(&pcp_stock->lock, flags)) + goto uncharge_counter; + + /* + * Use a high/low watermark here, in the spirit of pcp->{batch, high}. + * If the stock would exceed counter->batch, stock is trimmed to the low + * watermark of counter->batch / 2 so that sequential uncharges don't + * all trigger a hierarchy walk. + */ + stocked =3D pcp_stock->nr_pages + overage; + if (stocked > high) { + pcp_stock->nr_pages =3D low; + to_flush =3D stocked - low; + } else { + pcp_stock->nr_pages =3D stocked; + to_flush =3D 0; + } + raw_spin_unlock_irqrestore(&pcp_stock->lock, flags); + + if (!to_flush) + return 0; + +uncharge_counter: + page_counter_uncharge(counter, to_flush); + return min(overage, to_flush); +} + /** * page_counter_try_charge - try to hierarchically charge pages * @counter: counter * @nr_pages: number of pages to charge - * @fail: points first counter to hit its limit, if any + * @fail: only written on failure; the first counter to hit its limit * @nr_charged: optional; on success, set to the number of pages actually - * charged to the hierarchy + * charged to the hierarchy. Set to 0 if stock was served. + * + * A successful charge may still have bumped failcnt on @counter or an + * ancestor, since the greedy attempt is retried at the requested size. * - * Returns %true on success, or %false and @fail if the counter or one - * of its ancestors has hit its configured limit. + * Returns %true on success, or %false and sets @fail if the counter or + * one of its ancestors has hit its configured limit. */ bool page_counter_try_charge(struct page_counter *counter, unsigned long nr_pages, struct page_counter **fail, unsigned long *nr_charged) { - struct page_counter *c; + struct page_counter *c, *failed_at; + unsigned long charge =3D nr_pages; bool protection =3D track_protection(counter); bool track_failcnt =3D counter->track_failcnt; =20 + /* The stock is skipped in NMI; a greedy charge would just be undone */ + if (!in_nmi()) + charge =3D max(counter->batch, nr_pages); + +retry: + if (page_counter_consume_stock(counter, nr_pages)) { + if (nr_charged) + *nr_charged =3D 0; + return true; + } + for (c =3D counter; c; c =3D c->parent) { long new; /* @@ -148,9 +249,9 @@ bool page_counter_try_charge(struct page_counter *count= er, * we either see the new limit or the setter sees the * counter has changed and retries. */ - new =3D atomic_long_add_return(nr_pages, &c->usage); + new =3D atomic_long_add_return(charge, &c->usage); if (new > c->max) { - atomic_long_sub(nr_pages, &c->usage); + atomic_long_sub(charge, &c->usage); /* * This is racy, but we can live with some * inaccuracy in the failcnt which is only used @@ -158,7 +259,7 @@ bool page_counter_try_charge(struct page_counter *count= er, */ if (track_failcnt) data_race(c->failcnt++); - *fail =3D c; + failed_at =3D c; goto failed; } if (protection) @@ -172,15 +273,25 @@ bool page_counter_try_charge(struct page_counter *cou= nter, } } =20 + if (charge > nr_pages) + charge -=3D page_counter_refill_stock(counter, charge - nr_pages); + if (nr_charged) - *nr_charged =3D nr_pages; + *nr_charged =3D charge; =20 return true; =20 failed: - for (c =3D counter; c !=3D *fail; c =3D c->parent) - page_counter_cancel(c, nr_pages); + for (c =3D counter; c !=3D failed_at; c =3D c->parent) + page_counter_cancel(c, charge); + + /* Retry the stock & charge with the exact number of pages requested */ + if (charge > nr_pages) { + charge =3D nr_pages; + goto retry; + } =20 + *fail =3D failed_at; return false; } =20 @@ -330,7 +441,7 @@ void page_counter_drain_cpu_stock(struct page_counter *= counter, int cpu) /** * page_counter_alloc_stock - allocate the percpu stock for a page_counter * @counter: counter to allocate percpu stock for - * @batch: maximum number of pages a CPU may cache + * @batch: number of pages to precharge and the stock's high watermark * * Failure to allocate is not fatal; @counter falls back to hierarchy char= ges. * The caller must not (un)charge @counter concurrently with this call, an= d this --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-ot1-f50.google.com (mail-ot1-f50.google.com [209.85.210.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E4E634A3850 for ; Mon, 31 Aug 2026 16:38:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194282; cv=none; b=IUVDKWjnGIxn3k6YF25WuJGXFogqU7OPwfgEy6xt6egedHEuJV9vZYZcqjVZw5KHyox7Jg+FeSUIvKsVcvn/zHVfbARN/xxEoa7mzJiF8qv8SVBc2GVN4DApxVbBSN+tgjjqvFr0y7OMfh3ZcjE1yadiFxOdnghRL+Y3CjtR2w4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194282; c=relaxed/simple; bh=GGcdgcidJRK4f6nVOmkrYxscF0UoppYT966YKJwy27I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kjqMvYRmXTH2Hyk192NKTFYu6mI9jCATTfPC4PNf8f8YcTYuuX+F+Sb3mSjVM5g4K2F2pO6AxV4F3SQ4t/8d32IaTWHNSSMKQ5S4HY33zfLSGsp+1F2s0zsdv4USwo0sriRkoBFkYpVsS1LbVtdBKt8StRdo+IKM1bcHMKmGg5U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=nXxPixKM; arc=none smtp.client-ip=209.85.210.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="nXxPixKM" Received: by mail-ot1-f50.google.com with SMTP id 46e09a7af769-7f4e729368fso3010946a34.0 for ; Mon, 31 Aug 2026 09:38:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194279; x=1788799079; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tZLPtkE4K+x7inkg5AOVUnb1Simer6GkDXLFxdMoHE0=; b=nXxPixKMBGtAS8SA87yCVSy8wlTGTGgLh3py9zTibbLgoySarv+kJOAdCWTx4Ow7UB zcBIi/6e7InTjykVKfp8k+oeooGm4iHaAxucVM5Z40vNnZrWimD+sDYdcFeoa1mBaBny RumlAz4lowAWmcHEHSOxySgN6cm6IlsdgfUK6w8YpeH/4vTSBvohATI+FtXJYioEh1e+ iorkPyLZcMplD74VcGDyVJVdlnD6QCDX6t/h/Jfbe25kNTp0aYmnxOJ8tReUxsR/vAgM 35ZPYGa+ppJ06hX0Tj4U6P2yPJlNTc8ltPc62dSW42RFYN50VhKH+ia/bSpP21PakcVs kC2w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194279; x=1788799079; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=tZLPtkE4K+x7inkg5AOVUnb1Simer6GkDXLFxdMoHE0=; b=bxDjGMRfrrBkJQTyh95oc0b8FJW8Dbrdz/8xLIyr5pg3z3ou7ubPxRqbfB7zVO4q8K hF5YsI0ktGYNIYn/yKPfQUwnH0P/2OfuKMJwRbKBgOke6T1hy7WBiX1oG4o5WGS6FOE9 PgxTiax1onrdG2YsQp2H92UwxUrOnM9Lya+Oa1t6rSo1sW7XV7tX6LqzlqjSe/OLxHec 67Ca08vTtkU/a1UgytyS9/ysT5a5kJOy+GSxKsBGLqKFNheJsqESJcHPXhFIbrxWPdVh gFBPzh+odZPiswRSSYFeDqPkpRl8u3WAdRUCx1G24TiIDkH1TJYDSSTIsyskE3OuapxE HeBw== X-Forwarded-Encrypted: i=1; AHgh+RpVk+83ZpjQqP0221+OlV552Utq3Rga+AWJisAJHMyfg/7QGtXFlv8XqxaZMkbb0AZdd/E9RimG+55I5tE=@vger.kernel.org X-Gm-Message-State: AFuF++lOEqGwIg8FmAV837fFdgwabAKWXuzY2HV9KCkrcuUqGLu/+Wmr Y0i5FvUPW8f0pdN2LEHJH4++23blLYTktgC+GKRi41y3MvsyiNIsH4td X-Gm-Gg: AR+sD138n+OD3wyyweQAb1AvQAKsTCs8lBmE2n0aItIJG5lZCz1n0oxXEs9ol/VwqjD pTv+9X/QkkBtNtnaNwCybO00+LzPwXHiSD8ci715tGfzH9LfH1sZ9AZngyv2IiU53vUdcWlezEF dUoZYPPkXxmyYIPg/IQiPaEtslqwYtO3j2BAoJNbqM3owQjPAw+CUCeEEbObNn6kQXQWRce3kHE dJIXFX/m9+H7XyGJHdRP8DJMP4xM0Qh4tFGB/K+Xoy216MZVItF1lKCUhdOdX2OaAyAZAl/Bj+a lJviQMPaljHoSM6Aq/UmgIZHfVl9u/yIid137dVSeTGEEzdCTJQaAKnbWpPGAqcNAjaikqzC7jF E3d1UU9mxMVyAuypSXFZvHa+Jx252QtQBxzQXwagECNJQGoSOmH0z8mRg2lpi5cO9irDa+6F7Uc rq21SBwh80yu+AzYT1sEr1frc1hbc/FLSHoPVY4LvWvkVzru69W9ToZWzYjGWXieLPijYyWuTh1 ry2L/hBbskY9eShw8X7 X-Received: by 2002:a05:6820:4d02:b0:6b1:afc8:3162 with SMTP id 006d021491bc7-6b1c6642c52mr23408705eaf.7.1788194279553; Mon, 31 Aug 2026 09:37:59 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:2b::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-468a2fedbcasm10399867fac.2.2026.08.31.09.37.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:37:59 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 5/7] mm/page_counter: introduce an asynchronous drainer Date: Mon, 31 Aug 2026 09:37:49 -0700 Message-ID: <20260831163752.2193337-6-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The existing percpu memcg stock drainer schedules a stock drain worker per-cpu, one for every CPU containing the target memcg's stock. One issue with this design is the coarseness of the drain; when a worker runs on a CPU, it drains not only the target memcg's stock, but all other (up to 6) memcgs who are stocked on that CPU. Instead, use a per-page_counter drainer that iterates through each CPU and flushes any existing charges, leaving other unrelated memcgs' stock alone. Since that walks every possible CPU, the per-cpu helper now skips the remote lock when a stock looks empty. One benefit of having one worker flush through all CPUs is that duplicate drain requests when a worker is already queued are coalesced, since the asynchronous drainer takes no arguments and flushes all CPUs. We use the system_dfl_wq for this asynchronous drain worker, since memcg_wq is a percpu workqueue. Note that neither workqueue has WQ_MEM_RECLAIM. Previously the local CPU was drained inline because local_lock made it the only reachable stock. The lock is now a per-cpu raw_spinlock_t that any CPU can take, so drain any CPU's stock to return pages immediately. Suggested-by: Johannes Weiner Signed-off-by: Joshua Hahn --- include/linux/page_counter.h | 1 + mm/page_counter.c | 55 ++++++++++++++++++++++++++++++++---- 2 files changed, 50 insertions(+), 6 deletions(-) diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h index 428ca8e7b2da5..b10ef785f06de 100644 --- a/include/linux/page_counter.h +++ b/include/linux/page_counter.h @@ -114,6 +114,7 @@ static inline void page_counter_reset_watermark(struct = page_counter *counter) } =20 void page_counter_drain_cpu_stock(struct page_counter *counter, int cpu); +void page_counter_drain_stock_async(struct page_counter *counter); void page_counter_alloc_stock(struct page_counter *counter, unsigned long = batch); void page_counter_free_stock(struct page_counter *counter); =20 diff --git a/mm/page_counter.c b/mm/page_counter.c index a76949abf04e7..e6cfb5865ba75 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -12,6 +12,7 @@ #include #include #include +#include #include #include =20 @@ -135,7 +136,7 @@ static bool page_counter_consume_stock(struct page_coun= ter *counter, return false; =20 if (pcp_stock->nr_pages >=3D nr_pages) { - pcp_stock->nr_pages -=3D nr_pages; + WRITE_ONCE(pcp_stock->nr_pages, pcp_stock->nr_pages - nr_pages); charged =3D true; } =20 @@ -183,10 +184,10 @@ unsigned long page_counter_refill_stock(struct page_c= ounter *counter, */ stocked =3D pcp_stock->nr_pages + overage; if (stocked > high) { - pcp_stock->nr_pages =3D low; + WRITE_ONCE(pcp_stock->nr_pages, low); to_flush =3D stocked - low; } else { - pcp_stock->nr_pages =3D stocked; + WRITE_ONCE(pcp_stock->nr_pages, stocked); to_flush =3D 0; } raw_spin_unlock_irqrestore(&pcp_stock->lock, flags); @@ -429,15 +430,53 @@ void page_counter_drain_cpu_stock(struct page_counter= *counter, int cpu) return; =20 pcp_stock =3D per_cpu_ptr(stock, cpu); + + /* + * Skip the remote lock when empty. Racing with the charge path is why + * nr_pages uses WRITE_ONCE(); a stale read defers to the next drain. + */ + if (!READ_ONCE(pcp_stock->nr_pages)) + return; + raw_spin_lock_irqsave(&pcp_stock->lock, flags); nr_pages =3D pcp_stock->nr_pages; - pcp_stock->nr_pages =3D 0; + WRITE_ONCE(pcp_stock->nr_pages, 0); raw_spin_unlock_irqrestore(&pcp_stock->lock, flags); =20 if (nr_pages) page_counter_uncharge(counter, nr_pages); } =20 +static void page_counter_drain_work_fn(struct work_struct *work) +{ + struct page_counter *counter =3D container_of(work, struct page_counter, + drain_work); + int cpu; + + for_each_possible_cpu(cpu) + page_counter_drain_cpu_stock(counter, cpu); +} + +/** + * page_counter_drain_stock_async - schedule a page_counter stock drain + * @counter: page_counter to drain + * + * Drains any CPU's stock inline, then schedules a drain over every CPU and + * returns. Concurrent requests coalesce onto the same queued work. @count= er + * must outlive that work, which page_counter_free_stock() cancels. + * Must not be called from NMI context. + */ +void page_counter_drain_stock_async(struct page_counter *counter) +{ + /* Pairs with the smp_store_release() in page_counter_alloc_stock() */ + if (!smp_load_acquire(&counter->stock)) + return; + + /* Drain any CPU's stock to immediately return pages; migrating is OK */ + page_counter_drain_cpu_stock(counter, raw_smp_processor_id()); + queue_work(system_dfl_wq, &counter->drain_work); +} + /** * page_counter_alloc_stock - allocate the percpu stock for a page_counter * @counter: counter to allocate percpu stock for @@ -469,6 +508,7 @@ void page_counter_alloc_stock(struct page_counter *coun= ter, unsigned long batch) } =20 counter->batch =3D batch; + INIT_WORK(&counter->drain_work, page_counter_drain_work_fn); /* Publish stock only after percpu allocs / inits are finished */ smp_store_release(&counter->stock, stock); } @@ -477,8 +517,9 @@ void page_counter_alloc_stock(struct page_counter *coun= ter, unsigned long batch) * page_counter_free_stock - free @counter's percpu cached charge * @counter: page_counter whose stock to free * - * Caller must guarantee no (un)charge or drain of @counter is in flight o= r can - * start. memcg only calls this once the cgroup is dead and unreachable. + * Caller must guarantee no (un)charge or drain of @counter can start, and= must + * be ordered against the last CPU to touch the stock, since the drain pee= ks at + * nr_pages unlocked. memcg's RCU grace period before css_free provides bo= th. */ void page_counter_free_stock(struct page_counter *counter) { @@ -490,6 +531,8 @@ void page_counter_free_stock(struct page_counter *count= er) =20 /* Stop greedy over-charging before the stock goes away */ counter->batch =3D 0; + /* Make sure pending drainers don't run on freed page_counters */ + cancel_work_sync(&counter->drain_work); for_each_possible_cpu(cpu) page_counter_drain_cpu_stock(counter, cpu); =20 --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-oo1-f51.google.com (mail-oo1-f51.google.com [209.85.161.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E97CD4A3865 for ; Mon, 31 Aug 2026 16:38:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194284; cv=none; b=PleOHzE81Sy5ZQQJTlShgMdgWTB94zOqFNqy8s8e+nAojwhIc9gfUMvDHHmlvatpckOOJL/jo0oxvZCh1fw5o//Hpd1fMRlJonAma4ek+gb0D4rifVUE4BZxFDlYFKCY1x3m+jr4va9lSNIm1ssk/cxVvE1zOAkw0g8uiqbTXFc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194284; c=relaxed/simple; bh=U/TkUUz+sJRWRflpkXsQzp0UQSOV8Bu7uNvlbxcsYAo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XMAI8orF+lbN07pD15+IWX4f8EPcXj1T5gjdEDqGuydGQO7MOLfw3Y2vzy1e5I81V4ItR9NBx9OEDVMh7DpjnZ434FDypteYQfx6PimA1C/GD5x+EUIP6zgM+IM4wBtAxh8JQg+ib13RaSDcaXLyqdWJOiS0mKbYisQjXqJ3zRI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SHY9w46p; arc=none smtp.client-ip=209.85.161.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SHY9w46p" Received: by mail-oo1-f51.google.com with SMTP id 006d021491bc7-6b1bcd9e00bso1380184eaf.2 for ; Mon, 31 Aug 2026 09:38:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194281; x=1788799081; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=eQZlJ7zrV8AgxyRHbYvN1myWpv2l0XNczJpzFDxqb6c=; b=SHY9w46puYGV6wGbboRRWcTOGpZneWAFJm4xXjxg08zLSjqAafGPlIw+GCFH7dP0aA 3Q+wdcJP5aFUWY0xbs3oa46LojXz5EC01MZ5G0Hc0XXaashmET/n/1ZUrIM55j8YJcj8 easZ7bEfJu8kml46C5DGwfYJW+BtPr5UfXQCmaW1vQxW9yyNr9T302Oo6sjbash26rxb fLJQG6s7A3vS+QG9DHiTQIlAyk1voaazzuuFnJiyF0lP/hewAizjjGecb+DjSwOTzioy Hv1+Ow9uLXnXQTZVs/SJg/J4b+0ib/neUUDn3hvkCDP6meVxxMSryvPSGd0NE4COOCO1 Fz9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194281; x=1788799081; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=eQZlJ7zrV8AgxyRHbYvN1myWpv2l0XNczJpzFDxqb6c=; b=nSIKrq6/wu+H/scEk2TGYcwfEc093WZbjVrcKuk2alIqAb5W12CFwKk1Xw3RcG3E0Q GBDSN3SWR9T4me3aU46O0th7SssSksK7hw5mwU0P5aGAJeGJhOlWL8w5OhtWRYh9Z9az pOW9WLFRXm3JX7nnQj1w2hsI+AF/JEVut/wxl280Kt0K/0MO5ICvXf83+Az1KmwdvqWd jn2RKLcoxO+NCHcXYQwq8uNlklBbRsOVCXUIGI+0o1p36bYV6DGYwalwGIRv262/H/Qv trmoCvLB8X8Y69fzQMS6NnWvN20PlqqxImzqD7oPkcSd1ZJ4ZyjL7NRZEySxdKeuSect rGfA== X-Forwarded-Encrypted: i=1; AHgh+Rq3RnjYiM+uMlSgFOipYc9wKmo+jgJiwpPvTeZSE+DxcrCnMulFVfubVDZ7prAeK9UjPIiVc6B3F1Due3k=@vger.kernel.org X-Gm-Message-State: AFuF++n46BWHaUiWKXlm1wqwE/0B4DM91ddMWvWiRLgOvyUXi2uSjxxr gQNGIfVwFU6VjZrmDxsCZCvASBvowZnxA6ZYyNOrPgusLg8OKkTPNC2v X-Gm-Gg: AR+sD13WpPuPYZ0lwnu+ZTwZeM0zSu77pS5kiFtxmP8M4rbo1C4fH8lyOUu1KUSP5xe FN7FVjtCXwFULY0Z5eRpzOp34Yo20lEwQPSan4Im8zN21NZ+ykwkQwuGn+i9TFzFWk89xD7vfDN qoQgVGaxpnj4csN0k5Y7dc8QpmvQQsgM75VvXoxRPIhemQhgzXkbxIVZLocvDRuD0crfsh9pBLZ oJptYjR6YNHVKPLc6fqsd8jkU/69kxaSPkAc0zwBoRxqQBCisVtSwdmJMxD58Gm2WyA36M8X+RV p9WARzQ0c6NGNLeU02UJGvfDMDwxeU9ZVmu/Dk7mTev2h+T3l2Whc1QQ2dZRFoemxwe2Q4P16DQ 6LsBwqG78NbmCKgs5eCYWXwpFFYtZRakPG18ZL+DpmH39YT+4vx6ka/adXxdGAoYkYyYymmHVQh mBzcutTVzvsrPRFfKDMc5OGtrC/dSuNoGxIGs7/pdJqeAu1eME2Y9x2QFdDhO+Gde0x7cbqtHKX DyuAnkMIFwgcBeV5Zs= X-Received: by 2002:a05:6820:4d04:b0:6b1:7289:fc89 with SMTP id 006d021491bc7-6b1c659cf28mr24365399eaf.8.1788194280716; Mon, 31 Aug 2026 09:38:00 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:25::]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-468a4effec0sm10484590fac.12.2026.08.31.09.38.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:38:00 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 6/7] mm/memcontrol: convert memcg to use page_counter_stock Date: Mon, 31 Aug 2026 09:37:50 -0700 Message-ID: <20260831163752.2193337-7-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Now that page_counter transparently handles stock usage and refills, switch memcg to use page_counter_stock. stock is allocated per-CPU for every non-root memcg, which is 16 bytes per-cpu per-memcg and allocated with GFP_KERNEL_ACCOUNT. The !allow_spinning special case in try_charge_memcg goes away. It clamped batch to nr_pages so a charge would not have to refill the stock and potentially do an expensive flush. We now refill with a trylock and retry the exact size if the greedy charge fails, so a non-spinning caller is never made to wait or fail early. Also, while we no longer have a 7-memcg cap, it also means that each CPU can now cache an unbounded number of pages. System-wide, pre-charged but unused memory goes from NR_MEMCG_STOCK * batch * ncpus to nr_memcgs * batch * ncpus. We can also now drain stock on isolated CPUs as well, since draining is no longer a per-cpu local operation. This leaves one user-visible change for legacy cgroup v1 (memsw). Because the stock becomes private to the memory page_counter rather than shared by memory and memsw, memsw.usage - memory.usage no longer equals swap usage. Userspace programs using the legacy cgroup and deriving swap usage using this method can underflow. The next patch gives memsw its own stock, narrowing this to a transient drift. With all of these changes made, also remove all newly unused memcg code. obj_stock is untouched and is still needed. FLUSHING_CACHED_CHARGE and the memcg_wq are preserved so obj_stock can use them as well. Suggested-by: Johannes Weiner Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 285 +++++------------------------------------------- 1 file changed, 30 insertions(+), 255 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 71410084fa7fc..5678486cc55b0 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2027,35 +2027,7 @@ void mem_cgroup_print_oom_group(struct mem_cgroup *m= emcg) pr_cont(" are going to be killed due to memory.oom.group set\n"); } =20 -/* - * The value of NR_MEMCG_STOCK is selected to keep the cached memcgs and t= heir - * nr_pages in a single cacheline. This may change in future. - */ -#define NR_MEMCG_STOCK 7 - -/* - * Watermarks for a charge stock slot, in the spirit of pcp->high and - * pcp->batch: MEMCG_STOCK_HIGH is the high watermark at which a slot is - * trimmed, and it is trimmed down to MEMCG_STOCK_LOW rather than emptied. - */ -#define MEMCG_STOCK_LOW (MEMCG_CHARGE_BATCH / 2) -#define MEMCG_STOCK_HIGH (MEMCG_CHARGE_BATCH) - #define FLUSHING_CACHED_CHARGE 0 -struct memcg_stock_pcp { - local_trylock_t lock; - uint8_t nr_pages[NR_MEMCG_STOCK]; - struct mem_cgroup *cached[NR_MEMCG_STOCK]; - - struct work_struct work; - unsigned long flags; - uint8_t drain_idx; -}; - -static DEFINE_PER_CPU_ALIGNED(struct memcg_stock_pcp, memcg_stock) =3D { - .lock =3D INIT_LOCAL_TRYLOCK(lock), -}; - /* * NR_OBJ_STOCK is sized so the entire hot path of obj_stock_pcp * (lock, accounting metadata, nr_bytes[] and cached[]) fits within a @@ -2103,52 +2075,6 @@ static void drain_obj_stock(struct obj_stock_pcp *st= ock); static bool obj_stock_flush_required(struct obj_stock_pcp *stock, struct mem_cgroup *root_memcg); =20 -/** - * consume_stock: Try to consume stocked charge on this cpu. - * @memcg: memcg to consume from. - * @nr_pages: how many pages to charge. - * - * Consume the cached charge if enough nr_pages are present otherwise retu= rn - * failure. Also return failure for charge request larger than - * MEMCG_CHARGE_BATCH or if the local lock is already taken. - * - * returns true if successful, false otherwise. - */ -static bool consume_stock(struct mem_cgroup *memcg, unsigned int nr_pages) -{ - struct memcg_stock_pcp *stock; - uint8_t stock_pages; - bool ret =3D false; - int i; - - if (nr_pages > MEMCG_CHARGE_BATCH || - !local_trylock(&memcg_stock.lock)) - return ret; - - stock =3D this_cpu_ptr(&memcg_stock); - - for (i =3D 0; i < NR_MEMCG_STOCK; ++i) { - if (memcg !=3D READ_ONCE(stock->cached[i])) - continue; - - stock_pages =3D READ_ONCE(stock->nr_pages[i]); - if (stock_pages >=3D nr_pages) { - stock_pages -=3D nr_pages; - WRITE_ONCE(stock->nr_pages[i], stock_pages); - if (!stock_pages) { - css_put(&memcg->css); - WRITE_ONCE(stock->cached[i], NULL); - } - ret =3D true; - } - break; - } - - local_unlock(&memcg_stock.lock); - - return ret; -} - static void memcg_uncharge(struct mem_cgroup *memcg, unsigned int nr_pages) { page_counter_uncharge(&memcg->memory, nr_pages); @@ -2156,51 +2082,6 @@ static void memcg_uncharge(struct mem_cgroup *memcg,= unsigned int nr_pages) page_counter_uncharge(&memcg->memsw, nr_pages); } =20 -/* - * Returns stocks cached in percpu and reset cached information. - */ -static void drain_stock(struct memcg_stock_pcp *stock, int i) -{ - struct mem_cgroup *old =3D READ_ONCE(stock->cached[i]); - uint8_t stock_pages; - - if (!old) - return; - - stock_pages =3D READ_ONCE(stock->nr_pages[i]); - if (stock_pages) { - memcg_uncharge(old, stock_pages); - WRITE_ONCE(stock->nr_pages[i], 0); - } - - css_put(&old->css); - WRITE_ONCE(stock->cached[i], NULL); -} - -static void drain_stock_fully(struct memcg_stock_pcp *stock) -{ - int i; - - for (i =3D 0; i < NR_MEMCG_STOCK; ++i) - drain_stock(stock, i); -} - -static void drain_local_memcg_stock(struct work_struct *dummy) -{ - struct memcg_stock_pcp *stock; - - if (WARN_ONCE(!in_task(), "drain in non-task context")) - return; - - local_lock(&memcg_stock.lock); - - stock =3D this_cpu_ptr(&memcg_stock); - drain_stock_fully(stock); - clear_bit(FLUSHING_CACHED_CHARGE, &stock->flags); - - local_unlock(&memcg_stock.lock); -} - static void drain_local_obj_stock(struct work_struct *dummy) { struct obj_stock_pcp *stock; @@ -2217,92 +2098,6 @@ static void drain_local_obj_stock(struct work_struct= *dummy) local_unlock(&obj_stock.lock); } =20 -static void refill_stock(struct mem_cgroup *memcg, unsigned int nr_pages) -{ - struct memcg_stock_pcp *stock; - struct mem_cgroup *cached; - unsigned int stock_pages; - bool success =3D false; - int empty_slot =3D -1; - int i; - - /* - * nr_pages[] is a uint8_t and a slot's count is capped at - * MEMCG_STOCK_HIGH. Raising MEMCG_CHARGE_BATCH beyond 127 would need - * more careful handling of nr_pages[] in struct memcg_stock_pcp. - */ - BUILD_BUG_ON(MEMCG_CHARGE_BATCH > S8_MAX); - BUILD_BUG_ON(MEMCG_STOCK_HIGH > U8_MAX); - - VM_WARN_ON_ONCE(mem_cgroup_is_root(memcg)); - - if (nr_pages > MEMCG_CHARGE_BATCH || - !local_trylock(&memcg_stock.lock)) { - /* - * In case of larger than batch refill or unlikely failure to - * lock the percpu memcg_stock.lock, uncharge memcg directly. - */ - memcg_uncharge(memcg, nr_pages); - return; - } - - stock =3D this_cpu_ptr(&memcg_stock); - for (i =3D 0; i < NR_MEMCG_STOCK; ++i) { - cached =3D READ_ONCE(stock->cached[i]); - if (!cached && empty_slot =3D=3D -1) - empty_slot =3D i; - if (memcg =3D=3D READ_ONCE(stock->cached[i])) { - stock_pages =3D READ_ONCE(stock->nr_pages[i]) + nr_pages; - if (stock_pages > MEMCG_STOCK_HIGH) { - memcg_uncharge(memcg, - stock_pages - MEMCG_STOCK_LOW); - stock_pages =3D MEMCG_STOCK_LOW; - } - WRITE_ONCE(stock->nr_pages[i], stock_pages); - success =3D true; - break; - } - } - - if (!success) { - i =3D empty_slot; - if (i =3D=3D -1) { - i =3D stock->drain_idx++; - if (stock->drain_idx =3D=3D NR_MEMCG_STOCK) - stock->drain_idx =3D 0; - drain_stock(stock, i); - } - css_get(&memcg->css); - WRITE_ONCE(stock->cached[i], memcg); - WRITE_ONCE(stock->nr_pages[i], nr_pages); - } - - local_unlock(&memcg_stock.lock); -} - -static bool is_memcg_drain_needed(struct memcg_stock_pcp *stock, - struct mem_cgroup *root_memcg) -{ - struct mem_cgroup *memcg; - bool flush =3D false; - int i; - - rcu_read_lock(); - for (i =3D 0; i < NR_MEMCG_STOCK; ++i) { - memcg =3D READ_ONCE(stock->cached[i]); - if (!memcg) - continue; - - if (READ_ONCE(stock->nr_pages[i]) && - mem_cgroup_is_descendant(memcg, root_memcg)) { - flush =3D true; - break; - } - } - rcu_read_unlock(); - return flush; -} - static bool schedule_drain_work(int cpu, struct work_struct *work) { /* @@ -2325,34 +2120,22 @@ static bool schedule_drain_work(int cpu, struct wor= k_struct *work) */ void drain_all_stock(struct mem_cgroup *root_memcg) { + struct mem_cgroup *memcg; int cpu, curcpu; =20 /* If someone's already draining, avoid adding running more workers. */ if (!mutex_trylock(&percpu_charge_mutex)) return; - /* - * Notify other cpus that system-wide "drain" is running - * We do not care about races with the cpu hotplug because cpu down - * as well as workers from this path always operate on the local - * per-cpu data. CPU up doesn't touch memcg_stock at all. - */ + + for_each_mem_cgroup_tree(memcg, root_memcg) + page_counter_drain_stock_async(&memcg->memory); + + /* Hotplug races are OK; workers only touch their own cpu's obj_stock */ migrate_disable(); curcpu =3D smp_processor_id(); for_each_online_cpu(cpu) { - struct memcg_stock_pcp *memcg_st =3D &per_cpu(memcg_stock, cpu); struct obj_stock_pcp *obj_st =3D &per_cpu(obj_stock, cpu); =20 - if (!test_bit(FLUSHING_CACHED_CHARGE, &memcg_st->flags) && - is_memcg_drain_needed(memcg_st, root_memcg) && - !test_and_set_bit(FLUSHING_CACHED_CHARGE, - &memcg_st->flags)) { - if (cpu =3D=3D curcpu) - drain_local_memcg_stock(&memcg_st->work); - else if (!schedule_drain_work(cpu, &memcg_st->work)) - clear_bit(FLUSHING_CACHED_CHARGE, - &memcg_st->flags); - } - if (!test_bit(FLUSHING_CACHED_CHARGE, &obj_st->flags) && obj_stock_flush_required(obj_st, root_memcg) && !test_and_set_bit(FLUSHING_CACHED_CHARGE, @@ -2370,20 +2153,21 @@ void drain_all_stock(struct mem_cgroup *root_memcg) =20 static int memcg_hotplug_cpu_dead(unsigned int cpu) { - struct memcg_stock_pcp *memcg_st =3D &per_cpu(memcg_stock, cpu); + struct mem_cgroup *memcg; struct obj_stock_pcp *obj_st =3D &per_cpu(obj_stock, cpu); =20 /* no need for the local lock */ drain_obj_stock(obj_st); - drain_stock_fully(memcg_st); + + for_each_mem_cgroup_tree(memcg, NULL) + page_counter_drain_cpu_stock(&memcg->memory, cpu); =20 /* * A drain work queued before the CPU went away is executed by an * unbound worker on some other CPU and clears that CPU's flag, so - * clear the flags here to make these stocks drainable again once - * the CPU comes back online. + * clear the flag here to make this stock drainable again once the CPU + * comes back online. */ - clear_bit(FLUSHING_CACHED_CHARGE, &memcg_st->flags); clear_bit(FLUSHING_CACHED_CHARGE, &obj_st->flags); =20 return 0; @@ -2658,7 +2442,6 @@ void __mem_cgroup_handle_over_high(gfp_t gfp_mask) static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask, unsigned int nr_pages) { - unsigned int batch =3D max(MEMCG_CHARGE_BATCH, nr_pages); int nr_retries =3D MAX_RECLAIM_RETRIES; struct mem_cgroup *mem_over_limit; struct page_counter *counter; @@ -2672,35 +2455,22 @@ static int try_charge_memcg(struct mem_cgroup *memc= g, gfp_t gfp_mask, bool allow_spinning =3D gfpflags_allow_spinning(gfp_mask); =20 retry: - if (consume_stock(memcg, nr_pages)) - return 0; - - if (!allow_spinning) - /* Avoid the refill and flush of the older stock */ - batch =3D nr_pages; - reclaim_options =3D MEMCG_RECLAIM_MAY_SWAP; if (do_memsw_account() && - !page_counter_try_charge(&memcg->memsw, batch, &counter, NULL)) { + !page_counter_try_charge(&memcg->memsw, nr_pages, &counter, NULL)) { mem_over_limit =3D mem_cgroup_from_counter(counter, memsw); reclaim_options &=3D ~MEMCG_RECLAIM_MAY_SWAP; goto reclaim; } =20 - if (page_counter_try_charge(&memcg->memory, batch, &counter, - &nr_charged)) - goto done_restock; + if (page_counter_try_charge(&memcg->memory, nr_pages, &counter, &nr_charg= ed)) + goto check_high; =20 if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, batch); + page_counter_uncharge(&memcg->memsw, nr_pages); mem_over_limit =3D mem_cgroup_from_counter(counter, memory); =20 reclaim: - if (batch > nr_pages) { - batch =3D nr_pages; - goto retry; - } - /* * Prevent unbounded recursion when reclaim operations need to * allocate memory. This might exceed the limits temporarily, @@ -2809,10 +2579,9 @@ static int try_charge_memcg(struct mem_cgroup *memcg= , gfp_t gfp_mask, =20 return 0; =20 -done_restock: - if (batch > nr_pages) - refill_stock(memcg, batch - nr_pages); - +check_high: + if (!nr_charged) + return 0; /* * If the hierarchy is above the normal consumption range, schedule * reclaim on returning to userland. We can perform reclaim here @@ -3154,8 +2923,11 @@ static void obj_cgroup_uncharge_pages(struct obj_cgr= oup *objcg, =20 account_kmem_nmi_safe(memcg, -nr_pages); memcg1_account_kmem(memcg, -nr_pages); - if (!mem_cgroup_is_root(memcg)) - refill_stock(memcg, nr_pages); + if (!mem_cgroup_is_root(memcg)) { + page_counter_refill_stock(&memcg->memory, nr_pages); + if (do_memsw_account()) + page_counter_uncharge(&memcg->memsw, nr_pages); + } =20 css_put(&memcg->css); } @@ -4157,6 +3929,7 @@ static void __mem_cgroup_free(struct mem_cgroup *memc= g) =20 static void mem_cgroup_free(struct mem_cgroup *memcg) { + page_counter_free_stock(&memcg->memory); lru_gen_exit_memcg(memcg); memcg_wb_domain_exit(memcg); __mem_cgroup_free(memcg); @@ -4324,6 +4097,10 @@ static int mem_cgroup_css_online(struct cgroup_subsy= s_state *css) refcount_set(&memcg->id.ref, 1); css_get(css); =20 + /* stock allocation failure is nonfatal; fall back to direct charges */ + if (!mem_cgroup_is_root(memcg)) + page_counter_alloc_stock(&memcg->memory, MEMCG_CHARGE_BATCH); + /* * Ensure mem_cgroup_from_private_id() works once we're fully online. * @@ -5665,7 +5442,7 @@ void mem_cgroup_sk_uncharge(const struct sock *sk, un= signed int nr_pages) =20 mod_memcg_state(memcg, MEMCG_SOCK, -nr_pages); =20 - refill_stock(memcg, nr_pages); + page_counter_refill_stock(&memcg->memory, nr_pages); } =20 void mem_cgroup_flush_workqueue(void) @@ -5719,8 +5496,6 @@ int __init mem_cgroup_init(void) WARN_ON(!memcg_wq); =20 for_each_possible_cpu(cpu) { - INIT_WORK(&per_cpu_ptr(&memcg_stock, cpu)->work, - drain_local_memcg_stock); INIT_WORK(&per_cpu_ptr(&obj_stock, cpu)->work, drain_local_obj_stock); } --=20 2.53.0-Meta From nobody Sat Sep 26 14:39:01 2026 Received: from mail-oo1-f46.google.com (mail-oo1-f46.google.com [209.85.161.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47F264ADD90 for ; Mon, 31 Aug 2026 16:38:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194285; cv=none; b=oItRTwp4fUOEoLICXE8Rac9SOf8cpuGKCsKEWMth0L8pHoIpyUup622cgpR6GbTKcP1HWXB7/OreIhBsWIIzd6NzfgINq+pOFeB5BknmYczvMJEVyHHSV3baaQq4wzZ/oupZQIVGi+CTBgkG4b0PEPwsbDfGeuPAJRI4iT5yTE8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788194285; c=relaxed/simple; bh=K89uTa5EpQ9kFfWWW7y8lAp+J3EDDR87cMYtPD5wSV4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=l4C3J+clb9Xb62j/OStZAIssFaC0obE26wpraoqOeUr0drSZu4DRX64OFaYLENl/NpOh0P0rX8dIgwp4T1m4mUyYXGPYqY1+06Ld4mWi/J+o8Iv1NKVKl/S5ppD5hLQwscO43VJ2cPayhNiIRyi9pM/SFCrKtDLF1I/deYKQU1g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=DBhemhuv; arc=none smtp.client-ip=209.85.161.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="DBhemhuv" Received: by mail-oo1-f46.google.com with SMTP id 006d021491bc7-6b19e291cc6so1512446eaf.2 for ; Mon, 31 Aug 2026 09:38:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788194282; x=1788799082; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=KCEcxBXECIK58MqzfUHKxW4lG6L2Egx8+PoUClZnYOs=; b=DBhemhuvtEj1fp2lqdOnDfIWwtM6Jr80nvYYO85lqw5BbYm9aBFkR0esPGOtiMQU4x Z0+e8f7kAlqLejMdrIcQp/9qUQb/Pp/spA+z9l5/FVn8Z1CbyQvI5Yml+t2B6zXz7SpM UTfj/xwdzlGhPGxyLgsS/03r6AoUoxEDKuXN9JJLFbix/fwRvGl8FxkpS6IHCyLbkq+b YNvM9zXm6VpM2XZTMV8gsMPR1d3JHY7J1hNxzABnMS/VxTzJs+bijhmCTSvKpa6F0Gqc 13h9IIYJsuPPSvsP3/WHNwYPsbRofeAYgN5749BbjBJBrZXuxbVS9NxhRLpCfmQ8o8PF FVpQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788194282; x=1788799082; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=KCEcxBXECIK58MqzfUHKxW4lG6L2Egx8+PoUClZnYOs=; b=Xl+wYEbR7Ayk2+o9W2l2Pe5bAWbLvCe7x0aOgw3bTXR1CvhuhvG8tPrsnnA0RYKcYg jO49jHb4l9UAIaS3pAamNg/nJvEXcDRLmmB3mmOmagDdMuenKnl+svTj3novnOOGsv/l Q6pA/trYiKOujLDbv9SsTeGUve7KqOUPZhOTH49dBLVc9MKsPUvxFsIc0wd2xd3WxA+E SJrQlZt6FVO/teGYyM7aTEvix1OY1Cyhnc75ZAJ1InjONa9nm08fML6felRkEFiU06rX xazP4Y6KtlTpV9PSGh8Kmgjt3f4MTggxDSnX7XBcEbOBHdn3IM7TORqOJwNLp3JtaXwm vadw== X-Forwarded-Encrypted: i=1; AHgh+RqfDtwU9Kw9Dl5UiLx0LM6aPHx8Le+9xr2w41umS7mDjPSjeXpLlAIPYd4cRBb7amPHTdjWSQBCvLqYjdk=@vger.kernel.org X-Gm-Message-State: AFuF++mmCPWddAQO7zRyHhLIZj++Iuqpw3Z/M3hA6pVQ2rmzOlLerSkf oSHdGJPuPmWjpuu32QvgaS2dahGRvr5EwLwrUMsrPuXj8CKE1GniKYCL X-Gm-Gg: AR+sD11N2JaoH0NuZxqP1qD8QxoSD6gQcmXfPO7zcW8gSztfqLpEFHpla0IxU0DNr9N zQtynjYt1Sh1jek7zEqWPcwZb65UcV/6ELMJoEQ678Mw8HM07zSuUVKuiLwtLbt69QnTLKJu47K 0NpB82K1/RkVtV8Wua+q7JWMv1rvK5MtjM/9RDRpjbbNr96P9PelBXtO/VzIQIlmtsebPV8f7/E WNgMcz8WXl/+OA6SjcLoV0hR3Gdgqt8VnumdeRQR8cgWvtrlR9+jU4nnG0CM2qrs7Phdxt1HUvA UpC5T4+5Y8jDtYHliTnEr+BE63Yhs4VqAeaId/tZmiwFL0wKf4WAbxrQ/7A2Mnbpc1u0yqOvz/d tileK5lyaiTF+83ZMSll5yVak6LVGc9JJX1aG0KKQCyfRGXpLMdaA1PLbho28+gwog5t5Nmh+bR Gmt1ra3fBmMd8JwHAWZFOXg74rR/V0meqX88XPNIM+qMUtq2Tzh4qRdTjUKeEkjjEF3FC365/tZ Ohno3L0RTrmXhWPZg== X-Received: by 2002:a05:6820:1791:b0:6b1:d18b:4048 with SMTP id 006d021491bc7-6b372d21358mr2106016eaf.10.1788194281822; Mon, 31 Aug 2026 09:38:01 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:e::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7f4fa7e97d5sm8453832a34.10.2026.08.31.09.38.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 09:38:01 -0700 (PDT) From: Joshua Hahn To: hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org Cc: roman.gushchin@linux.dev, muchun.song@linux.dev, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, dev@lankhorst.se, mripard@kernel.org, nat@pixelcluster.dev, tj@kernel.org, mkoutny@suse.com, osalvador@suse.de, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org, kernel-team@meta.com Subject: [PATCH v5 7/7] mm/memcontrol: add stock to the memsw page_counter Date: Mon, 31 Aug 2026 09:37:51 -0700 Message-ID: <20260831163752.2193337-8-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> References: <20260831163752.2193337-1-joshua.hahnjy@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Before this series, each memcg had one stock shared by all its page_counters (memory + memsw). Now that the memcg stock was folded into the page_counter level, give memsw its own page_counter_stock so that it can benefit from caching charges as well. Note that while the allocation is conditional on do_memsw_account(), the freeing is not; the freer will only free non-NULL stocks. This matters because do_memsw_account() could have changed in between the allocation and the free. This narrows the memsw skew introduced by the previous patch. memsw is now charged in the same batches as memory, so the two no longer diverge systematically, but can still see transient drifts since each keeps its own per-cpu stock. The drift is bound by MEMCG_CHARGE_BATCH * nr_possible_cpus. In the unlikely scenario that one of the stocks does not get allocated, there will be different granularities of charging for the cgroup's lifetime (one charging in batch granularity, the other just in nr_pages). Signed-off-by: Joshua Hahn --- mm/memcontrol.c | 20 +++++++++++++++----- 1 file changed, 15 insertions(+), 5 deletions(-) diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 5678486cc55b0..33e4ffbd48a8b 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -2127,8 +2127,10 @@ void drain_all_stock(struct mem_cgroup *root_memcg) if (!mutex_trylock(&percpu_charge_mutex)) return; =20 - for_each_mem_cgroup_tree(memcg, root_memcg) + for_each_mem_cgroup_tree(memcg, root_memcg) { page_counter_drain_stock_async(&memcg->memory); + page_counter_drain_stock_async(&memcg->memsw); + } =20 /* Hotplug races are OK; workers only touch their own cpu's obj_stock */ migrate_disable(); @@ -2159,8 +2161,10 @@ static int memcg_hotplug_cpu_dead(unsigned int cpu) /* no need for the local lock */ drain_obj_stock(obj_st); =20 - for_each_mem_cgroup_tree(memcg, NULL) + for_each_mem_cgroup_tree(memcg, NULL) { page_counter_drain_cpu_stock(&memcg->memory, cpu); + page_counter_drain_cpu_stock(&memcg->memsw, cpu); + } =20 /* * A drain work queued before the CPU went away is executed by an @@ -2467,7 +2471,7 @@ static int try_charge_memcg(struct mem_cgroup *memcg,= gfp_t gfp_mask, goto check_high; =20 if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, nr_pages); + page_counter_refill_stock(&memcg->memsw, nr_pages); mem_over_limit =3D mem_cgroup_from_counter(counter, memory); =20 reclaim: @@ -2926,7 +2930,7 @@ static void obj_cgroup_uncharge_pages(struct obj_cgro= up *objcg, if (!mem_cgroup_is_root(memcg)) { page_counter_refill_stock(&memcg->memory, nr_pages); if (do_memsw_account()) - page_counter_uncharge(&memcg->memsw, nr_pages); + page_counter_refill_stock(&memcg->memsw, nr_pages); } =20 css_put(&memcg->css); @@ -3930,6 +3934,8 @@ static void __mem_cgroup_free(struct mem_cgroup *memc= g) static void mem_cgroup_free(struct mem_cgroup *memcg) { page_counter_free_stock(&memcg->memory); + /* memsw and swap are the same counter; only memsw is ever stocked */ + page_counter_free_stock(&memcg->memsw); lru_gen_exit_memcg(memcg); memcg_wb_domain_exit(memcg); __mem_cgroup_free(memcg); @@ -4098,8 +4104,12 @@ static int mem_cgroup_css_online(struct cgroup_subsy= s_state *css) css_get(css); =20 /* stock allocation failure is nonfatal; fall back to direct charges */ - if (!mem_cgroup_is_root(memcg)) + if (!mem_cgroup_is_root(memcg)) { page_counter_alloc_stock(&memcg->memory, MEMCG_CHARGE_BATCH); + if (do_memsw_account()) + page_counter_alloc_stock(&memcg->memsw, + MEMCG_CHARGE_BATCH); + } =20 /* * Ensure mem_cgroup_from_private_id() works once we're fully online. --=20 2.53.0-Meta