From nobody Sun Sep 27 03:46:19 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 11A7627AC57 for ; Fri, 18 Sep 2026 00:12:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789690341; cv=none; b=bwhQKNKEs+lSdqJ5YSX7evaxFQGNAroroeUGIlIMNoVvpIx6xWAPGAqrAcypZJ0U0/y6L8zlHWOZ6lh6k9EPu5MHqaLri7mVkoCgPQXkmUthcS4NrDkFDXQ+rFLFaCAAr68mC7UMsTEb7EMJ8+8GZpxRVeQbWGLPjZaG4vc0a/8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789690341; c=relaxed/simple; bh=MiigpDHTrS/YH+wsTyUGR8pYU9kRznEBA5jNJoFFlHQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=IAZz0xn6uYfdJIcJjfFAWXlxZFmdB53U2ArH8PdzwLI9WwXeo1FOGJKDtQlsoTemjEoGKBK4vdhZI6gcdQWb8HbdJcl/jggzUhzv2A1++CUfm3i2qjlnFBJHOE9I7SWkpJ5EKAQMRP0wlw71n8O5CACYTvtWZDRVyJ1J/Z7NLQw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=UrepoYds; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="UrepoYds" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-93a0fd41a7bso20367085a.1 for ; Thu, 17 Sep 2026 17:12:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789690334; x=1790295134; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ReC5OTam6MTJ6u4FECt/nJW9XaNR1Mg3esV3ZyKWZpE=; b=UrepoYdsfS2OILDDIOzCRdtk96VJLCe44aLqBs7EhEJm3hqF46Cu99XssurIzIqLqo y/4p4XkA2cjepzHSyS5ImiJXehWrSUgA2lcvG2bCCdirRuO01ZZ3NgrS2i2exGis5ASH oV7xlr745TZz7WJhcCD39XmUEWK4x+MiQspJ+OllS0giY9zIiIZ2sOoD4N8QGSD1A9j2 SAuJlE6kkaKMagCicLbpBVdss0kl1fWWQGFprbW41LWFXDKJcAX2MhCDsf2RYdqKvbae Q1eEoqcdEkes8YExSFobfMMNZaN431Kr0SrZyfr69E4u418WZ9/jVvPLUJpZwx4mtuij /gUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789690334; x=1790295134; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ReC5OTam6MTJ6u4FECt/nJW9XaNR1Mg3esV3ZyKWZpE=; b=a1uy4YbwUg2bVXATT1AybIdrvLYQvSfn/r/WJxz3DnyE7shtVc6JLIJ6CLQU7SYeDH FKDJgjEJgLDgEQk8zi7gptbZ98NXV10f0ixcjmcuEtwR7mhRM3ut+l6+ttO3ki2ipsel Upa3Bctnc1hOBCv1dUOvAwUCzL5UWA58owjGyRjTy1+Mdead04s1c9o/nSeJuDkPphmU 183LHrBDhOOra33j7Siq0SeYh2DLYDUh3QhpaCDjzaGXTBEs8dTTQ2GEZAkax3tsjQvE mXL6HZ+Y81DRy169p0oAfDtTRsqE04FlEejQB0Q/ZMejv2iK4R/LFgR9R95CEJscJ3fm zkbQ== X-Gm-Message-State: AFuF++moHyhLJ1X0yFOmpAWW7Anl2YTOboO5vL/e1P2qKy1PcUrltgMe L1nquHlGVrPyMgdclqaeT9XCtskSysW0HAhKlaKzRChCQdI7bO2qYE/iw8N3Yf7ICpE= X-Gm-Gg: AYBFou1lDLJsCiDunTj8c6U1aB/iDXpSjVE2dAAiMGvmR79oWh3hdI0c8TBx9dN5uPp vWEty0NALOCJv8+/s4znNaD/qTPHdt3w8aQfusm3pTZ6UVYAdJITMNPuDfEgXLdJJeUMIbmIxqV ClefCOcXsjUzniwyrlezvd4cFp+OUApcdOw96lWHEzg3YTQX7wu3gpP6lqKkonGY/tvRbt1xTFn sWbabuGu9qChcSF2PTg+R8QS3vR5Ir2Su3q1NnXGyHDkVI0NzAeFA3663ysnQotNwbjQMS++0b5 ZD3Tax98+lf1Uzt3xURjWtKRfl3cHPjenYnMcka1MJsdsuisD/QpGS1iN7IQgLc2Cm25k/Lhxe7 AKrtFfpVAc9geMtHzIytZTZ2NlUmmJRKq6UQks4SP51XzT8ZK+9BfcA8wDYYHDnoEKGo8qzx4jm XqAICniVUHsFjszhdyU719JkqZkghcityu2UJd4+Qsh5nallshZANEfxN1NHibacyCenjarF//a SbhVbqUtulfocARTXy33OF8RfivN5yuck9cY5rbKxobTS6/lmrcRO0ACe9X X-Received: by 2002:a05:620a:31a0:b0:93b:d7a0:d9e4 with SMTP id af79cd13be357-93bdca1640cmr106094985a.62.1789690334389; Thu, 17 Sep 2026 17:12:14 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93b81cff7cbsm583762585a.41.2026.09.17.17.12.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 17 Sep 2026 17:12:14 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, Matthew Wilcox Subject: [PATCH v2 1/2] mm/mempolicy: use SRCU for the weighted interleave state Date: Thu, 17 Sep 2026 20:12:02 -0400 Message-ID: <20260918001203.3389165-2-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260918001203.3389165-1-gourry@gourry.net> References: <20260918001203.3389165-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" alloc_pages_bulk_weighted_interleave() copies iw_table into a scratch array on every call so it can walk the weights outside of RCU. The copy exists only because the loop may sleep in the page allocator and so cannot hold rcu_read_lock(). Use SRCU to pin the global iw_table object and use it in-place instead. Retire through both flavors - call_srcu() for the sleeping readers, then kfree_rcu() for the reference-less ones - so writers no longer block on synchronize_rcu() either. Tested in a VM with KASAN, PROVE_LOCKING and DEBUG_OBJECTS_RCU_HEAD, with a udelay() injected into the read section to widen the race against concurrent sysfs weight writers, and placement checked against the configured weights. Every retired state reached its callback. Swapping the deferred free for a bare kfree() in the same test reports a use-after-free immediately. Suggested-by: Andrew Morton Suggested-by: Matthew Wilcox Assisted-by: LLM Signed-off-by: Gregory Price (Meta) Acked-by: David Hildenbrand (Arm) --- mm/mempolicy.c | 70 ++++++++++++++++++++++++-------------------------- 1 file changed, 34 insertions(+), 36 deletions(-) diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 060a0eb26917..2643915dc966 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -112,6 +112,7 @@ #include #include #include +#include =20 #include #include @@ -157,6 +158,7 @@ static const int weightiness =3D 32; */ struct weighted_interleave_state { bool mode_auto; + struct rcu_head rcu; u8 iw_table[]; }; static struct weighted_interleave_state __rcu *wi_state; @@ -168,6 +170,24 @@ static unsigned int *node_bw_table; */ static DEFINE_MUTEX(wi_state_lock); =20 +/* Readers that sleep while walking iw_table hold this instead */ +DEFINE_STATIC_SRCU_FAST(wi_srcu); + +static void wi_state_free_rcu(struct rcu_head *head) +{ + struct weighted_interleave_state *state =3D + container_of(head, struct weighted_interleave_state, rcu); + + kfree_rcu(state, rcu); +} + +/* Retire through both flavors: sleeping readers use SRCU, the rest RCU */ +static void wi_state_retire(struct weighted_interleave_state *state) +{ + if (state) + call_srcu(&wi_srcu, &state->rcu, wi_state_free_rcu); +} + static u8 get_il_weight(int node) { struct weighted_interleave_state *state; @@ -266,10 +286,7 @@ int mempolicy_set_node_perf(unsigned int node, struct = access_coordinate *coords) rcu_assign_pointer(wi_state, new_wi_state); =20 mutex_unlock(&wi_state_lock); - if (old_wi_state) { - synchronize_rcu(); - kfree(old_wi_state); - } + wi_state_retire(old_wi_state); out: kfree(old_bw); return 0; @@ -2644,7 +2661,8 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, unsigned long nr_allocated =3D 0; unsigned long rounds; unsigned long node_pages, delta; - u8 *weights, weight; + struct srcu_ctr __percpu *scp; + u8 *table, weight; unsigned int weight_total =3D 0; unsigned long rem_pages =3D nr_pages; nodemask_t nodes; @@ -2688,25 +2706,14 @@ static unsigned long alloc_pages_bulk_weighted_inte= rleave(gfp_t gfp, me->il_weight =3D 0; prev_node =3D node; =20 - /* create a local copy of node weights to operate on outside rcu */ - weights =3D kmalloc(nr_node_ids, gfp & GFP_RECLAIM_MASK); - if (!weights) - return total_allocated; - - rcu_read_lock(); - state =3D rcu_dereference(wi_state); - if (state) { - memcpy(weights, state->iw_table, nr_node_ids * sizeof(u8)); - rcu_read_unlock(); - } else { - rcu_read_unlock(); - for (i =3D 0; i < nr_node_ids; i++) - weights[i] =3D 1; - } + /* The page allocator may sleep, pin the weight table with SRCU */ + scp =3D srcu_read_lock_fast(&wi_srcu); + state =3D srcu_dereference(wi_state, &wi_srcu); + table =3D state ? state->iw_table : NULL; =20 /* calculate total, detect system default usage */ for_each_node_mask(node, nodes) - weight_total +=3D weights[node]; + weight_total +=3D table ? table[node] : 1; =20 /* * Calculate rounds/partial rounds to minimize __alloc_pages_bulk calls. @@ -2718,10 +2725,10 @@ static unsigned long alloc_pages_bulk_weighted_inte= rleave(gfp_t gfp, rounds =3D rem_pages / weight_total; delta =3D rem_pages % weight_total; resume_node =3D next_node_in(prev_node, nodes); - resume_weight =3D weights[resume_node]; + resume_weight =3D table ? table[resume_node] : 1; for (i =3D 0; i < nnodes; i++) { node =3D next_node_in(prev_node, nodes); - weight =3D weights[node]; + weight =3D table ? table[node] : 1; node_pages =3D weight * rounds; /* If a delta exists, add this node's portion of the delta */ if (delta > weight) { @@ -2747,7 +2754,7 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, } me->il_prev =3D resume_node; me->il_weight =3D resume_weight; - kfree(weights); + srcu_read_unlock_fast(&wi_srcu, scp); return total_allocated; } =20 @@ -3673,10 +3680,7 @@ static ssize_t node_store(struct kobject *kobj, stru= ct kobj_attribute *attr, =20 rcu_assign_pointer(wi_state, new_wi_state); mutex_unlock(&wi_state_lock); - if (old_wi_state) { - synchronize_rcu(); - kfree(old_wi_state); - } + wi_state_retire(old_wi_state); return count; } =20 @@ -3742,10 +3746,7 @@ static ssize_t weighted_interleave_auto_store(struct= kobject *kobj, update_wi_state: rcu_assign_pointer(wi_state, new_wi_state); mutex_unlock(&wi_state_lock); - if (old_wi_state) { - synchronize_rcu(); - kfree(old_wi_state); - } + wi_state_retire(old_wi_state); return count; } =20 @@ -3789,10 +3790,7 @@ static void wi_state_free(void) rcu_assign_pointer(wi_state, NULL); mutex_unlock(&wi_state_lock); =20 - if (old_wi_state) { - synchronize_rcu(); - kfree(old_wi_state); - } + wi_state_retire(old_wi_state); } =20 static struct kobj_attribute wi_auto_attr =3D { --=20 2.55.0 From nobody Sun Sep 27 03:46:19 2026 Received: from mail-qk1-f172.google.com (mail-qk1-f172.google.com [209.85.222.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ECE7C22576E for ; Fri, 18 Sep 2026 00:12:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789690344; cv=none; b=iMJcanvFkJjCuapP/IOp/r+e5itMOP1zmaXpMz9c5KfWaKAHbQZebXfx5TviRLFhGc3rrvYjq1iSKWJEgdxjMhpaXjhfNCT4rDDhj85SfoNwSCMHrsRKn+IxdYXtEOO0w4XdLTvgfe4OrZRzXpt3TWKhg5Jd6tKXHyckO/qOUvQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789690344; c=relaxed/simple; bh=D5qZupghknuQR+vMpyE0sB6nwdn+i5Gj/gtRJz5IWpw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pE2NGhSzxeRL/r138E0rG+fE0g2rpjZR4BmiZljZ4bYoFTpx2UepGDn1dYgk3vWwuCv3EDLWad61EYd1HzH2xKlNyomkyyXPZQZds52jiAYYsIr2lxQta72e+vLD58UisYTifNO6H8Bpaq815M/uoC5dPuO6SZyFOQZdfOliNug= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=ORKjGM+7; arc=none smtp.client-ip=209.85.222.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="ORKjGM+7" Received: by mail-qk1-f172.google.com with SMTP id af79cd13be357-939f5f82829so118138685a.1 for ; Thu, 17 Sep 2026 17:12:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789690336; x=1790295136; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PrpI8WWxUlqmCCZ/KlwzqkhVdxs2Qx8EkwLTiy0oOLg=; b=ORKjGM+72PHVL7J+pJgzVORT2MAuseO6HMQqHBosiXE0I3I67C++wpU3pU/DFluU1O UVlN094VYplG0f4xTYymDDC0Hi68S4Xv4o998tDA7cH/9DuQWaWvysbkZOZVWHPZhC3+ 3j08I/Bef1sXw/93JdLtHcR/TnJS22g9FLSljoBAD16IK0ka7UP+bLElcBdw2ZJqwa0C gbNI/BdSvbxPK7LtAQxh7Uh83nNYXOU2hESikyymn/ch0jmFLE6YYIwbCwprhMeueRaR 60c3Y1vlSb59WW894nCo9h0jRRH+5Hqa/dFc2xuetUT80GETxmpOEO3F6PnxLa6Ovobe uGmw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789690336; x=1790295136; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PrpI8WWxUlqmCCZ/KlwzqkhVdxs2Qx8EkwLTiy0oOLg=; b=mwchSCv+NUiUI3SFrYqyCoBHQRAio3JaRFPARyDVDg/g7CEeZVOAkd+cRBEXSSqiWa 2xOlxESThUy7BgKQNH66Aeui7BRWGXLAMrXzIAdsoBbiqwFEoOam03fyp3GVBbljK8G7 YOk5Ns4nN6yLfKwF0JMEJ3ZShnv+slYLSPZH/rbu+47SKsnut6MTdoQv0agkcZuVMf/d fQg416Z144EY/1mkwufp8nNVgT163yjvsJUt+C/VCTjFEI/TTeQ5xKHdE2VldSmGkfpC VWZtoXskPyrBSgrVzboIXUgn/TSzM52v3zCGf+sHyiGnWVlGWs6bkujVBmG8vUGTV/ws /FuQ== X-Gm-Message-State: AFuF++nl0ASKk9zx0A5n7yJ0iVeeymhXwGWIwXupFcmF8olVb0DYjcyq jTVajxx1UpmI7/+bNGj1REPhioCof7lfsUGwnbOfZXZzz7+d+7cdu24cRrsf3GuGPDo= X-Gm-Gg: AYBFou3VD8XxYqUe1BEpJPJV/rYumpWjzS15xF2E9kcKLD0vLDEyYBsUzPfLSkJofuD 9tDvbPMh+PGACI4519vQ8o7e2Vb4sAxK9QuUfEY13N2x4E2i722SBOoAwV+8LFrLts7U9P7QWAA Y//XEeW428MES6q6qZDGGYcULgyesYL+esfGXSEH2E65Yj70GB/2r7qkktpZcE5cWHGEaGGwKv3 lfGcG6qO/dgPtjd0sA/9yTi2DBHq0XVhC5OwdfccmuUpzhhYzmhdBFirznvVd6RKUZKz38sptgv MSPvEK3DnXG4A5gK4RRu3YURLdW6miyf+//3My+qmi5+JmtQtM/1dBqVSnu64WF7ezJGywPJ8j+ WJjY5HiSAL7r+gZ66z+NTh04dBfOaN5sMXzDm1WwLhwkyeZxqihqRWfeweyfk86C1OHRdmWc6XD X22UqErbzBdYLHxFy93uylxiepTzqwddZ99YcIf6Ac0aXjeS/WDv2La+phx793jTNLgn3UdO5Y3 15UJuEh8IsEG14wZLLifTNkIYuEF8jSGbzr1FGZLObnuHWCld1Sd5CLyi/S X-Received: by 2002:a05:620a:19a4:b0:939:e6d9:b2c with SMTP id af79cd13be357-93bc62a9fbbmr758877685a.16.1789690335926; Thu, 17 Sep 2026 17:12:15 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93b81cff7cbsm583762585a.41.2026.09.17.17.12.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 17 Sep 2026 17:12:15 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com Subject: [PATCH v2 2/2] mm/mempolicy: stop copying the nodemask in the interleave paths Date: Thu, 17 Sep 2026 20:12:03 -0400 Message-ID: <20260918001203.3389165-3-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260918001203.3389165-1-gourry@gourry.net> References: <20260918001203.3389165-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The interleave node selectors copy pol->nodes onto the stack so the mask cannot change while they walk it. nodemask_t is 128 bytes at MAX_NUMNODES=3D1024, and two of the three run per folio fault. The copy only buys consistency between the node count and the walk. Drop the consistency and just bounds check the walk instead. The nodelist access is racy by design, but safe as long as we handle the scenario where a cpuset rebind causes a torn nodemask read to perceive the nodemask as empty (weight_total =3D=3D 0). If an empty nodelist or weight is perceived, fall back to numa_node_id(), which is what the functions already did when the copy came back empty Otherwise, iterating the nodelist during the actual allocation loop is perfectly safe - a concurrent rebind may simply cause a skew in in the distribution of memory (or fail and fall back the same as any other error condition). weighted_interleave_nid() counts the nodes as we sum the weights. We use that node count to limit the maximum skew a single node can host. interleave_nid() walks with next_node_in() rather than next_node(), so a mask that shrank mid-walk wraps to a node still in the policy. alloc_pages_bulk_weighted_interleave() derives per-node counts from a weight total summed over the mask, so a changing mask can make them exceed the request. Clamp each chunk to the space left in page_array. A cpuset cookie will not work here: two of these take VMA policies, which mpol_rebind_mm() rebinds under mmap_write_lock(), not mems_allowed_seq. Cost is distribution accuracy during a rebind - but the copy never corrected this anyway, it was just a safety mechanism to prevent div/0 and overrunning the alloc request buffer. Remove read_once_policy_nodemask(), now unused. -fstack-usage at MAX_NUMNODES=3D1024: weighted_interleave_nid 184 -> 56 interleave_nid 168 -> 32 alloc_pages_bulk_mempolicy_noprof 360 -> 136 Assisted-by: LLM Signed-off-by: Gregory Price (Meta) Acked-by: David Hildenbrand (Arm) Reviewed-by: Rakie Kim --- mm/mempolicy.c | 93 ++++++++++++++++++++++++++++++-------------------- 1 file changed, 56 insertions(+), 37 deletions(-) diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 2643915dc966..fd97fb0289bc 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -2197,34 +2197,15 @@ unsigned int mempolicy_slab_node(void) } } =20 -static unsigned int read_once_policy_nodemask(struct mempolicy *pol, - nodemask_t *mask) -{ - /* - * barrier stabilizes the nodemask locally so that it can be iterated - * over safely without concern for changes. Allocators validate node - * selection does not violate mems_allowed, so this is safe. - */ - barrier(); - memcpy(mask, &pol->nodes, sizeof(nodemask_t)); - barrier(); - return nodes_weight(*mask); -} - static unsigned int weighted_interleave_nid(struct mempolicy *pol, pgoff_t= ilx) { struct weighted_interleave_state *state; - nodemask_t nodemask; - unsigned int target, nr_nodes; + unsigned int target, nnodes =3D 0; u8 *table =3D NULL; unsigned int weight_total =3D 0; u8 weight; int nid =3D 0; =20 - nr_nodes =3D read_once_policy_nodemask(pol, &nodemask); - if (!nr_nodes) - return numa_node_id(); - rcu_read_lock(); =20 state =3D rcu_dereference(wi_state); @@ -2232,22 +2213,45 @@ static unsigned int weighted_interleave_nid(struct = mempolicy *pol, pgoff_t ilx) if (state) table =3D state->iw_table; =20 - /* calculate the total weight */ - for_each_node_mask(nid, nodemask) + /* calculate the total weight and the node count */ + for_each_node_mask(nid, pol->nodes) { weight_total +=3D table ? table[nid] : 1; + nnodes++; + } + + /* the mask is empty */ + if (!weight_total) { + rcu_read_unlock(); + return numa_node_id(); + } =20 /* Calculate the node offset based on totals */ target =3D ilx % weight_total; - nid =3D first_node(nodemask); - while (target) { + nid =3D first_node(pol->nodes); + + /* + * The target was calculated in a separate loop, and a concurrent + * rebind can change the contents of pol->nodes as we calculate. + * Access is safe, in the worst case we suddenly perceive an empty + * nodemask and return numa_node_id() below - otherwise we may + * simply cause a skew in allocations. + * + * Clamp this loop to a single pass (nnodes) to keep the walk + * bounded by node count. + */ + while (target && nnodes-- && nid < MAX_NUMNODES) { /* detect system default usage */ weight =3D table ? table[nid] : 1; if (target < weight) break; target -=3D weight; - nid =3D next_node_in(nid, nodemask); + nid =3D next_node_in(nid, pol->nodes); } rcu_read_unlock(); + + /* the mask emptied under the walk */ + if (nid >=3D MAX_NUMNODES) + return numa_node_id(); return nid; } =20 @@ -2258,18 +2262,23 @@ static unsigned int weighted_interleave_nid(struct = mempolicy *pol, pgoff_t ilx) */ static unsigned int interleave_nid(struct mempolicy *pol, pgoff_t ilx) { - nodemask_t nodemask; unsigned int target, nnodes; int i; int nid; =20 - nnodes =3D read_once_policy_nodemask(pol, &nodemask); + nnodes =3D nodes_weight(pol->nodes); if (!nnodes) return numa_node_id(); target =3D ilx % nnodes; - nid =3D first_node(nodemask); - for (i =3D 0; i < target; i++) - nid =3D next_node(nid, nodemask); + nid =3D first_node(pol->nodes); + + /* A concurrent cpuset rebind may cause us to see an empty nodemask */ + for (i =3D 0; i < target && nid < MAX_NUMNODES; i++) + nid =3D next_node_in(nid, pol->nodes); + + /* the mask emptied under the walk */ + if (nid >=3D MAX_NUMNODES) + return numa_node_id(); return nid; } =20 @@ -2665,7 +2674,6 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, u8 *table, weight; unsigned int weight_total =3D 0; unsigned long rem_pages =3D nr_pages; - nodemask_t nodes; int nnodes, node; int resume_node =3D MAX_NUMNODES - 1; u8 resume_weight =3D 0; @@ -2675,10 +2683,10 @@ static unsigned long alloc_pages_bulk_weighted_inte= rleave(gfp_t gfp, if (!nr_pages) return 0; =20 - /* read the nodes onto the stack, retry if done during rebind */ + /* count the nodes, retry if a rebind happened during the read */ do { cpuset_mems_cookie =3D read_mems_allowed_begin(); - nnodes =3D read_once_policy_nodemask(pol, &nodes); + nnodes =3D nodes_weight(pol->nodes); } while (read_mems_allowed_retry(cpuset_mems_cookie)); =20 /* if the nodemask has become invalid, we cannot do anything */ @@ -2688,7 +2696,7 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, /* Continue allocating from most recent node and adjust the nr_pages */ node =3D me->il_prev; weight =3D me->il_weight; - if (weight && node_isset(node, nodes)) { + if (weight && node_isset(node, pol->nodes)) { node_pages =3D min(rem_pages, weight); nr_allocated =3D __alloc_pages_bulk(gfp, node, NULL, node_pages, page_array); @@ -2712,9 +2720,13 @@ static unsigned long alloc_pages_bulk_weighted_inter= leave(gfp_t gfp, table =3D state ? state->iw_table : NULL; =20 /* calculate total, detect system default usage */ - for_each_node_mask(node, nodes) + for_each_node_mask(node, pol->nodes) weight_total +=3D table ? table[node] : 1; =20 + /* the mask emptied since it was counted */ + if (!weight_total) + goto out; + /* * Calculate rounds/partial rounds to minimize __alloc_pages_bulk calls. * Track which node weighted interleave should resume from. @@ -2724,10 +2736,14 @@ static unsigned long alloc_pages_bulk_weighted_inte= rleave(gfp_t gfp, */ rounds =3D rem_pages / weight_total; delta =3D rem_pages % weight_total; - resume_node =3D next_node_in(prev_node, nodes); + resume_node =3D next_node_in(prev_node, pol->nodes); + if (resume_node >=3D MAX_NUMNODES) + goto out; resume_weight =3D table ? table[resume_node] : 1; for (i =3D 0; i < nnodes; i++) { - node =3D next_node_in(prev_node, nodes); + node =3D next_node_in(prev_node, pol->nodes); + if (node >=3D MAX_NUMNODES) + break; weight =3D table ? table[node] : 1; node_pages =3D weight * rounds; /* If a delta exists, add this node's portion of the delta */ @@ -2744,6 +2760,8 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, /* node_pages can be 0 if an allocation fails and rounds =3D=3D 0 */ if (!node_pages) break; + /* a rebind can invalidate the counts: never overrun page_array */ + node_pages =3D min(node_pages, nr_pages - total_allocated); nr_allocated =3D __alloc_pages_bulk(gfp, node, NULL, node_pages, page_array); page_array +=3D nr_allocated; @@ -2754,6 +2772,7 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, } me->il_prev =3D resume_node; me->il_weight =3D resume_weight; +out: srcu_read_unlock_fast(&wi_srcu, scp); return total_allocated; } --=20 2.55.0