From nobody Fri Oct 2 02:30:25 2026 Received: from invmail4.hynix.com (exvmail4.hynix.com [166.125.252.92]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 1C0223D47BF; Thu, 6 Aug 2026 08:25:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=166.125.252.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004713; cv=none; b=e+RbHQsaEdEJ8ov+Eo6k5MH8YeFTOiOEDeT4/7ZtF0yWrHgiwN02yg58CxobbVQ7VqVumgcExvz5hZtXrV8kf7OQlwUECBDMmdfQGN4mWH6AZR+wdfg7ARNmSX2qwsFy8Nqg8quznpob2k406e+9qHV9yrF775aKhRtdQyxyTw0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004713; c=relaxed/simple; bh=pOtWcI/7SvRNfth9GmfHdzerexk9s5vTUW60UNf+nOU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GYT+2qTQ2aFss3uKE5uuVak4mfozcQr8p77ZzJ+yVzIIBcc8//WDfhu5pMaHS43DuKhAlDG2jC962dPzRCpXqiu/j+zyfY4q1YqKdR4S7TPXoxDSIHpFyUa0vEJWGcGoU49JtLaKzbBLrUCJdRDz1uOV/K2hQQtGz6SI6xBEkqA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com; spf=pass smtp.mailfrom=sk.com; arc=none smtp.client-ip=166.125.252.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sk.com X-AuditID: a67dfc5b-c45ff70000001609-98-6a74414b2e1a From: Rakie Kim To: akpm@linux-foundation.org Cc: gourry@gourry.net, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-cxl@vger.kernel.org, nvdimm@lists.linux.dev, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, byungchul@sk.com, ying.huang@linux.alibaba.com, apopple@nvidia.com, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, dave@stgolabs.net, jic23@kernel.org, dave.jiang@intel.com, alison.schofield@intel.com, vishal.l.verma@intel.com, ira.weiny@intel.com, harry@kernel.org, kernel_team@skhynix.com, honggyu.kim@sk.com, yunjeong.mun@sk.com, rakie.kim@sk.com Subject: [PATCH 1/4] mm/numa: introduce nearest_nodes_nodemask() Date: Thu, 6 Aug 2026 17:09:32 +0900 Message-ID: <20260806080936.421-2-rakie.kim@sk.com> X-Mailer: git-send-email 2.52.0.windows.1 In-Reply-To: <20260806080936.421-1-rakie.kim@sk.com> References: <20260806080936.421-1-rakie.kim@sk.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrDIsWRmVeSWpSXmKPExsXC9ZZnoa63Y0mWwdsvEhZz1q9hs7j7+AKb xa4bIRYnbjayWay+uYbR4vnWX4wWP+8eZ7e4fmslo8X+p89ZLB40rWKyOL51HrvFulOH2CzO zzrFYnF51xw2i3tr/rNavHnsZvGtT9rifp+Dxcoff1gtjqzfzmQx+dICNouOl/dZLG5NOMZk sXpNhsXso/fYHSQ9ds66y+6xYFOpR3fbZXaPzSu0PBbvecnksWlVJ5vHpk+T2D1OzPjN4rHz oaXHi80zGT16m9+xeUydXe+xfstVFo/Pm+QC+KK4bFJSczLLUov07RK4Mjq2LmUu2CJZcXq5 ZQPjBZEuRg4OCQETiYc9gV2MnGDmtSvdjCBhNgEliWN7Y0DCIgKyElP/nmfpYuTiYBZYxCpx 4vNZVpCEsICDxNVHG5lBbBYBVYmnHw4xgti8AsYSXx/uZIGYqSmxbuMtMJsTaP73O4vAbCGg mj+/trFD1AtKnJz5BCzOLCAv0bx1NjPIMgmBr+wSVydNZIYYJClxcMUNlgmM/LOQ9MxC0rOA kWkVo1BmXlluYmaOiV5GZV5mhV5yfu4mRmCcLqv9E72D8dOF4EOMAhyMSjy8F4yLs4RYE8uK K3MPMUpwMCuJ8LIeLMoS4k1JrKxKLcqPLyrNSS0+xCjNwaIkzmv0rTxFSCA9sSQ1OzW1ILUI JsvEwSnVwJig88x1h1aPwm5V1YZ81qmv2t3WcDqb10/QzF/izHNM75m19YLSXPP3M78IXmvi Z1i/cGWk54HrTTn/r/CfLjhW9zpEX2RCVsb9HebX1oUsPcmlftZ2rkJx/8nFd54dkt8dLDVl tfzKyvuMewMd/rw705t2cf31DtVvd15rXu89fSQ9qNDv4EklluKMREMt5qLiRAB1lAvwzwIA AA== X-Brightmail-Tracker: H4sIAAAAAAAAA02Ra0hTYQCG+3bOzo5HB6e56KRlOAxJ8oYJn1Ax/WFfCtGVLIRceXLTOW1T USHyEuWFTEtBp4W4Lt5QnM7UDOnkPUrJ1DSveWclkuVtmjkl8N/D+7y8f14SEzXiNqRCFcWq VTKlhKBwKtcx0NnfOyrUbbrFEhZUlhNweKKbgA1fL8FP2fkEbB9IJGDZQDmAM4Y1AFeH2wSw f7AEwMVpIwabpmZwOJZUyoNthmcC+P5pBx9WdHIE7NJ24rCnoYCAI+WbfPhjwhcuZdjC0Qwp LFlZ50Oub4YPmytf8+CTz4UETJkbxeFgZisPlpXLoam2mID5LSMCqR2q1w4LUKE+GqXf7xGg 6mInpGuc4yF9aSqB9L8eC1B7rglH9eNeaLY6D6CHyfMEWvqGkG52gYdy8u+iyppeHC3q7c7R 16gTwaxSEcOqXU8FUfIUwwsssuZA7IdXXgmgW5wGLEiGPs70fUkHaYAkCVrCtL4NNMdi+hCT s9GFpwGKxOgiPtO++JFvFta0lOn9XoWZGaePMFMLHDCzkPZg/ozX4zubR5mKqsFtttjaXx4q 2mbRVmd9rVaw09/LdORNbucYfZhJNuRjmcBKu0tpd6lCwCsFYoUqJlymUHq6aMLkcSpFrMvN iHA92Pr05Z31rDrwu+c0B2gSSKyE3R6aUBFfFqOJC+cAQ2ISsZD/Th0qEgbL4uJZdcR1dbSS 1XDAlsQl+4V+V9ggER0ii2LDWDaSVf+3PNLCJgH4Uf0/67LcpUkOPngbcDLmHjzffmPfprj6 3lBzn6+rvzM1uQrcYqW3dAb1cs2e+LVHG6Py1OqzV/Wm540nPWs6x4pTm6iyiogA6JuYnS0J CLF31TWYUiW3H7g61LWuWnZ2U47HvP9acxfm7NM4WTY3b3Ryf3NRZLx8hljxGZLgGrnM3QlT a2T/AOU56VXPAgAA X-CFilter-Loop: Reflected Content-Type: text/plain; charset="utf-8" Add a NUMA helper, nearest_nodes_nodemask(), that returns every node in a given nodemask located at the minimum distance from a source node. Unlike nearest_node_nodemask(), which returns only a single node, this helper reports the complete set of nodes that share the closest distance. This is needed when several nodes are equally near and all of the nearest candidates must be considered together. The helper clears the output nodemask and sets every node that meets the minimum-distance condition. It returns 0 on success, or -EINVAL when the output argument is invalid. Signed-off-by: Rakie Kim --- include/linux/numa.h | 11 +++++++++++ mm/mempolicy.c | 41 +++++++++++++++++++++++++++++++++++++++++ 2 files changed, 52 insertions(+) diff --git a/include/linux/numa.h b/include/linux/numa.h index e6baaf6051bc..4f2a0c344122 100644 --- a/include/linux/numa.h +++ b/include/linux/numa.h @@ -33,6 +33,8 @@ int numa_nearest_node(int node, unsigned int state); =20 int nearest_node_nodemask(int node, nodemask_t *mask); =20 +int nearest_nodes_nodemask(int node, const nodemask_t *mask, nodemask_t *o= ut); + #ifndef memory_add_physaddr_to_nid int memory_add_physaddr_to_nid(u64 start); #endif @@ -54,6 +56,15 @@ static inline int nearest_node_nodemask(int node, nodema= sk_t *mask) return NUMA_NO_NODE; } =20 +static inline int nearest_nodes_nodemask(int node, const nodemask_t *mask, + nodemask_t *out) +{ + if (!out) + return -EINVAL; + nodes_clear(*out); + return 0; +} + static inline int memory_add_physaddr_to_nid(u64 start) { return 0; diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 4e4421b22b59..19417b0afc30 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -337,6 +337,47 @@ int nearest_node_nodemask(int node, nodemask_t *mask) } EXPORT_SYMBOL_GPL(nearest_node_nodemask); =20 +/** + * nearest_nodes_nodemask - Find all nodes in @mask that are nearest to @n= ode + * @node: The reference node ID to measure distance from + * @mask: The set of candidate nodes to compare against + * @out: Pointer to a nodemask that will store the nearest node(s) + * + * This function iterates over all nodes in @mask and measures the distance + * between each candidate node and the given @node using node_distance(). + * It finds the minimum distance and then records all nodes in @mask that + * share that same minimum distance into the output mask @out. + * + * For example, if multiple nodes have equal minimal distance to @node, all + * of them are included in @out. + * + * Return: 0 on success, or -EINVAL if @out is NULL. + */ +int nearest_nodes_nodemask(int node, const nodemask_t *mask, nodemask_t *o= ut) +{ + int dist, n, min_dist =3D INT_MAX; + + if (!out) + return -EINVAL; + + nodes_clear(*out); + + for_each_node_mask(n, *mask) { + dist =3D node_distance(node, n); + + if (dist < min_dist) { + min_dist =3D dist; + nodes_clear(*out); + node_set(n, *out); + } else if (dist =3D=3D min_dist) { + node_set(n, *out); + } + } + + return 0; +} +EXPORT_SYMBOL_GPL(nearest_nodes_nodemask); + struct mempolicy *get_task_policy(struct task_struct *p) { struct mempolicy *pol =3D p->mempolicy; --=20 2.25.1 From nobody Fri Oct 2 02:30:25 2026 Received: from invmail4.hynix.com (exvmail4.skhynix.com [166.125.252.92]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 1C1253DB331; Thu, 6 Aug 2026 08:25:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=166.125.252.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004715; cv=none; b=ORw1nXNfdhjS8inkFBE1uQzDcQoKuMn0XKNi51THK66PcEgRPl9c6QzU6dTLk5xuCYl0aKFQAZlbEa1T+bpSnXbI+Rh1bK0zg+QHv8Pl+qv3AkpkDPZpQ7SGAEC2ypxHg01lNX2+xiO+13UxrQgjUxhIuF0yUe0WdPixfaQ+rnM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004715; c=relaxed/simple; bh=pfk82g/J7NCtRDtZZ1249d5Q370rTH3tZXsBltyno/c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CpUbG+zcfXiukTDlvXIR2Es18RQ0eKqCCm9VMNK2qIk6ocMpP5RQzTOSSGyoBVSoITgc9vvcKYeHJGjgcKERPIZh8aAIZp2bHSrU2vYtIdN2hG00RIEn4dzD8/tBlykUYxLRkvpQz6T+3vUa4HIGPoJSro5Hiy17izEG37hixFI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com; spf=pass smtp.mailfrom=sk.com; arc=none smtp.client-ip=166.125.252.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sk.com X-AuditID: a67dfc5b-c45ff70000001609-a1-6a74414c113a From: Rakie Kim To: akpm@linux-foundation.org Cc: gourry@gourry.net, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-cxl@vger.kernel.org, nvdimm@lists.linux.dev, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, byungchul@sk.com, ying.huang@linux.alibaba.com, apopple@nvidia.com, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, dave@stgolabs.net, jic23@kernel.org, dave.jiang@intel.com, alison.schofield@intel.com, vishal.l.verma@intel.com, ira.weiny@intel.com, harry@kernel.org, kernel_team@skhynix.com, honggyu.kim@sk.com, yunjeong.mun@sk.com, rakie.kim@sk.com Subject: [PATCH 2/4] mm/memory-tiers: introduce package-aware topology management for NUMA nodes Date: Thu, 6 Aug 2026 17:09:33 +0900 Message-ID: <20260806080936.421-3-rakie.kim@sk.com> X-Mailer: git-send-email 2.52.0.windows.1 In-Reply-To: <20260806080936.421-1-rakie.kim@sk.com> References: <20260806080936.421-1-rakie.kim@sk.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Brightmail-Tracker: H4sIAAAAAAAAA02RW0hTcRzH+5+dm8PpcQ46KioMelDSLlr8gxCDihP4oOhTGjn00I5uU87U VBQ1ScwbXiF1LFHK25g1W6Z2c5jayBxJ5jSdssRJuUDF1LRyiuDbh+/t9/AjBeJ3qC/JqTJY XiVTSHEhKlx1bw2JupKRcnanTwI1PTocztktOByYjoNj1iIcdlt1AC4bdwDcnhsl4NeZTgDf Li2jcOFeFwJHjVoC6s0mHE40mVE4OaDB4bzuHwZ/2q/DzSo/aKuKhJ1buxgc7ulDYN3nFhyW rthQOFM9gsBunRw2v58nIn2Y/qY5gmkxZDLlJZME09sRzLS9WkEYQ9cDnDGs1RLM2MM/KNO/ eIlx9DYCprLYiTMNzQVMz/MvKLNuCIj2uCm8nMwquCyWPxORKJS32Zfx9PttSHa9vQYpBLpf oAy4kTQVTmtstcgRrw7/wMoASeKUlB55neCSJZQ/3bA3gZYBISmgWjF6bH0ccxneVCJtHl/B XYxSp+iiWcuBLqLCaL2lljjcDKL1z2ZQF7vt7//+1nrA4v3M7s4L4jDvRX9o/H6gC6hAutjY LHAdo6ltgq7o1mOHQz70UMc0Wg08m451mo51WgDSBcScKksp4xThofIcFZcdmpSmNID9vz7J 341/CdYssSZAkUDqLrKEqVPEmCxLnaM0AZoUSCUibIhPEYuSZTm5LJ92m89UsGoT8CNR6UnR +c27yWLqjiyDTWXZdJY/chHSzbcQnLCx3gn1VRNUXWWJ037Ld+pNhaNzwxqDWx0I18BtRQdY vDxB0FLe1fCJ/NNcdoFv4gbkU/tyE2KfXhMFaltnK2cfOx/dSLKbNBf894x/49sHo+eKtX4l HjEtyvIMiTJk8eMnz3mfi4MOOi9OO9g+ueA0U6XRUzZ7RNRGjRRVy2XnggW8WvYf40aRjNMC AAA= X-Brightmail-Tracker: H4sIAAAAAAAAA02ReUiTYQDGe79rn8PB1xr4oWW1CEvILK1eQcSC6qWkLAqjhBz6ofOY7kyF yoMktcQT1K0wB+ExlWZqHmEt84rlyqy8Lc0paomm5hHmQeB/D7/fw/PPQ+PCRsKelspUnEIm CRdTfIKf6+R/yOekKtQ1dZqEugoDBfuHLRSs+3oFvs/WUrC1O56Cpd0GAK1VSwAu9rfw4Jee YgBnRydw2PjDSsChhBIMtlQ95sE3j9pIWN5uomBHfjsBO+t0FBwwrJJwcvgMnE9zgINp3rD4 zwoJTZ+tJGyqqMFg1scCCt4fHyRgT3ozBksNIXC5uoiC2rcDPG9HVJvfz0MFRjVKTerkocoi Z6RvGMeQsSSZQsaZTB5qzV0mUO03DzRWmQfQw8SfFJrvRUg/No2hHO1dVPG8i0CzRkdf5jrf M4gLl2o4xWGvAH6IfthKRd3TY9HZwxlYHDD8AinAhmYZd3aqaYJMATRNMWK2+aX/OhYxu9ic vx1ECuDTOFNIsq2zZnJd7GAC2HbzOLWeCWY/G99r2eACxo0tt2TyNjcPsuXPeoj1bLO2v9BX uJGFa52VpWreZn8725Y3ssFxZjebWKXF04Ft/haVv0UVAKwEiKQyTYREGn7MRRkWEiOTRrsE RkYYwdqtT2+vZLwAvzvPmgBDA7GtwOKmDBWSEo0yJsIEWBoXiwTka0WoUBAkiYnlFJE3Fepw TmkCDjQhthOc8+MChEywRMWFcVwUp/hvMdrGPg6kXFp0scH33dBlmT+oz5tiV8um6h9ognVd Xe6T8rm5BPzyk9ChbaumUV//V324XdnCTnQx8EKzZPC4f03Dte/qesPMiPC0b6zccmqPT9K7 ZPOnpZUD0xdUbX5WkdqvSpOruZPotNBe7JWe6elx1Fkud+p1NV/de6JZesshK0IlExPKEMkR Z1yhlPwD4XMaDtICAAA= X-CFilter-Loop: Reflected Content-Type: text/plain; charset="utf-8" The NUMA distance model provides only relative latency values between nodes and has no notion of structural grouping. Memory policies based on distance alone therefore cannot tell which nodes are local to the same physical package and which belong to a different one, which limits how well they can keep placement and migration within a package. Introduce a package-aware topology layer that groups NUMA nodes into a "memory package": a set of CPU nodes and their local memory-only nodes (such as CXL or HBM). A subsystem that owns a node registers a resolver to supply its package, and a policy can query, for any node, the other nodes that belong to the same package. The grouping is built from the information firmware already provides, as nodes come online. A CPU node is placed in the package named by its firmware physical package id. A memory-only node is placed by an initiator CPU node when a driver supplies one, or otherwise by its nearest CPU node under the SLIT distance table; a SLIT-derived entry is provisional and is upgraded once a driver supplies an initiator. This makes the package association explicit so that policies can consume it, and the mapping can be revisited as firmware interfaces expose better information. A single package may hold more than one CPU node or more than one memory-only node, so the model does not assume a one-to-one package to node or package to CXL device relationship. The current package topology is exposed read-only under /sys/devices/system/package/packageN/: package_nodes (all NUMA nodes in the package), package_cpu_nodes (CPU nodes), package_mem_only_nodes (memory-only nodes), physical_package_id (physical package id). The interface is read-only by design: it reports the topology the kernel derived but does not let user space override it. A machine whose firmware describes the topology incorrectly should be fixed in firmware rather than papered over through sysfs. Topology is also validated for the symmetric shape that package-aware placement relies on - every package having the same number of CPU and memory-only nodes, at least two packages, and at least two nodes per package. The verdict is recomputed at boot and on node hotplug and is published for policies to consult. On any topology that does not meet these conditions the verdict is simply false, so a consumer degrades to its original, non-package-aware behavior instead of acting on a grouping that does not describe the machine. Signed-off-by: Rakie Kim --- .../ABI/testing/sysfs-devices-system-package | 35 + include/linux/memory-tiers.h | 113 ++ mm/memory-tiers.c | 1009 +++++++++++++++++ 3 files changed, 1157 insertions(+) create mode 100644 Documentation/ABI/testing/sysfs-devices-system-package diff --git a/Documentation/ABI/testing/sysfs-devices-system-package b/Docum= entation/ABI/testing/sysfs-devices-system-package new file mode 100644 index 000000000000..6500f9e5ff19 --- /dev/null +++ b/Documentation/ABI/testing/sysfs-devices-system-package @@ -0,0 +1,35 @@ +What: /sys/devices/system/package/ +Date: August 2026 +Contact: Linux memory management mailing list +Description: Memory package topology + + A "memory package" groups the NUMA nodes associated with one + physical CPU package (socket): the nodes that have CPUs and + the nodes that only have memory (e.g. CXL or HBM). + + All attributes are read-only; the topology cannot be + overridden from user space. + +What: /sys/devices/system/package/packageN/package_nodes +Date: August 2026 +Contact: Linux memory management mailing list +Description: All NUMA nodes in this package, in nodelist format + (e.g. "0,2"). + +What: /sys/devices/system/package/packageN/package_cpu_nodes +Date: August 2026 +Contact: Linux memory management mailing list +Description: The nodes that have CPUs in this package, in nodelist + format. + +What: /sys/devices/system/package/packageN/package_mem_only_nodes +Date: August 2026 +Contact: Linux memory management mailing list +Description: The nodes that only have memory (e.g. CXL/HBM) in this + package, in nodelist format. + +What: /sys/devices/system/package/packageN/physical_package_id +Date: August 2026 +Contact: Linux memory management mailing list +Description: The physical package id this group corresponds to, as + reported by CPU topology. diff --git a/include/linux/memory-tiers.h b/include/linux/memory-tiers.h index 7999c58629ee..f8778c43429f 100644 --- a/include/linux/memory-tiers.h +++ b/include/linux/memory-tiers.h @@ -52,10 +52,25 @@ int mt_perf_to_adistance(struct access_coordinate *perf= , int *adist); struct memory_dev_type *mt_find_alloc_memory_type(int adist, struct list_head *memory_types); void mt_put_memory_types(struct list_head *memory_types); + +int register_mp_package_notifier(struct notifier_block *notifier); +void unregister_mp_package_notifier(struct notifier_block *notifier); +int mp_probe_package_id(int nid); +int mp_add_package_node_by_initiator(int nid, int initiator_nid); +int mp_add_package_node(int nid); +int mp_get_package_nodes(int nid, nodemask_t *out); +int mp_get_package_cpu_nodes(int nid, nodemask_t *out); +int mp_get_package_memory_only_nodes(int nid, nodemask_t *out); +bool mp_is_topology_symmetric(void); #ifdef CONFIG_NUMA_MIGRATION int next_demotion_node(int node, const nodemask_t *allowed_mask); void node_get_allowed_targets(pg_data_t *pgdat, nodemask_t *targets); bool node_is_toptier(int node); + +int mp_next_demotion_nodemask(int nid, nodemask_t *out); +int mp_next_demotion_node(int nid); +int mp_next_promotion_nodemask(int nid, nodemask_t *out); +int mp_next_promotion_node(int nid); #else static inline int next_demotion_node(int node, const nodemask_t *allowed_m= ask) { @@ -71,6 +86,30 @@ static inline bool node_is_toptier(int node) { return true; } + +static inline int mp_next_demotion_nodemask(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_next_demotion_node(int nid) +{ + return NUMA_NO_NODE; +} + +static inline int mp_next_promotion_nodemask(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_next_promotion_node(int nid) +{ + return NUMA_NO_NODE; +} #endif =20 #else @@ -151,5 +190,79 @@ static inline struct memory_dev_type *mt_find_alloc_me= mory_type(int adist, static inline void mt_put_memory_types(struct list_head *memory_types) { } + +static inline int register_mp_package_notifier(struct notifier_block *noti= fier) +{ + return 0; +} + +static inline void unregister_mp_package_notifier(struct notifier_block *n= otifier) +{ +} + +static inline int mp_probe_package_id(int nid) +{ + return NOTIFY_DONE; +} + +static inline int mp_add_package_node_by_initiator(int nid, int initiator_= nid) +{ + return 0; +} + +static inline int mp_add_package_node(int nid) +{ + return 0; +} + +static inline int mp_get_package_nodes(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_get_package_cpu_nodes(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_get_package_memory_only_nodes(int nid, nodemask_t *ou= t) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline bool mp_is_topology_symmetric(void) +{ + return false; +} + +static inline int mp_next_demotion_nodemask(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_next_demotion_node(int nid) +{ + return NUMA_NO_NODE; +} + +static inline int mp_next_promotion_nodemask(int nid, nodemask_t *out) +{ + if (out) + nodes_clear(*out); + return -ENOENT; +} + +static inline int mp_next_promotion_node(int nid) +{ + return NUMA_NO_NODE; +} #endif /* CONFIG_NUMA */ #endif /* _LINUX_MEMORY_TIERS_H */ diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c index 54851d8a195b..5932df315604 100644 --- a/mm/memory-tiers.c +++ b/mm/memory-tiers.c @@ -1,4 +1,5 @@ // SPDX-License-Identifier: GPL-2.0 +#include #include #include #include @@ -51,6 +52,11 @@ static const struct bus_type memory_tier_subsys =3D { .dev_name =3D "memory_tier", }; =20 +static const struct bus_type package_subsys =3D { + .name =3D "package", + .dev_name =3D "package", +}; + #ifdef CONFIG_NUMA_BALANCING /** * folio_use_access_time - check if a folio reuses cpupid for page access = time @@ -1007,3 +1013,1006 @@ static int __init numa_init_sysfs(void) subsys_initcall(numa_init_sysfs); #endif /* CONFIG_SYSFS */ #endif + +/** + * enum mp_nodes_type - Selector for which subset of a package to return + * @MP_NODES_ALL: All NUMA nodes that belong to the package. + * @MP_NODES_CPU: Only CPU nodes in the package. + * @MP_NODES_MEM_ONLY: Only memory-only nodes (e.g. CXL/HBM) in the packa= ge. + * + * Used internally to choose which nodemask to expose for a given package. + */ +enum mp_nodes_type { + MP_NODES_ALL, + MP_NODES_CPU, + MP_NODES_MEM_ONLY +}; + +/** + * struct memory_package - Per-physical-package container + * @package_id: Physical package id (from topology). + * @nodes: Nodemask of all member nodes in this package. + * @cpu_nodes: Nodemask of CPU nodes in this package. + * @memory_only_nodes: Nodemask of memory-only nodes in this package. + * @cpu_list: List head of CPU-type members. + * @memory_only_list: List head of memory-only members. + * @list: Linkage on the global @memory_packages list. + * @dev: sysfs device for this package. + * + * A memory_package groups NUMA nodes that share the same physical CPU pac= kage. + * The masks are used to implement package-local placement/demotion/promot= ion. + */ +struct memory_package { + int package_id; + nodemask_t nodes; + nodemask_t cpu_nodes; + nodemask_t memory_only_nodes; + struct list_head cpu_list; + struct list_head memory_only_list; + struct list_head list; + struct device dev; +}; + +/** + * enum mpn_source_flags - Source used to resolve a node's package members= hip + * @MPN_SRC_UNKNOWN: Unknown/unspecified. + * @MPN_SRC_CPU: Directly resolved from a CPU node (1:1). + * @MPN_SRC_INITIATOR: Resolved via an initiator CPU node provided by a = driver. + * @MPN_SRC_SLIT: Resolved via SLIT/nearest-node. + * + * These flags are informational; they describe how a given node was bound= to + * its package and help with policy decisions later. + */ +enum mpn_source_flags { + MPN_SRC_UNKNOWN =3D 0, + MPN_SRC_CPU =3D BIT(1), + MPN_SRC_INITIATOR =3D BIT(2), + MPN_SRC_SLIT =3D BIT(3) +}; + +/** + * struct memory_package_node - Per-node membership and preferences + * @nid: NUMA node id for this entry. + * @initiator_nid: CPU nid that served as the initiator when resolving = @nid. + * @package_id: Resolved package id that @nid belongs to. + * @source_flags: One of &enum mpn_source_flags describing the resolut= ion. + * @preferred: Opposite-type nearest candidates inside the same pac= kage. + * @package: Pointer to the owning &struct memory_package (NULL u= ntil bound). + * @package_entry: Linkage on the owning package's type list. + * + * Each NUMA node that participates in package-aware policy gets a wrapper= entry + * that caches package membership and the precomputed set of preferred tar= gets. + */ +struct memory_package_node { + int nid; + int initiator_nid; + int package_id; + int source_flags; + nodemask_t preferred; + struct memory_package *package; + struct list_head package_entry; +}; + +#define node_is_memory_only(_nid) \ + (node_state((_nid), N_MEMORY) && !node_state((_nid), N_CPU)) + +static BLOCKING_NOTIFIER_HEAD(mp_package_algorithms); + +static LIST_HEAD(memory_packages); +static struct memory_package_node *mpns[MAX_NUMNODES]; +static DEFINE_MUTEX(memory_package_lock); + +/* + * RCU snapshot of the package topology. The allocation path reads it + * often, so it is published for lockless reads instead of locking on + * every access. + */ +struct mp_snapshot { + struct rcu_head rcu; + int nr_packages; + int pkg_of[MAX_NUMNODES]; + struct { + nodemask_t nodes; + nodemask_t cpu_nodes; + nodemask_t memory_only_nodes; + } pkg[]; +}; + +static struct mp_snapshot __rcu *mp_snapshot; + +/** + * register_mp_package_notifier - Register a package resolution algorithm + * @notifier: Notifier called with the nid to resolve (see mp_probe_packag= e_id()). + * + * Drivers (e.g., CXL region/decoder code) register here to supply a packa= ge + * hint for newly appearing nodes. The notifier is invoked during nid->pac= kage + * resolution. + * + * Return: 0 on success, negative errno on failure. + */ +int register_mp_package_notifier(struct notifier_block *notifier) +{ + return blocking_notifier_chain_register(&mp_package_algorithms, notifier); +} +EXPORT_SYMBOL_GPL(register_mp_package_notifier); + +/** + * unregister_mp_package_notifier - Unregister a package resolution algori= thm + * @notifier: Notifier previously registered with register_mp_package_noti= fier(). + */ +void unregister_mp_package_notifier(struct notifier_block *notifier) +{ + blocking_notifier_chain_unregister(&mp_package_algorithms, notifier); +} +EXPORT_SYMBOL_GPL(unregister_mp_package_notifier); + +/** + * mp_probe_package_id - Invoke registered notifiers to resolve a node's p= ackage + * @nid: NUMA node id to resolve. + * + * Calls the blocking notifier chain to let subsystems provide an initiato= r or + * package id for @nid. + * + * Return: Notifier return code (>=3D0 typically); negative errno on failu= re. + */ +int mp_probe_package_id(int nid) +{ + return blocking_notifier_call_chain(&mp_package_algorithms, nid, NULL); +} +EXPORT_SYMBOL_GPL(mp_probe_package_id); + +static int mp_node_to_package_id(int nid) +{ + int package_id; + unsigned int first_cpu; + const struct cpumask *cpu_mask; + + if (nid < 0 || nid >=3D MAX_NUMNODES) + return -EINVAL; + + if (!node_state(nid, N_CPU)) + return -EINVAL; + + cpu_mask =3D cpumask_of_node(nid); + if (cpumask_empty(cpu_mask)) + return -EINVAL; + + first_cpu =3D cpumask_first(cpu_mask); + if (first_cpu >=3D nr_cpu_ids) + return -EINVAL; + + package_id =3D topology_physical_package_id(first_cpu); + if (package_id < 0) + return -EINVAL; + + return package_id; +} + +static void update_package_preferred(struct memory_package *mp) +{ + struct memory_package_node *mpn; + + lockdep_assert_held(&memory_package_lock); + + /* + * For each CPU node, compute its preferred set as the nearest + * memory-only node(s) within the same package. If the package has + * no memory-only nodes, fall back to a self-reference so callers + * never see an empty preferred set. + */ + list_for_each_entry(mpn, &mp->cpu_list, package_entry) { + nodes_clear(mpn->preferred); + if (!nodes_empty(mp->memory_only_nodes)) + nearest_nodes_nodemask(mpn->nid, &mp->memory_only_nodes, + &mpn->preferred); + else + node_set(mpn->nid, mpn->preferred); + } + + /* + * Symmetrically, for each memory-only node, compute its preferred set + * as the nearest CPU node(s) within the same package. If the package + * has no CPU nodes, fall back to a self-reference. + */ + list_for_each_entry(mpn, &mp->memory_only_list, package_entry) { + nodes_clear(mpn->preferred); + if (!nodes_empty(mp->cpu_nodes)) + nearest_nodes_nodemask(mpn->nid, &mp->cpu_nodes, + &mpn->preferred); + else + node_set(mpn->nid, mpn->preferred); + } +} + +static inline bool memory_package_is_empty(struct memory_package *mp) +{ + lockdep_assert_held(&memory_package_lock); + + return (nodes_empty(mp->cpu_nodes) && nodes_empty(mp->memory_only_nodes)); +} + +static inline bool package_node_is_valid(int nid) +{ + if (!mpns[nid]) + return false; + + if (nodes_empty(mpns[nid]->preferred) || (mpns[nid]->package =3D=3D NULL)) + return false; + + return true; +} + +static const struct attribute_group *memory_package_groups[]; + +/* Freed when the last reference to the package's sysfs device is dropped.= */ +static void memory_package_release(struct device *dev) +{ + struct memory_package *mp =3D container_of(dev, struct memory_package, de= v); + + kfree(mp); +} + +static struct memory_package *create_memory_package(int package_id) +{ + struct memory_package *mempackage; + int ret; + + mempackage =3D kzalloc_obj(*mempackage); + if (!mempackage) + return ERR_PTR(-ENOMEM); + + mempackage->package_id =3D package_id; + mempackage->nodes =3D NODE_MASK_NONE; + mempackage->cpu_nodes =3D NODE_MASK_NONE; + mempackage->memory_only_nodes =3D NODE_MASK_NONE; + INIT_LIST_HEAD(&mempackage->cpu_list); + INIT_LIST_HEAD(&mempackage->memory_only_list); + INIT_LIST_HEAD(&mempackage->list); + device_initialize(&mempackage->dev); + mempackage->dev.release =3D memory_package_release; + dev_set_drvdata(&mempackage->dev, mempackage); + mempackage->dev.bus =3D &package_subsys; + mempackage->dev.groups =3D memory_package_groups; + ret =3D dev_set_name(&mempackage->dev, "package%d", package_id); + if (ret) { + put_device(&mempackage->dev); + return ERR_PTR(ret); + } + + return mempackage; +} + +static struct memory_package *find_create_memory_package(int package_id) +{ + struct memory_package *mempackage, *existing; + int ret; + + mutex_lock(&memory_package_lock); + list_for_each_entry(mempackage, &memory_packages, list) { + if (mempackage->package_id =3D=3D package_id) { + mutex_unlock(&memory_package_lock); + return mempackage; + } + } + mutex_unlock(&memory_package_lock); + + mempackage =3D create_memory_package(package_id); + if (IS_ERR(mempackage)) + return mempackage; + + mutex_lock(&memory_package_lock); + list_for_each_entry(existing, &memory_packages, list) { + if (existing->package_id =3D=3D package_id) { + mutex_unlock(&memory_package_lock); + put_device(&mempackage->dev); + return existing; + } + } + list_add(&mempackage->list, &memory_packages); + mutex_unlock(&memory_package_lock); + + ret =3D device_add(&mempackage->dev); + if (ret) { + mutex_lock(&memory_package_lock); + list_del(&mempackage->list); + mutex_unlock(&memory_package_lock); + put_device(&mempackage->dev); + return ERR_PTR(ret); + } + + return mempackage; +} + +static void mp_snapshot_rebuild(void) +{ + struct mp_snapshot *new, *old; + struct memory_package *mp; + int nr =3D 0, i =3D 0, nid; + + lockdep_assert_held(&memory_package_lock); + + list_for_each_entry(mp, &memory_packages, list) + nr++; + + new =3D kvzalloc(struct_size(new, pkg, nr), GFP_KERNEL); + if (!new) + return; + + memset(new->pkg_of, 0xff, sizeof(new->pkg_of)); + + list_for_each_entry(mp, &memory_packages, list) { + new->pkg[i].nodes =3D mp->nodes; + new->pkg[i].cpu_nodes =3D mp->cpu_nodes; + new->pkg[i].memory_only_nodes =3D mp->memory_only_nodes; + for_each_node_mask(nid, mp->nodes) + new->pkg_of[nid] =3D i; + i++; + } + new->nr_packages =3D nr; + + old =3D rcu_replace_pointer(mp_snapshot, new, + lockdep_is_held(&memory_package_lock)); + if (old) + kvfree_rcu(old, rcu); +} + +static int bind_node_to_package(int nid) +{ + int package_id, pkg_id; + struct memory_package *mp; + nodemask_t nodes, cpu, mem; + + mutex_lock(&memory_package_lock); + if (!mpns[nid]) { + mutex_unlock(&memory_package_lock); + return -EINVAL; + } + package_id =3D mpns[nid]->package_id; + mutex_unlock(&memory_package_lock); + + mp =3D find_create_memory_package(package_id); + if (IS_ERR(mp)) + return PTR_ERR(mp); + + mutex_lock(&memory_package_lock); + if (!mpns[nid]) { + mutex_unlock(&memory_package_lock); + return -ENOENT; + } + mpns[nid]->package =3D mp; + node_set(mpns[nid]->nid, mp->nodes); + if (node_is_memory_only(mpns[nid]->nid)) { + node_set(mpns[nid]->nid, mp->memory_only_nodes); + list_add(&mpns[nid]->package_entry, &mp->memory_only_list); + } else { + node_set(mpns[nid]->nid, mp->cpu_nodes); + list_add(&mpns[nid]->package_entry, &mp->cpu_list); + } + update_package_preferred(mp); + mp_snapshot_rebuild(); + pkg_id =3D mp->package_id; + nodes =3D mp->nodes; + cpu =3D mp->cpu_nodes; + mem =3D mp->memory_only_nodes; + mutex_unlock(&memory_package_lock); + + pr_info("memory_package %d: nodes=3D%*pbl cpu=3D%*pbl memory_only=3D%*pbl= \n", + pkg_id, nodemask_pr_args(&nodes), + nodemask_pr_args(&cpu), nodemask_pr_args(&mem)); + + return 0; +} + +static void unbind_node_to_package(struct memory_package *mp, int nid) +{ + lockdep_assert_held(&memory_package_lock); + + node_clear(nid, mp->nodes); + if (node_state(nid, N_CPU)) + node_clear(nid, mp->cpu_nodes); + else + node_clear(nid, mp->memory_only_nodes); + + if (mpns[nid]) + list_del(&mpns[nid]->package_entry); + + update_package_preferred(mp); +} + +static struct memory_package_node *create_package_node(int nid, int initia= tor_nid) +{ + int cpu_nid, package_id; + int source_flags; + struct memory_package_node *mpn; + + if (node_state(nid, N_CPU)) { + cpu_nid =3D nid; + source_flags =3D MPN_SRC_CPU; + } else { + if (initiator_nid >=3D 0) { + cpu_nid =3D initiator_nid; + source_flags =3D MPN_SRC_INITIATOR; + } else { + /* + * No driver-supplied initiator: fall back to the + * nearest CPU node (via SLIT/numa_distance). + */ + cpu_nid =3D numa_nearest_node(nid, N_CPU); + source_flags =3D MPN_SRC_SLIT; + } + } + + package_id =3D mp_node_to_package_id(cpu_nid); + if (package_id < 0) + return ERR_PTR(-EINVAL); + + mpn =3D kzalloc_obj(*mpn); + if (!mpn) + return ERR_PTR(-ENOMEM); + + mpn->nid =3D nid; + mpn->initiator_nid =3D cpu_nid; + mpn->package_id =3D package_id; + mpn->source_flags =3D source_flags; + mpn->preferred =3D NODE_MASK_NONE; + mpn->package =3D NULL; + INIT_LIST_HEAD(&mpn->package_entry); + + return mpn; +} + +/* + * Topology symmetry status + * Indicates whether all packages have identical node structure + * (same number of CPU nodes and memory-only nodes). + */ +static bool topology_symmetric; + +static void validate_topology_symmetry(void); + +static struct memory_package *__destroy_package_node(int nid) +{ + struct memory_package_node *mpn; + struct memory_package *mp, *unreg_mp =3D NULL; + + lockdep_assert_held(&memory_package_lock); + + mpn =3D mpns[nid]; + if (!mpn) + return NULL; + + mp =3D mpn->package; + if (mp) { + unbind_node_to_package(mp, nid); + mpn->package =3D NULL; + + if (memory_package_is_empty(mp)) { + list_del(&mp->list); + unreg_mp =3D mp; + } + } + + mpns[nid] =3D NULL; + kfree(mpn); + mp_snapshot_rebuild(); + + return unreg_mp; +} + +static void destroy_package_node(int nid) +{ + struct memory_package *unreg_mp; + + mutex_lock(&memory_package_lock); + unreg_mp =3D __destroy_package_node(nid); + mutex_unlock(&memory_package_lock); + + if (unreg_mp) + device_unregister(&unreg_mp->dev); + + validate_topology_symmetry(); +} + +static int find_package_node(int nid, int initiator_nid) +{ + struct memory_package *unreg_mp =3D NULL; + int ret =3D nid; + + mutex_lock(&memory_package_lock); + if (!mpns[nid]) { + ret =3D NUMA_NO_NODE; + } else if (mpns[nid]->source_flags =3D=3D MPN_SRC_SLIT && initiator_nid >= =3D 0) { + /* + * SLIT-derived entries are provisional; if a driver later + * provides an explicit initiator, drop the provisional + * entry and rebuild with the stronger hint. + */ + unreg_mp =3D __destroy_package_node(nid); + ret =3D NUMA_NO_NODE; + } + mutex_unlock(&memory_package_lock); + + if (unreg_mp) + device_unregister(&unreg_mp->dev); + + return ret; +} + +static int find_create_package_node(int nid, int initiator_nid) +{ + int mpn_nid; + struct memory_package_node *mpn; + + mpn_nid =3D find_package_node(nid, initiator_nid); + if (mpn_nid !=3D NUMA_NO_NODE) + return mpn_nid; + + mpn =3D create_package_node(nid, initiator_nid); + if (IS_ERR(mpn)) + return PTR_ERR(mpn); + + guard(mutex)(&memory_package_lock); + if (mpns[nid]) { + kfree(mpn); + return nid; + } + mpns[nid] =3D mpn; + + return nid; +} + +static int create_node_with_package(int nid) +{ + int ret; + + ret =3D find_create_package_node(nid, NUMA_NO_NODE); + if (ret < 0) + return ret; + + ret =3D bind_node_to_package(nid); + if (ret) + return ret; + + validate_topology_symmetry(); + return 0; +} + +/** + * mp_add_package_node_by_initiator - Add a node with an initiator + * @nid: Target NUMA node to add. + * @initiator_nid: CPU nid used to resolve @nid's package (>=3D0). + * + * Ensures that a &struct memory_package_node exists for @nid and that its + * package_id is determined using @initiator_nid when provided. Binding to= the + * package is not performed here. + * + * Return: 0 on success; negative errno on failure. + */ +int mp_add_package_node_by_initiator(int nid, int initiator_nid) +{ + int ret; + + ret =3D find_create_package_node(nid, initiator_nid); + if (ret < 0) + return ret; + + return 0; +} +EXPORT_SYMBOL_GPL(mp_add_package_node_by_initiator); + +/** + * mp_add_package_node - Add a node, resolving package automatically + * @nid: Target NUMA node to add. + * + * Wrapper over mp_add_package_node_by_initiator() that requests automatic + * initiator resolution (e.g., nearest CPU). + * + * Return: 0 on success; negative errno on failure. + */ +int mp_add_package_node(int nid) +{ + return mp_add_package_node_by_initiator(nid, NUMA_NO_NODE); +} +EXPORT_SYMBOL_GPL(mp_add_package_node); + +static int __mp_get_package_nodemask(int nid, enum mp_nodes_type node_type, + nodemask_t *out) +{ + struct mp_snapshot *snap; + int pkg; + + if (!out) + return -EINVAL; + + nodes_clear(*out); + + if (nid < 0 || nid >=3D MAX_NUMNODES) + return -EINVAL; + + guard(rcu)(); + + snap =3D rcu_dereference(mp_snapshot); + if (!snap) + return -ENOENT; + + pkg =3D snap->pkg_of[nid]; + if (pkg < 0) + return -ENOENT; + + switch (node_type) { + case MP_NODES_ALL: + nodes_copy(*out, snap->pkg[pkg].nodes); + break; + case MP_NODES_CPU: + nodes_copy(*out, snap->pkg[pkg].cpu_nodes); + break; + case MP_NODES_MEM_ONLY: + nodes_copy(*out, snap->pkg[pkg].memory_only_nodes); + break; + default: + return -EINVAL; + } + + return 0; +} + +#ifdef CONFIG_NUMA_MIGRATION +static int __mp_get_preferred_nodemask(int nid, enum mp_nodes_type node_ty= pe, + nodemask_t *out) +{ + int ret =3D 0; + + /* No hot-path callers: the mutex is fine here. */ + guard(mutex)(&memory_package_lock); + + if (!out) { + ret =3D -EINVAL; + goto out; + } + + nodes_clear(*out); + + if (nid < 0 || nid >=3D MAX_NUMNODES) { + ret =3D -EINVAL; + goto out; + } + + if (node_type =3D=3D MP_NODES_CPU) { + if (node_is_memory_only(nid)) { + ret =3D -EINVAL; + goto out; + } + } else if (node_type =3D=3D MP_NODES_MEM_ONLY) { + if (!node_is_memory_only(nid)) { + ret =3D -EINVAL; + goto out; + } + } else { + ret =3D -EINVAL; + goto out; + } + + if (!package_node_is_valid(nid)) { + ret =3D -ENOENT; + goto out; + } + + nodes_copy(*out, mpns[nid]->preferred); + +out: + return ret; +} + +/** + * mp_next_demotion_nodemask - Demotion candidates within a package + * @nid: CPU node from which memory would be demoted. + * @out: Output nodemask of nearest memory-only targets in the same packag= e. + * + * Return: 0 on success; negative errno if @nid is invalid or not initiali= zed. + */ +int mp_next_demotion_nodemask(int nid, nodemask_t *out) +{ + return __mp_get_preferred_nodemask(nid, MP_NODES_CPU, out); +} +EXPORT_SYMBOL_GPL(mp_next_demotion_nodemask); + +/** + * mp_next_demotion_node - Pick one demotion target + * @nid: CPU node from which memory would be demoted. + * + * Picks one target (random among the nearest) from mp_next_demotion_nodem= ask(). + * + * Return: target nid on success, or NUMA_NO_NODE if no candidate is avail= able. + */ +int mp_next_demotion_node(int nid) +{ + int target_nid; + nodemask_t target_nodemask; + + if (mp_next_demotion_nodemask(nid, &target_nodemask)) + return NUMA_NO_NODE; + if (nodes_empty(target_nodemask)) + return NUMA_NO_NODE; + + target_nid =3D node_random(&target_nodemask); + + return target_nid; +} +EXPORT_SYMBOL_GPL(mp_next_demotion_node); + +/** + * mp_next_promotion_nodemask - Promotion candidates within a package + * @nid: Memory-only node towards which promotion seeks CPU locality. + * @out: Output nodemask of nearest CPU targets in the same package. + * + * Return: 0 on success; negative errno if @nid is invalid or not initiali= zed. + */ +int mp_next_promotion_nodemask(int nid, nodemask_t *out) +{ + return __mp_get_preferred_nodemask(nid, MP_NODES_MEM_ONLY, out); +} +EXPORT_SYMBOL_GPL(mp_next_promotion_nodemask); + +/** + * mp_next_promotion_node - Pick one promotion target + * @nid: Memory-only node to be promoted towards CPUs. + * + * Picks one target (random among the nearest) from mp_next_promotion_node= mask(). + * + * Return: target nid on success, or NUMA_NO_NODE if no candidate is avail= able. + */ +int mp_next_promotion_node(int nid) +{ + int target_nid; + nodemask_t target_nodemask; + + if (mp_next_promotion_nodemask(nid, &target_nodemask)) + return NUMA_NO_NODE; + if (nodes_empty(target_nodemask)) + return NUMA_NO_NODE; + + target_nid =3D node_random(&target_nodemask); + + return target_nid; +} +EXPORT_SYMBOL_GPL(mp_next_promotion_node); +#endif /* CONFIG_NUMA_MIGRATION */ + +/** + * mp_get_package_nodes - Return all members of @nid's package + * @nid: Any NUMA node in the package. + * @out: Output nodemask to receive all members. + * + * Return: 0 on success; negative errno if @nid is invalid or not initiali= zed. + */ +int mp_get_package_nodes(int nid, nodemask_t *out) +{ + return __mp_get_package_nodemask(nid, MP_NODES_ALL, out); +} +EXPORT_SYMBOL_GPL(mp_get_package_nodes); + +/** + * mp_get_package_cpu_nodes - Return CPU members of @nid's package + * @nid: Any NUMA node in the package. + * @out: Output nodemask to receive CPU members. + * + * Return: 0 on success; negative errno if @nid is invalid or not initiali= zed. + */ +int mp_get_package_cpu_nodes(int nid, nodemask_t *out) +{ + return __mp_get_package_nodemask(nid, MP_NODES_CPU, out); +} +EXPORT_SYMBOL_GPL(mp_get_package_cpu_nodes); + +/** + * mp_get_package_memory_only_nodes - Return memory-only members of @nid's= package + * @nid: Any NUMA node in the package. + * @out: Output nodemask to receive memory-only members. + * + * Return: 0 on success; negative errno if @nid is invalid or not initiali= zed. + */ +int mp_get_package_memory_only_nodes(int nid, nodemask_t *out) +{ + return __mp_get_package_nodemask(nid, MP_NODES_MEM_ONLY, out); +} +EXPORT_SYMBOL_GPL(mp_get_package_memory_only_nodes); + +static int __meminit mp_hotplug_callback(struct notifier_block *nb, + unsigned long action, void *_arg) +{ + int nid; + struct node_notify *nn =3D _arg; + + nid =3D nn->nid; + if (nid < 0) + return notifier_from_errno(0); + + switch (action) { + case NODE_REMOVED_LAST_MEMORY: + destroy_package_node(nid); + break; + + case NODE_ADDED_FIRST_MEMORY: + create_node_with_package(nid); + break; + + default: + break; + } + + return notifier_from_errno(0); +} + +/* + * sysfs interface for memory_package topology + * Read-only attributes: + * - package_nodes: All NUMA nodes in this package (node mask) + * - package_cpu_nodes: CPU nodes in this package (node mask) + * - package_mem_only_nodes: Memory-only nodes in this package (node mas= k) + * - physical_package_id: Physical package ID + */ + +static ssize_t package_nodes_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct memory_package *mp =3D dev_get_drvdata(dev); + + guard(mutex)(&memory_package_lock); + return sysfs_emit(buf, "%*pbl\n", nodemask_pr_args(&mp->nodes)); +} +static DEVICE_ATTR_RO(package_nodes); + +static ssize_t package_cpu_nodes_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct memory_package *mp =3D dev_get_drvdata(dev); + + guard(mutex)(&memory_package_lock); + return sysfs_emit(buf, "%*pbl\n", nodemask_pr_args(&mp->cpu_nodes)); +} +static DEVICE_ATTR_RO(package_cpu_nodes); + +static ssize_t package_mem_only_nodes_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct memory_package *mp =3D dev_get_drvdata(dev); + + guard(mutex)(&memory_package_lock); + return sysfs_emit(buf, "%*pbl\n", nodemask_pr_args(&mp->memory_only_nodes= )); +} +static DEVICE_ATTR_RO(package_mem_only_nodes); + +static ssize_t physical_package_id_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct memory_package *mp =3D dev_get_drvdata(dev); + + return sysfs_emit(buf, "%d\n", mp->package_id); +} +static DEVICE_ATTR_RO(physical_package_id); + +static struct attribute *memory_package_attrs[] =3D { + &dev_attr_package_nodes.attr, + &dev_attr_package_cpu_nodes.attr, + &dev_attr_package_mem_only_nodes.attr, + &dev_attr_physical_package_id.attr, + NULL, +}; + +static const struct attribute_group memory_package_group =3D { + .attrs =3D memory_package_attrs, +}; + +static const struct attribute_group *memory_package_groups[] =3D { + &memory_package_group, + NULL, +}; + +static int __init memory_package_sysfs_init(void) +{ + int ret; + + ret =3D subsys_system_register(&package_subsys, NULL); + if (ret) { + pr_err("memory_package subsys_system_register failed: %d\n", ret); + return ret; + } + return 0; +} +core_initcall(memory_package_sysfs_init); + +/** + * mp_is_topology_symmetric - Check if all packages have identical node st= ructure + * + * Returns true if all memory packages have the same number of CPU nodes + * and memory-only nodes, indicating a symmetric topology. + * + * The topology_symmetric variable is protected by memory_package_lock for + * writes (in validate_topology_symmetry). + */ +bool mp_is_topology_symmetric(void) +{ + return READ_ONCE(topology_symmetric); +} +EXPORT_SYMBOL_GPL(mp_is_topology_symmetric); + +/* Walk the package list and compare node counts; the caller holds the loc= k. */ +static bool __check_topology_symmetry(void) +{ + struct memory_package *mp; + int ref_cpu_node_count =3D 0, ref_memory_only_node_count =3D 0; + int package_count =3D 0; + + lockdep_assert_held(&memory_package_lock); + + list_for_each_entry(mp, &memory_packages, list) { + int cpu_node_count =3D nodes_weight(mp->cpu_nodes); + int memory_only_node_count =3D nodes_weight(mp->memory_only_nodes); + int total_node_count =3D cpu_node_count + memory_only_node_count; + + package_count++; + + /* Each package must have at least 2 nodes */ + if (total_node_count < 2) + return false; + + if (package_count =3D=3D 1) { + ref_cpu_node_count =3D cpu_node_count; + ref_memory_only_node_count =3D memory_only_node_count; + } else if (cpu_node_count !=3D ref_cpu_node_count || + memory_only_node_count !=3D ref_memory_only_node_count) { + return false; + } + } + + if (package_count < 2) + return false; + + return true; +} + +/** + * validate_topology_symmetry - Validate topology for package-aware featur= es + * + * Validates topology symmetry by checking all memory packages have identi= cal + * node structure. Updates the global topology_symmetric variable. + * + * Topology is valid only when all conditions are met: + * 1. At least 2 packages exist (multi-package system) + * 2. All packages have identical cpu_node and memory_only_node counts + * 3. Each package has at least 2 total nodes (cpu_node + memory_only_node= >=3D 2) + * + * Valid configurations per package: + * - 1 cpu_node + 1 memory_only_node =3D 2 nodes (valid) + * - 2 cpu_nodes + 0 memory_only_node =3D 2 nodes (valid) + * - 0 cpu_node + 2 memory_only_nodes =3D 2 nodes (valid) + * - 1 cpu_node + 0 memory_only_node =3D 1 node (invalid - single node pac= kage) + */ +static void validate_topology_symmetry(void) +{ + guard(mutex)(&memory_package_lock); + + WRITE_ONCE(topology_symmetric, __check_topology_symmetry()); +} + +static int __init memory_package_init(void) +{ + int ret =3D 0, nid; + + for_each_online_node(nid) { + if (!node_state(nid, N_MEMORY)) + continue; + + ret =3D create_node_with_package(nid); + if (ret) + goto out; + } + + hotplug_node_notifier(mp_hotplug_callback, MEMTIER_HOTPLUG_PRI); + + validate_topology_symmetry(); + +out: + return ret; +} +late_initcall(memory_package_init); --=20 2.25.1 From nobody Fri Oct 2 02:30:25 2026 Received: from invmail4.hynix.com (exvmail4.hynix.com [166.125.252.92]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 1C37D3EEADF; Thu, 6 Aug 2026 08:25:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=166.125.252.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004713; cv=none; b=KylWTg8xcr4d9zQmknLCPTTWdnhGjBN5LChnfgMyDXo5cwSK9EhUF5CyBvcfQxy4Pm038qKZl/I5PVvU1SQDtXgBf28fzmFbR+EeMKBsesaKQ6owtdKb1OX1zQ/MX6Qwl4ob7HrXLKlYFpNvnV38rR5/VjYtZlIjTe8I7v/RgQ4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004713; c=relaxed/simple; bh=Np0F5SCXAYL/AAIQCaj0TJNePNy+UGo+UDdYyyTubLo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o1Y1I45C0PnwxfqhelLgIvI62haQ/UgAeJCahI1vU1YFS9GT+t7rWOcz/I4vMj077sePRXsZC529jDGt3JxJlShPKIBuNJHvCIWcedrXDlhis/JyEkEI5ereYr8RmwM3ugLuendZKrkBys0nIXymZA/GovwQplyDA6IsmSpQYIo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com; spf=pass smtp.mailfrom=sk.com; arc=none smtp.client-ip=166.125.252.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sk.com X-AuditID: a67dfc5b-c45ff70000001609-a9-6a74414d29b5 From: Rakie Kim To: akpm@linux-foundation.org Cc: gourry@gourry.net, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-cxl@vger.kernel.org, nvdimm@lists.linux.dev, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, byungchul@sk.com, ying.huang@linux.alibaba.com, apopple@nvidia.com, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, dave@stgolabs.net, jic23@kernel.org, dave.jiang@intel.com, alison.schofield@intel.com, vishal.l.verma@intel.com, ira.weiny@intel.com, harry@kernel.org, kernel_team@skhynix.com, honggyu.kim@sk.com, yunjeong.mun@sk.com, rakie.kim@sk.com Subject: [PATCH 3/4] mm/memory-tiers: register CXL nodes to memory packages via initiator Date: Thu, 6 Aug 2026 17:09:34 +0900 Message-ID: <20260806080936.421-4-rakie.kim@sk.com> X-Mailer: git-send-email 2.52.0.windows.1 In-Reply-To: <20260806080936.421-1-rakie.kim@sk.com> References: <20260806080936.421-1-rakie.kim@sk.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrHIsWRmVeSWpSXmKPExsXC9ZZnoa6vY0mWQf8lEYs569ewWdx9fIHN YteNEIsTNxvZLFbfXMNo8XzrL0aLn3ePs1tcv7WS0WL/0+csFg+aVjFZHN86j91i3alDbBbn Z51isbi8aw6bxb01/1kt3jx2s/jWJ21xv8/BYuWPP6wWR9ZvZ7KYfGkBm0XHy/ssFrcmHGOy WL0mw2L20XvsDpIeO2fdZfdYsKnUo7vtMrvH5hVaHov3vGTy2LSqk81j06dJ7B4nZvxm8dj5 0NLjxeaZjB69ze/YPKbOrvdYv+Uqi8fnTXIBfFFcNimpOZllqUX6dglcGXPen2YrOKFcceDN e9YGxm2yXYycHBICJhLLP25ggbF73s1h72Lk4GATUJI4tjcGJCwiICsx9e95oBIuDmaBRawS Jz6fZQVJCAtESmy6cIkdxGYRUJX4+24q2BxeAWOJFYu72CBmakqs23gLLM4JNP/7nUVgthBQ zZ9f29gh6gUlTs58AhZnFpCXaN46mxlkmYTAV3aJv/t2MEIMkpQ4uOIGywRG/llIemYh6VnA yLSKUSgzryw3MTPHRC+jMi+zQi85P3cTIzBWl9X+id7B+OlC8CFGAQ5GJR7eC8bFWUKsiWXF lbmHGCU4mJVEeFkPFmUJ8aYkVlalFuXHF5XmpBYfYpTmYFES5zX6Vp4iJJCeWJKanZpakFoE k2Xi4JRqYOyIt22vX3r6u1qfRktVcUqxSOvZlU91bYofvsp79uPqwavM145c3J250ITvjWtK S+yKSRbyD9uetvTVteSrHHgQ88cm5skmxtnPvzbvWKK8/rn3P+Vrs60PegYbHbvQnNdacdfz 9ras82HNrvs9tU7uXPBpzzwJxm8aL9RiLVepHahhEykoLldiKc5INNRiLipOBACQgDsF0QIA AA== X-Brightmail-Tracker: H4sIAAAAAAAAA02Ra0hTYQCG+c7ObcPRaQmdNBVWUQ1Sw4ovkNQ/9aERClmZUq48temcujl1 galJUivFLEPdClGwqUtxpqkk1UznhdK0Mi0v4T1FrYWZF8opgf9enufl/fPSPFEj7kTLlfGc SilViEkBLsjbG3bgpF98pOeHZm9oqDSRcGCki4QNn0/Ddw/0JGztSyNheZ8JwImaJQD/DFgp 2NtfCqBtfJoHX45N4HD4RhkGrTWPKdj0qI2AFe0WEnYWtOOwp8FAwkHTXwLOjByHC1nOcCjL F5YurhDQ8mmCgG8qn2PwfnchCW9NDeGwP7sFg+UmGVyuNZJQ3zxI+bqi+oIBChWaNehORg+F qo0SVPxiCkPmstskMv/MoVBr3jKO6r8dRZPV+QBlps+SaOELQsWT8xjK1aegymcfcWQzuwYy 5wXeEZxCnsCpPI6FC2SGuQ4ytnVX0quZOSIV1LroAJ9mmUPs3VkDpQM0TTJitqUxzI4dGRc2 d7UT1wEBzWOKCLbV9pawi21MCGvu6qbsGWf2sKuzubg9Cxkv1lisIzc297MVVf3rnL+2//tr 0XoWrXVWlmqpjf5Wti1/dJ3zGDc2vUbPywYOBZtUwSZVCLAy4ChXJkRL5YrD7uoomVYpT3K/ HBNtBmunliSv3KsDv3pOWABDA7GDsMtLHSkipAlqbbQFsDRP7CgkXqsiRcIIqfYap4q5qNIo OLUFONO4eLvQ/ywXLmKuSuO5KI6L5VT/LUbznVJBTsCYNaQ/qDh4HF+Ojm9Ql+j2zQelhdY1 BV+56XN9y/vEuN07+gZ2Wj1Hp4Iv8VOeQsqKP1wIGD7SJEDfM2Knp23ZHVGSxT4PIHE9U4WU yacexbkTg7L6GUnooLbXHxtyV5zT5LOBiakuAU9+yP1yL6g1LjU+TQlDRreIzDkxrpZJD0p4 KrX0H1WBXsDQAgAA X-CFilter-Loop: Reflected Content-Type: text/plain; charset="utf-8" A CXL memory node comes online without an explicit package association, and plain NUMA distance does not convey which physical package it belongs to. Without that association a CXL node cannot be grouped with the CPUs that front it. Register a package notifier per CXL region. When the region's memory node comes online, the notifier resolves an initiator CPU node - the NUMA node of the first memdev backing the region - and binds the memory node to that initiator's package. This gives the topology layer the CPU-side association that plain NUMA distance does not carry. The initiator nid comes from the firmware and driver description of the region's endpoint, so the association is only as accurate as that description; it gives a more direct host-side grouping than flat distance values. Signed-off-by: Rakie Kim --- drivers/cxl/core/region.c | 54 +++++++++++++++++++++++++++++++++++++++ drivers/cxl/cxl.h | 1 + drivers/dax/kmem.c | 3 +++ 3 files changed, 58 insertions(+) diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c index e50dc716d4e8..af66e2e06c62 100644 --- a/drivers/cxl/core/region.c +++ b/drivers/cxl/core/region.c @@ -2673,6 +2673,55 @@ static int cxl_region_calculate_adistance(struct not= ifier_block *nb, return NOTIFY_STOP; } =20 +/* + * Find a NUMA node to act as the initiator for this region: scan the + * region's endpoint targets and return the first one that resolves to a + * valid NUMA node. + */ +static int cxl_region_find_nearest_node(struct cxl_region *cxlr) +{ + struct cxl_region_params *p =3D &cxlr->params; + struct cxl_endpoint_decoder *cxled =3D NULL; + struct cxl_memdev *cxlmd =3D NULL; + int i, numa_node; + + for (i =3D 0; i < p->nr_targets; i++) { + cxled =3D p->targets[i]; + cxlmd =3D cxled_to_memdev(cxled); + numa_node =3D dev_to_node(&cxlmd->dev); + if (numa_node !=3D NUMA_NO_NODE) + return numa_node; + } + return NUMA_NO_NODE; +} + +/* + * Package notifier callback: when a new memory node is onlined via dax + * kmem, bind the node this CXL region backs to its memory package, using + * the nearest region target as the initiator. Notifications for other + * nodes are ignored. + */ +static int cxl_region_add_package_node(struct notifier_block *nb, + unsigned long dax_nid, void *data) +{ + int region_nid, nearest_nid, ret; + struct cxl_region *cxlr =3D container_of(nb, struct cxl_region, package_n= otifier); + + region_nid =3D phys_to_target_node(cxlr->params.res->start); + if (region_nid !=3D dax_nid) + return NOTIFY_DONE; + + nearest_nid =3D cxl_region_find_nearest_node(cxlr); + if (nearest_nid =3D=3D NUMA_NO_NODE) + return NOTIFY_DONE; + + ret =3D mp_add_package_node_by_initiator(dax_nid, nearest_nid); + if (ret) + return NOTIFY_DONE; + + return NOTIFY_OK; +} + /** * devm_cxl_add_region - Adds a region to a decoder * @cxlrd: root decoder @@ -3852,6 +3901,7 @@ static void shutdown_notifiers(void *_cxlr) =20 unregister_node_notifier(&cxlr->node_notifier); unregister_mt_adistance_algorithm(&cxlr->adist_notifier); + unregister_mp_package_notifier(&cxlr->package_notifier); } =20 static void remove_debugfs(void *dentry) @@ -4066,6 +4116,10 @@ static int cxl_region_probe(struct device *dev) cxlr->adist_notifier.priority =3D 100; register_mt_adistance_algorithm(&cxlr->adist_notifier); =20 + cxlr->package_notifier.notifier_call =3D cxl_region_add_package_node; + cxlr->package_notifier.priority =3D 100; + register_mp_package_notifier(&cxlr->package_notifier); + rc =3D devm_add_action_or_reset(&cxlr->dev, shutdown_notifiers, cxlr); if (rc) return rc; diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h index 1297594beaec..9281ca3a5f0f 100644 --- a/drivers/cxl/cxl.h +++ b/drivers/cxl/cxl.h @@ -477,6 +477,7 @@ struct cxl_region { struct access_coordinate coord[ACCESS_COORDINATE_MAX]; struct notifier_block node_notifier; struct notifier_block adist_notifier; + struct notifier_block package_notifier; }; =20 struct cxl_nvdimm_bridge { diff --git a/drivers/dax/kmem.c b/drivers/dax/kmem.c index 2cc8749bc871..1de23f196354 100644 --- a/drivers/dax/kmem.c +++ b/drivers/dax/kmem.c @@ -94,6 +94,9 @@ static int dev_dax_kmem_probe(struct dev_dax *dev_dax) if (IS_ERR(mtype)) return PTR_ERR(mtype); =20 + /* Resolve the memory package for this newly onlined kmem node. */ + mp_probe_package_id(numa_node); + for (i =3D 0; i < dev_dax->nr_range; i++) { struct range range; =20 --=20 2.25.1 From nobody Fri Oct 2 02:30:25 2026 Received: from invmail4.hynix.com (exvmail4.hynix.com [166.125.252.92]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 157AA3F0A96; Thu, 6 Aug 2026 08:25:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=166.125.252.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004719; cv=none; b=p/cWssh1va6Por//rYUOnFMNUn3nSJQCBDkDP4ZUgZSJuVL/zkBFgcLBPhl59xTfG4pcSrJMn7KkNhvKlBVM3EhSw0urbb7d8PA89A39UstcwRPc85fGGiCMlbtqNYj6uWJn3rywdXnyRSPgD9YB74zENLGLJhchJnSL9/1DY+4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786004719; c=relaxed/simple; bh=AuxpUL6xLeUPvp7JFNvkxOkrv6s/iFERNBvbv+WSVJ4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TMAS4CCuHUAjU2yIs3DLXgF7+s6xQUZPVwRqVtDG/pTjReqqz0vkaFHzyhOoAohHGTMbV0gtbMmu82Htnue1qP3qrZClPOM0incMc/ILoZWrEkT3TKz+I2tZyxHPoJl/EO8G8T7p5nq6Wbp5niuY6a9zMK1KwZwuxafanaCxadw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com; spf=pass smtp.mailfrom=sk.com; arc=none smtp.client-ip=166.125.252.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sk.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sk.com X-AuditID: a67dfc5b-c45ff70000001609-b2-6a74414ee24e From: Rakie Kim To: akpm@linux-foundation.org Cc: gourry@gourry.net, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-cxl@vger.kernel.org, nvdimm@lists.linux.dev, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, byungchul@sk.com, ying.huang@linux.alibaba.com, apopple@nvidia.com, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, dave@stgolabs.net, jic23@kernel.org, dave.jiang@intel.com, alison.schofield@intel.com, vishal.l.verma@intel.com, ira.weiny@intel.com, harry@kernel.org, kernel_team@skhynix.com, honggyu.kim@sk.com, yunjeong.mun@sk.com, rakie.kim@sk.com Subject: [PATCH 4/4] mm/mempolicy: enhance weighted interleave with package-aware locality Date: Thu, 6 Aug 2026 17:09:35 +0900 Message-ID: <20260806080936.421-5-rakie.kim@sk.com> X-Mailer: git-send-email 2.52.0.windows.1 In-Reply-To: <20260806080936.421-1-rakie.kim@sk.com> References: <20260806080936.421-1-rakie.kim@sk.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrLIsWRmVeSWpSXmKPExsXC9ZZnoa6fY0mWwcNFnBZz1q9hs7j7+AKb xa4bIRYnbjayWay+uYbR4vnWX4wWP+8eZ7e4fmslo8X+p89ZLB40rWKyOL51HrvFulOH2CzO zzrFYnF51xw2i3tr/rNavHnsZvGtT9rifp+Dxcoff1gtjqzfzmQx+dICNouOl/dZLG5NOMZk sXpNhsXso/fYHSQ9ds66y+6xYFOpR3fbZXaPzSu0PBbvecnksWlVJ5vHpk+T2D1OzPjN4rHz oaXHi80zGT16m9+xeUydXe+xfstVFo/Pm+QC+KK4bFJSczLLUov07RK4Mtb3v2QsuFFU0Xdz AlMD48ToLkZODgkBE4mX16czwthX33SydjFycLAJKEkc2xsDEhYRkJWY+vc8SxcjFwezwCJW iROfz4LVCAtESfz/LwBSwyKgKnHoy052EJtXwFhibc9kNoiRmhLrNt5iAbE5gcZ/v7MIzBYC qvnzaxtUvaDEyZlPwOLMAvISzVtnM4PskhD4yS7RvGoiE8QgSYmDK26wTGDkn4WkZxaSngWM TKsYhTLzynITM3NM9DIq8zIr9JLzczcxAiN1We2f6B2Mny4EH2IU4GBU4uG9YFycJcSaWFZc mXuIUYKDWUmEl/VgUZYQb0piZVVqUX58UWlOavEhRmkOFiVxXqNv5SlCAumJJanZqakFqUUw WSYOTqkGxq453pe+xS73a50vetIy7a3lv3f+KR/kJ11qCMh3YMxt2fRWoZ7lyZ1ZH37eLPYU eued8PXBzd5S60sxlimyN/f9ktPM0vvy0mFRW2taerrx5DtafZz7JDZFOPoFz83VD556QbQh 03xl3f43aoedO9blFs7+dT53e7lwXUPYl2nNlY9EZ3MrKrEUZyQaajEXFScCAEQtYk7QAgAA X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrDIsWRmVeSWpSXmKPExsXCNUM9RtfXsSTL4PJ9U4s569ewWdx9fIHN YteNEItzU2azWZy42chmsfrmGkaL51t/MVr8vHuc3eL6rZWMFp+fvWa22P/0OYvFg6ZVTBbH t85jtzg89ySrxbpTh9gszs86xWJxedccNot7a/6zWrx57GbxrU/a4n6fg8XKH39YLQ5de85q cWT9diaLyZcWsFl0vLzPYnFrwjEmi9VrMix+b1vBZjH76D12BzmPnbPusnss2FTq0d12md1j 8wotj8V7XjJ5bFrVyeax6dMkdo8TM36zeOx8aOnxYvNMRo/e5ndsHt9ue3gsfvGByWPq7HqP 9Vuusnh83iQXIBDFZZOSmpNZllqkb5fAlbG+/yVjwY2iir6bE5gaGCdGdzFyckgImEhcfdPJ 2sXIwcEmoCRxbG8MSFhEQFZi6t/zLF2MXBzMAotYJU58PgtWIywQJfH/vwBIDYuAqsShLzvZ QWxeAWOJtT2T2SBGakqs23iLBcTmBBr//c4iMFsIqObPr21Q9YISJ2c+AYszC8hLNG+dzTyB kWcWktQsJKkFjEyrGEUy88pyEzNzTPWKszMq8zIr9JLzczcxAuN0We2fiTsYv1x2P8QowMGo xMN7wbg4S4g1say4MvcQowQHs5IIL+vBoiwh3pTEyqrUovz4otKc1OJDjNIcLErivF7hqQlC AumJJanZqakFqUUwWSYOTqkGximFole3L7D0W+wy48dfljWXvi7Zsyn34qfAzt625deOlZid yNz+jGvGtQ26t/qvt3DXPF06O69s52wtro1f4iPeCMk9rtjiblcy4byk9Pqpvow7OZ1ZOz4K P0kyPqtRoGi/9sMxVmFxzTdS5dmvFJed3qgivfbph+cX08I6vsjqPLp4daoz92QlluKMREMt 5qLiRAA2UOG3zwIAAA== X-CFilter-Loop: Reflected Content-Type: text/plain; charset="utf-8" Weighted interleave places pages on nodes in proportion to per-node weights derived from each node's bandwidth. Within one package the weight given to a node matches the bandwidth a task sees from it; across packages it does not. The weights are set once from device bandwidth and applied the same way wherever the task runs, but memory reached over the interconnect to another package is slower than the same memory reached locally, so a node in another package is given a weight higher than the bandwidth it can deliver to the task. Flat weighted interleave then steers allocations onto the interconnect even when local capacity exists, degrading effective bandwidth. node0 node1 +-------+ +-------+ | CPU 0 |---------| CPU 1 | +-------+ +-------+ | DRAM0 | | DRAM1 | +---+---+ +---+---+ | | +---+---+ +---+---+ | CXL 0 | | CXL 1 | +-------+ +-------+ node2 node3 The numbers below are illustrative single-stream bandwidths (GB/s). Local DRAM sustains 300 and local CXL 150; any path that crosses to another package, over the interconnect, is capped at 100, so a node in another package delivers 100 whether it is DRAM or CXL. Local CXL (150) is still faster than any node in another package (100). The effective bandwidth each CPU sees is: node0 node1 node2 node3 from CPU 0: 300 100 150 100 from CPU 1: 100 300 100 150 Since a single per-node weight cannot encode the interconnect penalty, a reasonable set of global weights is taken from local device bandwidth (local DRAM : local CXL =3D 300 : 150 =3D 2 : 1): node0=3D2 node1=3D2 node2= =3D1 node3=3D1. node0 node1 node2 node3 global: 2 2 1 1 (same wherever the task runs) A task on CPU 0 gives node1 - remote DRAM, effective 100 - the same weight 2 as its own local node0 at 300. Worse, node1 is weighted above node2, the task's local CXL at effective 150, even though node2 is the faster of the two. The flat weights rank a slower interconnect-bound node above a faster local one, which is exactly backwards. Make weighted interleave package-aware. When enabled, node selection is restricted to the nodes of the task's current package, intersected with the policy nodemask: the task's pages are spread by weight across the package's nodes, and nodes outside the package are not part of the selection. The only mask-level fallback is the empty-intersection case - if the policy nodemask excludes every node of the current package, selection falls back to the package spanned by the policy's own nodes, so a misconfiguration never yields an empty candidate set. There is no spill to a remote package as a placement preference. Availability is still preferred over containment at allocation time: the resolved mask constrains node selection only, and the page allocator is invoked without a package nodemask, so when the selected node is exhausted the allocation is served from another node - exactly as plain weighted interleave already behaves - rather than forcing reclaim on the local package. The resolved mask is by construction a subset of the policy nodemask, which mempolicy already restricts to the task's cpuset; package mode can only narrow that set, never widen it, so cpusets and the task nodemask remain authoritative. node0 node1 node2 node3 from CPU 0: 2 0 1 0 from CPU 1: 0 2 0 1 Tasks on CPU 0 place pages on DRAM0(2) and CXL0(1) at 2:1, which matches their effective bandwidth of 300:150; tasks on CPU 1 place on DRAM1(2) and CXL1(1) the same way. This aligns allocation with per-package bandwidth, preserves NUMA locality, and keeps interleave traffic off the saturable cross-socket interconnect. The behavior is opt-in and off by default. A sysfs toggle at /sys/kernel/mm/mempolicy/weighted_interleave/package_mode turns it on or off at runtime; reading it reports the current setting. Enabling is refused on a topology that is not symmetric (see /sys/devices/system/package/), and if the topology stops being symmetric while the mode is on - for example after node hotplug - weighted interleave transparently degrades to its flat behavior while the configured value is preserved and takes effect again once the topology is symmetric. The following are the results with package-aware weighted interleave applied: System Configuration: - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids) 1) Throughput (System Bandwidth) - DRAM Only: 966 GB/s - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only) - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s) (38% increase compared to DRAM Only, 47% increase compared to Weighted Interleave) 2) Loaded Latency (Under High Bandwidth) - DRAM Only: 544 ns - Weighted Interleave: 545 ns - Package-Aware Weighted Interleave: 436 ns (20% reduction compared to both) Signed-off-by: Rakie Kim --- ...fs-kernel-mm-mempolicy-weighted-interleave | 17 ++ mm/mempolicy.c | 159 +++++++++++++++++- 2 files changed, 172 insertions(+), 4 deletions(-) diff --git a/Documentation/ABI/testing/sysfs-kernel-mm-mempolicy-weighted-i= nterleave b/Documentation/ABI/testing/sysfs-kernel-mm-mempolicy-weighted-in= terleave index 649c0e9b895c..d2ccba171c5e 100644 --- a/Documentation/ABI/testing/sysfs-kernel-mm-mempolicy-weighted-interlea= ve +++ b/Documentation/ABI/testing/sysfs-kernel-mm-mempolicy-weighted-interlea= ve @@ -52,3 +52,20 @@ Description: Auto-weighting configuration interface =20 Writing a new weight to a node directly via the nodeN interface will also automatically switch the system to manual mode. + +What: /sys/kernel/mm/mempolicy/weighted_interleave/package_mode +Date: August 2026 +Contact: Linux memory management mailing list +Description: Package-aware weighted interleave toggle + + 'true' restricts weighted interleave node selection to the + NUMA nodes of the package (CPU socket) the allocating task + is running on. 'false' (the default) uses the existing + weighted interleave behavior. + + Enabling is rejected with -EINVAL while the package topology + is not symmetric. + + Writing any true value string (e.g. Y or 1) enables the + restriction, any false value string (e.g. N or 0) disables + it. All other strings return -EINVAL. diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 19417b0afc30..66bccb9a0a19 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -117,6 +117,7 @@ #include #include #include +#include =20 #include "internal.h" =20 @@ -167,6 +168,8 @@ static unsigned int *node_bw_table; */ static DEFINE_MUTEX(wi_state_lock); =20 +static bool package_mode_enabled; + static u8 get_il_weight(int node) { struct weighted_interleave_state *state; @@ -180,6 +183,11 @@ static u8 get_il_weight(int node) return weight; } =20 +static bool wi_package_mode_enabled(void) +{ + return READ_ONCE(package_mode_enabled) && mp_is_topology_symmetric(); +} + /* * Convert bandwidth values into weighted interleave weights. * Call with wi_state_lock. @@ -2138,17 +2146,97 @@ bool apply_policy_zone(struct mempolicy *policy, en= um zone_type zone) return zone >=3D dynamic_policy_zone; } =20 +/** + * policy_resolve_package_nodes - Restrict policy nodes to the current pac= kage + * @policy: Target mempolicy whose user-selected nodes are in @policy->nod= es. + * @mask: Output nodemask. On success, contains policy->nodes limited to + * the package that should be used for the allocation. + * + * This helper combines two constraints to decide where within a package + * memory may be allocated: + * + * 1) The caller's package: derived via mp_get_package_nodes(numa_node_i= d()). + * 2) The user's preselected set @policy->nodes (cpusets/mempolicy). + * + * The function obtains the nodemask of the current CPU's package and + * intersects it with @policy->nodes. If the intersection is empty (e.g. t= he + * user excluded every node of the current package), it falls back to the + * node in @policy->nodes, derives that node's package, and intersects + * again. If the fallback also yields an empty set, @mask stays empty and a + * non-zero error is returned. + * + * Examples (packages: P0=3D{CPU:0, MEM:2}, P1=3D{CPU:1, MEM:3}): + * - policy->nodes =3D {0,1,2,3} + * on P0: mask =3D {0,2}; on P1: mask =3D {1,3}. + * - policy->nodes =3D {0,1,3} + * on P0: mask =3D {0} (only node 0 from P0 is allowed). + * - policy->nodes =3D {1,2,3} + * on P0: mask =3D {2} (only node 2 from P0 is allowed). + * - policy->nodes =3D {1,3} + * on P0: current package (P0) & policy =3D NULL -> fallback to poli= cy=3D1, + * package(1)=3DP1, mask =3D {1,3}. (User effectively opted = out of P0.) + * + * If the selected node is low on memory, the allocation may use another n= ode. + * + * Return: + * 0 on success with @mask set as above; + * -EINVAL if @policy/@mask is NULL; + * -ENOENT if even the fallback intersection is empty; + * Propagated error from mp_get_package_nodes() on failure. + */ +static int policy_resolve_package_nodes(struct mempolicy *policy, nodemask= _t *mask) +{ + nodemask_t package_mask; + int node, ret; + + if (!policy || !mask) + return -EINVAL; + + nodes_clear(*mask); + + node =3D numa_node_id(); + ret =3D mp_get_package_nodes(node, &package_mask); + if (ret) + return ret; + + nodes_and(*mask, package_mask, policy->nodes); + if (!nodes_empty(*mask)) + return 0; + + /* + * The user's nodemask excludes every node of the current package; + * fall back to the package spanned by the user's own first node. + */ + node =3D first_node(policy->nodes); + ret =3D mp_get_package_nodes(node, &package_mask); + if (ret) + return ret; + + nodes_and(*mask, package_mask, policy->nodes); + if (nodes_empty(*mask)) + return -ENOENT; + + return 0; +} + static unsigned int weighted_interleave_nodes(struct mempolicy *policy) { unsigned int node; unsigned int cpuset_mems_cookie; + nodemask_t mask; =20 retry: /* to prevent miscount use tsk->mems_allowed_seq to detect rebind */ cpuset_mems_cookie =3D read_mems_allowed_begin(); node =3D current->il_prev; - if (!current->il_weight || !node_isset(node, policy->nodes)) { - node =3D next_node_in(node, policy->nodes); + + /* Package mode off or unresolved: fall back to the full policy nodemask.= */ + if (!wi_package_mode_enabled() || + policy_resolve_package_nodes(policy, &mask)) + mask =3D policy->nodes; + + if (!current->il_weight || !node_isset(node, mask)) { + node =3D next_node_in(node, mask); if (read_mems_allowed_retry(cpuset_mems_cookie)) goto retry; if (node =3D=3D MAX_NUMNODES) @@ -2241,6 +2329,30 @@ static unsigned int read_once_policy_nodemask(struct= mempolicy *pol, return nodes_weight(*mask); } =20 +/* + * Package-aware counterpart of read_once_policy_nodemask(): resolve the + * current package's nodes intersected with the policy, falling back to the + * full policy nodemask when package mode is off or resolution fails. + */ +static unsigned int read_once_policy_package_nodemask(struct mempolicy *po= l, + nodemask_t *mask) +{ + nodemask_t package_mask; + + barrier(); + if (!wi_package_mode_enabled()) { + memcpy(mask, &pol->nodes, sizeof(nodemask_t)); + return nodes_weight(*mask); + } + if (policy_resolve_package_nodes(pol, &package_mask)) + memcpy(mask, &pol->nodes, sizeof(nodemask_t)); + else + memcpy(mask, &package_mask, sizeof(nodemask_t)); + barrier(); + + return nodes_weight(*mask); +} + static unsigned int weighted_interleave_nid(struct mempolicy *pol, pgoff_t= ilx) { struct weighted_interleave_state *state; @@ -2251,7 +2363,7 @@ static unsigned int weighted_interleave_nid(struct me= mpolicy *pol, pgoff_t ilx) u8 weight; int nid =3D 0; =20 - nr_nodes =3D read_once_policy_nodemask(pol, &nodemask); + nr_nodes =3D read_once_policy_package_nodemask(pol, &nodemask); if (!nr_nodes) return numa_node_id(); =20 @@ -2695,7 +2807,7 @@ static unsigned long alloc_pages_bulk_weighted_interl= eave(gfp_t gfp, /* read the nodes onto the stack, retry if done during rebind */ do { cpuset_mems_cookie =3D read_mems_allowed_begin(); - nnodes =3D read_once_policy_nodemask(pol, &nodes); + nnodes =3D read_once_policy_package_nodemask(pol, &nodes); } while (read_mems_allowed_retry(cpuset_mems_cookie)); =20 /* if the nodemask has become invalid, we cannot do anything */ @@ -3835,7 +3947,42 @@ static struct kobj_attribute wi_auto_attr =3D { .store =3D weighted_interleave_auto_store, }; =20 +static ssize_t package_mode_show(struct kobject *kobj, + struct kobj_attribute *attr, char *buf) +{ + return sysfs_emit(buf, "%s\n", str_true_false(READ_ONCE(package_mode_enab= led))); +} + +static ssize_t package_mode_store(struct kobject *kobj, + struct kobj_attribute *attr, const char *buf, size_t count) +{ + bool input; + int err; + + err =3D kstrtobool(buf, &input); + if (err) + return err; + + /* + * Disable package-aware weighted interleave on non-symmetric topologies. + * Non-symmetric topology (e.g., asymmetric CXL memory attachment) can + * cause performance degradation if package-aware allocation is used. + * Reject enable request if topology is not symmetric. + */ + if (input && !mp_is_topology_symmetric()) { + pr_warn("package_mode cannot be enabled on non-symmetric topology\n"); + return -EINVAL; + } + + WRITE_ONCE(package_mode_enabled, input); + return count; +} + +static struct kobj_attribute wi_package_mode_attr =3D + __ATTR(package_mode, 0664, package_mode_show, package_mode_store); + static void wi_cleanup(void) { + sysfs_remove_file(&wi_group->wi_kobj, &wi_package_mode_attr.attr); sysfs_remove_file(&wi_group->wi_kobj, &wi_auto_attr.attr); sysfs_wi_node_delete_all(); wi_state_free(); @@ -3941,6 +4088,10 @@ static int __init add_weighted_interleave_group(stru= ct kobject *mempolicy_kobj) if (err) goto err_put_kobj; =20 + err =3D sysfs_create_file(&wi_group->wi_kobj, &wi_package_mode_attr.attr); + if (err) + goto err_cleanup_kobj; + for_each_online_node(nid) { if (!node_state(nid, N_MEMORY)) continue; --=20 2.25.1