[PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave

Rakie Kim posted 4 patches 1 month, 3 weeks ago
.../ABI/testing/sysfs-devices-system-package  |   35 +
...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
drivers/cxl/core/region.c                     |   54 +
drivers/cxl/cxl.h                             |    1 +
drivers/dax/kmem.c                            |    3 +
include/linux/memory-tiers.h                  |  113 ++
include/linux/numa.h                          |   11 +
mm/memory-tiers.c                             | 1009 +++++++++++++++++
mm/mempolicy.c                                |  200 +++-
9 files changed, 1439 insertions(+), 4 deletions(-)
create mode 100644 Documentation/ABI/testing/sysfs-devices-system-package
[PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 3 weeks ago
Package-aware weighted interleave places a task's weighted-interleave
pages on the NUMA nodes of its local package, so that interleave traffic
does not have to cross the interconnect to another package. This keeps
each node's weight aligned with the bandwidth the task actually gets
from it, so effective bandwidth holds up on a system that has more than
one package. (A package is a CPU socket together with the memory
attached to it.)

Changes from RFC:
https://lore.kernel.org/all/20260316051258.246-1-rakie.kim@sk.com/
- Added an opt-in sysfs toggle (off by default) and a read-only sysfs
  view of the package topology
- Added topology validation with a clean fallback to plain weighted
  interleave on unsupported topologies
- Hardened the allocation, device-teardown, and node-hotplug paths

Weighted interleave places pages on nodes in proportion to per-node
weights that are set from each node's bandwidth. Within one package the
weight given to a node matches the bandwidth a task sees from it. Across
packages it no longer does: a memory node's physical bandwidth is fixed,
but the bandwidth a task effectively sees depends on which package its
CPU is in, because memory reached from another package, over the
interconnect between them, is slower than the same memory reached within
the package. The weights are set once from device bandwidth and applied
the same way wherever the task runs, so a node in another package is
given a weight higher than the bandwidth it can deliver to that task.
The kernel has no package abstraction and does not record which package
a node belongs to, so it cannot tell which node pairs are separated by
the interconnect.

          node0             node1
        +-------+         +-------+
        | CPU 0 |---------| CPU 1 |
        +-------+         +-------+
        | DRAM0 |         | DRAM1 |
        +---+---+         +---+---+
            |                 |
        +---+---+         +---+---+
        | CXL 0 |         | CXL 1 |
        +-------+         +-------+
          node2             node3

The numbers below are illustrative single-stream bandwidths (GB/s).
Local DRAM sustains 300 and local CXL 150; any path that crosses to
another package, over the interconnect, is capped at 100, so a node in
another package delivers 100 whether it is DRAM or CXL. Note that local
CXL (150) is still faster than any node in another package (100). The
effective bandwidth each CPU sees is therefore:

              node0  node1  node2  node3
from CPU 0:    300    100    150    100
from CPU 1:    100    300    100    150

Since a single per-node weight cannot encode the interconnect penalty,
a reasonable set of global weights is taken from local device bandwidth
(local DRAM : local CXL = 300 : 150 = 2 : 1): node0=2 node1=2 node2=1
node3=1.

Applied the same way to every source, these weights give the map:

              node0  node1  node2  node3
global:         2      2      1      1

A task on CPU 0 gives node1 - remote DRAM, effective 100 - the same
weight 2 as its own local node0 at 300. Worse, node1 is weighted above
node2, the task's local CXL at effective 150, even though node2 is the
faster of the two. The flat weights rank a slower interconnect-bound
node above a faster local one, which is exactly backwards.

This series makes weighted interleave package-aware. When it is on,
weighted interleave prefers the task's current package: while the
package's nodes have room, the task's pages are spread across them by
weight, so allocations stay off the interconnect. The rest of the
policy nodemask is used when the local package cannot serve the request
- when a node in it is under pressure and the page allocator falls back
along the zonelist, or when the policy nodemask happens to exclude every
node of the current package, in which case the package spanned by the
policy's own nodes is used instead. The nodes considered are always
within the policy nodemask, which mempolicy already narrows to the
task's cpuset, so cpusets and the task nodemask stay in control.

              node0  node1  node2  node3
from CPU 0:     2      0      1      0
from CPU 1:     0      2      0      1

A task on CPU 0 now places pages on node0 (weight 2) and node2
(weight 1) at 2:1, which matches their effective bandwidth of 300:150;
a task on CPU 1 places on node1 and node3 the same way. Placement
follows the bandwidth each task actually sees, NUMA locality is
preserved, and interleave traffic stays off the interconnect.

To make this possible the kernel needs a notion of which nodes share a
package. The NUMA distance model offers only relative latencies and no
structural grouping, which is especially limiting for CXL memory nodes
that come online without an explicit package association.

The series adds a package-aware topology layer that groups CPU and
memory-only nodes into a "memory package", built from the physical
package ids firmware reports and, for a memory-only node, an initiator
CPU node or SLIT distances. A package can contain more than one CPU node
or more than one memory-only node, so the layer maps a package to a set
of nodes rather than to a single node or a single CXL device.

The feature is off by default and opt-in through a sysfs toggle. The
package topology itself is exposed read-only under
/sys/devices/system/package/; there is deliberately no writable
override, since a machine whose firmware describes its topology
incorrectly should be fixed in firmware. On a topology that does not
have the symmetric shape the placement relies on, enabling is refused
and any active mode degrades cleanly to the original flat behavior.

Measured results:

System Configuration:
- Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)

1) Throughput (System Bandwidth)
   - DRAM Only: 966 GB/s
   - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
   - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
     (38% increase compared to DRAM Only,
      47% increase compared to Weighted Interleave)

2) Loaded Latency (Under High Bandwidth)
   - DRAM Only: 544 ns
   - Weighted Interleave: 545 ns
   - Package-Aware Weighted Interleave: 436 ns
     (20% reduction compared to both)

A small CXL driver change registers a CXL memory node into its package
as the node comes online, using the initiator the driver resolves for
the region; this is where the package layer gets the CPU-side
association that plain NUMA distance does not carry.

The memory_package layer offers a broader interface for grouping and
querying package topology - usable by memory tiering as well - and
package-aware weighted interleave uses the subset it needs.

[PATCH 1/4] mm/numa: introduce nearest_nodes_nodemask()
  Add a NUMA helper that returns every node sharing the minimum distance
  from a source node.

[PATCH 2/4] mm/memory-tiers: package-aware topology management
  Group NUMA nodes into memory packages from firmware topology data,
  expose the grouping read-only under /sys/devices/system/package/, and
  validate the symmetric shape that package-aware placement relies on.

[PATCH 3/4] mm/memory-tiers: register CXL nodes to packages
  Bind a CXL memory node to a package using an initiator CPU node.

[PATCH 4/4] mm/mempolicy: package-aware weighted interleave
  Prefer the current package for weighted interleave node selection,
  behind an opt-in package_mode sysfs toggle that is off by default.

Rakie Kim (4):
  mm/numa: introduce nearest_nodes_nodemask()
  mm/memory-tiers: introduce package-aware topology management for NUMA
    nodes
  mm/memory-tiers: register CXL nodes to memory packages via initiator
  mm/mempolicy: enhance weighted interleave with package-aware locality

 .../ABI/testing/sysfs-devices-system-package  |   35 +
 ...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
 drivers/cxl/core/region.c                     |   54 +
 drivers/cxl/cxl.h                             |    1 +
 drivers/dax/kmem.c                            |    3 +
 include/linux/memory-tiers.h                  |  113 ++
 include/linux/numa.h                          |   11 +
 mm/memory-tiers.c                             | 1009 +++++++++++++++++
 mm/mempolicy.c                                |  200 +++-
 9 files changed, 1439 insertions(+), 4 deletions(-)
 create mode 100644 Documentation/ABI/testing/sysfs-devices-system-package


base-commit: 8cd9520d35a6c38db6567e97dd93b1f11f185dc6
-- 
2.25.1
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Gregory Price 1 month, 2 weeks ago
On Thu, Aug 06, 2026 at 05:09:31PM +0900, Rakie Kim wrote:
> 
> To make this possible the kernel needs a notion of which nodes share a
> package. The NUMA distance model offers only relative latencies and no
> structural grouping, which is especially limiting for CXL memory nodes
> that come online without an explicit package association.
> 

The original attempt to handle cross-socket interleave tried to deal
with this with a matrix for weights, but this was deemed too invasive.

This brings back that matrix, but not for per-node weights - we're
basically just adding an addition weighting to filter on.

I have some concerns with the now additional filtering mechanism
introdced into the allocation stack, but fundamentally I think this is 
a *better* solution than a straight weight-matrix.

I will need to chew on this for a bit.  It seems there's non-trivial
sashiko bug reports here to address anyway.

> The series adds a package-aware topology layer that groups CPU and
> memory-only nodes into a "memory package", built from the physical
> package ids firmware reports and, for a memory-only node, an initiator
> CPU node or SLIT distances. A package can contain more than one CPU node
> or more than one memory-only node, so the layer maps a package to a set
> of nodes rather than to a single node or a single CXL device.
> 

When worded this way, "Package" sounds completely arbitrary and not a
useful distinction.  This really just sounds like an extention for the
existing fallback lists to take interconnects into account.

I wonder if abstract distance either:
  1) already gives you what you want (the secondary weight)
  2) can be twiddled in BIOS to give you what you want.

> The feature is off by default and opt-in through a sysfs toggle. The
> package topology itself is exposed read-only under
> /sys/devices/system/package/; there is deliberately no writable
> override, since a machine whose firmware describes its topology
> incorrectly should be fixed in firmware. On a topology that does not
> have the symmetric shape the placement relies on, enabling is refused
> and any active mode degrades cleanly to the original flat behavior.

The way this is written it sounds to me like this should just be the
default weighted interleave behavior.  We already know weighted
interleave does not jive well with multi-socket systems - this just
fixes that (in a more general sense, Socket => Package).

~Gregory
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Mon, 17 Aug 2026 12:19:51 -0400 Gregory Price <gourry@gourry.net> wrote:

> The original attempt to handle cross-socket interleave tried to deal
> with this with a matrix for weights, but this was deemed too invasive.
>
> This brings back that matrix, but not for per-node weights - we're
> basically just adding an addition weighting to filter on.

Yes. It uses additional information to filter the nodes weighted
interleave selects from.

> I have some concerns with the now additional filtering mechanism
> introdced into the allocation stack, but fundamentally I think this is
> a *better* solution than a straight weight-matrix.

About the cost of the filter: when the toggle is off, the filter does
not run. When it is on, node selection needs a nodemask filtering
step, but in my tests the overhead was negligible. I will look at
this part further and check whether there is more room to optimize.

> I will need to chew on this for a bit.  It seems there's non-trivial
> sashiko bug reports here to address anyway.

I am working on the sashiko findings now. The valid ones will be
fixed in the next version.

> When worded this way, "Package" sounds completely arbitrary and not a
> useful distinction.  This really just sounds like an extention for the
> existing fallback lists to take interconnects into account.

Other reviewers also pointed out that the "package" / "socket"
terminology is ambiguous. I think the description needs a full rework
so that it explains the current state better, and I plan to do that
in the next version.

> I wonder if abstract distance either:
>   1) already gives you what you want (the secondary weight)
>   2) can be twiddled in BIOS to give you what you want.

I thought about abstract distance a lot as well. My conclusion was
that adistance alone cannot tell whether nodes are in the same
package. This is an area I am still thinking about, and it needs
more thought.

About the BIOS information: as I reported before, on the two-package
server I tested, each package had its own CXL device, but the HMAT
reported one CPU node as the initiator of both CXL nodes:

https://lore.kernel.org/all/20260330025914.361-1-rakie.kim@sk.com/

I will post an update on that issue as well. What this series uses
is, in the end, BIOS information too. But some of that information
has errors and some of it looks reliable, so I think we also need to
sort out which is which, and understand why.

> The way this is written it sounds to me like this should just be the
> default weighted interleave behavior.  We already know weighted
> interleave does not jive well with multi-socket systems - this just
> fixes that (in a more general sense, Socket => Package).

I agree with your point. The default is off because it was requested
during the v1 review; Jonathan asked for it:

https://lore.kernel.org/all/20260325123350.00004d48@huawei.com/

Enabling is also not unconditional. I added a few constraints, so it
turns on only in a specific situation: when the packages have the
same node structure. I think the definition of these on/off
conditions is open, and it needs more discussion.

Thanks again for your time and review.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Gregory Price 1 month, 2 weeks ago
On Tue, Aug 18, 2026 at 03:01:58PM +0900, Rakie Kim wrote:
> On Mon, 17 Aug 2026 12:19:51 -0400 Gregory Price <gourry@gourry.net> wrote:
> 
> > I have some concerns with the now additional filtering mechanism
> > introdced into the allocation stack, but fundamentally I think this is
> > a *better* solution than a straight weight-matrix.
> 
> About the cost of the filter: when the toggle is off, the filter does
> not run. When it is on, node selection needs a nodemask filtering
> step, but in my tests the overhead was negligible. I will look at
> this part further and check whether there is more room to optimize.
> 

Not concerned about the performance, concerned about how complicated the
mempolicy - cgroup - zonelist - page_alloc interaction already is, and
then adding another filtering mechanism on top.

Today we have:

1) cpuset constrains mempolicy (nodemask remaps)
2) cpuset constrains zonelist walks
3) mempolicy nodemask constrains zonelist walks
4) memory-tiers.c nodemask constrains zonelist walks for demotion
5) zonelist membership constrains allocation access
6) a bunch of corner conditions that violate 1-3 for the sake of forward
   progress

now we're adding:

7) memory-tiers.c nodemask constrains mempolicy nodemask
   except when it doesn't, because fallbacks occurred hard enough

It's already un-intuitive how and when memory lands on certain nodes.

To be clear, I'm not saying this idea is bad - either as-is or in some
other form - just that adding another nodemask filtering path is making
it harder and harder to understand what lands where.

Mostly starting to wonder if we're reaching the point where the page
allocator needs to take something a little more descriptive than a
nodemask to dictate placement.

~Gregory
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 1 week ago
On Tue, 18 Aug 2026 09:30:36 -0400 Gregory Price <gourry@gourry.net> wrote:

> Not concerned about the performance, concerned about how complicated the
> mempolicy - cgroup - zonelist - page_alloc interaction already is, and
> then adding another filtering mechanism on top.
>
> Today we have:
>
> 1) cpuset constrains mempolicy (nodemask remaps)
> 2) cpuset constrains zonelist walks
> 3) mempolicy nodemask constrains zonelist walks
> 4) memory-tiers.c nodemask constrains zonelist walks for demotion
> 5) zonelist membership constrains allocation access
> 6) a bunch of corner conditions that violate 1-3 for the sake of forward
>    progress
>
> now we're adding:
>
> 7) memory-tiers.c nodemask constrains mempolicy nodemask
>    except when it doesn't, because fallbacks occurred hard enough
>
> It's already un-intuitive how and when memory lands on certain nodes.
>
> To be clear, I'm not saying this idea is bad - either as-is or in some
> other form - just that adding another nodemask filtering path is making
> it harder and harder to understand what lands where.

I understand what you are concerned about, and I share the concern.

I tried to keep the addition minimal and to make it act only in the
intended situation: the filter runs in weighted interleave node
selection, only while the toggle is on. It does not change cpusets,
the zonelist, or the allocator fallback. Still, it is true that this
adds one more filtering layer to the constraints you listed.

I will try to organize this part so that it is as easy to follow as
possible, and reinforce the documentation as well. I am not sure that
will be enough, though - it needs more thought.

> Mostly starting to wonder if we're reaching the point where the page
> allocator needs to take something a little more descriptive than a
> nodemask to dictate placement.

This is something I think about a lot as well. Many users simply let
the policy use all nodes, and in the situation this series targets,
that is exactly what costs performance. To fix it, the kernel needs
more information than a nodemask carries. But as you pointed out,
adding that information also adds complexity. This is my concern as
well, and I think how to handle it needs a discussion.

Thanks again for your time and review.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Gregory Price 1 month, 2 weeks ago
On Thu, Aug 06, 2026 at 05:09:31PM +0900, Rakie Kim wrote:
> Package-aware weighted interleave places a task's weighted-interleave
> pages on the NUMA nodes of its local package, so that interleave traffic
> does not have to cross the interconnect to another package. This keeps
> each node's weight aligned with the bandwidth the task actually gets
> from it, so effective bandwidth holds up on a system that has more than
> one package. (A package is a CPU socket together with the memory
> attached to it.)
>

The only major question I have before I dig into the actual patches is
whether you actually need this on a *per-task* basis, rather than simply
a *per-process* basis - because that's all this really buys you.

You can already do what is described here by simply using a combination
of cpuset and mempolicy

cpuset.mems = 0,2
mempolicy = weighted interleave --all

If this is actually required on a per-task basis, then I agree this
concept is reasonable.  I just want to make we're grounded on a real
use case before we go adding this complexity.

~Gregory
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Wed, 12 Aug 2026 22:37:06 -0400 Gregory Price <gourry@gourry.net> wrote:
> On Thu, Aug 06, 2026 at 05:09:31PM +0900, Rakie Kim wrote:
> > Package-aware weighted interleave places a task's weighted-interleave
> > pages on the NUMA nodes of its local package, so that interleave traffic
> > does not have to cross the interconnect to another package. This keeps
> > each node's weight aligned with the bandwidth the task actually gets
> > from it, so effective bandwidth holds up on a system that has more than
> > one package. (A package is a CPU socket together with the memory
> > attached to it.)
> >
>
> The only major question I have before I dig into the actual patches is
> whether you actually need this on a *per-task* basis, rather than simply
> a *per-process* basis - because that's all this really buys you.
>
> You can already do what is described here by simply using a combination
> of cpuset and mempolicy
>
> cpuset.mems = 0,2
> mempolicy = weighted interleave --all
>
> If this is actually required on a per-task basis, then I agree this
> concept is reasonable.  I just want to make we're grounded on a real
> use case before we go adding this complexity.
>
> ~Gregory

Hello Gregory,

Thank you for reviewing this series. My description was not precise
enough, and it seems to have confused several people. Let me go over
the example again.

I labelled the rows "from CPU 0" and "from CPU 1", which reads as two
separate tasks, each pinned to one place. What the rows were meant to
show is which package the allocation is requested from:

                  node0  node1  node2  node3
from package 0:     2      0      1      0
from package 1:     0      2      0      1

Both rows come from one policy at the same time. Which row applies is
decided per allocation, by the package the requesting CPU is in.

The cpuset.mems = 0,2 you describe is the node set of package 0, so
the tasks in that cgroup use node0 and node2 wherever they run.

The feature I am proposing does not fix a package that way. It uses
the nodes of the package the request comes from.

To build the feature I am proposing with cpuset, I think you would
need a cgroup per package and would have to place each thread in the
one matching the package its CPU is in.

cpuset.mems does not follow the CPU, so unless cpuset.cpus is set
alongside it, a thread running on a CPU of package 1 would use node0
and node2. For a program with tens or hundreds of threads, keeping
that placement right does not look easy.

To put it plainly, what this series aims for is that when a process
with many threads runs across several packages, each thread's
allocations are weighted-interleaved within the nodes of the package
that thread is on.

There is one process and one policy, but the nodes actually used
differ with where each thread sits. That is why I think this has to be
per task rather than per process.

The workloads I had in mind are memory-intensive server workloads:
vector databases as used in AI applications, or in-memory key-value
caching services that sit in front of large services and keep their
data in memory.

They run with a great many threads, spread across several packages and
running at the same time, and they want as much memory bandwidth as
they can get.

That is also why the numbers were measured with MLC. It produces high
bandwidth from many threads, much like those workloads, while still
giving a quantitative result.

The runs used "numactl -w all mlc" with no CPU binding, so the threads
were spread over both packages, and in that configuration the
bandwidth was higher than with the existing weighted interleave.

That is the reasoning behind my thinking that this belongs at the task
level. In the next version I will strengthen the documentation,
including why this feature is needed.

Thank you for taking the time to review this and for the question. It
made it much clearer to me what the cover letter still needs to
explain.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Gregory Price 1 month, 2 weeks ago
On Thu, Aug 13, 2026 at 03:23:48PM +0900, Rakie Kim wrote:
> On Wed, 12 Aug 2026 22:37:06 -0400 Gregory Price <gourry@gourry.net> wrote:
> 
> The cpuset.mems = 0,2 you describe is the node set of package 0, so
> the tasks in that cgroup use node0 and node2 wherever they run.
> 

I think you are over-complicating the explanation and that's making it
hard for folks to reason about this.

My best understanding here is for multi-socket systems, weighted
interleave as-designed is inherently sub-optimal for scaled workloads
that utilize multi-socket ("package") memory resources (DRAM, CXL...)

I think there is also an assumption that total memory utilization is
less than the total capacity - otherwise some of the assumptions here
break, but that comes with the general weighted interleave story.

I'm getting back from vacation, I will take a closer look this week.

~Gregory
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Sun, 16 Aug 2026 20:46:16 -0400 Gregory Price <gourry@gourry.net> wrote:

Hello Gregory,

Thanks for your review and feedback.

> I think you are over-complicating the explanation and that's making it
> hard for folks to reason about this.

You are right. I will keep my explanations focused on the key points
from here on, and the next cover letter will be written that way as
well.

> My best understanding here is for multi-socket systems, weighted
> interleave as-designed is inherently sub-optimal for scaled workloads
> that utilize multi-socket ("package") memory resources (DRAM, CXL...)

Your summary is correct. That is the problem this series addresses.

> I think there is also an assumption that total memory utilization is
> less than the total capacity - otherwise some of the assumptions here
> break, but that comes with the general weighted interleave story.

Your assumption is correct. It is the same as the existing weighted
interleave: when the nodes are full, memory is allocated from other
nodes.

> I'm getting back from vacation, I will take a closer look this week.

Thank you. I look forward to your comments.

Thanks again for your time and review.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Joshua Hahn 1 month, 2 weeks ago
On Sun, 16 Aug 2026 20:46:16 -0400 Gregory Price <gourry@gourry.net> wrote:

> On Thu, Aug 13, 2026 at 03:23:48PM +0900, Rakie Kim wrote:
> > On Wed, 12 Aug 2026 22:37:06 -0400 Gregory Price <gourry@gourry.net> wrote:
> > 
> > The cpuset.mems = 0,2 you describe is the node set of package 0, so
> > the tasks in that cgroup use node0 and node2 wherever they run.
> > 
> 
> I think you are over-complicating the explanation and that's making it
> hard for folks to reason about this.
> 
> My best understanding here is for multi-socket systems, weighted
> interleave as-designed is inherently sub-optimal for scaled workloads
> that utilize multi-socket ("package") memory resources (DRAM, CXL...)

I was thinking about this and I think Gregory is right here.

I wonder how much of the framing around this series can be preserved /
simplified if we just say "prevent tasks from allocating memory
cross-socket".

I also wonder if instead of limiting this to weighted interleave,
this can sit on top of other mpols and just nodemasks against
cross-socket nodes.

> I think there is also an assumption that total memory utilization is
> less than the total capacity - otherwise some of the assumptions here
> break, but that comes with the general weighted interleave story.
> 
> I'm getting back from vacation, I will take a closer look this week.

I am also out on vacation this week : -)
I'll take a closer look along with the code next week as well.

Thanks, Rakie and Gregory!
Joshua
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Sun, 16 Aug 2026 21:52:24 -0700 Joshua Hahn <joshua.hahnjy@gmail.com> wrote:

Hello Joshua,

Thanks for your feedback.

> On Sun, 16 Aug 2026 20:46:16 -0400 Gregory Price <gourry@gourry.net> wrote:
>
> > I think you are over-complicating the explanation and that's making it
> > hard for folks to reason about this.
> >
> > My best understanding here is for multi-socket systems, weighted
> > interleave as-designed is inherently sub-optimal for scaled workloads
> > that utilize multi-socket ("package") memory resources (DRAM, CXL...)
>
> I was thinking about this and I think Gregory is right here.

I agree. From now on I will keep my explanations focused on the key
points.

> I wonder how much of the framing around this series can be preserved /
> simplified if we just say "prevent tasks from allocating memory
> cross-socket".

I think this is a good suggestion. Even if it is not this exact
wording, I will look for a shorter and better description that
includes your idea.

> I also wonder if instead of limiting this to weighted interleave,
> this can sit on top of other mpols and just nodemasks against
> cross-socket nodes.

It does not have to apply only to weighted interleave. Patch 2 simply
groups NUMA nodes by the package they belong to. So the same grouping
can also be used for interleave, not just weighted interleave. I have
not applied it to interleave because I could not find a clear use
case for it yet.

> I am also out on vacation this week : -)
> I'll take a closer look along with the code next week as well.

Thank you. I look forward to your comments.

Thanks again for your time and review.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
On Thu, Aug 06, 2026 at 05:09:31PM +0900, Rakie Kim wrote:
> Package-aware weighted interleave places a task's weighted-interleave
> pages on the NUMA nodes of its local package, so that interleave traffic
> does not have to cross the interconnect to another package. This keeps
> each node's weight aligned with the bandwidth the task actually gets
> from it, so effective bandwidth holds up on a system that has more than
> one package. (A package is a CPU socket together with the memory
> attached to it.)
>
> Changes from RFC:
> https://lore.kernel.org/all/20260316051258.246-1-rakie.kim@sk.com/
> - Added an opt-in sysfs toggle (off by default) and a read-only sysfs
>   view of the package topology
> - Added topology validation with a clean fallback to plain weighted
>   interleave on unsupported topologies
> - Hardened the allocation, device-teardown, and node-hotplug paths

Please put change logs under the cover letter :) in mm we put the cover
letter in the actual upstream commit so it's better to keep separate for
reviewers.

I see the RFC was from march and tied to an LSF session I think?

While I don't want to be too pedantic, I think you should only really
un-RFC in a situation where you have a good sense that the relevant
maintainers are happy with the _concept_.

Looking at the RFC thread it's not clear that David was OK with this on the
mm side, though I see you got some feedback from Jonathan on the CXL driver
side.

So I wonder whether next respin this should be re-RFC'd unless you get
clear feedback that we want to go in this direction?

>
> Weighted interleave places pages on nodes in proportion to per-node
> weights that are set from each node's bandwidth. Within one package the
> weight given to a node matches the bandwidth a task sees from it. Across
> packages it no longer does: a memory node's physical bandwidth is fixed,
> but the bandwidth a task effectively sees depends on which package its
> CPU is in, because memory reached from another package, over the
> interconnect between them, is slower than the same memory reached within
> the package. The weights are set once from device bandwidth and applied
> the same way wherever the task runs, so a node in another package is
> given a weight higher than the bandwidth it can deliver to that task.
> The kernel has no package abstraction and does not record which package
> a node belongs to, so it cannot tell which node pairs are separated by
> the interconnect.

Please please - break up huge paragraphs like this :)

It's 2026 so I have to mention that if you've used AI to assist with
writing it (which is fine) please do a pass over it manually to curb AI's
tendency to be overly verbose + definitely try to break up paragraphs into
smaller charts at least :)

>
>           node0             node1
>         +-------+         +-------+
>         | CPU 0 |---------| CPU 1 |
>         +-------+         +-------+
>         | DRAM0 |         | DRAM1 |
>         +---+---+         +---+---+
>             |                 |
>         +---+---+         +---+---+
>         | CXL 0 |         | CXL 1 |
>         +-------+         +-------+
>           node2             node3

...though I _love_ ASCII diagrams so this is great ;)

>
> The numbers below are illustrative single-stream bandwidths (GB/s).
> Local DRAM sustains 300 and local CXL 150; any path that crosses to
> another package, over the interconnect, is capped at 100, so a node in
> another package delivers 100 whether it is DRAM or CXL. Note that local
> CXL (150) is still faster than any node in another package (100). The
> effective bandwidth each CPU sees is therefore:
>
>               node0  node1  node2  node3
> from CPU 0:    300    100    150    100
> from CPU 1:    100    300    100    150
>
> Since a single per-node weight cannot encode the interconnect penalty,
> a reasonable set of global weights is taken from local device bandwidth
> (local DRAM : local CXL = 300 : 150 = 2 : 1): node0=2 node1=2 node2=1
> node3=1.

Also great that you provide the receipts on actual observed real-world
numbers that's great.

NUMA isn't my area so I can't comment here on the technical details but
thanks for providing this :)


>
> Applied the same way to every source, these weights give the map:
>
>               node0  node1  node2  node3
> global:         2      2      1      1
>
> A task on CPU 0 gives node1 - remote DRAM, effective 100 - the same
> weight 2 as its own local node0 at 300. Worse, node1 is weighted above
> node2, the task's local CXL at effective 150, even though node2 is the
> faster of the two. The flat weights rank a slower interconnect-bound
> node above a faster local one, which is exactly backwards.
>
> This series makes weighted interleave package-aware. When it is on,
> weighted interleave prefers the task's current package: while the
> package's nodes have room, the task's pages are spread across them by
> weight, so allocations stay off the interconnect. The rest of the
> policy nodemask is used when the local package cannot serve the request
> - when a node in it is under pressure and the page allocator falls back
> along the zonelist, or when the policy nodemask happens to exclude every
> node of the current package, in which case the package spanned by the
> policy's own nodes is used instead. The nodes considered are always
> within the policy nodemask, which mempolicy already narrows to the
> task's cpuset, so cpusets and the task nodemask stay in control.
>
>               node0  node1  node2  node3
> from CPU 0:     2      0      1      0
> from CPU 1:     0      2      0      1
>
> A task on CPU 0 now places pages on node0 (weight 2) and node2
> (weight 1) at 2:1, which matches their effective bandwidth of 300:150;
> a task on CPU 1 places on node1 and node3 the same way. Placement
> follows the bandwidth each task actually sees, NUMA locality is
> preserved, and interleave traffic stays off the interconnect.
>
> To make this possible the kernel needs a notion of which nodes share a
> package. The NUMA distance model offers only relative latencies and no
> structural grouping, which is especially limiting for CXL memory nodes
> that come online without an explicit package association.
>
> The series adds a package-aware topology layer that groups CPU and
> memory-only nodes into a "memory package", built from the physical
> package ids firmware reports and, for a memory-only node, an initiator
> CPU node or SLIT distances. A package can contain more than one CPU node
> or more than one memory-only node, so the layer maps a package to a set
> of nodes rather than to a single node or a single CXL device.
>
> The feature is off by default and opt-in through a sysfs toggle. The
> package topology itself is exposed read-only under
> /sys/devices/system/package/; there is deliberately no writable
> override, since a machine whose firmware describes its topology
> incorrectly should be fixed in firmware. On a topology that does not
> have the symmetric shape the placement relies on, enabling is refused
> and any active mode degrades cleanly to the original flat behavior.
>
> Measured results:
>
> System Configuration:
> - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
>
> 1) Throughput (System Bandwidth)
>    - DRAM Only: 966 GB/s
>    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
>    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
>      (38% increase compared to DRAM Only,
>       47% increase compared to Weighted Interleave)
>
> 2) Loaded Latency (Under High Bandwidth)
>    - DRAM Only: 544 ns
>    - Weighted Interleave: 545 ns
>    - Package-Aware Weighted Interleave: 436 ns
>      (20% reduction compared to both)
>
> A small CXL driver change registers a CXL memory node into its package
> as the node comes online, using the initiator the driver resolves for
> the region; this is where the package layer gets the CPU-side
> association that plain NUMA distance does not carry.
>
> The memory_package layer offers a broader interface for grouping and
> querying package topology - usable by memory tiering as well - and
> package-aware weighted interleave uses the subset it needs.
>
> [PATCH 1/4] mm/numa: introduce nearest_nodes_nodemask()
>   Add a NUMA helper that returns every node sharing the minimum distance
>   from a source node.
>
> [PATCH 2/4] mm/memory-tiers: package-aware topology management
>   Group NUMA nodes into memory packages from firmware topology data,
>   expose the grouping read-only under /sys/devices/system/package/, and
>   validate the symmetric shape that package-aware placement relies on.
>
> [PATCH 3/4] mm/memory-tiers: register CXL nodes to packages
>   Bind a CXL memory node to a package using an initiator CPU node.
>
> [PATCH 4/4] mm/mempolicy: package-aware weighted interleave
>   Prefer the current package for weighted interleave node selection,
>   behind an opt-in package_mode sysfs toggle that is off by default.
>
> Rakie Kim (4):
>   mm/numa: introduce nearest_nodes_nodemask()
>   mm/memory-tiers: introduce package-aware topology management for NUMA
>     nodes
>   mm/memory-tiers: register CXL nodes to memory packages via initiator
>   mm/mempolicy: enhance weighted interleave with package-aware locality
>
>  .../ABI/testing/sysfs-devices-system-package  |   35 +
>  ...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
>  drivers/cxl/core/region.c                     |   54 +
>  drivers/cxl/cxl.h                             |    1 +
>  drivers/dax/kmem.c                            |    3 +
>  include/linux/memory-tiers.h                  |  113 ++
>  include/linux/numa.h                          |   11 +
>  mm/memory-tiers.c                             | 1009 +++++++++++++++++
>  mm/mempolicy.c                                |  200 +++-
>  9 files changed, 1439 insertions(+), 4 deletions(-)
>  create mode 100644 Documentation/ABI/testing/sysfs-devices-system-package
>
>
> base-commit: 8cd9520d35a6c38db6567e97dd93b1f11f185dc6
> --
> 2.25.1
>

--
Cheers, Lorenzo
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Wed, 12 Aug 2026 08:16:23 +0100 "Lorenzo Stoakes (ARM)" <ljs@kernel.org> wrote:

Hello Lorenzo,

Thank you for taking the time to review this series and for the
advice on how to submit it.

> On Thu, Aug 06, 2026 at 05:09:31PM +0900, Rakie Kim wrote:
> > Package-aware weighted interleave places a task's weighted-interleave
> > pages on the NUMA nodes of its local package, so that interleave traffic
> > does not have to cross the interconnect to another package. This keeps
> > each node's weight aligned with the bandwidth the task actually gets
> > from it, so effective bandwidth holds up on a system that has more than
> > one package. (A package is a CPU socket together with the memory
> > attached to it.)
> >
> > Changes from RFC:
> > https://lore.kernel.org/all/20260316051258.246-1-rakie.kim@sk.com/
> > - Added an opt-in sysfs toggle (off by default) and a read-only sysfs
> >   view of the package topology
> > - Added topology validation with a clean fallback to plain weighted
> >   interleave on unsupported topologies
> > - Hardened the allocation, device-teardown, and node-hotplug paths
> 
> Please put change logs under the cover letter :) in mm we put the cover
> letter in the actual upstream commit so it's better to keep separate for
> reviewers.
>

Thank you, I did not know that. I will keep the change log out of the
cover letter body from the next version.


> I see the RFC was from march and tied to an LSF session I think?
> 
> While I don't want to be too pedantic, I think you should only really
> un-RFC in a situation where you have a good sense that the relevant
> maintainers are happy with the _concept_.
> 
> Looking at the RFC thread it's not clear that David was OK with this on the
> mm side, though I see you got some feedback from Jonathan on the CXL driver
> side.
> 
> So I wonder whether next respin this should be re-RFC'd unless you get
> clear feedback that we want to go in this direction?
>

You are right about the LSF session: the RFC was posted with that
proposal in mind, but I could not attend for personal reasons, so the
discussion I had hoped for there did not happen.

Your point about the concept is fair as well, so I will post the next
version as an RFC again.


> >
> > Weighted interleave places pages on nodes in proportion to per-node
> > weights that are set from each node's bandwidth. Within one package the
> > weight given to a node matches the bandwidth a task sees from it. Across
> > packages it no longer does: a memory node's physical bandwidth is fixed,
> > but the bandwidth a task effectively sees depends on which package its
> > CPU is in, because memory reached from another package, over the
> > interconnect between them, is slower than the same memory reached within
> > the package. The weights are set once from device bandwidth and applied
> > the same way wherever the task runs, so a node in another package is
> > given a weight higher than the bandwidth it can deliver to that task.
> > The kernel has no package abstraction and does not record which package
> > a node belongs to, so it cannot tell which node pairs are separated by
> > the interconnect.
> 
> Please please - break up huge paragraphs like this :)
> 
> It's 2026 so I have to mention that if you've used AI to assist with
> writing it (which is fine) please do a pass over it manually to curb AI's
> tendency to be overly verbose + definitely try to break up paragraphs into
> smaller charts at least :)
>

I agree. Reading it again, my sentences are too long and the
explanation is more verbose than it needs to be, which makes the
cover letter hard to read. This needs to be fixed.

I will rewrite the cover letter for the next version in a form that
is easier to read.


> >
> >           node0             node1
> >         +-------+         +-------+
> >         | CPU 0 |---------| CPU 1 |
> >         +-------+         +-------+
> >         | DRAM0 |         | DRAM1 |
> >         +---+---+         +---+---+
> >             |                 |
> >         +---+---+         +---+---+
> >         | CXL 0 |         | CXL 1 |
> >         +-------+         +-------+
> >           node2             node3
> 
> ...though I _love_ ASCII diagrams so this is great ;)
>

Thank you.


> >
> > The numbers below are illustrative single-stream bandwidths (GB/s).
> > Local DRAM sustains 300 and local CXL 150; any path that crosses to
> > another package, over the interconnect, is capped at 100, so a node in
> > another package delivers 100 whether it is DRAM or CXL. Note that local
> > CXL (150) is still faster than any node in another package (100). The
> > effective bandwidth each CPU sees is therefore:
> >
> >               node0  node1  node2  node3
> > from CPU 0:    300    100    150    100
> > from CPU 1:    100    300    100    150
> >
> > Since a single per-node weight cannot encode the interconnect penalty,
> > a reasonable set of global weights is taken from local device bandwidth
> > (local DRAM : local CXL = 300 : 150 = 2 : 1): node0=2 node1=2 node2=1
> > node3=1.
> 
> Also great that you provide the receipts on actual observed real-world
> numbers that's great.
> 
> NUMA isn't my area so I can't comment here on the technical details but
> thanks for providing this :)
>

Thank you. I will keep reporting measured results with the next
versions.


> 
> >
> > Applied the same way to every source, these weights give the map:
> >
> >               node0  node1  node2  node3
> > global:         2      2      1      1
> >
> > A task on CPU 0 gives node1 - remote DRAM, effective 100 - the same
> > weight 2 as its own local node0 at 300. Worse, node1 is weighted above
> > node2, the task's local CXL at effective 150, even though node2 is the
> > faster of the two. The flat weights rank a slower interconnect-bound
> > node above a faster local one, which is exactly backwards.
> >
> > This series makes weighted interleave package-aware. When it is on,
> > weighted interleave prefers the task's current package: while the
> > package's nodes have room, the task's pages are spread across them by
> > weight, so allocations stay off the interconnect. The rest of the
> > policy nodemask is used when the local package cannot serve the request
> > - when a node in it is under pressure and the page allocator falls back
> > along the zonelist, or when the policy nodemask happens to exclude every
> > node of the current package, in which case the package spanned by the
> > policy's own nodes is used instead. The nodes considered are always
> > within the policy nodemask, which mempolicy already narrows to the
> > task's cpuset, so cpusets and the task nodemask stay in control.
> >
> >               node0  node1  node2  node3
> > from CPU 0:     2      0      1      0
> > from CPU 1:     0      2      0      1
> >
> > A task on CPU 0 now places pages on node0 (weight 2) and node2
> > (weight 1) at 2:1, which matches their effective bandwidth of 300:150;
> > a task on CPU 1 places on node1 and node3 the same way. Placement
> > follows the bandwidth each task actually sees, NUMA locality is
> > preserved, and interleave traffic stays off the interconnect.
> >
> > To make this possible the kernel needs a notion of which nodes share a
> > package. The NUMA distance model offers only relative latencies and no
> > structural grouping, which is especially limiting for CXL memory nodes
> > that come online without an explicit package association.
> >
> > The series adds a package-aware topology layer that groups CPU and
> > memory-only nodes into a "memory package", built from the physical
> > package ids firmware reports and, for a memory-only node, an initiator
> > CPU node or SLIT distances. A package can contain more than one CPU node
> > or more than one memory-only node, so the layer maps a package to a set
> > of nodes rather than to a single node or a single CXL device.
> >
> > The feature is off by default and opt-in through a sysfs toggle. The
> > package topology itself is exposed read-only under
> > /sys/devices/system/package/; there is deliberately no writable
> > override, since a machine whose firmware describes its topology
> > incorrectly should be fixed in firmware. On a topology that does not
> > have the symmetric shape the placement relies on, enabling is refused
> > and any active mode degrades cleanly to the original flat behavior.
> >
> > Measured results:
> >
> > System Configuration:
> > - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> >
> > 1) Throughput (System Bandwidth)
> >    - DRAM Only: 966 GB/s
> >    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
> >    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
> >      (38% increase compared to DRAM Only,
> >       47% increase compared to Weighted Interleave)
> >
> > 2) Loaded Latency (Under High Bandwidth)
> >    - DRAM Only: 544 ns
> >    - Weighted Interleave: 545 ns
> >    - Package-Aware Weighted Interleave: 436 ns
> >      (20% reduction compared to both)
> >
> > A small CXL driver change registers a CXL memory node into its package
> > as the node comes online, using the initiator the driver resolves for
> > the region; this is where the package layer gets the CPU-side
> > association that plain NUMA distance does not carry.
> >
> > The memory_package layer offers a broader interface for grouping and
> > querying package topology - usable by memory tiering as well - and
> > package-aware weighted interleave uses the subset it needs.
> >
> > [PATCH 1/4] mm/numa: introduce nearest_nodes_nodemask()
> >   Add a NUMA helper that returns every node sharing the minimum distance
> >   from a source node.
> >
> > [PATCH 2/4] mm/memory-tiers: package-aware topology management
> >   Group NUMA nodes into memory packages from firmware topology data,
> >   expose the grouping read-only under /sys/devices/system/package/, and
> >   validate the symmetric shape that package-aware placement relies on.
> >
> > [PATCH 3/4] mm/memory-tiers: register CXL nodes to packages
> >   Bind a CXL memory node to a package using an initiator CPU node.
> >
> > [PATCH 4/4] mm/mempolicy: package-aware weighted interleave
> >   Prefer the current package for weighted interleave node selection,
> >   behind an opt-in package_mode sysfs toggle that is off by default.
> >
> > Rakie Kim (4):
> >   mm/numa: introduce nearest_nodes_nodemask()
> >   mm/memory-tiers: introduce package-aware topology management for NUMA
> >     nodes
> >   mm/memory-tiers: register CXL nodes to memory packages via initiator
> >   mm/mempolicy: enhance weighted interleave with package-aware locality
> >
> >  .../ABI/testing/sysfs-devices-system-package  |   35 +
> >  ...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
> >  drivers/cxl/core/region.c                     |   54 +
> >  drivers/cxl/cxl.h                             |    1 +
> >  drivers/dax/kmem.c                            |    3 +
> >  include/linux/memory-tiers.h                  |  113 ++
> >  include/linux/numa.h                          |   11 +
> >  mm/memory-tiers.c                             | 1009 +++++++++++++++++
> >  mm/mempolicy.c                                |  200 +++-
> >  9 files changed, 1439 insertions(+), 4 deletions(-)
> >  create mode 100644 Documentation/ABI/testing/sysfs-devices-system-package
> >
> >
> > base-commit: 8cd9520d35a6c38db6567e97dd93b1f11f185dc6
> > --
> > 2.25.1
> >
>
> --
> Cheers, Lorenzo

Thank you again for the excellent advice. What you pointed out matters
for how a patch is delivered, so I will keep it in mind for the next
version and the ones after it.

Thanks again for your time and review.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Joshua Hahn 1 month, 3 weeks ago
On Thu,  6 Aug 2026 17:09:31 +0900 Rakie Kim <rakie.kim@sk.com> wrote:

> Package-aware weighted interleave places a task's weighted-interleave
> pages on the NUMA nodes of its local package, so that interleave traffic
> does not have to cross the interconnect to another package. This keeps
> each node's weight aligned with the bandwidth the task actually gets
> from it, so effective bandwidth holds up on a system that has more than
> one package. (A package is a CPU socket together with the memory
> attached to it.)
> 
> Changes from RFC:
> https://lore.kernel.org/all/20260316051258.246-1-rakie.kim@sk.com/
> - Added an opt-in sysfs toggle (off by default) and a read-only sysfs
>   view of the package topology
> - Added topology validation with a clean fallback to plain weighted
>   interleave on unsupported topologies
> - Hardened the allocation, device-teardown, and node-hotplug paths

Hello Rakie,

I hope you are doing well! Sorry for the late repsonse.

I have a few thoughts, some of which are carry-overs from the RFC
discussion we had before! I think there are still some open questions,
and I wanted to get your opinion on some of them.

My first question is whether we want cross-socket allocations at all.
The examples you gave seem to line up with node-restricted interleave,
as opposed to cross-socket interleave. I think the wording that you
use to describe the feature in 4/4 (which I will copy below)

> The resolved mask is by construction a subset of the policy nodemask, which
> mempolicy already restricts to the task's cpuset; package mode can only
> narrow that set, never widen it, so cpusets and the task nodemask remain
> authoritative.

is 100% the right way to treat these package-aware (socket-aware)
interleaving allocations, but the example below

[...snip...]

> Applied the same way to every source, these weights give the map:
> 
>               node0  node1  node2  node3
> global:         2      2      1      1

[...snip...]

>               node0  node1  node2  node3
> from CPU 0:     2      0      1      0
> from CPU 1:     0      2      0      1

Is essentially the existing weighted interleave mechanism with a 
nodemask/cpuset applied. With that said, I think a more interesting and
illustrative example would be if the user truly would want to allow some
allocations to go through cross-socket, but be able to control the
ratio at which these slip through.

              node0  node1  node2  node3
from CPU 0:     3      1      2      0
from CPU 1:     0      3      1      2

Maybe even more illustrative of the true capabilities of this series
would be if you have an asymmetric system where you bind some
host-level monitoring / logging workloads to one node (say, node0) and
want that to be able to cross through to the other socket, but not the
other way around:

              node0  node1  node2  node3
from CPU 0:     3      1      2      0
from CPU 1:     0      2      0      1

Anyways, these are just super hypothetical scenarios and I don't even
know if the configuration that I'm listing would really be beneficial
for the system. I think that coming up with some illustrative usecases
which are now made possible by this series could help motivate why we
would want to interleave across sockets. 

> A task on CPU 0 now places pages on node0 (weight 2) and node2
> (weight 1) at 2:1, which matches their effective bandwidth of 300:150;
> a task on CPU 1 places on node1 and node3 the same way. Placement
> follows the bandwidth each task actually sees, NUMA locality is
> preserved, and interleave traffic stays off the interconnect.
> 
> To make this possible the kernel needs a notion of which nodes share a
> package. The NUMA distance model offers only relative latencies and no
> structural grouping, which is especially limiting for CXL memory nodes
> that come online without an explicit package association.
> 
> The series adds a package-aware topology layer that groups CPU and
> memory-only nodes into a "memory package", built from the physical
> package ids firmware reports and, for a memory-only node, an initiator
> CPU node or SLIT distances. A package can contain more than one CPU node
> or more than one memory-only node, so the layer maps a package to a set
> of nodes rather than to a single node or a single CXL device.
> 
> The feature is off by default and opt-in through a sysfs toggle. The
> package topology itself is exposed read-only under
> /sys/devices/system/package/; there is deliberately no writable
> override, since a machine whose firmware describes its topology
> incorrectly should be fixed in firmware. On a topology that does not
> have the symmetric shape the placement relies on, enabling is refused
> and any active mode degrades cleanly to the original flat behavior.

I was also hoping to see what this interface looks like and maybe
discuss how we should relay the information to the users, since this
seems to be a new addition from the RFC.

> Measured results:
> 
> System Configuration:
> - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)

I think a description of this system's topology would help me understand
the results below a bit better : -)

> 1) Throughput (System Bandwidth)
>    - DRAM Only: 966 GB/s
>    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
>    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
>      (38% increase compared to DRAM Only,
>       47% increase compared to Weighted Interleave)
> 
> 2) Loaded Latency (Under High Bandwidth)
>    - DRAM Only: 544 ns
>    - Weighted Interleave: 545 ns
>    - Package-Aware Weighted Interleave: 436 ns
>      (20% reduction compared to both)

Really awesome results!

> A small CXL driver change registers a CXL memory node into its package
> as the node comes online, using the initiator the driver resolves for
> the region; this is where the package layer gets the CPU-side
> association that plain NUMA distance does not carry.
> 
> The memory_package layer offers a broader interface for grouping and
> querying package topology - usable by memory tiering as well - and
> package-aware weighted interleave uses the subset it needs.

I was hoping you could expand on this a bit more. Aside from the
alloction-time placement strategy, did you have other ideas in mind for
who could ingest the package information to make tiering decisions?

I definitely think this series makes a lot of sense and I am
hoping to hear more about it. Thank you, I hope you have a great day!

Joshua
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Tue, 11 Aug 2026 08:29:51 -0700 Joshua Hahn <joshua.hahnjy@gmail.com> wrote:
> On Thu,  6 Aug 2026 17:09:31 +0900 Rakie Kim <rakie.kim@sk.com> wrote:
> 
> > Package-aware weighted interleave places a task's weighted-interleave
> > pages on the NUMA nodes of its local package, so that interleave traffic
> > does not have to cross the interconnect to another package. This keeps
> > each node's weight aligned with the bandwidth the task actually gets
> > from it, so effective bandwidth holds up on a system that has more than
> > one package. (A package is a CPU socket together with the memory
> > attached to it.)
> > 
> > Changes from RFC:
> > https://lore.kernel.org/all/20260316051258.246-1-rakie.kim@sk.com/
> > - Added an opt-in sysfs toggle (off by default) and a read-only sysfs
> >   view of the package topology
> > - Added topology validation with a clean fallback to plain weighted
> >   interleave on unsupported topologies
> > - Hardened the allocation, device-teardown, and node-hotplug paths
> 
> Hello Rakie,
>
> I hope you are doing well! Sorry for the late repsonse.
>
> I have a few thoughts, some of which are carry-overs from the RFC
> discussion we had before! I think there are still some open questions,
> and I wanted to get your opinion on some of them.
>

Hello Joshua,

I am doing well, thank you, and I hope you are too. Thank you for
taking the time to review this series and for following up on the
questions from the RFC discussion.

> My first question is whether we want cross-socket allocations at all.
> The examples you gave seem to line up with node-restricted interleave,
> as opposed to cross-socket interleave. I think the wording that you
> use to describe the feature in 4/4 (which I will copy below)
> 
> > The resolved mask is by construction a subset of the policy nodemask, which
> > mempolicy already restricts to the task's cpuset; package mode can only
> > narrow that set, never widen it, so cpusets and the task nodemask remain
> > authoritative.
> 
> is 100% the right way to treat these package-aware (socket-aware)
> interleaving allocations, but the example below
> 
> [...snip...]
> 
> > Applied the same way to every source, these weights give the map:
> > 
> >               node0  node1  node2  node3
> > global:         2      2      1      1
> 
> [...snip...]
> 
> >               node0  node1  node2  node3
> > from CPU 0:     2      0      1      0
> > from CPU 1:     0      2      0      1
> 
> Is essentially the existing weighted interleave mechanism with a
> nodemask/cpuset applied.

The example I gave was not explained well enough, and I can see how
it reads as a manually applied nodemask.

A nodemask or a cpuset names a fixed set of nodes, while package mode
expresses a rule: use the nodes of the package the allocation is
requested from. The mask is resolved per allocation from the
requesting CPU, so a single policy gives {0,2} to a thread on package
0 and {1,3} to a thread on package 1 at the same time. One nodemask
cannot do that, since it is the same set for everyone who uses the
policy.

There is also the question of how a user would build such a nodemask.
The package a CXL node belongs to is not visible today: on the
systems I tested, the firmware reports node1 as the initiator for
both CXL nodes. The topology layer in this series is what makes that
association available, and the read-only view under
/sys/devices/system/package/ lets the user check it.

> With that said, I think a more interesting and
> illustrative example would be if the user truly would want to allow some
> allocations to go through cross-socket, but be able to control the
> ratio at which these slip through.
>
>               node0  node1  node2  node3
> from CPU 0:     3      1      2      0
> from CPU 1:     0      3      1      2
>
> Maybe even more illustrative of the true capabilities of this series
> would be if you have an asymmetric system where you bind some
> host-level monitoring / logging workloads to one node (say, node0) and
> want that to be able to cross through to the other socket, but not the
> other way around:
> 
>               node0  node1  node2  node3
> from CPU 0:     3      1      2      0
> from CPU 1:     0      2      0      1
> 
> Anyways, these are just super hypothetical scenarios and I don't even
> know if the configuration that I'm listing would really be beneficial
> for the system. I think that coming up with some illustrative usecases
> which are now made possible by this series could help motivate why we
> would want to interleave across sockets.
>

These maps are an interesting idea, and I would like to look at them
with you.

This series only narrows the candidate nodes; the weights themselves
stay global, so every source that reaches a node uses the same weight
for it. Both of your maps give a node a different weight depending on
which package the allocation comes from, so the weight table would
have to become per source rather than a single global one.

Encoding the weights that way came up in an earlier stage of this
work, and it was mentioned again briefly in the RFC thread. As I
recall, the difficulty then was less the placement logic than how a
user would drive it: weights would have to be configured for every
source, so both the interface and the structure behind it grow
considerably.

That does not make your suggestion less interesting to me. I think it
could work well once there are clear scenarios for it, and the
grouping added here is what such a table would be built on, since a
per source weight only has meaning when the kernel knows which
package each node belongs to. What I am unsure about is folding it
into this series, whose aim is the narrower one of raising effective
bandwidth by keeping interleave traffic within a package. Allowing a
controlled amount of cross-package traffic points the other way, so I
think it is a topic we could discuss separately, with the use cases
worked out first.

I agree that such use cases would make the direction much stronger,
and I will think about whether there are cases where allowing a
controlled amount of cross-package traffic would help.


> > A task on CPU 0 now places pages on node0 (weight 2) and node2
> > (weight 1) at 2:1, which matches their effective bandwidth of 300:150;
> > a task on CPU 1 places on node1 and node3 the same way. Placement
> > follows the bandwidth each task actually sees, NUMA locality is
> > preserved, and interleave traffic stays off the interconnect.
> > 
> > To make this possible the kernel needs a notion of which nodes share a
> > package. The NUMA distance model offers only relative latencies and no
> > structural grouping, which is especially limiting for CXL memory nodes
> > that come online without an explicit package association.
> > 
> > The series adds a package-aware topology layer that groups CPU and
> > memory-only nodes into a "memory package", built from the physical
> > package ids firmware reports and, for a memory-only node, an initiator
> > CPU node or SLIT distances. A package can contain more than one CPU node
> > or more than one memory-only node, so the layer maps a package to a set
> > of nodes rather than to a single node or a single CXL device.
> > 
> > The feature is off by default and opt-in through a sysfs toggle. The
> > package topology itself is exposed read-only under
> > /sys/devices/system/package/; there is deliberately no writable
> > override, since a machine whose firmware describes its topology
> > incorrectly should be fixed in firmware. On a topology that does not
> > have the symmetric shape the placement relies on, enabling is refused
> > and any active mode degrades cleanly to the original flat behavior.
> 
> I was also hoping to see what this interface looks like and maybe
> discuss how we should relay the information to the users, since this
> seems to be a new addition from the RFC.
>

Sure. The toggle lives with the existing weighted interleave knobs.
package_mode defaults to false, so nothing changes until the
operator explicitly enables it:

/sys/kernel/mm/mempolicy/weighted_interleave
|-- auto
|-- node0
|-- node1
|-- node2
|-- node3
`-- package_mode -> true/false

The package topology view is read-only and lives under
/sys/devices/system/package/. This is how it looks on the system I
am currently using:

/sys/devices/system/package
|-- package0
|   |-- package_cpu_nodes -> 0
|   |-- package_mem_only_nodes -> 2
|   |-- package_nodes -> 0,2
|   `-- physical_package_id -> 0
`-- package1
    |-- package_cpu_nodes -> 1
    |-- package_mem_only_nodes -> 3
    |-- package_nodes -> 1,3
    `-- physical_package_id -> 1

package_nodes shows every node grouped into that package, and the
cpu/mem_only files split them by type, so an operator can check how
the kernel grouped the topology before turning package_mode on. I
will update the documentation in the next version to describe this
interface and how to use it.


> > Measured results:
> > 
> > System Configuration:
> > - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> 
> I think a description of this system's topology would help me understand
> the results below a bit better : -)
>

That is a fair point. The system used for the measurements is
configured as follows:

- Processor:                 Dual-Socket Intel Xeon 6980P
                             (Granite Rapids)
- Local memory (per socket): 12 channels, DDR5-6400
- CXL memory (per socket):   8 channels, DDR5-6400

It boots as two CPU+DRAM nodes and two CXL memory-only nodes, which
is the topology shown in the sysfs output above. I will add this
description to the measured results in the next version.


> > 1) Throughput (System Bandwidth)
> >    - DRAM Only: 966 GB/s
> >    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
> >    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
> >      (38% increase compared to DRAM Only,
> >       47% increase compared to Weighted Interleave)
> > 
> > 2) Loaded Latency (Under High Bandwidth)
> >    - DRAM Only: 544 ns
> >    - Weighted Interleave: 545 ns
> >    - Package-Aware Weighted Interleave: 436 ns
> >      (20% reduction compared to both)
> 
> Really awesome results!
>

Thank you.


> > A small CXL driver change registers a CXL memory node into its package
> > as the node comes online, using the initiator the driver resolves for
> > the region; this is where the package layer gets the CPU-side
> > association that plain NUMA distance does not carry.
> > 
> > The memory_package layer offers a broader interface for grouping and
> > querying package topology - usable by memory tiering as well - and
> > package-aware weighted interleave uses the subset it needs.
> 
> I was hoping you could expand on this a bit more. Aside from the
> alloction-time placement strategy, did you have other ideas in mind for
> who could ingest the package information to make tiering decisions?
>

The case I had in mind is demotion and promotion target selection.
With the package information, tiering could keep those decisions
within a package: choosing the memory-only nodes of the task's
package as demotion targets, and symmetrically preferring the
package's CPU nodes when promoting, so that both hot and cold pages
stay close to the CPUs that use them.

To support this, the layer already exposes per-node "preferred" node
queries: for a CPU node it reports the nearest memory-only nodes in
the same package, and for a memory-only node the nearest CPU nodes.
Nothing consumes them yet; I kept them out of the placement path so
that tiering can adopt them separately when there is a real user.

> I definitely think this series makes a lot of sense and I am
> hoping to hear more about it. Thank you, I hope you have a great day!
>
> Joshua

Thank you again for the careful review and for the questions; they
were a great help in seeing what the cover letter needs to explain
better. I hope you have a great day too.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Joshua Hahn 1 month, 2 weeks ago
> Hello Joshua,
> 
> I am doing well, thank you, and I hope you are too. Thank you for
> taking the time to review this series and for following up on the
> questions from the RFC discussion.
> 
> > My first question is whether we want cross-socket allocations at all.
> > The examples you gave seem to line up with node-restricted interleave,
> > as opposed to cross-socket interleave. I think the wording that you
> > use to describe the feature in 4/4 (which I will copy below)
> > 
> > > The resolved mask is by construction a subset of the policy nodemask, which
> > > mempolicy already restricts to the task's cpuset; package mode can only
> > > narrow that set, never widen it, so cpusets and the task nodemask remain
> > > authoritative.
> > 
> > is 100% the right way to treat these package-aware (socket-aware)
> > interleaving allocations, but the example below
> > 
> > [...snip...]
> > 
> > > Applied the same way to every source, these weights give the map:
> > > 
> > >               node0  node1  node2  node3
> > > global:         2      2      1      1
> > 
> > [...snip...]
> > 
> > >               node0  node1  node2  node3
> > > from CPU 0:     2      0      1      0
> > > from CPU 1:     0      2      0      1
> > 
> > Is essentially the existing weighted interleave mechanism with a
> > nodemask/cpuset applied.
> 
> The example I gave was not explained well enough, and I can see how
> it reads as a manually applied nodemask.
> 
> A nodemask or a cpuset names a fixed set of nodes, while package mode
> expresses a rule: use the nodes of the package the allocation is
> requested from. The mask is resolved per allocation from the
> requesting CPU, so a single policy gives {0,2} to a thread on package
> 0 and {1,3} to a thread on package 1 at the same time. One nodemask
> cannot do that, since it is the same set for everyone who uses the
> policy.

Ah! I'm sorry. It seems I totally misunderstood the intent of the
series. I think that my brain short-circuted to the discussion at
LSFMMBPF from 2025, where I think we discussed having a real 2-D
grid with weights per-node, per-CPU. I think my confusion is responsible
for the examples below, which as I understand it now, are not the intent
of the series.

> There is also the question of how a user would build such a nodemask.
> The package a CXL node belongs to is not visible today: on the
> systems I tested, the firmware reports node1 as the initiator for
> both CXL nodes. The topology layer in this series is what makes that
> association available, and the read-only view under
> /sys/devices/system/package/ lets the user check it.

That makes sense. Now I really see the goal of the series and it makes
a lot more sense. Thank you for the clarification. 

> > With that said, I think a more interesting and
> > illustrative example would be if the user truly would want to allow some
> > allocations to go through cross-socket, but be able to control the
> > ratio at which these slip through.
> >
> >               node0  node1  node2  node3
> > from CPU 0:     3      1      2      0
> > from CPU 1:     0      3      1      2
> >
> > Maybe even more illustrative of the true capabilities of this series
> > would be if you have an asymmetric system where you bind some
> > host-level monitoring / logging workloads to one node (say, node0) and
> > want that to be able to cross through to the other socket, but not the
> > other way around:
> > 
> >               node0  node1  node2  node3
> > from CPU 0:     3      1      2      0
> > from CPU 1:     0      2      0      1
> > 
> > Anyways, these are just super hypothetical scenarios and I don't even
> > know if the configuration that I'm listing would really be beneficial
> > for the system. I think that coming up with some illustrative usecases
> > which are now made possible by this series could help motivate why we
> > would want to interleave across sockets.
> >
> 
> These maps are an interesting idea, and I would like to look at them
> with you.
> 
> This series only narrows the candidate nodes; the weights themselves
> stay global, so every source that reaches a node uses the same weight
> for it. Both of your maps give a node a different weight depending on
> which package the allocation comes from, so the weight table would
> have to become per source rather than a single global one.
> 
> Encoding the weights that way came up in an earlier stage of this
> work, and it was mentioned again briefly in the RFC thread. As I
> recall, the difficulty then was less the placement logic than how a
> user would drive it: weights would have to be configured for every
> source, so both the interface and the structure behind it grow
> considerably.

Yeah, I can imagine it is quite a lot of tuning that users have to do.
So I'm 100% on board for the goal of this series to make the existing
weighted interleave mechanism respect the initiator's POV. Sorry for
making you explain all of this, this confusion is just due to my
misunderstanding.

> That does not make your suggestion less interesting to me. I think it
> could work well once there are clear scenarios for it, and the
> grouping added here is what such a table would be built on, since a
> per source weight only has meaning when the kernel knows which
> package each node belongs to. What I am unsure about is folding it
> into this series, whose aim is the narrower one of raising effective
> bandwidth by keeping interleave traffic within a package. Allowing a
> controlled amount of cross-package traffic points the other way, so I
> think it is a topic we could discuss separately, with the use cases
> worked out first.

Thanks! Actually I think we can wait on this until we have real
usecases where we prefer to make cross-socket allocations.

> > I was also hoping to see what this interface looks like and maybe
> > discuss how we should relay the information to the users, since this
> > seems to be a new addition from the RFC.
> >
> 
> Sure. The toggle lives with the existing weighted interleave knobs.
> package_mode defaults to false, so nothing changes until the
> operator explicitly enables it:
> 
> /sys/kernel/mm/mempolicy/weighted_interleave
> |-- auto
> |-- node0
> |-- node1
> |-- node2
> |-- node3
> `-- package_mode -> true/false
> 
> The package topology view is read-only and lives under
> /sys/devices/system/package/. This is how it looks on the system I
> am currently using:
> 
> /sys/devices/system/package
> |-- package0
> |   |-- package_cpu_nodes -> 0
> |   |-- package_mem_only_nodes -> 2
> |   |-- package_nodes -> 0,2
> |   `-- physical_package_id -> 0
> `-- package1
>     |-- package_cpu_nodes -> 1
>     |-- package_mem_only_nodes -> 3
>     |-- package_nodes -> 1,3
>     `-- physical_package_id -> 1
> 
> package_nodes shows every node grouped into that package, and the
> cpu/mem_only files split them by type, so an operator can check how
> the kernel grouped the topology before turning package_mode on. I
> will update the documentation in the next version to describe this
> interface and how to use it.

Great, I think this would be a great addition to add to the cover
letter and also add as documentation, since it is user-facing. 

> > > Measured results:
> > > 
> > > System Configuration:
> > > - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> > 
> > I think a description of this system's topology would help me understand
> > the results below a bit better : -)
> >
> 
> That is a fair point. The system used for the measurements is
> configured as follows:
> 
> - Processor:                 Dual-Socket Intel Xeon 6980P
>                              (Granite Rapids)
> - Local memory (per socket): 12 channels, DDR5-6400
> - CXL memory (per socket):   8 channels, DDR5-6400
> 
> It boots as two CPU+DRAM nodes and two CXL memory-only nodes, which
> is the topology shown in the sysfs output above. I will add this
> description to the measured results in the next version.

Thanks. Notably I wanted to see if the DDR generation was the same
across DRAM and CXL. 

> The case I had in mind is demotion and promotion target selection.
> With the package information, tiering could keep those decisions
> within a package: choosing the memory-only nodes of the task's
> package as demotion targets, and symmetrically preferring the
> package's CPU nodes when promoting, so that both hot and cold pages
> stay close to the CPUs that use them.

Yeah, I like this idea a lot.

For demotion, we would just chnage the fallback zonelist based on the
sockets.

I think we actually get promotions for free, since if this series is
doing a good job of allocating memory close to the consuming CPU, and
the demotions prevent the memory from moving cross-socket, initiators
should only promote (NUMAB2 promotion) memory that is socket-local.

> To support this, the layer already exposes per-node "preferred" node
> queries: for a CPU node it reports the nearest memory-only nodes in
> the same package, and for a memory-only node the nearest CPU nodes.
> Nothing consumes them yet; I kept them out of the placement path so
> that tiering can adopt them separately when there is a real user.
> 
> > I definitely think this series makes a lot of sense and I am
> > hoping to hear more about it. Thank you, I hope you have a great day!
> >
> > Joshua
> 
> Thank you again for the careful review and for the questions; they
> were a great help in seeing what the cover letter needs to explain
> better. I hope you have a great day too.

Thank you Rakie. I don't think the cover letter was misleading,
it was just my fault for short-circuiting and thinking the series was
about adding per-socket per-node weights, as opposed to the
restriction that you're adding to the existing weights.

If I may add one more comment, I think 2/4 is a bit hard to review.
A 1k line patch is not so easy to see the full picture, I think it would
make it less intimidating to review if it could be split up into
smaller patches. Just my 2c : -)

Thanks again. I hope you have a great day!
Joshua
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 2 weeks ago
On Wed, 12 Aug 2026 07:49:55 -0700 Joshua Hahn <joshua.hahnjy@gmail.com> wrote:

Hello Joshua,

Thank you for coming back to this so quickly.

> > Hello Joshua,
> > 
> > I am doing well, thank you, and I hope you are too. Thank you for
> > taking the time to review this series and for following up on the
> > questions from the RFC discussion.
> > 
> > > My first question is whether we want cross-socket allocations at all.
> > > The examples you gave seem to line up with node-restricted interleave,
> > > as opposed to cross-socket interleave. I think the wording that you
> > > use to describe the feature in 4/4 (which I will copy below)
> > > 
> > > > The resolved mask is by construction a subset of the policy nodemask, which
> > > > mempolicy already restricts to the task's cpuset; package mode can only
> > > > narrow that set, never widen it, so cpusets and the task nodemask remain
> > > > authoritative.
> > > 
> > > is 100% the right way to treat these package-aware (socket-aware)
> > > interleaving allocations, but the example below
> > > 
> > > [...snip...]
> > > 
> > > > Applied the same way to every source, these weights give the map:
> > > > 
> > > >               node0  node1  node2  node3
> > > > global:         2      2      1      1
> > > 
> > > [...snip...]
> > > 
> > > >               node0  node1  node2  node3
> > > > from CPU 0:     2      0      1      0
> > > > from CPU 1:     0      2      0      1
> > > 
> > > Is essentially the existing weighted interleave mechanism with a
> > > nodemask/cpuset applied.
> > 
> > The example I gave was not explained well enough, and I can see how
> > it reads as a manually applied nodemask.
> > 
> > A nodemask or a cpuset names a fixed set of nodes, while package mode
> > expresses a rule: use the nodes of the package the allocation is
> > requested from. The mask is resolved per allocation from the
> > requesting CPU, so a single policy gives {0,2} to a thread on package
> > 0 and {1,3} to a thread on package 1 at the same time. One nodemask
> > cannot do that, since it is the same set for everyone who uses the
> > policy.
> 
> Ah! I'm sorry. It seems I totally misunderstood the intent of the
> series. I think that my brain short-circuted to the discussion at
> LSFMMBPF from 2025, where I think we discussed having a real 2-D
> grid with weights per-node, per-CPU. I think my confusion is responsible
> for the examples below, which as I understand it now, are not the intent
> of the series.
>

My explanation was not enough and that is what caused the confusion.
The 2-D grid is close enough to this work that the two are easy to
place together, and thanks to your questions I could fill in a good
deal of what the cover letter was missing.


> > There is also the question of how a user would build such a nodemask.
> > The package a CXL node belongs to is not visible today: on the
> > systems I tested, the firmware reports node1 as the initiator for
> > both CXL nodes. The topology layer in this series is what makes that
> > association available, and the read-only view under
> > /sys/devices/system/package/ lets the user check it.
> 
> That makes sense. Now I really see the goal of the series and it makes
> a lot more sense. Thank you for the clarification.
>

Thank you.


> > > With that said, I think a more interesting and
> > > illustrative example would be if the user truly would want to allow some
> > > allocations to go through cross-socket, but be able to control the
> > > ratio at which these slip through.
> > >
> > >               node0  node1  node2  node3
> > > from CPU 0:     3      1      2      0
> > > from CPU 1:     0      3      1      2
> > >
> > > Maybe even more illustrative of the true capabilities of this series
> > > would be if you have an asymmetric system where you bind some
> > > host-level monitoring / logging workloads to one node (say, node0) and
> > > want that to be able to cross through to the other socket, but not the
> > > other way around:
> > > 
> > >               node0  node1  node2  node3
> > > from CPU 0:     3      1      2      0
> > > from CPU 1:     0      2      0      1
> > > 
> > > Anyways, these are just super hypothetical scenarios and I don't even
> > > know if the configuration that I'm listing would really be beneficial
> > > for the system. I think that coming up with some illustrative usecases
> > > which are now made possible by this series could help motivate why we
> > > would want to interleave across sockets.
> > >
> > 
> > These maps are an interesting idea, and I would like to look at them
> > with you.
> > 
> > This series only narrows the candidate nodes; the weights themselves
> > stay global, so every source that reaches a node uses the same weight
> > for it. Both of your maps give a node a different weight depending on
> > which package the allocation comes from, so the weight table would
> > have to become per source rather than a single global one.
> > 
> > Encoding the weights that way came up in an earlier stage of this
> > work, and it was mentioned again briefly in the RFC thread. As I
> > recall, the difficulty then was less the placement logic than how a
> > user would drive it: weights would have to be configured for every
> > source, so both the interface and the structure behind it grow
> > considerably.
> 
> Yeah, I can imagine it is quite a lot of tuning that users have to do.
> So I'm 100% on board for the goal of this series to make the existing
> weighted interleave mechanism respect the initiator's POV. Sorry for
> making you explain all of this, this confusion is just due to my
> misunderstanding.
>

Thank you. "Respect the initiator's point of view" describes the goal
better than what I wrote, so I would like to use that framing in the
next cover letter.


> > That does not make your suggestion less interesting to me. I think it
> > could work well once there are clear scenarios for it, and the
> > grouping added here is what such a table would be built on, since a
> > per source weight only has meaning when the kernel knows which
> > package each node belongs to. What I am unsure about is folding it
> > into this series, whose aim is the narrower one of raising effective
> > bandwidth by keeping interleave traffic within a package. Allowing a
> > controlled amount of cross-package traffic points the other way, so I
> > think it is a topic we could discuss separately, with the use cases
> > worked out first.
> 
> Thanks! Actually I think we can wait on this until we have real
> usecases where we prefer to make cross-socket allocations.
>

Agreed. I will also keep thinking about what such use cases would
look like.


> > > I was also hoping to see what this interface looks like and maybe
> > > discuss how we should relay the information to the users, since this
> > > seems to be a new addition from the RFC.
> > >
> > 
> > Sure. The toggle lives with the existing weighted interleave knobs.
> > package_mode defaults to false, so nothing changes until the
> > operator explicitly enables it:
> > 
> > /sys/kernel/mm/mempolicy/weighted_interleave
> > |-- auto
> > |-- node0
> > |-- node1
> > |-- node2
> > |-- node3
> > `-- package_mode -> true/false
> > 
> > The package topology view is read-only and lives under
> > /sys/devices/system/package/. This is how it looks on the system I
> > am currently using:
> > 
> > /sys/devices/system/package
> > |-- package0
> > |   |-- package_cpu_nodes -> 0
> > |   |-- package_mem_only_nodes -> 2
> > |   |-- package_nodes -> 0,2
> > |   `-- physical_package_id -> 0
> > `-- package1
> >     |-- package_cpu_nodes -> 1
> >     |-- package_mem_only_nodes -> 3
> >     |-- package_nodes -> 1,3
> >     `-- physical_package_id -> 1
> > 
> > package_nodes shows every node grouped into that package, and the
> > cpu/mem_only files split them by type, so an operator can check how
> > the kernel grouped the topology before turning package_mode on. I
> > will update the documentation in the next version to describe this
> > interface and how to use it.
> 
> Great, I think this would be a great addition to add to the cover
> letter and also add as documentation, since it is user-facing.
>

I will put it in both. Andrew also asked for documentation aimed at
the operator, so the next version will describe this interface and
how to use it there as well.


> > > > Measured results:
> > > > 
> > > > System Configuration:
> > > > - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> > > 
> > > I think a description of this system's topology would help me understand
> > > the results below a bit better : -)
> > >
> > 
> > That is a fair point. The system used for the measurements is
> > configured as follows:
> > 
> > - Processor:                 Dual-Socket Intel Xeon 6980P
> >                              (Granite Rapids)
> > - Local memory (per socket): 12 channels, DDR5-6400
> > - CXL memory (per socket):   8 channels, DDR5-6400
> > 
> > It boots as two CPU+DRAM nodes and two CXL memory-only nodes, which
> > is the topology shown in the sysfs output above. I will add this
> > description to the measured results in the next version.
> 
> Thanks. Notably I wanted to see if the DDR generation was the same
> across DRAM and CXL.
>

Both sides are DDR5-6400, so the difference in the results comes from
the path rather than from the memory itself. I will make that clear
when I describe the system in the next version.


> > The case I had in mind is demotion and promotion target selection.
> > With the package information, tiering could keep those decisions
> > within a package: choosing the memory-only nodes of the task's
> > package as demotion targets, and symmetrically preferring the
> > package's CPU nodes when promoting, so that both hot and cold pages
> > stay close to the CPUs that use them.
> 
> Yeah, I like this idea a lot.
> 
> For demotion, we would just chnage the fallback zonelist based on the
> sockets.
> 
> I think we actually get promotions for free, since if this series is
> doing a good job of allocating memory close to the consuming CPU, and
> the demotions prevent the memory from moving cross-socket, initiators
> should only promote (NUMAB2 promotion) memory that is socket-local.
>

Thank you for the suggestion. Changing the demotion order by package
sounds like the natural first step, and I will look into it once the
placement side has settled.


> > To support this, the layer already exposes per-node "preferred" node
> > queries: for a CPU node it reports the nearest memory-only nodes in
> > the same package, and for a memory-only node the nearest CPU nodes.
> > Nothing consumes them yet; I kept them out of the placement path so
> > that tiering can adopt them separately when there is a real user.
> > 
> > > I definitely think this series makes a lot of sense and I am
> > > hoping to hear more about it. Thank you, I hope you have a great day!
> > >
> > > Joshua
> > 
> > Thank you again for the careful review and for the questions; they
> > were a great help in seeing what the cover letter needs to explain
> > better. I hope you have a great day too.
> 
> Thank you Rakie. I don't think the cover letter was misleading,
> it was just my fault for short-circuiting and thinking the series was
> about adding per-socket per-node weights, as opposed to the
> restriction that you're adding to the existing weights.
>

Thank you for saying so. Either way, your questions gave me a chance
to look again at what the cover letter was not saying clearly.

> If I may add one more comment, I think 2/4 is a bit hard to review.
> A 1k line patch is not so easy to see the full picture, I think it would
> make it less intimidating to review if it could be split up into
> smaller patches. Just my 2c : -)
>

You are right, it is too much to take in at once. The patch became
large and complex because several features ended up in a single
commit. I will separate them as much as I can in the next version.

> Thanks again. I hope you have a great day!
> Joshua

Thank you again for the review and for the discussion. I hope you
have a great day too.

Rakie Kim
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Andrew Morton 1 month, 3 weeks ago
On Thu,  6 Aug 2026 17:09:31 +0900 Rakie Kim <rakie.kim@sk.com> wrote:

> Package-aware weighted interleave places a task's weighted-interleave
> pages on the NUMA nodes of its local package, so that interleave traffic
> does not have to cross the interconnect to another package. This keeps
> each node's weight aligned with the bandwidth the task actually gets
> from it, so effective bandwidth holds up on a system that has more than
> one package. (A package is a CPU socket together with the memory
> attached to it.)

"package" is not a familiar term in MM.  It would be helpful if the
[0/N] were to carefully and fully define/describe the new term before
using it 40 times!

> Measured results:
> 
> System Configuration:
> - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> 
> 1) Throughput (System Bandwidth)
>    - DRAM Only: 966 GB/s
>    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
>    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
>      (38% increase compared to DRAM Only,
>       47% increase compared to Weighted Interleave)
> 
> 2) Loaded Latency (Under High Bandwidth)
>    - DRAM Only: 544 ns
>    - Weighted Interleave: 545 ns
>    - Package-Aware Weighted Interleave: 436 ns
>      (20% reduction compared to both)

Well that sounds nice.

>  .../ABI/testing/sysfs-devices-system-package  |   35 +
>  ...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
>  drivers/cxl/core/region.c                     |   54 +
>  drivers/cxl/cxl.h                             |    1 +
>  drivers/dax/kmem.c                            |    3 +
>  include/linux/memory-tiers.h                  |  113 ++
>  include/linux/numa.h                          |   11 +
>  mm/memory-tiers.c                             | 1009 +++++++++++++++++
>  mm/mempolicy.c                                |  200 +++-

Are some user-facing Documentation/ updates appropriate?

The Documentation/ABI things are rather dry and information-free.  How
about some documentation for the operator who is wondering "should I
use this and if so why and how"?


The runtime sysfs on/off tunable is interesting.  I hear from google
operations people that every new feature should have such an "off"
switch so that if development send them a new thing and they think it's
problematic, they can disable it in order to quickly get back to the
old regime.  Perhaps that was your motivation, perhaps not.  Can you
please describe?


I see you've been emailed the Sashiko report, which appears substantial.
	https://sashiko.dev/#/patchset/20260806080936.421-1-rakie.kim@sk.com
Re: [PATCH 0/4] mm/mempolicy: introduce package-aware weighted interleave
Posted by Rakie Kim 1 month, 3 weeks ago
On Thu, 6 Aug 2026 14:38:39 -0700 Andrew Morton <akpm@linux-foundation.org> wrote:
> On Thu,  6 Aug 2026 17:09:31 +0900 Rakie Kim <rakie.kim@sk.com> wrote:
>

Hello Andrew,

Thank you for taking the time to review this series.

> > Package-aware weighted interleave places a task's weighted-interleave
> > pages on the NUMA nodes of its local package, so that interleave traffic
> > does not have to cross the interconnect to another package. This keeps
> > each node's weight aligned with the bandwidth the task actually gets
> > from it, so effective bandwidth holds up on a system that has more than
> > one package. (A package is a CPU socket together with the memory
> > attached to it.)
>
> "package" is not a familiar term in MM.  It would be helpful if the
> [0/N] were to carefully and fully define/describe the new term before
> using it 40 times!
>

You are right. I think my explanation was not sufficient. When I
prepared this series, I went back and forth between "socket" and
other candidate terms, and settled on "package" because modern
processors can contain multiple NUMA nodes and dies within a single
physical socket, so "socket" felt misleading. I did not explain this
reasoning in the cover letter. In the next version, I will define
the term at the top of the cover letter before it is used.

> > Measured results:
> >
> > System Configuration:
> > - Processor: Dual-Socket Intel Xeon 6980P (Granite Rapids)
> >
> > 1) Throughput (System Bandwidth)
> >    - DRAM Only: 966 GB/s
> >    - Weighted Interleave: 903 GB/s (7% decrease compared to DRAM Only)
> >    - Package-Aware Weighted Interleave: 1329 GB/s (1.33 TB/s)
> >      (38% increase compared to DRAM Only,
> >       47% increase compared to Weighted Interleave)
> >
> > 2) Loaded Latency (Under High Bandwidth)
> >    - DRAM Only: 544 ns
> >    - Weighted Interleave: 545 ns
> >    - Package-Aware Weighted Interleave: 436 ns
> >      (20% reduction compared to both)
>
> Well that sounds nice.
>

Thank you.

> >  .../ABI/testing/sysfs-devices-system-package  |   35 +
> >  ...fs-kernel-mm-mempolicy-weighted-interleave |   17 +
> >  drivers/cxl/core/region.c                     |   54 +
> >  drivers/cxl/cxl.h                             |    1 +
> >  drivers/dax/kmem.c                            |    3 +
> >  include/linux/memory-tiers.h                  |  113 ++
> >  include/linux/numa.h                          |   11 +
> >  mm/memory-tiers.c                             | 1009 +++++++++++++++++
> >  mm/mempolicy.c                                |  200 +++-
>
> Are some user-facing Documentation/ updates appropriate?
>
> The Documentation/ABI things are rather dry and information-free.  How
> about some documentation for the operator who is wondering "should I
> use this and if so why and how"?
>

I agree with your point. The current ABI entries only describe the
sysfs files themselves. In the next version, I will strengthen the
documentation content so that it answers exactly those questions for
an operator: whether this feature fits their system, why it helps,
and how to enable and verify it.

>
> The runtime sysfs on/off tunable is interesting.  I hear from google
> operations people that every new feature should have such an "off"
> switch so that if development send them a new thing and they think it's
> problematic, they can disable it in order to quickly get back to the
> old regime.  Perhaps that was your motivation, perhaps not.  Can you
> please describe?
>

You are right that this was one of the purposes. The switch was
provided with two goals in mind. First, as you describe, it is an
"off" switch for the new feature: if it behaves unexpectedly in
production, the operator can disable it at runtime and immediately
return to the old behavior, without a reboot. Second, it is for
users who want to keep using the existing weighted interleave as it
is: the feature is off by default, and nothing changes for them
unless they explicitly turn it on. I will describe this motivation
in the documentation as well.

>
> I see you've been emailed the Sashiko report, which appears substantial.
> 	https://sashiko.dev/#/patchset/20260806080936.421-1-rakie.kim@sk.com
>

Yes, I have received the report, and it seems to provide a lot of
good information. I am reviewing each finding against the code, and
I plan to reflect it in the next version as much as possible.

Thanks again for your time and review.

Rakie Kim