[PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities

Florian Bezdeka posted 3 patches 1 month, 1 week ago
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 21 +++++++++++++++++----
kernel/irq/manage.c                               | 10 ++++++++++
lib/cpumask.c                                     | 10 ++++++----
3 files changed, 33 insertions(+), 8 deletions(-)
[PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Florian Bezdeka 1 month, 1 week ago
Hi all,

I'm trying to demonstrate a real problem for PREEMPT_RT here, using the
stmmac driver as example. There are more drivers "affected" but let's
ignore that for one moment. Let's discuss the underlying problem first.

To achieve the best throughput most network devices are based on
multiple queues. Each queue (pair) is equipped with one device IRQ. To
reach the maximum throughput, spreading / balancing them between all the
available CPUs makes sense.

While some IRQ chips - and with that some architectures - implement the
necessary spreading (or balancing) at IRQ chip level others don't do that.
If a device driver wants to make sure that balancing happens as intended
it has to implement that on his own (again).

The typical shortcoming of those implementations: They do not honor RT
relevant settings like the smp_default_affinity or isolated CPU cores.
Device IRQs are balanced over "all" or "all online CPUs".

That's not a problem for "normal" systems, but for RT systems - or
systems running cpu-isolating workloads - it is.

IRQ affinities can be controlled via /proc/irq/<n>/smp_affinity{_list}
for existing IRQs and via /proc/irq/default_smp_affinity for "new" or 
"not yet registered" IRQs. Those settings - as written by user space - 
must be honored. Always.

In our case the settings were bypassed by the following sequence:

    - system boot (all CPUs available, no isolation yet)
    - deployment of the first RT application
        - writing a new default smp affinity (remove RT isolated cores)
        - migrating away all that were targeting the now isolated cores
        - creating a cgroup with RT cores as the only usable cores
    - RT application running fine for some time
    - stmmac network interface went up for the first time
        - driver balances IRQs over all CPUs, ignoring the existing
	  affinities
    - Too much IRQ traffic on RT cores

Even with series applied there is (at least) one problem remaining:

The /proc/irq/<n> interface is populated on the first request_irq() call. 
That means that we can not control affinities from userspace until the 
IRQ gets requested.

The problem for network interfaces: We have to bring up the interfaces
once, to be able to control such affinities.

That raises the question why request_irq() is called on "link up" time,
while the low level vector allocation takes place during device probing.
At least that seems to be the common pattern. Can someone tell me why
this is done this way? Shouldn't we call request_irq() at the same time?

So, let's hope that all of this was short enough that somebody reads it
and precise enough to make the problem clear. The idea behind this series
is not fixing or applying the series as is. I'm expecting a discussion
first. So, input welcome!

Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
---
Florian Bezdeka (3):
      cpumask: Honor irq_default_affinity in cpumask_local_spread()
      genirq: Honor existing IRQ affinities when setting affinity hints
      net: stmmac: Migrate IRQ balancing to cpumask_local_spread()

 drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 21 +++++++++++++++++----
 kernel/irq/manage.c                               | 10 ++++++++++
 lib/cpumask.c                                     | 10 ++++++----
 3 files changed, 33 insertions(+), 8 deletions(-)
---
base-commit: aa2e13ae8d3cbe2c15ef4f7e971b2de0832794aa
change-id: 20260810-flo-net-7-2-make-stmmac-default-affinity-aware-a145a4941d8f

Best regards,
-- 
Florian Bezdeka <florian.bezdeka@siemens.com>
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Sebastian Andrzej Siewior 1 month, 1 week ago
On 2026-08-19 16:30:29 [+0200], Florian Bezdeka wrote:
…
> The typical shortcoming of those implementations: They do not honor RT
> relevant settings like the smp_default_affinity or isolated CPU cores.

"irqaffinity" if you refer to the boot command argument.
"default_smp_affinity" if you refer to the proc file.

> Device IRQs are balanced over "all" or "all online CPUs".
…
> That raises the question why request_irq() is called on "link up" time,
> while the low level vector allocation takes place during device probing.
> At least that seems to be the common pattern. Can someone tell me why
> this is done this way? Shouldn't we call request_irq() at the same time?

The IRQ vector is created while the system enumerates the IRQ-chips.
Once the devices are enumerated (such as the NICs) the devices is linked
with its IRQ. I think an exception are MSI-X devices which could ask for
one or more interrupt and then (at device's probe time) the PCI core
will link the requested amount of interrupts so their actual number
could change.
The driver _could_ request a "managed interrupt" which would be mapped
to a specific CPU. The difference to a "regular interrupt" is that if
that CPU goes down, the interrupt is not "moved" to another CPU. Instead
is remains off and the driver needs to deal with this (this is common
for NVME devices).

If the device is not programmed (as in IP address has been assigned,
link is up) then it should not create any interrupts. So it might be
reasonable to not request an interrupt either.
I *think* uarts do the same.

Sebastian
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Florian Bezdeka 1 month, 1 week ago
On Thu, 2026-08-20 at 17:12 +0200, Sebastian Andrzej Siewior wrote:
> On 2026-08-19 16:30:29 [+0200], Florian Bezdeka wrote:
> …
> > The typical shortcoming of those implementations: They do not honor RT
> > relevant settings like the smp_default_affinity or isolated CPU cores.
> 
> "irqaffinity" if you refer to the boot command argument.
> "default_smp_affinity" if you refer to the proc file.

The latter. Sorry.

All the system configuration has to happen during runtime, not during
boot time. There is no "fixed" or "static" configuration that is known
at boot time here.

> 
> > Device IRQs are balanced over "all" or "all online CPUs".
> …
> > That raises the question why request_irq() is called on "link up" time,
> > while the low level vector allocation takes place during device probing.
> > At least that seems to be the common pattern. Can someone tell me why
> > this is done this way? Shouldn't we call request_irq() at the same time?
> 
> The IRQ vector is created while the system enumerates the IRQ-chips.
> Once the devices are enumerated (such as the NICs) the devices is linked
> with its IRQ. I think an exception are MSI-X devices which could ask for
> one or more interrupt and then (at device's probe time) the PCI core
> will link the requested amount of interrupts so their actual number
> could change.
> The driver _could_ request a "managed interrupt" which would be mapped
> to a specific CPU. The difference to a "regular interrupt" is that if
> that CPU goes down, the interrupt is not "moved" to another CPU. Instead
> is remains off and the driver needs to deal with this (this is common
> for NVME devices).

I can't see yet how managed interrupts could help here. 

Those device IRQs can happily be migrated, co-located and of course we
want them to be configurable by userspace (/proc/irq/<n>/ interface),
which is also not possible for managed IRQs.

> 
> If the device is not programmed (as in IP address has been assigned,
> link is up) then it should not create any interrupts. So it might be
> reasonable to not request an interrupt either.
> I *think* uarts do the same.
> 

I'm fine with that - and thanks for explaining the current
implementation again, it matches my understanding. But: There must be a
way that userspace can configure affinities for "un-requested" IRQs
already.

As already mentioned the /proc/irq/<n> interface gets populated on
request_irq() time, which might be too late to be able to set an
affinity before the first IRQ arrives.

Please note that the default_smp_affinity can also change during
runtime, so that we have to honor the current value each time we do the
balancing / spreading.

Florian
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Andrew Lunn 1 month, 1 week ago
> That raises the question why request_irq() is called on "link up" time,
> while the low level vector allocation takes place during device probing.

If the interface is admin down, the hardware should not be generating
any interrupts. So there is no need to request them.

> At least that seems to be the common pattern. Can someone tell me why
> this is done this way? Shouldn't we call request_irq() at the same time?

If you want to change anything, move the low level vector allocation
into open(). But you need to be careful of EPROBE_DEFER. If the
interrupt controller has not loaded yet, i _guess_ the low level
vector allocation will return EPROBE_DEFER, and the MAC driver will
try to probe again later. If you get EPROBE_DEFER in open(), there is
nothing you can do about it.

	Andrew
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Florian Bezdeka 1 month, 1 week ago
Hi Andrew,

On Thu, 2026-08-20 at 01:54 +0200, Andrew Lunn wrote:
> > That raises the question why request_irq() is called on "link up" time,
> > while the low level vector allocation takes place during device probing.
> 
> If the interface is admin down, the hardware should not be generating
> any interrupts. So there is no need to request them.
> 
> > At least that seems to be the common pattern. Can someone tell me why
> > this is done this way? Shouldn't we call request_irq() at the same time?
> 
> If you want to change anything, move the low level vector allocation
> into open(). But you need to be careful of EPROBE_DEFER. If the
> interrupt controller has not loaded yet, i _guess_ the low level
> vector allocation will return EPROBE_DEFER, and the MAC driver will
> try to probe again later. If you get EPROBE_DEFER in open(), there is
> nothing you can do about it.
> 
> 	Andrew

Moving (=delaying) the vector allocation into open() would not help
here. The /proc/irq/<n> interface is populated on request_irq(),
normally done inside open() as well. Up to this point userspace is not
able to set any affinities.

This is one of the shortcomings mentioned in the cover letter. Once
/proc/irq/<n> appears, it might already be too late for a proper
affinity setting, the IRQ might have fired already. In addition there is
no notification mechanism that informs userspace about newly requested
IRQs.

It's more of the opposite that might help: Moving request_irq() into the
probe() phase. I'm wondering why the pattern is
    vector alloc => probe()
    request irq => open()

Seems that we would not "waste" too much resources when we move
request_irq() calls into probe() as well. But as always: I might miss
something.

Florian
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Andrew Lunn 1 month, 1 week ago
The problem with using stmmac as an example is that it is very old. It
was added to Linux in 2009. I guess most of the SoCs at that time were
single or dual core. CPU affinity was not something developers thought
about for that class of SoC.

Over time the number of cores has gone up and stmmac has got faster
link speeds. But is the architecture correct? Is the driver following
best practices?

I would suggest you look at more modern MAC drivers and see how they
do CPU affinity, etc. Is there anything which can be learned from them
and implemented in stmmac?

    Andrew
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Florian Bezdeka 1 month, 1 week ago
On Thu, 2026-08-20 at 02:10 +0200, Andrew Lunn wrote:
> The problem with using stmmac as an example is that it is very old. It
> was added to Linux in 2009. I guess most of the SoCs at that time were
> single or dual core. CPU affinity was not something developers thought
> about for that class of SoC.

Agree, stmmac might not be the best example, but this is the "driver
under test" here and it shows a pattern that is implemented in ~35
network drivers as well.

Users of irq_set_affinity_hint() in 7.2, see [1], but there is lot more
that needs a revisit, like irq_set_affinity_and_hint().

> 
> Over time the number of cores has gone up and stmmac has got faster
> link speeds. But is the architecture correct? Is the driver following
> best practices?

This are exactly the questions that I would like to discuss. Seems we're
getting closer.

> 
> I would suggest you look at more modern MAC drivers and see how they
> do CPU affinity, etc. Is there anything which can be learned from them
> and implemented in stmmac?

I'm still not sure if there is "the pattern" how modern drivers (btw:
which drivers do you consider "modern"?) implement that.

What I can tell: Most calls to irq_set_affinity_hint() (and friends) are
a real problem on RT enabled systems as they simply overwrite any
userspace settings.

Florian

[1] https://elixir.bootlin.com/linux/v7.2/A/ident/irq_set_affinity_hint
Re: [PATCH RFC 0/3] genirq: Allow drivers to respect userspace IRQ affinities
Posted by Yury Norov 1 month, 1 week ago
On Wed, Aug 19, 2026 at 04:30:29PM +0200, Florian Bezdeka wrote:
> Hi all,
> 
> I'm trying to demonstrate a real problem for PREEMPT_RT here, using the
> stmmac driver as example. There are more drivers "affected" but let's
> ignore that for one moment. Let's discuss the underlying problem first.
> 
> To achieve the best throughput most network devices are based on
> multiple queues. Each queue (pair) is equipped with one device IRQ. To
> reach the maximum throughput, spreading / balancing them between all the
> available CPUs makes sense.
> 
> While some IRQ chips - and with that some architectures - implement the
> necessary spreading (or balancing) at IRQ chip level others don't do that.
> If a device driver wants to make sure that balancing happens as intended
> it has to implement that on his own (again).
> 
> The typical shortcoming of those implementations: They do not honor RT
> relevant settings like the smp_default_affinity or isolated CPU cores.
> Device IRQs are balanced over "all" or "all online CPUs".
> 
> That's not a problem for "normal" systems, but for RT systems - or
> systems running cpu-isolating workloads - it is.
> 
> IRQ affinities can be controlled via /proc/irq/<n>/smp_affinity{_list}
> for existing IRQs and via /proc/irq/default_smp_affinity for "new" or 
> "not yet registered" IRQs. Those settings - as written by user space - 
> must be honored. Always.
> 
> In our case the settings were bypassed by the following sequence:
> 
>     - system boot (all CPUs available, no isolation yet)
>     - deployment of the first RT application
>         - writing a new default smp affinity (remove RT isolated cores)
>         - migrating away all that were targeting the now isolated cores
>         - creating a cgroup with RT cores as the only usable cores
>     - RT application running fine for some time
>     - stmmac network interface went up for the first time
>         - driver balances IRQs over all CPUs, ignoring the existing
> 	  affinities
>     - Too much IRQ traffic on RT cores
> 
> Even with series applied there is (at least) one problem remaining:
> 
> The /proc/irq/<n> interface is populated on the first request_irq() call. 
> That means that we can not control affinities from userspace until the 
> IRQ gets requested.
> 
> The problem for network interfaces: We have to bring up the interfaces
> once, to be able to control such affinities.
> 
> That raises the question why request_irq() is called on "link up" time,
> while the low level vector allocation takes place during device probing.
> At least that seems to be the common pattern. Can someone tell me why
> this is done this way? Shouldn't we call request_irq() at the same time?
> 
> So, let's hope that all of this was short enough that somebody reads it

I did!

> and precise enough to make the problem clear. The idea behind this series
> is not fixing or applying the series as is. I'm expecting a discussion
> first. So, input welcome!

Not sure I understand the full scope, but it's not because of your
description. And I want to learn more about the RT business. I'll
comment the cpumasks part, and please keep me in CC.

Thanks,
Yury

> Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
> ---
> Florian Bezdeka (3):
>       cpumask: Honor irq_default_affinity in cpumask_local_spread()
>       genirq: Honor existing IRQ affinities when setting affinity hints
>       net: stmmac: Migrate IRQ balancing to cpumask_local_spread()
> 
>  drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 21 +++++++++++++++++----
>  kernel/irq/manage.c                               | 10 ++++++++++
>  lib/cpumask.c                                     | 10 ++++++----
>  3 files changed, 33 insertions(+), 8 deletions(-)
> ---
> base-commit: aa2e13ae8d3cbe2c15ef4f7e971b2de0832794aa
> change-id: 20260810-flo-net-7-2-make-stmmac-default-affinity-aware-a145a4941d8f
> 
> Best regards,
> -- 
> Florian Bezdeka <florian.bezdeka@siemens.com>