[PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq

Marco Crivellari posted 6 patches 4 days, 16 hours ago
drivers/net/ethernet/ibm/ibmvnic.c           | 6 +++---
drivers/net/ethernet/ti/icssg/icssg_prueth.c | 2 +-
drivers/net/ethernet/ti/icssg/icssg_stats.c  | 2 +-
drivers/net/thunderbolt/main.c               | 7 ++++---
drivers/net/usb/pegasus.c                    | 9 +++++----
drivers/net/usb/r8152.c                      | 7 ++++---
6 files changed, 18 insertions(+), 15 deletions(-)
[PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq
Posted by Marco Crivellari 4 days, 16 hours ago
Hello,

Currently the code uses the per-cpu workqueue system_long_wq to schedule
long running works.

Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Another good reason to have this unbound,
is the "queue_delayed_work()" function, used to enqueue the work item.
More details on this will follow in the next section.

Recently, a new unbound workqueue specific for long running work has been
added:

    c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")

~~~ Details about queue_delayed_work ~~~

system_long_wq is a per-cpu workqueue and it is used as a parameter of
queue_delayed_work(). This function schedule an item that it will later
be enqueued (once the timer will fire). __queue_delayed_work() does the job
receiving as "cpu" WORK_CPU_UNBOUND:

    if (housekeeping_enabled(HK_TYPE_TIMER)) {
    //      [....]
    } else {
            if (likely(cpu == WORK_CPU_UNBOUND))
                    add_timer_global(timer);
            else
                    add_timer_on(timer, cpu);
    }

The timer is global, so can fire everywhere, and the work item will be
enqueued where the timer fired.

Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change the
workqueue with the new system_dfl_long_wq, so that the used workqueue is
now unbound and can benefit from scheduler task placement.

Thanks!

---
Changes in v3:
- rebased on v7.2-rc4

- added "net: ti: icssg-prueth: Move long delayed work on system_dfl_long_wq"

- changed also the workqueue in ibmvnic_reset().

Link to v2: https://lore.kernel.org/all/20260706134033.244295-1-marco.crivellari@suse.com/

Changes in v2:
- rebased on v7.2-rc2

- dropped the RFC prefix, kept Ack and review tags

Link to v1: https://lore.kernel.org/all/20260511092846.120141-1-marco.crivellari@suse.com/

Marco Crivellari (6):
  ibmvnic: Move long delayed work on system_dfl_long_wq
  net: ti: icssg-stats: Move long delayed work on system_dfl_long_wq
  net: ti: icssg-prueth: Move long delayed work on system_dfl_long_wq
  net: thunderbolt: Move long delayed work on system_dfl_long_wq
  net: usb: pegasus: Move long delayed work on system_dfl_long_wq
  net: usb: r8152: Move long delayed work on system_dfl_long_wq

 drivers/net/ethernet/ibm/ibmvnic.c           | 6 +++---
 drivers/net/ethernet/ti/icssg/icssg_prueth.c | 2 +-
 drivers/net/ethernet/ti/icssg/icssg_stats.c  | 2 +-
 drivers/net/thunderbolt/main.c               | 7 ++++---
 drivers/net/usb/pegasus.c                    | 9 +++++----
 drivers/net/usb/r8152.c                      | 7 ++++---
 6 files changed, 18 insertions(+), 15 deletions(-)

-- 
2.54.0
Re: [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq
Posted by Jacob Keller 4 days, 3 hours ago
On 7/20/2026 3:08 AM, Marco Crivellari wrote:
> Hello,
> 
> Currently the code uses the per-cpu workqueue system_long_wq to schedule
> long running works.
> 
> Unbound works could benefit from scheduler task placement, to optimize
> performance and power consumption. Another good reason to have this unbound,
> is the "queue_delayed_work()" function, used to enqueue the work item.
> More details on this will follow in the next section.
> 
> Recently, a new unbound workqueue specific for long running work has been
> added:
> 
>     c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
> 
> ~~~ Details about queue_delayed_work ~~~
> 
> system_long_wq is a per-cpu workqueue and it is used as a parameter of
> queue_delayed_work(). This function schedule an item that it will later
> be enqueued (once the timer will fire). __queue_delayed_work() does the job
> receiving as "cpu" WORK_CPU_UNBOUND:
> 
>     if (housekeeping_enabled(HK_TYPE_TIMER)) {
>     //      [....]
>     } else {
>             if (likely(cpu == WORK_CPU_UNBOUND))
>                     add_timer_global(timer);
>             else
>                     add_timer_on(timer, cpu);
>     }
> 
> The timer is global, so can fire everywhere, and the work item will be
> enqueued where the timer fired.
> 
> Since the workqueue work doesn't rely on per-cpu variables, there is no
> obvious reason that justify the use of a per-cpu workqueue. So change the
> workqueue with the new system_dfl_long_wq, so that the used workqueue is
> now unbound and can benefit from scheduler task placement.
> 


Ok. So if I am understanding this correctly, the current code uses
system_long_wq which is per-CPU, but is fired using an unbound timer. As
a result, whichever CPU the timer triggers on will be the one which
selects the work queue. From there, the work item will be enqueued to
that work queue and remain on that work queue until resolving with no
way for scheduler to adjust it?

With the new change, we schedule on the system_dfl_long_wq which *isn't*
per CPU, so the scheduler is free to move the task around and
reschedule. As a result we get better overall behavior with more input
from the scheduler, instead of effective randomness from the timer which
is then forced so that such long running task cannot migrate?

That sounds like a pretty good improvement for the cases where the
queued work doesn't depend on any per-cpu behavior. Nice!

I am not sure I can speak to any of the individual drivers here since I
wouldn't know whether moving that particular work item would be
affected.. so feel free to take this review with a grain of salt :)

Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Re: [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq
Posted by Marco Crivellari 3 days, 17 hours ago
Hi,

On Tue, Jul 21, 2026 at 12:35 AM Jacob Keller <jacob.e.keller@intel.com> wrote:
> [...]
> > system_long_wq is a per-cpu workqueue and it is used as a parameter of
> > queue_delayed_work(). This function schedule an item that it will later
> > be enqueued (once the timer will fire). __queue_delayed_work() does the job
> > receiving as "cpu" WORK_CPU_UNBOUND:
> >
> >     if (housekeeping_enabled(HK_TYPE_TIMER)) {
> >     //      [....]
> >     } else {
> >             if (likely(cpu == WORK_CPU_UNBOUND))
> >                     add_timer_global(timer);
> >             else
> >                     add_timer_on(timer, cpu);
> >     }
> >
> > The timer is global, so can fire everywhere, and the work item will be
> > enqueued where the timer fired.
> >
> > Since the workqueue work doesn't rely on per-cpu variables, there is no
> > obvious reason that justify the use of a per-cpu workqueue. So change the
> > workqueue with the new system_dfl_long_wq, so that the used workqueue is
> > now unbound and can benefit from scheduler task placement.
> >
>
>
> Ok. So if I am understanding this correctly, the current code uses
> system_long_wq which is per-CPU, but is fired using an unbound timer. As
> a result, whichever CPU the timer triggers on will be the one which
> selects the work queue. From there, the work item will be enqueued to
> that work queue and remain on that work queue until resolving with no
> way for scheduler to adjust it?
>
> With the new change, we schedule on the system_dfl_long_wq which *isn't*
> per CPU, so the scheduler is free to move the task around and
> reschedule. As a result we get better overall behavior with more input
> from the scheduler, instead of effective randomness from the timer which
> is then forced so that such long running task cannot migrate?
>
> That sounds like a pretty good improvement for the cases where the
> queued work doesn't depend on any per-cpu behavior. Nice!

Yes, that's pretty much it!

> I am not sure I can speak to any of the individual drivers here since I
> wouldn't know whether moving that particular work item would be
> affected.. so feel free to take this review with a grain of salt :)
>
> Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>

Sure, thank you!

-- 

Marco Crivellari

SUSE Labs