[PATCH] net: iterate online nodes in skb_defer_free_flush()

Kris Pan posted 1 patch 2 weeks ago
There is a newer version of this series
net/core/dev.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
[PATCH] net: iterate online nodes in skb_defer_free_flush()
Posted by Kris Pan 2 weeks ago
skb_attempt_defer_free() only queues skbs on the current CPU's node
(numa_node_id() of a running CPU), which is always online, so the
flush loop never needs to visit nodes that are merely possible.

for_each_node() walks node_possible_map.  On machines where the
possible map is much larger than the online map -- e.g. a POWER10
LPAR with 32 possible but 1 online node -- the flush loop touches 31
cold, always-empty per-node lists on every softirq pass, showing up
as skb_defer_free_flush() and _find_next_bit() overhead.

Use for_each_online_node() to iterate only node_online_map.

Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%.

Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
Reported-by: kernel test robot <oliver.sang@intel.com>
Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com
Signed-off-by: Kris Pan <kris.pan@intel.com>
---
 net/core/dev.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index 290e0f099e6bf..b528b6a986fcf 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -6907,7 +6907,7 @@ static void skb_defer_free_flush(void)
 	struct skb_defer_node *sdn;
 	int node;
 
-	for_each_node(node) {
+	for_each_online_node(node) {
 		sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node;
 
 		if (llist_empty(&sdn->defer_list))
-- 
2.43.0
Re: [PATCH] net: iterate online nodes in skb_defer_free_flush()
Posted by Eric Dumazet 2 weeks ago
On Thu, Sep 10, 2026 at 7:07 PM Kris Pan <kris.pan@intel.com> wrote:
>
> skb_attempt_defer_free() only queues skbs on the current CPU's node
> (numa_node_id() of a running CPU), which is always online, so the
> flush loop never needs to visit nodes that are merely possible.
>
> for_each_node() walks node_possible_map.  On machines where the
> possible map is much larger than the online map -- e.g. a POWER10
> LPAR with 32 possible but 1 online node -- the flush loop touches 31
> cold, always-empty per-node lists on every softirq pass, showing up
> as skb_defer_free_flush() and _find_next_bit() overhead.
>
> Use for_each_online_node() to iterate only node_online_map.
>
> Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
> nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%.
>
> Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
> Reported-by: kernel test robot <oliver.sang@intel.com>
> Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com
> Signed-off-by: Kris Pan <kris.pan@intel.com>
> ---
>  net/core/dev.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/net/core/dev.c b/net/core/dev.c
> index 290e0f099e6bf..b528b6a986fcf 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -6907,7 +6907,7 @@ static void skb_defer_free_flush(void)
>         struct skb_defer_node *sdn;
>         int node;
>
> -       for_each_node(node) {
> +       for_each_online_node(node) {

SGTM, but you probably could have added

Suggested-by: Adrian Tomasov <atomasov@redhat.com>

(Assuming you read
https://lore.kernel.org/oe-lkp/20260910143041.18106-1-atomasov@redhat.com/)

Reviewed-by: Eric Dumazet <edumazet@google.com>

Thanks.
[PATCH v2] net: iterate online nodes in skb_defer_free_flush()
Posted by Kris Pan 2 weeks ago
skb_attempt_defer_free() only queues skbs on the current CPU's node
(numa_node_id() of a running CPU), which is always online, so the
flush loop never needs to visit nodes that are merely possible.

for_each_node() walks node_possible_map.  On machines where the
possible map is much larger than the online map -- e.g. a POWER10
LPAR with 32 possible but 1 online node -- the flush loop touches 31
cold, always-empty per-node lists on every softirq pass, showing up
as skb_defer_free_flush() and _find_next_bit() overhead.

Use for_each_online_node() to iterate only node_online_map.

Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%.

Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
Reported-by: kernel test robot <oliver.sang@intel.com>
Suggested-by: Adrian Tomasov <atomasov@redhat.com>
Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com
Signed-off-by: Kris Pan <kris.pan@intel.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
---
 net/core/dev.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index 290e0f099e6bf..b528b6a986fcf 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -6907,7 +6907,7 @@ static void skb_defer_free_flush(void)
 	struct skb_defer_node *sdn;
 	int node;
 
-	for_each_node(node) {
+	for_each_online_node(node) {
 		sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node;
 
 		if (llist_empty(&sdn->defer_list))
-- 
2.43.0
Re: [PATCH v2] net: iterate online nodes in skb_defer_free_flush()
Posted by Jakub Kicinski 1 week, 2 days ago
On Fri, 11 Sep 2026 10:38:04 +0800 Kris Pan wrote:
> skb_attempt_defer_free() only queues skbs on the current CPU's node
> (numa_node_id() of a running CPU), which is always online, so the
> flush loop never needs to visit nodes that are merely possible.
> 
> for_each_node() walks node_possible_map.  On machines where the
> possible map is much larger than the online map -- e.g. a POWER10
> LPAR with 32 possible but 1 online node -- the flush loop touches 31
> cold, always-empty per-node lists on every softirq pass, showing up
> as skb_defer_free_flush() and _find_next_bit() overhead.
> 
> Use for_each_online_node() to iterate only node_online_map.
> 
> Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
> nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%.

Clashiko confirms this is racy:

https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911023804.3503989-1-kris.pan@intel.com
-- 
pw-bot: cr
Re: [PATCH v2] net: iterate online nodes in skb_defer_free_flush()
Posted by Kris Pan 1 week, 2 days ago
Right, the node-offline race is real.  I'll fix it in a v3 so the flush
doesn't depend on node_online_map, e.g. a per-CPU mask of pending nodes
set by the producer and drained by the flush.

Thanks.
Re: [PATCH v2] net: iterate online nodes in skb_defer_free_flush()
Posted by Eric Dumazet 1 week, 2 days ago
On Tue, Sep 15, 2026 at 4:56 PM Kris Pan <kris.pan@intel.com> wrote:
>
> Right, the node-offline race is real.  I'll fix it in a v3 so the flush
> doesn't depend on node_online_map, e.g. a per-CPU mask of pending nodes
> set by the producer and drained by the flush.

Certainly not.

We should not add a per-CPU active-node bitmask updated by
skb_attempt_defer_free(),
as writing to a shared bitmask from remote CPUs would re-introduce the
cross-NUMA
cache line bouncing that commit 5628f3fe3b16 eliminated.

Instead, keep for_each_online_node(node) in the skb_defer_free_flush()
fast path,
and handle cleanup in the cold dev_cpu_dead(unsigned int oldcpu) hotplug path.