net/core/dev.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-)
skb_attempt_defer_free() only queues skbs on the current CPU's node
(numa_node_id() of a running CPU), which is always online, so the
flush loop never needs to visit nodes that are merely possible.
for_each_node() walks node_possible_map. On machines where the
possible map is much larger than the online map -- e.g. a POWER10
LPAR with 32 possible but 1 online node -- the flush loop touches 31
cold, always-empty per-node lists on every softirq pass, showing up
as skb_defer_free_flush() and _find_next_bit() overhead.
Use for_each_online_node() to iterate only node_online_map.
Loopback UDP throughput in a QEMU guest with 32 possible / 1 online
nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%.
Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
Reported-by: kernel test robot <oliver.sang@intel.com>
Suggested-by: Adrian Tomasov <atomasov@redhat.com>
Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com
Signed-off-by: Kris Pan <kris.pan@intel.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
---
net/core/dev.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index 290e0f099e6bf..b528b6a986fcf 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -6907,7 +6907,7 @@ static void skb_defer_free_flush(void)
struct skb_defer_node *sdn;
int node;
- for_each_node(node) {
+ for_each_online_node(node) {
sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node;
if (llist_empty(&sdn->defer_list))
--
2.43.0
On Fri, 11 Sep 2026 10:38:04 +0800 Kris Pan wrote: > skb_attempt_defer_free() only queues skbs on the current CPU's node > (numa_node_id() of a running CPU), which is always online, so the > flush loop never needs to visit nodes that are merely possible. > > for_each_node() walks node_possible_map. On machines where the > possible map is much larger than the online map -- e.g. a POWER10 > LPAR with 32 possible but 1 online node -- the flush loop touches 31 > cold, always-empty per-node lists on every softirq pass, showing up > as skb_defer_free_flush() and _find_next_bit() overhead. > > Use for_each_online_node() to iterate only node_online_map. > > Loopback UDP throughput in a QEMU guest with 32 possible / 1 online > nodes (bench_udp, 8 senders, 6 interleaved runs) improves by ~5%. Clashiko confirms this is racy: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911023804.3503989-1-kris.pan@intel.com -- pw-bot: cr
Right, the node-offline race is real. I'll fix it in a v3 so the flush doesn't depend on node_online_map, e.g. a per-CPU mask of pending nodes set by the producer and drained by the flush. Thanks.
On Tue, Sep 15, 2026 at 4:56 PM Kris Pan <kris.pan@intel.com> wrote: > > Right, the node-offline race is real. I'll fix it in a v3 so the flush > doesn't depend on node_online_map, e.g. a per-CPU mask of pending nodes > set by the producer and drained by the flush. Certainly not. We should not add a per-CPU active-node bitmask updated by skb_attempt_defer_free(), as writing to a shared bitmask from remote CPUs would re-introduce the cross-NUMA cache line bouncing that commit 5628f3fe3b16 eliminated. Instead, keep for_each_online_node(node) in the skb_defer_free_flush() fast path, and handle cleanup in the cold dev_cpu_dead(unsigned int oldcpu) hotplug path.
© 2016 - 2026 Red Hat, Inc.