[PATCH net-next v3 0/3] net: hash uncached route lists by device

Chris J Arges posted 3 patches 1 week ago
net/ipv4/route.c                              |  36 +++++++--
net/ipv6/route.c                              | 101 ++++++++++++++++++--------
tools/testing/selftests/net/vrf-xfrm-tests.sh |  35 +++++++++
3 files changed, 133 insertions(+), 39 deletions(-)
[PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Chris J Arges 1 week ago
We have observed hung tasks blocked on rtnl_mutex while network namespaces
were being removed. The namespaces contained many network devices, and the
host had accumulated a large population of entries on the global per-CPU
uncached route lists. A perf profile collected during one incident
attributed most of the cleanup worker's samples to rt_flush_dev():

```
99.92% kworker/u384:3-  worker_thread
  `-88.71% process_one_work
      `-81.02% cleanup_net
          `-81.00% unregister_netdevice_many_notify
              `-79.42% notifier_call_chain
                  `-78.05% fib_netdev_event
                      `-77.92% rt_flush_dev
```

For each device, rt_flush_dev() visits every possible CPU and scans the
global uncached route population while its caller holds rtnl_mutex. If N is
the number of devices, C the number of possible CPUs, and R the number of
uncached routes, the cost is O(N * (C + R)).

During namespace cleanup, other processes that issue RTNETLINK operations
requiring the RTNL lock can stall until cleanup releases the lock.

A minimal reproducer is available here:
https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm

This series replaces each per-CPU uncached route list with a hash table
using the network device as its key. Each table uses 64 buckets.
IPv6 routes need additional handling because dst.dev and
rt6i_idev->dev can refer to different devices. Routes are keyed by
rt6i_idev->dev when available. Device teardown scans one bucket for
ordinary devices and all buckets for loopback and L3 master devices.

We measured user-visible RTNL latency on a 192-CPU x86-64 host. The test
added approximately 80,000 uncached routes across 256 devices simulating a
distribution we saw in production with 6 devices having 4k to 20k routes,
and all others holding ~100 routes. The devices being removed owned none
of these routes.

During asynchronous namespace cleanup, the test repeatedly sends an
idempotent RTM_NEWLINK request that requires RTNL. It then records the
worst request-to-acknowledgment latency in each observation window.

Results from this test show the median latency for the RTM_NEWLINK request
to complete after waiting for unregsiter batch show between 68-75%
reduction in latency when using the patch.

We also measured end-to-end route insertion cost separately on the same
machine. The test inserted 100,000 routes per round for 30 rounds after
three warmups, while pinned to one CPU. Median insertion cost was
2,069 ns/op without hashing and 2,066 ns/op with hashing. This test found
no measurable insertion regression.

The hash approach adds no per-route fields. On x86-64, the tables add
approximately 3 KiB per possible CPU with 64 buckets.

Patch 1 hashes IPv4 uncached routes by network device.
Patch 2 applies the hashing design to IPv6 and handles routes whose device
references differ.
Patch 3 adds a selftest for the IPv6 case.

Signed-off-by: Chris J Arges <carges@cloudflare.com>
---
Changes in v3:
- Remove IPv4 and IPv6 Kconfig options; use fixed 64-bucket tables.
- Use rcu_assign_pointer in rt6_uncached_list_flush
- Key IPv6 routes by rt6i_idev and scan all buckets for loopback/VRF
- RCT all the things
- Link to v2: https://patch.msgid.link/20260914-hash-bucket-route-lists-v2-0-29f6297d8a5a@cloudflare.com

Changes in v2:
- Add IPv4 and IPv6 Kconfig options for the uncached-route hash size.
- Keep 64 buckets as the default and document the per-CPU memory tradeoff.
- Link to v1: https://patch.msgid.link/20260826-hash-bucket-route-lists-v1-0-fa9b9f30eb74@cloudflare.com

To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: David Ahern <dsahern@kernel.org>
To: Ido Schimmel <idosch@nvidia.com>
To: Shuah Khan <shuah@kernel.org>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-kselftest@vger.kernel.org

---
Chris J Arges (3):
      ipv4: hash uncached routes by device
      ipv6: hash uncached routes by device
      selftests: net: cover IPv6 uncached route device mismatch

 net/ipv4/route.c                              |  36 +++++++--
 net/ipv6/route.c                              | 101 ++++++++++++++++++--------
 tools/testing/selftests/net/vrf-xfrm-tests.sh |  35 +++++++++
 3 files changed, 133 insertions(+), 39 deletions(-)
---
base-commit: 26ee8cd69d46a14b37ba5e512084fe80d730127a
change-id: 20260820-hash-bucket-route-lists-b8cc27ccd53c

Best regards,
--  
Chris J Arges <carges@cloudflare.com>
Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Kuniyuki Iwashima 1 week ago
From: Chris J Arges <carges@cloudflare.com>
Date: Thu, 17 Sep 2026 14:38:21 -0500
> We have observed hung tasks blocked on rtnl_mutex while network namespaces
> were being removed. The namespaces contained many network devices, and the
> host had accumulated a large population of entries on the global per-CPU
> uncached route lists. A perf profile collected during one incident
> attributed most of the cleanup worker's samples to rt_flush_dev():
> 
> ```
> 99.92% kworker/u384:3-  worker_thread
>   `-88.71% process_one_work
>       `-81.02% cleanup_net
>           `-81.00% unregister_netdevice_many_notify
>               `-79.42% notifier_call_chain
>                   `-78.05% fib_netdev_event
>                       `-77.92% rt_flush_dev
> ```
> 
> For each device, rt_flush_dev() visits every possible CPU and scans the
> global uncached route population while its caller holds rtnl_mutex. If N is
> the number of devices, C the number of possible CPUs, and R the number of
> uncached routes, the cost is O(N * (C + R)).
> 
> During namespace cleanup, other processes that issue RTNETLINK operations
> requiring the RTNL lock can stall until cleanup releases the lock.
> 
> A minimal reproducer is available here:
> https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> 
> This series replaces each per-CPU uncached route list with a hash table
> using the network device as its key. Each table uses 64 buckets.

This sounds a bit overkill.  Also, this series still leaves
O(N * C) loops.

Given unregistering a single device is less common than
destroying netns, I think the right approach should be to
make the route flush once in cleanup_net() + outside RTNL.

Could you try this change ? (only compile-tested)

---8<---
diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index 46b4c67e2966..d8ce7dc0fbdc 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -489,6 +489,7 @@ struct pernet_operations {
 	 */
 	int (*init)(struct net *net);
 	void (*pre_exit)(struct net *net);
+	void (*pre_exit_batch)(struct list_head *net_exit_list);
 	void (*exit)(struct net *net);
 	void (*exit_batch)(struct list_head *net_exit_list);
 	/* Following method is called with RTNL held. */
diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
index da5f881fbd3b..7fc9bf45f3b6 100644
--- a/net/core/net_namespace.c
+++ b/net/core/net_namespace.c
@@ -160,6 +160,9 @@ static void ops_pre_exit_list(const struct pernet_operations *ops,
 		list_for_each_entry(net, net_exit_list, exit_list)
 			ops->pre_exit(net);
 	}
+
+	if (ops->pre_exit_batch)
+		ops->pre_exit_batch(net_exit_list);
 }
 
 static void ops_exit_rtnl_list(const struct list_head *ops_list,
diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
index 8a3dc04e8cac..b8d76b6279e1 100644
--- a/net/ipv4/fib_frontend.c
+++ b/net/ipv4/fib_frontend.c
@@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
 	nl_fib_lookup_exit(net);
 }
 
+static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+	rt_flush_dev(NULL);
+}
+
 static void __net_exit fib_net_exit_rtnl(struct net *net,
 					 struct list_head *dev_kill_list)
 {
@@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
 static struct pernet_operations fib_net_ops = {
 	.init = fib_net_init,
 	.pre_exit = fib_net_pre_exit,
+	.pre_exit_batch = fib_net_pre_exit_batch,
 	.exit_rtnl = fib_net_exit_rtnl,
 	.exit = fib_net_exit,
 };
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d7da2f1acbb5..d35b66b33bbc 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1554,14 +1554,28 @@ struct uncached_list {
 
 static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
 
+static void rt_replace_uncached_list(struct rtable *rt)
+{
+	struct net_device *dev = dst_dev(&rt->dst);
+
+	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+	netdev_ref_replace(dev, blackhole_netdev,
+			   &rt->dst.dev_tracker, GFP_ATOMIC);
+}
+
 void rt_add_uncached_list(struct rtable *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
 
-	rt->dst.rt_uncached_list = ul;
-
 	spin_lock_bh(&ul->lock);
-	list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
+		rt_replace_uncached_list(rt);
+	} else {
+		rt->dst.rt_uncached_list = ul;
+		list_add_tail(&rt->dst.rt_uncached, &ul->head);
+	}
+
 	spin_unlock_bh(&ul->lock);
 }
 
@@ -1587,6 +1601,9 @@ void rt_flush_dev(struct net_device *dev)
 	struct rtable *rt, *safe;
 	int cpu;
 
+	if (dev && !check_net(dev_net(dev)))
+		return;
+
 	for_each_possible_cpu(cpu) {
 		struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
 
@@ -1595,11 +1612,11 @@ void rt_flush_dev(struct net_device *dev)
 
 		spin_lock_bh(&ul->lock);
 		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
-			if (rt->dst.dev != dev)
+			if (rt->dst.dev != dev &&
+			    (dev || check_net(dev_net(rt->dst.dev))))
 				continue;
-			rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
-			netdev_ref_replace(dev, blackhole_netdev,
-					   &rt->dst.dev_tracker, GFP_ATOMIC);
+
+			rt_replace_uncached_list(rt);
 			list_del_init(&rt->dst.rt_uncached);
 		}
 		spin_unlock_bh(&ul->lock);
diff --git a/net/ipv6/route.c b/net/ipv6/route.c
index 7535b09068a0..28233197e1e1 100644
--- a/net/ipv6/route.c
+++ b/net/ipv6/route.c
@@ -135,14 +135,35 @@ struct uncached_list {
 
 static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt6_uncached_list);
 
+static void rt6_uncached_list_replace(struct rt6_info *rt)
+{
+	struct net_device *dev = dst_dev(&rt->dst);
+	struct inet6_dev *rt_idev = rt->rt6i_idev;
+
+	if (rt_idev) {
+		rt->rt6i_idev = in6_dev_get(blackhole_netdev);
+		in6_dev_put(rt_idev);
+	}
+
+	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+	netdev_ref_replace(dev, blackhole_netdev,
+			   &rt->dst.dev_tracker,
+			   GFP_ATOMIC);
+}
+
 void rt6_uncached_list_add(struct rt6_info *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
 
-	rt->dst.rt_uncached_list = ul;
-
 	spin_lock_bh(&ul->lock);
-	list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
+		rt6_uncached_list_replace(rt);
+	} else {
+		rt->dst.rt_uncached_list = ul;
+		list_add_tail(&rt->dst.rt_uncached, &ul->head);
+	}
+
 	spin_unlock_bh(&ul->lock);
 }
 
@@ -161,6 +182,9 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
 {
 	int cpu;
 
+	if (dev && !check_net(dev_net(dev)))
+		return;
+
 	for_each_possible_cpu(cpu) {
 		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);
 		struct rt6_info *rt, *safe;
@@ -172,23 +196,17 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
 		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
 			struct inet6_dev *rt_idev = rt->rt6i_idev;
 			struct net_device *rt_dev = rt->dst.dev;
-			bool handled = false;
 
-			if (rt_idev && rt_idev->dev == dev) {
-				rt->rt6i_idev = in6_dev_get(blackhole_netdev);
-				in6_dev_put(rt_idev);
-				handled = true;
+			if (dev) {
+				if (rt_dev != dev &&
+				    (!rt_idev || rt_idev->dev != dev))
+					continue;
+			} else if (check_net(dev_net(rt_dev))) {
+				continue;
 			}
 
-			if (rt_dev == dev) {
-				rt->dst.dev = blackhole_netdev;
-				netdev_ref_replace(rt_dev, blackhole_netdev,
-						   &rt->dst.dev_tracker,
-						   GFP_ATOMIC);
-				handled = true;
-			}
-			if (handled)
-				list_del_init(&rt->dst.rt_uncached);
+			rt6_uncached_list_replace(rt);
+			list_del_init(&rt->dst.rt_uncached);
 		}
 		spin_unlock_bh(&ul->lock);
 	}
@@ -6795,6 +6813,11 @@ static int __net_init ip6_route_net_init(struct net *net)
 	goto out;
 }
 
+static void __net_exit ip6_route_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+	rt6_uncached_list_flush_dev(NULL);
+}
+
 static void __net_exit ip6_route_net_exit(struct net *net)
 {
 	kfree(net->ipv6.fib6_null_entry);
@@ -6833,6 +6856,7 @@ static void __net_exit ip6_route_net_exit_late(struct net *net)
 
 static struct pernet_operations ip6_route_net_ops = {
 	.init = ip6_route_net_init,
+	.pre_exit_batch = ip6_route_net_pre_exit_batch,
 	.exit = ip6_route_net_exit,
 };
 
---8<---
Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Chris Arges 1 week ago
On 2026-09-17 22:10:22, Kuniyuki Iwashima wrote:
> From: Chris J Arges <carges@cloudflare.com>
> Date: Thu, 17 Sep 2026 14:38:21 -0500
> > We have observed hung tasks blocked on rtnl_mutex while network namespaces
> > were being removed. The namespaces contained many network devices, and the
> > host had accumulated a large population of entries on the global per-CPU
> > uncached route lists. A perf profile collected during one incident
> > attributed most of the cleanup worker's samples to rt_flush_dev():
> > 
> > ```
> > 99.92% kworker/u384:3-  worker_thread
> >   `-88.71% process_one_work
> >       `-81.02% cleanup_net
> >           `-81.00% unregister_netdevice_many_notify
> >               `-79.42% notifier_call_chain
> >                   `-78.05% fib_netdev_event
> >                       `-77.92% rt_flush_dev
> > ```
> > 
> > For each device, rt_flush_dev() visits every possible CPU and scans the
> > global uncached route population while its caller holds rtnl_mutex. If N is
> > the number of devices, C the number of possible CPUs, and R the number of
> > uncached routes, the cost is O(N * (C + R)).
> > 
> > During namespace cleanup, other processes that issue RTNETLINK operations
> > requiring the RTNL lock can stall until cleanup releases the lock.
> > 
> > A minimal reproducer is available here:
> > https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> > 
> > This series replaces each per-CPU uncached route list with a hash table
> > using the network device as its key. Each table uses 64 buckets.
> 
> This sounds a bit overkill.  Also, this series still leaves
> O(N * C) loops.
> 
> Given unregistering a single device is less common than
> destroying netns, I think the right approach should be to
> make the route flush once in cleanup_net() + outside RTNL.
> 
> Could you try this change ? (only compile-tested)
> 
Excellent, I'll test this and report back.
Thanks,
--chris


> ---8<---
> diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
> index 46b4c67e2966..d8ce7dc0fbdc 100644
> --- a/include/net/net_namespace.h
> +++ b/include/net/net_namespace.h
> @@ -489,6 +489,7 @@ struct pernet_operations {
>  	 */
>  	int (*init)(struct net *net);
>  	void (*pre_exit)(struct net *net);
> +	void (*pre_exit_batch)(struct list_head *net_exit_list);
>  	void (*exit)(struct net *net);
>  	void (*exit_batch)(struct list_head *net_exit_list);
>  	/* Following method is called with RTNL held. */
> diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
> index da5f881fbd3b..7fc9bf45f3b6 100644
> --- a/net/core/net_namespace.c
> +++ b/net/core/net_namespace.c
> @@ -160,6 +160,9 @@ static void ops_pre_exit_list(const struct pernet_operations *ops,
>  		list_for_each_entry(net, net_exit_list, exit_list)
>  			ops->pre_exit(net);
>  	}
> +
> +	if (ops->pre_exit_batch)
> +		ops->pre_exit_batch(net_exit_list);
>  }
>  
>  static void ops_exit_rtnl_list(const struct list_head *ops_list,
> diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
> index 8a3dc04e8cac..b8d76b6279e1 100644
> --- a/net/ipv4/fib_frontend.c
> +++ b/net/ipv4/fib_frontend.c
> @@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
>  	nl_fib_lookup_exit(net);
>  }
>  
> +static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
> +{
> +	rt_flush_dev(NULL);
> +}
> +
>  static void __net_exit fib_net_exit_rtnl(struct net *net,
>  					 struct list_head *dev_kill_list)
>  {
> @@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
>  static struct pernet_operations fib_net_ops = {
>  	.init = fib_net_init,
>  	.pre_exit = fib_net_pre_exit,
> +	.pre_exit_batch = fib_net_pre_exit_batch,
>  	.exit_rtnl = fib_net_exit_rtnl,
>  	.exit = fib_net_exit,
>  };
> diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> index d7da2f1acbb5..d35b66b33bbc 100644
> --- a/net/ipv4/route.c
> +++ b/net/ipv4/route.c
> @@ -1554,14 +1554,28 @@ struct uncached_list {
>  
>  static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
>  
> +static void rt_replace_uncached_list(struct rtable *rt)
> +{
> +	struct net_device *dev = dst_dev(&rt->dst);
> +
> +	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
> +	netdev_ref_replace(dev, blackhole_netdev,
> +			   &rt->dst.dev_tracker, GFP_ATOMIC);
> +}
> +
>  void rt_add_uncached_list(struct rtable *rt)
>  {
>  	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
>  
> -	rt->dst.rt_uncached_list = ul;
> -
>  	spin_lock_bh(&ul->lock);
> -	list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +
> +	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> +		rt_replace_uncached_list(rt);
> +	} else {
> +		rt->dst.rt_uncached_list = ul;
> +		list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +	}
> +
>  	spin_unlock_bh(&ul->lock);
>  }
>  
> @@ -1587,6 +1601,9 @@ void rt_flush_dev(struct net_device *dev)
>  	struct rtable *rt, *safe;
>  	int cpu;
>  
> +	if (dev && !check_net(dev_net(dev)))
> +		return;
> +
>  	for_each_possible_cpu(cpu) {
>  		struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
>  
> @@ -1595,11 +1612,11 @@ void rt_flush_dev(struct net_device *dev)
>  
>  		spin_lock_bh(&ul->lock);
>  		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
> -			if (rt->dst.dev != dev)
> +			if (rt->dst.dev != dev &&
> +			    (dev || check_net(dev_net(rt->dst.dev))))
>  				continue;
> -			rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
> -			netdev_ref_replace(dev, blackhole_netdev,
> -					   &rt->dst.dev_tracker, GFP_ATOMIC);
> +
> +			rt_replace_uncached_list(rt);
>  			list_del_init(&rt->dst.rt_uncached);
>  		}
>  		spin_unlock_bh(&ul->lock);
> diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> index 7535b09068a0..28233197e1e1 100644
> --- a/net/ipv6/route.c
> +++ b/net/ipv6/route.c
> @@ -135,14 +135,35 @@ struct uncached_list {
>  
>  static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt6_uncached_list);
>  
> +static void rt6_uncached_list_replace(struct rt6_info *rt)
> +{
> +	struct net_device *dev = dst_dev(&rt->dst);
> +	struct inet6_dev *rt_idev = rt->rt6i_idev;
> +
> +	if (rt_idev) {
> +		rt->rt6i_idev = in6_dev_get(blackhole_netdev);
> +		in6_dev_put(rt_idev);
> +	}
> +
> +	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
> +	netdev_ref_replace(dev, blackhole_netdev,
> +			   &rt->dst.dev_tracker,
> +			   GFP_ATOMIC);
> +}
> +
>  void rt6_uncached_list_add(struct rt6_info *rt)
>  {
>  	struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
>  
> -	rt->dst.rt_uncached_list = ul;
> -
>  	spin_lock_bh(&ul->lock);
> -	list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +
> +	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> +		rt6_uncached_list_replace(rt);
> +	} else {
> +		rt->dst.rt_uncached_list = ul;
> +		list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +	}
> +
>  	spin_unlock_bh(&ul->lock);
>  }
>  
> @@ -161,6 +182,9 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
>  {
>  	int cpu;
>  
> +	if (dev && !check_net(dev_net(dev)))
> +		return;
> +
>  	for_each_possible_cpu(cpu) {
>  		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);
>  		struct rt6_info *rt, *safe;
> @@ -172,23 +196,17 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
>  		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
>  			struct inet6_dev *rt_idev = rt->rt6i_idev;
>  			struct net_device *rt_dev = rt->dst.dev;
> -			bool handled = false;
>  
> -			if (rt_idev && rt_idev->dev == dev) {
> -				rt->rt6i_idev = in6_dev_get(blackhole_netdev);
> -				in6_dev_put(rt_idev);
> -				handled = true;
> +			if (dev) {
> +				if (rt_dev != dev &&
> +				    (!rt_idev || rt_idev->dev != dev))
> +					continue;
> +			} else if (check_net(dev_net(rt_dev))) {
> +				continue;
>  			}
>  
> -			if (rt_dev == dev) {
> -				rt->dst.dev = blackhole_netdev;
> -				netdev_ref_replace(rt_dev, blackhole_netdev,
> -						   &rt->dst.dev_tracker,
> -						   GFP_ATOMIC);
> -				handled = true;
> -			}
> -			if (handled)
> -				list_del_init(&rt->dst.rt_uncached);
> +			rt6_uncached_list_replace(rt);
> +			list_del_init(&rt->dst.rt_uncached);
>  		}
>  		spin_unlock_bh(&ul->lock);
>  	}
> @@ -6795,6 +6813,11 @@ static int __net_init ip6_route_net_init(struct net *net)
>  	goto out;
>  }
>  
> +static void __net_exit ip6_route_net_pre_exit_batch(struct list_head *net_exit_list)
> +{
> +	rt6_uncached_list_flush_dev(NULL);
> +}
> +
>  static void __net_exit ip6_route_net_exit(struct net *net)
>  {
>  	kfree(net->ipv6.fib6_null_entry);
> @@ -6833,6 +6856,7 @@ static void __net_exit ip6_route_net_exit_late(struct net *net)
>  
>  static struct pernet_operations ip6_route_net_ops = {
>  	.init = ip6_route_net_init,
> +	.pre_exit_batch = ip6_route_net_pre_exit_batch,
>  	.exit = ip6_route_net_exit,
>  };
>  
> ---8<---
Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Kuniyuki Iwashima 6 days, 20 hours ago
From: Chris Arges <carges@cloudflare.com>
Date: Thu, 17 Sep 2026 19:38:40 -0500
> On 2026-09-17 22:10:22, Kuniyuki Iwashima wrote:
> > From: Chris J Arges <carges@cloudflare.com>
> > Date: Thu, 17 Sep 2026 14:38:21 -0500
> > > We have observed hung tasks blocked on rtnl_mutex while network namespaces
> > > were being removed. The namespaces contained many network devices, and the
> > > host had accumulated a large population of entries on the global per-CPU
> > > uncached route lists. A perf profile collected during one incident
> > > attributed most of the cleanup worker's samples to rt_flush_dev():
> > > 
> > > ```
> > > 99.92% kworker/u384:3-  worker_thread
> > >   `-88.71% process_one_work
> > >       `-81.02% cleanup_net
> > >           `-81.00% unregister_netdevice_many_notify
> > >               `-79.42% notifier_call_chain
> > >                   `-78.05% fib_netdev_event
> > >                       `-77.92% rt_flush_dev
> > > ```
> > > 
> > > For each device, rt_flush_dev() visits every possible CPU and scans the
> > > global uncached route population while its caller holds rtnl_mutex. If N is
> > > the number of devices, C the number of possible CPUs, and R the number of
> > > uncached routes, the cost is O(N * (C + R)).
> > > 
> > > During namespace cleanup, other processes that issue RTNETLINK operations
> > > requiring the RTNL lock can stall until cleanup releases the lock.
> > > 
> > > A minimal reproducer is available here:
> > > https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> > > 
> > > This series replaces each per-CPU uncached route list with a hash table
> > > using the network device as its key. Each table uses 64 buckets.
> > 
> > This sounds a bit overkill.  Also, this series still leaves
> > O(N * C) loops.
> > 
> > Given unregistering a single device is less common than
> > destroying netns, I think the right approach should be to
> > make the route flush once in cleanup_net() + outside RTNL.
> > 
> > Could you try this change ? (only compile-tested)
> > 
> Excellent, I'll test this and report back.

I found a pre-existing issue, which affects the previous
diff, so on top of it, please apply this patch

  https://lore.kernel.org/netdev/20260918041439.2575935-1-kuniyu@google.com/T/#u

and this diff :

---8<---
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d35b66b33bbc..c12e20e07749 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1567,12 +1567,13 @@ void rt_add_uncached_list(struct rtable *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
 
+	rt->dst.rt_uncached_list = ul;
+
 	spin_lock_bh(&ul->lock);
 
 	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
 		rt_replace_uncached_list(rt);
 	} else {
-		rt->dst.rt_uncached_list = ul;
 		list_add_tail(&rt->dst.rt_uncached, &ul->head);
 	}
 
diff --git a/net/ipv6/route.c b/net/ipv6/route.c
index 6cffe8440b44..f22793abbbb8 100644
--- a/net/ipv6/route.c
+++ b/net/ipv6/route.c
@@ -155,12 +155,13 @@ void rt6_uncached_list_add(struct rt6_info *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
 
+	rt->dst.rt_uncached_list = ul;
+
 	spin_lock_bh(&ul->lock);
 
 	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
 		rt6_uncached_list_replace(rt);
 	} else {
-		rt->dst.rt_uncached_list = ul;
 		list_add_tail(&rt->dst.rt_uncached, &ul->head);
 	}
 
---8<---

Thanks !
Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Chris Arges 6 days, 6 hours ago
On 2026-09-18 04:19:34, Kuniyuki Iwashima wrote:
> From: Chris Arges <carges@cloudflare.com>
> Date: Thu, 17 Sep 2026 19:38:40 -0500
> > On 2026-09-17 22:10:22, Kuniyuki Iwashima wrote:
> > > From: Chris J Arges <carges@cloudflare.com>
> > > Date: Thu, 17 Sep 2026 14:38:21 -0500
> > > > We have observed hung tasks blocked on rtnl_mutex while network namespaces
> > > > were being removed. The namespaces contained many network devices, and the
> > > > host had accumulated a large population of entries on the global per-CPU
> > > > uncached route lists. A perf profile collected during one incident
> > > > attributed most of the cleanup worker's samples to rt_flush_dev():
> > > > 
> > > > ```
> > > > 99.92% kworker/u384:3-  worker_thread
> > > >   `-88.71% process_one_work
> > > >       `-81.02% cleanup_net
> > > >           `-81.00% unregister_netdevice_many_notify
> > > >               `-79.42% notifier_call_chain
> > > >                   `-78.05% fib_netdev_event
> > > >                       `-77.92% rt_flush_dev
> > > > ```
> > > > 
> > > > For each device, rt_flush_dev() visits every possible CPU and scans the
> > > > global uncached route population while its caller holds rtnl_mutex. If N is
> > > > the number of devices, C the number of possible CPUs, and R the number of
> > > > uncached routes, the cost is O(N * (C + R)).
> > > > 
> > > > During namespace cleanup, other processes that issue RTNETLINK operations
> > > > requiring the RTNL lock can stall until cleanup releases the lock.
> > > > 
> > > > A minimal reproducer is available here:
> > > > https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> > > > 
> > > > This series replaces each per-CPU uncached route list with a hash table
> > > > using the network device as its key. Each table uses 64 buckets.
> > > 
> > > This sounds a bit overkill.  Also, this series still leaves
> > > O(N * C) loops.
> > > 
> > > Given unregistering a single device is less common than
> > > destroying netns, I think the right approach should be to
> > > make the route flush once in cleanup_net() + outside RTNL.
> > > 
> > > Could you try this change ? (only compile-tested)
> > > 
> > Excellent, I'll test this and report back.
> 
> I found a pre-existing issue, which affects the previous
> diff, so on top of it, please apply this patch
> 
>   https://lore.kernel.org/netdev/20260918041439.2575935-1-kuniyu@google.com/T/#u
> 
> and this diff :
> 
> ---8<---
> diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> index d35b66b33bbc..c12e20e07749 100644
> --- a/net/ipv4/route.c
> +++ b/net/ipv4/route.c
> @@ -1567,12 +1567,13 @@ void rt_add_uncached_list(struct rtable *rt)
>  {
>  	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
>  
> +	rt->dst.rt_uncached_list = ul;
> +
>  	spin_lock_bh(&ul->lock);
>  
>  	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
>  		rt_replace_uncached_list(rt);
>  	} else {
> -		rt->dst.rt_uncached_list = ul;
>  		list_add_tail(&rt->dst.rt_uncached, &ul->head);
>  	}
>  
> diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> index 6cffe8440b44..f22793abbbb8 100644
> --- a/net/ipv6/route.c
> +++ b/net/ipv6/route.c
> @@ -155,12 +155,13 @@ void rt6_uncached_list_add(struct rt6_info *rt)
>  {
>  	struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
>  
> +	rt->dst.rt_uncached_list = ul;
> +
>  	spin_lock_bh(&ul->lock);
>  
>  	if (!check_net(dst_dev_net_rcu(&rt->dst))) {
>  		rt6_uncached_list_replace(rt);
>  	} else {
> -		rt->dst.rt_uncached_list = ul;
>  		list_add_tail(&rt->dst.rt_uncached, &ul->head);
>  	}
>  
> ---8<---

Kuniyuki,

I was able to test this diff, the previous diff you sent plus the fixup
mentioned above. I was able to confirm even greater reduction in contention
as measured by how much latency an unrelated process takes when waiting for
cleanup_net to complete. This makes sense since we don't even need to hold the
lock when processing those routing entries with your patch.

Some rough average latency numbers with 36 devices, 160k routes, 8 vCPUs:
- main: 137ms
- my hashing proposal: 13ms
- your patchset: 1.8ms

I'd be happy to retest any proposed patches.

Thanks,
--chris
Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device
Posted by Kuniyuki Iwashima 6 days, 5 hours ago
On Fri, Sep 18, 2026 at 11:24 AM Chris Arges <carges@cloudflare.com> wrote:
>
> On 2026-09-18 04:19:34, Kuniyuki Iwashima wrote:
> > From: Chris Arges <carges@cloudflare.com>
> > Date: Thu, 17 Sep 2026 19:38:40 -0500
> > > On 2026-09-17 22:10:22, Kuniyuki Iwashima wrote:
> > > > From: Chris J Arges <carges@cloudflare.com>
> > > > Date: Thu, 17 Sep 2026 14:38:21 -0500
> > > > > We have observed hung tasks blocked on rtnl_mutex while network namespaces
> > > > > were being removed. The namespaces contained many network devices, and the
> > > > > host had accumulated a large population of entries on the global per-CPU
> > > > > uncached route lists. A perf profile collected during one incident
> > > > > attributed most of the cleanup worker's samples to rt_flush_dev():
> > > > >
> > > > > ```
> > > > > 99.92% kworker/u384:3-  worker_thread
> > > > >   `-88.71% process_one_work
> > > > >       `-81.02% cleanup_net
> > > > >           `-81.00% unregister_netdevice_many_notify
> > > > >               `-79.42% notifier_call_chain
> > > > >                   `-78.05% fib_netdev_event
> > > > >                       `-77.92% rt_flush_dev
> > > > > ```
> > > > >
> > > > > For each device, rt_flush_dev() visits every possible CPU and scans the
> > > > > global uncached route population while its caller holds rtnl_mutex. If N is
> > > > > the number of devices, C the number of possible CPUs, and R the number of
> > > > > uncached routes, the cost is O(N * (C + R)).
> > > > >
> > > > > During namespace cleanup, other processes that issue RTNETLINK operations
> > > > > requiring the RTNL lock can stall until cleanup releases the lock.
> > > > >
> > > > > A minimal reproducer is available here:
> > > > > https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> > > > >
> > > > > This series replaces each per-CPU uncached route list with a hash table
> > > > > using the network device as its key. Each table uses 64 buckets.
> > > >
> > > > This sounds a bit overkill.  Also, this series still leaves
> > > > O(N * C) loops.
> > > >
> > > > Given unregistering a single device is less common than
> > > > destroying netns, I think the right approach should be to
> > > > make the route flush once in cleanup_net() + outside RTNL.
> > > >
> > > > Could you try this change ? (only compile-tested)
> > > >
> > > Excellent, I'll test this and report back.
> >
> > I found a pre-existing issue, which affects the previous
> > diff, so on top of it, please apply this patch
> >
> >   https://lore.kernel.org/netdev/20260918041439.2575935-1-kuniyu@google.com/T/#u
> >
> > and this diff :
> >
> > ---8<---
> > diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> > index d35b66b33bbc..c12e20e07749 100644
> > --- a/net/ipv4/route.c
> > +++ b/net/ipv4/route.c
> > @@ -1567,12 +1567,13 @@ void rt_add_uncached_list(struct rtable *rt)
> >  {
> >       struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
> >
> > +     rt->dst.rt_uncached_list = ul;
> > +
> >       spin_lock_bh(&ul->lock);
> >
> >       if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> >               rt_replace_uncached_list(rt);
> >       } else {
> > -             rt->dst.rt_uncached_list = ul;
> >               list_add_tail(&rt->dst.rt_uncached, &ul->head);
> >       }
> >
> > diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> > index 6cffe8440b44..f22793abbbb8 100644
> > --- a/net/ipv6/route.c
> > +++ b/net/ipv6/route.c
> > @@ -155,12 +155,13 @@ void rt6_uncached_list_add(struct rt6_info *rt)
> >  {
> >       struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
> >
> > +     rt->dst.rt_uncached_list = ul;
> > +
> >       spin_lock_bh(&ul->lock);
> >
> >       if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> >               rt6_uncached_list_replace(rt);
> >       } else {
> > -             rt->dst.rt_uncached_list = ul;
> >               list_add_tail(&rt->dst.rt_uncached, &ul->head);
> >       }
> >
> > ---8<---
>
> Kuniyuki,
>
> I was able to test this diff, the previous diff you sent plus the fixup
> mentioned above. I was able to confirm even greater reduction in contention
> as measured by how much latency an unrelated process takes when waiting for
> cleanup_net to complete. This makes sense since we don't even need to hold the
> lock when processing those routing entries with your patch.
>
> Some rough average latency numbers with 36 devices, 160k routes, 8 vCPUs:
> - main: 137ms
> - my hashing proposal: 13ms
> - your patchset: 1.8ms
>
> I'd be happy to retest any proposed patches.

Great, thank you for testing !

I will post patches officially once my fix lands in net-next
(so should be after next Thursday)

Thanks !