[PATCH v2] sched/topology: Free NUMA masks on topology allocation failure

Fengyu Wang posted 1 patch 1 month, 2 weeks ago
kernel/sched/topology.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)
[PATCH v2] sched/topology: Free NUMA masks on topology allocation failure
Posted by Fengyu Wang 1 month, 2 weeks ago
sched_init_numa() publishes sched_domains_numa_masks before it
allocates the topology array.  When that allocation fails, the early
return leaves the masks published while sched_domains_numa_levels is
still zero: nothing dereferences them, but nothing can free them
either, and the topology they were built for is never installed.

Free the masks on that path, and publish them only once the topology
array they were built for has been allocated.

Fixes: cb83b629bae0 ("sched/numa: Rewrite the CONFIG_NUMA sched domain support")
Signed-off-by: Fengyu Wang <wangfengyu@hygon.cn>
---
v2:
 - Publish sched_domains_numa_masks only after the topology array has
   been allocated, instead of publishing it early and unpublishing it
   on the failure path.  This drops the rcu_assign_pointer(NULL) and
   the synchronize_rcu() from the error path (Tim Chen).

v1: https://lore.kernel.org/lkml/20260731081413.5505-1-wangfengyu@hygon.cn/

Tested by hardcoding tl to NULL right after the kzalloc() to force the
failure path; the masks are released and the machine boots normally.

 kernel/sched/topology.c | 12 ++++++++++--
 1 file changed, 10 insertions(+), 2 deletions(-)

diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 21e816ad23ee..50457f720808 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -2392,15 +2392,23 @@ void sched_init_numa(int offline_node)
 			}
 		}
 	}
-	rcu_assign_pointer(sched_domains_numa_masks, masks);
 
 	/* Compute default topology size */
 	for (i = 0; sched_domain_topology[i].mask; i++);
 
 	tl = kzalloc((i + nr_levels + 1) *
 			sizeof(struct sched_domain_topology_level), GFP_KERNEL);
-	if (!tl)
+	if (!tl) {
+		for (i = 0; i < nr_levels; i++) {
+			for_each_node(j)
+				kfree(masks[i][j]);
+			kfree(masks[i]);
+		}
+		kfree(masks);
 		return;
+	}
+
+	rcu_assign_pointer(sched_domains_numa_masks, masks);
 
 	/*
 	 * Copy the default topology bits..
-- 
2.34.1
Re: [PATCH v2] sched/topology: Free NUMA masks on topology allocation failure
Posted by Valentin Schneider 1 month, 2 weeks ago
On 12/08/26 14:22, Fengyu Wang wrote:
> sched_init_numa() publishes sched_domains_numa_masks before it
> allocates the topology array.  When that allocation fails, the early
> return leaves the masks published while sched_domains_numa_levels is
> still zero: nothing dereferences them, but nothing can free them
> either, and the topology they were built for is never installed.
>
> Free the masks on that path, and publish them only once the topology
> array they were built for has been allocated.
>

sched_init_numa() still returns nothing, and an allocation failure in
scheduler topology code is pretty much always paired with a dumpster fire,
but I guess this may help you get further in the booting process of a
borked kernel...

> Fixes: cb83b629bae0 ("sched/numa: Rewrite the CONFIG_NUMA sched domain support")
> Signed-off-by: Fengyu Wang <wangfengyu@hygon.cn>

Reviewed-by: Valentin Schneider <vschneid@redhat.com>
Re: [PATCH v2] sched/topology: Free NUMA masks on topology allocation failure
Posted by Tim Chen 1 month, 2 weeks ago
On Wed, 2026-08-12 at 14:22 +0800, Fengyu Wang wrote:
> sched_init_numa() publishes sched_domains_numa_masks before it
> allocates the topology array.  When that allocation fails, the early
> return leaves the masks published while sched_domains_numa_levels is
> still zero: nothing dereferences them, but nothing can free them
> either, and the topology they were built for is never installed.
> 
> Free the masks on that path, and publish them only once the topology
> array they were built for has been allocated.
> 

Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>

Tim

> Fixes: cb83b629bae0 ("sched/numa: Rewrite the CONFIG_NUMA sched domain support")
> Signed-off-by: Fengyu Wang <wangfengyu@hygon.cn>
> ---
> v2:
>  - Publish sched_domains_numa_masks only after the topology array has
>    been allocated, instead of publishing it early and unpublishing it
>    on the failure path.  This drops the rcu_assign_pointer(NULL) and
>    the synchronize_rcu() from the error path (Tim Chen).
> 
> v1: https://lore.kernel.org/lkml/20260731081413.5505-1-wangfengyu@hygon.cn/
> 
> Tested by hardcoding tl to NULL right after the kzalloc() to force the
> failure path; the masks are released and the machine boots normally.
> 
>  kernel/sched/topology.c | 12 ++++++++++--
>  1 file changed, 10 insertions(+), 2 deletions(-)
> 
> diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
> index 21e816ad23ee..50457f720808 100644
> --- a/kernel/sched/topology.c
> +++ b/kernel/sched/topology.c
> @@ -2392,15 +2392,23 @@ void sched_init_numa(int offline_node)
>  			}
>  		}
>  	}
> -	rcu_assign_pointer(sched_domains_numa_masks, masks);
>  
>  	/* Compute default topology size */
>  	for (i = 0; sched_domain_topology[i].mask; i++);
>  
>  	tl = kzalloc((i + nr_levels + 1) *
>  			sizeof(struct sched_domain_topology_level), GFP_KERNEL);
> -	if (!tl)
> +	if (!tl) {
> +		for (i = 0; i < nr_levels; i++) {
> +			for_each_node(j)
> +				kfree(masks[i][j]);
> +			kfree(masks[i]);
> +		}
> +		kfree(masks);
>  		return;
> +	}
> +
> +	rcu_assign_pointer(sched_domains_numa_masks, masks);
>  
>  	/*
>  	 * Copy the default topology bits..