[PATCH] kthread: Report cpumask allocation failure without warning

Quchaosheng posted 1 patch 1 week, 2 days ago
kernel/kthread.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
[PATCH] kthread: Report cpumask allocation failure without warning
Posted by Quchaosheng 1 week, 2 days ago
kthread_affine_node() warns when zalloc_cpumask_var() fails:

  if (!zalloc_cpumask_var(&affinity, GFP_KERNEL)) {
          WARN_ON_ONCE(1);
          return;
  }

The allocation uses GFP_KERNEL, so it can fail under memory pressure or
fault injection.  A failed allocation is a recoverable condition and not a
kernel bug, so the warning is noise.  syzbot reports it for a WireGuard
NAPI thread:

  WARNING: kernel/kthread.c:359 at kthread_affine_node+0x200/0x2e8
  CPU: 0 PID: 5207 Comm: napi/wg2-0
  Call Trace:
   alloc_cpumask_var_node+0xfc/0x138
   zalloc_cpumask_var
   kthread_affine_node+0x148/0x2e8
   kthread+0x29c/0x3d4
   ret_from_fork+0x10/0x20

The failure is not silent though: the early return also skips
list_add_tail() of kthread::affinity_node, so the thread never joins
kthread_affinity_list and kthreads_online_cpu() will not fix up its
affinity on a later CPU hotplug.  Keep the failure visible with
pr_warn_once() instead of dropping the message.

Reported-by: syzbot+37ca7ae3e98cb65c3209@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=37ca7ae3e98cb65c3209
Signed-off-by: Quchaosheng <quchaosheng000406@163.com>
---
 kernel/kthread.c | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)

diff --git a/kernel/kthread.c b/kernel/kthread.c
index a3f95c904..b59fa7c7e 100644
--- a/kernel/kthread.c
+++ b/kernel/kthread.c
@@ -356,7 +356,14 @@ static void kthread_affine_node(void)
 		return;
 
 	if (!zalloc_cpumask_var(&affinity, GFP_KERNEL)) {
-		WARN_ON_ONCE(1);
+		/*
+		 * The thread stays out of kthread_affinity_list, so a later
+		 * CPU hotplug will not fix up its affinity. Report it, but do
+		 * not warn: the allocation can fail under memory pressure or
+		 * fault injection, and that is not a kernel bug.
+		 */
+		pr_warn_once("kthread: %s: no cpumask, node affinity not set\n",
+			     current->comm);
 		return;
 	}
 
-- 
2.43.0
Re: [PATCH] kthread: Report cpumask allocation failure without warning
Posted by Bradley Morgan 1 week, 1 day ago
On 16 September 2026 02:59:48 BST, Quchaosheng <quchaosheng000406@163.com>
wrote:
>kthread_affine_node() warns when zalloc_cpumask_var() fails:
>
>  if (!zalloc_cpumask_var(&affinity, GFP_KERNEL)) {
>          WARN_ON_ONCE(1);
>          return;
>  }
>
>The allocation uses GFP_KERNEL, so it can fail under memory pressure or
>fault injection.  A failed allocation is a recoverable condition and not a
>kernel bug, so the warning is noise.  syzbot reports it for a WireGuard
>NAPI thread:
>
>  WARNING: kernel/kthread.c:359 at kthread_affine_node+0x200/0x2e8
>  CPU: 0 PID: 5207 Comm: napi/wg2-0
>  Call Trace:
>   alloc_cpumask_var_node+0xfc/0x138
>   zalloc_cpumask_var
>   kthread_affine_node+0x148/0x2e8
>   kthread+0x29c/0x3d4
>   ret_from_fork+0x10/0x20
>
>The failure is not silent though: the early return also skips
>list_add_tail() of kthread::affinity_node, so the thread never joins
>kthread_affinity_list and kthreads_online_cpu() will not fix up its
>affinity on a later CPU hotplug.  Keep the failure visible with
>pr_warn_once() instead of dropping the message.

I swear there was a patch for this I  kinda NAKed.

>
>Reported-by: syzbot+37ca7ae3e98cb65c3209@syzkaller.appspotmail.com
>Closes: https://syzkaller.appspot.com/bug?extid=37ca7ae3e98cb65c3209
>Signed-off-by: Quchaosheng <quchaosheng000406@163.com>
>---
> kernel/kthread.c | 9 ++++++++-
> 1 file changed, 8 insertions(+), 1 deletion(-)
>
>diff --git a/kernel/kthread.c b/kernel/kthread.c
>index a3f95c904..b59fa7c7e 100644
>--- a/kernel/kthread.c
>+++ b/kernel/kthread.c
>@@ -356,7 +356,14 @@ static void kthread_affine_node(void)
> 		return;
> 
> 	if (!zalloc_cpumask_var(&affinity, GFP_KERNEL)) {
>-		WARN_ON_ONCE(1);
>+		/*
>+		 * The thread stays out of kthread_affinity_list, so a later
>+		 * CPU hotplug will not fix up its affinity. Report it, but do
>+		 * not warn: the allocation can fail under memory pressure or
>+		 * fault injection, and that is not a kernel bug.
>+		 */

For the comment length, I have to ask did you use AI to develop this?

>+		pr_warn_once("kthread: %s: no cpumask, node affinity not set\n",
>+			     current->comm);

Ummmmmm. I mean, okay then? But what effect does this make? (Except from
the message).

Under memory pressure I presume

1: your memory needs replacing, you get what you get
2: you did this on purpose, you deserve whatever you get.

And under fault injection is just a test.

I really think this code is sane as is, 

NAK.

> 		return;
> 	}
> 
>

--- Thanks!
https://lore.kernel.org/all/EE579805-42F2-4C58-B752-F28779EEB717@grrlz.net/
Re: [PATCH] kthread: Report cpumask allocation failure without warning
Posted by Quchaosheng 1 week, 1 day ago
Hi Bradley,

Thanks for the review.

On the AI question, since you asked directly: yes. I used an AI coding
assistant while preparing this, including the comment and the changelog. I
reviewed the change and tested it myself, but per
Documentation/process/coding-assistants.rst the submission should have carried
an "Assisted-by: LLM ..." tag and it did not. That is my mistake and I will add
the tag on any future revision.

On the NAK itself, I would like to answer your "what effect does this make
(except from the message)" question, because there is more to it than the text.

WARN_ON_ONCE(1) is not only a message. __warn() taints the kernel with
TAINT_WARN, prints modules and dumps the stack, and calls
check_panic_on_warn() (kernel/panic.c). So on any kernel running with
panic_on_warn=1, a failed GFP_KERNEL allocation in the thread creation path
becomes a full panic. That is the harm: the allocation failure is recoverable,
but the WARN reports it as a kernel bug, taints the box and can take it down.
pr_warn_once() keeps the report and drops the taint and the panic.

You also wrote on 14 September:

    Perhaps you could warn -> info? Because people may prefer to know if it's
    broken.

which I read as the same direction. I kept pr_warn_once() rather than
pr_info() so the line still stands out in the log.

On "under memory pressure you get what you get" and "fault injection is just a
test": I agree with both. Nothing is going to work well at that point. My only
argument for changing the severity is the taint/panic consequence above, not
the message length.

You said on 14 September that you were not comfortable taking a position until
the maintainers had input. Frederic wrote this code and has already commented
on the thread, so I am happy to leave the call to him.

Thanks,
Quchaosheng