[PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask

Wang Shuaiwei posted 1 patch 3 weeks ago
include/linux/interrupt.h | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
[PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by Wang Shuaiwei 3 weeks ago
During CPU-hotplug stress testing combined with an I/O
workload, the following message is observed:

  "NOHZ tick-stop error: local softirq work is pending, handler #10!!!"

On CPU offline, ksoftirqd is parked at CPUHP_AP_SMPBOOT_THREADS while
the CPU still takes interrupts. A block completion raised after that
point stays pending, and if the dying CPU goes idle before teardown,
report_idle_softirq() sees it, keeps the tick alive and prints the
error above.

Such pending block completions on the outgoing CPU are drained later by
blk_softirq_cpu_dead() (CPUHP_BLOCK_SOFTIRQ_DEAD), so a pending
BLOCK_SOFTIRQ is harmless here. Add it to SOFTIRQ_HOTPLUG_SAFE_MASK
like the other teardown-covered vectors.

Signed-off-by: Wang Shuaiwei <wangshuaiwei1@xiaomi.com>
---
 include/linux/interrupt.h | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/include/linux/interrupt.h b/include/linux/interrupt.h
index 3bf969ad8fe0..52bb684090c0 100644
--- a/include/linux/interrupt.h
+++ b/include/linux/interrupt.h
@@ -573,9 +573,11 @@ enum
  * _ IRQ_POLL: irq_poll_cpu_dead() migrates the queue
  *
  * _ (HR)TIMER_SOFTIRQ: (hr)timers_dead_cpu() migrates the queue
+ *
+ * _ BLOCK_SOFTIRQ: blk_softirq_cpu_dead() completes the remaining requests
  */
-#define SOFTIRQ_HOTPLUG_SAFE_MASK (BIT(TIMER_SOFTIRQ) | BIT(IRQ_POLL_SOFTIRQ) |\
-				   BIT(HRTIMER_SOFTIRQ) | BIT(RCU_SOFTIRQ))
+#define SOFTIRQ_HOTPLUG_SAFE_MASK (BIT(TIMER_SOFTIRQ) | BIT(BLOCK_SOFTIRQ) |\
+				   BIT(IRQ_POLL_SOFTIRQ) | BIT(HRTIMER_SOFTIRQ) | BIT(RCU_SOFTIRQ))
 
 
 /* map softirq index to softirq name. update 'softirq_to_name' in
-- 
2.43.0
Re: [PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by Frederic Weisbecker 1 week ago
Le Fri, Sep 04, 2026 at 02:45:06PM +0800, Wang Shuaiwei a écrit :
> During CPU-hotplug stress testing combined with an I/O
> workload, the following message is observed:
> 
>   "NOHZ tick-stop error: local softirq work is pending, handler #10!!!"
> 
> On CPU offline, ksoftirqd is parked at CPUHP_AP_SMPBOOT_THREADS while
> the CPU still takes interrupts. A block completion raised after that
> point stays pending, and if the dying CPU goes idle before teardown,
> report_idle_softirq() sees it, keeps the tick alive and prints the
> error above.
> 
> Such pending block completions on the outgoing CPU are drained later by
> blk_softirq_cpu_dead() (CPUHP_BLOCK_SOFTIRQ_DEAD), so a pending
> BLOCK_SOFTIRQ is harmless here. Add it to SOFTIRQ_HOTPLUG_SAFE_MASK
> like the other teardown-covered vectors.
> 
> Signed-off-by: Wang Shuaiwei <wangshuaiwei1@xiaomi.com>

Reviewed-by: Frederic Weisbecker <frederic@kernel.org>

-- 
Frederic Weisbecker
SUSE Labs
Re: [PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by Sebastian Andrzej Siewior 1 week, 1 day ago
On 2026-09-04 14:45:06 [+0800], Wang Shuaiwei wrote:
> During CPU-hotplug stress testing combined with an I/O
> workload, the following message is observed:
> 
>   "NOHZ tick-stop error: local softirq work is pending, handler #10!!!"
> 
> On CPU offline, ksoftirqd is parked at CPUHP_AP_SMPBOOT_THREADS while
> the CPU still takes interrupts. A block completion raised after that
> point stays pending, and if the dying CPU goes idle before teardown,
> report_idle_softirq() sees it, keeps the tick alive and prints the
> error above.
> 
> Such pending block completions on the outgoing CPU are drained later by
> blk_softirq_cpu_dead() (CPUHP_BLOCK_SOFTIRQ_DEAD), so a pending
> BLOCK_SOFTIRQ is harmless here. Add it to SOFTIRQ_HOTPLUG_SAFE_MASK
> like the other teardown-covered vectors.
> 
> Signed-off-by: Wang Shuaiwei <wangshuaiwei1@xiaomi.com>

Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>

Should this go via nohz, softirq or block?

Sebastian
Re: [PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by Wang Shuaiwei 1 week ago
On Thu, 17 Sep 2026 11:51:10 +0200, Sebastian Andrzej Siewior wrote:
> Should this go via nohz, softirq or block?

I'd prefer it to go through the nohz tree since the change is in
SOFTIRQ_HOTPLUG_SAFE_MASK which lives in the NOHZ tick-stop path.

Jens, could you please have a look at this? The patch marks BLOCK_SOFTIRQ
as hotplug-safe, relying on blk_softirq_cpu_dead() to drain pending
completions during CPU offline. Just want to make sure the block layer
is comfortable with that assumption.

Thanks,
Shuaiwei
Re: [PATCH] tick/block: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by Wang Shuaiwei 1 week, 1 day ago
On Fri, 4 Sep 2026 14:45:06 +0800, Wang Shuaiwei wrote:
> During CPU-hotplug stress testing combined with an I/O
> workload, the following message is observed:
> 
>   "NOHZ tick-stop error: local softirq work is pending, handler #10!!!"
> 
> On CPU offline, ksoftirqd is parked at CPUHP_AP_SMPBOOT_THREADS while
> the CPU still takes interrupts. A block completion raised after that
> ...

Hi,

Just wanted to follow up on this patch.
I'd appreciate any feedback or suggestions when you get a chance.

Thanks,
Shuaiwei
[tip: timers/nohz] tick/nohz: Add BLOCK_SOFTIRQ to the hotplug safe mask
Posted by tip-bot2 for Wang Shuaiwei 6 days, 8 hours ago
The following commit has been merged into the timers/nohz branch of tip:

Commit-ID:     a59a940ff89fcaf305362443335f99010ff3b9cf
Gitweb:        https://git.kernel.org/tip/a59a940ff89fcaf305362443335f99010ff3b9cf
Author:        Wang Shuaiwei <wangshuaiwei1@xiaomi.com>
AuthorDate:    Fri, 04 Sep 2026 14:45:06 +08:00
Committer:     Thomas Gleixner <tglx@kernel.org>
CommitterDate: Sat, 19 Sep 2026 23:01:14 +02:00

tick/nohz: Add BLOCK_SOFTIRQ to the hotplug safe mask

During CPU-hotplug stress testing combined with an I/O workload, the
following message is observed:

  "NOHZ tick-stop error: local softirq work is pending, handler #10!!!"

On CPU offline, ksoftirqd is parked at CPUHP_AP_SMPBOOT_THREADS while the
CPU still takes interrupts. A block completion raised after that point
stays pending, and if the dying CPU goes idle before teardown,
report_idle_softirq() sees it, keeps the tick alive and prints the error
above.

Such pending block completions on the outgoing CPU are drained later by
blk_softirq_cpu_dead() (CPUHP_BLOCK_SOFTIRQ_DEAD), so a pending
BLOCK_SOFTIRQ is harmless here. Add it to SOFTIRQ_HOTPLUG_SAFE_MASK like
the other teardown-covered vectors.

Signed-off-by: Wang Shuaiwei <wangshuaiwei1@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Link: https://patch.msgid.link/20260904064506.671203-1-wangshuaiwei1@xiaomi.com
---
 include/linux/interrupt.h | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/include/linux/interrupt.h b/include/linux/interrupt.h
index 3bf969a..52bb684 100644
--- a/include/linux/interrupt.h
+++ b/include/linux/interrupt.h
@@ -573,9 +573,11 @@ enum
  * _ IRQ_POLL: irq_poll_cpu_dead() migrates the queue
  *
  * _ (HR)TIMER_SOFTIRQ: (hr)timers_dead_cpu() migrates the queue
+ *
+ * _ BLOCK_SOFTIRQ: blk_softirq_cpu_dead() completes the remaining requests
  */
-#define SOFTIRQ_HOTPLUG_SAFE_MASK (BIT(TIMER_SOFTIRQ) | BIT(IRQ_POLL_SOFTIRQ) |\
-				   BIT(HRTIMER_SOFTIRQ) | BIT(RCU_SOFTIRQ))
+#define SOFTIRQ_HOTPLUG_SAFE_MASK (BIT(TIMER_SOFTIRQ) | BIT(BLOCK_SOFTIRQ) |\
+				   BIT(IRQ_POLL_SOFTIRQ) | BIT(HRTIMER_SOFTIRQ) | BIT(RCU_SOFTIRQ))
 
 
 /* map softirq index to softirq name. update 'softirq_to_name' in