[PATCH v2 0/3] Add telemetry tracepoints for CPU hogs, distress, and BH budget yields

Aaron Tomlin posted 3 patches 3 weeks ago
include/trace/events/workqueue.h | 147 +++++++++++++++++++++++++++++++
kernel/workqueue.c               |  35 +++++++-
2 files changed, 180 insertions(+), 2 deletions(-)
[PATCH v2 0/3] Add telemetry tracepoints for CPU hogs, distress, and BH budget yields
Posted by Aaron Tomlin 3 weeks ago
Hi Tejun, Lai,

While the workqueue subsystem maintains rich internal telemetry via
pwq->stats[], wq_cpu_intensive_thresh_us, and distress mechanisms, several
critical state transitions and execution anomalies currently lack real-time
event notifications.

Across production fleets and low-latency networking workloads, polling
pwq->stats[] or running drgn scripts is impractical for detecting
intermittent stalls. Tail-latency spikes and packet drops often stem from
softirq overruns or latency-critical work items queueing behind CPU-bound
tasks. These event-driven tracepoints allow zero-overhead eBPF tools and
latency profilers to capture stack traces and kernel context at the exact
moment a starvation event or softirq budget exhaustion occurs.

This patch series introduces lightweight tracepoints for these key
operational boundaries:

Patch 1 adds workqueue_cpu_intensive tracepoint. When a concurrency-managed
worker runs for longer than wq_cpu_intensive_thresh_us without sleeping,
wq_worker_tick() marks it as WORKER_CPU_INTENSIVE and kicks the pool to
prevent queue starvation. The tracepoint captures the offending work
function, workqueue name, CPU, and elapsed duration.

Patch 2 adds workqueue_mayday and workqueue_rescued tracepoints. One when
worker allocation stalls trigger mayday distress, and another when pending
work items are handed off to the rescuer thread to ensure forward progress.

Patch 3 adds workqueue_bh_budget_yield tracepoint. Bottom-Half (BH)
workqueues enforce execution limits in softirq context (BH_WORKER_JIFFIES
and BH_WORKER_RESTARTS). When a BH worker hits these limits while pending
work remains, it yields and re-raises the softirq. The tracepoint can be
used to identify softirq budget saturation and track whether yielding
occurred due to time slice expiration or restart counts.

Changes since v1:

 - Fixed 32-bit architecture build/link failures reported by including
   <linux/math64.h> and using div_u64() (Tejun Heo)

 - Kept the fast-path threshold comparison in wq_worker_tick() untouched to
   avoid performing division on every timer tick

 - Calculated elapsed duration using div_u64() on the cold path only after
   marking the worker CPU_INTENSIVE and releasing pool->lock

 - Eliminated the race where re-reading jiffies after the loop could
   misattribute an exhausted restart budget to a timeout; timeout is now
   derived directly inside the loop from the condition that actually
   terminated it (Tejun Heo)

 - Corrected restart counting. Explicitly tracked and reported the actual
   number of restarts executed rather than total loop iterations
   (Tejun Heo)

 - Link to v1: https://lore.kernel.org/lkml/20260829230517.42468-1-atomlin@atomlin.com/

Aaron Tomlin (3):
  workqueue: Add workqueue_cpu_intensive tracepoint
  workqueue: Add workqueue_mayday and workqueue_rescued tracepoints
  workqueue: Add workqueue_bh_budget_yield tracepoint

 include/trace/events/workqueue.h | 147 +++++++++++++++++++++++++++++++
 kernel/workqueue.c               |  35 +++++++-
 2 files changed, 180 insertions(+), 2 deletions(-)

-- 
2.55.0
Re: [PATCH v2 0/3] Add telemetry tracepoints for CPU hogs, distress, and BH budget yields
Posted by Tejun Heo 3 weeks ago
Hello, Aaron.

> Aaron Tomlin (3):
>   workqueue: Add workqueue_cpu_intensive tracepoint
>   workqueue: Add workqueue_mayday and workqueue_rescued tracepoints
>   workqueue: Add workqueue_bh_budget_yield tracepoint

Applied 1-3 to wq/for-7.4 with the mayday interval references removed and
"executing CPU" changed to "pool CPU" in the BH tracepoint description.
The comment adjustment is below.

Thanks.

-- 
tejun

diff --git a/include/trace/events/workqueue.h b/include/trace/events/workqueue.h
index 12ac24586835..eb97d4dbb5b4 100644
--- a/include/trace/events/workqueue.h
+++ b/include/trace/events/workqueue.h
@@ -169,9 +169,8 @@ TRACE_EVENT(workqueue_cpu_intensive,
  * workqueue_mayday - called when a pool_workqueue sends mayday to rescuer
  * @pwq:	pointer to struct pool_workqueue
  *
- * This event occurs when a worker pool fails to create a new worker
- * within MAYDAY_INTERVAL and requests the workqueue's rescuer thread to
- * process pending works.
+ * This event occurs when a worker pool fails to create a new worker and
+ * requests the workqueue's rescuer thread to process pending works.
  */
 TRACE_EVENT(workqueue_mayday,
Re: [PATCH v2 0/3] Add telemetry tracepoints for CPU hogs, distress, and BH budget yields
Posted by Aaron Tomlin 2 weeks, 6 days ago
On Fri, Sep 04, 2026 at 10:22:45AM -1000, Tejun Heo wrote:
> Hello, Aaron.
> 
> > Aaron Tomlin (3):
> >   workqueue: Add workqueue_cpu_intensive tracepoint
> >   workqueue: Add workqueue_mayday and workqueue_rescued tracepoints
> >   workqueue: Add workqueue_bh_budget_yield tracepoint
> 
> Applied 1-3 to wq/for-7.4 with the mayday interval references removed and
> "executing CPU" changed to "pool CPU" in the BH tracepoint description.
> The comment adjustment is below.
> 
> Thanks.
> 
> -- 
> tejun

Thank you Tejun.

-- 
Aaron Tomlin