include/trace/events/workqueue.h | 147 +++++++++++++++++++++++++++++++ kernel/workqueue.c | 35 +++++++- 2 files changed, 180 insertions(+), 2 deletions(-)
Hi Tejun, Lai, While the workqueue subsystem maintains rich internal telemetry via pwq->stats[], wq_cpu_intensive_thresh_us, and distress mechanisms, several critical state transitions and execution anomalies currently lack real-time event notifications. Across production fleets and low-latency networking workloads, polling pwq->stats[] or running drgn scripts is impractical for detecting intermittent stalls. Tail-latency spikes and packet drops often stem from softirq overruns or latency-critical work items queueing behind CPU-bound tasks. These event-driven tracepoints allow zero-overhead eBPF tools and latency profilers to capture stack traces and kernel context at the exact moment a starvation event or softirq budget exhaustion occurs. This patch series introduces lightweight tracepoints for these key operational boundaries: Patch 1 adds workqueue_cpu_intensive tracepoint. When a concurrency-managed worker runs for longer than wq_cpu_intensive_thresh_us without sleeping, wq_worker_tick() marks it as WORKER_CPU_INTENSIVE and kicks the pool to prevent queue starvation. The tracepoint captures the offending work function, workqueue name, CPU, and elapsed duration. Patch 2 adds workqueue_mayday and workqueue_rescued tracepoints. One when worker allocation stalls trigger mayday distress, and another when pending work items are handed off to the rescuer thread to ensure forward progress. Patch 3 adds workqueue_bh_budget_yield tracepoint. Bottom-Half (BH) workqueues enforce execution limits in softirq context (BH_WORKER_JIFFIES and BH_WORKER_RESTARTS). When a BH worker hits these limits while pending work remains, it yields and re-raises the softirq. The tracepoint can be used to identify softirq budget saturation and track whether yielding occurred due to time slice expiration or restart counts. Changes since v1: - Fixed 32-bit architecture build/link failures reported by including <linux/math64.h> and using div_u64() (Tejun Heo) - Kept the fast-path threshold comparison in wq_worker_tick() untouched to avoid performing division on every timer tick - Calculated elapsed duration using div_u64() on the cold path only after marking the worker CPU_INTENSIVE and releasing pool->lock - Eliminated the race where re-reading jiffies after the loop could misattribute an exhausted restart budget to a timeout; timeout is now derived directly inside the loop from the condition that actually terminated it (Tejun Heo) - Corrected restart counting. Explicitly tracked and reported the actual number of restarts executed rather than total loop iterations (Tejun Heo) - Link to v1: https://lore.kernel.org/lkml/20260829230517.42468-1-atomlin@atomlin.com/ Aaron Tomlin (3): workqueue: Add workqueue_cpu_intensive tracepoint workqueue: Add workqueue_mayday and workqueue_rescued tracepoints workqueue: Add workqueue_bh_budget_yield tracepoint include/trace/events/workqueue.h | 147 +++++++++++++++++++++++++++++++ kernel/workqueue.c | 35 +++++++- 2 files changed, 180 insertions(+), 2 deletions(-) -- 2.55.0
Hello, Aaron. > Aaron Tomlin (3): > workqueue: Add workqueue_cpu_intensive tracepoint > workqueue: Add workqueue_mayday and workqueue_rescued tracepoints > workqueue: Add workqueue_bh_budget_yield tracepoint Applied 1-3 to wq/for-7.4 with the mayday interval references removed and "executing CPU" changed to "pool CPU" in the BH tracepoint description. The comment adjustment is below. Thanks. -- tejun diff --git a/include/trace/events/workqueue.h b/include/trace/events/workqueue.h index 12ac24586835..eb97d4dbb5b4 100644 --- a/include/trace/events/workqueue.h +++ b/include/trace/events/workqueue.h @@ -169,9 +169,8 @@ TRACE_EVENT(workqueue_cpu_intensive, * workqueue_mayday - called when a pool_workqueue sends mayday to rescuer * @pwq: pointer to struct pool_workqueue * - * This event occurs when a worker pool fails to create a new worker - * within MAYDAY_INTERVAL and requests the workqueue's rescuer thread to - * process pending works. + * This event occurs when a worker pool fails to create a new worker and + * requests the workqueue's rescuer thread to process pending works. */ TRACE_EVENT(workqueue_mayday,
On Fri, Sep 04, 2026 at 10:22:45AM -1000, Tejun Heo wrote: > Hello, Aaron. > > > Aaron Tomlin (3): > > workqueue: Add workqueue_cpu_intensive tracepoint > > workqueue: Add workqueue_mayday and workqueue_rescued tracepoints > > workqueue: Add workqueue_bh_budget_yield tracepoint > > Applied 1-3 to wq/for-7.4 with the mayday interval references removed and > "executing CPU" changed to "pool CPU" in the BH tracepoint description. > The comment adjustment is below. > > Thanks. > > -- > tejun Thank you Tejun. -- Aaron Tomlin
© 2016 - 2026 Red Hat, Inc.