From nobody Sat Jul 25 05:29:26 2026 Received: from fanzine2.igalia.com (fanzine2.igalia.com [213.97.179.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7BD043264FB for ; Fri, 17 Jul 2026 09:02:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=213.97.179.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784278948; cv=none; b=pWIsBerYlnjVKpO+EM5ShpGUHjxdxQDQmYlv2fr1n+GObUKdBalUk/IhZoH76/TzjeL6aNKT/5TkI54PrvuSXNMx2TwbUP09qA3wbG3cor6EsHYL/FE+6/kucEHad7+Cr9aCUElFAs/qmCotUUbEO56a0Y4K6U8Un0rtRwVZF7o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784278948; c=relaxed/simple; bh=z817GlJvpW/v0SEwuM5Zg46xOqYuV76xBfbmT/tJJSg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ff792fLYKGEzhYTT85wjRuXnfu9nsDg1Hirrxe4grLQ0B6cDr4GjeoUMLYsRY8DsPwnBE1s01r2qiLThDk5Zvv+mKoV3fHyku6fch8C6noqwUsCVcigQutJ/zVxGL5j1mRZmfukdR0yCYtlI3d5447+DKnBeueTa7zS2pdIRsbg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=igalia.com; spf=pass smtp.mailfrom=igalia.com; dkim=pass (2048-bit key) header.d=igalia.com header.i=@igalia.com header.b=j2D9t+nN; arc=none smtp.client-ip=213.97.179.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=igalia.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=igalia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=igalia.com header.i=@igalia.com header.b="j2D9t+nN" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=igalia.com; s=20170329; h=Content-Transfer-Encoding:MIME-Version:Message-ID:Date:Subject: Cc:To:From:From:Reply-To; bh=twaV3rxbqyRlXZsIkPLUkgNmSIZGFN45/0Jworau1do=; b= j2D9t+nNJUwlAKjXelAjMRXXRFH625oyJ5j9ZPqHNLvl4RIh/u8SHhU3LaWi9nRSqbBWkOIJCBVqU Nh5a/zf4SHm+iXk29lfQnpfq0NQyzZxJIxHI+asBRU2xt/enfcLP4Z4b+Iudm5AESu3HIC2Y8Z1mC FLBli35JcQosxUN6cLdOuL3LdJvlIoEaMsyMhGj6o6M0Yr32vSzbmrHzoiiGursEm/CMrLZM4rNbQ SzVrwptW9yrgGeqKHJCIcTrCRof5dDlhXPg4cpeh/HrmUPDYMWxKHLtwS0uUxUrltqFbbpAtPiTQj BSXFFYFn+m+GUuckn35iUlZPwo0hF82mHQ==; Received: from [90.240.106.137] (helo=localhost) by fanzine2.igalia.com with esmtpsa (Cipher TLS1.3:ECDHE_SECP256R1__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim) id 1wkeSS-00GJcS-Go; Fri, 17 Jul 2026 11:02:12 +0200 From: Tvrtko Ursulin To: dri-devel@lists.freedesktop.org Cc: Boris Brezillon , Steven Price , Liviu Dudau , Chia-I Wu , Matthew Brost , kernel-dev@igalia.com, linux-kernel@vger.kernel.org, Tvrtko Ursulin , Chia-I Wu , Tejun Heo Subject: [RFC v3 1/2] workqueue: Add support for real-time workers Date: Fri, 17 Jul 2026 10:02:08 +0100 Message-ID: <20260717090209.26931-2-tvrtko.ursulin@igalia.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260717090209.26931-1-tvrtko.ursulin@igalia.com> References: <20260717090209.26931-1-tvrtko.ursulin@igalia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" For use cases such as the DRM scheduler submitting work to the GPU on behalf of low latency userspace applications, where latter have sufficient privileges to have had successfully obtained realtime Vulkan global priority, competing with random background CPU load can create large latency spikes which gets in the way of a smooth user experience. For these situations the existing WQ_HIGHPRI does not bring a noticeable improvement and a stronger hint is needed. Lets add WQ_RTPRI which creates workers with a SCHED_FIFO scheduling class to improve this. We use a minimum priority level since we only care about winning the contest against normal background CPU load. We also limit the number of instantiated threads, both per workqueue to a maximum of two, or one on a dual processor machine, and system-wide to one less than the number of CPUs. The idea of that is to prevent system wide starvation caused by potentially misbehaving work items. Signed-off-by: Tvrtko Ursulin Cc: Boris Brezillon Cc: Chia-I Wu Cc: Liviu Dudau Cc: Matthew Brost Cc: Steven Price Cc: Tejun Heo --- v2: * Limit WQ_RTPRI to unbound workqueues and make it have strict CPU affinitity. (Tejun) * Fixed commit message typos. (AI) * Fixed sysfs handling, max_active setting and user modified nice application. (AI) v3: * Fix worker->pool null pointer dereference race by moving the global decrement to detach_dying_workers(). * Rebase for upstream changes. --- include/linux/workqueue.h | 23 ++++-- kernel/workqueue.c | 152 +++++++++++++++++++++++++++++--------- 2 files changed, 135 insertions(+), 40 deletions(-) diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h index a283766a192a..e9ab53568e0c 100644 --- a/include/linux/workqueue.h +++ b/include/linux/workqueue.h @@ -140,6 +140,13 @@ enum wq_affn_scope { WQ_AFFN_NR_TYPES, }; =20 +enum wq_priority { + WQ_PRIO_NORMAL =3D 0, + WQ_PRIO_HIGH =3D 1, + WQ_PRIO_RT =3D 2, + NUM_WQ_PRIO, /* Keep last */ +}; + /** * struct workqueue_attrs - A struct for workqueue attributes. * @@ -147,7 +154,12 @@ enum wq_affn_scope { */ struct workqueue_attrs { /** - * @nice: nice level + * @prio: priority level + */ + enum wq_priority prio; + + /** + * @nice: nice level for WQ_PRIO_HIGH */ int nice; =20 @@ -374,8 +386,9 @@ enum wq_flags { WQ_FREEZABLE =3D 1 << 2, /* freeze during suspend */ WQ_MEM_RECLAIM =3D 1 << 3, /* may be used for memory reclaim */ WQ_HIGHPRI =3D 1 << 4, /* high priority */ - WQ_CPU_INTENSIVE =3D 1 << 5, /* cpu intensive workqueue */ - WQ_SYSFS =3D 1 << 6, /* visible in sysfs, see workqueue_sysfs_register()= */ + WQ_RTPRI =3D 1 << 5, /* real-time priority, valid only with WQ_UNBOUND */ + WQ_CPU_INTENSIVE =3D 1 << 6, /* cpu intensive workqueue */ + WQ_SYSFS =3D 1 << 7, /* visible in sysfs, see workqueue_sysfs_register()= */ =20 /* * Per-cpu workqueues are generally preferred because they tend to @@ -402,8 +415,8 @@ enum wq_flags { * * http://thread.gmane.org/gmane.linux.kernel/1480396 */ - WQ_POWER_EFFICIENT =3D 1 << 7, - WQ_PERCPU =3D 1 << 8, /* bound to a specific cpu */ + WQ_POWER_EFFICIENT =3D 1 << 8, + WQ_PERCPU =3D 1 << 9, /* bound to a specific cpu */ =20 __WQ_DESTROYING =3D 1 << 15, /* internal: workqueue is destroying */ __WQ_DRAINING =3D 1 << 16, /* internal: workqueue is draining */ diff --git a/kernel/workqueue.c b/kernel/workqueue.c index 78068ae8f28a..30fed46a3164 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -509,10 +509,10 @@ static DEFINE_IDR(worker_pool_idr); /* PR: idr of all= pools */ static DEFINE_HASHTABLE(unbound_pool_hash, UNBOUND_POOL_HASH_ORDER); =20 /* I: attributes used when instantiating standard unbound pools on demand = */ -static struct workqueue_attrs *unbound_std_wq_attrs[NR_STD_WORKER_POOLS]; +static struct workqueue_attrs *unbound_std_wq_attrs[NR_STD_WORKER_POOLS + = 1]; =20 /* I: attributes used when instantiating ordered pools on demand */ -static struct workqueue_attrs *ordered_wq_attrs[NR_STD_WORKER_POOLS]; +static struct workqueue_attrs *ordered_wq_attrs[NR_STD_WORKER_POOLS + 1]; =20 /* * I: kthread_worker to release pwq's. pwq release needs to be bounced to a @@ -546,6 +546,8 @@ EXPORT_SYMBOL_GPL(system_bh_highpri_wq); struct workqueue_struct *system_dfl_long_wq __ro_after_init; EXPORT_SYMBOL_GPL(system_dfl_long_wq); =20 +static atomic_t total_rtpri_workers =3D ATOMIC_INIT(0); + static int worker_thread(void *__worker); static void workqueue_sysfs_unregister(struct workqueue_struct *wq); static void show_pwq(struct pool_workqueue *pwq); @@ -1236,9 +1238,7 @@ static bool assign_work(struct work_struct *work, str= uct worker *worker, =20 static struct irq_work *bh_pool_irq_work(struct worker_pool *pool) { - int high =3D pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL ? 1 : 0; - - return &per_cpu(bh_pool_irq_works, pool->cpu)[high]; + return &per_cpu(bh_pool_irq_works, pool->cpu)[pool->attrs->prio]; } =20 static void kick_bh_pool(struct worker_pool *pool) @@ -1251,7 +1251,7 @@ static void kick_bh_pool(struct worker_pool *pool) return; } #endif - if (pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio >=3D WQ_PRIO_HIGH) raise_softirq_irqoff(HI_SOFTIRQ); else raise_softirq_irqoff(TASKLET_SOFTIRQ); @@ -2809,13 +2809,16 @@ static int format_worker_id(char *buf, size_t size,= struct worker *worker, worker->rescue_wq->name); =20 if (pool) { - if (pool->cpu >=3D 0) + if (pool->cpu >=3D 0) { + const char *suffix[NUM_WQ_PRIO] =3D { "", "H", "R" }; + return scnprintf(buf, size, "kworker/%d:%d%s", pool->cpu, worker->id, - pool->attrs->nice < 0 ? "H" : ""); - else + suffix[pool->attrs->prio]); + } else { return scnprintf(buf, size, "kworker/u%d:%d", pool->id, worker->id); + } } else { return scnprintf(buf, size, "kworker/dying"); } @@ -2838,12 +2841,30 @@ static struct worker *create_worker(struct worker_p= ool *pool) struct worker *worker; int id; =20 + /* + * Do not consume all CPUs with RT workers to avoid scheduler + * starvation. + */ + if (pool->attrs->prio =3D=3D WQ_PRIO_RT) { + unsigned int max =3D num_online_cpus(); + + if (max > 2) + max =3D max - 1; + else + max =3D 1; + + if (atomic_inc_return(&total_rtpri_workers) > max) { + atomic_dec(&total_rtpri_workers); + return NULL; + } + } + /* ID is needed to determine kthread name */ id =3D ida_alloc(&pool->worker_ida, GFP_KERNEL); if (id < 0) { pr_err_once("workqueue: Failed to allocate a worker ID: %pe\n", ERR_PTR(id)); - return NULL; + goto fail_ida; } =20 worker =3D alloc_worker(pool->node); @@ -2871,7 +2892,11 @@ static struct worker *create_worker(struct worker_po= ol *pool) goto fail; } =20 - set_user_nice(worker->task, pool->attrs->nice); + if (pool->attrs->prio =3D=3D WQ_PRIO_RT) + sched_set_fifo_low(worker->task); + else + set_user_nice(worker->task, pool->attrs->nice); + kthread_bind_mask(worker->task, pool_allowed_cpus(pool)); } =20 @@ -2896,6 +2921,9 @@ static struct worker *create_worker(struct worker_poo= l *pool) =20 return worker; =20 +fail_ida: + if (pool->attrs->prio =3D=3D WQ_PRIO_RT) + atomic_dec(&total_rtpri_workers); fail: ida_free(&pool->worker_ida, id); kfree(worker); @@ -2906,8 +2934,11 @@ static void detach_dying_workers(struct list_head *c= ull_list) { struct worker *worker; =20 - list_for_each_entry(worker, cull_list, entry) + list_for_each_entry(worker, cull_list, entry) { + if (worker->pool->attrs->prio =3D=3D WQ_PRIO_RT) + atomic_dec(&total_rtpri_workers); detach_worker(worker); + } } =20 static void reap_dying_workers(struct list_head *cull_list) @@ -3773,7 +3804,7 @@ static void drain_dead_softirq_workfn(struct work_str= uct *work) * don't hog this CPU's BH. */ if (repeat) { - if (pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio >=3D WQ_PRIO_HIGH) queue_work(system_bh_highpri_wq, work); else queue_work(system_bh_wq, work); @@ -3805,7 +3836,7 @@ void workqueue_softirq_dead(unsigned int cpu) dead_work.pool =3D pool; init_completion(&dead_work.done); =20 - if (pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio >=3D WQ_PRIO_HIGH) queue_work(system_bh_highpri_wq, &dead_work.work); else queue_work(system_bh_wq, &dead_work.work); @@ -4780,6 +4811,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(= void) static void copy_workqueue_attrs(struct workqueue_attrs *to, const struct workqueue_attrs *from) { + to->prio =3D from->prio; to->nice =3D from->nice; cpumask_copy(to->cpumask, from->cpumask); cpumask_copy(to->__pod_cpumask, from->__pod_cpumask); @@ -4811,6 +4843,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs = *attrs) { u32 hash =3D 0; =20 + hash =3D jhash_1word(attrs->prio, hash); hash =3D jhash_1word(attrs->nice, hash); hash =3D jhash_1word(attrs->affn_strict, hash); hash =3D jhash(cpumask_bits(attrs->__pod_cpumask), @@ -4825,6 +4858,8 @@ static u32 wqattrs_hash(const struct workqueue_attrs = *attrs) static bool wqattrs_equal(const struct workqueue_attrs *a, const struct workqueue_attrs *b) { + if (a->prio !=3D b->prio) + return false; if (a->nice !=3D b->nice) return false; if (a->affn_strict !=3D b->affn_strict) @@ -5601,11 +5636,17 @@ static void unbound_wq_update_pwq(struct workqueue_= struct *wq, int cpu) =20 static int alloc_and_link_pwqs(struct workqueue_struct *wq) { - bool highpri =3D wq->flags & WQ_HIGHPRI; - int cpu, ret; + int prio, cpu, ret; =20 lockdep_assert_held(&wq_pool_mutex); =20 + if (wq->flags & WQ_RTPRI) + prio =3D WQ_PRIO_RT; + else if (wq->flags & WQ_HIGHPRI) + prio =3D WQ_PRIO_HIGH; + else + prio =3D WQ_PRIO_NORMAL; + wq->cpu_pwq =3D alloc_percpu(struct pool_workqueue *); if (!wq->cpu_pwq) goto enomem; @@ -5622,7 +5663,7 @@ static int alloc_and_link_pwqs(struct workqueue_struc= t *wq) struct pool_workqueue **pwq_p; struct worker_pool *pool; =20 - pool =3D &(per_cpu_ptr(pools, cpu)[highpri]); + pool =3D &(per_cpu_ptr(pools, cpu)[prio]); pwq_p =3D per_cpu_ptr(wq->cpu_pwq, cpu); =20 *pwq_p =3D kmem_cache_alloc_node(pwq_cache, GFP_KERNEL, @@ -5642,14 +5683,14 @@ static int alloc_and_link_pwqs(struct workqueue_str= uct *wq) if (wq->flags & __WQ_ORDERED) { struct pool_workqueue *dfl_pwq; =20 - ret =3D apply_workqueue_attrs_locked(wq, ordered_wq_attrs[highpri]); + ret =3D apply_workqueue_attrs_locked(wq, ordered_wq_attrs[prio]); /* there should only be single pwq for ordering guarantee */ dfl_pwq =3D rcu_access_pointer(wq->dfl_pwq); WARN(!ret && (wq->pwqs.next !=3D &dfl_pwq->pwqs_node || wq->pwqs.prev !=3D &dfl_pwq->pwqs_node), "ordering guarantee broken for workqueue %s\n", wq->name); } else { - ret =3D apply_workqueue_attrs_locked(wq, unbound_std_wq_attrs[highpri]); + ret =3D apply_workqueue_attrs_locked(wq, unbound_std_wq_attrs[prio]); } =20 if (ret) @@ -5814,6 +5855,12 @@ static struct workqueue_struct *__alloc_workqueue(co= nst char *fmt, return NULL; } =20 + if (flags & WQ_RTPRI) { + if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) !=3D + WQ_UNBOUND)) + return NULL; + } + /* see the comment above the definition of WQ_POWER_EFFICIENT */ if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient) flags =3D (flags & ~WQ_PERCPU) | WQ_UNBOUND; @@ -5857,7 +5904,17 @@ static struct workqueue_struct *__alloc_workqueue(co= nst char *fmt, flags &=3D ~WQ_PERCPU; } =20 - if (flags & WQ_BH) { + if (flags & WQ_RTPRI) { + /* + * RT workqueues are limited to max half of possible CPUs to + * avoid scheduling starvation and have strict CPU affinity for + * low latency execution. + */ + max_active =3D min_t(int, max_active, + DIV_ROUND_UP(num_possible_cpus(), 2)); + wq->unbound_attrs->affn_scope =3D WQ_AFFN_CPU; + wq->unbound_attrs->affn_strict =3D true; + } else if (flags & WQ_BH) { /* * BH workqueues always share a single execution context per CPU * and don't impose any max_active limit. @@ -6359,17 +6416,26 @@ void print_worker_info(const char *log_lvl, struct = task_struct *task) } } =20 +static void pr_cont_bh_suffix(enum wq_priority prio) +{ + if (prio =3D=3D WQ_PRIO_RT) + pr_cont("bh-rt"); + else if (prio =3D=3D WQ_PRIO_HIGH) + pr_cont("bh-hi"); + else + pr_cont("bh"); +} + static void pr_cont_pool_info(struct worker_pool *pool) { pr_cont(" cpus=3D%*pbl", nr_cpumask_bits, pool->attrs->cpumask); if (pool->node !=3D NUMA_NO_NODE) pr_cont(" node=3D%d", pool->node); - pr_cont(" flags=3D0x%x", pool->flags); + pr_cont(" flags=3D0x%x ", pool->flags); if (pool->flags & POOL_BH) - pr_cont(" bh%s", - pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pr_cont_bh_suffix(pool->attrs->prio); else - pr_cont(" nice=3D%d", pool->attrs->nice); + pr_cont("prio=3D%d", pool->attrs->prio); } =20 static void pr_cont_worker_id(struct worker *worker) @@ -6377,8 +6443,7 @@ static void pr_cont_worker_id(struct worker *worker) struct worker_pool *pool =3D worker->pool; =20 if (pool->flags & POOL_BH) - pr_cont("bh%s", - pool->attrs->nice =3D=3D HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pr_cont_bh_suffix(pool->attrs->prio); else pr_cont("%d%s", task_pid_nr(worker->task), worker->rescue_wq ? "(RESCUER)" : ""); @@ -7302,7 +7367,11 @@ static ssize_t wq_nice_show(struct device *dev, stru= ct device_attribute *attr, int written; =20 mutex_lock(&wq->mutex); - written =3D scnprintf(buf, PAGE_SIZE, "%d\n", wq->unbound_attrs->nice); + if (wq->unbound_attrs->prio !=3D WQ_PRIO_RT) + written =3D scnprintf(buf, PAGE_SIZE, "%d\n", + wq->unbound_attrs->nice); + else + written =3D -EINVAL; mutex_unlock(&wq->mutex); =20 return written; @@ -7328,13 +7397,20 @@ static ssize_t wq_nice_store(struct device *dev, st= ruct device_attribute *attr, { struct workqueue_struct *wq =3D dev_to_wq(dev); struct workqueue_attrs *attrs; - int ret =3D -ENOMEM; + int ret; =20 mutex_lock(&wq_pool_mutex); =20 attrs =3D wq_sysfs_prep_attrs(wq); - if (!attrs) + if (!attrs) { + ret =3D -ENOMEM; goto out_unlock; + } + + if (attrs->prio =3D=3D WQ_PRIO_RT) { + ret =3D -EINVAL; + goto out_unlock; + } =20 if (sscanf(buf, "%d", &attrs->nice) =3D=3D 1 && attrs->nice >=3D MIN_NICE && attrs->nice <=3D MAX_NICE) @@ -7942,12 +8018,13 @@ static void __init restrict_unbound_cpumask(const c= har *name, const struct cpuma cpumask_and(wq_unbound_cpumask, wq_unbound_cpumask, mask); } =20 -static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu,= int nice) +static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu,= enum wq_priority prio, int nice) { BUG_ON(init_worker_pool(pool)); pool->cpu =3D cpu; cpumask_copy(pool->attrs->cpumask, cpumask_of(cpu)); cpumask_copy(pool->attrs->__pod_cpumask, cpumask_of(cpu)); + pool->attrs->prio =3D prio; pool->attrs->nice =3D nice; pool->attrs->affn_strict =3D true; pool->node =3D cpu_to_node(cpu); @@ -7971,7 +8048,8 @@ static void __init init_cpu_worker_pool(struct worker= _pool *pool, int cpu, int n void __init workqueue_init_early(void) { struct wq_pod_type *pt =3D &wq_pod_types[WQ_AFFN_SYSTEM]; - int std_nice[NR_STD_WORKER_POOLS] =3D { 0, HIGHPRI_NICE_LEVEL }; + int std_prio[NR_STD_WORKER_POOLS + 1] =3D { 0, WQ_PRIO_HIGH, WQ_PRIO_RT }; + int std_nice[NR_STD_WORKER_POOLS + 1] =3D { 0, HIGHPRI_NICE_LEVEL, 0 }; void (*irq_work_fns[NR_STD_WORKER_POOLS])(struct irq_work *) =3D { bh_pool_kick_normal, bh_pool_kick_highpri }; int i, cpu; @@ -8023,22 +8101,25 @@ void __init workqueue_init_early(void) =20 i =3D 0; for_each_bh_worker_pool(pool, cpu) { - init_cpu_worker_pool(pool, cpu, std_nice[i]); + init_cpu_worker_pool(pool, cpu, std_prio[i], std_nice[i]); pool->flags |=3D POOL_BH; init_irq_work(bh_pool_irq_work(pool), irq_work_fns[i]); i++; } =20 i =3D 0; - for_each_cpu_worker_pool(pool, cpu) - init_cpu_worker_pool(pool, cpu, std_nice[i++]); + for_each_cpu_worker_pool(pool, cpu) { + init_cpu_worker_pool(pool, cpu, std_prio[i], std_nice[i]); + i++; + } } =20 /* create default unbound and ordered wq attrs */ - for (i =3D 0; i < NR_STD_WORKER_POOLS; i++) { + for (i =3D 0; i < NR_STD_WORKER_POOLS + 1; i++) { struct workqueue_attrs *attrs; =20 BUG_ON(!(attrs =3D alloc_workqueue_attrs())); + attrs->prio =3D std_prio[i]; attrs->nice =3D std_nice[i]; unbound_std_wq_attrs[i] =3D attrs; =20 @@ -8047,6 +8128,7 @@ void __init workqueue_init_early(void) * guaranteed by max_active which is enforced by pwqs. */ BUG_ON(!(attrs =3D alloc_workqueue_attrs())); + attrs->prio =3D std_prio[i]; attrs->nice =3D std_nice[i]; attrs->ordered =3D true; ordered_wq_attrs[i] =3D attrs; --=20 2.54.0 From nobody Sat Jul 25 05:29:26 2026 Received: from fanzine2.igalia.com (fanzine2.igalia.com [213.97.179.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1D9FA322533 for ; Fri, 17 Jul 2026 09:02:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=213.97.179.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784278948; cv=none; b=sRMVseKapx4bI1E7pb1WSoc/c2fDJQoeL9aYpz043vN4/IEXTFux4b0jLk57Eb76v+InQrmC+Yh8brzvtb61XvKrkUbYqX9Zubey8fTHCJaqgjxA0JTZceB8pOjbUGn1xZ7N3KhfqbASNV2gb4FTEkYdYFH2E1fWkSK/IVbfqJI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784278948; c=relaxed/simple; bh=WX8ToMMxjMLGCftSpXYZKi00G6B/c6fSZjSq5BJkZ3o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JhIE2x/zzaL4St+VfygNbxZPudJ4F6CyMrmh8U/b/cLzT4CBQCJt8xTpLhxnuU2AT73oEH5Ayjde2E0lgRmdyJBAAnb/aRj/q85LvPnO8BaSARYvOw4gddi9JIzw5Pf2Xh6Mg6eLvPv1zD6Ki+d19jGhzDxxBLXH6bKRRWiq2YM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=igalia.com; spf=pass smtp.mailfrom=igalia.com; dkim=pass (2048-bit key) header.d=igalia.com header.i=@igalia.com header.b=nB5TXsZ3; arc=none smtp.client-ip=213.97.179.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=igalia.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=igalia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=igalia.com header.i=@igalia.com header.b="nB5TXsZ3" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=igalia.com; s=20170329; h=Content-Transfer-Encoding:MIME-Version:Message-ID:Date:Subject: Cc:To:From:From:Reply-To; bh=Dr07AZPFq9i5ZADmfRFGp2kFen4Vgevol2vCzyBZcb8=; b= nB5TXsZ3jBYgi8/epPWL+5Ps7YfvePP/oEmvRT92FAMRnRHn+EkX4YuOjiFGeVyhqT01EqGDP+56k zcrA/l3K7tP0skuDPMJ5CCUaFjCLXmNNaPn7KbikQMTe5dp+BdxO7tasHZQt0tKthCyBzpJWOpV2e mh6UQ8iJWHtWStVTR9W4gWk0dhAsVcmHFf+qaaPZmwqFuP0xva3SoxW8tQnq91XsFh0cwZGZj+PZ9 dZBKGuKHUvoK1nP4k5pJvKcnJLzqzFxLXIY1LmZ8LpiclOcS+8iemaWxamWstM+fFAk5ix7XCFmM0 HMgCgs2LN2w4jtvB/okmamYz86LU0MWi4w==; Received: from [90.240.106.137] (helo=localhost) by fanzine2.igalia.com with esmtpsa (Cipher TLS1.3:ECDHE_SECP256R1__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim) id 1wkeST-00GJcU-8v; Fri, 17 Jul 2026 11:02:13 +0200 From: Tvrtko Ursulin To: dri-devel@lists.freedesktop.org Cc: Boris Brezillon , Steven Price , Liviu Dudau , Chia-I Wu , Matthew Brost , kernel-dev@igalia.com, linux-kernel@vger.kernel.org, Tvrtko Ursulin , Chia-I Wu , Tejun Heo Subject: [RFC v3 2/2] drm/panthor: Create per queue priority workqueues Date: Fri, 17 Jul 2026 10:02:09 +0100 Message-ID: <20260717090209.26931-3-tvrtko.ursulin@igalia.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260717090209.26931-1-tvrtko.ursulin@igalia.com> References: <20260717090209.26931-1-tvrtko.ursulin@igalia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Split the single workqueue shared between the driver internal logic and DRM scheduler use into separate ones, where the DRM scheduler one is created per GPU priority level using the appropriate mapping to workqueue priorities. Low and medium GPU priority are served by a normal workqueue, high is server by a WQ_HIGHPRI instance, while realtime GPU priority is using the newly added WQ_RTPRI flag for lowest possible latency. These workqueues are device global and for all three we set the maximum concurrency to two in order to keep the GPU optimally fed with work. Signed-off-by: Tvrtko Ursulin Cc: Boris Brezillon Cc: Chia-I Wu Cc: Liviu Dudau Cc: Matthew Brost Cc: Steven Price Cc: Tejun Heo --- drivers/gpu/drm/panthor/panthor_sched.c | 38 +++++++++++++++++++++---- 1 file changed, 33 insertions(+), 5 deletions(-) diff --git a/drivers/gpu/drm/panthor/panthor_sched.c b/drivers/gpu/drm/pant= hor/panthor_sched.c index 5832dccfc093..70e84ebb6c66 100644 --- a/drivers/gpu/drm/panthor/panthor_sched.c +++ b/drivers/gpu/drm/panthor/panthor_sched.c @@ -152,11 +152,18 @@ struct panthor_scheduler { * * Used for the scheduler tick, group update or other kind of FW * event processing that can't be handled in the threaded interrupt - * path. Also passed to the drm_gpu_scheduler instances embedded - * in panthor_queue. + * path. */ struct workqueue_struct *wq; =20 + /** + * @submit_wq: Per priority workqueues for the DRM scheduler + * + * Passed to the drm_gpu_scheduler instances embedded + * in panthor_queue based on the queue priority. + */ + struct workqueue_struct *submit_wq[PANTHOR_CSG_PRIORITY_COUNT]; + /** * @heap_alloc_wq: Workqueue used to schedule tiler_oom works. * @@ -3500,7 +3507,6 @@ group_create_queue(struct panthor_group *group, { struct drm_sched_init_args sched_args =3D { .ops =3D &panthor_queue_sched_ops, - .submit_wq =3D group->ptdev->scheduler->wq, /* * The credit limit argument tells us the total number of * instructions across all CS slots in the ringbuffer, with @@ -3593,8 +3599,14 @@ group_create_queue(struct panthor_group *group, goto err_free_queue; } =20 + if (group->priority >=3D ARRAY_SIZE(group->ptdev->scheduler->submit_wq) || + !group->ptdev->scheduler->submit_wq[group->priority]) { + ret =3D -EINVAL; + goto err_free_queue; + } + sched_args.name =3D queue->name; - + sched_args.submit_wq =3D group->ptdev->scheduler->submit_wq[group->priori= ty]; ret =3D drm_sched_init(&queue->scheduler, &sched_args); if (ret) goto err_free_queue; @@ -4084,6 +4096,15 @@ static void panthor_sched_fini(struct drm_device *dd= ev, void *res) if (!sched || !sched->csg_slot_count) return; =20 + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]); + + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]); + + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]); + if (sched->wq) destroy_workqueue(sched->wq); =20 @@ -4185,7 +4206,14 @@ int panthor_sched_init(struct panthor_device *ptdev) */ sched->heap_alloc_wq =3D alloc_workqueue("panthor-heap-alloc", WQ_UNBOUND= , 0); sched->wq =3D alloc_workqueue("panthor-csf-sched", WQ_MEM_RECLAIM | WQ_UN= BOUND, 0); - if (!sched->wq || !sched->heap_alloc_wq) { + sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] =3D alloc_workqueue("pantho= r-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] =3D sched->submit_wq[PANTHOR_C= SG_PRIORITY_MEDIUM]; + sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] =3D alloc_workqueue("panthor-= drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] =3D alloc_workqueue("panthor-dr= m-rt", WQ_RTPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + if (!sched->wq || !sched->heap_alloc_wq || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) { panthor_sched_fini(&ptdev->base, sched); drm_err(&ptdev->base, "Failed to allocate the workqueues"); return -ENOMEM; --=20 2.54.0