From nobody Fri Sep 25 00:41:18 2026 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1DF7B3A4F3B for ; Fri, 18 Sep 2026 14:25:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741550; cv=none; b=AxoFpSFpHzxOVIQf+NC7je/69aRkOTRSSA5CJhP/IUbNCqO0SN/XS+nfSwDzOg5idEH17SUxALRHy5ph+OLzfrxiEnYsXi7aSO7cPY5xDSyoHcROmpnkCl0w/JeJAhT/XaG6I7Z+iedNceCVkNgLimuY+iBPxDAFNL0Fc75V2go= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741550; c=relaxed/simple; bh=5Hwqbfy8RdpERP0rf8VELLjud0A0y8xSKq9Vf04oooQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=uIwLrWRwvsX3jequEqtwHkg45B4KMhktmNKbKHbE7mHOgkcAOkVjnhYB3/30GOOHIQVO8zMNdROGp0fD8z16Nx58O6P0TPqMMPvyZMOg1+AQqtDciK51+58ux7S5+dMSA9B7G5IZqDy4G4d7pglx9Huf4eUW54JTYSgm2xokLM4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=hcsmyg4o; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="hcsmyg4o" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=4yh9QPWkCyce8lBJo2JXbcB5jmBdkfD2PJs0cugwfys=; b=hcsmyg4oaYWZtOq+JKE+mTGVv+ ACXBUZzCTrknpdkMKX+GsWP1FDdrGrsvgO/D3ttVeMHQ7xPhXtRWrbTnYwy4/qASg5drwXlXu+cYY FzrnsHjQsEbRGQ3y/efYlFUnA26tSs4YFTV+ooEpoKpjonWS/oOnWPsfi2TN3ej0J3mMbTGVKySXE peEs8PHm6XMeArJ6CIq18Q1blS+GOlMzncjMgSG9TYCQvTLo3goaaWgC0mHdiZCqQml4F1LJkBPAw ClVS4VaeuQO7oWF80duRI+L5+Rz8y7qrw3O/UxkNMSzawzVSJbdXd2XsrYwjdv/Llu6DN9rf2Xu4c 7DxpWlyQ==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1x7ZX5-006n91-0o; Fri, 18 Sep 2026 14:25:43 +0000 From: Breno Leitao Date: Fri, 18 Sep 2026 07:25:33 -0700 Subject: [PATCH RFC 1/3] workqueue: Maintain both max_active limits Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260918-wq_final-v1-1-5c43c08a26bc@debian.org> References: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> In-Reply-To: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> To: Tejun Heo , Lai Jiangshan Cc: linux-kernel@vger.kernel.org, marco.crivellari@suse.com, Breno Leitao , kernel-team@meta.com X-Mailer: b4 0.16-dev-f8e9d X-Developer-Signature: v=1; a=openpgp-sha256; l=3696; i=leitao@debian.org; h=from:subject:message-id; bh=5Hwqbfy8RdpERP0rf8VELLjud0A0y8xSKq9Vf04oooQ=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqrUng5+RGpg3uGcDqpWjdd+L+VaMHTAZTQ4mwh XaY2Xqtp6CJAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCaq1J4AAKCRA1o5Of/Hh3 bU/BD/4gDeshxZm/8qYUboPQzr/rQjSQB966JZibu2EW7dh+WqdBVS5ylUp9sqIVYA9F1PJOUbJ llnDK4pPDM4CslUYuQ50al5Eu6wdnBe8MYq32jrioT7fQ0vbiLtvzOPXgcvcqcXg1szFDctAES9 i8sc1u4XnVRmVcChOXbbddnI0KpIFlK/bDh5ZrOdTbw1tXKvg3Mf82DO3JH33xo038ruyahE+KL hav0BY8BrZMOkGcghK4AgDfhxX+kWAmHyJlSCDWzt46D8B2CT7NLP9ZMpGhJ69/iJVrWoqi4JjJ 2OEtEqo//88pvXlFJn+eRvc7xURGl1rYDfCgx+t8sA3bjcHS3ethCKQ9bKCUNe2ckyoTurMR1oP iBoMb0t8UCiDpL0oHybwvM4UsjJ8/UyKGIwLk5l7Lq0v9qhwPXF60hEHuTGSJ0Pn2mT9bCbIREB y79Mt90cTQwhzVrtsa6wSqVZbEdxOKMpwUmaEf5hvqG4LOC3Si76jeSyGHcq5Is+CIOqmxlTTFl rNm2TgcVlxtkgvG7sRa8cdb45X+grK3bs+S7cTyTvbV8OPiPu6vQ1OJIcfdg5yMailZv2v5ErtB PEMeMOWpb4WL0bhW++606F6EdqXWLbk/zYXpzniGUmGGVVc/7x6azH5F+5qUGbFelg3V7LcJadJ PmUM4QiwOG10HJQ== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao I've done commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own max_active"), but, later I found the solution was not that simple. I want to keep only one active, either percpu_max_active or max_active. But there is no instant at which the backend flips, apply_wqattrs_commit() swaps the slots one CPU at a time: for_each_possible_cpu(cpu) ctx->pwq_tbl[cpu] =3D install_pwq(ctx->wq, cpu, ctx->pwq_tbl[cpu]); __queue_work() reads a slot under rcu_read_lock() plus pool->lock, it never takes wq->mutex. So partway through that loop, CPU 0 already has a concurrency-managed pwq while CPU 7 still has a pod-backed one, and work can be queued to either. So, let's keep both values the same, which is silly for now. Maybe we should revert commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own max_active"). Not convinced yet. Link: https://lore.kernel.org/all/amESSqf0TMmzhFGz@slm.duckdns.org/ Signed-off-by: Breno Leitao --- kernel/workqueue.c | 39 ++++++++++++++++++--------------------- 1 file changed, 18 insertions(+), 21 deletions(-) diff --git a/kernel/workqueue.c b/kernel/workqueue.c index e618108c6127da..1c4f8bdd1cd509 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -6055,24 +6055,25 @@ static void wq_adjust_max_active(struct workqueue_s= truct *wq) } =20 /* - * Update the limit and then kick inactive work items if more active + * Update both limits and then kick inactive work items if more active * work items are allowed. This doesn't break work item ordering * because new work items are always queued behind existing inactive * work items if there are any. + * + * Which one a pwq honours follows its pool, see pwq_tryinc_nr_active(). + * Keeping both current means a pwq is never metered against a limit + * that was never set. */ - if (wq->flags & WQ_UNBOUND) { - if (wq->max_active =3D=3D new_max && wq->min_active =3D=3D new_min) - return; + if (wq->max_active =3D=3D new_max && wq->min_active =3D=3D new_min && + wq->percpu_max_active =3D=3D new_max) + return; =20 - WRITE_ONCE(wq->max_active, new_max); - WRITE_ONCE(wq->min_active, new_min); - wq_update_node_max_active(wq, -1); - } else { - if (wq->percpu_max_active =3D=3D new_max) - return; + WRITE_ONCE(wq->max_active, new_max); + WRITE_ONCE(wq->min_active, new_min); + WRITE_ONCE(wq->percpu_max_active, new_max); =20 - WRITE_ONCE(wq->percpu_max_active, new_max); - } + if (wq->flags & WQ_UNBOUND) + wq_update_node_max_active(wq, -1); =20 if (new_max =3D=3D 0) return; @@ -6169,14 +6170,11 @@ static struct workqueue_struct *__alloc_workqueue(c= onst char *fmt, =20 /* init wq */ wq->flags =3D flags; - if (flags & WQ_UNBOUND) { - wq->max_active =3D max_active; - wq->min_active =3D min(max_active, WQ_DFL_MIN_ACTIVE); - wq->saved_min_active =3D wq->min_active; - } else { - wq->percpu_max_active =3D max_active; - } + wq->max_active =3D max_active; + wq->min_active =3D min(max_active, WQ_DFL_MIN_ACTIVE); + wq->percpu_max_active =3D max_active; wq->saved_max_active =3D max_active; + wq->saved_min_active =3D wq->min_active; mutex_init(&wq->mutex); atomic_set(&wq->nr_pwqs_to_flush, 0); INIT_LIST_HEAD(&wq->pwqs); @@ -6452,8 +6450,7 @@ void workqueue_set_max_active(struct workqueue_struct= *wq, int max_active) mutex_lock(&wq->mutex); =20 wq->saved_max_active =3D max_active; - if (wq->flags & WQ_UNBOUND) - wq->saved_min_active =3D min(wq->saved_min_active, max_active); + wq->saved_min_active =3D min(wq->saved_min_active, max_active); =20 wq_adjust_max_active(wq); =20 --=20 2.53.0-Meta From nobody Fri Sep 25 00:41:18 2026 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 256993AD53F for ; Fri, 18 Sep 2026 14:25:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741553; cv=none; b=Qf2zM1QcTm6mFqMMSXiYuzWETobs1n4DJKWiPlePUcC7LgsLhr2NAC/brWLcxm9YnRSnmRR632PixdVFSgQj3VxTqJwkgD4lxNjtkbBdstH6Q1Jl8vhsYmlyzlhZ8n67elTod6yttDH323kdKEiIv/Kjj8FzrMi/kTmm9TcW0YI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741553; c=relaxed/simple; bh=X5RcxNz2QGiWT9Zse8Ys3prPWoakI8lQ0XS1HkPc+4E=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Bd4ZtwxWUyE5huBwKZEu8Udt4HC3bvFnrzsLlTppEk/OERm0Qx6U7YkQOwbopSzMdkdE5B8qVPtT9eryRjPD5Gd2NbYvO+kiwJcRhpitbt9FfVgVIKclcho5QC/7/d3r5Lo09I4wYZuG8fe7oDgQBEoSygoRe9OMpzbMFtk8vZg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=QkO32hVj; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="QkO32hVj" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=rt8Z/CWD/C96rMtWvjjrIBWSSM8KoD63SSMFo3bWtps=; b=QkO32hVjfc+/G0/g/NXGebIVc7 ezbPXutfyVjTWSZCVuoskfL4x7rtcyTPEao46jt81Cn0r5gaEY8rRWlq2p+ysVCp2p0L5kuzFkJGI d9B2h2Ev3CIzH/bN5fpanZ8uL7sQdewrGdX3Bi2+pv2egzKphjMPHR+SiLmYW3wwune91ojgORVAy FtHYqEFgxg04fZur6vRIvOsC7AKE8m6Zs2Yy2WXknZgcsmuy10+sg0FdqNOuxDrkgqZlcsmV+2Efi dNhqOrc1lNr/6vFrTFUBKAdV4kz7DcMH8wDmjMpSfaiZhbO9vZuox7lXHaX2rb7l6G5yaxfw973Pl pGpi0ylA==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1x7ZX9-006n9A-0n; Fri, 18 Sep 2026 14:25:47 +0000 From: Breno Leitao Date: Fri, 18 Sep 2026 07:25:34 -0700 Subject: [PATCH RFC 2/3] workqueue: Add a concurrency_managed workqueue attribute Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260918-wq_final-v1-2-5c43c08a26bc@debian.org> References: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> In-Reply-To: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> To: Tejun Heo , Lai Jiangshan Cc: linux-kernel@vger.kernel.org, marco.crivellari@suse.com, Breno Leitao , kernel-team@meta.com X-Mailer: b4 0.16-dev-f8e9d X-Developer-Signature: v=1; a=openpgp-sha256; l=4395; i=leitao@debian.org; h=from:subject:message-id; bh=X5RcxNz2QGiWT9Zse8Ys3prPWoakI8lQ0XS1HkPc+4E=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqrUngwJLoy6f396TtUgvc3oqu7FT/cXB7Cc+rp Uo+UjJ+cGeJAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCaq1J4AAKCRA1o5Of/Hh3 bVXrD/wMpThM6/OzJybgaam5GIPFVP2rZsRosvoUidSeko17zVp0tKPIxn/Jfu6eOf8aTnrTevE uiIg7XWDtSh1VQkMXVk1F1T8shbxUuBtJKs7UqNMGj5FfcyDtad0IQ4k9rGqF0eq0xJGL12YaOt MifTHji1PUj8QFB7iSPq6uSyyT3k6ufliT6UpPFEmqP4cwAAbljrZBMgOV6CHj6I8+Iuz0SBQ5s 7bR88wSQN2WPc+twLcIqnzYGjXVqbjs51im1cnoF53OFEuACJG0I7hsz1F+Rj8kq2BZLs/zK8q7 VvTP7j9z48AkQ/yDf5F0nPCG4/4ZoARJPhaMLTrTRyRzM3OCQlO/g+VDk+CBDQqGag9UgWNxNgH nzgnqzY6yWfDdvHqgw8BXR+xW7ZVEHGfyt1ozzfSsfzSaOXdBPu9TbmB8BrPmMLDB+VInLg40uw v/ldxBeQJoP7Jo7HBUJC9b17MUYi/FJp56rIkenPi1/WvU12BnJdNdX8dnoK7ED/lamMdVLCXtV i39EN8S+02u7+j1dQxX6hD1Ao2k14RHuZT1zVAlDnlqeeG0YC+nD+ggmdbYS/a3zzfFLG36cW2c LMhdS1R+HtgJclN/wtYDzNOeh0ITSATokNNqHqlhORCjj9Ls2KBlSEyq228jO8LUo6QKESWCykm qwnqZW8samPdsAg== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao Tejun wanted to have a CM affinity (WQ_AFFN_CPU_CM), but I am proposing to have an attribute field, that we can set with WQ_AFFN_CPU, depending on whether we want to have CM or not. Create workqueue_attrs->concurrency_managed, and use it to decide if we do CM. It always comes with WQ_AFFN_CPU and affn_strict, since alloc_wq_std_attrs() sets the three together for a WQ_PERCPU workqueue, which is the only thing that sets it. It sits below the divider, so wqattrs_clear_for_pool() clears it and it does not become part of the pool hash. This is not set by sysfs (for now?!), and apply_workqueue_attrs() refuses attrs that carry it, so it is fixed when the workqueue is created, for now. So, a workqueue has 4 "attributes" with this patch: * cpumask - which CPUs the workqueue may use at all * affn_scope - cpu / smt / cache / cache_shard / numa / system * affn_strict - whether the pod boundary is a hard bind or a hint * concurrency_managed - CM enabled or now. Signed-off-by: Breno Leitao --- include/linux/workqueue.h | 10 ++++++++++ kernel/workqueue.c | 19 ++++++++++++++++--- 2 files changed, 26 insertions(+), 3 deletions(-) diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h index a283766a192aaf..a7fc9b0f72af2e 100644 --- a/include/linux/workqueue.h +++ b/include/linux/workqueue.h @@ -204,6 +204,16 @@ struct workqueue_attrs { */ enum wq_affn_scope affn_scope; =20 + /** + * @concurrency_managed: use the concurrency managed per-cpu pools + * + * Those keep at most one worker running per CPU and account max_active + * per CPU rather than per node. Set from %WQ_PERCPU when the workqueue + * is created and fixed for its lifetime, so it always comes with + * %WQ_AFFN_CPU and @affn_strict. + */ + bool concurrency_managed; + /** * @ordered: work items must be executed one by one in queueing order */ diff --git a/kernel/workqueue.c b/kernel/workqueue.c index 1c4f8bdd1cd509..f86e8d0770ca22 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -5026,6 +5026,7 @@ static void copy_workqueue_attrs(struct workqueue_att= rs *to, * get_unbound_pool() explicitly clears the fields. */ to->affn_scope =3D from->affn_scope; + to->concurrency_managed =3D from->concurrency_managed; to->ordered =3D from->ordered; } =20 @@ -5036,6 +5037,7 @@ static void copy_workqueue_attrs(struct workqueue_att= rs *to, static void wqattrs_clear_for_pool(struct workqueue_attrs *attrs) { attrs->affn_scope =3D WQ_AFFN_NR_TYPES; + attrs->concurrency_managed =3D false; attrs->ordered =3D false; if (attrs->affn_strict) cpumask_copy(attrs->cpumask, cpu_possible_mask); @@ -5607,9 +5609,9 @@ static struct pool_workqueue *alloc_pwq(struct workqu= eue_struct *wq, =20 lockdep_assert_held(&wq_pool_mutex); =20 - WARN_ON_ONCE((wq->flags & WQ_PERCPU) && cpu < 0); + WARN_ON_ONCE(attrs->concurrency_managed && cpu < 0); =20 - if (cpu >=3D 0 && (wq->flags & WQ_PERCPU)) { + if (cpu >=3D 0 && attrs->concurrency_managed) { pool =3D get_percpu_pool(wq, cpu); } else { pool =3D get_unbound_pool(attrs); @@ -5737,7 +5739,7 @@ apply_wqattrs_prepare(struct workqueue_struct *wq, copy_workqueue_attrs(new_attrs, attrs); wqattrs_actualize_cpumask(new_attrs, unbound_cpumask); cpumask_copy(new_attrs->__pod_cpumask, new_attrs->cpumask); - if (!(wq->flags & WQ_PERCPU)) { + if (!new_attrs->concurrency_managed) { ctx->dfl_pwq =3D alloc_unbound_pwq(wq, new_attrs); if (!ctx->dfl_pwq) goto out_free; @@ -5843,6 +5845,10 @@ int apply_workqueue_attrs(struct workqueue_struct *w= q, if (WARN_ON(!(wq->flags & WQ_UNBOUND))) return -EINVAL; =20 + /* concurrency management comes from WQ_PERCPU, it is not applied */ + if (WARN_ON(attrs->concurrency_managed)) + return -EINVAL; + mutex_lock(&wq_pool_mutex); ret =3D apply_workqueue_attrs_locked(wq, attrs); mutex_unlock(&wq_pool_mutex); @@ -5934,6 +5940,13 @@ static struct workqueue_attrs *alloc_wq_std_attrs(st= ruct workqueue_struct *wq) if (wq->flags & __WQ_ORDERED) attrs->ordered =3D true; =20 + /* a percpu workqueue wants a concurrency managed pwq on every CPU */ + if (wq->flags & WQ_PERCPU) { + attrs->affn_scope =3D WQ_AFFN_CPU; + attrs->affn_strict =3D true; + attrs->concurrency_managed =3D true; + } + return attrs; } =20 --=20 2.53.0-Meta From nobody Fri Sep 25 00:41:18 2026 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D7EC23876A1 for ; Fri, 18 Sep 2026 14:25:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741558; cv=none; b=Ct+1FgF3bkHtPOohs54RFFRrYMObmjA598e3scBim/HLwDbH7bjsf1HD3L6qRH143NxX5ELG8V80l0AGzD47hFlW5+s8IzY9OARjG++IhjKyPNqZ9JiSZlIe1M4EudbB0706ihhx++NQSQy5uozkkSUNolct7VLiK2hciJ2s4ic= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789741558; c=relaxed/simple; bh=iTP33CeJKSOUaA/Zn+hMtqQTpJaAZgbgUaQmX+iMFoU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=o2PrQyP38neNBZJGaWo5sbHm0sXKnCEa+6t57NCKcrxs+MuKVCbd4YfgM9EVEpdzavM5TdrP4VS9xzhJhZbTTmDID1UQKU4TTyx83ma32eX3HE3paR1M8cS38uv8EUWnVjTVR0ZyQFUUnRlUK4preoufsQkcik5enkrMXRhl/6s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=s2UA+mvK; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="s2UA+mvK" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=6NMW/Nho9XykeGyvUq9sJtzumS4a/caq3j8CGCoh14c=; b=s2UA+mvK+RbpxcInGx+KuYOnxX +R8AhDRa5YnNcPygQQN50E/QqegF4AcIYeko+OdscXdP6v1rVLfs3RhVAtIYh95YR+4mYozCnZ61u cOcGg+ZObm9oWn5f0kSPGsD0yw7VC022Exh/x36zJ2n3AKvlh1vvWp+e9DxGtAGkwKebt3BQs84mH kF0oFH5S+x9KuWOGwqPHPyooMCvSMmfFoyFIJtHgj66f1AqqYLHQe01xdTJ12SbMyKtB0mFF+MVji QKBeByIIWnL9zyjqEbFqPHO7bffQ0DoQXLXkf70nzcgpqI9cYL6QxEN6wEm7I922hQdkqSAIqo5Gb ef9FdhiQ==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1x7ZXC-006n9S-1v; Fri, 18 Sep 2026 14:25:50 +0000 From: Breno Leitao Date: Fri, 18 Sep 2026 07:25:35 -0700 Subject: [PATCH RFC 3/3] workqueue: Back every workqueue with the unbound machinery Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260918-wq_final-v1-3-5c43c08a26bc@debian.org> References: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> In-Reply-To: <20260918-wq_final-v1-0-5c43c08a26bc@debian.org> To: Tejun Heo , Lai Jiangshan Cc: linux-kernel@vger.kernel.org, marco.crivellari@suse.com, Breno Leitao , kernel-team@meta.com X-Mailer: b4 0.16-dev-f8e9d X-Developer-Signature: v=1; a=openpgp-sha256; l=7438; i=leitao@debian.org; h=from:subject:message-id; bh=iTP33CeJKSOUaA/Zn+hMtqQTpJaAZgbgUaQmX+iMFoU=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqrUngbNYbvrcs+33qNAzq0PcOOWAQFRoExrebk srJHLB02u+JAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCaq1J4AAKCRA1o5Of/Hh3 bdVaD/4zMPzA/UkNQX+nrxJVMSFOYqp/gD7D58glzMD21ebpqS3MgQiPU6UzlQqaEKFjgtp6K4X 8zA2lFTx+gt2Hqmz6t3GjhyKde/+j2+MW5pJE/dVZ1Z3rsN5VJ+rd3nuBhvBouMDtGMgxe94ID7 5aZltYbCa+19m8/y78Nr+lvP9EwHHy+31Nq8ahXk+JGTAiKP8MABun/CkfqBJOTrPerBa9oVACe vfNs5iUFf4x1Bki8wcYlzL7LEEWlckoAaHdkyE6i2ZCFIFTt691yDOAiV9WcDG6SUG5ZzTiasvH U3JWR5dU6jDsln74+wK7EP4uOYGoEhqn/UtIZ3r/MNrIXcjTAIR0CkGj1qsQf7P9mjeQayw90tL YyUbZJ51cMNLXMSpqc1HY67I04rdYxlIZO67b9H8uLSONzgbmtRQWSQTwARu3JUGfg/dKc3FSYe W+zv/RbhM55+MjDrn9FmJRlcSyenjT3YgGCQzuzkeY0fh2yn6oG46qQKAeuflTxxFu4TOpDcoiM 3Atj0TyXTpAcGCtQrGQrJmHkNQEwXozR+N5y5jHnKfrb/B4ri0CH/ivK4ng+8ab96+90IAKw2+I Ih8EFGk+7boxNQBDPPGxiaFXbl04hWUjhRvePOWFch5ZQmxfQ3sV1Fev0lwS8/x4zdcvw1rA5Je DbLpFbT1YKUjWyg== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao Now all workqueues will be WQ_UNBOUND, including WQ_PERCPU, and the attributes will define the affinity and concurrency management. It gives WQ_PERCPU a single meaning. It no longer selects per-cpu pools; it only asserts "this workqueue is concurrency managed and may not leave that backend." This opens up space for a lot of optimizations down the line, but I want to stop here to discuss if this is the right approach. Signed-off-by: Breno Leitao --- kernel/workqueue.c | 51 +++++++++++++++++++++++++++++---------------------- 1 file changed, 29 insertions(+), 22 deletions(-) diff --git a/kernel/workqueue.c b/kernel/workqueue.c index f86e8d0770ca22..da4d3e6cc5ee91 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -1669,7 +1669,7 @@ static bool is_percpu_pool(struct worker_pool *pool) static struct wq_node_nr_active *wq_node_nr_active(struct workqueue_struct= *wq, int node) { - BUG_ON(!(wq->flags & WQ_UNBOUND)); + BUG_ON(wq->flags & WQ_PERCPU); =20 if (node =3D=3D NUMA_NO_NODE) node =3D nr_node_ids; @@ -2453,10 +2453,10 @@ static void __queue_work(int cpu, struct workqueue_= struct *wq, retry: /* pwq which will be used unless @work is executing elsewhere */ if (req_cpu =3D=3D WORK_CPU_UNBOUND) { - if (wq->flags & WQ_UNBOUND) - cpu =3D wq_select_unbound_cpu(raw_smp_processor_id()); - else + if (wq->flags & WQ_PERCPU) cpu =3D raw_smp_processor_id(); + else + cpu =3D wq_select_unbound_cpu(raw_smp_processor_id()); } =20 pwq =3D rcu_dereference(*per_cpu_ptr(wq->cpu_pwq, cpu)); @@ -2500,7 +2500,7 @@ static void __queue_work(int cpu, struct workqueue_st= ruct *wq, * on it, so the retrying is guaranteed to make forward-progress. */ if (unlikely(!pwq->refcnt)) { - if (wq->flags & WQ_UNBOUND) { + if (!(wq->flags & WQ_PERCPU)) { raw_spin_unlock(&pool->lock); cpu_relax(); goto retry; @@ -2669,7 +2669,7 @@ bool queue_work_node(int node, struct workqueue_struc= t *wq, * workqueue_select_cpu_near would need to be updated to allow for * some round robin type logic. */ - WARN_ON_ONCE(!(wq->flags & WQ_UNBOUND)); + WARN_ON_ONCE(wq->flags & WQ_PERCPU); =20 local_irq_save(irq_flags); =20 @@ -5297,7 +5297,7 @@ static void rcu_free_wq(struct rcu_head *rcu) struct workqueue_struct *wq =3D container_of(rcu, struct workqueue_struct, rcu); =20 - if (wq->flags & WQ_UNBOUND) + if (!(wq->flags & WQ_PERCPU)) free_node_nr_active(wq->node_nr_active); =20 free_flush_pnodes(wq); @@ -5799,7 +5799,7 @@ static void apply_wqattrs_commit(struct apply_wqattrs= _ctx *ctx) ctx->dfl_pwq =3D install_pwq(ctx->wq, -1, ctx->dfl_pwq); =20 /* update node_nr_active->max, which only unbound workqueues have */ - if (ctx->wq->flags & WQ_UNBOUND) + if (!(ctx->wq->flags & WQ_PERCPU)) wq_update_node_max_active(ctx->wq, -1); =20 mutex_unlock(&ctx->wq->mutex); @@ -5841,8 +5841,8 @@ int apply_workqueue_attrs(struct workqueue_struct *wq, { int ret; =20 - /* only unbound workqueues can change attributes */ - if (WARN_ON(!(wq->flags & WQ_UNBOUND))) + /* a percpu workqueue is pinned to the concurrency managed backend */ + if (WARN_ON(wq->flags & WQ_PERCPU)) return -EINVAL; =20 /* concurrency management comes from WQ_PERCPU, it is not applied */ @@ -5882,7 +5882,7 @@ static void unbound_wq_update_pwq(struct workqueue_st= ruct *wq, int cpu) =20 lockdep_assert_held(&wq_pool_mutex); =20 - if (!(wq->flags & WQ_UNBOUND) || wq->attrs->ordered) + if (wq->attrs->ordered || wq->attrs->concurrency_managed) return; =20 /* @@ -6085,7 +6085,7 @@ static void wq_adjust_max_active(struct workqueue_str= uct *wq) WRITE_ONCE(wq->min_active, new_min); WRITE_ONCE(wq->percpu_max_active, new_max); =20 - if (wq->flags & WQ_UNBOUND) + if (!(wq->flags & WQ_PERCPU)) wq_update_node_max_active(wq, -1); =20 if (new_max =3D=3D 0) @@ -6170,6 +6170,13 @@ static struct workqueue_struct *__alloc_workqueue(co= nst char *fmt, flags &=3D ~WQ_PERCPU; } =20 + /* + * Every workqueue is backed by the unbound machinery now. WQ_PERCPU no + * longer picks a backend, it only says the workqueue is concurrency + * managed and may not leave that backend. + */ + flags |=3D WQ_UNBOUND; + if (flags & WQ_BH) { /* * BH workqueues always share a single execution context per CPU @@ -6200,7 +6207,7 @@ static struct workqueue_struct *__alloc_workqueue(con= st char *fmt, if (alloc_flush_pnodes(wq) < 0) goto err_free_wq; =20 - if (flags & WQ_UNBOUND) { + if (!(flags & WQ_PERCPU)) { if (alloc_node_nr_active(wq->node_nr_active) < 0) goto err_free_wq; } @@ -6239,7 +6246,7 @@ static struct workqueue_struct *__alloc_workqueue(con= st char *fmt, */ if (pwq_release_worker) kthread_flush_worker(pwq_release_worker); - if (wq->flags & WQ_UNBOUND) + if (!(wq->flags & WQ_PERCPU)) free_node_nr_active(wq->node_nr_active); err_free_wq: free_workqueue_attrs(wq->attrs); @@ -6488,8 +6495,7 @@ EXPORT_SYMBOL_GPL(workqueue_set_max_active); void workqueue_set_min_active(struct workqueue_struct *wq, int min_active) { /* min_active is only meaningful for non-ordered unbound workqueues */ - if (WARN_ON((wq->flags & (WQ_BH | WQ_UNBOUND | __WQ_ORDERED)) !=3D - WQ_UNBOUND)) + if (WARN_ON(wq->flags & (WQ_BH | WQ_PERCPU | __WQ_ORDERED))) return; =20 mutex_lock(&wq->mutex); @@ -7185,7 +7191,7 @@ int workqueue_online_cpu(unsigned int cpu) list_for_each_entry(wq, &workqueues, list) { struct workqueue_attrs *attrs =3D wq->attrs; =20 - if (wq->flags & WQ_UNBOUND) { + if (!(wq->flags & WQ_PERCPU)) { const struct wq_pod_type *pt =3D wqattrs_pod_type(attrs); int tcpu; =20 @@ -7220,7 +7226,7 @@ int workqueue_offline_cpu(unsigned int cpu) list_for_each_entry(wq, &workqueues, list) { struct workqueue_attrs *attrs =3D wq->attrs; =20 - if (wq->flags & WQ_UNBOUND) { + if (!(wq->flags & WQ_PERCPU)) { const struct wq_pod_type *pt =3D wqattrs_pod_type(attrs); int tcpu; =20 @@ -7395,7 +7401,8 @@ static int workqueue_apply_unbound_cpumask(const cpum= ask_var_t unbound_cpumask) lockdep_assert_held(&wq_pool_mutex); =20 list_for_each_entry(wq, &workqueues, list) { - if (!(wq->flags & WQ_UNBOUND) || (wq->flags & __WQ_DESTROYING)) + if ((wq->flags & __WQ_DESTROYING) || + wq->attrs->concurrency_managed) continue; =20 ctx =3D apply_wqattrs_prepare(wq, wq->attrs, unbound_cpumask); @@ -7556,7 +7563,7 @@ static ssize_t per_cpu_show(struct device *dev, struc= t device_attribute *attr, { struct workqueue_struct *wq =3D dev_to_wq(dev); =20 - return scnprintf(buf, PAGE_SIZE, "%d\n", (bool)!(wq->flags & WQ_UNBOUND)); + return scnprintf(buf, PAGE_SIZE, "%d\n", (bool)(wq->flags & WQ_PERCPU)); } static DEVICE_ATTR_RO(per_cpu); =20 @@ -7932,7 +7939,7 @@ int workqueue_sysfs_register(struct workqueue_struct = *wq) return ret; } =20 - if (wq->flags & WQ_UNBOUND) { + if (!(wq->flags & WQ_PERCPU)) { struct device_attribute *attr; =20 for (attr =3D wq_sysfs_unbound_attrs; attr->attr.name; attr++) { @@ -8800,7 +8807,7 @@ void __init workqueue_init_topology(void) list_for_each_entry(wq, &workqueues, list) { for_each_online_cpu(cpu) unbound_wq_update_pwq(wq, cpu); - if (wq->flags & WQ_UNBOUND) { + if (!(wq->flags & WQ_PERCPU)) { mutex_lock(&wq->mutex); wq_update_node_max_active(wq, -1); mutex_unlock(&wq->mutex); --=20 2.53.0-Meta