From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7D23039D6DA; Wed, 2 Sep 2026 19:29:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377350; cv=none; b=WlKjtUay56pHInfGIfsIbSFkp7JNySMK7XuUuuQaIfWMH6uc0koVdXUlUU40SfNQz0t9wbDp+l0EU7guqOntIXaEXl6EO00E75RvHd2LuylaNJp+X4CqbULFW/UF4B+demxZBdiXX7LHjoMcxracxAyMTClMPKu2uJ2thGOfKPQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377350; c=relaxed/simple; bh=ab7jO6lz+XT+Bcp3r5CMowTvoccAdGfKY8J7aGSj2o8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FXXMZRyk3y6djMR0WeG5JLpwK97DqX6VZGwf/NXSHcQnlR+yYdiusA+oeJ21/w3TPqk7gxpYcmWfRNcfMIFdv+h6+m17WfsAHzqZoup3HlLRHdDtH6tfu3YGADgmo+WugDr2zhbHhV3R1zItp1skaVnCFpy3Y/Tel2V5JrfNhSI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=iqCbQh3Y; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="iqCbQh3Y" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B86B91F00A3D; Wed, 2 Sep 2026 19:29:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377346; bh=UntNAVVE5PcFzEZfDLpbTIVU+V+Nb8NqweMsWPYgcwc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=iqCbQh3YjeV03Yz/TvlFbHs49UF910IqRbRPZRsQsKtEl7FhXI0IP5JvGg6G/qWWc 8e4veG7Dv19TjHGCD1ZuCrTxSyLxLzy5NmttVTfS7DlE32P5mUycZOINui3yXvG3tS MJqO9LJ3E1MT42sW0yyf1jWQWumK8HDFtLAPvurPO3uIiEZvgxqB4RBF4aSOsNVZqC Zbk2hQmf2mFFt2ZZZEi6xelGiQZ9dOfpUiexH2Efn9fJ82aIhLiGaYkUByh0/Q1KEN IEcr/Mbgh71DEBglJxu7djxZnj7I6qSUqonyTZyFTBilJ8YooK1RQWyaBPsj6d1+Ts 0ToZMDtb6Nf6Q== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:46 -0400 Subject: [PATCH RFC v2 1/8] SUNRPC: Use atomic_t for XID allocation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-1-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1882; i=cel@kernel.org; h=from:subject:message-id; bh=ab7jO6lz+XT+Bcp3r5CMowTvoccAdGfKY8J7aGSj2o8=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkAeWFgSuvKZvq5qDSzHfsgKcZmKJvXvpJtC FQyW28h2mSJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ ly6PEACp7Jfnzys31rtXXC9B3r0tFm8EcwmzdWzDxB3Ge0+BO8hRkXWbXsJoTdkrtgIEePgLP7T FGkbihSCGtxXw1Mz7ZxtHFVCzjX4ouvqsSyS2AQab1ZnACs0EiR/tjjWU6OjUJSyFZiPyW6yl72 XrJUcG8jQVzL40Fx4wePFwVNpv5ibCifEsy8bAaKyj3oTvOYRx/NHS9pI0ucgcwkcsqL/D2xCAK HIBfJz4hWlUyEscTR5nkPuRSdbQtaD+1fUr5MlUZNJjvxi4GC0v8LyG4vAT/RhG/yShbVmwIVaj OKPioU7V8Dnhtw7lexX344sINP0RJrNHfMD1xLkuclj7xq+tOCqqyKIyT3HTo1T9OfEulv8hzq5 gW8Q63OcqCBaC8gQgUVqxUQ/8TFmrkgminNNyQQN+lSMUMdWjd5vQyRoYV8dDV/TGWg6vQGCYVY ibeOovAjx9zCu1dZD/ihRtgujzx9bbx7qFO90/crUI4Xwsq5Ajy5YIAAAqRgorQq90F0oi9Ni0Z TtZT1HcIphr//YWWIi/MRtc8bJ282hl2kxRP5hOLpNSmyC49zLPKj7rczF5VRejhesBXk3wEg+9 hv9+aBl+GnhrextchaCF0s5iQtJ6EzR2olua7Aa00DSbags4NsipuT7P/xBXPCec/yTNpczc8jZ wSYgqB9b3B2IYew== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprt_alloc_xid() acquires reserve_lock to increment a simple counter. Under a high-IOPS NFSv3 workload on 100GbE RDMA, profiling shows 1.06% of system-wide CPU cycles contending on this lock in xprt_request_init, as ~150 RPC worker threads serialize on the counter. reserve_lock protects the slot table and backlog queue, but XID allocation is an independent operation that does not require synchronization with either. Signed-off-by: Chuck Lever --- include/linux/sunrpc/xprt.h | 2 +- net/sunrpc/xprt.c | 9 ++------- 2 files changed, 3 insertions(+), 8 deletions(-) diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h index a82045804d34..0d6c3f6bf97e 100644 --- a/include/linux/sunrpc/xprt.h +++ b/include/linux/sunrpc/xprt.h @@ -273,7 +273,7 @@ struct rpc_xprt { spinlock_t transport_lock; /* lock transport info */ spinlock_t reserve_lock; /* lock slot table */ spinlock_t queue_lock; /* send/receive queue lock */ - u32 xid; /* Next XID value to use */ + atomic_t xid; /* Most recently issued XID */ struct rpc_task * snd_task; /* Task blocked in send */ =20 struct list_head xmit_queue; /* Send queue */ diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c index 48a3618cbb29..186c14f0f928 100644 --- a/net/sunrpc/xprt.c +++ b/net/sunrpc/xprt.c @@ -1882,18 +1882,13 @@ xprt_init_connect_cookie(struct rpc_rqst *req, stru= ct rpc_xprt *xprt) static __be32 xprt_alloc_xid(struct rpc_xprt *xprt) { - __be32 xid; - - spin_lock(&xprt->reserve_lock); - xid =3D (__force __be32)xprt->xid++; - spin_unlock(&xprt->reserve_lock); - return xid; + return (__force __be32)atomic_inc_return(&xprt->xid); } =20 static void xprt_init_xid(struct rpc_xprt *xprt) { - xprt->xid =3D get_random_u32(); + atomic_set(&xprt->xid, get_random_u32()); } =20 static void --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3C58A3921C0; Wed, 2 Sep 2026 19:29:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377361; cv=none; b=qQ0qPCNlzlVrF2o4Lu3IaIj/dMtZK0oFLFL7dcX/eyscCea3BsWAvjYq0U6K8gKtlHRK8FvTkRrZDCzOHjb+XEqjqVuvGXGSHCkvqdyvSJWKJJEoxJ7vASHbvwL+FjsC5lGZ5oXWtBr5ln0zLeuY0+s17RWmZqozXi67hV+WujY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377361; c=relaxed/simple; bh=kcw7JM5jtqpHMeCt4LZb3tnxTxZ899QMpy4Qy7hECnI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=gTyYzxUU014uu8F98c199e/YX45hzeoqjQozfKApypZ91iNCIrI7peg0ckvqMsdV83T+KEXZ+5pxbKTZu5PyC9L+D4FM43KSRwjVPSTClWsfYiipAVgm4Fv/OB80iKVdNW/cgmXShYDK82isRRhAcEZrxponi16f2997jq98hlM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NAJTnxzh; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NAJTnxzh" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 7E97A1F00A3A; Wed, 2 Sep 2026 19:29:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377347; bh=+eYTodiaUqlicVvz2NVCtjLK1G4eOJJQEJOrtPVpfso=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=NAJTnxzhn6EQL2a90E2Av0JOJ7g+u/XvmW+LcPTkuyLyOpj+uOf1paQKGeuMyLfrW LbUTDr9dgltyEAvfc8U+cySAQhntqiuxZBcu2fAskozmaxFHTlCh+G7fULqooxscZN HNYvNjvpzWQPt4x/8amF9YrXViwezRQTmx+ugvT+bhVh7WDMLA06fmdgHGLOTh9U2p +IYcDD2rJ0v4cBOTiYJp6sTn3J0hKEmhEfLxCjeQgnJbgakRd/oXyyBONzj/9ByxPH vMkdnap6SXkUV/2khI3FZP3IjVFikzFdzjQMiqESsa0dDYskUZhACZG1AW91adCIOy Qm3IqdKnhhdMg== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:47 -0400 Subject: [PATCH RFC v2 2/8] SUNRPC: Split recv_lock out of xprt->queue_lock Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-2-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=15822; i=cel@kernel.org; h=from:subject:message-id; bh=kcw7JM5jtqpHMeCt4LZb3tnxTxZ899QMpy4Qy7hECnI=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkAlVFUPH0B6+sOQfT9XEWu71on3M96lhWgk JFSpoeUZeyJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l7UXEAC6rB2BhRChU2fcMFdWDF6h1baCtd8VuGUK/nmNAiIkuMOJggVQXmSRVaIkRWsE5M/yRR3 JUqfbaEQoMjeauQx82y4uAIB4XpZha6IiFhPOIRS8fGpq4Nq2ZOCMLOAbtc3z6DxlJmqZ2T2XIA 8L9um5BieevuiL1jWWt+JoZhypsjmnsncNXpRz3HLyVJnPq2heglO3+lIzV86eEo3vzTSiNmxbR NLg++6eHjT3f0UqRf81yLTw7TYXikeF4jKYhNqJhBc0d2wpwdtuk01Fl6mQ7A3Nqk1eedS6RzFN HoCedUdElyWjhZIJuNxAl8/HFUmmdRg/Z6BwPZbqjiRgrFFnTgXKNne+LBpuPXeYtl+Jnzpbwx8 rgnkT2SRb9m1t0EGgtW1tk245BLbOUEwOCq3VwzU11nrVSp+ga9KEpzln3rNQjxFlaBpytjEHt0 XQfpaBX3medLRykHdng0HiY2HfjI0iXMDxX0mb9ACzytFMYivfOijhN8+98De3GWZcid5riZ4vG Fdn2fO9dR1iyIrZJzy9SqQLOWrfCMvkKV9Wrw8DAI0+luCCRC2KzJHzK1CkVSiAmFSnC9vnFllT ZlzqQKj6toWA5aGcBtNDI9TLBQ4ZXUwmjdvKENOhftWGN3gW+2fp9L+RbZBk3XiDmxGOhpK3but GvzXFIHlFO0fXwQ== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprt->queue_lock protects two independent structures: the recv_queue rb-tree for reply matching and the xmit_queue list for transmit draining. No hot path touches both in one critical section, yet every RPC submit and completion contends on the same lock. Under a 4KB NFSv3 READ workload on 100GbE RDMA, 53% of non-idle CPU cycles are spent in native_queued_spin_lock_slowpath: the CQ completion worker running rpcrdma_reply_handler serializes against ~150 kworker threads enqueuing receives and transmits. Introduce xprt->recv_lock for the receive path -- recv_queue operations, request lookup, receive-side pinning, and completion -- leaving queue_lock to the xmit_queue and the xprt_transmit drain loop. A request is pinned under the lock of the queue it was found through, so xprt_request_dequeue_xprt() drains pins once under each lock before it dequeues from that queue. The transmit dequeue has to see a zero pin count under queue_lock: it frees the send buffer's bvec that a transmitter is iterating, and it aborts a partial send only if the request is still first in the queue. A receive-side unpin runs under recv_lock rather than the lock that publishes RPC_TASK_MSG_PIN_WAIT, so xprt_unpin_rqst() wakes the waiter whenever the count reaches zero instead of testing the flag. Also move the rq_private_buf memcpy in xprt_request_enqueue_receive above the lock acquisition: until the rb-tree insert publishes the request, the reply handler cannot see it, so the copy is safe unlocked and the submitter's critical section shrinks to the insert alone. Signed-off-by: Chuck Lever --- include/linux/sunrpc/xprt.h | 6 ++- net/sunrpc/svcsock.c | 6 +-- net/sunrpc/xprt.c | 76 +++++++++++++++++++-------= ---- net/sunrpc/xprtrdma/rpc_rdma.c | 14 +++--- net/sunrpc/xprtrdma/svc_rdma_backchannel.c | 8 ++-- net/sunrpc/xprtsock.c | 18 +++---- 6 files changed, 77 insertions(+), 51 deletions(-) diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h index 0d6c3f6bf97e..ed1e28b74f02 100644 --- a/include/linux/sunrpc/xprt.h +++ b/include/linux/sunrpc/xprt.h @@ -272,7 +272,7 @@ struct rpc_xprt { atomic_long_t queuelen; spinlock_t transport_lock; /* lock transport info */ spinlock_t reserve_lock; /* lock slot table */ - spinlock_t queue_lock; /* send/receive queue lock */ + spinlock_t queue_lock; /* send queue lock */ atomic_t xid; /* Most recently issued XID */ struct rpc_task * snd_task; /* Task blocked in send */ =20 @@ -292,6 +292,10 @@ struct rpc_xprt { * backchannel rpc_rqst's */ #endif /* CONFIG_SUNRPC_BACKCHANNEL */ =20 + /* + * Receive stuff + */ + spinlock_t recv_lock; /* receive queue lock */ struct rb_root recv_queue; /* Receive queue */ =20 struct { diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c index 50e5e7f5b762..8939ba604385 100644 --- a/net/sunrpc/svcsock.c +++ b/net/sunrpc/svcsock.c @@ -1102,7 +1102,7 @@ static int receive_cb_reply(struct svc_sock *svsk, st= ruct svc_rqst *rqstp) =20 if (!bc_xprt) return -EAGAIN; - spin_lock(&bc_xprt->queue_lock); + spin_lock(&bc_xprt->recv_lock); req =3D xprt_lookup_rqst(bc_xprt, xid); if (!req) goto unlock_eagain; @@ -1120,10 +1120,10 @@ static int receive_cb_reply(struct svc_sock *svsk, = struct svc_rqst *rqstp) memcpy(dst->iov_base, src->iov_base, src->iov_len); xprt_complete_rqst(req->rq_task, rqstp->rq_arg.len); rqstp->rq_arg.len =3D 0; - spin_unlock(&bc_xprt->queue_lock); + spin_unlock(&bc_xprt->recv_lock); return 0; unlock_eagain: - spin_unlock(&bc_xprt->queue_lock); + spin_unlock(&bc_xprt->recv_lock); return -EAGAIN; } =20 diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c index 186c14f0f928..883123ec70b0 100644 --- a/net/sunrpc/xprt.c +++ b/net/sunrpc/xprt.c @@ -1061,7 +1061,7 @@ xprt_request_rb_remove(struct rpc_xprt *xprt, struct = rpc_rqst *req) * @xprt: transport on which the original request was transmitted * @xid: RPC XID of incoming reply * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ struct rpc_rqst *xprt_lookup_rqst(struct rpc_xprt *xprt, __be32 xid) { @@ -1092,8 +1092,9 @@ xprt_is_pinned_rqst(struct rpc_rqst *req) * xprt_pin_rqst - Pin a request on the transport receive list * @req: Request to pin * - * Caller must ensure this is atomic with the call to xprt_lookup_rqst() - * so should be holding xprt->queue_lock. + * Caller must hold the lock that protects the queue through which + * it found the request: xprt->recv_lock for the receive path, + * xprt->queue_lock for the transmit drain path. */ void xprt_pin_rqst(struct rpc_rqst *req) { @@ -1105,14 +1106,10 @@ EXPORT_SYMBOL_GPL(xprt_pin_rqst); * xprt_unpin_rqst - Unpin a request on the transport receive list * @req: Request to pin * - * Caller should be holding xprt->queue_lock. + * Caller holds the lock it held for the matching xprt_pin_rqst(). */ void xprt_unpin_rqst(struct rpc_rqst *req) { - if (!test_bit(RPC_TASK_MSG_PIN_WAIT, &req->rq_task->tk_runstate)) { - atomic_dec(&req->rq_pin); - return; - } if (atomic_dec_and_test(&req->rq_pin)) wake_up_var(&req->rq_pin); } @@ -1123,6 +1120,26 @@ static void xprt_wait_on_pinned_rqst(struct rpc_rqst= *req) wait_var_event(&req->rq_pin, !xprt_is_pinned_rqst(req)); } =20 +/* + * A pin is taken under the lock of the queue the request was found + * through, so a zero count observed under @lock rules out any pinner + * that came through that queue. The lock is dropped to wait, and the + * re-test under it catches a pinner that arrived in the gap. + */ +static void xprt_request_drain_pins(struct rpc_task *task, spinlock_t *loc= k) + __must_hold(lock) +{ + struct rpc_rqst *req =3D task->tk_rqstp; + + while (xprt_is_pinned_rqst(req)) { + set_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); + spin_unlock(lock); + xprt_wait_on_pinned_rqst(req); + spin_lock(lock); + clear_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); + } +} + static bool xprt_request_data_received(struct rpc_task *task) { @@ -1155,16 +1172,16 @@ xprt_request_enqueue_receive(struct rpc_task *task) ret =3D xprt_request_prepare(task->tk_rqstp, &req->rq_rcv_buf); if (ret) return ret; - spin_lock(&xprt->queue_lock); - - /* Update the softirq receive buffer */ + /* Reply handlers cannot find the request until the rb-tree + * insert below publishes it, so the copy needs no lock. + */ memcpy(&req->rq_private_buf, &req->rq_rcv_buf, sizeof(req->rq_private_buf)); =20 - /* Add request to the receive list */ + spin_lock(&xprt->recv_lock); xprt_request_rb_insert(xprt, req); set_bit(RPC_TASK_NEED_RECV, &task->tk_runstate); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 /* Turn off autodisconnect */ timer_delete_sync(&xprt->timer); @@ -1175,7 +1192,7 @@ xprt_request_enqueue_receive(struct rpc_task *task) * xprt_request_dequeue_receive_locked - Remove a request from the receive= queue * @task: RPC task * - * Caller must hold xprt->queue_lock. + * Caller must hold xprt->recv_lock. */ static void xprt_request_dequeue_receive_locked(struct rpc_task *task) @@ -1190,7 +1207,7 @@ xprt_request_dequeue_receive_locked(struct rpc_task *= task) * xprt_update_rtt - Update RPC RTT statistics * @task: RPC request that recently completed * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ void xprt_update_rtt(struct rpc_task *task) { @@ -1212,7 +1229,7 @@ EXPORT_SYMBOL_GPL(xprt_update_rtt); * @task: RPC request that recently completed * @copied: actual number of bytes received from the transport * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ void xprt_complete_rqst(struct rpc_task *task, int copied) { @@ -1309,7 +1326,7 @@ void xprt_request_wait_receive(struct rpc_task *task) * The spinlock ensures atomicity between the test of * req->rq_reply_bytes_recvd, and the call to rpc_sleep_on(). */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); if (test_bit(RPC_TASK_NEED_RECV, &task->tk_runstate)) { xprt->ops->wait_for_reply_request(task); /* @@ -1321,7 +1338,7 @@ void xprt_request_wait_receive(struct rpc_task *task) rpc_wake_up_queued_task_set_status(&xprt->pending, task, -ENOTCONN); } - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 static bool @@ -1439,7 +1456,12 @@ xprt_request_dequeue_transmit(struct rpc_task *task) * @task: pointer to rpc_task * * Remove a task from the transmit and receive queues, and ensure that - * it is not pinned by the receive work item. + * it is not pinned by any concurrent work item. + * + * The transmit dequeue frees the send buffer's bvec and may abort a + * partial send, so it must not run while a transmitter holds a pin. + * Drain pins under each queue's lock before leaving that queue; once + * the request is off both, no new pin can be taken. */ void xprt_request_dequeue_xprt(struct rpc_task *task) @@ -1451,16 +1473,15 @@ xprt_request_dequeue_xprt(struct rpc_task *task) test_bit(RPC_TASK_NEED_RECV, &task->tk_runstate) || xprt_is_pinned_rqst(req)) { spin_lock(&xprt->queue_lock); - while (xprt_is_pinned_rqst(req)) { - set_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); - spin_unlock(&xprt->queue_lock); - xprt_wait_on_pinned_rqst(req); - spin_lock(&xprt->queue_lock); - clear_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); - } + xprt_request_drain_pins(task, &xprt->queue_lock); xprt_request_dequeue_transmit_locked(task); - xprt_request_dequeue_receive_locked(task); spin_unlock(&xprt->queue_lock); + + spin_lock(&xprt->recv_lock); + xprt_request_drain_pins(task, &xprt->recv_lock); + xprt_request_dequeue_receive_locked(task); + spin_unlock(&xprt->recv_lock); + xdr_free_bvec(&req->rq_rcv_buf); } } @@ -2038,6 +2059,7 @@ static void xprt_init(struct rpc_xprt *xprt, struct n= et *net) spin_lock_init(&xprt->transport_lock); spin_lock_init(&xprt->reserve_lock); spin_lock_init(&xprt->queue_lock); + spin_lock_init(&xprt->recv_lock); =20 INIT_LIST_HEAD(&xprt->free); xprt->recv_queue =3D RB_ROOT; diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c index 1285f04cdac1..a82d3d9bc7ae 100644 --- a/net/sunrpc/xprtrdma/rpc_rdma.c +++ b/net/sunrpc/xprtrdma/rpc_rdma.c @@ -1321,9 +1321,9 @@ void rpcrdma_unpin_rqst(struct rpcrdma_rep *rep) req->rl_reply =3D NULL; rep->rr_rqst =3D NULL; =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_unpin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 /** @@ -1363,10 +1363,10 @@ void rpcrdma_complete_rqst(struct rpcrdma_rep *rep) goto out_badheader; =20 out: - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(rqst->rq_task, status); xprt_unpin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); return; =20 out_badheader: @@ -1492,12 +1492,12 @@ void rpcrdma_reply_handler(struct rpcrdma_rep *rep) /* Match incoming rpcrdma_rep to an rpcrdma_req to * get context for handling any incoming chunks. */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); rqst =3D xprt_lookup_rqst(xprt, rep->rr_xid); if (!rqst) goto out_norqst; xprt_pin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 if (buf->rb_credits !=3D credits) rpcrdma_update_cwnd(r_xprt, credits); @@ -1524,7 +1524,7 @@ void rpcrdma_reply_handler(struct rpcrdma_rep *rep) return; =20 out_norqst: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); trace_xprtrdma_reply_rqst_err(rep); rpcrdma_rep_put(buf, rep); goto out_post; diff --git a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c b/net/sunrpc/xprtrd= ma/svc_rdma_backchannel.c index e5a78b761012..3c7b85427f33 100644 --- a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c +++ b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c @@ -28,7 +28,7 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, struct rpc_rqst *req; u32 credits; =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); req =3D xprt_lookup_rqst(xprt, *rdma_resp); if (!req) goto out_unlock; @@ -39,7 +39,7 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, goto out_unlock; memcpy(dst->iov_base, src->iov_base, src->iov_len); xprt_pin_rqst(req); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 credits =3D be32_to_cpup(rdma_resp + 2); if (credits =3D=3D 0) @@ -50,13 +50,13 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, xprt->cwnd =3D credits << RPC_CWNDSHIFT; spin_unlock(&xprt->transport_lock); =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(req->rq_task, rcvbuf->len); xprt_unpin_rqst(req); rcvbuf->len =3D 0; =20 out_unlock: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 /* Send a reverse-direction RPC Call. diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c index 7f60723fa64d..1454da9575b3 100644 --- a/net/sunrpc/xprtsock.c +++ b/net/sunrpc/xprtsock.c @@ -673,25 +673,25 @@ xs_read_stream_reply(struct sock_xprt *transport, str= uct msghdr *msg, int flags) ssize_t ret =3D 0; =20 /* Look up and lock the request corresponding to the given XID */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); req =3D xprt_lookup_rqst(xprt, transport->recv.xid); if (!req || (transport->recv.copied && !req->rq_private_buf.len)) { msg->msg_flags |=3D MSG_TRUNC; goto out; } xprt_pin_rqst(req); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 ret =3D xs_read_stream_request(transport, msg, flags, req); =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); if (msg->msg_flags & (MSG_EOR|MSG_TRUNC)) xprt_complete_rqst(req->rq_task, transport->recv.copied); else req->rq_private_buf.len =3D transport->recv.copied; xprt_unpin_rqst(req); out: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); return ret; } =20 @@ -1398,13 +1398,13 @@ static void xs_udp_data_read_skb(struct rpc_xprt *x= prt, return; =20 /* Look up and lock the request corresponding to the given XID */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); rovr =3D xprt_lookup_rqst(xprt, *xp); if (!rovr) goto out_unlock; xprt_pin_rqst(rovr); xprt_update_rtt(rovr->rq_task); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); task =3D rovr->rq_task; =20 if ((copied =3D rovr->rq_private_buf.buflen) > repsize) @@ -1412,7 +1412,7 @@ static void xs_udp_data_read_skb(struct rpc_xprt *xpr= t, =20 /* Suck it into the iovec, verify checksum if not done by hw. */ if (csum_partial_copy_to_xdr(&rovr->rq_private_buf, skb)) { - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); __UDPX_INC_STATS(sk, UDP_MIB_INERRORS); goto out_unpin; } @@ -1421,13 +1421,13 @@ static void xs_udp_data_read_skb(struct rpc_xprt *x= prt, spin_lock(&xprt->transport_lock); xprt_adjust_cwnd(xprt, task, copied); spin_unlock(&xprt->transport_lock); - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(task, copied); __UDPX_INC_STATS(sk, UDP_MIB_INDATAGRAMS); out_unpin: xprt_unpin_rqst(rovr); out_unlock: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 static void xs_udp_data_receive(struct sock_xprt *transport) --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7BF3F39EF12; Wed, 2 Sep 2026 19:29:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377364; cv=none; b=NTEA7xTFmNA4VILp0CLz1VSFGm1no7FSbylBrC5X+9JPCsVPoLzqMTSLP4KSQrYa6/5Ubc+WovgPl/K/Dt2Fz/wxIh0v2Mo54A3T43cNp757SCAhTYX3rLMu2D9mqWHjB8W/7Yny9OKor5C1A7IJ2P6j7Q/Cwq4CkyT3GZgwu3Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377364; c=relaxed/simple; bh=8U0XBPRVxSfOU0obQjD1fxaP88zM2TcU5L4hYzBluvY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=M2o9aYjyMmxEssCmGBBls/ebmKS7h6QdpCbVn7zTNyldsbsXugKNxXwEQfHxn2dbriQDUX2NjBaiqip6X8uYyAvdnblEcUYtXLOpqlN3u2uw3h3N5/X+OZK1LLWhDgDZA9wwhmBNsEkHjzp05t6x0B7fX1Q/yyFcuJNUKTR5TBo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=jsz9pgy4; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="jsz9pgy4" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 576AE1F00ADB; Wed, 2 Sep 2026 19:29:07 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377347; bh=Oqjf36SDCKS8VzDfnH6ImhAmZ1KHIl0D7LfeOiVeArU=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=jsz9pgy44s5jK9Y2y4/YAXiy5MO8jxZYQYPRwjI2g0AevKdQQX7deLJJXFaizYAUU yWODiwzzbsd/sxEtLnIXex3eyseuT3+720gfoJAdXv1pqpCl6UtzlEztAKskzy/eaD TWX35QTRdwddwnQUQEt69VzOpfdvZRm2Uuj/RtzHIGeXMr+rB/JzpzjK/GCeykJ1M/ fuRm+sY/vr7WmdqLorobbf0UJ/sI1Z+Qjni2K8crOfAOw6f8mQcZKzrnPTe6IyAUhP S5cfrgmID7zMN7zd96J7rkJmgKKTqQcVA1WW9aDNYE9broJ6E7vKnz+W9/2ECra3a1 s5l1ao363lMdQ== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:48 -0400 Subject: [PATCH RFC v2 3/8] SUNRPC: Set WQ_SYSFS on rpciod and xprtiod Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-3-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1248; i=cel@kernel.org; h=from:subject:message-id; bh=8U0XBPRVxSfOU0obQjD1fxaP88zM2TcU5L4hYzBluvY=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkAhBl6UAuyLi+KpB+/ch9lI+aI4JqgbZp0X KBAopwQNamJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l/R4D/42tBmmj06EEIB22lTVPWqgLgbbtgF8+INJsb2G/+x7ezyxQUbuGbDGqV3br5hvaSAaIMb ZCc/G1MdvrGZ1BS209jCQA9OoKR5vju8lHHVNUJtuMKKwyl5H2kCjVRJmpsKcgpc9TbSie+fiF+ efOdJKU4ysZc9QF9Imd4fen8qVurJiUh7FEhwYBebOPikL7KaPYuQ5fzJErPrloL8OFK/cSOSFu r/f5F5PJ+yGfd1GeV76YG4li04SiijtJmJ4TzAPODI+hu9fSts7PabFXlQBmOpov3QAq58qJYS1 i92VQjtlwscjOV9E82RLfjx2UzROiAloylYpZMT52FtLaMSCr6d/yjL6yBohj4zA6JBpzytyJBc Q8R98wTfY+z0xhzGDdSz86MO3XQ55juFdqnradvNzhDSM3HOv7DCD4ueI42RcMr2pUc/IMiozpV MoTjaUIFgIRW3/dwjtEsZEVDMeIWIbgamS3VYr5skZWbMxwFcb2ntvDONVyFuqf5NHunIeX8B6G mP1rp5NE/uqedhCHXerYPCikyAqZDByDO6X6EtO/uJ/1dgFJrfEcoVKK7WDpfXmEu6COa9nc6ie efV8zw0CQqHrjpgJ8nwBPcClsVILZG3BLNmNCdmICXcnif9W7GHKTMvgDJRuYxk4RBSHU55n9sY ncELwmij1zstWtg== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 An unbound workqueue's attributes can be changed at run time only through /sys/devices/virtual/workqueue/, and a workqueue appears there only when it is created with WQ_SYSFS. rpciod and xprtiod are created without it, so their affinity scope cannot be tuned without reloading the sunrpc module. Create both with WQ_SYSFS. Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index 016f16ca5779..e81419aa553c 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1273,16 +1273,17 @@ void rpciod_down(void) */ static int rpciod_start(void) { + const unsigned int wq_flags =3D WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS; struct workqueue_struct *wq; =20 /* * Create the rpciod thread and wait for it to start. */ - wq =3D alloc_workqueue("rpciod", WQ_MEM_RECLAIM | WQ_UNBOUND, 0); + wq =3D alloc_workqueue("rpciod", wq_flags, 0); if (!wq) goto out_failed; rpciod_workqueue =3D wq; - wq =3D alloc_workqueue("xprtiod", WQ_UNBOUND | WQ_MEM_RECLAIM, 0); + wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) goto free_rpciod; xprtiod_workqueue =3D wq; --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1978D39659A; Wed, 2 Sep 2026 19:29:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377362; cv=none; b=UENrOAzhlvaHT/ZOvhvHErE6GiWHWwXw5VswqiWHVU5HcPJZsD1Cc3PViuprQl2rbdngc2kt7OrPMOkggeFE5AkN/TJo8LCUAd537mEKXwir7I6SfHficUW4n2BuMy3GHsgbDmqQvb98JxZ2K4/Oi7Pfbq5v8fY9Es2tzxNmGYU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377362; c=relaxed/simple; bh=O3TNdnVPoa4rGbmi2tla4aohyiGvRmbOQcuo0rGfqKE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=EWX0gC8+YklMzUSv9qdtcbERiRju0NME/P7cv7An8bnRTFbHMpFTk25hU4j2bG4uPzj9ZuucTggQUq/LLTXcWfgAJFBbZ118J1TpryCge4ZhOB47goPbEQC5hG+vA1yFdyIZw3cCulZ8bDr4ilneNFwlTd8lkj2nZ8TL6lEwvl8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=E3uurmS+; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="E3uurmS+" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1BD6B1F00A3E; Wed, 2 Sep 2026 19:29:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377348; bh=2Zj70FxOcUP0kW1TVIjjFDiwHX+7yaJeILXFSkVDfkg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=E3uurmS+zQ5o2UMdKcqhYNEqtDL4cYXvdfjzc1/Th6Fmdga4fFHSwKkSGC/P4184P 9YYrX5CF2ACkAo3HbcWmgQisz/QapiAh7M0jjjY5eGOl7bV1z9bdcwjXkkGKvcW4lf NKEWQUnewUMZnbfwtE5abhXRFrknubUX6estY3S6onz4X6Mzt0l0FCzp5vwFpQhRLU 5lTVOsI/AoI5owNfDm2qhRH/5ZvFOu6hxqnE/dAsgPtqfTvXEmVZr5ybWoNIlgv0+p Y2VFLrG9Tc2iLjsydyJh0vk7FUfx4aYIAiClj5sugl3r64GFgnZDwGDPltFsyXPtSF 4PbLJ0bGi87SA== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:49 -0400 Subject: [PATCH RFC v2 4/8] NFS: Set WQ_SYSFS on nfsiod Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-4-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1001; i=cel@kernel.org; h=from:subject:message-id; bh=O3TNdnVPoa4rGbmi2tla4aohyiGvRmbOQcuo0rGfqKE=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkAaXKA35xguWdIO/e6pfBSBFjxh08mpsOrY N+Y7QYczUmJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l8eoD/9A8+juOd72lLi5emkXSQ/DSEoNSJxur07U9V27h6RGDAHhe1Zm9eVrDhAI22QGpQEtG8P FTvduuIaEuclhpLNSV+qTxgFoxJpkdgJrx/oVy0841LYdrVE1wsWcu6nR/MOWhZZ693MGo/U/ct jCk4BmmnIRa5GbotbDXXDU1I8G5z/qV30D2gCdtlhBN36nCDxRXrb3Fr8qtX6t/jrrUQ3QV19wd 5AhCJwYZ3ugMi2aeIvc0UCPv0JH/s/4ieERED2j8Nq9X9e3bPnVEEu0WDGyp3RAo7gM3WpRcAUf gJZMODuji1Usr+TwW8J1qFiqzzbVIDzqmf3Orn+CCLC7iWLou+rpvtp4dla/oACVxp/FzlQyf0D v+6SkQk7rEnusg296fNcv0VvlpglhmN4NyZAo84+iLuqSEi1QdW8Px57JYVDci7NzARHXY8t0SU NXuxLDCl9VN5AaX+FovYU29R8nizd2K4dZ6ogOb4H9EXAhf4sKOZaa9bSlG0j3mdcQBz33/SN3l Nmdfx0vBeZEYM6fc+DxsaQO2Z9fEPclqz+H/pRCgTBKiH3PXVsfC49fS93q+YRF819bDefxI9TI roXhTwPbqr08qnjNDQBlgPHRuleBOjPoRKCUsPzDDn2W3GHoffVNwyV9U31wCyoGS+Afxj07bNN 7C3M7JbCRVB+lRg== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 An unbound workqueue's attributes can be changed at run time only through /sys/devices/virtual/workqueue/, and a workqueue appears there only when it is created with WQ_SYSFS. nfsiod is created without it, so its affinity scope cannot be tuned without reloading the nfs module. Create nfsiod with WQ_SYSFS. Signed-off-by: Chuck Lever --- fs/nfs/inode.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 3022454f7698..107a2135029d 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -2619,7 +2619,8 @@ static void nfsiod_stop(void) static int nfsiod_start(void) { dprintk("RPC: creating workqueue nfsiod\n"); - nfsiod_workqueue =3D alloc_workqueue("nfsiod", WQ_MEM_RECLAIM | WQ_UNBOUN= D, 0); + nfsiod_workqueue =3D alloc_workqueue("nfsiod", + WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS, 0); if (nfsiod_workqueue =3D=3D NULL) return -ENOMEM; #if IS_ENABLED(CONFIG_NFS_LOCALIO) --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F068F3A75B1; Wed, 2 Sep 2026 19:29:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377374; cv=none; b=IMZzrJmldpxdB1Q3KaoFxhVvcVvWpFWqe00CdNW5fTnSZvYFWMJ9CMOuVnbpOfAg2my2E7LBSSFoV3bAwkeKa/+Tx7eODpeVPpxc1etDCBXb18We9yAlpMCYneMY2AyJKfdYGWf0lbmRgW6GA+0TPOJ2Z7WTFAznAp1npKR+6YE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377374; c=relaxed/simple; bh=j/QSHRUfrO/KT+a3fsQHrV0wDYoDF5qm9xBGqxvYjUg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Cu6rTU3PjQ7N9hQxpSeaXfGTSxT0iLesi4E4noiD6JdS+EhhPihSc6g+oJPBtvGxL74QVc+LrdwJZ4rR6X7dho5Qer9U4dBpa/1cqPtZ1Bq5os7g+jkEIdsRd4S2OAbKCHkiKxlzLZsmS/5eWIFsFwj+Bzo0bsaI3rZwfRnIxCg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=aqEkmehQ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="aqEkmehQ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DD79F1F00A3F; Wed, 2 Sep 2026 19:29:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377349; bh=2AhqExJ5ImvCkUoJzOUvUEsFmXKytKRmnDUDouDSiM8=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=aqEkmehQNMAO7ivXKssNPNDy1QKMEYEAzi3pqKRrmrsRAQKGJscX8LgsOTlSLd/gW X1ZsuG98zIrmYKvm+zRn3fu2JF/684PC8mc1xswDCdTSTzgbceU7OxpBYn+cp+B6gx MRyo5gePyVD65t7vnK/0ciQXNJjnzast+n/Y4ROttsdlfQdHgyMqoAQQD0kvFaYPT5 sYDyRZulIGSzy1okMbD86xLBRBHdBySTbGK5Ml1MTeHNSbBY/Niu69aSx3vcTRIiYy l7zf6wspT5bLONqeQ8CXdnYi6fFFdzDNNLCSH2yiCRsh+pFaiCc+/CaO36ryU/o3ds x2du5Byc/97Xg== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:50 -0400 Subject: [PATCH RFC v2 5/8] workqueue: add workqueue_set_affn_scope() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-5-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=3168; i=cel@kernel.org; h=from:subject:message-id; bh=j/QSHRUfrO/KT+a3fsQHrV0wDYoDF5qm9xBGqxvYjUg=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkALBUtrjLkaUKNHiXr2tZCl+DvW3s9kcB1o mglhUSQ5wmJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l35qD/9SHPbQ6vTgb5vn1aIb0Ug9KgZH4MqWoWq4Pw44bFFFVZS3xztLsPrK9fmiyokSybt96Dj mygMlzXaN2YCDDzwwMUF5kpYeuW+QJA6HVAJvFL1KW6o0p7YuOJ4+hgXXJwGyyFiZLSoUE2NYr4 3a5WikTyytsKBTNCrDKRdGLAA9a/M2qre880CjWMxo9DFrMr/EHZrTXlhJhiHFvYrm7Tfh2/qwv Nl7zJwxugaliKWod0ZykfXswiXT3CzHOBsHtZT88O4k2htNJIVM0/m87Q2brRcGsPQb4CYJQeDy +FfIv/ivBOvNNyAeqE5LAigrxmsmwyovknX5c7SV4qxskUTVRzvjPhUELVwi/coubOQVI4hPv75 Bpw0vvPJjcZx6DQQnYusNhC+Z13fbdh6Ipvf0H0MD3WWMvxY7Ll8nzph7UuCriuSmeAF2wIPHUw +p/OvAbN7kHlykPH46alkZe1GocIx8Xjn/3zFBvaGRb7iSDqrq1X9LVvoGCs+hszOLPw5p+XvaY nj226XADf6bzsM7e0QF6l64hZk02iViCwp7Kvkzz4NlmqjM7moaLERiKTWhxWFwom0gPkF6Th7C DP3mokf4qSmHFWgfoVubUkw4Ig8+P4iI88PDu5g3OFkfSV5OwCan/2lGuk5KVr1gkpyvelI3RkD GWvoqiu9Rg69GeQ== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 An unbound workqueue's affinity scope can be changed only through sysfs. A module whose workqueue contends on the pool lock at the default scope has no in-kernel way to select a finer one, because alloc_workqueue_attrs() and apply_workqueue_attrs() are not exported. Exporting them would also invite a caller to apply freshly allocated attributes, which resets the nice level, cpumask, and strict affinity the workqueue already carries. Add workqueue_set_affn_scope(), which copies the workqueue's current attributes, replaces only the scope, and applies the result under wq_pool_mutex, as the sysfs affinity_scope store does. Export it so that SUNRPC and NFS can set the scope of their workqueues when they create them. Signed-off-by: Chuck Lever --- include/linux/workqueue.h | 2 ++ kernel/workqueue.c | 37 +++++++++++++++++++++++++++++++++++++ 2 files changed, 39 insertions(+) diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h index c8a36423cb34..585c32d8dc98 100644 --- a/include/linux/workqueue.h +++ b/include/linux/workqueue.h @@ -618,6 +618,8 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(vo= id); void free_workqueue_attrs(struct workqueue_attrs *attrs); int apply_workqueue_attrs(struct workqueue_struct *wq, const struct workqueue_attrs *attrs); +int workqueue_set_affn_scope(struct workqueue_struct *wq, + enum wq_affn_scope affn_scope); extern int workqueue_unbound_housekeeping_update(const struct cpumask *hk); =20 extern bool queue_work_on(int cpu, struct workqueue_struct *wq, diff --git a/kernel/workqueue.c b/kernel/workqueue.c index 3c034cbc5bb3..0c2b8e87cd59 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -5610,6 +5610,43 @@ int apply_workqueue_attrs(struct workqueue_struct *w= q, return ret; } =20 +/** + * workqueue_set_affn_scope - change the affinity scope of an unbound work= queue + * @wq: the target unbound workqueue + * @affn_scope: the new scope, or %WQ_AFFN_DFL for the system default + * + * Reapply @wq's current attributes with only the affinity scope + * replaced, so the nice level, cpumask, and strict affinity the + * workqueue already carries survive. Pool-workqueue replacement + * proceeds as for apply_workqueue_attrs(). + * + * Context: Process context. Takes wq_pool_mutex and performs + * GFP_KERNEL allocations. + * + * Return: 0 on success and -errno on failure. + */ +int workqueue_set_affn_scope(struct workqueue_struct *wq, + enum wq_affn_scope affn_scope) +{ + struct workqueue_attrs *attrs; + int ret =3D -ENOMEM; + + if ((unsigned int)affn_scope >=3D WQ_AFFN_NR_TYPES) + return -EINVAL; + + mutex_lock(&wq_pool_mutex); + attrs =3D alloc_workqueue_attrs(); + if (attrs) { + copy_workqueue_attrs(attrs, wq->attrs); + attrs->affn_scope =3D affn_scope; + ret =3D apply_workqueue_attrs_locked(wq, attrs); + } + mutex_unlock(&wq_pool_mutex); + free_workqueue_attrs(attrs); + return ret; +} +EXPORT_SYMBOL_GPL(workqueue_set_affn_scope); + /** * unbound_wq_update_pwq - update a pwq slot for CPU hot[un]plug * @wq: the target workqueue --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E6D113A5430; Wed, 2 Sep 2026 19:29:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377364; cv=none; b=HoLkreqe2Nr2TanDme5PO8yoyjPxe/zVGBPdPPvIMkb/BgNZlIkWqvKv6vep9UzbYJ6f+YMQF640SetZOq0x8/AJKw+jHnwPwUo1tkr58WIj/2OzOGLqOQdyFe2N8nzObK51D2c/3AQaeNjrIoEIJZmXLF4TGTOdNqo86V53yf4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377364; c=relaxed/simple; bh=l59kSsqmbSZKGr3a44xK8gQJ6X/HbwJ7DhjoK/lpYxk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Dd6Qw5WMDclUTGMnJZ3xSDleUs8AQ0ylcdHxf3WDjz/ehfhS0VwRCVKoiZI4NesNtLUy7Fguys46rOORJbh53tLZHj++vJwmQ3ExBshNjmWbPA1nqY+BMiSHbieh5vHW6XNXkTVmX6yzfbj7dgLU/l8efLFy4p8CeqvMmbF7Gbk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=GvMZDG49; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="GvMZDG49" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A9BE61F00AC4; Wed, 2 Sep 2026 19:29:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377350; bh=F+29SWxzE93hqeHEyD2uBD9Y99JfUFtqyMqz2TIF/bU=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=GvMZDG49+gshP3N0gr7l7iWiIo8xAaIMSQhpHGdi1w7jJKn0oVfkTxdGUbqbHU2bt 7XSiBPAn0cPs3FOknNGg9EL/RCikuJV+C6xwB5IXnR8qv6o3/e0Akhlb2Of+RMcnea w5vqyT8YuvWPREY/spKtvFiHVqPq7yUu6ZSs5WnPioq6MGUkXjkmuMUruB7HKCwIkl Xxjw9GXmatyPR0+Ht6HYcnDZI82bUU91C5GVYf0jWvwNMniOeU1PA1GokwGiskw5ut Jh90el0iyOI0MmGBxWZYo4mS5nTfEbsW3qf71/OOmFdtsVVfIE/BhNzmpJLzkBrwz/ I0nz6NxgnmipQ== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:51 -0400 Subject: [PATCH RFC v2 6/8] SUNRPC: Reduce rpciod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-6-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1957; i=cel@kernel.org; h=from:subject:message-id; bh=l59kSsqmbSZKGr3a44xK8gQJ6X/HbwJ7DhjoK/lpYxk=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkA4sLiyTjxed1Izc24igQEahb7QkOSFINBz +X/iqeEX2CJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ lybuD/9tp8KGHNRFVt/2vYUlP6AunckvY/o4VXVQ3KXIiJL1RLPhi/Ly9jmyYrYRlM63Dn0d8vx VcVFVc+bkXIIv/1mLdecdpf7yAxVY9MCQc8z9fcwXAlkzuV8/onpWskafje89nekFATmVWSZKKN ynNMG40WEtKVfjUd3e0hunstTScjpYnTm3FKDEgGSbKzrGUbJtWTKeTV839PVTmKmsE8OGX/dSW w9NQcirpwpHC+5dzaC0y8TsU0Ii7U2TY/gCLuf4JlOmxEUs8bmmFN0LVIUwE3pkQ481bis0S7bI f8TNtLB11M/Hw1eVKVxgWJkTDjb4xLOZXFSFUyPN2L1ykwJeSB9a/SZkFPtgv6STf7ZJnmOuLQV y2qmYbNOt8TOKc3n7sJg3M+A6F1LrE3wV0k06XEE69nUGfCzFWYJ3281ji0WUYBw/3ag0Nz2WbS r6HNW8q4ZjIyw2U1KAxqiGacA2rOkHxev57RyBzZjuT9f10dXp9unPxpQMzgtgZj+HLDca4gyw3 Bb/hssu88VSex6mc0L4Us2LE1QfDFkof8ddVUsgKki1ZFpLNTxi8ap+i+OMzgUBcfRAPbpArH73 5pM5dO2w7xNMcc1dce6levRe0SPsZrXfF11wm2BTt0xN9hnsnBus4dPTwD9d32Q8UGS6ozQPrE0 v+19BAI7gTYebQA== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 rpciod drives the RPC client state machine. Under heavy NFS workloads, multiple CPUs queue RPC task completions concurrently and contend on the UNBOUND worker pool lock. perf profiles on a 12-core system show 30-40% of cycles lost to native_queued_spin_lock_slowpath in the rpciod pool at the WQ_AFFN_CACHE scope (one pool per LLC). The WQ_AFFN_CACHE_SHARD default helps little here, because its 8-core shards split this system into just two pools of six cores each. Set WQ_AFFN_SMT on rpciod so each SMT group gets its own pool and lock. Most UNBOUND workqueues never contend on the pool lock and profit from a coarser scope's cache locality. rpciod's sustained completion traffic makes the lock a first-order bottleneck, so the override belongs on this workqueue rather than in the system-wide default. The cost is one pool per SMT group, or per CPU on a system without SMT, and each pool keeps up to two idle kworkers rather than culling its last ones. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index e81419aa553c..2a1938b9e41e 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1268,6 +1268,17 @@ void rpciod_down(void) module_put(THIS_MODULE); } =20 +static void rpc_set_wq_smt_affinity(struct workqueue_struct *wq, + const char *name) +{ + int err; + + err =3D workqueue_set_affn_scope(wq, WQ_AFFN_SMT); + if (err) + pr_warn("%s: failed to set SMT affinity scope: %d\n", + name, err); +} + /* * Start up the rpciod workqueue. */ @@ -1282,6 +1293,7 @@ static int rpciod_start(void) wq =3D alloc_workqueue("rpciod", wq_flags, 0); if (!wq) goto out_failed; + rpc_set_wq_smt_affinity(wq, "rpciod"); rpciod_workqueue =3D wq; wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BA2A13939DA; Wed, 2 Sep 2026 19:29:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377374; cv=none; b=lMdRnjHYO7tIxNGovBmxnTLnWw7arh4RU1KRHVkhEkmxo9EmNpJBEBYrWKFbZPSWNkTt22b0JvULPWKr22QpyfN7FvVYca5Z/Ixph7mywKfoB5APMCRogHsF6J3WuEwR3g2Um2+VSA9C05fTfe61ge8h4PKhCFsbt3WOQTfgIBk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377374; c=relaxed/simple; bh=IteBc8RWWPlFxkX4q9XD157hPNKYOGmPLb5Z0/H5yeg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=lXrii2s0gSoMhh0xcYXukbwQttNMrLp9ppeYIR58D8fzQGe3Yc8LOf+mtpIESn4lV085RjXephjzLvqxvabO0qKL7DxPT3WwccXvUZrRLaAslAeTda6eeBBIZKbLqOqwFhUoG1hDNIZRODf4dAav7gMi2ovDtgTCCGKgaArIBrg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ARyYgqsv; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ARyYgqsv" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 727FD1F00ADE; Wed, 2 Sep 2026 19:29:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377351; bh=rfUk8VokEiAvKG3hrqEJQmi4sGGt5LkfdsOmziny2Zc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=ARyYgqsvyefh6IA6+W8xki5v5AccHNaVjGT6LGh1f2YBMmyyiB6Rj2wKtTpBZBPNy 3FEswVoqaQwtBGtm/RJ9f6UCJiBDb1+SoVamWqmsJesx3th65wcxnImxbxS45CJhhi 20CGjrrmes5ZciaAeLsxpd/mEs5c6ufzhIKtYTUEls9GLX+iJPuoonmnZuFPN69C/1 KWahyhu65CgE+tRwBJXezBGYmR3O9pPzgma9yHDpVWwhEr/jPnDOm9TL1TFzQJ+er8 yez3+w0W8KKEVF/eGkP8qTeCXyMS2YDiPFPpGyGjiJ35YXJj8TMseFStnFYegI0eFq Jl+48vEkywSqg== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:52 -0400 Subject: [PATCH RFC v2 7/8] NFS: Reduce nfsiod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-7-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1750; i=cel@kernel.org; h=from:subject:message-id; bh=IteBc8RWWPlFxkX4q9XD157hPNKYOGmPLb5Z0/H5yeg=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkA9qyELWPYcWMAO7bHXhZbE6fBpkgM4ochl njDI7anxP2JAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l/hHD/47yl8mayPhvFTdq6ccpyHXMrByHpditJiYadKrnGGb/Q7Dgg0vpVxf8JTy5H8NUufNCmy 4ItHDDi+LSPdyBDa5l4U155/Ucz69soY75EI4jML85+GYDyiIWeu9BETnJkC/FyMKswQHxFOadY fIJeDcrSKt7D19K8KzaHOPsuzNHeB6B44MSVxzulKoixc4C6XjoRhulTb19eiUynq+Fia5XGbDi xwZX4dj5hiUiae/iY3tk3s/jFFN0ihZdzqCL26CbcYi76XIcYu+Ya2dNfsdHuOq3xWu/YF/gU2+ 6eNJgx6+lXwKkVbyA3Q0qszDbs2lIr2EG3DFQ2iLXPrJCMFg4OOCC0Zdl5I2N+rnmD/KMKniSEB 2pPfQz/7fWDBVvUz/8BYId4hfk1Ut9/9mETpA1PD4q85DzH1td7Elj95GLyCK8sUGTTMB4gEnDX Ws4t4niROSdYRl8CZvpJZRgE1imk1+wK1axaqWFU0GbGeCFSd+7sofDChJRi5nyzYj3VThR18c7 zoyfqYvGUrdCXjjH0fAhK7hbi9OC1nsk/czKwQwJdXIpC/cSQEn6TOe4+NagebQZup8JQCD/51j JROdsgf8Gxrq8STHlrDhHLDHnv2r7UwVnPPgMSgxbG7idUFA4R7XD3BySqabmyDPSmXF3EcheFo Fc20xSgOE4V0pgg== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 The default affinity scope for unbound workqueues is now WQ_AFFN_CACHE_SHARD, which splits each LLC into shards of about eight cores. On a single-socket system whose LLC fits in one shard, every NFS I/O completion serializes on one nfsiod pool lock. Profiling 4KB random writes over NFSv3/RDMA with nconnect=3D3 shows that lock consuming 17% of CPU cycles: 8% dequeuing work and 9% enqueuing follow-on work from rpciod and nfsiod workers. Set nfsiod's affinity scope to WQ_AFFN_SMT so each SMT group gets its own pool and queue_work_on() contends only with sibling threads. Enqueue contention disappears and dequeue contention drops to 1.4%. Throughput is unchanged because the workload is transport-limited, but the freed cycles cut submission latency variance by 67% (slat stdev 31.6 us to 10.5 us), IOPS stdev by 31%, and p99.9 completion latency by 11%. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- fs/nfs/inode.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 107a2135029d..21c4560696bd 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -2618,11 +2618,16 @@ static void nfsiod_stop(void) */ static int nfsiod_start(void) { + int err; + dprintk("RPC: creating workqueue nfsiod\n"); nfsiod_workqueue =3D alloc_workqueue("nfsiod", WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS, 0); if (nfsiod_workqueue =3D=3D NULL) return -ENOMEM; + err =3D workqueue_set_affn_scope(nfsiod_workqueue, WQ_AFFN_SMT); + if (err) + pr_warn("nfsiod: failed to set SMT affinity scope: %d\n", err); #if IS_ENABLED(CONFIG_NFS_LOCALIO) /* * localio writes need to use a normal (non-memreclaim) workqueue. --=20 2.55.0 From nobody Sat Sep 26 09:19:12 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D5E303A6F04; Wed, 2 Sep 2026 19:29:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377360; cv=none; b=nmQkQVouav8YNXt7Gr4LSwM6DBb73Tg7dh2C/ymkubczi/zx6fkYhIeSHtjf4cpJRGWvbaiFmGCMAlzhDbDAd+CfKSEiCrHz0JAU9FCXJdKzM2awN+3tjWj9+FDqPTSdnwdpIQwFVEz+oLcpP7VEnLJo8hZYkd8xkldUIogbG24= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788377360; c=relaxed/simple; bh=3Jw7o41F/A1SDrtWghxIVFeAWgYLWjA00RvC+YvDL10=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=RaS/JOjqFQ6odh2MZadbc5cW/CmeNtKgbqI8rt6atp7oy42vMw8c891U8fDaQDUYbLPf/yHkJ6Bvf61hIkMTxX6WYefrGr32kTQR+H1V1n1l77uDLFoEg5rTT+nCOI7OYRp5vK9igMJ3XmbHzTu42+nnKNbLXITu568Sb3p1Ie8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=iDxGrzBh; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="iDxGrzBh" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 432A41F00ACF; Wed, 2 Sep 2026 19:29:11 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788377351; bh=8D76IdjG5RCkGqgJJYSjBkybiBKnaTAZ9CnZKgv9Dfk=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=iDxGrzBhrtV3et7EFjOYAdajuSriWFxS7fdB/ThScSmJINVacHODWrQO6PdntvRby BwVqsgq9oHUiNgfryrdSC60lKgl7dui+/u/4bp0dcB7vR6ozSlm/qVdG/zFAsxE5c4 ve6r2uGpC1irh3X//5MoncbGaVvrct6Fi2Zm2R1d23YqNY5g4/JOsreBytaHMKMecO fI5LSPS417iSE62PAPSjGxsQmXHSlEMQs4zaCFAQnbLTCcpNwaSODueO1//xGMIoPZ jtwE7DTWL9gC/ECk722IYEIeKnzZ2aGEZo4MNd0QfnUtvzc+Z4T4DHngW3ST0knOFk RSFaqHdON1RNQ== From: Chuck Lever Date: Wed, 02 Sep 2026 15:28:53 -0400 Subject: [PATCH RFC v2 8/8] SUNRPC: Reduce xprtiod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260902-performance-v2-8-b71c0c082f9d@kernel.org> References: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> In-Reply-To: <20260902-performance-v2-0-b71c0c082f9d@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=875; i=cel@kernel.org; h=from:subject:message-id; bh=3Jw7o41F/A1SDrtWghxIVFeAWgYLWjA00RvC+YvDL10=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqmHkATSl09RG7ifjXrwhFtzjpVY+n9p9D2pRqE lTPq/TDJcyJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCaph5AAAKCRAzarMzb2Z/ l9a0D/9V2sEgivDgG9qZnZXwL0oQXVmyPPsI03zXmS220U97QgMiQ0CDOtcfazyHqWeuDH66w3/ zxrVOgpFBu2vAlCcyBj1VdbAFAIEksXOO4FMf7pLmklI/7QYJCt1fHWWa3ifcutzVM2Run/le+9 eduY+7oNmaQvOnyf9u1aLLsfv2K4nNYWek6P98/OdXfqaLJPAHAvEbNro6vSbqEAAZ6/LV3tnPD gG2hf+eqx07grSqRUjuMu6Bf2Lay6GoiKeAVjU54PmX5tEgf6X+HPEF2mCcyalnAnBNr9nwCwlj hAaqel94xPeIDApHu8uNePP+zh1jvomKjj032zY1FDuc9IksvkWD9fM58fViM0M6k07UW9IP2YN C09UITnuk53n7VB7DSiJ8Izvf0qvyywkwoOJF7HX0boGmVsKp0Vn2hr4FgcZ6P5rO3vbam8ERVV TSlktofh3YHfFW5Hd9CwkgrcPKmKhJ/o/3tKLda6zq2GydM6FTySlvVrwijXikXvzJhRxk1p6VI 2++rNclefqVBreuhjb/t6bKIZ8jZCa9YX9KS2zBna0JVK+Z4O5dHhDQOOPL2md3wsRNFSCRhhji K6OHo/i0Szbnu2GEhh6YnI2APqZJFRFMizUmmIPsq9AjNU/fNSKvtFqRNacjJa9ylmAOJZy8Afo vnOcWpg6D4+Fs/A== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprtiod handles transport-level operations: socket receive processing, error recovery, and connection lifecycle. On systems driving heavy NFS traffic, these operations contend on the UNBOUND worker pool lock just as rpciod does. Set WQ_AFFN_SMT on the xprtiod workqueue to give each SMT group its own pool and lock. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 1 + 1 file changed, 1 insertion(+) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index 2a1938b9e41e..e678326e6b0d 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1298,6 +1298,7 @@ static int rpciod_start(void) wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) goto free_rpciod; + rpc_set_wq_smt_affinity(wq, "xprtiod"); xprtiod_workqueue =3D wq; return 1; free_rpciod: --=20 2.55.0