From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 563F8471279; Mon, 31 Aug 2026 18:22:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200530; cv=none; b=msplIMFFnx6Ry/eobw70F5mSYEjO4MTGOMIHXLEQ2farzCLJQWdYV6UEaLV3ThpxZiiNvqKuEVBgSFtjR1gqfyG+sFj4dlmDhOqX/EeviIju66QG7sbMLeA7lnSpwDRiWVqEtfe7Svt9CP6Yr7MAOK9UWNneUXOr/UaZE84LT90= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200530; c=relaxed/simple; bh=ab7jO6lz+XT+Bcp3r5CMowTvoccAdGfKY8J7aGSj2o8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=u/x6K1trxHTigNS6sD9dAZ2zNg5BV+6i0627hR6reHT2BBgaqlasQU4rQsCx98w8POmIe6IlzHrBOYErG3307yJxHnshKsoTqqbM+FnbcdjhNNysQm8UfGr+EBFC47IAhd+dEMIr8Reo0dDjJpn1nuqQb1fi3/ol7Gi0YgKwiTc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bReXodlq; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bReXodlq" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A4BE21F00A3F; Mon, 31 Aug 2026 18:22:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200529; bh=UntNAVVE5PcFzEZfDLpbTIVU+V+Nb8NqweMsWPYgcwc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=bReXodlq25mp/gElXliLrUsPb+wRRGhh1X8ya5lkUo418ChxWXgJJpQZVncJ5zMgm 8nHl5KFEmGFNtCzZy/UY3zllU7GSXfM95PljV/BOTo74uoe89qc8HMSISVJXb5QW7A XhiqgU5yiSQ/N95QMeTObstW3SrFvtxiQyEPdHcUW/yaaSW2+RaYKN4NqCXl2i9Beo heXnY+OHfFEUND2IJZw2ODK4h9UhYhhYCUgquDogys7NygoFdEiLOCbPS1S4BFIxYQ id/VxzPV50wc4PUwMQzguIL0SFSc2oAfgrFw5vK2qpfy7Uek4g0c/lupZt5FpAv5fU HPiQ3Gw0UkJlQ== From: Chuck Lever Date: Mon, 31 Aug 2026 14:21:57 -0400 Subject: [PATCH RFC 1/8] SUNRPC: Use atomic_t for XID allocation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-1-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1882; i=cel@kernel.org; h=from:subject:message-id; bh=ab7jO6lz+XT+Bcp3r5CMowTvoccAdGfKY8J7aGSj2o8=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPctMWEWCC5MfvrJN8bSSrTeVCWbkEI1pj8 OzYvj4qjb2JAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ lx+YD/kB6vVbL7dpS/6cajJoFDwTigMndT/3/XvsiUD4N1cCzPlDce1+UVpFaI9qKekfHVgJedc cAR9TEsl4VEpUQlH1skk1AXPtvS7wU5nUEMMKadZqewVOwQ9dlzLXKptkcPfq05avbjz/H6nqHQ 3oDGpbzTj1opbVTF6+ShfVSI2w10hpoDub1WgL7VfsxNKVftKc3XVFab63L0bcYcUl1njjp6v1C jHO9wHwsY8AnMF2oN/W+fSimiEwvVWVFPo1ZCMgEKhw+0oceBTHfYHv/obiCcjPuY5A2OUsHqX5 UlkquBq0JkXa2JIw49s1vKOVDqAEYcJw9Cb8ptb67I9zYI0a+6hFP6CiDNr4y4reAug0ZRjZ35o w3DJAN/2Skj1pkypq/6iF6rxZna8DxY5h3Mo2h7XxUTabVQaMwDj6hWAdDt00Qar1wE7QAJNle2 581p+Vlvrn/anXX1/dEC4bSwta5VcqvuS+cam41yFwWNICCBJDDypkiiWNRslpBqsTqPjrWtBbv FU+rWzHscGdOsqzoCgK7oBwh08LTDj05NDBJPioPyb+7fio+aGKY6z5F1ZqlQQ8ZFwwwLCy4C28 iiACbkU9D7+UF31nb1Rokd8IkgIN2yIM78IFS9d+EYJw7CcFmarcznAxgpFkNKEGoOVoOSTo/dC nHiQ41DwpyyRDlg== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprt_alloc_xid() acquires reserve_lock to increment a simple counter. Under a high-IOPS NFSv3 workload on 100GbE RDMA, profiling shows 1.06% of system-wide CPU cycles contending on this lock in xprt_request_init, as ~150 RPC worker threads serialize on the counter. reserve_lock protects the slot table and backlog queue, but XID allocation is an independent operation that does not require synchronization with either. Signed-off-by: Chuck Lever --- include/linux/sunrpc/xprt.h | 2 +- net/sunrpc/xprt.c | 9 ++------- 2 files changed, 3 insertions(+), 8 deletions(-) diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h index a82045804d34..0d6c3f6bf97e 100644 --- a/include/linux/sunrpc/xprt.h +++ b/include/linux/sunrpc/xprt.h @@ -273,7 +273,7 @@ struct rpc_xprt { spinlock_t transport_lock; /* lock transport info */ spinlock_t reserve_lock; /* lock slot table */ spinlock_t queue_lock; /* send/receive queue lock */ - u32 xid; /* Next XID value to use */ + atomic_t xid; /* Most recently issued XID */ struct rpc_task * snd_task; /* Task blocked in send */ =20 struct list_head xmit_queue; /* Send queue */ diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c index 48a3618cbb29..186c14f0f928 100644 --- a/net/sunrpc/xprt.c +++ b/net/sunrpc/xprt.c @@ -1882,18 +1882,13 @@ xprt_init_connect_cookie(struct rpc_rqst *req, stru= ct rpc_xprt *xprt) static __be32 xprt_alloc_xid(struct rpc_xprt *xprt) { - __be32 xid; - - spin_lock(&xprt->reserve_lock); - xid =3D (__force __be32)xprt->xid++; - spin_unlock(&xprt->reserve_lock); - return xid; + return (__force __be32)atomic_inc_return(&xprt->xid); } =20 static void xprt_init_xid(struct rpc_xprt *xprt) { - xprt->xid =3D get_random_u32(); + atomic_set(&xprt->xid, get_random_u32()); } =20 static void --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2898C4A3D5A; Mon, 31 Aug 2026 18:22:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200531; cv=none; b=ghIn5q7Km/sin0jywtVXo6axJa9a0zhLu8JWOrAoN6DoRSaM0tJvCyEvgaJ3GycYVI1lk4rY0NveifAR7KRXVCwYBgS6jZachLgqYeoJU8P/9DQeXCCFh4Z2PsjICM2oKdxOUgflNSuCadGEePe1x/OGoO+G/3/kntWy+Xh0Lxk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200531; c=relaxed/simple; bh=NMoveCuj/U3vFYbbxpy7pgcMhDCKjoWIC99DsA8Pox8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FiU2cRtRI6g/OdV6sUcCvswE6+H3NYz0EoaaxkwxamckNFkSvg/Qsqe0yOTeoQ36uYblu5cqlQsl9ubzC9Hsz4odV/8Pr7BveBmIvf4jH+FbNEayFyv/IPMyKE4cMpQK3nzAZGIRcG+SgpdSbX/goa4n+mJ/HRSZH220ukjmwss= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=btEKD86j; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="btEKD86j" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 69EB71F00ACA; Mon, 31 Aug 2026 18:22:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200530; bh=Pu4GPT9xCofiWopDGkZSuP+4u3elKLBzQ0R5BEM+3Fg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=btEKD86jgNOsJL0oKrub4fuY4KwHZw9+y5aQNynDhNnAwDVKq5pB5r8rA4WQP7fWw YIJ3mPXStwIUvuEMwK+5md4pvvytTBP8NinMwwy5w49tYazsxVWJMEU4sDOZtG2eSF TYRMYgR67dhU2eWzPSPx1gIDU18gGKmOzyr6yyYO1LvWtJTyMMgvEahm/wtZyuHJ9Y NPACmCSCd/MrZczU7NFhBUN9ahJHTEFkmqP7VYcKX2nDvLBmaU67Udot8O7cTY+3Ym CoDCyJ0IYq+rARxTx3IsWXB4ozAat9pjEX9Xm1rKsD3GmfmBsC8KtquSXl9abdB+JA PZKyTGj5gF5rA== From: Chuck Lever Date: Mon, 31 Aug 2026 14:21:58 -0400 Subject: [PATCH RFC 2/8] SUNRPC: Execute initial async RPC states in caller's context Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-2-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=2178; i=cel@kernel.org; h=from:subject:message-id; bh=NMoveCuj/U3vFYbbxpy7pgcMhDCKjoWIC99DsA8Pox8=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPOC0GWRQVNLqFBploRctZQAImd8K0Laag/ SbEtcrdA1iJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ l+spD/9mVO49lQZzoGwus61XbLGxoq5OZrEgMGEM1vdSGGtjN4/zD3ihSi0utsfD49cWoQIBk7L 6vzqdZhkPjA243jpnAbtbwjyLm2r7q3JwCtZvkJ/f4rI/tmV6cQPmFG+i+VRo+AtMytBEV5Qvc5 Aisl6cN0bmgYYyQNXrS4jGD6Tr3reu7VgXJQ2Jy6C2zcpISUqQtZtpwBIEYahdo7NSkpYUUSNKb e74q9tnXIm4I9dt/L0jH94DL3BPoCFr3V/2RSinLSRLzurZKXO2668NtT7sHdFjg6xkIJdIEaE+ P4zzmZ8MEKkin3klXi1wOGOvBy/5GqA+NEiHJMiyl8zqfIGHeTMPZ1bGMjDWV9NEZZgv/SgbWC7 SLv7iT7L4aDUBNdbnAg4oq1Y/V6t1Cute5TtSiK3a8/Xx3IIdHv45oCxY+5zJ12uPFE9hPcSnQN 0CcrAiiaffZ4zhg09xrInEDaYr6YwbSkCjdf3ZTD3IT2wPZwOVAeoZVWmKls8SzbkmCDGZ4NGlo Y0MnaCK5bfpJJcXhIiO5A4ddZh14dlz3qh72RTiV4EWdb5cRG2zak59PPZ5vvABBLZCwJKW4est gnxfDJDjS/idV5FoxDE9+80mX6RljdMATQkIquxVjLfrVDjjcTyzwOynuwj93UkQVNaotC8sutu tis7w6m/z2aFObw== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 Each async RPC task is dispatched to the rpciod workqueue before its first FSM state runs, though the initial states (call_start through call_allocate) typically complete without blocking. Under high concurrency this adds a workqueue enqueue, dequeue, and context switch per RPC and contributes to rpciod pool lock contention. Have rpc_execute() call __rpc_execute() directly for async tasks as well. The initial states then run in the caller's context, and __rpc_execute() returns to the caller when the task first sleeps on a wait queue. Later wake-ups still dispatch to rpciod via rpc_make_runnable(). Commit d6a1ed08c6ac ("SUNRPC: Reduce asynchronous RPC task stack usage") moved async dispatch onto rpciod to bound caller stack depth. Inline execution preserves that bound because __rpc_execute() still runs each state from the same stack frame. Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 16 +++++++--------- 1 file changed, 7 insertions(+), 9 deletions(-) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index 016f16ca5779..9b2800f6bfc7 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1003,8 +1003,9 @@ static void __rpc_execute(struct rpc_task *task) current_restore_flags(pflags, PF_MEMALLOC); } =20 -/* - * User-visible entry point to the scheduler. +/** + * rpc_execute - Consumer entry point to the RPC scheduler + * @task: RPC task to be scheduled * * This may be called recursively if e.g. an async NFS task updates * the attributes and finds that dirty pages must be flushed. @@ -1014,15 +1015,12 @@ static void __rpc_execute(struct rpc_task *task) */ void rpc_execute(struct rpc_task *task) { - bool is_async =3D RPC_IS_ASYNC(task); + unsigned int pflags =3D memalloc_nofs_save(); =20 rpc_set_active(task); - rpc_make_runnable(rpciod_workqueue, task); - if (!is_async) { - unsigned int pflags =3D memalloc_nofs_save(); - __rpc_execute(task); - memalloc_nofs_restore(pflags); - } + rpc_test_and_set_running(task); + __rpc_execute(task); + memalloc_nofs_restore(pflags); } =20 static void rpc_async_schedule(struct work_struct *work) --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 21E8F4A3F13; Mon, 31 Aug 2026 18:22:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200532; cv=none; b=ELmsQ1L+pgTqzbCEN6yHqKIL7om86nm68GHbsnBqUyMbstJUhJM2B+iloQPBaNKFHK3nWCPFOLIt7bjFBhSmHjZ30kAHgO+6SUDWWEOXoltWnnW4z3nNR3FfgQIIn+y4peHwwPikSev/82oWCsT1NhcSLT1uiuzRmiqNlJ0A/fs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200532; c=relaxed/simple; bh=vVF+1nLfEsYCYEe7bJXvQYnkeffWIDla56CaCSEaT38=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=IFtUUcB1I8ssxmwESn37FHVxiaw6JTI2MPiN5XUrNWY05uHXPcWz+hg7L+dk7xQ9A5CS66liPvTzuQPsN24IINoKTngNcj2c0cblO+u29ylAwvjwCp50tfhoZ0FT2HfDClGtSdnIURgtULa2ibWRORHjTHIYZqlzLoyT0W+g1mM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=jknZyUO0; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="jknZyUO0" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2E0A51F000E9; Mon, 31 Aug 2026 18:22:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200530; bh=glbK/Pyh1W0SQOn60xP0+QoSWP4AQc/SXX426gdzA+8=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=jknZyUO0ypQB/AVtXXsD4Db9PVskgRKnP6OPUZNAn+iNspvJI7O6XVi+e5Bt6r4xy FcM0nAymbGUovDlbJLVG/GSXW7BkOF5lo5QkD4k4InGGSZXI5Y7EcTDzKJvRGhQreR sVK7/y8NZSLBZ6l3hU8ghAXt+r7/R9RNdz/l4eBE1inNJaqRnLw7wD2NwDol0ehfJ5 0dEgAgH1f13b9MLlU/Y33Z+3aOT+QDGJiayctK2Lib+2fHuZ3O2n7M+VU7RIUDJPsF BJ4UARXbQQQfxEHLj2tkLSDGbhsdTTNdKgPuoxCoBhm3kpg4p00cm7+ZGNRG8pwAqM tXy5ryFih1zFQ== From: Chuck Lever Date: Mon, 31 Aug 2026 14:21:59 -0400 Subject: [PATCH RFC 3/8] SUNRPC: Split recv_lock out of xprt->queue_lock Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-3-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=14272; i=cel@kernel.org; h=from:subject:message-id; bh=vVF+1nLfEsYCYEe7bJXvQYnkeffWIDla56CaCSEaT38=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPgPjC1te6pqQ9vWJE3vSbVY8ZL0j/Rr3fe yMgGcEdxpmJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ l9CGEACzdUlO41//NKheRpD7LlfD8KDYFEYMP7NrLXkssTFdaMDUhiKapMp9rGiQSyMURYVTV0T Tq9qqQH3sVGOes3Tzl0zZehdv1JFaM23S9YDPGBdBlmH4T0ZFIBkIyrDkP34X8MWiLB1RUa5TH2 FFtoOLHP6+6K7YLxoyEWikQ12uD86RwcVKzQwsXEmhY+Jp/RAATDupVkJNVL6pVKielDsL6pXRZ QyeFPpPBaKL0gw2u8119Es0a54/ypo72cCPCcA9EESQoInu6a1s+8Jhni/CpfSvCjyjP+WNAzGQ 655QD5hjxVd890isqiwxwL8RQRdvH3u5iHFUDt0vNv6MfZ+GOBeJvFhfz7s6eZgTiJhFL8svK0V IDYlTwA13xxWT9zgTLQlxBm7CZs17W2gZKzwMX6gi+fDaWuyn4wABeK2rc/B/0DU3SZ70LpPLWF kMYddtwYDmd7wIv1DsoZHNC7xivHZzBpWLanclU1gADWLI2jrE1tVBYiQm7IlJxvjE+LIr+SSxP ki8Vbd5n2De9oKJ4PG0IFqmBc6AONunDX1xlOYM9+oA8qLAqtZMYMdZQGmyaBaC7ak2V0jyMwa1 oqyapyl/TrJPf2OYVtaU9CLAs97iRlWmzUIW1kaGc/zKD12/98Kcm82L69hZtH4Jn3g1FdJZGYq xFvWwRqPIF/nEoA== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprt->queue_lock protects two independent structures: the recv_queue rb-tree for reply matching and the xmit_queue list for transmit draining. No hot path touches both in one critical section, yet every RPC submit and completion contends on the same lock. Under a 4KB NFSv3 READ workload on 100GbE RDMA, 53% of non-idle CPU cycles are spent in native_queued_spin_lock_slowpath: the CQ completion worker running rpcrdma_reply_handler serializes against ~150 kworker threads enqueuing receives and transmits. Introduce xprt->recv_lock for the receive path -- recv_queue operations, request lookup, receive-side pinning, and completion -- leaving queue_lock to the xmit_queue and the xprt_transmit drain loop. Also move the rq_private_buf memcpy in xprt_request_enqueue_receive above the lock acquisition: until the rb-tree insert publishes the request, the reply handler cannot see it, so the copy is safe unlocked and the submitter's critical section shrinks to the insert alone. Signed-off-by: Chuck Lever --- include/linux/sunrpc/xprt.h | 6 +++- net/sunrpc/svcsock.c | 6 ++-- net/sunrpc/xprt.c | 54 ++++++++++++++++----------= ---- net/sunrpc/xprtrdma/rpc_rdma.c | 14 ++++---- net/sunrpc/xprtrdma/svc_rdma_backchannel.c | 8 ++--- net/sunrpc/xprtsock.c | 18 +++++----- 6 files changed, 57 insertions(+), 49 deletions(-) diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h index 0d6c3f6bf97e..ed1e28b74f02 100644 --- a/include/linux/sunrpc/xprt.h +++ b/include/linux/sunrpc/xprt.h @@ -272,7 +272,7 @@ struct rpc_xprt { atomic_long_t queuelen; spinlock_t transport_lock; /* lock transport info */ spinlock_t reserve_lock; /* lock slot table */ - spinlock_t queue_lock; /* send/receive queue lock */ + spinlock_t queue_lock; /* send queue lock */ atomic_t xid; /* Most recently issued XID */ struct rpc_task * snd_task; /* Task blocked in send */ =20 @@ -292,6 +292,10 @@ struct rpc_xprt { * backchannel rpc_rqst's */ #endif /* CONFIG_SUNRPC_BACKCHANNEL */ =20 + /* + * Receive stuff + */ + spinlock_t recv_lock; /* receive queue lock */ struct rb_root recv_queue; /* Receive queue */ =20 struct { diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c index 50e5e7f5b762..8939ba604385 100644 --- a/net/sunrpc/svcsock.c +++ b/net/sunrpc/svcsock.c @@ -1102,7 +1102,7 @@ static int receive_cb_reply(struct svc_sock *svsk, st= ruct svc_rqst *rqstp) =20 if (!bc_xprt) return -EAGAIN; - spin_lock(&bc_xprt->queue_lock); + spin_lock(&bc_xprt->recv_lock); req =3D xprt_lookup_rqst(bc_xprt, xid); if (!req) goto unlock_eagain; @@ -1120,10 +1120,10 @@ static int receive_cb_reply(struct svc_sock *svsk, = struct svc_rqst *rqstp) memcpy(dst->iov_base, src->iov_base, src->iov_len); xprt_complete_rqst(req->rq_task, rqstp->rq_arg.len); rqstp->rq_arg.len =3D 0; - spin_unlock(&bc_xprt->queue_lock); + spin_unlock(&bc_xprt->recv_lock); return 0; unlock_eagain: - spin_unlock(&bc_xprt->queue_lock); + spin_unlock(&bc_xprt->recv_lock); return -EAGAIN; } =20 diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c index 186c14f0f928..42c66464d8f3 100644 --- a/net/sunrpc/xprt.c +++ b/net/sunrpc/xprt.c @@ -1061,7 +1061,7 @@ xprt_request_rb_remove(struct rpc_xprt *xprt, struct = rpc_rqst *req) * @xprt: transport on which the original request was transmitted * @xid: RPC XID of incoming reply * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ struct rpc_rqst *xprt_lookup_rqst(struct rpc_xprt *xprt, __be32 xid) { @@ -1092,8 +1092,9 @@ xprt_is_pinned_rqst(struct rpc_rqst *req) * xprt_pin_rqst - Pin a request on the transport receive list * @req: Request to pin * - * Caller must ensure this is atomic with the call to xprt_lookup_rqst() - * so should be holding xprt->queue_lock. + * Caller must hold the lock that protects the queue through which + * it found the request: xprt->recv_lock for the receive path, + * xprt->queue_lock for the transmit drain path. */ void xprt_pin_rqst(struct rpc_rqst *req) { @@ -1105,14 +1106,10 @@ EXPORT_SYMBOL_GPL(xprt_pin_rqst); * xprt_unpin_rqst - Unpin a request on the transport receive list * @req: Request to pin * - * Caller should be holding xprt->queue_lock. + * Caller holds the lock it held for the matching xprt_pin_rqst(). */ void xprt_unpin_rqst(struct rpc_rqst *req) { - if (!test_bit(RPC_TASK_MSG_PIN_WAIT, &req->rq_task->tk_runstate)) { - atomic_dec(&req->rq_pin); - return; - } if (atomic_dec_and_test(&req->rq_pin)) wake_up_var(&req->rq_pin); } @@ -1155,16 +1152,16 @@ xprt_request_enqueue_receive(struct rpc_task *task) ret =3D xprt_request_prepare(task->tk_rqstp, &req->rq_rcv_buf); if (ret) return ret; - spin_lock(&xprt->queue_lock); - - /* Update the softirq receive buffer */ + /* Reply handlers cannot find the request until the rb-tree + * insert below publishes it, so the copy needs no lock. + */ memcpy(&req->rq_private_buf, &req->rq_rcv_buf, sizeof(req->rq_private_buf)); =20 - /* Add request to the receive list */ + spin_lock(&xprt->recv_lock); xprt_request_rb_insert(xprt, req); set_bit(RPC_TASK_NEED_RECV, &task->tk_runstate); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 /* Turn off autodisconnect */ timer_delete_sync(&xprt->timer); @@ -1175,7 +1172,7 @@ xprt_request_enqueue_receive(struct rpc_task *task) * xprt_request_dequeue_receive_locked - Remove a request from the receive= queue * @task: RPC task * - * Caller must hold xprt->queue_lock. + * Caller must hold xprt->recv_lock. */ static void xprt_request_dequeue_receive_locked(struct rpc_task *task) @@ -1190,7 +1187,7 @@ xprt_request_dequeue_receive_locked(struct rpc_task *= task) * xprt_update_rtt - Update RPC RTT statistics * @task: RPC request that recently completed * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ void xprt_update_rtt(struct rpc_task *task) { @@ -1212,7 +1209,7 @@ EXPORT_SYMBOL_GPL(xprt_update_rtt); * @task: RPC request that recently completed * @copied: actual number of bytes received from the transport * - * Caller holds xprt->queue_lock. + * Caller holds xprt->recv_lock. */ void xprt_complete_rqst(struct rpc_task *task, int copied) { @@ -1309,7 +1306,7 @@ void xprt_request_wait_receive(struct rpc_task *task) * The spinlock ensures atomicity between the test of * req->rq_reply_bytes_recvd, and the call to rpc_sleep_on(). */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); if (test_bit(RPC_TASK_NEED_RECV, &task->tk_runstate)) { xprt->ops->wait_for_reply_request(task); /* @@ -1321,7 +1318,7 @@ void xprt_request_wait_receive(struct rpc_task *task) rpc_wake_up_queued_task_set_status(&xprt->pending, task, -ENOTCONN); } - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 static bool @@ -1439,7 +1436,11 @@ xprt_request_dequeue_transmit(struct rpc_task *task) * @task: pointer to rpc_task * * Remove a task from the transmit and receive queues, and ensure that - * it is not pinned by the receive work item. + * it is not pinned by any concurrent work item. + * + * Dequeuing from both queues prevents new pins: xprt_lookup_rqst + * and xprt_transmit can no longer find the request. The wait for + * in-flight pins to drain then needs neither lock. */ void xprt_request_dequeue_xprt(struct rpc_task *task) @@ -1451,16 +1452,18 @@ xprt_request_dequeue_xprt(struct rpc_task *task) test_bit(RPC_TASK_NEED_RECV, &task->tk_runstate) || xprt_is_pinned_rqst(req)) { spin_lock(&xprt->queue_lock); - while (xprt_is_pinned_rqst(req)) { + xprt_request_dequeue_transmit_locked(task); + spin_unlock(&xprt->queue_lock); + + spin_lock(&xprt->recv_lock); + xprt_request_dequeue_receive_locked(task); + spin_unlock(&xprt->recv_lock); + + if (xprt_is_pinned_rqst(req)) { set_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); - spin_unlock(&xprt->queue_lock); xprt_wait_on_pinned_rqst(req); - spin_lock(&xprt->queue_lock); clear_bit(RPC_TASK_MSG_PIN_WAIT, &task->tk_runstate); } - xprt_request_dequeue_transmit_locked(task); - xprt_request_dequeue_receive_locked(task); - spin_unlock(&xprt->queue_lock); xdr_free_bvec(&req->rq_rcv_buf); } } @@ -2038,6 +2041,7 @@ static void xprt_init(struct rpc_xprt *xprt, struct n= et *net) spin_lock_init(&xprt->transport_lock); spin_lock_init(&xprt->reserve_lock); spin_lock_init(&xprt->queue_lock); + spin_lock_init(&xprt->recv_lock); =20 INIT_LIST_HEAD(&xprt->free); xprt->recv_queue =3D RB_ROOT; diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c index 1285f04cdac1..a82d3d9bc7ae 100644 --- a/net/sunrpc/xprtrdma/rpc_rdma.c +++ b/net/sunrpc/xprtrdma/rpc_rdma.c @@ -1321,9 +1321,9 @@ void rpcrdma_unpin_rqst(struct rpcrdma_rep *rep) req->rl_reply =3D NULL; rep->rr_rqst =3D NULL; =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_unpin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 /** @@ -1363,10 +1363,10 @@ void rpcrdma_complete_rqst(struct rpcrdma_rep *rep) goto out_badheader; =20 out: - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(rqst->rq_task, status); xprt_unpin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); return; =20 out_badheader: @@ -1492,12 +1492,12 @@ void rpcrdma_reply_handler(struct rpcrdma_rep *rep) /* Match incoming rpcrdma_rep to an rpcrdma_req to * get context for handling any incoming chunks. */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); rqst =3D xprt_lookup_rqst(xprt, rep->rr_xid); if (!rqst) goto out_norqst; xprt_pin_rqst(rqst); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 if (buf->rb_credits !=3D credits) rpcrdma_update_cwnd(r_xprt, credits); @@ -1524,7 +1524,7 @@ void rpcrdma_reply_handler(struct rpcrdma_rep *rep) return; =20 out_norqst: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); trace_xprtrdma_reply_rqst_err(rep); rpcrdma_rep_put(buf, rep); goto out_post; diff --git a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c b/net/sunrpc/xprtrd= ma/svc_rdma_backchannel.c index e5a78b761012..3c7b85427f33 100644 --- a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c +++ b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c @@ -28,7 +28,7 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, struct rpc_rqst *req; u32 credits; =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); req =3D xprt_lookup_rqst(xprt, *rdma_resp); if (!req) goto out_unlock; @@ -39,7 +39,7 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, goto out_unlock; memcpy(dst->iov_base, src->iov_base, src->iov_len); xprt_pin_rqst(req); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 credits =3D be32_to_cpup(rdma_resp + 2); if (credits =3D=3D 0) @@ -50,13 +50,13 @@ void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp, xprt->cwnd =3D credits << RPC_CWNDSHIFT; spin_unlock(&xprt->transport_lock); =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(req->rq_task, rcvbuf->len); xprt_unpin_rqst(req); rcvbuf->len =3D 0; =20 out_unlock: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 /* Send a reverse-direction RPC Call. diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c index 7f60723fa64d..1454da9575b3 100644 --- a/net/sunrpc/xprtsock.c +++ b/net/sunrpc/xprtsock.c @@ -673,25 +673,25 @@ xs_read_stream_reply(struct sock_xprt *transport, str= uct msghdr *msg, int flags) ssize_t ret =3D 0; =20 /* Look up and lock the request corresponding to the given XID */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); req =3D xprt_lookup_rqst(xprt, transport->recv.xid); if (!req || (transport->recv.copied && !req->rq_private_buf.len)) { msg->msg_flags |=3D MSG_TRUNC; goto out; } xprt_pin_rqst(req); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); =20 ret =3D xs_read_stream_request(transport, msg, flags, req); =20 - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); if (msg->msg_flags & (MSG_EOR|MSG_TRUNC)) xprt_complete_rqst(req->rq_task, transport->recv.copied); else req->rq_private_buf.len =3D transport->recv.copied; xprt_unpin_rqst(req); out: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); return ret; } =20 @@ -1398,13 +1398,13 @@ static void xs_udp_data_read_skb(struct rpc_xprt *x= prt, return; =20 /* Look up and lock the request corresponding to the given XID */ - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); rovr =3D xprt_lookup_rqst(xprt, *xp); if (!rovr) goto out_unlock; xprt_pin_rqst(rovr); xprt_update_rtt(rovr->rq_task); - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); task =3D rovr->rq_task; =20 if ((copied =3D rovr->rq_private_buf.buflen) > repsize) @@ -1412,7 +1412,7 @@ static void xs_udp_data_read_skb(struct rpc_xprt *xpr= t, =20 /* Suck it into the iovec, verify checksum if not done by hw. */ if (csum_partial_copy_to_xdr(&rovr->rq_private_buf, skb)) { - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); __UDPX_INC_STATS(sk, UDP_MIB_INERRORS); goto out_unpin; } @@ -1421,13 +1421,13 @@ static void xs_udp_data_read_skb(struct rpc_xprt *x= prt, spin_lock(&xprt->transport_lock); xprt_adjust_cwnd(xprt, task, copied); spin_unlock(&xprt->transport_lock); - spin_lock(&xprt->queue_lock); + spin_lock(&xprt->recv_lock); xprt_complete_rqst(task, copied); __UDPX_INC_STATS(sk, UDP_MIB_INDATAGRAMS); out_unpin: xprt_unpin_rqst(rovr); out_unlock: - spin_unlock(&xprt->queue_lock); + spin_unlock(&xprt->recv_lock); } =20 static void xs_udp_data_receive(struct sock_xprt *transport) --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9F7F84A3F1F; Mon, 31 Aug 2026 18:22:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200533; cv=none; b=QVAHFyQNlVMjQuTeLQxI5CQZtukG/D978pimCGgaD5O55bEbhXOB372o+cLBM6GHj8Tn4T7i187FDhgr3jfZbvcNWvJz7c++iZlIqcQp9k9ZkvoNvBUgrssfKfBSKPDv9Ko+Kncny+V7sWm22ozcyFxtWdrdmsssxCy7puZyM2Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200533; c=relaxed/simple; bh=tSL0uX5WiMXqa9KcVZAW5G1vvE7oLcVy4eyodGJqHX8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OAl+8HxQ9O91lySoFyVgSsMXYaohhb1VxCDeQqmYTGCAU7ZannmkcIv9MrxgjAzQ7qLxFUUq43R66r1Kmq4WyHdc4yfNCE7+spHUe8HVTOnCXMErjpx9fsz5Sa+d6LLVouS75cjl3BY7rWCmcqxN+eyHSCPlENaYhFisvhWyyO4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PDP0n0BO; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PDP0n0BO" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 00D841F00A3F; Mon, 31 Aug 2026 18:22:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200531; bh=/8h3ZyhxcBc9+MXGp+kCBkAD/N+CQiquY4vElNXs4Sg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=PDP0n0BOsT6lp9ONvnDCVMuRExMSo0EBlKmp20JVBFHkv88FsWgYGbJKQxf7x9aFC +kRlRxpSuTFoko0U8RdID4TPDQmWTA9HAl/F4/SipEKnwy0FEwrOwE6RuCjfzQtd0X jLVFnaHrKOjuCtR3utQYUckZaq0Q6rBGuW2MfLji3Pu0GszXWXvvUO0fyemCZm5fsz srumqCGSAqAq+ncYheOAg1G1TDA1FOR76phj9tLwOF9GQtniY9mRSaC4D4d4Cs5yVP a58kdqXIoTB2AIRVAELuTbV5P5hvwZZ2pA+yqwdwGFS5Vc/oeUvTN8cdNeIu+Kypiz ygCaKPW9CL+vg== From: Chuck Lever Date: Mon, 31 Aug 2026 14:22:00 -0400 Subject: [PATCH RFC 4/8] Set WQ_SYSFS on key NFS-related workqueues Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-4-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1677; i=cel@kernel.org; h=from:subject:message-id; bh=tSL0uX5WiMXqa9KcVZAW5G1vvE7oLcVy4eyodGJqHX8=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPUk2WAUnwrhHYs0lkjHwtVJSUin5p4QXyZ MaVjHYBURaJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ lwvSD/43BvorEcNglmPSncbfa79si+LpJ1OF+6XfuUqJPkfMTRojinLnugM/ePaOe79BYOuFWxW lhkzXwT2ve74Geue08TTZzLrCUAv0tCFcTXkrhgeT4DlZ0PywZr5r9D2pwbf4aObU+pRi3n2GRa YBRfDGWYPYrlONUEawGaAV6/Yz245zfkJniSPpVhmxitM08hWptoPUjea6X7pjqa2XFzoCPSfOw 5qY0lmVo7AApzNKNlJfK60vwML5bT8jQgVgBJCXuvNVsnVH+1Sv1lM9axjr2rKSTTWywyheQyeX cujnGIJKj0Me8ONzoX5OzPJyQB7JobD2/Y0cEAqIFgIOMljtu9WEbBHok2oT9v8kwUW/EXXFAX8 6OSiS+WgyS+wtIaL308rchVKvS17Gyu/QGpTRYXsQnefJ++vUwWdBgL6zlLSYi/uIZckhtGzgga U3R1OMXfl0LR0jXUe+Zf8KEaHPBE2BM9VF8gq0PGZPHi1t+Q8tqp5WXMuLO8YMHJmNIomV1ZQ1F Mkn74NDi+2G2MWfAki3x1e85DzzT6UoLkFEhk5deXQmoSjjYSLW/5ZiBzBn9bC+v6p83Mob2YcI 5AZBz1frnfU+5GJAUf0WzRcf/n5g+rL/2dAiu0d11t+rAVoIkv9Z5ZYHXY2dSXYYYQzRZ++H4dS 6l4B9XBz3VrkKBw== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 WQ_SYSFS exposes settable workqueue attributes under /sys/devices/virtual/workqueue/. For NFS workloads, this is useful for tuning affinity scope at runtime without reloading modules. Signed-off-by: Chuck Lever --- fs/nfs/inode.c | 3 ++- net/sunrpc/sched.c | 5 +++-- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 3022454f7698..107a2135029d 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -2619,7 +2619,8 @@ static void nfsiod_stop(void) static int nfsiod_start(void) { dprintk("RPC: creating workqueue nfsiod\n"); - nfsiod_workqueue =3D alloc_workqueue("nfsiod", WQ_MEM_RECLAIM | WQ_UNBOUN= D, 0); + nfsiod_workqueue =3D alloc_workqueue("nfsiod", + WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS, 0); if (nfsiod_workqueue =3D=3D NULL) return -ENOMEM; #if IS_ENABLED(CONFIG_NFS_LOCALIO) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index 9b2800f6bfc7..c31cf55b933f 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1271,16 +1271,17 @@ void rpciod_down(void) */ static int rpciod_start(void) { + const unsigned int wq_flags =3D WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS; struct workqueue_struct *wq; =20 /* * Create the rpciod thread and wait for it to start. */ - wq =3D alloc_workqueue("rpciod", WQ_MEM_RECLAIM | WQ_UNBOUND, 0); + wq =3D alloc_workqueue("rpciod", wq_flags, 0); if (!wq) goto out_failed; rpciod_workqueue =3D wq; - wq =3D alloc_workqueue("xprtiod", WQ_UNBOUND | WQ_MEM_RECLAIM, 0); + wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) goto free_rpciod; xprtiod_workqueue =3D wq; --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A30284A3F3C; Mon, 31 Aug 2026 18:22:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200534; cv=none; b=CoB308hUZyKyRmmJanXhyFDp1UFQhfz+OINeKoOSUB9LKO8TqfP2QfIjUTE78XDVTKiexzlObF0rQ4x1l1Df/5Ue8fLpPrBs5dhlKzxcOSJOe72quZsFK3gnoN5hkbPkdPlisePDD5dlHh84MPKwboDWf0UgzOhZHYwiHKldz/k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200534; c=relaxed/simple; bh=wz0MMyyCVaHC2FjQqTHwJ7v4CqEft8XGFBczTPxg0gc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ahJHUOn25jg5FgCy5ty+QJSyIpqLz5Gl1tjV6TUBhSnQAEeskYCSQ4if9LF8tQ8ybHJt4bnvMDLQR0FyserbD1jbx0xidm1E5Qs3p8KOafDx6l2fOdrxHnW/k0Fy0kRWGNOylDYCg9qn+COtLGhIReCAuo4QplxZySR6kWj8QxY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=IoPkQP6b; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="IoPkQP6b" Received: by smtp.kernel.org (Postfix) with ESMTPSA id BF2881F00A3D; Mon, 31 Aug 2026 18:22:11 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200532; bh=N5/NvHDvb3vrI/xCSkZr3CvmB/0BWoH3Q+mf/sS77qk=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=IoPkQP6bbTr2j8P4eUXGv/jtSoKApzuIa3olJamVyWdUAPIo/djbKp6F/NagL5l7l IEhY4rrJEyE4KsMHNgq1dQ3aIyOidrBnXjs8/kChBLZGspKZ7t0uzR/bTIdkcWeIsA jZh5hEGGmpYO6HP6sOvQh08Wd4Q/Ue3+kNLd80DRVkxqs+PHp+qJzzxrDljlRSZYKk ChpSbnuWidS+3bXBokEn3MBPkHaNU7Vdjt4dj7LYpyUCa/DMQlyV2kK7dBRx8D7Q4H 66DrcBX4CNwlcssW1FYooiMDiUwZ6MxiKOoaGGqb0uZJumojK+38tJJ+xviEB+pws8 0eJ9m1RT8TdbQ== From: Chuck Lever Date: Mon, 31 Aug 2026 14:22:01 -0400 Subject: [PATCH RFC 5/8] workqueue: Export the functions needed for WQ attribute modification Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-5-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1253; i=cel@kernel.org; h=from:subject:message-id; bh=wz0MMyyCVaHC2FjQqTHwJ7v4CqEft8XGFBczTPxg0gc=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPuhAhLHX7AWO8Ot4Np0II83BTgr62QTWg9 kZMKeVQjVGJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ lzCeD/4gztX07Z05+X9r+YuO8v22geySVkE3TEXIXsCeW5QmpcqaWCXnMKg6STUkYjBvDQ9vQNE dt5PQQ1pqsh44J9NiVg/sVBtdPXWvKBB+fvLQFN79Mn8Ee9NRzebJWv7l05GNxlCGoUvQibYQ9V eho8uiLGBXqARlC1+NGmeIu/Y9k8pXLB1oP6p2836u1LcDHuHM5gByqLrjUJDyrSqctHRVEmSsz eiY8U4t3vo2bwj4iJi7iDlXWl/SD77P9XLmPAQ0rrCSNEzpeA80rkvUMZuddRhY1/5q9dmnk3GD QKUmuPXvCPPqOjRdmAj9CO4hL1I0llgQOO5ECzheschTOf1AVg/rFqGvRYvicsPJgHwAfaJfo3N DFe6cILKBYaoezka8S76X9ABKZ6CIMnbGjp71vUCI+It36Z108SN7AYep2C9c6L5z0iFn1z2q1l TpxunTcuMOBuhp5FUio9OyVipCBSHZPQ6+pQYFu19TgjubqqTZB+iQ4tMk5ZMWSb3JywsdqUqc2 svsfeg/Y87qIX6nC5amoBvWTfpzOKeJ/7vGfQS6BbIYW8hY3J2fbE/u/M+4asM9erRBcTzjvi4e EpS83qcftyDfZeBZzE80OGKtY4EGNzAdRlz5opjdDxwTPRQaJmT6d3lMT3xrzrBElgQXVjl7bOZ Sj92Ekwj+l8uKFg== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 Modules that create unbound workqueues currently have no way to modify workqueue attributes at run time. Export the allocation, application, and teardown functions so that modules such as sunrpc can adjust affinity scope without reloading. Signed-off-by: Chuck Lever --- kernel/workqueue.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/kernel/workqueue.c b/kernel/workqueue.c index 3c034cbc5bb3..fd76a875cf2d 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -4813,6 +4813,7 @@ void free_workqueue_attrs(struct workqueue_attrs *att= rs) kfree(attrs); } } +EXPORT_SYMBOL_GPL(free_workqueue_attrs); =20 /** * alloc_workqueue_attrs - allocate a workqueue_attrs @@ -4841,6 +4842,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(= void) free_workqueue_attrs(attrs); return NULL; } +EXPORT_SYMBOL_GPL(alloc_workqueue_attrs_noprof); =20 static void copy_workqueue_attrs(struct workqueue_attrs *to, const struct workqueue_attrs *from) @@ -5609,6 +5611,7 @@ int apply_workqueue_attrs(struct workqueue_struct *wq, =20 return ret; } +EXPORT_SYMBOL_GPL(apply_workqueue_attrs); =20 /** * unbound_wq_update_pwq - update a pwq slot for CPU hot[un]plug --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 327684457C0; Mon, 31 Aug 2026 18:22:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200534; cv=none; b=ZEPG1tuxZk2SwLPnGYzeygWQ1WTM4PQLn8aqiKxbamjN1WtlGJA5NpNxI2xKYsu5kRXUkSubfdUB+D8iToAATahz7/2ewHIomwrw0cuzofj6w+IcTknL4YyVzdFwLDDb/qoBzW0XlxPbWz2zV4wLzbE1UB6tb51ryBexySgk0+s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200534; c=relaxed/simple; bh=dq1K12sW81ZZDaOE6GK1QC+Xtcdmi0Mgy2uQbeewkGg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Y61vTZCL1L/iT7b/nJUsAAk0gHqKn1v2xqcL0taGtsr6x0Jp2E43rw5V1pXn/q9Lzbs4v25AjZSYCqCW1O2LtGmxOvUpd9RgI48kUYJi/iifU40DYlS3AWf52PnFE6X9U8kLqwOEuAJh3La1yum4RA4KpZGbkYcZiaMGwYe9Yu4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ddnYDz7q; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ddnYDz7q" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 827D21F00ADB; Mon, 31 Aug 2026 18:22:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200533; bh=1KLBJGuRObI2XbOROtGtBccuic6ok9txtKzBKn2YGZ0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=ddnYDz7qG/marazDz04oMr/aLqcrhTwgJyqoBlNXis64Zo4xr0LaIC6wVKd4mbWod 2W4YCX6F3MS1EoCV97Q+12jRm+rLlbcUnYYUlfMlGNKWkUG+T7KJKkEgkBgBXn/OGp lMZQWSNwPVJAVbyPna61O9QYMgl0Wz7tMa98R6VIgSIJlEe+afV+jMZGurjSY6qcA5 oACS2rExNX1rjQv/zKgs/m4R/q+FMb9ppatTCEhqPwd+rvwiwG6q8Hxedfk2CbtWIG lKXnsA3x1UM80Xchn9cDn6KltNXrssIm2n2R8C5/whTbZlpR7vh6DcTwqRwAuAKRvO OWCglpyKCld+w== From: Chuck Lever Date: Mon, 31 Aug 2026 14:22:02 -0400 Subject: [PATCH RFC 6/8] SUNRPC: Reduce rpciod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-6-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=2114; i=cel@kernel.org; h=from:subject:message-id; bh=dq1K12sW81ZZDaOE6GK1QC+Xtcdmi0Mgy2uQbeewkGg=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPSA8/uSXNJdaVPVpYuQocZE7GZh9gifGsf Mvw5RPDVXOJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ l14BD/wP3+aKMjx5r9lfpl81bv5C5oYEMBB9TWhRPFsnUFCuuFZQjjYnirV2OsuQQQd4nLLixxi +Lbqg3lisU2EnOrhVaAAHNs+DrZgmHONXhiowvVin2wuvxWxSHFDNV14ZP91Lx7v+35TrpZ6XzG 8bTpbmibC6dH1gRKWFR0vzJyu0eI580tilzjqe0fQv6yoE4DDry018GyR6wNA2ZsAlZ0Ft6dcU6 ote7XlXHixlqZX+T/q7UB88q1SsjrmhBJVd17XIF5cQuHY95B72bcfMWEHif0oy3oWPWbw59VKe aA5S68YGNiQq8Z3ZlRTxa+oMdwiN/0iCh/Ylzv8ov8b24otMWQ3aJ1sYyYFyjIYlGsD2PJL/1fP WXzlt9L05auTYBRztHL5tyb5RkG7MXMauxYP/XvFKhOnNjixiuqXjsGWtEUiTdt9wkq0iScKfRX ZbbypAEaHBvicsv3k1+9DusVZik7SIBEt+Ao+kCM/Np9dyeojk56Ch0sSg8xAH6BgdtNyuVZMhE t7MjCtCdaOTsBJoLIQ4Ez6wQENQPzT3Od1j40Bjo/gH9IlVVLe5VOvacublLS+QSVSOC0T00c1J KXPfDZuFGF9TpD8dc4YzjZ1eU0V45kXyQeWDaC7j2ZJ0WwXuu6rQWXvdTe3bvVjuBBESnYZiXUg DBdlBD+u2kTYefA== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 rpciod drives the RPC client state machine. Under heavy NFS workloads, multiple CPUs queue RPC task completions concurrently and contend on the UNBOUND worker pool lock. perf profiles on a 12-core system show 30-40% of cycles lost to native_queued_spin_lock_slowpath in the rpciod pool at the WQ_AFFN_CACHE scope (one pool per LLC). The WQ_AFFN_CACHE_SHARD default helps little here, because its 8-core shards split this system into just two pools of six cores each. Set WQ_AFFN_SMT on rpciod so each SMT group gets its own pool and lock. Most UNBOUND workqueues never contend on the pool lock and profit from a coarser scope's cache locality. rpciod's sustained completion traffic makes the lock a first-order bottleneck, so the override belongs on this workqueue rather than in the system-wide default. Idle kworkers are culled, so the extra pools cost little on large systems. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index c31cf55b933f..b84af2c0b104 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1266,6 +1266,25 @@ void rpciod_down(void) module_put(THIS_MODULE); } =20 +static void rpc_set_wq_smt_affinity(struct workqueue_struct *wq, + const char *name) +{ + struct workqueue_attrs *attrs; + int err; + + attrs =3D alloc_workqueue_attrs(); + if (!attrs) { + pr_warn("%s: failed to allocate workqueue attrs\n", name); + return; + } + attrs->affn_scope =3D WQ_AFFN_SMT; + err =3D apply_workqueue_attrs(wq, attrs); + free_workqueue_attrs(attrs); + if (err) + pr_warn("%s: failed to set SMT affinity scope: %d\n", + name, err); +} + /* * Start up the rpciod workqueue. */ @@ -1280,6 +1299,7 @@ static int rpciod_start(void) wq =3D alloc_workqueue("rpciod", wq_flags, 0); if (!wq) goto out_failed; + rpc_set_wq_smt_affinity(wq, "rpciod"); rpciod_workqueue =3D wq; wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0BA5A39A4D6; Mon, 31 Aug 2026 18:22:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200535; cv=none; b=cQpbl/Q27gxdbN4oT0pS/RIuqT55LAFDqyhRafHOH/G2cNiZhj9tIvAZ1SZ+99X/UaZXPjg7kIrbt5kbwC2sFdQoz9dMwrpkoiraYvaCQXUuk4nbbNi6hIu/RMZKVlaATLXgU/2XmvRT+Re7hQAZeWx26uFrXdqlRt4kwFYaMl0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200535; c=relaxed/simple; bh=RW896w+ixR26FNVdNSEmTu+JhQhaAIq+agG+jv57MiU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Yj1iIZ410s/NVHKJ/+XtrDKWlB6ieiFKHU1KNXCFwVZKsb43g5i44Ow4atPqn7AGQn3i18otGlxwLJ1N+CHqTaZsNcxt6bO4Ue+dxIH4we22z0FlVsQk8BFMADRUA3AZPuuVPNB4e2XxPJV7oiubC0tgVET8HypIbDuiQHsD4zU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bIm6lb0u; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bIm6lb0u" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 49DB91F00A3E; Mon, 31 Aug 2026 18:22:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200533; bh=z0MZq52InHRy26UM/UyBfExpVeGJ7udY3VsN7YkGOTw=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=bIm6lb0uG2DwtXmqYa6N48hHhpyGyafsTOEc3SMc/VvQDTbKDw1N6z2LKz/nuyEKm IrM3uaq0r38CfGzoFXyqKoLFRgcyqsOVOty29ty3vr7OYcWQr9p6MtdxYN6YhLJ39P 3SFJEmGrot8HYITgPU8eQP6ZedfgpLt2Ss7QO2jSSsnJI0G3EFkJjeBxuzkY29LigY Y23/L8p8GxvbCbE5wttLxvCR5/wJiTN9rVPbkutqWZbk7LJnShiKWS1VNqCHWwh0NZ iDbla5tJw3hr1IECTYZMfLvSBG9MLSObpt2naKmj0EWfVVe6k+VIHrymjvYwXYI47+ QJAWmGvO4dwtA== From: Chuck Lever Date: Mon, 31 Aug 2026 14:22:03 -0400 Subject: [PATCH RFC 7/8] NFS: Reduce nfsiod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-7-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=1997; i=cel@kernel.org; h=from:subject:message-id; bh=RW896w+ixR26FNVdNSEmTu+JhQhaAIq+agG+jv57MiU=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZPU7rmo/CML+GJu7O7UTT3DDB2RMhyHHtT9 cAtrKRxFAaJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ lyOFD/0coyl6dcjFgytEgMk1+jFynDB4V8LSGVhIS9lvrBfdv0sMq2z4j92yOdiwIui674JtqHJ wVoFA9feekH9nEOgOZMwfgtHUeu6vLHl0jHDcxPSrtMXpipqjqdycUgtfrcXaNTUoYNPoHJGdPT 80TCvrszhxjtEjfKjG+LPda4cgDwitmfxfOyQr0u3/x96Cb3yYuM7sQgA2qHyz7GD8sKp4zlmob dG5Ezm5L7RQB8tBaRigWrEtjhfUxhU6V0IhD97lyXemmIVisF6MXqTPsUsS51DzHbhHGbh+HazB shZ/8WqCzHf52dR6bYmnI+I6LAJGyimuuRMyR1y3B68sqCCX2To2x8kWx94QIJCr36k8fyA7P5S RyaxUYYKtYv1i3PXA6TA7+7y4R7JrRhlqkaH9hQ45E7swJ/y1Jw55T0vHBjH8g2qkty0i0K/0LN k10SznHsmacFLdOPQXrEeWdcr5r6zQuuN5zm9exKEpGCuxbYTGSPpiuNbacwIz13QxpLYKh7B7s BR4jN3Dw5+w0mEmttx6rZTgBgySeVAtuWka/RhsVyRokO3QmE5uVv9RGLzUXa/XQQS0lUYkCZ/U 34SUClCqTwpX/FxCKPlLFGxWNp7jSwkYK0RA8pavH+xBg9O6z58t55UrUyl6NerGv5gzYkDeVcZ XImN5UVNtx93aNQ== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 The default affinity scope for unbound workqueues is now WQ_AFFN_CACHE_SHARD, which splits each LLC into shards of about eight cores. On a single-socket system whose LLC fits in one shard, every NFS I/O completion serializes on one nfsiod pool lock. Profiling 4KB random writes over NFSv3/RDMA with nconnect=3D3 shows that lock consuming 17% of CPU cycles: 8% dequeuing work and 9% enqueuing follow-on work from rpciod and nfsiod workers. Set nfsiod's affinity scope to WQ_AFFN_SMT so each CPU gets its own pool and queue_work_on() no longer takes a lock on another CPU. Enqueue contention disappears and dequeue contention drops to 1.4%. Throughput is unchanged because the workload is transport-limited, but the freed cycles cut submission latency variance by 67% (slat stdev 31.6 us to 10.5 us), IOPS stdev by 31%, and p99.9 completion latency by 11%. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- fs/nfs/inode.c | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 107a2135029d..8e0d2bebba54 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -2618,11 +2618,26 @@ static void nfsiod_stop(void) */ static int nfsiod_start(void) { + struct workqueue_attrs *attrs; + dprintk("RPC: creating workqueue nfsiod\n"); nfsiod_workqueue =3D alloc_workqueue("nfsiod", WQ_MEM_RECLAIM | WQ_UNBOUND | WQ_SYSFS, 0); if (nfsiod_workqueue =3D=3D NULL) return -ENOMEM; + attrs =3D alloc_workqueue_attrs(); + if (attrs) { + int err; + + attrs->affn_scope =3D WQ_AFFN_SMT; + err =3D apply_workqueue_attrs(nfsiod_workqueue, attrs); + free_workqueue_attrs(attrs); + if (err) + pr_warn("nfsiod: failed to set SMT affinity scope: %d\n", + err); + } else { + pr_warn("nfsiod: failed to allocate workqueue attrs\n"); + } #if IS_ENABLED(CONFIG_NFS_LOCALIO) /* * localio writes need to use a normal (non-memreclaim) workqueue. --=20 2.55.0 From nobody Sat Sep 26 14:37:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A56F54A4414; Mon, 31 Aug 2026 18:22:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200536; cv=none; b=heSNNznIhA11UokJup3LEK6MQzdinPITd/YISVU/8jzyAoc+9bPTcAnUF2JXsMfe//t8DFwGB2v09A+49cC4q7wUjEAHtLR9F9F9zajjAuJyhCaZILGKyubOc7r7aHp3C5POPmY/VpvgiiV9eIa4awUPvnQmw4elJuGoZ/D4f8M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788200536; c=relaxed/simple; bh=xT7HizUrknGMjf/pPX9I2cRTBh8bgQVEBDDV51nL1X0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Zo1M4OVMhvHd/duZEQzMikz5hvHUfJjqzbTH4p9ARtJJEooUf02K2xMPWhVBRG/XMR90KDvqhqAZQN8Z6l7c2qy7m9NiDTI61WD+GCMCvb1RD1F5MXn3STnJzHTmmIanw0JTdgdn9mRCvkN1cnaHTCTgw8dSEHCDJhE5EUfeD4s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=BLN5ppeS; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="BLN5ppeS" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0E55B1F000E9; Mon, 31 Aug 2026 18:22:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788200534; bh=NFZHnOZfgrGz8mBkIoJ8Uq/DZjW0+VctiqeMxOFdM9Y=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=BLN5ppeS7uHIJGvwBLObqcmTRLLiJCrgy5huyttjTZ7nSa0cWAQvUC79rG0EHcgbc tdsD/ykZZH0N6NmO8DmS+yufqykwHJC5WAVbxlokhxn0Q2k8Lc5QyPTJeXiau5HuKY dWx/l9X4esOtVmFzmUDz2Re0KMdAFu7QTC7jKi/1Idk/KRZcq7ZTF5ojeVPZo9IQWU wOCgh244yCksso/pT+D5V10SQit0m+n7zmXCr4elWYwBRZ9a73OdtGVgnROUDqMqJh jamfU16ZmaDazmmkNEuQbCD2YOpMD59ftUEZdtahCWH9MKEs2vNm/dayIlEuElNMWR EAQ0tsDFGwTlQ== From: Chuck Lever Date: Mon, 31 Aug 2026 14:22:04 -0400 Subject: [PATCH RFC 8/8] SUNRPC: Reduce xprtiod workqueue contention Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260831-performance-v1-8-8d9fd9b67f96@kernel.org> References: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> In-Reply-To: <20260831-performance-v1-0-8d9fd9b67f96@kernel.org> To: Trond Myklebust , Anna Schumaker , Tejun Heo Cc: Lai Jiangshan , linux-nfs@vger.kernel.org, open list , Chuck Lever X-Mailer: b4 0.16-dev-da966 X-Developer-Signature: v=1; a=openpgp-sha256; l=875; i=cel@kernel.org; h=from:subject:message-id; bh=xT7HizUrknGMjf/pPX9I2cRTBh8bgQVEBDDV51nL1X0=; b=owEBbQKS/ZANAwAKATNqszNvZn+XAcsmYgBqlcZP5nKqeYag30AG0GTN1esiHQX73ffO8sLq5 +TNUUk2dcmJAjMEAAEKAB0WIQQosuWwEobfJDzyPv4zarMzb2Z/lwUCapXGTwAKCRAzarMzb2Z/ l/gzD/wPbRSVBV0PBWuiS4P+BZ8p0IcfWe90a7g9R5eWpeunnk25H9hv5KpIJa0mAy81necntm0 RmLJNM1ETHohLRq18VSkHGOAbm0GHg1DMV8y1y7vBhhMs0qy8tVILrw4cl/bXWmGDEq7qHQfxyB dRMWvqTaKFzE2L4QC/k2jkA1fKlg3III6fD78hcdynguog69aFMpw4514mkvBOuBFXf29APnnGy AodrIMN0N09fhPpXzSK+S/Z6oYhl2FYWx6Z15O2B+M1wTEOp27bZC/5zC9qVmy6opNrKa0+qqGf 8K10a1E7QqsGTbF2xam4glfAqyOfqyixoM6S7BanYnWUFtO6XIEkQ9X6Vgm3t/bAWA3WX1y73/E yJynqGLbfR4RZw6Q7MrE3Mcbi8Eu5crORIQk3Zqa4pypvD/IPK3Wyrzc0HhXsWRjZ6pvCAO0V9m Z0joktucneaNVgde7Cc98xWw+daRQLIzOf1E8yZr0v7VD5Fggl/8fNhEgJ8GiL56ojhJt/AqRK3 2no4w+l7yQ0BnfHx96pmGDhxxk+N+Nunw600cHQuJ+rh3ZljeohZNiLvFhHkXgLREQahds/n4+r AejBhizsaTudlHD2np4sZmk8d+bXI9f3p6/FxvH7XhK3bXZ9gTOax4dba4IqEJVj83pjymEvmMl BJatB6Ulltj65Dw== X-Developer-Key: i=cel@kernel.org; a=openpgp; fpr=28B2E5B01286DF243CF23EFE336AB3336F667F97 xprtiod handles transport-level operations: socket receive processing, error recovery, and connection lifecycle. On systems driving heavy NFS traffic, these operations contend on the UNBOUND worker pool lock just as rpciod does. Set WQ_AFFN_SMT on the xprtiod workqueue to give each SMT group its own pool and lock. Suggested-by: Tejun Heo Signed-off-by: Chuck Lever --- net/sunrpc/sched.c | 1 + 1 file changed, 1 insertion(+) diff --git a/net/sunrpc/sched.c b/net/sunrpc/sched.c index b84af2c0b104..3236dfdb4096 100644 --- a/net/sunrpc/sched.c +++ b/net/sunrpc/sched.c @@ -1304,6 +1304,7 @@ static int rpciod_start(void) wq =3D alloc_workqueue("xprtiod", wq_flags, 0); if (!wq) goto free_rpciod; + rpc_set_wq_smt_affinity(wq, "xprtiod"); xprtiod_workqueue =3D wq; return 1; free_rpciod: --=20 2.55.0