From nobody Tue Sep 29 07:39:37 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5964E2F12CE; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; cv=none; b=MB883M+J0vf+GBn57Z0YK04M3p2GX8tOOxp5rMvYJCaVSz1BJR3ZR1Ljrxi9S+1ltxPkT+rmPVYIfq9/57gt+JWHR7W8GTx7I1iXmXcwtpn8depipEs8f78WtuhQ2wCMJKRHtcsrle1+QmU6x47hxGWJTMJAi3d+WDpIw749sHE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; c=relaxed/simple; bh=NRy+vc4hWJQrUiBYfp3itTG4myVbZh4BfWueZTePwcI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=j98W3F5M+xfMPeOmSx1gcOiMhGekFDLlQKP2w4M+mzYMGpxeRdPLbIvPafEyF2g1A+yYLqJuaJj7p5P0IpXHUK1eHWUfYRsupVj1x03lFX5lZYX6Lgr2sgPJMQw3sFinyVKlM9aInNvHl4R7J+t8Sz3gly7QSqoEU8KNCb3PWfw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=f4M66gt1; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="f4M66gt1" Received: by smtp.kernel.org (Postfix) with ESMTPS id 0953AC2BCF7; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786425477; bh=NRy+vc4hWJQrUiBYfp3itTG4myVbZh4BfWueZTePwcI=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=f4M66gt1x93Hgm0oqoWKVAnKCi8H/Y/jvPWhjVdYOAits+5jIQ47R3NmNxWBpXKfA w9EwjASyLW3ZwXXfavnW/oPNVFfkw+6R8m7VQn2d4KEBWwavfONEnTdgMUhbUrP/cv b6GzjxKQvs/wpLD+g5NGduqM0xnvakKRRYqOXUp69gNMT2BOoeD/QcYGz/FJ9OqgUX kQrD8z7xb+wXAfo1YAqhqIKm06Uf5C82cmOA9PiGIi/5QuNZr5sP8pMkg6jRtXcjLu pHTypmOjUph9FLM6/IaGu7svd6XQBlFG9oNkfKiG7Z9mldbxHeGc5ugq6m89/Mrrm0 CddlsP1fWy1Ag== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id DCC3BC5AC67; Tue, 11 Aug 2026 05:17:56 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Tue, 11 Aug 2026 13:17:45 +0800 Subject: [PATCH v3 1/5] ceph: use READ_ONCE/WRITE_ONCE for oldest_tid Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-ceph-mdsc-mutex-optimization-v3-1-d031114419f4@clyso.com> References: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> In-Reply-To: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786425473; l=3350; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=exqvyxSW2xd5rzPWwAc/Vr8hxgWOg5BerR4rCyEJz+w=; b=BB94b/P0thWS3JT87Rohqa2H4/zOu7aOMr0w8HtppR8ZR2gugWyAmAeDTYzi1hU3m/vJMR2WC SPvCBm3SQsRCKEgnKQxTOa86A+xxZ1NJP13yJmciimfufJGDRgmfBJn X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li The oldest_client_tid sent in the MDS request header is advisory: a stale value is harmless --- at worst the MDS may resend a reply the client already has, or trim its completed-requests table slightly earlier or later than optimal, both of which the protocol handles correctly. The field is monotonic and does not need lock serialization to stay correct. Switch all accesses to READ_ONCE() and WRITE_ONCE() to prevent the compiler from tearing or inventing loads, documenting that these lockless accesses are intentional. This removes the last reason the MDS request-send path had to be serialized under mdsc->mutex. Signed-off-by: Xiubo Li --- fs/ceph/mds_client.c | 20 +++++++------------- 1 file changed, 7 insertions(+), 13 deletions(-) diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c index f365e434e35d..2c51edd0e9cd 100644 --- a/fs/ceph/mds_client.c +++ b/fs/ceph/mds_client.c @@ -1240,8 +1240,8 @@ static void __register_request(struct ceph_mds_client= *mdsc, if (!req->r_mnt_idmap) req->r_mnt_idmap =3D &nop_mnt_idmap; =20 - if (mdsc->oldest_tid =3D=3D 0 && req->r_op !=3D CEPH_MDS_OP_SETFILELOCK) - mdsc->oldest_tid =3D req->r_tid; + if (READ_ONCE(mdsc->oldest_tid) =3D=3D 0 && req->r_op !=3D CEPH_MDS_OP_SE= TFILELOCK) + WRITE_ONCE(mdsc->oldest_tid, req->r_tid); =20 if (dir) { struct ceph_inode_info *ci =3D ceph_inode(dir); @@ -1262,14 +1262,14 @@ static void __unregister_request(struct ceph_mds_cl= ient *mdsc, /* Never leave an unregistered request on an unsafe list! */ list_del_init(&req->r_unsafe_item); =20 - if (req->r_tid =3D=3D mdsc->oldest_tid) { + if (req->r_tid =3D=3D READ_ONCE(mdsc->oldest_tid)) { struct rb_node *p =3D rb_next(&req->r_node); - mdsc->oldest_tid =3D 0; + WRITE_ONCE(mdsc->oldest_tid, 0); while (p) { struct ceph_mds_request *next_req =3D rb_entry(p, struct ceph_mds_request, r_node); if (next_req->r_op !=3D CEPH_MDS_OP_SETFILELOCK) { - mdsc->oldest_tid =3D next_req->r_tid; + WRITE_ONCE(mdsc->oldest_tid, next_req->r_tid); break; } p =3D rb_next(p); @@ -1698,7 +1698,7 @@ create_session_full_msg(struct ceph_mds_client *mdsc,= int op, u64 seq) ceph_encode_32(&p, 0); =20 /* version =3D=3D 7, oldest_client_tid */ - ceph_encode_64(&p, mdsc->oldest_tid); + ceph_encode_64(&p, READ_ONCE(mdsc->oldest_tid)); =20 msg->front.iov_len =3D p - msg->front.iov_base; msg->hdr.front_len =3D cpu_to_le32(msg->front.iov_len); @@ -2764,7 +2764,7 @@ static struct ceph_mds_request *__get_oldest_req(stru= ct ceph_mds_client *mdsc) =20 static inline u64 __get_oldest_tid(struct ceph_mds_client *mdsc) { - return mdsc->oldest_tid; + return READ_ONCE(mdsc->oldest_tid); } =20 #if IS_ENABLED(CONFIG_FS_ENCRYPTION) @@ -3443,9 +3443,6 @@ static void complete_request(struct ceph_mds_client *= mdsc, complete_all(&req->r_completion); } =20 -/* - * called under mdsc->mutex - */ static int __prepare_send_request(struct ceph_mds_session *session, struct ceph_mds_request *req, bool drop_cap_releases) @@ -3560,9 +3557,6 @@ static int __prepare_send_request(struct ceph_mds_ses= sion *session, return 0; } =20 -/* - * called under mdsc->mutex - */ static int __send_request(struct ceph_mds_session *session, struct ceph_mds_request *req, bool drop_cap_releases) --=20 2.53.0 From nobody Tue Sep 29 07:39:37 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 59770314B96; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; cv=none; b=R60uX1BoJKg3NcnAgh1B0lDxzOfgCzyCle9fnoK5iUi52mOEQ5gHQjfx1FUNqBwDbX4hnFAtKqRrS5Mn4vSzAfNDBW/XEfj72gz2tbwPr5ZDUfHRzCt36wTe94YFaUGo4WN977Gk2ax+6O2PrXxwncxTiSFVO3vPZcfpxwF8BIY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; c=relaxed/simple; bh=3cValD5m5/uxaRHosmqshs8uecw1utQlbHxHOY+tWw4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WQ2/dHgxu0lZEFidsWMrbCx4UkcDTSqGTnrQEzjnV94TSaQGUvEtIdGVxG0uPzbYpyV679fIJzZvf0fy0jENJj+3iXvN3qOr/NDxvamiGorOBV3//2VwUlM818siTHE6Z+ECZ27/uhCjwEbUc6UUmYY8Ak4B9LpAbw146L42/ew= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=nBPhB+aT; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="nBPhB+aT" Received: by smtp.kernel.org (Postfix) with ESMTPS id 16E03C2BCF5; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786425477; bh=3cValD5m5/uxaRHosmqshs8uecw1utQlbHxHOY+tWw4=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=nBPhB+aT4nGerEbEiMASRd28gYvtUixDQ2I3CJSL6RlQ5HYHwAabfaCG6IWS81JHq /1b/zqOSyRVoMo/s5pfx03g6th55cWO91xHvUXrjaAG/hoR2+HraiC1manbpf3ftJ6 WWYfKrC1Enl0gBelSK1NmRhXqGnvym0ANEFIwRbmq4j1hWegDu7Gvgj5+vD89r/sTh ipZfTYRSQqgory3r1RsnR6G0oi+MtFy24IcXb37sK8u5Xee5qCuhJ15FN7BGNjS1+f n6BcuedjVWhZdhTVVqXIOYYhqY7QhnWAHFkFR+c90oeltE7QzpcKkc2QybXff95aA3 ipvUb9+YfvYew== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id EC598C5B56A; Tue, 11 Aug 2026 05:17:56 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Tue, 11 Aug 2026 13:17:46 +0800 Subject: [PATCH v3 2/5] ceph: replace the request_tree rbtree with an xarray keyed by r_tid Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-ceph-mdsc-mutex-optimization-v3-2-d031114419f4@clyso.com> References: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> In-Reply-To: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786425473; l=11768; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=pPoiHEBnnXznTML12FQlX2UyRdvmh5GajVHPyfd3h90=; b=d5YWoBAlK4EpMBdDj3DEPaPHyCqtwsgIQDmpMQiVX1kVEBwJHhRXpNZ9LIQ1L5SZcB2AANcyW HUMW0h5xzLrA1hOCybBYBm8ijwkRk26YYubvFwzAVMYcliLa7PdYIBj X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li Replace the red-black tree that indexes pending MDS requests by transaction ID with an xarray. The xarray provides O(1) keyed lookups versus the rbtree's O(log N), and its internal RCU locking eliminates the need for an external mutex during lookups. Iteration is also simpler, with the xarray's native iterators replacing open-coded rb_first() / rb_next() walks. Because xarray indices are unsigned long, a u64 transaction ID would be truncated on 32-bit platforms. Restrict CEPH_FS to 64BIT since there are no 32-bit users. Validate the xarray insertion and clean up the structure at shutdown. Signed-off-by: Xiubo Li --- fs/ceph/Kconfig | 1 + fs/ceph/debugfs.c | 6 +-- fs/ceph/mds_client.c | 112 +++++++++++++++++++++++------------------------= ---- fs/ceph/mds_client.h | 3 +- 4 files changed, 55 insertions(+), 67 deletions(-) diff --git a/fs/ceph/Kconfig b/fs/ceph/Kconfig index 3d64a316ca31..1d95f9c7e271 100644 --- a/fs/ceph/Kconfig +++ b/fs/ceph/Kconfig @@ -2,6 +2,7 @@ config CEPH_FS tristate "Ceph distributed file system" depends on INET + depends on 64BIT select CEPH_LIB select NETFS_SUPPORT select FS_ENCRYPTION_ALGS if FS_ENCRYPTION diff --git a/fs/ceph/debugfs.c b/fs/ceph/debugfs.c index 18eb5da03411..491ead3fe1c6 100644 --- a/fs/ceph/debugfs.c +++ b/fs/ceph/debugfs.c @@ -87,12 +87,12 @@ static int mdsc_show(struct seq_file *s, void *p) struct ceph_fs_client *fsc =3D s->private; struct ceph_mds_client *mdsc =3D fsc->mdsc; struct ceph_mds_request *req; - struct rb_node *rp; + unsigned long idx; char *path; =20 mutex_lock(&mdsc->mutex); - for (rp =3D rb_first(&mdsc->request_tree); rp; rp =3D rb_next(rp)) { - req =3D rb_entry(rp, struct ceph_mds_request, r_node); + idx =3D 0; + xa_for_each(&mdsc->request_tree, idx, req) { =20 if (req->r_request && req->r_session) seq_printf(s, "%lld\tmds%d\t", req->r_tid, diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c index 2c51edd0e9cd..bf63f5c1fca2 100644 --- a/fs/ceph/mds_client.c +++ b/fs/ceph/mds_client.c @@ -1188,7 +1188,6 @@ void ceph_mdsc_release_request(struct kref *kref) kmem_cache_free(ceph_mds_request_cachep, req); } =20 -DEFINE_RB_FUNCS(request, struct ceph_mds_request, r_tid, r_node) =20 /* * lookup session, bump ref if found. @@ -1200,7 +1199,7 @@ lookup_get_request(struct ceph_mds_client *mdsc, u64 = tid) { struct ceph_mds_request *req; =20 - req =3D lookup_request(&mdsc->request_tree, tid); + req =3D xa_load(&mdsc->request_tree, tid); if (req) ceph_mdsc_get_request(req); =20 @@ -1234,7 +1233,13 @@ static void __register_request(struct ceph_mds_clien= t *mdsc, } doutc(cl, "%p tid %lld\n", req, req->r_tid); ceph_mdsc_get_request(req); - insert_request(&mdsc->request_tree, req); + if (xa_is_err(xa_store(&mdsc->request_tree, req->r_tid, req, + GFP_NOFS))) { + pr_err_client(cl, "%p tid %lld: xa_store failed\n", + req, req->r_tid); + req->r_err =3D -ENOMEM; + return; + } =20 req->r_cred =3D get_current_cred(); if (!req->r_mnt_idmap) @@ -1263,20 +1268,19 @@ static void __unregister_request(struct ceph_mds_cl= ient *mdsc, list_del_init(&req->r_unsafe_item); =20 if (req->r_tid =3D=3D READ_ONCE(mdsc->oldest_tid)) { - struct rb_node *p =3D rb_next(&req->r_node); + unsigned long tidx =3D req->r_tid + 1; + struct ceph_mds_request *next_req; + WRITE_ONCE(mdsc->oldest_tid, 0); - while (p) { - struct ceph_mds_request *next_req =3D - rb_entry(p, struct ceph_mds_request, r_node); + xa_for_each_start(&mdsc->request_tree, tidx, next_req, tidx) { if (next_req->r_op !=3D CEPH_MDS_OP_SETFILELOCK) { WRITE_ONCE(mdsc->oldest_tid, next_req->r_tid); break; } - p =3D rb_next(p); } } =20 - erase_request(&mdsc->request_tree, req); + xa_erase(&mdsc->request_tree, req->r_tid); =20 if (req->r_unsafe_dir) { struct ceph_inode_info *ci =3D ceph_inode(req->r_unsafe_dir); @@ -1833,7 +1837,7 @@ static void cleanup_session_requests(struct ceph_mds_= client *mdsc, { struct ceph_client *cl =3D mdsc->fsc->client; struct ceph_mds_request *req; - struct rb_node *p; + unsigned long idx; =20 doutc(cl, "mds%d\n", session->s_mds); mutex_lock(&mdsc->mutex); @@ -1849,10 +1853,8 @@ static void cleanup_session_requests(struct ceph_mds= _client *mdsc, __unregister_request(mdsc, req); } /* zero r_attempts, so kick_requests() will re-send requests */ - p =3D rb_first(&mdsc->request_tree); - while (p) { - req =3D rb_entry(p, struct ceph_mds_request, r_node); - p =3D rb_next(p); + idx =3D 0; + xa_for_each(&mdsc->request_tree, idx, req) { if (req->r_session && req->r_session->s_mds =3D=3D session->s_mds) req->r_attempts =3D 0; @@ -2736,7 +2738,6 @@ ceph_mdsc_create_request(struct ceph_mds_client *mdsc= , int op, int mode) req->r_fmode =3D -1; req->r_feature_needed =3D -1; kref_init(&req->r_kref); - RB_CLEAR_NODE(&req->r_node); INIT_LIST_HEAD(&req->r_wait); init_completion(&req->r_completion); init_completion(&req->r_safe_completion); @@ -2756,10 +2757,9 @@ ceph_mdsc_create_request(struct ceph_mds_client *mds= c, int op, int mode) */ static struct ceph_mds_request *__get_oldest_req(struct ceph_mds_client *m= dsc) { - if (RB_EMPTY_ROOT(&mdsc->request_tree)) - return NULL; - return rb_entry(rb_first(&mdsc->request_tree), - struct ceph_mds_request, r_node); + unsigned long idx =3D 0; + + return xa_find(&mdsc->request_tree, &idx, ULONG_MAX, XA_PRESENT); } =20 static inline u64 __get_oldest_tid(struct ceph_mds_client *mdsc) @@ -3820,12 +3820,11 @@ static void kick_requests(struct ceph_mds_client *m= dsc, int mds) { struct ceph_client *cl =3D mdsc->fsc->client; struct ceph_mds_request *req; - struct rb_node *p =3D rb_first(&mdsc->request_tree); + unsigned long idx; =20 doutc(cl, "kick_requests mds%d\n", mds); - while (p) { - req =3D rb_entry(p, struct ceph_mds_request, r_node); - p =3D rb_next(p); + idx =3D 0; + xa_for_each(&mdsc->request_tree, idx, req) { if (test_bit(CEPH_MDS_R_GOT_UNSAFE, &req->r_req_flags)) continue; if (req->r_attempts > 0) @@ -4668,7 +4667,7 @@ static void replay_unsafe_requests(struct ceph_mds_cl= ient *mdsc, struct ceph_mds_session *session) { struct ceph_mds_request *req, *nreq; - struct rb_node *p; + unsigned long idx; =20 doutc(mdsc->fsc->client, "mds%d\n", session->s_mds); =20 @@ -4680,10 +4679,8 @@ static void replay_unsafe_requests(struct ceph_mds_c= lient *mdsc, * also re-send old requests when MDS enters reconnect stage. So that MDS * can process completed request in clientreplay stage. */ - p =3D rb_first(&mdsc->request_tree); - while (p) { - req =3D rb_entry(p, struct ceph_mds_request, r_node); - p =3D rb_next(p); + idx =3D 0; + xa_for_each(&mdsc->request_tree, idx, req) { if (test_bit(CEPH_MDS_R_GOT_UNSAFE, &req->r_req_flags)) continue; if (req->r_attempts =3D=3D 0) @@ -5528,7 +5525,7 @@ static void ceph_mdsc_reset_workfn(struct work_struct= *work) */ { struct ceph_mds_request *req; - struct rb_node *rn; + unsigned long idx; u64 last_tid; =20 mutex_lock(&mdsc->mutex); @@ -5536,14 +5533,12 @@ static void ceph_mdsc_reset_workfn(struct work_stru= ct *work) mutex_unlock(&mdsc->mutex); =20 mutex_lock(&mdsc->mutex); - rn =3D rb_first(&mdsc->request_tree); - while (rn) { - req =3D rb_entry(rn, struct ceph_mds_request, r_node); - if (req->r_tid > last_tid) - break; + idx =3D 0; + while ((req =3D xa_find(&mdsc->request_tree, &idx, last_tid, + XA_PRESENT))) { if (req->r_op =3D=3D CEPH_MDS_OP_SETFILELOCK || !(req->r_op & CEPH_MDS_OP_WRITE)) { - rn =3D rb_next(rn); + idx++; continue; } ceph_mdsc_get_request(req); @@ -5556,7 +5551,7 @@ static void ceph_mdsc_reset_workfn(struct work_struct= *work) ceph_mdsc_put_request(req); if (time_after(jiffies, drain_deadline)) break; - rn =3D rb_first(&mdsc->request_tree); + idx =3D 0; /* restart: tree may have changed */ } mutex_unlock(&mdsc->mutex); =20 @@ -6282,7 +6277,7 @@ int ceph_mdsc_init(struct ceph_fs_client *fsc) mdsc->snap_realms =3D RB_ROOT; INIT_LIST_HEAD(&mdsc->snap_empty); spin_lock_init(&mdsc->snap_empty_lock); - mdsc->request_tree =3D RB_ROOT; + xa_init(&mdsc->request_tree); INIT_DELAYED_WORK(&mdsc->delayed_work, delayed_work); mdsc->last_renew_caps =3D jiffies; INIT_LIST_HEAD(&mdsc->cap_delay_list); @@ -6604,34 +6599,33 @@ static void flush_mdlog_and_wait_mdsc_unsafe_reques= ts(struct ceph_mds_client *md u64 want_tid) { struct ceph_client *cl =3D mdsc->fsc->client; - struct ceph_mds_request *req =3D NULL, *nextreq; + struct ceph_mds_request *req; struct ceph_mds_session *last_session =3D NULL; - struct rb_node *n; + unsigned long idx; =20 mutex_lock(&mdsc->mutex); doutc(cl, "want %lld\n", want_tid); -restart: - req =3D __get_oldest_req(mdsc); - while (req && req->r_tid <=3D want_tid) { - /* find next request */ - n =3D rb_next(&req->r_node); - if (n) - nextreq =3D rb_entry(n, struct ceph_mds_request, r_node); - else - nextreq =3D NULL; - if (req->r_op !=3D CEPH_MDS_OP_SETFILELOCK && - (req->r_op & CEPH_MDS_OP_WRITE)) { + idx =3D 0; + while ((req =3D xa_find(&mdsc->request_tree, &idx, want_tid, + XA_PRESENT))) { + u64 next_tid =3D req->r_tid + 1; + + if (req->r_op =3D=3D CEPH_MDS_OP_SETFILELOCK || + !(req->r_op & CEPH_MDS_OP_WRITE)) { + idx =3D next_tid; + continue; + } + + { struct ceph_mds_session *s =3D req->r_session; =20 if (!s) { - req =3D nextreq; + idx =3D next_tid; continue; } =20 /* write op */ ceph_mdsc_get_request(req); - if (nextreq) - ceph_mdsc_get_request(nextreq); s =3D ceph_get_mds_session(s); mutex_unlock(&mdsc->mutex); =20 @@ -6649,16 +6643,9 @@ static void flush_mdlog_and_wait_mdsc_unsafe_request= s(struct ceph_mds_client *md =20 mutex_lock(&mdsc->mutex); ceph_mdsc_put_request(req); - if (!nextreq) - break; /* next dne before, so we're done! */ - if (RB_EMPTY_NODE(&nextreq->r_node)) { - /* next request was removed from tree */ - ceph_mdsc_put_request(nextreq); - goto restart; - } - ceph_mdsc_put_request(nextreq); /* won't go away */ + /* restart from the next tid; tree may have changed */ + idx =3D next_tid; } - req =3D nextreq; } mutex_unlock(&mdsc->mutex); ceph_put_mds_session(last_session); @@ -6817,6 +6804,7 @@ static void ceph_mdsc_stop(struct ceph_mds_client *md= sc) if (mdsc->mdsmap) ceph_mdsmap_destroy(mdsc->mdsmap); kfree(mdsc->sessions); + xa_destroy(&mdsc->request_tree); ceph_caps_finalize(mdsc); =20 if (mdsc->s_cap_auths) { diff --git a/fs/ceph/mds_client.h b/fs/ceph/mds_client.h index 0ece4c9e3529..976dc9ffac17 100644 --- a/fs/ceph/mds_client.h +++ b/fs/ceph/mds_client.h @@ -330,7 +330,6 @@ typedef int (*ceph_mds_request_wait_callback_t) (struct= ceph_mds_client *mdsc, */ struct ceph_mds_request { u64 r_tid; /* transaction id */ - struct rb_node r_node; struct ceph_mds_client *r_mdsc; =20 struct kref r_kref; @@ -535,7 +534,7 @@ struct ceph_mds_client { u64 last_tid; /* most recent mds request */ u64 oldest_tid; /* oldest incomplete mds request, excluding setfilelock requests */ - struct rb_root request_tree; /* pending mds requests */ + struct xarray request_tree; /* pending mds requests */ struct delayed_work delayed_work; /* delayed work */ unsigned long last_renew_caps; /* last time we renewed our caps */ struct list_head cap_delay_list; /* caps with delayed release */ --=20 2.53.0 From nobody Tue Sep 29 07:39:37 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5980A31717D; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; cv=none; b=SV7RddktaWh9S4kUSAOpzIBd4+bFWIDFBeVLwwTJTMgopdEukravw+1G3NK2jg5UtwMsNPQSQTQGoPBfV8y+XTR9nVPeMawT7GZ/3rFQDC/qjdNUyPftCZS/iaIFZn6v1JZNoQamJcTVWW3dUMNFxwh9H5DE/PBbplbbNbPqmqE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; c=relaxed/simple; bh=YJelZ7kf8NdTQa4sh+hlr//q0TCx1ycHMJtwSwZPTDk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=u2aWyI+KEAnXkJf54PkB+JES5VKhUJDBnSDctkmBrX5ln6O3hke4INq70JpZg8i9pKTa6uEkHPqIyhxGnre72AujqkxEX8UCGtG+I9gw84JQia9hHUdoVn5np6Z02rV0T4mhdsOTr3YaWlF5was/WgQ/1Al7En0Hp2uA4UG9+FI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ZTXrT8u1; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ZTXrT8u1" Received: by smtp.kernel.org (Postfix) with ESMTPS id 1FA7EC2BCFA; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786425477; bh=YJelZ7kf8NdTQa4sh+hlr//q0TCx1ycHMJtwSwZPTDk=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ZTXrT8u17kP3NewWHLlOfw4YjJO/0P4QV3RMHkTDLcSMDQfe9/kFmvNi8DbjDxvuk t6I2ChdXEGtmZxH8UT08eCQzOuwCxaNCiFDaSpGhUUBtWEKEPuY1RCEUbN3ZUFLU6h u3Lio6eA/edxXeqgaZxkC0XdpER50xmLgmyzjF6Zt91CaNtnLgIIvgPFbXbBqSzgIM Wkni+EFuQ7fmq8EBOAW67KdGQMI/Sit0JOWlX+Mh+91c4/BaYa9nfZW2+aafXt7Vxx m53AWQ886PDHKg00ky6hsBN9H6jzqxfrOI3m0sMiWbDY/qaiC0FmBp/3XDnlqWIMIV EiAgeyLUB7gPw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 08112C5AD55; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Tue, 11 Aug 2026 13:17:47 +0800 Subject: [PATCH v3 3/5] ceph: add wait_list_lock for wait-list serialization Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-ceph-mdsc-mutex-optimization-v3-3-d031114419f4@clyso.com> References: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> In-Reply-To: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786425473; l=4398; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=Jfam+nJ/bsXm8oYLqWWTI70QiKEhx69kaUtydPwxx54=; b=LQp0pJADQv7MsFthHxayVOrMsuO4ch5pe4nuh/wfOgGQvmb2M7ipA/R8W0du8qVb0z855SWXe q1LhL7FtStKCNkvU8DmgjsFLJkMEz2Ew2fBniy6nieaFZWPg5blTZw/ X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li The per-MDS session wait list and the global waiting-for-map list are currently serialized by mdsc->mutex, even though the list operations themselves don't need the mutex's broader protection. Introduce a dedicated spinlock to guard these lists so that waking and kicking waiters can run outside the mutex. Reviewed-by: Viacheslav Dubeyko Signed-off-by: Xiubo Li --- fs/ceph/mds_client.c | 18 +++++++++++++++++- fs/ceph/mds_client.h | 3 +++ 2 files changed, 20 insertions(+), 1 deletion(-) diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c index bf63f5c1fca2..2d2994b4399b 100644 --- a/fs/ceph/mds_client.c +++ b/fs/ceph/mds_client.c @@ -3618,7 +3618,9 @@ static void __do_request(struct ceph_mds_client *mdsc, doutc(cl, "no mdsmap, waiting for map\n"); trace_ceph_mdsc_suspend_request(mdsc, session, req, ceph_mdsc_suspend_reason_no_mdsmap); + spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &mdsc->waiting_for_map); + spin_unlock(&mdsc->wait_list_lock); return; } if (!(mdsc->fsc->mount_options->flags & @@ -3641,7 +3643,9 @@ static void __do_request(struct ceph_mds_client *mdsc, doutc(cl, "no mds or not active, waiting for map\n"); trace_ceph_mdsc_suspend_request(mdsc, session, req, ceph_mdsc_suspend_reason_no_active_mds); + spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &mdsc->waiting_for_map); + spin_unlock(&mdsc->wait_list_lock); return; } =20 @@ -3689,9 +3693,12 @@ static void __do_request(struct ceph_mds_client *mds= c, if (ceph_test_mount_opt(mdsc->fsc, CLEANRECOVER)) { trace_ceph_mdsc_suspend_request(mdsc, session, req, ceph_mdsc_suspend_reason_rejected); + spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &mdsc->waiting_for_map); - } else + spin_unlock(&mdsc->wait_list_lock); + } else { err =3D -EACCES; + } goto out_session; } =20 @@ -3706,7 +3713,9 @@ static void __do_request(struct ceph_mds_client *mdsc, } trace_ceph_mdsc_suspend_request(mdsc, session, req, ceph_mdsc_suspend_reason_session); + spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &session->s_waiting); + spin_unlock(&mdsc->wait_list_lock); goto out_session; } =20 @@ -3799,7 +3808,9 @@ static void __wake_requests(struct ceph_mds_client *m= dsc, struct ceph_mds_request *req; LIST_HEAD(tmp_list); =20 + spin_lock(&mdsc->wait_list_lock); list_splice_init(head, &tmp_list); + spin_unlock(&mdsc->wait_list_lock); =20 while (!list_empty(&tmp_list)) { req =3D list_entry(tmp_list.next, @@ -3832,7 +3843,9 @@ static void kick_requests(struct ceph_mds_client *mds= c, int mds) if (req->r_session && req->r_session->s_mds =3D=3D mds) { doutc(cl, " kicking tid %llu\n", req->r_tid); + spin_lock(&mdsc->wait_list_lock); list_del_init(&req->r_wait); + spin_unlock(&mdsc->wait_list_lock); trace_ceph_mdsc_resume_request(mdsc, req); __do_request(mdsc, req); } @@ -6277,6 +6290,7 @@ int ceph_mdsc_init(struct ceph_fs_client *fsc) mdsc->snap_realms =3D RB_ROOT; INIT_LIST_HEAD(&mdsc->snap_empty); spin_lock_init(&mdsc->snap_empty_lock); + spin_lock_init(&mdsc->wait_list_lock); xa_init(&mdsc->request_tree); INIT_DELAYED_WORK(&mdsc->delayed_work, delayed_work); mdsc->last_renew_caps =3D jiffies; @@ -6359,7 +6373,9 @@ static void wait_requests(struct ceph_mds_client *mds= c) mutex_lock(&mdsc->mutex); while ((req =3D __get_oldest_req(mdsc))) { doutc(cl, "timed out on tid %llu\n", req->r_tid); + spin_lock(&mdsc->wait_list_lock); list_del_init(&req->r_wait); + spin_unlock(&mdsc->wait_list_lock); __unregister_request(mdsc, req); } } diff --git a/fs/ceph/mds_client.h b/fs/ceph/mds_client.h index 976dc9ffac17..19d2eae9da5b 100644 --- a/fs/ceph/mds_client.h +++ b/fs/ceph/mds_client.h @@ -530,6 +530,9 @@ struct ceph_mds_client { struct list_head snap_empty; int num_snap_realms; spinlock_t snap_empty_lock; /* protect snap_empty */ + spinlock_t wait_list_lock; /* protect waiting_for_map + * and s_waiting lists + */ =20 u64 last_tid; /* most recent mds request */ u64 oldest_tid; /* oldest incomplete mds request, --=20 2.53.0 From nobody Tue Sep 29 07:39:37 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 595C92E06D2; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; cv=none; b=mgS+HW1BE8OaY1UMFRqnIt56NS8ZgN8EPDLHeYbyZJLpT0ky1vm+hwcaoyv77E+7C694+o0VTzbtjsfclvzj16NbzPxsbkKQbsEw7jQ0re6FnkiVAv+/+al9J0SvU8vM9sl/fpOSGKgm4yM9P7pAcGNuvBhZfrpX8VrR766pz4U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; c=relaxed/simple; bh=UuGkmM89NH+7Jbzo5H/Rzfqgh6m9HRlCoP64h2SCn4o=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=PjkeZtic7CIGT81QZCvs3g45hpeb4h7n3orz7ybttqrYZlYnoIJaCtAKoIXsaFZ0fgjrlCraGKqy4qDSLstXoGSKbVwLfJbm4OiXAfh8ucAtxtbTjIVbLzvg9HXbfxZOJl2Sg0EFekpVn+Ibjj32LMEaDgOqiQla76jQMXAqocE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=BURUu4N6; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="BURUu4N6" Received: by smtp.kernel.org (Postfix) with ESMTPS id 2E19EC2BCFC; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786425477; bh=UuGkmM89NH+7Jbzo5H/Rzfqgh6m9HRlCoP64h2SCn4o=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=BURUu4N6dLfyjY2KMHKfLprtL1HUf5b+4H3PXzgPhtN44WS3BN/i+iY5f5yAnHjIK OsDWf41g+ylISPKTJsf37xmT4P1NGeud+tvVbp9GbNhJNoCAjhqkCfSuZXdIyN4R52 MSDkEDDelYiBSIy4eAGZTe6yOL9Xl68e9OD97C7AlCOoki+EuWEhOBaelyeYUlWmq/ ZBDpm4ERsWN3BwZHQ350JKJ5zufUXBk2tzgDh7D7aHsA3CNRKC2TG2XNnL83pVjFpQ hu0pswz1N/HG4cPpjYZUTiZ3p4QaCAwudbmENRojZQ3aY14jIF5R4ytz+ySTP136rR fdO6TzAU8Ri4Q== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1A750C5B570; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Tue, 11 Aug 2026 13:17:48 +0800 Subject: [PATCH v3 4/5] ceph: move mdsc->mutex into __do_request() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-ceph-mdsc-mutex-optimization-v3-4-d031114419f4@clyso.com> References: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> In-Reply-To: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786425473; l=10269; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=ZiU/ETqcSu60DlAGSZPWXHglQAGqMwVQOvzb9DzX8Qk=; b=PmbM5C4Q13BX03yiwvvneRG4m4SCwx3+DkaGOLwZniXXEjRt27QLSwBxynER01pKG80lQIYUW WygrqRWDCaRA0tw3F1JVdVuoh5L3rGVxK5dgqskgEKICQSqeyxIKD5I X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li Currently every caller must hold mdsc->mutex when invoking the request-send machinery. Move the mutex acquisition inside __do_request() so that callers can fire off a request without first serializing on the global lock. The mutex is released before the network send phase and re-acquired only for cleanup, so dentry traversal and message construction run concurrently across CPUs. This is the primary source of the observed 2x stat throughput improvement: the per-request send path shrinks from hundreds of microseconds to tens of microseconds once it no longer waits on the mutex. Wait-list draining, session-state wake-ups, and request kicking are reworked to either use the new wait-list spinlock or collect candidates under the mutex and process them outside it. Signed-off-by: Xiubo Li --- fs/ceph/mds_client.c | 69 ++++++++++++++++++++++++++++++++----------------= ---- 1 file changed, 42 insertions(+), 27 deletions(-) diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c index 2d2994b4399b..76dfd2c86392 100644 --- a/fs/ceph/mds_client.c +++ b/fs/ceph/mds_client.c @@ -3586,9 +3586,12 @@ static void __do_request(struct ceph_mds_client *mds= c, int err =3D 0; bool random; =20 + mutex_lock(&mdsc->mutex); + if (req->r_err || test_bit(CEPH_MDS_R_GOT_RESULT, &req->r_req_flags)) { if (test_bit(CEPH_MDS_R_ABORTED, &req->r_req_flags)) __unregister_request(mdsc, req); + mutex_unlock(&mdsc->mutex); return; } =20 @@ -3621,6 +3624,7 @@ static void __do_request(struct ceph_mds_client *mdsc, spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &mdsc->waiting_for_map); spin_unlock(&mdsc->wait_list_lock); + mutex_unlock(&mdsc->mutex); return; } if (!(mdsc->fsc->mount_options->flags & @@ -3646,6 +3650,7 @@ static void __do_request(struct ceph_mds_client *mdsc, spin_lock(&mdsc->wait_list_lock); list_add(&req->r_wait, &mdsc->waiting_for_map); spin_unlock(&mdsc->wait_list_lock); + mutex_unlock(&mdsc->mutex); return; } =20 @@ -3719,6 +3724,8 @@ static void __do_request(struct ceph_mds_client *mdsc, goto out_session; } =20 + mutex_unlock(&mdsc->mutex); + /* send request */ req->r_resend_mds =3D -1; /* forget any previous mds hint */ =20 @@ -3749,6 +3756,7 @@ static void __do_request(struct ceph_mds_client *mdsc, err =3D wait_on_bit(&di->flags, CEPH_DENTRY_ASYNC_CREATE_BIT, TASK_KILLABLE); if (err) { + mutex_lock(&mdsc->mutex); mutex_lock(&req->r_fill_mutex); set_bit(CEPH_MDS_R_ABORTED, &req->r_req_flags); mutex_unlock(&req->r_fill_mutex); @@ -3786,6 +3794,8 @@ static void __do_request(struct ceph_mds_client *mdsc, =20 err =3D __send_request(session, req, false); =20 + mutex_lock(&mdsc->mutex); + out_session: ceph_put_mds_session(session); finish: @@ -3795,12 +3805,10 @@ static void __do_request(struct ceph_mds_client *md= sc, complete_request(mdsc, req); __unregister_request(mdsc, req); } + mutex_unlock(&mdsc->mutex); return; } =20 -/* - * called under mdsc->mutex - */ static void __wake_requests(struct ceph_mds_client *mdsc, struct list_head *head) { @@ -3830,10 +3838,14 @@ static void __wake_requests(struct ceph_mds_client = *mdsc, static void kick_requests(struct ceph_mds_client *mdsc, int mds) { struct ceph_client *cl =3D mdsc->fsc->client; - struct ceph_mds_request *req; + struct ceph_mds_request *req, *nreq; unsigned long idx; + LIST_HEAD(kick_list); =20 doutc(cl, "kick_requests mds%d\n", mds); + + /* collect matching requests under the mutex */ + mutex_lock(&mdsc->mutex); idx =3D 0; xa_for_each(&mdsc->request_tree, idx, req) { if (test_bit(CEPH_MDS_R_GOT_UNSAFE, &req->r_req_flags)) @@ -3842,14 +3854,23 @@ static void kick_requests(struct ceph_mds_client *m= dsc, int mds) continue; /* only new requests */ if (req->r_session && req->r_session->s_mds =3D=3D mds) { - doutc(cl, " kicking tid %llu\n", req->r_tid); + ceph_mdsc_get_request(req); spin_lock(&mdsc->wait_list_lock); list_del_init(&req->r_wait); spin_unlock(&mdsc->wait_list_lock); - trace_ceph_mdsc_resume_request(mdsc, req); - __do_request(mdsc, req); + list_add_tail(&req->r_wait, &kick_list); } } + mutex_unlock(&mdsc->mutex); + + /* replay without the mutex */ + list_for_each_entry_safe(req, nreq, &kick_list, r_wait) { + doutc(cl, " kicking tid %llu\n", req->r_tid); + trace_ceph_mdsc_resume_request(mdsc, req); + __do_request(mdsc, req); + list_del_init(&req->r_wait); + ceph_mdsc_put_request(req); + } } =20 int ceph_mdsc_submit_request(struct ceph_mds_client *mdsc, struct inode *d= ir, @@ -3909,10 +3930,11 @@ int ceph_mdsc_submit_request(struct ceph_mds_client= *mdsc, struct inode *dir, doutc(cl, "submit_request on %p for inode %p\n", req, dir); mutex_lock(&mdsc->mutex); __register_request(mdsc, req, dir); + mutex_unlock(&mdsc->mutex); + trace_ceph_mdsc_submit_request(mdsc, req); __do_request(mdsc, req); err =3D req->r_err; - mutex_unlock(&mdsc->mutex); return err; } =20 @@ -4295,13 +4317,14 @@ static void handle_forward(struct ceph_mds_client *= mdsc, req->r_num_fwd =3D fwd_seq; req->r_resend_mds =3D next_mds; put_request_session(req); - __do_request(mdsc, req); } mutex_unlock(&mdsc->mutex); =20 /* kick calling process */ if (aborted) complete_request(mdsc, req); + else if (!test_bit(CEPH_MDS_R_ABORTED, &req->r_req_flags)) + __do_request(mdsc, req); ceph_mdsc_put_request(req); return; =20 @@ -4625,11 +4648,9 @@ static void handle_session(struct ceph_mds_session *= session, =20 mutex_unlock(&session->s_mutex); if (wake) { - mutex_lock(&mdsc->mutex); __wake_requests(mdsc, &session->s_waiting); if (wake =3D=3D 2) kick_requests(mdsc, mds); - mutex_unlock(&mdsc->mutex); } if (op =3D=3D CEPH_SESSION_CLOSE) ceph_put_mds_session(session); @@ -5264,9 +5285,7 @@ static int send_mds_reconnect(struct ceph_mds_client = *mdsc, =20 mutex_unlock(&session->s_mutex); =20 - mutex_lock(&mdsc->mutex); __wake_requests(mdsc, &session->s_waiting); - mutex_unlock(&mdsc->mutex); =20 up_read(&mdsc->snap_rwsem); ceph_pagelist_release(recon_state.pagelist); @@ -5691,8 +5710,8 @@ static void ceph_mdsc_reset_workfn(struct work_struct= *work) } sessions[i]->s_state =3D CEPH_MDS_SESSION_CLOSED; __unregister_session(mdsc, sessions[i]); - __wake_requests(mdsc, &sessions[i]->s_waiting); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &sessions[i]->s_waiting); =20 mutex_lock(&sessions[i]->s_mutex); cleanup_session_requests(mdsc, sessions[i]); @@ -5703,9 +5722,7 @@ static void ceph_mdsc_reset_workfn(struct work_struct= *work) =20 ceph_put_mds_session(sessions[i]); =20 - mutex_lock(&mdsc->mutex); kick_requests(mdsc, mds); - mutex_unlock(&mdsc->mutex); =20 torn_down++; pr_info_client(cl, "mds%d session reset complete\n", mds); @@ -5821,8 +5838,8 @@ static void check_new_map(struct ceph_mds_client *mds= c, /* force close session for stopped mds */ ceph_get_mds_session(s); __unregister_session(mdsc, s); - __wake_requests(mdsc, &s->s_waiting); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &s->s_waiting); =20 mutex_lock(&s->s_mutex); cleanup_session_requests(mdsc, s); @@ -5831,8 +5848,8 @@ static void check_new_map(struct ceph_mds_client *mds= c, =20 ceph_put_mds_session(s); =20 - mutex_lock(&mdsc->mutex); kick_requests(mdsc, i); + mutex_lock(&mdsc->mutex); continue; } =20 @@ -5876,8 +5893,8 @@ static void check_new_map(struct ceph_mds_client *mds= c, oldstate !=3D CEPH_MDS_STATE_STARTING) pr_info_client(cl, "mds%d recovery completed\n", s->s_mds); - kick_requests(mdsc, i); mutex_unlock(&mdsc->mutex); + kick_requests(mdsc, i); mutex_lock(&s->s_mutex); mutex_lock(&mdsc->mutex); ceph_kick_flushing_caps(mdsc, s); @@ -6785,8 +6802,8 @@ void ceph_mdsc_force_umount(struct ceph_mds_client *m= dsc) =20 if (session->s_state =3D=3D CEPH_MDS_SESSION_REJECTED) __unregister_session(mdsc, session); - __wake_requests(mdsc, &session->s_waiting); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &session->s_waiting); =20 mutex_lock(&session->s_mutex); __close_session(mdsc, session); @@ -6797,11 +6814,11 @@ void ceph_mdsc_force_umount(struct ceph_mds_client = *mdsc) mutex_unlock(&session->s_mutex); ceph_put_mds_session(session); =20 - mutex_lock(&mdsc->mutex); kick_requests(mdsc, mds); + mutex_lock(&mdsc->mutex); } - __wake_requests(mdsc, &mdsc->waiting_for_map); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &mdsc->waiting_for_map); } =20 static void ceph_mdsc_stop(struct ceph_mds_client *mdsc) @@ -6943,8 +6960,8 @@ void ceph_mdsc_handle_fsmap(struct ceph_mds_client *m= dsc, struct ceph_msg *msg) err_out: mutex_lock(&mdsc->mutex); mdsc->mdsmap_err =3D err; - __wake_requests(mdsc, &mdsc->waiting_for_map); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &mdsc->waiting_for_map); } =20 /* @@ -6995,11 +7012,11 @@ void ceph_mdsc_handle_mdsmap(struct ceph_mds_client= *mdsc, struct ceph_msg *msg) mdsc->fsc->max_file_size =3D min((loff_t)mdsc->mdsmap->m_max_file_size, MAX_LFS_FILESIZE); =20 + mutex_unlock(&mdsc->mutex); __wake_requests(mdsc, &mdsc->waiting_for_map); ceph_monc_got_map(&mdsc->fsc->client->monc, CEPH_SUB_MDSMAP, mdsc->mdsmap->m_epoch); =20 - mutex_unlock(&mdsc->mutex); schedule_delayed(mdsc, 0); return; =20 @@ -7097,8 +7114,8 @@ static void mds_peer_reset(struct ceph_connection *co= n) ceph_get_mds_session(s); s->s_state =3D CEPH_MDS_SESSION_CLOSED; __unregister_session(mdsc, s); - __wake_requests(mdsc, &s->s_waiting); mutex_unlock(&mdsc->mutex); + __wake_requests(mdsc, &s->s_waiting); =20 mutex_lock(&s->s_mutex); cleanup_session_requests(mdsc, s); @@ -7107,9 +7124,7 @@ static void mds_peer_reset(struct ceph_connection *co= n) =20 wake_up_all(&mdsc->session_close_wq); =20 - mutex_lock(&mdsc->mutex); kick_requests(mdsc, s->s_mds); - mutex_unlock(&mdsc->mutex); =20 ceph_put_mds_session(s); break; --=20 2.53.0 From nobody Tue Sep 29 07:39:37 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8F47A33C513; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; cv=none; b=h/uLwHfwldjjonFYgWCmWSe30F0Rtcmexh5undFlzhxOv6X2NwJ+8lQ/YHqNe8d8HlMRZ92D3nu/rmjzbnYix9b4vkwhKyIx/m5sEg5vSDU+W0Tfl6O0H9T8MYOCJqHfM4rbxhgyCNCXXsPKfm7hicYKJhWw7q6RiFdQ3vimHbU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786425477; c=relaxed/simple; bh=xZIGNd2ACYvNHPTnMaupN5d/PGOCAJY3rA5hJeos6rQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=AJm2JHcD0szNkpg+54vHKXCoUo66fPhg+CZM+gKchx0GYxAoKCn/vgHD4DFZH15urj5wk7DYualAUUHa8Ug9wsGowepQOMQdSGLSx2g/cX5wAJMKnTrRaXJrVeFsEB3cUD4Zf5UGKUmD0khXss3eEsqlqUTVcehT+6svL8M5eAw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ogunXcs1; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ogunXcs1" Received: by smtp.kernel.org (Postfix) with ESMTPS id 3A580C2BCFD; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786425477; bh=xZIGNd2ACYvNHPTnMaupN5d/PGOCAJY3rA5hJeos6rQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ogunXcs1RuCtlbZbvauHb4effVdL72a3eNn6X+LzT1wSod0sxsuKwu4OZHAslHNWs rrZ+QZtPLmvmwEDOqwmYaCnFaz5cQKJbkAAItEQGpWoFspcrcODxqzzEeZMSF6Ngpr 2z6eoojbOSc2LuT73FaGXtKub03QX1yWaEhvtefGtFoeMfvfme1hWnRh8HMaoCg544 5GvUZuDPf4vCFI2VAlWsbiFMlOpftOdHnpDlQHGEeeOHbgrb1IufTVJaLp5kYfJkjN EstEh3BbG1UeWEq3MEz17pnG6ObmRhkaoqsWWlFt7uChnLAyYujJJ/rGqRnQzur1pU KnI0W3dUHgfwg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 27DA1C5AD7B; Tue, 11 Aug 2026 05:17:57 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Tue, 11 Aug 2026 13:17:49 +0800 Subject: [PATCH v3 5/5] ceph: narrow mdsc->mutex scope in replay_unsafe_requests Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-ceph-mdsc-mutex-optimization-v3-5-d031114419f4@clyso.com> References: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> In-Reply-To: <20260811-ceph-mdsc-mutex-optimization-v3-0-d031114419f4@clyso.com> To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786425473; l=2687; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=W6+X1JGkKuujMu/zlmW/Tgr6nBaTcr4f6OKmGGGk9no=; b=uwbDXWhPuBN6FGdXC6ib3qyjVDLPPXQOVzii07mS+ePQIwUInUDSzejN2a8yc98v3ukFdxe+W R49ezb7Oxq/CVV3OebxSAXU0DqJPGf5GukIXsdmK9zdsuWE121YqVvd X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li Currently replay_unsafe_requests() holds mdsc->mutex across the entire function, including the __send_request() calls. Since __send_request() is lockless and the async cap-release helper schedules deferred work, neither needs the mutex. Collect the unsafe-list entries and the matching old xarray entries into local lists under mdsc->mutex, taking a reference on each, then replay them outside the mutex. Taking a reference ensures a concurrent reply handler can complete and unregister a request without invalidating the local list or the iterator. Signed-off-by: Xiubo Li --- fs/ceph/mds_client.c | 34 ++++++++++++++++++++++++++++++---- 1 file changed, 30 insertions(+), 4 deletions(-) diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c index 76dfd2c86392..d8eb9154ae26 100644 --- a/fs/ceph/mds_client.c +++ b/fs/ceph/mds_client.c @@ -4702,12 +4702,23 @@ static void replay_unsafe_requests(struct ceph_mds_= client *mdsc, { struct ceph_mds_request *req, *nreq; unsigned long idx; + LIST_HEAD(unsafe_list); + LIST_HEAD(old_list); =20 doutc(mdsc->fsc->client, "mds%d\n", session->s_mds); =20 + /* + * Collect unsafe and old requests under mdsc->mutex, then + * replay them without it: __send_request() is lockless and + * ceph_mdsc_release_dir_caps_async() schedules work. + */ mutex_lock(&mdsc->mutex); - list_for_each_entry_safe(req, nreq, &session->s_unsafe, r_unsafe_item) - __send_request(session, req, true); + + list_for_each_entry_safe(req, nreq, &session->s_unsafe, + r_unsafe_item) { + ceph_mdsc_get_request(req); + list_move(&req->r_unsafe_item, &unsafe_list); + } =20 /* * also re-send old requests when MDS enters reconnect stage. So that MDS @@ -4724,11 +4735,26 @@ static void replay_unsafe_requests(struct ceph_mds_= client *mdsc, if (req->r_session->s_mds !=3D session->s_mds) continue; =20 - ceph_mdsc_release_dir_caps_async(req); + ceph_mdsc_get_request(req); + list_add_tail(&req->r_unsafe_item, &old_list); + } =20 + mutex_unlock(&mdsc->mutex); + + /* replay unsafe requests */ + list_for_each_entry_safe(req, nreq, &unsafe_list, r_unsafe_item) { __send_request(session, req, true); + list_del_init(&req->r_unsafe_item); + ceph_mdsc_put_request(req); + } + + /* replay old requests */ + list_for_each_entry_safe(req, nreq, &old_list, r_unsafe_item) { + ceph_mdsc_release_dir_caps_async(req); + __send_request(session, req, true); + list_del_init(&req->r_unsafe_item); + ceph_mdsc_put_request(req); } - mutex_unlock(&mdsc->mutex); } =20 static int send_reconnect_partial(struct ceph_reconnect_state *recon_state) --=20 2.53.0