From nobody Mon Sep 28 13:17:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 41E77175A9C; Fri, 21 Aug 2026 06:37:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787294277; cv=none; b=Q1Qra0l3TiFHej1zrE5WEyluJJnRGLB1y3xKXwSLfFOBwp24w6SZ7kT1RGlGg07ZzCExJKVOyZPLQDtlvDSaj/pbRsOaDfhkoRf5fQLYw6VCV4XF4fRjDk3e2miVT07UM4xbqRU/ccfW2C6RNfOpyo6LEifxp34YswFcuF4FAMc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787294277; c=relaxed/simple; bh=MY8KvlUpORqCkdMKJJ6n1UIF8KV/S73TfZGZRcCI66Y=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:To:Cc; b=Y/aW19wjcAWwAmRv55Y9pcvqZbl974b+7SG1vVOgp62+TuZCqAQU+4fmBost7eIvH7jwPn1O0DIrpBgQZFiG+k3T6My+XWSap0okTX8+02LFTKF1tkzBDSQp8MEh4ONJBvxyOeqvll25nUsvJmr+7yoXvtZ7Fu3v7FpY3k4GRkg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=EpykeFEH; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="EpykeFEH" Received: by smtp.kernel.org (Postfix) with ESMTPS id B8438C19425; Fri, 21 Aug 2026 06:37:56 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1787294276; bh=MY8KvlUpORqCkdMKJJ6n1UIF8KV/S73TfZGZRcCI66Y=; h=From:Date:Subject:To:Cc:Reply-To:From; b=EpykeFEHt1ZpLIE8J5xDLZBtVSemzYzQQDAkY6EzCEbY1hx2a1uOV8+GFkppS8keb OZbs3qwjzHEBw2j0eDht0zRw60imrPNqgwYZTHjmVWScYfTobkxk5wOKMtZ5rdNvcM vMkmaBI3Ox6oW49ZD9YUb19kt8MgCx/3ILzzKfYnyBn7BkH9FQ/Fc12KZX7aIS86PC 1I6GSx7y2JWAb6XK8nr/Z6K2ZZqDEy672swxx5jU2nxUALteMfjwvWMVhcaGiP2ccM JgMTFYNRLbe/zGcbjzwncXG0xPoHm/2x2O/VdMux59mZl5bZN5HAivtqfB5XnZgeQv oI3UKhBNxkTxQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 93AB0C5DF87; Fri, 21 Aug 2026 06:37:56 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Thu, 20 Aug 2026 23:37:54 -0700 Subject: [PATCH v4] ceph: add 'lazyio' mount option to kclient Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260820-lazyio-v4-1-601bee1ea2c8@clyso.com> X-B4-Tracking: v=1; b=H4sIAAAAAAAC/2XNSw6CMBgE4KuQrq3p30JbXHkP44I+kBqkhiIRC Xe3YAJGl5PMNzOiYFtnAzokI2pt74LzTQzpLkG6KpqLxc7EjCihnHCa4bp4Dc5jnktRCqazTJQ olu+tLd1zGTqdPzk81NXqbtZzo3Kh8+2wPPUw9/5Ge8CAjQJGhBQyN+ao6yH4vfY3NI/2dGOCw MpoZLnmXEieSpKpX8Y2JiFfGYuMAgAxTOhU2W82TdMbTL1lhxwBAAA= X-Change-ID: 20260625-lazyio-6987f73c557f To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1787294274; l=21050; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=GbOEd8WI47DuNGb4LjIlYqplcC/o7uhJgziZYT8chkQ=; b=tNta7ZiOwvEHYbbl0hnvGsISYGNL29j4Ek3PSdo6DGSJmvcTLwzj17vO5QWcti/ArGkRJv3RH F0B48AMafSyCI33BVChfapa3xf1RTlIvDOl2PuEzaz1AWCXDRUiOB+q X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li Add a 'lazyio' mount option to the kernel Ceph client that enables LazyIO globally for all regular file opens on a mount. This is the kclient equivalent of the 'client_force_lazyio=3Dtrue' config option in the ceph-fuse userspace client. When 'lazyio' is specified, CEPH_FILE_MODE_LAZY is automatically added to every regular file's fmode at open time in ceph_open() and ceph_atomic_open(), causing the I/O paths to request CEPH_CAP_FILE_LAZYIO from the MDS. This permits buffered I/O via the page cache even when multiple clients have the file open for write =E2=80=94 beneficial for HPC workloads that can tolerate relaxed cache coherency. The mount option is exposed as 'lazyio' / 'nolazyio' via the VFS fsparam_flag_no mechanism and supports remount. Link: https://tracker.ceph.com/issues/77594 Signed-off-by: Xiubo Li --- The following is a test report: =3D=3D=3D LazyIO Multi-Client Read Test =3D=3D=3D (B writes =E2=86=92 MDS revokes CACHE from A; lazyio keeps LAZYIO =E2=86=92= buffered reads) nolazyio: 200 MB in 0.63s -> 319.8 MB/s lazyio : 200 MB in 0.24s -> 840.3 MB/s =3D=3D=3D IO500 tool ior test: =3D=3D=3D nolazyio lazyio Write: 74.49 MB/s -> 185.24 MB/s --- Changed in V4: - Base the CACHE/BUFFER -> LAZYIO used-cap substitution on implemented, not issued, so ceph_check_caps() holds the revoke ACK until writeback and invalidation complete; also queue writeback when revoking LAZYIO with dirty buffers. - In try_get_cap_refs(), let LAZYIO substitute for missing CACHE/BUFFER for callers that request LAZYIO, without dropping issued CACHE/BUFFER refs when LAZYIO is unavailable. - In handle_cap_grant(), invalidate on CACHE-then-LAZYIO revoke sequences and flush dirty data before ACKing when LAZYIO was covering it. - In ceph_page_mkwrite(), wait for BUFFER/LAZYIO to be (re)granted before dirtying the folio instead of proceeding with WR alone. - Request LAZYIO from the netfs read path (ceph_init_request) and ceph_renew_caps(). - Link to v3: https://patch.msgid.link/20260819-lazyio-v3-1-21110d37c4be@cl= yso.com Changes in v3: - Rework try_get_cap_refs() so a missing LAZYIO never makes a want unsatisfiable and silently drops CACHE/BUFFER refs (degrading to sync I/O= ). - Add ceph_adjust_caps_used_for_lazyio() and report LAZYIO in used caps to the MDS when it stands in for CACHE/BUFFER. - In handle_cap_grant(), flush dirty data before acking an LAZYIO revoke, and invalidate when either CACHE or LAZYIO is revoked and neither remains. - Define CEPH_O_LAZY as 00020000 (matching src/include/ceph_fs.h) and send it in the open request flags. - Link to v2: https://patch.msgid.link/20260701-lazyio-v2-1-9c667864805b@cl= yso.com Changes in v2: - Move LAZYIO fmode addition from ceph_init_file_info() to ceph_open() befo= re ceph_caps_for_mode(), so the MDS sees LAZYIO in the wanted caps at open t= ime. Also cover ceph_atomic_open() for the writer path. - Add ceph_adjust_caps_used_for_lazyio() to substitute LAZYIO for CACHE/BUF= FER in used caps reported to MDS when those caps are not issued. - Gate try_get_cap_refs() LAZYIO substitution on want & CEPH_CAP_FILE_LAZYIO so a non-lazy fd cannot have its consistency guarantees weakened by a laz= y fd on the same inode. - Separate revocation handling in handle_cap_grant(): BUFFER revocation alw= ays triggers writeback; LAZYIO revocation triggers writeback only when dirty = data is held (i_wrbuffer_ref/i_wb_ref); clean cached pages fall through to the existing invalidation path. - Add LAZYIO to ceph_init_request() so readahead works for lazy fds. - Fix fmode propagation: use fmode instead of req->r_fmode in ceph_open() M= DS path, and set req->r_fmode |=3D CEPH_FILE_MODE_LAZY for both ceph_open() = and ceph_atomic_open(). - Remove dead #ifdef O_LAZY blocks in ceph_flags_to_mode() and ceph_renew_c= aps(). - Link to v1: https://patch.msgid.link/20260625-lazyio-v1-1-db13078789dd@cl= yso.com To: Ilya Dryomov To: Alex Markuze To: Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org Cc: linux-kernel@vger.kernel.org --- fs/ceph/addr.c | 52 +++++++++++++++++++ fs/ceph/caps.c | 116 +++++++++++++++++++++++++++++++++++++++= ---- fs/ceph/file.c | 33 ++++++++++-- fs/ceph/super.c | 15 ++++++ fs/ceph/super.h | 1 + fs/ceph/util.c | 4 -- include/linux/ceph/ceph_fs.h | 1 + 7 files changed, 202 insertions(+), 20 deletions(-) diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c index 0a86f672cc09..ce105c4b02f4 100644 --- a/fs/ceph/addr.c +++ b/fs/ceph/addr.c @@ -491,6 +491,10 @@ static int ceph_init_request(struct netfs_io_request *= rreq, struct file *file) rreq->netfs_priv =3D priv; return 0; } + + /* If this is a lazy fd, also try to get LAZYIO caps */ + if (fi->fmode & CEPH_FILE_MODE_LAZY) + want |=3D CEPH_CAP_FILE_LAZYIO; } =20 /* @@ -2055,6 +2059,30 @@ static vm_fault_t ceph_filemap_fault(struct vm_fault= *vmf) return ret; } =20 +/* + * Return true if the MDS has issued us a cap that covers buffered + * dirtying: either BUFFER, or LAZYIO (as long as LAZYIO itself is not + * being revoked). This mirrors the conditions under which + * try_get_cap_refs() will hand out a BUFFER or LAZYIO ref. + */ +static bool ceph_have_dirtyable_caps(struct inode *inode) +{ + struct ceph_inode_info *ci =3D ceph_inode(inode); + int have, implemented; + bool ret =3D false; + + spin_lock(&ci->i_ceph_lock); + have =3D __ceph_caps_issued(ci, &implemented); + if (have & CEPH_CAP_FILE_BUFFER) { + ret =3D true; + } else if ((have & CEPH_CAP_FILE_LAZYIO) && + !((implemented & ~have) & CEPH_CAP_FILE_LAZYIO)) { + ret =3D true; + } + spin_unlock(&ci->i_ceph_lock); + return ret; +} + static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf) { struct vm_area_struct *vma =3D vmf->vma; @@ -2093,6 +2121,7 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *= vmf) else want =3D CEPH_CAP_FILE_BUFFER; =20 +retry_caps: got =3D 0; err =3D ceph_get_caps(vma->vm_file, CEPH_CAP_FILE_WR, want, off + len, &g= ot); if (err < 0) @@ -2101,6 +2130,29 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault = *vmf) doutc(cl, "%llx.%llx %llu~%zd got cap refs on %s\n", ceph_vinop(inode), off, len, ceph_cap_string(got)); =20 + /* + * ceph_write_iter() makes the same check and falls back to + * synchronous writes, but a page fault has no such fallback: + * dirtying the folio without BUFFER or LAZYIO would leave dirty + * data uncovered by any issued cap (e.g. while LAZYIO is being + * revoked, or after it has been released). Wait for the MDS to + * (re)grant a covering cap, matching how the exclude gate in + * try_get_cap_refs() blocks buffered writes while BUFFER is + * revoking. + */ + if ((fi->fmode & CEPH_FILE_MODE_LAZY) && + (got & (CEPH_CAP_FILE_BUFFER | CEPH_CAP_FILE_LAZYIO)) =3D=3D 0) { + ceph_put_cap_refs(ci, got); + got =3D 0; + doutc(cl, "%llx.%llx %llu~%zd waiting for BUFFER or LAZYIO\n", + ceph_vinop(inode), off, len); + err =3D wait_event_killable(ci->i_cap_wq, + ceph_have_dirtyable_caps(inode)); + if (err) + goto out_free; + goto retry_caps; + } + /* Update time before taking folio lock */ file_update_time(vma->vm_file); inode_inc_iversion_raw(inode); diff --git a/fs/ceph/caps.c b/fs/ceph/caps.c index d51454e995a8..612283d84399 100644 --- a/fs/ceph/caps.c +++ b/fs/ceph/caps.c @@ -999,6 +999,36 @@ int __ceph_caps_used(struct ceph_inode_info *ci) return used; } =20 +/* + * Substitute LAZYIO for CACHE/BUFFER when they are not issued. + * If we have LAZYIO but not CACHE/BUFFER, report LAZYIO as used instead + * so the MDS knows we're fine with the weaker consistency guarantee. + * + * Base the substitution on "implemented" rather than "issued": while + * LAZYIO is being revoked, "issued" no longer contains it but + * "implemented" still does. If used reverted to CACHE/BUFFER at that + * point, ceph_check_caps() would see (revoking & cap_used) =3D=3D 0 and + * ACK the revoke while dirty or stale pages were still present. Only + * once the revoke is ACKed does "implemented" drop LAZYIO. + */ +static inline int ceph_adjust_caps_used_for_lazyio(int used, int issued, + int implemented) +{ + if (!(used & (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_BUFFER))) + return used; + if (!(implemented & CEPH_CAP_FILE_LAZYIO)) + return used; + if (!(issued & CEPH_CAP_FILE_CACHE)) { + used &=3D ~CEPH_CAP_FILE_CACHE; + used |=3D CEPH_CAP_FILE_LAZYIO; + } + if (!(issued & CEPH_CAP_FILE_BUFFER)) { + used &=3D ~CEPH_CAP_FILE_BUFFER; + used |=3D CEPH_CAP_FILE_LAZYIO; + } + return used; +} + #define FMODE_WAIT_BIAS 1000 =20 /* @@ -2049,6 +2079,10 @@ void ceph_check_caps(struct ceph_inode_info *ci, int= flags) * usually because they have outstanding references). */ issued =3D __ceph_caps_issued(ci, &implemented); + + /* substitute LAZYIO for CACHE/BUFFER when they are not issued */ + used =3D ceph_adjust_caps_used_for_lazyio(used, issued, implemented); + revoking =3D implemented & ~issued; =20 want =3D file_wanted; @@ -2164,10 +2198,13 @@ void ceph_check_caps(struct ceph_inode_info *ci, in= t flags) * at most 5 seconds. That means the MDS needs to wait at * most 5 seconds to finished the Fb capability's revocation. * - * Let's queue a writeback for it. + * Let's queue a writeback for it. The same applies when + * LAZYIO is revoked while it was covering for BUFFER + * (dirty pages exist, but BUFFER isn't issued). */ if (S_ISREG(inode->i_mode) && ci->i_wrbuffer_ref && - (revoking & CEPH_CAP_FILE_BUFFER)) + (revoking & (CEPH_CAP_FILE_BUFFER | + CEPH_CAP_FILE_LAZYIO))) queue_writeback =3D true; } =20 @@ -2905,9 +2942,45 @@ static int try_get_cap_refs(struct inode *inode, int= need, int want, } snap_rwsem_locked =3D true; } - if ((have & want) =3D=3D want) + /* + * Allow LAZYIO to act as a substitute for + * CACHE or BUFFER when those caps are not + * issued, but only for callers that + * explicitly requested LAZYIO. This + * prevents a non-lazy fd from having its + * CACHE/BUFFER wants satisfied by LAZYIO + * on an inode where a different fd is lazy. + * + * A missing LAZYIO cap, however, must never + * cost us the CACHE/BUFFER refs that are + * actually issued: the MDS only grants + * LAZYIO for files opened with CEPH_O_LAZY + * and it can be revoked at any time. If we + * made the whole want unsatisfiable without + * it, the I/O paths would silently drop + * CACHE/BUFFER (e.g. take no wrbuffer refs) + * and degrade to synchronous writes. + */ + if ((have & (want & ~CEPH_CAP_FILE_LAZYIO)) =3D=3D + (want & ~CEPH_CAP_FILE_LAZYIO)) { + *got =3D need | (want & ~exclude & + ~CEPH_CAP_FILE_LAZYIO); + if ((want & CEPH_CAP_FILE_LAZYIO) && + (have & CEPH_CAP_FILE_LAZYIO) && + !(exclude & CEPH_CAP_FILE_LAZYIO)) + *got |=3D CEPH_CAP_FILE_LAZYIO; + } else if ((want & CEPH_CAP_FILE_LAZYIO) && + (have & CEPH_CAP_FILE_LAZYIO) && + !(exclude & CEPH_CAP_FILE_LAZYIO) && + ((have & want) =3D=3D + (want & ~(CEPH_CAP_FILE_CACHE | + CEPH_CAP_FILE_BUFFER)))) { + /* + * LAZYIO substitutes for missing CACHE/BUFFER; + * it is already included via (want & ~exclude). + */ *got =3D need | (want & ~exclude); - else + } else *got =3D need; ceph_take_cap_refs(ci, *got, true); ret =3D 1; @@ -3525,13 +3598,20 @@ static void handle_cap_grant(struct inode *inode, =20 =20 /* - * If CACHE is being revoked, and we have no dirty buffers, - * try to invalidate (once). (If there are dirty buffers, we - * will invalidate _after_ writeback.) + * Check the revocation of *both* CACHE and LAZYIO, because + * CACHE may have been revoked earlier and cap->issued no + * longer contains it -- at that point only LAZYIO was + * covering us. If LAZYIO is now also being revoked and no + * cache cap remains, we must invalidate the page cache. + * Without this, a CACHE-revoked-then-LAZYIO-revoked sequence + * leaves stale pages in memory until the next periodic + * check_caps (up to 60s). Also invalidate when we have no + * dirty buffers (if dirty, invalidate after writeback). */ if (S_ISREG(inode->i_mode) && /* don't invalidate readdir cache */ - ((cap->issued & ~newcaps) & CEPH_CAP_FILE_CACHE) && - (newcaps & CEPH_CAP_FILE_LAZYIO) =3D=3D 0 && + ((cap->issued & ~newcaps) & + (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO)) && + !(newcaps & (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO)) && !(ci->i_wrbuffer_ref || ci->i_wb_ref)) { if (try_nonblocking_invalidate(inode)) { /* there were locked pages.. invalidate later @@ -3675,6 +3755,7 @@ static void handle_cap_grant(struct inode *inode, /* check cap bits */ wanted =3D __ceph_caps_wanted(ci); used =3D __ceph_caps_used(ci); + used =3D ceph_adjust_caps_used_for_lazyio(used, cap->issued, cap->impleme= nted); dirty =3D __ceph_caps_dirty(ci); doutc(cl, " my wanted =3D %s, used =3D %s, dirty %s\n", ceph_cap_string(wanted), ceph_cap_string(used), @@ -3702,13 +3783,26 @@ static void handle_cap_grant(struct inode *inode, doutc(cl, "revocation: %s -> %s (revoking %s)\n", ceph_cap_string(cap->issued), ceph_cap_string(newcaps), ceph_cap_string(revoking)); + /* + * If BUFFER is being revoked and we have dirty data, + * trigger writeback before acking. When LAZYIO was + * covering for BUFFER (BUFFER not issued, dirty refs + * held), also trigger writeback. Clean cached pages + * under LAZYIO are handled by queue_invalidate below. + */ if (S_ISREG(inode->i_mode) && (revoking & used & CEPH_CAP_FILE_BUFFER)) { writeback =3D true; /* initiate writeback; will delay ack */ revoke_wait =3D true; + } else if (S_ISREG(inode->i_mode) && + (revoking & used & CEPH_CAP_FILE_LAZYIO) && + (ci->i_wrbuffer_ref || ci->i_wb_ref)) { + /* LAZYIO was covering for dirty data =E2=80=94 flush first */ + writeback =3D true; + revoke_wait =3D true; } else if (queue_invalidate && - revoking =3D=3D CEPH_CAP_FILE_CACHE && - (newcaps & CEPH_CAP_FILE_LAZYIO) =3D=3D 0) { + (revoking & (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO)) && + !(newcaps & (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO))) { revoke_wait =3D true; /* do nothing yet, invalidation will be queued */ } else if (cap =3D=3D ci->i_auth_cap) { check_caps =3D 1; /* check auth cap only */ diff --git a/fs/ceph/file.c b/fs/ceph/file.c index d54d71669176..4d5afd1e0277 100644 --- a/fs/ceph/file.c +++ b/fs/ceph/file.c @@ -346,10 +346,6 @@ int ceph_renew_caps(struct inode *inode, int fmode) flags =3D O_RDONLY; else if (wanted & CEPH_CAP_FILE_WR) flags =3D O_WRONLY; -#ifdef O_LAZY - if (wanted & CEPH_CAP_FILE_LAZYIO) - flags |=3D O_LAZY; -#endif =20 req =3D prepare_open_request(inode->i_sb, flags, 0); if (IS_ERR(req)) { @@ -357,6 +353,10 @@ int ceph_renew_caps(struct inode *inode, int fmode) goto out; } =20 + if (wanted & CEPH_CAP_FILE_LAZYIO) { + req->r_fmode |=3D CEPH_FILE_MODE_LAZY; + req->r_args.open.flags |=3D cpu_to_le32(CEPH_O_LAZY); + } req->r_inode =3D inode; ihold(inode); req->r_num_caps =3D 1; @@ -408,6 +408,19 @@ int ceph_open(struct inode *inode, struct file *file) doutc(cl, "%p %llx.%llx file %p flags %d (%d)\n", inode, ceph_vinop(inode), file, flags, file->f_flags); fmode =3D ceph_flags_to_mode(flags); + + /* + * If lazyio mount option is set, enable lazyio for all regular + * files. Skip snapped files: snap caps never include LAZYIO, + * so including it in wanted would force an unnecessary MDS + * round-trip for every open of a snapped file. + */ + if (S_ISREG(inode->i_mode) && + ceph_snap(inode) =3D=3D CEPH_NOSNAP && + (fsc->mount_options->flags & CEPH_MOUNT_OPT_LAZYIO)) { + fmode |=3D CEPH_FILE_MODE_LAZY; + } + wanted =3D ceph_caps_for_mode(fmode); =20 if (fmode & CEPH_FILE_MODE_WR) @@ -484,13 +497,16 @@ int ceph_open(struct inode *inode, struct file *file) err =3D PTR_ERR(req); goto out; } + req->r_fmode |=3D fmode & CEPH_FILE_MODE_LAZY; + if (fmode & CEPH_FILE_MODE_LAZY) + req->r_args.open.flags |=3D cpu_to_le32(CEPH_O_LAZY); req->r_inode =3D inode; ihold(inode); =20 req->r_num_caps =3D 1; err =3D ceph_mdsc_do_request(mdsc, NULL, req); if (!err) - err =3D ceph_init_file(inode, file, req->r_fmode); + err =3D ceph_init_file(inode, file, fmode); ceph_mdsc_put_request(req); doutc(cl, "open result=3D%d on %llx.%llx\n", err, ceph_vinop(inode)); out: @@ -834,6 +850,9 @@ int ceph_atomic_open(struct inode *dir, struct dentry *= dentry, } else { int fmode =3D ceph_flags_to_mode(flags); =20 + if (fsc->mount_options->flags & CEPH_MOUNT_OPT_LAZYIO) + fmode |=3D CEPH_FILE_MODE_LAZY; + mask =3D MAY_READ; if (fmode & CEPH_FILE_MODE_WR) mask |=3D MAY_WRITE; @@ -876,6 +895,10 @@ int ceph_atomic_open(struct inode *dir, struct dentry = *dentry, err =3D PTR_ERR(req); goto out_ctx; } + if (fsc->mount_options->flags & CEPH_MOUNT_OPT_LAZYIO) { + req->r_fmode |=3D CEPH_FILE_MODE_LAZY; + req->r_args.open.flags |=3D cpu_to_le32(CEPH_O_LAZY); + } req->r_dentry =3D dget(dentry); req->r_num_caps =3D 2; mask =3D CEPH_STAT_CAP_INODE | CEPH_CAP_AUTH_SHARED; diff --git a/fs/ceph/super.c b/fs/ceph/super.c index c05fbd4237f8..0bbd38933f0e 100644 --- a/fs/ceph/super.c +++ b/fs/ceph/super.c @@ -177,6 +177,7 @@ enum { Opt_wsync, Opt_pagecache, Opt_sparseread, + Opt_lazyio, }; =20 enum ceph_recover_session_mode { @@ -203,6 +204,7 @@ static const struct fs_parameter_spec ceph_mount_parame= ters[] =3D { fsparam_flag_no ("fsc", Opt_fscache), // fsc|nofsc fsparam_string ("fsc", Opt_fscache), // fsc=3D... fsparam_flag_no ("ino32", Opt_ino32), + fsparam_flag_no ("lazyio", Opt_lazyio), fsparam_string ("mds_namespace", Opt_mds_namespace), fsparam_string ("mon_addr", Opt_mon_addr), fsparam_flag_no ("poolperm", Opt_poolperm), @@ -593,6 +595,12 @@ static int ceph_parse_mount_param(struct fs_context *f= c, else fsopt->flags |=3D CEPH_MOUNT_OPT_SPARSEREAD; break; + case Opt_lazyio: + if (result.negated) + fsopt->flags &=3D ~CEPH_MOUNT_OPT_LAZYIO; + else + fsopt->flags |=3D CEPH_MOUNT_OPT_LAZYIO; + break; case Opt_test_dummy_encryption: #ifdef CONFIG_FS_ENCRYPTION fscrypt_free_dummy_policy(&fsopt->dummy_enc_policy); @@ -749,6 +757,8 @@ static int ceph_show_options(struct seq_file *m, struct= dentry *root) seq_puts(m, ",nopagecache"); if (fsopt->flags & CEPH_MOUNT_OPT_SPARSEREAD) seq_puts(m, ",sparseread"); + if (fsopt->flags & CEPH_MOUNT_OPT_LAZYIO) + seq_puts(m, ",lazyio"); =20 fscrypt_show_test_dummy_encryption(m, ',', root->d_sb); =20 @@ -1410,6 +1420,11 @@ static int ceph_reconfigure_fc(struct fs_context *fc) else ceph_clear_mount_opt(fsc, SPARSEREAD); =20 + if (fsopt->flags & CEPH_MOUNT_OPT_LAZYIO) + ceph_set_mount_opt(fsc, LAZYIO); + else + ceph_clear_mount_opt(fsc, LAZYIO); + if (strcmp_null(fsc->mount_options->mon_addr, fsopt->mon_addr)) { kfree(fsc->mount_options->mon_addr); fsc->mount_options->mon_addr =3D fsopt->mon_addr; diff --git a/fs/ceph/super.h b/fs/ceph/super.h index afc89ce91804..aec2eb4d0256 100644 --- a/fs/ceph/super.h +++ b/fs/ceph/super.h @@ -45,6 +45,7 @@ #define CEPH_MOUNT_OPT_ASYNC_DIROPS (1<<15) /* allow async directory op= s */ #define CEPH_MOUNT_OPT_NOPAGECACHE (1<<16) /* bypass pagecache altoget= her */ #define CEPH_MOUNT_OPT_SPARSEREAD (1<<17) /* always do sparse reads */ +#define CEPH_MOUNT_OPT_LAZYIO (1<<18) /* force lazyio for all fil= e opens */ =20 #define CEPH_MOUNT_OPT_DEFAULT \ (CEPH_MOUNT_OPT_DCACHE | \ diff --git a/fs/ceph/util.c b/fs/ceph/util.c index 2c34875675bf..be3db3f20344 100644 --- a/fs/ceph/util.c +++ b/fs/ceph/util.c @@ -73,10 +73,6 @@ int ceph_flags_to_mode(int flags) mode =3D CEPH_FILE_MODE_RDWR; break; } -#ifdef O_LAZY - if (flags & O_LAZY) - mode |=3D CEPH_FILE_MODE_LAZY; -#endif =20 return mode; } diff --git a/include/linux/ceph/ceph_fs.h b/include/linux/ceph/ceph_fs.h index 69ac3e55a3fe..01fd5f6647c8 100644 --- a/include/linux/ceph/ceph_fs.h +++ b/include/linux/ceph/ceph_fs.h @@ -414,6 +414,7 @@ extern const char *ceph_mds_op_name(int op); #define CEPH_O_CREAT 00000100 #define CEPH_O_EXCL 00000200 #define CEPH_O_TRUNC 00001000 +#define CEPH_O_LAZY 00020000 #define CEPH_O_DIRECTORY 00200000 #define CEPH_O_NOFOLLOW 00400000 =20 --- base-commit: 9fc75b71fdd38465c76c6f6a884cdd4ae3c72d90 change-id: 20260625-lazyio-6987f73c557f Best regards, -- =20 Xiubo Li