From nobody Sat Jul 25 02:35:07 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B3E4412274; Mon, 20 Jul 2026 12:39:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784551179; cv=none; b=p8RlYNa9i5p4r7grEVCKYv9leaJsWwt9jYwRm+azY0en0ehjSVJBa1K0GrXTKQGcA2WZSYh+WcNIXG2xqHcugU4Syil73Ahdb9TrbrnMEOz/DNoYQqZhMF4Lec7gb4F5i2eagubPay5+c7/wSeQ4PrZwFEAncdlYMaoOt1OLR0w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784551179; c=relaxed/simple; bh=TsD4vh+Uvz3r3g/6HzQVavPJYOiyZvbfqnIBytRZoWk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:To:Cc; b=j0sYsuxZ9cDfLDJQ2zEAP0OAhM9X6bKbFUYh+yiXmUlS4mZgxbW6GOBCRXUY2B/us7Yvg3KOBop30uS9v9LmRT8aLZK0orEdAAiPO8AgfRWtZIi3ccwgBa21u2aRTWA2MkEmSS3qilP99NnwIYi8uI4z3tVZ273NISmYiIN2j8o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=asg/KliM; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="asg/KliM" Received: by smtp.kernel.org (Postfix) with ESMTPS id C9BC4C2BCB8; Mon, 20 Jul 2026 12:39:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784551178; bh=TsD4vh+Uvz3r3g/6HzQVavPJYOiyZvbfqnIBytRZoWk=; h=From:Date:Subject:To:Cc:Reply-To:From; b=asg/KliMhJssvBffz7SWYVPpMgPzMOzkMaagAi2cLg1QhuJw0sy8woUPN7DBmfy4I rNirHwijAkiph4WKhELmJJYgHiuoXu+Uz3p78Z2FvcLvKbq+kf8lrHjfia8nbufRU7 SiTsLNQHskTvcAXZgYBq7/MAKEt4Ujqhy28hNYK6kg5rO8s3w5xC9IV7wjNJyZu+Mi UT6EGhozNT9OOYwEuDQJXVkfaAsaAOr1X8CCxFpZT11bo5q6bHKr7b9t/M/dP5gt70 LPfCCPtP6bo6dykCHpM7UzLI5fh8af+dGyYHohHheqXXFFtJsJmtAi41GrGoku9RcB s20BY2trSQGRA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id B6D33C4452A; Mon, 20 Jul 2026 12:39:38 +0000 (UTC) From: Xiubo Li via B4 Relay Date: Mon, 20 Jul 2026 20:39:24 +0800 Subject: [PATCH] ceph: parallelize object copy in ceph_do_objects_copy() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260720-b4-ceph-copyfrom-v1-1-9a4c229df6e2@clyso.com> X-B4-Tracking: v=1; b=H4sIAAAAAAAC/yXMQQqDMBBA0avIrB2IoUzFq4gLM451CpqQaGkR7 260y7f4f4ckUSVBU+wQ5aNJ/ZJRlQXw1C8vQR2ywRpL5mkNugeyhAnZh98Y/YxcU22cHYgqgpy FKKN+72Xb/Z029xZerw8cxwlGjrXYdAAAAA== X-Change-ID: 20260720-b4-ceph-copyfrom-c8680b2d6616 To: Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko Cc: ceph-devel@vger.kernel.org, linux-kernel@vger.kernel.org, Xiubo Li X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784551170; l=9188; i=xiubo.li@clyso.com; s=20260625; h=from:subject:message-id; bh=tBKXfCJ+NgJ+kSb6ibcSYY9ADa5eQQQSOPM+cAzJ68Y=; b=lLJhX844uCwzYJMues2rQhaN+mW0jEx8/TKpFwSdDCXZh6Dtm6LVBoohNC+AMIeMwEhTemzMb mdBHAOXmAB2D92Tdny88+GmPKW+v18diP2t9lx5QvPdNICTq+B6v/nK X-Developer-Key: i=xiubo.li@clyso.com; a=ed25519; pk=V3NGr0AgAopiUhaLY51ipBkLN5LlcLhjOEfLEq1RoZ8= X-Endpoint-Received: by B4 Relay for xiubo.li@clyso.com/20260625 with auth_id=840 X-Original-From: Xiubo Li Reply-To: xiubo.li@clyso.com From: Xiubo Li The current ceph_do_objects_copy() submits COPY_FROM2 requests serially (submit-wait-submit-wait), which means the total latency scales linearly with the number of objects being copied. Convert to a sliding-window parallel model: submit up to CEPH_COPY_FROM_MAX_INFLIGHT (16) requests concurrently, then wait for them in FIFO order. Because all in-flight requests make progress while we wait for the oldest one, the total wait time is MAX(latency_i) rather than SUM(latency_i). Error handling uses a 'truncate at first failure' strategy: all submitted requests are allowed to complete, results are scanned in offset order, and the copy is truncated at the first failing object. Objects beyond the truncation point that were successfully written by the OSD become orphaned beyond EOF, which is harmless =E2=80=94 they are unreachable through reads and are cleaned up on file deletion. Signed-off-by: Xiubo Li --- fs/ceph/file.c | 193 +++++++++++++++++++++++++++++++++++++++++++++--------= ---- 1 file changed, 152 insertions(+), 41 deletions(-) diff --git a/fs/ceph/file.c b/fs/ceph/file.c index 9d89d7fc1095..32147c719c35 100644 --- a/fs/ceph/file.c +++ b/fs/ceph/file.c @@ -2965,6 +2965,20 @@ ceph_alloc_copyfrom_request(struct ceph_osd_client *= osdc, return ERR_PTR(ret); } =20 +/* + * Default maximum number of in-flight COPY_FROM2 requests. Can be + * overridden at module load time via the copyfrom_max_inflight parameter. + * + * Higher values improve throughput over high-latency links, but too many + * concurrent requests can saturate OSD disk queues, especially in small + * clusters. Tune this to match the number of OSDs and their concurrency + * capability. For most clusters 16--64 is a reasonable range. + */ +static unsigned int copyfrom_max_inflight =3D 16; +module_param(copyfrom_max_inflight, uint, 0644); +MODULE_PARM_DESC(copyfrom_max_inflight, + "Maximum in-flight COPY_FROM2 requests per copy_file_range call"); + static ssize_t ceph_do_objects_copy(struct ceph_inode_info *src_ci, u64 *s= rc_off, struct ceph_inode_info *dst_ci, u64 *dst_off, struct ceph_fs_client *fsc, @@ -2973,13 +2987,19 @@ static ssize_t ceph_do_objects_copy(struct ceph_ino= de_info *src_ci, u64 *src_off struct ceph_object_locator src_oloc, dst_oloc; struct ceph_object_id src_oid, dst_oid; struct ceph_osd_client *osdc; + struct ceph_osd_request **reqs =3D NULL; struct ceph_osd_request *req; ssize_t bytes =3D 0; u64 src_objnum, src_objoff, dst_objnum, dst_objoff; u32 src_objlen, dst_objlen; u32 object_size =3D src_ci->i_layout.object_size; struct ceph_client *cl =3D fsc->client; - int ret; + u64 orig_src_off =3D *src_off; + u64 orig_dst_off =3D *dst_off; + int num_objects, head =3D 0, tail =3D 0, inflight =3D 0; + int first_fail_obj =3D -1, first_fail_err =3D 0; + int slot, ret; + bool have_eopnotsupp =3D false; =20 src_oloc.pool =3D src_ci->i_layout.pool_id; src_oloc.pool_ns =3D ceph_try_get_string(src_ci->i_layout.pool_ns); @@ -2987,54 +3007,145 @@ static ssize_t ceph_do_objects_copy(struct ceph_in= ode_info *src_ci, u64 *src_off dst_oloc.pool_ns =3D ceph_try_get_string(dst_ci->i_layout.pool_ns); osdc =3D &fsc->client->osdc; =20 - while (len >=3D object_size) { - ceph_calc_file_object_mapping(&src_ci->i_layout, *src_off, - object_size, &src_objnum, - &src_objoff, &src_objlen); - ceph_calc_file_object_mapping(&dst_ci->i_layout, *dst_off, - object_size, &dst_objnum, - &dst_objoff, &dst_objlen); - ceph_oid_init(&src_oid); - ceph_oid_printf(&src_oid, "%llx.%08llx", - src_ci->i_vino.ino, src_objnum); - ceph_oid_init(&dst_oid); - ceph_oid_printf(&dst_oid, "%llx.%08llx", - dst_ci->i_vino.ino, dst_objnum); - /* Do an object remote copy */ - req =3D ceph_alloc_copyfrom_request(osdc, src_ci->i_vino.snap, - &src_oid, &src_oloc, - &dst_oid, &dst_oloc, - dst_ci->i_truncate_seq, - dst_ci->i_truncate_size); - if (IS_ERR(req)) - ret =3D PTR_ERR(req); - else { + num_objects =3D len / object_size; + if (!num_objects) + goto out; + + reqs =3D kvmalloc_array(copyfrom_max_inflight, sizeof(*reqs), GFP_KERNEL); + if (!reqs) { + bytes =3D -ENOMEM; + goto out; + } + + /* + * Sliding window: submit requests up to MAX_INFLIGHT, then wait for + * the oldest in-flight request to complete before submitting more. + * The reqs[] ring buffer is indexed by (slot % MAX_INFLIGHT), so + * memory is bounded to MAX_INFLIGHT regardless of num_objects. + * First-failure and EOPNOTSUPP detection happen inline =E2=80=94 no post= -scan + * needed. + */ + while (head < num_objects) { + /* Submit new requests while the window has room */ + while (tail < num_objects && + first_fail_obj =3D=3D -1 && + inflight < copyfrom_max_inflight) { + u64 object_src_off =3D orig_src_off + + (u64)tail * object_size; + u64 object_dst_off =3D orig_dst_off + + (u64)tail * object_size; + + ceph_calc_file_object_mapping(&src_ci->i_layout, + object_src_off, + object_size, + &src_objnum, + &src_objoff, + &src_objlen); + ceph_calc_file_object_mapping(&dst_ci->i_layout, + object_dst_off, + object_size, + &dst_objnum, + &dst_objoff, + &dst_objlen); + ceph_oid_init(&src_oid); + ceph_oid_printf(&src_oid, "%llx.%08llx", + src_ci->i_vino.ino, src_objnum); + ceph_oid_init(&dst_oid); + ceph_oid_printf(&dst_oid, "%llx.%08llx", + dst_ci->i_vino.ino, dst_objnum); + + /* Do an object remote copy */ + slot =3D tail % copyfrom_max_inflight; + req =3D ceph_alloc_copyfrom_request(osdc, + src_ci->i_vino.snap, + &src_oid, &src_oloc, + &dst_oid, &dst_oloc, + dst_ci->i_truncate_seq, + dst_ci->i_truncate_size); + if (IS_ERR(req)) { + /* + * Allocation failed: treat as first failure + * and stop submitting. Drain inflight + * requests to check for EOPNOTSUPP. + */ + if (first_fail_obj =3D=3D -1) { + first_fail_obj =3D tail; + first_fail_err =3D PTR_ERR(req); + } + reqs[slot] =3D NULL; + tail++; + continue; + } ceph_osdc_start_request(osdc, req); - ret =3D ceph_osdc_wait_request(osdc, req); + reqs[slot] =3D req; + tail++; + inflight++; + } + + /* + * Wait for the oldest in-flight request (FIFO order). + * This is required for correctness: we must determine the + * first failure in object-offset order to know how many + * bytes were successfully copied. It does not hurt + * performance because all requests in the window are + * submitted concurrently -- later requests complete in + * the background while we wait, so the wall-clock time + * is dominated by the slowest request, not the sum. + */ + slot =3D head % copyfrom_max_inflight; + if (reqs[slot]) { + ret =3D ceph_osdc_wait_request(osdc, reqs[slot]); ceph_update_copyfrom_metrics(&fsc->mdsc->metric, - req->r_start_latency, - req->r_end_latency, + reqs[slot]->r_start_latency, + reqs[slot]->r_end_latency, object_size, ret); - ceph_osdc_put_request(req); - } - if (ret) { - if (ret =3D=3D -EOPNOTSUPP) { - fsc->have_copy_from2 =3D false; - pr_notice_client(cl, - "OSDs don't support copy-from2; disabling copy offload\n"); + ceph_osdc_put_request(reqs[slot]); + inflight--; + + if (ret =3D=3D -EOPNOTSUPP) + have_eopnotsupp =3D true; + + if (ret < 0 && first_fail_obj =3D=3D -1) { + first_fail_obj =3D head; + first_fail_err =3D ret; } - doutc(cl, "returned %d\n", ret); - if (bytes <=3D 0) - bytes =3D ret; - goto out; + if (ret) + doutc(cl, "object %d returned %d\n", head, ret); } - len -=3D object_size; - bytes +=3D object_size; - *src_off +=3D object_size; - *dst_off +=3D object_size; + /* + * If reqs[slot] was NULL, allocation for this object had + * failed earlier; first_fail_obj was already set in the + * submit loop -- just drain past it. + */ + head++; } =20 + /* + * Determine bytes copied: all objects before the first failure + * succeeded. Any objects beyond the truncation point that were + * successfully written by the OSD become unreachable orphans + * (beyond EOF) and are harmless. + */ + if (first_fail_obj =3D=3D -1) { + bytes =3D (ssize_t)num_objects * object_size; + } else if (first_fail_obj =3D=3D 0) { + bytes =3D first_fail_err; + } else { + bytes =3D (ssize_t)first_fail_obj * object_size; + } + + if (have_eopnotsupp) { + fsc->have_copy_from2 =3D false; + pr_notice_client(cl, + "OSDs don't support copy-from2; disabling copy offload\n"); + } + + if (bytes > 0) { + *src_off =3D orig_src_off + bytes; + *dst_off =3D orig_dst_off + bytes; + } out: + kvfree(reqs); ceph_oloc_destroy(&src_oloc); ceph_oloc_destroy(&dst_oloc); return bytes; --- base-commit: fc67edb66b3c9924c4e0bb366a92b32ea13c526a change-id: 20260720-b4-ceph-copyfrom-c8680b2d6616 Best regards, -- =20 Xiubo Li