fs/remap_range.c | 25 +++++++++++++++++++++---- include/uapi/linux/fs.h | 5 ++++- 2 files changed, 25 insertions(+), 5 deletions(-)
On success FIDEDUPERANGE reports the requested length in bytes_deduped
even when the filesystem shortens a destination range and deduplicates
fewer bytes. This predates the VFS hoisting of the ioctl (the btrfs
ioctl behaved the same way), and changing the default would change an
ABI that deployed consumers such as duperemove depend on: they advance
their offsets by bytes_deduped and expect the historical semantics.
Add a flag to opt into the truthful behaviour. With
FILE_DEDUPE_RANGE_REPORT_PROGRESS set, bytes_deduped in each
destination's info is an advance hint for the next call on that
destination:
- if status is an error, bytes_deduped is 0;
- if status is FILE_DEDUPE_RANGE_DIFFERS, bytes_deduped is a safe
advance step: one filesystem block, capped to the requested length
so a sub-block request ending at EOF cannot be advanced past its
end, and falling back to the requested length when the top-level
inode reports a degenerate block size (stacked filesystems);
- if status is FILE_DEDUPE_RANGE_SAME, bytes_deduped is the number of
bytes actually deduplicated;
- in both success cases a zero value means no further work is
possible.
Unknown flag bits are rejected. The flags field shares an anonymous
union with the old reserved2 name, so existing source keeps compiling
and the binary layout is unchanged; old kernels require the field to
be zero, so new callers setting the flag on old kernels get -EINVAL
rather than silently the old semantics.
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Link: https://lore.kernel.org/linux-fsdevel/20260805071414.3414870-1-matthias.goergens@gmail.com/
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
fs/remap_range.c | 25 +++++++++++++++++++++----
include/uapi/linux/fs.h | 5 ++++-
2 files changed, 25 insertions(+), 5 deletions(-)
diff --git a/fs/remap_range.c b/fs/remap_range.c
index 26afbbbfb10c2..ac88a81c12739 100644
--- a/fs/remap_range.c
+++ b/fs/remap_range.c
@@ -503,7 +503,7 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
if (!(file->f_mode & FMODE_READ))
return -EINVAL;
- if (same->reserved1 || same->reserved2)
+ if (same->reserved1 || (same->flags & ~FILE_DEDUPE_RANGE_REPORT_PROGRESS))
return -EINVAL;
off = same->src_offset;
@@ -551,12 +551,29 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
deduped = vfs_dedupe_file_range_one(file, off, fd_file(dst_fd),
info->dest_offset, len,
REMAP_FILE_CAN_SHORTEN);
- if (deduped == -EBADE)
+ if (deduped == -EBADE) {
info->status = FILE_DEDUPE_RANGE_DIFFERS;
- else if (deduped < 0)
+ if (same->flags & FILE_DEDUPE_RANGE_REPORT_PROGRESS) {
+ u64 step = i_blocksize(src);
+
+ /*
+ * Stacked filesystems (e.g. overlayfs) can
+ * report a degenerate block size here; a
+ * one-byte step would only misalign the
+ * next call, so advance past the whole
+ * request instead.
+ */
+ if (step <= 1)
+ step = len;
+ info->bytes_deduped = min_t(u64, step, len);
+ }
+ } else if (deduped < 0) {
info->status = deduped;
- else
+ } else if (same->flags & FILE_DEDUPE_RANGE_REPORT_PROGRESS) {
+ info->bytes_deduped = deduped;
+ } else {
info->bytes_deduped = len;
+ }
next_loop:
if (fatal_signal_pending(current))
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index bd87262f2e349..471f698beaa93 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -178,13 +178,16 @@ struct file_dedupe_range_info {
__u32 reserved; /* must be zero */
};
+/* flags for struct file_dedupe_range */
+#define FILE_DEDUPE_RANGE_REPORT_PROGRESS (1U << 0)
+
/* from struct btrfs_ioctl_file_extent_same_args */
struct file_dedupe_range {
__u64 src_offset; /* in - start of extent in source */
__u64 src_length; /* in - length of extent */
__u16 dest_count; /* in - total elements in info array */
__u16 reserved1; /* must be zero */
- __u32 reserved2; /* must be zero */
+ __u32 flags; /* in - FILE_DEDUPE_RANGE_* flags */
struct file_dedupe_range_info info[];
};
--
2.55.0
On Mon, Aug 17, 2026 at 07:56:09PM +0800, Matthias Goergens wrote: > On success FIDEDUPERANGE reports the requested length in bytes_deduped > even when the filesystem shortens a destination range and deduplicates > fewer bytes. This predates the VFS hoisting of the ioctl (the btrfs > ioctl behaved the same way), and changing the default would change an > ABI that deployed consumers such as duperemove depend on: they advance > their offsets by bytes_deduped and expect the historical semantics. > > Add a flag to opt into the truthful behaviour. With What is "truthful"? It just is different. And you completely fail to explain why it is useful here, instead spewing a weird AI-like monologue just duplicating the patch content. Start with why you care, i.e. what application or type of application cares how much actually was deduplicated, and how you define the user visible behavior of that having happened.
Christoph, fair enough, and sorry for the delay. v4 follows with the changelog rewritten to start from the callers: dedupe tools advance their file offsets by bytes_deduped, the kernel can shorten a request to a block boundary but reports the requested length, and rmlint on an ordinary invocation therefore reports a pair as fully deduplicated with the last 1696 bytes of it not shared. Checking the v4 text against the code turned up two changes of substance. The DIFFERS advance hint from v3 is gone. The comparison covers the whole shortened range and stops at the first mismatch anywhere in it, so on DIFFERS the kernel has no mismatch offset to report; a one-block hint would tell a caller to skip a block that may match. v4 reports 0 on DIFFERS and leaves subdividing to the caller, which is what rmlint already does. Darrick, on your sketch specifically: that is why I dropped it; if you still want a hint there, I would rather it be the compared length than one block. Patch 1 is new: dax_dedupe_file_range_compare() returns the positive iomap_iter() count instead of the comparison error, and XFS passes that up as the remap result. Today the ioctl masks it by reporting the requested length; with the flag it would surface as a successful one-byte dedupe, so it needs fixing first. Also corrected from v3: its text described the flags field as a union with the old reserved2 name, but the diff was a plain rename. v4 has the union; both spellings compile and the struct size is unchanged. And the changelog now attributes the shortening to generic_remap_checks(), which runs before the comparison; v3 named generic_remap_check_len(), which runs after it. Two things I looked at and left alone, so that they are on record: ocfs2 returns 0 after copying inline data for a whole-file request, so under the flag that case reports SAME with 0 (FICLONERANGE already gets -EINVAL there today); and kernels before 4.5 ignored the reserved field in the btrfs ioctl, so only 4.5 and later reject the flag. The fstests test for the flag (generic/806 v2 on the fstests list) was written for the v3 semantics and needs a v3 of its own; that and the ioctl_fideduperange(2) man-page update follow once this settles. Changes since v3 (https://lore.kernel.org/linux-fsdevel/20260817115609.3586664-2-matthias.goergens@gmail.com/): - Changelog rewritten to start from the callers and the measured rmlint case (Christoph). - DIFFERS no longer carries an advance hint; bytes_deduped is 0 there. - New patch 1 fixing the DAX comparator's return value. - The flags field is an anonymous union with reserved2, as the v3 text already claimed; v3's diff was a plain rename. - Shortening attributed to generic_remap_checks(); v3 named generic_remap_check_len(), which runs after the comparison. Matthias Goergens (2): dax: return the comparison error from dax_dedupe_file_range_compare() vfs: add FILE_DEDUPE_RANGE_REPORT_PROGRESS flag to FIDEDUPERANGE fs/dax.c | 2 +- fs/remap_range.c | 4 +++- include/uapi/linux/fs.h | 8 +++++++- 3 files changed, 11 insertions(+), 3 deletions(-) -- 2.55.0
dax_dedupe_file_range_compare() returns ret, the positive result of the
last iomap_iter() call, when dax_range_compare_iter() fails. The caller
treats any non-zero return as the result of the range preparation, and
xfs_file_remap_range() returns it as the remap result, so a failed
comparison on a DAX file reports success with a small positive length
instead of the error.
Return the error itself.
Fixes: 0e79e3736d54 ("fsdax: dedupe: iter two files at the same time")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
fs/dax.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/fs/dax.c b/fs/dax.c
index 1fbba0d21c13d..de11bbbb6a384 100644
--- a/fs/dax.c
+++ b/fs/dax.c
@@ -2264,7 +2264,7 @@ int dax_dedupe_file_range_compare(struct inode *src, loff_t srcoff,
status = dax_range_compare_iter(&src_iter, &dst_iter,
min(src_iter.len, dst_iter.len), same);
if (status < 0)
- return ret;
+ return status;
src_iter.status = dst_iter.status = status;
}
return ret;
--
2.55.0
Deduplication tools such as duperemove, bees and rmlint find matching
ranges in two files, call FIDEDUPERANGE on each match and advance their
file offsets by the bytes_deduped the kernel returns. They rely on that
value to know where to continue.
The kernel does not give them a value they can act on.
vfs_dedupe_file_range() passes REMAP_FILE_CAN_SHORTEN, so
generic_remap_checks() rounds a request whose length is not block
aligned down to a block multiple unless it ends at both files' EOF, but
the ioctl then reports the length it asked for (the request, capped at 1
GiB per call) in bytes_deduped, not the shortened one. The caller cannot
tell that the tail of its request was left alone. Measured with rmlint
2.10.3 on btrfs with 4 KiB blocks: rmlint --dedupe on a 100000-byte file
against a 250000-byte file with the same prefix issues one call and is
told bytes_deduped=100000 with status SAME, while FIEMAP shows 24 shared
blocks, 98304 bytes. rmlint's loop ends because bytes_deduped equals the
file size, so it reports the pair fully deduplicated with 1696 bytes not
shared. duperemove (process_dedupes()) and bees advance the same way,
and jdupes advances by its own requested length without reading the
field, so all of them skip such a tail without noticing.
Add a flag that a caller sets to get a value it can act on. With
FILE_DEDUPE_RANGE_REPORT_PROGRESS set in file_dedupe_range.flags,
bytes_deduped in each destination's info is the length the filesystem
reports as deduplicated when status is FILE_DEDUPE_RANGE_SAME, and 0
when status is FILE_DEDUPE_RANGE_DIFFERS or an error. A caller advances
by it as it advances today, and must treat 0 as "stop or subdivide"
rather than retry unchanged. One cause of a SAME result of 0 is a
request shorter than a block that does not end at both files' EOF,
which the generic range preparation shortens to nothing before any
remapping. On DIFFERS the kernel has no usable progress or mismatch
offset to report, so it reports 0 and leaves subdividing the range to
the caller, as rmlint already does.
The default cannot change. Reporting the shortened length by default
was done once, in commit 4a57a8400075 ("vf/remap: return the amount of
bytes actually deduplicated"), and reverted the next day because
generic/517 expected the old value and the effect on deployed callers
was unknown. That effect is now known: duperemove re-queues a request
while its status is 0 and has no check for bytes_deduped == 0, so a 0
with status SAME on a sub-block request would make it re-issue the
same request forever.
Without the flag nothing changes. Unknown flag bits are rejected. The
flags field is an anonymous union with the old reserved2 name, so
existing source that spells .reserved2 keeps compiling and the layout
is unchanged. Kernels since 4.5, when the VFS took over the ioctl,
reject a non-zero field with -EINVAL, so a new caller cannot get the
old semantics by accident and can fall back to a call without the
flag.
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Link: https://lore.kernel.org/linux-fsdevel/20260805071414.3414870-1-matthias.goergens@gmail.com/
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
See the cover letter for the changes since v3. Note for C++ callers: a
positional initialiser of struct file_dedupe_range now needs braces
around the union member under -Wmissing-braces; designated initialisers
with either .reserved2 or .flags are unaffected.
fs/remap_range.c | 4 +++-
include/uapi/linux/fs.h | 8 +++++++-
2 files changed, 10 insertions(+), 2 deletions(-)
diff --git a/fs/remap_range.c b/fs/remap_range.c
index 26afbbbfb10c2..63f1b6f90c161 100644
--- a/fs/remap_range.c
+++ b/fs/remap_range.c
@@ -503,7 +503,7 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
if (!(file->f_mode & FMODE_READ))
return -EINVAL;
- if (same->reserved1 || same->reserved2)
+ if (same->reserved1 || (same->flags & ~FILE_DEDUPE_RANGE_REPORT_PROGRESS))
return -EINVAL;
off = same->src_offset;
@@ -555,6 +555,8 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
info->status = FILE_DEDUPE_RANGE_DIFFERS;
else if (deduped < 0)
info->status = deduped;
+ else if (same->flags & FILE_DEDUPE_RANGE_REPORT_PROGRESS)
+ info->bytes_deduped = deduped;
else
info->bytes_deduped = len;
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a5..4e40855e1ef8e 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -178,13 +178,19 @@ struct file_dedupe_range_info {
__u32 reserved; /* must be zero */
};
+/* flags for struct file_dedupe_range */
+#define FILE_DEDUPE_RANGE_REPORT_PROGRESS (1U << 0)
+
/* from struct btrfs_ioctl_file_extent_same_args */
struct file_dedupe_range {
__u64 src_offset; /* in - start of extent in source */
__u64 src_length; /* in - length of extent */
__u16 dest_count; /* in - total elements in info array */
__u16 reserved1; /* must be zero */
- __u32 reserved2; /* must be zero */
+ union {
+ __u32 reserved2; /* must be zero (older callers) */
+ __u32 flags; /* in - FILE_DEDUPE_RANGE_* flags */
+ };
struct file_dedupe_range_info info[];
};
--
2.55.0
On Fri, Sep 18, 2026 at 12:37 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> Deduplication tools such as duperemove, bees and rmlint find matching
> ranges in two files, call FIDEDUPERANGE on each match and advance their
> file offsets by the bytes_deduped the kernel returns. They rely on that
> value to know where to continue.
>
> The kernel does not give them a value they can act on.
> vfs_dedupe_file_range() passes REMAP_FILE_CAN_SHORTEN, so
> generic_remap_checks() rounds a request whose length is not block
> aligned down to a block multiple unless it ends at both files' EOF, but
> the ioctl then reports the length it asked for (the request, capped at 1
> GiB per call) in bytes_deduped, not the shortened one. The caller cannot
> tell that the tail of its request was left alone. Measured with rmlint
> 2.10.3 on btrfs with 4 KiB blocks: rmlint --dedupe on a 100000-byte file
> against a 250000-byte file with the same prefix issues one call and is
> told bytes_deduped=100000 with status SAME, while FIEMAP shows 24 shared
> blocks, 98304 bytes. rmlint's loop ends because bytes_deduped equals the
> file size, so it reports the pair fully deduplicated with 1696 bytes not
> shared. duperemove (process_dedupes()) and bees advance the same way,
> and jdupes advances by its own requested length without reading the
> field, so all of them skip such a tail without noticing.
>
> Add a flag that a caller sets to get a value it can act on. With
> FILE_DEDUPE_RANGE_REPORT_PROGRESS set in file_dedupe_range.flags,
> bytes_deduped in each destination's info is the length the filesystem
> reports as deduplicated when status is FILE_DEDUPE_RANGE_SAME, and 0
> when status is FILE_DEDUPE_RANGE_DIFFERS or an error. A caller advances
> by it as it advances today, and must treat 0 as "stop or subdivide"
> rather than retry unchanged. One cause of a SAME result of 0 is a
> request shorter than a block that does not end at both files' EOF,
> which the generic range preparation shortens to nothing before any
> remapping. On DIFFERS the kernel has no usable progress or mismatch
> offset to report, so it reports 0 and leaves subdividing the range to
> the caller, as rmlint already does.
>
> The default cannot change. Reporting the shortened length by default
> was done once, in commit 4a57a8400075 ("vf/remap: return the amount of
> bytes actually deduplicated"), and reverted the next day because
> generic/517 expected the old value and the effect on deployed callers
> was unknown. That effect is now known: duperemove re-queues a request
> while its status is 0 and has no check for bytes_deduped == 0, so a 0
> with status SAME on a sub-block request would make it re-issue the
> same request forever.
>
> Without the flag nothing changes. Unknown flag bits are rejected. The
> flags field is an anonymous union with the old reserved2 name, so
> existing source that spells .reserved2 keeps compiling and the layout
I don't why you think that old programs compile should not break with
new headers. I don't think this is a promise from Linux UAPI.
> is unchanged. Kernels since 4.5, when the VFS took over the ioctl,
> reject a non-zero field with -EINVAL, so a new caller cannot get the
> old semantics by accident and can fall back to a call without the
> flag.
>
> Suggested-by: Darrick J. Wong <djwong@kernel.org>
> Link: https://lore.kernel.org/linux-fsdevel/20260805071414.3414870-1-matthias.goergens@gmail.com/
> Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
> ---
> See the cover letter for the changes since v3. Note for C++ callers: a
> positional initialiser of struct file_dedupe_range now needs braces
> around the union member under -Wmissing-braces; designated initialisers
> with either .reserved2 or .flags are unaffected.
I think you created a problem yourself and then solved it.
The union is not needed.
>
> fs/remap_range.c | 4 +++-
> include/uapi/linux/fs.h | 8 +++++++-
> 2 files changed, 10 insertions(+), 2 deletions(-)
>
> diff --git a/fs/remap_range.c b/fs/remap_range.c
> index 26afbbbfb10c2..63f1b6f90c161 100644
> --- a/fs/remap_range.c
> +++ b/fs/remap_range.c
> @@ -503,7 +503,7 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
> if (!(file->f_mode & FMODE_READ))
> return -EINVAL;
>
> - if (same->reserved1 || same->reserved2)
> + if (same->reserved1 || (same->flags & ~FILE_DEDUPE_RANGE_REPORT_PROGRESS))
> return -EINVAL;
>
> off = same->src_offset;
> @@ -555,6 +555,8 @@ int vfs_dedupe_file_range(struct file *file, struct file_dedupe_range *same)
> info->status = FILE_DEDUPE_RANGE_DIFFERS;
> else if (deduped < 0)
> info->status = deduped;
> + else if (same->flags & FILE_DEDUPE_RANGE_REPORT_PROGRESS)
> + info->bytes_deduped = deduped;
> else
> info->bytes_deduped = len;
>
> diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
> index 34c6f219462a5..4e40855e1ef8e 100644
> --- a/include/uapi/linux/fs.h
> +++ b/include/uapi/linux/fs.h
> @@ -178,13 +178,19 @@ struct file_dedupe_range_info {
> __u32 reserved; /* must be zero */
> };
>
> +/* flags for struct file_dedupe_range */
> +#define FILE_DEDUPE_RANGE_REPORT_PROGRESS (1U << 0)
> +
This thought has crossed my mind - perhaps should be dismissed.
These are vfs internal flags to the filesystem dedupe/remap method:
#define REMAP_FILE_DEDUP (1 << 0)
#define REMAP_FILE_CAN_SHORTEN (1 << 1)
The above is a dedupe ioctl UAPI flag, which is not needed to be passed
to the filesystem (right?), so it's fine that they live in two
different namespaces.
In the future, it could be that a dedupe UAPI flag will need to be propagated
into the filesystem. Maybe even a flag that would be common to clone and
dedupe.
For now, using bit 1 for both internal and uapi is fine and I don't see
a reason to change that, just a point to consider for the future.
> /* from struct btrfs_ioctl_file_extent_same_args */
> struct file_dedupe_range {
> __u64 src_offset; /* in - start of extent in source */
> __u64 src_length; /* in - length of extent */
> __u16 dest_count; /* in - total elements in info array */
> __u16 reserved1; /* must be zero */
> - __u32 reserved2; /* must be zero */
> + union {
> + __u32 reserved2; /* must be zero (older callers) */
> + __u32 flags; /* in - FILE_DEDUPE_RANGE_* flags */
> + };
> struct file_dedupe_range_info info[];
> };
Please drop the union.
The text referring to "older callers" in this context is not useful.
It is quite obvious that the old kernel will not accept new flags in a
properly written UAPI.
Thanks,
Amir.
© 2016 - 2026 Red Hat, Inc.