From nobody Fri Oct 2 12:19:51 2026 Received: from pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.246.68.102]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D06EA3368A4; Fri, 31 Jul 2026 20:39:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.246.68.102 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785530371; cv=none; b=TjilBLodkAJqVZNM9Pf3KFjFtyU0tJxfxc5jKV10aBR2Ro61nJhU0pzB2lvsIgYLrikpPNbqdbDifQiXqeRXW11YkEu8qMVrJV/8EjdB3gFkeyvZpMCNY98AvidIJ/rob7r+exCBad9S7qG72RMgdOFzev9wlz9+sbPJaJyDTrE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785530371; c=relaxed/simple; bh=JmyLqMQOCQhrSgHHrKOoWCVisbxlpu7jpYOhnApJxLU=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=OqZVfHOm4KFrMRYyj0+kY/5pcc5SItTi2Qc6mPsKpnac31bQoQ42/S0oB25x81f5oza8B12XjvcHJODHTaNmYIw8q2lirCAL5enUCnEfIqC4FMSJI9jEL4Sz1dnS0VwRa+km51/6dqO1s4jyn4tb/DN/jzr22ie7u9Zb/e8vZm4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=tdCapOgM; arc=none smtp.client-ip=44.246.68.102 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="tdCapOgM" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1785530369; x=1817066369; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=1XxHHgDxFaSg+PKoRdxaXxYeDVMEsgZ+UBXj89p6PiM=; b=tdCapOgM5gu2/3Hwm+qwfDzfIsATlewc6u6XakXfkDD1ao5NWqIZs+dh kRjEZEMozrgrmEb1gGnb06VGhiiB5H8BPbD9N6TpADy1gGZuxHH/32HX/ nusGhC66LKFdMiUjcCtJc5E4iwIKoF+kyPk8Do3GvtikSapgFGpHxG4pz SMIqyJdS2j4Ys4hcL5Fkjh1oDJckGoQQlVCDrFwVRiL1SpADAViDpJxXu FlvfxfCsNe2KWoALLlwsOsm67gXanwpW0021JTV5+UOw8hvaxZtGbthaZ 7rE1hl6OjhqIaaaK2Pid3M+OvlHBja6PV6ab2fTGJISzvFjO00ffomd2s Q==; X-CSE-ConnectionGUID: yT/zqkz6TiqzE/kEwCpw/Q== X-CSE-MsgGUID: 8hB2ugaATHKcSp/4s5+Q0g== X-IronPort-AV: E=Sophos;i="6.25,197,1779148800"; d="scan'208";a="24812508" Received: from ip-10-5-0-115.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.0.115]) by internal-pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 31 Jul 2026 20:39:29 +0000 Received: from EX19MTAUWC001.ant.amazon.com [205.251.233.53:6666] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.53.230:2525] with esmtp (Farcaster) id dc5acb33-2963-492f-a241-07f8fe5799a1; Fri, 31 Jul 2026 20:39:29 +0000 (UTC) X-Farcaster-Flow-ID: dc5acb33-2963-492f-a241-07f8fe5799a1 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWC001.ant.amazon.com (10.250.64.174) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 31 Jul 2026 20:39:29 +0000 Received: from dev-dsk-jamz-1e-e35f4cd9.us-east-1.amazon.com (10.189.35.140) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 31 Jul 2026 20:39:28 +0000 From: Jimmy Zuber To: Miklos Szeredi , Shuah Khan CC: , , Subject: [PATCH 1/2] fuse: zero the partial EOF page when extending a file Date: Fri, 31 Jul 2026 20:38:41 +0000 Message-ID: <20260731203842.540798-2-jamz@amazon.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260731203842.540798-1-jamz@amazon.com> References: <20260731203842.540798-1-jamz@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D044UWA001.ant.amazon.com (10.13.139.100) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Extending a fuse file past a non-page-aligned EOF does not zero the tail of the old last page. If that page is cached and was dirtied beyond the old EOF -- e.g. an application mmap()ed the EOF page and stored into the region past EOF, which is undefined until the file grows -- the now in-bounds tail is exposed to subsequent reads as stale data instead of zeros, in violation of POSIX file-extension semantics. Other filesystems zero this via pagecache_isize_extended(), but that helper is a no-op for fuse: it returns early when i_blocksize() >=3D PAGE_SIZE, and a non-fuseblk fuse mount has s_blocksize =3D=3D PAGE_SIZE (the server-suppl= ied st_blksize only sets fi->cached_i_blkbits, not i_blkbits). The NFS client hit the same problem and open-codes the zeroing in nfs_truncate_last_folio(); add the equivalent fuse_zero_partial_eof_folio() and call it from the three paths that extend a file: a buffered write, a size-extending setattr/truncate, and a size-extending fallocate (fuse_write_update_attr(), fuse_do_setattr() and fuse_file_fallocate()). writeback_cache connections are unaffected, as their writes go through iomap_file_buffered_write(), which zeroes post-EOF folios. The bug is observable on a non-writeback_cache server that returns FOPEN_KEEP_CACHE on writable files (without FOPEN_DIRECT_IO), and is caught by the new write_extend_eof fuse selftest. Signed-off-by: Jimmy Zuber --- fs/fuse/dir.c | 3 +++ fs/fuse/file.c | 56 ++++++++++++++++++++++++++++++++++++++++++++++++ fs/fuse/fuse_i.h | 1 + 3 files changed, 60 insertions(+) diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c index 795e92037ce7..f6614ccef186 100644 --- a/fs/fuse/dir.c +++ b/fs/fuse/dir.c @@ -2282,6 +2282,9 @@ int fuse_do_setattr(struct mnt_idmap *idmap, struct d= entry *dentry, */ if ((is_truncate || !is_wb) && S_ISREG(inode->i_mode) && oldsize !=3D outarg.attr.size) { + if (outarg.attr.size > oldsize) + fuse_zero_partial_eof_folio(inode, oldsize, + outarg.attr.size); truncate_pagecache(inode, outarg.attr.size); invalidate_inode_pages2(mapping); } diff --git a/fs/fuse/file.c b/fs/fuse/file.c index cb8da4c06d17..a9063b4e9217 100644 --- a/fs/fuse/file.c +++ b/fs/fuse/file.c @@ -21,6 +21,8 @@ #include #include #include +#include +#include =20 static int fuse_send_open(struct fuse_mount *fm, u64 nodeid, unsigned int open_flags, int opcode, @@ -1200,20 +1202,64 @@ static ssize_t fuse_send_write(struct fuse_io_args = *ia, loff_t pos, return err ?: ia->write.out.size; } =20 +/* + * An operation extended i_size past a non-folio-aligned old EOF at @from, + * turning [@from, @to) into a hole that must read back as zero. If the o= ld + * last folio is cached and was dirtied beyond the old EOF (e.g. mmap stor= es + * into the post-EOF region, which are undefined until the file grows), ze= ro + * that tail so it is not exposed as stale data (xfstests generic/363). + * + * pagecache_isize_extended() cannot be used: it bails out for + * i_blocksize() >=3D PAGE_SIZE, and a non-fuseblk mount has + * s_blocksize =3D=3D PAGE_SIZE, so the zeroing has to be done here. + * Callers hold i_rwsem, serialising this against concurrent writes and + * truncates; it must not run under fi->lock, as it locks the folio. + */ +void fuse_zero_partial_eof_folio(struct inode *inode, loff_t from, loff_t = to) +{ + struct folio *folio; + size_t offset, end; + + if (from >=3D to) + return; + + folio =3D filemap_lock_folio(inode->i_mapping, from >> PAGE_SHIFT); + if (IS_ERR(folio)) + return; + + if (folio_mkclean(folio)) + folio_mark_dirty(folio); + + if (folio_test_dirty(folio)) { + offset =3D offset_in_folio(folio, from); + end =3D min_t(loff_t, to - folio_pos(folio), folio_size(folio)); + folio_zero_segment(folio, offset, end); + } + + folio_unlock(folio); + folio_put(folio); +} + bool fuse_write_update_attr(struct inode *inode, loff_t pos, ssize_t writt= en) { struct fuse_conn *fc =3D get_fuse_conn(inode); struct fuse_inode *fi =3D get_fuse_inode(inode); bool ret =3D false; + loff_t old_size =3D 0; =20 spin_lock(&fi->lock); fi->attr_version =3D atomic64_inc_return(&fc->attr_version); if (written > 0 && pos > inode->i_size) { + old_size =3D inode->i_size; i_size_write(inode, pos); ret =3D true; } spin_unlock(&fi->lock); =20 + /* [old_size, pos - written) is the hole this write opened past EOF. */ + if (ret) + fuse_zero_partial_eof_folio(inode, old_size, pos - written); + fuse_invalidate_attr_mask(inode, FUSE_STATX_MODSIZE); =20 return ret; @@ -2913,8 +2959,18 @@ static long fuse_file_fallocate(struct file *file, i= nt mode, loff_t offset, =20 /* we could have extended the file */ if (!(mode & FALLOC_FL_KEEP_SIZE)) { + loff_t oldsize =3D i_size_read(inode); + if (fuse_write_update_attr(inode, offset + length, length)) file_update_time(file); + /* + * fuse_write_update_attr() already zeroes up to @offset when + * the write started past the old EOF; this additionally covers + * a fallocate whose range starts at or before it. fallocate + * writes no data, so the whole extension must read as zero; the + * overlap is a no-op. + */ + fuse_zero_partial_eof_folio(inode, oldsize, offset + length); } =20 if (mode & (FALLOC_FL_PUNCH_HOLE | FALLOC_FL_ZERO_RANGE)) diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h index 85f738c53122..ee3b91b56fef 100644 --- a/fs/fuse/fuse_i.h +++ b/fs/fuse/fuse_i.h @@ -1183,6 +1183,7 @@ long fuse_ioctl_common(struct file *file, unsigned in= t cmd, __poll_t fuse_file_poll(struct file *file, poll_table *wait); =20 bool fuse_write_update_attr(struct inode *inode, loff_t pos, ssize_t writt= en); +void fuse_zero_partial_eof_folio(struct inode *inode, loff_t from, loff_t = to); =20 int fuse_flush_times(struct inode *inode, struct fuse_file *ff); int fuse_write_inode(struct inode *inode, struct writeback_control *wbc); --=20 2.50.1 From nobody Fri Oct 2 12:19:51 2026 Received: from pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.245.243.92]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 929FD37997A; Fri, 31 Jul 2026 20:39:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.245.243.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785530400; cv=none; b=YFhXC5S8+FZ3WgHoCzRU1nxHISUj++qncx0L3gh8DFSNUZPQ650zizAYjGMrqazQUkwrWV6EMzNBvplP2Z/WXyOLHnU6M0HNs3j2HFzUYHWjZgJvOBOTz6x2B5u4x9j2jz617yoLPdAYtNSxVpyrRXn1uqIlSsM69C7p1x9+rjc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785530400; c=relaxed/simple; bh=7jou5J1V5WhQ8JPJPMA+mGod2e8BORQN/EQKB8JEnpI=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=QtDJOId/wooaOsodHBaimwv1gBR3ppNCrbMYHUPPZnINjaD1YnD7bWV9IazPP2pkykvTJW/nIsA68DCR5vNtA5lAY8EVpIXk4QXC1Ek4WpSFX/9vWwfZiOTbc53/qlGyQxEM9v7CYKB5CMwCw4aQrG55A1lqZSjFCySpBcn0Y2M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=ENsGIcMh; arc=none smtp.client-ip=44.245.243.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="ENsGIcMh" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1785530398; x=1817066398; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=UY9++tBrj+ufFcnIOLtX0X/zdoGWUMqnOXF6hmcyc7k=; b=ENsGIcMh2HB2k0Lpzi5X+hPgwvNR0XjFUZItb0ZQ9KD9tywD/N9EWl9j XI5Qd9Lww7NhQna4TZp1Wu1SXBxOPrJeaBDjwodhtqljG1mpglFyMVtKv +nFTebDikVMBlv+3xF6/L8JL6DRsA4nOO9Wrv7gOJogWw7P7uPecY02ry N4GdLdokcwIzZKSqR29JUQUC49qCAjs8HQS6Hlob81idjlbs8oCj3/eGY Eq+bzLL2X/ImXyAlCTJGUQgj81wi0CVeQsBXHT95dR6fiHBcy56fHkI7+ qfFrPDOrOqD29/Cmg1QHylIsCmuSP3oZEo76IfIDkmBJpGsHQ8Y2/SSGB A==; X-CSE-ConnectionGUID: n4PvU6OJQ1SIQT0KC6nEKA== X-CSE-MsgGUID: LnErrtT4SJqbSQ1TtRnchQ== X-IronPort-AV: E=Sophos;i="6.25,197,1779148800"; d="scan'208";a="24292015" Received: from ip-10-5-9-48.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.9.48]) by internal-pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 31 Jul 2026 20:39:58 +0000 Received: from EX19MTAUWB002.ant.amazon.com [205.251.233.48:25417] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.53.248:2525] with esmtp (Farcaster) id ee6a3ae2-b6d0-4ebd-bfc2-cd2c8d8c2707; Fri, 31 Jul 2026 20:39:57 +0000 (UTC) X-Farcaster-Flow-ID: ee6a3ae2-b6d0-4ebd-bfc2-cd2c8d8c2707 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWB002.ant.amazon.com (10.250.64.231) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 31 Jul 2026 20:39:57 +0000 Received: from dev-dsk-jamz-1e-e35f4cd9.us-east-1.amazon.com (10.189.35.140) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 31 Jul 2026 20:39:56 +0000 From: Jimmy Zuber To: Miklos Szeredi , Shuah Khan CC: , , Subject: [PATCH 2/2] selftests/fuse: test post-EOF page zeroing when a file is extended Date: Fri, 31 Jul 2026 20:38:42 +0000 Message-ID: <20260731203842.540798-3-jamz@amazon.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260731203842.540798-1-jamz@amazon.com> References: <20260731203842.540798-1-jamz@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D044UWB003.ant.amazon.com (10.13.139.168) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Add a regression test for the bug where extending a file left the tail of the old partial EOF page exposing stale mmap-dirtied data instead of zeros. The test is a self-contained raw /dev/fuse server (no libfuse dependency) that runs without writeback_cache and returns FOPEN_KEEP_CACHE, the configuration in which the bug is visible. Its backing data is always zero in the hole, so any non-zero byte a read sees is stale page-cache data. All offsets are relative to the runtime page size. Four cases: - write_extend: pollute the post-EOF tail, extend past it by writing into a later page, and verify the tail reads back as zero; - ftruncate_extend: same, but extend via ftruncate(); - fallocate_extend: same, but extend via fallocate() at the old EOF; - extend_into_eof_page_preserves_data: an extending write landing inside the old EOF page must not be clobbered by the zeroing. Each case fails without the fix and passes with it. Signed-off-by: Jimmy Zuber --- .../selftests/filesystems/fuse/.gitignore | 1 + .../selftests/filesystems/fuse/Makefile | 3 + .../filesystems/fuse/write_extend_eof_test.c | 368 ++++++++++++++++++ 3 files changed, 372 insertions(+) create mode 100644 tools/testing/selftests/filesystems/fuse/write_extend_e= of_test.c diff --git a/tools/testing/selftests/filesystems/fuse/.gitignore b/tools/te= sting/selftests/filesystems/fuse/.gitignore index 3e72e742d08e..fb51603fe419 100644 --- a/tools/testing/selftests/filesystems/fuse/.gitignore +++ b/tools/testing/selftests/filesystems/fuse/.gitignore @@ -1,3 +1,4 @@ # SPDX-License-Identifier: GPL-2.0-only fuse_mnt fusectl_test +write_extend_eof_test diff --git a/tools/testing/selftests/filesystems/fuse/Makefile b/tools/test= ing/selftests/filesystems/fuse/Makefile index 612aad69a93a..0c2c613af0ac 100644 --- a/tools/testing/selftests/filesystems/fuse/Makefile +++ b/tools/testing/selftests/filesystems/fuse/Makefile @@ -3,10 +3,13 @@ CFLAGS +=3D -Wall -O2 -g $(KHDR_INCLUDES) =20 TEST_GEN_PROGS :=3D fusectl_test +TEST_GEN_PROGS +=3D write_extend_eof_test TEST_GEN_FILES :=3D fuse_mnt =20 include ../../lib.mk =20 +$(OUTPUT)/write_extend_eof_test: LDLIBS +=3D -lpthread + VAR_CFLAGS :=3D $(shell pkg-config fuse --cflags 2>/dev/null) ifeq ($(VAR_CFLAGS),) VAR_CFLAGS :=3D -D_FILE_OFFSET_BITS=3D64 -I/usr/include/fuse diff --git a/tools/testing/selftests/filesystems/fuse/write_extend_eof_test= .c b/tools/testing/selftests/filesystems/fuse/write_extend_eof_test.c new file mode 100644 index 000000000000..ca6ce6eca382 --- /dev/null +++ b/tools/testing/selftests/filesystems/fuse/write_extend_eof_test.c @@ -0,0 +1,368 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Regression test for the fuse write-extend partial-EOF-page zeroing bug. + * + * A buffered write that extends i_size past a non-page-aligned EOF must z= ero + * the tail of the old last page. If an application has mmap'd that page = and + * stored into the post-EOF region (undefined until the file grows), the + * now-in-bounds tail must read back as zero, not as the stale stored byte= s. + * + * The bug is exposed on a non-writeback_cache server that keeps the page = cache + * across the write (FOPEN_KEEP_CACHE without FOPEN_DIRECT_IO). This test= is a + * raw /dev/fuse server in that mode; the backing data is always zero in t= he + * hole, so any non-zero byte a read sees is stale page-cache data. + * + * Requires root to mount fuse. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "../../kselftest_harness.h" + +#define FUSE_ROOT_ID 1 +#define FILE_INO 2 +#define MAX_WRITE (128 * 1024) +#define BACKING_SIZE (4 * 1024 * 1024) +#define POLLUTE 0xee + +/* Server-side state, shared with the responder thread. */ +struct server { + int fd; + unsigned char backing[BACKING_SIZE]; /* authoritative bytes */ + uint64_t size; +}; + +static void reply(int fd, uint64_t unique, int error, void *data, size_t l= en) +{ + struct fuse_out_header oh =3D { + .len =3D sizeof(oh) + (data ? len : 0), + .error =3D error, + .unique =3D unique, + }; + struct iovec iov[2] =3D { { &oh, sizeof(oh) }, { data, len } }; + + /* Errors here are teardown races (device closed on unmount); ignore. */ + if (writev(fd, iov, data ? 2 : 1) < 0) + return; +} + +static void fill_attr(struct fuse_attr *a, uint64_t ino, uint32_t mode, + uint64_t size) +{ + memset(a, 0, sizeof(*a)); + a->ino =3D ino; + a->mode =3D mode; + a->nlink =3D 1; + a->size =3D size; + a->blksize =3D sysconf(_SC_PAGESIZE); +} + +static void *server_thread(void *arg) +{ + struct server *s =3D arg; + static char buf[MAX_WRITE + 4096]; + + for (;;) { + ssize_t n =3D read(s->fd, buf, sizeof(buf)); + struct fuse_in_header *ih =3D (void *)buf; + + if (n < 0) { + if (errno =3D=3D EINTR || errno =3D=3D EAGAIN) + continue; + return NULL; /* device closed on unmount */ + } + if (n < (ssize_t)sizeof(*ih)) + continue; + + switch (ih->opcode) { + case FUSE_INIT: { + struct fuse_init_in *in =3D (void *)(ih + 1); + struct fuse_init_out out =3D {0}; + + /* No FUSE_WRITEBACK_CACHE: the exposed configuration. */ + out.major =3D FUSE_KERNEL_VERSION; + out.minor =3D FUSE_KERNEL_MINOR_VERSION; + out.max_readahead =3D in->max_readahead; + out.max_write =3D MAX_WRITE; + out.max_background =3D 16; + out.congestion_threshold =3D 12; + out.flags =3D FUSE_MAX_PAGES; + out.max_pages =3D MAX_WRITE / sysconf(_SC_PAGESIZE); + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_GETATTR: { + struct fuse_attr_out out =3D {0}; + int root =3D ih->nodeid =3D=3D FUSE_ROOT_ID; + + out.attr_valid =3D 3600; + fill_attr(&out.attr, ih->nodeid, + root ? (S_IFDIR | 0755) : (S_IFREG | 0644), + root ? 0 : s->size); + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_LOOKUP: { + struct fuse_entry_out out =3D {0}; + + out.nodeid =3D FILE_INO; + out.attr_valid =3D 3600; + out.entry_valid =3D 3600; + fill_attr(&out.attr, FILE_INO, S_IFREG | 0644, s->size); + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_OPEN: + case FUSE_OPENDIR: { + struct fuse_open_out out =3D {0}; + + /* Keep the cache across the write, but not direct I/O. */ + out.open_flags =3D FOPEN_KEEP_CACHE; + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_READ: { + struct fuse_read_in *in =3D (void *)(ih + 1); + uint64_t off =3D in->offset; + uint32_t size =3D in->size; + + if (off >=3D BACKING_SIZE) + size =3D 0; + else if (off + size > BACKING_SIZE) + size =3D BACKING_SIZE - off; + reply(s->fd, ih->unique, 0, s->backing + off, size); + break; + } + case FUSE_WRITE: { + struct fuse_write_in *in =3D (void *)(ih + 1); + struct fuse_write_out out =3D {0}; + uint64_t off =3D in->offset; + uint32_t size =3D in->size; + + if (off < BACKING_SIZE) { + uint32_t c =3D size; + + if (off + c > BACKING_SIZE) + c =3D BACKING_SIZE - off; + memcpy(s->backing + off, in + 1, c); + if (off + c > s->size) + s->size =3D off + c; + } + out.size =3D size; + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_SETATTR: { + struct fuse_setattr_in *in =3D (void *)(ih + 1); + struct fuse_attr_out out =3D {0}; + + if ((in->valid & FATTR_SIZE) && in->size <=3D BACKING_SIZE) { + if (in->size > s->size) + memset(s->backing + s->size, 0, + in->size - s->size); + s->size =3D in->size; + } + out.attr_valid =3D 3600; + fill_attr(&out.attr, ih->nodeid, S_IFREG | 0644, s->size); + reply(s->fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_FALLOCATE: { + struct fuse_fallocate_in *in =3D (void *)(ih + 1); + uint64_t end =3D in->offset + in->length; + + /* Only plain (size-extending) fallocate is used here. */ + if (!(in->mode & FALLOC_FL_KEEP_SIZE) && + end <=3D BACKING_SIZE && end > s->size) { + memset(s->backing + s->size, 0, end - s->size); + s->size =3D end; + } + reply(s->fd, ih->unique, 0, NULL, 0); + break; + } + case FUSE_FLUSH: + case FUSE_RELEASE: + case FUSE_RELEASEDIR: + case FUSE_FSYNC: + case FUSE_ACCESS: + reply(s->fd, ih->unique, 0, NULL, 0); + break; + case FUSE_FORGET: + break; + default: + reply(s->fd, ih->unique, -EOPNOTSUPP, NULL, 0); + break; + } + } +} + +FIXTURE(fuse) +{ + struct server *srv; + pthread_t thread; + char dir[64]; + long page; /* runtime page size */ + off_t eof; /* mid-page EOF, page-relative */ + int fd; /* open test file */ + char *map; /* mmap of the EOF page */ + int mounted; +}; + +FIXTURE_SETUP(fuse) +{ + char opts[128]; + pthread_t t; + + if (geteuid() !=3D 0) + SKIP(return, "need root to mount fuse"); + + self->page =3D sysconf(_SC_PAGESIZE); + self->fd =3D -1; + self->map =3D MAP_FAILED; + + self->srv =3D mmap(NULL, sizeof(*self->srv), PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_ANONYMOUS, -1, 0); + ASSERT_NE(MAP_FAILED, self->srv); + + self->srv->fd =3D open("/dev/fuse", O_RDWR); + ASSERT_GE(self->srv->fd, 0); + + strcpy(self->dir, "/tmp/fuse_weof_XXXXXX"); + ASSERT_NE(NULL, mkdtemp(self->dir)); + + snprintf(opts, sizeof(opts), + "fd=3D%d,rootmode=3D40000,user_id=3D0,group_id=3D0", + self->srv->fd); + ASSERT_EQ(0, mount("fuse", self->dir, "fuse", 0, opts)); + self->mounted =3D 1; + + ASSERT_EQ(0, pthread_create(&t, NULL, server_thread, self->srv)); + self->thread =3D t; +} + +FIXTURE_TEARDOWN(fuse) +{ + if (self->map !=3D MAP_FAILED) + munmap(self->map, self->page); + if (self->fd >=3D 0) + close(self->fd); + if (self->mounted) + umount2(self->dir, MNT_DETACH); + if (self->srv && self->srv !=3D MAP_FAILED) { + if (self->srv->fd > 0) + close(self->srv->fd); + munmap(self->srv, sizeof(*self->srv)); + } + if (self->dir[0]) + rmdir(self->dir); +} + +/* + * Create the test file with a mid-page EOF and mmap-store POLLUTE into its + * post-EOF tail (a legal store, undefined until the file grows). Leaves = the + * file open and the EOF page mapped in the fixture for the caller to exte= nd. + */ +static void pollute_eof_tail(struct __test_metadata *_metadata, + FIXTURE_DATA(fuse) * self) +{ + off_t eof =3D 2 * self->page + self->page / 4; + char path[128]; + char *buf; + + snprintf(path, sizeof(path), "%s/file", self->dir); + self->fd =3D open(path, O_RDWR | O_CREAT | O_TRUNC, 0644); + ASSERT_GE(self->fd, 0); + self->eof =3D eof; + + buf =3D malloc(eof); + ASSERT_NE(NULL, buf); + memset(buf, 'A', eof); + ASSERT_EQ(eof, pwrite(self->fd, buf, eof, 0)); + free(buf); + + self->map =3D mmap(NULL, self->page, PROT_READ | PROT_WRITE, MAP_SHARED, + self->fd, eof & ~(self->page - 1)); + ASSERT_NE(MAP_FAILED, self->map); + memset(self->map + (eof & (self->page - 1)), POLLUTE, + self->page - (eof & (self->page - 1))); +} + +/* Assert the old post-EOF tail [eof, end of its page) now reads back as z= ero. */ +static void assert_tail_zeroed(struct __test_metadata *_metadata, + FIXTURE_DATA(fuse) * self) +{ + off_t base =3D self->eof & ~(self->page - 1); + char *tail =3D malloc(self->page); + int i; + + ASSERT_NE(NULL, tail); + ASSERT_EQ(self->page, pread(self->fd, tail, self->page, base)); + for (i =3D self->eof & (self->page - 1); i < self->page; i++) + ASSERT_EQ(0, tail[i]); + free(tail); +} + +/* Basic: pollute the post-EOF tail, extend past it by a later write. */ +TEST_F(fuse, write_extend) +{ + pollute_eof_tail(_metadata, self); + ASSERT_EQ(4, pwrite(self->fd, "data", 4, 5 * self->page + self->page / 3)= ); + assert_tail_zeroed(_metadata, self); +} + +/* Extend via ftruncate() rather than a write. */ +TEST_F(fuse, ftruncate_extend) +{ + pollute_eof_tail(_metadata, self); + ASSERT_EQ(0, ftruncate(self->fd, 8 * self->page)); + assert_tail_zeroed(_metadata, self); +} + +/* Extend via fallocate() starting at the old EOF. */ +TEST_F(fuse, fallocate_extend) +{ + pollute_eof_tail(_metadata, self); + ASSERT_EQ(0, fallocate(self->fd, 0, self->eof, 4 * self->page)); + assert_tail_zeroed(_metadata, self); +} + +/* A write landing inside the old EOF page must not clobber its own data. = */ +TEST_F(fuse, extend_into_eof_page_preserves_data) +{ + off_t base, wr; + char *buf, *rd; + int i; + + pollute_eof_tail(_metadata, self); + base =3D self->eof & ~(self->page - 1); + wr =3D base + 3 * self->page / 4; /* starts in the EOF page */ + + buf =3D malloc(2 * self->page); + ASSERT_NE(NULL, buf); + memset(buf, 'B', 2 * self->page); + ASSERT_EQ(2 * self->page, pwrite(self->fd, buf, 2 * self->page, wr)); + free(buf); + + rd =3D malloc(self->page); + ASSERT_NE(NULL, rd); + ASSERT_EQ(self->page, pread(self->fd, rd, self->page, base)); + /* [eof, wr) is hole -> zero; [wr, page) is written data -> 'B'. */ + for (i =3D self->eof & (self->page - 1); i < wr - base; i++) + ASSERT_EQ(0, rd[i]); + for (i =3D wr - base; i < self->page; i++) + ASSERT_EQ('B', rd[i]); + free(rd); +} + +TEST_HARNESS_MAIN --=20 2.50.1