From nobody Mon Sep 28 16:21:15 2026 Received: from pdx-out-006.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-006.esa.us-west-2.outbound.mail-perimeter.amazon.com [52.26.1.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 436C9242910; Thu, 20 Aug 2026 13:02:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=52.26.1.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787230963; cv=none; b=cWiutemtlcCiV4arpL0kFyeKn8/9u3kMLLhvsd0tlSbe8bvxyaObDaceUHAdcOyQuLiVRxm+CFAv2ne9NjgRYexIAbJUGsGnJ1lSpOx5nxmp56WxaETqoI7G8xzNkRx9Tgxnqw4L9r5W7UWSXvBwyban/eaCRJZhPS4LcH78yjY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787230963; c=relaxed/simple; bh=3BOSagCYHSXYSSlAI3BXVEaLc7IJaiOllCs2CcwHhyc=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=axonnfRxiW2MyBjMdaMR7TovJn8UoDTv9bjS/tbV8NBWyBlJrKlDMqKFmkhVmd6/HKutB9Qfe5UsA+sygATFd53ONpwjEctnf+zMtX9W5OU2RP9Z+2v8FBVXI0MRU2hoCArc/Rzx7cwt4b78hPSyEJS4ltAIy0hollbdzAzJdM8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=shL2wc0C; arc=none smtp.client-ip=52.26.1.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="shL2wc0C" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1787230962; x=1818766962; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=TmhlGf2UjnP+JtX2La7bxsJ2rBRB1E5qckNV4lWaN1E=; b=shL2wc0C/7G2kGS488jcqAce68fJoLT06RmtJQ1tBzWi+KtBj5faaB6E Q0ZNUSzC1CcwZoJfCxZdMr+q+M+7pRSO58vgJVcgXCUgYnV15V54qPMvC JnQ4WKg8Gp3j7YTomOANf0Ipw0nX85aQRcx5ubCJQe3mOMTBJ4CWh/nPj +3dHPsZ77f7q9FUhD67XrDP+cZRsDczh/2vmJ3N6gza/lFzIqqLCJqlcH +rdgOU2WNT+SKt3T5T+73m1XcP5AVMibussHcTgQZEj2ua0g3ERDVMkEt IdvO9ik0N/zQWKoZLaS/lEk1DXvx2l2oSfhmc86nbU4wQEu7OwgRAiTC9 w==; X-CSE-ConnectionGUID: QI4sMvMXRVem7H/WweT9xQ== X-CSE-MsgGUID: 5xnput/SScCTHmYrnkEpKQ== X-IronPort-AV: E=Sophos;i="6.25,233,1779148800"; d="scan'208";a="26501818" Received: from ip-10-5-6-203.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.6.203]) by internal-pdx-out-006.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 20 Aug 2026 13:02:42 +0000 Received: from EX19MTAUWB001.ant.amazon.com [205.251.233.104:9293] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.23.75:2525] with esmtp (Farcaster) id 6e438dc1-2af6-4d16-9ca6-4222d902327f; Thu, 20 Aug 2026 13:02:41 +0000 (UTC) X-Farcaster-Flow-ID: 6e438dc1-2af6-4d16-9ca6-4222d902327f Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWB001.ant.amazon.com (10.250.64.248) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Thu, 20 Aug 2026 13:02:41 +0000 Received: from dev-dsk-jamz-1e-e35f4cd9.us-east-1.amazon.com (10.189.35.140) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Thu, 20 Aug 2026 13:02:40 +0000 From: Jimmy Zuber To: Miklos Szeredi CC: , , , , Shuah Khan , Subject: [PATCH v3 1/2] fuse: allow FUSE_SYNCFS for privileged userspace servers Date: Thu, 20 Aug 2026 13:01:57 +0000 Message-ID: <20260820130158.254808-2-jamz@amazon.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260820130158.254808-1-jamz@amazon.com> References: <20260820130158.254808-1-jamz@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D038UWB002.ant.amazon.com (10.13.139.185) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Propagating syncfs()/sync() to a FUSE server via FUSE_SYNCFS lets the server flush its own cached or intermediate state when userspace asks the filesystem to sync. This is currently enabled only for virtiofs and fuseblk, because an untrusted server can use it to stall sync() indefinitely (see commit 2d82ab251ef0 ("virtiofs: propagate sync() to file server"), and commit d3906d8f3cee ("fuse: enable FUSE_SYNCFS for all fuseblk servers")). Both of those mount types require host privilege to set up, so the server is trusted not to abuse it. There is nothing virtiofs- or block-specific about wanting to handle syncfs(), though. A plain /dev/fuse server is just as entitled to participate in the sync() path -- so that data it has buffered reaches stable storage when the user asks for it -- provided it is equally trusted. The trust property that virtiofs and fuseblk satisfy is that neither can be mounted without CAP_SYS_ADMIN in the initial user namespace (neither sets FS_USERNS_MOUNT). A plain fuse mount does set FS_USERNS_MOUNT, so its existence guarantees no such privilege; the privilege has to be checked rather than assumed. Add an opt-in INIT flag, FUSE_HAS_SYNCFS, and honor it only when the server opened /dev/fuse with CAP_SYS_ADMIN in the initial user namespace, recorded at mount time via file_ns_capable(). This is the same privilege that mounting virtiofs or fuseblk requires, applied to the process that actually services (and could stall) the connection. Checking the device opener's capability -- rather than the mount's user namespace -- avoids treating an unprivileged server that merely happens to run in the initial user namespace (e.g. a normal sshfs mount) as trusted. The flag is only advertised to servers that pass this check, so an unprivileged server is never invited to opt in (and is ignored by fuse_syncfs_enable() if it sets the flag anyway). Signed-off-by: Jimmy Zuber Assisted-by: Claude:claude-opus-4-8 [Claude-Code] --- fs/fuse/fuse_i.h | 9 +++++++++ fs/fuse/inode.c | 28 ++++++++++++++++++++++++++++ include/uapi/linux/fuse.h | 12 +++++++++++- 3 files changed, 48 insertions(+), 1 deletion(-) diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h index c8d4c5f3af7e..97d356588d00 100644 --- a/fs/fuse/fuse_i.h +++ b/fs/fuse/fuse_i.h @@ -391,6 +391,7 @@ struct fuse_fs_context { bool no_control:1; bool no_force_umount:1; bool legacy_opts_show:1; + bool syncfs_capable:1; enum fuse_dax_mode dax_mode; unsigned int max_read; unsigned int blksize; @@ -675,6 +676,14 @@ struct fuse_conn { /** @sync_fs: Propagate syncfs() to server */ unsigned int sync_fs:1; =20 + /** + * @syncfs_capable: the privilege required to honor FUSE_HAS_SYNCFS was + * present when /dev/fuse was opened (CAP_SYS_ADMIN in the initial user + * namespace), i.e. the same privilege that mounting virtiofs/fuseblk + * requires. + */ + unsigned int syncfs_capable:1; + /** @init_security: Initialize security xattrs when creating a new inode = */ unsigned int init_security:1; =20 diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c index 1c6ee01c6796..970d4d2e9833 100644 --- a/fs/fuse/inode.c +++ b/fs/fuse/inode.c @@ -803,6 +803,15 @@ static int fuse_opt_fd(struct fs_context *fsc, struct = file *file) if (file->f_cred->user_ns !=3D fsc->user_ns) return invalfc(fsc, "wrong user namespace for fuse device"); =20 + /* + * Record whether the server opened /dev/fuse with CAP_SYS_ADMIN in the + * initial user namespace -- the same privilege that mounting virtiofs + * or fuseblk requires. Only such servers are trusted to receive + * FUSE_SYNCFS (see fuse_syncfs_enable()). + */ + ctx->syncfs_capable =3D file_ns_capable(file, &init_user_ns, + CAP_SYS_ADMIN); + ctx->fud =3D fuse_dev_grab(file); =20 return 0; @@ -1269,6 +1278,16 @@ struct fuse_init_args { struct fuse_mount *fm; }; =20 +/* + * A server can stall syncfs()/sync(), so only honor FUSE_HAS_SYNCFS for + * servers that opened /dev/fuse with CAP_SYS_ADMIN in the initial user + * namespace -- the same privilege required to mount virtiofs or fuseblk. + */ +static bool fuse_syncfs_enable(struct fuse_conn *fc, u64 flags) +{ + return (flags & FUSE_HAS_SYNCFS) && fc->syncfs_capable; +} + static void process_init_reply(struct fuse_args *args, int error) { struct fuse_init_args *ia =3D container_of(args, typeof(*ia), args); @@ -1410,6 +1429,9 @@ static void process_init_reply(struct fuse_args *args= , int error) =20 if (flags & FUSE_REQUEST_TIMEOUT) timeout =3D arg->request_timeout; + + if (fuse_syncfs_enable(fc, flags)) + fc->sync_fs =3D 1; } else { ra_pages =3D fc->max_read / PAGE_SIZE; fc->no_lock =3D 1; @@ -1479,6 +1501,11 @@ static struct fuse_init_args *fuse_new_init(struct f= use_mount *fm) flags |=3D FUSE_SUBMOUNTS; if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH)) flags |=3D FUSE_PASSTHROUGH; + /* Only offered to sufficiently privileged servers; see + * fuse_syncfs_enable(). + */ + if (fm->fc->syncfs_capable) + flags |=3D FUSE_HAS_SYNCFS; =20 /* * This is just an information flag for fuse server. No need to check @@ -1774,6 +1801,7 @@ int fuse_fill_super_common(struct super_block *sb, st= ruct fuse_fs_context *ctx) =20 fc->default_permissions =3D ctx->default_permissions; fc->allow_other =3D ctx->allow_other; + fc->syncfs_capable =3D ctx->syncfs_capable; fc->user_id =3D ctx->user_id; fc->group_id =3D ctx->group_id; fc->legacy_opts_show =3D ctx->legacy_opts_show; diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h index 7435e09c87fe..10a7f31c4bdf 100644 --- a/include/uapi/linux/fuse.h +++ b/include/uapi/linux/fuse.h @@ -248,6 +248,9 @@ * - add bufpool offset field to fuse_uring_ent_in_out struct * - add FUSE_URING_ZERO_COPY, FUSE_URING_ENT_ZERO_COPY, and * FOPEN_IO_URING_ZERO_COPY flag + * + * 7.47 + * - add FUSE_HAS_SYNCFS opt-in flag for privileged userspace servers */ =20 #ifndef _LINUX_FUSE_H @@ -283,7 +286,7 @@ #define FUSE_KERNEL_VERSION 7 =20 /** Minor version number of this interface */ -#define FUSE_KERNEL_MINOR_VERSION 46 +#define FUSE_KERNEL_MINOR_VERSION 47 =20 /** The node ID of the root inode */ #define FUSE_ROOT_ID 1 @@ -464,6 +467,12 @@ struct fuse_file_lock { * FUSE_REQUEST_TIMEOUT: kernel supports timing out requests. * init_out.request_timeout contains the timeout (in secs) * FUSE_HAS_IO_URING_BUFPOOL: kernel supports io-uring buffer pools + * FUSE_HAS_SYNCFS: server requests that syncfs()/sync() be propagated as + * FUSE_SYNCFS requests. Since an untrusted server can use this + * to stall sync(), it is only honored when /dev/fuse was opened + * with CAP_SYS_ADMIN in the initial user namespace (the same + * privilege that mounting virtiofs or fuseblk requires). + * Insufficiently privileged servers ignore it. */ #define FUSE_ASYNC_READ (1 << 0) #define FUSE_POSIX_LOCKS (1 << 1) @@ -512,6 +521,7 @@ struct fuse_file_lock { #define FUSE_OVER_IO_URING (1ULL << 41) #define FUSE_REQUEST_TIMEOUT (1ULL << 42) #define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43) +#define FUSE_HAS_SYNCFS (1ULL << 44) =20 /** * CUSE INIT request/reply flags --=20 2.50.1 From nobody Mon Sep 28 16:21:15 2026 Received: from pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.245.243.92]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C1BBA40B6C9; Thu, 20 Aug 2026 13:03:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.245.243.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787230999; cv=none; b=lXMrJ5iTJ7VlOxKEBb8J6lMx/LT4BCu9NTykTjX7eF87Snc3hfRdGOsoVHh+iDqRbZq+6WGNh892rHuByiozPpLhEC1FuSkNo2NnlOZrhVpuU08t/GqBYRwr3WBXtXQi++gIJNQEvi4jNiHmLGGOxU50nNgz3VlJtBnWVWGb8Uo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787230999; c=relaxed/simple; bh=TC/gEjcaDahOGzLQ0qBz0TKbqHmaEx0xaOk0vlqdkRQ=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=uT1h7XSIQcq9m8oiDJG/Af3x9kT3FBNaDo5s9VgIB1tf70sF2DRxTeNqqMhqlIw1gC5K4o2/Zab/FzZ0jAiUNNtgI3ijDR1gcF/EQr//yjUQ+W9mcDJhqtaNHg1Bz0GtJqqAV5vkkuPEVv+lTtoMmwnETrHEr6cotBi6U17mZ1s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=sXVzWcrV; arc=none smtp.client-ip=44.245.243.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="sXVzWcrV" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1787230997; x=1818766997; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=GyV3eF5AwmF0ksd3fZmNP5deV2Ca8CiMD7eozPEEj9A=; b=sXVzWcrVXZ3X3BgynLWGcRCTwb3FfoPNEV3pPQPOEoubjdlxBglLGEVl hEI9ckj4u6Pv2KR09dJ4l40wDkfnYRJFWjLbR5eMrnPH5huHA6H2HUJir uPx7utw0PVXFRvvLyicirXg6pCXxYb1/104hWcVoV+pooMk8sYwFPJTSH RHTIO7OaJ8RgaL4htHyLKpFZmYSWpqtQjr4izzY+0cNVG4Ve0fdC4Xo/9 hujd6WwJog0TT8ccB4JwAZpMM2PUG3qS8NXpOZL5F1XB9x3STFo/ny4PS 79Tk0BhDerLiGQ4Sl+4WSbVqmCwpBw4JjwJyyVeqgRNWX3Mp8F27hdNMs A==; X-CSE-ConnectionGUID: FexE2yvGS6G//J+zfKd2BQ== X-CSE-MsgGUID: D5cS9HBqRpGlv7rZ/Mjb5A== X-IronPort-AV: E=Sophos;i="6.25,233,1779148800"; d="scan'208";a="25974355" Received: from ip-10-5-0-115.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.0.115]) by internal-pdx-out-001.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 20 Aug 2026 13:03:17 +0000 Received: from EX19MTAUWC002.ant.amazon.com [205.251.233.51:31619] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.23.75:2525] with esmtp (Farcaster) id 07032202-e1ff-451c-b203-6c7da7d62e05; Thu, 20 Aug 2026 13:03:16 +0000 (UTC) X-Farcaster-Flow-ID: 07032202-e1ff-451c-b203-6c7da7d62e05 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWC002.ant.amazon.com (10.250.64.143) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Thu, 20 Aug 2026 13:03:15 +0000 Received: from dev-dsk-jamz-1e-e35f4cd9.us-east-1.amazon.com (10.189.35.140) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Thu, 20 Aug 2026 13:03:14 +0000 From: Jimmy Zuber To: Miklos Szeredi CC: , , , , Shuah Khan , Subject: [PATCH v3 2/2] selftests/fuse: add test for FUSE_HAS_SYNCFS privilege gating Date: Thu, 20 Aug 2026 13:01:58 +0000 Message-ID: <20260820130158.254808-3-jamz@amazon.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260820130158.254808-1-jamz@amazon.com> References: <20260820130158.254808-1-jamz@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D036UWB002.ant.amazon.com (10.13.139.139) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Add a selftest that talks the raw FUSE protocol over /dev/fuse (rather than via libfuse, which negotiates INIT internally) so it can both choose whether to advertise FUSE_HAS_SYNCFS and directly observe whether a FUSE_SYNCFS opcode is forwarded by the kernel. Three cases are covered: T1: host-root mount, server sets FUSE_HAS_SYNCFS -> FUSE_SYNCFS must reach the server. T2: host-root mount, server does not opt in -> FUSE_SYNCFS must not be sent (back-compat). T3: server opts in but opened /dev/fuse without CAP_SYS_ADMIN while still in the initial user namespace -> FUSE_SYNCFS must be withheld. This is the case that distinguishes gating on the server's privilege from gating on the mount's user namespace. Signed-off-by: Jimmy Zuber Assisted-by: Claude:claude-opus-4-8 [Claude-Code] --- .../selftests/filesystems/fuse/.gitignore | 1 + .../selftests/filesystems/fuse/Makefile | 2 +- .../selftests/filesystems/fuse/test_syncfs.c | 309 ++++++++++++++++++ 3 files changed, 311 insertions(+), 1 deletion(-) create mode 100644 tools/testing/selftests/filesystems/fuse/test_syncfs.c diff --git a/tools/testing/selftests/filesystems/fuse/.gitignore b/tools/te= sting/selftests/filesystems/fuse/.gitignore index 3e72e742d08e..4e92f363c74a 100644 --- a/tools/testing/selftests/filesystems/fuse/.gitignore +++ b/tools/testing/selftests/filesystems/fuse/.gitignore @@ -1,3 +1,4 @@ # SPDX-License-Identifier: GPL-2.0-only fuse_mnt fusectl_test +test_syncfs diff --git a/tools/testing/selftests/filesystems/fuse/Makefile b/tools/test= ing/selftests/filesystems/fuse/Makefile index 612aad69a93a..cbba01635226 100644 --- a/tools/testing/selftests/filesystems/fuse/Makefile +++ b/tools/testing/selftests/filesystems/fuse/Makefile @@ -2,7 +2,7 @@ =20 CFLAGS +=3D -Wall -O2 -g $(KHDR_INCLUDES) =20 -TEST_GEN_PROGS :=3D fusectl_test +TEST_GEN_PROGS :=3D fusectl_test test_syncfs TEST_GEN_FILES :=3D fuse_mnt =20 include ../../lib.mk diff --git a/tools/testing/selftests/filesystems/fuse/test_syncfs.c b/tools= /testing/selftests/filesystems/fuse/test_syncfs.c new file mode 100644 index 000000000000..d3ca2613b45c --- /dev/null +++ b/tools/testing/selftests/filesystems/fuse/test_syncfs.c @@ -0,0 +1,309 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Test that FUSE_SYNCFS is propagated to a userspace server only when the + * server opts in with FUSE_HAS_SYNCFS *and* opened /dev/fuse with + * CAP_SYS_ADMIN in the initial user namespace. + * + * Unlike the libfuse-based selftests, this talks the raw FUSE wire protoc= ol + * over /dev/fuse so it can (a) choose whether to advertise FUSE_HAS_SYNCF= S in + * the INIT reply and (b) observe directly whether a FUSE_SYNCFS opcode ar= rives. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "../../kselftest.h" + +#define FUSE_ROOT_ID 1 + +/* + * Add or drop CAP_SYS_ADMIN in the effective set via raw capget/capset + * (avoids a libcap dependency). Used to construct a server that opens + * /dev/fuse without the privilege FUSE_HAS_SYNCFS requires. + */ +static int set_sysadmin(int on) +{ + struct __user_cap_header_struct hdr =3D { + .version =3D _LINUX_CAPABILITY_VERSION_3, + .pid =3D 0, + }; + struct __user_cap_data_struct data[_LINUX_CAPABILITY_U32S_3] =3D {}; + unsigned int idx =3D CAP_SYS_ADMIN >> 5; + __u32 bit =3D 1U << (CAP_SYS_ADMIN & 31); + + if (syscall(SYS_capget, &hdr, data)) + return -1; + if (on) + data[idx].effective |=3D bit; + else + data[idx].effective &=3D ~bit; + return syscall(SYS_capset, &hdr, data); +} + +/* + * eventfd the server child writes once when it receives FUSE_SYNCFS; the + * parent poll()s it to observe (or rule out) propagation. + */ +static int syncfs_evfd; + +static void reply(int fd, uint64_t unique, int error, void *data, size_t l= en) +{ + struct fuse_out_header oh =3D { + .len =3D sizeof(oh) + len, + .error =3D error, + .unique =3D unique, + }; + struct iovec iov[2] =3D { + { &oh, sizeof(oh) }, + { data, len }, + }; + + if (writev(fd, iov, data ? 2 : 1) < 0) + ksft_print_msg("server writev failed: %s\n", strerror(errno)); +} + +static void fill_attr(struct fuse_attr *a, uint64_t ino, uint32_t mode, + uint32_t nlink) +{ + memset(a, 0, sizeof(*a)); + a->ino =3D ino; + a->mode =3D mode; + a->nlink =3D nlink; + a->blksize =3D 4096; +} + +/* + * Minimal FUSE server. Advertises FUSE_HAS_SYNCFS in its INIT reply iff + * @advertise is set. Signals syncfs_evfd when a FUSE_SYNCFS opcode arriv= es. + */ +#define SERVER_MAX_WRITE 65536 +static void run_server(int fd, int advertise) +{ + /* + * The kernel rejects reads (EINVAL) whose buffer is smaller than + * max_write + header, so size generously for the max_write we + * advertise in the INIT reply below. + */ + static char buf[SERVER_MAX_WRITE + 4096]; + + for (;;) { + ssize_t n =3D read(fd, buf, sizeof(buf)); + struct fuse_in_header *ih =3D (void *)buf; + + if (n < 0) { + if (errno =3D=3D EINTR || errno =3D=3D EAGAIN) + continue; + return; /* device closed on unmount/abort */ + } + if (n < (ssize_t)sizeof(*ih)) + continue; + + switch (ih->opcode) { + case FUSE_INIT: { + struct fuse_init_in *in =3D (void *)(ih + 1); + struct fuse_init_out out =3D {0}; + uint64_t flags =3D FUSE_INIT_EXT; + + out.major =3D FUSE_KERNEL_VERSION; + out.minor =3D FUSE_KERNEL_MINOR_VERSION; + out.max_readahead =3D in->max_readahead; + out.max_write =3D SERVER_MAX_WRITE; + out.max_background =3D 16; + out.congestion_threshold =3D 12; + if (advertise) + flags |=3D FUSE_HAS_SYNCFS; + out.flags =3D flags; + out.flags2 =3D flags >> 32; + reply(fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_GETATTR: { + struct fuse_attr_out out =3D {0}; + + fill_attr(&out.attr, FUSE_ROOT_ID, S_IFDIR | 0755, 2); + reply(fd, ih->unique, 0, &out, sizeof(out)); + break; + } + case FUSE_SYNCFS: { + uint64_t one =3D 1; + + if (write(syncfs_evfd, &one, sizeof(one)) < 0) + ksft_print_msg("server eventfd write failed: %s\n", + strerror(errno)); + reply(fd, ih->unique, 0, NULL, 0); + break; + } + default: + /* + * Anything else (e.g. OPENDIR from opening the mount + * root) is not needed to drive this test; -ENOSYS lets + * the kernel proceed. + */ + reply(fd, ih->unique, -ENOSYS, NULL, 0); + break; + } + } +} + +/* + * Mount a fuse fs backed by a forked server, issue syncfs(), and report + * whether the server observed FUSE_SYNCFS. Returns 0 on success, -1 if t= he + * environment could not support the test (caller should skip). + * + * If @unpriv_open is set, /dev/fuse is opened with CAP_SYS_ADMIN dropped + * (regained only for the mount() syscall), so the device opener -- the + * server principal FUSE_HAS_SYNCFS is gated on -- lacks the privilege even + * though it remains in the initial user namespace. + */ +static int do_mount_and_syncfs(const char *mnt, int advertise, int unpriv_= open, + int *seen) +{ + struct pollfd pfd =3D { .events =3D POLLIN }; + char opts[256]; + int fd, mfd =3D -1, i; + pid_t pid; + + syncfs_evfd =3D eventfd(0, EFD_CLOEXEC); + if (syncfs_evfd < 0) + return -1; + + if (unpriv_open && set_sysadmin(0)) + goto out_evfd; + fd =3D open("/dev/fuse", O_RDWR); + if (unpriv_open && set_sysadmin(1)) { + if (fd >=3D 0) + close(fd); + goto out_evfd; + } + if (fd < 0) + goto out_evfd; + + mkdir(mnt, 0755); + snprintf(opts, sizeof(opts), + "fd=3D%d,rootmode=3D40000,user_id=3D%d,group_id=3D%d", + fd, getuid(), getgid()); + + if (mount("fuse", mnt, "fuse", 0, opts) < 0) + goto out_fd; + + pid =3D fork(); + if (pid < 0) + goto out_umount; + if (pid =3D=3D 0) { + run_server(fd, advertise); + _exit(0); + } + + /* + * The parent does not service the fuse fd; the child does. Close our + * copy so the kernel sees a single server, and so that if the child + * dies the connection aborts instead of hanging us forever. + */ + close(fd); + + /* + * mount() returns before the server has answered FUSE_INIT, so the + * first open() can race and fail with ENOTCONN; retry until the + * handshake settles. + */ + for (i =3D 0; i < 1000; i++) { + mfd =3D open(mnt, O_RDONLY | O_DIRECTORY); + if (mfd >=3D 0) + break; + usleep(1000); + } + if (mfd >=3D 0) { + syncfs(mfd); + close(mfd); + } + + /* + * No waiting is needed: the server writes syncfs_evfd before it replies + * to FUSE_SYNCFS, and that reply is what unblocks the synchronous + * syncfs() above. So once syncfs() has returned, the eventfd is already + * signalled if the opcode was propagated, and will never be otherwise. + * poll() with a zero timeout therefore decides both cases immediately. + */ + pfd.fd =3D syncfs_evfd; + *seen =3D poll(&pfd, 1, 0) > 0 && (pfd.revents & POLLIN); + + kill(pid, SIGKILL); + waitpid(pid, NULL, 0); + umount2(mnt, MNT_DETACH); + close(syncfs_evfd); + return 0; + +out_umount: + umount2(mnt, MNT_DETACH); +out_fd: + close(fd); +out_evfd: + close(syncfs_evfd); + return -1; +} + +int main(void) +{ + char mnt[] =3D "/tmp/fuse_syncfs_XXXXXX"; + int seen, ret; + + ksft_print_header(); + ksft_set_plan(3); + + /* Hard watchdog: never let a stuck syncfs hang the test runner. */ + signal(SIGALRM, SIG_DFL); + alarm(60); + + if (geteuid() !=3D 0) + ksft_exit_skip("test requires root to mount fuse\n"); + + if (!mkdtemp(mnt)) + ksft_exit_fail_msg("mkdtemp failed\n"); + + /* T1: host-root mount, server opts in -> syncfs must reach server. */ + ret =3D do_mount_and_syncfs(mnt, 1, 0, &seen); + if (ret < 0) + ksft_test_result_skip("T1: could not mount fuse\n"); + else + ksft_test_result(seen, + "T1 host-root + FUSE_HAS_SYNCFS: server receives FUSE_SYNCFS\n"); + + /* T2: host-root mount, server does NOT opt in -> no FUSE_SYNCFS. */ + ret =3D do_mount_and_syncfs(mnt, 0, 0, &seen); + if (ret < 0) + ksft_test_result_skip("T2: could not mount fuse\n"); + else + ksft_test_result(!seen, + "T2 host-root, no opt-in: server does NOT receive FUSE_SYNCFS\n"); + + /* + * T3: server opts in but opened /dev/fuse without CAP_SYS_ADMIN while + * still in the initial user namespace -> kernel must withhold + * FUSE_SYNCFS. This is the case that distinguishes gating on the + * server's privilege from gating on the mount's user namespace. + */ + ret =3D do_mount_and_syncfs(mnt, 1, 1, &seen); + if (ret < 0) + ksft_test_result_skip("T3: could not mount fuse unprivileged\n"); + else + ksft_test_result(!seen, + "T3 init_userns, opener lacks CAP_SYS_ADMIN: FUSE_SYNCFS withheld\n"); + + rmdir(mnt); + ksft_finished(); +} --=20 2.50.1