From nobody Sat Jul 25 04:17:19 2026 Received: from mail-wm1-f50.google.com (mail-wm1-f50.google.com [209.85.128.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F1CE2344D92 for ; Sat, 18 Jul 2026 19:15:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784402160; cv=none; b=UXsSvFnRq2WwYIH/X4wa7yiwPTVHQPQVLEGNiAOK2gtaMII9sO5pFTZDF+rxr0yKg1wpVHuJzfbNBKCRpP/X2KMzSys2sEYdbMzJwIM7SQqQkZ6vNch7Ugzh5TogrTgCPO/AdivaeKuMTGpKx4dxIRGH4wNYmQsY3rGBj1Wi8lE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784402160; c=relaxed/simple; bh=dUfONPU9riLiIRrViYYJAhG4mXBGTpEnKqbvgYwE/+U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dgQGt2edK6psiOPBLpexQD9Y4fi473jR7CUc70KsAnCP4Hz7Cnu1aDOwzC3GZPwNz9TujA4lwqsEUd9MKFZcuV1ojtuGgA5mSr5/3cXT9YKsRpYIxSBMynvo1Iwz0SjoJ66KpFs6DgJl7PB7c6qKKAkkRbi7mla7iEnSQEgybdU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=eVukKhcN; arc=none smtp.client-ip=209.85.128.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="eVukKhcN" Received: by mail-wm1-f50.google.com with SMTP id 5b1f17b1804b1-4953de5be0aso22324545e9.0 for ; Sat, 18 Jul 2026 12:15:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784402156; x=1785006956; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LpnnUCjvFhQvZ/hdXCpoqY+iircQ7nkPeqN838t5fO0=; b=eVukKhcNiFmVjVT2lO5UfxU7R7rUk6Bg6YNlOM0p1DD3jcr+L0m0CafcPvGHKmO7tW kS+meqo8K3daCJPE7E7FH/bfRIGkTwrxZAFye/igJsr64VgLsQ5F+vwHGI4KnI3uDkIM jeDseMHwHykeJ05g6kWLbc3XP9LI6IQSWQQ3whC613nIl0lRclaDPGcealG6iUJ67eJe ampn6kFwHxjOR7RN//p0nj5jnY5l5iCcuz6sqRJQXzbkeTT7gE/4K0LaLpAkLLpsENJK 0/1+sxW7g20jzSOSlSbht9hE9MQPbGyi5Uis8nmDyVzw7EdACLwp9CXwPkkfL8x45sD5 1VYA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784402156; x=1785006956; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=LpnnUCjvFhQvZ/hdXCpoqY+iircQ7nkPeqN838t5fO0=; b=oByLXtRYNaykDBwmkP7kxDGcoAlxE9qUyszCyu37dWw2Jd9Y2BbaGpl1DOeiVhYm/w aIPCrBweQT6GFszVjZ00JYXa+BNB1MPoPAq8WH34D6DrmOx+UVKmFpJYc9YF6E6Dm0rf q71A7AHXiZrpewwNjHIIGPq1Eri0McQ3iNUt+0bgdoYHskwbTcLUq22Wre3I+NnNq+9T iqmn0E3IdQ4FH3kNx0dKP479S5rICJ6RroPN1oXfz7XLLcCeJFQqZvW7DpF/4It7Y7/K gZR0QJ7C9LdVPzJ+Yc5aSLFj1Y2I6LPa/mePBrudUcn+pZ3tuSWvkfteGzzTVmITddO0 Mtrw== X-Forwarded-Encrypted: i=1; AHgh+RpjPV0TIBGrce5RcSBLxdtNghNv48jG3ucqhvCzhR6m5Ij+nJkizuoR8aGBH9FhgSUq5Ybm7n9L+9e9dKw=@vger.kernel.org X-Gm-Message-State: AOJu0YxNpPrZrJW90KK0ejacAo13fZ4HtKZW1msvVGq7nHwbRej+gSef iqXqEf1qLwRTQlwqJsAunC8YOis9udIVNTIURU/JnyNz6YqVsFdYEnq/ X-Gm-Gg: AfdE7ckt6/GX+57pu0muljHUh8odYBfJ+IURJNgvtNGJ1yJzf8Mz6ym2gITdbmM9+QK 9CETxPfNmgGHwW/2P1DjbHDac8+QCxGcA1CS0kgEBy0vbF/u0uw3NZNcsRxF+iBZ1kfS3R1yN4t MI3EsSD9adJmEgZAwn3JnXGr54tXWyqUNRRpCM6oTvKN3CV7uuCmPeAISSH5ZcozsZnEB1yQ8vy ki3QDQqVDdY87/ikbvu/qr0dC/Wq5JWvP6pM21id29RyKoAisSEollxoeVGCE1XLFiZC6SNWdSR +/26YGD7E7qARNMt77DZdqB53EvDker74IvPnaClP1l9DYgMrTx25cmjDDkYjqoRRto75LIbM0F vCHQt8uQU2BC8zJaZocP3ecgOLzp382peWnQ8DdZ3nrLVmSDb8HQhR6hNhpHvzqw5+8alZ6Ctrb JxxlJVMw== X-Received: by 2002:a05:600c:46c3:b0:495:522f:d996 with SMTP id 5b1f17b1804b1-495522fdb00mr35405945e9.18.1784402156071; Sat, 18 Jul 2026 12:15:56 -0700 (PDT) Received: from spark.Home ([2001:8a0:7280:4000:c2b:a5ab:a31c:e0a5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2e8529sm140037005e9.11.2026.07.18.12.15.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 12:15:55 -0700 (PDT) From: Eric Curtin To: Alexander Viro , Christian Brauner Cc: Jan Kara , Jonathan Corbet , Shuah Khan , Eric Biggers , "Theodore Y . Ts'o" , Gao Xiang , Chao Yu , fsverity@lists.linux.dev, linux-erofs@lists.ozlabs.org, linux-fsdevel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, Eric Curtin Subject: [RFC PATCH 1/2] init: support mounting the root filesystem from an image file Date: Sat, 18 Jul 2026 20:15:50 +0100 Message-ID: <20260718191551.1703670-2-ericcurtin17@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718191551.1703670-1-ericcurtin17@gmail.com> References: <20260718191551.1703670-1-ericcurtin17@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Image-based Linux systems (bootc-style OS updaters, ChromeOS/Android-like A/B schemes, embedded appliances) commonly keep one or more immutable root filesystem images as plain files on a writable "carrier" filesystem and pick one at boot. Today that always requires an initramfs, even when the initramfs has nothing else to do: the kernel can only mount a block device (or NFS/CIFS/ubi/mtd) as root, so userspace must mount the carrier, loop-mount the image and switch_root into it. Add a rootimage=3D parameter naming an image file on the filesystem specified by root=3D. When set, root=3D only designates the carrier: after mounting it as usual, mount the image file read-only on /image and make that the root that prepare_namespace() pivots into. The image is mounted through the filesystem's file-backed mount support (available in erofs since v6.12), so no loop device is involved. rootimagefstype=3D and rootimageflags=3D mirror rootfstype=3D/rootflags=3D for the image mount, wh= ile "ro"/"rw"/rootflags=3D keep applying to the carrier. The carrier mount is detached again by default; the image mount keeps its superblock pinned, so userspace can still mount it later. Since systems following this model virtually always need the carrier mounted (it holds their writable state), rootimagesrcdir=3D optionally names a directory inside the image where the carrier mount is moved instead. This is the file-backed analogue of what dm-mod.create=3D (CONFIG_DM_INIT) already does for device-mapper targets: moving a fixed, declarative bit of early-boot setup from initramfs userspace into the kernel so that small and verified systems can boot with no initramfs at all. Assisted-by: opencode:claude-fable-5 Signed-off-by: Eric Curtin --- .../admin-guide/kernel-parameters.txt | 29 ++++ init/do_mounts.c | 158 ++++++++++++++++-- 2 files changed, 174 insertions(+), 13 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentatio= n/admin-guide/kernel-parameters.txt index b5493a7f8..5dbd56098 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -6699,6 +6699,35 @@ Kernel parameters =20 rootfstype=3D [KNL] Set root filesystem type =20 + rootimage=3D [KNL] Mount the actual root filesystem from this + filesystem image file (absolute path) located on the + filesystem specified by root=3D, instead of using that + filesystem as the root directly. This allows + image-based systems that keep their (typically + read-only, e.g. erofs) root filesystem images as + plain files on a carrier filesystem to boot without + an initramfs. The image is mounted read-only via + the filesystem's file-backed mount support (no loop + device is set up); "ro", "rw" and rootflags=3D keep + applying to the carrier filesystem. Once the image + is mounted, the carrier mount is detached (but kept + busy by the image mount itself) unless + rootimagesrcdir=3D is also given. + + rootimageflags=3D [KNL] Set root image filesystem mount option string, + used together with rootimage=3D. + + rootimagefstype=3D [KNL] Set root image filesystem type, used together + with rootimage=3D. Default is to try all filesystems + known to the kernel that can mount a block device or + an image file. + + rootimagesrcdir=3D [KNL] Directory (absolute path) inside the root + image where the mount of the carrier filesystem + holding the image (i.e. the filesystem specified by + root=3D) is moved to, instead of detaching it. Used + together with rootimage=3D. + rootwait [KNL] Wait (indefinitely) for root device to show up. Useful for devices that are detected asynchronously (e.g. USB and MMC devices). diff --git a/init/do_mounts.c b/init/do_mounts.c index 95e0b3a0f..1b96ef30b 100644 --- a/init/do_mounts.c +++ b/init/do_mounts.c @@ -122,13 +122,51 @@ __setup("rootflags=3D", root_data_setup); __setup("rootfstype=3D", fs_names_setup); __setup("rootdelay=3D", root_delay_setup); =20 +/* + * rootimage=3D mounts the actual root filesystem from an image file locat= ed + * on the filesystem specified by root=3D instead of using that filesystem + * as the root directly. + */ +static char * __initdata root_image; +static int __init root_image_setup(char *str) +{ + root_image =3D str; + return 1; +} + +static char * __initdata root_image_fs_names; +static int __init root_image_fs_names_setup(char *str) +{ + root_image_fs_names =3D str; + return 1; +} + +static char * __initdata root_image_mount_data; +static int __init root_image_data_setup(char *str) +{ + root_image_mount_data =3D str; + return 1; +} + +static char * __initdata root_image_srcdir; +static int __init root_image_srcdir_setup(char *str) +{ + root_image_srcdir =3D str; + return 1; +} + +__setup("rootimage=3D", root_image_setup); +__setup("rootimagefstype=3D", root_image_fs_names_setup); +__setup("rootimageflags=3D", root_image_data_setup); +__setup("rootimagesrcdir=3D", root_image_srcdir_setup); + /* This can return zero length strings. Caller should check */ -static int __init split_fs_names(char *page, size_t size) +static int __init split_fs_names(char *page, size_t size, char *names) { int count =3D 1; char *p =3D page; =20 - strscpy(p, root_fs_names, size); + strscpy(p, names, size); while (*p++) { if (p[-1] =3D=3D ',') { p[-1] =3D '\0'; @@ -139,8 +177,9 @@ static int __init split_fs_names(char *page, size_t siz= e) return count; } =20 -static int __init do_mount_root(const char *name, const char *fs, - const int flags, const void *data) +static int __init do_mount_root(const char *name, const char *dir, + const char *fs, const int flags, + const void *data) { struct super_block *s; char *data_page =3D NULL; @@ -154,11 +193,11 @@ static int __init do_mount_root(const char *name, con= st char *fs, strscpy_pad(data_page, data, PAGE_SIZE); } =20 - ret =3D init_mount(name, "/root", fs, flags, data_page); + ret =3D init_mount(name, dir, fs, flags, data_page); if (ret) goto out; =20 - init_chdir("/root"); + init_chdir(dir); s =3D current->fs->pwd.dentry->d_sb; ROOT_DEV =3D s->s_dev; printk(KERN_INFO @@ -185,7 +224,7 @@ void __init mount_root_generic(char *name, char *pretty= _name, int flags) scnprintf(b, BDEVNAME_SIZE, "unknown-block(%u,%u)", MAJOR(ROOT_DEV), MINOR(ROOT_DEV)); if (root_fs_names) - num_fs =3D split_fs_names(fs_names, PAGE_SIZE); + num_fs =3D split_fs_names(fs_names, PAGE_SIZE, root_fs_names); else num_fs =3D list_bdev_fs_names(fs_names, PAGE_SIZE); retry: @@ -194,7 +233,7 @@ void __init mount_root_generic(char *name, char *pretty= _name, int flags) =20 if (!*p) continue; - err =3D do_mount_root(name, p, flags, root_mount_data); + err =3D do_mount_root(name, "/root", p, flags, root_mount_data); switch (err) { case 0: goto out; @@ -266,7 +305,8 @@ static void __init mount_nfs_root(void) */ timeout =3D NFSROOT_TIMEOUT_MIN; for (try =3D 1; ; try++) { - if (!do_mount_root(root_dev, "nfs", root_mountflags, root_data)) + if (!do_mount_root(root_dev, "/root", "nfs", root_mountflags, + root_data)) return; if (try > NFSROOT_RETRY_MAX) break; @@ -303,7 +343,7 @@ static void __init mount_cifs_root(void) =20 timeout =3D CIFSROOT_TIMEOUT_MIN; for (try =3D 1; ; try++) { - if (!do_mount_root(root_dev, "cifs", root_mountflags, + if (!do_mount_root(root_dev, "/root", "cifs", root_mountflags, root_data)) return; if (try > CIFSROOT_RETRY_MAX) @@ -345,7 +385,7 @@ static int __init mount_nodev_root(char *root_device_na= me) fs_names =3D kmalloc(PAGE_SIZE, GFP_KERNEL); if (!fs_names) return -EINVAL; - num_fs =3D split_fs_names(fs_names, PAGE_SIZE); + num_fs =3D split_fs_names(fs_names, PAGE_SIZE, root_fs_names); =20 for (i =3D 0, fstype =3D fs_names; i < num_fs; i++, fstype +=3D strlen(fstype) + 1) { @@ -353,8 +393,8 @@ static int __init mount_nodev_root(char *root_device_na= me) continue; if (!fs_is_nodev(fstype)) continue; - err =3D do_mount_root(root_device_name, fstype, root_mountflags, - root_mount_data); + err =3D do_mount_root(root_device_name, "/root", fstype, + root_mountflags, root_mount_data); if (!err) break; } @@ -402,6 +442,96 @@ void __init mount_root(char *root_device_name) } } =20 +/* + * Mount the actual root filesystem from the image file rootimage=3D on the + * filesystem that was just mounted from root=3D (the "carrier"), so that + * image-based systems can boot without an initramfs. + * + * Called with the carrier mounted at /root and the cwd there. The image + * is always mounted read-only; "ro"/"rw"/rootflags=3D keep applying to the + * carrier. On success the cwd is the image's root, ready for the pivot + * in prepare_namespace(). The carrier mount is moved to rootimagesrcdir= =3D + * inside the image if set, and detached otherwise; either way the image + * mount keeps the carrier superblock pinned. + */ +static void __init mount_root_image(void) +{ + unsigned long flags =3D MS_RDONLY | MS_SILENT; + char *path, *fs_names, *p; + struct file *file; + int num_fs, i, err; + + if (root_image[0] !=3D '/') + panic("VFS: rootimage=3D must be an absolute path"); + + path =3D kmalloc(PATH_MAX, GFP_KERNEL); + fs_names =3D kmalloc(PAGE_SIZE, GFP_KERNEL); + if (!path || !fs_names) + panic("VFS: unable to mount root image: not enough memory"); + + if (snprintf(path, PATH_MAX, "/root%s", root_image) >=3D PATH_MAX) + panic("VFS: rootimage=3D path too long"); + + /* + * Nothing that could modify the carrier runs yet, so the image + * cannot change between this open and the mount below. + */ + file =3D filp_open(path, O_RDONLY | O_LARGEFILE, 0); + if (IS_ERR(file)) + panic("VFS: unable to open root image %s: error %ld", + root_image, PTR_ERR(file)); + + err =3D init_mkdir("/image", 0700); + if (err < 0 && err !=3D -EEXIST) + panic("VFS: unable to create /image: error %d", err); + + if (root_image_fs_names) + num_fs =3D split_fs_names(fs_names, PAGE_SIZE, + root_image_fs_names); + else + num_fs =3D list_bdev_fs_names(fs_names, PAGE_SIZE); + + for (i =3D 0, p =3D fs_names; i < num_fs; i++, p +=3D strlen(p) + 1) { + if (!*p) + continue; + err =3D do_mount_root(path, "/image", p, flags, + root_image_mount_data); + switch (err) { + case 0: + goto mounted; + case -EACCES: + case -EINVAL: + case -ENOTBLK: + continue; + } + panic("VFS: unable to mount root image %s: error %d", + root_image, err); + } + panic("VFS: no filesystem could mount root image %s", root_image); + +mounted: + fput(file); + + if (root_image_srcdir) { + if (root_image_srcdir[0] !=3D '/') + panic("VFS: rootimagesrcdir=3D must be an absolute path"); + if (snprintf(path, PATH_MAX, ".%s", root_image_srcdir) >=3D + PATH_MAX) + panic("VFS: rootimagesrcdir=3D path too long"); + err =3D init_mount("/root", path, NULL, MS_MOVE, NULL); + if (err) + pr_err("VFS: failed to move the root image's carrier filesystem to %s: = error %d\n", + root_image_srcdir, err); + } else { + err =3D -EINVAL; + } + if (err) + init_umount("/root", MNT_DETACH); + + kfree(path); + kfree(fs_names); +} + /* wait for any asynchronous scanning to complete */ static void __init wait_for_root(char *root_device_name) { @@ -481,6 +611,8 @@ void __init prepare_namespace(void) if (root_wait) wait_for_root(saved_root_name); mount_root(saved_root_name); + if (root_image) + mount_root_image(); devtmpfs_mount(); =20 if (init_pivot_root(".", ".")) { --=20 2.43.0 From nobody Sat Jul 25 04:17:19 2026 Received: from mail-wm1-f41.google.com (mail-wm1-f41.google.com [209.85.128.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74540360ED0 for ; Sat, 18 Jul 2026 19:15:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784402161; cv=none; b=TieowxVZVeU2F8A8oI+IEVZt+OoL4HOJQSWxFq0P5KVOxF6Kt7px5CMrJdYMhQvDx/St1jBEfS2RETi16jvhrZq7CaLl22OYemKgJZH/jwrzxHvjEOy65WvZ4PNbMG+ds/WSuArSiJDKfevf8ycvAIbyJcuNQ93D3+q6VkmKxpM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784402161; c=relaxed/simple; bh=JaIyalHpyLApnPkNPXoy+faaGxYp0kiUgkaXuqNN+a8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kI0TzRyaLF/cfa7AFRaiHR0SWPUygdlNm/vYTorTO69fV4eGo6tQcbmWK1db6ErBIPpR07pNBRZMtfHGFQTIj3MWJaxZSNgU8d6dr26bLgVSjGRgLhh4H6FNfmMuBAvZiROZbvstYN55qNqi53uxU9d6S5/i90zYWP0ns77IW5E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jQS5coay; arc=none smtp.client-ip=209.85.128.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jQS5coay" Received: by mail-wm1-f41.google.com with SMTP id 5b1f17b1804b1-493f6de72faso17910985e9.0 for ; Sat, 18 Jul 2026 12:15:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784402158; x=1785006958; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=n/RuzzR56AKowWPeBaHkPa2me5UyA+I6b2NoNXGG3tc=; b=jQS5coaylvMHpThtcAzmeg8WaIdOmTmVY+x9jWSw9pcTbiz4O90Yi65DXiU8GNKXQV +JrkUat+l5x7hMF7xdd9I7UB7iPYo2cXi41bRrW6fVM6MjJQdHA95H14fKJ1u3kqz9/k 415Ou+UdgXp9nByzJvn2nawDVLvpOQ/w8shuZD3wJEwBGM0t81pfpgI4IhpnpfgHkFpm hT2jdlxtPi+SFUwpe5o/Ax1ycItQGORx+upKo28n1u5VOTEiIwGAUvnPt2lXwqbwcRIm y03I+YBm//LsWtcuyExi+eCF+vP56e6NfvLEuYolrodRrxmFrEV8mH1P/a3DChoAbi84 MIGw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784402158; x=1785006958; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=n/RuzzR56AKowWPeBaHkPa2me5UyA+I6b2NoNXGG3tc=; b=YZf3WQmDBPv2SFDd00B7floIAQ2VR4UxKr+4tAmmbDoVNh5f785w/RxN49yWFRPoXm hXDOTSy5zD3UGf3pa/LMIkoUmrAauvtRZzJ384IVMHdkSABw8jTmONSZCbcYjpC+T1uu UW2vzujDqPX6PKufubWa+0qb/+WdVcUV6imxcLPuBzQp+qw4Ph0cU5kPmtwvHsow+qaI n1ruJPsyb8PypYZ/y1cLYQYPKDuSn0yHMu2/NnvT0VcEwedYLzMx3YEurSzyNX/0w62E G5qa7gaAHMI6XNh6LvCUi8QNl/Ma4WNzueS75Zx7OB6MtLTFPX2M7thlawzwroOoYoSP 1bjw== X-Forwarded-Encrypted: i=1; AHgh+RpBQARGGuMNXOm+WicxPw3AyCu+0odjgacpb/YCLJlmdWQMmze56BePv9yEP2kvNgJqnZdp0DJ3vUogC3A=@vger.kernel.org X-Gm-Message-State: AOJu0YzJpwiq7XoF4yMjR+d3sGvOpoLOnfUwbR9T18T7ZOMkWS362feb fky7cMLidIxyAw08osiyz0zSpqELbn9II7Tl9Z5uq4F1pG/3JrHTBuSe X-Gm-Gg: AfdE7cnJGmDkWHrupewPTEp1jGQaVF4jLb//Hb6gJYoEwF7HiNTPvgdwRKFe7u2VPmx /G7q64y8UPNMYFoRJ4ndGraAfsouxlgGxolGe5sxayOZKvE2Lt3xz65f8kTOigorrEaUybPL+MO JjgP1JjmgEepXDNtJjzKwuKvXCAUcAB7U+uIKa5OQs6fjiDYy73lGj4VGJwvotP7hib9Yxd0t04 fG79b20SIhm5woJWIcUfjyjp1TS3Hj+ywvtsbsZ8QaO88kCzGtIgEz2r2Wf+qOE4sggH4bCk2a1 DUdeD8+x9NjflPCmTRRAmny0nQP4zf2aJLPXp0de54+s6RLNvOBVsgjOeggyO/v7I8uZxhf97Or 64e+6NQgB2bStzXPX3wqrKPF1qr2uF/90xbVjzLPfxSFUk7J5dft68ukFu/PQCYGm8pdYVpJRti m3RUo/vA== X-Received: by 2002:a05:600c:4e8b:b0:495:501a:fcf8 with SMTP id 5b1f17b1804b1-495501afdc8mr41543145e9.9.1784402157300; Sat, 18 Jul 2026 12:15:57 -0700 (PDT) Received: from spark.Home ([2001:8a0:7280:4000:c2b:a5ab:a31c:e0a5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2e8529sm140037005e9.11.2026.07.18.12.15.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 12:15:56 -0700 (PDT) From: Eric Curtin To: Alexander Viro , Christian Brauner Cc: Jan Kara , Jonathan Corbet , Shuah Khan , Eric Biggers , "Theodore Y . Ts'o" , Gao Xiang , Chao Yu , fsverity@lists.linux.dev, linux-erofs@lists.ozlabs.org, linux-fsdevel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, Eric Curtin Subject: [RFC PATCH 2/2] init: support pinning the root image's fsverity digest Date: Sat, 18 Jul 2026 20:15:51 +0100 Message-ID: <20260718191551.1703670-3-ericcurtin17@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718191551.1703670-1-ericcurtin17@gmail.com> References: <20260718191551.1703670-1-ericcurtin17@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When the root filesystem is mounted from an image file with rootimage=3D, the carrier filesystem holding the image is typically writable and therefore untrusted. Systems that seal their root images with fsverity currently need an initramfs for the sole purpose of checking that the image carries the expected fsverity digest before mounting it. Add rootimageverity=3D:, which requires the rootimage=3D file to have fsverity enabled with exactly this file digest and fails the boot otherwise, using the same fsverity_get_digest() interface that IMA and overlayfs already use for digest pinning. Combined with a trusted kernel command line (e.g. a signed unified kernel image, or a TPM-measured bootloader configuration), this extends the chain of trust to every byte of the root filesystem without any userspace boot stage: the digest pins the image's Merkle tree, and fsverity keeps verifying all data read from the image against it at runtime, so post-boot tampering with the carrier filesystem is detected as well. It is the file-backed counterpart of setting up a dm-verity target for a partition-backed root via dm-mod.create=3D. Verification is done on the file the kernel is about to mount: it is opened before mounting (which also loads the fsverity information) and kept open across the mount, and no userspace exists yet that could race a replacement in between. Assisted-by: opencode:claude-fable-5 Signed-off-by: Eric Curtin --- .../admin-guide/kernel-parameters.txt | 13 ++++ init/do_mounts.c | 63 +++++++++++++++++++ 2 files changed, 76 insertions(+) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentatio= n/admin-guide/kernel-parameters.txt index 5dbd56098..105ebb171 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -6728,6 +6728,19 @@ Kernel parameters root=3D) is moved to, instead of detaching it. Used together with rootimage=3D. =20 + rootimageverity=3D [KNL] Require the root image specified by + rootimage=3D to have fsverity enabled with this file + digest, given as :, + e.g. sha256:dd1b3fa9... The boot is aborted if the + image carries no or a different fsverity digest. + Because fsverity keeps verifying data read from the + image against its Merkle tree at runtime, a trusted + (e.g. signed or TPM-measured) kernel command line + extends the chain of trust to the complete root + filesystem contents without an initramfs. Requires + CONFIG_FS_VERITY and a carrier filesystem with + fsverity support. + rootwait [KNL] Wait (indefinitely) for root device to show up. Useful for devices that are detected asynchronously (e.g. USB and MMC devices). diff --git a/init/do_mounts.c b/init/do_mounts.c index 1b96ef30b..4eb792b27 100644 --- a/init/do_mounts.c +++ b/init/do_mounts.c @@ -24,6 +24,9 @@ #include #include #include +#include +#include +#include #include =20 #include "do_mounts.h" @@ -155,10 +158,18 @@ static int __init root_image_srcdir_setup(char *str) return 1; } =20 +static char * __initdata root_image_verity; +static int __init root_image_verity_setup(char *str) +{ + root_image_verity =3D str; + return 1; +} + __setup("rootimage=3D", root_image_setup); __setup("rootimagefstype=3D", root_image_fs_names_setup); __setup("rootimageflags=3D", root_image_data_setup); __setup("rootimagesrcdir=3D", root_image_srcdir_setup); +__setup("rootimageverity=3D", root_image_verity_setup); =20 /* This can return zero length strings. Caller should check */ static int __init split_fs_names(char *page, size_t size, char *names) @@ -442,6 +453,55 @@ void __init mount_root(char *root_device_name) } } =20 +#ifdef CONFIG_FS_VERITY +/* + * Require the root image to carry the fsverity file digest given by + * rootimageverity=3D:. @file must have been + * opened so that its fsverity information is loaded. Any deviation + * fails the boot: with a trusted command line this pins the complete + * image contents, which fsverity keeps verifying against the image's + * Merkle tree as they are read. + */ +static void __init verify_root_image(struct file *file) +{ + u8 want[FS_VERITY_MAX_DIGEST_SIZE], got[FS_VERITY_MAX_DIGEST_SIZE]; + enum hash_algo want_algo, got_algo; + int want_size, got_size, i; + char *hex; + + hex =3D strchr(root_image_verity, ':'); + if (!hex) + panic("VFS: rootimageverity=3D expects :"); + *hex++ =3D '\0'; + i =3D match_string(hash_algo_name, HASH_ALGO__LAST, root_image_verity); + if (i < 0) + panic("VFS: rootimageverity=3D: unknown hash algorithm \"%s\"", + root_image_verity); + want_algo =3D i; + want_size =3D hash_digest_size[want_algo]; + if (strlen(hex) !=3D 2 * want_size || hex2bin(want, hex, want_size)) + panic("VFS: rootimageverity=3D: expected %d-byte hex digest", + want_size); + + got_size =3D fsverity_get_digest(file_inode(file), got, NULL, &got_algo); + if (!got_size) + panic("VFS: root image does not have fsverity enabled"); + if (got_algo !=3D want_algo || got_size !=3D want_size || + memcmp(want, got, want_size)) + panic("VFS: root image fsverity digest mismatch: expected %s:%*phN, got = %s:%*phN", + hash_algo_name[want_algo], want_size, want, + hash_algo_name[got_algo], got_size, got); + + pr_info("VFS: verified root image fsverity digest %s:%*phN\n", + hash_algo_name[want_algo], want_size, want); +} +#else /* !CONFIG_FS_VERITY */ +static void __init verify_root_image(struct file *file) +{ + panic("VFS: rootimageverity=3D requires CONFIG_FS_VERITY"); +} +#endif /* !CONFIG_FS_VERITY */ + /* * Mount the actual root filesystem from the image file rootimage=3D on the * filesystem that was just mounted from root=3D (the "carrier"), so that @@ -481,6 +541,9 @@ static void __init mount_root_image(void) panic("VFS: unable to open root image %s: error %ld", root_image, PTR_ERR(file)); =20 + if (root_image_verity) + verify_root_image(file); + err =3D init_mkdir("/image", 0700); if (err < 0 && err !=3D -EEXIST) panic("VFS: unable to create /image: error %d", err); --=20 2.43.0