From nobody Sat Jul 25 23:03:44 2026 Received: from mail-pj1-f53.google.com (mail-pj1-f53.google.com [209.85.216.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5392D3126C2 for ; Sun, 12 Jul 2026 04:08:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829334; cv=none; b=XNfXKIV5j58sSFYaH6HVM+dc4G69HD8DeKSEzBt6+5eZ28XCWVpKrXARvlWIkbDyBTsJDb7askegn+x3ay/oQTFJtlOXlHIrhZtzsGi+78/N09s/Se4gHHxDPxB3mZbH7HJZHaKdCJBicAa6Y2bwTXNneuQqgKzD9SsrjNT3dbI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829334; c=relaxed/simple; bh=1q6mbB/bkD8j4HZw6McdfhEuY8vqwQWBAE158sj4Y8w=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=JZbJex8loTZJnxmxDMwUgIUhb8hdnBCPW/x5oJBZnVjNy8HPKzZHE92muycnWE7xzjGkSrkeeOeHY3NpuJdmC9Kkg9P0VH9NV8GIpMcrHcQxTs69SurZ2r4SkEXiSzMHZRmfcbUlEEZamM5COhxcjHTGiHdv0HqRBpuiKcuHufE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=cbokcfpl; arc=none smtp.client-ip=209.85.216.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="cbokcfpl" Received: by mail-pj1-f53.google.com with SMTP id 98e67ed59e1d1-383cb94f742so1807010a91.3 for ; Sat, 11 Jul 2026 21:08:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783829332; x=1784434132; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=1vSs3PpxD7fzzGCyEVeEQofuGfZwATqltGOSkAKOV+Y=; b=cbokcfpllFKwZICTa8uyEtG3gK/7FawCR3BdoMCN2eYnV/UaEACwX8YYUUTY937kqu XrhvbfA812Px1NShMdXdMLykuPbxRC5P+2PoPv3ARp/vRKOFfCsUPO71qaLla4K85qfO ztngsdEEwdwI/6rv+0V2rsNJiSerwBjniugxklJfOUOAXlFU04Zch8ec5bzDlUImavL3 Id7FFDrNo72KepusQwMZi7n2BOfd/7MWmiPrCIH09//1ijwcKdTqEj093c8nw4Rw8EKl zjzAQu36pZUYE9gZQ+3xDntkT/jY0pnlf+I4CD2rhkkarpitFXpDnOZcUQt1p+r15IRj wyqQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783829332; x=1784434132; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=1vSs3PpxD7fzzGCyEVeEQofuGfZwATqltGOSkAKOV+Y=; b=l7kDV4VdkPKRQ6LeTzLp3UvdcwoWfAX7WPT1MySFaik2hjwL9C6KJcaHbVSTuVuYUn iLNMvoy1yqL4oqqsa/hHSCLzig1Eij6cNYuGmW4834KeaOqGgfuASRg8UtKwD4IIjU8E J+S8qCDn3BtkVUtPnxtOKMhk0Mp+v7N2Auf+F97nF8XBHC968rMQ6+4vtDUz3ZFXdKSY 4WxMacT+Ojo/3FeeD5ODVUiuyXI3LLOxeMh9if0H5V1l82Hwpswk/l8TJPf3OwYo5sZc X+tCTW3AA4FaJL/yPlaXRWnd0hkWo6nGTn4wqJVnYJ1bi631WgbQtfevz8xNDfN/4phw VF8g== X-Forwarded-Encrypted: i=1; AHgh+RrS1ZugDhrRyRbDelyAvCwRyn/E3J8yejbKG6C/fSw9tlrx6jLHl7BCxX0ADYf1xHvYRoq/9zJKD1SdVlU=@vger.kernel.org X-Gm-Message-State: AOJu0Ywi5PwVLTFlyupp981g/b3LUCXdzmsZkRsvY6fnffnTsIAwDUhy +9OyaaS6+2yhE7U+H7rzs0VOrXF0Uy4FASbMXWH1CDF7mtnhpHmh+aYN X-Gm-Gg: AfdE7cljEEyPTVU21vndX91Zv5BuE1KmXEi00PGNckbyqjaidPybOSjnwvhYA/z2Ht/ 5d36pNDDWE2uFywuUdArGDQgJknuMAWGyCf67kXDuKpI6GQj0T95hUY+ZsFXLfK/3hZ/GI3TyQC HZeb8OZTANzbme6lRK6F+phYB46JMvhvF8kh+ulaJmp6up9yCtAwsGgqSnIDFAyuXQ4QE2Qd+5w 4S2IQjiC63KofQQPwtyKtlQSOhCncbSi9q1any09/2gkHZF6bON8n6cWvQg8jN8uD3hDgVoDn7k iZ1f7juP0wvGEpNy53oe6UInPmUzjaU76n4srvv8G/mRD6mxbOYyRIkmYz+KkWlvvlq2entRR/4 YQ6gljvcaMYxBOXs50j1FrlEA5LhE6dZH6WKtn9ROWDeZvYRPSoQtNWVBNgkVfmvD35mz6TkVBy vvUgmVukoYiwoo0w/zsJ02A/tJ7n8= X-Received: by 2002:a17:90a:e7c4:b0:38d:eaec:4396 with SMTP id 98e67ed59e1d1-38deaec5594mr631416a91.11.1783829331759; Sat, 11 Jul 2026 21:08:51 -0700 (PDT) Received: from [127.0.0.2] ([98.35.8.117]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31174839f89sm56808928eec.10.2026.07.11.21.08.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 11 Jul 2026 21:08:51 -0700 (PDT) From: Farid Zakaria Date: Sat, 11 Jul 2026 21:08:14 -0700 Subject: [PATCH v2 1/5] exec: stash a bpf-selected interpreter in struct linux_binprm Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260711-binfmt-misc-bpf-v2-v2-1-d6591ceaf207@gmail.com> References: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> In-Reply-To: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> To: Christian Brauner , Alexei Starovoitov , Daniel Borkmann , Martin KaFai Lau , Shuah Khan Cc: Andrii Nakryiko , Kees Cook , Alexander Viro , Jan Kara , Jonathan Corbet , Jann Horn , John Ericson , linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Farid Zakaria X-Mailer: b4 0.14.3 From: Christian Brauner The upcoming bpf-backed binfmt_misc handlers select the interpreter for a binary programmatically at exec time. The selection runs before load_misc_binary() has copied the binary path from bprm->interp into the argument vector, so the selecting program cannot go through bprm_change_interp() directly without clobbering argv[1]. Stage the selected path in the bprm instead. The bprm is exclusively owned by the task doing the exec so no synchronization is needed. The consumer frees and clears the field once the exec attempt that set it is finished; free_bprm() covers all error paths. Signed-off-by: Christian Brauner (Amutable) --- fs/exec.c | 1 + include/linux/binfmts.h | 1 + 2 files changed, 2 insertions(+) diff --git a/fs/exec.c b/fs/exec.c index b92fe7db1..7c9e28f54 100644 --- a/fs/exec.c +++ b/fs/exec.c @@ -1418,6 +1418,7 @@ static void free_bprm(struct linux_binprm *bprm) /* If a binfmt changed the interp, free it. */ if (bprm->interp !=3D bprm->filename) kfree(bprm->interp); + kfree(bprm->bpf_interp); kfree(bprm->fdpath); kfree(bprm); } diff --git a/include/linux/binfmts.h b/include/linux/binfmts.h index 7e7333b7b..1dbec6905 100644 --- a/include/linux/binfmts.h +++ b/include/linux/binfmts.h @@ -65,6 +65,7 @@ struct linux_binprm { of the time same as filename, but could be different for binfmt_{misc,script} */ const char *fdpath; /* generated filename for execveat */ + const char *bpf_interp; /* interpreter selected by a bpf handler */ unsigned interp_flags; int execfd; /* File descriptor of the executable */ unsigned long exec; --=20 2.51.2 From nobody Sat Jul 25 23:03:44 2026 Received: from mail-pj1-f41.google.com (mail-pj1-f41.google.com [209.85.216.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A935436C581 for ; Sun, 12 Jul 2026 04:08:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829335; cv=none; b=V67uLA8KV7mRwq1GU2Q9ZaCm1OnKqZTzHWl6OmQ5EgBRFsiLxEvsgKHruV9svbDIdcx3GhbaD6pjf8vD1DjFfrd2dUCD0SjVtmxBhmR5nd/xlL9zmnO7mwiagFm6xl/2XA05OcBqLvyD2QWbBLrF9hpbUNIRnm6m036+v1vbRpY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829335; c=relaxed/simple; bh=D/zjIHpqzbNt6li9XSQgyI/JUXAjkj9YKqZasbkGUDY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=EG4Tj/PRAZu3iUwhoe+RejkQoM5w1ozyp7qGDnElAUY7cZO7berRGe8Q+ULLxfxh1sH3G8SWKRJ5QV7hcckSRUjGLpHKoofFuKS7ZKoSA25frPY+XROCyfaMgaljA9VFaJVuMcsMsq0DlmlF86ivnK0rAZECEQMESYEtJvqNPyQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=N8Hd8Ghd; arc=none smtp.client-ip=209.85.216.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="N8Hd8Ghd" Received: by mail-pj1-f41.google.com with SMTP id 98e67ed59e1d1-38511175ad3so1793594a91.2 for ; Sat, 11 Jul 2026 21:08:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783829333; x=1784434133; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nSAL3gg/+UGNm6w8OO6Kj1lCK77Tqi4uiJZac+rrH54=; b=N8Hd8GhdoQzmqrT9fQy6JfgSgNr5c2UfDa8DASbmsE7mG2P8n92CWbgZ3d8VPeT1JY e+ATqzNCQeRLG3J1nSIZ1M4EjCZaGfqedEaecDmt8mrC8eKVuSqNuTSu+BhvM2N4eNqx CefMEhf7xYhU5dUlie4o8NVEZC2B8BoJRCS0GB9BnLzr16hDVCsOujz+VOjvAWdagRwg RvPZdwTjQkwOXMvrWmBuoZif4UOHz1OQjGQMAtvqnnFJdDzW1ITMQlxtKr+Tz4SYVJ4+ UVC7Jfjgyt7ymVbX/ogFpEsBl3dug+tALFT8aQ3rjuKsApoHSdOx6P9JmdjAE6xpMNXL JpQA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783829333; x=1784434133; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=nSAL3gg/+UGNm6w8OO6Kj1lCK77Tqi4uiJZac+rrH54=; b=CDK5IW4yNIHnF46Q2iYscMPPKZ1e90ddIYinRgyHKXcXZgdK4UhQgBfXfpkktIi6Pl zy1oaBB6/wnH6+eJNBpXFvsacFp1nOP8eaV30+vYrIaGIZjPayRP1hKapQXNR7uGvnqM x1B6dg7T0O3/0nz8RgKORvNUluao3wCHHolIUaC2Hb+sBxAct5kB2XeL41zD9+w1DDpl 6as8mF3MpKYLpvQO9ItBCuM2615JF5K0mUhDxNqxeCjrjFjDBhvBj0Ar4p7RI6F8K73N rDZK6zz/Uxq1M8H1JdBqoU6cKyyoHnpuasddi0jvj1VGvrQzmcWmcE/yD5L31qpczlFK UAQA== X-Forwarded-Encrypted: i=1; AHgh+Rpv2i7u6m6pqx04El1hOAPAnNr6jZsuQ/mIYezvlwyFqzhT9rErIrRjEX9WpO41v/iMBi//t8t5APSzlxY=@vger.kernel.org X-Gm-Message-State: AOJu0Yxt37r7X00mbHGQd+s+lRXByHMECP9xIUQ2q6jl4FQ0vtTWbj/U HlNkMcbG1Qwz7EHRj3M4dI5pj71EtbP5ooTY4FX4qiDZ3cBG2hc1FxdS X-Gm-Gg: AfdE7cmCBW1ZUH6Oaj40DMO3DAzr2uI2ujeNfYsQOV5PQ9/nndJUP8IAx51BxynYiNq EE0DDcDq0nkbPBwUFk/sNUo5jj7iQab8ezrZQWoi7xsDA7zAsq31R5DpNBAf72TSCGR6lvio/N7 cL+S2WiBgOjuwfDIfOo9MPzV6lC+DLSwjMyb78W/tBs+7RH77vAdErqTweXNcGrpVM9WuVDWQSF v8+pjqFw86ehfRH+RufbIhDAG3saSXCtXJJrOuStl4qvVQpwwMg3hMnOhs4HDtq5+265rbgZZSv BTAwt7+bYXgbrNlqRdLMzZX/nrY0ITAq9ZJ2A3GbEmf8wH+jTKDUp+q8GsVsw5UUblbe93ERw0z l2FLBPlooya0itGXCpBEXvjvpGVjcIoAODtZA41K3r3DZfXatCwcWhT8w3/rvGLqs6KxcBByEKw Kd8zb93jH4eWBBjfTG X-Received: by 2002:a17:90b:54c4:b0:38d:b36e:982d with SMTP id 98e67ed59e1d1-38dc777d7ddmr4186175a91.29.1783829333037; Sat, 11 Jul 2026 21:08:53 -0700 (PDT) Received: from [127.0.0.2] ([98.35.8.117]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31174839f89sm56808928eec.10.2026.07.11.21.08.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 11 Jul 2026 21:08:52 -0700 (PDT) From: Farid Zakaria Date: Sat, 11 Jul 2026 21:08:15 -0700 Subject: [PATCH v2 2/5] binfmt_misc: add binfmt_misc_ops bpf struct_ops Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260711-binfmt-misc-bpf-v2-v2-2-d6591ceaf207@gmail.com> References: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> In-Reply-To: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> To: Christian Brauner , Alexei Starovoitov , Daniel Borkmann , Martin KaFai Lau , Shuah Khan Cc: Andrii Nakryiko , Kees Cook , Alexander Viro , Jan Kara , Jonathan Corbet , Jann Horn , John Ericson , linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Farid Zakaria X-Mailer: b4 0.14.3 From: Christian Brauner Add the bpf plumbing for binary type handlers whose matching and interpreter selection are implemented by a bpf program instead of a fixed magic/extension and a fixed interpreter string recorded at registration time. This serves relocatable binary formats where the interpreter must be computed per binary, e.g. relative to the location of the binary itself, as discussed for hermetic Nix-style executables. A handler is an instance of the new binfmt_misc_ops struct_ops with a single op: int (*load)(struct linux_binprm *bprm); and a name that binfmt_misc entries reference it by. struct_ops is the sanctioned mechanism for this kind of user-supplied policy callback: program types, attach types, and the uapi helper list are frozen, and every recently added subsystem hook (bpf qdisc, SMC handshake control, io_uring loop ops, sched_ext) is a struct_ops user. The op receives the bprm as a trusted BTF pointer, so a program can match on the header in bprm->buf, read arbitrary file content via bpf_dynptr_from_file() to parse e.g. ELF program headers, and inspect the binary's location. No dedicated program type, ctx blob, or uapi helper is needed. The load program is required to be sleepable. Reliable file reads at exec time fault in the file's pages; a non-sleepable program would be limited to whatever happens to be resident in the page cache and would fail sporadically on cold caches. This also constrains the caller: binfmt_misc must invoke the op from sleepable context, which a later patch takes care of. The interpreter is selected through the new kfunc: int bpf_binprm_set_interp(struct linux_binprm *bprm, const char *path, size_t path__sz); which enforces an absolute path shorter than PATH_MAX and stages a copy in the bprm. The bprm is exclusively owned by the task doing the exec, so no shared or per-CPU state is involved and nothing here can race. The kfunc is registered for struct_ops programs with a filter limiting it to binfmt_misc_ops programs. Registering an ops instance (updating the struct_ops map or attaching its link) publishes the handler under its name in a registry keyed by the registering task's user namespace. Lookups walk the user namespace hierarchy upwards, mirroring how binfmt_misc instances themselves are resolved in load_binfmt_misc(). Consumers take a reference on the ops via bpf_struct_ops_get() which pins the underlying map and programs, so an activated handler keeps working even if the map is deleted or the registering container goes away; deregistration only prevents new activations, exactly like unregistering a tcp congestion ops with live users. Link: https://lore.kernel.org/20260704211409.1978485-1-farid.m.zakaria@gmai= l.com Signed-off-by: Christian Brauner (Amutable) --- fs/Kconfig.binfmt | 14 +++ fs/Makefile | 1 + fs/binfmt_misc_bpf.c | 275 ++++++++++++++++++++++++++++++++++++++++= ++++ include/linux/binfmt_misc.h | 49 ++++++++ 4 files changed, 339 insertions(+) diff --git a/fs/Kconfig.binfmt b/fs/Kconfig.binfmt index 1949e25c7..daeac4889 100644 --- a/fs/Kconfig.binfmt +++ b/fs/Kconfig.binfmt @@ -168,6 +168,20 @@ config BINFMT_MISC you have use for it; the module is called binfmt_misc. If you don't know what to answer at this point, say Y. =20 +config BINFMT_MISC_BPF + bool "BPF-selected interpreters for misc binaries" + depends on BINFMT_MISC=3Dy + depends on BPF_SYSCALL && BPF_JIT && DEBUG_INFO_BTF + help + Allow binfmt_misc binary type handlers to be implemented as bpf + struct_ops programs. Instead of matching a fixed magic and + redirecting to a fixed interpreter recorded at registration time + such handlers match binaries programmatically and compute the + interpreter to use per binary, e.g. relative to the location of + the binary itself. + + If you don't know what to answer at this point, say N. + config COREDUMP bool "Enable core dump support" if EXPERT default y diff --git a/fs/Makefile b/fs/Makefile index 89a8a9d20..499c6670f 100644 --- a/fs/Makefile +++ b/fs/Makefile @@ -33,6 +33,7 @@ obj-$(CONFIG_FS_ENCRYPTION) +=3D crypto/ obj-$(CONFIG_FS_VERITY) +=3D verity/ obj-$(CONFIG_FILE_LOCKING) +=3D locks.o obj-$(CONFIG_BINFMT_MISC) +=3D binfmt_misc.o +obj-$(CONFIG_BINFMT_MISC_BPF) +=3D binfmt_misc_bpf.o obj-$(CONFIG_BINFMT_SCRIPT) +=3D binfmt_script.o obj-$(CONFIG_BINFMT_ELF) +=3D binfmt_elf.o obj-$(CONFIG_COMPAT_BINFMT_ELF) +=3D compat_binfmt_elf.o diff --git a/fs/binfmt_misc_bpf.c b/fs/binfmt_misc_bpf.c new file mode 100644 index 000000000..72da0964d --- /dev/null +++ b/fs/binfmt_misc_bpf.c @@ -0,0 +1,275 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * BPF-backed binary type handlers for binfmt_misc. + * + * A handler is a struct binfmt_misc_ops struct_ops map. Loading and + * registering it makes the handler available under its name in the user + * namespace it was registered in. A binfmt_misc 'B' entry activates it: + * + * echo ':entry:B:::::' > /register + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +struct bm_bpf_ops_reg { + struct list_head list; + const struct binfmt_misc_ops *ops; + struct bpf_link *link; + struct user_namespace *user_ns; +}; + +static DEFINE_MUTEX(bm_bpf_ops_lock); +static LIST_HEAD(bm_bpf_ops_list); + +static struct bpf_struct_ops bpf_binfmt_misc_ops; + +static struct bm_bpf_ops_reg *bm_bpf_ops_find(const struct user_namespace = *user_ns, + const char *name) +{ + struct bm_bpf_ops_reg *reg; + + lockdep_assert_held(&bm_bpf_ops_lock); + + list_for_each_entry(reg, &bm_bpf_ops_list, list) { + if (reg->user_ns =3D=3D user_ns && !strcmp(reg->ops->name, name)) + return reg; + } + return NULL; +} + +/** + * binfmt_misc_get_ops - look up a bpf binary type handler by name + * @user_ns: user namespace of the binfmt_misc instance + * @name: name the handler was registered under + * + * Search @user_ns and its ancestors for a handler named @name, mirroring + * the instance lookup in load_binfmt_misc(). The returned handler stays + * callable until binfmt_misc_put_ops() even if the backing struct_ops map + * is detached or deleted in the meantime. + * + * Return: the handler on success, NULL on failure + */ +const struct binfmt_misc_ops *binfmt_misc_get_ops(struct user_namespace *u= ser_ns, + const char *name) +{ + const struct user_namespace *ns; + struct bm_bpf_ops_reg *reg; + + guard(mutex)(&bm_bpf_ops_lock); + + for (ns =3D user_ns; ns; ns =3D ns->parent) { + reg =3D bm_bpf_ops_find(ns, name); + if (!reg) + continue; + if (!bpf_struct_ops_get(reg->ops)) + return NULL; + return reg->ops; + } + return NULL; +} + +void binfmt_misc_put_ops(const struct binfmt_misc_ops *ops) +{ + bpf_struct_ops_put(ops); +} + +bool bpf_prog_is_binfmt_misc_ops(const struct bpf_prog *prog) +{ + return prog->type =3D=3D BPF_PROG_TYPE_STRUCT_OPS && + prog->aux->st_ops =3D=3D &bpf_binfmt_misc_ops; +} + +__bpf_kfunc_start_defs(); + +/** + * bpf_binprm_set_interp - select the interpreter for the current exec + * @bprm: binary that is being executed + * @path: absolute path to the interpreter + * @path__sz: size of the @path buffer, including the terminating NUL + * + * To be called from the load program of a struct binfmt_misc_ops handler + * before returning a positive value. The path is opened with the + * credentials of the task doing the exec after the program returns. + * + * Return: 0 on success, a negative errno on failure + */ +__bpf_kfunc int bpf_binprm_set_interp(struct linux_binprm *bprm, + const char *path, size_t path__sz) +{ + size_t len; + char *interp; + + if (!path__sz) + return -EINVAL; + len =3D strnlen(path, path__sz); + if (len =3D=3D path__sz) + return -EINVAL; + if (path[0] !=3D '/') + return -EINVAL; + if (len >=3D PATH_MAX) + return -ENAMETOOLONG; + + interp =3D kmemdup_nul(path, len, GFP_KERNEL); + if (!interp) + return -ENOMEM; + + kfree(bprm->bpf_interp); + bprm->bpf_interp =3D interp; + return 0; +} + +__bpf_kfunc_end_defs(); + +BTF_KFUNCS_START(bm_bpf_kfunc_ids) +BTF_ID_FLAGS(func, bpf_binprm_set_interp, KF_SLEEPABLE) +BTF_KFUNCS_END(bm_bpf_kfunc_ids) + +static int bm_bpf_kfunc_filter(const struct bpf_prog *prog, u32 kfunc_id) +{ + if (!btf_id_set8_contains(&bm_bpf_kfunc_ids, kfunc_id)) + return 0; + if (bpf_prog_is_binfmt_misc_ops(prog)) + return 0; + return -EACCES; +} + +static const struct btf_kfunc_id_set bm_bpf_kfunc_set =3D { + .owner =3D THIS_MODULE, + .set =3D &bm_bpf_kfunc_ids, + .filter =3D bm_bpf_kfunc_filter, +}; + +static int bm_bpf_ops__load(struct linux_binprm *bprm) +{ + return 0; +} + +static struct binfmt_misc_ops bm_bpf_ops_stubs =3D { + .load =3D bm_bpf_ops__load, +}; + +static int bm_bpf_init(struct btf *btf) +{ + return register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS, + &bm_bpf_kfunc_set); +} + +static int bm_bpf_check_member(const struct btf_type *t, + const struct btf_member *member, + const struct bpf_prog *prog) +{ + u32 moff =3D __btf_member_bit_offset(t, member) / 8; + + switch (moff) { + case offsetof(struct binfmt_misc_ops, load): + /* Reliable file reads at exec time require sleeping. */ + if (!prog->sleepable) + return -EINVAL; + break; + } + return 0; +} + +static int bm_bpf_init_member(const struct btf_type *t, + const struct btf_member *member, + void *kdata, const void *udata) +{ + const struct binfmt_misc_ops *uops =3D udata; + struct binfmt_misc_ops *ops =3D kdata; + u32 moff =3D __btf_member_bit_offset(t, member) / 8; + + switch (moff) { + case offsetof(struct binfmt_misc_ops, name): + if (bpf_obj_name_cpy(ops->name, uops->name, + sizeof(ops->name)) <=3D 0) + return -EINVAL; + return 1; + } + return 0; +} + +static int bm_bpf_validate(void *kdata) +{ + struct binfmt_misc_ops *ops =3D kdata; + + if (!ops->load) + return -EINVAL; + return 0; +} + +static int bm_bpf_reg(void *kdata, struct bpf_link *link) +{ + struct binfmt_misc_ops *ops =3D kdata; + struct bm_bpf_ops_reg *reg; + + reg =3D kzalloc_obj(*reg, GFP_KERNEL_ACCOUNT); + if (!reg) + return -ENOMEM; + + reg->ops =3D ops; + reg->link =3D link; + reg->user_ns =3D get_user_ns(current_user_ns()); + + guard(mutex)(&bm_bpf_ops_lock); + + if (bm_bpf_ops_find(reg->user_ns, ops->name)) { + put_user_ns(reg->user_ns); + kfree(reg); + return -EEXIST; + } + + list_add(®->list, &bm_bpf_ops_list); + return 0; +} + +static void bm_bpf_unreg(void *kdata, struct bpf_link *link) +{ + struct bm_bpf_ops_reg *reg; + + guard(mutex)(&bm_bpf_ops_lock); + + list_for_each_entry(reg, &bm_bpf_ops_list, list) { + if (reg->ops =3D=3D kdata && reg->link =3D=3D link) { + list_del(®->list); + put_user_ns(reg->user_ns); + kfree(reg); + return; + } + } +} + +static const struct bpf_verifier_ops bm_bpf_verifier_ops =3D { + .get_func_proto =3D bpf_base_func_proto, + .is_valid_access =3D bpf_tracing_btf_ctx_access, +}; + +static struct bpf_struct_ops bpf_binfmt_misc_ops =3D { + .verifier_ops =3D &bm_bpf_verifier_ops, + .init =3D bm_bpf_init, + .check_member =3D bm_bpf_check_member, + .init_member =3D bm_bpf_init_member, + .validate =3D bm_bpf_validate, + .reg =3D bm_bpf_reg, + .unreg =3D bm_bpf_unreg, + .cfi_stubs =3D &bm_bpf_ops_stubs, + .name =3D "binfmt_misc_ops", + .owner =3D THIS_MODULE, +}; + +static int __init bm_bpf_struct_ops_init(void) +{ + return register_bpf_struct_ops(&bpf_binfmt_misc_ops, binfmt_misc_ops); +} +late_initcall(bm_bpf_struct_ops_init); diff --git a/include/linux/binfmt_misc.h b/include/linux/binfmt_misc.h new file mode 100644 index 000000000..e1d26c430 --- /dev/null +++ b/include/linux/binfmt_misc.h @@ -0,0 +1,49 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _LINUX_BINFMT_MISC_H +#define _LINUX_BINFMT_MISC_H + +#include + +struct bpf_prog; +struct linux_binprm; +struct user_namespace; + +#define BINFMT_MISC_OPS_NAME_MAX 16 + +/** + * struct binfmt_misc_ops - bpf-backed binary type handler + * @load: match @bprm and select an interpreter via bpf_binprm_set_interp(= ); + * returns > 0 if the binary was handled, 0 to fall through to the + * handlers registered after this one, a negative errno to fail the + * exec; -ENOEXEC does not fail the exec but moves on to the + * remaining binary formats + * @name: name that 'B' entries reference the handler by + */ +struct binfmt_misc_ops { + int (*load)(struct linux_binprm *bprm); + char name[BINFMT_MISC_OPS_NAME_MAX]; +}; + +#ifdef CONFIG_BINFMT_MISC_BPF +const struct binfmt_misc_ops *binfmt_misc_get_ops(struct user_namespace *u= ser_ns, + const char *name); +void binfmt_misc_put_ops(const struct binfmt_misc_ops *ops); +bool bpf_prog_is_binfmt_misc_ops(const struct bpf_prog *prog); +#else +static inline const struct binfmt_misc_ops * +binfmt_misc_get_ops(struct user_namespace *user_ns, const char *name) +{ + return NULL; +} + +static inline void binfmt_misc_put_ops(const struct binfmt_misc_ops *ops) +{ +} + +static inline bool bpf_prog_is_binfmt_misc_ops(const struct bpf_prog *prog) +{ + return false; +} +#endif /* CONFIG_BINFMT_MISC_BPF */ + +#endif /* _LINUX_BINFMT_MISC_H */ --=20 2.51.2 From nobody Sat Jul 25 23:03:44 2026 Received: from mail-pj1-f49.google.com (mail-pj1-f49.google.com [209.85.216.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 52DE036DA18 for ; Sun, 12 Jul 2026 04:08:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829338; cv=none; b=rAdJb/wZ+xVdNukQXnuUyJLmPolScVoHJsyIHD2qj4Suyg5vhjijoZ4zoKfr1kMXJrz6PKLSdYqzo4XINjV1MVsAYUZ+nBVEqXg4DkNMaqaKhOmfZiO3n5WDWlYh8oDMxY6caL6FgFyh+g7wT43iS3CktW4hxQkEhbMvqGYDGeQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829338; c=relaxed/simple; bh=tem2K9TU/zJyLS+Wwme5cI666FVV7MGdzn0q1YhpdBo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=KH1Qt97u2UT//wX8QeGP1e8WUrKbHkr4TkX39K4sfdf8goQDU1I8140q6NQmDYu56RmPWyhd1vAb2SnW59THIkCOxbf2RSYPbnAsyVNp2zuktvYJLN/YYbc4uNedHgzrsKQxM6S3sIwqGeou6McyKihKYYoKNY6ZxQQz/Ypa9rI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=fc8m2Wf1; arc=none smtp.client-ip=209.85.216.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="fc8m2Wf1" Received: by mail-pj1-f49.google.com with SMTP id 98e67ed59e1d1-38175907a56so2493623a91.0 for ; Sat, 11 Jul 2026 21:08:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783829335; x=1784434135; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=aH3kKOMcwn56DSlHVteprYGWMctwEqxrrgXOQeiC+bw=; b=fc8m2Wf1j8reMWn1MsB4PVkDPbV+LncGso1FWNy2IMyiJpk6A8+aNGpV4s4ZNoggp7 yiL1p/fWVYUUQZQxcRHYNKw9pbl/K6QsXsVtAExZO+xcxWvgXrJCG4fXkdFrAT5+tjPv 34lJBNyuVWa6ji4/AtKrt0cS1GP/SsftARDEsGaXy1i9yQvBS2N6hwSeJX/F08Thx8lS uWPL1U2BsomqSLekliaYbtOjL4M5ihbzIyWxP8s08T+DgBOMLfP5mRyARcwQSbFSDhRq JXCMsgmx5En9jFkpiyNCoQxcbCxDzLlgFPzxmWtfGqgeA4LiNzitAiuVxOVu/0/6xOrG PRPQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783829335; x=1784434135; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=aH3kKOMcwn56DSlHVteprYGWMctwEqxrrgXOQeiC+bw=; b=XhhjdDb6atFR7MYYkBTOqPe7g1k1qzB3Tuxt396ktIp2C6LrUZv+WCCO9W/RaACyaJ pKHZzLnIt/DKI9tunSKVHNR24k1lNrLM3xZAT9F5omliCqh8ufuJG4Ub9dLXKSHaC84x MaEazVSvA8HOG4mUWwKPaH2qWUqpptZOY7rgiDRIBiYVdNplRYaZk5u3hs5oo+PDs+Y1 +tEa+P9dO/2I7WeaSIp01JjonJEVgwHclHDm1y89M78Nft6hPwx/x6OVphp0X3rAzEg/ R7ZOALyP7a/b+rXZr+oMj8vxmaAdx6idHDKY2vdVuvSsLZ8Jd8MKrSEij5+svqQm31U6 79yA== X-Forwarded-Encrypted: i=1; AHgh+Rp2YYmzhxyeCDS+LunOgBIHVZn+8Ul92VWwO5xQJs8to3Bv7iUy2s7+mChGaAvBPb19Ic83h7rH5QbM3EY=@vger.kernel.org X-Gm-Message-State: AOJu0YwxOFRqrkAxhCEGIN8GTcSAZ3Yg7r5NyrjRFYh5m8AGotU8Pawe 0VlY2vHuPmjoOErQC2E1v7/7nMLQV13eeDSpQ1wjJsg1Ro0OGRgshl16WMcipdre X-Gm-Gg: AfdE7cnSA6K1wVsCfwnr5IbAaNIGTtRgZsbSB92MuWAFITH15lRSqkJP4eLTeyRfbW7 VvV1/iXp2kgpfuTggbZkvk95/dy2QkgPXTYI1XhdQcz79y70VCxYDrwe3CitDPat1PugpBIvHsT L/Wzdq+NNSDVaIR/KaKV/QPm7x7gZW0gx45XWYcH24YO6tvwACwdi8S0hSYygu+lTJUUtxclkws +Bdoxs+07/fX7+c6xnfrjNWNAxumepPCSweqnzcXunuUZI0CMUKA8G6vpcy7Agml8iLETX/HhUw tm81Vw6Xp7QKBsPLKfezbVoFwyvuXBYqlYTqrN/d3/HnGhDlgYe8/eUToZILPh2wyAFh5JXHRKa th2LalMEOC/S16GHktSbXS+fRIga4jeFm7Tu9DME1pPfdyTACVKEISLMwKrs6CsiDVFFEywTw0N W+g9FvgpomAfk5ZQSb X-Received: by 2002:a17:90b:3f84:b0:387:d5bd:622e with SMTP id 98e67ed59e1d1-38dc83734a0mr3348857a91.17.1783829334478; Sat, 11 Jul 2026 21:08:54 -0700 (PDT) Received: from [127.0.0.2] ([98.35.8.117]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31174839f89sm56808928eec.10.2026.07.11.21.08.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 11 Jul 2026 21:08:53 -0700 (PDT) From: Farid Zakaria Date: Sat, 11 Jul 2026 21:08:16 -0700 Subject: [PATCH v2 3/5] binfmt_misc: wire up bpf-backed 'B' entries Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260711-binfmt-misc-bpf-v2-v2-3-d6591ceaf207@gmail.com> References: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> In-Reply-To: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> To: Christian Brauner , Alexei Starovoitov , Daniel Borkmann , Martin KaFai Lau , Shuah Khan Cc: Andrii Nakryiko , Kees Cook , Alexander Viro , Jan Kara , Jonathan Corbet , Jann Horn , John Ericson , linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Farid Zakaria X-Mailer: b4 0.14.3 From: Christian Brauner Activate a registered binfmt_misc_ops handler through the existing text interface with the new 'B' entry type: echo ':name:B:::::' > /register The offset field carries the handler name; magic, mask, and interpreter must be empty since the program supplies both the matching and the interpreter. Reusing the register file keeps the existing permission model intact: activating a handler requires the same write access to a binfmt_misc instance as any other registration, and the per user namespace instance semantics apply unchanged. A 'B' entry in a container's own instance shadows the host's handlers just like any other entry, and the privilege needed to shadow e.g. all ELF binaries is the same as for a static 'M' entry matching \x7fELF today; the only novelty is that matching becomes programmable. The entry takes its own reference on the ops for its whole lifetime. It is dropped in put_binfmt_handler() next to the MISC_FMT_OPEN_FILE interp_file cleanup, which the existing users refcount already defers past any concurrent load_misc_binary(), and explicitly on the registration failure path where the users refcount is not live yet. The program runs from load_misc_binary(), never from the matching walk under entries_lock: the read_lock disables preemption while the load program must be able to sleep to read file content. search_binfmt_handler() therefore only nominates the next enabled 'B' entry and the program decides outside the lock. A declining program (returning 0) falls through to handlers registered after it via a skip cursor and a rescan; entries registered or removed between rescans can shift the cursor, so a program may be consulted twice in that window, which is harmless since matching must be free of side effects. Returning a positive value without having selected an interpreter terminates the binfmt_misc scan with -ENOEXEC so the remaining binary formats still get a shot. The 'C' and 'F' flags are rejected for 'B' entries. 'F' exists to pre-open a fixed interpreter at registration time in the registrar's context which is meaningless for a per-exec computed path. 'C' honors the suid bits of the matched binary while executing the interpreter; combined with a program-chosen interpreter that would let a user namespace root pick what runs with a setuid binary's credentials. The computed path itself cannot widen access: it is opened with open_exec() under the caller's credentials with the usual LSM and noexec checks, identical to a statically registered interpreter. Link: https://lore.kernel.org/20260704211409.1978485-1-farid.m.zakaria@gmai= l.com Signed-off-by: Christian Brauner (Amutable) Co-developed-by: Farid Zakaria Signed-off-by: Farid Zakaria Assisted-by: Claude:Opus-4.8 --- Documentation/admin-guide/binfmt-misc.rst | 40 +++++++- fs/binfmt_misc.c | 149 ++++++++++++++++++++++++++= +--- 2 files changed, 175 insertions(+), 14 deletions(-) diff --git a/Documentation/admin-guide/binfmt-misc.rst b/Documentation/admi= n-guide/binfmt-misc.rst index 306ef48f5..de948dae7 100644 --- a/Documentation/admin-guide/binfmt-misc.rst +++ b/Documentation/admin-guide/binfmt-misc.rst @@ -26,11 +26,13 @@ Here is what the fields mean: name below ``/proc/sys/fs/binfmt_misc``; cannot contain slashes ``/`` f= or obvious reasons. - ``type`` - is the type of recognition. Give ``M`` for magic and ``E`` for extensio= n. + is the type of recognition. Give ``M`` for magic, ``E`` for extension a= nd + ``B`` for a bpf-backed handler (see below). - ``offset`` is the offset of the magic/mask in the file, counted in bytes. This defaults to 0 if you omit it (i.e. you write ``:name:type::magic...``). - Ignored when using filename extension matching. + Ignored when using filename extension matching. For ``B`` entries this + field carries the name of the bpf handler instead. - ``magic`` is the byte sequence binfmt_misc is matching for. The magic string may contain hex-encoded characters like ``\x0a`` or ``\xA4``. Note that= you @@ -97,6 +99,40 @@ There are some restrictions: offset+size(magic) has to be less than 128 - the interpreter string may not exceed 127 characters =20 + +bpf-backed handlers +------------------- + +With ``CONFIG_BINFMT_MISC_BPF`` both the matching and the interpreter +selection can be delegated to a bpf program. A handler is an instance of t= he +``binfmt_misc_ops`` struct_ops with a sleepable ``load`` program and a +``name``. Once the struct_ops map is registered the handler can be activat= ed +with a ``B`` entry that references it by name and carries neither magic, +mask, nor interpreter:: + + echo ':qemu:B:my_handler::::' > register + +At exec time the ``load`` program receives the ``linux_binprm`` of the +binary. It can match on the header in ``bprm->buf``, read the file itself, +e.g. to parse ELF program headers, and derive the interpreter from the +binary's location. It selects the interpreter by calling the +``bpf_binprm_set_interp()`` kfunc with an absolute path and returning a +positive value. Returning ``0`` falls through to the handlers registered +after this one, a negative errno fails the exec with that error; +``-ENOEXEC`` ends the binfmt_misc search but lets the remaining binary +formats have a go. The interpreter is opened with the credentials of the +task doing the exec, exactly as a statically registered interpreter would +be. + +Handlers are looked up in the user namespace the struct_ops map was +registered in, falling back to ancestor namespaces, mirroring how +binfmt_misc instances themselves are looked up. The entry keeps the handler +alive; deleting the struct_ops map only prevents new registrations. + +The ``C`` and ``F`` flags cannot be combined with ``B`` entries: there is = no +fixed interpreter to pre-open and a program-selected interpreter must never +inherit the credentials of a setuid binary. + To use binfmt_misc you have to mount it first. You can mount it with ``mount -t binfmt_misc none /proc/sys/fs/binfmt_misc`` command, or you can= add a line ``none /proc/sys/fs/binfmt_misc binfmt_misc defaults 0 0`` to your diff --git a/fs/binfmt_misc.c b/fs/binfmt_misc.c index fcaad14f8..4ece75f95 100644 --- a/fs/binfmt_misc.c +++ b/fs/binfmt_misc.c @@ -11,6 +11,7 @@ #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt =20 #include +#include #include #include #include @@ -38,6 +39,7 @@ enum binfmt_misc_entry_bits { MISC_FMT_ENABLED_BIT =3D 0, MISC_FMT_MAGIC_BIT =3D 1, + MISC_FMT_BPF_BIT =3D 2, }; =20 /* Entry behavior flags, fixed at registration time. */ @@ -59,6 +61,8 @@ struct binfmt_misc_entry { char *name; struct dentry *dentry; struct file *interp_file; + const struct binfmt_misc_ops *bpf_ops; /* bpf-backed handler ('B') */ + const char *bpf_ops_name; refcount_t users; /* sync removal with load_misc_binary() */ struct rcu_head rcu; char buf[]; /* register string, fields point in here */ @@ -109,16 +113,21 @@ static bool entry_matches_extension(const struct binf= mt_misc_entry *e, * search_binfmt_handler - search for a binary handler for @bprm * @misc: handle to binfmt_misc instance * @bprm: binary for which we are looking for a handler + * @bpf_skip: number of bpf-backed handlers to skip over * * Search for a binary type handler for @bprm in the list of registered bi= nary - * type handlers. + * type handlers. A bpf-backed handler cannot be matched here as its progr= am + * must run in sleepable context; it is returned as a candidate and the + * program decides in load_misc_binary(). @bpf_skip resumes the search aft= er + * the first @bpf_skip candidates declined. * * The caller must hold the RCU read lock. * * Return: binary type list entry on success, NULL on failure */ static struct binfmt_misc_entry * -search_binfmt_handler(struct binfmt_misc *misc, struct linux_binprm *bprm) +search_binfmt_handler(struct binfmt_misc *misc, struct linux_binprm *bprm, + unsigned int bpf_skip) { char *dot =3D strrchr(bprm->interp, '.'); const char *ext =3D dot ? dot + 1 : NULL; @@ -130,6 +139,15 @@ search_binfmt_handler(struct binfmt_misc *misc, struct= linux_binprm *bprm) if (!test_bit(MISC_FMT_ENABLED_BIT, &e->flags)) continue; =20 + /* A bpf handler is decided in load_misc_binary(). */ + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) { + if (bpf_skip) { + bpf_skip--; + continue; + } + return e; + } + if (test_bit(MISC_FMT_MAGIC_BIT, &e->flags)) { if (entry_matches_magic(e, bprm)) return e; @@ -146,6 +164,7 @@ search_binfmt_handler(struct binfmt_misc *misc, struct = linux_binprm *bprm) * get_binfmt_handler - try to find a binary type handler * @misc: handle to binfmt_misc instance * @bprm: binary for which we are looking for a handler + * @bpf_skip: number of bpf-backed handlers to skip over * * Try to find a binfmt handler for the binary type. If one is found take a * reference to protect against removal via bm_{entry,status}_write(). The @@ -156,13 +175,14 @@ search_binfmt_handler(struct binfmt_misc *misc, struc= t linux_binprm *bprm) * Return: binary type list entry on success, NULL on failure */ static struct binfmt_misc_entry *get_binfmt_handler(struct binfmt_misc *mi= sc, - struct linux_binprm *bprm) + struct linux_binprm *bprm, + unsigned int bpf_skip) { struct binfmt_misc_entry *e; =20 guard(rcu)(); do { - e =3D search_binfmt_handler(misc, bprm); + e =3D search_binfmt_handler(misc, bprm, bpf_skip); } while (e && !refcount_inc_not_zero(&e->users)); return e; } @@ -182,6 +202,8 @@ static void put_binfmt_handler(struct binfmt_misc_entry= *e) exe_file_allow_write_access(e->interp_file); filp_close(e->interp_file, NULL); } + if (e->bpf_ops) + binfmt_misc_put_ops(e->bpf_ops); /* Lockless walkers may still dereference this entry. */ kfree_rcu(e, rcu); } @@ -224,13 +246,16 @@ static int load_misc_binary(struct linux_binprm *bprm) struct binfmt_misc_entry *fmt __free(put_binfmt_handler) =3D NULL; struct file *interp_file; struct binfmt_misc *misc; + const char *interpreter; + unsigned int bpf_skip =3D 0; int retval; =20 misc =3D current_binfmt_misc(); if (!READ_ONCE(misc->enabled)) return -ENOEXEC; =20 - fmt =3D get_binfmt_handler(misc, bprm); +retry: + fmt =3D get_binfmt_handler(misc, bprm, bpf_skip); if (!fmt) return -ENOEXEC; =20 @@ -238,6 +263,37 @@ static int load_misc_binary(struct linux_binprm *bprm) if (bprm->interp_flags & BINPRM_FLAGS_PATH_INACCESSIBLE) return -ENOENT; =20 + if (test_bit(MISC_FMT_BPF_BIT, &fmt->flags)) { + /* + * A bpf-backed handler matches and picks the interpreter in + * sleepable context: > 0 means it handled the binary, 0 falls + * through to the handlers after this one and a negative errno + * fails the exec. + */ + retval =3D fmt->bpf_ops->load(bprm); + if (retval < 0) { + /* Keep a program-supplied error within errno range. */ + if (retval < -MAX_ERRNO) + retval =3D -ENOEXEC; + return retval; + } + if (!retval) { + /* Declined: move on to the handlers after this one. */ + kfree(bprm->bpf_interp); + bprm->bpf_interp =3D NULL; + put_binfmt_handler(fmt); + fmt =3D NULL; + bpf_skip++; + goto retry; + } + /* Selecting an interpreter is part of the contract. */ + if (!bprm->bpf_interp) + return -ENOEXEC; + interpreter =3D bprm->bpf_interp; + } else { + interpreter =3D fmt->interpreter; + } + if (fmt->flags & MISC_FMT_PRESERVE_ARGV0) { bprm->interp_flags |=3D BINPRM_FLAGS_PRESERVE_ARGV0; } else { @@ -256,13 +312,13 @@ static int load_misc_binary(struct linux_binprm *bprm) bprm->argc++; =20 /* add the interp as argv[0] */ - retval =3D copy_string_kernel(fmt->interpreter, bprm); + retval =3D copy_string_kernel(interpreter, bprm); if (retval < 0) return retval; bprm->argc++; =20 /* Update interp in case binfmt_script needs it. */ - retval =3D bprm_change_interp(fmt->interpreter, bprm); + retval =3D bprm_change_interp(interpreter, bprm); if (retval < 0) return retval; =20 @@ -277,7 +333,7 @@ static int load_misc_binary(struct linux_binprm *bprm) } } } else { - interp_file =3D open_exec(fmt->interpreter); + interp_file =3D open_exec(interpreter); } if (IS_ERR(interp_file)) return PTR_ERR(interp_file); @@ -428,6 +484,38 @@ static char *parse_extension_fields(struct binfmt_misc= _entry *e, char *p, return p; } =20 +/* Parse the 'offset' field of a 'B' entry: the bpf handler name. */ +static char *parse_bpf_fields(struct binfmt_misc_entry *e, char *p, char d= el) +{ + char *s; + + /* The 'offset' field carries the bpf handler name. */ + s =3D strchr(p, del); + if (!s) + return NULL; + *s++ =3D '\0'; + e->bpf_ops_name =3D p; + if (!e->bpf_ops_name[0] || + strlen(e->bpf_ops_name) >=3D BINFMT_MISC_OPS_NAME_MAX) + return NULL; + p =3D s; + pr_debug("register: bpf handler: {%s}\n", e->bpf_ops_name); + + /* The 'magic' field must be empty. */ + s =3D strchr(p, del); + if (!s || s !=3D p) + return NULL; + *s++ =3D '\0'; + p =3D s; + + /* The 'mask' field must be empty. */ + s =3D strchr(p, del); + if (!s || s !=3D p) + return NULL; + *s++ =3D '\0'; + return s; +} + /* * This registers a new binary format, it recognises the syntax * ':name:type:offset:magic:mask:interpreter:flags' @@ -492,13 +580,21 @@ static struct binfmt_misc_entry *create_entry(const c= har __user *buffer, pr_debug("register: type: M (magic)\n"); e->flags =3D BIT(MISC_FMT_ENABLED_BIT) | BIT(MISC_FMT_MAGIC_BIT); break; + case 'B': + if (!IS_ENABLED(CONFIG_BINFMT_MISC_BPF)) + return ERR_PTR(-EINVAL); + pr_debug("register: type: B (bpf)\n"); + e->flags =3D BIT(MISC_FMT_ENABLED_BIT) | BIT(MISC_FMT_BPF_BIT); + break; default: return ERR_PTR(-EINVAL); } if (*p++ !=3D del) return ERR_PTR(-EINVAL); =20 - if (test_bit(MISC_FMT_MAGIC_BIT, &e->flags)) + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) + p =3D parse_bpf_fields(e, p, del); + else if (test_bit(MISC_FMT_MAGIC_BIT, &e->flags)) p =3D parse_magic_fields(e, p, del); else p =3D parse_extension_fields(e, p, del); @@ -511,8 +607,13 @@ static struct binfmt_misc_entry *create_entry(const ch= ar __user *buffer, if (!p) return ERR_PTR(-EINVAL); *p++ =3D '\0'; - if (!e->interpreter[0]) + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) { + /* The program selects the interpreter at exec time. */ + if (e->interpreter[0]) + return ERR_PTR(-EINVAL); + } else if (!e->interpreter[0]) { return ERR_PTR(-EINVAL); + } pr_debug("register: interpreter: {%s}\n", e->interpreter); =20 /* Parse the 'flags' field. */ @@ -522,6 +623,14 @@ static struct binfmt_misc_entry *create_entry(const ch= ar __user *buffer, if (p !=3D buf + count) return ERR_PTR(-EINVAL); =20 + /* + * A program-selected interpreter cannot be pre-opened and must not + * inherit the credentials of a setuid binary it was chosen for. + */ + if (test_bit(MISC_FMT_BPF_BIT, &e->flags) && + (e->flags & (MISC_FMT_CREDENTIALS | MISC_FMT_OPEN_FILE))) + return ERR_PTR(-EINVAL); + return no_free_ptr(e); } =20 @@ -575,7 +684,10 @@ static int bm_entry_show(struct seq_file *m, void *unu= sed) else seq_puts(m, "disabled\n"); =20 - seq_printf(m, "interpreter %s\n", e->interpreter); + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) + seq_printf(m, "bpf %s\n", e->bpf_ops_name); + else + seq_printf(m, "interpreter %s\n", e->interpreter); =20 /* print the special flags */ seq_puts(m, "flags: "); @@ -589,7 +701,9 @@ static int bm_entry_show(struct seq_file *m, void *unus= ed) seq_putc(m, 'F'); seq_putc(m, '\n'); =20 - if (!test_bit(MISC_FMT_MAGIC_BIT, &e->flags)) { + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) { + /* No magic or extension to print for a bpf handler. */ + } else if (!test_bit(MISC_FMT_MAGIC_BIT, &e->flags)) { seq_printf(m, "extension .%s\n", e->magic); } else { seq_printf(m, "offset %i\nmagic ", e->offset); @@ -840,6 +954,15 @@ static ssize_t bm_register_write(struct file *file, co= nst char __user *buffer, if (IS_ERR(e)) return PTR_ERR(e); =20 + if (test_bit(MISC_FMT_BPF_BIT, &e->flags)) { + e->bpf_ops =3D binfmt_misc_get_ops(sb->s_user_ns, e->bpf_ops_name); + if (!e->bpf_ops) { + pr_notice("register: no bpf handler named %s\n", + e->bpf_ops_name); + return -ENOENT; + } + } + if (e->flags & MISC_FMT_OPEN_FILE) { /* * Now that we support unprivileged binfmt_misc mounts make @@ -864,6 +987,8 @@ static ssize_t bm_register_write(struct file *file, con= st char __user *buffer, exe_file_allow_write_access(f); filp_close(f, NULL); } + if (e->bpf_ops) + binfmt_misc_put_ops(e->bpf_ops); return err; } =20 --=20 2.51.2 From nobody Sat Jul 25 23:03:44 2026 Received: from mail-pj1-f50.google.com (mail-pj1-f50.google.com [209.85.216.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 99BDB370AD1 for ; Sun, 12 Jul 2026 04:08:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829338; cv=none; b=GbtVZXsBOn0ANe2BqwfgosxogDOpoT+f8KwSXjjKshopu4HDqf3+AkU6ydhHOqx62pj+3AsthbQqjaTX2+cOJjvXS4glCxT5zZEOfgSzxvlGslENHA6VoG+PEs1wqQDXBTcGfQb/KQTpvUDfrcrTMSytVDqTQUePZm7BY5WyKo4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829338; c=relaxed/simple; bh=tLYfqud9Alqsuxe2P8hqfcpiiFUye1isw4lqTTAq7GM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=D/TwQnB5SaOT172lsiEOFlj/R0ohBR5x8JIpLLgXWNbDMVw3GXqFDrjF9HlWhU0iOQhbCTJbgtGeoDlOkINKRIYCuCj+GxTSvID1RTglwxiyoV/ULFHy3GFJbg8OT5kdhmgcf5IYFZwzmN46fEcv2b/oCeeg9nGZPdXnfkZY7dg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=YTkMz0G2; arc=none smtp.client-ip=209.85.216.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="YTkMz0G2" Received: by mail-pj1-f50.google.com with SMTP id 98e67ed59e1d1-381b831d535so2591320a91.0 for ; Sat, 11 Jul 2026 21:08:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783829336; x=1784434136; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QOcBwZZCQda0ksnOyTTPJoQylHUy+3FzULiJTSWI84w=; b=YTkMz0G2lYG3k+HO6Zfj+E3IkAtsfUbdK9IHOYMjpSdmm4TnpG+7vES9umMhfv8NCJ N3C6+LbQqBfvD+FfZDlhRSb8h+7DkD5bHI6tbEqtTdKJlCfAYMtdVR1gzVW8LU/NOnWB JKddzULTjjGuMMrXyflrBLLez5v/ddd35O/RURSsx/a3/aySno6DUV0mvDI4gh3UR0J5 Re0Do7cxZ/6U5vnv2/qk0PymakVQth/plYR90RlXb0Yp4spekAZ2Yw+AQiKIDhPXqxUQ 8b5Ilpa54+ugO9bR350EPNf+KodfVSDcytrkRnrWD0Xo7hVvR8tAvm7ZRdGnElLFFZyC 7Fqg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783829336; x=1784434136; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=QOcBwZZCQda0ksnOyTTPJoQylHUy+3FzULiJTSWI84w=; b=WLzY3jPFW8GFed87PwWk+innwekTaEhWc2sUMhMlawt+Wnvq0iRtT8nOCcPojXINwF xdvUUglA+v1L7ssnRuGtIOewSurkeHl7uBXBvo7XfjP7V1Oz3wfglpxY8WKV61Cte6Hj y03fFIrm2r1nHtkEcBZXMQVFV2ch5An2LBbWUuos++H8ou9aCOouNJOVXIflaU+bn8BK BCORi45HUyPcCX8i+tlLAszarYfNkZZ4HGmBGC0e2EWxiEyTLafxBG1UsGl4S5gZ+llo YtDx//LeqOAp/a5NfBHpBJm35gMSEsXVF6WniqenloFUAvGNOyrVv8bPBSodT1sJY/hQ S6Qw== X-Forwarded-Encrypted: i=1; AHgh+RoAMAyl8RqLjYnITLtPBoW5zYCdXLgtFIO6aCnQI9nxt37/XY5MUMbUznspp1uOFlApgTwump/UpMarcmY=@vger.kernel.org X-Gm-Message-State: AOJu0Yx0nzZbQqT5eKcy0ofxRCyspZfqO5rUUcuCBSgjSU2ksCsSzfxu QkDYC8Xm5cK87PgO2fxq4Cgj+u1LHQzOu5d/PaJ2b/bEIxvnRBhMcuXl X-Gm-Gg: AfdE7cnVQmK1/B15VZR8z/3CKVrb+hHCgiayWFhBBIUxIbm8OawZBJiYotCwsaHzDW9 +2MgMVTjYGUdTfX3TxhLJqKbmr2+bw0w/N2RdFGpf1pQc7CSwO4F2hJNzNRry/MQk9ppwCw95rN SlIGEO7R+jUJLNKxIPy5cpZdgNXzilIQhBqaDwUwwp1CRvR4bAH57cLOMvI1gnwnYzP52dtFCAP rzPt16b8ClzyaW8AnBr+XEpEkd75rYVz1azSWRDj/Jf+YTontnugzJ9M1z6Cp6DBrRHq6X4bsvV hYU45p3QOq+fN8inFs4Xv91yrf+xvGeEXuFbMj70aQoanoHglE5BeT0BOgUZ6iuh0TI1zJ8tj32 uWpVBMf4+3a0vC0QDIHguzVofs+g+gZZwTACz5aUkKmPyeziGLoJHkEgR3M7v9JaEMNhShWkcr5 Zwlas7PB3q+pcOdG8r X-Received: by 2002:a17:90b:39ad:b0:381:abcc:c8e9 with SMTP id 98e67ed59e1d1-38dc73bcfecmr4632607a91.7.1783829335997; Sat, 11 Jul 2026 21:08:55 -0700 (PDT) Received: from [127.0.0.2] ([98.35.8.117]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31174839f89sm56808928eec.10.2026.07.11.21.08.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 11 Jul 2026 21:08:55 -0700 (PDT) From: Farid Zakaria Date: Sat, 11 Jul 2026 21:08:17 -0700 Subject: [PATCH v2 4/5] bpf: allow fs kfuncs for binfmt_misc_ops programs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260711-binfmt-misc-bpf-v2-v2-4-d6591ceaf207@gmail.com> References: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> In-Reply-To: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> To: Christian Brauner , Alexei Starovoitov , Daniel Borkmann , Martin KaFai Lau , Shuah Khan Cc: Andrii Nakryiko , Kees Cook , Alexander Viro , Jan Kara , Jonathan Corbet , Jann Horn , John Ericson , linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Farid Zakaria X-Mailer: b4 0.14.3 From: Christian Brauner The fs kfuncs are currently exclusive to LSM programs. A binfmt_misc load program needs a subset of them to do anything interesting: to compute an interpreter relative to the binary's location it wants bpf_path_d_path() on bprm->file->f_path, and matching on per-binary metadata wants bpf_get_file_xattr() and friends. Register the fs kfunc set for struct_ops programs as well and extend the filter to admit binfmt_misc_ops programs. The xattr setters stay exclusive to LSM programs: a binary type handler decides how to run a binary, it has no business modifying filesystem state. This only takes effect in builds that have the fs kfunc set at all, i.e. CONFIG_BPF_LSM. Without it a binfmt_misc handler is limited to bprm fields and the file-backed dynptr, which are provided by the common kfunc set. Link: https://lore.kernel.org/20260704211409.1978485-1-farid.m.zakaria@gmai= l.com Signed-off-by: Christian Brauner (Amutable) --- fs/bpf_fs_kfuncs.c | 23 ++++++++++++++++++++--- 1 file changed, 20 insertions(+), 3 deletions(-) diff --git a/fs/bpf_fs_kfuncs.c b/fs/bpf_fs_kfuncs.c index 768aca2dc..aa1fe988b 100644 --- a/fs/bpf_fs_kfuncs.c +++ b/fs/bpf_fs_kfuncs.c @@ -1,6 +1,7 @@ // SPDX-License-Identifier: GPL-2.0 /* Copyright (c) 2024 Google LLC. */ =20 +#include #include #include #include @@ -387,10 +388,20 @@ BTF_ID_FLAGS(func, bpf_remove_dentry_xattr, KF_SLEEPA= BLE) BTF_ID_FLAGS(func, bpf_real_inode, KF_SLEEPABLE | KF_RET_NULL) BTF_KFUNCS_END(bpf_fs_kfunc_set_ids) =20 +/* Side-effecting kfuncs that stay exclusive to LSM programs. */ +BTF_SET_START(bpf_fs_kfunc_lsm_only_ids) +BTF_ID(func, bpf_set_dentry_xattr) +BTF_ID(func, bpf_remove_dentry_xattr) +BTF_SET_END(bpf_fs_kfunc_lsm_only_ids) + static int bpf_fs_kfuncs_filter(const struct bpf_prog *prog, u32 kfunc_id) { - if (!btf_id_set8_contains(&bpf_fs_kfunc_set_ids, kfunc_id) || - prog->type =3D=3D BPF_PROG_TYPE_LSM) + if (!btf_id_set8_contains(&bpf_fs_kfunc_set_ids, kfunc_id)) + return 0; + if (prog->type =3D=3D BPF_PROG_TYPE_LSM) + return 0; + if (bpf_prog_is_binfmt_misc_ops(prog) && + !btf_id_set_contains(&bpf_fs_kfunc_lsm_only_ids, kfunc_id)) return 0; return -EACCES; } @@ -433,7 +444,13 @@ static const struct btf_kfunc_id_set bpf_fs_kfunc_set = =3D { =20 static int __init bpf_fs_kfuncs_init(void) { - return register_btf_kfunc_id_set(BPF_PROG_TYPE_LSM, &bpf_fs_kfunc_set); + int ret; + + ret =3D register_btf_kfunc_id_set(BPF_PROG_TYPE_LSM, &bpf_fs_kfunc_set); + if (ret || !IS_ENABLED(CONFIG_BINFMT_MISC_BPF)) + return ret; + return register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS, + &bpf_fs_kfunc_set); } =20 late_initcall(bpf_fs_kfuncs_init); --=20 2.51.2 From nobody Sat Jul 25 23:03:44 2026 Received: from mail-pj1-f46.google.com (mail-pj1-f46.google.com [209.85.216.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 43122370D69 for ; Sun, 12 Jul 2026 04:08:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829340; cv=none; b=TKPsQjhbHFLvlTdUr2yzfjbXouPPyE5LP6cel1YBC9A6kyvO/2chccS1hKX1t2eOU4iNiHMSseWxgbDe2kTP2Cjc+12jYpPis9HL2c5GkRPvB0ZXC0Ijcl+TkInoitoM3sTU+xyMpuvHkK7UEvXndkpajVoqLeqxaow5C2L1iY4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783829340; c=relaxed/simple; bh=UwONQdSIR58AzteSQrmv9kMXhoLqyKx9woojMAMVUUM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tfZIyq0S/CLtNpmoJM0ItBaxBdvmhUXyWZ+xroNCnpW5xH75SyN0f+boVXKInf1Zvi0X+HS9FyTv04emrQYGXAp657pDi4Lly8dQgr1AEnHZp4Yu+Aici0vwYn8k8k8H0YgG3dtqHXncWjW2l293WuEix9aFETiQ7fgoJRCEWJc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=USdSJQm+; arc=none smtp.client-ip=209.85.216.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="USdSJQm+" Received: by mail-pj1-f46.google.com with SMTP id 98e67ed59e1d1-38df0038497so23666a91.3 for ; Sat, 11 Jul 2026 21:08:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783829338; x=1784434138; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SNMnN6PRNArZolS5V8JiE9hnU6pWFtcLrBR1+3UF0vs=; b=USdSJQm+ZaSXWnLEkdgAN3JwEpOQLBgXgvOrpyo9cVscUsdSTzureuKDaIUbc/KQWM dVMCvNBxtADz/+wsWceW2uXWtyYnU4RHPa4fkJgIPB7WKcK3MPbVy8AmUqwViYklovXS X1GH8ZV+7ijOH53cl1dXCB+36oSL/KfLXOqgplrqPKd3Yu+u9BnCHSU8MCTg9vVDhL3Q vaPNqClixUdWO64mjfQ8aso64bRSjclvhwPvoP8zBbLyrfc11sqXcX64cck0of7hry0o JZuvRIHZ07ba7peHwu3FLSQwzxWhOLZmDuH6v+wav0E9jYdBvZ5t8UaQs8QrP1Vtn3cj arjw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783829338; x=1784434138; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SNMnN6PRNArZolS5V8JiE9hnU6pWFtcLrBR1+3UF0vs=; b=Yb7qZWwpk4uOsUA4C9vfqKtr2x62T+sHmKuY4wQFWms5xiUrtMBEpd8laFqQ2nPivB yr8fUY94YrToKj8E2NDV28nhkOMeCEpswj9+ERSP7rtacveyrVlXRcDsWRh6zmSVonL2 Bd5JXBdNbbQPz+1SteqEKO2k+IWDK/BJ98NzG0ZLDSt4MjdjJ9iWDVAyBO+6xAeoE7di 8ENKk2PKU3QjGyRFhinPOQN8bh5XK1MFYRitWHFIVY9uop05+9v1FQwIUtNzLnMpyGYD abKQHvwB4ZSOL1Z8tQy1RxZaIGD3K3oVqGU9fQ3gUArKmrm12Eryeq7i5ofD19AfFEUs tt3Q== X-Forwarded-Encrypted: i=1; AHgh+RrMFN1mnUlIWDGlmYztztRwPTHP+zLZzX7AkEzKvrpNQ8vx/2CdXqHnenKHSRlmtJUZfoLfXvn46f8KCns=@vger.kernel.org X-Gm-Message-State: AOJu0YxLP4LKDO6HsKBBmwTKLD2KjhCLqnOt54MULV5VtYyX/Wzm3yRn jiCRDYSK3VAnEcGjpAHYUepBV+QwNPe/dxaog98EmgH5R5e3zd+IWi7k8YFLpJK4 X-Gm-Gg: AfdE7clLxpUQw7Zjkq6Z94S5DY8okW14rntrUc1tl0N+9PCk3wpAm++hIirwNf7yY6t 4LOJQGGAqZkYRUNTdCX6u86m3dQoDcE2vVLeIlw06GtOKZhLSKOD/LHXwavpO9b7YBhwfL+Mgef d2XY2fWxboIUtvXH9I3EZBOwOmidrfu1xgf1BzpIWvsBP9W+YTdfVS4EmnCqzVPd6sqsXisoPyu Iu+PS8XNk2FdqPGa+lRLMJMlmDIXDo/LqQJQ46m/UTbX/DQuOL17EjQbJjmziHU2DQzzbDN/ll+ kDffpqm2/gcMiFbkTyZRrFdr7XeXI2kr4xnoToVd3GYdMFFqAKdQgRL2QZ+LVnvq4K6Ha7F2O4K b6SFooQ43hT2x5LgZeQATDl7/ZnWPc/P1S2H6/Ivk/8I9UhJIHF9xLBtler6du/n6UhY1pnmPLq rxvk61lN9DY4xkzMuP X-Received: by 2002:a17:90a:e70f:b0:387:e0cb:7ee with SMTP id 98e67ed59e1d1-38dc781ce28mr3727842a91.34.1783829337465; Sat, 11 Jul 2026 21:08:57 -0700 (PDT) Received: from [127.0.0.2] ([98.35.8.117]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31174839f89sm56808928eec.10.2026.07.11.21.08.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 11 Jul 2026 21:08:56 -0700 (PDT) From: Farid Zakaria Date: Sat, 11 Jul 2026 21:08:18 -0700 Subject: [PATCH v2 5/5] selftests/exec: add binfmt_misc bpf-backed handler test Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260711-binfmt-misc-bpf-v2-v2-5-d6591ceaf207@gmail.com> References: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> In-Reply-To: <20260711-binfmt-misc-bpf-v2-v2-0-d6591ceaf207@gmail.com> To: Christian Brauner , Alexei Starovoitov , Daniel Borkmann , Martin KaFai Lau , Shuah Khan Cc: Andrii Nakryiko , Kees Cook , Alexander Viro , Jan Kara , Jonathan Corbet , Jann Horn , John Ericson , linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Farid Zakaria X-Mailer: b4 0.14.3 Exercise the bpf-backed ('B') binfmt_misc handlers end to end. A handler is a struct binfmt_misc_ops struct_ops map; the test loads and attaches it (which publishes it by name), activates it with a 'B' entry, and checks that a matched binary is routed to the interpreter the program selected via bpf_binprm_set_interp(). Two self-contained cases are covered: - bpf_interp: match a synthetic aarch64 ELF header from bprm->buf and route it to a fixed interpreter chosen by the program. - nix_origin: resolve a "$ORIGIN/..."-relative PT_INTERP to an interpreter co-located with the binary -- the relocatable-loader case the kernel ELF loader cannot express. The relocatable binary is linked with PT_INTERP set to the literal "$ORIGIN/binfmt_bpf_interp" (-Wl,--dynamic-linker), which the kernel cannot resolve on its own. Both route to a small test interpreter that prints a marker, proving the program-selected interpreter actually ran. The bpf objects are compiled against the running kernel's BTF: the Makefile generates vmlinux.h with bpftool and the harness links libbpf. Override CLANG/BPFTOOL/VMLINUX_BTF/LIBBPF_CFLAGS/LIBBPF_LDLIBS as needed. Signed-off-by: Farid Zakaria Assisted-by: Claude:Opus-4.8 --- tools/testing/selftests/exec/Makefile | 37 ++++ tools/testing/selftests/exec/binfmt_bpf_app.c | 12 ++ tools/testing/selftests/exec/binfmt_bpf_interp.c | 15 ++ tools/testing/selftests/exec/binfmt_misc_bpf.c | 260 +++++++++++++++++++= ++++ tools/testing/selftests/exec/bpf_interp.bpf.c | 52 +++++ tools/testing/selftests/exec/nix_origin.bpf.c | 179 ++++++++++++++++ 6 files changed, 555 insertions(+) diff --git a/tools/testing/selftests/exec/Makefile b/tools/testing/selftest= s/exec/Makefile index 45a3cfc43..5423e3f7f 100644 --- a/tools/testing/selftests/exec/Makefile +++ b/tools/testing/selftests/exec/Makefile @@ -21,6 +21,12 @@ TEST_GEN_PROGS +=3D recursion-depth TEST_GEN_PROGS +=3D null-argv TEST_GEN_PROGS +=3D check-exec =20 +# binfmt_misc bpf-backed ('B') handler test: a libbpf harness plus its +# struct_ops objects and the test interpreter/app it routes between. +TEST_GEN_PROGS +=3D binfmt_misc_bpf +TEST_GEN_FILES +=3D bpf_interp.bpf.o nix_origin.bpf.o +TEST_GEN_FILES +=3D binfmt_bpf_interp binfmt_bpf_app + EXTRA_CLEAN :=3D $(OUTPUT)/subdir.moved $(OUTPUT)/execveat.moved $(OUTPUT)= /xxxxx* \ $(OUTPUT)/S_I*.test =20 @@ -55,3 +61,34 @@ $(OUTPUT)/script-exec.inc: $(CHECK_EXEC_SAMPLES)/script-= exec.inc cp $< $@ $(OUTPUT)/script-noexec.inc: $(CHECK_EXEC_SAMPLES)/script-noexec.inc cp $< $@ + +# --- binfmt_misc bpf ('B') handler test --------------------------------- +# The struct_ops bpf objects are compiled against the running kernel's BTF. +# Override CLANG/BPFTOOL/VMLINUX_BTF if they aren't on PATH or the vmlinux +# BTF lives elsewhere, and LIBBPF_CFLAGS/LDLIBS to point at a libbpf insta= ll. +CLANG ?=3D clang +BPFTOOL ?=3D bpftool +VMLINUX_BTF ?=3D /sys/kernel/btf/vmlinux +BPF_CFLAGS ?=3D -I$(OUTPUT) +LIBBPF_CFLAGS ?=3D +LIBBPF_LDLIBS ?=3D -lbpf -lelf -lz + +$(OUTPUT)/vmlinux.h: + $(BPFTOOL) btf dump file $(VMLINUX_BTF) format c > $@ + sed -i '/__ksym;$$/d' $@ + +$(OUTPUT)/%.bpf.o: %.bpf.c $(OUTPUT)/vmlinux.h + $(CLANG) -g -O2 -target bpf -mcpu=3Dv3 $(BPF_CFLAGS) -c $< -o $@ + +$(OUTPUT)/binfmt_misc_bpf: binfmt_misc_bpf.c + $(CC) $(CFLAGS) $(LIBBPF_CFLAGS) $(LDFLAGS) $< $(LIBBPF_LDLIBS) -o $@ + +$(OUTPUT)/binfmt_bpf_interp: binfmt_bpf_interp.c + $(CC) $(CFLAGS) $(LDFLAGS) $< -o $@ + +# PT_INTERP is set to the literal "$ORIGIN/binfmt_bpf_interp"; the nix_ori= gin +# handler resolves it relative to the binary at run time. +$(OUTPUT)/binfmt_bpf_app: binfmt_bpf_app.c + $(CC) $(CFLAGS) $(LDFLAGS) -Wl,--dynamic-linker,'$$ORIGIN/binfmt_bpf_inte= rp' $< -o $@ + +EXTRA_CLEAN +=3D $(OUTPUT)/vmlinux.h $(OUTPUT)/bpf_interp.bpf.o $(OUTPUT)/= nix_origin.bpf.o diff --git a/tools/testing/selftests/exec/binfmt_bpf_app.c b/tools/testing/= selftests/exec/binfmt_bpf_app.c new file mode 100644 index 000000000..472270f14 --- /dev/null +++ b/tools/testing/selftests/exec/binfmt_bpf_app.c @@ -0,0 +1,12 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * A relocatable binary for the binfmt_misc_bpf $ORIGIN case. The Makefile + * links it with PT_INTERP set to the literal "$ORIGIN/binfmt_bpf_interp" + * (-Wl,--dynamic-linker), which the kernel ELF loader cannot resolve. The + * nix_origin bpf handler resolves it relative to this binary's directory = and + * routes execution to the co-located interpreter. + */ +int main(void) +{ + return 0; +} diff --git a/tools/testing/selftests/exec/binfmt_bpf_interp.c b/tools/testi= ng/selftests/exec/binfmt_bpf_interp.c new file mode 100644 index 000000000..2db205f09 --- /dev/null +++ b/tools/testing/selftests/exec/binfmt_bpf_interp.c @@ -0,0 +1,15 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Test interpreter for the binfmt_misc_bpf selftest. A bpf-backed 'B' han= dler + * routes a matched binary here; printing this marker proves the program's + * chosen interpreter actually ran. + */ +#include + +int main(int argc, char **argv) +{ + (void)argc; + (void)argv; + write(1, "BPF_INTERP_RAN\n", 15); + return 0; +} diff --git a/tools/testing/selftests/exec/binfmt_misc_bpf.c b/tools/testing= /selftests/exec/binfmt_misc_bpf.c new file mode 100644 index 000000000..b3e017ade --- /dev/null +++ b/tools/testing/selftests/exec/binfmt_misc_bpf.c @@ -0,0 +1,260 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Selftest for binfmt_misc bpf-backed ('B') handlers. + * + * A handler is a struct binfmt_misc_ops struct_ops map. Attaching it publ= ishes + * it by name in the caller's user namespace; a 'B' entry activates it: + * + * echo ':name:B:::::' > /proc/sys/fs/binfmt_misc/register + * + * Two self-contained cases are exercised: + * + * 1. bpf_interp: the handler matches a synthetic aarch64 ELF header and + * routes it to a fixed interpreter chosen by the program. + * 2. nix_origin: the handler resolves a "$ORIGIN/..."-relative PT_INTER= P to + * an interpreter co-located with the binary (the relocatable-loader = case + * the kernel ELF loader cannot express). + * + * Both route to a test interpreter that prints BPF_INTERP_RAN, proving the + * program's chosen interpreter actually ran. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include + +#include + +#define INTERP_PATH "/tmp/binfmt_bpf_interp" +#define AARCH64_PATH "/tmp/binfmt_bpf_aarch64" +#define RELOC_DIR "/tmp/binfmt_reloc" +#define BINFMT_REG "/proc/sys/fs/binfmt_misc/register" +#define EXPECT "BPF_INTERP_RAN" + +static char testdir[512]; /* directory holding this test's built artifacts= */ + +static int copy_file(const char *src, const char *dst) +{ + char buf[4096]; + int in, out; + ssize_t n; + + in =3D open(src, O_RDONLY); + if (in < 0) + return -1; + out =3D open(dst, O_WRONLY | O_CREAT | O_TRUNC, 0755); + if (out < 0) { + close(in); + return -1; + } + while ((n =3D read(in, buf, sizeof(buf))) > 0) { + if (write(out, buf, n) !=3D n) { + close(in); + close(out); + return -1; + } + } + close(in); + close(out); + return n < 0 ? -1 : 0; +} + +/* A minimal 64-bit little-endian aarch64 ELF header, padded to the read s= ize. */ +static int create_fake_aarch64(const char *path) +{ + unsigned char hdr[256] =3D {0}; + int fd; + + hdr[0] =3D 0x7f; hdr[1] =3D 'E'; hdr[2] =3D 'L'; hdr[3] =3D 'F'; + hdr[4] =3D 2; /* ELFCLASS64 */ + hdr[5] =3D 1; /* ELFDATA2LSB */ + hdr[6] =3D 1; /* EV_CURRENT */ + hdr[16] =3D 2; /* e_type =3D ET_EXEC */ + hdr[18] =3D 183 & 0xff; /* e_machine =3D EM_AARCH64 */ + hdr[19] =3D (183 >> 8) & 0xff; + hdr[20] =3D 1; /* e_version */ + + fd =3D open(path, O_WRONLY | O_CREAT | O_TRUNC, 0755); + if (fd < 0) + return -1; + if (write(fd, hdr, sizeof(hdr)) !=3D (ssize_t)sizeof(hdr)) { + close(fd); + return -1; + } + close(fd); + return 0; +} + +static int register_entry(const char *name, const char *handler) +{ + char rule[128]; + int fd; + ssize_t n; + + snprintf(rule, sizeof(rule), ":%s:B:%s::::", name, handler); + fd =3D open(BINFMT_REG, O_WRONLY); + if (fd < 0) + return -1; + n =3D write(fd, rule, strlen(rule)); + close(fd); + return n < 0 ? -1 : 0; +} + +static void unregister_entry(const char *name) +{ + char path[128]; + int fd; + + snprintf(path, sizeof(path), "/proc/sys/fs/binfmt_misc/%s", name); + fd =3D open(path, O_WRONLY); + if (fd >=3D 0) { + if (write(fd, "-1", 2) < 0) + ; /* best effort */ + close(fd); + } +} + +static int check_output(const char *cmd, const char *expected) +{ + char buf[128]; + FILE *fp; + + fp =3D popen(cmd, "r"); + if (!fp) + return -1; + if (!fgets(buf, sizeof(buf), fp)) { + pclose(fp); + return -1; + } + pclose(fp); + return strncmp(buf, expected, strlen(expected)) ? -1 : 0; +} + +/* + * Load @objfile, attach its struct_ops map @handler (which publishes the + * handler), activate a 'B' entry named @entry that references it, run @ta= rget + * and check it produced @expect. + */ +static int run_case(const char *objfile, const char *handler, + const char *entry, const char *target, const char *expect) +{ + struct bpf_object *obj; + struct bpf_map *map; + struct bpf_link *link; + int ret =3D -1; + + obj =3D bpf_object__open_file(objfile, NULL); + if (!obj || libbpf_get_error(obj)) { + fprintf(stderr, "open %s failed\n", objfile); + return -1; + } + if (bpf_object__load(obj)) { + fprintf(stderr, "load %s failed (check dmesg for the verifier log)\n", + objfile); + goto close; + } + map =3D bpf_object__find_map_by_name(obj, handler); + if (!map) { + fprintf(stderr, "no struct_ops map '%s' in %s\n", handler, objfile); + goto close; + } + link =3D bpf_map__attach_struct_ops(map); + if (!link || libbpf_get_error(link)) { + fprintf(stderr, "attach struct_ops '%s' failed\n", handler); + goto close; + } + if (register_entry(entry, handler)) { + fprintf(stderr, "register 'B' entry '%s' failed\n", entry); + goto detach; + } + ret =3D check_output(target, expect); + unregister_entry(entry); +detach: + bpf_link__destroy(link); +close: + bpf_object__close(obj); + return ret; +} + +int main(void) +{ + char src[600], obj[600], appdst[600], interpdst[600]; + char exe[512]; + ssize_t n; + int fail =3D 0; + struct stat st; + + if (getuid() !=3D 0) { + fprintf(stderr, "Skipping: test must be run as root\n"); + return 4; /* KSFT_SKIP */ + } + + n =3D readlink("/proc/self/exe", exe, sizeof(exe) - 1); + if (n < 0) { + perror("readlink"); + return 1; + } + exe[n] =3D '\0'; + snprintf(testdir, sizeof(testdir), "%s", dirname(exe)); + + if (stat("/sys/fs/bpf", &st) < 0) + mkdir("/sys/fs/bpf", 0755); + mount("bpf", "/sys/fs/bpf", "bpf", 0, NULL); + if (access(BINFMT_REG, F_OK) < 0) + mount("binfmt_misc", "/proc/sys/fs/binfmt_misc", "binfmt_misc", 0, NULL); + + /* Shared test interpreter. */ + snprintf(src, sizeof(src), "%s/binfmt_bpf_interp", testdir); + if (copy_file(src, INTERP_PATH)) { + fprintf(stderr, "cannot install %s\n", INTERP_PATH); + return 1; + } + + /* Case 1: match a synthetic aarch64 header -> fixed interpreter. */ + printf("[*] case 1: match aarch64 header -> program-chosen interpreter\n"= ); + if (create_fake_aarch64(AARCH64_PATH)) { + fprintf(stderr, "cannot create %s\n", AARCH64_PATH); + return 1; + } + snprintf(obj, sizeof(obj), "%s/bpf_interp.bpf.o", testdir); + if (run_case(obj, "bpf_interp", "test_bpf_interp", AARCH64_PATH, EXPECT) = =3D=3D 0) + printf("[+] case 1 passed\n"); + else { + printf("[-] case 1 FAILED\n"); + fail =3D 1; + } + unlink(AARCH64_PATH); + + /* Case 2: $ORIGIN-relative PT_INTERP -> co-located interpreter. */ + printf("[*] case 2: $ORIGIN interpreter resolved relative to the binary\n= "); + mkdir(RELOC_DIR, 0755); + snprintf(appdst, sizeof(appdst), "%s/app", RELOC_DIR); + snprintf(interpdst, sizeof(interpdst), "%s/binfmt_bpf_interp", RELOC_DIR); + snprintf(src, sizeof(src), "%s/binfmt_bpf_app", testdir); + if (copy_file(src, appdst) || + copy_file(INTERP_PATH, interpdst)) { + fprintf(stderr, "cannot set up %s\n", RELOC_DIR); + fail =3D 1; + } else { + snprintf(obj, sizeof(obj), "%s/nix_origin.bpf.o", testdir); + if (run_case(obj, "nix_origin", "test_bpf_origin", appdst, EXPECT) =3D= =3D 0) + printf("[+] case 2 passed\n"); + else { + printf("[-] case 2 FAILED\n"); + fail =3D 1; + } + } + unlink(appdst); + unlink(interpdst); + rmdir(RELOC_DIR); + unlink(INTERP_PATH); + + if (!fail) + printf("[*] all binfmt_misc bpf cases passed\n"); + return fail; +} diff --git a/tools/testing/selftests/exec/bpf_interp.bpf.c b/tools/testing/= selftests/exec/bpf_interp.bpf.c new file mode 100644 index 000000000..ca98c3716 --- /dev/null +++ b/tools/testing/selftests/exec/bpf_interp.bpf.c @@ -0,0 +1,52 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * binfmt_misc_ops handler for the selftest's fixed-interpreter case: matc= h a + * 64-bit aarch64 ELF header from the prefetched buffer and route it to a = fixed + * interpreter chosen by the program. This is the portable, self-contained + * equivalent of routing a foreign binary to an emulator: it matches + * programmatically and computes the interpreter, but points at a test bin= ary + * the harness installs rather than a system emulator. + */ +#include "vmlinux.h" +#include +#include + +char _license[] SEC("license") =3D "GPL"; + +#define EI_CLASS 4 +#define ELFCLASS64 2 +#define EM_AARCH64 183 + +extern int bpf_binprm_set_interp(struct linux_binprm *bprm, const char *pa= th, + size_t path__sz) __ksym; + +SEC("struct_ops.s/load") +int BPF_PROG(bpf_interp_load, struct linux_binprm *bprm) +{ + /* + * Keep the path on the (writable) stack: bpf_binprm_set_interp() takes + * a sized memory arg and the verifier rejects a read-only .rodata + * buffer for it. The harness installs the interpreter at this path. + */ + char interp[] =3D "/tmp/binfmt_bpf_interp"; + __u16 machine; + + if (bprm->buf[0] !=3D 0x7f || bprm->buf[1] !=3D 'E' || + bprm->buf[2] !=3D 'L' || bprm->buf[3] !=3D 'F' || + bprm->buf[EI_CLASS] !=3D ELFCLASS64) + return 0; + + /* e_machine is a 16-bit little-endian field at offset 18. */ + machine =3D (__u8)bprm->buf[18] | ((__u16)(__u8)bprm->buf[19] << 8); + if (machine !=3D EM_AARCH64) + return 0; + + /* @path__sz includes the terminating NUL. */ + return bpf_binprm_set_interp(bprm, interp, sizeof(interp)) ?: 1; +} + +SEC(".struct_ops.link") +struct binfmt_misc_ops bpf_interp =3D { + .load =3D (void *)bpf_interp_load, + .name =3D "bpf_interp", +}; diff --git a/tools/testing/selftests/exec/nix_origin.bpf.c b/tools/testing/= selftests/exec/nix_origin.bpf.c new file mode 100644 index 000000000..a853d64f4 --- /dev/null +++ b/tools/testing/selftests/exec/nix_origin.bpf.c @@ -0,0 +1,179 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * nix_origin.bpf.c - $ORIGIN-relative PT_INTERP resolution + * + * A binfmt_misc_ops handler that makes relocatable (Nix-style) ELF + * binaries work: if PT_INTERP starts with "$ORIGIN/", the loader is + * resolved relative to the directory of the binary being executed and + * selected via bpf_binprm_set_interp(). Anything else is declined so + * regular binaries pass through untouched. + * + * Activate with: + * bpftool struct_ops register nix_origin.bpf.o /sys/fs/bpf + * echo ':nix-origin:B:nix_origin::::' > /proc/sys/fs/binfmt_misc/regist= er + */ +#include "vmlinux.h" +#include +#include + +char _license[] SEC("license") =3D "GPL"; + +#define PATH_MAX 4096 +#define EI_CLASS 4 +#define ELFCLASSXX 2 /* ELFCLASS64; flip to 1 for 32-bit */ +#define PT_INTERP 3 +#define MAX_PHDRS 64 + +#define ORIGIN "$ORIGIN" +#define ORIGIN_LEN (sizeof(ORIGIN) - 1) + +#define ENOENT 2 +#define ENAMETOOLONG 36 + +extern int bpf_dynptr_from_file(struct file *file, __u32 flags, + struct bpf_dynptr *ptr__uninit) __ksym; +extern int bpf_dynptr_file_discard(struct bpf_dynptr *dynptr) __ksym; +extern int bpf_path_d_path(const struct path *path, char *buf, + size_t buf__sz) __ksym; +extern int bpf_binprm_set_interp(struct linux_binprm *bprm, const char *pa= th, + size_t path__sz) __ksym; + +struct scratch { + char interp[PATH_MAX]; /* PT_INTERP as embedded in the binary */ + char path[PATH_MAX]; /* d_path of the binary, becomes the result */ +}; + +/* Keyed by pid: execs run concurrently and the program can sleep. */ +struct { + __uint(type, BPF_MAP_TYPE_HASH); + __uint(max_entries, 512); + __type(key, __u64); + __type(value, struct scratch); +} scratch_map SEC(".maps"); + +static const struct scratch zero_scratch; + +SEC("struct_ops.s/load") +int BPF_PROG(nix_origin_load, struct linux_binprm *bprm) +{ + __u32 isz, sfx, rsz, slash; + struct elf64_phdr phdr; + struct elf64_hdr ehdr; + struct bpf_dynptr dp; + struct scratch *sc; + bool found =3D false; + __u64 id; + int ret =3D 0, len, i; + + /* Cheap reject from the prefetched header. */ + if (bprm->buf[0] !=3D 0x7f || bprm->buf[1] !=3D 'E' || + bprm->buf[2] !=3D 'L' || bprm->buf[3] !=3D 'F' || + bprm->buf[EI_CLASS] !=3D ELFCLASSXX) + return 0; + + if (bpf_dynptr_from_file(bprm->file, 0, &dp)) + goto out; + + if (bpf_dynptr_read(&ehdr, sizeof(ehdr), &dp, 0, 0)) + goto out; + if (ehdr.e_phentsize !=3D sizeof(struct elf64_phdr)) + goto out; + + bpf_for(i, 0, ehdr.e_phnum) { + if (i >=3D MAX_PHDRS) + break; + if (bpf_dynptr_read(&phdr, sizeof(phdr), &dp, + ehdr.e_phoff + i * sizeof(phdr), 0)) + goto out; + if (phdr.p_type =3D=3D PT_INTERP) { + found =3D true; + break; + } + } + if (!found) + goto out; + + isz =3D phdr.p_filesz; + if (isz <=3D ORIGIN_LEN + 1 || isz >=3D sizeof(sc->interp)) + goto out; + /* + * The range check above compiles to a test on a zero-extended copy of + * the u64 p_filesz, so the verifier does not carry the bound to the + * dynptr_read() length below ("unbounded memory access"). Mask isz to + * the buffer size (a power of two) and force the masked value to be + * materialized with a barrier so the read uses the bounded register. + */ + isz &=3D sizeof(sc->interp) - 1; + barrier_var(isz); + + id =3D bpf_get_current_pid_tgid(); + if (bpf_map_update_elem(&scratch_map, &id, &zero_scratch, BPF_ANY)) + goto out; + sc =3D bpf_map_lookup_elem(&scratch_map, &id); + if (!sc) + goto out_del; + + if (bpf_dynptr_read(sc->interp, isz, &dp, phdr.p_offset, 0)) + goto out_del; + if (sc->interp[isz - 1] !=3D '\0') + goto out_del; + + /* Not "$ORIGIN/..."? Decline, the elf loader owns it. */ + if (sc->interp[0] !=3D '$' || sc->interp[1] !=3D 'O' || + sc->interp[2] !=3D 'R' || sc->interp[3] !=3D 'I' || + sc->interp[4] !=3D 'G' || sc->interp[5] !=3D 'I' || + sc->interp[6] !=3D 'N' || sc->interp[7] !=3D '/') + goto out_del; + + /* + * From here on the binary is ours: resolution failures fail the + * exec instead of falling back to binfmt_elf, which would resolve + * the literal "$ORIGIN/..." relative to the caller's cwd. + */ + ret =3D -ENOENT; + len =3D bpf_path_d_path(&bprm->file->f_path, sc->path, sizeof(sc->path)); + if (len <=3D 0 || len > sizeof(sc->path)) + goto out_del; + /* Unreachable or unlinked ("... (deleted)") binaries can't resolve. */ + if (sc->path[0] !=3D '/') + goto out_del; + + /* $ORIGIN =3D dirname of the binary. */ + slash =3D 0; + bpf_for(i, 1, len - 1) { + if (i >=3D sizeof(sc->path)) + break; + if (sc->path[i] =3D=3D '/') + slash =3D i; + } + + /* Splice the suffix (leading '/' and NUL included) onto the dir. */ + sfx =3D isz - ORIGIN_LEN; + rsz =3D slash + sfx; + if (rsz > sizeof(sc->path)) { + ret =3D -ENAMETOOLONG; + goto out_del; + } + bpf_for(i, 0, sfx) { + __u32 s =3D ORIGIN_LEN + i, d =3D slash + i; + + if (s >=3D sizeof(sc->interp) || d >=3D sizeof(sc->path)) + break; + sc->path[d] =3D sc->interp[s]; + } + + ret =3D bpf_binprm_set_interp(bprm, sc->path, rsz); + if (!ret) + ret =3D 1; +out_del: + bpf_map_delete_elem(&scratch_map, &id); +out: + bpf_dynptr_file_discard(&dp); + return ret; +} + +SEC(".struct_ops.link") +struct binfmt_misc_ops nix_origin =3D { + .load =3D (void *)nix_origin_load, + .name =3D "nix_origin", +}; --=20 2.51.2