From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-008.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-008.esa.us-west-2.outbound.mail-perimeter.amazon.com [52.42.203.116]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 75A7B438023; Wed, 23 Sep 2026 05:40:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=52.42.203.116 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142007; cv=none; b=KptKm/OVRrEMVzDVwhst3Hh6trqb0UZR3v6C7qY5qBtJW0PYPq7IRSMoRcuVU8/2NhOc/z+5yAtsYMFM2d+Sz16p3q/PxwaEwCb3pd9elQyd67Gs3Y1su9SDL94YdYztxZV/UoWvNxb1DZt7riW2WXJSvZlFamwysGHQPFvqqj8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142007; c=relaxed/simple; bh=ry4jQGL2wKU9mNSo2ehUuXrA6vdswJ5kRMdi0BaXGPY=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=mtEe9y+/aYW/myTVnG+naDcy2CLyppXwYNLFYhh6FYS14/1YDi7e9JqxW1IDq9y6IRyUuJil8LvjzYq5FH2CS8y48Igx6QuTFTw3VPj2NRfs2Hfeyss9YXffWi1qHcZp3p2FLsv8E8c0JMVa4UEz0BqXRLTT4dih5roNrTjGGPk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=DzviD7TV; arc=none smtp.client-ip=52.42.203.116 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="DzviD7TV" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142005; x=1821678005; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=/T3L6Erhez9jj6KkQNlY833QrYTQpIlvZMQo38B3kLc=; b=DzviD7TVLx4okwPnnjwCKrSMeNiBtH4j9gI/o+XB8uei0u3mY6Dpy0ZW oSNTWi5+xtVQBDLW1zOXUeZwOmJMQOtASH1u2tJpqAFKs89pOGc+9ekHT ZNDfxPqzFibd3ruz3vRnB6U6MHJ+dC+mq9/7SqSm0bzLrubxwfMkdJvM8 jpbuoqX+QkeLbGjpsLcywnRlXyUZMXdOikxbXeE6KVE9NyAAlUKQXegov 58OL9oC92Tu9hsHFErhVTTdRTpAQDPmSJEyVgXp2P5nXAxQCpqmW8u9UO X2GPpxgX+Y8qeJe+gEJXYrG/9o9lV0chwCeDVRvlMOfs9Xr32vJ2547g5 Q==; X-CSE-ConnectionGUID: L+gsmlieQC26TX0aL6CBFQ== X-CSE-MsgGUID: tfhBL2k4TFuylusL659bKg== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29438704" Received: from ip-10-5-0-115.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.0.115]) by internal-pdx-out-008.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:40:05 +0000 Received: from EX19MTAUWA002.ant.amazon.com [205.251.233.234:17433] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.1.232:2525] with esmtp (Farcaster) id 62ed8f0f-62c8-4742-83aa-d70b362ac146; Wed, 23 Sep 2026 05:40:04 +0000 (UTC) X-Farcaster-Flow-ID: 62ed8f0f-62c8-4742-83aa-d70b362ac146 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWA002.ant.amazon.com (10.250.64.202) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:04 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:04 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 1/6] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Date: Wed, 23 Sep 2026 05:39:43 +0000 Message-ID: <20260923053948.30617-2-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D038UWB003.ant.amazon.com (10.13.139.157) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Make btf_parse_module() take the vmlinux BTF as an argument instead of fetching it with bpf_get_btf_vmlinux(), and add a data_owned flag: when set, the passed .BTF data is an already kvmalloc()ed copy that the new btf takes ownership of on success (on failure the caller keeps it). Factor the sysfs file creation out of the module notifier into btf_module_sysfs_add() and the teardown into btf_module_free(), and set btf_mod->module right after the allocation rather than under the mutex. No functional change. This prepares for CONFIG_DEBUG_INFO_BTF=3Dm, where a module can be loaded before the vmlinux BTF is available: its .BTF is then copied and exposed in sysfs first and parsed later, at which point the parser must take the copy as is so that the sysfs file keeps pointing at valid data. Signed-off-by: Jay Wang --- kernel/bpf/btf.c | 105 +++++++++++++++++++++++++++++------------------ 1 file changed, 64 insertions(+), 41 deletions(-) diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index 4a1fa4fbdf4e..a6634237dc89 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -6516,16 +6516,20 @@ __u32 btf_relocate_id(const struct btf *btf, __u32 = id) =20 #ifdef CONFIG_DEBUG_INFO_BTF_MODULES =20 -static struct btf *btf_parse_module(const char *module_name, const void *d= ata, - unsigned int data_size, void *base_data, - unsigned int base_data_size) +/* + * Parse split module BTF against @vmlinux_btf. @data is the module's .BTF + * section; if @data_owned, it is an already kvmalloc()ed copy that the new + * btf takes ownership of on success (on failure the caller keeps it). + */ +static struct btf *btf_parse_module(const char *module_name, struct btf *v= mlinux_btf, + void *data, unsigned int data_size, bool data_owned, + void *base_data, unsigned int base_data_size) { - struct btf *btf =3D NULL, *vmlinux_btf, *base_btf =3D NULL; + struct btf *btf =3D NULL, *base_btf =3D NULL; struct btf_verifier_env *env =3D NULL; struct bpf_verifier_log *log; int err =3D 0; =20 - vmlinux_btf =3D bpf_get_btf_vmlinux(); if (IS_ERR(vmlinux_btf)) return vmlinux_btf; if (!vmlinux_btf) @@ -6562,7 +6566,10 @@ static struct btf *btf_parse_module(const char *modu= le_name, const void *data, btf->named_start_id =3D 0; strscpy(btf->name, module_name); =20 - btf->data =3D kvmemdup(data, data_size, GFP_KERNEL | __GFP_NOWARN); + if (data_owned) + btf->data =3D data; + else + btf->data =3D kvmemdup(data, data_size, GFP_KERNEL | __GFP_NOWARN); if (!btf->data) { err =3D -ENOMEM; goto errout; @@ -6605,7 +6612,8 @@ static struct btf *btf_parse_module(const char *modul= e_name, const void *data, if (!IS_ERR(base_btf) && base_btf !=3D vmlinux_btf) btf_free(base_btf); if (btf) { - kvfree(btf->data); + if (!data_owned) + kvfree(btf->data); kvfree(btf->types); kfree(btf); } @@ -8610,6 +8618,48 @@ static DEFINE_MUTEX(btf_module_mutex); =20 static void purge_cand_cache(struct btf *btf); =20 +static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *na= me, + void *data, size_t data_size) +{ + struct bin_attribute *attr; + int err; + + if (!IS_ENABLED(CONFIG_SYSFS)) + return 0; + + attr =3D kzalloc_obj(*attr); + if (!attr) + return -ENOMEM; + + sysfs_bin_attr_init(attr); + attr->attr.name =3D name; + attr->attr.mode =3D 0444; + attr->size =3D data_size; + attr->private =3D data; + attr->read =3D sysfs_bin_attr_simple_read; + + err =3D sysfs_create_bin_file(btf_kobj, attr); + if (err) { + pr_warn("failed to register module [%s] BTF in sysfs: %d\n", + name, err); + kfree(attr); + return err; + } + + btf_mod->sysfs_attr =3D attr; + return 0; +} + +static void btf_module_free(struct btf_module *btf_mod) +{ + if (btf_mod->sysfs_attr) + sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr); + purge_cand_cache(btf_mod->btf); + btf_put(btf_mod->btf); + kfree(btf_mod->sysfs_attr); + kfree(btf_mod); +} + static int btf_module_notify(struct notifier_block *nb, unsigned long op, void *module) { @@ -8630,7 +8680,10 @@ static int btf_module_notify(struct notifier_block *= nb, unsigned long op, err =3D -ENOMEM; goto out; } - btf =3D btf_parse_module(mod->name, mod->btf_data, mod->btf_data_size, + btf_mod->module =3D module; + + btf =3D btf_parse_module(mod->name, bpf_get_btf_vmlinux(), + mod->btf_data, mod->btf_data_size, false, mod->btf_base_data, mod->btf_base_data_size); if (IS_ERR(btf)) { kfree(btf_mod); @@ -8652,37 +8705,12 @@ static int btf_module_notify(struct notifier_block = *nb, unsigned long op, =20 purge_cand_cache(NULL); mutex_lock(&btf_module_mutex); - btf_mod->module =3D module; btf_mod->btf =3D btf; list_add(&btf_mod->list, &btf_modules); mutex_unlock(&btf_module_mutex); =20 - if (IS_ENABLED(CONFIG_SYSFS)) { - struct bin_attribute *attr; - - attr =3D kzalloc_obj(*attr); - if (!attr) - goto out; - - sysfs_bin_attr_init(attr); - attr->attr.name =3D btf->name; - attr->attr.mode =3D 0444; - attr->size =3D btf->data_size; - attr->private =3D btf->data; - attr->read =3D sysfs_bin_attr_simple_read; - - err =3D sysfs_create_bin_file(btf_kobj, attr); - if (err) { - pr_warn("failed to register module [%s] BTF in sysfs: %d\n", - mod->name, err); - kfree(attr); - err =3D 0; - goto out; - } - - btf_mod->sysfs_attr =3D attr; - } - + /* not fatal, the module BTF is usable without the sysfs file */ + btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size); break; case MODULE_STATE_LIVE: mutex_lock(&btf_module_mutex); @@ -8709,12 +8737,7 @@ static int btf_module_notify(struct notifier_block *= nb, unsigned long op, */ btf_free_id(btf_mod->btf); list_del(&btf_mod->list); - if (btf_mod->sysfs_attr) - sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr); - purge_cand_cache(btf_mod->btf); - btf_put(btf_mod->btf); - kfree(btf_mod->sysfs_attr); - kfree(btf_mod); + btf_module_free(btf_mod); break; } mutex_unlock(&btf_module_mutex); --=20 2.47.3 From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-004.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-004.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.246.77.92]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA6E321A459; Wed, 23 Sep 2026 05:40:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.246.77.92 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142022; cv=none; b=JhURkVErdc67O750bXxfLlFNm3alLyKcps1SKpetES0+Frfew784I3BxzE6dqSBf0MQsYKkZleXGtVbEKc9okIuzwCr69m+I8hRiznaaLisCx3iPpFUOzzAkhlK1ufBLZKuTYhirKYKfmxF6eWRCXGeo0r7yLbA7tSmzaNyrj9A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142022; c=relaxed/simple; bh=hDWO4kKtFW0ws5P8AqsAJnLaI0zPunaAAMiFodM6fPs=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gQotEb6UIj4UXVnC9fWGr0HouFb5ztawn9vd4Q3wWK5YOMwF8OhcjHYnuekKf5wXoMWxFafj4EiE5io9b6f7BZgmoLIAJNabzqYWJbDQHFxjXYb56z6GsyP+PP3QYjsWrSqPDgVidDiLZrJ6xkkZShxCSH/GG4rDXAhjtEcd6oU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=YGjjWEst; arc=none smtp.client-ip=44.246.77.92 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="YGjjWEst" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142020; x=1821678020; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=oYdFmiyHha0KnlTs+SZuSj16NqNIECti33IxnnQ3JPY=; b=YGjjWEstgaHx5ilCKCX4WRCmAC6cbhJKETdhoGreLxPp+9Z5GhsHJJEC 9nczxidqs7LGaDMQzkcMNOtQTYHJjOTLd5+k7YtKOUiWyjSoVsVdePVT/ vzrgn17weyJ18mFwrU6H0/EWX5vCWfzKEKGcdBvRTQodcLuHWKUVorjNN IxQamSicvsCeTFOIXCQENx2AWUqPnhtL/Dt6hWjKwIVXoSgMfND6gw052 lSC4zI9A4QLi2h5LlnM5DuC+5A9Tc5u3Ehaxdp772cocv1pKSL/7xV20v 0mVkcaODKBLlh4dUqYU6pfQOIFZbk0niaYe5Ki6daDdxnXcnfTZKFsutC A==; X-CSE-ConnectionGUID: krfUjUrZT1SyzXdQGrDikQ== X-CSE-MsgGUID: 1eaHeKSaT1CYavLe1CkrUA== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29402263" Received: from ip-10-5-6-203.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.6.203]) by internal-pdx-out-004.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:40:20 +0000 Received: from EX19MTAUWB001.ant.amazon.com [205.251.233.51:8335] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.49.70:2525] with esmtp (Farcaster) id b07ed87e-ffeb-4ce0-bea5-2fd6a0aba4f1; Wed, 23 Sep 2026 05:40:20 +0000 (UTC) X-Farcaster-Flow-ID: b07ed87e-ffeb-4ce0-bea5-2fd6a0aba4f1 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWB001.ant.amazon.com (10.250.64.248) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:20 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:19 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 2/6] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Date: Wed, 23 Sep 2026 05:39:44 +0000 Message-ID: <20260923053948.30617-3-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D036UWC001.ant.amazon.com (10.13.139.233) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Split __register_btf_kfunc_id_set(), register_btf_id_dtor_kfuncs() and __register_bpf_struct_ops() into the part that looks up the BTF for the owner and the part that adds the registration to a given BTF: btf_kfunc_id_set_add(), btf_dtor_kfuncs_add() and btf_struct_ops_add(). In is_valid_value_type(), look up bpf_struct_ops_common_value in the btf the function was given rather than in the btf_vmlinux global. The id is a vmlinux id and a module BTF resolves it through its base, so the result is the same; the function already uses the passed btf for every other lookup. No functional change. With CONFIG_DEBUG_INFO_BTF=3Dm, registrations made from initcalls before the vmlinux BTF is available are queued and applied later by the BTF parsing code, which needs the add-to-this-btf half on its own; the struct_ops ones are applied before the parsed vmlinux BTF is published, i.e. while btf_vmlinux is still NULL. Signed-off-by: Jay Wang --- kernel/bpf/bpf_struct_ops.c | 3 +- kernel/bpf/btf.c | 91 ++++++++++++++++++++++--------------- 2 files changed, 57 insertions(+), 37 deletions(-) diff --git a/kernel/bpf/bpf_struct_ops.c b/kernel/bpf/bpf_struct_ops.c index 1178acd72296..bf3004908d15 100644 --- a/kernel/bpf/bpf_struct_ops.c +++ b/kernel/bpf/bpf_struct_ops.c @@ -103,7 +103,8 @@ static bool is_valid_value_type(struct btf *btf, s32 va= lue_id, } member =3D btf_type_member(vt); mt =3D btf_type_by_id(btf, member->type); - common_value_type =3D btf_type_by_id(btf_vmlinux, + /* a vmlinux id resolves through the base BTF of a module BTF too */ + common_value_type =3D btf_type_by_id(btf, st_ops_ids[IDX_ST_OPS_COMMON_VALUE_ID]); if (mt !=3D common_value_type) { pr_warn("The first member of %s should be bpf_struct_ops_common_value\n", diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index a6634237dc89..c3c1421208b4 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -9313,11 +9313,26 @@ u32 *btf_kfunc_is_modify_return(const struct btf *b= tf, u32 kfunc_btf_id, return btf_kfunc_id_set_contains(btf, BTF_KFUNC_HOOK_FMODRET, kfunc_btf_i= d); } =20 +static int btf_kfunc_id_set_add(struct btf *btf, enum btf_kfunc_hook hook, + const struct btf_kfunc_id_set *kset) +{ + int ret, i; + + for (i =3D 0; i < kset->set->cnt; i++) { + ret =3D btf_check_kfunc_protos(btf, btf_relocate_id(btf, kset->set->pair= s[i].id), + kset->set->pairs[i].flags); + if (ret) + return ret; + } + + return btf_populate_kfunc_set(btf, hook, kset); +} + static int __register_btf_kfunc_id_set(enum btf_kfunc_hook hook, const struct btf_kfunc_id_set *kset) { struct btf *btf; - int ret, i; + int ret; =20 btf =3D btf_get_module_btf(kset->owner); if (!btf) @@ -9325,16 +9340,7 @@ static int __register_btf_kfunc_id_set(enum btf_kfun= c_hook hook, if (IS_ERR(btf)) return PTR_ERR(btf); =20 - for (i =3D 0; i < kset->set->cnt; i++) { - ret =3D btf_check_kfunc_protos(btf, btf_relocate_id(btf, kset->set->pair= s[i].id), - kset->set->pairs[i].flags); - if (ret) - goto err_out; - } - - ret =3D btf_populate_kfunc_set(btf, hook, kset); - -err_out: + ret =3D btf_kfunc_id_set_add(btf, hook, kset); btf_put(btf); return ret; } @@ -9426,21 +9432,13 @@ static int btf_check_dtor_kfuncs(struct btf *btf, c= onst struct btf_id_dtor_kfunc return 0; } =20 -/* This function must be invoked only from initcalls/module init functions= */ -int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32= add_cnt, - struct module *owner) +static int btf_dtor_kfuncs_add(struct btf *btf, const struct btf_id_dtor_k= func *dtors, + u32 add_cnt) { struct btf_id_dtor_kfunc_tab *tab; - struct btf *btf; u32 tab_cnt, i; int ret; =20 - btf =3D btf_get_module_btf(owner); - if (!btf) - return check_btf_kconfigs(owner, "dtor kfuncs"); - if (IS_ERR(btf)) - return PTR_ERR(btf); - if (add_cnt >=3D BTF_DTOR_KFUNC_MAX_CNT) { pr_err("cannot register more than %d kfunc destructors\n", BTF_DTOR_KFUN= C_MAX_CNT); ret =3D -E2BIG; @@ -9497,6 +9495,23 @@ int register_btf_id_dtor_kfuncs(const struct btf_id_= dtor_kfunc *dtors, u32 add_c end: if (ret) btf_free_dtor_kfunc_tab(btf); + return ret; +} + +/* This function must be invoked only from initcalls/module init functions= */ +int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32= add_cnt, + struct module *owner) +{ + struct btf *btf; + int ret; + + btf =3D btf_get_module_btf(owner); + if (!btf) + return check_btf_kconfigs(owner, "dtor kfuncs"); + if (IS_ERR(btf)) + return PTR_ERR(btf); + + ret =3D btf_dtor_kfuncs_add(btf, dtors, add_cnt); btf_put(btf); return ret; } @@ -10135,32 +10150,36 @@ bpf_struct_ops_find(struct btf *btf, u32 type_id) return NULL; } =20 -int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops) +static int btf_struct_ops_add(struct btf *btf, struct bpf_struct_ops *st_o= ps) { struct bpf_verifier_log *log; - struct btf *btf; - int err =3D 0; - - btf =3D btf_get_module_btf(st_ops->owner); - if (!btf) - return check_btf_kconfigs(st_ops->owner, "struct_ops"); - if (IS_ERR(btf)) - return PTR_ERR(btf); + int err; =20 log =3D kzalloc_obj(*log, GFP_KERNEL | __GFP_NOWARN); - if (!log) { - err =3D -ENOMEM; - goto errout; - } + if (!log) + return -ENOMEM; =20 log->level =3D BPF_LOG_KERNEL; =20 err =3D btf_add_struct_ops(btf, st_ops, log); =20 -errout: kfree(log); - btf_put(btf); + return err; +} =20 +int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops) +{ + struct btf *btf; + int err; + + btf =3D btf_get_module_btf(st_ops->owner); + if (!btf) + return check_btf_kconfigs(st_ops->owner, "struct_ops"); + if (IS_ERR(btf)) + return PTR_ERR(btf); + + err =3D btf_struct_ops_add(btf, st_ops); + btf_put(btf); return err; } EXPORT_SYMBOL_GPL(__register_bpf_struct_ops); --=20 2.47.3 From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com [34.218.115.239]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9625121A459; Wed, 23 Sep 2026 05:40:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=34.218.115.239 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142038; cv=none; b=UVUiuM25ubndXaSeAbQ9bQH7cB0jOc5DRHneemZ2elkGTKtHg0tWz/vBFzjW7v/TuEeyy//Y3BqMF0d2xuUDPJRPaUhCGoYDyaP05nYyt3gloFvR/u7+TZP3ZOe8ytrbgkz2m5fWdwF76fcM59JjSUF38zr7Bqip/IlQlON67jo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142038; c=relaxed/simple; bh=Tfn8qvtEMA3bVae84++iK1W+N2jp6kiBt73YYSQudGM=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=b37LIfOsjl5y7fE1d18yTEzWXZwD1ihg8KRixHcvANy2Ef8Wvn1VkImi08jaPRjIF5SZhU7CxqH1/eByplcKCzbc0LRJvpaN+G1ES1jht0yb+WCbxh2+WbJNWRx6s3SKxqEGfngDvX3jl1x/KAIS1OAV4kF8J1u1UAwwyAiakPU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=dWTNutNd; arc=none smtp.client-ip=34.218.115.239 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="dWTNutNd" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142036; x=1821678036; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=mNOFMcm0ALQtx8uNz3E+ErHcgScbjyXVRxmwiRpaAqs=; b=dWTNutNdYX5MvjpmeEqbwtcNvivNU1edMSQjbUF5flQjjwkqIYGb1HV/ 1M1J4z7koUCEuDBoSX9rCVz+SqGVn/39SuZEaE46JjziSMwv9Zh0Bi9QJ DYXAKYWLkmrwQy54JLcRJMYjGBMQGlDX7cFhYb3gQwOjaFvjP7/vkYEln NG/qBassS/nYcshneYdUvKWXAOMlbrh2Qi2qeJT/B6ewBXZEvyudGSI5T tnBSyp1O9J5jKKkRaMKpDrTiB4i3eDu4aBOL2pSHZgy7fzyxByIoqSWgv WM3ErQsTus/dAca8nPiPmh92b9yLyoXXZgX+MKG1Qn0Abxjq7mA//PezH Q==; X-CSE-ConnectionGUID: SBWtaKhPSbOjzZ252NTu3w== X-CSE-MsgGUID: XhJ/DhJTTymXN7+atpYlng== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29195618" Received: from ip-10-5-9-48.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.9.48]) by internal-pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:40:36 +0000 Received: from EX19MTAUWC002.ant.amazon.com [205.251.233.111:23925] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.60.208:2525] with esmtp (Farcaster) id 8069c48d-6f98-4c16-adc6-07f7fb2f3ce0; Wed, 23 Sep 2026 05:40:35 +0000 (UTC) X-Farcaster-Flow-ID: 8069c48d-6f98-4c16-adc6-07f7fb2f3ce0 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWC002.ant.amazon.com (10.250.64.143) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:35 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:35 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 3/6] bpf: fetch the vmlinux BTF where kernel types enter a program Date: Wed, 23 Sep 2026 05:39:45 +0000 Message-ID: <20260923053948.30617-4-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D038UWB004.ant.amazon.com (10.13.139.177) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" bpf_check() fetches the vmlinux BTF up front for every program, whether the program uses kernel types or not. With the upcoming CONFIG_DEBUG_INFO_BTF=3Dm that fetch loads a module and parses 5 MiB of BTF, and since systemd loads socket filters at boot, it would happen on every system, whether anything uses BTF or not. Stop fetching up front and fetch at the points where kernel types enter the verifier state instead: - bpf_add_kfunc_call(), for the first kfunc call of a program; - check_pseudo_btf_id(), for ldimm64 of a kernel variable; - check_ptr_to_map_access(), for accessing a map pointer's fields; - check_helper_call(), when the helper's prototype takes or returns a PTR_TO_BTF_ID (helper_uses_vmlinux_btf()). Together with the existing fetch in bpf_prog_load() for attach_btf and the struct_ops map creation, every PTR_TO_BTF_ID register a program can hold originates from one of these sites. A program that uses none of them, such as a socket filter, no longer touches the vmlinux BTF. bpf_snprintf_btf() and bpf_seq_printf_btf() run in program context and cannot afford a fetch that may sleep. Add bpf_peek_btf_vmlinux(), which returns the parsed vmlinux BTF or NULL without parsing anything, and use it there; the helpers fail with -EINVAL if the BTF is not parsed, as they do on a kernel without BTF. With CONFIG_DEBUG_INFO_BTF=3Dy the vmlinux BTF is parsed at boot by the first kfunc registration, so nothing changes. Without BTF, a helper that takes or returns a kernel pointer is now rejected with -ENOTSUPP at the call rather than with -EINVAL for its zero return type id. Signed-off-by: Jay Wang --- include/linux/bpf.h | 1 + kernel/bpf/verifier.c | 54 +++++++++++++++++++++++++++++++++++----- kernel/trace/bpf_trace.c | 3 ++- 3 files changed, 51 insertions(+), 7 deletions(-) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index e7c5e203eddd..a3c4caad5dfc 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -3165,6 +3165,7 @@ static inline s32 bpf_call_args_imm(s16 idx) #endif =20 struct btf *bpf_get_btf_vmlinux(void); +struct btf *bpf_peek_btf_vmlinux(void); =20 /* Map specifics */ struct xdp_frame; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index a7c9e2d8965d..2425ea74b61d 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -2873,7 +2873,8 @@ int bpf_add_kfunc_call(struct bpf_verifier_env *env, = u32 func_id, u16 offset) tab =3D prog_aux->kfunc_tab; btf_tab =3D prog_aux->kfunc_btf_tab; if (!tab) { - if (!btf_vmlinux) { + /* with CONFIG_DEBUG_INFO_BTF=3Dm this is where the vmlinux BTF gets loa= ded */ + if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) { verbose(env, "calling kernel function is not supported without CONFIG_D= EBUG_INFO_BTF\n"); return -ENOTSUPP; } @@ -6257,7 +6258,8 @@ static int check_ptr_to_map_access(struct bpf_verifie= r_env *env, u32 btf_id; int ret; =20 - if (!btf_vmlinux) { + /* with CONFIG_DEBUG_INFO_BTF=3Dm this is where the vmlinux BTF gets load= ed */ + if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) { verbose(env, "map_ptr access not supported without CONFIG_DEBUG_INFO_BTF= \n"); return -ENOTSUPP; } @@ -11568,6 +11570,20 @@ static int release_reg(struct bpf_verifier_env *en= v, struct bpf_reg_state *reg, return err; } =20 +/* Does calling @fn bring kernel BTF types into the program state? */ +static bool helper_uses_vmlinux_btf(const struct bpf_func_proto *fn) +{ + int i; + + if (base_type(fn->ret_type) =3D=3D RET_PTR_TO_BTF_ID) + return true; + for (i =3D 0; i < MAX_BPF_FUNC_ARGS; i++) { + if (base_type(fn->arg_type[i]) =3D=3D ARG_PTR_TO_BTF_ID) + return true; + } + return false; +} + static int check_helper_call(struct bpf_verifier_env *env, struct bpf_insn= *insn, int *insn_idx_p) { @@ -11637,6 +11653,16 @@ static int check_helper_call(struct bpf_verifier_e= nv *env, struct bpf_insn *insn return err; } =20 + /* + * Helpers that take or return kernel BTF pointers need the vmlinux + * BTF; with CONFIG_DEBUG_INFO_BTF=3Dm this is where it gets loaded. + */ + if (helper_uses_vmlinux_btf(fn) && IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))= { + verbose(env, "helper %s#%d is not supported without vmlinux BTF\n", + func_id_name(func_id), func_id); + return -ENOTSUPP; + } + if (fn->might_sleep && !in_sleepable_context(env)) { verbose(env, "sleepable helper %s#%d in %s\n", func_id_name(func_id), fu= nc_id, non_sleepable_context_description(env)); @@ -19235,12 +19261,13 @@ static int check_pseudo_btf_id(struct bpf_verifie= r_env *env, return -EINVAL; } } else { - if (!btf_vmlinux) { + /* with CONFIG_DEBUG_INFO_BTF=3Dm this is where the vmlinux BTF gets loa= ded */ + btf =3D bpf_get_btf_vmlinux(); + if (IS_ERR_OR_NULL(btf)) { verbose(env, "kernel is missing BTF, make sure CONFIG_DEBUG_INFO_BTF=3D= y is specified in Kconfig.\n"); return -EINVAL; } - btf_get(btf_vmlinux); - btf =3D btf_vmlinux; + btf_get(btf); } =20 err =3D __check_pseudo_btf_id(env, insn, aux, btf); @@ -21153,6 +21180,17 @@ struct btf *bpf_get_btf_vmlinux(void) return btf; } =20 +/* + * The vmlinux BTF if it has been parsed already, else NULL. Unlike + * bpf_get_btf_vmlinux() this never loads or parses anything, so it is safe + * to call from a running BPF program. + */ +struct btf *bpf_peek_btf_vmlinux(void) +{ + /* Pairs with the smp_store_release() in bpf_get_btf_vmlinux() */ + return smp_load_acquire(&btf_vmlinux); +} + /* * The add_fd_from_fd_array() is executed only if fd_array_cnt is non-zero= . In * this case expect that every file descriptor in the array is either a ma= p or @@ -21724,7 +21762,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_at= tr *attr, bpfptr_t uattr, if (ret) goto err_prep; =20 - bpf_get_btf_vmlinux(); + /* + * The vmlinux BTF is not fetched up front: with CONFIG_DEBUG_INFO_BTF=3Dm + * it is loaded on demand, at the points where kernel types enter the + * program (attach_btf, kfuncs, ksyms, map pointers, BTF-typed helpers). + */ =20 /* Serialize verification of unprivileged programs. */ if (!is_priv) diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c index 195f78db9bda..c022b2877f0b 100644 --- a/kernel/trace/bpf_trace.c +++ b/kernel/trace/bpf_trace.c @@ -1015,7 +1015,8 @@ static int bpf_btf_printf_prepare(struct btf_ptr *ptr= , u32 btf_ptr_size, if (btf_ptr_size !=3D sizeof(struct btf_ptr)) return -EINVAL; =20 - *btf =3D bpf_get_btf_vmlinux(); + /* Called from a running program: only use the BTF if it is parsed. */ + *btf =3D bpf_peek_btf_vmlinux(); =20 if (IS_ERR_OR_NULL(*btf)) return IS_ERR(*btf) ? PTR_ERR(*btf) : -EINVAL; --=20 2.47.3 From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.246.68.102]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C09F1418A2D; Wed, 23 Sep 2026 05:40:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.246.68.102 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142054; cv=none; b=RKMifVzzay6gsiSnsi6f55knOHY5gWxIHYAiI0IGbOZ/zg3nLZFjdIMwvsm1QnkHBnokLVQjtUiYz6hn6TfeJJsG6Nmm/NqurAHRt0Se9AHHNYYmGSNnzrNTcbhBTyexKHT/qgAyPyE5p9ucfRHeykPWIOB/q13ieoEFzePXWHI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142054; c=relaxed/simple; bh=Hw5gkQ2orDU/8pZftu6havBJSOrFzrWOnv5UYWuVba0=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=KzZVT2PIP7c3t9rDgfQDOaqNLCsJtpHuzGm2EIIL5I0uyVWPTi7OcuLvEqFWHCO8SuLrgRAxZRLjLhELWJTOVkKHb6ce9U2ByM2NsLvVSgiLiIenLP00vGgj0dTz4K06OvmAiN9+5YWysxf6ri6JuC1/5X7Jj+zUE4GsgVnWWHc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=lQPhofLI; arc=none smtp.client-ip=44.246.68.102 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="lQPhofLI" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142052; x=1821678052; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=4ExtyunNGCNzKwBq3DLx8FFZogSxkEtEWhjPwJuhR+g=; b=lQPhofLIvt+nBXH3lQAL0RSYqcvebSXmBO0fOvQTBQY09bCGPKv1Zd6F ijRlN1qsRqKEmBBWYA4Cmh7y8PhyGJRmgMMSHlSIpKp1J682MY0SqR8tv ONctXxY9FMqlnU5VjH9ebcbNSpUq2v8QEhxAncdO8iLZoxKjrd+7HHhI7 u6DTsmm2hh2ASSvcf4cW7YPqQzMb4YiS9fpXX6+pbJn0jUKR8p84rp7mc hZfh4HaJXy+Ny8DFx3nTKlwIBRytSpBfmK7bmsplMgBVpeliJ2h9948aZ uS9H6k0Sa5Ge0ImIxHwK04Hh6yBlTxd/Pqvs680v3xtuAGUUCcjI2UEnI g==; X-CSE-ConnectionGUID: lps+hww6TV+xy4EWY+sQoQ== X-CSE-MsgGUID: IRDXTvs9SXeWKIu7bELnzg== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29318518" Received: from ip-10-5-6-203.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.6.203]) by internal-pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:40:52 +0000 Received: from EX19MTAUWA001.ant.amazon.com [205.251.233.182:10800] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.46.137:2525] with esmtp (Farcaster) id 719cb39e-d607-431e-b29a-ce0313174c3e; Wed, 23 Sep 2026 05:40:52 +0000 (UTC) X-Farcaster-Flow-ID: 719cb39e-d607-431e-b29a-ce0313174c3e Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWA001.ant.amazon.com (10.250.64.204) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:51 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:40:51 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 4/6] bpf: take the vmlinux BTF from the btf_vmlinux module Date: Wed, 23 Sep 2026 05:39:46 +0000 Message-ID: <20260923053948.30617-5-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D038UWB001.ant.amazon.com (10.13.139.148) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Add the runtime side of delivering the vmlinux BTF as a module: with CONFIG_DEBUG_INFO_BTF=3Dm the BTF is carried by a module named btf_vmlinux and installed by the BTF module notifier when it loads, and whoever needs the BTF first loads the module. Nothing in this patch is reachable yet: CONFIG_DEBUG_INFO_BTF is still a bool and every new path is under IS_MODULE(CONFIG_DEBUG_INFO_BTF); the kbuild side and the Kconfig change follow. With CONFIG_DEBUG_INFO_BTF=3Dy the vmlinux BTF, 5.4 MiB on x86-64 with a distribution config, is part of the kernel image and resident from boot whether anything uses it or not. Most systems never do. Carrying it in a module that is loaded on first use makes the memory a cost of using BTF rather than of having a kernel that supports it. btf_vmlinux_data() hands out the raw vmlinux BTF: from __start_BTF with =3Dy, or with =3Dm from a vmalloc_user() copy that the notifier makes when the btf_vmlinux module loads. If the copy is not there and the caller asked to load, it calls request_module("btf_vmlinux"); the notifier installs the copy before init_module() returns, so the data is either present afterwards or the module is not available (yet). The result is not cached, a later call retries. btf_parse_vmlinux() and bpf_get_btf_vmlinux() use it; the latter loads outside btf_vmlinux_lock so that the notifier is never blocked by the caller, and returns NULL like a kernel without BTF when the module cannot be loaded. The verifier trusts the BTF as the description of this kernel's types, so the notifier only accepts a payload whose size and SHA-256 match the values linked into the kernel as .BTF.meta (struct btf_vmlinux_meta, filled in by scripts/gen-btf.sh in a later patch). A carrier from another build is refused with -EINVAL even if vermagic lets it load, and a second carrier is ignored. The copy is never freed: as with =3Dy, the BTF stays for the lifetime of the kernel, and the carrier has no exit. The module notifier, btf_parse_module() and the btf_data fields in struct module are compiled for CONFIG_DEBUG_INFO_BTF_MODULES or =3Dm; with =3Dm and no module BTF, the notifier only recognizes the carrier. /sys/kernel/btf/vmlinux exists from boot with its final size, which is known from .BTF.meta before the BTF is loaded; the first read() or mmap() loads it. mmap() uses remap_vmalloc_range() on the copy. This keeps stat() working before the load, which is what the btf_sysfs selftest does. BPF_BTF_GET_NEXT_ID loads the BTF too: kernel BTFs get their ids when the vmlinux BTF is parsed, and whoever enumerates BTF ids wants them. Signed-off-by: Jay Wang --- include/linux/btf.h | 2 + include/linux/module.h | 2 +- kernel/bpf/btf.c | 168 +++++++++++++++++++++++++++++++++++++++-- kernel/bpf/syscall.c | 6 ++ kernel/bpf/sysfs_btf.c | 80 +++++++++++++++++++- kernel/bpf/verifier.c | 43 +++++++---- kernel/module/main.c | 4 +- 7 files changed, 277 insertions(+), 28 deletions(-) diff --git a/include/linux/btf.h b/include/linux/btf.h index ddd0f4f32d24..0bf10811fe53 100644 --- a/include/linux/btf.h +++ b/include/linux/btf.h @@ -581,6 +581,8 @@ __u32 *btf_field_iter_next(struct btf_field_iter *it); const char *btf_name_by_offset(const struct btf *btf, u32 offset); const char *btf_str_by_offset(const struct btf *btf, u32 offset); struct btf *btf_parse_vmlinux(void); +void *btf_vmlinux_data(u32 *size, bool load); +u32 btf_vmlinux_size(void); struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog); u32 *btf_kfunc_flags(const struct btf *btf, u32 kfunc_btf_id, const struct= bpf_prog *prog); int btf_kfunc_check_flag(const struct btf *btf, u32 kfunc_btf_id, u32 flag= ); diff --git a/include/linux/module.h b/include/linux/module.h index 96cc98568eea..82734996a862 100644 --- a/include/linux/module.h +++ b/include/linux/module.h @@ -497,7 +497,7 @@ struct module { unsigned int num_bpf_raw_events; struct bpf_raw_event_map *bpf_raw_events; #endif -#ifdef CONFIG_DEBUG_INFO_BTF_MODULES +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_IN= FO_BTF) unsigned int btf_data_size; unsigned int btf_base_data_size; void *btf_data; diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index c3c1421208b4..50eb7a95fd82 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -29,6 +29,7 @@ #include #include #include +#include #include =20 #include @@ -6092,10 +6093,86 @@ static struct btf *btf_parse(const union bpf_attr *= attr, bpfptr_t uattr, return ERR_PTR(err); } =20 +#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF) extern char __start_BTF[]; extern char __stop_BTF[]; +#endif extern struct btf *btf_vmlinux; =20 +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +/* + * With CONFIG_DEBUG_INFO_BTF=3Dm the vmlinux BTF is not part of the kernel + * image. The btf_vmlinux module carries it in its .BTF section; when the + * module loads, btf_module_notify() copies the section here. The copy is + * made with vmalloc_user() so that /sys/kernel/btf/vmlinux can be mmap()ed + * as with the built-in BTF. Set once, never cleared: like the built-in + * BTF, once present it stays for the lifetime of the kernel. + * + * The size and SHA-256 of the BTF are linked into the kernel as .BTF.meta + * by scripts/gen-btf.sh: the size makes /sys/kernel/btf/vmlinux report its + * size before the BTF is loaded, the hash makes sure only the BTF this + * kernel was built with is accepted. + */ +struct btf_vmlinux_meta { + u32 size; + u8 sha256[SHA256_DIGEST_SIZE]; +} __packed; + +extern const struct btf_vmlinux_meta __start_BTF_meta[]; +#define btf_vmlinux_meta (__start_BTF_meta[0]) + +static void *btf_vmlinux_raw; +#endif + +/** + * btf_vmlinux_data - get the raw vmlinux BTF + * @size: where to store the size of the BTF + * @load: with CONFIG_DEBUG_INFO_BTF=3Dm, load the btf_vmlinux module if t= he + * BTF is not present yet; may sleep + * + * Return: the raw BTF, or NULL if it is not available. + */ +void *btf_vmlinux_data(u32 *size, bool load) +{ +#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF) + *size =3D __stop_BTF - __start_BTF; + return __start_BTF; +#elif IS_MODULE(CONFIG_DEBUG_INFO_BTF) + /* Pairs with the smp_store_release() in btf_vmlinux_module_coming() */ + void *data =3D smp_load_acquire(&btf_vmlinux_raw); + + if (!data && load) { + /* + * The module notifier installs the BTF before init_module() + * returns, so it is either there after this or the module is + * not available (yet). Not cached: a later call retries, + * e.g. once the module becomes reachable on the root fs. + */ + request_module("btf_vmlinux"); + /* Same pairing as above */ + data =3D smp_load_acquire(&btf_vmlinux_raw); + } + *size =3D btf_vmlinux_meta.size; + return data; +#else + return NULL; +#endif +} + +/** + * btf_vmlinux_size - size of the vmlinux BTF, known even before it is loa= ded + */ +u32 btf_vmlinux_size(void) +{ +#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF) + return __stop_BTF - __start_BTF; +#elif IS_MODULE(CONFIG_DEBUG_INFO_BTF) + return btf_vmlinux_meta.size; +#else + return 0; +#endif +} + #define BPF_MAP_TYPE(_id, _ops) #define BPF_LINK_TYPE(_id, _name) static union { @@ -6479,15 +6556,22 @@ struct btf *btf_parse_vmlinux(void) struct btf_verifier_env *env =3D NULL; struct bpf_verifier_log *log; struct btf *btf; + void *data; + u32 size; int err; =20 + /* The caller made sure the BTF is present, see bpf_get_btf_vmlinux() */ + data =3D btf_vmlinux_data(&size, false); + if (!data) + return ERR_PTR(-ENOENT); + env =3D kzalloc_obj(*env, GFP_KERNEL | __GFP_NOWARN); if (!env) return ERR_PTR(-ENOMEM); =20 log =3D &env->log; log->level =3D BPF_LOG_KERNEL; - btf =3D btf_parse_base(env, "vmlinux", __start_BTF, __stop_BTF - __start_= BTF); + btf =3D btf_parse_base(env, "vmlinux", data, size); if (IS_ERR(btf)) goto err_out; =20 @@ -6514,7 +6598,7 @@ __u32 btf_relocate_id(const struct btf *btf, __u32 id) return btf->base_id_map[id]; } =20 -#ifdef CONFIG_DEBUG_INFO_BTF_MODULES +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_IN= FO_BTF) =20 /* * Parse split module BTF against @vmlinux_btf. @data is the module's .BTF @@ -6620,7 +6704,7 @@ static struct btf *btf_parse_module(const char *modul= e_name, struct btf *vmlinux return ERR_PTR(err); } =20 -#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */ +#endif /* CONFIG_DEBUG_INFO_BTF_MODULES || CONFIG_DEBUG_INFO_BTF=3Dm */ =20 struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog) { @@ -8604,7 +8688,16 @@ enum { BTF_MODULE_F_LIVE =3D (1 << 0), }; =20 -#ifdef CONFIG_DEBUG_INFO_BTF_MODULES +/* + * The module notifier registers module BTF (CONFIG_DEBUG_INFO_BTF_MODULES) + * and picks up the vmlinux BTF from the btf_vmlinux module + * (CONFIG_DEBUG_INFO_BTF=3Dm). + */ +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_IN= FO_BTF) +#define BTF_MODULE_NOTIFIER 1 +#endif + +#ifdef BTF_MODULE_NOTIFIER struct btf_module { struct list_head list; struct module *module; @@ -8660,6 +8753,52 @@ static void btf_module_free(struct btf_module *btf_m= od) kfree(btf_mod); } =20 +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +/* + * The btf_vmlinux module carries the vmlinux BTF in its .BTF section + * (scripts/gen-btf.sh). Keep a copy; the module is only the carrier and + * has no BTF of its own. + */ +static int btf_vmlinux_module_coming(struct module *mod) +{ + u8 sha256sum[SHA256_DIGEST_SIZE]; + void *data; + + if (btf_vmlinux_raw) + return 0; + + /* + * The verifier trusts the BTF as the description of this kernel's + * types, so a BTF from a different build must not get in even if + * the module otherwise loads (same release string, same vermagic). + */ + if (mod->btf_data_size !=3D btf_vmlinux_meta.size) { + pr_err("module [%s]: BTF size %u does not match this kernel (%u)\n", + mod->name, mod->btf_data_size, btf_vmlinux_meta.size); + return -EINVAL; + } + sha256(mod->btf_data, mod->btf_data_size, sha256sum); + if (memcmp(sha256sum, btf_vmlinux_meta.sha256, sizeof(sha256sum))) { + pr_err("module [%s]: BTF does not match this kernel\n", mod->name); + return -EINVAL; + } + + data =3D vmalloc_user(mod->btf_data_size); + if (!data) + return -ENOMEM; + memcpy(data, mod->btf_data, mod->btf_data_size); + + /* Pairs with the smp_load_acquire() in btf_vmlinux_data() */ + smp_store_release(&btf_vmlinux_raw, data); + return 0; +} +#else +static int btf_vmlinux_module_coming(struct module *mod) +{ + return 0; +} +#endif + static int btf_module_notify(struct notifier_block *nb, unsigned long op, void *module) { @@ -8668,9 +8807,17 @@ static int btf_module_notify(struct notifier_block *= nb, unsigned long op, struct btf *btf; int err =3D 0; =20 - if (mod->btf_data_size =3D=3D 0 || - (op !=3D MODULE_STATE_COMING && op !=3D MODULE_STATE_LIVE && - op !=3D MODULE_STATE_GOING)) + if (op !=3D MODULE_STATE_COMING && op !=3D MODULE_STATE_LIVE && + op !=3D MODULE_STATE_GOING) + goto out; + + if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && !strcmp(mod->name, "btf_vmlinux")= ) { + if (op =3D=3D MODULE_STATE_COMING) + err =3D btf_vmlinux_module_coming(mod); + goto out; + } + + if (!IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || mod->btf_data_size =3D= =3D 0) goto out; =20 switch (op) { @@ -8758,7 +8905,7 @@ static int __init btf_module_init(void) } =20 fs_initcall(btf_module_init); -#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */ +#endif /* BTF_MODULE_NOTIFIER */ =20 struct module *btf_try_get_module(const struct btf *btf) { @@ -9715,6 +9862,11 @@ static void purge_cand_cache(struct btf *btf) __purge_cand_cache(btf, module_cand_cache, MODULE_CAND_CACHE_SIZE); mutex_unlock(&cand_cache_mutex); } +#elif defined(BTF_MODULE_NOTIFIER) +/* CONFIG_DEBUG_INFO_BTF=3Dm without module BTF: nothing is ever cached */ +static void purge_cand_cache(struct btf *btf) +{ +} #endif =20 static struct bpf_cand_cache * diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index 74496fd716d3..9e9eaa798813 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -6406,6 +6406,12 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uatt= r, unsigned int size, &map_idr, &map_idr_lock); break; case BPF_BTF_GET_NEXT_ID: + /* + * With CONFIG_DEBUG_INFO_BTF=3Dm the kernel BTFs get ids when the + * vmlinux BTF is loaded; whoever enumerates them wants them. + */ + if (IS_MODULE(CONFIG_DEBUG_INFO_BTF)) + bpf_get_btf_vmlinux(); err =3D bpf_obj_get_next_id(&attr, uattr.user, &btf_idr, &btf_idr_lock); break; diff --git a/kernel/bpf/sysfs_btf.c b/kernel/bpf/sysfs_btf.c index 9cbe15ce3540..03f47734987f 100644 --- a/kernel/bpf/sysfs_btf.c +++ b/kernel/bpf/sysfs_btf.c @@ -9,8 +9,13 @@ #include #include #include +#include #include +#include =20 +struct kobject *btf_kobj; + +#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF) /* See scripts/link-vmlinux.sh, gen_btf() func for details */ extern char __start_BTF[]; extern char __stop_BTF[]; @@ -49,12 +54,81 @@ static struct bin_attribute bin_attr_btf_vmlinux __ro_a= fter_init =3D { .mmap =3D btf_sysfs_vmlinux_mmap, }; =20 -struct kobject *btf_kobj; - -static int __init btf_vmlinux_init(void) +static void __init btf_sysfs_vmlinux_init(void) { bin_attr_btf_vmlinux.private =3D __start_BTF; bin_attr_btf_vmlinux.size =3D __stop_BTF - __start_BTF; +} + +#else /* CONFIG_DEBUG_INFO_BTF=3Dm */ + +/* + * The BTF is carried by the btf_vmlinux module and only loaded when + * something needs it. Its size is known from the start, so the file has + * its final size from boot; the first read() or mmap() loads the BTF. + */ +static ssize_t btf_sysfs_vmlinux_read(struct file *filp, struct kobject *k= obj, + const struct bin_attribute *attr, + char *buf, loff_t off, size_t count) +{ + void *data; + u32 size; + + /* Loads the module, parses the BTF and registers module BTFs. */ + if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) + return -ENODEV; + data =3D btf_vmlinux_data(&size, false); + if (!data) + return -ENODEV; + + /* sysfs clamps @off and @count to attr->size, which is @size */ + memcpy(buf, data + off, count); + return count; +} + +static int btf_sysfs_vmlinux_mmap(struct file *filp, struct kobject *kobj, + const struct bin_attribute *attr, + struct vm_area_struct *vma) +{ + size_t vm_size =3D vma->vm_end - vma->vm_start; + void *data; + u32 size; + + if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) + return -ENODEV; + data =3D btf_vmlinux_data(&size, false); + if (!data) + return -ENODEV; + + if (vma->vm_pgoff) + return -EINVAL; + + if (vma->vm_flags & (VM_WRITE | VM_EXEC | VM_MAYSHARE)) + return -EACCES; + + if (vm_size > PAGE_ALIGN(size)) + return -EINVAL; + + vm_flags_mod(vma, VM_DONTDUMP, VM_MAYEXEC | VM_MAYWRITE); + /* the copy was made with vmalloc_user() for this purpose */ + return remap_vmalloc_range(vma, data, 0); +} + +static struct bin_attribute bin_attr_btf_vmlinux __ro_after_init =3D { + .attr =3D { .name =3D "vmlinux", .mode =3D 0444, }, + .read =3D btf_sysfs_vmlinux_read, + .mmap =3D btf_sysfs_vmlinux_mmap, +}; + +static void __init btf_sysfs_vmlinux_init(void) +{ + bin_attr_btf_vmlinux.size =3D btf_vmlinux_size(); +} +#endif + +static int __init btf_vmlinux_init(void) +{ + btf_sysfs_vmlinux_init(); =20 if (bin_attr_btf_vmlinux.size =3D=3D 0) return 0; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index 2425ea74b61d..a7b73bc146a8 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -21157,26 +21157,41 @@ int bpf_check_attach_btf_id_multi(struct btf *btf= , struct bpf_prog *prog, u32 bt return 0; } =20 +/* + * Returns the parsed vmlinux BTF, NULL if the kernel has none, or an ERR_= PTR + * if it is malformed. With CONFIG_DEBUG_INFO_BTF=3Dm the BTF lives in the + * btf_vmlinux module; the first caller loads it and parses it. May sleep. + */ struct btf *bpf_get_btf_vmlinux(void) { /* Pairs with the smp_store_release() on the parse path below. */ struct btf *btf =3D smp_load_acquire(&btf_vmlinux); + u32 size; =20 - if (!btf && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) { - mutex_lock(&btf_vmlinux_lock); - btf =3D btf_vmlinux; - if (!btf) { - btf =3D btf_parse_vmlinux(); - /* - * Order the parsed BTF contents and the globals the - * parse populated (e.g. bpf_ctx_convert.t) before - * the pointer publication. Pairs with the acquire - * on the lockless fast path above. - */ - smp_store_release(&btf_vmlinux, btf); - } - mutex_unlock(&btf_vmlinux_lock); + if (btf || !IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) + return btf; + + /* + * Loading the module may take a while and its notifier must not be + * blocked by us, so do it outside btf_vmlinux_lock. Not available: + * behave like a kernel without BTF, and retry next time. + */ + if (!btf_vmlinux_data(&size, true)) + return NULL; + + mutex_lock(&btf_vmlinux_lock); + btf =3D btf_vmlinux; + if (!btf) { + btf =3D btf_parse_vmlinux(); + /* + * Order the parsed BTF contents and the globals the + * parse populated (e.g. bpf_ctx_convert.t) before + * the pointer publication. Pairs with the acquire + * on the lockless fast path above. + */ + smp_store_release(&btf_vmlinux, btf); } + mutex_unlock(&btf_vmlinux_lock); return btf; } =20 diff --git a/kernel/module/main.c b/kernel/module/main.c index d0e1e0bd2ad0..694c4bc7e679 100644 --- a/kernel/module/main.c +++ b/kernel/module/main.c @@ -2718,7 +2718,7 @@ static int find_module_sections(struct module *mod, s= truct load_info *info) sizeof(*mod->bpf_raw_events), &mod->num_bpf_raw_events); #endif -#ifdef CONFIG_DEBUG_INFO_BTF_MODULES +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_IN= FO_BTF) mod->btf_data =3D any_section_objs(info, ".BTF", 1, &mod->btf_data_size); mod->btf_base_data =3D any_section_objs(info, ".BTF.base", 1, &mod->btf_base_data_size); @@ -3172,7 +3172,7 @@ static noinline int do_init_module(struct module *mod) mod->mem[type].size =3D 0; } =20 -#ifdef CONFIG_DEBUG_INFO_BTF_MODULES +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_IN= FO_BTF) /* .BTF is not SHF_ALLOC and will get removed, so sanitize pointers */ mod->btf_data =3D NULL; mod->btf_base_data =3D NULL; --=20 2.47.3 From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.246.68.102]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4DD073C1977; Wed, 23 Sep 2026 05:41:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.246.68.102 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142070; cv=none; b=SRfWmAfuwiKrk/WrTl1Mfmal1exTwdd+gxTnE9EBRJSdYGURSEClBAlnc2KaaWHDFyBFYiS29QkAN746/d8d7iYu8X9M8sbNdh+xaBs7pm+Et7O5/5vmC0gIu+skBFwdfh5EgKwFUYvgxGG+s5uu8iYtxKcKt6fvd9KZkoN4rjs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142070; c=relaxed/simple; bh=3mrZ06axOfwxF5syrWegesLOKvJNpg2qVDvYk0jNGsU=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=c24I0nrRgM13CD+nqkFUI36LgZAAGvTaOveKJfnFfI9ZhvguPTr0cmKHu6DPYVjAZy2+Lhp9tSxyicMKOe/zVG2A4FM4ozsnuTPYAC4y7uPIXv/lJ/18BusgLLLKMsB39FxBl4Ky01EkEDKDNQIVJVh65c2+lGIrnFxTmDBp32Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=O5RoG8ht; arc=none smtp.client-ip=44.246.68.102 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="O5RoG8ht" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142068; x=1821678068; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=m57+EWci8nagxSm0zqha+4oBDWC0wPRFTb81AvfWLH0=; b=O5RoG8ht5CidcuK8nd679J5jpleudXpyl1wepQ0AAbTB52QSXZCO4opg Do2ZMKrXA/RxZse3EYZT3C80BZdV1e/DkmR/hAOG5pJc764gQLW9k4onV OvAJi1F1EmYfZGSMrCK+woWP8bd+KuvbrjLv7Ln3XDc2bT0RDeIlib0FG Q8mck9RdwFeyMXju69gal0V1uEsNTuzLr/9rqvd6mLWy3HMs64rsJwZVV 9TMOTWq6jRCbU12AeiVGUAFeW0+elYO3gwXtzg0ifeb99eRtiOVSMuai8 +aCLN3MUTkbp+fShmeeSN6SFaaFz0aKIwPiYDUOPXPe4gbwZ1YWodOHBt g==; X-CSE-ConnectionGUID: uBpdzH8jRB+ntUvuCudFNg== X-CSE-MsgGUID: j2KDraPMRpiavy0EO1dhBQ== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29318535" Received: from ip-10-5-9-48.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.9.48]) by internal-pdx-out-003.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:41:07 +0000 Received: from EX19MTAUWC001.ant.amazon.com [205.251.233.105:24310] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.16.195:2525] with esmtp (Farcaster) id ff2f438e-ffd1-4b3a-87e3-1a5ba2fd6b64; Wed, 23 Sep 2026 05:41:07 +0000 (UTC) X-Farcaster-Flow-ID: ff2f438e-ffd1-4b3a-87e3-1a5ba2fd6b64 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWC001.ant.amazon.com (10.250.64.174) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:41:07 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:41:07 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 5/6] bpf: defer registrations until the vmlinux BTF is available Date: Wed, 23 Sep 2026 05:39:47 +0000 Message-ID: <20260923053948.30617-6-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D033UWC003.ant.amazon.com (10.13.139.217) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" With CONFIG_DEBUG_INFO_BTF=3Dm the vmlinux BTF is loaded on first use. For that to save anything, nothing may pull it in at boot. The verifier no longer does since the previous patches; two things still do: - register_btf_kfunc_id_set(), register_btf_id_dtor_kfuncs() and register_bpf_struct_ops() for vmlinux run from initcalls and need the parsed BTF. Queue them instead (btf_defer_reg()) and apply them in btf_parse_vmlinux(), before the BTF is published, so that no program can see a vmlinux BTF without its kfuncs and struct_ops. Applying a struct_ops runs its ->init(), which registers the kfunc sets of its hook; those land back on the queue, so it is drained in a loop until a pass adds nothing, and only then are new registrations applied directly. The dtor arrays are copied: every caller in the tree builds them on the stack of its initcall. - Module BTF is split BTF against the vmlinux BTF and used to be parsed in the module notifier. The notifier cannot load btf_vmlinux (that would nest a module load in a module load), so a module loaded before the vmlinux BTF keeps a copy of its .BTF and .BTF.base, is exposed in /sys/kernel/btf right away (the raw bytes need no parsing) and gets a list entry with btf =3D=3D NULL. Its kfunc, dtor kfunc and struct_ops registrations wait on that entry. When the vmlinux BTF arrives, btf_parse_deferred_modules() parses the kept copies, gives them ids and applies the waiting registrations; the copy is the one btf_parse_module() makes anyway, so the sysfs file keeps pointing at valid data. Registrations are applied after dropping btf_module_mutex because they walk the module list (btf_check_kfunc_name()). A module whose BTF turns out to mismatch at that point is already running and keeps running without BTF, with a warning. With =3Dy such a module would have been refused at load time unless CONFIG_MODULE_ALLOW_BTF_MISMATCH; that check only applies to modules loaded after the vmlinux BTF. Walkers of the module BTF list skip entries whose BTF is not parsed yet. With =3Dy the BTF is present from boot, so the queues never fill and the notifier takes the existing path. Still nothing is reachable until the Kconfig symbol becomes a tristate. Signed-off-by: Jay Wang --- include/linux/btf.h | 5 + kernel/bpf/btf.c | 426 +++++++++++++++++++++++++++++++++++++++++- kernel/bpf/verifier.c | 9 +- 3 files changed, 433 insertions(+), 7 deletions(-) diff --git a/include/linux/btf.h b/include/linux/btf.h index 0bf10811fe53..3b99d6386dec 100644 --- a/include/linux/btf.h +++ b/include/linux/btf.h @@ -583,6 +583,11 @@ const char *btf_str_by_offset(const struct btf *btf, u= 32 offset); struct btf *btf_parse_vmlinux(void); void *btf_vmlinux_data(u32 *size, bool load); u32 btf_vmlinux_size(void); +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +void btf_parse_deferred_modules(void); +#else +static inline void btf_parse_deferred_modules(void) {} +#endif struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog); u32 *btf_kfunc_flags(const struct btf *btf, u32 kfunc_btf_id, const struct= bpf_prog *prog); int btf_kfunc_check_flag(const struct btf *btf, u32 kfunc_btf_id, u32 flag= ); diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index 50eb7a95fd82..cbba20a908e9 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -6551,6 +6551,8 @@ static struct btf *btf_parse_base(struct btf_verifier= _env *env, const char *name return ERR_PTR(err); } =20 +static void btf_apply_deferred_vmlinux_regs(struct btf *btf); + struct btf *btf_parse_vmlinux(void) { struct btf_verifier_env *env =3D NULL; @@ -6581,7 +6583,15 @@ struct btf *btf_parse_vmlinux(void) if (err) { btf_free(btf); btf =3D ERR_PTR(err); + goto err_out; } + + /* + * With CONFIG_DEBUG_INFO_BTF=3Dm, kfunc, dtor kfunc and struct_ops + * registrations for vmlinux made before the BTF was available were + * queued; apply them now, before the BTF becomes visible to anyone. + */ + btf_apply_deferred_vmlinux_regs(btf); err_out: btf_verifier_env_free(env); return btf; @@ -8697,13 +8707,62 @@ enum { #define BTF_MODULE_NOTIFIER 1 #endif =20 +/* + * CONFIG_DEBUG_INFO_BTF=3Dm: a kfunc, dtor kfunc or struct_ops registrati= on + * made while the BTF it applies to is not available yet. Kept until the = BTF + * arrives, see btf_defer_reg() and btf_apply_deferred_regs(). + */ +enum btf_deferred_reg_kind { + BTF_DEFERRED_KFUNC_SET, + BTF_DEFERRED_DTOR_KFUNCS, + BTF_DEFERRED_STRUCT_OPS, +}; + +struct btf_deferred_reg { + struct list_head list; + enum btf_deferred_reg_kind kind; + /* set while the registration is queued for a module BTF */ + struct btf *btf; + struct module *module; + union { + struct { + enum btf_kfunc_hook hook; + const struct btf_kfunc_id_set *kset; + } kfunc; + struct { + const struct btf_id_dtor_kfunc *dtors; + u32 cnt; + } dtor; + struct bpf_struct_ops *st_ops; + }; +}; + #ifdef BTF_MODULE_NOTIFIER +static void btf_free_deferred_regs(struct list_head *regs); +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +static void btf_free_deferred_reg(struct btf_deferred_reg *reg); +static void btf_apply_deferred_regs(struct list_head *regs); +#endif + struct btf_module { struct list_head list; struct module *module; struct btf *btf; struct bin_attribute *sysfs_attr; int flags; + /* + * CONFIG_DEBUG_INFO_BTF=3Dm: a module loaded before the vmlinux BTF is + * available cannot have its BTF parsed yet. Its .BTF and .BTF.base + * sections are copied here and parsed once the vmlinux BTF arrives + * (btf_parse_deferred_modules()); @btf is NULL until then. + * Registrations of the module's kfuncs, dtor kfuncs and struct_ops + * wait in @deferred_regs. + */ + void *data; + void *base_data; + u32 data_size; + u32 base_data_size; + struct list_head deferred_regs; }; =20 static LIST_HEAD(btf_modules); @@ -8747,8 +8806,14 @@ static void btf_module_free(struct btf_module *btf_m= od) { if (btf_mod->sysfs_attr) sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr); - purge_cand_cache(btf_mod->btf); - btf_put(btf_mod->btf); + if (btf_mod->btf) { + purge_cand_cache(btf_mod->btf); + btf_put(btf_mod->btf); + } else { + kvfree(btf_mod->data); + kvfree(btf_mod->base_data); + } + btf_free_deferred_regs(&btf_mod->deferred_regs); kfree(btf_mod->sysfs_attr); kfree(btf_mod); } @@ -8792,11 +8857,55 @@ static int btf_vmlinux_module_coming(struct module = *mod) smp_store_release(&btf_vmlinux_raw, data); return 0; } + +/* + * The vmlinux BTF is not available yet and must not be loaded from the + * module notifier (that would nest a module load into a module load). Ke= ep + * the module's BTF for btf_parse_deferred_modules(). The .BTF data can be + * exposed in sysfs right away, it needs no parsing. + */ +static int btf_module_defer(struct btf_module *btf_mod, struct module *mod) +{ + int err; + + btf_mod->data =3D kvmemdup(mod->btf_data, mod->btf_data_size, + GFP_KERNEL | __GFP_NOWARN); + if (!btf_mod->data) + return -ENOMEM; + btf_mod->data_size =3D mod->btf_data_size; + + if (mod->btf_base_data) { + btf_mod->base_data =3D kvmemdup(mod->btf_base_data, + mod->btf_base_data_size, + GFP_KERNEL | __GFP_NOWARN); + if (!btf_mod->base_data) { + kvfree(btf_mod->data); + return -ENOMEM; + } + btf_mod->base_data_size =3D mod->btf_base_data_size; + } + + err =3D btf_module_sysfs_add(btf_mod, mod->name, btf_mod->data, + btf_mod->data_size); + if (err) { + kvfree(btf_mod->data); + kvfree(btf_mod->base_data); + return err; + } + + list_add(&btf_mod->list, &btf_modules); + return 0; +} #else static int btf_vmlinux_module_coming(struct module *mod) { return 0; } + +static int btf_module_defer(struct btf_module *btf_mod, struct module *mod) +{ + return 0; +} #endif =20 static int btf_module_notify(struct notifier_block *nb, unsigned long op, @@ -8828,6 +8937,24 @@ static int btf_module_notify(struct notifier_block *= nb, unsigned long op, goto out; } btf_mod->module =3D module; + INIT_LIST_HEAD(&btf_mod->deferred_regs); + + if (IS_MODULE(CONFIG_DEBUG_INFO_BTF)) { + mutex_lock(&btf_module_mutex); + /* Pairs with the publication in bpf_get_btf_vmlinux() */ + if (!smp_load_acquire(&btf_vmlinux)) { + err =3D btf_module_defer(btf_mod, mod); + mutex_unlock(&btf_module_mutex); + if (err) { + pr_warn("failed to keep module [%s] BTF: %d\n", + mod->name, err); + kfree(btf_mod); + err =3D 0; + } + goto out; + } + mutex_unlock(&btf_module_mutex); + } =20 btf =3D btf_parse_module(mod->name, bpf_get_btf_vmlinux(), mod->btf_data, mod->btf_data_size, false, @@ -8882,7 +9009,8 @@ static int btf_module_notify(struct notifier_block *n= b, unsigned long op, * btf_try_get_module() on such BTFs will fail. This may * be called again on btf_put(), but it's ok to do so. */ - btf_free_id(btf_mod->btf); + if (btf_mod->btf) + btf_free_id(btf_mod->btf); list_del(&btf_mod->list); btf_module_free(btf_mod); break; @@ -8905,6 +9033,87 @@ static int __init btf_module_init(void) } =20 fs_initcall(btf_module_init); + +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +/* + * CONFIG_DEBUG_INFO_BTF=3Dm: the vmlinux BTF has just become available. = Parse + * the BTF of the modules that were loaded before it, and apply the + * registrations that waited for them. Called from bpf_get_btf_vmlinux() + * once btf_vmlinux is published, with no locks held. + */ +void btf_parse_deferred_modules(void) +{ + /* Pairs with the publication in bpf_get_btf_vmlinux() */ + struct btf *vmlinux_btf =3D smp_load_acquire(&btf_vmlinux); + struct btf_deferred_reg *reg, *rtmp; + struct btf_module *btf_mod, *tmp; + bool parsed =3D false; + LIST_HEAD(regs); + struct btf *btf; + int err; + + if (IS_ERR_OR_NULL(vmlinux_btf)) + return; + + mutex_lock(&btf_module_mutex); + list_for_each_entry_safe(btf_mod, tmp, &btf_modules, list) { + if (btf_mod->btf) + continue; + + btf =3D btf_parse_module(btf_mod->module->name, vmlinux_btf, + btf_mod->data, btf_mod->data_size, true, + btf_mod->base_data, btf_mod->base_data_size); + err =3D PTR_ERR_OR_ZERO(btf); + if (!err) { + err =3D btf_alloc_id(btf); + if (err) { + /* btf owns the data now, btf_free() drops it */ + btf_mod->data =3D NULL; + btf_free(btf); + } + } + if (err) { + /* + * The module is loaded and stays. Unlike at load time + * there is no way to reject it, so drop its BTF. + */ + pr_warn("failed to validate module [%s] BTF: %d\n", + btf_mod->module->name, err); + list_del(&btf_mod->list); + btf_module_free(btf_mod); + continue; + } + + /* btf->data is btf_mod->data now, the sysfs file keeps working */ + btf_mod->data =3D NULL; + kvfree(btf_mod->base_data); + btf_mod->base_data =3D NULL; + btf_mod->btf =3D btf; + parsed =3D true; + + /* + * Registrations are applied after dropping the mutex (they + * walk btf_modules); pin what they need until then. + */ + list_for_each_entry_safe(reg, rtmp, &btf_mod->deferred_regs, list) { + list_del(®->list); + if (!try_module_get(btf_mod->module)) { + btf_free_deferred_reg(reg); + continue; + } + btf_get(btf); + reg->btf =3D btf; + reg->module =3D btf_mod->module; + list_add_tail(®->list, ®s); + } + } + mutex_unlock(&btf_module_mutex); + + if (parsed) + purge_cand_cache(NULL); + btf_apply_deferred_regs(®s); +} +#endif /* IS_MODULE(CONFIG_DEBUG_INFO_BTF) */ #endif /* BTF_MODULE_NOTIFIER */ =20 struct module *btf_try_get_module(const struct btf *btf) @@ -8957,8 +9166,11 @@ struct btf *btf_get_module_btf(const struct module *= module) if (btf_mod->module !=3D module) continue; =20 - btf_get(btf_mod->btf); - btf =3D btf_mod->btf; + /* NULL while waiting for the vmlinux BTF (CONFIG_DEBUG_INFO_BTF=3Dm) */ + if (btf_mod->btf) { + btf_get(btf_mod->btf); + btf =3D btf_mod->btf; + } break; } mutex_unlock(&btf_module_mutex); @@ -9135,7 +9347,8 @@ static int btf_check_kfunc_name(struct btf *btf, cons= t char *func_name, u32 kind #ifdef CONFIG_DEBUG_INFO_BTF_MODULES guard(mutex)(&btf_module_mutex); list_for_each_entry_safe(btf_mod, tmp, &btf_modules, list) { - if (btf_mod->btf =3D=3D btf) + /* skip ourselves and, with CONFIG_DEBUG_INFO_BTF=3Dm, unparsed BTF */ + if (btf_mod->btf =3D=3D btf || !btf_mod->btf) continue; id =3D btf_find_by_name_kind(btf_mod->btf, func_name, kind); if (id >=3D 0) { @@ -9475,12 +9688,22 @@ static int btf_kfunc_id_set_add(struct btf *btf, en= um btf_kfunc_hook hook, return btf_populate_kfunc_set(btf, hook, kset); } =20 +static int btf_defer_reg(struct module *owner, const struct btf_deferred_r= eg *tmpl); + static int __register_btf_kfunc_id_set(enum btf_kfunc_hook hook, const struct btf_kfunc_id_set *kset) { + struct btf_deferred_reg tmpl =3D { + .kind =3D BTF_DEFERRED_KFUNC_SET, + .kfunc =3D { .hook =3D hook, .kset =3D kset }, + }; struct btf *btf; int ret; =20 + ret =3D btf_defer_reg(kset->owner, &tmpl); + if (ret) + return ret > 0 ? 0 : ret; + btf =3D btf_get_module_btf(kset->owner); if (!btf) return check_btf_kconfigs(kset->owner, "kfunc"); @@ -9649,9 +9872,17 @@ static int btf_dtor_kfuncs_add(struct btf *btf, cons= t struct btf_id_dtor_kfunc * int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32= add_cnt, struct module *owner) { + struct btf_deferred_reg tmpl =3D { + .kind =3D BTF_DEFERRED_DTOR_KFUNCS, + .dtor =3D { .dtors =3D dtors, .cnt =3D add_cnt }, + }; struct btf *btf; int ret; =20 + ret =3D btf_defer_reg(owner, &tmpl); + if (ret) + return ret > 0 ? 0 : ret; + btf =3D btf_get_module_btf(owner); if (!btf) return check_btf_kconfigs(owner, "dtor kfuncs"); @@ -10321,9 +10552,17 @@ static int btf_struct_ops_add(struct btf *btf, str= uct bpf_struct_ops *st_ops) =20 int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops) { + struct btf_deferred_reg tmpl =3D { + .kind =3D BTF_DEFERRED_STRUCT_OPS, + .st_ops =3D st_ops, + }; struct btf *btf; int err; =20 + err =3D btf_defer_reg(st_ops->owner, &tmpl); + if (err) + return err > 0 ? 0 : err; + btf =3D btf_get_module_btf(st_ops->owner); if (!btf) return check_btf_kconfigs(st_ops->owner, "struct_ops"); @@ -10335,8 +10574,183 @@ int __register_bpf_struct_ops(struct bpf_struct_o= ps *st_ops) return err; } EXPORT_SYMBOL_GPL(__register_bpf_struct_ops); +#else +static int btf_struct_ops_add(struct btf *btf, struct bpf_struct_ops *st_o= ps) +{ + return -EOPNOTSUPP; +} +#endif + +/* + * CONFIG_DEBUG_INFO_BTF=3Dm: registrations made before the BTF they apply= to + * is available. Registrations for vmlinux wait in btf_vmlinux_deferred_r= egs + * until btf_parse_vmlinux() applies them; registrations for a module wait= in + * its struct btf_module until btf_parse_deferred_modules() does. Both li= sts + * are protected by btf_module_mutex. + */ +#ifdef BTF_MODULE_NOTIFIER +static LIST_HEAD(btf_vmlinux_deferred_regs); +/* Set when the vmlinux BTF is parsed; new registrations apply directly */ +static bool btf_vmlinux_regs_closed; + +/* + * Queue @tmpl if the BTF for @owner is not available yet. Returns 1 if t= he + * registration was queued and is to be considered done, 0 if the caller h= as + * to apply it, or -ENOMEM. + */ +static int btf_defer_reg(struct module *owner, const struct btf_deferred_r= eg *tmpl) +{ + struct list_head *head =3D NULL; + struct btf_deferred_reg *reg; + struct btf_module *btf_mod; + + if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF)) + return 0; + + guard(mutex)(&btf_module_mutex); + if (!owner) { + if (!btf_vmlinux_regs_closed) + head =3D &btf_vmlinux_deferred_regs; + } else { + list_for_each_entry(btf_mod, &btf_modules, list) { + if (btf_mod->module !=3D owner) + continue; + if (!btf_mod->btf) + head =3D &btf_mod->deferred_regs; + break; + } + } + if (!head) + return 0; + + reg =3D kmemdup(tmpl, sizeof(*reg), GFP_KERNEL); + if (!reg) + return -ENOMEM; + /* + * kfunc id sets and struct_ops are static data of their owner, but + * the dtor arrays are commonly built on the stack of the initcall. + */ + if (reg->kind =3D=3D BTF_DEFERRED_DTOR_KFUNCS) { + reg->dtor.dtors =3D kmemdup_array(tmpl->dtor.dtors, tmpl->dtor.cnt, + sizeof(*tmpl->dtor.dtors), GFP_KERNEL); + if (!reg->dtor.dtors) { + kfree(reg); + return -ENOMEM; + } + } + list_add_tail(®->list, head); + return 1; +} + +static void btf_free_deferred_reg(struct btf_deferred_reg *reg) +{ + if (reg->kind =3D=3D BTF_DEFERRED_DTOR_KFUNCS) + kfree(reg->dtor.dtors); + kfree(reg); +} + +static const char *btf_deferred_reg_name(const struct btf_deferred_reg *re= g) +{ + switch (reg->kind) { + case BTF_DEFERRED_KFUNC_SET: return "kfunc set"; + case BTF_DEFERRED_DTOR_KFUNCS: return "dtor kfuncs"; + case BTF_DEFERRED_STRUCT_OPS: return "struct_ops"; + } + return "?"; +} + +static int btf_apply_deferred_reg(struct btf *btf, const struct btf_deferr= ed_reg *reg) +{ + switch (reg->kind) { + case BTF_DEFERRED_KFUNC_SET: + return btf_kfunc_id_set_add(btf, reg->kfunc.hook, reg->kfunc.kset); + case BTF_DEFERRED_DTOR_KFUNCS: + return btf_dtor_kfuncs_add(btf, reg->dtor.dtors, reg->dtor.cnt); + case BTF_DEFERRED_STRUCT_OPS: + return btf_struct_ops_add(btf, reg->st_ops); + } + return -EINVAL; +} + +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +/* Apply and free the registrations in @regs; each is pinned to its BTF an= d module. */ +static void btf_apply_deferred_regs(struct list_head *regs) +{ + struct btf_deferred_reg *reg, *tmp; + int err; + + list_for_each_entry_safe(reg, tmp, regs, list) { + err =3D btf_apply_deferred_reg(reg->btf, reg); + if (err) + pr_warn("failed to register deferred %s for module [%s] BTF: %d\n", + btf_deferred_reg_name(reg), reg->btf->name, err); + btf_put(reg->btf); + module_put(reg->module); + list_del(®->list); + btf_free_deferred_reg(reg); + } +} #endif =20 +static void btf_free_deferred_regs(struct list_head *regs) +{ + struct btf_deferred_reg *reg, *tmp; + + list_for_each_entry_safe(reg, tmp, regs, list) { + list_del(®->list); + btf_free_deferred_reg(reg); + } +} + +/* + * The vmlinux BTF has just been parsed; apply the registrations that wait= ed + * for it. Runs under btf_vmlinux_lock, before @btf is published, so noth= ing + * can observe a vmlinux BTF without its kfuncs and struct_ops. + * + * Applying a registration can queue further ones: a struct_ops ->init() + * registers the kfuncs of its hook. Those must not go through + * bpf_get_btf_vmlinux() (we hold its lock), so the queue stays open until + * a pass applies nothing new, and only then are registrations applied dir= ectly. + */ +static void btf_apply_deferred_vmlinux_regs(struct btf *btf) +{ + struct btf_deferred_reg *reg, *tmp; + LIST_HEAD(regs); + int err; + + if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF)) + return; + + mutex_lock(&btf_module_mutex); + while (!list_empty(&btf_vmlinux_deferred_regs)) { + list_splice_init(&btf_vmlinux_deferred_regs, ®s); + mutex_unlock(&btf_module_mutex); + + list_for_each_entry_safe(reg, tmp, ®s, list) { + err =3D btf_apply_deferred_reg(btf, reg); + if (err) + pr_warn("failed to register deferred %s for vmlinux BTF: %d\n", + btf_deferred_reg_name(reg), err); + list_del(®->list); + btf_free_deferred_reg(reg); + } + + mutex_lock(&btf_module_mutex); + } + btf_vmlinux_regs_closed =3D true; + mutex_unlock(&btf_module_mutex); +} +#else +static int btf_defer_reg(struct module *owner, const struct btf_deferred_r= eg *tmpl) +{ + return 0; +} + +static void btf_apply_deferred_vmlinux_regs(struct btf *btf) +{ +} +#endif /* BTF_MODULE_NOTIFIER */ + bool btf_param_match_suffix(const struct btf *btf, const struct btf_param *arg, const char *suffix) diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index a7b73bc146a8..b6f094d5306a 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -21160,12 +21160,14 @@ int bpf_check_attach_btf_id_multi(struct btf *btf= , struct bpf_prog *prog, u32 bt /* * Returns the parsed vmlinux BTF, NULL if the kernel has none, or an ERR_= PTR * if it is malformed. With CONFIG_DEBUG_INFO_BTF=3Dm the BTF lives in the - * btf_vmlinux module; the first caller loads it and parses it. May sleep. + * btf_vmlinux module; the first caller loads it, parses it and then regis= ters + * the BTF of the modules that were loaded before it. May sleep. */ struct btf *bpf_get_btf_vmlinux(void) { /* Pairs with the smp_store_release() on the parse path below. */ struct btf *btf =3D smp_load_acquire(&btf_vmlinux); + bool parsed =3D false; u32 size; =20 if (btf || !IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) @@ -21190,8 +21192,13 @@ struct btf *bpf_get_btf_vmlinux(void) * on the lockless fast path above. */ smp_store_release(&btf_vmlinux, btf); + parsed =3D true; } mutex_unlock(&btf_vmlinux_lock); + + if (parsed && IS_MODULE(CONFIG_DEBUG_INFO_BTF)) + btf_parse_deferred_modules(); + return btf; } =20 --=20 2.47.3 From nobody Thu Sep 24 13:37:12 2026 Received: from pdx-out-007.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-007.esa.us-west-2.outbound.mail-perimeter.amazon.com [52.34.181.151]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD5843C1977; Wed, 23 Sep 2026 05:41:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=52.34.181.151 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142085; cv=none; b=pRs6hQKa6VWid9C/csApg4I+EIsQ4IAXb8il/izCsTzl/m1CejNwxtpXtgl2qeQn/cQNMDnwfu2THLK6HRq7UMzrim8/3XfoP5PohKUvikpC05+Uk0/7iayDOPCVRXsWJ/sidfHWP6Id6VY8YK+XsJKoNWR53ToaD5KQ8C3ZJZs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790142085; c=relaxed/simple; bh=Ojy5qHrjikKKwR93jbWNyZrhvbt0Jp+kI1q/xVRg37I=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=C4l0R4Kks6GaV3ba5dOHrEdoBjocRnOCcOVAvje9JyV4iSndcMKP6Gu/oT8DRi7JaTgcqZ/1vkOWj+n4rPyNnJAkSkN0ld9nyLYqqB33JjjonU3Oeza8dfju8GTO1b5Pv5Y+AypADCtWZwS7hpcj9zMjC+X0++BlkpAIOsYXhtY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com; spf=pass smtp.mailfrom=amazon.com; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b=MbM9kgi0; arc=none smtp.client-ip=52.34.181.151 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.com header.i=@amazon.com header.b="MbM9kgi0" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.com; i=@amazon.com; q=dns/txt; s=amazoncorp2; t=1790142083; x=1821678083; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=usAUkm2oNPIPlWlGTUyS8j3XcpVsC1Gdc0Kkt71uYak=; b=MbM9kgi0E5Gh5t9XkygXGDFAQmlz2GJhUjkHz/N6UDrwbeC34Ttapd+U ewPSYbx1SpumyJr4VvmgxJ28lKNDD+GBCMjKI9UVfOpk6S+GoTU2jHb20 i45Nte1T73xlDUPHzaIAETz+8EfaZLJRIKLa8BX/2AEmn+ycEou8J/vXj CNF5UfZ1YeJwMIFlVPbzGaby7V+JUGlLAklVTbEvpIp773yc80KPunN+F pfDnZWZzx9NM+pEezITTPjuwLTMvgKFLxjd6r6c5lAQ2cRIV1He+4FP3v DGw9G2mbMvlpcRv0WrNU3MyuB/BGbeuIeFQNOaQLbbnAEZGraUc64owWD Q==; X-CSE-ConnectionGUID: Fi1O6+OPSgaMM6ONJ4F+Hg== X-CSE-MsgGUID: nrxdIa+MQm6Fcrsm3xloQA== X-IronPort-AV: E=Sophos;i="6.27,117,1787011200"; d="scan'208";a="29398036" Received: from ip-10-5-9-48.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.9.48]) by internal-pdx-out-007.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 23 Sep 2026 05:41:23 +0000 Received: from EX19MTAUWB002.ant.amazon.com [205.251.233.111:25629] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.20.120:2525] with esmtp (Farcaster) id 36533df6-7c9e-4377-a345-731bd5588139; Wed, 23 Sep 2026 05:41:23 +0000 (UTC) X-Farcaster-Flow-ID: 36533df6-7c9e-4377-a345-731bd5588139 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWB002.ant.amazon.com (10.250.64.231) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:41:22 +0000 Received: from dev-dsk-wanjay-2c-d25651b4.us-west-2.amazon.com (172.19.198.4) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.49; Wed, 23 Sep 2026 05:41:22 +0000 From: Jay Wang To: , Alexei Starovoitov , "Daniel Borkmann" , Andrii Nakryiko , "Eduard Zingerman" , Kumar Kartikeya Dwivedi CC: Alan Maguire , Martin KaFai Lau , Yonghong Song , "Nathan Chancellor" , Nicolas Schier , , Luis Chamberlain , "Petr Pavlu" , , Arnd Bergmann , , Hazem Mohamed Abuelfotoh , Bjoern Doebel , Subject: [PATCH bpf-next 6/6] kbuild, bpf: allow building the vmlinux BTF as a module Date: Wed, 23 Sep 2026 05:39:48 +0000 Message-ID: <20260923053948.30617-7-wanjay@amazon.com> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260923053948.30617-1-wanjay@amazon.com> References: <20260923053948.30617-1-wanjay@amazon.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D035UWB004.ant.amazon.com (10.13.138.104) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Make CONFIG_DEBUG_INFO_BTF a tristate. With =3Dm the vmlinux BTF is not part of the kernel image: it is carried by a new module, btf_vmlinux, and loaded the first time something needs it. Nothing that works with =3Dy stops working; the 5.4 MiB of read-only data (distribution config) is simply not there on systems where nothing uses it. The only way to save that memory today is CONFIG_DEBUG_INFO_BTF=3Dn, which a distribution cannot ship: one binary goes to every user, and off takes BTF away from the users of CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext or bpf-lsm. Whether BTF is used is a property of the workload, not of the build, so let the first user decide. The BTF is generated as before, but with =3Dm the .BTF section is linked into vmlinux as a non-loadable section (like .comment), so the vmlinux ELF still carries it for module BTF generation and tooling while the image does not. .BTF_ids stays loadable, the verifier needs it once the BTF is loaded. The object that carries .BTF also carries .BTF.meta, the size and SHA-256 of the BTF (struct btf_vmlinux_meta, checked by the module notifier); the first link, which the BTF is generated from, gets a zeroed .BTF.meta of the same size. kernel/bpf/btf_vmlinux.c is an empty carrier module; scripts/gen-btf.sh gives it the vmlinux .BTF as its own .BTF section instead of generating split BTF for it, so modules depend on vmlinux with =3Dm as they do with CONFIG_DEBUG_INFO_BTF_MODULES. Makefiles that compiled kfunc objects with obj-$(CONFIG_DEBUG_INFO_BTF) now treat m as y, and the #ifdef CONFIG_DEBUG_INFO_BTF sites that must also apply with =3Dm (the .BTF_ids tables, type tags, tracepoint and syscall BTF ids) use IS_ENABLED(): the generated BTF and its id tables are the same for =3Dy and =3Dm, only the delivery of the blob differs. CONFIG_BPF_PRELOAD is not selectable with =3Dm: its iterator programs attach through the vmlinux BTF, so every bpffs mount (systemd does one at boot) would load it and defeat the point. The module has no exit: once loaded the BTF stays, as with =3Dy. The runtime side -- loading the module on first use, checking it against .BTF.meta, deferring kfunc and struct_ops registrations and module BTF until it arrives -- is in the preceding patches; this one makes it selectable. Tested with 1 GiB of memory, same tree, =3Dy vs =3Dm, both with CONFIG_DEBUG_INFO_BTF_MODULES=3Dy: - MemTotal is ~5.4 MB higher with =3Dm while the BTF is unused: the size of the .BTF section. - stat() of /sys/kernel/btf/vmlinux reports the BTF size before it is loaded, as the btf_sysfs selftest expects. - With BTF in use, MemFree is the same within run-to-run noise. - Modules loaded before the trigger (ext4, nf_conntrack and its kfuncs, xfrm_interface) appear in /sys/kernel/btf immediately and get BTF ids once the BTF is loaded; a socket filter loads without loading the module; a kprobe program calling bpf_get_current_task_btf(), opening /sys/kernel/btf/vmlinux or BPF_BTF_GET_NEXT_ID each load it. - A carrier module with one byte of its .BTF changed is refused with "BTF does not match this kernel" and leaves no state behind. - After the load: fstat/read/mmap of /sys/kernel/btf/vmlinux, a struct_ops map for tcp_congestion_ops, a syscall program calling the bpf_task_from_pid()/bpf_task_release() kfuncs, and modules loaded afterwards (nf_nat) all work as with =3Dy. - =3Dm without DEBUG_INFO_BTF_MODULES, and =3Dy, build and pass the same tests. Signed-off-by: Jay Wang --- Documentation/bpf/btf.rst | 35 ++++++++++++ Makefile | 8 ++- include/asm-generic/vmlinux.lds.h | 30 +++++++++- include/linux/btf_ids.h | 2 +- include/linux/compiler_types.h | 2 +- include/trace/trace_events.h | 2 +- kernel/bpf/Makefile | 6 +- kernel/bpf/btf_vmlinux.c | 23 ++++++++ kernel/bpf/preload/Kconfig | 4 ++ kernel/trace/trace_syscalls.c | 6 +- lib/Kconfig.debug | 13 ++++- net/netfilter/Makefile | 6 +- net/xfrm/Makefile | 4 +- scripts/Makefile.modfinal | 14 +++-- scripts/gen-btf.sh | 93 +++++++++++++++++++++++++++++-- scripts/link-vmlinux.sh | 25 +++++++-- 16 files changed, 244 insertions(+), 29 deletions(-) create mode 100644 kernel/bpf/btf_vmlinux.c diff --git a/Documentation/bpf/btf.rst b/Documentation/bpf/btf.rst index 004aa1058d85..a835231187ef 100644 --- a/Documentation/bpf/btf.rst +++ b/Documentation/bpf/btf.rst @@ -1197,6 +1197,41 @@ format.:: .long 58 .long 8206 # Line 8 Col 14 =20 +6.1 Kernel BTF +-------------- + +With CONFIG_DEBUG_INFO_BTF=3Dy the BTF of the kernel is generated at link = time +from its DWARF and placed in the .BTF section of vmlinux, which is read-on= ly +data of the kernel image. It is available as /sys/kernel/btf/vmlinux and, = if +CONFIG_DEBUG_INFO_BTF_MODULES is set, module BTF is generated as split BTF +against it and available as /sys/kernel/btf/. + +With CONFIG_DEBUG_INFO_BTF=3Dm the same BTF is generated, but it is not pa= rt of +the kernel image (the vmlinux ELF file still carries it in a non-loadable = .BTF +section for tooling and module BTF generation). It is delivered by the +btf_vmlinux module, which the kernel loads on demand the first time the BT= F is +needed: when /sys/kernel/btf/vmlinux is read or mmap()ed, when kernel BTF +objects are enumerated (BPF_BTF_GET_NEXT_ID), or when a BPF program needs +kernel type information (an attach_btf_id, a kfunc call, a ksym, a map poi= nter +or a helper that takes or returns a kernel BTF pointer). Until then no mem= ory +is used for it, and afterwards nothing differs from =3Dy. In particular: + + * /sys/kernel/btf/vmlinux exists from boot with its final size. + * Modules loaded before the vmlinux BTF are exposed in /sys/kernel/btf r= ight + away, their BTF is parsed and gets a BTF id once the vmlinux BTF is + loaded, together with their kfunc and struct_ops registrations. + * kfunc, dtor kfunc and struct_ops registrations of the kernel itself are + applied before the BTF becomes visible. + * The kernel only accepts the BTF it was built with: the size and SHA-25= 6 of + the BTF are linked into the kernel and checked against the module. + * Once loaded the BTF stays; the module cannot be unloaded. + +If the module is not available (not installed, or the root file system is = not +mounted yet), the kernel behaves as one built without BTF and retries next +time. CONFIG_BPF_PRELOAD is not available with =3Dm: its iterators attach = through +the vmlinux BTF, so mounting bpffs would load it. bpf_snprintf_btf() and b= pf_seq_printf_btf() only use the BTF if it has +already been parsed, as they run in program context. + 7. Testing =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =20 diff --git a/Makefile b/Makefile index 66654fa71655..0a37decd9d01 100644 --- a/Makefile +++ b/Makefile @@ -1208,7 +1208,8 @@ endif # include additional Makefiles when needed include-y :=3D scripts/Makefile.warn include-$(CONFIG_DEBUG_INFO) +=3D scripts/Makefile.debug -include-$(CONFIG_DEBUG_INFO_BTF)+=3D scripts/Makefile.btf +# CONFIG_DEBUG_INFO_BTF is a tristate; BTF is generated for both y and m +include-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D scripts/Makefile.btf include-$(CONFIG_KASAN) +=3D scripts/Makefile.kasan include-$(CONFIG_KCSAN) +=3D scripts/Makefile.kcsan include-$(CONFIG_KMSAN) +=3D scripts/Makefile.kmsan @@ -1744,8 +1745,9 @@ endif # =20 # *.ko are usually independent of vmlinux, but CONFIG_DEBUG_INFO_BTF_MODUL= ES -# is an exception. -ifdef CONFIG_DEBUG_INFO_BTF_MODULES +# is an exception, and so is the btf_vmlinux module with CONFIG_DEBUG_INFO= _BTF=3Dm, +# which carries the vmlinux BTF. +ifneq ($(CONFIG_DEBUG_INFO_BTF_MODULES)$(filter m,$(CONFIG_DEBUG_INFO_BTF)= ),) KBUILD_BUILTIN :=3D y modules: vmlinux endif diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinu= x.lds.h index b2988aa12f66..9e2f4861fef2 100644 --- a/include/asm-generic/vmlinux.lds.h +++ b/include/asm-generic/vmlinux.lds.h @@ -674,8 +674,17 @@ =20 /* * .BTF + * + * With CONFIG_DEBUG_INFO_BTF=3Dy the vmlinux BTF is loaded as read-only d= ata and + * bounded by __start_BTF/__stop_BTF. With CONFIG_DEBUG_INFO_BTF=3Dm it i= s still + * emitted into the vmlinux ELF file so that module BTF generation and too= ling + * can read it, but as a non-loadable section (see BTF_NOLOAD in ELF_DETAI= LS): + * the btf_vmlinux module carries a copy and provides it on demand at runt= ime. + * What is loaded instead is .BTF.meta, the size and hash of that BTF (see + * scripts/gen-btf.sh), empty in the first link that the BTF is generated + * from. .BTF_ids is needed by the kernel in both cases. */ -#ifdef CONFIG_DEBUG_INFO_BTF +#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF) #define BTF \ . =3D ALIGN(PAGE_SIZE); \ .BTF : AT(ADDR(.BTF) - LOAD_OFFSET) { \ @@ -685,10 +694,28 @@ .BTF_ids : AT(ADDR(.BTF_ids) - LOAD_OFFSET) { \ *(.BTF_ids) \ } +#elif IS_MODULE(CONFIG_DEBUG_INFO_BTF) +#define BTF \ + . =3D ALIGN(8); \ + .BTF.meta : AT(ADDR(.BTF.meta) - LOAD_OFFSET) { \ + BOUNDED_SECTION_BY(.BTF.meta, _BTF_meta) \ + } \ + . =3D ALIGN(PAGE_SIZE); \ + .BTF_ids : AT(ADDR(.BTF_ids) - LOAD_OFFSET) { \ + *(.BTF_ids) \ + } #else #define BTF #endif =20 +#if IS_MODULE(CONFIG_DEBUG_INFO_BTF) +/* quoted: BTF is a macro, an unquoted .BTF here would expand it */ +#define BTF_NOLOAD \ + ".BTF" 0 : { *(".BTF") } +#else +#define BTF_NOLOAD +#endif + /* * Init task */ @@ -849,6 +876,7 @@ /* Required sections not related to debugging. */ #define ELF_DETAILS \ .comment 0 : { *(.comment) } \ + BTF_NOLOAD \ .symtab 0 : { *(.symtab) } \ .strtab 0 : { *(.strtab) } \ .shstrtab 0 : { *(.shstrtab) } \ diff --git a/include/linux/btf_ids.h b/include/linux/btf_ids.h index 8b5a9ee92513..c665afff100e 100644 --- a/include/linux/btf_ids.h +++ b/include/linux/btf_ids.h @@ -22,7 +22,7 @@ struct btf_id_set8 { } pairs[]; }; =20 -#ifdef CONFIG_DEBUG_INFO_BTF +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF) =20 #include /* for __PASTE */ #include /* for __maybe_unused */ diff --git a/include/linux/compiler_types.h b/include/linux/compiler_types.h index c5921f139007..a90a99849cee 100644 --- a/include/linux/compiler_types.h +++ b/include/linux/compiler_types.h @@ -34,7 +34,7 @@ * Skipped when running bindgen due to a libclang issue; * see https://github.com/rust-lang/rust-bindgen/issues/2244. */ -#if defined(CONFIG_DEBUG_INFO_BTF) && defined(CONFIG_PAHOLE_HAS_BTF_TAG) &= & \ +#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF) && defined(CONFIG_PAHOLE_HAS_BTF_TAG= ) && \ __has_attribute(btf_type_tag) && !defined(__BINDGEN__) # define BTF_TYPE_TAG(value) __attribute__((btf_type_tag(#value))) #else diff --git a/include/trace/trace_events.h b/include/trace/trace_events.h index 93011f800d0f..2a0098929771 100644 --- a/include/trace/trace_events.h +++ b/include/trace/trace_events.h @@ -398,7 +398,7 @@ static inline notrace int trace_event_get_offsets_##cal= l( \ #define _TRACE_PERF_INIT(call) #endif /* CONFIG_PERF_EVENTS */ =20 -#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF) +#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF) /* * Per-template BTF id list, populated at link time by resolve_btfids: * [0] FUNC __bpf_trace_ (the BPF dispatcher) diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile index 9a92c348bbda..8ab46f496fa4 100644 --- a/kernel/bpf/Makefile +++ b/kernel/bpf/Makefile @@ -41,7 +41,11 @@ ifeq ($(CONFIG_INET),y) obj-$(CONFIG_BPF_SYSCALL) +=3D reuseport_array.o endif ifeq ($(CONFIG_SYSFS),y) -obj-$(CONFIG_DEBUG_INFO_BTF) +=3D sysfs_btf.o +obj-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D sysfs_btf.o +endif +# With CONFIG_DEBUG_INFO_BTF=3Dm the vmlinux BTF is carried by this module +ifeq ($(CONFIG_DEBUG_INFO_BTF),m) +obj-m +=3D btf_vmlinux.o endif ifeq ($(CONFIG_BPF_JIT),y) obj-$(CONFIG_BPF_SYSCALL) +=3D bpf_struct_ops.o diff --git a/kernel/bpf/btf_vmlinux.c b/kernel/bpf/btf_vmlinux.c new file mode 100644 index 000000000000..8d89b4bb3c43 --- /dev/null +++ b/kernel/bpf/btf_vmlinux.c @@ -0,0 +1,23 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Carrier module for the vmlinux BTF when CONFIG_DEBUG_INFO_BTF=3Dm. + * + * This module has no code of its own. Its .BTF section is a copy of the + * vmlinux BTF (see scripts/gen-btf.sh), which the BTF module notifier in + * kernel/bpf/btf.c recognizes by module name and installs as the vmlinux = BTF. + * The kernel loads it on demand, the first time the vmlinux BTF is needed. + * + * There is deliberately no module_exit(): once the BTF is in use it cannot + * be taken away again, exactly as with CONFIG_DEBUG_INFO_BTF=3Dy. + */ +#include +#include + +static int __init btf_vmlinux_init(void) +{ + return 0; +} +module_init(btf_vmlinux_init); + +MODULE_DESCRIPTION("BTF type information for vmlinux"); +MODULE_LICENSE("GPL"); diff --git a/kernel/bpf/preload/Kconfig b/kernel/bpf/preload/Kconfig index aef7b0bc96d6..b1600bdce7a0 100644 --- a/kernel/bpf/preload/Kconfig +++ b/kernel/bpf/preload/Kconfig @@ -6,6 +6,10 @@ menuconfig BPF_PRELOAD # The dependency on !COMPILE_TEST prevents it from being enabled # in allmodconfig or allyesconfig configurations depends on !COMPILE_TEST + # The preloaded iterators attach through the vmlinux BTF, so with + # CONFIG_DEBUG_INFO_BTF=3Dm every bpffs mount would load the BTF, which + # defeats the point of =3Dm on any system that mounts bpffs at boot. + depends on DEBUG_INFO_BTF!=3Dm help This builds kernel module with several embedded BPF programs that are pinned into BPF FS mount point as human readable files that are diff --git a/kernel/trace/trace_syscalls.c b/kernel/trace/trace_syscalls.c index e35744049e3f..7a0d59c308c2 100644 --- a/kernel/trace/trace_syscalls.c +++ b/kernel/trace/trace_syscalls.c @@ -1304,7 +1304,7 @@ struct trace_event_functions exit_syscall_print_funcs= =3D { .trace =3D print_syscall_exit, }; =20 -#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF) +#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF) /* BTF id lists for the shared sys_enter/sys_exit dispatcher tracepoints. = */ BTF_ID_LIST(syscall_enter_btf_ids) BTF_ID(func, __bpf_trace_sys_enter) @@ -1321,7 +1321,7 @@ struct trace_event_class __refdata event_class_syscal= l_enter =3D { .fields_array =3D syscall_enter_fields_array, .get_fields =3D syscall_get_enter_fields, .raw_init =3D init_syscall_trace, -#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF) +#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF) .btf_ids =3D syscall_enter_btf_ids, #endif }; @@ -1336,7 +1336,7 @@ struct trace_event_class __refdata event_class_syscal= l_exit =3D { }, .fields =3D LIST_HEAD_INIT(event_class_syscall_exit.fields), .raw_init =3D init_syscall_trace, -#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF) +#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF) .btf_ids =3D syscall_exit_btf_ids, #endif }; diff --git a/lib/Kconfig.debug b/lib/Kconfig.debug index 134b15a44625..5307caa39176 100644 --- a/lib/Kconfig.debug +++ b/lib/Kconfig.debug @@ -396,7 +396,7 @@ config DEBUG_INFO_SPLIT Incompatible with older versions of ccache. =20 config DEBUG_INFO_BTF - bool "Generate BTF type information" + tristate "Generate BTF type information" depends on !DEBUG_INFO_SPLIT && !DEBUG_INFO_REDUCED depends on !GCC_PLUGIN_RANDSTRUCT || COMPILE_TEST depends on BPF_SYSCALL @@ -408,6 +408,17 @@ config DEBUG_INFO_BTF Turning this on requires pahole v1.22 or later, which will convert DWARF type info into equivalent deduplicated BTF type info. =20 + If built as a module (=3Dm), the vmlinux BTF is not part of the + kernel image. It is carried by the btf_vmlinux module, which is + loaded on demand the first time the BTF is needed: when a BPF + program requires kernel type information, or when + /sys/kernel/btf/vmlinux is opened. Until then, no memory is + spent on it. The BTF is still emitted into the vmlinux ELF file + (as a non-loadable section) so tooling and module BTF generation + work as before. Module BTF (DEBUG_INFO_BTF_MODULES) is registered + when the vmlinux BTF becomes available. Not compatible with + BPF_PRELOAD, whose iterators would load the BTF at every bpffs mount. + config PAHOLE_HAS_BTF_TAG def_bool PAHOLE_VERSION >=3D 123 depends on CC_IS_CLANG diff --git a/net/netfilter/Makefile b/net/netfilter/Makefile index 6bf74d488a29..a2c7f00794d2 100644 --- a/net/netfilter/Makefile +++ b/net/netfilter/Makefile @@ -18,7 +18,7 @@ nf_conntrack-$(CONFIG_NF_CT_PROTO_GRE) +=3D nf_conntrack_= proto_gre.o ifeq ($(CONFIG_NF_CONNTRACK),m) nf_conntrack-$(CONFIG_DEBUG_INFO_BTF_MODULES) +=3D nf_conntrack_bpf.o else ifeq ($(CONFIG_NF_CONNTRACK),y) -nf_conntrack-$(CONFIG_DEBUG_INFO_BTF) +=3D nf_conntrack_bpf.o +nf_conntrack-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D nf_conntrack_bpf.o endif =20 obj-$(CONFIG_NETFILTER) =3D netfilter.o @@ -65,7 +65,7 @@ nf_nat-$(CONFIG_NF_NAT_OVS) +=3D nf_nat_ovs.o ifeq ($(CONFIG_NF_NAT),m) nf_nat-$(CONFIG_DEBUG_INFO_BTF_MODULES) +=3D nf_nat_bpf.o else ifeq ($(CONFIG_NF_NAT),y) -nf_nat-$(CONFIG_DEBUG_INFO_BTF) +=3D nf_nat_bpf.o +nf_nat-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D nf_nat_bpf.o endif =20 # NAT helpers @@ -147,7 +147,7 @@ nf_flow_table-$(CONFIG_NF_FLOW_TABLE_PROCFS) +=3D nf_fl= ow_table_procfs.o ifeq ($(CONFIG_NF_FLOW_TABLE),m) nf_flow_table-$(CONFIG_DEBUG_INFO_BTF_MODULES) +=3D nf_flow_table_bpf.o else ifeq ($(CONFIG_NF_FLOW_TABLE),y) -nf_flow_table-$(CONFIG_DEBUG_INFO_BTF) +=3D nf_flow_table_bpf.o +nf_flow_table-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D nf_flow_table_bpf= .o endif =20 obj-$(CONFIG_NF_FLOW_TABLE_INET) +=3D nf_flow_table_inet.o diff --git a/net/xfrm/Makefile b/net/xfrm/Makefile index 5a1787587cb3..b7f6e5046a0e 100644 --- a/net/xfrm/Makefile +++ b/net/xfrm/Makefile @@ -8,7 +8,7 @@ xfrm_interface-$(CONFIG_XFRM_INTERFACE) +=3D xfrm_interface= _core.o ifeq ($(CONFIG_XFRM_INTERFACE),m) xfrm_interface-$(CONFIG_DEBUG_INFO_BTF_MODULES) +=3D xfrm_interface_bpf.o else ifeq ($(CONFIG_XFRM_INTERFACE),y) -xfrm_interface-$(CONFIG_DEBUG_INFO_BTF) +=3D xfrm_interface_bpf.o +xfrm_interface-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D xfrm_interface_b= pf.o endif =20 obj-$(CONFIG_XFRM) :=3D xfrm_policy.o xfrm_state.o xfrm_hash.o \ @@ -23,4 +23,4 @@ obj-$(CONFIG_XFRM_IPCOMP) +=3D xfrm_ipcomp.o obj-$(CONFIG_XFRM_INTERFACE) +=3D xfrm_interface.o obj-$(CONFIG_XFRM_IPTFS) +=3D xfrm_iptfs.o obj-$(CONFIG_XFRM_ESPINTCP) +=3D espintcp.o -obj-$(CONFIG_DEBUG_INFO_BTF) +=3D xfrm_state_bpf.o +obj-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) +=3D xfrm_state_bpf.o diff --git a/scripts/Makefile.modfinal b/scripts/Makefile.modfinal index 01a37ec872b9..ad182f84b5fc 100644 --- a/scripts/Makefile.modfinal +++ b/scripts/Makefile.modfinal @@ -46,12 +46,18 @@ quiet_cmd_btf_ko =3D BTF [M] $@ $(CONFIG_SHELL) $(srctree)/scripts/gen-btf.sh --btf_base $(objtree)/vmli= nux $@; \ fi; =20 -# Re-generate module BTFs if either module's .ko or vmlinux changed -%.ko: %.o %.mod.o .module-common.o $(objtree)/scripts/module.lds $(and $(C= ONFIG_DEBUG_INFO_BTF_MODULES),$(KBUILD_BUILTIN),$(objtree)/vmlinux) FORCE - +$(call if_changed,ld_ko_o) +# Modules that get a .BTF section: all of them with CONFIG_DEBUG_INFO_BTF_= MODULES, +# otherwise only the vmlinux BTF carrier module with CONFIG_DEBUG_INFO_BTF= =3Dm. ifdef CONFIG_DEBUG_INFO_BTF_MODULES - +$(if $(newer-prereqs),$(call cmd,btf_ko)) +btf-modules :=3D $(modules:%.o=3D%.ko) +else ifeq ($(CONFIG_DEBUG_INFO_BTF),m) +btf-modules :=3D $(filter %/btf_vmlinux.ko,$(modules:%.o=3D%.ko)) endif + +# Re-generate module BTFs if either module's .ko or vmlinux changed +%.ko: %.o %.mod.o .module-common.o $(objtree)/scripts/module.lds $(and $(b= tf-modules),$(KBUILD_BUILTIN),$(objtree)/vmlinux) FORCE + +$(call if_changed,ld_ko_o) + +$(if $(and $(filter $@,$(btf-modules)),$(newer-prereqs)),$(call cmd,btf_= ko)) +$(call cmd,check_tracepoint) =20 targets +=3D $(modules:%.o=3D%.ko) $(modules:%.o=3D%.mod.o) .module-common= .o diff --git a/scripts/gen-btf.sh b/scripts/gen-btf.sh index 8ca96eb10a69..7fa3189a3ded 100755 --- a/scripts/gen-btf.sh +++ b/scripts/gen-btf.sh @@ -22,16 +22,26 @@ # - ${1}.btf.o ready for linking into vmlinux # - ${1}.BTF_ids with .BTF_ids data blob # This output is consumed by scripts/link-vmlinux.sh +# +# With CONFIG_DEBUG_INFO_BTF=3Dm the .BTF section in ${1}.btf.o is not +# allocatable, so the vmlinux ELF file carries the BTF but the kernel image +# does not. ${1}.btf.o then also carries .BTF.meta, the size and SHA-256 = of +# the BTF for the kernel (struct btf_vmlinux_meta); "--placeholder ${1}" +# produces a ${1}.btf.o with a zeroed .BTF.meta and no .BTF for the first +# vmlinux link, which the BTF is generated from. The btf_vmlinux module g= ets +# no BTF of its own; its .BTF section is a copy of the vmlinux BTF, extrac= ted +# from --btf_base. =20 set -e =20 usage() { - echo "Usage: $0 [--btf_base ] " + echo "Usage: $0 [--btf_base ] [--placeholder] " exit 1 } =20 BTF_BASE=3D"" +PLACEHOLDER=3D"" =20 while [ $# -gt 0 ]; do case "$1" in @@ -39,6 +49,10 @@ while [ $# -gt 0 ]; do BTF_BASE=3D"$2" shift 2 ;; + --placeholder) + PLACEHOLDER=3D1 + shift + ;; -*) echo "Unknown option: $1" >&2 usage @@ -60,6 +74,10 @@ is_enabled() { grep -q "^$1=3Dy" ${objtree}/include/config/auto.conf } =20 +is_module() { + grep -q "^$1=3Dm" ${objtree}/include/config/auto.conf +} + case "${KBUILD_VERBOSE}" in *1*) set -x @@ -79,6 +97,30 @@ gen_btf_data() --btf ${btf1} "${ELF_FILE}" } =20 +# Write one byte with value $1 (0..255) +put_byte() +{ + printf "\\$(printf '%03o' "$1")" +} + +# CONFIG_DEBUG_INFO_BTF=3Dm: write struct btf_vmlinux_meta { u32 size; u8 +# sha256[32]; } for the BTF in $1 to $2, in the target's byte order. +gen_btf_meta() +{ + size=3D$(${CONFIG_SHELL} "${srctree}/scripts/file-size.sh" "$1") + sha256=3D$(sha256sum < "$1" | cut -d' ' -f1) + { + if is_enabled CONFIG_CPU_BIG_ENDIAN; then + for shift in 24 16 8 0; do put_byte $(( (size >> shift) & 255 )); done + else + for shift in 0 8 16 24; do put_byte $(( (size >> shift) & 255 )); done + fi + for byte in $(echo "${sha256}" | sed 's/../& /g'); do + put_byte $(( 0x${byte} )) + done + } > "$2" +} + gen_btf_o() { btf_data=3D${ELF_FILE}.btf.o @@ -88,9 +130,23 @@ gen_btf_o() # deletes all symbols including __start_BTF and __stop_BTF, which will # be redefined in the linker script. echo "" | ${CC} ${CLANG_FLAGS} ${KBUILD_CPPFLAGS} ${KBUILD_CFLAGS} -fno-l= to -c -x c -o ${btf_data} - - ${OBJCOPY} --add-section .BTF=3D${ELF_FILE}.BTF \ - --set-section-flags .BTF=3Dalloc,readonly ${btf_data} - ${OBJCOPY} --only-section=3D.BTF --strip-all ${btf_data} + if is_module CONFIG_DEBUG_INFO_BTF; then + # CONFIG_DEBUG_INFO_BTF=3Dm: .BTF stays non-allocatable, kept in the + # vmlinux ELF file for tooling but not loaded; the btf_vmlinux + # module provides it at runtime. What is loaded is .BTF.meta, + # its size and hash, so that /sys/kernel/btf/vmlinux has the right + # size from boot and only the matching BTF is accepted. + gen_btf_meta ${ELF_FILE}.BTF ${ELF_FILE}.BTF.meta + ${OBJCOPY} --add-section .BTF=3D${ELF_FILE}.BTF \ + --set-section-flags .BTF=3Dreadonly \ + --add-section .BTF.meta=3D${ELF_FILE}.BTF.meta \ + --set-section-flags .BTF.meta=3Dalloc,readonly ${btf_data} + ${OBJCOPY} --only-section=3D.BTF --only-section=3D.BTF.meta --strip-all = ${btf_data} + else + ${OBJCOPY} --add-section .BTF=3D${ELF_FILE}.BTF \ + --set-section-flags .BTF=3Dalloc,readonly ${btf_data} + ${OBJCOPY} --only-section=3D.BTF --strip-all ${btf_data} + fi =20 # Change e_type to ET_REL so that it can be used to link final vmlinux. # GNU ld 2.35+ and lld do not allow an ET_EXEC input. @@ -121,6 +177,7 @@ cleanup() { rm -f "${ELF_FILE}.BTF.1" rm -f "${ELF_FILE}.BTF" + rm -f "${ELF_FILE}.BTF.meta" if [ "${BTFGEN_MODE}" =3D "module" ]; then rm -f "${ELF_FILE}.BTF.base" rm -f "${ELF_FILE}.BTF_ids" @@ -133,6 +190,34 @@ if [ -n "${BTF_BASE}" ]; then BTFGEN_MODE=3D"module" fi =20 +if [ -n "${PLACEHOLDER}" ]; then + btf_data=3D${ELF_FILE}.btf.o + echo "" | ${CC} ${CLANG_FLAGS} ${KBUILD_CPPFLAGS} ${KBUILD_CFLAGS} -fno-l= to -c -x c -o ${btf_data} - + head -c 36 /dev/zero > ${ELF_FILE}.BTF.meta + ${OBJCOPY} --add-section .BTF.meta=3D${ELF_FILE}.BTF.meta \ + --set-section-flags .BTF.meta=3Dalloc,readonly ${btf_data} + ${OBJCOPY} --only-section=3D.BTF.meta --strip-all ${btf_data} + exit 0 +fi + +# CONFIG_DEBUG_INFO_BTF=3Dm: the btf_vmlinux module carries the vmlinux BTF +# itself. Its own types are of no interest, so instead of generating split +# BTF for it, copy the (non-loadable) .BTF section of vmlinux into the mod= ule. +# The kernel recognizes the module by name and treats its .BTF as base BTF. +case "${BTFGEN_MODE}:${ELF_FILE}" in +module:*/btf_vmlinux.ko) + if is_module CONFIG_DEBUG_INFO_BTF; then + # -O binary only emits allocatable sections; make .BTF one for + # the extraction. ${BTF_BASE} itself is not modified. + ${OBJCOPY} -O binary --only-section=3D.BTF \ + --set-section-flags .BTF=3Dalloc,load,readonly \ + "${BTF_BASE}" "${ELF_FILE}.BTF" + ${OBJCOPY} --add-section .BTF=3D"${ELF_FILE}.BTF" "${ELF_FILE}" + exit 0 + fi + ;; +esac + gen_btf_data =20 case "${BTFGEN_MODE}" in diff --git a/scripts/link-vmlinux.sh b/scripts/link-vmlinux.sh index ab0b8125c8cb..a9a066267ef6 100755 --- a/scripts/link-vmlinux.sh +++ b/scripts/link-vmlinux.sh @@ -37,6 +37,15 @@ is_enabled() { grep -q "^$1=3Dy" include/config/auto.conf } =20 +is_module() { + grep -q "^$1=3Dm" include/config/auto.conf +} + +# =3Dy or =3Dm +is_set() { + grep -q "^$1=3D[ym]" include/config/auto.conf +} + # Nice output in kbuild format # Will be suppressed by "make -s" info() @@ -211,17 +220,25 @@ if is_enabled CONFIG_KALLSYMS; then kallsyms .tmp_vmlinux0.syms .tmp_vmlinux0.kallsyms fi =20 -if is_enabled CONFIG_KALLSYMS || is_enabled CONFIG_DEBUG_INFO_BTF; then +if is_module CONFIG_DEBUG_INFO_BTF; then + # The kernel refers to the size and hash of its BTF, which only the + # BTF generated from the first link can provide; link a placeholder + # of the same layout until then. + ${CONFIG_SHELL} ${srctree}/scripts/gen-btf.sh --placeholder .tmp_vmlinux0 + btf_vmlinux_bin_o=3D.tmp_vmlinux0.btf.o +fi + +if is_enabled CONFIG_KALLSYMS || is_set CONFIG_DEBUG_INFO_BTF; then =20 # The kallsyms linking does not need debug symbols, but the BTF does. - if ! is_enabled CONFIG_DEBUG_INFO_BTF; then + if ! is_set CONFIG_DEBUG_INFO_BTF; then strip_debug=3D1 fi =20 vmlinux_link .tmp_vmlinux1 fi =20 -if is_enabled CONFIG_DEBUG_INFO_BTF; then +if is_set CONFIG_DEBUG_INFO_BTF; then info BTF .tmp_vmlinux1 if ! ${CONFIG_SHELL} ${srctree}/scripts/gen-btf.sh .tmp_vmlinux1; then echo >&2 "Failed to generate BTF for vmlinux" @@ -287,7 +304,7 @@ fi =20 vmlinux_link "${VMLINUX}" =20 -if is_enabled CONFIG_DEBUG_INFO_BTF; then +if is_set CONFIG_DEBUG_INFO_BTF; then info BTFIDS ${VMLINUX} ${RESOLVE_BTFIDS} --patch_btfids ${btfids_vmlinux} ${VMLINUX} fi --=20 2.47.3