From nobody Mon Sep 28 08:07:25 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 37277367298; Mon, 24 Aug 2026 14:38:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582292; cv=none; b=O8meRb84VLcXkXsaiPkRZeBoWH7neHqSdISai2DuoIpwG4w1MKFq4C+ayO003yH2TXXWygWYYEWwyEZvfbqfk1Mgh1XkrrM+HJdWeW+igNyh5f62Y42GUI5PzPjWvt13oEC2oxKkB+6Sy+8axR0c25EsPJwjQa0nDqdYPg9SwRM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582292; c=relaxed/simple; bh=P3AdyzdShDrUb/YgX9pMKwXXddwQSm7F29np86e13z4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JHS+TYKrJfuDCD81SSCrLw/oGW2tsa1j5dyOp9KilIv5bGYml8p+wtRRKz5vKLuuLaGEqovkl40+GNYxDLOSuEs/W+A3HrN4lGn5lio70aeRZHdAf51zvwmeWMXmpf4MIXRR8B1LkkAGZByVmZKOkiHV/hIg9yknzz455gog4oc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=MkPPT7bD; arc=none smtp.client-ip=117.135.210.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="MkPPT7bD" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=l+ yg31nd4nBh5hM/Aoz4/e9I3HPD3hBY2IDhBbqEslg=; b=MkPPT7bDm9o2bXvAnC 0yV5vGg5DGKFV2+ZJT87YgdgX/J+l+wWqwY8AGrpDB27Dv+/Y2FjMpuKdiBa9fc5 huI+U9U71G8IYI+Z4Powkx5XOoa59pjOdQfmdqlDNkL2r4UEKIsWTHeh6przAeZ8 Md89TtbvyMG7zrOxDjmghBOzE= Received: from nec8-i7 (unknown []) by gzsmtp2 (Coremail) with SMTP id PSgvCgDn+osAV4xqIt5SMw--.54939S3; Mon, 24 Aug 2026 22:36:50 +0800 (CST) From: chenyuan_fl@163.com To: bpf@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Shuah Khan , Nuoqi Gui , Yuan Chen Subject: [PATCH 1/4] bpf: Cancel special fields in resizable hashtab on recycle Date: Mon, 24 Aug 2026 22:36:18 +0800 Message-ID: <20260824143621.2098856-2-chenyuan_fl@163.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260824143621.2098856-1-chenyuan_fl@163.com> References: <20260824143621.2098856-1-chenyuan_fl@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PSgvCgDn+osAV4xqIt5SMw--.54939S3 X-Coremail-Antispam: 1Uf129KBjvJXoW3WF4ftF43Kry5Kw4xAr1kKrg_yoWxWw18pF Z5Wr13Cr1kJrnIqrZ0yw4vkrWrZ3s5tr4YkFZ5GryF934rXF97Jr1fJa97uF1jyF1vvFnY qF4IqFW3ua1UC37anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x07jalksUUUUU= X-CM-SenderInfo: xfkh05pxdqswro6rljoofrz/xtbDAQKTUWqMVwKt8AAA3v Content-Type: text/plain; charset="utf-8" From: Yuan Chen rhtab_delete_elem() and rhtab_map_update_existing() eagerly call bpf_obj_free_fields() when an element is deleted or its value is replaced. This runs kptr destructors in the caller's execution context, which is unsafe for BPF programs running in NMI context (e.g. perf_event programs attached to hardware PMU overflows): referenced kptr destructors may take locks or otherwise cannot run in NMI. Commit a3a81d247651 ("bpf: Cancel special fields on map value recycle") switched the hash map and array recycle paths to bpf_obj_cancel_fields(), which only cancels NMI-safe fields (timer, workqueue, task_work), but it missed the resizable hashtab. rhtab_map_update_existing() even documents the intended "cancel" semantics while still calling bpf_obj_free_fields(). Fix the resizable hashtab the same way: * rhtab_delete_elem() and rhtab_map_update_existing() now cancel only NMI-safe fields. Referenced kptrs stay attached to the recycled element and are destroyed by rhtab_mem_dtor() once the element is eventually freed, keeping the reference accounting balanced. * rhtab_map_update_elem() initializes the special fields of a freshly allocated element. The bpf memory allocator may return a recycled element that still owns a referenced kptr, and check_and_init_map_value() would zero that slot, dropping the reference without releasing it. rhtab_init_map_value() initializes the remaining fields (spin lock, timer, workqueue, task_work, refcount) but leaves kptr slots untouched, matching the hash map semantics. Verified with a selftest: a perf_event (NMI) program overwrites a rhtab element that holds a referenced task kptr, and a second phase deletes and re-inserts the element to exercise the recycle path. Before the patch the NMI update eagerly released the kptr and the recycle path zeroed the inherited slot; after the patch the kptr is inherited on both paths and the probe observes it non-NULL. Fixes: a3a81d247651 ("bpf: Cancel special fields on map value recycle") Signed-off-by: Yuan Chen --- kernel/bpf/hashtab.c | 70 +++++++++++++++++++++++++++++++++++++------- 1 file changed, 60 insertions(+), 10 deletions(-) diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c index d40cb5dd446c..0df8db27cd8c 100644 --- a/kernel/bpf/hashtab.c +++ b/kernel/bpf/hashtab.c @@ -2864,14 +2864,56 @@ static int rhtab_map_alloc_check(union bpf_attr *at= tr) return htab_map_alloc_check(attr); } =20 -static void rhtab_check_and_free_fields(struct bpf_rhtab *rhtab, - struct rhtab_elem *elem) +static void rhtab_cancel_fields(struct bpf_rhtab *rhtab, + struct rhtab_elem *elem) { if (IS_ERR_OR_NULL(rhtab->map.record)) return; =20 - bpf_obj_free_fields(rhtab->map.record, - rhtab_elem_value(elem, rhtab->map.key_size)); + /* + * Only cancel NMI-safe fields (timer, workqueue, task_work) here. + * RHASH values can also carry referenced kptrs (and per-cpu kptrs), + * whose destructors must not run from arbitrary BPF execution + * contexts (e.g. NMI); leave them attached to the recycled element + * and let rhtab_mem_dtor() destroy them once the element is + * eventually freed. This matches the hash map semantics introduced + * by a3a81d247651 ("bpf: Cancel special fields on map value + * recycle"). + */ + bpf_map_free_internal_structs(&rhtab->map, + rhtab_elem_value(elem, rhtab->map.key_size)); +} + +/* + * Initialize special fields of a freshly allocated rhtab element, but keep + * kptr fields untouched. A recycled element may carry a referenced kptr f= rom + * its previous life: the delete path only cancels NMI-safe fields (matchi= ng + * the hash map semantics), so the kptr reference stays owned by the eleme= nt + * until rhtab_mem_dtor() destroys it. Zeroing it here (as + * check_and_init_map_value() would) would drop the reference without + * releasing it. + */ +static void rhtab_init_map_value(struct bpf_map *map, void *value) +{ + struct btf_record *rec =3D map->record; + int i; + + if (IS_ERR_OR_NULL(rec)) + return; + + for (i =3D 0; i < rec->cnt; i++) { + struct btf_field *field =3D &rec->fields[i]; + void *field_ptr =3D value + field->offset; + + switch (field->type) { + case BPF_KPTR_UNREF: + case BPF_KPTR_REF: + case BPF_KPTR_PERCPU: + continue; + default: + bpf_obj_init_field(field, field_ptr); + } + } } =20 static void rhtab_mem_dtor(void *obj, void *ctx) @@ -2963,8 +3005,8 @@ static int rhtab_delete_elem(struct bpf_rhtab *rhtab,= struct rhtab_elem *elem, v rhtab_read_elem_value(&rhtab->map, copy, elem, flags); check_and_init_map_value(&rhtab->map, copy); } - /* Release internal structs: kptr, bpf_timer, task_work, wq */ - rhtab_check_and_free_fields(rhtab, elem); + /* Cancel NMI-safe fields; full destruction happens in rhtab_mem_dtor */ + rhtab_cancel_fields(rhtab, elem); bpf_mem_cache_free_rcu(&rhtab->ma, elem); return 0; } @@ -3022,10 +3064,11 @@ static long rhtab_map_update_existing(struct bpf_ma= p *map, struct rhtab_elem *el * BPF_F_LOCK, matching arraymap semantics. * * copy_map_value() skips special-field offsets, so old timers/ - * kptrs/etc. still sit in the slot. Cancel them after the copy - * to match arraymap's update semantics. + * kptrs/etc. still sit in the slot. Cancel the NMI-safe ones after + * the copy to match arraymap's update semantics; referenced kptrs + * stay attached and are destroyed by rhtab_mem_dtor(). */ - rhtab_check_and_free_fields(rhtab, elem); + rhtab_cancel_fields(rhtab, elem); return 0; } =20 @@ -3066,7 +3109,14 @@ static long rhtab_map_update_elem(struct bpf_map *ma= p, void *key, void *value, u =20 memcpy(elem->data, key, map->key_size); copy_map_value(map, rhtab_elem_value(elem, map->key_size), value); - check_and_init_map_value(map, rhtab_elem_value(elem, map->key_size)); + /* + * Initialize special fields of the (possibly recycled) element, but + * leave kptr slots alone: a recycled element may still own a + * referenced kptr that rhtab_mem_dtor() will release, so zeroing it + * here would leak the reference. Fresh memory from the bpf mem + * allocator is zeroed, so skipping the kptr init is safe there too. + */ + rhtab_init_map_value(map, rhtab_elem_value(elem, map->key_size)); =20 /* Prevent deadlock for NMI programs attempting to take bucket lock */ bpf_disable_instrumentation(); --=20 2.54.0 From nobody Mon Sep 28 08:07:25 2026 Received: from m16.mail.163.com (m16.mail.163.com [220.197.31.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 199D3438011; Mon, 24 Aug 2026 14:38:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=220.197.31.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582300; cv=none; b=qq5oD5u7Qq/reYja6JcH6vqmIr7bd1dkIMGYa+g+Ox2Y0u4tWIuheoIn8fJQbRUUDErFzdegW/pFW3faWr13r0c6ouPD3Gm6PlG1kWW5DtNZnzAjO0jADyl/J4EAWu/kN7Ou6zTF2V9y7BBuBDkqUao945gsAC5+BOXE/fUhjDE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582300; c=relaxed/simple; bh=yiUiflDVanlKSVYyi+qbFEeVwBkElRInG2Qf/yLZgbw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XTJZ9mm1j1hRtY0yH0RzAQiXp3zeIZcJKR9dH75ZeFx8x+ayVcDNGuC1fPeLVl4PVmdkvWrCnFaefQZ0Y4lEKYzK3dAHxbCxns+g/y0GocvcV04qpaggg+12QNw4a2m8o+KJxrTlnZ5HK44x53EbMfq+qRrC8DO3skIIWbi6K4w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=Xxio33sw; arc=none smtp.client-ip=220.197.31.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="Xxio33sw" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=3E TUOZnkD0p5Jk119NJQbzsZZOFcufeLhAP1I1izjpY=; b=Xxio33swPz5Fhapz34 0thUa1EdnAAF7nDO2Di02YvIL44LaR7+QsPK5+QU2BilxOsqgSuHkAn845i8GCsA nOiaVtAjKTgfihivrVehZzjBXgAa+RG/tsF0pzTqgVmWIrqKZeCVxomR8HWXd80M XPGMufncJIfx/9eDJn4O5DyzY= Received: from nec8-i7 (unknown []) by gzsmtp2 (Coremail) with SMTP id PSgvCgDn+osAV4xqIt5SMw--.54939S4; Mon, 24 Aug 2026 22:36:50 +0800 (CST) From: chenyuan_fl@163.com To: bpf@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Shuah Khan , Nuoqi Gui , Yuan Chen Subject: [PATCH 2/4] bpf: Fix use-after-free of program BTF in mem-alloc destructor Date: Mon, 24 Aug 2026 22:36:19 +0800 Message-ID: <20260824143621.2098856-3-chenyuan_fl@163.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260824143621.2098856-1-chenyuan_fl@163.com> References: <20260824143621.2098856-1-chenyuan_fl@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PSgvCgDn+osAV4xqIt5SMw--.54939S4 X-Coremail-Antispam: 1Uf129KBjvJXoWxZw4kZFWDGFy3Kr4rCw45Wrg_yoWrJFy5pF 4xCr4akw4ktrs7CwsxWwsxCry5Wan7ZF1UuFyfWw1Ygw4rXr1DXr409FWa9FW5Cr4rtws5 Ar1jgFZxGry5AFDanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x07jwZ2fUUUUU= X-CM-SenderInfo: xfkh05pxdqswro6rljoofrz/xtbDAQKTUWqMVwKt+AAA3n Content-Type: text/plain; charset="utf-8" From: Yuan Chen bpf_ma_set_dtor() duplicates the map's btf_record for the bpf_mem_alloc destructor. For kptr fields backed by the program BTF (MEM_ALLOC kptrs, e.g. objects allocated with bpf_obj_new()/bpf_percpu_obj_new()), btf_record_dup() only borrows the reference, matching what btf_parse_fields() did for the map's own record. The duplicated record, however, is released later from the deferred bpf_mem_alloc destructor workqueue (free_mem_alloc_deferred), by which time the program BTF may already have been freed: bpf_map_free() drops the map's own reference, and the RCU callback can run before the workqueue. Reading field->kptr.btf in btf_record_free() (via btf_is_kernel()) is then a use-after-free, detected by KASAN as "slab-use-after-free in btf_is_kernel" when a map with a MEM_ALLOC kptr field is destroyed. Hold a reference on program BTF for the lifetime of the duplicated record and drop it right before the record is freed. The last btf_put() only schedules the object for RCU destruction, so btf_record_free() can still safely read the field descriptors. The rhtab kptr selftests exercise this path on every map teardown and triggered the bug under KASAN; with this fix they pass cleanly. Fixes: 1df97a7453ee ("bpf: Register dtor for freeing special fields") Signed-off-by: Yuan Chen --- kernel/bpf/hashtab.c | 45 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 45 insertions(+) diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c index 0df8db27cd8c..b8df2bc9a9a0 100644 --- a/kernel/bpf/hashtab.c +++ b/kernel/bpf/hashtab.c @@ -493,10 +493,54 @@ static void htab_pcpu_mem_dtor(void *obj, void *ctx) bpf_obj_free_fields(hrec->record, per_cpu_ptr(pptr, cpu)); } =20 +/* + * bpf_ma_set_dtor() duplicates the map's btf_record. For kptr fields whose + * btf is the program BTF (MEM_ALLOC kptrs, e.g. objects allocated with + * bpf_obj_new()/bpf_percpu_obj_new()) btf_record_dup() only borrows the + * reference, like btf_parse_fields() did for the map's own record. The + * duplicated record is released later from the deferred bpf_mem_alloc + * destructor workqueue, by which time the program BTF may already have be= en + * freed (the map dropped its own reference in bpf_map_free()), so reading + * field->kptr.btf there would be a use-after-free. + * + * Hold a reference on non-kernel (program) BTF for the lifetime of the + * duplicated record and release it before the record is freed. After the + * last btf_put() the object is only destroyed after an RCU grace period, = so + * btf_record_free() can still safely read the field descriptors. + */ +static void htab_record_prog_btf_ref(struct btf_record *rec, bool get) +{ + int i; + + if (IS_ERR_OR_NULL(rec)) + return; + + for (i =3D 0; i < rec->cnt; i++) { + const struct btf_field *field =3D &rec->fields[i]; + + switch (field->type) { + case BPF_KPTR_UNREF: + case BPF_KPTR_REF: + case BPF_KPTR_PERCPU: + case BPF_UPTR: + if (field->kptr.btf && !btf_is_kernel(field->kptr.btf)) { + if (get) + btf_get(field->kptr.btf); + else + btf_put(field->kptr.btf); + } + break; + default: + break; + } + } +} + static void htab_dtor_ctx_free(void *ctx) { struct htab_btf_record *hrec =3D ctx; =20 + htab_record_prog_btf_ref(hrec->record, false); btf_record_free(hrec->record); kfree(ctx); } @@ -521,6 +565,7 @@ static int bpf_ma_set_dtor(struct bpf_map *map, struct = bpf_mem_alloc *ma, kfree(hrec); return err; } + htab_record_prog_btf_ref(hrec->record, true); bpf_mem_alloc_set_dtor(ma, dtor, htab_dtor_ctx_free, hrec); return 0; } --=20 2.54.0 From nobody Mon Sep 28 08:07:25 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.3]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B7DAF363C5F; Mon, 24 Aug 2026 14:38:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.3 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582292; cv=none; b=f55i+MH1QSsns1y004xDXKt/sODxGMCa7tEJGCIifZNYpLl+/TrkD39nZyFca3BEY1A6KdxGDSlIL75NPQ4KYnUsuDN1zc3Lslp1JZxx6xq+36SCshktOLTICBfpf4SCc+XGx3E0GOvGw6UF22OYTnDAmT80PMqvmnSFRocvF7s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582292; c=relaxed/simple; bh=8TpAxauxOe5fg0/lVEVtXpzKKFdzA6e7fIzOJyWXBWY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uKq2c8XxeEYiv4IN7rwgDC9xVSDjxiNYdtF25qgU30rDBEBhEolmHUOE/iDaIaEA/J+8sSBLYZMbPWw9BwtGnozt/l0JYuCqSjP7yH3SbXw86d0zk+uUJAIkmjqVf54XYXoD4o1UmpCXoCOebf3Mo12VFJYgDGoBWxWVnUJX4KM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=E3bLIyD8; arc=none smtp.client-ip=117.135.210.3 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="E3bLIyD8" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=ct +peqGCPMUcfHs6HCx5ClFNeJruycb7Blf1DOzAfzE=; b=E3bLIyD8hu9ifPsZJx NkK9TflrhEhvcd9Ebv8G5Xye2Yhvi78qf1dgN95kiwAhHtjSKj4bM2PK8SwhCyFR r6svFSOIfs0LghMFdMEAGI8ZbLhq3y9Ne2jUDDGDdZYgXsZze0tC4nzbc0YTh8rk o5mwo05y78wd+jRnrymkl6VS4= Received: from nec8-i7 (unknown []) by gzsmtp2 (Coremail) with SMTP id PSgvCgDn+osAV4xqIt5SMw--.54939S5; Mon, 24 Aug 2026 22:36:51 +0800 (CST) From: chenyuan_fl@163.com To: bpf@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Shuah Khan , Nuoqi Gui , Yuan Chen Subject: [PATCH 3/4] selftests/bpf: Test rhtab kptr recycle from NMI context Date: Mon, 24 Aug 2026 22:36:20 +0800 Message-ID: <20260824143621.2098856-4-chenyuan_fl@163.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260824143621.2098856-1-chenyuan_fl@163.com> References: <20260824143621.2098856-1-chenyuan_fl@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PSgvCgDn+osAV4xqIt5SMw--.54939S5 X-Coremail-Antispam: 1Uf129KBjvJXoW3CryDKw1xKFyrJF4UWFyUWrg_yoWDWryUpa 9Y9ry5tr4Fqwn8XrsYva18J3yF9ws5ZFW5CrZ2gryYvr1qgrn7JFyxKF4YkF9ayF4DXryS vwsxKFW3GFWUArDanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x07jWfO7UUUUU= X-CM-SenderInfo: xfkh05pxdqswro6rljoofrz/xtbC5QOTUWqMVwPdBAAA3P Content-Type: text/plain; charset="utf-8" From: Yuan Chen A perf_event program running in NMI context overwrites a rhtab element whose value holds a referenced task kptr. The old kptr must stay attached to the element (cancel semantics, matching hash maps); before the rhtab recycle fix the NMI update eagerly released it and the probe observed NULL. The test asserts the NMI program actually ran, so the probe result is meaningful. A second phase deletes and re-inserts the element 2000 times. The re-insertion may recycle the freed element, which still owns the kptr; before the fix the alloc path zeroed the inherited slot via check_and_init_map_value(), leaking the reference, and the probe never observed a non-NULL pointer. The test requires at least one recycle to inherit the kptr, and also verifies that plain (non-special) value bytes still round-trip through the recycled element on every iteration. The NMI phase is skipped when no hardware PMU is available. Signed-off-by: Yuan Chen --- .../selftests/bpf/prog_tests/rhtab_kptr.c | 146 ++++++++++++++++++ .../testing/selftests/bpf/progs/rhtab_kptr.c | 132 ++++++++++++++++ 2 files changed, 278 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/rhtab_kptr.c create mode 100644 tools/testing/selftests/bpf/progs/rhtab_kptr.c diff --git a/tools/testing/selftests/bpf/prog_tests/rhtab_kptr.c b/tools/te= sting/selftests/bpf/prog_tests/rhtab_kptr.c new file mode 100644 index 000000000000..13158d74cbc1 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/rhtab_kptr.c @@ -0,0 +1,146 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft Co., Ltd. */ + +#include +#include +#include +#include +#include "rhtab_kptr.skel.h" + +static __u64 read_counter(struct rhtab_kptr *skel, u32 idx) +{ + __u64 vals[libbpf_num_possible_cpus()]; + __u64 sum =3D 0; + int i, err; + + err =3D bpf_map_lookup_elem(bpf_map__fd(skel->maps.counters), &idx, vals); + if (!ASSERT_OK(err, "lookup_counter")) + return 0; + for (i =3D 0; i < libbpf_num_possible_cpus(); i++) + sum +=3D vals[i]; + return sum; +} + +void test_rhtab_kptr(void) +{ + struct perf_event_attr attr =3D { + .type =3D PERF_TYPE_HARDWARE, + .config =3D PERF_COUNT_HW_CPU_CYCLES, + .freq =3D 1, + .sample_freq =3D read_perf_max_sample_freq(), + .size =3D sizeof(struct perf_event_attr), + }; + LIBBPF_OPTS(bpf_test_run_opts, topts); + struct rhtab_kptr *skel; + __u32 key =3D 0; + __u64 zero =3D 0; + __u64 nonnull_before; + int pmu_fd, i, err; + + skel =3D rhtab_kptr__open_and_load(); + if (!ASSERT_OK_PTR(skel, "open_and_load")) + return; + + /* Create the element and stash a referenced task kptr in it. */ + if (!ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.rhtab), + &key, &zero, BPF_ANY), "create_elem")) + goto out; + if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.init_el= em), + &topts), "test_run_init") || + !ASSERT_EQ(topts.retval, 0, "init_ret")) + goto out; + + pmu_fd =3D syscall(__NR_perf_event_open, &attr, -1, 0, -1, 0); + if (pmu_fd >=3D 0) { + skel->links.nmi_update =3D bpf_program__attach_perf_event(skel->progs.nm= i_update, + pmu_fd); + if (!ASSERT_OK_PTR(skel->links.nmi_update, "attach_perf_event")) { + close(pmu_fd); + goto out; + } + + /* Let the NMI handler overwrite the element, and make sure it + * actually ran before probing (otherwise the probe would pass + * vacuously even on an unfixed kernel). + */ + for (i =3D 0; i < 20 && read_counter(skel, 1) =3D=3D 0; i++) + usleep(100000); + ASSERT_GT(read_counter(skel, 1), 0, "nmi_update_ran"); + + bpf_link__destroy(skel->links.nmi_update); + skel->links.nmi_update =3D NULL; + close(pmu_fd); + + /* + * The old kptr must still be attached to the element: the + * NMI update path only cancels NMI-safe fields, mirroring + * hash map semantics. Before the fix the kptr was released + * from the NMI context and the probe below would see NULL. + */ + topts.retval =3D 0; + if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.probe_= elem), + &topts), "test_run_probe") || + !ASSERT_EQ(topts.retval, 0, "probe_ret")) + goto out; + + ASSERT_EQ(read_counter(skel, 2), 1, "xchg_non_null"); + ASSERT_EQ(read_counter(skel, 3), 0, "xchg_null"); + } else { + test__skip(); + } + + /* + * Now exercise the delete/re-insert recycle path. The delete only + * cancels NMI-safe fields, so the freed element still owns the kptr. + * If the re-insertion recycles that element, the kptr must be + * inherited; zeroing it (as check_and_init_map_value() did before + * the fix) leaks the reference and probe_elem() observes NULL. + * Fresh memory handed out by the allocator is zeroed, so NULL probes + * are expected too; only require that the inherited kptr survives at + * least one recycle. + */ + nonnull_before =3D read_counter(skel, 2); + for (i =3D 0; i < 2000; i++) { + topts.retval =3D 0; + err =3D bpf_prog_test_run_opts(bpf_program__fd(skel->progs.init_elem), + &topts); + if (err || topts.retval) { + /* Element may be gone; recreate and retry once. */ + if (!ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.rhtab), + &key, &zero, BPF_ANY), + "recreate_elem")) + goto out; + topts.retval =3D 0; + err =3D bpf_prog_test_run_opts(bpf_program__fd(skel->progs.init_elem), + &topts); + } + if (!ASSERT_OK(err, "test_run_init_loop") || + !ASSERT_EQ(topts.retval, 0, "init_loop_ret")) + goto out; + + topts.retval =3D 0; + if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.del_el= em), + &topts), "test_run_del")) + goto out; + topts.retval =3D 0; + if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.upd_el= em), + &topts), "test_run_upd")) + goto out; + topts.retval =3D 0; + if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.probe_= elem), + &topts), "test_run_probe")) + goto out; + } + + /* + * Plain (non-special) value bytes must survive the recycle path: + * every probe must observe the magic value written by upd_elem() in + * the same iteration, regardless of whether the element memory was + * recycled or freshly allocated. + */ + ASSERT_EQ(read_counter(skel, 4), 2000, "recycle_magic_roundtrip"); + + ASSERT_GT(read_counter(skel, 2), nonnull_before, "recycle_xchg_non_null"); +out: + rhtab_kptr__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/rhtab_kptr.c b/tools/testing= /selftests/bpf/progs/rhtab_kptr.c new file mode 100644 index 000000000000..fd6bd63cb405 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/rhtab_kptr.c @@ -0,0 +1,132 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft Co., Ltd. */ + +/* + * Verify that the rhtab update/delete recycle paths do not eagerly destroy + * referenced kptrs. rhtab must match the hash map semantics introduced by + * commit a3a81d247651 ("bpf: Cancel special fields on map value recycle"): + * only NMI-safe fields (timer, workqueue, task_work) are cancelled on + * update/delete, while kptrs stay attached to the recycled element until = it + * is eventually freed. + * + * Two paths are exercised: + * 1. a perf_event (NMI) program overwrites an existing element; without = the + * fix the NMI update releases the old kptr and probe_elem() observes + * NULL; + * 2. the element is deleted and re-inserted; the re-insertion may recycle + * the freed element, and zeroing the inherited kptr slot (as + * check_and_init_map_value() did before the fix) would drop the + * reference without releasing it. probe_elem() must observe the + * inherited non-NULL pointer, and plain (non-special) value bytes must + * still round-trip through the recycled element. + */ +#include +#include + +char LICENSE[] SEC("license") =3D "GPL"; + +struct val_t { + struct task_struct __kptr * tsk; + __u32 magic; +}; + +struct { + __uint(type, BPF_MAP_TYPE_RHASH); + __uint(max_entries, 16); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, __u32); + __type(value, struct val_t); +} rhtab SEC(".maps"); + +struct { + __uint(type, BPF_MAP_TYPE_PERCPU_ARRAY); + __uint(max_entries, 5); + __type(key, __u32); + __type(value, __u64); +} counters SEC(".maps"); + +/* 0: init ok, 1: nmi update ok, 2: probe xchg non-NULL, 3: probe xchg NUL= L, + * 4: probe saw expected magic value + */ +static __always_inline void bump(u32 idx) +{ + u64 *v =3D bpf_map_lookup_elem(&counters, &idx); + + if (v) + (*v)++; +} + +extern struct task_struct *bpf_task_acquire(struct task_struct *p) __ksym; +extern void bpf_task_release(struct task_struct *p) __ksym; + +SEC("perf_event") +int nmi_update(struct bpf_perf_event_data *ctx) +{ + struct val_t val =3D {}; + u32 key =3D 0; + + if (bpf_map_update_elem(&rhtab, &key, &val, BPF_ANY) =3D=3D 0) + bump(1); + return 0; +} + +SEC("syscall") +int init_elem(void *ctx) +{ + struct val_t *val; + struct task_struct *task, *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&rhtab, &key); + if (!val) + return 1; + task =3D bpf_task_acquire(bpf_get_current_task_btf()); + if (!task) + return 2; + old =3D bpf_kptr_xchg(&val->tsk, task); + if (old) + bpf_task_release(old); + bump(0); + return 0; +} + +SEC("syscall") +int del_elem(void *ctx) +{ + u32 key =3D 0; + + bpf_map_delete_elem(&rhtab, &key); + return 0; +} + +SEC("syscall") +int upd_elem(void *ctx) +{ + struct val_t val =3D { .magic =3D 0x52484153 }; /* "RHAS" */ + u32 key =3D 0; + + bpf_map_update_elem(&rhtab, &key, &val, BPF_ANY); + return 0; +} + +SEC("syscall") +int probe_elem(void *ctx) +{ + struct val_t *val; + struct task_struct *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&rhtab, &key); + if (!val) + return 1; + old =3D bpf_kptr_xchg(&val->tsk, NULL); + if (old) { + bpf_task_release(old); + bump(2); + } else { + bump(3); + } + if (val->magic =3D=3D 0x52484153) + bump(4); + return 0; +} --=20 2.54.0 From nobody Mon Sep 28 08:07:25 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.3]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 092E74399E5; Mon, 24 Aug 2026 14:38:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.3 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582297; cv=none; b=RpW7HJgVtJgAlp0nuYDDKKMu3ZLoE/v1b/DH2JlQwJoBezNvrO8VQOaSpjOuTL+3hsbLw/4kNddjcxMohpZ3f0s5Xydkj4gaRuXeS/4tGGrcvSR9Ogasf7yScdUANPCP4xSIFNcW0YgiVo9a8WR2FEeacN+cbatUOVQgGOq55RI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787582297; c=relaxed/simple; bh=klWq5A4m8MvX34YD61Xji+aUUgui42kzirk9tGbrTDk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=j++QgYBAEzgX+FGSPhriomqX+9C9XRdD6nP5lapswCd0+IPY2GZZFcm4IGGRhAlmArmWdQrgRY8pIa4xX3Zl5Z+4jcXKzOPxI4IL8SNspUoyb2lp4n10SslrmuQ/jBHrPgtPlHW1PzUiBiBy5XmzmVKdgTa0ER7mPXPsDsp/gqM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=GaL9kVq1; arc=none smtp.client-ip=117.135.210.3 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="GaL9kVq1" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=0Q jPRk861kKD6Bxm9sd9ewmj9GbHTqQYhsF2ap3X5rw=; b=GaL9kVq1qwSNoQNCyN Lb5VV5+2qRf+dG84MJ06fL09AOEJCIWlXLfQ5XiHB2nvgOIj1IaY62/XpvV4xbLC pkGbVNLuyuTfUIjib+lSg9M6ai52li1a3k9bN6HNxdM7OxXPOBHNQfaYNvQOBUFK yBD70z4Eg6wDPpKEveLtjoH/U= Received: from nec8-i7 (unknown []) by gzsmtp2 (Coremail) with SMTP id PSgvCgDn+osAV4xqIt5SMw--.54939S6; Mon, 24 Aug 2026 22:36:52 +0800 (CST) From: chenyuan_fl@163.com To: bpf@vger.kernel.org Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Shuah Khan , Nuoqi Gui , Yuan Chen Subject: [PATCH 4/4] selftests/bpf: Test rhtab special-field combinations Date: Mon, 24 Aug 2026 22:36:21 +0800 Message-ID: <20260824143621.2098856-5-chenyuan_fl@163.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260824143621.2098856-1-chenyuan_fl@163.com> References: <20260824143621.2098856-1-chenyuan_fl@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PSgvCgDn+osAV4xqIt5SMw--.54939S6 X-Coremail-Antispam: 1Uf129KBjvAXoW3Cr4UAryDAFyUWryDXF1UZFb_yoW8Xry5Zo Z7Wr45Za18GryvgrWkWFykCF1rWayqgasrXF1Yv39rXa4IkFyUCr9rCrWxX3Wxu3W8trWU uas09w4fZr1fJF45n29KB7ZKAUJUUUU8529EdanIXcx71UUUUU7v73VFW2AGmfu7bjvjm3 AaLaJ3UbIYCTnIWIevJa73UjIFyTuYvjxU2J3vUUUUU X-CM-SenderInfo: xfkh05pxdqswro6rljoofrz/xtbDAQSUUmqMVwSuBwAA3f Content-Type: text/plain; charset="utf-8" From: Yuan Chen BPF_MAP_TYPE_RHASH allows spin locks, timers, workqueues, task_work, kptrs (referenced, untrusted, per-cpu) and refcounts in map values. The recycle fix only changes kptr slot handling, so verify each field combination end to end: * lock_kptr: bpf_spin_lock + referenced kptr + plain data in one value. BPF_F_LOCK syscall updates/lookups must work before and after many delete/re-insert recycle cycles, the referenced kptr must be inherited on recycled elements (zeroing it would leak the reference), and the plain bytes must round-trip every iteration. * timer: arm a bpf_timer and verify it fires, delete the element and verify the timer is cancelled, then re-insert (possibly recycling the freed element) and arm a fresh timer again. * kptr_untrusted: the untrusted kptr must survive the recycle like a referenced one. * kptr_percpu: the per-cpu kptr reference must survive the recycle (zeroing it would leak the reference). On the unfixed kernel the three kptr subtests fail at the recycle assertions while the lock and timer paths still pass, isolating the behavior change to kptr slots only. Signed-off-by: Yuan Chen --- .../selftests/bpf/prog_tests/rhtab_fields.c | 213 ++++++++++++ .../selftests/bpf/progs/rhtab_fields.c | 305 ++++++++++++++++++ 2 files changed, 518 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/rhtab_fields.c create mode 100644 tools/testing/selftests/bpf/progs/rhtab_fields.c diff --git a/tools/testing/selftests/bpf/prog_tests/rhtab_fields.c b/tools/= testing/selftests/bpf/prog_tests/rhtab_fields.c new file mode 100644 index 000000000000..29de05bcbd4b --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/rhtab_fields.c @@ -0,0 +1,213 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft Co., Ltd. */ + +#include +#include +#include "rhtab_fields.skel.h" + +#define RECYCLE_LOOPS 2000 +#define LK_MAGIC 0x52484142 + +/* Userspace view of the BPF value types (layouts must match the progs). */ +struct lock_kptr_val_user { + __u32 lock; + __u32 pad; + __u64 tsk; + __u32 magic; + __u32 pad2; +}; + +static __u64 read_counter(struct rhtab_fields *skel, u32 idx) +{ + __u64 vals[libbpf_num_possible_cpus()]; + __u64 sum =3D 0; + int i, err; + + err =3D bpf_map_lookup_elem(bpf_map__fd(skel->maps.counters), &idx, vals); + if (!ASSERT_OK(err, "lookup_counter")) + return 0; + for (i =3D 0; i < libbpf_num_possible_cpus(); i++) + sum +=3D vals[i]; + return sum; +} + +/* Returns the program retval; asserts the test_run itself succeeded. */ +static int run_prog(struct rhtab_fields *skel, const char *name) +{ + LIBBPF_OPTS(bpf_test_run_opts, topts); + struct bpf_program *prog; + int err; + + prog =3D bpf_object__find_program_by_name(skel->obj, name); + if (!ASSERT_OK_PTR(prog, name)) + return -1; + err =3D bpf_prog_test_run_opts(bpf_program__fd(prog), &topts); + if (!ASSERT_OK(err, name)) + return -1; + return topts.retval; +} + +static void recycle_loop(struct rhtab_fields *skel, int map_fd, + const char *init, const char *del, + const char *upd, const char *probe) +{ + u64 zero =3D 0; + u32 key =3D 0; + int i; + + for (i =3D 0; i < RECYCLE_LOOPS; i++) { + if (run_prog(skel, init) !=3D 0) { + /* Element may be gone; recreate and retry once. */ + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &zero, BPF_ANY), + "recreate_elem")) + return; + if (!ASSERT_OK(run_prog(skel, init), init)) + return; + } + if (!ASSERT_OK(run_prog(skel, del), del)) + return; + if (!ASSERT_OK(run_prog(skel, upd), upd)) + return; + if (!ASSERT_OK(run_prog(skel, probe), probe)) + return; + } +} + +static void subtest_lock_kptr(struct rhtab_fields *skel) +{ + struct lock_kptr_val_user val =3D {}; + struct lock_kptr_val_user out =3D {}; + u64 nonnull_before; + u32 key =3D 0; + int map_fd; + + map_fd =3D bpf_map__fd(skel->maps.lkmap); + + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &val, BPF_ANY), + "create_elem")) + return; + + /* Spin lock must be usable from the syscall path (BPF_F_LOCK). */ + val.magic =3D LK_MAGIC; + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &val, BPF_F_LOCK), + "locked_update")) + return; + if (!ASSERT_OK(bpf_map_lookup_elem_flags(map_fd, &key, &out, BPF_F_LOCK), + "locked_lookup")) + return; + ASSERT_EQ(out.magic, LK_MAGIC, "locked_lookup_magic"); + + /* + * Delete/re-insert recycle cycles: the referenced kptr must be + * inherited on recycled elements (zeroing it would leak the + * reference) and the plain magic bytes must round-trip every time. + */ + nonnull_before =3D read_counter(skel, 1); + recycle_loop(skel, map_fd, "lk_init", "lk_del", "lk_upd", "lk_probe"); + ASSERT_GT(read_counter(skel, 1), nonnull_before, "recycle_xchg_non_null"); + ASSERT_EQ(read_counter(skel, 3), RECYCLE_LOOPS, "recycle_magic_roundtrip"= ); + + /* The spin lock must still work after many recycles. */ + val.magic =3D LK_MAGIC + 1; + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &val, BPF_F_LOCK), + "post_recycle_locked_update")) + return; + memset(&out, 0, sizeof(out)); + if (!ASSERT_OK(bpf_map_lookup_elem_flags(map_fd, &key, &out, BPF_F_LOCK), + "post_recycle_locked_lookup")) + return; + ASSERT_EQ(out.magic, LK_MAGIC + 1, "post_recycle_locked_magic"); +} + +static void subtest_timer(struct rhtab_fields *skel) +{ + u64 zero =3D 0; + u32 key =3D 0; + int fired, map_fd; + + map_fd =3D bpf_map__fd(skel->maps.tmap); + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &zero, BPF_ANY), + "create_elem")) + return; + + if (!ASSERT_OK(run_prog(skel, "arm_timer"), "arm_timer_first")) + return; + usleep(300000); + if (!ASSERT_GT(skel->bss->timer_fired, 0, "timer_fired_first")) + return; + + /* Deleting the element must cancel the timer. */ + fired =3D skel->bss->timer_fired; + if (!ASSERT_OK(bpf_map_delete_elem(map_fd, &key), "delete_elem")) + return; + usleep(300000); + ASSERT_EQ(skel->bss->timer_fired, fired, "timer_cancelled_after_delete"); + + /* + * Re-insert (may recycle the freed element): the timer field must be + * re-initialized so a fresh timer can be armed again. + */ + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &zero, BPF_ANY), + "recreate_elem")) + return; + if (!ASSERT_OK(run_prog(skel, "arm_timer"), "arm_timer_second")) + return; + usleep(300000); + ASSERT_GT(skel->bss->timer_fired, fired, "timer_fired_second"); +} + +static void subtest_kptr_untrusted(struct rhtab_fields *skel) +{ + u64 nonnull_before; + u64 zero =3D 0; + u32 key =3D 0; + int map_fd; + + map_fd =3D bpf_map__fd(skel->maps.umap); + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &zero, BPF_ANY), + "create_elem")) + return; + + /* The untrusted kptr must survive the recycle like a referenced one. */ + nonnull_before =3D read_counter(skel, 5); + recycle_loop(skel, map_fd, "u_init", "u_del", "u_upd", "u_probe"); + ASSERT_GT(read_counter(skel, 5), nonnull_before, "recycle_unref_non_null"= ); +} + +static void subtest_kptr_percpu(struct rhtab_fields *skel) +{ + u64 nonnull_before; + u64 zero =3D 0; + u32 key =3D 0; + int map_fd; + + map_fd =3D bpf_map__fd(skel->maps.pcmap); + if (!ASSERT_OK(bpf_map_update_elem(map_fd, &key, &zero, BPF_ANY), + "create_elem")) + return; + + /* The per-cpu kptr reference must survive the recycle (no leak). */ + nonnull_before =3D read_counter(skel, 7); + recycle_loop(skel, map_fd, "pc_init", "pc_del", "pc_upd", "pc_probe"); + ASSERT_GT(read_counter(skel, 7), nonnull_before, "recycle_pcpu_non_null"); +} + +void test_rhtab_fields(void) +{ + struct rhtab_fields *skel; + + skel =3D rhtab_fields__open_and_load(); + if (!ASSERT_OK_PTR(skel, "open_and_load")) + return; + + if (test__start_subtest("lock_kptr")) + subtest_lock_kptr(skel); + if (test__start_subtest("timer")) + subtest_timer(skel); + if (test__start_subtest("kptr_untrusted")) + subtest_kptr_untrusted(skel); + if (test__start_subtest("kptr_percpu")) + subtest_kptr_percpu(skel); + + rhtab_fields__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/rhtab_fields.c b/tools/testi= ng/selftests/bpf/progs/rhtab_fields.c new file mode 100644 index 000000000000..85335f19f172 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/rhtab_fields.c @@ -0,0 +1,305 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft Co., Ltd. */ + +/* + * Combined special-field tests for BPF_MAP_TYPE_RHASH. Each map carries a + * different field combination and is exercised through delete/re-insert + * cycles so the bpf memory allocator recycles element memory: + * + * 1. lkmap: bpf_spin_lock + referenced kptr + plain data in one value. + * After every recycle the spin lock must still be usable (initialized= by + * the alloc path), the referenced kptr must be inherited instead of + * zeroed (zeroing would leak the reference), and the plain bytes must + * round-trip. + * 2. tmap: bpf_timer. The delete path must cancel the timer, and a recyc= led + * element must be able to arm a fresh timer again. + * 3. umap: untrusted (unreferenced) kptr. The inherited pointer must be + * preserved on recycle, matching hash map behavior. + * 4. pcmap: per-cpu kptr. Like the referenced kptr, the per-cpu reference + * must not be dropped on recycle. + */ + +#include +#include +#include "bpf_experimental.h" + +char LICENSE[] SEC("license") =3D "GPL"; + +struct lock_kptr_val { + struct bpf_spin_lock lock; + struct task_struct __kptr * tsk; + __u32 magic; +}; + +struct { + __uint(type, BPF_MAP_TYPE_RHASH); + __uint(max_entries, 16); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, __u32); + __type(value, struct lock_kptr_val); +} lkmap SEC(".maps"); + +struct timer_val { + struct bpf_timer timer; + __u64 data; +}; + +struct { + __uint(type, BPF_MAP_TYPE_RHASH); + __uint(max_entries, 16); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, __u32); + __type(value, struct timer_val); +} tmap SEC(".maps"); + +struct unref_val { + struct task_struct __kptr_untrusted * tsk; +}; + +struct { + __uint(type, BPF_MAP_TYPE_RHASH); + __uint(max_entries, 16); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, __u32); + __type(value, struct unref_val); +} umap SEC(".maps"); + +struct pcval { + __u64 v; +}; + +struct pcpu_val { + struct pcval __percpu_kptr * pc; +}; + +struct { + __uint(type, BPF_MAP_TYPE_RHASH); + __uint(max_entries, 16); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, __u32); + __type(value, struct pcpu_val); +} pcmap SEC(".maps"); + +struct { + __uint(type, BPF_MAP_TYPE_PERCPU_ARRAY); + __uint(max_entries, 9); + __type(key, __u32); + __type(value, __u64); +} counters SEC(".maps"); + +/* 0: lk init ok, 1: lk probe xchg non-NULL, 2: lk probe xchg NULL, + * 3: lk probe magic ok, 4: u init ok, 5: u probe ptr non-NULL, + * 6: pc init ok, 7: pc probe xchg non-NULL, 8: pc probe xchg NULL + */ +static __always_inline void bump(u32 idx) +{ + u64 *v =3D bpf_map_lookup_elem(&counters, &idx); + + if (v) + (*v)++; +} + +extern struct task_struct *bpf_task_acquire(struct task_struct *p) __ksym; +extern void bpf_task_release(struct task_struct *p) __ksym; + +int timer_fired; + +/* Map 1: spin lock + referenced kptr + plain data. */ + +SEC("syscall") +int lk_init(void *ctx) +{ + struct lock_kptr_val *val; + struct task_struct *task, *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&lkmap, &key); + if (!val) + return 1; + task =3D bpf_task_acquire(bpf_get_current_task_btf()); + if (!task) + return 2; + old =3D bpf_kptr_xchg(&val->tsk, task); + if (old) + bpf_task_release(old); + bump(0); + return 0; +} + +SEC("syscall") +int lk_del(void *ctx) +{ + u32 key =3D 0; + + bpf_map_delete_elem(&lkmap, &key); + return 0; +} + +SEC("syscall") +int lk_upd(void *ctx) +{ + struct lock_kptr_val val =3D { .magic =3D 0x52484142 }; + u32 key =3D 0; + + bpf_map_update_elem(&lkmap, &key, &val, BPF_ANY); + return 0; +} + +SEC("syscall") +int lk_probe(void *ctx) +{ + struct lock_kptr_val *val; + struct task_struct *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&lkmap, &key); + if (!val) + return 1; + old =3D bpf_kptr_xchg(&val->tsk, NULL); + if (old) { + bpf_task_release(old); + bump(1); + } else { + bump(2); + } + if (val->magic =3D=3D 0x52484142) + bump(3); + return 0; +} + +/* Map 2: bpf_timer. */ + +static int timer_cb(void *map, void *key, struct timer_val *value) +{ + timer_fired++; + return 0; +} + +SEC("syscall") +int arm_timer(void *ctx) +{ + struct timer_val *val; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&tmap, &key); + if (!val) + return 1; + /* 1 =3D=3D CLOCK_MONOTONIC */ + if (bpf_timer_init(&val->timer, &tmap, 1)) + return 2; + bpf_timer_set_callback(&val->timer, timer_cb); + if (bpf_timer_start(&val->timer, 50000, 0)) + return 3; + return 0; +} + +/* Map 3: untrusted kptr. */ + +SEC("syscall") +int u_init(void *ctx) +{ + struct unref_val *val; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&umap, &key); + if (!val) + return 1; + val->tsk =3D bpf_get_current_task_btf(); + bump(4); + return 0; +} + +SEC("syscall") +int u_del(void *ctx) +{ + u32 key =3D 0; + + bpf_map_delete_elem(&umap, &key); + return 0; +} + +SEC("syscall") +int u_upd(void *ctx) +{ + struct unref_val val =3D {}; + u32 key =3D 0; + + bpf_map_update_elem(&umap, &key, &val, BPF_ANY); + return 0; +} + +SEC("syscall") +int u_probe(void *ctx) +{ + struct unref_val *val; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&umap, &key); + if (!val) + return 1; + if (val->tsk) + bump(5); + val->tsk =3D NULL; + return 0; +} + +/* Map 4: per-cpu kptr. */ + +SEC("syscall") +int pc_init(void *ctx) +{ + struct pcpu_val *val; + struct pcval *p, *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&pcmap, &key); + if (!val) + return 1; + p =3D bpf_percpu_obj_new(struct pcval); + if (!p) + return 2; + old =3D bpf_kptr_xchg(&val->pc, p); + if (old) + bpf_percpu_obj_drop(old); + bump(6); + return 0; +} + +SEC("syscall") +int pc_del(void *ctx) +{ + u32 key =3D 0; + + bpf_map_delete_elem(&pcmap, &key); + return 0; +} + +SEC("syscall") +int pc_upd(void *ctx) +{ + struct pcpu_val val =3D {}; + u32 key =3D 0; + + bpf_map_update_elem(&pcmap, &key, &val, BPF_ANY); + return 0; +} + +SEC("syscall") +int pc_probe(void *ctx) +{ + struct pcpu_val *val; + struct pcval *old; + u32 key =3D 0; + + val =3D bpf_map_lookup_elem(&pcmap, &key); + if (!val) + return 1; + old =3D bpf_kptr_xchg(&val->pc, NULL); + if (old) { + bpf_percpu_obj_drop(old); + bump(7); + } else { + bump(8); + } + return 0; +} --=20 2.54.0