From nobody Thu Sep 24 21:47:03 2026 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 53F1745D93E for ; Sun, 20 Sep 2026 16:32:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789921975; cv=none; b=YOyAUCwMd+zLXrQam1+FC2QAqg0WJfiJetXLScsO8vMCnCI7rPf4ZRH47ZD/eqH0dshEsYci1+gkX8rK3iJp2gk22sum96S5ZnRYy1N5Fr1VWCW7sHjzwf1xtOX+Az7LxOgg9VFA5hlV1QCNE9xFIlBpmb5sZiP2a6zkQVPi498= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789921975; c=relaxed/simple; bh=ISMr9k/4ayfp7cIsuOtuFI3mL3SoCez2K155WDMQWyU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rGHvwgoxDbDH9ZF4l4ZgsBhxs8WiYRETiW/mV2hSikX9JCgOqolhBKcM8k5XjV6XSykn/TBfn3S0toYKf/PJuJ7//qRrMXdQvtXAQubwrmh7wj2OIhaCpczwVXDumNUwGW9kGFOWVH9m9MPb/5LH85UTg0AjhmkK860OOG5Rv5g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SB4EHrZy; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SB4EHrZy" Received: by mail-pj2-f12.google.com with SMTP id d9443c01a7336-2db1ca069c8so22610975ad.3 for ; Sun, 20 Sep 2026 09:32:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789921971; x=1790526771; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VxrYw5XYqojCDttriJ4CNjE+Rxifl6ZmijYb7xfnw6E=; b=SB4EHrZydpEUlLfwAIrxMb/vSHzMILHeky8mJrePffr9xV/4uoNxd0IcWezRAtp6aI mp2KK8qQ6ud98djfGG8C2EpEDn2Dq2ClxMhmEb/NdpaZvjFw1SDgbtx5b6NiaM1ma7Q1 YthFBhKJd9H7GrInfl4D6n0raeu89h7pOdGeQ2yhAi+yYxZ30f+QEaNwJH96gKBb9UeG izxgqDA+eu/plJfl9VPUI+GCCjJ4xK5Xtg9Gn+SrkELJSLDtrNYvXMRqwu8avXrpyFCi 2urqYrT2e6/hkyjk7nu3j9t3N7n8HiYVKD+AksT6BMz15d8N05wbQtrn/5CDNtx+UgiQ NMjg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789921971; x=1790526771; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VxrYw5XYqojCDttriJ4CNjE+Rxifl6ZmijYb7xfnw6E=; b=mI58bGJFu9dEedFGrju6sU3I8jub6fotK5dC/6tDv1fbos3FZU71/c7xPCFgcVvU9c 8swYnizWHLAcmJjrIFAxfKo6vMk+1mGZjCkuQ9alkegPgLdSljXePNVTyBwahooniYAC RTlv3Wh6I4R6Z4mbb+yDzxsjav93NnuZefPjLWUXaKBmr8mc7W/xZmYdKmzPZralN5Sr nPBM5LuBm0zrGkF7r4FwnOtuOr2sqCNI8HlJ663oKEnQ5x9jzBFHN/aFC7vHbaNJ9bIC erOBu0GqH2ElTC49yHrzVLaCK7j83DyAlESi0TASUJIZlwDCqnZYOpovg5eDroNzQfCS ZIOg== X-Gm-Message-State: AFuF++ld+FFvGs8QJLCXYGjOACjJoBStlecCOf+kAbGKS8pPXSM4MxMk q0UCRBDr5PeXGBBhKsnbO2g21nXx3KsUxQ2RpD/f+Y8Ss9tI7eGP6z0D X-Gm-Gg: AYBFou0WLc55KdAyK1swO9IGFok6Y4fc1VxnYVfT73CuaKLQwRT5ZibFXVJUr3pOkV+ hzUsgpOW228N3tpsj+4sr84JeTTARfkd7SSylE76d7bS00jKolgvnORqvqleMB87AdPiH2VSYld HXJJznjV6Fvi5yjFyF4guZGcdcE3dnp4XKueZmKIHSbcj4cGPaGfVJCapHUnM2vhS2yBeGaeVB9 ou6Zxxc5xnIFz7WeLhsjd8/eWFhUaBJg6sA41zcgH54rlGOEg2AQZjJ14qmnrNuz7rtfW7mXvmj dfAmVvoAy7W566n7fEfuh03ep2AY0VpZTSu63tuz1Iyd5lSMyX+znpaBp5UTgo6GOm0xvvxqSZA 8yV39spa2wM55JSth+UdpwU+WqbxFZpf9IZH+HtLFYWqEszxecRvRk5TJOxjFNua3QF87eo0nWi I6VACjDuHMKJe9DdzCnbLarQGuwh9BjBDCDd5vw8GA9u/lLef78or2J0sCCffFM/ulbpuitDtGx 2sX8T/lFEvZRKvk3+yqHU6cruec02V+ X-Received: by 2002:a17:903:380d:b0:2dd:c100:b2ce with SMTP id d9443c01a7336-2ddc100b384mr67923315ad.57.1789921970736; Sun, 20 Sep 2026 09:32:50 -0700 (PDT) Received: from 192.168.50.3 ([198.176.50.208]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17e17e0sm21784355ad.70.2026.09.20.09.32.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 20 Sep 2026 09:32:49 -0700 (PDT) From: Weiming Shi To: Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Shuah Khan Cc: linux-kernel@vger.kernel.org, bpf@vger.kernel.org, netdev@vger.kernel.org, linux-kselftest@vger.kernel.org, =?UTF-8?q?Toke=20H=C3=B8iland-J=C3=B8rgensen?= , Peter Oskolkov , Xiang Mei , stable@vger.kernel.org Subject: [PATCH v3 1/3] bpf: propagate cb_access from freplace programs Date: Mon, 21 Sep 2026 00:32:09 +0800 Message-ID: <20260920163211.795547-2-bestswngs@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260920163211.795547-1-bestswngs@gmail.com> References: <20260920163211.795547-1-bestswngs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" bpf_prog_run_save_cb() and bpf_prog_run_clear_cb() use the target program's cb_access flag to decide whether skb->cb must be saved or cleared. An extension program can introduce ctx->cb[] access behind a target which does not access it itself, leaving protocol-owned control block contents visible to the replacement and preventing the wrapper from restoring them. Make cb_access independently addressable and propagate it to the target before activating a replacement. Wait for wrappers which observed the old value, and snapshot the flag once per invocation so save and restore decisions remain paired. Cc: stable@vger.kernel.org Fixes: be8704ff07d2 ("bpf: Introduce dynamic program extensions") Suggested-by: Daniel Borkmann Assisted-by: LLM Signed-off-by: Weiming Shi --- include/linux/bpf.h | 2 +- include/linux/filter.h | 7 ++++--- kernel/bpf/syscall.c | 6 ++++++ 3 files changed, 11 insertions(+), 4 deletions(-) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index e57af902560c3..606e7cf397af2 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -1864,11 +1864,11 @@ struct bpf_prog_aux { =20 struct bpf_prog { u16 pages; /* Number of allocated pages */ + bool cb_access; /* Is control block accessed? */ u32 jited:1, /* Is our filter JIT'ed? */ jit_requested:1,/* archs need to JIT the prog */ jit_required:1, /* program strictly requires JIT compiler */ gpl_compatible:1, /* Is filter GPL compatible? */ - cb_access:1, /* Is control block accessed? */ dst_needed:1, /* Do we need dst entry? */ blinding_requested:1, /* needs constant blinding */ blinded:1, /* Was blinded */ diff --git a/include/linux/filter.h b/include/linux/filter.h index 39decde7fc730..788c2d625db4a 100644 --- a/include/linux/filter.h +++ b/include/linux/filter.h @@ -1047,16 +1047,17 @@ static inline u32 __bpf_prog_run_save_cb(const stru= ct bpf_prog *prog, const struct sk_buff *skb =3D ctx; u8 *cb_data =3D bpf_skb_cb(skb); u8 cb_saved[BPF_SKB_CB_LEN]; + bool cb_access =3D READ_ONCE(prog->cb_access); u32 res; =20 - if (unlikely(prog->cb_access)) { + if (unlikely(cb_access)) { memcpy(cb_saved, cb_data, sizeof(cb_saved)); memset(cb_data, 0, sizeof(cb_saved)); } =20 res =3D bpf_prog_run(prog, skb); =20 - if (unlikely(prog->cb_access)) + if (unlikely(cb_access)) memcpy(cb_data, cb_saved, sizeof(cb_saved)); =20 return res; @@ -1079,7 +1080,7 @@ static inline u32 bpf_prog_run_clear_cb(const struct = bpf_prog *prog, u8 *cb_data =3D bpf_skb_cb(skb); u32 res; =20 - if (unlikely(prog->cb_access)) + if (unlikely(READ_ONCE(prog->cb_access))) memset(cb_data, 0, BPF_SKB_CB_LEN); =20 res =3D bpf_prog_run_pin_on_cpu(prog, skb); diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index c7bc9ba9b331f..43a29e47c8d01 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -3812,6 +3812,12 @@ static int bpf_tracing_prog_attach(struct bpf_prog *= prog, if (err) goto out_unlock; =20 + if (prog->type =3D=3D BPF_PROG_TYPE_EXT && READ_ONCE(prog->cb_access)) { + WRITE_ONCE(tgt_prog->cb_access, true); + /* Drain runs that observed cb_access=3Dfalse before enabling freplace. = */ + synchronize_rcu(); + } + err =3D bpf_trampoline_link_prog(&link->link.node, tr, tgt_prog); if (err) { bpf_link_cleanup(&link_primer); --=20 2.55.0 From nobody Thu Sep 24 21:47:03 2026 Received: from mail-pj2-f21.google.com (mail-pj2-f21.google.com [74.125.227.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C9BC546AA66 for ; Sun, 20 Sep 2026 16:33:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789921993; cv=none; b=F9DJkfvgxehMpkGTFXJjniDfLucqQrmCPpbNtXVN0ZglqdqPtdCDPSbHbGxj2E/nUU8l4iDMph6eDMR98ylu3xtDyWQrHugMbDrz+CgWMc6CeGu8cg3rT0KnTzHdqfcgzYVjGlRWMDPlPOO8kDy/mBvsIYPPNcl+GjEBB0CkPtI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789921993; c=relaxed/simple; bh=9b28pbK9azbLidjba42geuga0SJ1WJgd6Ehnda9bR/E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=plFmA3SiP2GyLou5ev3na8KGf09IU1yrv+TdO4pt/EUrmA8xpc65Io3Mz2jss6zCQSNd0n25s0RBLtbU/R9yu5PnpjlUgGIL6pVvRUx6ZPZpe7cabd5SwfT1MKeChec9TNaEEz+xmuWd+0MLuxl90vIO2ZAByPRWQDJEho+4L3I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=OGsfOe/a; arc=none smtp.client-ip=74.125.227.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="OGsfOe/a" Received: by mail-pj2-f21.google.com with SMTP id d9443c01a7336-2d747f0b25dso24657915ad.2 for ; Sun, 20 Sep 2026 09:33:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789921988; x=1790526788; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=x1bOr4cY3cVT9fuwETlIQTXCDoyIrnhNiHPjwyjuD7g=; b=OGsfOe/aK26NYgpKtUmW0yGyX65PH7MsC1nC6TRTGKppgSDd1jkdJs9OoH3Ti1gXMf 9Q6r7FCKC0zxtidIIyCMQ8dcJ/tdU8hiPhIu8gV6FST8XKKP5c/7LLbHk3L6ECO/KrSP Ss6Off+S8/r/jN5fhTAOgh3WRmvku53FC4AhigE8EMODf+1u+U1pzsfTBv8J1FCt6Lg4 T0twFn4ATG0YjrjkQRFLoUOeVaDfwOBn29HZGxSeqLIdAIZ0Tf+RoZh1bAP9JyUAq2DS 0yQQpx8uxDGVkf7acYD1mxoHH4GzXJnlLnVQUejTgcy9jq4OSo7Hl45u4CLYRDwXi0RF RJVA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789921988; x=1790526788; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=x1bOr4cY3cVT9fuwETlIQTXCDoyIrnhNiHPjwyjuD7g=; b=wJd03c9+uF1uwOGQcynOjoDSuVrpTdVxEL64kf5qtu2+yKYKeja1dgk8lL/UZttp1O WUiRTV11nFSmr7zuD5xw9mbrw3KzRhEhkfjicPzLbmYbtauNumnBrSA//Xe+8ugQVGtT Q7yUDaBtoKyzcDQe53lgnKN3c8XXwGTctcm1i0A2EwbY+mEZhPEsaN7nQkURlfKQq5q8 rQdphbqmee4PdrpNE2IAnx5IGJRj/dIJQCV+2tTJB0RmUEAIPFn1TFCfFUHJSH7YZyS4 TRpl+J60+wnMui5AFB/r3JGwzRUyKfqVRYnTqXoYJ1Yj5NLjveXjxIabD1W1xMFbsZxp ezwQ== X-Gm-Message-State: AFuF++kZ07l0w4zv2OMkOE8i1FkYF5L0XLOcvjlSTUbI3SiAgrr1UF9c IdNF06vYykEZuXgAEv1xinvwbN28ZPm/plmrvgohUTngx7Jfx83DAewJ X-Gm-Gg: AYBFou20iQxhiaCuO6+J21EhBY/cl9cGJMoNwmGTqQTImsuNICGykFXA82s2dmggBQP CYNyp1++T/vCOixexPHsSMt0qIoP1NcYfQCr0qcqT03gXQTHhlG4JhXG97nlaltEtCTsBX1l37j JvVMhCOdEFef8cAAajJ4OORZ9pDrXJIEGUiLc5CEWC3puZBi0d0ye0K61KQp47tsJA0aQ98V1U7 ussL9BYjCIr223Z9mdIDrSBo051GieH+2G8nklNfQrAPq2o/WtwOuLuSSudN/gTd20T4DoAcdvP DZBIQXDJhJgdSw6tt8vEtCMo5Gkyrx0Lbi2eIndqmvH8dMub4vkmzIiJloya8gui+K/Cv6qvyUk jZNcOj7hW1YrDTs/doEBqQsibb8TgjYfsZxWLEHny8rVZeaaLNaOFvlbXpys7HB71YhlK+PGGVs eT1dKq/YSffF51OSaVDK6TNUN77xEtThGkVzZ/H7b33M9qEIcK13cmFnPEsqFPFjRC+gtUcqQkY 2rTQw1PImDcyv0kZDLzfppWCiQHN6LQ X-Received: by 2002:a17:902:fd8c:b0:2dd:ad74:ac22 with SMTP id d9443c01a7336-2ddb1bd6e1bmr144054675ad.29.1789921987745; Sun, 20 Sep 2026 09:33:07 -0700 (PDT) Received: from 192.168.50.3 ([198.176.50.208]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17e17e0sm21784355ad.70.2026.09.20.09.32.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 20 Sep 2026 09:33:07 -0700 (PDT) From: Weiming Shi To: Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Shuah Khan Cc: linux-kernel@vger.kernel.org, bpf@vger.kernel.org, netdev@vger.kernel.org, linux-kselftest@vger.kernel.org, =?UTF-8?q?Toke=20H=C3=B8iland-J=C3=B8rgensen?= , Peter Oskolkov , Xiang Mei , stable@vger.kernel.org Subject: [PATCH v3 2/3] bpf: clear stale IPv4 options after LWT encapsulation Date: Mon, 21 Sep 2026 00:32:10 +0800 Message-ID: <20260920163211.795547-3-bestswngs@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260920163211.795547-1-bestswngs@gmail.com> References: <20260920163211.795547-1-bestswngs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" bpf_lwt_push_ip_encap() rebases the network header after prepending an IP header, but leaves IPCB(skb)->opt describing the inner IPv4 header. An ingress LWT route can consequently make an ICMP error use stale option offsets when constructing its reply. Mark each completed LWT IP encapsulation in the run's BPF network context and reset the protocol control block after bpf_prog_run_save_cb() restores it. This also covers an skb which was already marked encapsulated before entering LWT. A failed SEG6 encapsulation does not set the marker, so it no longer causes valid IPv6 metadata to be cleared. Select the reset layout from the protocol callback which consumes the packet: the original family for BPF_OK and unsupported redirects, or the new family for supported reroute and redirect paths. Preserve the ingress interface and L3-slave state from the restored original control block and initialize the IPv6 next-header offset when needed. Save and restore the marker around nested LWT runs. Cc: stable@vger.kernel.org Fixes: 52f278774e79 ("bpf: implement BPF_LWT_ENCAP_IP mode in bpf_lwt_push_= encap") Reported-by: Xiang Mei Link: https://lore.kernel.org/bpf/20260915170147.3943392-2-bestswngs@gmail.= com/ Suggested-by: Daniel Borkmann Link: https://lore.kernel.org/bpf/48990076-414c-4196-99b9-86fce41b8054@ioge= arbox.net/ Assisted-by: LLM Signed-off-by: Weiming Shi --- include/linux/filter.h | 1 + net/core/lwt_bpf.c | 47 ++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 48 insertions(+) diff --git a/include/linux/filter.h b/include/linux/filter.h index 788c2d625db4a..c8ca526d0661d 100644 --- a/include/linux/filter.h +++ b/include/linux/filter.h @@ -848,6 +848,7 @@ struct bpf_nh_params { #define BPF_RI_F_CPU_MAP_INIT BIT(2) #define BPF_RI_F_DEV_MAP_INIT BIT(3) #define BPF_RI_F_XSK_MAP_INIT BIT(4) +#define BPF_RI_F_LWT_IP_ENCAP BIT(5) =20 struct bpf_redirect_info { u64 tgt_index; diff --git a/net/core/lwt_bpf.c b/net/core/lwt_bpf.c index da49364ec63de..d585484a3a766 100644 --- a/net/core/lwt_bpf.c +++ b/net/core/lwt_bpf.c @@ -36,10 +36,44 @@ static inline struct bpf_lwt *bpf_lwt_lwtunnel(struct l= wtunnel_state *lwt) #define NO_REDIRECT false #define CAN_REDIRECT true =20 +static void bpf_lwt_reset_cb(struct sk_buff *skb, __be16 orig_proto, + bool use_new_proto) +{ + __be16 cb_proto =3D use_new_proto ? skb->protocol : orig_proto; + int iif =3D skb->skb_iif; + bool l3slave =3D false; + + /* VRF may have replaced skb_iif with the master device index. */ + if (orig_proto =3D=3D htons(ETH_P_IP)) { + iif =3D IPCB(skb)->iif; + l3slave =3D ipv4_l3mdev_skb(IPCB(skb)->flags); + } else if (orig_proto =3D=3D htons(ETH_P_IPV6)) { + iif =3D IP6CB(skb)->iif; + l3slave =3D ipv6_l3mdev_skb(IP6CB(skb)->flags); + } + + if (cb_proto =3D=3D htons(ETH_P_IP)) { + memset(IPCB(skb), 0, sizeof(*IPCB(skb))); + IPCB(skb)->iif =3D iif; + if (l3slave) + IPCB(skb)->flags |=3D IPSKB_L3SLAVE; + } else if (cb_proto =3D=3D htons(ETH_P_IPV6)) { + memset(IP6CB(skb), 0, sizeof(*IP6CB(skb))); + IP6CB(skb)->iif =3D iif; + IP6CB(skb)->nhoff =3D offsetof(struct ipv6hdr, nexthdr); + if (l3slave) + IP6CB(skb)->flags |=3D IP6SKB_L3SLAVE; + } +} + static int run_lwt_bpf(struct sk_buff *skb, struct bpf_lwt_prog *lwt, struct dst_entry *dst, bool can_redirect) { struct bpf_net_context __bpf_net_ctx, *bpf_net_ctx; + struct bpf_redirect_info *ri; + bool lwt_ip_encap, nested_lwt_ip_encap; + __be16 orig_proto =3D skb->protocol; + bool use_new_proto; int ret; =20 /* Disabling BH is needed to protect per-CPU bpf_redirect_info between @@ -47,8 +81,20 @@ static int run_lwt_bpf(struct sk_buff *skb, struct bpf_l= wt_prog *lwt, */ local_bh_disable(); bpf_net_ctx =3D bpf_net_ctx_set(&__bpf_net_ctx); + ri =3D bpf_net_ctx_get_ri(); + nested_lwt_ip_encap =3D ri->kern_flags & BPF_RI_F_LWT_IP_ENCAP; + ri->kern_flags &=3D ~BPF_RI_F_LWT_IP_ENCAP; bpf_compute_data_pointers(skb); ret =3D bpf_prog_run_save_cb(lwt->prog, skb); + lwt_ip_encap =3D ri->kern_flags & BPF_RI_F_LWT_IP_ENCAP; + ri->kern_flags &=3D ~BPF_RI_F_LWT_IP_ENCAP; + if (nested_lwt_ip_encap) + ri->kern_flags |=3D BPF_RI_F_LWT_IP_ENCAP; + use_new_proto =3D (ret =3D=3D BPF_LWT_REROUTE && + lwt->prog->type !=3D BPF_PROG_TYPE_LWT_OUT) || + (ret =3D=3D BPF_REDIRECT && can_redirect); + if (lwt_ip_encap) + bpf_lwt_reset_cb(skb, orig_proto, use_new_proto); =20 switch (ret) { case BPF_OK: @@ -668,6 +714,7 @@ int bpf_lwt_push_ip_encap(struct sk_buff *skb, void *hd= r, u32 len, bool ingress) } else { skb->protocol =3D htons(ETH_P_IPV6); } + bpf_net_ctx_get_ri()->kern_flags |=3D BPF_RI_F_LWT_IP_ENCAP; =20 if (skb_is_gso(skb)) return handle_gso_encap(skb, ipv4, len); --=20 2.55.0 From nobody Thu Sep 24 21:47:03 2026 Received: from mail-pj2-f13.google.com (mail-pj2-f13.google.com [74.125.227.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 404A146C843 for ; Sun, 20 Sep 2026 16:33:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789922009; cv=none; b=JjaXQ9OFpVt+SiqqMaVOMb4FX5MST4bxjPAq/ejetlYjFq5gqZ31lwle6iItih9s7J1/dqYn3cWegm94fSz0xpmt/ezIHzOB5V8uyMsKGkbVNmHPNWLwJ+ZiQbqUPJtukyWZWkNy3zsObJt0t98kMlO3v68Pm6cvgsYUPIAWhro= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789922009; c=relaxed/simple; bh=M6eq7518NcqGSls1Q2O5LYBYseZpRAZf6EGOdd7HnfY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NCpaUJv5078Ccl2r+FVbV1DYF4u827tBbLc/HnrLZQPzv+b0whotlDIODLJytNbLPrdqiYPOYU5l9U4ysm3KD+F06siSSdEiUQINB3o3Aqj56UFg9+hkej3Jca+Q4+muBU2BJXXEerpVYpgzbv6BW5ZpBWuGYk6NYax2bGJae40= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=A4VK5fpl; arc=none smtp.client-ip=74.125.227.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="A4VK5fpl" Received: by mail-pj2-f13.google.com with SMTP id d9443c01a7336-2d91ede8035so26744085ad.3 for ; Sun, 20 Sep 2026 09:33:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789922003; x=1790526803; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=5MGj7x4NYijNiXI1Yu7fch5YTGUoaVlAca3kPtJVJv4=; b=A4VK5fpljIkREq65ZNv6coL7o7F3QgolNvH7tGDFQ7Ty4ghfuOAJl97+Q3emj1LI5+ iC39b/AbKF0bJOvIwOPvu6kOgtLtR8aluIbYBRRJFddhGPzCa1fw2fUDIyQst9pntn4k F+Nyq626xgJMSa2jFarI7SNBBawAQBS6GAq1jPBS6RHsxucA90e5qgGVncw+a6xFXqKJ L5MVKSKzJxzDkzAQM0tVLtX22vV3HZjRI8POq/t5K+NL3eHhx4bDBefKhR0Tde0CYRPC Fb5mrI4fHyBlVqSqj9KgRgTGA19e9+Qf1bCJjOW4fUC4FofBzF+tVMe5dq9nL6z4pvHh m/0w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789922003; x=1790526803; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=5MGj7x4NYijNiXI1Yu7fch5YTGUoaVlAca3kPtJVJv4=; b=REjCmkOQjLoQ2kdPRZeViApP9XlnptYyGSMQ0x20gCxxz25Wl8YX0iRT8JxtjUlVC0 9KD2DRMgKgUyg/0VzjSbvwLDb+P/07qHJZ3l+yvJjrGA7ulGCwMSHvN/LQPiUvA0Bp0z nthR3FvB9svfrW6ixMAje2dzPm9xrYn7MGGDpfQlMMpv3TgIuZbbX+CRKeuDXJ/nKBh/ gOrbgTmkh0BvZFHagAl9g8tolq9b/0SuZdIvwn7z3vHfFKtsqxhdPCy28HtrHhN/2My5 vanfGc2G95Xq4Wvy0d8F8YD3beoVAOvst0bnNNR1j+eW/Yk+AWjDSedS9LfjTi/iYc1A ysyA== X-Gm-Message-State: AFuF++m49SI4sCO4DqDn4APJo2GdrqV9MJ3HHEKzVuKHdlDWOET+ig2W wBbmJ/h/OBanRIpGTqNho8MLq+R9txc1bbRMoadlJYzIjiKOfsoOVqWE X-Gm-Gg: AYBFou1JQTyhR3J9JHiWr2tNCy1W96b/NNTHzCYk2iv0u2WAkYjuvAuMh7Z8ALI2R8J Uh8DfITUJJKICqEfzVBBBCcSgNf2CY9EOw8D9qLd5OGjeL7lcAGCoRK3HsHvyBype9QaWO1mMVd qmKwoo76X/uFnSQE30nlP6PWV0Blt9+FMIiVi+VEPRTPiAFBQY1gRJePUFjyiU8eIrDZZVfNwob vIb4Zcnh1vI+gGCKdfNVOzWeA4zhzB9O4Op+GU3S+sEFd4jvy90Ekapvxlc/aI8wlgFYRlhvAiC 35w9xPkV7zXl7TyB5A/P3B6xxWhkjaNM5LXyWw4Zr0JlgGe3pAq3S9FD2Pe/2OkvNa03iYwjb27 Lj2oSgdJRFu/+FD1KTYQUbh0CQ0dYqpOqDv3+6J+jBpgLOIFWqgzJHhzE3FuYTe+wG6nbPfv9Jl ACUDTW4sQ2im96yrdZm/dyCrgpqPjGCxQDt4bkdXsDUVerl+xSEPLBHcR+AEVLhyyfVxgY8oEEk nN0eHZh8I7MaMHKd/4DZpDrikCl+Mpfg2qZgkCGYnA= X-Received: by 2002:a17:903:2312:b0:2cc:6018:f030 with SMTP id d9443c01a7336-2ddb1b899famr139374895ad.14.1789922002512; Sun, 20 Sep 2026 09:33:22 -0700 (PDT) Received: from 192.168.50.3 ([198.176.50.208]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17e17e0sm21784355ad.70.2026.09.20.09.33.08 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 20 Sep 2026 09:33:21 -0700 (PDT) From: Weiming Shi To: Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Shuah Khan Cc: linux-kernel@vger.kernel.org, bpf@vger.kernel.org, netdev@vger.kernel.org, linux-kselftest@vger.kernel.org, =?UTF-8?q?Toke=20H=C3=B8iland-J=C3=B8rgensen?= , Peter Oskolkov , Xiang Mei Subject: [PATCH v3 3/3] selftests/bpf: cover stale CB after LWT IP encapsulation Date: Mon, 21 Sep 2026 00:32:11 +0800 Message-ID: <20260920163211.795547-4-bestswngs@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260920163211.795547-1-bestswngs@gmail.com> References: <20260920163211.795547-1-bestswngs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add regression coverage for stale protocol control-block contents after bpf_lwt_push_ip_encap(). Inject an IPv4 packet carrying a Record Route option through an ingress LWT program and verify that the resulting ICMP Time Exceeded packet has no options copied from the inner header. Exercise both direct ctx->cb[] access and access introduced by a freplace program. The latter also verifies that cb_access is propagated to the target so the replacement observes cleared BPF scratch space. Run both variants on VRF ingress as well, checking that ICMP source selection still uses the original ingress slave rather than the VRF master after the protocol control block is reset. Add an ingress TC pre-encapsulation case which enters LWT with skb->encapsulation already set and a compiled outer Record Route option. This covers the true-to-true transition that an encapsulation-bit edge check misses. Signed-off-by: Weiming Shi --- .../selftests/bpf/prog_tests/lwt_ip_encap.c | 288 ++++++++++++++++++ .../bpf/progs/lwt_ip_encap_stale_cb.c | 100 ++++++ .../progs/lwt_ip_encap_stale_cb_freplace.c | 32 ++ 3 files changed, 420 insertions(+) create mode 100644 tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb= .c create mode 100644 tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb= _freplace.c diff --git a/tools/testing/selftests/bpf/prog_tests/lwt_ip_encap.c b/tools/= testing/selftests/bpf/prog_tests/lwt_ip_encap.c index 39e8a3b8b6afb..0cc36b4d75b83 100644 --- a/tools/testing/selftests/bpf/prog_tests/lwt_ip_encap.c +++ b/tools/testing/selftests/bpf/prog_tests/lwt_ip_encap.c @@ -1,8 +1,17 @@ // SPDX-License-Identifier: GPL-2.0-only +#include +#include +#include +#include +#include #include +#include +#include =20 #include "network_helpers.h" #include "test_progs.h" +#include "lwt_ip_encap_stale_cb.skel.h" +#include "lwt_ip_encap_stale_cb_freplace.skel.h" #include "test_lwt_ip_encap.skel.h" =20 #define BPF_FILE "test_lwt_ip_encap.bpf.o" @@ -686,3 +695,282 @@ void test_lwt_ip_encap_vxlan_ipv6(void) { lwt_ip_encap_vxlan(IPV6_ENCAP); } + +#define STALE_CB_NETNS "lwt-ip-encap-stale-cb" +#define STALE_CB_DST "10.9.9.0/24" +#define STALE_CB_PIN_FMT "/sys/fs/bpf/lwt_ip_encap_stale_cb_%d" +#define STALE_CB_PKT_LEN 64 + +static __u16 stale_cb_csum(const void *data, size_t len) +{ + const __u16 *word =3D data; + __u32 sum =3D 0; + + while (len > 1) { + sum +=3D *word++; + len -=3D sizeof(*word); + } + if (len) + sum +=3D *(const __u8 *)word; + while (sum >> 16) + sum =3D (sum & 0xffff) + (sum >> 16); + + return ~sum; +} + +static int stale_cb_get_mac(const char *ifname, __u8 mac[ETH_ALEN]) +{ + struct ifreq ifr =3D {}; + int fd; + + fd =3D socket(AF_INET, SOCK_DGRAM, 0); + if (fd < 0) + return -errno; + strncpy(ifr.ifr_name, ifname, sizeof(ifr.ifr_name) - 1); + if (ioctl(fd, SIOCGIFHWADDR, &ifr)) { + int err =3D -errno; + + close(fd); + return err; + } + memcpy(mac, ifr.ifr_hwaddr.sa_data, ETH_ALEN); + close(fd); + return 0; +} + +static int stale_cb_open_packet_socket(int ifindex) +{ + struct sockaddr_ll addr =3D { + .sll_family =3D AF_PACKET, + .sll_protocol =3D htons(ETH_P_ALL), + .sll_ifindex =3D ifindex, + }; + struct timeval timeout =3D { .tv_sec =3D 2 }; + int fd; + + fd =3D socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL)); + if (fd < 0) + return -errno; + if (bind(fd, (struct sockaddr *)&addr, sizeof(addr)) || + setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &timeout, sizeof(timeout))) { + int err =3D -errno; + + close(fd); + return err; + } + + return fd; +} + +static int stale_cb_send_packet(int fd, const __u8 src_mac[ETH_ALEN], + const __u8 dst_mac[ETH_ALEN]) +{ + __u8 frame[ETH_HLEN + STALE_CB_PKT_LEN] =3D {}; + struct ethhdr *eth =3D (struct ethhdr *)frame; + struct iphdr *iph =3D (struct iphdr *)(frame + ETH_HLEN); + __u8 *opt =3D (__u8 *)(iph + 1); + + memcpy(eth->h_source, src_mac, ETH_ALEN); + memcpy(eth->h_dest, dst_mac, ETH_ALEN); + eth->h_proto =3D htons(ETH_P_IP); + + iph->version =3D 4; + iph->ihl =3D 7; + iph->tos =3D 8; + iph->tot_len =3D htons(STALE_CB_PKT_LEN); + iph->id =3D htons(0x1234); + iph->ttl =3D 64; + iph->protocol =3D IPPROTO_UDP; + iph->saddr =3D inet_addr("10.0.0.2"); + iph->daddr =3D inet_addr("10.9.9.9"); + opt[0] =3D IPOPT_RR; + opt[1] =3D 8; + opt[2] =3D 4; + iph->check =3D stale_cb_csum(iph, iph->ihl * 4); + + memset(frame + ETH_HLEN + iph->ihl * 4, 0x41, + STALE_CB_PKT_LEN - iph->ihl * 4); + if (send(fd, frame, sizeof(frame), 0) !=3D sizeof(frame)) + return -errno; + + return 0; +} + +static int stale_cb_icmp_ihl(int fd, __u32 *saddr) +{ + __u8 packet[512]; + ssize_t len; + + while ((len =3D recv(fd, packet, sizeof(packet), 0)) >=3D 0) { + const struct ethhdr *eth =3D (const struct ethhdr *)packet; + const struct iphdr *iph; + const struct icmphdr *icmph; + size_t ip_len; + + if (len < ETH_HLEN + sizeof(*iph) || + eth->h_proto !=3D htons(ETH_P_IP)) + continue; + iph =3D (const struct iphdr *)(packet + ETH_HLEN); + ip_len =3D iph->ihl * 4; + if (iph->ihl < 5 || len < ETH_HLEN + ip_len + sizeof(*icmph) || + iph->protocol !=3D IPPROTO_ICMP) + continue; + icmph =3D (const struct icmphdr *)((const __u8 *)iph + ip_len); + if (icmph->type =3D=3D ICMP_TIME_EXCEEDED) { + *saddr =3D iph->saddr; + return iph->ihl; + } + } + + return -errno; +} + +static void lwt_ip_encap_stale_cb(bool use_freplace, bool use_vrf, + bool pre_encap) +{ + LIBBPF_OPTS(bpf_tc_hook, tc_hook, + .attach_point =3D BPF_TC_INGRESS, + ); + LIBBPF_OPTS(bpf_tc_opts, tc_opts, + .handle =3D 1, + .priority =3D 1, + ); + struct lwt_ip_encap_stale_cb_freplace *freplace_skel =3D NULL; + struct lwt_ip_encap_stale_cb *skel =3D NULL; + struct bpf_program *target, *replacement; + struct bpf_link *freplace_link =3D NULL; + struct netns_obj *netns =3D NULL; + char pin_path[128]; + __u8 mac0[ETH_ALEN], mac1[ETH_ALEN]; + __u32 saddr =3D 0; + bool tc_hook_created =3D false; + int ifindex, packet_fd =3D -1, prog_fd, err, ihl; + + skel =3D lwt_ip_encap_stale_cb__open_and_load(); + if (!ASSERT_OK_PTR(skel, "open_and_load target")) + goto out; + target =3D use_freplace ? skel->progs.lwt_in_freplace_target : + skel->progs.lwt_in_direct; + prog_fd =3D bpf_program__fd(target); + + if (use_freplace) { + freplace_skel =3D lwt_ip_encap_stale_cb_freplace__open(); + if (!ASSERT_OK_PTR(freplace_skel, "open freplace")) + goto out; + replacement =3D freplace_skel->progs.replace_add_ip_encap; + err =3D bpf_program__set_attach_target(replacement, prog_fd, + "add_ip_encap"); + if (!ASSERT_OK(err, "set freplace target")) + goto out; + err =3D lwt_ip_encap_stale_cb_freplace__load(freplace_skel); + if (!ASSERT_OK(err, "load freplace")) + goto out; + freplace_link =3D bpf_program__attach_freplace(replacement, prog_fd, + "add_ip_encap"); + if (!ASSERT_OK_PTR(freplace_link, "attach freplace")) + goto out; + } + + snprintf(pin_path, sizeof(pin_path), STALE_CB_PIN_FMT, getpid()); + unlink(pin_path); + err =3D bpf_program__pin(target, pin_path); + if (!ASSERT_OK(err, "pin target")) + goto out; + + netns =3D netns_new(STALE_CB_NETNS, true); + if (!ASSERT_OK_PTR(netns, "create netns")) + goto out_unpin; + + SYS(out_netns, "ip link add vh0 type veth peer name vh1"); + if (use_vrf) { + SYS(out_netns, "ip link add vrf0 type vrf table 1001"); + SYS(out_netns, "ip link set vrf0 up"); + SYS(out_netns, "ip link set vh1 master vrf0"); + SYS(out_netns, "ip addr add 10.1.0.1/32 dev vrf0"); + SYS(out_netns, "sysctl -wq net.ipv4.conf.vrf0.rp_filter=3D0"); + SYS(out_netns, "sysctl -wq net.ipv4.icmp_errors_use_inbound_ifaddr=3D1"); + } + SYS(out_netns, "ip link set vh0 up"); + SYS(out_netns, "ip link set vh1 up"); + SYS(out_netns, "ip addr add 10.0.0.1/24 dev vh1"); + SYS(out_netns, "sysctl -wq net.ipv4.ip_forward=3D1"); + SYS(out_netns, "sysctl -wq net.ipv4.conf.all.rp_filter=3D0"); + SYS(out_netns, "sysctl -wq net.ipv4.conf.vh1.rp_filter=3D0"); + SYS(out_netns, "sysctl -wq net.ipv4.conf.all.accept_local=3D1"); + + if (pre_encap) { + tc_hook.ifindex =3D if_nametoindex("vh1"); + if (!ASSERT_GT(tc_hook.ifindex, 0, "vh1 ifindex")) + goto out_netns; + err =3D bpf_tc_hook_create(&tc_hook); + if (!ASSERT_OK(err, "create vh1 ingress hook")) + goto out_netns; + tc_hook_created =3D true; + tc_opts.prog_fd =3D bpf_program__fd(skel->progs.tc_pre_encap); + err =3D bpf_tc_attach(&tc_hook, &tc_opts); + if (!ASSERT_OK(err, "attach pre-encapsulation program")) + goto out_netns; + } + + if (!ASSERT_OK(stale_cb_get_mac("vh0", mac0), "get vh0 mac") || + !ASSERT_OK(stale_cb_get_mac("vh1", mac1), "get vh1 mac")) + goto out_netns; + SYS(out_netns, + "ip neigh replace 10.0.0.2 lladdr %02x:%02x:%02x:%02x:%02x:%02x nud p= ermanent dev vh1", + mac0[0], mac0[1], mac0[2], mac0[3], mac0[4], mac0[5]); + SYS(out_netns, + "ip route add %s encap bpf in pinned %s via 10.0.0.2 dev vh1 %s", + STALE_CB_DST, pin_path, use_vrf ? "vrf vrf0" : ""); + + ifindex =3D if_nametoindex("vh0"); + if (!ASSERT_GT(ifindex, 0, "vh0 ifindex")) + goto out_netns; + packet_fd =3D stale_cb_open_packet_socket(ifindex); + if (!ASSERT_OK_FD(packet_fd, "open packet socket")) + goto out_netns; + if (!ASSERT_OK(stale_cb_send_packet(packet_fd, mac0, mac1), + "send crafted packet")) + goto out_netns; + + ihl =3D stale_cb_icmp_ihl(packet_fd, &saddr); + if (!ASSERT_EQ(ihl, 5, "ICMP IPv4 header length")) + goto out_netns; + if (use_vrf && !ASSERT_EQ(saddr, inet_addr("10.0.0.1"), + "ICMP source is ingress slave address")) + goto out_netns; + if (use_freplace) { + ASSERT_TRUE(freplace_skel->bss->freplace_ran, "freplace ran"); + ASSERT_TRUE(freplace_skel->bss->freplace_cb_zero, + "freplace cb was cleared"); + } else { + ASSERT_TRUE(skel->bss->direct_ran, "direct program ran"); + ASSERT_TRUE(skel->bss->direct_cb_zero, "direct cb was cleared"); + } + +out_netns: + if (packet_fd >=3D 0) + close(packet_fd); + if (tc_hook_created) + bpf_tc_hook_destroy(&tc_hook); + netns_free(netns); +out_unpin: + unlink(pin_path); +out: + bpf_link__destroy(freplace_link); + lwt_ip_encap_stale_cb_freplace__destroy(freplace_skel); + lwt_ip_encap_stale_cb__destroy(skel); +} + +void test_lwt_ip_encap_stale_cb(void) +{ + if (test__start_subtest("direct-cb-access")) + lwt_ip_encap_stale_cb(false, false, false); + if (test__start_subtest("freplace-cb-access")) + lwt_ip_encap_stale_cb(true, false, false); + if (test__start_subtest("vrf-direct-cb-access")) + lwt_ip_encap_stale_cb(false, true, false); + if (test__start_subtest("vrf-freplace-cb-access")) + lwt_ip_encap_stale_cb(true, true, false); + if (test__start_subtest("already-encapsulated")) + lwt_ip_encap_stale_cb(false, false, true); +} diff --git a/tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb.c b/to= ols/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb.c new file mode 100644 index 0000000000000..7638db379118e --- /dev/null +++ b/tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb.c @@ -0,0 +1,100 @@ +// SPDX-License-Identifier: GPL-2.0 +#include "vmlinux.h" +#include +#include +#include "bpf_tracing_net.h" + +#define IPOPT_RR 7 +#define IPOPT_MINOFF 4 + +bool direct_cb_zero; +bool direct_ran; + +struct tc_outer_ipv4 { + struct iphdr iph; + __u8 options[8]; +}; + +static __always_inline __u16 fold_csum(__u64 csum) +{ + csum =3D (csum & 0xffffffff) + (csum >> 32); + csum =3D (csum & 0xffff) + (csum >> 16); + csum =3D (csum & 0xffff) + (csum >> 16); + + return ~csum; +} + +SEC("tc") +int tc_pre_encap(struct __sk_buff *skb) +{ + struct tc_outer_ipv4 outer =3D {}; + struct iphdr inner; + __s64 csum; + + if (skb->protocol !=3D bpf_htons(ETH_P_IP)) + return TC_ACT_OK; + if (bpf_skb_load_bytes(skb, ETH_HLEN, &inner, sizeof(inner))) + return TC_ACT_SHOT; + + outer.iph.version =3D 4; + outer.iph.ihl =3D sizeof(outer) / 4; + outer.iph.tos =3D 8; + outer.iph.tot_len =3D bpf_htons(bpf_ntohs(inner.tot_len) + sizeof(outer)); + outer.iph.id =3D bpf_htons(0x2345); + outer.iph.ttl =3D 64; + outer.iph.protocol =3D IPPROTO_IPIP; + outer.iph.saddr =3D bpf_htonl(0x0a000002); /* 10.0.0.2 */ + outer.iph.daddr =3D bpf_htonl(0x0a090909); /* 10.9.9.9 */ + outer.options[0] =3D IPOPT_RR; + outer.options[1] =3D sizeof(outer.options); + outer.options[2] =3D IPOPT_MINOFF; + csum =3D bpf_csum_diff(NULL, 0, (__be32 *)&outer, sizeof(outer), 0); + if (csum < 0) + return TC_ACT_SHOT; + outer.iph.check =3D fold_csum(csum); + + if (bpf_skb_adjust_room(skb, sizeof(outer), BPF_ADJ_ROOM_MAC, + BPF_F_ADJ_ROOM_FIXED_GSO | + BPF_F_ADJ_ROOM_ENCAP_L3_IPV4)) + return TC_ACT_SHOT; + if (bpf_skb_store_bytes(skb, ETH_HLEN, &outer, sizeof(outer), + BPF_F_INVALIDATE_HASH)) + return TC_ACT_SHOT; + + return TC_ACT_OK; +} + +__noinline int add_ip_encap(struct __sk_buff *skb) +{ + struct iphdr iph =3D {}; + + iph.version =3D 4; + iph.ihl =3D 5; + iph.ttl =3D 1; + iph.protocol =3D 4; /* IPPROTO_IPIP */ + iph.tot_len =3D bpf_htons(skb->len + sizeof(iph)); + iph.saddr =3D bpf_htonl(0x0a000002); /* 10.0.0.2 */ + iph.daddr =3D bpf_htonl(0x0a090909); /* 10.9.9.9 */ + + if (bpf_lwt_push_encap(skb, BPF_LWT_ENCAP_IP, &iph, sizeof(iph))) + return BPF_DROP; + + return BPF_OK; +} + +SEC("lwt_in") +int lwt_in_direct(struct __sk_buff *skb) +{ + direct_cb_zero =3D !(skb->cb[0] | skb->cb[1] | skb->cb[2] | + skb->cb[3] | skb->cb[4]); + direct_ran =3D true; + return add_ip_encap(skb); +} + +SEC("lwt_in") +int lwt_in_freplace_target(struct __sk_buff *skb) +{ + return add_ip_encap(skb); +} + +char _license[] SEC("license") =3D "GPL"; diff --git a/tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb_frepla= ce.c b/tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb_freplace.c new file mode 100644 index 0000000000000..6358f76e47c8b --- /dev/null +++ b/tools/testing/selftests/bpf/progs/lwt_ip_encap_stale_cb_freplace.c @@ -0,0 +1,32 @@ +// SPDX-License-Identifier: GPL-2.0 +#include "vmlinux.h" +#include +#include + +bool freplace_cb_zero; +bool freplace_ran; + +SEC("freplace/add_ip_encap") +int replace_add_ip_encap(struct __sk_buff *skb) +{ + struct iphdr iph =3D {}; + + freplace_cb_zero =3D !(skb->cb[0] | skb->cb[1] | skb->cb[2] | + skb->cb[3] | skb->cb[4]); + freplace_ran =3D true; + + iph.version =3D 4; + iph.ihl =3D 5; + iph.ttl =3D 1; + iph.protocol =3D 4; /* IPPROTO_IPIP */ + iph.tot_len =3D bpf_htons(skb->len + sizeof(iph)); + iph.saddr =3D bpf_htonl(0x0a000002); /* 10.0.0.2 */ + iph.daddr =3D bpf_htonl(0x0a090909); /* 10.9.9.9 */ + + if (bpf_lwt_push_encap(skb, BPF_LWT_ENCAP_IP, &iph, sizeof(iph))) + return BPF_DROP; + + return BPF_OK; +} + +char _license[] SEC("license") =3D "GPL"; --=20 2.55.0