[PATCH net v3 0/2] net/mlx5e: Prevent stale XSK buffer release on refill retries

Jerome Tollet posted 2 patches 1 month, 1 week ago
There is a newer version of this series
drivers/net/ethernet/mellanox/mlx5/core/en_rx.c | 14 ++++++++++----
1 file changed, 10 insertions(+), 4 deletions(-)
[PATCH net v3 0/2] net/mlx5e: Prevent stale XSK buffer release on refill retries
Posted by Jerome Tollet 1 month, 1 week ago
Prevent duplicate XSK buffer release when a deferred RX refill fails and
the same WQE is retried.

Patch 1 fixes legacy cyclic RQ. It is unchanged from v2 and retains
Dragos' Reviewed-by tag.

Patch 2 fixes the analogous striding-RQ MPWQE path pointed out by the
Sashiko automated review. Targeted fault injection forced three
consecutive allocation failures for one MPWQE. Stock freed the same 16
XSK buffer pointers three times; the fix freed each pointer only once,
made subsequent releases no-ops, and preserved the successful
allocation path's bitmap reset.

Changes in v3:
- Turn the submission into a two-patch series.
- Add the striding-RQ MPWQE fix and its dedicated Fixes tag.
- Rebase onto net/main at 4e15e89faac9.

v2: https://lore.kernel.org/netdev/20260820151558.11015-1-jtollet@cisco.com/

Jerome Tollet (2):
  net/mlx5e: Prevent stale XSK buffer release on refill retry
  net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry

 drivers/net/ethernet/mellanox/mlx5/core/en_rx.c | 14 ++++++++++----
 1 file changed, 10 insertions(+), 4 deletions(-)

-- 
2.55.0
[PATCH net v4 0/2] net/mlx5e: Prevent stale XSK buffer release on refill retries
Posted by Jerome Tollet 1 month ago
Prevent duplicate XSK buffer release when a deferred RX refill fails and
the same WQE is retried.

Patch 1 fixes legacy cyclic RQ. It is unchanged from v3 and retains
Dragos' Reviewed-by tag.

Patch 2 fixes the analogous striding-RQ MPWQE path. Following Dragos'
review, it now fills skip_release_bitmap in the common error path of
mlx5e_xsk_alloc_rx_mpwqe(), consistently with mlx5e_alloc_rx_mpwqe().

Targeted fault injection covered both an early allocation failure and a
partial 8-of-16-buffer unwind. With three consecutive failures for one
MPWQE, the original 16 XSK buffers were released only once, retries saw a
full bitmap, and a later successful allocation cleared it. A clean
20-second AF_XDP zero-copy pressure run exercised 1,575,262 buffer
allocation failures without invalid descriptors, WQE errors, or kernel
warnings.

Changes in v4:
- Keep patch 1 unchanged.
- Move the MPWQE bitmap update from the release loop to the common XSK
  allocator error path, as suggested by Dragos.
- Validate both early and partial allocation failures and run an
  additional zero-copy pressure test.

v3: https://lore.kernel.org/netdev/cover.1787347981.git.jtollet@cisco.com/

Jerome Tollet (2):
  net/mlx5e: Prevent stale XSK buffer release on refill retry
  net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry

 drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c | 2 ++
 drivers/net/ethernet/mellanox/mlx5/core/en_rx.c     | 7 +++++--
 2 files changed, 7 insertions(+), 2 deletions(-)


base-commit: 4e15e89faac9f308baeb01f46c13a051814d2449
-- 
2.55.0
[PATCH net v4 1/2] net/mlx5e: Prevent stale XSK buffer release on refill retry
Posted by Jerome Tollet 1 month ago
When an XDP redirect to an AF_XDP socket fails because its RX ring is
full, the XSK core frees the buffer. During the subsequent batched refill
of a legacy cyclic RQ, mlx5e also releases the WQE's XSK buffer before
allocating a replacement. If that refill succeeds only partially, a WQE
left without a replacement retains its old buffer pointer.

The buffer can meanwhile be allocated to another WQE. A later refill
retry can then free the live buffer through the stale pointer and publish
the same UMEM frame twice.

Mark the WQE as released immediately after the driver-side free. The flag
is already cleared when a replacement buffer is assigned, so refill
retries no longer release stale pointers.

The failure is silent and produces no kernel warning or splat. A
standalone legacy cyclic-RQ zero-copy libxsk reproducer, using 64-byte UDP
traffic offered at 12 Mpps, detected it: stock stopped after 2,854,914
packets in 4.094 seconds, with 4,542 xdp_rx_ring_full events and 64
ownership/double-publication errors. With this change it processed
356,904,225 packets in 30 seconds despite 571,405 xdp_rx_ring_full events,
with no ownership or data errors.

Fixes: 3f93f82988bc ("net/mlx5e: RX, Defer page release in legacy rq for better recycling")
Cc: stable@vger.kernel.org
Suggested-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Jerome Tollet <jtollet@cisco.com>
---
 drivers/net/ethernet/mellanox/mlx5/core/en_rx.c | 7 +++++--
 1 file changed, 5 insertions(+), 2 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
index 206cf9db3466..7bd0606a5253 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
@@ -410,8 +410,11 @@ static inline void mlx5e_free_rx_wqe(struct mlx5e_rq *rq,
 
 static void mlx5e_xsk_free_rx_wqe(struct mlx5e_wqe_frag_info *wi)
 {
-	if (!(wi->flags & BIT(MLX5E_WQE_FRAG_SKIP_RELEASE)))
-		xsk_buff_free(*wi->xskp);
+	if (wi->flags & BIT(MLX5E_WQE_FRAG_SKIP_RELEASE))
+		return;
+
+	xsk_buff_free(*wi->xskp);
+	wi->flags |= BIT(MLX5E_WQE_FRAG_SKIP_RELEASE);
 }
 
 static void mlx5e_dealloc_rx_wqe(struct mlx5e_rq *rq, u16 ix)
-- 
2.55.0
[PATCH net v4 2/2] net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry
Posted by Jerome Tollet 1 month ago
With AF_XDP on a striding RQ, mlx5e defers releasing XSK buffers until
an MPWQE is refilled. If XSK allocation then returns -ENOMEM,
actual_wq_head is not advanced and a later NAPI poll retries the same
WQE.

mlx5e_free_rx_mpwqe() leaves each released slot marked as releasable. On
retry it can therefore call xsk_buff_free() again through stale pointers
after the frames have returned to the XSK pool and been reallocated.

Set all skip_release_bitmap bits in the common error path of
mlx5e_xsk_alloc_rx_mpwqe(). This matches mlx5e_alloc_rx_mpwqe(). A
successful allocation already clears the bitmap after replacing every
buffer, so retries become idempotent without changing the success path.

Fault injection forced three consecutive failures for one selected MPWQE.
Both an early allocation failure and a partial 8-of-16-buffer unwind
released the original 16 XSK buffers only once. Each error left a full
bitmap, the following NAPI retry skipped the release, and a later
successful allocation cleared it. A 20-second AF_XDP zero-copy pressure
run exercised 1,575,262 buffer allocation failures without invalid
descriptors, WQE errors, or kernel warnings.

Fixes: 4c2a13236807 ("net/mlx5e: RX, Defer page release in striding rq for better recycling")
Cc: stable@vger.kernel.org
Signed-off-by: Jerome Tollet <jtollet@cisco.com>
---
 drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
index 4f984f6a2cb9..55ec6387ab28 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
@@ -3,6 +3,7 @@
 
 #include "rx.h"
 #include "en/xdp.h"
+#include <linux/bitmap.h>
 #include <net/xdp_sock_drv.h>
 #include <linux/filter.h>
 
@@ -156,6 +157,7 @@ err_reuse_batch:
 		xsk_buff_free(xsk_buffs[batch]);
 
 err:
+	bitmap_fill(wi->skip_release_bitmap, rq->mpwqe.pages_per_wqe);
 	rq->stats->buff_alloc_err++;
 	return -ENOMEM;
 }
-- 
2.55.0
Re: [PATCH net v4 2/2] net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry
Posted by Jerome Tollet 1 month ago
Adding William Tu to Cc, as suggested by get_maintainer.pl.
Apologies for the omission.

Regards,
Jerome
Re: [PATCH net v4 2/2] net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry
Posted by Jerome Tollet 1 month ago
Hi Dragos,

Thanks for the review and for suggesting the error-path approach.

Regards,
Jerome
Re: [PATCH net v4 2/2] net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry
Posted by Dragos Tatulea 1 month ago

On 24.08.26 16:16, Jerome Tollet wrote:
> With AF_XDP on a striding RQ, mlx5e defers releasing XSK buffers until
> an MPWQE is refilled. If XSK allocation then returns -ENOMEM,
> actual_wq_head is not advanced and a later NAPI poll retries the same
> WQE.
> 
> mlx5e_free_rx_mpwqe() leaves each released slot marked as releasable. On
> retry it can therefore call xsk_buff_free() again through stale pointers
> after the frames have returned to the XSK pool and been reallocated.
> 
> Set all skip_release_bitmap bits in the common error path of
> mlx5e_xsk_alloc_rx_mpwqe(). This matches mlx5e_alloc_rx_mpwqe(). A
> successful allocation already clears the bitmap after replacing every
> buffer, so retries become idempotent without changing the success path.
> 
> Fault injection forced three consecutive failures for one selected MPWQE.
> Both an early allocation failure and a partial 8-of-16-buffer unwind
> released the original 16 XSK buffers only once. Each error left a full
> bitmap, the following NAPI retry skipped the release, and a later
> successful allocation cleared it. A 20-second AF_XDP zero-copy pressure
> run exercised 1,575,262 buffer allocation failures without invalid
> descriptors, WQE errors, or kernel warnings.
> 
> Fixes: 4c2a13236807 ("net/mlx5e: RX, Defer page release in striding rq for better recycling")
> Cc: stable@vger.kernel.org
> Signed-off-by: Jerome Tollet <jtollet@cisco.com>
> ---
>  drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c | 2 ++
>  1 file changed, 2 insertions(+)
> 
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
> index 4f984f6a2cb9..55ec6387ab28 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en/xsk/rx.c
> @@ -3,6 +3,7 @@
>  
>  #include "rx.h"
>  #include "en/xdp.h"
> +#include <linux/bitmap.h>
>  #include <net/xdp_sock_drv.h>
>  #include <linux/filter.h>
>  
> @@ -156,6 +157,7 @@ err_reuse_batch:
>  		xsk_buff_free(xsk_buffs[batch]);
>  
>  err:
> +	bitmap_fill(wi->skip_release_bitmap, rq->mpwqe.pages_per_wqe);
>  	rq->stats->buff_alloc_err++;
>  	return -ENOMEM;
>  }

Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>

Thanks,
Dragos