From nobody Sat Sep 26 21:13:33 2026 Received: from mail-wm1-f43.google.com (mail-wm1-f43.google.com [209.85.128.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 012343921CD for ; Sun, 30 Aug 2026 07:57:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076691; cv=none; b=YoY2jRTtf/bEFLQFgWj1FfSJ8/PF0a44oziBTzWF/0V+tEBsVM5VM7qrRnOee6VLcSEtUBB9EFFx5WfzNB+jFyyMxtfXlVdG2IcGXQ8du9VUOs5oBP4PWZ3yijsO9L88D2RLshCxzo9pab/vDEBFMz+Md29QCKZ43CKAu1kgSQ0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076691; c=relaxed/simple; bh=QPzvEd+OiJWLrw1OvP16PqVcJmY7BMYLmGhlwyLx5kw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gCqP7Yc0VCRwv8DDJNDDz6FMWHAdEce3QXED+Nwa1RPS6lQS7mi0P8YrgQT88d0PoM1x5ybDXtlsQcqYh1VHRHMiI6Vvko557dVwlXcQnOuvHb61sMj0ovTocU33Sv1s7nlFTXqDcrdqRQqeWAh551b7GUdaatvnSJvgemTKWno= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=D67UlqSP; arc=none smtp.client-ip=209.85.128.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="D67UlqSP" Received: by mail-wm1-f43.google.com with SMTP id 5b1f17b1804b1-4995b0343c1so23284105e9.3 for ; Sun, 30 Aug 2026 00:57:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788076677; x=1788681477; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Q7KBSEmc46+2bR21q/TsshXTy+Zk9y7NLL/izGkQXcU=; b=D67UlqSPDGckdO5W/I36gA8D5OaYNu1cBIezAazZA/hLF1qf23AhN8roYqnULoOkJg zGcuZm6dw+wpWRhPbnedOhkdbh+xAdoT8jeKA7Dt0JUnjIqzHviJbEhRVtuFbUih0p9G KEcN+EfvXB9JFVkxv0D0pZ4zLOMD1pPe1eURblnueJndMAnRhwFc2UOWHoU4iIlS/1Yx 4tNi+mjy0SHDqbku6cue4OMfh4NFm57h4u7kCapBHYrVVLZix8Y+qxXG5pqeL64hSDNl LtOMwyDXKUM7v35dzNx+I7WnAjd4rDiC6yH34XanbxBuPs7XU80EcbbKKChzK/a63tFs HvRQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788076677; x=1788681477; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Q7KBSEmc46+2bR21q/TsshXTy+Zk9y7NLL/izGkQXcU=; b=lUYyZV6hN3I13YhI1qZJXU/kAXgvG91h9Yq4vMIVXIICDy98Mzqow/RWuIA3P8KVrZ kSNhiAAc865bFq8zc7U6FwY7gcKH3amjnEbY0Rc4oCzSaE8qU2Y0hJrvFvTCQzWKtGPA /eIUtlWUoPsCY0OoQrALXQIIOBh0kaIUVl5A53gRmu4YTAgyd1ezHDQhHcH3wmObt/nQ Yxtyav01QOJBQk2rJB8elJEDYm2qddHQ+YVCqRIq/iAWOvpeERJAH4JwGkdX96HAIoYK 6AhDJ0AM2WLHwcpRV4u6Bh1H5/It+ufZN4v6sfbSe/WfTA5QfelP/N0glMgsbnakO7ZL qWiw== X-Forwarded-Encrypted: i=1; AHgh+RpE+Gzkl1aesRSnO59ZO9wCICWVogND1sVa1SIkitb4om/llK8naJVfyO+2ryZG3c+vxBppH1tEGc+3cl4=@vger.kernel.org X-Gm-Message-State: AFuF++nmclSG4mppeJmUlVTPayona3s1I6X9CxX91S1fm3wtkcnN7Zm4 xiPyjl94Lsk+gUDXyTN+4CvEZ8lEc4/ngABu20kFjBrPxKcWJ76u51g= X-Gm-Gg: AR+sD10yM6y0siFi3HgJxfrC8YWNXj1SRVm7VATIQTmdbNqMOuJxheu17JnVLHG42em zmSMZl/RcMJAL3pDhIxV1GtawOoD9zCB5Kaukzs7ze7h5oz17dDoyrgUx7ZQyevXIUmZ8RaVGJn 23IlK0WJHyC0HzpQqnF+xr0bmOmDcxa81sEx9GcoQzxFhrds5jK44XYZ2jI8vXiau1JxahGMb5r Pt88jPohM4wQOeLZbWmFQkWr/PB7xUscVM/HD/6u9FSQDKJsN38KH0Ay5rrdQNQOlLJsOVjS/4M 5WMzDf3gOn8kqzF+Yzd3/xzutZ6G/5DgcDNWfMOAWqQogEyGzdsPISrKmUxFjHOI4j/G1358Uwz GFkr1jcUf+aW/FszfYskqijQ+flTyjZUhje+STne1HQT4PK4jbvYAs6/e7myXOLBYaHSY8veAKp aM/zh5sDXHagaLi1kspC4s1QaZqJPhxlwO9jV9UrUtyGhn7E8Qc6M= X-Received: by 2002:a05:600c:81ca:b0:49b:47b3:d6d with SMTP id 5b1f17b1804b1-49b91c3b7c4mr283650335e9.10.1788076676788; Sun, 30 Aug 2026 00:57:56 -0700 (PDT) Received: from fedora ([46.8.219.5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49cd53a4678sm14757875e9.13.2026.08.30.00.57.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 30 Aug 2026 00:57:56 -0700 (PDT) From: Vitaliy Sochnev To: Lorenzo Bianconi , netdev@vger.kernel.org Cc: Andrew Lunn , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , linux-mediatek@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Vitaliy Sochnev Subject: [PATCH net 1/4] net: airoha: handle RX_NO_CPU_DSCP interrupt, not just RX_DONE Date: Sun, 30 Aug 2026 10:57:14 +0100 Message-ID: <20260830095717.37218-2-sochnev.v.74@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260830095717.37218-1-sochnev.v.74@gmail.com> References: <20260830095717.37218-1-sochnev.v.74@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The QDMA hardware raises a dedicated interrupt (NO_CPU_DSCP, one bit per ring in QDMA_CSR_INT_ENABLE2/3) when an RX ring runs out of free CPU descriptors. airoha_qdma_hw_init() already unmasks this interrupt for every ring (INT_RX1_MASK()/INT_RX2_MASK() OR it together with the RX_DONE bits before writing the enable register), but airoha_irq_handler() only ever extracts the RX_DONE bits from the same status word - the NO_CPU_DSCP bits are read and acknowledged (cleared) along with everything else at the top of the handler, then silently dropped. This matters because once a ring is genuinely drained to zero posted descriptors, no further RX_DONE interrupt can fire for it: hardware has nothing left to receive a frame into, so NAPI is never rescheduled and airoha_qdma_fill_rx_queue() (which reposts descriptors) is never called again. The ring is stuck until the interface is brought down and back up. This is most visible on rings that carry low, bursty volumes of protocol control traffic, in particular RX ring 4, to which airoha_fe_vip_setup() force-routes ~15 unrelated VIP-classified protocols (BOOTP, PPPoE Discovery, ISAKMP, DHCPv6, SIP, LLDP, PPP LCP/IPCP/CHAP/PAP/IPv6CP, ...) via PATN_FCPU_EN_MASK, all sharing the same RX_DSCP_NUM() default of 16 descriptors. A short burst on that ring (e.g. a DHCP lease renewal exchange, or the LCP/IPCP/CHAP/PAP negotiation that follows a PPPoE PADO) can drain it faster than the CPU reposts descriptors, after which every one of those protocols silently stops being received on that device until it is reconfigured - with no error, warning, or netdev/ethtool counter indicating why. Fix airoha_irq_handler() to treat NO_CPU_DSCP the same as RX_DONE for the purpose of scheduling NAPI: airoha_qdma_rx_process() already calls airoha_qdma_fill_rx_queue() unconditionally at the end of every poll, even when zero descriptors were reaped, so scheduling NAPI in response to NO_CPU_DSCP is sufficient to make an emptied ring recover on its own. airoha_qdma_rx_napi_poll() is updated to re-enable the NO_CPU_DSCP bit alongside RX_DONE when napi_complete() runs, mirroring the existing disable/enable dance so the interrupt isn't left masked after its first use. One open question worth flagging explicitly: if the underlying no-free-descriptor condition re-latches this bit immediately after the ack write (rather than only on the next empty->non-empty transition), a ring that airoha_qdma_fill_rx_queue() genuinely cannot repost into (e.g. page_pool_dev_alloc_frag() returning NULL under memory pressure) would turn this into a self-reasserting interrupt storm on the hard IRQ path: mask -> napi_schedule() -> poll reaps 0, refills 0 -> napi_complete() -> unmask -> NO_CPU_DSCP fires again immediately. I don't have documentation confirming which behavior this bit actually has. Regardless of the answer, masking NO_CPU_DSCP until a refill actually succeeds - the natural-looking alternative - is worse: a fully memory-starved ring can never fire RX_DONE either (nothing was posted for hw to complete), so that would leave it permanently dead once the memory pressure clears rather than self-healing. A bounded storm tied to genuine memory pressure, if that's what this is, seems preferable to a ring with no way back either way. Fixes: 23290c7bc190 ("net: airoha: Introduce Airoha NPU support") Link: https://github.com/openwrt/openwrt/issues/24715 Signed-off-by: Vitaliy Sochnev Acked-by: Lorenzo Bianconi --- drivers/net/ethernet/airoha/airoha_eth.c | 28 +++++++++++++++++++----- 1 file changed, 23 insertions(+), 5 deletions(-) diff --git a/drivers/net/ethernet/airoha/airoha_eth.c b/drivers/net/etherne= t/airoha/airoha_eth.c index 64619e9a704d..a3e5aaeb75b3 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.c +++ b/drivers/net/ethernet/airoha/airoha_eth.c @@ -784,13 +784,15 @@ static int airoha_qdma_rx_napi_poll(struct napi_struc= t *napi, int budget) int i, qid =3D q - &qdma->q_rx[0]; int intr_reg =3D qid < RX_DONE_HIGH_OFFSET ? QDMA_INT_REG_IDX1 : QDMA_INT_REG_IDX2; + u32 bit =3D qid % RX_DONE_HIGH_OFFSET; =20 for (i =3D 0; i < ARRAY_SIZE(qdma->irq_banks); i++) { if (!(BIT(qid) & RX_IRQ_BANK_PIN_MASK(i))) continue; =20 airoha_qdma_irq_enable(&qdma->irq_banks[i], intr_reg, - BIT(qid % RX_DONE_HIGH_OFFSET)); + BIT(bit) | + BIT(bit + RX_NO_CPU_DSCP_LOW_OFFSET)); } } =20 @@ -1468,16 +1470,32 @@ static irqreturn_t airoha_irq_handler(int irq, void= *dev_instance) if (!test_bit(DEV_STATE_INITIALIZED, &qdma->eth->state)) return IRQ_NONE; =20 - rx_intr1 =3D intr[1] & RX_DONE_LOW_INT_MASK; + /* A ring can also raise NO_CPU_DSCP when it runs out of free RX + * descriptors (e.g. a burst of VIP-classified control traffic + * forced onto a small ring). Once a ring is fully drained no more + * RX_DONE interrupts can fire for it, since there are no free + * descriptors left for hardware to receive into, so without this + * NAPI is never rescheduled and the ring never gets refilled again. + * Treat NO_CPU_DSCP the same as RX_DONE for scheduling NAPI: + * airoha_qdma_rx_process() unconditionally calls + * airoha_qdma_fill_rx_queue() at the end of every poll, even when + * zero descriptors were reaped, so this alone is enough to recover + * the ring. + */ + rx_intr1 =3D intr[1] & (RX_DONE_LOW_INT_MASK | RX_NO_CPU_DSCP_LOW_INT_MAS= K); if (rx_intr1) { airoha_qdma_irq_disable(irq_bank, QDMA_INT_REG_IDX1, rx_intr1); - rx_intr_mask |=3D rx_intr1; + rx_intr_mask |=3D (rx_intr1 & RX_DONE_LOW_INT_MASK) | + ((rx_intr1 & RX_NO_CPU_DSCP_LOW_INT_MASK) >> + RX_NO_CPU_DSCP_LOW_OFFSET); } =20 - rx_intr2 =3D intr[2] & RX_DONE_HIGH_INT_MASK; + rx_intr2 =3D intr[2] & (RX_DONE_HIGH_INT_MASK | RX_NO_CPU_DSCP_HIGH_INT_M= ASK); if (rx_intr2) { airoha_qdma_irq_disable(irq_bank, QDMA_INT_REG_IDX2, rx_intr2); - rx_intr_mask |=3D (rx_intr2 << 16); + rx_intr_mask |=3D ((rx_intr2 & RX_DONE_HIGH_INT_MASK) | + ((rx_intr2 & RX_NO_CPU_DSCP_HIGH_INT_MASK) >> + RX_NO_CPU_DSCP_LOW_OFFSET)) << 16; } =20 for (i =3D 0; rx_intr_mask && i < ARRAY_SIZE(qdma->q_rx); i++) { --=20 2.55.0 From nobody Sat Sep 26 21:13:33 2026 Received: from mail-wr1-f43.google.com (mail-wr1-f43.google.com [209.85.221.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 87E193932CA for ; Sun, 30 Aug 2026 07:58:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076700; cv=none; b=cTdblLVtPgADl8udMbuMmj/MH08AZBKYNhd+JNcHkJTnfrVq7Ypu2s4spqp+YJrb7aFVkp6cjat2iv2tOsbhm0Sk3lToYhesR584sITRfeThR2Pp94pGvYdUFnx1iVb+O6qQpv0XVO4ijCLZWxHjL3h4aKZGN3Tiixf2j5hV2xw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076700; c=relaxed/simple; bh=x8JllFo9rS1hBK/YFwAJzvhldCMKsOtTSJ/EGfuN8Jo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ugHT2pRH9nY12txVhJb0JLsxXj79xLn2e9xVR3F4jS+sfvEldK1BOEWtfEHkn75kzF5c+sxF3w4wdU1XO1UrI5BgTY2jHkvCWg9Or5Of+n+8DCCg/TyiAF/G+Vd11RwWAPW9ewAB6idEJrCACMH+5NrHohvWlyk+/0ME3kIadKg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=aVvIQUO6; arc=none smtp.client-ip=209.85.221.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="aVvIQUO6" Received: by mail-wr1-f43.google.com with SMTP id ffacd0b85a97d-48433f36a21so496068f8f.1 for ; Sun, 30 Aug 2026 00:58:11 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788076686; x=1788681486; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=lK3jDNN57cOcYB2qd2k1oPSR5oRC9dvZDkaFcSv62Xg=; b=aVvIQUO6fLe0CX+imaZmpzNf3NcGDHcoWtmwKXZJOCFi14dWBWEDdH05M9PApVvGD/ 9YkUOLE5kcN2Y/KIE9scsgUDghYq8+Sb8q+L93XQwK67VbtyOiAa6FB0Xb9zW6DEnoJa +1DsvNmsa05wS6f0YYgJSERpUwsfZlngjfZ6smc4sXH0X7kKk8oS3OftHKc4SNxMKuhv AHlg87ijA4MteN8029IBHtOzG7lWfCzEyGygaj8friCGSsFaIg6z/FsneKlngKd0ePwh FsffMjBJSvk4+LQFiIJzY6GeaM37JcnxHGY0obn8iUvy3k4bbxByTAgQc1eEDbeufoP1 UIng== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788076686; x=1788681486; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=lK3jDNN57cOcYB2qd2k1oPSR5oRC9dvZDkaFcSv62Xg=; b=XI0e97NvCVkeS6frKPk9m2kZm1NvNKPMmmPErjB0wIEj9hsHmiZLm5KsqDFJS9qYUG bnE3EGlU+Vs+Fwb9VgxX2r1EIbEuYBCKSux/4XpJZCaOXW2IoW4v2+6qSGe/xRl/14JO Pb9AFUIEZC8eei09CWxjicOrQSgJFBneVQbMbPz1LviPqpbLQ9J5ooJypf+gxfUGowxB dydQzTUcLsQnVKgGv5UD+jI7e9P9AF1pbATZU70KGc8CJqYbaHYyQrcjjimq4P9BovLq bifaTSZXUPupViLOV42WX5fXIqX0s3vc9U2QzgfhPbxh8ImHWUTlkLj/AUMpw/G8X5LK Mmvg== X-Forwarded-Encrypted: i=1; AHgh+RrjFNdbeIwm9Ma/hPjNeBz83GfLnNu+l8sQvl+vKPCIv60hX0SGjHBCaS2SgZWX0MyvzENEKKQgHqWxDlg=@vger.kernel.org X-Gm-Message-State: AFuF++miG454/TuDOdvX9x+/zY5I2hd0KZI1SBShRzMeWu6HpI2C1uND DyNP/4wv/Lj91eVvPfu3jG49Je5b5m+nhejVlyAjLNNJqgD8ZRmJ5Fw= X-Gm-Gg: AR+sD136uU19ZZdwaCk2h1tpYRypQDX5R+zNcmuGJa3XZzyUfDGWpCdsdhloAynAmqT e5RaIGiccmZ8JYErL4D2Vbd5ULsPwr0m5tUUhrikZUfiONyhgOm1g4w3arZyO4Cl5eZVuQxF63C JRWkMPRPqFatHcgv+em+otN7u3V9aZazL9we/8Fo7ll+N63oDUF4ozOkI2qyV6E6fCyveAKdTF7 UPjC2y6O3D9BxU7kAtOcSXkOU3AAYfPgb+l7YSEvQhCsJcUjr4FXwNg2uQRQxgIHMtvjRxfanin WC6RGyx+x+3HCEArWCYPRMUG2iuaA8wiQGD49KUiDT3ge58wAH1uou5woIRuJ6f2176SoDFSFCw CyBpNQJGwYe9c+vupwFXZhpJFdTLDfjKcVK3jMX7IDVi2ezMlsOrnlfDnQyHuNuKl94A8HmIz2t QgYr7f+P1mEUoiJCxlIWgqAFQ9VmtECp4UKdfAm34mbJBi/bMn0OnQ7MTGIauPzg== X-Received: by 2002:a05:600c:4f49:b0:495:5d6d:9cc1 with SMTP id 5b1f17b1804b1-49cd52b2c93mr18490055e9.0.1788076685856; Sun, 30 Aug 2026 00:58:05 -0700 (PDT) Received: from fedora ([46.8.219.5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49cd53a4678sm14757875e9.13.2026.08.30.00.58.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 30 Aug 2026 00:58:05 -0700 (PDT) From: Vitaliy Sochnev To: Lorenzo Bianconi , netdev@vger.kernel.org Cc: Andrew Lunn , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , linux-mediatek@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Vitaliy Sochnev Subject: [PATCH net-next 2/4] net: airoha: recover RX ring after hw completion race Date: Sun, 30 Aug 2026 10:57:15 +0100 Message-ID: <20260830095717.37218-3-sochnev.v.74@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260830095717.37218-1-sochnev.v.74@gmail.com> References: <20260830095717.37218-1-sochnev.v.74@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On AN7583, hardware can write a completed RX descriptor past the software-posted boundary (q->head) before airoha_qdma_fill_rx_queue() has actually posted a fresh buffer there, as if hardware advances its own completion pointer independently of software's posting bookkeeping. Since airoha_qdma_rx_process() consumes the ring strictly sequentially starting at q->tail, the consumer stalls forever waiting on a descriptor that hardware never writes, even though real, completed frames are sitting further along in the ring. In practice this is reachable during PPPoE/DHCP negotiation bursts on the small shared "force to CPU" ring, and the interface silently stops receiving on it. Detect this without inspecting ring/descriptor memory content at all: REG_RX_DMA_IDX is hardware's own completion counter, independent of what has or hasn't been posted. If it keeps advancing across polls while the software consumer (tail) does not move, hardware is making progress the consumer can never observe - the ring is stuck. Idle rings, where hardware isn't advancing either, are correctly left alone. An earlier version of this recovery instead scanned ahead in the ring for a DONE descriptor and trusted its content; that caused a real OOM panic once it wandered into genuinely uninitialized DMA memory that coincidentally had the DONE bit set. Comparing a hardware register cannot misfire that way. Once a stall is confirmed, defer to a work item (register access here can sleep) that disables RX DMA, waits for it to actually go idle, resyncs the ring via the existing cleanup_rx_queue()/fill_rx_queue() pair - which only ever touches the software-owned [tail, head) window and rewrites both RX_CPU_IDX and RX_DMA_IDX from it - and re-enables RX DMA. This deliberately drops whatever was in flight on the ring rather than trying to identify and preserve the specific descriptor hardware used; a prior attempt at the latter caused a page_pool double-free when the assumptions about which page was safe to free turned out not to hold in all cases. GLOBAL_CFG_RX_DMA_EN_MASK in REG_QDMA_GLOBAL_CFG is per-QDMA-instance, not per-ring, so recovering one ring briefly pauses RX DMA on every ring behind that QDMA (bounded by the 50ms busy-wait below). No per-ring equivalent exists in the register map; this is the same bit airoha_qdma_start()/stop() already use for whole-device up/down. Log the actual measured duration of that pause alongside the recovery message, rather than just citing the 50ms read_poll_timeout() upper bound: that's the real cost paid by every other ring on the same QDMA instance each time recovery fires, and it's worth having the real number instead of the theoretical ceiling. One cost worth calling out explicitly: airoha_qdma_rx_check_stall() adds an uncached MMIO read of RX_DMA_IDX on the RX path, once per airoha_qdma_rx_process() call that ends via the non-DONE break - i.e. essentially every poll, on every ring, even though the condition it detects is rare and specific to one ring. Gating the read on whether the *previous* poll for this ring also reaped zero descriptors would keep it off rings that are actively receiving, and I traced that through for correctness: once the race actually happens, q->tail freezes permanently (consumption here is strictly sequential), so `done` is 0 on every poll after that point, not just some - the gate would only cost about one extra poll before AIROHA_RX_STALL_THRESHOLD is reached, not a suppression or false-negative risk. I haven't implemented that gating here for lack of profiling data justifying the added per-queue state against the (also unmeasured) cost of the current unconditional read; happy to add it if it turns out to matter in practice. Fixes: 23290c7bc190 ("net: airoha: Introduce Airoha NPU support") Signed-off-by: Vitaliy Sochnev --- drivers/net/ethernet/airoha/airoha_eth.c | 135 ++++++++++++++++++++++- drivers/net/ethernet/airoha/airoha_eth.h | 16 +++ 2 files changed, 150 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/airoha/airoha_eth.c b/drivers/net/etherne= t/airoha/airoha_eth.c index a3e5aaeb75b3..b53fe5b17653 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.c +++ b/drivers/net/ethernet/airoha/airoha_eth.c @@ -3,12 +3,14 @@ * Copyright (c) 2024 AIROHA Inc * Author: Lorenzo Bianconi */ +#include #include #include #include #include #include #include +#include #include #include #include @@ -657,6 +659,47 @@ airoha_qdma_get_gdm_dev(struct airoha_eth *eth, struct= airoha_qdma_desc *desc) return port->devs[d] ? port->devs[d] : ERR_PTR(-ENODEV); } =20 +/* number of consecutive polls where hw completion (RX_DMA_IDX) advances + * while the sw consumer (tail) doesn't, before declaring the ring stuck + */ +#define AIROHA_RX_STALL_THRESHOLD 3 + +/* Detect an RX ring where hw's own completion pointer (RX_DMA_IDX) keeps + * moving while the sw consumer (q->tail) doesn't - i.e. hw has written + * further descriptors somewhere in the ring, but the strictly sequential + * consumer can never reach them because the one at q->head, which it is + * waiting on, was never marked DONE. This happens when hw writes a + * completed descriptor past q->head before airoha_qdma_fill_rx_queue() + * has posted a fresh buffer there, decoupled from sw's own posting + * bookkeeping. + * + * Deliberately does not inspect ring/descriptor memory content to detect + * this: an earlier version scanned ahead for a DONE descriptor and trusted + * its content, which caused a real OOM panic after it wandered into + * genuinely uninitialized DMA memory that coincidentally had the DONE bit + * set. RX_DMA_IDX is a hw register with a well-defined value regardless of + * ring content, so this can't misfire on garbage memory, and idle rings + * (no hw progress either) are naturally left alone. + */ +static void airoha_qdma_rx_check_stall(struct airoha_queue *q) +{ + struct airoha_qdma *qdma =3D q->qdma; + int qid =3D q - &qdma->q_rx[0]; + u32 dma_idx =3D airoha_qdma_get(qdma, REG_RX_DMA_IDX(qid), + RX_RING_DMA_IDX_MASK); + + if (q->stall_tail =3D=3D q->tail && dma_idx !=3D q->stall_dma_idx) { + if (++q->stall_count >=3D AIROHA_RX_STALL_THRESHOLD && + !test_and_set_bit(qid, qdma->rx_recover_mask)) + schedule_work(&qdma->rx_recover_work); + } else { + q->stall_count =3D 0; + } + + q->stall_tail =3D q->tail; + q->stall_dma_idx =3D dma_idx; +} + static int airoha_qdma_rx_process(struct airoha_queue *q, int budget) { enum dma_data_direction dir =3D page_pool_get_dma_dir(q->page_pool); @@ -675,8 +718,10 @@ static int airoha_qdma_rx_process(struct airoha_queue = *q, int budget) struct page *page; =20 desc_ctrl =3D le32_to_cpu(READ_ONCE(desc->ctrl)); - if (!(desc_ctrl & QDMA_DESC_DONE_MASK)) + if (!(desc_ctrl & QDMA_DESC_DONE_MASK)) { + airoha_qdma_rx_check_stall(q); break; + } =20 dma_rmb(); =20 @@ -894,6 +939,75 @@ static void airoha_qdma_cleanup_rx_queue(struct airoha= _queue *q) FIELD_PREP(RX_RING_DMA_IDX_MASK, q->tail)); } =20 +static void airoha_qdma_rx_recover_work(struct work_struct *work) +{ + struct airoha_qdma *qdma =3D container_of(work, struct airoha_qdma, + rx_recover_work); + int qid; + + for_each_set_bit(qid, qdma->rx_recover_mask, AIROHA_NUM_RX_RING) { + struct airoha_queue *q =3D &qdma->q_rx[qid]; + ktime_t rx_dma_off_ts; + s64 rx_dma_off_us; + u32 status; + + if (!q->ndesc) + goto next; + + napi_disable(&q->napi); + + /* GLOBAL_CFG_RX_DMA_EN_MASK is per-QDMA, not per-ring, so + * this pauses every RX ring on this QDMA instance, not just + * the stalled one - track how long for, since that's the + * real-world cost of recovery on unrelated rings. + */ + rx_dma_off_ts =3D ktime_get(); + + airoha_qdma_clear(qdma, REG_QDMA_GLOBAL_CFG, + GLOBAL_CFG_RX_DMA_EN_MASK); + if (read_poll_timeout(airoha_qdma_rr, status, + !(status & GLOBAL_CFG_RX_DMA_BUSY_MASK), + USEC_PER_MSEC, 50 * USEC_PER_MSEC, true, + qdma, REG_QDMA_GLOBAL_CFG)) + dev_warn(qdma->eth->dev, + "qid=3D%d RX DMA busy timeout during recovery\n", + qid); + + /* Drop whatever is currently in flight on this ring and + * re-arm it from a known-clean state. cleanup_rx_queue() + * only ever touches the sw-owned [tail, head) window and + * resyncs both RX_CPU_IDX and RX_DMA_IDX to it, which is + * what un-wedges a ring where hw wrote past the sw head + * without the consumer ever advancing - no need to figure + * out which descriptor hw actually used. + */ + airoha_qdma_cleanup_rx_queue(q); + if (q->skb) { + /* discard whatever scatter-gather frame was + * mid-assembly when the stall was hit, cleanup_rx_queue() + * above only resyncs the ring, not this + */ + dev_kfree_skb(q->skb); + q->skb =3D NULL; + } + airoha_qdma_fill_rx_queue(q); + + airoha_qdma_set(qdma, REG_QDMA_GLOBAL_CFG, + GLOBAL_CFG_RX_DMA_EN_MASK); + rx_dma_off_us =3D ktime_us_delta(ktime_get(), rx_dma_off_ts); + + q->stall_count =3D 0; + napi_enable(&q->napi); + napi_schedule(&q->napi); + + dev_warn_ratelimited(qdma->eth->dev, + "qid=3D%d RX ring recovered after hw stall (RX DMA paused for %ll= d us on this QDMA instance)\n", + qid, rx_dma_off_us); +next: + clear_bit(qid, qdma->rx_recover_mask); + } +} + static int airoha_qdma_init_rx(struct airoha_qdma *qdma) { int i; @@ -1594,6 +1708,8 @@ static void airoha_qdma_cleanup(struct airoha_eth *et= h, { int i; =20 + cancel_work_sync(&qdma->rx_recover_work); + if (test_bit(DEV_STATE_INITIALIZED, ð->state)) { u32 status; =20 @@ -1651,6 +1767,15 @@ static int airoha_hw_init(struct platform_device *pd= ev, if (err) return err; =20 + /* INIT_WORK() every instance up front, before any of them can fail + * init and jump to the error path below, since that path tears down + * every eth->qdma[] slot unconditionally, including ones this loop + * never reached. + */ + for (i =3D 0; i < ARRAY_SIZE(eth->qdma); i++) + INIT_WORK(ð->qdma[i].rx_recover_work, + airoha_qdma_rx_recover_work); + for (i =3D 0; i < ARRAY_SIZE(eth->qdma); i++) { err =3D airoha_qdma_init(pdev, eth, ð->qdma[i]); if (err) @@ -1699,6 +1824,14 @@ static void airoha_qdma_stop_napi(struct airoha_qdma= *qdma) { int i; =20 + /* Make sure rx_recover_work is neither running nor able to re-arm + * before any napi_disable() below: it also calls napi_disable()/ + * napi_enable() on q_rx[].napi, and napi_disable() on an + * already-disabled NAPI spins in napi_disable_locked() forever, + * since only napi_enable() clears the state it waits on. + */ + disable_work_sync(&qdma->rx_recover_work); + for (i =3D 0; i < ARRAY_SIZE(qdma->q_tx_irq); i++) napi_disable(&qdma->q_tx_irq[i].napi); =20 diff --git a/drivers/net/ethernet/airoha/airoha_eth.h b/drivers/net/etherne= t/airoha/airoha_eth.h index fa9a8edce22f..483d6b59c351 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.h +++ b/drivers/net/ethernet/airoha/airoha_eth.h @@ -207,6 +207,15 @@ struct airoha_queue { bool txq_stopped; bool flushing; =20 + /* RX hw stall detection: last REG_RX_DMA_IDX/tail snapshot taken + * whenever the head-of-line descriptor isn't DONE, and how many + * consecutive times hw made progress (DMA_IDX moved) while the + * consumer (tail) didn't. See airoha_qdma_rx_check_stall(). + */ + u32 stall_dma_idx; + u16 stall_tail; + u8 stall_count; + struct napi_struct napi; struct page_pool *page_pool; struct sk_buff *skb; @@ -567,6 +576,13 @@ struct airoha_qdma { struct airoha_queue q_tx[AIROHA_NUM_TX_RING]; struct airoha_queue q_rx[AIROHA_NUM_RX_RING]; =20 + /* recovery for RX rings whose hw completion pointer (RX_DMA_IDX) + * keeps moving while the sw consumer is stuck; see + * airoha_qdma_rx_check_stall() and airoha_qdma_rx_recover_work(). + */ + struct work_struct rx_recover_work; + DECLARE_BITMAP(rx_recover_mask, AIROHA_NUM_RX_RING); + DECLARE_BITMAP(qos_channel_map, AIROHA_NUM_QOS_CHANNELS); }; =20 --=20 2.55.0 From nobody Sat Sep 26 21:13:33 2026 Received: from mail-wm1-f47.google.com (mail-wm1-f47.google.com [209.85.128.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A0140390C88 for ; Sun, 30 Aug 2026 07:58:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076702; cv=none; b=mE5fVX16itrHjKVq7sbi5guGjvY4LDj+H/0NMUfFLuyHzzNtqrXzD5dkHBX3a8WLtzT5oxAkOupo2Owx2pKiihEkj3aImdOsm9n9IKqX8egcwvzQDfa0t4XTxnVfBAoYE41rLshyxhCixc8Tr4zXtIAFiClfXHHQRvbtwMwU2bY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076702; c=relaxed/simple; bh=u2Q45OOdAmGx3TrPjohDXloU8EdNrjaZU3VoBzuVIkQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sFqRlVpnjCtK6cz0NXyFzSaqIV8ciUJ2r+gx/hwFwMcgBAr58ny4wAJ6zbMSEvxMnopHCVvop32cEPVk5wtRGTYkuVlJA5bR1Obeav3qWcR1wzM6Rvg+4a9u+1DMSMVyns+w26CboOit2g/6JqrjYgFUPKdobfpsKrB9gFyfqmU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jfRwptkM; arc=none smtp.client-ip=209.85.128.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jfRwptkM" Received: by mail-wm1-f47.google.com with SMTP id 5b1f17b1804b1-4998b5a63e2so23031715e9.1 for ; Sun, 30 Aug 2026 00:58:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788076691; x=1788681491; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cAJxRQQXvVgRQc+qrTy5F1Tifzrj9A/VOza63N0XwI4=; b=jfRwptkMjcY6631c6cZ5RiIyDyRz8PK2laaurQLBAx8/kKmf//53eNX1sN1zUq+8OM SjrpURF0h6bVGqotw7eAiC00NG+TZ+i9JaoZqctWABd1801qxQ5tWQPsQYds+fEBgy7G L79+ji4JJxe0jUz6zvBfhcw5clQK7Hum9J+v/e7FmW9m5gOCzZh7H6iCpfIexUlLxy3m lWHS8ni4FxZlsf40jvlXcInVT2U8RuCYrKuoff6xs1IYxCMCkE1SKWtLh0+iC/8m5/9c j9XpzOdDPNnSaCZGJ6K+Ay43RAcu4MgTYJx8ahauCI+rWa4tIXA5i1qW/hWpLYtpCpp7 Wzfg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788076691; x=1788681491; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=cAJxRQQXvVgRQc+qrTy5F1Tifzrj9A/VOza63N0XwI4=; b=XDk+a7unCfx00KqVEMH/F8lnRmJIfAkFZefm9znLADaNaGjtctPh61bn/SE6Xo++FN b08KQPE4uEAopA+rZyfc/0AeCE4wpG/sl2ftgL1kwHsOkjeo68e0RiJ3Diwxusl8UXe7 JobqtCRAuQ/LE0cni4PjmFRuf1evEK4ASI3yNR/YyU/hVfypzj56ZSPC0T3EE/wVYAkt U8HXR1NoknqDQtfJa0m4KJswVsQ1rEfNDmm384h1ethqRMi4pmBdHjCZehkWFXVo6HC8 FFQ9aGLHA35UC2c31AveE2ayoku2MA0Q6cPLgEh3PsF55gXw2Q6sWewIlJvbzDroejfu dFfg== X-Forwarded-Encrypted: i=1; AHgh+RoAOEtgNT6cCiNXwmE+g+Sl6+rUV6z+m3ZEtg+MVWVBv28rWc7PNI3xz5tyv9RFiQneSr7wBWOCqLQqwr0=@vger.kernel.org X-Gm-Message-State: AFuF++m6N9uethP92S61HVbB8vqLbIVw9AZPxyeMcSehuaMHl6AzBdDG 3tEj1y+scuvOQDxb1uxjz5pPYxgvX5lec+isbS3yyroCZdaq8pBwKCw= X-Gm-Gg: AR+sD11iFqdW9gFMJ+L4oaIaEVVOfyuTY2Az6tOENSAxN7c0bn0I22CbFDjYq9N582n WFkas0WAP9+upSnncZerw9FQZWuHXGyyADozexOnqUZsuzG8DXeIjTpwRqRAFwhWzp46YLh+kil KdUUOGrJ81K0Cdraw6IDhtkVxy6DN6lxj4pxp8SchU5TOQYhoVgC/Bg0moFUWlxaTLO+gbfL37u GJ0/HLNur6pFGIKRCAqxI8JvaiAQWVJu7+TyKOSjXu6xjBYcndcHadxNUjVJ8nQ+7pM1xHC0jyR jx2WQh0xjVdPph1ap/lTBXO2VWP7TYR7B1BfWiNn8vGQKZaEqMmA2oPY1gM9A/k7zN0kLrURPtd rL3A/7rZbDi015mmfIRRkJ5Qh21zxfAJ2Do3ges7deI46Ecz7psIsxGzVethZXjQU6S02fJshf1 ckDHW+d+kNJe68qoEQ4f954XybdyEhyt3h2FpByqjUmUHLRBZujpA= X-Received: by 2002:a05:600c:3546:b0:499:dbc0:370d with SMTP id 5b1f17b1804b1-49b91c1dad9mr266282045e9.2.1788076691084; Sun, 30 Aug 2026 00:58:11 -0700 (PDT) Received: from fedora ([46.8.219.5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49cd53a4678sm14757875e9.13.2026.08.30.00.58.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 30 Aug 2026 00:58:10 -0700 (PDT) From: Vitaliy Sochnev To: Lorenzo Bianconi , netdev@vger.kernel.org Cc: Andrew Lunn , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , linux-mediatek@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Vitaliy Sochnev Subject: [PATCH net-next 3/4] net: airoha: add rx_stall_recover ethtool counter Date: Sun, 30 Aug 2026 10:57:16 +0100 Message-ID: <20260830095717.37218-4-sochnev.v.74@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260830095717.37218-1-sochnev.v.74@gmail.com> References: <20260830095717.37218-1-sochnev.v.74@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Make the RX-ring hw-stall recovery added by the previous commit observable without grepping dmesg for its dev_warn_ratelimited(): add a custom ethtool -S statistic, rx_stall_recover, incremented once per completed recovery. It's tracked per-QDMA-instance rather than per-netdev, since the stalled ring can carry traffic for more than one netdev at once (VIP classification shares ring 4 across several protocols/ports on this hardware) - there's no single netdev to attribute an individual event to, so all netdevs behind the same QDMA instance report the same aggregate count. Kept separate from the actual recovery fix since this adds new ethtool ABI and has no bearing on correctness; happy to have it queued independently if that's preferred. Signed-off-by: Vitaliy Sochnev --- drivers/net/ethernet/airoha/airoha_eth.c | 44 ++++++++++++++++++++++++ drivers/net/ethernet/airoha/airoha_eth.h | 5 +++ 2 files changed, 49 insertions(+) diff --git a/drivers/net/ethernet/airoha/airoha_eth.c b/drivers/net/etherne= t/airoha/airoha_eth.c index b53fe5b17653..efb1dd69cc16 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.c +++ b/drivers/net/ethernet/airoha/airoha_eth.c @@ -1000,6 +1000,7 @@ static void airoha_qdma_rx_recover_work(struct work_s= truct *work) napi_enable(&q->napi); napi_schedule(&q->napi); =20 + qdma->rx_recover_count++; dev_warn_ratelimited(qdma->eth->dev, "qid=3D%d RX ring recovered after hw stall (RX DMA paused for %ll= d us on this QDMA instance)\n", qid, rx_dma_off_us); @@ -2600,6 +2601,46 @@ static void airoha_ethtool_get_drvinfo(struct net_de= vice *netdev, strscpy(info->bus_info, dev_name(eth->dev), sizeof(info->bus_info)); } =20 +static const char airoha_ethtool_stats_str[][ETH_GSTRING_LEN] =3D { + "rx_stall_recover", +}; + +static void airoha_ethtool_get_strings(struct net_device *netdev, u32 sset, + u8 *data) +{ + int i; + + if (sset !=3D ETH_SS_STATS) + return; + + for (i =3D 0; i < ARRAY_SIZE(airoha_ethtool_stats_str); i++) + ethtool_puts(&data, airoha_ethtool_stats_str[i]); +} + +static int airoha_ethtool_get_sset_count(struct net_device *netdev, int ss= et) +{ + return sset =3D=3D ETH_SS_STATS ? + ARRAY_SIZE(airoha_ethtool_stats_str) : -EOPNOTSUPP; +} + +static void airoha_ethtool_get_ethtool_stats(struct net_device *netdev, + struct ethtool_stats *stats, + u64 *data) +{ + struct airoha_gdm_dev *dev =3D netdev_priv(netdev); + struct airoha_qdma *qdma; + + rcu_read_lock(); + qdma =3D rcu_dereference(dev->qdma); + /* aggregate recovery count for the whole qdma instance: the + * stalled ring can carry traffic for more than one netdev (VIP + * classification shares ring 4 across several protocols/ports), + * so there's no single netdev to attribute an individual event to + */ + data[0] =3D qdma ? qdma->rx_recover_count : 0; + rcu_read_unlock(); +} + static void airoha_ethtool_get_mac_stats(struct net_device *netdev, struct ethtool_eth_mac_stats *stats) { @@ -3502,6 +3543,9 @@ static const struct net_device_ops airoha_netdev_ops = =3D { =20 static const struct ethtool_ops airoha_ethtool_ops =3D { .get_drvinfo =3D airoha_ethtool_get_drvinfo, + .get_strings =3D airoha_ethtool_get_strings, + .get_sset_count =3D airoha_ethtool_get_sset_count, + .get_ethtool_stats =3D airoha_ethtool_get_ethtool_stats, .get_eth_mac_stats =3D airoha_ethtool_get_mac_stats, .get_rmon_stats =3D airoha_ethtool_get_rmon_stats, .get_link_ksettings =3D phy_ethtool_get_link_ksettings, diff --git a/drivers/net/ethernet/airoha/airoha_eth.h b/drivers/net/etherne= t/airoha/airoha_eth.h index 483d6b59c351..d6591a779743 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.h +++ b/drivers/net/ethernet/airoha/airoha_eth.h @@ -582,6 +582,11 @@ struct airoha_qdma { */ struct work_struct rx_recover_work; DECLARE_BITMAP(rx_recover_mask, AIROHA_NUM_RX_RING); + /* count of completed hw-stall recoveries, exposed via ethtool -S + * so a recovery event (and the packets it drops) is observable + * without grepping dmesg for the dev_warn_ratelimited() above + */ + u32 rx_recover_count; =20 DECLARE_BITMAP(qos_channel_map, AIROHA_NUM_QOS_CHANNELS); }; --=20 2.55.0 From nobody Sat Sep 26 21:13:33 2026 Received: from mail-wm1-f52.google.com (mail-wm1-f52.google.com [209.85.128.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC67839282C for ; Sun, 30 Aug 2026 07:58:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076707; cv=none; b=pZnIHLSkbyCFCpF3D5cOMs9I2ynkLMqkjecv/EBDGLzpQ0OqL3x7MCyJIz/eaJ2VVkwY+e7YvxSyeKdGCJNJLOVGMbz1BnTfsRldc8kwTRIjPN9drPcrVyJjnL9uKs++oyo0XhyS8yEdanLcoIsSSrbnTxbG0B6nGWRiVdvkB64= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788076707; c=relaxed/simple; bh=jegzOw+EXJIP4W/35cQGumTCQXocudJkdz7vv0uyPE8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kgzwoY/Al75vlNkrj+eJUyogmjvwmNuA4oMKDCCQCCBK+/ApBEipXru0oUYky2vPBDz0BIPLim8zC3Xrjwt5qDbHtU0VTu7ksEJd4z8Hm17cux20uzkqUrXaeAulVDtbeXoPClbKwI3MPkJK7EovgGTd8nlctor7gXdOyusiLyI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=dIMGFTD3; arc=none smtp.client-ip=209.85.128.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="dIMGFTD3" Received: by mail-wm1-f52.google.com with SMTP id 5b1f17b1804b1-4995b0343c1so23285235e9.3 for ; Sun, 30 Aug 2026 00:58:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788076696; x=1788681496; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=+j1AD6ZP4+NwsaE6HRfrB5G8mDdyvb/fqBpbDqu3VZE=; b=dIMGFTD34cL9Vp75CFV88Ji/ud6Q61woZTjg2Bh1TNJjNi7ef1uJWzbmNQy1TDX2aM yeK8qjhzkGiZZAuNS6xwGRmUe4kNbgKCL0FbhRTTm2vt8ZSWLnrXo92j/gy03NsPtT17 GrUwmZmIUkczBgUU52rALxI3qa4QbaqK9K6zy7glSUgz5QMZwaXb1uI+vui44B27WzrS 604b3ER5kPVzt5NGyOcZY73Pl6FfaXUTL+EMmy45RHIjD4MPVUzJaju34ozBTshwE17q zFtSfHbyhFlJUlWIMjUQ+8hahc8Vbl4A1K5a/RjQQcz/vwciaWgMbyqgtjtO+Wp38sTW ddsQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788076696; x=1788681496; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=+j1AD6ZP4+NwsaE6HRfrB5G8mDdyvb/fqBpbDqu3VZE=; b=J+ee3WCs3z442Duatn6uKfJKsUc3XvqxHvQcS7OfAiGJghqIIoxQWFpgOdLo1qZezO R9SHO2NAOJn007R9oEfdlCg9tiBT/b/C/Fnc6PaUPvgy7XGobLIpZHDHO7qS9mg6cqCA UHmpF1ettFgqp4YoYCfCbSlS/7kkHf+NTi1huwQtF1BxtuEW2X5R1RxLRHk2PiRh28Um sdNBP6uC9kip0CR3JqwVZEiy2Jl4IpLCfE4kTd1qs8e/TnLyCfw1dlPO79fQO9b6nv4j H2KoQMx5TgHtXcgp/GFxwYBO5oWA/RhT/rjy1ph0L7CVgjy6dGDdCKgKzIba+Tmh859r IT2Q== X-Forwarded-Encrypted: i=1; AHgh+Rr1JSJpQJMHqHP5ncdQNR3DNKjg/V7lgGhpkKl7JIsQe5mq0Fat5dASxCYtydKm5Sff+6YiGPyMTt6WWlc=@vger.kernel.org X-Gm-Message-State: AFuF++kTjg/RnCEAl52+CN3G96g3v+1k1tWPuDZlpYJrm57Tx8e+JrTa /aEVrcffkt/Q4M/7u4NndtzfAi2qXXzLhyoB7Mk3dW1T96nGgGqWeeV02HnUVW0= X-Gm-Gg: AR+sD12kuG1vO7r1J5/OlEET/8pEQoSl7t/1/+kJaNq0fEIYZZd+Q/iUgalDDKWocNS Vc1xQw+tvibVRr7Y4+WSb3LIqptl5M6SsM+uA0Nb62SxN0e3ylLyYmwUXGxZXSw+L3zQMU3pkBn j7IEJNkduEkIuGbd9OBn9x7XFORdaplrhtO6KFRbizga726O8rnKBYFmBh4Z5m8nTXguRdkIlWs 946Gj5wS3WQNuxvN4EWMxRYhzlXb//YR84KOerM81kmbW5sK1eATDFHBbENfAcgz3nZ38/79hIm 1aLW1eHY8q8GS77omZhA1VF0l/ZC93+bsm3vA9ugTZHHUj6qRv9nN3nkcstxfLm8kA8U52fJzHo 8gFHqVg+uRIWrQC1ygYfjajMRSZ/vm7++2+fuxXV/9cJfZYv+8KItX9u1GuDTEf7ze4SZaPuDkZ 6qNBYwYXeUMI/Sd4j7MwNtwcE3lOwqSvehRh8seIv5cb/Q3YoDlmI= X-Received: by 2002:a05:600c:8b35:b0:499:7a19:408b with SMTP id 5b1f17b1804b1-49b91c3b71dmr267281245e9.11.1788076696142; Sun, 30 Aug 2026 00:58:16 -0700 (PDT) Received: from fedora ([46.8.219.5]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49cd53a4678sm14757875e9.13.2026.08.30.00.58.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 30 Aug 2026 00:58:15 -0700 (PDT) From: Vitaliy Sochnev To: Lorenzo Bianconi , netdev@vger.kernel.org Cc: Andrew Lunn , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , linux-mediatek@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Vitaliy Sochnev Subject: [PATCH net-next 4/4] net: airoha: grow RX ring 4 to 128 descriptors Date: Sun, 30 Aug 2026 10:57:17 +0100 Message-ID: <20260830095717.37218-5-sochnev.v.74@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260830095717.37218-1-sochnev.v.74@gmail.com> References: <20260830095717.37218-1-sochnev.v.74@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Ring 4 is the shared "force to CPU" ring for a wide set of protocols (BOOTP, PPPoE Discovery, PPP LCP/IPCP/CHAP/IPv6CP/PAP, ISAKMP, DHCPv6, SIP, LLDP, ...) and currently falls into the 16-descriptor default in RX_DSCP_NUM(), same as most other non-hashed rings. The hw completion race recovered by airoha_qdma_rx_check_stall()/rx_recover_work() in the previous commits appears strongly correlated with a ring going from idle to receiving its first frame(s) - exactly the access pattern this shared ring sees under protocol negotiation bursts. Give it the same 128-descriptor allowance already used for rings 2/11/15, matching the other rings that see bursty, non-hashed traffic, to reduce how often that condition is hit in the first place. Ring 4 and the VIP classification that forces these protocols onto it are shared driver-wide, with no DT/hardware property distinguishing one chip variant's ring 4 from another's, so this isn't scoped to AN7583 specifically even though that's where the race was found and reproduced. 24h+ stress runs forcing repeated PPPoE/DHCP renegotiation on an AN7581 board, both with and without the two preceding fixes, completed 570+ forced reconnect cycles each with no regressions from the larger ring. Signed-off-by: Vitaliy Sochnev --- drivers/net/ethernet/airoha/airoha_eth.h | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/airoha/airoha_eth.h b/drivers/net/etherne= t/airoha/airoha_eth.h index d6591a779743..cd75c16d8d0c 100644 --- a/drivers/net/ethernet/airoha/airoha_eth.h +++ b/drivers/net/ethernet/airoha/airoha_eth.h @@ -41,6 +41,7 @@ #define TX_DSCP_NUM 1024 #define RX_DSCP_NUM(_n) \ ((_n) =3D=3D 2 ? 128 : \ + (_n) =3D=3D 4 ? 128 : \ (_n) =3D=3D 11 ? 128 : \ (_n) =3D=3D 15 ? 128 : \ (_n) =3D=3D 0 ? 1024 : 16) --=20 2.55.0