From nobody Mon Sep 28 11:39:38 2026 Received: from mail-ed1-f49.google.com (mail-ed1-f49.google.com [209.85.208.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA0213C5837 for ; Sat, 22 Aug 2026 15:53:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787413992; cv=none; b=AsBWZFbOSUMOR7PNrbjW1G/XLfjE+8GYBN4kDK8wvU/Bu8bP01xN1Sz5lzYHUbrEBDXNcnif0AdI+k80xRSevJyCqymeVD33wBkIMK1HSfm01PvU9USsG6cwTVTDVBe67e41NDFS4ITCoSFUyMKy2rfMPvzExwb51d10g18CSV4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787413992; c=relaxed/simple; bh=bwFb8+8zfaie+LAvs8/7XUIsL9MQzN7tXA0AU8G3HCs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B4BS95q1XXZgXLNfR8mAGvKt2H9bLAoiVjrlz2OcjvWrxbm6rbyDQWDA+suceV9mR36egDDYWoqP0+t2GjalWvmwAal8pa3LT3PqVM9tVGuRlL5joBQT2+jkEeVU2Sk9qu9PutL4i44VKENUOOwJ2kYaJQoDS5y2mU8PEBXAzro= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la; spf=pass smtp.mailfrom=lex.la; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b=CQsgwUzc; arc=none smtp.client-ip=209.85.208.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lex.la Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b="CQsgwUzc" Received: by mail-ed1-f49.google.com with SMTP id 4fb4d7f45d1cf-6a422090b2fso3095978a12.0 for ; Sat, 22 Aug 2026 08:53:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=lex.la; s=google; t=1787413989; x=1788018789; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=IThcaUItiGTT1yNVv9ohDD3OMYEp3sYsB42Cn7lxU0E=; b=CQsgwUzcWtIsP99cWomQoR75dBMdxvYmegtT1H4M3qROq85FM3Q7oYe6+lIQGfr9WJ J+d5z5gbQnXnic4ic/KFVJAyyIhqX6GtUPieOhj/+wSJ8/9WFsQCSzxIG4H23Uej55QL ePLDGhrOwoA+qJWdJOSciYGsNqTO+a69Jr9OKiOLZYa0CXz3Ug3qEkHrm5/7vR+YIeqj TchKKbRf6ydAxEf0KdcctSz8MGwYHcCNLm9CRT38PGgVnJ1qtvcKxMvPhf7dsgG+1FlS P429gCZWA9InGzLkaEdGL866bmC9b+Vv8Q0Eu5nDAM+k/pZjnVgelP3zbeHTFHcrLoiK 0Xyw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787413989; x=1788018789; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=IThcaUItiGTT1yNVv9ohDD3OMYEp3sYsB42Cn7lxU0E=; b=Y6Vf2Brr3f9sjq0Xt9gGL8Uey7YOlOvWM96U6oapjEe34PWpjnuRjZPtrEQglWgtLt 5qfuB1Vxx0k+7vnOeY4M+cxkHMrOxMSe/hUNvMv36YEaueqP7DOMNLS7tKLp/2ckpvlq OPKFxKQjKhRsuftWn2t/5YpqpqiTYX1gcl9ap4Vm2ZHWHf8SB1XlpIGE9noPwU2Unei1 R4lA2K49fRQoZyZg4fCSmEZK6+P0sfmDPOq8LpI02ysKDFuT8/Yf2Oy7mfO48astUhn6 ljxoyb9pen5rQmmEpvovYn49WOtA1fw6kzgmCVfpH1/KqRYN3jVq6qRy3Cu7sKy+mIUn sppQ== X-Forwarded-Encrypted: i=1; AHgh+Rr/oEtSxDu4elGkJ13yzZMCQyvJQ53f8P7gTIoZmjfi/QYkrlu50NDFnqdHU9etsoHw9z/yim8kWXlpTSA=@vger.kernel.org X-Gm-Message-State: AFuF++n0vMzGA+mCDBMA6odqE2EPSN15kXkyRLyZYSsadSSuAZO5s4yP IE29Kg7RuXWwE0dFHjQUG1b5NS+ojEqbdgSG/+WACiEiLKR/NIZ8DGJjwGtQ66vcqOM= X-Gm-Gg: AR+sD12mCh6GmofYFQa+NG1sV+jWG9b/rDHg25QidoA1w9MyJ56JXcwL0kdbXIMSWO2 jnt3QHT2zTvj96LjJsl6hbvRdEMWm3hpBRUuHwQII+dgR4wCpjMBmoEnS/cDVLhY87hMR4ORZJq qVW0v1bi4ixWW72YaGrO//8CNQY+0Z9eKvXxIoSlMQJymyVetR/nkngcr9ZPrMkxRzh3jibyM5B xgA8eibKERkArKsxE8BCH3mSRMmZmF5jSVgerRGQcDozkrTFpsS/cqK+NaTkvmQoQ/Ve06mmASN JUGcIVqL2XFBG+6HNmfVId38Opavq2YT4SfU/SMmVO0np+1R9hvg91Akz8rsAQr2nQUPBz2iMdL X42J/MVrllMsAz9GvTdIxjGy5Gh1gU9Allx/jFDGBjqf14PECDJvopS7uV0iPGZEdFtKd5Ag1et r1IigKJdAOyUjJJM1DwQU+PQDw8MFydJHKFp6UFYP8SAgKS1wqMxnMwdO9alLnEtvhKog2rQ== X-Received: by 2002:a05:6402:444a:b0:698:351c:979c with SMTP id 4fb4d7f45d1cf-6a42f16f016mr17302770a12.2.1787413988818; Sat, 22 Aug 2026 08:53:08 -0700 (PDT) Received: from ownbook.home.lex.la ([84.17.55.225]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-6a3ff16c546sm12396250a12.21.2026.08.22.08.53.07 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 22 Aug 2026 08:53:08 -0700 (PDT) From: Aleksei Sviridkin To: Andrew Lunn , Vladimir Oltean , Heiner Kallweit , Russell King Cc: "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , netdev@vger.kernel.org, linux-kernel@vger.kernel.org, Aleksei Sviridkin Subject: [PATCH net-next 1/3] net: phylink: unwind the PHY binding when bringup fails late Date: Sat, 22 Aug 2026 18:52:57 +0300 Message-ID: <20260822155259.87146-2-f@lex.la> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260822155259.87146-1-f@lex.la> References: <20260822155259.87146-1-f@lex.la> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" phylink_bringup_phy() records the PHY in pl->phydev before its last fallible step: on a MAC whose phylink ops implement LPI, phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers unwind with phy_detach(), which knows nothing about pl->phydev, so a pointer to a PHY that is no longer attached outlives the failed connect. What that costs depends on how the caller got here. phylink_connect_phy() and the SFP path go through phylink_attach_phy(), which refuses to attach while pl->phydev is set and turns a transient MDIO error into a permanent -EBUSY. phylink_fwnode_phy_connect() has no such check, so a later connect overwrites the stale pointer and hides the problem. A disconnect does not: phylink_disconnect_phy() hands that pointer to phy_disconnect(), and the second phy_detach() on the same PHY drops a device reference and two module references that were only ever taken once. Clear the binding on the failure path, the same three fields phylink_disconnect_phy() clears, under the same locks. The PHY-side fields are left to phy_detach(), which every caller already runs on this path. Signed-off-by: Aleksei Sviridkin --- Reachability The failing step needs pl->mac_supports_eee_ops, i.e. a MAC whose phylink ops implement the LPI callbacks; mt7530 is one, and on the board I tested ethtool --show-eee returns -EOPNOTSUPP, which is what phylink reports when mac_supports_eee_ops is set and mac_supports_eee is not, so that tail runs on every bringup there. The error itself is an MDIO transaction failure inside phy_eee_rx_clock_stop(), which cannot be produced deliberately, so this patch is compile-tested and the series it belongs to ran on hardware with it in place. The double-detach path needs a port that outlives a failed connect, which is what patch 3 introduces; before that, DSA destroyed the port immediately and the stale pointer went with it. drivers/net/phy/phylink.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/drivers/net/phy/phylink.c b/drivers/net/phy/phylink.c index 5b8e95690..9d403ff1b 100644 --- a/drivers/net/phy/phylink.c +++ b/drivers/net/phy/phylink.c @@ -2197,6 +2197,18 @@ static int phylink_bringup_phy(struct phylink *pl, s= truct phy_device *phy, if (ret =3D=3D 0 && phy_interrupt_is_valid(phy)) phy_request_interrupt(phy); =20 + if (ret) { + mutex_lock(&pl->phydev_mutex); + mutex_lock(&phy->lock); + mutex_lock(&pl->state_mutex); + pl->phydev =3D NULL; + pl->phy_enable_tx_lpi =3D false; + pl->mac_tx_clk_stop =3D false; + mutex_unlock(&pl->state_mutex); + mutex_unlock(&phy->lock); + mutex_unlock(&pl->phydev_mutex); + } + return ret; } =20 --=20 2.43.0 From nobody Mon Sep 28 11:39:38 2026 Received: from mail-ed1-f43.google.com (mail-ed1-f43.google.com [209.85.208.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6A8D43F86E0 for ; Sat, 22 Aug 2026 15:53:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787413994; cv=none; b=nYn8PLFgYa7cJMKkOY97MZwNLXi0uSogry93JPRcZnF14iGMMDMd7XvWjabzGuL+tJOAfnhQZC8Z51NrLTbro+MPYalFZEIr4jGIUhCO6+kvKSe5KpYYsIAGuL1aRrXsc3OMwH/DTYFdUUVLX1HB01TXdv7ehGI3TsY7SL5LINg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787413994; c=relaxed/simple; bh=PEdLvA4GpT1cwRKppTTyjQ9Hm086QYTXZOhbR/7hS28=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=siXzJbU3VuzD7v5L4UbZc+LPzIYWA1BYLEmMwWdTU4W9QpLo/TGtw2MZcYtp2khvOl+HrwxCEJM4wyJn142bLGDanxnjiDcAgOwR4S8grRJmy0MwPHV7q8ePkGzJJqDTQCm9FX1M4eZ36LFkZDkTFInfScmrxwGxOUoI0KyhnAE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la; spf=pass smtp.mailfrom=lex.la; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b=Ollze9ua; arc=none smtp.client-ip=209.85.208.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lex.la Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b="Ollze9ua" Received: by mail-ed1-f43.google.com with SMTP id 4fb4d7f45d1cf-6a1a546a6bbso3645672a12.1 for ; Sat, 22 Aug 2026 08:53:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=lex.la; s=google; t=1787413990; x=1788018790; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LtVl3RA330rX+Ddr3fAI1EuluyZ5pYws23mx3vWwVMo=; b=Ollze9uaPiuPgW88DZAMfAqt+pbmNmQcs2dJyUeGbOVPIWbsbblvnTG1rvxw+8Z0z6 6cLAW7PQN7ADk3wq22payJdxt5f2n+jCYqFyGr1wpC1qO8IEMbGSmxwfqVqbjNr378Gx bWy5LDzV0PliYE7WGgnb++f9IcJxSImnERCRAKj/sbmCm+JUJJWphFBtP2CewvzBCFxg 1sIkLpEkCqzxezTOUkuH8dtG9joY2cO/Hto31q2K4NpDJptHCEVS90lGltZBkYiepiJ9 63x1C7jtEg/9lXYGJZyVfLUhtNPc/uWasuXPbVPVKLdRIVJEAWPeWiP8kKKusADrYW1R ZP1w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787413990; x=1788018790; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=LtVl3RA330rX+Ddr3fAI1EuluyZ5pYws23mx3vWwVMo=; b=dxT5cR5Z28yE51a1O++jxIIzeA/9l+Ae8t+tRZ17fN1TGwCR/9c6OJp0Ev3/OtEyCm xqUK4jptT/3ZVnqyovkFQ9LC4cjo5l4t/U3rQRNGPrrPxpAFMG8F3WY6tcOxJPaf4Q/W yfbJEyjt/twpD6RLt/swKUXFkD1yAjiTaipJz2RfzOyEy32FIO0YJihah/gNqQ8tAVla vUJIajlnvVuUJhoCW8pgnrLbsFXi0fOkngfoV8So0U0ZDvD4LU6HE1CgbWvFX+Zog/HH vXxxn08IqdB61Rdo5mY3xebrzAVeRB/NWzdGdVPuWzTUREWv+jt26P5qAC5clpHP1Lqq C7+Q== X-Forwarded-Encrypted: i=1; AHgh+RqJNURY1LyJlz5oD15kwAPdDo7vjAATjDhJ0Aevr4XMV6P+CkJOlyi35iYjxGVAzS1et3uUs2T8xbmsPhw=@vger.kernel.org X-Gm-Message-State: AFuF++mlLfksI/BE0Pd1x/e2bpknPlYDEke+cacGKVZ0CxLvezUuJkTE FPGYWAjABPpTJADGmUGCuY1jZ7/QaBiNicbSjt13nITRxt5zEsAX+uN+b5CbP7sfzFc= X-Gm-Gg: AR+sD13C5exqLwz3n/53xgNm6bQMThYi1cUK8vNyLSGUeDDs15//2dadmBhN1Y+8fA/ Wv/2cplZoRIOvpkPN95RbyW9kUXirKg/LrDGwcdkyq44nFk9UvIZl11xy/IFALZp4HKrBJY+4kb 04OBjYL9hulAYns0TPqPfX8NXsGMkYWb/RmM1jCw24f3cFjZ29seZLL1NgweUNJ8jh1qVfMIy8K 1iZcXQ74aftyMevOp95q8hnFvC3LOnvEmJYhsZAWUzJIPaomghPhz3eyIhD4T6GzVMC2OfpN5WK SG7dDhw8aQrWKo9l6C0yRXBI1s518kSsNh9EQnLaQknqxntz9CGVWiFYDrVQXKkUml34a29wMcx 31GfjHssEFoKStTQ59bHciLTsX301Q7chQ/0u/TnmGOf76HMzBZcd791ASvrGTOCY7H98K0OOgq HRUMQHl4ynHmmn+t2FjL022RUANdQpOm2YsQAQPCnzkmRQrbuE0x5MfU867E1Hc+RGVuWMELJ/R gtBYuip X-Received: by 2002:a05:6402:2186:b0:6a1:2400:baea with SMTP id 4fb4d7f45d1cf-6a42f18a70cmr13414465a12.10.1787413990540; Sat, 22 Aug 2026 08:53:10 -0700 (PDT) Received: from ownbook.home.lex.la ([84.17.55.225]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-6a3ff16c546sm12396250a12.21.2026.08.22.08.53.08 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 22 Aug 2026 08:53:09 -0700 (PDT) From: Aleksei Sviridkin To: Andrew Lunn , Vladimir Oltean , Heiner Kallweit , Russell King Cc: "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , netdev@vger.kernel.org, linux-kernel@vger.kernel.org, Aleksei Sviridkin Subject: [PATCH net-next 2/3] net: phy: restore the interrupt after a generic-driver bind cycle Date: Sat, 22 Aug 2026 18:52:58 +0300 Message-ID: <20260822155259.87146-3-f@lex.la> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260822155259.87146-1-f@lex.la> References: <20260822155259.87146-1-f@lex.la> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" fwnode_mdiobus_phy_device_register() resolves the interrupt declared for a PHY once, at MDIO bus registration. If no specific driver is available when the PHY is attached, the generic driver binds and phy_probe() parks the device in polling mode, since the generic driver has no interrupt callbacks. phy_detach() releases the generic driver so a specific driver can bind later, but nothing brings the interrupt back: the firmware node is never re-read after bus registration, so the specific driver attaches with irq =3D=3D PHY_POLL, phy_request_interrupt() is never reached, and the PHY is polled for the rest of the uptime with nothing in the logs but the "irq=3DPOLL" attach line. A DSA switch probing before the rootfs is mounted produces exactly that cycle for a PHY whose driver is a module: the generic driver binds and fails validation during switch setup, and the real driver binds at ifup. Observed on an MT7981B board with an Airoha EN8811H on an MT7531 port: the device tree declares the INT_B line, yet the attach says irq=3DPOLL and the interrupt is never claimed. Save the interrupt when the generic driver binds and give it back when that driver is released. The restore runs before the device becomes bindable again, so a concurrently arriving specific driver cannot observe or overwrite the intermediate state. Only the value the generic-driver cycle took is restored. A PHY already parked in polling mode before that cycle, by a failed phy_request_interrupt() or by a driver that chose PHY_POLL in its own probe, had PHY_POLL saved, so the restore is skipped. The PHY_F_NO_IRQ and no-interrupt-support checks in phy_attach_direct() still apply to whichever driver binds next. Signed-off-by: Aleksei Sviridkin --- Testing MT7981B board, mt7530 switch, Airoha EN8811H whose INT_B line is in the device tree, driver in a module on the rootfs, together with the next patch: the attach line reports irq=3D15 instead of irq=3DPOLL, the EINT is claimed, its counter advances on link changes forced from the link partner, and there is no interrupt storm. The SoC's internal PHY on the same board, which has no interrupt in its bus table, keeps irq=3DPOLL through the same boot, so the save-restore pair does not resurrect an interrupt the device never had. Consistent across reboots. Hardware testing was done on 6.18 with this exact shape of the change; on net-next the files are compile-tested. The saved value uses zero as "nothing saved"; no registration path produces a valid interrupt number of zero, and non-positive values are never restored. drivers/net/phy/phy_device.c | 13 +++++++++++++ include/linux/phy.h | 6 ++++++ 2 files changed, 19 insertions(+) diff --git a/drivers/net/phy/phy_device.c b/drivers/net/phy/phy_device.c index 94b2e85e0..6047dce61 100644 --- a/drivers/net/phy/phy_device.c +++ b/drivers/net/phy/phy_device.c @@ -1780,6 +1780,7 @@ int phy_attach_direct(struct net_device *dev, struct = phy_device *phydev, else d->driver =3D &genphy_driver.mdiodrv.driver; =20 + phydev->genphy_saved_irq =3D phydev->irq; phydev->is_genphy_driven =3D 1; } =20 @@ -1897,6 +1898,9 @@ int phy_attach_direct(struct net_device *dev, struct = phy_device *phydev, error_module_put: module_put(d->driver->owner); phydev->is_genphy_driven =3D 0; + if (phydev->genphy_saved_irq > 0 && phydev->irq =3D=3D PHY_POLL) + phydev->irq =3D phydev->genphy_saved_irq; + phydev->genphy_saved_irq =3D 0; d->driver =3D NULL; error_put_device: put_device(d); @@ -1965,6 +1969,15 @@ void phy_detach(struct phy_device *phydev) * real driver could be loaded */ if (phydev->is_genphy_driven) { + /* Give back the interrupt phy_probe() parked when the generic + * driver bound, before the device becomes bindable again. A + * PHY that was in polling mode for any other reason had + * PHY_POLL saved, and the restore is skipped. + */ + if (phydev->genphy_saved_irq > 0 && phydev->irq =3D=3D PHY_POLL) + phydev->irq =3D phydev->genphy_saved_irq; + phydev->genphy_saved_irq =3D 0; + device_release_driver(&phydev->mdio.dev); phydev->is_genphy_driven =3D 0; } diff --git a/include/linux/phy.h b/include/linux/phy.h index 5f8d65868..43e20b19e 100644 --- a/include/linux/phy.h +++ b/include/linux/phy.h @@ -591,6 +591,10 @@ struct phy_oatc14_sqi_capability { * - Bits [31:24] are reserved for defining generic * PHY driver behavior. * @irq: IRQ number of the PHY's interrupt (-1 if none) + * @genphy_saved_irq: value of @irq before the generic driver bound, given + * back when that driver is released; zero outside a + * generic bind cycle, and non-positive values are + * never restored * @phylink: Pointer to phylink instance for this PHY * @sfp_bus_attached: Flag indicating whether the SFP bus has been attached * @sfp_bus: SFP bus attached to this PHY's fiber port @@ -762,6 +766,8 @@ struct phy_device { */ int irq; =20 + int genphy_saved_irq; + /* private data pointer */ /* For use by PHYs to maintain extra state */ void *priv; --=20 2.43.0 From nobody Mon Sep 28 11:39:38 2026 Received: from mail-ed1-f41.google.com (mail-ed1-f41.google.com [209.85.208.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BD9993F9F38 for ; Sat, 22 Aug 2026 15:53:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787414008; cv=none; b=AoIYeNe8nXLI8WLRUSqexUXhq0og7fiaelaDX1N3WWw9e2pzFLPq6Egvmr/8nG0kdiZP+8OwTfKGUIzttw8Bk9Rr9+j3MUX4XoS1+rtNGUg80jAbdIaGDs9C6EbMjYY6Fu4a5RQ0089WMJLohGhxp9eS8Fn3PyGn5oXGy841odA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787414008; c=relaxed/simple; bh=dZIHWtcFQLBecZAhEuwWGEx82qTtwRuHcNk7vLDZ0H8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s1XGW3hUhORpfRllc91VV1bSFPDoI+ayT/HdmnR8Dt1AhWSi6joTqYDEBbD4G8paa1XnuF15Sq7Hs9ZSkBl0AFJsjC6dMd07LAtfB7fiiCA3w2uA17Fqy6npCD9rtZ7IKXOV/OKPDi2nus+CC6yL0Ew3laQzFZW/s+qglPG5pK0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la; spf=pass smtp.mailfrom=lex.la; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b=BAE2U5YA; arc=none smtp.client-ip=209.85.208.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=lex.la Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lex.la Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=lex.la header.i=@lex.la header.b="BAE2U5YA" Received: by mail-ed1-f41.google.com with SMTP id 4fb4d7f45d1cf-6a1a546a6bbso3645695a12.1 for ; Sat, 22 Aug 2026 08:53:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=lex.la; s=google; t=1787413992; x=1788018792; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7sx3EpEYcOkpXa9JzxMZLwohk3KOOS93sfAyPQWLR+Q=; b=BAE2U5YAKdkXQeBxmoX0BzzQUnvnjhgDQykZ5ZdSSvuq+0zh1cQnuMXqAbacrI44I0 HtsDHIYDVq810B60ErpnBVL2hhRev+DNMOQcG0JSvpZtF0NL7ectoWU2Z5Gmr46vJKUc SUzXZZlkc1hHPdMFynXaWI7x8zK0rRuvnUERcpFAwh+qMEAUjK6xRYx4xAJdLos6F6XS UmVwTlupk4/7gNVQarjuyybXNwdFxs3yQvKK14aUFHRqUaMnhZK1gcOAbfjddQP4GE2T NVj8QECxoh8ODYd9hVKD7LdABsmoSbhB5Ht8Z6Uynqw7m/4bmEn+llps2sR/L1+zlwXA /+JA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787413992; x=1788018792; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7sx3EpEYcOkpXa9JzxMZLwohk3KOOS93sfAyPQWLR+Q=; b=qb0NoJBG68I4P7HWcpAidUnOarqurR+RzCURQlCWhUemyY+hPsHaYWRXS4lLORSUFL OEa9nx+kLVrmV/NHCKZOMSuPk/si8d7Ppi+uJ2DfjcoaWk/tJVN3SzEwTQb8i5IecZMK JZLkA48JNVzJSxVFg1vpYrO4fuviViUTyVaYR3Rd34SeOZDaMGdhZCcCAaB4Pola9oK7 7iHMZOhWGlzt9K9JDw7JeO+FvvKTvsI+npGms6Yo3Nf8T2p37990LpG5EGplPTsmRm0U PwvBE7McxTZGNcFBAM6Pb1F+c2PG5oXyWqUOHKckx4aft8gzd2EZurwxlv9CMHWdoC/g YD3A== X-Forwarded-Encrypted: i=1; AHgh+RrfnHoBbBD4PI6WeSTl/f1yS7GJ1cGAgqJX1AuJQOHpEnoiwn7nCvXAeaudnjMRJn0G53D9Np22PbWWtkE=@vger.kernel.org X-Gm-Message-State: AFuF++nZF8+fW07JI7MbXELfhlcoM+qrSaLb1dlInR5yYXbeP9/OWnTb ymfBImfxMYI76aTa8DjGcP8dkgEo9oZYRDAO4LGDTp7BNWqWnTYAM1rh2fu3RALm6u4= X-Gm-Gg: AR+sD13nxdkcLXldmCoRY+lsC2AS/+zIePWGLzWuAmPdZycrgIc+watCmTyfpFGUQXB AixtGXnfKXaVS++g1pvjHeAxU/FhRlWlCSjbxkkIqDmPbuPTBMfjftez5vBUs3pqpuKQCKEujaw cSy7r8+h8kn1ehOHuyD0B3AE1gZc9xc/4qu8Tz6C5ZyqWS2feen2J21gLC5/WaXwmXYfxdJ91r1 xlbQ6JtZNSUM9pzMAKBfvW+dbFd71rWZQTRQiXhNv3jKx0xf/PHP2O/lRxopGsgifjcVR+pbA2k vTO4kygoQ1RZEQ75TKA7Ti9q2xx7yLwou+vEZxjk6S1NYoHQ2NlRDI+TfrTTntzos/VLuv2hIZw FoXQAUXBn0u3vZosbl6MwLZCC86yo5815lFnBGe37xyifYhfgZ+CMs0VDHO6EmuvyrSYK5JsUQH cVt/4cTdc1LpW/dzQBjo97WQCmQ+oiMcwF/2eR591Uqkzkxud7RsJBmtbwAdT3j9b8I+lRdw== X-Received: by 2002:a05:6402:1f02:b0:6a1:28d8:2306 with SMTP id 4fb4d7f45d1cf-6a42f0f7c19mr16106289a12.1.1787413991910; Sat, 22 Aug 2026 08:53:11 -0700 (PDT) Received: from ownbook.home.lex.la ([84.17.55.225]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-6a3ff16c546sm12396250a12.21.2026.08.22.08.53.10 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 22 Aug 2026 08:53:11 -0700 (PDT) From: Aleksei Sviridkin To: Andrew Lunn , Vladimir Oltean , Heiner Kallweit , Russell King Cc: "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , netdev@vger.kernel.org, linux-kernel@vger.kernel.org, Aleksei Sviridkin Subject: [PATCH net-next 3/3] net: dsa: connect a late-arriving PHY at ifup Date: Sat, 22 Aug 2026 18:52:59 +0300 Message-ID: <20260822155259.87146-4-f@lex.la> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260822155259.87146-1-f@lex.la> References: <20260822155259.87146-1-f@lex.la> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A PHY whose driver loads firmware at probe has no driver bound while that module still sits in an unmounted rootfs. phy_attach_direct() falls back to the generic driver, whose feature set lacks the modes the port is wired for, and the port is dropped for the rest of the uptime: mt7530-mdio mdio-bus:1f lan4: validation of 2500base-x [...] failed: -EIN= VAL mt7530-mdio mdio-bus:1f lan4: error -22 setting up PHY for tree 0, switch= 0, port 5 The same PHY wired to a MAC on the same SoC comes up. The difference is when the connect happens: the MAC driver connects from ndo_open, DSA connects during setup, at 1.9 s, before any rootfs exists. Keep the port when the connect fails on a PHY that has no driver of its own, and connect it from the open path instead, ahead of enabling the port, so a port in MLO_AN_PHY mode is never started without a PHY. A PHY whose driver bound in the meantime connects normally. Deferring the switch probe instead does not converge. Every retry re-runs the port setup and flaps the other user ports. Measured on an MT7981B board at two deadlines: the probe gave up at 22.2 s and the module arrived at 27.6 s with a 20 s deadline, 47.3 s and 52.7 s with a 45 s one. Arrival tracks the deadline at a constant offset, so raising the deadline moves the target with it. The module is loaded by userspace, and userspace is what the retries keep from running. The hotplug traffic each retry generates is the suspected mechanism; that part was not measured. Building the PHY driver into the kernel does not help either. It loads firmware with request_firmware_direct(), which does not fall back to the usermode helper, and passes the error straight out of probe rather than -EPROBE_DEFER, so a built-in driver fails once against the unmounted rootfs and is never retried. Baking the blobs in through CONFIG_EXTRA_FIRMWARE puts 144 KiB into every kernel built from that configuration. That call belongs to whoever configures the kernel. The driverless check races the driver arrival itself: the module can bind between the failed connect and the check reading the driver pointer. That failed attempt ran with the generic driver bound and says nothing about the driver present now, so the connect is retried once, and only for a PHY that had no driver when the attempt started. A port whose PHY was bound all along keeps its single attempt, so no port pays a second attach, reset toggle and validation warning for a window it can never be in. The predicate covers every PHY named in the port's description that the generic driver serves, including ones that are not waiting for a module. A port whose PHY cannot satisfy the configured phy-mode is registered now instead of being dropped at setup, and each ifup on it fails with a warning. Each failed ifup binds the generic driver, re-reads its abilities over MDIO and releases it again. That cycle used to run once at setup; on such a port it now runs per attempt, and a PHY that declares a reset line sees it toggled each time. Signed-off-by: Aleksei Sviridkin --- Why not the existing phy_detach() path phy_detach() already releases the generic driver so a real one can bind later, and the obvious question is why DSA cannot use it. It is not reachable here. dsa_port_setup_as_unused() never creates a netdevice for the port, so there is no ndo_open to hook and no detach to trigger. Turning that into create-then-detach is a larger and riskier change than keeping the port and retrying the connect that already exists. Why the check reads mdio.dev.driver and not phy_driver_is_genphy() The flag behind that helper is set in phy_attach_direct() and cleared in phy_detach(), which the failed connect has already run by the time DSA looks, so it reads false on exactly the ports this patch is for. Sampling before the connect does not work either: nothing is bound to any PHY at that point, so the check would match every port. Reading the bound driver from DSA is a reach into the device model. If that is the wrong layer, a small phylib accessor fits here and I can add one. Relation to earlier attempts The 2021 RFC "Make the PHY library stop being so greedy when binding the generic PHY driver" went at the same problem class from the phylib side, adding device_pending_probe() so phy_attach_direct() could hold off on the generic driver while a specific driver's probe was still pending. It was turned down as belonging in the driver core rather than in phylib, and nothing equivalent has landed since. This patch stays out of that argument: it touches neither probe deferral nor driver matching, and changes only what DSA does with a connect that already failed. The phy_port work does not cover this case. It represents port topology and runs from phy_probe(), after a driver is bound; the decision this patch depends on happens earlier, in the fallback inside phy_attach_direct(). Blast radius Ports whose PHY has a driver at setup time never take the new path. Their connect succeeds and the branch is not reached. A port kept across a failed connect is what makes patch 1 of this series necessary: without it the failed bringup leaves a pointer to a detached PHY behind, and the teardown of a port that was never opened would detach that PHY a second time. The retry doubles the connect attempt on ports that fail with a driver already bound, including ports that end up dropped anyway. Each attempt is a full attach/detach cycle, so the reset line is toggled once more and phylink prints its validation warning twice. phylink_bringup_phy() requests the PHY interrupt only after validation succeeds. On a board whose interrupt description is wrong, that description stays dormant as long as the connect always fails, and goes live the moment this patch makes the connect work. The patch is the trigger there, not the cause. Patch 2 of this series is what makes a correct description survive to that point at all. Testing MT7981B board, mt7530 switch, Airoha EN8811H on port 5, driver in a module on the rootfs. Without the patch the port is dropped at 1.9 s and stays gone. With it the port survives setup, the PHY driver binds at 6.3 s, and the port attaches it and joins the bridge on the first ifup at 16.9 s. Boot time is unchanged, and the late connect leaves the phylink instances of the other ports alone. The retry taken when a driver binds during the failed connect closes a window of microseconds; it cannot be exercised deliberately on hardware and is compile-tested, as is this patch on net-next. The hardware testing was done on a 6.18 backport carrying everything here except that retry. net/dsa/user.c | 98 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 98 insertions(+) diff --git a/net/dsa/user.c b/net/dsa/user.c index 041f9060c..0cdc1c9b1 100644 --- a/net/dsa/user.c +++ b/net/dsa/user.c @@ -397,6 +397,49 @@ void dsa_user_host_uc_uninstall(struct net_device *dev) dsa_port_standalone_host_fdb_del(dp, dev->dev_addr, 0); } =20 +/* Returns the port's PHY node with a reference taken, or NULL if the port= has + * no PHY node. Looks at the same properties as phylink_fwnode_phy_connect= (). + */ +static struct fwnode_handle *dsa_user_phy_fwnode(const struct dsa_port *dp) +{ + struct fwnode_handle *phy_fwnode; + + phy_fwnode =3D fwnode_get_phy_node(of_fwnode_handle(dp->dn)); + + return IS_ERR(phy_fwnode) ? NULL : phy_fwnode; +} + +/* Connect a PHY that dsa_user_phy_setup() left behind because it had no + * driver of its own back then. Runs before the port is enabled, so that a + * port in MLO_AN_PHY mode is not started without its PHY. + */ +static int dsa_user_late_phy_connect(struct net_device *dev) +{ + struct dsa_port *dp =3D dsa_user_to_port(dev); + struct fwnode_handle *phy_fwnode; + struct dsa_switch *ds =3D dp->ds; + u32 phy_flags =3D 0; + int err; + + if (dev->phydev) + return 0; + + phy_fwnode =3D dsa_user_phy_fwnode(dp); + if (!phy_fwnode) + return 0; + + fwnode_handle_put(phy_fwnode); + + if (ds->ops->get_phy_flags) + phy_flags =3D ds->ops->get_phy_flags(ds, dp->index); + + err =3D phylink_of_phy_connect(dp->pl, dp->dn, phy_flags); + if (err) + netdev_warn(dev, "could not connect PHY: %pe\n", ERR_PTR(err)); + + return err; +} + static int dsa_user_open(struct net_device *dev) { struct net_device *conduit =3D dsa_user_to_conduit(dev); @@ -413,6 +456,10 @@ static int dsa_user_open(struct net_device *dev) if (err) goto out; =20 + err =3D dsa_user_late_phy_connect(dev); + if (err) + goto out_del_host_uc; + err =3D dsa_port_enable_rt(dp, dev->phydev); if (err) goto out_del_host_uc; @@ -2651,11 +2698,37 @@ static int dsa_user_phy_connect(struct net_device *= user_dev, int addr, return phylink_connect_phy(dp->pl, user_dev->phydev); } =20 +/* Whether the port's PHY is served by the generic driver rather than by o= ne + * of its own. Only meaningful after a failed connect, which releases the + * generic driver again. + */ +static bool dsa_user_phy_lacks_driver(const struct dsa_port *dp) +{ + struct fwnode_handle *phy_fwnode; + struct phy_device *phydev; + bool lacks_driver; + + phy_fwnode =3D dsa_user_phy_fwnode(dp); + if (!phy_fwnode) + return false; + + phydev =3D fwnode_phy_find_device(phy_fwnode); + fwnode_handle_put(phy_fwnode); + if (!phydev) + return false; + + lacks_driver =3D !READ_ONCE(phydev->mdio.dev.driver); + put_device(&phydev->mdio.dev); + + return lacks_driver; +} + static int dsa_user_phy_setup(struct net_device *user_dev) { struct dsa_port *dp =3D dsa_user_to_port(user_dev); struct device_node *port_dn =3D dp->dn; struct dsa_switch *ds =3D dp->ds; + bool had_driver, retried =3D false; u32 phy_flags =3D 0; int ret; =20 @@ -2678,6 +2751,8 @@ static int dsa_user_phy_setup(struct net_device *user= _dev) if (ds->ops->get_phy_flags) phy_flags =3D ds->ops->get_phy_flags(ds, dp->index); =20 + had_driver =3D !dsa_user_phy_lacks_driver(dp); +connect: ret =3D phylink_of_phy_connect(dp->pl, port_dn, phy_flags); if (ret =3D=3D -ENODEV && ds->user_mii_bus) { /* We could not connect to a designated PHY or SFP, so try to @@ -2686,6 +2761,29 @@ static int dsa_user_phy_setup(struct net_device *use= r_dev) ret =3D dsa_user_phy_connect(user_dev, dp->index, phy_flags); } if (ret) { + if (dsa_user_phy_lacks_driver(dp)) { + /* Not known to be reachable from the internal MDIO bus + * fallback, which assigns user_dev->phydev before it + * connects, but do not hand the open path a leftover. + */ + user_dev->phydev =3D NULL; + + netdev_info(user_dev, + "PHY has no driver, connecting it at open\n"); + return 0; + } + + /* A driver that was not there before this attempt is one that + * bound while it ran: the failure came from the generic driver + * and says nothing about this one, so try once more. A port + * that had its driver all along keeps the single attempt. + */ + if (!had_driver && !retried) { + had_driver =3D true; + retried =3D true; + goto connect; + } + netdev_err(user_dev, "failed to connect to PHY: %pe\n", ERR_PTR(ret)); dsa_port_phylink_destroy(dp); --=20 2.43.0