drivers/net/phy/phy_device.c | 13 +++++ drivers/net/phy/phylink.c | 12 +++++ include/linux/phy.h | 6 +++ net/dsa/user.c | 98 ++++++++++++++++++++++++++++++++++++ 4 files changed, 129 insertions(+)
A PHY that loads firmware at probe keeps its driver in a module on the rootfs: built in, request_firmware_direct() fails against a rootfs that is not mounted yet and the error comes straight out of probe. A DSA switch probes long before that module can load, and three things go wrong, one per layer. phylink can fail its PHY bringup after it has already recorded the PHY in pl->phydev, leaving a pointer to a PHY that phy_detach() has since released. Depending on the caller that is either a permanent -EBUSY or a stale pointer handed to phy_disconnect() later, which detaches the same PHY twice (patch 1). Keeping a port across a failed connect makes that window reachable, so it comes first. phylib binds the generic driver during the attach, and once the failed connect unwinds, the interrupt the firmware node declared is gone: phy_probe() parked the PHY in polling mode and nothing after MDIO bus registration ever brings the irq back (patch 2). DSA drops the user port when the connect at setup fails, so the port never exists, no matter that the driver shows up seconds later (patch 3). With the series applied the port survives setup and connects its PHY on the first ifup after the module loads, with the interrupt the device tree declares. Tested on an MT7981B board (mt7530 switch, Airoha EN8811H with its INT_B line in the device tree) running a 6.18 backport of everything here except the retry in patch 3: the attach line reports a real interrupt instead of irq=POLL, the EINT is claimed and its counter advances on link changes forced from the link partner, and the port passes traffic. On net-next all three files are compile-tested; the board runs an OpenWrt 6.18 kernel. Aleksei Sviridkin (3): net: phylink: unwind the PHY binding when bringup fails late net: phy: restore the interrupt after a generic-driver bind cycle net: dsa: connect a late-arriving PHY at ifup drivers/net/phy/phy_device.c | 13 +++++ drivers/net/phy/phylink.c | 12 +++++ include/linux/phy.h | 6 +++ net/dsa/user.c | 98 ++++++++++++++++++++++++++++++++++++ 4 files changed, 129 insertions(+) -- 2.43.0
Hi Aleksei
I've had time to think about this, and now have a architecture to
solve the problem which i think it better.
It splits into two parts, getting the PHY firmware downloaded and
registered with phylib, and the phylink handling "hotplug" PHYs.
When power is applied to the "PHY", or after a reset, it is not
actually a PHY. It is a microcontroller sat in its bootloader waiting
for firmware to be downloaded. At that point, it has no PHY
functionality. So lets represent it this way in DT:
davinci_mdio: mdio@5c030000 {
reg = <0x5c030000 0x1000>;
#address-cells = <1>;
#size-cells = <0>;
reset-gpios = <&gpio2 5 1>;
reset-delay-us = <2>;
ethphy0: ethernet-phy@1 {
reg = <1>;
};
mcu: mcu@3 {
compatible = "airoha,en8811h-mcu";
reg = <3>;
}
The compatible here makes it an MDIO device, not a PHY device. The
MDIO subsystem will load an MDIO driver for that compatible, and the
driver can then access device 3 on the MDIO bus. That driver will then
poll the filesystem for the firmware and download it. It might need to
do that in a thread, rather than probe(), i don't know.
Once the firmware starts, we have a PHY. And thinking ahead a bit,
there is no reason this MCU is for a single PHY, it could be a quad
PHY. We need to be able to represent this PHY in DT:
davinci_mdio: mdio@5c030000 {
reg = <0x5c030000 0x1000>;
#address-cells = <1>;
#size-cells = <0>;
reset-gpios = <&gpio2 5 1>;
reset-delay-us = <2>;
ethphy0: ethernet-phy@1 {
reg = <1>;
};
mcu: mcu@3 {
compatible = "airoha,en8811h-mcu";
reg = <3>;
mdio {
ethphy3: ethernet-phy@3 {
reg = <3>;
};
};
};
Have the MDIO device create a new MDIO bus, with pass through
operations to access the underlying MDIO bus, but just for one
address. For all other addresses return -ENODEV. When you register
this MDIO bus, it will get scanned and the PHY found. Since the PHY is
now actually up and running phylib is happy, its usual semantics are
true, the device is ready to go as soon a probe() returns.
As you pointed out, there are currently 3 devices which need to
download firmware. I _guess_ 3/4 of the code can be shared, so please
put must of it into a library, and only have code for actually
downloading to the PHY in the driver.
Then there is a phylink part. This is inspired by how SFP works. We
need some property in the MAC node which indicates the PHY is going to
arrive late. I'm not sure 'hotplug' is the correct description here,
since we know it is there, it is described in DT, it cannot be
exchanged for something else. For the moment, lets just call this
property 'slow-to-probe'. phylink_of_phy_connect() will look for this
property. If it finds 'slow-to-probe', there must also be a phy-handle
pointing to the PHY. phylink then sets itself up to handle this slow
PHY. It needs to poll the phy-handle until it resolves. It can then
call its own phylink_connect_phy() function to connect up the PHY.
As with an SFP, ksetting_get() should return no link modes if the PHY
is not connected yet. ksetting_set() will automatically return EINVAL
when asked to enable a link mode, since none are supported.
eee_get/eee_set should do the same. Since this is how SFPs work, it
should not be too hard to make user space understand an interface can
start out not supporting anything, and then later have various link
modes, autoneg etc.
Please have a think about this architecture, and see if you can find
any holes in it.
Andrew
© 2016 - 2026 Red Hat, Inc.