[PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips

Birger Koblitz posted 13 patches 1 month, 3 weeks ago
There is a newer version of this series
MAINTAINERS                        |    6 +
drivers/net/phy/ax88796b.c         |  180 ++++++
drivers/net/usb/Kconfig            |   10 +-
drivers/net/usb/Makefile           |    3 +-
drivers/net/usb/ax88179_178a.c     |  664 ++------------------
drivers/net/usb/ax88179_lib.c      |  515 ++++++++++++++++
drivers/net/usb/ax88179_lib.h      |  358 +++++++++++
drivers/net/usb/ax88179a_devices.c | 1201 ++++++++++++++++++++++++++++++++++++
8 files changed, 2313 insertions(+), 624 deletions(-)
[PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Birger Koblitz 1 month, 3 weeks ago
This adds support for the current generation of ASIX network adapter chips,
which are based on the AX88179A. This includes the AX88179A/B (1GBit-PHY),
AX88772D/E (100MBit) and AX88279 (2.5GBit).

The AX179A-based chips all provide both a CDC-NCM compatible USB interface,
and a proprietary vendor interface with more features. By default, the
proprietary vendor interface is not active and Linux will load the CDC-NCM
driver to support the devices. If the ax88179_178a module is configured by
the OS to have precedence over CDC-NCM, then this driver will switch the
device to use the vendor interface, and the device will be controlled by
the ax88179_178a driver when the device is probed again after an automatic
reset of the device bringing up the vendor interface.

The following hardware was tested:
Delock 66046 2.5GBit adapter (AX88279, FW: 1.2.0.0)
TP-Link UE306 1GBit adapter (AX88179B, FW: 1.3.0.0)
Renkforce RF-4708614 1GBit adapter (AX88179A, FW: 1.0.4.0)
UGREEN CR110 100MBit adapter (AX88722E, FW: 1.3.0.0)

The driver supports the following features
- EEE
- TCP segmentation offload
- VLAN filtering/tagging offload 
  (NETIF_F_HW_VLAN_CTAG_FILTER, NETIF_F_HW_VLAN_CTAG_RX/TX)
- RX/TX checksum offload
- FC/Pause configuration
- EEPROM read access

The code is based on the ASIX 4.1.0 out-of-tree driver published under
the GPL,, the aqc111 driver which provides support for the AX88279A,
and some tracing of USB-transfers of the Windows-driver.

Signed-off-by: Birger Koblitz <mail@birger-koblitz.de>
Tested-by: Jianhui Xu <neuromoments@gmail.com>
---
Changes in v6:
- Use genphy_read_status() in asix_ax88279_read_status()
- Use SGMII and 2500BaseX interfaces for PHY
- Fix speed determination of MAC/PHY link
- Fix bulk transfer configuration to use enums for bulk configuration types
- Link to v5: https://lore.kernel.org/r/20260802-ax88179a-v5-0-dcb9fea4acd4@birger-koblitz.de

Changes in v5:
- Fixed read_status() and config_aneg() in PHY driver
- Introduced netdev2data() as convenience function
- netdev_info instances and debugging relicts removed
- Several instances of badly written/formatted code that was copy and pasted from
  the original driver corrected
- select PHYLINK added to Kconfig when phylib dependency added
- PHYLINK selects PHYLIB, so not needed to specify separately
- Link to v4: https://lore.kernel.org/r/20260731-ax88179a-v4-0-2cf1f71b1dd2@birger-koblitz.de

Changes in v4:
- Split driver into library part and part2 for AX88179 and AX88179A-based
  controllers
- Driver renamed ax88179
- Improved phylink use: use phylink standard functions for speed and EEE-settings,
  correct MAC capabilities, removed ax88179_status() irq-urb callback
- Fixes in PHY driver for AX88179A integrated PHYs
- Link to v3: https://lore.kernel.org/r/20260724-ax88179a-v3-0-bdde4f905883@birger-koblitz.de

Changes in v3:
- Add PHY drivers for the PHYs of the AX88179A-based controllers
- Use phylink for the AX88179A-based chips
- Link to v2: https://lore.kernel.org/r/20260708-ax88179a-v2-0-0800fedb2e16@birger-koblitz.de

Changes in v2:
- Correctly use net-next prefix
- Fix compilation issue in HW support patch
- Split MMD support patch into patches for EEE/new chip support
- Do not use ADVERTISE_RESV but private flag definition
- Fix pause configuration to keep track of settings when autoneg disabled
- Fix issue with unitialized variable reported by kernel test robot <lkp@intel.com>
- Avoid white-space changes

- Link to v1: https://lore.kernel.org/r/20260701-ax88179a-v1-0-13685df67515@birger-koblitz.de

---
Birger Koblitz (13):
      ax88179_178a: Fix endianness of pause watermark register
      ax88179_178a: Split driver into library and device specific code
      ax88179_178a: Add netdev2data() convenience function
      ax88179_178a: Add HW support for AX179A-based chips
      ax88179_178a: Add EEE configuration support for AX88179A MACs
      ax88179_178a: Add EEE configuration support for AX88179A PHYs
      ax88179_178a: Add VLAN offload support for AX88179A
      ax88179_178a: Add AX179A/AX279 multicast configuration
      ax88179_178a: Add Suspend/resume support for AX88179A/772D/279
      ax88179_178a: Add ethtool get_drvinfo
      ax88179_178a: Update driver name and information
      ax88179_178a: Add support for AX88179A/772D/279 EEPROM access
      ax88796b: Add support for AX88772D, AX88179A and AX88279

 MAINTAINERS                        |    6 +
 drivers/net/phy/ax88796b.c         |  180 ++++++
 drivers/net/usb/Kconfig            |   10 +-
 drivers/net/usb/Makefile           |    3 +-
 drivers/net/usb/ax88179_178a.c     |  664 ++------------------
 drivers/net/usb/ax88179_lib.c      |  515 ++++++++++++++++
 drivers/net/usb/ax88179_lib.h      |  358 +++++++++++
 drivers/net/usb/ax88179a_devices.c | 1201 ++++++++++++++++++++++++++++++++++++
 8 files changed, 2313 insertions(+), 624 deletions(-)
---
base-commit: 1a9edf8be190decb17227e3cba540513d93ebb85
change-id: 20260630-ax88179a-a1d89fe21730

Best regards,
-- 
Birger Koblitz <mail@birger-koblitz.de>
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Jianhui Xu 1 month, 3 weeks ago
Hi Birger,

I tested v6 on the same ASIX AX88179B adapter (USB 0b95:1790,
bcdDevice 0x0200, firmware 1.3.0.0).

The 13 patches applied to net-next commit
df13c1df8147675470213ffff29dd5762fa321f5 and built successfully as
7.2.0-rc3-ax88179b-v6. The full build and focused W=1 builds for ax88179.o
and ax88796b.o were clean.

All 13 fresh direct-kernel QEMU starts completed a new DHCPDISCOVER at
1000baseT/Full without reloading the driver. This includes five functional
runs and the eight independent suspend/resume runs described below, so
I did not reproduce the v5 cold zero-RX failure.

However, I reproduced the intermittent 100-Mbit carrier-without-RX problem
in two of the five functional runs. In both failures, advertise 0x008
negotiated 100baseT/Full and reported carrier, but ARP remained incomplete,
bound pings to both the gateway and test host failed, and the RX counter
did not move (60 to 60 and 61 to 61) while TX increased. The same
transition passed in the other three runs.

In both failed runs, the subsequent advertise 0x002 transition negotiated
10baseT/Full and passed traffic, and restoring the default advertisement
negotiated 1000baseT/Full and passed traffic.

Default 1000baseT/Full, 10baseT/Full-only, restored 1000baseT/Full, EEE
disable/restore, pause enable/restore, and a read-only EEPROM query
otherwise passed in all five functional runs. In the fifth run,
I additionally unloaded and reloaded ax88179 and ax88796b; DHCP, both bound
traffic paths, and RX growth passed afterward. QEMU USB detach/reattach
also removed the device, reprobed it, reacquired DHCP, and passed both
traffic paths with RX growth.

I also ran eight independent QEMU ACPI S3 suspend/resume trials: two with
Wake-on-LAN disabled and six with magic-packet wake configured. QEMU's
monitor confirmed every guest was paused in S3, and I resumed each guest
with system_wakeup. All eight returned SSH and 1000baseT/Full carrier,
passed bound gateway and test-host traffic immediately after resume, and
showed RX growth immediately and again during the 10- and 20-second delayed
checks. Thus I did not reproduce the v5 post-resume frozen-RX state in
these eight v6 trials.

I then investigated the reproducible 100baseT/Full failure further. The
immediate failure mechanism is that the adapter's MAC loses
AX_MEDIUM_RECEIVE_EN (0x0100) after link configuration.

With a diagnostic register dump, a successful 100baseT/Full transition
reported:

medium mode: 0x0102
RX_CTL: 0x0198
MAC path: 0x03
bulk-in: 05 c0 04 06 0f

A failed transition reported:

medium mode: 0x0002
RX_CTL: 0x0198
MAC path: 0x03
bulk-in: 05 c0 04 06 0f

Thus the only captured difference was AX_MEDIUM_RECEIVE_EN being clear.
ftrace from a separate failure also showed that bulk-IN URBs stopped
completing after the link transition even though usbnet_bh continued to
run.

An immediate readback in mac_link_up() was not sufficient. An exact build
with that diagnostic reproduced zero RX after an EEE restore, showing that
AX_MEDIUM_RECEIVE_EN can be lost after mac_link_up() has returned.

As an experiment, I therefore added a delayed check one second after
link-up. If carrier is still present and AX_MEDIUM_RECEIVE_EN is clear, the
worker restores the bit and verifies it by readback. The work is cancelled
on link-down and synchronously cancelled during stop, suspend, and detach.

In the first run with this final experimental patch, the worker directly
detected and repaired the condition twice, after the 100baseT/Full link-up
in stress cycles 03 and 07. Both cycles then passed bound gateway and
test-host traffic. All ten 100baseT/Full -> 1000baseT/Full stress cycles
passed, as did cold DHCP, 10 Mbit/s, EEE, pause, module reload, and USB
detach/reattach.

A second fresh run passed another ten stress cycles without a failure.
Separate deep-S3 cycles with both wol d and wol g passed immediate,
10-second, and 20-second post-resume traffic and RX checks.

The experimental patch builds from v6 head b61cb69fb19f0 and passes focused
W=1 builds plus strict checkpatch (0 errors, 0 warnings, 0 checks). For
reference, the experimental diff is:

    diff --git a/drivers/net/usb/ax88179_lib.h b/drivers/net/usb/ax88179_lib.h
    --- a/drivers/net/usb/ax88179_lib.h
    +++ b/drivers/net/usb/ax88179_lib.h
    @@ -319,6 +319,7 @@ struct ax88179_data {
     	struct phy_device *phydev;
     	struct phylink *phylink;
     	struct phylink_config phylink_config;
    +	struct delayed_work rx_check;
     	int (*resume)(struct usb_interface *intf);
     	int (*suspend)(struct usb_interface *intf, pm_message_t message);
     };
    diff --git a/drivers/net/usb/ax88179a_devices.c b/drivers/net/usb/ax88179a_devices.c
    --- a/drivers/net/usb/ax88179a_devices.c
    +++ b/drivers/net/usb/ax88179a_devices.c
    @@ -124,6 +124,7 @@ static int ax88179a_suspend(struct usb_interface *intf, pm_message_t message)
     	u8 tmp8;
     
     	priv = dev->driver_priv;
    +	cancel_delayed_work_sync(&priv->rx_check);
     	ax88179_set_pm_mode(dev, true);
     
     	if (netif_running(dev->net)) {
    @@ -417,16 +418,56 @@ static int ax88179a_init_phy(struct usbnet *dev)
     	return 0;
     }
     
    +static void ax88179a_rx_check(struct work_struct *work)
    +{
    +	struct ax88179_data *data;
    +	struct usbnet *dev;
    +	u16 mode;
    +	int i, ret;
    +
    +	data = container_of(to_delayed_work(work), struct ax88179_data,
    +			    rx_check);
    +	dev = netdev_priv(to_net_dev(data->phylink_config.dev));
    +
    +	if (!netif_device_present(dev->net) || !netif_running(dev->net) ||
    +	    !netif_carrier_ok(dev->net))
    +		return;
    +
    +	ret = ax88179_read_cmd(dev, AX_ACCESS_MAC, AX_MEDIUM_STATUS_MODE,
    +			       2, 2, &mode);
    +	if (ret != 2 || (mode & AX_MEDIUM_RECEIVE_EN))
    +		return;
    +
    +	netdev_warn(dev->net, "RX disabled after link configuration, restoring\n");
    +	for (i = 0; i < 3; i++) {
    +		mode |= AX_MEDIUM_RECEIVE_EN;
    +		ret = ax88179_write_cmd(dev, AX_ACCESS_MAC,
    +					AX_MEDIUM_STATUS_MODE, 2, 2, &mode);
    +		if (ret != 2)
    +			continue;
    +
    +		ret = ax88179_read_cmd(dev, AX_ACCESS_MAC,
    +				       AX_MEDIUM_STATUS_MODE, 2, 2, &mode);
    +		if (ret == 2 && (mode & AX_MEDIUM_RECEIVE_EN))
    +			return;
    +	}
    +
    +	netdev_err(dev->net, "failed to restore RX after link configuration\n");
    +}
    +
     static void ax88179a_mac_config(struct phylink_config *config, unsigned int mode,
     				const struct phylink_link_state *state)
     {
     	/* Nothing to do */
     }
     
     static void ax88179a_mac_link_down(struct phylink_config *config,
     				   unsigned int mode, phy_interface_t interface)
     {
    -	/* Nothing to do */
    +	struct usbnet *dev = netdev_priv(to_net_dev(config->dev));
    +	struct ax88179_data *data = dev->driver_priv;
    +
    +	cancel_delayed_work(&data->rx_check);
     }
     
     static void ax88179a_mac_link_up(struct phylink_config *config,
    @@ -544,6 +585,8 @@ static void ax88179a_mac_link_up(struct phylink_config *config,
     
     	tmp8 = AX_MAC_RX_PATH_READY | AX_MAC_TX_PATH_READY;
     	ax88179_write_cmd(dev, AX_ACCESS_MAC, AX88179A_MAC_PATH, 1, 1, &tmp8);
    +
    +	mod_delayed_work(system_wq, &ax179_data->rx_check, HZ);
     }
     
     static void ax88179a_mac_disable_tx_lpi(struct phylink_config *config)
    @@ -741,6 +784,7 @@ static int ax88179a_bind(struct usbnet *dev, struct usb_interface *intf)
     		return -ENOMEM;
     
     	dev->driver_priv = ax179_data;
    +	INIT_DELAYED_WORK(&ax179_data->rx_check, ax88179a_rx_check);
     
     	ret = ax88179_read_cmd(dev, AX_ACCESS_MAC, AX_CHIP_STATUS,
     			       1, 1, &ax179_data->chip_version);
    @@ -840,6 +884,7 @@ static void ax88179a_unbind(struct usbnet *dev, struct usb_interface *intf)
     	u16 tmp16;
     	u8 tmp8;
     
    +	cancel_delayed_work_sync(&ax179_data->rx_check);
     	/* Configure RX control register => stop operation */
     	tmp16 = AX_RX_CTL_STOP;
     	ax88179_write_cmd(dev, AX_ACCESS_MAC, AX_RX_CTL, 2, 2, &tmp16);
    @@ -1149,6 +1194,7 @@ static int ax88179a_stop(struct usbnet *dev)
     	u16 reg16;
     	u8 reg8;
     
    +	cancel_delayed_work_sync(&ax179_data->rx_check);
     	ax88179_read_cmd(dev, AX_ACCESS_MAC, AX_MEDIUM_STATUS_MODE, 2, 2, &reg16);
     	reg16 &= ~AX_MEDIUM_RECEIVE_EN;
     	ax88179_write_cmd(dev, AX_ACCESS_MAC, AX_MEDIUM_STATUS_MODE, 2, 2, &reg16);

Because the unmodified v6 series still reproduces the intermittent
100baseT/Full carrier-without-RX failure, I cannot add a Tested-by for v6.

Thanks,
Jianhui
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Birger Koblitz 1 month, 3 weeks ago
Hi Jianhui,

thanks for testing this so thoroughly, again!
On 09/08/2026 03:36, Jianhui Xu wrote:
> Hi Birger,
> 
> I tested v6 on the same ASIX AX88179B adapter (USB 0b95:1790,
> bcdDevice 0x0200, firmware 1.3.0.0).
> 
> The 13 patches applied to net-next commit
> df13c1df8147675470213ffff29dd5762fa321f5 and built successfully as
> 7.2.0-rc3-ax88179b-v6. The full build and focused W=1 builds for ax88179.o
> and ax88796b.o were clean.
> 
> All 13 fresh direct-kernel QEMU starts completed a new DHCPDISCOVER at
> 1000baseT/Full without reloading the driver. This includes five functional
> runs and the eight independent suspend/resume runs described below, so
> I did not reproduce the v5 cold zero-RX failure.
> 
> However, I reproduced the intermittent 100-Mbit carrier-without-RX problem
> in two of the five functional runs. In both failures, advertise 0x008
> negotiated 100baseT/Full and reported carrier, but ARP remained incomplete,
> bound pings to both the gateway and test host failed, and the RX counter
> did not move (60 to 60 and 61 to 61) while TX increased. The same
> transition passed in the other three runs.
So the bottom line is that everything works, except that there are spurious
failures with 100baseT/Full links where RX is disabled after the link
was established. I have so far not seen this myself, but will test with more
and different adapters as link partners in order to reproduce the issue. I tested
the 100MBit connections mainly with a AX88772E 100MBit adapter (UGREEN CR110),
which has the same firmware (1.3.0.0) as your and my AX88179B adapter. You did
not mention which device is used on the other side of the Ethernet link (or maybe
I missed that), could you specify this?

> An immediate readback in mac_link_up() was not sufficient. An exact build
> with that diagnostic reproduced zero RX after an EEE restore, showing that
> AX_MEDIUM_RECEIVE_EN can be lost after mac_link_up() has returned.
The only way this could be coming from the driver that I see is via a call to
ax88179a_stop(), which would clear exactly that bit.
Have you traced this and can exclude that this function is called somehow?
If this is not the case, this would mean there is a bug in the firmware
of the adapter which clears AX_MEDIUM_RECEIVE_EN in some cases for 100baseT/Full,
which seems to be what you also seem to suspect based on your proposed patch.

> As an experiment, I therefore added a delayed check one second after
> link-up. If carrier is still present and AX_MEDIUM_RECEIVE_EN is clear, the
> worker restores the bit and verifies it by readback. The work is cancelled
> on link-down and synchronously cancelled during stop, suspend, and detach.
If this can indeed be attributed to a bug in the firmware of the adapters, I would
add your patch to the series with an "Authored-by" you, as this sounds like a
good solution for this issue.

Birger
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Jianhui Xu 1 month, 2 weeks ago
Sorry for the late reply. I think the previous experimental workaround may
have some race conditions, so I was seeking for a better solution.

> I tested the 100MBit connections mainly with a AX88772E 100MBit adapter
> (UGREEN CR110), which has the same firmware (1.3.0.0) as your and my
> AX88179B adapter.

My adapter is an AX88179B with firmware v1.3.0.3, not v1.3.0.0.

> You did not mention which device is used on the other side of the
> Ethernet link (or maybe I missed that), could you specify this?

The link partner was the Ethernet port of a ZTE ZXHN F7005MV3 gateway, not
another USB Ethernet adapter. ethtool reported autonegotiation support and
no advertised pause frames.

> The only way this could be coming from the driver that I see is via a call
> to ax88179a_stop(), which would clear exactly that bit.
> Have you traced this and can exclude that this function is called somehow?

Yes. I added a dynamic kprobe on ax88179_write_cmd(), filtered to
AX_MEDIUM_STATUS_MODE writes, and recorded the value and caller stack.

During 30 100baseT/Full-to-1000baseT/Full cycles, four transitions to
100baseT/Full lost RX. In all four cases:

- ax88179a_mac_link_up() first wrote 0x0102;
- there was no intervening Linux write to AX_MEDIUM_STATUS_MODE;
- about one second later the delayed worker read the register with
  AX_MEDIUM_RECEIVE_EN clear and restored 0x0102.

All 71 traced writes to AX_MEDIUM_STATUS_MODE had
AX_MEDIUM_RECEIVE_EN set. The complete 1,209-entry trace contained no
ax88179a_stop() or ax88179_change_mtu() caller.

The earlier failed-state dump also retained AX_RX_CTL at 0x0198 with carrier
up. I therefore think ax88179a_stop() can be excluded as the direct source
of these clears.

I also repeated the test with the same AX88179B and ZTE link partner using
the ASIX vendor driver on the Arch host. All 30 further
100baseT/Full-to-1000baseT/Full cycles passed. All 4,415 carrier-up
100baseT/Full samples retained receive-enable at 0x0132. A separate trace
recorded 30 writes of 0x0132 and 30 writes of 0x0133, with no write clearing
AX_MEDIUM_RECEIVE_EN.

Dense sampling did show the adapter changing 0x0133 to 0x0033, or 0x0132 to
0x0032, while the link was down during renegotiation, without a corresponding
vendor-driver write. The vendor link-setting path then restored 0x0132 or
0x0133 before the link became stably up.

So the device can clear AX_MEDIUM_RECEIVE_EN without a corresponding host
write. In the failing v6 case, the clear likewise was not caused by a Linux
write and appears to be an autonomous device-side change.

I do not think this proves an unconditional device-side bug, though. The v6
driver failed four times in 30 cycles after writing 0x0102, while the vendor
driver had no failures after writing 0x0132.

I therefore tested a focused v6 variant that writes 0x0132 instead of 0x0102
at 100baseT/Full. Failures still occurred, so retaining the RX/TX
flow-control bits alone is not sufficient to prevent the problem.

This points to some other difference in the vendor driver's link-setting
sequence, possibly register ordering or timing. The physical xHCI host in
my tests versus QEMU's emulated xHCI is another uncontrolled difference.

At this point this looks like a device-side quirk exposed by the driver's
link-setting sequence, but I cannot distinguish firmware behavior from
autonomous MAC hardware behavior.

> If this can indeed be attributed to a bug in the firmware of the
> adapters, I would add your patch to the series with an "Authored-by" you,
> as this sounds like a good solution for this issue.

I found several race conditions in the patch. For example,
`cancel_delayed_work(...)` only cancels pending work. If the callback has
already started running, it may still be running when
`cancel_delayed_work(...)` returns. Therefore, if another execution context
performs a read-modify-write operation on `MEDIUM_STATUS`, there can be
a race: the link is brought down, but the worker subsequently writes the
RX-enable bit back. This particular issue can be fixed by using
`cancel_delayed_work_sync(...)`, but there are still other races.

One possible solution would be to add a mutex to serialize accesses to the
medium register. However, I am reluctant to add too much synchronization
machinery for what is essentially a defensive workaround, especially given
how infrequently these network configuration operations occur.

I see two options:
1. Leave the code mostly as it is, changing only
   `cancel_delayed_work(...)` to `cancel_delayed_work_sync(...)`. This keeps
   the main code path simple and clear, at the cost of leaving a few rare
   corner cases unresolved.
2. Add stronger synchronization to eliminate these races completely, at the
   cost of making the code considerably more complex for cases that are
   unlikely to occur in practice.

What do you think?

Thanks,
Jianhui
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Birger Koblitz 1 month, 2 weeks ago
On 8/10/26 03:35, Jianhui Xu wrote:
>> The only way this could be coming from the driver that I see is via a call
>> to ax88179a_stop(), which would clear exactly that bit.
>> Have you traced this and can exclude that this function is called somehow?
> 
> Yes. I added a dynamic kprobe on ax88179_write_cmd(), filtered to
> AX_MEDIUM_STATUS_MODE writes, and recorded the value and caller stack.
> 
> During 30 100baseT/Full-to-1000baseT/Full cycles, four transitions to
> 100baseT/Full lost RX. In all four cases:
> 
> - ax88179a_mac_link_up() first wrote 0x0102;
> - there was no intervening Linux write to AX_MEDIUM_STATUS_MODE;
> - about one second later the delayed worker read the register with
>    AX_MEDIUM_RECEIVE_EN clear and restored 0x0102.
> 
> All 71 traced writes to AX_MEDIUM_STATUS_MODE had
> AX_MEDIUM_RECEIVE_EN set. The complete 1,209-entry trace contained no
> ax88179a_stop() or ax88179_change_mtu() caller.
> 
[...]
> I do not think this proves an unconditional device-side bug, though. The v6
> driver failed four times in 30 cycles after writing 0x0102, while the vendor
> driver had no failures after writing 0x0132.
I have finally understood what is happening: There is a race condition between
the controller of the AX88179A trying to set up and optimize the link and
phylink trying to configure the link on the mac-side. When a link change
is requested by phylink triggering re-configuring the PHY, the PHY is continued
to be polled by phylink. At this point, the PHY may report that the link is up
before the controller is actually finished configuring it. mac_link_up() is called
by phylink, but the controller overwrites the AX_MEDIUM_RECEIVE_EN
bit that is set by mac_link_up() when it continues with its configuration.

The solution is simple: do not poll the PHY with phylink, but wait until the
controller decides the link is completely configured, at which point an interrupt
USB-URB is sent. Then handle this interrupt in phylink in order to read the final PHY
configuration and only then call mac_link_up().

I will provide a v7 with an additional phylink function phylink_mac_interrupt()
being introduced as suggested by Andrew, which is called by ax88179a_status()
in response to usbnet receiving the link change interrupt. I tested changing
the link a couple of dozen times and it always worked, now.

Birger
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Andrew Lunn 1 month, 2 weeks ago
> I have finally understood what is happening: There is a race condition between
> the controller of the AX88179A trying to set up and optimize the link and
> phylink trying to configure the link on the mac-side.

This is why i dislike any hardware/firmware which thinks it is smarter
than Linux and directly access the hardware. Such features often break
stuff, because it does not respect the mutex Linux uses to serialise
access to the device.

> I will provide a v7 with an additional phylink function phylink_mac_interrupt()
> being introduced as suggested by Andrew, which is called by ax88179a_status()
> in response to usbnet receiving the link change interrupt. I tested changing
> the link a couple of dozen times and it always worked, now.

This will make it safer, but might not stop all the problems. Does the
PHY have LEDs? Does it have temperature sensors? Features like this
operate asynchronously to link state. The user can configure them at
any time. Such register writes might collide with what the firmware is
doing. I don't suppose you have physical access to the MDIO bus and
can put a logic analyser on it? Are there bus transactions happening
during normal operation which are not from Linux?

	Andrew
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Birger Koblitz 1 month, 2 weeks ago
On 8/10/26 15:25, Andrew Lunn wrote:
>> I have finally understood what is happening: There is a race condition between
>> the controller of the AX88179A trying to set up and optimize the link and
>> phylink trying to configure the link on the mac-side.
> 
> This is why i dislike any hardware/firmware which thinks it is smarter
> than Linux and directly access the hardware. Such features often break
> stuff, because it does not respect the mutex Linux uses to serialise
> access to the device.
The hardware by default uses CDC_NCM and even needs to be switched to the
vendor mode in order to allow any manual configuration of MAC and PHY.
As much as possible is supposed to work automagically. Linux support looks
very much like an afterthought.

> 
>> I will provide a v7 with an additional phylink function phylink_mac_interrupt()
>> being introduced as suggested by Andrew, which is called by ax88179a_status()
>> in response to usbnet receiving the link change interrupt. I tested changing
>> the link a couple of dozen times and it always worked, now.
> 
> This will make it safer, but might not stop all the problems. Does the
> PHY have LEDs? Does it have temperature sensors? Features like this
> operate asynchronously to link state. The user can configure them at
> any time. Such register writes might collide with what the firmware is
> doing. I don't suppose you have physical access to the MDIO bus and
> can put a logic analyser on it? Are there bus transactions happening
> during normal operation which are not from Linux?
> 
The device has LEDs, but they are entirely controlled by firmware. The behaviour
can be configured from the EEPROM. Writing to the EEPROM involves stopping
the firmware on the device. There are no temperature sensors, or maybe there are,
but the firmware hides them from us.

There is no physical access to the MDIO bus, as MAC and PHY are in the same
chip. I hope that otherwise we are safe from firmware interference, though.

Birger
Re: [PATCH net-next v6 00/13] ax88179_178a: Add support for AX88179A-based chips
Posted by Andrew Lunn 1 month, 2 weeks ago
> The device has LEDs, but they are entirely controlled by firmware. The behaviour
> can be configured from the EEPROM. Writing to the EEPROM involves stopping
> the firmware on the device. There are no temperature sensors, or maybe there are,
> but the firmware hides them from us.

That keeps thinks simpler.

> There is no physical access to the MDIO bus, as MAC and PHY are in the same
> chip. I hope that otherwise we are safe from firmware interference, though.

It is the sort of thing which you find out in time. As you have seen,
there can be race conditions, and there could be issues hiding until
specific timing is hit to trigger them. But there is little we can do
about that, so lets just go with what you have.

      Andrew