drivers/net/dsa/b53/b53_common.c | 118 ++++++++++++++++++++++++++----- include/net/dsa.h | 3 + net/dsa/port.c | 21 ++++-- net/dsa/user.c | 4 +- 4 files changed, 120 insertions(+), 26 deletions(-)
Since v5.15, a standalone port on a b53 switch cannot receive its own
tagged traffic: the switch VID lookup is always active, an 8021q
upper's VID never reaches the VLAN table, and every tagged frame
resolves to an empty member set and is discarded. The common victim is
a VLAN-tagged PPPoE WAN, where the PADI goes out and the tagged PADO
never reaches the CPU.
My first attempt disabled the VLAN table while not filtering:
https://lore.kernel.org/all/20260805072641.402-1-strst.gs@gmail.com/
Jonas pointed out that this moves the ARL to shared VLAN learning and
desynchronizes the hardware table from the bridge fdb, and I withdrew
it. I then measured the alternatives on an RT-N18U (BCM53011 rev 5),
with the outbound direction of the same link as a positive control on
every run:
- With the table enabled, no ingress VID check setting delivers the
frame: VC4_NO_ING_VID_CHK, VC4_ING_VID_VIO_FWD and
VC4_ING_VID_VIO_TO_IMP all give 0, and clearing VC0_DROP_VID_MISS
changes nothing. The frame does not die at ingress admission, it
dies when forwarding resolves the VID against an empty member set.
- With the table disabled, a static fdb entry with VID 100 is lost
from the hardware ARL no matter how the driver drives the ARL
registers: keeping ARLTBL_IVL_SVL_SELECT at IVL does not preserve
it, and neither does additionally keeping the VID learning bits in
VLAN_CTRL0 set.
So on this hardware, delivering the frame and keeping VID-keyed ARL
entries are mutually exclusive unless the VID is in the table. This
series therefore programs the table, narrowed to what is actually
needed: a standalone port only needs the VIDs its 8021q uppers use,
which is one table write per upper instead of entries for all 4096
VIDs.
I tried to keep the fix inside b53, but the driver cannot solve this
alone: without NETIF_F_HW_VLAN_CTAG_FILTER the 8021q layer never calls
.ndo_vlan_rx_add_vid, so the VIDs never reach the driver, and DSA
manages that feature bit. The one existing way to get it,
ds->needs_standalone_vlan_filtering, does not work here. It was
measured insufficient, because f089652b6b16 makes .port_vlan_add skip
the hardware write while not filtering, and its other effect is one
b53 cannot take: with vlan_filtering_is_global, the forced
vlan_filtering=1 in dsa_port_reset_vlan_filtering() would flip the
whole switch into VLAN filtering when any port leaves a VLAN-unaware
bridge. hellcreek relies on exactly those semantics, so patch 1 adds a
narrower opt-in that only delivers the VIDs and leaves vlan_filtering
alone, and patch 2 uses it in b53 and programs entries that carry
standalone members, masked so bridge VLANs stay without effect while
not filtering.
Tested on the RT-N18U: the standalone upper receives 7 of 7 probe
frames with vlan_filtering staying 0, the static fdb entry with a VID
now survives a vlan_filtering toggle since the table and the ARL mode
are never touched, uppers keep working across bridge join and leave
and across a vlan_filtering toggle including on ports that were
bridged while the toggle happened, deleting an upper or bridging its
port verifiably stops delivery of that VID to the CPU, and the PPPoE
session from the original report establishes. 802.1ad uppers keep
working as software VLANs, since this switch does not parse 0x88a8,
and stacked QinQ over an offloaded upper works too.
Semih Baskan (2):
net: dsa: let drivers offload 8021q uppers on standalone ports
net: dsa: b53: offload 8021q uppers on standalone ports
drivers/net/dsa/b53/b53_common.c | 118 ++++++++++++++++++++++++++-----
include/net/dsa.h | 3 +
net/dsa/port.c | 21 ++++--
net/dsa/user.c | 4 +-
4 files changed, 120 insertions(+), 26 deletions(-)
Hi, On Thu, Aug 6, 2026 at 9:31 AM Semih Baskan <strst.gs@gmail.com> wrote: > > Since v5.15, a standalone port on a b53 switch cannot receive its own > tagged traffic: the switch VID lookup is always active, an 8021q > upper's VID never reaches the VLAN table, and every tagged frame > resolves to an empty member set and is discarded. The common victim is > a VLAN-tagged PPPoE WAN, where the PADI goes out and the tagged PADO > never reaches the CPU. > > My first attempt disabled the VLAN table while not filtering: > > https://lore.kernel.org/all/20260805072641.402-1-strst.gs@gmail.com/ > > Jonas pointed out that this moves the ARL to shared VLAN learning and > desynchronizes the hardware table from the bridge fdb, and I withdrew > it. I then measured the alternatives on an RT-N18U (BCM53011 rev 5), > with the outbound direction of the same link as a positive control on > every run: > > - With the table enabled, no ingress VID check setting delivers the > frame: VC4_NO_ING_VID_CHK, VC4_ING_VID_VIO_FWD and > VC4_ING_VID_VIO_TO_IMP all give 0, and clearing VC0_DROP_VID_MISS > changes nothing. The frame does not die at ingress admission, it > dies when forwarding resolves the VID against an empty member set. Note that this is working on switches other than bcm5301x (at least on bcm63268 and bcm53115), so this seems to be a bcm5301x specific issue. Unfortunately I do not have a device with such a switch. Though it only works for standalone ports, it does not allow forwarding between ports. Though this isn't the first time a bcm5310x issue showed up with packets not properly trapped to CPU. Rafal, Florian, did you ever figure out the issue? It feels like there is something missing with the CPU port configuration. Which port are you using as CPU port? I see several device trees in-tree using port 5, but according to the register definitions in OpenMDK, the only valid port for both BRCM_HDR and GLOBAL_CONFIG's FRM_MGMT_PORT is imp0 / port 8 [1]. So I now wonder if using port 5 as CPU port only appears to work (i.e. enabling the header does), but anything that is supposed to trap to CPU tries to forward to 8, which is disabled. Or does not forward at all, because the FRM_MGMT_PORT is configured to an invalid value. In addition to that, I see that b53_brcm_hdr_setup() does not clear GC_FRM_MGMT_PORT_M, so if it defaults/was programmed to anything before, it may become a wrong value. If you are using port 5 as CPU port, can you try switching to port 8 / gmac2? > - With the table disabled, a static fdb entry with VID 100 is lost > from the hardware ARL no matter how the driver drives the ARL > registers: keeping ARLTBL_IVL_SVL_SELECT at IVL does not preserve > it, and neither does additionally keeping the VID learning bits in > VLAN_CTRL0 set. ARLTBL_IVL_SVL_SELECT is only implemented on bcm5302x / bcm58xx / bcm53134 (and maybe some other newer switches), so no wonder it doesn't do anything for you (it also has some additional dependencies which aren't implemented in b53, so this is essentially dead code). [1] https://github.com/Broadcom/OpenMDK/blob/v2.11.0/cdk/PKG/chip/bcm53010/bcm53010_a0_defs.h Best regards, Jonas
Hi Jonas,
> If you are using port 5 as CPU port, can you try switching to port 8 / gmac2?
You are right. I tested it today on the RT-N18U and switching the CPU
port to port 8 makes tagged standalone RX work on an unpatched driver,
with nothing in the VLAN table. Details below.
User ports attach to port 5 (gmac0) here, as in every bcm5301x device
tree in-tree. I enabled port@8 (gmac2) as the only CPU port in the dts,
left the driver completely stock, and rebuilt:
- the board comes up with the whole management path over gmac2, so
the port 8 to gmac2 path works on BCM47081,
- the tagged probe that always failed on port 5 delivers 7 of 7
(outbound positive control 7 of 7, delivery attributed to eth2 by
interface counters, eth0 stayed at 0),
- the PPPoE session from the original report establishes.
I also instrumented b53_brcm_hdr_setup() to read back GLOBAL_CONFIG,
and the register documentation in OpenMDK explains the mechanism you
suspected. GMNGCFG.FRM_MNGP on bcm53010 [1]:
00 = no IMP port
01 = reserved
10 = IMP0 only: all traffic to CPU from LAN and WAN ports goes to IMP0
11 = dual IMP: LAN-port CPU traffic goes to IMP0, WAN-port traffic to
IMP1, and "In polar, IMP0 is Port 8 and IMP1 is Port 5."
With CPU port 5 the driver ORs GC_FRM_MGMT_PORT_M, which is the field
mask, so the register reads back 0xc2: FRM_MNGP=11, dual IMP. In that
mode the tagged frames reach the CPU on neither IMP: I also tested an
intermediate build with port@8 enabled as a second CPU port while the
user ports stayed on port 5, which is what plain bcm-ns.dtsi describes
since it does not disable port@7/8, and delivery still failed with the
gmac2 counters at zero. With port 8 as the only CPU port the driver
programs 0x82: FRM_MNGP=10, IMP0 only, and delivery works. So your
no-clear observation is correct and the value it produces is worse
than a stale leftover: the port 5 branch cannot program anything
better, because there is no "IMP1 only" encoding, and LAN-class
management traps can never arrive on port 5 on this chip.
One smaller register note: BRCM_HDR_CTRL on bcm53010 does have per-port
bits (bit0 port 8, bit1 port 5, bit2 port 7) [1], so the Broadcom header
itself is valid on port 5. That is why tagging works there at all; the
port-8-only limitation is in the IMP routing, not the header.
A data point that fits the same picture: with CPU on port 5, BPDUs sent
into switch port 0 do reach the CPU. Port 0 is WAN-class in the chip's
management routing, so its traps go to IMP1, port 5. The trap
destination depends on the port class, which is probably why this
half-works and has been so confusing historically.
> ARLTBL_IVL_SVL_SELECT is only implemented on bcm5302x / bcm58xx /
> bcm53134 (and maybe some other newer switches), so no wonder it
> doesn't do anything for you
Thanks, noted. That makes the second measurement stronger rather than
weaker: on this chip there is no knob at all that preserves VID-keyed
ARL entries with the table disabled.
One property of the port 8 setup worth knowing before anyone reads it
as the fix: delivery is indiscriminate. After deleting the 8021q upper
I still see the tagged frames of that VID on the CPU with tcpdump.
Every unknown VID from the wire reaches the CPU, always, which is the
pre-5.15 behavior with its unfiltered nature included. The series
delivers only the VIDs that uppers subscribe, and stops delivering when
they go away.
So as I see it there are two valid fixes on different timescales.
Moving bcm5301x device trees to gmac2 fixes the trap path at the root,
but it changes the conduit for every board, needs per-board validation,
and cannot go to stable. The series fixes the deployed port 5 topology
selectively and is backportable. They do not conflict; the VLAN entries
are correct and harmless under either CPU port. I am happy to help test
a device tree migration on the RT-N18U if you want to pursue that
separately.
Also for completeness: a runtime test of your suggestion via conduit
reassignment is not possible, b53 does not implement
port_change_conduit, so I tested through the device tree.
[1] https://github.com/Broadcom/OpenMDK/blob/v2.11.0/cdk/PKG/chip/bcm53010/bcm53010_a0_defs.h
Best regards,
Semih
© 2016 - 2026 Red Hat, Inc.