drivers/usb/host/xhci-ring.c | 200 ++++++++++++++++++----------------- 1 file changed, 102 insertions(+), 98 deletions(-)
Hi, This series is motivated by a recently found rare regression due to my commit from last year and the solution suggested by Mathias Nyman. https://lore.kernel.org/linux-usb/al_hchyOdPoPWKEo@spiral/ I think it's a good solution not only for this specific case, but also in general, because the next event after Missed Service Error almost always references some TD - exceptions are Ring Underrun, which we have special handling for, and Stopped - Length Invalid, which would be a rare occurrence, not currently supported anyway, and possible to support within the proposed framework by making find_td_by_dma() calculate accurate 'missed_tds' value while still returning NULL. The first two patches fix bugs, because I found another obscure one. The next two patches prepare for the last one by simplifying things. The last patch implements the big change, the whole matching/skipping loop is replaced with a more straightforward and robust version. New functionality is paid for with a net increase of 4 LOC, not too bad. I gave this a bit of testing and it seems to be working, including weird cases like: Missed Service Error retires a waiting TD with error_mid_td, then skipping is triggered by another Transaction Error immediately afterwards, and it turns out that two TDs were missed. [ 1665.224921] xhci_hcd 0000:0a:00.0: Transfer error for slot 1 ep 2 on endpoint [ 1665.225151] xhci_hcd 0000:0a:00.0: Missed Service Error for slot 1 ep 2, skip 1, try now 0 [ 1665.225156] xhci_hcd 0000:0a:00.0: Missing TD completion event after mid TD error [ 1665.225305] xhci_hcd 0000:0a:00.0: Transfer error for slot 1 ep 2 on endpoint [ 1665.225308] xhci_hcd 0000:0a:00.0: Skipped 2 TDs on slot 1 ep 2 comp_code 4, TD found 1, skip flag 0 Additional testing of patch 1 in isolation would be appreciated from the reporter of the regression (Cc). That patch would go to v6.18. Regards, Michal Michal Pecio (5): usb: xhci: Handle bogus TRB pointers in Missed Service Error events usb: xhci: Guarantee URB giveback on Ring Underrun/Overrun usb: xhci: Don't set the skip flag on non-isoc endpoints usb: xhci: Shorten the TD skipping loop usb: xhci: Rework and improve the TD matching and skipping logic drivers/usb/host/xhci-ring.c | 200 ++++++++++++++++++----------------- 1 file changed, 102 insertions(+), 98 deletions(-) -- 2.48.1
At 2026-08-04 12:01:10 +0200, Michal Pecio wrote: ... > Additional testing of patch 1 in isolation would be appreciated from > the reporter of the regression (Cc). That patch would go to v6.18. That's me! I'm new here so please forgive me if I'm asking stupid questions: I found the patch applies cleanly to the linux-6.18.y branch in the stable repo; am I in the right place? And then am I looking for anything in particular in the logs while running it, or do you simply want an observation on whether I can reproduce the bug or not? Thanks.
On Wed, 5 Aug 2026 12:06:48 -0700, Bart Nagel wrote: > At 2026-08-04 12:01:10 +0200, Michal Pecio wrote: > ... > > Additional testing of patch 1 in isolation would be appreciated from > > the reporter of the regression (Cc). That patch would go to v6.18. > > That's me! I'm new here so please forgive me if I'm asking stupid > questions: > > I found the patch applies cleanly to the linux-6.18.y branch in the > stable repo; am I in the right place? That's fine, there were hardly any changes in this area this year so 6.18 should be about equivalent to any 7.x. And it's the only one of them which will stay supported (here) for a few years. > And then am I looking for anything in particular in the logs while > running it, or do you simply want an observation on whether I can > reproduce the bug or not? At this point just see if it works. These patches don't affect HW behavior, they only ignore bogus events to avoid creating problems. And this patch should ignore every case ignored by the old one. Regard, Michal
At 2026-08-05 22:21:20 +0200, Michal Pecio wrote: > On Wed, 5 Aug 2026 12:06:48 -0700, Bart Nagel wrote: > > At 2026-08-04 12:01:10 +0200, Michal Pecio wrote: > > ... > > > Additional testing of patch 1 in isolation would be appreciated from > > > the reporter of the regression (Cc). That patch would go to v6.18. > > > > That's me! I'm new here so please forgive me if I'm asking stupid > > questions: > > > > I found the patch applies cleanly to the linux-6.18.y branch in the > > stable repo; am I in the right place? > > That's fine, there were hardly any changes in this area this year > so 6.18 should be about equivalent to any 7.x. And it's the only one > of them which will stay supported (here) for a few years. > > > And then am I looking for anything in particular in the logs while > > running it, or do you simply want an observation on whether I can > > reproduce the bug or not? > > At this point just see if it works. These patches don't affect HW > behavior, they only ignore bogus events to avoid creating problems. > And this patch should ignore every case ignored by the old one. OK, I ran it with high system load (which as discussed earlier made the bug pop up more readily) for a little over 12 hours, and no freeze. In case logs are interesting, they are temporarily (~48h) here: dmesg log with handle_tx_event, but with "spurious event" messages filtered out otherwise the file is 400 times the size: https://tmpfiles.org/wvw5Fxo00DQt/dmesg-6.18.42-patch-2026-08-04-no-spurious.txt Filtered debugfs trace as directed by Michal earlier in the process: https://tmpfiles.org/w2wTFcoQ02zV/trace-6.18.42-patch-2026-08-04.txt Thanks for all your work.
© 2016 - 2026 Red Hat, Inc.