> -----Original Message----- > From: Alex Williamson <alex.williamson@nvidia.com> > Sent: 07 July 2026 23:13 > To: Michał Winiarski <michal.winiarski@intel.com>; Shameer Kolothum Thodi > <skolothumtho@nvidia.com> > Cc: Satyanarayana K V P <satyanarayana.k.v.p@intel.com>; qemu- > devel@nongnu.org; Michal Wajdeczko <michal.wajdeczko@intel.com>; > Matthew Brost <matthew.brost@intel.com>; Cédric Le Goater > <clg@redhat.com> > Subject: Re: [RFC 0/1] Forward AER errors to guest > > On Tue, 7 Jul 2026 23:05:33 +0200 > Michał Winiarski <michal.winiarski@intel.com> wrote: > > > On Fri, Jul 03, 2026 at 07:59:46AM -0600, Alex Williamson wrote: > > > > > > > > > On Fri, Jul 3, 2026, at 5:13 AM, Satyanarayana K V P wrote: > > > > Today, vfio-pci unconditionally stops the VM when any error event is > > > > reported by device. This prevents guest-driven error handling and > recovery > > > > for platforms that support PCIe AER. > > > > > > > > This series adds an optional vfio-pci extension parameter, > > > > "x-forward-aer=on", to forward AER errors to the guest instead of > > > > forcing an immediate VM stop. If the endpoint supports AER, the error is > > > > forwarded directly; otherwise, checks the upstream PCIe bridge and > > > > forwards the error there when supported.If neither device nor PCIe > bridge > > > > supports AER, the VM is immediately stopped when the error is > reported. > > > > > > The error eventfd is signaled for err_detected in the host, this is > > > the beginning of the host error handling and the point at which > > > drivers should stop accessing the device until the resume callback. > > > Letting the VMM continue at this point does the opposite of that. > > > Thanks, > > > > > > Alex > > > > Hi, > > > > Today - the default error handler used by VFIO is not implementing the > > .resume() callback and just signals err_trigger and returns > > PCI_ERS_RESULT_CAN_RECOVER as part of .error_detected(). > > Qemu registers a callback to be called upon err_trigger signal (using > > one of the poll/select syscall variants), the callback calls > > vm_stop(RUN_STATE_INTERNAL_ERROR) - this is a runstate that can't > > continue, the VM is effectively killed (qemu needs to trigger VM reset > > to recover). > > Yes, this is the extent of the implementation currently. > > > The AER forward was proposed as a relatively simple opt-in mechanism > > to avoid hitting RUN_STATE_INTERNAL_ERROR and be able to handle > > errors for SR-IOV VFs using VFIO. For uncorrectable errors with > > SR-IOV, PF driver can be the entity that handles the device recovery. > > The VF driver needs to stop using the device until it is recovered by > > PF, and reinitialize any state that was lost as part of the recovery > > action. Forwarding AER would allow the VM to move forward without > > introducing any new uAPI. > > Interesting, nothing in the proposal suggested this was explicitly for > SR-IOV VFs. > > > If we would want to build a generic error recovery mechanism, we would > > need to duplicate more of the err_handler_t callbacks as VFIO uAPI. > > > > If the recovery is handled by other entity running on the host (VFIO > > variant? Or for SR-IOV, the PF driver), we could also go into a > > different runstate (RUN_STATE_PAUSED? something that doesn't require > > full VM reset) and add a VFIO uAPI that can propagate the .resume() > > into userspace. > > And a .slot_reset() to actually restore state if PF is reset (we > currently rely on that happening on .release()), and we'll also need to > teardown things like interrupts (also currently reliant on release). > > > If we would want the SW running inside the VM to handle the > > recovery... it becomes more involved, and it requires a more > > substantial solution that wasn't meant to be addressed here. > > Note that the current error eventfd is triggered on any uncorrected > error, fatal or non-fatal. How does the VM know which it is? There > are certainly hacks we can do to limp along a little further with the > current limited implementation, but nothing that seems to actually get > us closer to a general purpose, supportable feature. > > > So - we're letting the VMM continue but we're also informing it that > > something went wrong with the device, so that the SW running inside > > the VM can do its own action. > > If it's a virtual function, the driver can potentially fully recover > > after the host does its thing. > > If it's a regular native function, it will probably need to go into > > PCI_ERS_RESULT_DISCONNECT, which, for most usecases is still a better > > scenario then having to reset the entire VM. > > We've been discussing this internally as well. The VM can perform its > own recovery, resetting and re-initializing devices from a VM > perspective, but it can't "handle" the host recovery process. In fact, > I think at best it could pause and observe the host process, resuming > and injecting an error after host recovery completes. > > > If we would want to absolutely make sure that the users are not able > > to issue any IO across error recovery, we can probably unmap (or > > zero-map) the BARs at VFIO driver level, similar to what's done > > during reset. > > Yes, the shape I imagined would immediately block all access to the > device on .err_detected(), zap BARs, fail access, and signal the > existing error eventfd. The VMM would need to handle the potential > sigbus it might see before noticing the eventfd, determine whether it > maps to the mmio space of a device and call a new device feature ioctl > to test whether the device is in an error state (hand waving how QEMU > gets from the trap handler to a point where the VM is paused and we can > poll an ioctl). Flag bits on the GET of that feature ioctl might > indicate Fatal/Non-Fatal, InProgress, DeviceReset, and Failed. Perhaps > even a sequence number and/or SET on the feature ioctl might implement > a W1C style acknowledgment - some mechanism by which the user can track > whether they've missed an event. QEMU could autonomously perform a > surprise hot-unplug on Fatal or Failed, or otherwise pause the VM for > InProgress to clear and inject an appropriate error when it does. > APEI/GHES handling in the guest might be particularly useful to > indicate to the VM whether a device reset has already been performed. > In fact, the whole process doesn't seem too dissimilar to firmware > first error handling that runs underneath the host OS on bare metal. As Alex mentioned, we have been discussing this internally. I was planning to look further into the observer-based approach outlined above, including the required vfio-pci and QEMU changes. However, I will be on PTO for the next two weeks, so I will only be able to pick this up after I return. By all means, please go ahead if this is urgent for you. I can catch up with any progress when I am back. Thanks, Shameer
© 2016 - 2026 Red Hat, Inc.