RE: [RFC 0/1] Forward AER errors to guest

Shameer Kolothum Thodi posted 1 patch 2 weeks, 4 days ago
Only 0 patches received!
RE: [RFC 0/1] Forward AER errors to guest
Posted by Shameer Kolothum Thodi 2 weeks, 4 days ago

> -----Original Message-----
> From: Alex Williamson <alex.williamson@nvidia.com>
> Sent: 07 July 2026 23:13
> To: Michał Winiarski <michal.winiarski@intel.com>; Shameer Kolothum Thodi
> <skolothumtho@nvidia.com>
> Cc: Satyanarayana K V P <satyanarayana.k.v.p@intel.com>; qemu-
> devel@nongnu.org; Michal Wajdeczko <michal.wajdeczko@intel.com>;
> Matthew Brost <matthew.brost@intel.com>; Cédric Le Goater
> <clg@redhat.com>
> Subject: Re: [RFC 0/1] Forward AER errors to guest
> 
> On Tue, 7 Jul 2026 23:05:33 +0200
> Michał Winiarski <michal.winiarski@intel.com> wrote:
> 
> > On Fri, Jul 03, 2026 at 07:59:46AM -0600, Alex Williamson wrote:
> > >
> > >
> > > On Fri, Jul 3, 2026, at 5:13 AM, Satyanarayana K V P wrote:
> > > > Today, vfio-pci unconditionally stops the VM when any error event is
> > > > reported by device. This prevents guest-driven error handling and
> recovery
> > > > for platforms that support PCIe AER.
> > > >
> > > > This series adds an optional vfio-pci extension parameter,
> > > > "x-forward-aer=on", to forward AER errors to the guest instead of
> > > > forcing an immediate VM stop. If the endpoint supports AER, the error is
> > > > forwarded directly; otherwise, checks the upstream PCIe bridge and
> > > > forwards the error there when supported.If neither device nor PCIe
> bridge
> > > > supports AER, the VM is immediately stopped when the error is
> reported.
> > >
> > > The error eventfd is signaled for err_detected in the host, this is
> > > the beginning of the host error handling and the point at which
> > > drivers should stop accessing the device until the resume callback.
> > >  Letting the VMM continue at this point does the opposite of that.
> > > Thanks,
> > >
> > > Alex
> >
> > Hi,
> >
> > Today - the default error handler used by VFIO is not implementing the
> > .resume() callback and just signals err_trigger and returns
> > PCI_ERS_RESULT_CAN_RECOVER as part of .error_detected().
> > Qemu registers a callback to be called upon err_trigger signal (using
> > one of the poll/select syscall variants), the callback calls
> > vm_stop(RUN_STATE_INTERNAL_ERROR) - this is a runstate that can't
> > continue, the VM is effectively killed (qemu needs to trigger VM reset
> > to recover).
> 
> Yes, this is the extent of the implementation currently.
> 
> > The AER forward was proposed as a relatively simple opt-in mechanism
> > to avoid hitting RUN_STATE_INTERNAL_ERROR and be able to handle
> > errors for SR-IOV VFs using VFIO. For uncorrectable errors with
> > SR-IOV, PF driver can be the entity that handles the device recovery.
> > The VF driver needs to stop using the device until it is recovered by
> > PF, and reinitialize any state that was lost as part of the recovery
> > action. Forwarding AER would allow the VM to move forward without
> > introducing any new uAPI.
> 
> Interesting, nothing in the proposal suggested this was explicitly for
> SR-IOV VFs.
> 
> > If we would want to build a generic error recovery mechanism, we would
> > need to duplicate more of the err_handler_t callbacks as VFIO uAPI.
> >
> > If the recovery is handled by other entity running on the host (VFIO
> > variant? Or for SR-IOV, the PF driver), we could also go into a
> > different runstate (RUN_STATE_PAUSED? something that doesn't require
> > full VM reset) and add a VFIO uAPI that can propagate the .resume()
> > into userspace.
> 
> And a .slot_reset() to actually restore state if PF is reset (we
> currently rely on that happening on .release()), and we'll also need to
> teardown things like interrupts (also currently reliant on release).
> 
> > If we would want the SW running inside the VM to handle the
> > recovery... it becomes more involved, and it requires a more
> > substantial solution that wasn't meant to be addressed here.
> 
> Note that the current error eventfd is triggered on any uncorrected
> error, fatal or non-fatal.  How does the VM know which it is?  There
> are certainly hacks we can do to limp along a little further with the
> current limited implementation, but nothing that seems to actually get
> us closer to a general purpose, supportable feature.
> 
> > So - we're letting the VMM continue but we're also informing it that
> > something went wrong with the device, so that the SW running inside
> > the VM can do its own action.
> > If it's a virtual function, the driver can potentially fully recover
> > after the host does its thing.
> > If it's a regular native function, it will probably need to go into
> > PCI_ERS_RESULT_DISCONNECT, which, for most usecases is still a better
> > scenario then having to reset the entire VM.
> 
> We've been discussing this internally as well.  The VM can perform its
> own recovery, resetting and re-initializing devices from a VM
> perspective, but it can't "handle" the host recovery process.  In fact,
> I think at best it could pause and observe the host process, resuming
> and injecting an error after host recovery completes.
> 
> > If we would want to absolutely make sure that the users are not able
> > to issue any IO across error recovery, we can probably unmap (or
> > zero-map) the BARs at VFIO driver level, similar to what's done
> > during reset.
> 
> Yes, the shape I imagined would immediately block all access to the
> device on .err_detected(), zap BARs, fail access, and signal the
> existing error eventfd.  The VMM would need to handle the potential
> sigbus it might see before noticing the eventfd, determine whether it
> maps to the mmio space of a device and call a new device feature ioctl
> to test whether the device is in an error state (hand waving how QEMU
> gets from the trap handler to a point where the VM is paused and we can
> poll an ioctl).  Flag bits on the GET of that feature ioctl might
> indicate Fatal/Non-Fatal, InProgress, DeviceReset, and Failed.  Perhaps
> even a sequence number and/or SET on the feature ioctl might implement
> a W1C style acknowledgment - some mechanism by which the user can track
> whether they've missed an event.  QEMU could autonomously perform a
> surprise hot-unplug on Fatal or Failed, or otherwise pause the VM for
> InProgress to clear and inject an appropriate error when it does.
> APEI/GHES handling in the guest might be particularly useful to
> indicate to the VM whether a device reset has already been performed.
> In fact, the whole process doesn't seem too dissimilar to firmware
> first error handling that runs underneath the host OS on bare metal.

As Alex mentioned, we have been discussing this internally. I was
planning to look further into the observer-based approach outlined
above, including the required vfio-pci and QEMU changes.

However, I will be on PTO for the next two weeks, so I will only be
able to pick this up after I return. By all means, please go ahead
if this is urgent for you. I can catch up with any progress when
I am back.

Thanks,
Shameer