[RFC PATCH v1 0/1] vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset

Pranjal Shrivastava posted 1 patch 1 month, 3 weeks ago
drivers/vfio/pci/vfio_pci_core.c | 88 ++++++++++++++++++++------------
1 file changed, 56 insertions(+), 32 deletions(-)
[RFC PATCH v1 0/1] vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset
Posted by Pranjal Shrivastava 1 month, 3 weeks ago
Introduce PCI .reset_prepare and .reset_done handlers to safely revoke
active userspace mappings and exported DMABUFs during sysfs-triggered
device resets.

We are seeing a situation where system health and monitoring daemons 
(at times erroneously) issue device resets via sysfs for devices bound
to vfio-pci:

  echo 1 > /sys/bus/pci/devices/0000:01:00.0/reset

However, because vfio-pci does not implement the .reset_prepare and
.reset_done error handlers, this hardware reset occurs completely unnoticed
by the VFIO driver.

Consequently, active traditional userspace BAR mappings and exported DMABUFs
are never zapped or revoked. Importers of the DMABUFs (e.g., RDMA drivers)
continue to issue DMAs (such as PCIe Memory Writes) toward the Endpoint. 
These transactions are silently dropped by the root port or trigger CTOs 
while higher-level actions (e.g., RDMA reg_mr) continue to succeed.

We'd like to fix this by implementing the PCI reset ops for vfio-pci
that revoke the DMABUFs and zap the BARs while holding the memory lock
allowing concurrent user accesses to sleep and fault back in once the reset
completes.

Note: I've tried to handle the locking as a first attempt here, there
might've been some cases that were missed. Also, for the RFC, the
drivers that implement their own pci_error_handlers (like nvgrace) are 
not altered for now.

Quick Note about Matt's DMABUF mmap Series
==========================================
While this patch is aimed for the current upstream code, I believe with
Matt's refactor [1] these ops might change slightly. If we have consensus
on this patch, I'd send another patch based to Matt based on their series
for them to include it in their next version.

[1] https://lore.kernel.org/all/20260715174737.15287-1-matt@ozlabs.org/

Thanks,
Praan

Pranjal Shrivastava (1):
  vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset

 drivers/vfio/pci/vfio_pci_core.c | 88 ++++++++++++++++++++------------
 1 file changed, 56 insertions(+), 32 deletions(-)

-- 
2.55.0.679.g6767b8d81c-goog
Re: [RFC PATCH v1 0/1] vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset
Posted by Alex Williamson 1 month, 2 weeks ago
On Fri,  7 Aug 2026 20:14:04 +0000
Pranjal Shrivastava <praan@google.com> wrote:

> Introduce PCI .reset_prepare and .reset_done handlers to safely revoke
> active userspace mappings and exported DMABUFs during sysfs-triggered
> device resets.
> 
> We are seeing a situation where system health and monitoring daemons 
> (at times erroneously) issue device resets via sysfs for devices bound
> to vfio-pci:
> 
>   echo 1 > /sys/bus/pci/devices/0000:01:00.0/reset
> 
> However, because vfio-pci does not implement the .reset_prepare and
> .reset_done error handlers, this hardware reset occurs completely unnoticed
> by the VFIO driver.
> 
> Consequently, active traditional userspace BAR mappings and exported DMABUFs
> are never zapped or revoked. Importers of the DMABUFs (e.g., RDMA drivers)
> continue to issue DMAs (such as PCIe Memory Writes) toward the Endpoint. 
> These transactions are silently dropped by the root port or trigger CTOs 
> while higher-level actions (e.g., RDMA reg_mr) continue to succeed.
> 
> We'd like to fix this by implementing the PCI reset ops for vfio-pci
> that revoke the DMABUFs and zap the BARs while holding the memory lock
> allowing concurrent user accesses to sleep and fault back in once the reset
> completes.

That sounds like a nice, serene solution, but that's not actually what
happens.  Due to the write vs read memory_lock semaphore, CPU faults
are stalled.  On the other hand, DMA mappings via IOMMUFD/dmabuf are
lost.  They require the userspace driver to be involved to perform the
unmap/remap.

Potentially this is all better than letting the device generate a
machine check as it's still trying to run across the reset, but let's
not pretend this is just a hiccup for the device that will continue
running after the rogue reset.  Thanks,

Alex
 
> Note: I've tried to handle the locking as a first attempt here, there
> might've been some cases that were missed. Also, for the RFC, the
> drivers that implement their own pci_error_handlers (like nvgrace) are 
> not altered for now.
> 
> Quick Note about Matt's DMABUF mmap Series
> ==========================================
> While this patch is aimed for the current upstream code, I believe with
> Matt's refactor [1] these ops might change slightly. If we have consensus
> on this patch, I'd send another patch based to Matt based on their series
> for them to include it in their next version.
> 
> [1] https://lore.kernel.org/all/20260715174737.15287-1-matt@ozlabs.org/
> 
> Thanks,
> Praan
> 
> Pranjal Shrivastava (1):
>   vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset
> 
>  drivers/vfio/pci/vfio_pci_core.c | 88 ++++++++++++++++++++------------
>  1 file changed, 56 insertions(+), 32 deletions(-)
>
Re: [RFC PATCH v1 0/1] vfio/pci: Revoke BARs and DMABUFs during sysfs-triggered PCI reset
Posted by Pranjal Shrivastava 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 10:03:52AM -0600, Alex Williamson wrote:

Hi Alex,

> On Fri,  7 Aug 2026 20:14:04 +0000
> Pranjal Shrivastava <praan@google.com> wrote:
> 
> > Introduce PCI .reset_prepare and .reset_done handlers to safely revoke
> > active userspace mappings and exported DMABUFs during sysfs-triggered
> > device resets.
> > 
> > We are seeing a situation where system health and monitoring daemons 
> > (at times erroneously) issue device resets via sysfs for devices bound
> > to vfio-pci:
> > 
> >   echo 1 > /sys/bus/pci/devices/0000:01:00.0/reset
> > 
> > However, because vfio-pci does not implement the .reset_prepare and
> > .reset_done error handlers, this hardware reset occurs completely unnoticed
> > by the VFIO driver.
> > 
> > Consequently, active traditional userspace BAR mappings and exported DMABUFs
> > are never zapped or revoked. Importers of the DMABUFs (e.g., RDMA drivers)
> > continue to issue DMAs (such as PCIe Memory Writes) toward the Endpoint. 
> > These transactions are silently dropped by the root port or trigger CTOs 
> > while higher-level actions (e.g., RDMA reg_mr) continue to succeed.
> > 
> > We'd like to fix this by implementing the PCI reset ops for vfio-pci
> > that revoke the DMABUFs and zap the BARs while holding the memory lock
> > allowing concurrent user accesses to sleep and fault back in once the reset
> > completes.
> 
> That sounds like a nice, serene solution, but that's not actually what
> happens.  Due to the write vs read memory_lock semaphore, CPU faults
> are stalled.  On the other hand, DMA mappings via IOMMUFD/dmabuf are
> lost.  They require the userspace driver to be involved to perform the
> unmap/remap.
> 
> Potentially this is all better than letting the device generate a
> machine check as it's still trying to run across the reset, but let's
> not pretend this is just a hiccup for the device that will continue
> running after the rogue reset.  Thanks,
> 

I tend to agree. My intention is definitely not to pretend this is a seamless
hiccup or allow the device/user to carry on as if nothing happened. In fact,
the very problem today with exported DMABUFs is that the user/importer *does*
silently continue to register DMABUFs with the RDMA subsystem and attempts
issuing DMAs to a reset device because nothing told them the state was gone.

I don't mind permanently tearing down the CPU mappings as well along with
revoking the DMABUFs. That way, we fail loudly and force userspace to unmap
and re-initialize (with a dev_warn() explaining that an out-of-band reset
occurred).

What do you think about that approach?

Thanks,
Praan