[RFC PATCH 00/12] virtio: support devices that own their virtqueue memory

Alexander Graf posted 12 patches 1 month, 2 weeks ago
There is a newer version of this series
Documentation/driver-api/virtio/index.rst     |    1 +
.../driver-api/virtio/virtio-dmb.rst          |  803 ++++++++
drivers/vdpa/vdpa.c                           |    2 +-
drivers/virtio/Kconfig                        |   29 +
drivers/virtio/Makefile                       |    3 +-
drivers/virtio/virtio.c                       |   15 +-
drivers/virtio/virtio_dmb.c                   | 1719 +++++++++++++++++
drivers/virtio/virtio_dmb.h                   |   34 +
drivers/virtio/virtio_dmb_test.c              |  279 +++
drivers/virtio/virtio_pci_modern.c            |  100 +-
drivers/virtio/virtio_pci_modern_dev.c        |   23 +-
drivers/virtio/virtio_ring.c                  |  216 ++-
include/linux/virtio.h                        |    3 +
include/linux/virtio_config.h                 |   41 +
include/linux/virtio_pci_modern.h             |    1 +
include/uapi/linux/virtio_config.h            |   17 +-
include/uapi/linux/virtio_pci.h               |   10 +
17 files changed, 3254 insertions(+), 42 deletions(-)
create mode 100644 Documentation/driver-api/virtio/virtio-dmb.rst
create mode 100644 drivers/virtio/virtio_dmb.c
create mode 100644 drivers/virtio/virtio_dmb.h
create mode 100644 drivers/virtio/virtio_dmb_test.c
[RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Alexander Graf 1 month, 2 weeks ago
Virtio drivers use guest memory to back virtqueues and their buffers.
That means a VMM needs to be able to map guest memory. That is ok in the
normal virt case. It gets icky with confidential computing (where we use
swiotlb as workaround) and it defeats the purpose of isolated vhost-user
backing devices, because they end up with full RAM access to the guest.

So instead, I'm proposing an extension to virtio which allows it to give
each virtio device its own dedicated memory region to communicate with the
host, called DMB (Device Memory Buffer). A trusted hypervisor can force
DMB to be present, which then enables safer, more isolated and resilient
communication between guest and host.

With DMB, the device provides a shared memory region that both parties
agree is the full memory map both have access to. All memory offsets
that previously would have been into guest RAM, are then offsets into
this shared memory buffer region. One nice property of this is that it
is a generic mechanism in the virtio transport layer, so higher level
drivers work unmodified.

I was exploring to use swiotlb instead to create individual pools. But
that approach has multiple downsides:

1. Swiotlb is an OS primitive which is not available in all Operating
Systems. DMB however lives in the virtio transport layer, which means we
can add support for it in any OS independent of generic layers. This
helps with Windows support.

2. We munge DMA space together. DMB provides a separate DMA space per
virtio device. This means we can for example implement a device in
vhost-user and give the implementing process only visibility to the DMB
region, not all of guest memory. That reduces the exposure the
vhost-user provider has, improving security.

3. Devices can opt-in. A hypervisor can choose to use standard virtio
semantics for self-implemented devices (e.g. NSM), while requiring DMB
for devices implemented by less trustworthy providers. The
non-trustworthy devices do not get any visibility into the trustworthy
ones, even with DMB in place for both.

== Limitations ==

  - Only PCI is wired up.
  - Feature bit 44 and the shared memory id register at offset 0x40 of the
    PCI common configuration are provisional: the OASIS technical
    committee has the specification and has allocated neither.

  https://lore.kernel.org/virtio-comment/20260804161202.38619-1-graf@amazon.com/

  - The device I ran this against is not public, so you cannot reproduce
    the numbers below. The KUnit test you can.

== Testing ==

Without patches 10 and 12 a receive refill livelocks and queues starve
each other: 29,319,791 receive softirqs in five seconds with not one
packet received, against 4 with them, and 2 of 7 receive queues that never
see a buffer, against none. Earlier revisions moved 512 MiB of O_DIRECT
block I/O and 2.7 GB of verified vsock through a region with no error.

The KUnit test for patch 12's guarantee you can run yourself, and two of
its five cases fail if I take the fix out:

  tools/testing/kunit/kunit.py run --arch=x86_64 \
      --kconfig_add CONFIG_VIRTIO_MMIO=y \
      --kconfig_add CONFIG_VIRTIO_DMB=y virtio_dmb

Patch 1 can be taken on its own: a stale worked example in a vdpa
comment. Patch 10 fixes something older than this series too, a failed
mapping arriving as -EIO from a packed ring, failing an I/O that on a
split ring is only back-pressure, but it does not apply alone: it wants
patch 2, and patch 9 for the file its documentation hunk edits. Both
carry Fixes:. Patch 12 wants 10 first.

I wrote this series with an AI coding assistant, which drafted the code,
the changelogs and this cover letter. I reviewed and reworked all of it,
and every commit carries an Assisted-by: trailer.

Alex

Alexander Graf (12):
  vdpa: correct the VIRTIO_DEVICE_F_MASK example value
  virtio_ring: validate premapped addresses through the device's map
  virtio: add the VIRTIO_F_DMB feature bit
  virtio_pci: read the device memory buffer shared memory id
  virtio_pci: create virtqueues with the device's mapping token
  virtio: add a device memory buffer region allocator
  virtio: locate the device memory buffer after feature negotiation
  virtio_pci: support VIRTIO_F_DMB
  Documentation: virtio: describe the device memory buffer
  virtio_ring: report a bounded pool's exhaustion as -ENOSPC
  virtio: expose device memory buffer occupancy over debugfs
  virtio: guarantee a virtqueue can publish its first descriptor chain

 Documentation/driver-api/virtio/index.rst     |    1 +
 .../driver-api/virtio/virtio-dmb.rst          |  803 ++++++++
 drivers/vdpa/vdpa.c                           |    2 +-
 drivers/virtio/Kconfig                        |   29 +
 drivers/virtio/Makefile                       |    3 +-
 drivers/virtio/virtio.c                       |   15 +-
 drivers/virtio/virtio_dmb.c                   | 1719 +++++++++++++++++
 drivers/virtio/virtio_dmb.h                   |   34 +
 drivers/virtio/virtio_dmb_test.c              |  279 +++
 drivers/virtio/virtio_pci_modern.c            |  100 +-
 drivers/virtio/virtio_pci_modern_dev.c        |   23 +-
 drivers/virtio/virtio_ring.c                  |  216 ++-
 include/linux/virtio.h                        |    3 +
 include/linux/virtio_config.h                 |   41 +
 include/linux/virtio_pci_modern.h             |    1 +
 include/uapi/linux/virtio_config.h            |   17 +-
 include/uapi/linux/virtio_pci.h               |   10 +
 17 files changed, 3254 insertions(+), 42 deletions(-)
 create mode 100644 Documentation/driver-api/virtio/virtio-dmb.rst
 create mode 100644 drivers/virtio/virtio_dmb.c
 create mode 100644 drivers/virtio/virtio_dmb.h
 create mode 100644 drivers/virtio/virtio_dmb_test.c


base-commit: fc02acf6ac0ccde0c805c2daa9148683cdd01ba8
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Michael S. Tsirkin 1 month, 2 weeks ago
On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
> Virtio drivers use guest memory to back virtqueues and their buffers.
> That means a VMM needs to be able to map guest memory. That is ok in the
> normal virt case. It gets icky with confidential computing (where we use
> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
> backing devices, because they end up with full RAM access to the guest.
> 
> So instead, I'm proposing an extension to virtio which allows it to give
> each virtio device its own dedicated memory region to communicate with the
> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
> DMB to be present, which then enables safer, more isolated and resilient
> communication between guest and host.
> 
> With DMB, the device provides a shared memory region that both parties
> agree is the full memory map both have access to. All memory offsets
> that previously would have been into guest RAM, are then offsets into
> this shared memory buffer region. One nice property of this is that it
> is a generic mechanism in the virtio transport layer, so higher level
> drivers work unmodified.
> 
> I was exploring to use swiotlb instead to create individual pools. But
> that approach has multiple downsides:
> 
> 1. Swiotlb is an OS primitive which is not available in all Operating
> Systems. DMB however lives in the virtio transport layer, which means we
> can add support for it in any OS independent of generic layers. This
> helps with Windows support.


How does it help, if you are going to put a pool in
the driver, put a pool in the driver. Maybe with virtio mem to
simplify allocation.

> 2. We munge DMA space together. DMB provides a separate DMA space per
> virtio device. This means we can for example implement a device in
> vhost-user and give the implementing process only visibility to the DMB
> region, not all of guest memory. That reduces the exposure the
> vhost-user provider has, improving security.
> 
> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
> semantics for self-implemented devices (e.g. NSM), while requiring DMB
> for devices implemented by less trustworthy providers. The
> non-trustworthy devices do not get any visibility into the trustworthy
> ones, even with DMB in place for both.



So I am not sure whether the implication is that it's purely a software
construct. But if it is, can we extend virtio iommu to
add a way to discover and enforce trust boundaries?
And maybe translate offsets to BARs, if that is desired?



It seems to be that the result would be that we don't need fiddly
special casing in virtio ring specifically, and a lot of things like
pre-mapped dma will begin to work.



> == Limitations ==
> 
>   - Only PCI is wired up.
>   - Feature bit 44 and the shared memory id register at offset 0x40 of the
>     PCI common configuration are provisional: the OASIS technical
>     committee has the specification and has allocated neither.
> 
>   https://lore.kernel.org/virtio-comment/20260804161202.38619-1-graf@amazon.com/
> 
>   - The device I ran this against is not public, so you cannot reproduce
>     the numbers below. The KUnit test you can.
> 
> == Testing ==
> 
> Without patches 10 and 12 a receive refill livelocks and queues starve
> each other: 29,319,791 receive softirqs in five seconds with not one
> packet received, against 4 with them, and 2 of 7 receive queues that never
> see a buffer, against none. Earlier revisions moved 512 MiB of O_DIRECT
> block I/O and 2.7 GB of verified vsock through a region with no error.
> 
> The KUnit test for patch 12's guarantee you can run yourself, and two of
> its five cases fail if I take the fix out:
> 
>   tools/testing/kunit/kunit.py run --arch=x86_64 \
>       --kconfig_add CONFIG_VIRTIO_MMIO=y \
>       --kconfig_add CONFIG_VIRTIO_DMB=y virtio_dmb
> 
> Patch 1 can be taken on its own: a stale worked example in a vdpa
> comment. Patch 10 fixes something older than this series too, a failed
> mapping arriving as -EIO from a packed ring, failing an I/O that on a
> split ring is only back-pressure, but it does not apply alone: it wants
> patch 2, and patch 9 for the file its documentation hunk edits. Both
> carry Fixes:. Patch 12 wants 10 first.
> 
> I wrote this series with an AI coding assistant, which drafted the code,
> the changelogs and this cover letter. I reviewed and reworked all of it,
> and every commit carries an Assisted-by: trailer.
> 
> Alex
> 
> Alexander Graf (12):
>   vdpa: correct the VIRTIO_DEVICE_F_MASK example value
>   virtio_ring: validate premapped addresses through the device's map
>   virtio: add the VIRTIO_F_DMB feature bit
>   virtio_pci: read the device memory buffer shared memory id
>   virtio_pci: create virtqueues with the device's mapping token
>   virtio: add a device memory buffer region allocator
>   virtio: locate the device memory buffer after feature negotiation
>   virtio_pci: support VIRTIO_F_DMB
>   Documentation: virtio: describe the device memory buffer
>   virtio_ring: report a bounded pool's exhaustion as -ENOSPC
>   virtio: expose device memory buffer occupancy over debugfs
>   virtio: guarantee a virtqueue can publish its first descriptor chain
> 
>  Documentation/driver-api/virtio/index.rst     |    1 +
>  .../driver-api/virtio/virtio-dmb.rst          |  803 ++++++++
>  drivers/vdpa/vdpa.c                           |    2 +-
>  drivers/virtio/Kconfig                        |   29 +
>  drivers/virtio/Makefile                       |    3 +-
>  drivers/virtio/virtio.c                       |   15 +-
>  drivers/virtio/virtio_dmb.c                   | 1719 +++++++++++++++++
>  drivers/virtio/virtio_dmb.h                   |   34 +
>  drivers/virtio/virtio_dmb_test.c              |  279 +++
>  drivers/virtio/virtio_pci_modern.c            |  100 +-
>  drivers/virtio/virtio_pci_modern_dev.c        |   23 +-
>  drivers/virtio/virtio_ring.c                  |  216 ++-
>  include/linux/virtio.h                        |    3 +
>  include/linux/virtio_config.h                 |   41 +
>  include/linux/virtio_pci_modern.h             |    1 +
>  include/uapi/linux/virtio_config.h            |   17 +-
>  include/uapi/linux/virtio_pci.h               |   10 +
>  17 files changed, 3254 insertions(+), 42 deletions(-)
>  create mode 100644 Documentation/driver-api/virtio/virtio-dmb.rst
>  create mode 100644 drivers/virtio/virtio_dmb.c
>  create mode 100644 drivers/virtio/virtio_dmb.h
>  create mode 100644 drivers/virtio/virtio_dmb_test.c
> 
> 
> base-commit: fc02acf6ac0ccde0c805c2daa9148683cdd01ba8
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Graf (AWS), Alexander 1 month, 2 weeks ago
Hey Michael,

Thanks a bunch for having a detailed and quick look!

On 10.08.26 08:23, Michael S. Tsirkin wrote:
> On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
>> Virtio drivers use guest memory to back virtqueues and their buffers.
>> That means a VMM needs to be able to map guest memory. That is ok in the
>> normal virt case. It gets icky with confidential computing (where we use
>> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
>> backing devices, because they end up with full RAM access to the guest.
>>
>> So instead, I'm proposing an extension to virtio which allows it to give
>> each virtio device its own dedicated memory region to communicate with the
>> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
>> DMB to be present, which then enables safer, more isolated and resilient
>> communication between guest and host.
>>
>> With DMB, the device provides a shared memory region that both parties
>> agree is the full memory map both have access to. All memory offsets
>> that previously would have been into guest RAM, are then offsets into
>> this shared memory buffer region. One nice property of this is that it
>> is a generic mechanism in the virtio transport layer, so higher level
>> drivers work unmodified.
>>
>> I was exploring to use swiotlb instead to create individual pools. But
>> that approach has multiple downsides:
>>
>> 1. Swiotlb is an OS primitive which is not available in all Operating
>> Systems. DMB however lives in the virtio transport layer, which means we
>> can add support for it in any OS independent of generic layers. This
>> helps with Windows support.
>
> How does it help, if you are going to put a pool in
> the driver, put a pool in the driver. Maybe with virtio mem to
> simplify allocation.


I'm not sure I understand the suggestion :)


>
>> 2. We munge DMA space together. DMB provides a separate DMA space per
>> virtio device. This means we can for example implement a device in
>> vhost-user and give the implementing process only visibility to the DMB
>> region, not all of guest memory. That reduces the exposure the
>> vhost-user provider has, improving security.
>>
>> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
>> semantics for self-implemented devices (e.g. NSM), while requiring DMB
>> for devices implemented by less trustworthy providers. The
>> non-trustworthy devices do not get any visibility into the trustworthy
>> ones, even with DMB in place for both.
>
> So I am not sure whether the implication is that it's purely a software
> construct. But if it is, can we extend virtio iommu to
> add a way to discover and enforce trust boundaries?
> And maybe translate offsets to BARs, if that is desired?
>
> It seems to be that the result would be that we don't need fiddly
> special casing in virtio ring specifically, and a lot of things like
> pre-mapped dma will begin to work.


On thing I'm trying to avoid is dynamicity. Anything that dynamically 
changes visibility or needs state tracking is something that can go 
wrong. By keeping everything self-contained within the guest, I can 
reason about what is visible and what is not easily. Especially for 
confidential computing, IMHO static wins over dynamic in general.

Or did I misunderstand your suggestion?

As to pure software construct: With CXL, you can implement the exact 
same protocol on real hardware as well. All it takes is cache coherency 
of the BAR (or whatever the transport uses) region.



Alex
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Michael S. Tsirkin 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 07:39:02AM +0000, Graf (AWS), Alexander wrote:
> Hey Michael,
> 
> Thanks a bunch for having a detailed and quick look!
> 
> On 10.08.26 08:23, Michael S. Tsirkin wrote:
> > On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
> >> Virtio drivers use guest memory to back virtqueues and their buffers.
> >> That means a VMM needs to be able to map guest memory. That is ok in the
> >> normal virt case. It gets icky with confidential computing (where we use
> >> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
> >> backing devices, because they end up with full RAM access to the guest.
> >>
> >> So instead, I'm proposing an extension to virtio which allows it to give
> >> each virtio device its own dedicated memory region to communicate with the
> >> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
> >> DMB to be present, which then enables safer, more isolated and resilient
> >> communication between guest and host.
> >>
> >> With DMB, the device provides a shared memory region that both parties
> >> agree is the full memory map both have access to. All memory offsets
> >> that previously would have been into guest RAM, are then offsets into
> >> this shared memory buffer region. One nice property of this is that it
> >> is a generic mechanism in the virtio transport layer, so higher level
> >> drivers work unmodified.
> >>
> >> I was exploring to use swiotlb instead to create individual pools. But
> >> that approach has multiple downsides:
> >>
> >> 1. Swiotlb is an OS primitive which is not available in all Operating
> >> Systems. DMB however lives in the virtio transport layer, which means we
> >> can add support for it in any OS independent of generic layers. This
> >> helps with Windows support.
> >
> > How does it help, if you are going to put a pool in
> > the driver, put a pool in the driver. Maybe with virtio mem to
> > simplify allocation.
> 
> 
> I'm not sure I understand the suggestion :)
> 


I'm not sure what the problem is for windows :)

But if you want a chunk of contiguos memory that windows
does not poke at without a driver, virtio mem is that :)

> >
> >> 2. We munge DMA space together. DMB provides a separate DMA space per
> >> virtio device. This means we can for example implement a device in
> >> vhost-user and give the implementing process only visibility to the DMB
> >> region, not all of guest memory. That reduces the exposure the
> >> vhost-user provider has, improving security.
> >>
> >> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
> >> semantics for self-implemented devices (e.g. NSM), while requiring DMB
> >> for devices implemented by less trustworthy providers. The
> >> non-trustworthy devices do not get any visibility into the trustworthy
> >> ones, even with DMB in place for both.
> >
> > So I am not sure whether the implication is that it's purely a software
> > construct. But if it is, can we extend virtio iommu to
> > add a way to discover and enforce trust boundaries?
> > And maybe translate offsets to BARs, if that is desired?
> >
> > It seems to be that the result would be that we don't need fiddly
> > special casing in virtio ring specifically, and a lot of things like
> > pre-mapped dma will begin to work.
> 
> 
> On thing I'm trying to avoid is dynamicity. Anything that dynamically 
> changes visibility or needs state tracking is something that can go 
> wrong. By keeping everything self-contained within the guest, I can 
> reason about what is visible and what is not easily. Especially for 
> confidential computing, IMHO static wins over dynamic in general.
> 
> Or did I misunderstand your suggestion?


I get this part. But we can absolutely make it static.


Let's start by replicating your functionality with virtio-iommu.  So:

-device is behind virtio-iommu
-virtio-iommu tells guest "this device can only consume memory from that range,"
-and maybe: ...and translation is 1:1 with this offset

if guest does not acknowledge it will just fail?


this covers the proposal here simply by using a dedicated range
per device, right?



But look what we can easily add later: share a range
between devices, now you can move data between them with zero copies.


Isn't that better?


> As to pure software construct: With CXL, you can implement the exact 
> same protocol on real hardware as well. All it takes is cache coherency 
> of the BAR (or whatever the transport uses) region.
> 
> 
> 
> Alex

Yes this is what I thought originally, that you intend to
do it in hardware. But re-reading it, it begins to look like
that's not really the case, it's only "theoretically possible"?



-- 
MST
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Graf (AWS), Alexander 1 month, 2 weeks ago
On 10.08.26 10:04, Michael S. Tsirkin wrote:
> On Mon, Aug 10, 2026 at 07:39:02AM +0000, Graf (AWS), Alexander wrote:
>> Hey Michael,
>>
>> Thanks a bunch for having a detailed and quick look!
>>
>> On 10.08.26 08:23, Michael S. Tsirkin wrote:
>>> On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
>>>> Virtio drivers use guest memory to back virtqueues and their buffers.
>>>> That means a VMM needs to be able to map guest memory. That is ok in the
>>>> normal virt case. It gets icky with confidential computing (where we use
>>>> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
>>>> backing devices, because they end up with full RAM access to the guest.
>>>>
>>>> So instead, I'm proposing an extension to virtio which allows it to give
>>>> each virtio device its own dedicated memory region to communicate with the
>>>> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
>>>> DMB to be present, which then enables safer, more isolated and resilient
>>>> communication between guest and host.
>>>>
>>>> With DMB, the device provides a shared memory region that both parties
>>>> agree is the full memory map both have access to. All memory offsets
>>>> that previously would have been into guest RAM, are then offsets into
>>>> this shared memory buffer region. One nice property of this is that it
>>>> is a generic mechanism in the virtio transport layer, so higher level
>>>> drivers work unmodified.
>>>>
>>>> I was exploring to use swiotlb instead to create individual pools. But
>>>> that approach has multiple downsides:
>>>>
>>>> 1. Swiotlb is an OS primitive which is not available in all Operating
>>>> Systems. DMB however lives in the virtio transport layer, which means we
>>>> can add support for it in any OS independent of generic layers. This
>>>> helps with Windows support.
>>> How does it help, if you are going to put a pool in
>>> the driver, put a pool in the driver. Maybe with virtio mem to
>>> simplify allocation.
>>
>> I'm not sure I understand the suggestion :)
>>
>
> I'm not sure what the problem is for windows :)
>
> But if you want a chunk of contiguos memory that windows
> does not poke at without a driver, virtio mem is that :)
>
>>>> 2. We munge DMA space together. DMB provides a separate DMA space per
>>>> virtio device. This means we can for example implement a device in
>>>> vhost-user and give the implementing process only visibility to the DMB
>>>> region, not all of guest memory. That reduces the exposure the
>>>> vhost-user provider has, improving security.
>>>>
>>>> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
>>>> semantics for self-implemented devices (e.g. NSM), while requiring DMB
>>>> for devices implemented by less trustworthy providers. The
>>>> non-trustworthy devices do not get any visibility into the trustworthy
>>>> ones, even with DMB in place for both.
>>> So I am not sure whether the implication is that it's purely a software
>>> construct. But if it is, can we extend virtio iommu to
>>> add a way to discover and enforce trust boundaries?
>>> And maybe translate offsets to BARs, if that is desired?
>>>
>>> It seems to be that the result would be that we don't need fiddly
>>> special casing in virtio ring specifically, and a lot of things like
>>> pre-mapped dma will begin to work.
>>
>> On thing I'm trying to avoid is dynamicity. Anything that dynamically
>> changes visibility or needs state tracking is something that can go
>> wrong. By keeping everything self-contained within the guest, I can
>> reason about what is visible and what is not easily. Especially for
>> confidential computing, IMHO static wins over dynamic in general.
>>
>> Or did I misunderstand your suggestion?
>
> I get this part. But we can absolutely make it static.
>
>
> Let's start by replicating your functionality with virtio-iommu.  So:
>
> -device is behind virtio-iommu
> -virtio-iommu tells guest "this device can only consume memory from that range,"
> -and maybe: ...and translation is 1:1 with this offset
>
> if guest does not acknowledge it will just fail?
>
>
> this covers the proposal here simply by using a dedicated range
> per device, right?
>
>
>
> But look what we can easily add later: share a range
> between devices, now you can move data between them with zero copies.
>
>
> Isn't that better?


I don't know if "Add device emulation for virtio-mem and virtio-iommu, 
including all of its state machine and tracking" is necessarily an 
improvement, but let me try and prototype it.


>
>
>> As to pure software construct: With CXL, you can implement the exact
>> same protocol on real hardware as well. All it takes is cache coherency
>> of the BAR (or whatever the transport uses) region.
>>
>>
>>
>> Alex
> Yes this is what I thought originally, that you intend to
> do it in hardware. But re-reading it, it begins to look like
> that's not really the case, it's only "theoretically possible"?


Today I only have backends in software. I think it would be a pretty 
cool feature even for actual hardware or for 
not-as-obvious-but-still-cache-coherent-software such as EL3/SMM. I 
don't have concrete plans for them, but I like generic building blocks 
over too targeted and tied to a single use case :)


Alex
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Graf (AWS), Alexander 1 month, 2 weeks ago
On 10.08.26 10:25, Graf (AWS), Alexander wrote:
> On 10.08.26 10:04, Michael S. Tsirkin wrote:
>> On Mon, Aug 10, 2026 at 07:39:02AM +0000, Graf (AWS), Alexander wrote:
>>> Hey Michael,
>>>
>>> Thanks a bunch for having a detailed and quick look!
>>>
>>> On 10.08.26 08:23, Michael S. Tsirkin wrote:
>>>> On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
>>>>> Virtio drivers use guest memory to back virtqueues and their buffers.
>>>>> That means a VMM needs to be able to map guest memory. That is ok in the
>>>>> normal virt case. It gets icky with confidential computing (where we use
>>>>> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
>>>>> backing devices, because they end up with full RAM access to the guest.
>>>>>
>>>>> So instead, I'm proposing an extension to virtio which allows it to give
>>>>> each virtio device its own dedicated memory region to communicate with the
>>>>> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
>>>>> DMB to be present, which then enables safer, more isolated and resilient
>>>>> communication between guest and host.
>>>>>
>>>>> With DMB, the device provides a shared memory region that both parties
>>>>> agree is the full memory map both have access to. All memory offsets
>>>>> that previously would have been into guest RAM, are then offsets into
>>>>> this shared memory buffer region. One nice property of this is that it
>>>>> is a generic mechanism in the virtio transport layer, so higher level
>>>>> drivers work unmodified.
>>>>>
>>>>> I was exploring to use swiotlb instead to create individual pools. But
>>>>> that approach has multiple downsides:
>>>>>
>>>>> 1. Swiotlb is an OS primitive which is not available in all Operating
>>>>> Systems. DMB however lives in the virtio transport layer, which means we
>>>>> can add support for it in any OS independent of generic layers. This
>>>>> helps with Windows support.
>>>> How does it help, if you are going to put a pool in
>>>> the driver, put a pool in the driver. Maybe with virtio mem to
>>>> simplify allocation.
>>> I'm not sure I understand the suggestion :)
>>>
>> I'm not sure what the problem is for windows :)
>>
>> But if you want a chunk of contiguos memory that windows
>> does not poke at without a driver, virtio mem is that :)
>>
>>>>> 2. We munge DMA space together. DMB provides a separate DMA space per
>>>>> virtio device. This means we can for example implement a device in
>>>>> vhost-user and give the implementing process only visibility to the DMB
>>>>> region, not all of guest memory. That reduces the exposure the
>>>>> vhost-user provider has, improving security.
>>>>>
>>>>> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
>>>>> semantics for self-implemented devices (e.g. NSM), while requiring DMB
>>>>> for devices implemented by less trustworthy providers. The
>>>>> non-trustworthy devices do not get any visibility into the trustworthy
>>>>> ones, even with DMB in place for both.
>>>> So I am not sure whether the implication is that it's purely a software
>>>> construct. But if it is, can we extend virtio iommu to
>>>> add a way to discover and enforce trust boundaries?
>>>> And maybe translate offsets to BARs, if that is desired?
>>>>
>>>> It seems to be that the result would be that we don't need fiddly
>>>> special casing in virtio ring specifically, and a lot of things like
>>>> pre-mapped dma will begin to work.
>>> On thing I'm trying to avoid is dynamicity. Anything that dynamically
>>> changes visibility or needs state tracking is something that can go
>>> wrong. By keeping everything self-contained within the guest, I can
>>> reason about what is visible and what is not easily. Especially for
>>> confidential computing, IMHO static wins over dynamic in general.
>>>
>>> Or did I misunderstand your suggestion?
>> I get this part. But we can absolutely make it static.
>>
>>
>> Let's start by replicating your functionality with virtio-iommu.  So:
>>
>> -device is behind virtio-iommu
>> -virtio-iommu tells guest "this device can only consume memory from that range,"
>> -and maybe: ...and translation is 1:1 with this offset
>>
>> if guest does not acknowledge it will just fail?
>>
>>
>> this covers the proposal here simply by using a dedicated range
>> per device, right?
>>
>>
>>
>> But look what we can easily add later: share a range
>> between devices, now you can move data between them with zero copies.
>>
>>
>> Isn't that better?
>
> I don't know if "Add device emulation for virtio-mem and virtio-iommu,
> including all of its state machine and tracking" is necessarily an
> improvement, but let me try and prototype it.


I've been pondering about the virtio-mem idea a bit. I think you're 
trying to address 2 concerns:

1) RAM should be annotated as RAM.
2) We may in the future want to allow zero-copy DMA between virtio 
devices with DMB active

For 1, I tend to agree. That's why SHM is such a good fit. SHM already 
has RAM properties today and virtio devices use it that way. I don't 
think we need a special virtio-mem addition for that semantic.

However, virtio-mem brings in an interesting (and dangerous) other 
concept into the mix: Treating shared memory and private memory as a 
potentially common pool. Depending on how we implement the new 
virtio-mem mode, we may end up accidentally treating that shared memory 
as RAM.

That's a very undesirable design property: I want that a guest can rest 
fairly assured that it only every passes data into the shared memory 
area that it actively wants to. It's one of the properties that most of 
today's confidential compute technologies get wrong IMHO. I've seen way 
too many cases where you end up with a generic framework that happens to 
declare a page shared when in reality the page actually contains a mix 
of shared and private data. By clearly separating private and shared 
memory, I have a clear gate of when data passes between them and can 
easily reason about the secrecy of data.


for 2, I think it's an interesting use case. But it's nothing I'm 
worried about today. If it comes for free, I'm happy to take it. But 
looking at how much of a complicated monster we'd be creating with 
virtio-mem and virtio-iommu, I am convinced it's the wrong thing to 
optimize for :).

However, DMB as spec'ed does reference a "target SHM id". If we really 
later see a need for shared DMA, I think we can fairly easily create an 
extension that picks up your virtio-mem idea (or maybe something 
different) to implement a shared target identifier instead.


As for pure virtio-iommu, I genuinely fail to see how it is an 
improvement to either the code flow or the design principles I'm trying 
to achieve with this. We would still need a special purpose allocator. 
Or new zones, which Linux hates. And all we're gaining is yet another 
device that maintains useless state which eats up resources and which 
requires additional inter-connection between components to properly 
describe and tie up.

Maybe I'm missing the real point you're trying to make? :)


Alex
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Michael S. Tsirkin 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 07:14:09PM +0000, Graf (AWS), Alexander wrote:
> Maybe I'm missing the real point you're trying to make? :)

At least one point is not to have text like "intepret every address as
an offset" in the spec and actually make it an address.
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Graf (AWS), Alexander 1 month, 2 weeks ago
On 10.08.26 23:42, Michael S. Tsirkin wrote:
> On Mon, Aug 10, 2026 at 07:14:09PM +0000, Graf (AWS), Alexander wrote:
>> Maybe I'm missing the real point you're trying to make? :)
> At least one point is not to have text like "intepret every address as
> an offset" in the spec and actually make it an address.


I am genuinely unable to see how we can make it prettier.

What I'm trying to transmit is that each device should live in its own 
address space, with the host specifying which boundary an address space 
has. I agree that on surface it sounds a bit like an IOMMU. And yes, we 
can combine that with virtio-iommu, but that's really just punting the 
problem down one layer and makes it more difficult to recreate the ties 
that a native per-device mechanism already gives us.

With a pure virtio-iommu implementation, the most obvious way to create 
a private memory region (not RAM, because this region should never be 
used by anything as RAM) via an SHM region to the virtio-iommu device.

Now we're moving the wording in the spec from "Device DMA is routed to a 
private address space which is backed by the DMB SHM region" into 
"Virtio-iommu has the ability to target a special address space, backed 
by an SHM region". And we then need to find a way to size that region, 
dynamically slice it (for hotplug support) and communicate between all 
components where and how to use the virtio-iommu SHM region instead.

If we don't want to use a virtio-iommu SHM region, but punt it down one 
layer to virtio-mem, it gets even more difficult. We'd have all of the 
problems that we get when using virtio-iommu, but on top of that we now 
need to enlighten virtio-mem with a special "private" memory region that 
is not RAM. And then connect that back into virtio-iommu.

So while I understand your concern, I fail to see a path towards an 
actually viable, genuinely better (prettier / more maintainable / easier 
/ less error prone / etc) implementation with virtio-iommu.


Alex
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Stefan Hajnoczi 1 month, 2 weeks ago
On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
> Virtio drivers use guest memory to back virtqueues and their buffers.
> That means a VMM needs to be able to map guest memory. That is ok in the
> normal virt case. It gets icky with confidential computing (where we use
> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
> backing devices, because they end up with full RAM access to the guest.
> 
> So instead, I'm proposing an extension to virtio which allows it to give
> each virtio device its own dedicated memory region to communicate with the
> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
> DMB to be present, which then enables safer, more isolated and resilient
> communication between guest and host.
> 
> With DMB, the device provides a shared memory region that both parties
> agree is the full memory map both have access to. All memory offsets
> that previously would have been into guest RAM, are then offsets into
> this shared memory buffer region. One nice property of this is that it
> is a generic mechanism in the virtio transport layer, so higher level
> drivers work unmodified.

This is similar to VIRTIO's Shared Memory Regions. A problem with this
kind of approach is that guest software that depends on zero-copy,
O_DIRECT, the ability to mmap, etc may break when passing an address
from one device to another device.

For example, a guest userspace application writing to a virtio-blk
device with O_DIRECT can pass any source memory buffer. One VIRTIO
device will be unable to address another VIRTIO device's DMB. This
problem also extends to vhost-user where one back-end cannot access
another back-end's DMB or Shared Memory Regions.

Maybe your use case will never hit this problem because you can rely on
the guest software never to assume zero-copy/O_DIRECT/etc works.
virtiofs hit it with its DAX Shared Memory Regions.

I'm mentioned it in case this is something you want to think about
before deploying this approach.

> 
> I was exploring to use swiotlb instead to create individual pools. But
> that approach has multiple downsides:
> 
> 1. Swiotlb is an OS primitive which is not available in all Operating
> Systems. DMB however lives in the virtio transport layer, which means we
> can add support for it in any OS independent of generic layers. This
> helps with Windows support.
> 
> 2. We munge DMA space together. DMB provides a separate DMA space per
> virtio device. This means we can for example implement a device in
> vhost-user and give the implementing process only visibility to the DMB
> region, not all of guest memory. That reduces the exposure the
> vhost-user provider has, improving security.

Connor Kite is working on a different approach for vhost-user memory
isolation here:
https://lore.kernel.org/qemu-devel/20260723-vhost-user-isolated-memory-v1-0-6b97c439eb28@gmail.com/T/#t

It involves a bounce buffer at the VMM level. Unmodified vhost-user
back-ends never sees guest RAM. Guest drivers are also unmodified. The
cost of doing this is that the VMM has to intercept kick and call
eventfds in order to copy between the bounce buffer and guest RAM.

I don't see DMB or vhost-user memory isolation as conflicting features.
There can be two ways of solving the same problem with different
trade-offs. I just wanted to share a link to Connor's ongoing work.

> 
> 3. Devices can opt-in. A hypervisor can choose to use standard virtio
> semantics for self-implemented devices (e.g. NSM), while requiring DMB
> for devices implemented by less trustworthy providers. The
> non-trustworthy devices do not get any visibility into the trustworthy
> ones, even with DMB in place for both.
> 
> == Limitations ==
> 
>   - Only PCI is wired up.
>   - Feature bit 44 and the shared memory id register at offset 0x40 of the
>     PCI common configuration are provisional: the OASIS technical
>     committee has the specification and has allocated neither.
> 
>   https://lore.kernel.org/virtio-comment/20260804161202.38619-1-graf@amazon.com/
> 
>   - The device I ran this against is not public, so you cannot reproduce
>     the numbers below. The KUnit test you can.
> 
> == Testing ==
> 
> Without patches 10 and 12 a receive refill livelocks and queues starve
> each other: 29,319,791 receive softirqs in five seconds with not one
> packet received, against 4 with them, and 2 of 7 receive queues that never
> see a buffer, against none. Earlier revisions moved 512 MiB of O_DIRECT
> block I/O and 2.7 GB of verified vsock through a region with no error.
> 
> The KUnit test for patch 12's guarantee you can run yourself, and two of
> its five cases fail if I take the fix out:
> 
>   tools/testing/kunit/kunit.py run --arch=x86_64 \
>       --kconfig_add CONFIG_VIRTIO_MMIO=y \
>       --kconfig_add CONFIG_VIRTIO_DMB=y virtio_dmb
> 
> Patch 1 can be taken on its own: a stale worked example in a vdpa
> comment. Patch 10 fixes something older than this series too, a failed
> mapping arriving as -EIO from a packed ring, failing an I/O that on a
> split ring is only back-pressure, but it does not apply alone: it wants
> patch 2, and patch 9 for the file its documentation hunk edits. Both
> carry Fixes:. Patch 12 wants 10 first.
> 
> I wrote this series with an AI coding assistant, which drafted the code,
> the changelogs and this cover letter. I reviewed and reworked all of it,
> and every commit carries an Assisted-by: trailer.
> 
> Alex
> 
> Alexander Graf (12):
>   vdpa: correct the VIRTIO_DEVICE_F_MASK example value
>   virtio_ring: validate premapped addresses through the device's map
>   virtio: add the VIRTIO_F_DMB feature bit
>   virtio_pci: read the device memory buffer shared memory id
>   virtio_pci: create virtqueues with the device's mapping token
>   virtio: add a device memory buffer region allocator
>   virtio: locate the device memory buffer after feature negotiation
>   virtio_pci: support VIRTIO_F_DMB
>   Documentation: virtio: describe the device memory buffer
>   virtio_ring: report a bounded pool's exhaustion as -ENOSPC
>   virtio: expose device memory buffer occupancy over debugfs
>   virtio: guarantee a virtqueue can publish its first descriptor chain
> 
>  Documentation/driver-api/virtio/index.rst     |    1 +
>  .../driver-api/virtio/virtio-dmb.rst          |  803 ++++++++
>  drivers/vdpa/vdpa.c                           |    2 +-
>  drivers/virtio/Kconfig                        |   29 +
>  drivers/virtio/Makefile                       |    3 +-
>  drivers/virtio/virtio.c                       |   15 +-
>  drivers/virtio/virtio_dmb.c                   | 1719 +++++++++++++++++
>  drivers/virtio/virtio_dmb.h                   |   34 +
>  drivers/virtio/virtio_dmb_test.c              |  279 +++
>  drivers/virtio/virtio_pci_modern.c            |  100 +-
>  drivers/virtio/virtio_pci_modern_dev.c        |   23 +-
>  drivers/virtio/virtio_ring.c                  |  216 ++-
>  include/linux/virtio.h                        |    3 +
>  include/linux/virtio_config.h                 |   41 +
>  include/linux/virtio_pci_modern.h             |    1 +
>  include/uapi/linux/virtio_config.h            |   17 +-
>  include/uapi/linux/virtio_pci.h               |   10 +
>  17 files changed, 3254 insertions(+), 42 deletions(-)
>  create mode 100644 Documentation/driver-api/virtio/virtio-dmb.rst
>  create mode 100644 drivers/virtio/virtio_dmb.c
>  create mode 100644 drivers/virtio/virtio_dmb.h
>  create mode 100644 drivers/virtio/virtio_dmb_test.c
> 
> 
> base-commit: fc02acf6ac0ccde0c805c2daa9148683cdd01ba8
> 
Re: [RFC PATCH 00/12] virtio: support devices that own their virtqueue memory
Posted by Graf (AWS), Alexander 1 month, 2 weeks ago
On 10.08.26 22:39, Stefan Hajnoczi wrote:
> On Sun, Aug 09, 2026 at 06:19:58PM +0000, Alexander Graf wrote:
>> Virtio drivers use guest memory to back virtqueues and their buffers.
>> That means a VMM needs to be able to map guest memory. That is ok in the
>> normal virt case. It gets icky with confidential computing (where we use
>> swiotlb as workaround) and it defeats the purpose of isolated vhost-user
>> backing devices, because they end up with full RAM access to the guest.
>>
>> So instead, I'm proposing an extension to virtio which allows it to give
>> each virtio device its own dedicated memory region to communicate with the
>> host, called DMB (Device Memory Buffer). A trusted hypervisor can force
>> DMB to be present, which then enables safer, more isolated and resilient
>> communication between guest and host.
>>
>> With DMB, the device provides a shared memory region that both parties
>> agree is the full memory map both have access to. All memory offsets
>> that previously would have been into guest RAM, are then offsets into
>> this shared memory buffer region. One nice property of this is that it
>> is a generic mechanism in the virtio transport layer, so higher level
>> drivers work unmodified.
> This is similar to VIRTIO's Shared Memory Regions. A problem with this
> kind of approach is that guest software that depends on zero-copy,
> O_DIRECT, the ability to mmap, etc may break when passing an address
> from one device to another device.
>
> For example, a guest userspace application writing to a virtio-blk
> device with O_DIRECT can pass any source memory buffer. One VIRTIO
> device will be unable to address another VIRTIO device's DMB. This
> problem also extends to vhost-user where one back-end cannot access
> another back-end's DMB or Shared Memory Regions.
>
> Maybe your use case will never hit this problem because you can rely on
> the guest software never to assume zero-copy/O_DIRECT/etc works.
> virtiofs hit it with its DAX Shared Memory Regions.


Virtiofs is a bit trickier. It assumes reverse mapping order (host maps 
memory into guest address space's SHM region) which is something I'm not 
looking to support. The main reason you did run into it is because the 
page cache is actually implemented by the SHM region, so Linux 
(rightfully) assumes that it has direct access to it.

For normal virtio device operations (like O_DIRECT on virtio-blk), I am 
not sure whether we can ever run into a case where anything in the 
driver assumes that both devices share the same IOVA space. Thanks a lot 
for the heads-up though, I'll take a deeper look to make sure that this 
is the case. But even if it is, IMHO we can consider it a guest bug and 
fix it in the guest code when DMB is active.


> I'm mentioned it in case this is something you want to think about
> before deploying this approach.
>
>> I was exploring to use swiotlb instead to create individual pools. But
>> that approach has multiple downsides:
>>
>> 1. Swiotlb is an OS primitive which is not available in all Operating
>> Systems. DMB however lives in the virtio transport layer, which means we
>> can add support for it in any OS independent of generic layers. This
>> helps with Windows support.
>>
>> 2. We munge DMA space together. DMB provides a separate DMA space per
>> virtio device. This means we can for example implement a device in
>> vhost-user and give the implementing process only visibility to the DMB
>> region, not all of guest memory. That reduces the exposure the
>> vhost-user provider has, improving security.
> Connor Kite is working on a different approach for vhost-user memory
> isolation here:
> https://lore.kernel.org/qemu-devel/20260723-vhost-user-isolated-memory-v1-0-6b97c439eb28@gmail.com/T/#t
>
> It involves a bounce buffer at the VMM level. Unmodified vhost-user
> back-ends never sees guest RAM. Guest drivers are also unmodified. The
> cost of doing this is that the VMM has to intercept kick and call
> eventfds in order to copy between the bounce buffer and guest RAM.
>
> I don't see DMB or vhost-user memory isolation as conflicting features.
> There can be two ways of solving the same problem with different
> trade-offs. I just wanted to share a link to Connor's ongoing work.


Thanks a bunch :). My design motivation is a bit different, but it's 
great to see more people interested in isolation!


Alex