[PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device

Zhenzhong Duan posted 6 patches 1 month, 3 weeks ago
Patches applied successfully (tree, apply log)
git fetch https://github.com/patchew-project/qemu tags/patchew/20260806073825.363468-1-zhenzhong.duan@intel.com
Maintainers: Yi Liu <yi.l.liu@intel.com>, Eric Auger <eric.auger@redhat.com>, Zhenzhong Duan <zhenzhong.duan@intel.com>, Peter Maydell <peter.maydell@linaro.org>, "Michael S. Tsirkin" <mst@redhat.com>, Jason Wang <jasowangio@gmail.com>, "Clément Mathieu--Drif" <clement.mathieu--drif@bull.com>, Paolo Bonzini <pbonzini@redhat.com>, Richard Henderson <richard.henderson@linaro.org>, Alex Williamson <alex@shazbot.org>, "Cédric Le Goater" <clg@redhat.com>
There is a newer version of this series
hw/i386/intel_iommu_accel.h    |  16 ++-
hw/i386/intel_iommu_internal.h |   3 +
include/hw/i386/intel_iommu.h  |   6 +
include/system/iommufd.h       |   7 +-
backends/iommufd.c             |  28 +++-
hw/arm/smmuv3-accel.c          |   6 +-
hw/i386/intel_iommu.c          |  11 +-
hw/i386/intel_iommu_accel.c    | 245 ++++++++++++++++++++++++++++++++-
hw/vfio/iommufd.c              |   2 +-
backends/trace-events          |   3 +-
hw/i386/trace-events           |   2 +
11 files changed, 308 insertions(+), 21 deletions(-)
[PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Zhenzhong Duan 1 month, 3 weeks ago
Hi,

When svm=on,pasid-bits=N,fsts=on is configured for virtual vtd, we enable
guest's support for vSVM. In this case, the host VTD may generate
recoverable first stage page fault event, QEMU read the event and inject
it to guest.

After guest handles the event, it sends a page group response, QEMU gets
the response and pass it to host VTD.

This series adds QEMU support for receiving such host events through the
FAULTQ interface and propagating them to the guest, then catching guest's
responses and propagating them to host.

GIT branch: https://github.com/yiliu1765/qemu/tree/zhenzhong/iommufd_prq.v3

Tests:
Tested with DSA wq attached to an user app which triggers IO page fault on
process page table durging DMA.

Thanks
Zhenzhong

Changelog:
v3:
- convert to PCI pasid before passing to vtd_pri_request_page (Clement)
- expose vtd_pri_request_page and drop accel_ops pointer (Clement)
- add a cleanup patch "intel_iommu_accel: Guard VTDAccelPASIDCacheEntry definition" (Philippe)

v2:
- define fault_fd int type (Liuyi)
- update comments on FAULTQ_BUF_SIZE (Liuyi)
- drop param vtd_hiod in vtd_destroy_old_fs_faultq() (Liuyi)
- s/vtd_propagate_page_group_response_accel/vtd_accel_propagate_page_group_response (Liuyi)
- call vtd_prq_response_notify() directly instead of through pri_notifier (Liuyi)
- replace "fault_id != 0" check with "fault_fd < 0" check (Liuyi)


Zhenzhong Duan (6):
  intel_iommu_accel: Guard VTDAccelPASIDCacheEntry definition
  backends/iommufd: Introduce iommufd_backend_alloc_faultq
  backends/iommufd: Extend iommufd_backend_alloc_hwpt() with fault_id
  intel_iommu_accel: Add PRQ injection for passthrough device
  intel_iommu_accel: Accept PRQ response for passthrough device
  intel_iommu_accel: teardown FAULTQ resources in bottom half

 hw/i386/intel_iommu_accel.h    |  16 ++-
 hw/i386/intel_iommu_internal.h |   3 +
 include/hw/i386/intel_iommu.h  |   6 +
 include/system/iommufd.h       |   7 +-
 backends/iommufd.c             |  28 +++-
 hw/arm/smmuv3-accel.c          |   6 +-
 hw/i386/intel_iommu.c          |  11 +-
 hw/i386/intel_iommu_accel.c    | 245 ++++++++++++++++++++++++++++++++-
 hw/vfio/iommufd.c              |   2 +-
 backends/trace-events          |   3 +-
 hw/i386/trace-events           |   2 +
 11 files changed, 308 insertions(+), 21 deletions(-)

-- 
2.52.0
Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Clément MATHIEU--DRIF 1 month, 2 weeks ago
Hi Zhenzhong,

I will read this series early next week.
thanks for the respin!

btw, related to this work: have ever seen
vfio_device_get_aw_bits returning 47 whereas
the iommu supports 48? This leads the following
aw test to fail:

```
ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
if (ret < 0) {
    return false;
}
if (s->aw_bits > ret) {
    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
    return false;
}
```

I got this when trying to nest VMs.

On Thu, 2026-08-06 at 03:38 -0400, Zhenzhong Duan wrote:
> Caution: External email. Do not open attachments or click links, unless this email comes from a known sender and you know the content is safe.
> 
> 
> Hi,
> 
> When svm=on,pasid-bits=N,fsts=on is configured for virtual vtd, we enable  
> guest's support for vSVM. In this case, the host VTD may generate  
> recoverable first stage page fault event, QEMU read the event and inject  
> it to guest.
> 
> After guest handles the event, it sends a page group response, QEMU gets  
> the response and pass it to host VTD.
> 
> This series adds QEMU support for receiving such host events through the  
> FAULTQ interface and propagating them to the guest, then catching guest's  
> responses and propagating them to host.
> 
> GIT branch: [https://github.com/yiliu1765/qemu/tree/zhenzhong/iommufd_prq.v3](https://github.com/yiliu1765/qemu/tree/zhenzhong/iommufd_prq.v3)
> 
> Tests:  
> Tested with DSA wq attached to an user app which triggers IO page fault on  
> process page table durging DMA.
> 
> Thanks  
> Zhenzhong
> 
> Changelog:  
> v3:  
> - convert to PCI pasid before passing to vtd_pri_request_page (Clement)  
> - expose vtd_pri_request_page and drop accel_ops pointer (Clement)  
> - add a cleanup patch "intel_iommu_accel: Guard VTDAccelPASIDCacheEntry definition" (Philippe)
> 
> v2:  
> - define fault_fd int type (Liuyi)  
> - update comments on FAULTQ_BUF_SIZE (Liuyi)  
> - drop param vtd_hiod in vtd_destroy_old_fs_faultq() (Liuyi)  
> - s/vtd_propagate_page_group_response_accel/vtd_accel_propagate_page_group_response (Liuyi)  
> - call vtd_prq_response_notify() directly instead of through pri_notifier (Liuyi)  
> - replace "fault_id != 0" check with "fault_fd < 0" check (Liuyi)
> 
> 
> Zhenzhong Duan (6):  
>   intel_iommu_accel: Guard VTDAccelPASIDCacheEntry definition  
>   backends/iommufd: Introduce iommufd_backend_alloc_faultq  
>   backends/iommufd: Extend iommufd_backend_alloc_hwpt() with fault_id  
>   intel_iommu_accel: Add PRQ injection for passthrough device  
>   intel_iommu_accel: Accept PRQ response for passthrough device  
>   intel_iommu_accel: teardown FAULTQ resources in bottom half
> 
>  hw/i386/intel_iommu_accel.h    |  16 ++-  
>  hw/i386/intel_iommu_internal.h |   3 +  
>  include/hw/i386/intel_iommu.h  |   6 +  
>  include/system/iommufd.h       |   7 +-  
>  backends/iommufd.c             |  28 +++-  
>  hw/arm/smmuv3-accel.c          |   6 +-  
>  hw/i386/intel_iommu.c          |  11 +-  
>  hw/i386/intel_iommu_accel.c    | 245 ++++++++++++++++++++++++++++++++-  
>  hw/vfio/iommufd.c              |   2 +-  
>  backends/trace-events          |   3 +-  
>  hw/i386/trace-events           |   2 +  
>  11 files changed, 308 insertions(+), 21 deletions(-)
> 
> --  
> 2.52.0
> 
RE: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Duan, Zhenzhong 1 month, 2 weeks ago
Hi Clement,

>-----Original Message-----
>From: Clément MATHIEU--DRIF <clement.mathieu--drif@bull.com>
>Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>device
>
>Hi Zhenzhong,
>
>I will read this series early next week.
>thanks for the respin!
>
>btw, related to this work: have ever seen
>vfio_device_get_aw_bits returning 47 whereas
>the iommu supports 48? This leads the following
>aw test to fail:
>
>```
>ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
>if (ret < 0) {
>    return false;
>}
>if (s->aw_bits > ret) {
>    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
>    return false;
>}
>```
>
>I got this when trying to nest VMs.

I am trying to reproduce, what device did you passthrough?

Thanks
Zhenzhong
Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Clément MATHIEU--DRIF 1 month, 2 weeks ago
On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
> Caution: External email. Do not open attachments or click links, unless this email comes from a known sender and you know the content is safe.
> 
> 
> Hi Clement,
> 
> 
> > -----Original Message-----
> > From: Clément MATHIEU--DRIF <[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > device
> > 
> > Hi Zhenzhong,
> > 
> > I will read this series early next week.
> > thanks for the respin!
> > 
> > btw, related to this work: have ever seen
> > vfio_device_get_aw_bits returning 47 whereas
> > the iommu supports 48? This leads the following
> > aw test to fail:
> > 
> > ```
> > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
> > if (ret < 0) {
> >    return false;
> > }
> > if (s->aw_bits > ret) {
> >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
> >    return false;
> > }
> > ```
> > 
> > I got this when trying to nest VMs.
> 
> 
> I am trying to reproduce, what device did you passthrough?

I used edu for testing.  
An I had aw-bits=48 on both vms.  

> 
> Thanks  
> Zhenzhong
RE: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Duan, Zhenzhong 1 month, 2 weeks ago

>-----Original Message-----
>From: Clément MATHIEU--DRIF <clement.mathieu--drif@bull.com>
>Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>device
>
>
>On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
>> Caution: External email. Do not open attachments or click links, unless this email
>comes from a known sender and you know the content is safe.
>>
>>
>> Hi Clement,
>>
>>
>> > -----Original Message-----
>> > From: Clément MATHIEU--DRIF <[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
>> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>> > device
>> >
>> > Hi Zhenzhong,
>> >
>> > I will read this series early next week.
>> > thanks for the respin!
>> >
>> > btw, related to this work: have ever seen
>> > vfio_device_get_aw_bits returning 47 whereas
>> > the iommu supports 48? This leads the following
>> > aw test to fail:
>> >
>> > ```
>> > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
>> > if (ret < 0) {
>> >    return false;
>> > }
>> > if (s->aw_bits > ret) {
>> >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
>> >    return false;
>> > }
>> > ```
>> >
>> > I got this when trying to nest VMs.
>>
>>
>> I am trying to reproduce, what device did you passthrough?
>
>I used edu for testing.
>An I had aw-bits=48 on both vms.

Reproduced the same with L1(fsts=on) and L2(fsts=off).

It's expected failure as first stage page table only supports 47bit IOVA ranges
and L2 which uses second stage page table supports 48bit IOVA ranges,
that's not supported.

This is documented in vtd spec 3.6:

Software using first-stage translation structures to translate an IO Virtual Address (IOVA) must use 
canonical addresses. Additionally, software must limit addresses to less than the minimum of MGAW 
and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM is 4-level and 56-bit 
when FSPM is 5-level).

Thanks
Zhenzhong
Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Clément MATHIEU--DRIF 1 month, 2 weeks ago
On Wed, 2026-08-12 at 13:26 +0000, Duan, Zhenzhong wrote:
> Caution: External email. Do not open attachments or click links, unless this email comes from a known sender and you know the content is safe.
> 
> 
> 
> > -----Original Message-----
> > From: Clément MATHIEU--DRIF <[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > device
> > 
> > 
> > On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
> > 
> > > Caution: External email. Do not open attachments or click links, unless this email
> > 
> > comes from a known sender and you know the content is safe.
> > 
> > > 
> > > 
> > > Hi Clement,
> > > 
> > > 
> > > 
> > > > -----Original Message-----
> > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com))>
> > 
> > > 
> > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > > > device
> > > > 
> > > > Hi Zhenzhong,
> > > > 
> > > > I will read this series early next week.
> > > > thanks for the respin!
> > > > 
> > > > btw, related to this work: have ever seen
> > > > vfio_device_get_aw_bits returning 47 whereas
> > > > the iommu supports 48? This leads the following
> > > > aw test to fail:
> > > > 
> > > > ```
> > > > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
> > > > if (ret < 0) {
> > > >    return false;
> > > > }
> > > > if (s->aw_bits > ret) {
> > > >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
> > > >    return false;
> > > > }
> > > > ```
> > > > 
> > > > I got this when trying to nest VMs.
> > > 
> > > 
> > > 
> > > I am trying to reproduce, what device did you passthrough?
> > 
> > 
> > I used edu for testing.
> > An I had aw-bits=48 on both vms.
> 
> 
> Reproduced the same with L1(fsts=on) and L2(fsts=off).

My config has fsts=on on both VMs.  
Which kernel are you using?  
Mine is RHEL10.2 (6.12.0-211.7.3.el10_2.x86_64)

> 
> It's expected failure as first stage page table only supports 47bit IOVA ranges  
> and L2 which uses second stage page table supports 48bit IOVA ranges,  
> that's not supported.
> 
> This is documented in vtd spec 3.6:
> 
> Software using first-stage translation structures to translate an IO Virtual Address (IOVA) must use  
> canonical addresses. Additionally, software must limit addresses to less than the minimum of MGAW  
> and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM is 4-level and 56-bit  
> when FSPM is 5-level).
> 
> Thanks  
> Zhenzhong
RE: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Duan, Zhenzhong 1 month, 1 week ago
Hi Clement,

>-----Original Message-----
>From: Clément MATHIEU--DRIF <clement.mathieu--drif@bull.com>
>Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>device
>
>
>On Wed, 2026-08-12 at 13:26 +0000, Duan, Zhenzhong wrote:
>> Caution: External email. Do not open attachments or click links, unless this email
>comes from a known sender and you know the content is safe.
>>
>>
>>
>> > -----Original Message-----
>> > From: Clément MATHIEU--DRIF <[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
>> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>> > device
>> >
>> >
>> > On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
>> >
>> > > Caution: External email. Do not open attachments or click links, unless this
>email
>> >
>> > comes from a known sender and you know the content is safe.
>> >
>> > >
>> > >
>> > > Hi Clement,
>> > >
>> > >
>> > >
>> > > > -----Original Message-----
>> > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
>> > >
>> >
>> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com))>
>> >
>> > >
>> > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
>passthrough
>> > > > device
>> > > >
>> > > > Hi Zhenzhong,
>> > > >
>> > > > I will read this series early next week.
>> > > > thanks for the respin!
>> > > >
>> > > > btw, related to this work: have ever seen
>> > > > vfio_device_get_aw_bits returning 47 whereas
>> > > > the iommu supports 48? This leads the following
>> > > > aw test to fail:
>> > > >
>> > > > ```
>> > > > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
>> > > > if (ret < 0) {
>> > > >    return false;
>> > > > }
>> > > > if (s->aw_bits > ret) {
>> > > >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
>> > > >    return false;
>> > > > }
>> > > > ```
>> > > >
>> > > > I got this when trying to nest VMs.
>> > >
>> > >
>> > >
>> > > I am trying to reproduce, what device did you passthrough?
>> >
>> >
>> > I used edu for testing.
>> > An I had aw-bits=48 on both vms.
>>
>>
>> Reproduced the same with L1(fsts=on) and L2(fsts=off).
>
>My config has fsts=on on both VMs.
>Which kernel are you using?
>Mine is RHEL10.2 (6.12.0-211.7.3.el10_2.x86_64)

In theory we will not run into that check if fsts=on on both VMs, on my env I see:

"qemu-system-x86_64: -device vfio-pci,host=00:04.0,iommufd=iommufd0: vfio 0000:00:04.0: Failed to allocate hwpt: Operation not supported"

which is expected as we want nesting support in L1 but it only supports fsts.

My kernel is 7.2.0-rc6, QEMU is v11.1.0

>
>>
>> It's expected failure as first stage page table only supports 47bit IOVA ranges
>> and L2 which uses second stage page table supports 48bit IOVA ranges,
>> that's not supported.
>>
>> This is documented in vtd spec 3.6:
>>
>> Software using first-stage translation structures to translate an IO Virtual
>Address (IOVA) must use
>> canonical addresses. Additionally, software must limit addresses to less than the
>minimum of MGAW
>> and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM
>is 4-level and 56-bit
>> when FSPM is 5-level).
>>
>> Thanks
>> Zhenzhong
Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Clément MATHIEU--DRIF 1 month, 1 week ago
On Fri, 2026-08-21 at 02:44 +0000, Duan, Zhenzhong wrote:
> Caution: External email. Do not open attachments or click links, unless this email comes from a known sender and you know the content is safe.
> 
> 
> Hi Clement,
> 
> 
> > -----Original Message-----
> > From: Clément MATHIEU--DRIF <[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > device
> > 
> > 
> > On Wed, 2026-08-12 at 13:26 +0000, Duan, Zhenzhong wrote:
> > 
> > > Caution: External email. Do not open attachments or click links, unless this email
> > 
> > comes from a known sender and you know the content is safe.
> > 
> > > 
> > > 
> > > 
> > > 
> > > > -----Original Message-----
> > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com))>
> > 
> > > 
> > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > > > device
> > > > 
> > > > 
> > > > On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
> > > > 
> > > > 
> > > > > Caution: External email. Do not open attachments or click links, unless this
> > > > 
> > > 
> > 
> > email
> > 
> > > 
> > > > 
> > > > comes from a known sender and you know the content is safe.
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > Hi Clement,
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > > -----Original Message-----
> > > > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > > > 
> > > > > 
> > > > 
> > > > 
> > > > [[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.com](mailto:drif@bull.com))](mailto:[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)))>
> > 
> > > 
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
> > > > > 
> > > > 
> > > 
> > 
> > passthrough
> > 
> > > 
> > > > 
> > > > > 
> > > > > > device
> > > > > > 
> > > > > > Hi Zhenzhong,
> > > > > > 
> > > > > > I will read this series early next week.
> > > > > > thanks for the respin!
> > > > > > 
> > > > > > btw, related to this work: have ever seen
> > > > > > vfio_device_get_aw_bits returning 47 whereas
> > > > > > the iommu supports 48? This leads the following
> > > > > > aw test to fail:
> > > > > > 
> > > > > > ```
> > > > > > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS, errp);
> > > > > > if (ret < 0) {
> > > > > >    return false;
> > > > > > }
> > > > > > if (s->aw_bits > ret) {
> > > > > >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
> > > > > >    return false;
> > > > > > }
> > > > > > ```
> > > > > > 
> > > > > > I got this when trying to nest VMs.
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > I am trying to reproduce, what device did you passthrough?
> > > > 
> > > > 
> > > > 
> > > > I used edu for testing.
> > > > An I had aw-bits=48 on both vms.
> > > 
> > > 
> > > 
> > > Reproduced the same with L1(fsts=on) and L2(fsts=off).
> > 
> > 
> > My config has fsts=on on both VMs.
> > Which kernel are you using?
> > Mine is RHEL10.2 (6.12.0-211.7.3.el10_2.x86_64)
> 
> 
> In theory we will not run into that check if fsts=on on both VMs, on my env I see:
> 
> "qemu-system-x86_64: -device vfio-pci,host=00:04.0,iommufd=iommufd0: vfio 0000:00:04.0: Failed to allocate hwpt: Operation not supported"
> 
> which is expected as we want nesting support in L1 but it only supports fsts.

Oh yep, I forgot to mention that I'm not attaching iommufd for this test.  
I agree with what you say when iommufd is present, but without it (weird config,  
I admit, but interesting for testing), I get:

```
qemu-system-x86_64: -device vfio-pci,host=00:01.0: vfio 0000:00:01.0: Failed to set vIOMMU: aw-bits 48 > host aw-bits 47
```

Just wondering why such a 1bit difference.  

> 
> My kernel is 7.2.0-rc6, QEMU is v11.1.0
> 
> 
> 
> > 
> > > 
> > > It's expected failure as first stage page table only supports 47bit IOVA ranges
> > > and L2 which uses second stage page table supports 48bit IOVA ranges,
> > > that's not supported.
> > > 
> > > This is documented in vtd spec 3.6:
> > > 
> > > Software using first-stage translation structures to translate an IO Virtual
> > 
> > Address (IOVA) must use
> > 
> > > canonical addresses. Additionally, software must limit addresses to less than the
> > 
> > minimum of MGAW
> > 
> > > and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM
> > 
> > is 4-level and 56-bit
> > 
> > > when FSPM is 5-level).
> > > 
> > > Thanks
> > > Zhenzhong
> >
> 
RE: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Duan, Zhenzhong 1 month, 1 week ago

>-----Original Message-----
>From: Clément MATHIEU--DRIF <clement.mathieu--drif@bull.com>
>Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>device
>
>
>On Fri, 2026-08-21 at 02:44 +0000, Duan, Zhenzhong wrote:
>> Caution: External email. Do not open attachments or click links, unless this email
>comes from a known sender and you know the content is safe.
>>
>>
>> Hi Clement,
>>
>>
>> > -----Original Message-----
>> > From: Clément MATHIEU--DRIF <[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
>> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
>> > device
>> >
>> >
>> > On Wed, 2026-08-12 at 13:26 +0000, Duan, Zhenzhong wrote:
>> >
>> > > Caution: External email. Do not open attachments or click links, unless this
>email
>> >
>> > comes from a known sender and you know the content is safe.
>> >
>> > >
>> > >
>> > >
>> > >
>> > > > -----Original Message-----
>> > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
>> > >
>> >
>> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com))>
>> >
>> > >
>> > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
>passthrough
>> > > > device
>> > > >
>> > > >
>> > > > On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
>> > > >
>> > > >
>> > > > > Caution: External email. Do not open attachments or click links, unless
>this
>> > > >
>> > >
>> >
>> > email
>> >
>> > >
>> > > >
>> > > > comes from a known sender and you know the content is safe.
>> > > >
>> > > >
>> > > > >
>> > > > >
>> > > > > Hi Clement,
>> > > > >
>> > > > >
>> > > > >
>> > > > >
>> > > > > > -----Original Message-----
>> > > > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
>> > > > >
>> > > > >
>> > > >
>> > > >
>> > > >
>[[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.com](mailto:drif@bull.c
>om))](mailto:[clement.mathieu--
>> > >
>> >
>> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--
>drif@bull.com](mailto:clement.mathieu--drif@bull.com)))>
>> >
>> > >
>> > > >
>> > > >
>> > > > >
>> > > > >
>> > > > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
>> > > > >
>> > > >
>> > >
>> >
>> > passthrough
>> >
>> > >
>> > > >
>> > > > >
>> > > > > > device
>> > > > > >
>> > > > > > Hi Zhenzhong,
>> > > > > >
>> > > > > > I will read this series early next week.
>> > > > > > thanks for the respin!
>> > > > > >
>> > > > > > btw, related to this work: have ever seen
>> > > > > > vfio_device_get_aw_bits returning 47 whereas
>> > > > > > the iommu supports 48? This leads the following
>> > > > > > aw test to fail:
>> > > > > >
>> > > > > > ```
>> > > > > > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS,
>errp);
>> > > > > > if (ret < 0) {
>> > > > > >    return false;
>> > > > > > }
>> > > > > > if (s->aw_bits > ret) {
>> > > > > >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
>> > > > > >    return false;
>> > > > > > }
>> > > > > > ```
>> > > > > >
>> > > > > > I got this when trying to nest VMs.
>> > > > >
>> > > > >
>> > > > >
>> > > > >
>> > > > > I am trying to reproduce, what device did you passthrough?
>> > > >
>> > > >
>> > > >
>> > > > I used edu for testing.
>> > > > An I had aw-bits=48 on both vms.
>> > >
>> > >
>> > >
>> > > Reproduced the same with L1(fsts=on) and L2(fsts=off).
>> >
>> >
>> > My config has fsts=on on both VMs.
>> > Which kernel are you using?
>> > Mine is RHEL10.2 (6.12.0-211.7.3.el10_2.x86_64)
>>
>>
>> In theory we will not run into that check if fsts=on on both VMs, on my env I see:
>>
>> "qemu-system-x86_64: -device vfio-pci,host=00:04.0,iommufd=iommufd0: vfio
>0000:00:04.0: Failed to allocate hwpt: Operation not supported"
>>
>> which is expected as we want nesting support in L1 but it only supports fsts.
>
>Oh yep, I forgot to mention that I'm not attaching iommufd for this test.
>I agree with what you say when iommufd is present, but without it (weird config,
>I admit, but interesting for testing), I get:
>
>```
>qemu-system-x86_64: -device vfio-pci,host=00:01.0: vfio 0000:00:01.0: Failed to
>set vIOMMU: aw-bits 48 > host aw-bits 47

Oh, I see, I reproduced the same.

>```
>
>Just wondering why such a 1bit difference.

host aw-bits 47 is what host kernel supports, not the pure hw limit.
For first stage page table, PT_FEAT_SIGN_EXTEND is set,
kernel limited max iova ranges to 47bit even though hw supports 48 bit, see below:

static __always_inline struct pt_range _pt_top_range(struct pt_common *common,
                                                     uintptr_t top_of_table)
{
...
        /*
         * The top range will default to the lower region only with sign extend.
         */
        range.max_vasz_lg2 = max_vasz_lg2;
        if (pt_feature(common, PT_FEAT_SIGN_EXTEND))
                max_vasz_lg2--;
...
}

So qemu gets host aw-bits 47. This is strictly following vtd spec:

Software using first-stage translation structures to translate an IO Virtual Address (IOVA) must use 
canonical addresses. Additionally, software must limit addresses to less than the minimum of MGAW 
and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM is 4-level and 56-bit 
when FSPM is 5-level).

BRs,
Zhenzhong
Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough device
Posted by Clément MATHIEU--DRIF 1 month ago
On Fri, 2026-08-21 at 07:15 +0000, Duan, Zhenzhong wrote:
> Caution: External email. Do not open attachments or click links, unless this email comes from a known sender and you know the content is safe.
> 
> 
> 
> > -----Original Message-----
> > From: Clément MATHIEU--DRIF <[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)>
> > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > device
> > 
> > 
> > On Fri, 2026-08-21 at 02:44 +0000, Duan, Zhenzhong wrote:
> > 
> > > Caution: External email. Do not open attachments or click links, unless this email
> > 
> > comes from a known sender and you know the content is safe.
> > 
> > > 
> > > 
> > > Hi Clement,
> > > 
> > > 
> > > 
> > > > -----Original Message-----
> > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com))>
> > 
> > > 
> > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for passthrough
> > > > device
> > > > 
> > > > 
> > > > On Wed, 2026-08-12 at 13:26 +0000, Duan, Zhenzhong wrote:
> > > > 
> > > > 
> > > > > Caution: External email. Do not open attachments or click links, unless this
> > > > 
> > > 
> > 
> > email
> > 
> > > 
> > > > 
> > > > comes from a known sender and you know the content is safe.
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > > -----Original Message-----
> > > > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > > > 
> > > > > 
> > > > 
> > > > 
> > > > [[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.com](mailto:drif@bull.com))](mailto:[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com)))>
> > 
> > > 
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
> > > > > 
> > > > 
> > > 
> > 
> > passthrough
> > 
> > > 
> > > > 
> > > > > 
> > > > > > device
> > > > > > 
> > > > > > 
> > > > > > On Mon, 2026-08-10 at 10:49 +0000, Duan, Zhenzhong wrote:
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > > Caution: External email. Do not open attachments or click links, unless
> > > > > > 
> > > > > 
> > > > 
> > > 
> > 
> > this
> > 
> > > 
> > > > 
> > > > > 
> > > > > > 
> > > > > 
> > > > > 
> > > > 
> > > > 
> > > > email
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > > 
> > > > > > comes from a known sender and you know the content is safe.
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > Hi Clement,
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > > -----Original Message-----
> > > > > > > > From: Clément MATHIEU--DRIF <[clement.mathieu--
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > 
> > > > 
> > > 
> > 
> > [[[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.com](mailto:drif@bull.com))](mailto:[[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.c](mailto:drif@bull.c)
> > om))](mailto:[clement.mathieu--
> > 
> > > 
> > > > 
> > > > > 
> > > > 
> > > > 
> > > > [[drif@bull.com](mailto:drif@bull.com)](mailto:[drif@bull.com](mailto:drif@bull.com))](mailto:[clement.mathieu--
> > > 
> > 
> > [drif@bull.com](mailto:drif@bull.com)](mailto:[clement.mathieu--drif@bull.com](mailto:clement.mathieu--drif@bull.com))))>
> > 
> > > 
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > > Subject: Re: [PATCH v3 0/6] intel_iommu: Enable PRQ support for
> > > > > > > 
> > > > > > > 
> > > > > > 
> > > > > > 
> > > > > 
> > > > > 
> > > > 
> > > > 
> > > > passthrough
> > > > 
> > > > 
> > > > > 
> > > > > 
> > > > > > 
> > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > > device
> > > > > > > > 
> > > > > > > > Hi Zhenzhong,
> > > > > > > > 
> > > > > > > > I will read this series early next week.
> > > > > > > > thanks for the respin!
> > > > > > > > 
> > > > > > > > btw, related to this work: have ever seen
> > > > > > > > vfio_device_get_aw_bits returning 47 whereas
> > > > > > > > the iommu supports 48? This leads the following
> > > > > > > > aw test to fail:
> > > > > > > > 
> > > > > > > > ```
> > > > > > > > ret = hiodc->get_cap(hiod, HOST_IOMMU_DEVICE_CAP_AW_BITS,
> > > > > > > 
> > > > > > 
> > > > > 
> > > > 
> > > 
> > 
> > errp);
> > 
> > > 
> > > > 
> > > > > 
> > > > > > 
> > > > > > > 
> > > > > > > > if (ret < 0) {
> > > > > > > >    return false;
> > > > > > > > }
> > > > > > > > if (s->aw_bits > ret) {
> > > > > > > >    error_setg(errp, "aw-bits %d > host aw-bits %d", s->aw_bits, ret);
> > > > > > > >    return false;
> > > > > > > > }
> > > > > > > > ```
> > > > > > > > 
> > > > > > > > I got this when trying to nest VMs.
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > 
> > > > > > > I am trying to reproduce, what device did you passthrough?
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > 
> > > > > > I used edu for testing.
> > > > > > An I had aw-bits=48 on both vms.
> > > > > 
> > > > > 
> > > > > 
> > > > > 
> > > > > Reproduced the same with L1(fsts=on) and L2(fsts=off).
> > > > 
> > > > 
> > > > 
> > > > My config has fsts=on on both VMs.
> > > > Which kernel are you using?
> > > > Mine is RHEL10.2 (6.12.0-211.7.3.el10_2.x86_64)
> > > 
> > > 
> > > 
> > > In theory we will not run into that check if fsts=on on both VMs, on my env I see:
> > > 
> > > "qemu-system-x86_64: -device vfio-pci,host=00:04.0,iommufd=iommufd0: vfio
> > 
> > 0000:00:04.0: Failed to allocate hwpt: Operation not supported"
> > 
> > > 
> > > which is expected as we want nesting support in L1 but it only supports fsts.
> > 
> > 
> > Oh yep, I forgot to mention that I'm not attaching iommufd for this test.
> > I agree with what you say when iommufd is present, but without it (weird config,
> > I admit, but interesting for testing), I get:
> > 
> > ```
> > qemu-system-x86_64: -device vfio-pci,host=00:01.0: vfio 0000:00:01.0: Failed to
> > set vIOMMU: aw-bits 48 > host aw-bits 47
> 
> 
> Oh, I see, I reproduced the same.
> 
> 
> > ```
> > 
> > Just wondering why such a 1bit difference.
> 
> 
> host aw-bits 47 is what host kernel supports, not the pure hw limit.  
> For first stage page table, PT_FEAT_SIGN_EXTEND is set,  
> kernel limited max iova ranges to 47bit even though hw supports 48 bit, see below:
> 
> static __always_inline struct pt_range _pt_top_range(struct pt_common *common,  
>                                                      uintptr_t top_of_table)  
> {  
> ...  
>         /*  
>          * The top range will default to the lower region only with sign extend.  
>          */  
>         range.max_vasz_lg2 = max_vasz_lg2;  
>         if (pt_feature(common, PT_FEAT_SIGN_EXTEND))  
>                 max_vasz_lg2--;  
> ...  
> }
> 
> So qemu gets host aw-bits 47. This is strictly following vtd spec:
> 
> Software using first-stage translation structures to translate an IO Virtual Address (IOVA) must use  
> canonical addresses. Additionally, software must limit addresses to less than the minimum of MGAW  
> and the lower canonical address width implied by FSPM (i.e., 47-bit when FSPM is 4-level and 56-bit  
> when FSPM is 5-level).

Oh, I see, good to know!  
That makes sense.  

Thanks!

cmd

> 
> BRs,  
> Zhenzhong