[PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem

Marc-André Lureau posted 12 patches 1 month ago
Patches applied successfully (tree, apply log)
git fetch https://github.com/patchew-project/qemu tags/patchew/20260604-rdm5-v5-0-5768e6a0943d@redhat.com
Maintainers: Paolo Bonzini <pbonzini@redhat.com>, "Dr. David Alan Gilbert" <dave@treblig.org>, "Maciej S. Szmigiero" <maciej.szmigiero@oracle.com>, Peter Xu <peterx@redhat.com>, Fabiano Rosas <farosas@suse.de>, Mark Kanda <mark.kanda@oracle.com>, Ben Chaney <bchaney@akamai.com>, Alex Williamson <alex@shazbot.org>, "Cédric Le Goater" <clg@redhat.com>, "Michael S. Tsirkin" <mst@redhat.com>, David Hildenbrand <david@kernel.org>, "Philippe Mathieu-Daudé" <philmd@mailo.com>, Zhao Liu <zhao1.liu@intel.com>, Eric Blake <eblake@redhat.com>, Markus Armbruster <armbru@redhat.com>, Marcelo Tosatti <mtosatti@redhat.com>
MAINTAINERS                                 |    4 +
qapi/machine.json                           |   56 ++
include/hw/vfio/vfio-container.h            |    2 +-
include/hw/vfio/vfio-cpr.h                  |    2 +-
include/hw/virtio/virtio-mem.h              |    3 -
include/monitor/hmp.h                       |    1 +
include/system/memory.h                     |  279 +-----
include/system/ram-discard-manager.h        |  358 ++++++++
include/system/ramblock.h                   |    5 +-
accel/kvm/kvm-all.c                         |    2 +-
hw/core/machine-hmp-cmds.c                  |   32 +
hw/vfio/cpr-legacy.c                        |    4 +-
hw/vfio/listener.c                          |   10 +-
hw/virtio/virtio-mem.c                      |  259 +-----
migration/ram.c                             |    6 +-
system/memory.c                             |   88 +-
system/memory_mapping.c                     |    4 +-
system/physmem.c                            |   27 +-
system/ram-block-attributes.c               |  328 +++----
system/ram-discard-manager.c                |  612 +++++++++++++
target/i386/kvm/tdx.c                       |    2 +-
tests/unit/test-ram-discard-manager-stubs.c |   48 ++
tests/unit/test-ram-discard-manager.c       | 1235 +++++++++++++++++++++++++++
hmp-commands-info.hx                        |   13 +
rust/bindings/system-sys/lib.rs             |    2 +-
system/meson.build                          |    1 +
system/trace-events                         |    2 +-
tests/unit/meson.build                      |    8 +-
28 files changed, 2569 insertions(+), 824 deletions(-)
[PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Marc-André Lureau 1 month ago
Hi,

This is an attempt to fix the incompatibility of virtio-mem with confidential
VMs. The solution implements what was discussed earlier with D. Hildenbrand:
https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238

The first patches are misc cleanups. Then some code refactoring to have split a
manager/source. And finally, the manager learns to deal with multiple sources.

This has been tested together with the Linux kernel series from
Zhenzhong Duan [1] for TDX guests.

(help fix https://issues.redhat.com/browse/RHEL-131968)

v5:
- explain why InterfaceClass in rust binding is not needed after
  moving RamDiscardSource
- drop extraneous added blank line
- replaced a runtime "if (has_rdm)"" with an assert()
- dropped "RFC: hw/virtio: start virtio-mem guest_memfd regions as
  shared", in favour of start-private memory approach from "[RFC PATCH
  0/6] Support virtio-mem memory hotplug in TDX guests" [1]
- demote "monitor: add 'info ramblock-attributes'" as RFC
- rebased and collected trailers

v4:
 - added "system/physmem: make ram_block_discard_range() handle guest_memfd"
 - added "monitor: add 'info ramblock-attributes' command"
 - added "RFC: hw/virtio: start virtio-mem guest_memfd regions as shared"
 - skip calling source in notify_populate (it may not have updated its
   internal state)
 - rebased, collected trailer tags

v3: issues found by Cédric
 - fix assertion error on shutdown, due to rcu-defer cleanup
 - fix API doc warnings

v2:
 - drop replay_{populated,discarded} from source, suggested by Peter Xu
 - add extra manager cleanup
 - add r-b tags for preliminary patches

- Link to v4: https://lore.kernel.org/qemu-devel/20260504-rdm5-v4-0-bdf61e57c1e1@redhat.com

[1]: https://lore.kernel.org/lkml/20260604093551.1511079-1-zhenzhong.duan@intel.com/

---
Marc-André Lureau (12):
      system/memory: split RamDiscardManager into source and manager
      system/memory: move RamDiscardManager to separate compilation unit
      system/memory: constify section arguments
      system/ram-discard-manager: implement replay via is_populated iteration
      virtio-mem: remove replay_populated/replay_discarded implementation
      system/ram-discard-manager: drop replay from source interface
      system/memory: implement RamDiscardManager multi-source aggregation
      system/physmem: destroy ram block attributes before RCU-deferred reclaim
      system/memory: add RamDiscardManager reference counting and cleanup
      tests: add unit tests for RamDiscardManager multi-source aggregation
      system/physmem: make ram_block_discard_range() handle guest_memfd
      RFC: monitor: add 'info ramblock-attributes' command

 MAINTAINERS                                 |    4 +
 qapi/machine.json                           |   56 ++
 include/hw/vfio/vfio-container.h            |    2 +-
 include/hw/vfio/vfio-cpr.h                  |    2 +-
 include/hw/virtio/virtio-mem.h              |    3 -
 include/monitor/hmp.h                       |    1 +
 include/system/memory.h                     |  279 +-----
 include/system/ram-discard-manager.h        |  358 ++++++++
 include/system/ramblock.h                   |    5 +-
 accel/kvm/kvm-all.c                         |    2 +-
 hw/core/machine-hmp-cmds.c                  |   32 +
 hw/vfio/cpr-legacy.c                        |    4 +-
 hw/vfio/listener.c                          |   10 +-
 hw/virtio/virtio-mem.c                      |  259 +-----
 migration/ram.c                             |    6 +-
 system/memory.c                             |   88 +-
 system/memory_mapping.c                     |    4 +-
 system/physmem.c                            |   27 +-
 system/ram-block-attributes.c               |  328 +++----
 system/ram-discard-manager.c                |  612 +++++++++++++
 target/i386/kvm/tdx.c                       |    2 +-
 tests/unit/test-ram-discard-manager-stubs.c |   48 ++
 tests/unit/test-ram-discard-manager.c       | 1235 +++++++++++++++++++++++++++
 hmp-commands-info.hx                        |   13 +
 rust/bindings/system-sys/lib.rs             |    2 +-
 system/meson.build                          |    1 +
 system/trace-events                         |    2 +-
 tests/unit/meson.build                      |    8 +-
 28 files changed, 2569 insertions(+), 824 deletions(-)
---
base-commit: 405c32d2b18a683ba36301351af75125d9afda08
change-id: 20260414-rdm5-b6df2366d603

Best regards,
--  
Marc-André Lureau <marcandre.lureau@redhat.com>


Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Marc-André Lureau 3 weeks ago
Hi

On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
<marcandre.lureau@redhat.com> wrote:
>
> Hi,
>
> This is an attempt to fix the incompatibility of virtio-mem with confidential
> VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
>
> The first patches are misc cleanups. Then some code refactoring to have split a
> manager/source. And finally, the manager learns to deal with multiple sources.
>
> This has been tested together with the Linux kernel series from
> Zhenzhong Duan [1] for TDX guests.
>
> (help fix https://issues.redhat.com/browse/RHEL-131968)

Can the patch 1-11 be queued or are we missing something?
(RFC patch 12 can be dropped for now)

thanks

>
> v5:
> - explain why InterfaceClass in rust binding is not needed after
>   moving RamDiscardSource
> - drop extraneous added blank line
> - replaced a runtime "if (has_rdm)"" with an assert()
> - dropped "RFC: hw/virtio: start virtio-mem guest_memfd regions as
>   shared", in favour of start-private memory approach from "[RFC PATCH
>   0/6] Support virtio-mem memory hotplug in TDX guests" [1]
> - demote "monitor: add 'info ramblock-attributes'" as RFC
> - rebased and collected trailers
>
> v4:
>  - added "system/physmem: make ram_block_discard_range() handle guest_memfd"
>  - added "monitor: add 'info ramblock-attributes' command"
>  - added "RFC: hw/virtio: start virtio-mem guest_memfd regions as shared"
>  - skip calling source in notify_populate (it may not have updated its
>    internal state)
>  - rebased, collected trailer tags
>
> v3: issues found by Cédric
>  - fix assertion error on shutdown, due to rcu-defer cleanup
>  - fix API doc warnings
>
> v2:
>  - drop replay_{populated,discarded} from source, suggested by Peter Xu
>  - add extra manager cleanup
>  - add r-b tags for preliminary patches
>
> - Link to v4: https://lore.kernel.org/qemu-devel/20260504-rdm5-v4-0-bdf61e57c1e1@redhat.com
>
> [1]: https://lore.kernel.org/lkml/20260604093551.1511079-1-zhenzhong.duan@intel.com/
>
> ---
> Marc-André Lureau (12):
>       system/memory: split RamDiscardManager into source and manager
>       system/memory: move RamDiscardManager to separate compilation unit
>       system/memory: constify section arguments
>       system/ram-discard-manager: implement replay via is_populated iteration
>       virtio-mem: remove replay_populated/replay_discarded implementation
>       system/ram-discard-manager: drop replay from source interface
>       system/memory: implement RamDiscardManager multi-source aggregation
>       system/physmem: destroy ram block attributes before RCU-deferred reclaim
>       system/memory: add RamDiscardManager reference counting and cleanup
>       tests: add unit tests for RamDiscardManager multi-source aggregation
>       system/physmem: make ram_block_discard_range() handle guest_memfd
>       RFC: monitor: add 'info ramblock-attributes' command
>
>  MAINTAINERS                                 |    4 +
>  qapi/machine.json                           |   56 ++
>  include/hw/vfio/vfio-container.h            |    2 +-
>  include/hw/vfio/vfio-cpr.h                  |    2 +-
>  include/hw/virtio/virtio-mem.h              |    3 -
>  include/monitor/hmp.h                       |    1 +
>  include/system/memory.h                     |  279 +-----
>  include/system/ram-discard-manager.h        |  358 ++++++++
>  include/system/ramblock.h                   |    5 +-
>  accel/kvm/kvm-all.c                         |    2 +-
>  hw/core/machine-hmp-cmds.c                  |   32 +
>  hw/vfio/cpr-legacy.c                        |    4 +-
>  hw/vfio/listener.c                          |   10 +-
>  hw/virtio/virtio-mem.c                      |  259 +-----
>  migration/ram.c                             |    6 +-
>  system/memory.c                             |   88 +-
>  system/memory_mapping.c                     |    4 +-
>  system/physmem.c                            |   27 +-
>  system/ram-block-attributes.c               |  328 +++----
>  system/ram-discard-manager.c                |  612 +++++++++++++
>  target/i386/kvm/tdx.c                       |    2 +-
>  tests/unit/test-ram-discard-manager-stubs.c |   48 ++
>  tests/unit/test-ram-discard-manager.c       | 1235 +++++++++++++++++++++++++++
>  hmp-commands-info.hx                        |   13 +
>  rust/bindings/system-sys/lib.rs             |    2 +-
>  system/meson.build                          |    1 +
>  system/trace-events                         |    2 +-
>  tests/unit/meson.build                      |    8 +-
>  28 files changed, 2569 insertions(+), 824 deletions(-)
> ---
> base-commit: 405c32d2b18a683ba36301351af75125d9afda08
> change-id: 20260414-rdm5-b6df2366d603
>
> Best regards,
> --
> Marc-André Lureau <marcandre.lureau@redhat.com>
>
>
Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Peter Xu 3 weeks ago
On Fri, Jun 19, 2026 at 12:11:48AM +0400, Marc-André Lureau wrote:
> Hi
> 
> On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
> <marcandre.lureau@redhat.com> wrote:
> >
> > Hi,
> >
> > This is an attempt to fix the incompatibility of virtio-mem with confidential
> > VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> > https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
> >
> > The first patches are misc cleanups. Then some code refactoring to have split a
> > manager/source. And finally, the manager learns to deal with multiple sources.
> >
> > This has been tested together with the Linux kernel series from
> > Zhenzhong Duan [1] for TDX guests.
> >
> > (help fix https://issues.redhat.com/browse/RHEL-131968)
> 
> Can the patch 1-11 be queued or are we missing something?
> (RFC patch 12 can be dropped for now)

Likely yes.. one thing to double check with you before I do: We don't need
the kernel series, do we?  Since when unplug, I expect with the truncation
approach that this series proposed, KVM will emit TDH.MEM.PAGE.REMOVE then
unaccept is done (?).

Say, what happens if we run QEMU with this series applied, but without the
kernel series?

What confused me a bit is the dependency of this series v.s. the kernel
one.  It seems to use different approaches, but then I don't understand why
this series was tested with the kernel change.

Thanks,

-- 
Peter Xu


Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Marc-André Lureau 2 weeks, 4 days ago
Hi Peter

On Fri, Jun 19, 2026 at 7:13 PM Peter Xu <peterx@redhat.com> wrote:
>
> On Fri, Jun 19, 2026 at 12:11:48AM +0400, Marc-André Lureau wrote:
> > Hi
> >
> > On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
> > <marcandre.lureau@redhat.com> wrote:
> > >
> > > Hi,
> > >
> > > This is an attempt to fix the incompatibility of virtio-mem with confidential
> > > VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> > > https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
> > >
> > > The first patches are misc cleanups. Then some code refactoring to have split a
> > > manager/source. And finally, the manager learns to deal with multiple sources.
> > >
> > > This has been tested together with the Linux kernel series from
> > > Zhenzhong Duan [1] for TDX guests.
> > >
> > > (help fix https://issues.redhat.com/browse/RHEL-131968)
> >
> > Can the patch 1-11 be queued or are we missing something?
> > (RFC patch 12 can be dropped for now)
>
> Likely yes.. one thing to double check with you before I do: We don't need
> the kernel series, do we?  Since when unplug, I expect with the truncation
> approach that this series proposed, KVM will emit TDH.MEM.PAGE.REMOVE then
> unaccept is done (?).

> Say, what happens if we run QEMU with this series applied, but without the
> kernel series?

The kernel series is needed at least for PAGE.ACCEPT. Without it, QEMU
will have KVM_RUN return EIO, and finish into assert (while tearing
down ioeventfd).

> What confused me a bit is the dependency of this series v.s. the kernel
> one.  It seems to use different approaches, but then I don't understand why
> this series was tested with the kernel change.

My understanding is that the kernel may perform TDG.MEM.PAGE.RELEASE.
That depends on TDX config TDCS_CONFIG_PAGE_RELEASE which qemu/kvm
doesnt currently control. I don't know whether this is then
redundant/needless with qemu doing discard on the guest_memfd..
Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Peter Xu 2 weeks, 4 days ago
On Mon, Jun 22, 2026 at 03:53:33PM +0400, Marc-André Lureau wrote:
> Hi Peter
> 
> On Fri, Jun 19, 2026 at 7:13 PM Peter Xu <peterx@redhat.com> wrote:
> >
> > On Fri, Jun 19, 2026 at 12:11:48AM +0400, Marc-André Lureau wrote:
> > > Hi
> > >
> > > On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
> > > <marcandre.lureau@redhat.com> wrote:
> > > >
> > > > Hi,
> > > >
> > > > This is an attempt to fix the incompatibility of virtio-mem with confidential
> > > > VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> > > > https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
> > > >
> > > > The first patches are misc cleanups. Then some code refactoring to have split a
> > > > manager/source. And finally, the manager learns to deal with multiple sources.
> > > >
> > > > This has been tested together with the Linux kernel series from
> > > > Zhenzhong Duan [1] for TDX guests.
> > > >
> > > > (help fix https://issues.redhat.com/browse/RHEL-131968)
> > >
> > > Can the patch 1-11 be queued or are we missing something?
> > > (RFC patch 12 can be dropped for now)
> >
> > Likely yes.. one thing to double check with you before I do: We don't need
> > the kernel series, do we?  Since when unplug, I expect with the truncation
> > approach that this series proposed, KVM will emit TDH.MEM.PAGE.REMOVE then
> > unaccept is done (?).
> 
> > Say, what happens if we run QEMU with this series applied, but without the
> > kernel series?
> 
> The kernel series is needed at least for PAGE.ACCEPT. Without it, QEMU
> will have KVM_RUN return EIO, and finish into assert (while tearing
> down ioeventfd).

Could you elaborate a bit more on why ACCEPT would fail?

I used to ask the event flow here:

https://lore.kernel.org/qemu-devel/agTjb3M8ElUAlfp1@x1.local/

If AUG existed, then why ACCEPT would fail?

PS: I didn't read a lot of what a Linux guest would do; I know there're
some lazy accept approach, but IIUC it's only a matter of time to ACCEPT,
not correctness.  My understanding is if we properly notify these new slots
with AUG, then it should be able to ACCEPT?

> 
> > What confused me a bit is the dependency of this series v.s. the kernel
> > one.  It seems to use different approaches, but then I don't understand why
> > this series was tested with the kernel change.
> 
> My understanding is that the kernel may perform TDG.MEM.PAGE.RELEASE.
> That depends on TDX config TDCS_CONFIG_PAGE_RELEASE which qemu/kvm
> doesnt currently control. I don't know whether this is then
> redundant/needless with qemu doing discard on the guest_memfd..

The problem is if this series depends on the kernel series, should we then
wait for the kernel solution be accepted first in case it'll change?

But obviously I still don't yet fully understand how this whole thing
works.. :(

Thanks,

-- 
Peter Xu


Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Marc-André Lureau 2 weeks, 4 days ago
Hi Peter

On Mon, Jun 22, 2026 at 11:28 PM Peter Xu <peterx@redhat.com> wrote:
>
> On Mon, Jun 22, 2026 at 03:53:33PM +0400, Marc-André Lureau wrote:
> > Hi Peter
> >
> > On Fri, Jun 19, 2026 at 7:13 PM Peter Xu <peterx@redhat.com> wrote:
> > >
> > > On Fri, Jun 19, 2026 at 12:11:48AM +0400, Marc-André Lureau wrote:
> > > > Hi
> > > >
> > > > On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
> > > > <marcandre.lureau@redhat.com> wrote:
> > > > >
> > > > > Hi,
> > > > >
> > > > > This is an attempt to fix the incompatibility of virtio-mem with confidential
> > > > > VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> > > > > https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
> > > > >
> > > > > The first patches are misc cleanups. Then some code refactoring to have split a
> > > > > manager/source. And finally, the manager learns to deal with multiple sources.
> > > > >
> > > > > This has been tested together with the Linux kernel series from
> > > > > Zhenzhong Duan [1] for TDX guests.
> > > > >
> > > > > (help fix https://issues.redhat.com/browse/RHEL-131968)
> > > >
> > > > Can the patch 1-11 be queued or are we missing something?
> > > > (RFC patch 12 can be dropped for now)
> > >
> > > Likely yes.. one thing to double check with you before I do: We don't need
> > > the kernel series, do we?  Since when unplug, I expect with the truncation
> > > approach that this series proposed, KVM will emit TDH.MEM.PAGE.REMOVE then
> > > unaccept is done (?).
> >
> > > Say, what happens if we run QEMU with this series applied, but without the
> > > kernel series?
> >
> > The kernel series is needed at least for PAGE.ACCEPT. Without it, QEMU
> > will have KVM_RUN return EIO, and finish into assert (while tearing
> > down ioeventfd).
>
> Could you elaborate a bit more on why ACCEPT would fail?
>
> I used to ask the event flow here:
>
> https://lore.kernel.org/qemu-devel/agTjb3M8ElUAlfp1@x1.local/
>
> If AUG existed, then why ACCEPT would fail?
>
> PS: I didn't read a lot of what a Linux guest would do; I know there're
> some lazy accept approach, but IIUC it's only a matter of time to ACCEPT,
> not correctness.  My understanding is if we properly notify these new slots
> with AUG, then it should be able to ACCEPT?

Yes, it will, but it needs the kernel patches to do it (or virtio-mem
will let the guest access un-accepted pages and qemu will crash). I
submitted a few proposals before Zhenzhong Duan proposed the last
iteration.

> >
> > > What confused me a bit is the dependency of this series v.s. the kernel
> > > one.  It seems to use different approaches, but then I don't understand why
> > > this series was tested with the kernel change.
> >
> > My understanding is that the kernel may perform TDG.MEM.PAGE.RELEASE.
> > That depends on TDX config TDCS_CONFIG_PAGE_RELEASE which qemu/kvm
> > doesnt currently control. I don't know whether this is then
> > redundant/needless with qemu doing discard on the guest_memfd..
>
> The problem is if this series depends on the kernel series, should we then
> wait for the kernel solution be accepted first in case it'll change?

I don't think qemu needs to wait for the kernel to be fixed.
Furthermore, this series is not tdx/sev/etc specific

>
> But obviously I still don't yet fully understand how this whole thing
> works.. :(

I am also slow, because I don't focus on this most of the time.
Someone with more experience and dedication would handle it better.
Re: [PATCH v5 00/12] Make RamDiscardManager work with multiple sources & virtio-mem
Posted by Peter Xu 2 weeks, 3 days ago
On Tue, Jun 23, 2026 at 12:00:04AM +0400, Marc-André Lureau wrote:
> Hi Peter
> 
> On Mon, Jun 22, 2026 at 11:28 PM Peter Xu <peterx@redhat.com> wrote:
> >
> > On Mon, Jun 22, 2026 at 03:53:33PM +0400, Marc-André Lureau wrote:
> > > Hi Peter
> > >
> > > On Fri, Jun 19, 2026 at 7:13 PM Peter Xu <peterx@redhat.com> wrote:
> > > >
> > > > On Fri, Jun 19, 2026 at 12:11:48AM +0400, Marc-André Lureau wrote:
> > > > > Hi
> > > > >
> > > > > On Thu, Jun 4, 2026 at 5:46 PM Marc-André Lureau
> > > > > <marcandre.lureau@redhat.com> wrote:
> > > > > >
> > > > > > Hi,
> > > > > >
> > > > > > This is an attempt to fix the incompatibility of virtio-mem with confidential
> > > > > > VMs. The solution implements what was discussed earlier with D. Hildenbrand:
> > > > > > https://patchwork.ozlabs.org/project/qemu-devel/patch/20250407074939.18657-5-chenyi.qiang@intel.com/#3502238
> > > > > >
> > > > > > The first patches are misc cleanups. Then some code refactoring to have split a
> > > > > > manager/source. And finally, the manager learns to deal with multiple sources.
> > > > > >
> > > > > > This has been tested together with the Linux kernel series from
> > > > > > Zhenzhong Duan [1] for TDX guests.
> > > > > >
> > > > > > (help fix https://issues.redhat.com/browse/RHEL-131968)
> > > > >
> > > > > Can the patch 1-11 be queued or are we missing something?
> > > > > (RFC patch 12 can be dropped for now)
> > > >
> > > > Likely yes.. one thing to double check with you before I do: We don't need
> > > > the kernel series, do we?  Since when unplug, I expect with the truncation
> > > > approach that this series proposed, KVM will emit TDH.MEM.PAGE.REMOVE then
> > > > unaccept is done (?).
> > >
> > > > Say, what happens if we run QEMU with this series applied, but without the
> > > > kernel series?
> > >
> > > The kernel series is needed at least for PAGE.ACCEPT. Without it, QEMU
> > > will have KVM_RUN return EIO, and finish into assert (while tearing
> > > down ioeventfd).
> >
> > Could you elaborate a bit more on why ACCEPT would fail?
> >
> > I used to ask the event flow here:
> >
> > https://lore.kernel.org/qemu-devel/agTjb3M8ElUAlfp1@x1.local/
> >
> > If AUG existed, then why ACCEPT would fail?
> >
> > PS: I didn't read a lot of what a Linux guest would do; I know there're
> > some lazy accept approach, but IIUC it's only a matter of time to ACCEPT,
> > not correctness.  My understanding is if we properly notify these new slots
> > with AUG, then it should be able to ACCEPT?
> 
> Yes, it will, but it needs the kernel patches to do it (or virtio-mem
> will let the guest access un-accepted pages and qemu will crash). I
> submitted a few proposals before Zhenzhong Duan proposed the last
> iteration.

OK, I think I misunderstood both the crash and also what Zhenzhong's series
is trying to do.  After a closer look, it seems to be a proposal for any
plug/unplug to work.  I'm surprised (if I get it right this time..) that TD
didn't seem to support DIMM hotplugs besides virtio-mem.

> 
> > >
> > > > What confused me a bit is the dependency of this series v.s. the kernel
> > > > one.  It seems to use different approaches, but then I don't understand why
> > > > this series was tested with the kernel change.
> > >
> > > My understanding is that the kernel may perform TDG.MEM.PAGE.RELEASE.
> > > That depends on TDX config TDCS_CONFIG_PAGE_RELEASE which qemu/kvm
> > > doesnt currently control. I don't know whether this is then
> > > redundant/needless with qemu doing discard on the guest_memfd..
> >
> > The problem is if this series depends on the kernel series, should we then
> > wait for the kernel solution be accepted first in case it'll change?
> 
> I don't think qemu needs to wait for the kernel to be fixed.
> Furthermore, this series is not tdx/sev/etc specific

AFAIU it is coco specific, otherwise we only always have 1 source, hence no
need for this series to provide this function.

Said that, I agree with you.. Looks like there's no major plan to add
anything specific to QEMU.

Patch 1-10 are all reviewed patches and correctly resolve the known >1 ram
sources and it's a design problem, I don't see how we can bypass that.

Patch 11 is very reasonable on its own, then I assume the guest ACCEPT
problem will need to evolve on its own.  Looks like the direction is
correct, and only some corner cases to think about (acpi coverage, lazy
accept, etc.) that I saw in the discussion.

> 
> >
> > But obviously I still don't yet fully understand how this whole thing
> > works.. :(
> 
> I am also slow, because I don't focus on this most of the time.
> Someone with more experience and dedication would handle it better.

I queued patch 1-11, will send a pull this week.  Thanks for all the
explanations.

-- 
Peter Xu