arch/arm64/kvm/mmu.c | 4 ++-- arch/arm64/kvm/nested.c | 4 ++-- arch/x86/kvm/mmu/mmu.c | 2 +- arch/x86/kvm/svm/sev.c | 45 +++++++++++++++++++++++++-------------------- include/linux/kvm_host.h | 6 ++---- virt/kvm/guest_memfd.c | 9 ++------- 6 files changed, 34 insertions(+), 36 deletions(-)
KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
VMSA / RMP handlers) hold this refcount across page fault handling.
Holding a page refcount across fault handling is problematic for guest_memfd.
In-place memory conversions between confidential computing shared and private
states inspect folio refcounts to ensure exclusive ownership by guest_memfd. A
concurrent guest page fault taking a reference on the folio causes conversions
to fail due to an elevated refcount.
guest_memfd already notifies KVM of page invalidations, so users of guest_memfd
within KVM only need to respect the MMU invalidation protocol to safely rely on
guest_memfd to ensure page presence.
This series first prepares the SEV-SNP handlers by treating unassigned RMP
entries as benign races on PSMASH failure (which can occur on concurrent
truncation) and dropping page references early in the RMP fault and VMSA reload
paths. It then updates kvm_gmem_get_pfn() to drop the folio reference internally
and stop returning a struct page pointer across x86 and arm64.
Removing struct page from kvm_gmem_get_pfn() also moves KVM closer toward
supporting memory backends that are not backed by struct page.
I really want in-place conversions to merge in time for 7.4 and so I went
ahead to try this, building off Sean's sample code [1].
[1] https://lore.kernel.org/all/an5RJYTwlYeym--O@google.com/
I also split the patch up so it's easier to review :)
Changes from v1:
+ sev_handle_rmp_fault() does need to adopt the MMU invalidation protocol,
please see reason in patch description.
v1: https://patch.msgid.link/20260818-gmem-no-return-page-v1-0-4f8d939efdbc@google.com
Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
Ackerley Tng (2):
KVM: SEV: Treat unassigned RMP entry as benign race on PSMASH failure
KVM: SEV: Drop page refcount early in VMSA reload
Sean Christopherson (2):
KVM: SEV: Drop page refcount early during RMP fault handling
KVM: guest_memfd: Stop returning struct page from PFN lookup
arch/arm64/kvm/mmu.c | 4 ++--
arch/arm64/kvm/nested.c | 4 ++--
arch/x86/kvm/mmu/mmu.c | 2 +-
arch/x86/kvm/svm/sev.c | 45 +++++++++++++++++++++++++--------------------
include/linux/kvm_host.h | 6 ++----
virt/kvm/guest_memfd.c | 9 ++-------
6 files changed, 34 insertions(+), 36 deletions(-)
---
base-commit: 1b731e5ded480bd1e5546aed35584238661ce72e
change-id: 20260818-gmem-no-return-page-614927a29f97
Best regards,
--
Ackerley Tng <ackerleytng@google.com>
On 8/18/26 11:15, Ackerley Tng wrote: > KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct > page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP > VMSA / RMP handlers) hold this refcount across page fault handling. > > Holding a page refcount across fault handling is problematic for guest_memfd. > In-place memory conversions between confidential computing shared and private > states inspect folio refcounts to ensure exclusive ownership by guest_memfd. A > concurrent guest page fault taking a reference on the folio causes conversions > to fail due to an elevated refcount. Right. Won't we still, at least temporarily, grab a reference while looking up the folio in the page cache, or will we be preventing that concurrent race with locking? > > guest_memfd already notifies KVM of page invalidations, so users of guest_memfd > within KVM only need to respect the MMU invalidation protocol to safely rely on > guest_memfd to ensure page presence. Yes, the invalidation protocol is the crucial part. If we get that wrong, we're in holy CVE land. For GUP-fast, there was a similar discussion with MMU notifiers, but to this day, KVM actually grabs+drops references. [...] > Removing struct page from kvm_gmem_get_pfn() also moves KVM closer toward > supporting memory backends that are not backed by struct page. Agreed, they should not be messing with the struct page at all. -- Cheers, David
On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote: > On 8/18/26 11:15, Ackerley Tng wrote: > > KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct > > page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP > > VMSA / RMP handlers) hold this refcount across page fault handling. > > > > Holding a page refcount across fault handling is problematic for guest_memfd. > > In-place memory conversions between confidential computing shared and private > > states inspect folio refcounts to ensure exclusive ownership by guest_memfd. A > > concurrent guest page fault taking a reference on the folio causes conversions > > to fail due to an elevated refcount. > > Right. Won't we still, at least temporarily, grab a reference while looking up > the folio in the page cache, or will we be preventing that concurrent race with > locking? The latter. What I want to aim for is that if the relevant guest_memfd range has never been mmap()'d and there are no memory failures, then conversion is guaranteed to not fail due to elevated refcounts. Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT* required to perform an in-place conversion.
On 8/18/26 21:55, Sean Christopherson wrote: > On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote: >> On 8/18/26 11:15, Ackerley Tng wrote: >>> KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct >>> page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP >>> VMSA / RMP handlers) hold this refcount across page fault handling. >>> >>> Holding a page refcount across fault handling is problematic for guest_memfd. >>> In-place memory conversions between confidential computing shared and private >>> states inspect folio refcounts to ensure exclusive ownership by guest_memfd. A >>> concurrent guest page fault taking a reference on the folio causes conversions >>> to fail due to an elevated refcount. >> >> Right. Won't we still, at least temporarily, grab a reference while looking up >> the folio in the page cache, or will we be preventing that concurrent race with >> locking? > > The latter. What I want to aim for is that if the relevant guest_memfd range > has never been mmap()'d and there are no memory failures, then conversion is > guaranteed to not fail due to elevated refcounts. > > Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT* > required to perform an in-place conversion. Having the VM access a page that is currently under conversion (triggered by the VM) should not be the common case, no? Except, prefaulting, of course. -- Cheers, David
On Wed, Aug 19, 2026, David Hildenbrand (Arm) wrote: > On 8/18/26 21:55, Sean Christopherson wrote: > > On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote: > >> On 8/18/26 11:15, Ackerley Tng wrote: > >>> KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct > >>> page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP > >>> VMSA / RMP handlers) hold this refcount across page fault handling. > >>> > >>> Holding a page refcount across fault handling is problematic for guest_memfd. > >>> In-place memory conversions between confidential computing shared and private > >>> states inspect folio refcounts to ensure exclusive ownership by guest_memfd. A > >>> concurrent guest page fault taking a reference on the folio causes conversions > >>> to fail due to an elevated refcount. > >> > >> Right. Won't we still, at least temporarily, grab a reference while looking up > >> the folio in the page cache, or will we be preventing that concurrent race with > >> locking? > > > > The latter. What I want to aim for is that if the relevant guest_memfd range > > has never been mmap()'d and there are no memory failures, then conversion is > > guaranteed to not fail due to elevated refcounts. > > > > Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT* > > required to perform an in-place conversion. > > Having the VM access a page that is currently under conversion (triggered by the > VM) should not be the common case, no? Except, prefaulting, of course. "not be the common case" is likely an understatement. In practice, I don't it will happen outside of guest bugs and KVM testcases. It's the testcases that I want to "unblock" though. If we commit to never having to pause vCPUs, even if the guest is misbehaving, then that gives us deterministic behavior we can validate, i.e. a way to detect similar regressions in the future. I don't expect any regressions would be super problematic, but being able to treat any failed conversion as a KVM bug (for the curated setup) mitigates the risk of death by a thousand cuts, i.e. reduces the risk of gradually degrading conversion performance because more and more transient references being taken by KVM.
© 2016 - 2026 Red Hat, Inc.