[PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup

Ackerley Tng posted 4 patches 1 month, 1 week ago
There is a newer version of this series
arch/arm64/kvm/mmu.c     |  4 ++--
arch/arm64/kvm/nested.c  |  4 ++--
arch/x86/kvm/mmu/mmu.c   |  2 +-
arch/x86/kvm/svm/sev.c   | 45 +++++++++++++++++++++++++--------------------
include/linux/kvm_host.h |  6 ++----
virt/kvm/guest_memfd.c   |  9 ++-------
6 files changed, 34 insertions(+), 36 deletions(-)
[PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup
Posted by Ackerley Tng 1 month, 1 week ago
KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
VMSA / RMP handlers) hold this refcount across page fault handling.

Holding a page refcount across fault handling is problematic for guest_memfd.
In-place memory conversions between confidential computing shared and private
states inspect folio refcounts to ensure exclusive ownership by guest_memfd.  A
concurrent guest page fault taking a reference on the folio causes conversions
to fail due to an elevated refcount.

guest_memfd already notifies KVM of page invalidations, so users of guest_memfd
within KVM only need to respect the MMU invalidation protocol to safely rely on
guest_memfd to ensure page presence.

This series first prepares the SEV-SNP handlers by treating unassigned RMP
entries as benign races on PSMASH failure (which can occur on concurrent
truncation) and dropping page references early in the RMP fault and VMSA reload
paths. It then updates kvm_gmem_get_pfn() to drop the folio reference internally
and stop returning a struct page pointer across x86 and arm64.

Removing struct page from kvm_gmem_get_pfn() also moves KVM closer toward
supporting memory backends that are not backed by struct page.

I really want in-place conversions to merge in time for 7.4 and so I went
ahead to try this, building off Sean's sample code [1].

[1] https://lore.kernel.org/all/an5RJYTwlYeym--O@google.com/

I also split the patch up so it's easier to review :)

Changes from v1:

+ sev_handle_rmp_fault() does need to adopt the MMU invalidation protocol,
  please see reason in patch description.

v1: https://patch.msgid.link/20260818-gmem-no-return-page-v1-0-4f8d939efdbc@google.com

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
Ackerley Tng (2):
      KVM: SEV: Treat unassigned RMP entry as benign race on PSMASH failure
      KVM: SEV: Drop page refcount early in VMSA reload

Sean Christopherson (2):
      KVM: SEV: Drop page refcount early during RMP fault handling
      KVM: guest_memfd: Stop returning struct page from PFN lookup

 arch/arm64/kvm/mmu.c     |  4 ++--
 arch/arm64/kvm/nested.c  |  4 ++--
 arch/x86/kvm/mmu/mmu.c   |  2 +-
 arch/x86/kvm/svm/sev.c   | 45 +++++++++++++++++++++++++--------------------
 include/linux/kvm_host.h |  6 ++----
 virt/kvm/guest_memfd.c   |  9 ++-------
 6 files changed, 34 insertions(+), 36 deletions(-)
---
base-commit: 1b731e5ded480bd1e5546aed35584238661ce72e
change-id: 20260818-gmem-no-return-page-614927a29f97

Best regards,
--  
Ackerley Tng <ackerleytng@google.com>
Re: [PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/18/26 11:15, Ackerley Tng wrote:
> KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
> page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
> VMSA / RMP handlers) hold this refcount across page fault handling.
> 
> Holding a page refcount across fault handling is problematic for guest_memfd.
> In-place memory conversions between confidential computing shared and private
> states inspect folio refcounts to ensure exclusive ownership by guest_memfd.  A
> concurrent guest page fault taking a reference on the folio causes conversions
> to fail due to an elevated refcount.

Right. Won't we still, at least temporarily, grab a reference while looking up
the folio in the page cache, or will we be preventing that concurrent race with
locking?

> 
> guest_memfd already notifies KVM of page invalidations, so users of guest_memfd
> within KVM only need to respect the MMU invalidation protocol to safely rely on
> guest_memfd to ensure page presence.

Yes, the invalidation protocol is the crucial part. If we get that wrong, we're
in holy CVE land.

For GUP-fast, there was a similar discussion with MMU notifiers, but to this
day, KVM actually grabs+drops references.

[...]

> Removing struct page from kvm_gmem_get_pfn() also moves KVM closer toward
> supporting memory backends that are not backed by struct page.

Agreed, they should not be messing with the struct page at all.

-- 
Cheers,

David
Re: [PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup
Posted by Sean Christopherson 1 month, 1 week ago
On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote:
> On 8/18/26 11:15, Ackerley Tng wrote:
> > KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
> > page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
> > VMSA / RMP handlers) hold this refcount across page fault handling.
> > 
> > Holding a page refcount across fault handling is problematic for guest_memfd.
> > In-place memory conversions between confidential computing shared and private
> > states inspect folio refcounts to ensure exclusive ownership by guest_memfd.  A
> > concurrent guest page fault taking a reference on the folio causes conversions
> > to fail due to an elevated refcount.
> 
> Right. Won't we still, at least temporarily, grab a reference while looking up
> the folio in the page cache, or will we be preventing that concurrent race with
> locking?

The latter.  What I want to aim for is that if the relevant guest_memfd range
has never been mmap()'d and there are no memory failures, then conversion is
guaranteed to not fail due to elevated refcounts.

Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT*
required to perform an in-place conversion.
Re: [PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/18/26 21:55, Sean Christopherson wrote:
> On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote:
>> On 8/18/26 11:15, Ackerley Tng wrote:
>>> KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
>>> page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
>>> VMSA / RMP handlers) hold this refcount across page fault handling.
>>>
>>> Holding a page refcount across fault handling is problematic for guest_memfd.
>>> In-place memory conversions between confidential computing shared and private
>>> states inspect folio refcounts to ensure exclusive ownership by guest_memfd.  A
>>> concurrent guest page fault taking a reference on the folio causes conversions
>>> to fail due to an elevated refcount.
>>
>> Right. Won't we still, at least temporarily, grab a reference while looking up
>> the folio in the page cache, or will we be preventing that concurrent race with
>> locking?
> 
> The latter.  What I want to aim for is that if the relevant guest_memfd range
> has never been mmap()'d and there are no memory failures, then conversion is
> guaranteed to not fail due to elevated refcounts.
> 
> Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT*
> required to perform an in-place conversion.

Having the VM access a page that is currently under conversion (triggered by the
VM) should not be the common case, no? Except, prefaulting, of course.

-- 
Cheers,

David
Re: [PATCH v2 0/4] Stop returning struct page from guest_memfd PFN lookup
Posted by Sean Christopherson 1 month, 1 week ago
On Wed, Aug 19, 2026, David Hildenbrand (Arm) wrote:
> On 8/18/26 21:55, Sean Christopherson wrote:
> > On Tue, Aug 18, 2026, David Hildenbrand (Arm) wrote:
> >> On 8/18/26 11:15, Ackerley Tng wrote:
> >>> KVM currently expects kvm_gmem_get_pfn() to return a refcounted struct
> >>> page. Callers (such as x86 TDP MMU, arm64 Stage-2 fault handler, and SEV-SNP
> >>> VMSA / RMP handlers) hold this refcount across page fault handling.
> >>>
> >>> Holding a page refcount across fault handling is problematic for guest_memfd.
> >>> In-place memory conversions between confidential computing shared and private
> >>> states inspect folio refcounts to ensure exclusive ownership by guest_memfd.  A
> >>> concurrent guest page fault taking a reference on the folio causes conversions
> >>> to fail due to an elevated refcount.
> >>
> >> Right. Won't we still, at least temporarily, grab a reference while looking up
> >> the folio in the page cache, or will we be preventing that concurrent race with
> >> locking?
> > 
> > The latter.  What I want to aim for is that if the relevant guest_memfd range
> > has never been mmap()'d and there are no memory failures, then conversion is
> > guaranteed to not fail due to elevated refcounts.
> > 
> > Or to put it a different way, I want KVM's ABI to be that pausing vCPU is *NOT*
> > required to perform an in-place conversion.
> 
> Having the VM access a page that is currently under conversion (triggered by the
> VM) should not be the common case, no? Except, prefaulting, of course.

"not be the common case" is likely an understatement.  In practice, I don't it
will happen outside of guest bugs and KVM testcases.

It's the testcases that I want to "unblock" though.  If we commit to never having
to pause vCPUs, even if the guest is misbehaving, then that gives us deterministic
behavior we can validate, i.e. a way to detect similar regressions in the future.

I don't expect any regressions would be super problematic, but being able to treat
any failed conversion as a KVM bug (for the curated setup) mitigates the risk of
death by a thousand cuts, i.e. reduces the risk of gradually degrading conversion
performance because more and more transient references being taken by KVM.