Documentation/virt/kvm/api.rst | 17 +- arch/arm64/include/asm/esr.h | 129 +++++--- arch/arm64/include/asm/kvm_emulate.h | 52 +--- arch/arm64/include/asm/kvm_pgtable.h | 5 +- arch/arm64/include/asm/kvm_pkvm.h | 2 +- arch/arm64/kvm/Kconfig | 1 + arch/arm64/kvm/arm.c | 1 + arch/arm64/kvm/hyp/nvhe/mem_protect.c | 10 +- arch/arm64/kvm/hyp/pgtable.c | 5 +- arch/arm64/kvm/mmu.c | 324 +++++++++++++++++---- arch/arm64/kvm/nested.c | 2 +- tools/testing/selftests/kvm/Makefile.kvm | 2 + .../selftests/kvm/arm64/nv_pre_fault_memory_test.c | 206 +++++++++++++ .../testing/selftests/kvm/pre_fault_memory_test.c | 152 ++++++++-- 14 files changed, 736 insertions(+), 172 deletions(-)
This series implements the KVM stage 2 page table pre-faulting feature for
arm64.
== Foundations ==
The series begins by establishing required foundations:
1. Updating kvm_s2_fault_desc to independently store the exception syndrome
register (ESR) value, and updating all code paths to use this value
exclusively.
This is needed so we can later generate a synthetic fault to perform the
pre-faulting - we need to be sure the code doesn't grab an incorrect ESR
from elsewhere.
2. Updating kvm_s2_fault_desc to independently store the kvm_s2_mmu and
updating all code paths to use this value exclusively.
Similarly this is needed so we can generate a synthetic fault against the
canonical stage 2 MMU (it would make no sense for it to touch nested shadow
page tables) - we need to be sure that the code doesn't grab an incorrect
MMU from elsewhere.
3. Updating the abort paths which consume kvm_s2_fault_desc to also return
a kvm_s2_fault_result data structure.
To perform pre-faulting the code must know the granule size of what was
just walked. So the abort paths have to tell us what that was.
4. Pass walk flags to kvm_pgtable_get_leaf() to permit walking page tables
under the MMU read lock.
This is Jack's patch, verbatim, which allows the use of the
KVM_PGTABLE_WALK_SHARED flag to walk page tables under the MMU read lock.
Pre-faulting requires it to be able to work in parallel as specified by the
API. The read lock precludes page tables being torn down behind our back.
== Implementation ==
Pre-faulting is implemented in kvm_arch_vcpu_pre_fault_memory() whose job
is to pre-fault the stage 2 page tables which map a specific GPA (which,
for arm64, is the guest's IPA).
This function is called by kvm_vcpu_pre_fault_memory() for each GPA in the
range, which itself is ultimately invoked by userland via the
KVM_PRE_FAULT_MEMORY ioctl.
The implementation is simple - try to walk to the stage 2 page table
mapping the GPA - if unmapped, fault it in through a synthetic page fault.
pKVM is not supported regardless of whether the VM is protected or
not.
This is because pKVM instantiates vCPUs upon run, but pre-faulting is
typically performed before a vCPU is run. It would be confusing and
inconsistent to error out on non-running vCPUs but to pre-fault running
ones.
Jack's original test suite is also included in the series.
== Credits ==
This series is based, with gratitude, on Jack Thomson's series and their
respins (links provided below) as well as the feedback he received.
The series includes Jack's v5 "KVM: arm64: Pass walk flags to
kvm_pgtable_get_leaf()" patch verbatim, and all three of his test patches,
two of which required minor fixups.
Link: https://patch.msgid.link/20260612162354.73378-1-jackabt.amazon@gmail.com/
Link: https://patch.msgid.link/20260113152643.18858-1-jackabt.amazon@gmail.com/
Link: https://patch.msgid.link/20251119154910.97716-1-jackabt.amazon@gmail.com/
Link: https://patch.msgid.link/20251013151502.6679-1-jackabt.amazon@gmail.com/
Link: https://patch.msgid.link/20250911134648.58945-1-jackabt.amazon@gmail.com/
== Reviewer Notes ==
I synced with maintainers on this who asked me to take a look, as there hadn't
been progress on the series for some time.
Am happy to rebase again after -rc1.
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
Jack Thomson (4):
KVM: arm64: Pass walk flags to kvm_pgtable_get_leaf()
KVM: selftests: Enable pre_fault_memory_test for arm64
KVM: selftests: Add option for different backing in pre-fault tests
KVM: selftests: Add nested pre-fault test for arm64
Lorenzo Stoakes (ARM) (4):
KVM: arm64: Propagate and use esr in s2fd when handling guest aborts
KVM: arm64: Propagate and use mmu in s2fd when handling guest aborts
KVM: arm64: Propagate and use kvm_s2_fault_result on S2 fault
KVM: arm64: Implement KVM_PRE_FAULT_MEMORY
Documentation/virt/kvm/api.rst | 17 +-
arch/arm64/include/asm/esr.h | 129 +++++---
arch/arm64/include/asm/kvm_emulate.h | 52 +---
arch/arm64/include/asm/kvm_pgtable.h | 5 +-
arch/arm64/include/asm/kvm_pkvm.h | 2 +-
arch/arm64/kvm/Kconfig | 1 +
arch/arm64/kvm/arm.c | 1 +
arch/arm64/kvm/hyp/nvhe/mem_protect.c | 10 +-
arch/arm64/kvm/hyp/pgtable.c | 5 +-
arch/arm64/kvm/mmu.c | 324 +++++++++++++++++----
arch/arm64/kvm/nested.c | 2 +-
tools/testing/selftests/kvm/Makefile.kvm | 2 +
.../selftests/kvm/arm64/nv_pre_fault_memory_test.c | 206 +++++++++++++
.../testing/selftests/kvm/pre_fault_memory_test.c | 152 ++++++++--
14 files changed, 736 insertions(+), 172 deletions(-)
---
base-commit: aa8e5dc6a7a2a1141ab40706a51010adcd0e57d2
change-id: 20260815-kvm-arm-prefault-9bb411b6897d
Best regards,
--
Lorenzo Stoakes (ARM) <ljs@kernel.org>
Hi Lorenzo, On Tue, 25 Aug 2026 at 17:01, Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > This series implements the KVM stage 2 page table pre-faulting feature for > arm64. I know there's likely going to be a V2, but FWIW, I tested this series as follows (finally managed to get protected mode running on my Mac): QEMU (arm64, -cpu max,pauth=off): - Protected-mode host, boots non-protected and protected guests - pre_fault_memory_test on a non-protected host, VHE and nVHE, all 13 guest modes - nv_pre_fault_memory_test with kvm-arm.mode=nested FVP (AEM, FEAT_NV2): - nv_pre_fault_memory_test with kvm-arm.mode=nested Apple M4 hardware, running as a KVM host at EL2 under Apple's nested virtualization (nVHE only): - pre_fault_memory_test, anonymous and anonymous_thp backing, over the five guest modes available there (the IPA is capped at 40 bits and the 64K granule isn't implemented) - with kvm-arm.mode=protected, pre_fault_memory_test skips on KVM_CAP_PRE_FAULT_MEMORY, so the pKVM exclusion holds on a live pKVM host Tested-by: Fuad Tabba < fuad.tabba@linux.dev> (happy to retest again when you respin) Cheers, /fuad > > == Foundations == > > The series begins by establishing required foundations: > > 1. Updating kvm_s2_fault_desc to independently store the exception syndrome > register (ESR) value, and updating all code paths to use this value > exclusively. > > This is needed so we can later generate a synthetic fault to perform the > pre-faulting - we need to be sure the code doesn't grab an incorrect ESR > from elsewhere. > > 2. Updating kvm_s2_fault_desc to independently store the kvm_s2_mmu and > updating all code paths to use this value exclusively. > > Similarly this is needed so we can generate a synthetic fault against the > canonical stage 2 MMU (it would make no sense for it to touch nested shadow > page tables) - we need to be sure that the code doesn't grab an incorrect > MMU from elsewhere. > > 3. Updating the abort paths which consume kvm_s2_fault_desc to also return > a kvm_s2_fault_result data structure. > > To perform pre-faulting the code must know the granule size of what was > just walked. So the abort paths have to tell us what that was. > > 4. Pass walk flags to kvm_pgtable_get_leaf() to permit walking page tables > under the MMU read lock. > > This is Jack's patch, verbatim, which allows the use of the > KVM_PGTABLE_WALK_SHARED flag to walk page tables under the MMU read lock. > > Pre-faulting requires it to be able to work in parallel as specified by the > API. The read lock precludes page tables being torn down behind our back. > > == Implementation == > > Pre-faulting is implemented in kvm_arch_vcpu_pre_fault_memory() whose job > is to pre-fault the stage 2 page tables which map a specific GPA (which, > for arm64, is the guest's IPA). > > This function is called by kvm_vcpu_pre_fault_memory() for each GPA in the > range, which itself is ultimately invoked by userland via the > KVM_PRE_FAULT_MEMORY ioctl. > > The implementation is simple - try to walk to the stage 2 page table > mapping the GPA - if unmapped, fault it in through a synthetic page fault. > > pKVM is not supported regardless of whether the VM is protected or > not. > > This is because pKVM instantiates vCPUs upon run, but pre-faulting is > typically performed before a vCPU is run. It would be confusing and > inconsistent to error out on non-running vCPUs but to pre-fault running > ones. > > Jack's original test suite is also included in the series. > > == Credits == > > This series is based, with gratitude, on Jack Thomson's series and their > respins (links provided below) as well as the feedback he received. > > The series includes Jack's v5 "KVM: arm64: Pass walk flags to > kvm_pgtable_get_leaf()" patch verbatim, and all three of his test patches, > two of which required minor fixups. > > Link: https://patch.msgid.link/20260612162354.73378-1-jackabt.amazon@gmail.com/ > Link: https://patch.msgid.link/20260113152643.18858-1-jackabt.amazon@gmail.com/ > Link: https://patch.msgid.link/20251119154910.97716-1-jackabt.amazon@gmail.com/ > Link: https://patch.msgid.link/20251013151502.6679-1-jackabt.amazon@gmail.com/ > Link: https://patch.msgid.link/20250911134648.58945-1-jackabt.amazon@gmail.com/ > > == Reviewer Notes == > > I synced with maintainers on this who asked me to take a look, as there hadn't > been progress on the series for some time. > > Am happy to rebase again after -rc1. > > Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org> > --- > Jack Thomson (4): > KVM: arm64: Pass walk flags to kvm_pgtable_get_leaf() > KVM: selftests: Enable pre_fault_memory_test for arm64 > KVM: selftests: Add option for different backing in pre-fault tests > KVM: selftests: Add nested pre-fault test for arm64 > > Lorenzo Stoakes (ARM) (4): > KVM: arm64: Propagate and use esr in s2fd when handling guest aborts > KVM: arm64: Propagate and use mmu in s2fd when handling guest aborts > KVM: arm64: Propagate and use kvm_s2_fault_result on S2 fault > KVM: arm64: Implement KVM_PRE_FAULT_MEMORY > > Documentation/virt/kvm/api.rst | 17 +- > arch/arm64/include/asm/esr.h | 129 +++++--- > arch/arm64/include/asm/kvm_emulate.h | 52 +--- > arch/arm64/include/asm/kvm_pgtable.h | 5 +- > arch/arm64/include/asm/kvm_pkvm.h | 2 +- > arch/arm64/kvm/Kconfig | 1 + > arch/arm64/kvm/arm.c | 1 + > arch/arm64/kvm/hyp/nvhe/mem_protect.c | 10 +- > arch/arm64/kvm/hyp/pgtable.c | 5 +- > arch/arm64/kvm/mmu.c | 324 +++++++++++++++++---- > arch/arm64/kvm/nested.c | 2 +- > tools/testing/selftests/kvm/Makefile.kvm | 2 + > .../selftests/kvm/arm64/nv_pre_fault_memory_test.c | 206 +++++++++++++ > .../testing/selftests/kvm/pre_fault_memory_test.c | 152 ++++++++-- > 14 files changed, 736 insertions(+), 172 deletions(-) > --- > base-commit: aa8e5dc6a7a2a1141ab40706a51010adcd0e57d2 > change-id: 20260815-kvm-arm-prefault-9bb411b6897d > > Best regards, > -- > Lorenzo Stoakes (ARM) <ljs@kernel.org> >
On Thu, Sep 10, 2026 at 07:44:25PM +0100, Fuad Tabba wrote: > Hi Lorenzo, > > On Tue, 25 Aug 2026 at 17:01, Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > > > This series implements the KVM stage 2 page table pre-faulting feature for > > arm64. > > I know there's likely going to be a V2, but FWIW, I tested this series > as follows (finally managed to get protected mode running on my Mac): > > QEMU (arm64, -cpu max,pauth=off): > - Protected-mode host, boots non-protected and protected guests > - pre_fault_memory_test on a non-protected host, VHE and nVHE, all 13 > guest modes > - nv_pre_fault_memory_test with kvm-arm.mode=nested > > FVP (AEM, FEAT_NV2): > - nv_pre_fault_memory_test with kvm-arm.mode=nested > > Apple M4 hardware, running as a KVM host at EL2 under Apple's nested > virtualization (nVHE only): > - pre_fault_memory_test, anonymous and anonymous_thp backing, over the five > guest modes available there (the IPA is capped at 40 bits and the 64K > granule isn't implemented) > - with kvm-arm.mode=protected, pre_fault_memory_test skips on > KVM_CAP_PRE_FAULT_MEMORY, so the pKVM exclusion holds on a live pKVM host > > Tested-by: Fuad Tabba < fuad.tabba@linux.dev> > (happy to retest again when you respin) Perfect, much appreciated thanks! :) I won't carry the tags forward to v2 just in case anything breaks there, so appreciate you taking another run at the v2 also thanks! :) > > Cheers, > /fuad -- Cheers, Lorenzo
© 2016 - 2026 Red Hat, Inc.