[RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage

Leonardo Bras posted 5 patches 3 weeks, 3 days ago
arch/arm64/include/asm/kvm_host.h    |  2 ++
arch/arm64/include/asm/kvm_mmu.h     |  6 ++++
arch/arm64/include/asm/kvm_nested.h  |  9 +++--
arch/arm64/include/asm/kvm_pgtable.h | 12 +++++--
arch/arm64/kvm/arm.c                 | 15 ++++++++
arch/arm64/kvm/hyp/pgtable.c         | 49 +++++++++++++++++++++-----
arch/arm64/kvm/mmu.c                 | 51 ++++++++++++++++++++++------
arch/arm64/kvm/nested.c              |  4 ++-
arch/arm64/kvm/ptdump.c              | 10 ++++--
9 files changed, 130 insertions(+), 28 deletions(-)
[RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage
Posted by Leonardo Bras 3 weeks, 3 days ago
This series have 2 main goals:

1 - Patches #1,#2,#3 : Change the PTE descriptor to use WD/WC/RO encodings
    making use of the DBM bit, adapting all usages, and
2 - Patches #4,#5 are an RFC on using HAFDBS on a guest to avoid resetting
    all PTEs to WC when dirty-logging starts, speeding-up startup.

(1) will also introduce a new walker for cleaning the dirty-bit, which will
clear the DBM bit if it's a block mapping (hugepage). This is needed as on
lazy-splitting we need to fault a write so we can do the lazy splitting.
This is needed for both the next patches, and for HDBSS & HACDBS
enablement.

On (2), I really just want feedback to understand if it's worth pursuing.
My main idea is that we can use HAFDBS _outside_ dirty-logging to only mark
dirty the pages that were actually written to. That is supposed to make it
faster to transverse the pagetables when we need to clean the dirty-bit, as
there is potentially less atomic writes to perform. The price paid for that
is disabling HAFDBS on every vcpu before we can start cleaning the pages,
done by a (new) vcpu request.

Please let me know of what do you think!

Thanks!
Leo

Leonardo Bras (5):
  KVM: arm64: pgtables: Change write bit from S2AP_W to DBM
  KVM: arm64: Add KVM_PGTABLE_PROT_DIRTY
  KVM: arm64: Introduce a dedicated walker for stage2 write-protect
  KVM: arm64: Add KVM_REQ_RELOAD_STAGE2
  KVM: arm64: Enable HAFDBS for guests not on migration

 arch/arm64/include/asm/kvm_host.h    |  2 ++
 arch/arm64/include/asm/kvm_mmu.h     |  6 ++++
 arch/arm64/include/asm/kvm_nested.h  |  9 +++--
 arch/arm64/include/asm/kvm_pgtable.h | 12 +++++--
 arch/arm64/kvm/arm.c                 | 15 ++++++++
 arch/arm64/kvm/hyp/pgtable.c         | 49 +++++++++++++++++++++-----
 arch/arm64/kvm/mmu.c                 | 51 ++++++++++++++++++++++------
 arch/arm64/kvm/nested.c              |  4 ++-
 arch/arm64/kvm/ptdump.c              | 10 ++++--
 9 files changed, 130 insertions(+), 28 deletions(-)


base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
-- 
2.55.0
Re: [RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage
Posted by Marc Zyngier 1 week, 6 days ago
On Tue, 01 Sep 2026 18:15:51 +0100,
Leonardo Bras <leo.bras@arm.com> wrote:
> 
> This series have 2 main goals:
> 
> 1 - Patches #1,#2,#3 : Change the PTE descriptor to use WD/WC/RO encodings
>     making use of the DBM bit, adapting all usages, and

What are WD and WC? I can sort of guess that this is write-dirty and
write clean, but that's not exactly obvious. More importantly, you
don't even explain *why* anything needs changing...

> 2 - Patches #4,#5 are an RFC on using HAFDBS on a guest to avoid resetting
>     all PTEs to WC when dirty-logging starts, speeding-up startup.
> 
> (1) will also introduce a new walker for cleaning the dirty-bit, which will
> clear the DBM bit if it's a block mapping (hugepage). This is needed as on
> lazy-splitting we need to fault a write so we can do the lazy splitting.
> This is needed for both the next patches, and for HDBSS & HACDBS
> enablement.

Again, this is incredibly opaque to the reviewer. What is the problem
you are trying to solve? This is what a cover letter is for.

	M.

-- 
Without deviation from the norm, progress is not possible.
Re: [RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage
Posted by Leonardo Bras 1 week, 3 days ago
On Sat, Sep 12, 2026 at 01:24:11PM +0100, Marc Zyngier wrote:
> On Tue, 01 Sep 2026 18:15:51 +0100,
> Leonardo Bras <leo.bras@arm.com> wrote:
> > 
> > This series have 2 main goals:
> > 
> > 1 - Patches #1,#2,#3 : Change the PTE descriptor to use WD/WC/RO encodings
> >     making use of the DBM bit, adapting all usages, and
> 
> What are WD and WC? I can sort of guess that this is write-dirty and
> write clean, but that's not exactly obvious. More importantly, you
> don't even explain *why* anything needs changing...
> 
> > 2 - Patches #4,#5 are an RFC on using HAFDBS on a guest to avoid resetting
> >     all PTEs to WC when dirty-logging starts, speeding-up startup.
> > 
> > (1) will also introduce a new walker for cleaning the dirty-bit, which will
> > clear the DBM bit if it's a block mapping (hugepage). This is needed as on
> > lazy-splitting we need to fault a write so we can do the lazy splitting.
> > This is needed for both the next patches, and for HDBSS & HACDBS
> > enablement.
> 
> Again, this is incredibly opaque to the reviewer. What is the problem
> you are trying to solve? This is what a cover letter is for.

Hi Marc, thanks for reviewing!

Okay, I will try to explain it better on the next version. What do you 
think of this text:

===================

This series have 2 main goals:
1 - Introduce a new PTE encoding (Patches #1, #2, #3)
2 - An idea to use HAFDBS on a guest to avoid resetting all PTEs to WC 
    when dirty-logging starts

Goal 1:

Patches #1, #2, #3: Before adding Stage-2 support to HAFDBS, HDBSS and 
HACDBS, we need to change the PTE descriptor encoding, as the DBM bit 
is required on mappings for those hardware engines to actually being 
able to update the PTEs. Currently what we have is:

- Read-Only  (RO): S2AP[1]=0
- Read-Write (RW): S2AP[1]=1 

and for them to work with the new features, we need to have: 

- Read-Only (RO):      DBM=0, S2AP[1]=0
- Writable-Clean (WC): DBM=1, S2AP[1]=0
- Writable-Dirty (WD): DBM=1, S2AP[1]=1

WC and WD are described in the Arm ARM, on R_XZFQH and R_BRFGY.

We also need to prepare for dealing with lazy splitting when HAFDBS is 
enabled: since it updates the PTE without taking a fault on guest write, it 
means we can't have lazy splitting if we mark all PTEs as WC. 
To address that, there is a suggestion to set, on dirty-track enable:

- All pages as WC, as they don't need splitting, and
- All blocks as RO, as they are required to fault to do lazy splitting

In order to have that, a new walker is introduced to have a different 
behavior depending on the entry's level.

This is needed for the Goal 2, as well as for HDBSS enablement.


Goal 2:

Patches #4,#5 are an RFC on using HAFDBS on a guest to avoid resetting
all PTEs to WC when dirty-logging starts, making it faster.

I really just want feedback to understand if it's worth pursuing. 

My main idea is that we can use HAFDBS _outside_ dirty-logging to only mark
dirty the pages that were actually written to. 

That is supposed to make it faster to transverse the pagetables when we 
need to clean the dirty-bit, as there is potentially less atomic writes to 
perform. 

The price paid for that is disabling HAFDBS on every vcpu before we can 
start cleaning the pages, during a dirty-track request.

Please let me know of what you think!

Thanks!
Leo