[PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM

Mostafa Saleh posted 3 patches 1 month, 3 weeks ago
arch/arm64/kvm/hyp/include/hyp/switch.h | 15 +----------
arch/arm64/kvm/hyp/nvhe/pkvm.c          | 14 ++++++++++
arch/arm64/kvm/hyp/nvhe/timer-sr.c      | 10 ++++----
include/kvm/arm_arch_timer.h            | 34 +++++++++++++++----------
4 files changed, 41 insertions(+), 32 deletions(-)
[PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Mostafa Saleh 1 month, 3 weeks ago
What started as a  small patch ended up as a 3 patch series thanks
to Sashiko.

First patch from Marc to consolidate the offset calculation,
follow up patches fix issues with non-protected VM and timer
offset and protected VM running with broken CNTVOFF_EL2.

The patches have all the information so I will keep this short.

v1: https://lore.kernel.org/all/20260806150105.4010701-1-smostafa@google.com/
Main changes:
- Add comment based on Fuad review
- Add prefix patches that also fix the non-protected hang properly.

Marc Zyngier (1):
  KVM: arm64: Make timer_get_offset() work in all contexts

Mostafa Saleh (2):
  KVM: arm64: Fix timer offsets for non-protected VMs
  KVM: arm64: Fix hvhe and broken CNTVOFF_EL2

 arch/arm64/kvm/hyp/include/hyp/switch.h | 15 +----------
 arch/arm64/kvm/hyp/nvhe/pkvm.c          | 14 ++++++++++
 arch/arm64/kvm/hyp/nvhe/timer-sr.c      | 10 ++++----
 include/kvm/arm_arch_timer.h            | 34 +++++++++++++++----------
 4 files changed, 41 insertions(+), 32 deletions(-)

-- 
2.55.0.654.g21b8a5bc05-goog
Re: [PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Oliver Upton 1 month, 3 weeks ago
On Sat, 08 Aug 2026 08:58:21 +0000, Mostafa Saleh wrote:
> What started as a  small patch ended up as a 3 patch series thanks
> to Sashiko.
> 
> First patch from Marc to consolidate the offset calculation,
> follow up patches fix issues with non-protected VM and timer
> offset and protected VM running with broken CNTVOFF_EL2.
> 
> [...]

Dropped the unintended SOB in patch 3 you mentioned.

Applied to next, thanks!

[1/3] KVM: arm64: Make timer_get_offset() work in all contexts
      https://git.kernel.org/kvmarm/kvmarm/c/2858600ecd01
[2/3] KVM: arm64: Fix timer offsets for non-protected VMs
      https://git.kernel.org/kvmarm/kvmarm/c/47d3eef780e3
[3/3] KVM: arm64: Fix hvhe and broken CNTVOFF_EL2
      https://git.kernel.org/kvmarm/kvmarm/c/2e813a6e8ebe

--
Best,
Oliver
Re: [PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Mostafa Saleh 1 month, 2 weeks ago
On Sat, Aug 8, 2026 at 7:44 PM Oliver Upton <oupton@kernel.org> wrote:
>
> On Sat, 08 Aug 2026 08:58:21 +0000, Mostafa Saleh wrote:
> > What started as a  small patch ended up as a 3 patch series thanks
> > to Sashiko.
> >
> > First patch from Marc to consolidate the offset calculation,
> > follow up patches fix issues with non-protected VM and timer
> > offset and protected VM running with broken CNTVOFF_EL2.
> >
> > [...]
>
> Dropped the unintended SOB in patch 3 you mentioned.
>
> Applied to next, thanks!
>
> [1/3] KVM: arm64: Make timer_get_offset() work in all contexts
>       https://git.kernel.org/kvmarm/kvmarm/c/2858600ecd01
> [2/3] KVM: arm64: Fix timer offsets for non-protected VMs
>       https://git.kernel.org/kvmarm/kvmarm/c/47d3eef780e3
> [3/3] KVM: arm64: Fix hvhe and broken CNTVOFF_EL2
>       https://git.kernel.org/kvmarm/kvmarm/c/2e813a6e8ebe
>

Thanks Oliver! I believe there is one more bug. I'm not sure where the
bug is or if it relates to the broken timers.
Before those patches I could not boot a protected VM because of the
panic, now after booting protected VMs, I sometimes get a system
reset.
I confirmed that cntvoff_el2 does not get written to non-zero, I also
removed the sysreg write completely (rely on xzr value at init) so my
guess is that the HW might be allergic to more than just non-zero
values in cntvoff_el2.
I do not have issues with non-protected VMs anymore.

Thanks,
Mostafa




> --
> Best,
> Oliver
Re: [PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Marc Zyngier 1 month, 2 weeks ago
On Tue, 11 Aug 2026 13:17:25 +0100,
Mostafa Saleh <smostafa@google.com> wrote:
> 
> On Sat, Aug 8, 2026 at 7:44 PM Oliver Upton <oupton@kernel.org> wrote:
> >
> > On Sat, 08 Aug 2026 08:58:21 +0000, Mostafa Saleh wrote:
> > > What started as a  small patch ended up as a 3 patch series thanks
> > > to Sashiko.
> > >
> > > First patch from Marc to consolidate the offset calculation,
> > > follow up patches fix issues with non-protected VM and timer
> > > offset and protected VM running with broken CNTVOFF_EL2.
> > >
> > > [...]
> >
> > Dropped the unintended SOB in patch 3 you mentioned.
> >
> > Applied to next, thanks!
> >
> > [1/3] KVM: arm64: Make timer_get_offset() work in all contexts
> >       https://git.kernel.org/kvmarm/kvmarm/c/2858600ecd01
> > [2/3] KVM: arm64: Fix timer offsets for non-protected VMs
> >       https://git.kernel.org/kvmarm/kvmarm/c/47d3eef780e3
> > [3/3] KVM: arm64: Fix hvhe and broken CNTVOFF_EL2
> >       https://git.kernel.org/kvmarm/kvmarm/c/2e813a6e8ebe
> >
> 
> Thanks Oliver! I believe there is one more bug. I'm not sure where the
> bug is or if it relates to the broken timers.
> Before those patches I could not boot a protected VM because of the
> panic, now after booting protected VMs, I sometimes get a system
> reset.

On this quality HW, this is usually an indication that you are taking
an exception in a tight loop.

> I confirmed that cntvoff_el2 does not get written to non-zero, I also
> removed the sysreg write completely (rely on xzr value at init) so my
> guess is that the HW might be allergic to more than just non-zero
> values in cntvoff_el2.

Is that in hVHE mode? Can you trap the access and route it to the
existing handling code?

> I do not have issues with non-protected VMs anymore.

Do these run with an offset or not?

	M.

-- 
Without deviation from the norm, progress is not possible.
Re: [PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Mostafa Saleh 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 02:18:15PM +0100, Marc Zyngier wrote:
> On Tue, 11 Aug 2026 13:17:25 +0100,
> Mostafa Saleh <smostafa@google.com> wrote:
> > 
> > On Sat, Aug 8, 2026 at 7:44 PM Oliver Upton <oupton@kernel.org> wrote:
> > >
> > > On Sat, 08 Aug 2026 08:58:21 +0000, Mostafa Saleh wrote:
> > > > What started as a  small patch ended up as a 3 patch series thanks
> > > > to Sashiko.
> > > >
> > > > First patch from Marc to consolidate the offset calculation,
> > > > follow up patches fix issues with non-protected VM and timer
> > > > offset and protected VM running with broken CNTVOFF_EL2.
> > > >
> > > > [...]
> > >
> > > Dropped the unintended SOB in patch 3 you mentioned.
> > >
> > > Applied to next, thanks!
> > >
> > > [1/3] KVM: arm64: Make timer_get_offset() work in all contexts
> > >       https://git.kernel.org/kvmarm/kvmarm/c/2858600ecd01
> > > [2/3] KVM: arm64: Fix timer offsets for non-protected VMs
> > >       https://git.kernel.org/kvmarm/kvmarm/c/47d3eef780e3
> > > [3/3] KVM: arm64: Fix hvhe and broken CNTVOFF_EL2
> > >       https://git.kernel.org/kvmarm/kvmarm/c/2e813a6e8ebe
> > >
> > 
> > Thanks Oliver! I believe there is one more bug. I'm not sure where the
> > bug is or if it relates to the broken timers.
> > Before those patches I could not boot a protected VM because of the
> > panic, now after booting protected VMs, I sometimes get a system
> > reset.
> 
> On this quality HW, this is usually an indication that you are taking
> an exception in a tight loop.
> 

I tried to add a poor man exception storm detection in the kernel
handlers (el1h_64_sync_handler and __gic_handle_irq), but it didn’t trigger
before the reset.

My hunch would be that there are some paths in protected VMs that takes
too long that the watchdog fires. I saw the resets mostly at either
userspace boot or VM teardown, otherwise the VM seems functional.

I will collect timestamps from hypervisor entry/exit and check
how large are those.

> > I confirmed that cntvoff_el2 does not get written to non-zero, I also
> > removed the sysreg write completely (rely on xzr value at init) so my
> > guess is that the HW might be allergic to more than just non-zero
> > values in cntvoff_el2.
> 
> Is that in hVHE mode? Can you trap the access and route it to the
> existing handling code?
> 

Yes, only hVHE. Protected nVHE works fine.

One interesting observation is that when starting a VM with a single
vcpu I don’t see the reset anymore compared to 4 cpus before.

Enabling traps for timer unconditionally for protected VMs still has
the same issue.

> > I do not have issues with non-protected VMs anymore.
> 
> Do these run with an offset or not?

Yes, they have the offset set from kvm_timer_vcpu_init() with
kvm_phys_timer_read().

Thanks,
Mostafa

> 
> 	M.
> 
> -- 
> Without deviation from the norm, progress is not possible.
Re: [PATCH v2 0/3] KVM: arm64: Fixes for timers and pKVM
Posted by Mostafa Saleh 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 6:22 PM Mostafa Saleh <smostafa@google.com> wrote:
>
> On Tue, Aug 11, 2026 at 02:18:15PM +0100, Marc Zyngier wrote:
> > On Tue, 11 Aug 2026 13:17:25 +0100,
> > Mostafa Saleh <smostafa@google.com> wrote:
> > >
> > > On Sat, Aug 8, 2026 at 7:44 PM Oliver Upton <oupton@kernel.org> wrote:
> > > >
> > > > On Sat, 08 Aug 2026 08:58:21 +0000, Mostafa Saleh wrote:
> > > > > What started as a  small patch ended up as a 3 patch series thanks
> > > > > to Sashiko.
> > > > >
> > > > > First patch from Marc to consolidate the offset calculation,
> > > > > follow up patches fix issues with non-protected VM and timer
> > > > > offset and protected VM running with broken CNTVOFF_EL2.
> > > > >
> > > > > [...]
> > > >
> > > > Dropped the unintended SOB in patch 3 you mentioned.
> > > >
> > > > Applied to next, thanks!
> > > >
> > > > [1/3] KVM: arm64: Make timer_get_offset() work in all contexts
> > > >       https://git.kernel.org/kvmarm/kvmarm/c/2858600ecd01
> > > > [2/3] KVM: arm64: Fix timer offsets for non-protected VMs
> > > >       https://git.kernel.org/kvmarm/kvmarm/c/47d3eef780e3
> > > > [3/3] KVM: arm64: Fix hvhe and broken CNTVOFF_EL2
> > > >       https://git.kernel.org/kvmarm/kvmarm/c/2e813a6e8ebe
> > > >
> > >
> > > Thanks Oliver! I believe there is one more bug. I'm not sure where the
> > > bug is or if it relates to the broken timers.
> > > Before those patches I could not boot a protected VM because of the
> > > panic, now after booting protected VMs, I sometimes get a system
> > > reset.
> >
> > On this quality HW, this is usually an indication that you are taking
> > an exception in a tight loop.
> >
>
> I tried to add a poor man exception storm detection in the kernel
> handlers (el1h_64_sync_handler and __gic_handle_irq), but it didn’t trigger
> before the reset.
>
> My hunch would be that there are some paths in protected VMs that takes
> too long that the watchdog fires. I saw the resets mostly at either
> userspace boot or VM teardown, otherwise the VM seems functional.
>
> I will collect timestamps from hypervisor entry/exit and check
> how large are those.
>

Max numbers I've seen doesn't exceed 2-3 ms, compared  to SMCs which
can take almost a second.
I'm not sure if the watchdog is the right conclusion, but without any
clue from the firmware logs, I am lost.

Also, correcting myself, I now see it happened with a single vCPU
although that took longer to reproduce.

Thanks,
Mostafa

> > > I confirmed that cntvoff_el2 does not get written to non-zero, I also
> > > removed the sysreg write completely (rely on xzr value at init) so my
> > > guess is that the HW might be allergic to more than just non-zero
> > > values in cntvoff_el2.
> >
> > Is that in hVHE mode? Can you trap the access and route it to the
> > existing handling code?
> >
>
> Yes, only hVHE. Protected nVHE works fine.
>
> One interesting observation is that when starting a VM with a single
> vcpu I don’t see the reset anymore compared to 4 cpus before.
>
> Enabling traps for timer unconditionally for protected VMs still has
> the same issue.
>
> > > I do not have issues with non-protected VMs anymore.
> >
> > Do these run with an offset or not?
>
> Yes, they have the offset set from kvm_timer_vcpu_init() with
> kvm_phys_timer_read().
>
> Thanks,
> Mostafa
>
> >
> >       M.
> >
> > --
> > Without deviation from the norm, progress is not possible.