From nobody Tue Aug 25 00:12:26 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org; dmarc=pass(p=none dis=none) header.from=infradead.org ARC-Seal: i=1; a=rsa-sha256; t=1783113769; cv=none; d=zohomail.com; s=zohoarc; b=Cov/EUtm1+t+WdP2qqdhdqLXHJ7d8ZJ7wU+5xacgaih2ZCxG9jC2YOBigC/NbZvdAtdbaGD58SKN2KTwnHduuLp6bxw4nCqd4F9LFLul6jXq/iRkaHNZvmcOgK7rej7vLNCe6/odHUdNwPgNykWg08ERbny/5QgUmvJuVCnd6H8= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783113769; h=Content-Type:Content-Transfer-Encoding:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To:Cc; bh=/GVIzOguJ6FEtu0Xi+CX41sxIMFfZk32Sup1CQOK2cA=; b=AR4QCr5+WpiY9gcoEhcdkYIIqoQaGYQQj+fYwtfT+4Hi7lbzIHW/jj636MgnUNrBdmSh3wJmJv/vsXpnuBWMuUpJhZ71s9w+c7MenlLaUqXNZFqAyFlFOxLgbHUvqGibcZ4iBL41IC2ODUum09M0myw6u9obBP/mjLbjFg0AHKo= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 1783113769407220.28829288680095; Fri, 3 Jul 2026 14:22:49 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1353760.1609526 (Exim 4.92) (envelope-from ) id 1wflKs-0006w3-Et; Fri, 03 Jul 2026 21:22:10 +0000 Received: by outflank-mailman (output) from mailman id 1353760.1609526; Fri, 03 Jul 2026 21:22:10 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1wflKs-0006sg-3X; Fri, 03 Jul 2026 21:22:10 +0000 Received: by outflank-mailman (input) for mailman id 1353760; Fri, 03 Jul 2026 21:22:07 +0000 Received: from mx.expurgate.net ([194.145.224.20]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1wflKo-0005mx-QX for xen-devel@lists.xenproject.org; Fri, 03 Jul 2026 21:22:07 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1wflKo-00BXCb-7G; Fri, 03 Jul 2026 23:22:06 +0200 Received: from [10.42.69.1] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6a48279b-2eae-0a2a0a5409dd-0a2a45019e62-26 for ; Fri, 03 Jul 2026 23:22:06 +0200 Received: from [90.155.50.34] (helo=casper.infradead.org) by tlsNG-d62444.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6a4827f9-400f-0a2a45010019-5a9b3222b94c-3 for ; Fri, 03 Jul 2026 23:22:01 +0200 Received: from [2001:8b0:10b:1::425] (helo=i7.infradead.org) by casper.infradead.org with esmtpsa (Exim 4.99.1 #2 (Red Hat Linux)) id 1wflKW-0000000AsY0-0gKK; Fri, 03 Jul 2026 21:21:48 +0000 Received: from dwoodhou by i7.infradead.org with local (Exim 4.99.2 #2 (Red Hat Linux)) id 1wflKW-00000001RO9-0J5N; Fri, 03 Jul 2026 22:21:48 +0100 X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=casper.20170209 header.d=infradead.org header.i="@infradead.org" header.h="Sender:Content-Transfer-Encoding:Content-Type:MIME-Version:References:In-Reply-To:Message-ID:Date:Subject:To:From" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=Sender:Content-Transfer-Encoding: Content-Type:MIME-Version:References:In-Reply-To:Message-ID:Date:Subject:To: From:Reply-To:Cc:Content-ID:Content-Description; bh=/GVIzOguJ6FEtu0Xi+CX41sxIMFfZk32Sup1CQOK2cA=; b=BaLfOYts8sUMPmGLNaou7hmwDv ZEfJ2OXH77XM1HGv9Qct/70aBrcfJbrFrtbzbbsDoOpGsFipDwRaLmNxMu2mtVcacWgZcyXNPob4E g7uxUNPd+zDqY6qBjNJ+x9SpZ1a0eWx9S8whOdlJn11Sdea2BzLYhL6nLX7Nria99A57xj169SUeP IB8A6a3ZpKbJ6hmsFt+zSw0qlRITMmBDgjsOQounx182YctWVblQukg2QqcyGsPz/AVV8mZRN0Xn0 G7G6WOhJrirf5aATGmlwhdiPBxIdukuDn/grl0vpdDfgN1+/dlXxuqFgaxEG4ywrZfLj+AV07Ve/h tySrJOhA==; From: David Woodhouse To: Paolo Bonzini , Jonathan Corbet , Shuah Khan , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Vitaly Kuznetsov , Juergen Gross , Boris Ostrovsky , David Woodhouse , Paul Durrant , Jonathan Cameron , Sascha Bischoff , Marc Zyngier , Joey Gouly , Jack Allister , Dongli Zhang , joe.jin@oracle.com, kvm@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, xen-devel@lists.xenproject.org, linux-kselftest@vger.kernel.org Subject: [PATCH v6 06/36] KVM: selftests: Add KVM/PV clock selftest to prove timer correction Date: Fri, 3 Jul 2026 22:17:45 +0100 Message-ID: <20260703212145.343527-7-dwmw2@infradead.org> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260703212145.343527-1-dwmw2@infradead.org> References: <20260703212145.343527-1-dwmw2@infradead.org> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Sender: David Woodhouse X-SRS-Rewrite: SMTP reverse-path rewritten from by casper.infradead.org. See http://www.infradead.org/rpr.html X-purgate-ID: tlsNG-d62444/1783113721-FFCCE1E0-C9E5487E/0/0 X-purgate-type: clean X-purgate-size: 16061 X-ZohoMail-DKIM: pass (identity @infradead.org) X-ZM-MESSAGEID: 1783113770658158500 From: Jack Allister A VM's KVM/PV clock has an inherent relationship to its TSC. When either the host system live-updates or the VM is live-migrated this pairing of the two clock sources should stay the same. In reality this is not the case without some correction taking place. The KVM_GET_CLOCK_GUEST/KVM_SET_CLOCK_GUEST ioctls can be used to perform a correction on the PVTI (PV time information) structure held by KVM to effectively fix up the kvmclock_offset prior to the guest VM resuming in either a live-update/migration scenario. This test proves that without the necessary fixup there is a perceived change in the guest TSC and KVM/PV clock relationship before and after a simulated LU/LM takes place, and that the correction eliminates it. The test: 1. Snapshots the PVTI at boot (PVTI0). 2. Induces a change in PVTI data (KVM_REQ_MASTERCLOCK_UPDATE). 3. Snapshots the PVTI after the change (PVTI1). 4. Requests correction via KVM_SET_CLOCK_GUEST using PVTI0. 5. Snapshots the PVTI after correction (PVTI2). Then samples the TSC at a single point in time and calculates the KVM clock using each PVTI snapshot. The corrected clock should match the boot clock to within =C2=B11ns. The test enumerates multiple TSC frequencies from 1GHz to 5GHz at 500MHz steps, crossing the 32-bit boundary, to exercise the scaling path at various ratios. The sleep duration between snapshots is configurable via the -s/--sleep command line option. Co-developed-by: David Woodhouse Signed-off-by: David Woodhouse Signed-off-by: Jack Allister Reviewed-by: Paul Durrant Cc: Dongli Zhang --- tools/testing/selftests/kvm/Makefile.kvm | 1 + .../testing/selftests/kvm/x86/pvclock_test.c | 443 ++++++++++++++++++ 2 files changed, 444 insertions(+) create mode 100644 tools/testing/selftests/kvm/x86/pvclock_test.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selft= ests/kvm/Makefile.kvm index 9118a5a51b89..fb935ae3bf38 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -105,6 +105,7 @@ TEST_GEN_PROGS_x86 +=3D x86/pmu_counters_test TEST_GEN_PROGS_x86 +=3D x86/pmu_event_filter_test TEST_GEN_PROGS_x86 +=3D x86/private_mem_conversions_test TEST_GEN_PROGS_x86 +=3D x86/private_mem_kvm_exits_test +TEST_GEN_PROGS_x86 +=3D x86/pvclock_test TEST_GEN_PROGS_x86 +=3D x86/set_boot_cpu_id TEST_GEN_PROGS_x86 +=3D x86/set_sregs_test TEST_GEN_PROGS_x86 +=3D x86/smaller_maxphyaddr_emulation_test diff --git a/tools/testing/selftests/kvm/x86/pvclock_test.c b/tools/testing= /selftests/kvm/x86/pvclock_test.c new file mode 100644 index 000000000000..f2b917ed5dea --- /dev/null +++ b/tools/testing/selftests/kvm/x86/pvclock_test.c @@ -0,0 +1,443 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * Copyright =C2=A9 Amazon.com, Inc. or its affiliates. + * + * Tests for pvclock API + * KVM_SET_CLOCK_GUEST/KVM_GET_CLOCK_GUEST + */ +#include +#include +#include +#include +#include + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +#include + +/* + * Reproduce the pvclock calculation the guest uses to convert TSC to + * nanoseconds. This must match the kernel's __pvclock_read_cycles(). + */ +static inline uint64_t pvclock_scale_delta(uint64_t delta, uint32_t mul, + int8_t shift) +{ + if (shift < 0) + delta >>=3D -shift; + else + delta <<=3D shift; + return ((__uint128_t)delta * mul) >> 32; +} + +static inline uint64_t pvclock_read_cycles(struct pvclock_vcpu_time_info *= src, + uint64_t tsc) +{ + uint64_t delta =3D tsc - src->tsc_timestamp; + + return src->system_time + pvclock_scale_delta(delta, + src->tsc_to_system_mul, + src->tsc_shift); +} + +static inline void pvti_snapshot(struct pvclock_vcpu_time_info *dst, + volatile struct pvclock_vcpu_time_info *src) +{ + uint32_t version; + + do { + version =3D src->version; + __asm__ __volatile__("" ::: "memory"); + *dst =3D *src; + __asm__ __volatile__("" ::: "memory"); + } while ((src->version & 1) || src->version !=3D version); +} + +enum { + STAGE_FIRST_BOOT, + STAGE_UNCORRECTED, + STAGE_CORRECTED +}; + +#define KVMCLOCK_GPA 0xc0000000ull +#define KVMCLOCK_SIZE sizeof(struct pvclock_vcpu_time_info) + +static void trigger_pvti_update(void) +{ + /* + * Toggle between KVM's old and new system time methods to coerce KVM + * into updating the fields in the PV time info struct. + */ + wrmsr(MSR_KVM_SYSTEM_TIME, KVMCLOCK_GPA | KVM_MSR_ENABLED); + wrmsr(MSR_KVM_SYSTEM_TIME_NEW, KVMCLOCK_GPA | KVM_MSR_ENABLED); +} + +static void guest_code(void) +{ + struct pvclock_vcpu_time_info *pvti =3D + (void *)(unsigned long)KVMCLOCK_GPA; + struct pvclock_vcpu_time_info pvti_boot; + struct pvclock_vcpu_time_info pvti_uncorrected; + struct pvclock_vcpu_time_info pvti_corrected; + uint64_t tsc_guest; + uint64_t clk_boot, clk_uncorrected, clk_corrected; + int64_t delta_corrected; + + /* Set up kvmclock and snapshot the initial pvclock parameters. */ + wrmsr(MSR_KVM_SYSTEM_TIME_NEW, KVMCLOCK_GPA | KVM_MSR_ENABLED); + pvti_snapshot(&pvti_boot, pvti); + GUEST_SYNC(STAGE_FIRST_BOOT); + + /* + * Trigger an update of the PVTI. Calculating the KVM clock using this + * updated structure will show a delta from the original. + */ + trigger_pvti_update(); + pvti_snapshot(&pvti_uncorrected, pvti); + GUEST_SYNC(STAGE_UNCORRECTED); + + /* + * Snapshot the corrected time (the host does KVM_SET_CLOCK_GUEST when + * handling STAGE_UNCORRECTED). + */ + pvti_snapshot(&pvti_corrected, pvti); + + /* + * Sample the TSC at a single point in time, then calculate the + * effective KVM clock using the PVTI from each stage. Verify that the + * corrected clock matches the boot clock to within =C2=B12ns. + */ + tsc_guest =3D rdtsc(); + + clk_boot =3D pvclock_read_cycles(&pvti_boot, tsc_guest); + clk_uncorrected =3D pvclock_read_cycles(&pvti_uncorrected, tsc_guest); + clk_corrected =3D pvclock_read_cycles(&pvti_corrected, tsc_guest); + + delta_corrected =3D clk_boot - clk_corrected; + + __GUEST_ASSERT(delta_corrected >=3D -2 && delta_corrected <=3D 2, + "corrected delta %ld out of range (boot=3D%lu uncorrected=3D%lu c= orrected=3D%lu)", + delta_corrected, clk_boot, clk_uncorrected, clk_corrected); + + GUEST_SYNC(STAGE_CORRECTED); +} + +static void run_test(struct kvm_vm *vm, struct kvm_vcpu *vcpu, + unsigned int sleep_sec) +{ + struct pvclock_vcpu_time_info pvti_before; + struct ucall uc; + + for (;;) { + vcpu_run(vcpu); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO); + + switch (get_ucall(vcpu, &uc)) { + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + case UCALL_SYNC: + break; + default: + TEST_FAIL("Unexpected ucall"); + } + + switch (uc.args[1]) { + case STAGE_FIRST_BOOT: + /* Save the pvclock parameters before the update. */ + vcpu_ioctl(vcpu, KVM_GET_CLOCK_GUEST, &pvti_before); + + /* Sleep to let the clocks diverge. */ + sleep(sleep_sec); + break; + + case STAGE_UNCORRECTED: + /* Restore the original pvclock parameters. */ + vcpu_ioctl(vcpu, KVM_SET_CLOCK_GUEST, &pvti_before); + break; + + case STAGE_CORRECTED: + /* Guest verified the delta in-guest. */ + return; + + default: + TEST_FAIL("Unknown stage %lu", uc.args[1]); + } + } +} + +static void configure_pvclock(struct kvm_vm *vm) +{ + unsigned int nr_pages; + + nr_pages =3D vm_calc_num_guest_pages(VM_MODE_DEFAULT, getpagesize()); + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, + KVMCLOCK_GPA, 1, nr_pages, 0); + virt_map(vm, KVMCLOCK_GPA, KVMCLOCK_GPA, nr_pages); +} + +static void run_at_frequency(uint64_t tsc_khz, unsigned int sleep_sec) +{ + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + + pr_info("Testing at TSC frequency %lu kHz\n", tsc_khz); + vm =3D vm_create_with_one_vcpu(&vcpu, guest_code); + configure_pvclock(vm); + vcpu_ioctl(vcpu, KVM_SET_TSC_KHZ, (void *)tsc_khz); + run_test(vm, vcpu, sleep_sec); + kvm_vm_free(vm); +} + +static void test_tsc_stable_bit(void); +static void test_clock_guest_with_offsets(void); + +static void usage(const char *name) +{ + printf("Usage: %s [options]\n" + " -s, --sleep SEC sleep duration between snapshots (default: = 2)\n" + " -h, --help show this help\n", name); +} + +int main(int argc, char *argv[]) +{ + static const struct option long_opts[] =3D { + { "sleep", required_argument, NULL, 's' }, + { "help", no_argument, NULL, 'h' }, + { NULL, 0, NULL, 0 }, + }; + unsigned int sleep_sec =3D 2; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + uint64_t host_khz; + uint64_t freq; + int opt; + + while ((opt =3D getopt_long(argc, argv, "s:h", long_opts, NULL)) !=3D -1)= { + switch (opt) { + case 's': + sleep_sec =3D atoi(optarg); + break; + case 'h': + default: + usage(argv[0]); + return opt =3D=3D 'h' ? 0 : 1; + } + } + + TEST_REQUIRE(sys_clocksource_is_based_on_tsc()); + TEST_REQUIRE(kvm_has_cap(KVM_CAP_TSC_CONTROL)); + + vm =3D vm_create_with_one_vcpu(&vcpu, guest_code); + configure_pvclock(vm); + + /* Check KVM_GET_CLOCK_GUEST is supported */ + { + struct pvclock_vcpu_time_info tmp; + int ret =3D __vcpu_ioctl(vcpu, KVM_GET_CLOCK_GUEST, &tmp); + TEST_REQUIRE(ret =3D=3D 0); + } + + /* First run at native frequency (no scaling). */ + run_test(vm, vcpu, sleep_sec); + + /* + * Then enumerate a range of TSC frequencies crossing the 32-bit + * boundary, to exercise the scaling path at various ratios. + */ + host_khz =3D __vcpu_ioctl(vcpu, KVM_GET_TSC_KHZ, NULL); + kvm_vm_free(vm); + + for (freq =3D 1000000; freq <=3D 5000000; freq +=3D 500000) { + if (freq =3D=3D host_khz) + continue; + run_at_frequency(freq, sleep_sec); + } + + test_tsc_stable_bit(); + test_clock_guest_with_offsets(); + + return 0; +} + +static volatile uint32_t vcpu_counter; +static void guest_code_stable_bit(void) +{ + uint32_t idx =3D __atomic_fetch_add(&vcpu_counter, 1, __ATOMIC_SEQ_CST); + uint64_t gpa =3D KVMCLOCK_GPA + idx * sizeof(struct pvclock_vcpu_time_inf= o); + + wrmsr(MSR_KVM_SYSTEM_TIME_NEW, gpa | KVM_MSR_ENABLED); + GUEST_SYNC(0); + GUEST_SYNC(0); + GUEST_SYNC(0); +} + +static void set_tsc_offset(struct kvm_vcpu *vcpu, uint64_t offset) +{ + struct kvm_device_attr attr =3D { + .group =3D KVM_VCPU_TSC_CTRL, + .attr =3D KVM_VCPU_TSC_OFFSET, + .addr =3D (__u64)(uintptr_t)&offset, + }; + + TEST_REQUIRE(__vcpu_has_device_attr(vcpu, KVM_VCPU_TSC_CTRL, + KVM_VCPU_TSC_OFFSET) =3D=3D 0); + vcpu_ioctl(vcpu, KVM_SET_DEVICE_ATTR, &attr); +} + +static void run_vcpu_once(struct kvm_vcpu *vcpu) +{ + struct ucall uc; + + vcpu_run(vcpu); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO); + switch (get_ucall(vcpu, &uc)) { + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + case UCALL_SYNC: + break; + default: + TEST_FAIL("Unexpected ucall"); + } +} + +static void test_tsc_stable_bit(void) +{ + struct pvclock_vcpu_time_info pvti; + struct kvm_vcpu *vcpus[2]; + struct kvm_vm *vm; + int ret; + + pr_info("Testing PVCLOCK_TSC_STABLE_BIT with matched/unmatched TSCs\n"); + + vm =3D vm_create_with_vcpus(2, guest_code_stable_bit, vcpus); + configure_pvclock(vm); + + /* + * Case 1: All TSCs matched (same frequency and offset). + * Master clock should be active, PVCLOCK_TSC_STABLE_BIT set. + */ + run_vcpu_once(vcpus[0]); + + ret =3D __vcpu_ioctl(vcpus[0], KVM_GET_CLOCK_GUEST, &pvti); + TEST_ASSERT(!ret, "GET_CLOCK_GUEST should succeed with matched TSCs"); + TEST_ASSERT(pvti.flags & PVCLOCK_TSC_STABLE_BIT, + "PVCLOCK_TSC_STABLE_BIT should be set with matched TSCs"); + + /* + * Case 2: Different TSC offset, same frequency. + * Master clock should still be active (frequency matches), but + * PVCLOCK_TSC_STABLE_BIT should be cleared (offsets differ). + */ + set_tsc_offset(vcpus[1], 12345678); + run_vcpu_once(vcpus[1]); + run_vcpu_once(vcpus[0]); + + ret =3D __vcpu_ioctl(vcpus[0], KVM_GET_CLOCK_GUEST, &pvti); + if (ret) { + /* Master clock disabled by offset mismatch =E2=80=94 old kernel */ + pr_info(" Skipping offset tests (master clock requires matched offsets)= \n"); + goto out_stable; + } + TEST_ASSERT(!(pvti.flags & PVCLOCK_TSC_STABLE_BIT), + "PVCLOCK_TSC_STABLE_BIT should be clear with offset-mismatched TSCs"= ); + + /* + * Case 3: Different TSC frequency. + * Master clock should be disabled entirely. + */ + vcpu_ioctl(vcpus[1], KVM_SET_TSC_KHZ, + (void *)(unsigned long)(__vcpu_ioctl(vcpus[1], KVM_GET_TSC_KHZ, NULL)= / 2)); + /* Write TSC to trigger kvm_synchronize_tsc / kvm_track_tsc_matching */ + vcpu_set_msr(vcpus[1], MSR_IA32_TSC, 0); + run_vcpu_once(vcpus[1]); + + ret =3D __vcpu_ioctl(vcpus[0], KVM_GET_CLOCK_GUEST, &pvti); + TEST_ASSERT(ret && errno =3D=3D EINVAL, + "GET_CLOCK_GUEST should fail with frequency-mismatched TSCs, got %d = (errno %d)", + ret, errno); + +out_stable: + kvm_vm_free(vm); +} + +static void test_clock_guest_with_offsets(void) +{ + struct pvclock_vcpu_time_info pvti0, pvti1, pvti1_after; + struct kvm_vcpu *vcpus[2]; + struct kvm_vm *vm; + int64_t delta; + int ret; + + pr_info("Testing KVM_[GS]ET_CLOCK_GUEST with different TSC offsets\n"); + + vm =3D vm_create_with_vcpus(2, guest_code_stable_bit, vcpus); + configure_pvclock(vm); + + /* Set different TSC offsets on the two vCPUs */ + set_tsc_offset(vcpus[0], 0); + set_tsc_offset(vcpus[1], 1000000000ull); + + /* Run both to establish kvmclock */ + run_vcpu_once(vcpus[0]); + run_vcpu_once(vcpus[1]); + + /* GET_CLOCK_GUEST on both =E2=80=94 should succeed (master clock active)= */ + ret =3D __vcpu_ioctl(vcpus[0], KVM_GET_CLOCK_GUEST, &pvti0); + if (ret) { + pr_info(" Skipping (master clock requires matched offsets on this kerne= l)\n"); + kvm_vm_free(vm); + return; + } + ret =3D __vcpu_ioctl(vcpus[1], KVM_GET_CLOCK_GUEST, &pvti1); + TEST_ASSERT(!ret, "GET_CLOCK_GUEST on vcpu1 failed"); + + /* The tsc_timestamps should differ (different offsets) */ + TEST_ASSERT(pvti0.tsc_timestamp !=3D pvti1.tsc_timestamp, + "tsc_timestamps should differ with different offsets"); + + /* Sleep to let time elapse, then restore vcpu0's clock */ + sleep(1); + vcpu_ioctl(vcpus[0], KVM_SET_CLOCK_GUEST, &pvti0); + + /* Run vcpu0 to process the clock update */ + run_vcpu_once(vcpus[0]); + + /* GET_CLOCK_GUEST on vcpu1 =E2=80=94 should reflect the correction */ + ret =3D __vcpu_ioctl(vcpus[1], KVM_GET_CLOCK_GUEST, &pvti1_after); + TEST_ASSERT(!ret, "GET_CLOCK_GUEST on vcpu1 after SET failed"); + + /* + * After SET on vcpu0, verify the correction worked by getting + * the clock on vcpu0 again. The mul/shift should be the same, + * and computing kvmclock at the same TSC should give the same + * result as the original (within =C2=B12ns). + */ + { + struct pvclock_vcpu_time_info pvti0_after; + uint64_t tsc_now, clk_from_old, clk_from_new; + + ret =3D __vcpu_ioctl(vcpus[0], KVM_GET_CLOCK_GUEST, &pvti0_after); + TEST_ASSERT(!ret, "GET_CLOCK_GUEST on vcpu0 after SET failed"); + + tsc_now =3D pvti0_after.tsc_timestamp; + clk_from_old =3D pvclock_read_cycles(&pvti0, tsc_now); + clk_from_new =3D pvclock_read_cycles(&pvti0_after, tsc_now); + + delta =3D (int64_t)clk_from_new - (int64_t)clk_from_old; + TEST_ASSERT(delta >=3D -2 && delta <=3D 2, + "clock correction delta should be <=3D2ns, got %ld ns", + delta); + } + + /* + * Also verify that vcpu1's clock is still accessible (master + * clock still active with different offsets). + */ + ret =3D __vcpu_ioctl(vcpus[1], KVM_GET_CLOCK_GUEST, &pvti1_after); + TEST_ASSERT(!ret, "GET_CLOCK_GUEST on vcpu1 after SET failed"); + + kvm_vm_free(vm); +} --=20 2.54.0