From nobody Fri Oct 2 12:24:37 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 23D9C44C67C for ; Fri, 31 Jul 2026 17:19:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785518373; cv=none; b=drgnlW/zvTdHGnmoSrdzHwRMk3IfPpQ6QQWqhWbDTw/y6PCBpCFTJTNpz0v4YJQv2wnyj7mT6ZF7MuFsKjCwL5DM+K++9TOjsMbgvFjD6GlKeLrhNdPxkB5wZBxVkKZlXvDcevMpUaooJIRfN79x1J+hPMxuChlcTuS7obqgRMw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785518373; c=relaxed/simple; bh=zMeBIHN0CGdZnKytc/O1sTgliErd+dm5vaI/CQ98LIA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=UtRLLpJ7SmfYuw8+T4BZ2e1c93OnlNkhURhOchlunOtKbM5od3ONcLIQEsxpdFDB4oIOhRF7wt9uWGxSLSTrmGlZZCCmO6MAiixgvNUn1AMXjBob5ORZ9CTSpiB740VE3y193NvBbeGKRVKL0nUOQeR7Joz0pSpjLd0S0pNtlGE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=n7K4by+n; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="n7K4by+n" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2cfa4e4684bso29365205ad.2 for ; Fri, 31 Jul 2026 10:19:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1785518369; x=1786123169; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=DkmcTN43qpor8TZvvRVUAHGm6+rjqfKAX0vCCdL92HM=; b=n7K4by+nGPDx5TFRT7vjR8fVTFok/p87gKJxXGNaaCcVeUlYZmtG1Y658E71F8ZZQw 8FrMicfByRp5QGEYWp9qQByQuP6vZsxzNwNk0+PjZd0aOveTzUOWGXc5Pgad4argvfw0 KkiJxIh0fkL7RIlY+9iFaC9RATCKJ9HaWqnniTB5t98kpAtjfD8T+ssyEZr/qJ8I9Lkj IuF4tXaHM3uMl5BQjDRJLXahDnftOU4L/bx9zOTWpV6k7rW3rZgzmhwq1+kZ3hTIUEGk BbMVojqhZPZa3O/46udPFCgu2rXBnufJouXg24pEv1cuX6NkQKxbTciOfo/NWHZ69nCx cdhA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785518369; x=1786123169; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=DkmcTN43qpor8TZvvRVUAHGm6+rjqfKAX0vCCdL92HM=; b=Ooadix+yEcGkbjs2pxYhBe97PYXyZ2WHT89O1zYBK3W/B2YPdOQZe8jGK1FQ7XDlCL umrGyXLmXbUEMqh4cDVws+B58zWWZGlgttG9vsoB84lM3zWfVESYhzf8T+fztep9X5rY SqH3c3lKZIX6npkhGAGjVuf8QN8+uN764o0neHdadg28wbgK5ysWtJd7FDhaprLRzkL8 9uW41rr3J9vyWB4D08dmOZzIBnOHb9f8Lh394OsmWk9oUfrh2HYRkPuwINXLsZWKqoDZ I/qu1dGSKV4stX78oX8UXrYJMuw2hWIMNj7J8tno365ZlrfMLqNXy0K8V8r5AUgu4E4h 3Knw== X-Forwarded-Encrypted: i=1; AHgh+RppKTNEYKH62VR2/gHmforvdiCuMVUCJcU0a5FzmPuLDSro50CnBVvIsCzuTj4+2B5urawUydfeZFtWEKQ=@vger.kernel.org X-Gm-Message-State: AOJu0Yxc/07F5rOhXuevnOM9x68yB7q5j8IW1ig7QzY9BKBpjSbcUFNU PEHZ+Vs9kny+pc90fEHff98mY0deUbfPWYTDrSOPevd/5CC2NMUKmnruwUxYOM5wuQi0rf7frQq sVVFSAQ== X-Received: from plht14.prod.google.com ([2002:a17:903:2f0e:b0:2ca:d66c:97b2]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:e5cd:b0:2c9:97a9:2098 with SMTP id d9443c01a7336-2d052421484mr8012735ad.44.1785518369216; Fri, 31 Jul 2026 10:19:29 -0700 (PDT) Reply-To: Sean Christopherson Date: Fri, 31 Jul 2026 10:19:25 -0700 In-Reply-To: <20260731171926.2629627-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260731171926.2629627-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.508.g3f0d502094-goog Message-ID: <20260731171926.2629627-2-seanjc@google.com> Subject: [PATCH v7 1/2] KVM: VMX: Bury all of the VMX preemption timer code under CONFIG_X86_64=y From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Binbin Wu , Chao Gao , Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Double down on using the VMX preemption timer only for 64-bit kernels, and bury the setup and runtime adjustment code, and all global variables, under CONFIG_X86_64=3Dy. This will allow addressing a widespread Intel erratum without running afoul of unused-but-set-variable and __udivdi3() errors on 32-bit kernels. No functional change intended. Reviewed-by: Binbin Wu Reviewed-by: Chao Gao Signed-off-by: Sean Christopherson --- arch/x86/kvm/vmx/vmx.c | 122 +++++++++++++++++++++++------------------ 1 file changed, 69 insertions(+), 53 deletions(-) diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c index e4b9ac7fed9f..a07faa066ef0 100644 --- a/arch/x86/kvm/vmx/vmx.c +++ b/arch/x86/kvm/vmx/vmx.c @@ -150,10 +150,12 @@ module_param(dump_invalid_vmcs, bool, 0644); #define KVM_VMX_TSC_MULTIPLIER_MAX 0xffffffffffffffffULL =20 /* Guest_tsc -> host_tsc conversion requires 64-bit division. */ +#ifdef CONFIG_X86_64 static int __read_mostly cpu_preemption_timer_multi; static bool __read_mostly enable_preemption_timer =3D 1; -#ifdef CONFIG_X86_64 module_param_named(preemption_timer, enable_preemption_timer, bool, S_IRUG= O); +#else +#define enable_preemption_timer false #endif =20 extern bool __read_mostly allow_smaller_maxphyaddr; @@ -7408,32 +7410,6 @@ static void vmx_refresh_guest_perf_global_control(st= ruct kvm_vcpu *vcpu) pmu->global_ctrl =3D vmcs_read64(GUEST_IA32_PERF_GLOBAL_CTRL); } =20 -static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediat= e_exit) -{ - struct vcpu_vmx *vmx =3D to_vmx(vcpu); - u64 tscl; - u32 delta_tsc; - - if (force_immediate_exit) { - vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, 0); - vmx->loaded_vmcs->hv_timer_soft_disabled =3D false; - } else if (vmx->hv_deadline_tsc !=3D -1) { - tscl =3D rdtsc(); - if (vmx->hv_deadline_tsc > tscl) - /* set_hv_timer ensures the delta fits in 32-bits */ - delta_tsc =3D (u32)((vmx->hv_deadline_tsc - tscl) >> - cpu_preemption_timer_multi); - else - delta_tsc =3D 0; - - vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc); - vmx->loaded_vmcs->hv_timer_soft_disabled =3D false; - } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) { - vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1); - vmx->loaded_vmcs->hv_timer_soft_disabled =3D true; - } -} - void noinstr vmx_update_host_rsp(struct vcpu_vmx *vmx, unsigned long host_= rsp) { if (unlikely(host_rsp !=3D vmx->loaded_vmcs->host_state.rsp)) { @@ -7519,6 +7495,8 @@ static noinstr void vmx_vcpu_enter_exit(struct kvm_vc= pu *vcpu, guest_state_exit_irqoff(); } =20 +static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediat= e_exit); + fastpath_t vmx_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags) { bool force_immediate_exit =3D run_flags & KVM_RUN_FORCE_IMMEDIATE_EXIT; @@ -8328,6 +8306,36 @@ static inline int u64_shl_div_u64(u64 a, unsigned in= t shift, return 0; } =20 +static __init void vmx_setup_preemption_timer(void) +{ + if (!cpu_has_vmx_preemption_timer()) + enable_preemption_timer =3D false; + + if (enable_preemption_timer) { + u64 use_timer_freq =3D 5000ULL * 1000 * 1000; + + cpu_preemption_timer_multi =3D + vmx_misc_preemption_timer_rate(vmcs_config.misc); + + if (tsc_khz) + use_timer_freq =3D (u64)tsc_khz * 1000; + use_timer_freq >>=3D cpu_preemption_timer_multi; + + /* + * KVM "disables" the preemption timer by setting it to its max + * value. Don't use the timer if it might cause spurious exits + * at a rate faster than 0.1 Hz (of uninterrupted guest time). + */ + if (use_timer_freq > 0xffffffffu / 10) + enable_preemption_timer =3D false; + } + + if (!enable_preemption_timer) { + vt_x86_ops.set_hv_timer =3D NULL; + vt_x86_ops.cancel_hv_timer =3D NULL; + } +} + int vmx_set_hv_timer(struct kvm_vcpu *vcpu, u64 guest_deadline_tsc, bool *expired) { @@ -8372,6 +8380,39 @@ void vmx_cancel_hv_timer(struct kvm_vcpu *vcpu) { to_vmx(vcpu)->hv_deadline_tsc =3D -1; } + +static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediat= e_exit) +{ + struct vcpu_vmx *vmx =3D to_vmx(vcpu); + u64 tscl; + u32 delta_tsc; + + if (force_immediate_exit) { + vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, 0); + vmx->loaded_vmcs->hv_timer_soft_disabled =3D false; + } else if (vmx->hv_deadline_tsc !=3D -1) { + tscl =3D rdtsc(); + if (vmx->hv_deadline_tsc > tscl) + /* set_hv_timer ensures the delta fits in 32-bits */ + delta_tsc =3D (u32)((vmx->hv_deadline_tsc - tscl) >> + cpu_preemption_timer_multi); + else + delta_tsc =3D 0; + + vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc); + vmx->loaded_vmcs->hv_timer_soft_disabled =3D false; + } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) { + vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1); + vmx->loaded_vmcs->hv_timer_soft_disabled =3D true; + } +} +#else +static __init void vmx_setup_preemption_timer(void) { } + +static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediat= e_exit) +{ + BUILD_BUG_ON(1); +} #endif =20 void vmx_update_cpu_dirty_logging(struct kvm_vcpu *vcpu) @@ -8734,32 +8775,7 @@ __init int vmx_hardware_setup(void) if (!enable_ept || !enable_ept_ad_bits || !cpu_has_vmx_pml()) enable_pml =3D 0; =20 - if (!cpu_has_vmx_preemption_timer()) - enable_preemption_timer =3D false; - - if (enable_preemption_timer) { - u64 use_timer_freq =3D 5000ULL * 1000 * 1000; - - cpu_preemption_timer_multi =3D - vmx_misc_preemption_timer_rate(vmcs_config.misc); - - if (tsc_khz) - use_timer_freq =3D (u64)tsc_khz * 1000; - use_timer_freq >>=3D cpu_preemption_timer_multi; - - /* - * KVM "disables" the preemption timer by setting it to its max - * value. Don't use the timer if it might cause spurious exits - * at a rate faster than 0.1 Hz (of uninterrupted guest time). - */ - if (use_timer_freq > 0xffffffffu / 10) - enable_preemption_timer =3D false; - } - - if (!enable_preemption_timer) { - vt_x86_ops.set_hv_timer =3D NULL; - vt_x86_ops.cancel_hv_timer =3D NULL; - } + vmx_setup_preemption_timer(); =20 kvm_caps.supported_mce_cap |=3D MCG_LMCE_P; kvm_caps.supported_mce_cap |=3D MCG_CMCI_P; --=20 2.55.0.508.g3f0d502094-goog From nobody Fri Oct 2 12:24:37 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D13B643F8AC for ; Fri, 31 Jul 2026 17:19:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785518376; cv=none; b=CkqyVTgfLpgheaf8LVj2SuLSJgrUGxLl49rikxEMv2aKN3rLB/+co1NRQWsGNPFS5Ud0pCQ3C39L7ASxxZueeybzRD5oSgCDSzP/W8YrbmEaG//a6hmSxo+db5OiLIJUdNzII7z6tgleQ5YftruLEfOrQWz8ANMPb+cbKUU91W8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785518376; c=relaxed/simple; bh=uYFUivxPYUYbmgO3jVkPEYX/2SJgZBAfO4E2lQNbJ5E=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=P83miY1+hVRpVnB8YpGFlIAozrQMAFSrcw4xBYe92y6/GxD2UY5dySErxuvFbH/liUW9zD4/5CHznsBQ0+zJ850vXMwY43bal5LjhSf/16zA5hksQANuSIGH0ainqg2QUtRrAhUnW87UNg7mrEWLZ77TgrWgrizOtrJ4mdD8H6I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=l5DHJfvg; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="l5DHJfvg" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2ccb6f6a3f4so15103425ad.1 for ; Fri, 31 Jul 2026 10:19:31 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1785518370; x=1786123170; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=xWrk4GpBsMY+nAsfOPZnb8P50AG697NWSPek8RL0lYI=; b=l5DHJfvgpQ0djOy9WNoY1qDFSNYbfXNf6SFFdZVIYo+t5mX19bjWzoMgRU3PigmVQd xxwmA13LNiI2THYEdJKiRlsp+IDPjdOxYAbN4GQHcbwkbUxgbh/AJmU/UhsQ/JUYWgFT 1ljnyo78G6HlLXKX4AcRw7CJLkg9/vWROxzaEZe+gmbrxh7cZMUvHLyJum88CCBzXKab CZA7GS4C8HdWwuhVUg3dytI4TeZOWadz3WfP9McwUJX6WwXBMZhZSAs+jLX3fReFOLRO yNsXcBnTgHdnCHM/j3AwXI6oHS0gKmEj89/XFwygKxHuZ6/jVO46P264rdSg2xpsYZ5A AbMg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785518370; x=1786123170; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=xWrk4GpBsMY+nAsfOPZnb8P50AG697NWSPek8RL0lYI=; b=FPbfz6TMSD0wByAZkhV9rTIKATooGxmuZFns+Cwto4p5E2YV9YVg1FTH1nozzsLjeo tYAMxYsNlmeuAyJUlQtvFZEbWnW9420urWGn0whStWZfkKourBNgLshSP3DPnQlblYrE GyIsgelwdm27NZ4h57bfSVzNw9qe+AS3phIZJad4ibmozNFUonY91atlUJ2UH982KTfq d0qIRxXGz3FGl/4Lsq0YBA/vz1xxs2hNWpuxjgrxtTuc1PTqgmRHpEClZ1nMvUFJtAA0 Wcel2doCbmoAPn4nlKsQHlNyoVFjLVLCxCWjqUN9i5dACQHrVpv/RRTvlC1VGbuDDnpC PIdA== X-Forwarded-Encrypted: i=1; AHgh+Ro7qQ37HlJlBshfa3EKrMHTlNrQA46FK6HpcnaKcRR6f14FFpOo4oylXfA5mxugB2aq3IED5oPzfLux31k=@vger.kernel.org X-Gm-Message-State: AOJu0YwDmsduq3mx9mupSE7sNhz8HKMXqtYtV9qQyIY3PSfPSW6S+Jyc 2oX5fINIPmpx/B0ve/MhymXFVRK/uZITERQBkK8Xh8L0xU1l7WdzIEIGjxn76lLSlwlGXf4D2T1 PokkkEw== X-Received: from plbml4.prod.google.com ([2002:a17:903:34c4:b0:2ca:ddb3:59fc]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:ea0c:b0:2cf:41ba:96c2 with SMTP id d9443c01a7336-2d05366b5dfmr3813555ad.12.1785518370302; Fri, 31 Jul 2026 10:19:30 -0700 (PDT) Reply-To: Sean Christopherson Date: Fri, 31 Jul 2026 10:19:26 -0700 In-Reply-To: <20260731171926.2629627-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260731171926.2629627-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.508.g3f0d502094-goog Message-ID: <20260731171926.2629627-3-seanjc@google.com> Subject: [PATCH v7 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Binbin Wu , Chao Gao , Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jim Mattson Due to a widespread Intel erratum (e.g. EMR158), programming the VMX-preemption timer with certain large values may cause the timer to expire earlier than expected. The recommended workaround is to cap the VMX-preemption timer value to strictly less than: 2^25 * CPUID.15H:EBX[31:0] / CPUID.15H:EAX[31:0]. Calculate the maximum "safe" preemption timer value during hardware setup based on CPUID 15H when available, and use the adjusted max value in all locations where KVM currently hardcodes the max architectural value, including in the subtle case where KVM soft-disables the timer. Don't apply the workaround when running as a VM, because absent explicit enumeration to state the bug is present (or not), it's L0's responsibility to faithfully emulate/virtualize the VMX preemption timer. WARN if the above logic would result in a max value of zero and fall back to the maximum architectural value, as the expectation is that real hardware will never provide problematic EAX/EBX values (which is another reason to ignore the erratum when running as a VM; there's less chance of a false positive on the WARN due to L0 providing an unanticipated ratio). Reported-by: Sean Christopherson Closes: https://lore.kernel.org/all/Zn9X0yFxZi_Mrlnt@google.com/ Suggested-by: Chao Gao Assisted-by: Gemini:Gemini-Next Reviewed-by: Chao Gao Signed-off-by: Jim Mattson Reviewed-by: Binbin Wu [sean: track inclusive max instead of exclusive limit, massage changelog] Signed-off-by: Sean Christopherson --- arch/x86/kvm/vmx/vmx.c | 40 +++++++++++++++++++++++++++++++++++----- 1 file changed, 35 insertions(+), 5 deletions(-) diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c index a07faa066ef0..6432d1329357 100644 --- a/arch/x86/kvm/vmx/vmx.c +++ b/arch/x86/kvm/vmx/vmx.c @@ -153,6 +153,7 @@ module_param(dump_invalid_vmcs, bool, 0644); #ifdef CONFIG_X86_64 static int __read_mostly cpu_preemption_timer_multi; static bool __read_mostly enable_preemption_timer =3D 1; +static u64 __ro_after_init preemption_timer_max_value; module_param_named(preemption_timer, enable_preemption_timer, bool, S_IRUG= O); #else #define enable_preemption_timer false @@ -8306,6 +8307,33 @@ static inline int u64_shl_div_u64(u64 a, unsigned in= t shift, return 0; } =20 +/* + * Workaround for a widespread Intel erratum (e.g. EMR158) where the + * VMX-preemption timer may expire earlier than expected when programmed + * with large values. The workaround is to cap the timer value to strictly + * less than 2^25 * CPUID.15H:EBX / CPUID.15H:EAX. + */ +static __init u64 calc_preemption_timer_max_value(void) +{ + const u64 ARCHITECTURAL_MAX_VALUE =3D UINT_MAX; + u32 eax, ebx, ecx, edx; + + if (cpu_feature_enabled(X86_FEATURE_HYPERVISOR)) + return ARCHITECTURAL_MAX_VALUE; + + if (cpuid_eax(0) < 0x15) + return ARCHITECTURAL_MAX_VALUE; + + cpuid(0x15, &eax, &ebx, &ecx, &edx); + if (!eax || !ebx) + return ARCHITECTURAL_MAX_VALUE; + + if (WARN_ON_ONCE(!(((u64)ebx << 25) / eax))) + return ARCHITECTURAL_MAX_VALUE; + + return min((((u64)ebx << 25) / eax) - 1, ARCHITECTURAL_MAX_VALUE); +} + static __init void vmx_setup_preemption_timer(void) { if (!cpu_has_vmx_preemption_timer()) @@ -8317,6 +8345,8 @@ static __init void vmx_setup_preemption_timer(void) cpu_preemption_timer_multi =3D vmx_misc_preemption_timer_rate(vmcs_config.misc); =20 + preemption_timer_max_value =3D calc_preemption_timer_max_value(); + if (tsc_khz) use_timer_freq =3D (u64)tsc_khz * 1000; use_timer_freq >>=3D cpu_preemption_timer_multi; @@ -8326,7 +8356,7 @@ static __init void vmx_setup_preemption_timer(void) * value. Don't use the timer if it might cause spurious exits * at a rate faster than 0.1 Hz (of uninterrupted guest time). */ - if (use_timer_freq > 0xffffffffu / 10) + if (use_timer_freq > preemption_timer_max_value / 10) enable_preemption_timer =3D false; } =20 @@ -8363,12 +8393,12 @@ int vmx_set_hv_timer(struct kvm_vcpu *vcpu, u64 gue= st_deadline_tsc, return -ERANGE; =20 /* - * If the delta tsc can't fit in the 32 bit after the multi shift, - * we can't use the preemption timer. + * If the delta tsc exceeds the preemption timer limit after the + * multi shift, we can't use the preemption timer. * It's possible that it fits on later vmentries, but checking * on every vmentry is costly so we just use an hrtimer. */ - if (delta_tsc >> (cpu_preemption_timer_multi + 32)) + if ((delta_tsc >> cpu_preemption_timer_multi) > preemption_timer_max_valu= e) return -ERANGE; =20 vmx->hv_deadline_tsc =3D tscl + delta_tsc; @@ -8402,7 +8432,7 @@ static void vmx_update_hv_timer(struct kvm_vcpu *vcpu= , bool force_immediate_exit vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc); vmx->loaded_vmcs->hv_timer_soft_disabled =3D false; } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) { - vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1); + vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, preemption_timer_max_value); vmx->loaded_vmcs->hv_timer_soft_disabled =3D true; } } --=20 2.55.0.508.g3f0d502094-goog