From nobody Thu Sep 24 13:36:57 2026 Received: from mail-pf1-f197.google.com (mail-pf1-f197.google.com [209.85.210.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E548712C534 for ; Thu, 24 Sep 2026 00:25:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209557; cv=none; b=kwG5+oDiRyCVd9DLSDCOB1hzsCOwEIec2D3Gb4d6BgnncgFOlg2jq6hSFtrRZ7NubQymRdslz64nNc+3M/XGcX4DSt6s+Kb3aWWfi1ufCcXBDRkPITwM2ceJmGNmL9Yf6c5klyhy93qp5pVu7U+QSjNGE2y86KMGKYKXMc2rnRM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209557; c=relaxed/simple; bh=20s750AZyJ6AuoaDZl//pQa5Coi9qzlUv/FwrTE1vgY=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=riMTYPEVTMVg34aroWxoFaIkCHVtjanzpuYSIpTD8KWLRYypMlZ0lXY2+8MWMMYNNxiNii5ZwtLAlsxFJQ40wbZai0xlrWITGSkWFcSNfV80GKAKCLuDYDvebNdAcotz5Dxaj9UZAc9mAbqOngyAAY8KxvUopj+vvyo2c3Um4uU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=QsJn1pQ2; arc=none smtp.client-ip=209.85.210.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="QsJn1pQ2" Received: by mail-pf1-f197.google.com with SMTP id d2e1a72fcca58-86a59faf521so1221199b3a.2 for ; Wed, 23 Sep 2026 17:25:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790209555; x=1790814355; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=uM6Oh64YX8bMZ7uBf0Og489cgMWoz35uyTpErdBKlDc=; b=QsJn1pQ2F6LmE6YNLsehgLfms4P97f4XhEc+zJLuzKR4nQF6BzTJr5QNu1DVBM41+7 857tirHqM59lZ3A03tvqiBQZK1sTXJiIlqAPx8ruBuxTKuOj3xgbJGWn0ccjiTrc5FHw bgxrYFJ1UlRShzqYcOYvHgGV4FjU6yRdpY9kKJ482HQRf7Rxvdma2GOig44cTq6iZrmb QjE0WVu1GfRUpHUp46tfDsg9f5lecs07gBvo4UXVV3XM2qIhEm8VnB9xlR20LjOzfmTh 3NlweknPxjmXb95Od59vS07DGzR/BC7uYA3ED9aC/KInPDpUSr75UIP5Nd/oYTOVRjUT 5l9Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790209555; x=1790814355; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=uM6Oh64YX8bMZ7uBf0Og489cgMWoz35uyTpErdBKlDc=; b=GyzpCjzuIXTRZNMJR+d1YKc/JbDcVpCdSWASEuJCkCwmK8/vXi1TMOGnNvyDbXYjR0 LFfIwffPblqeJ6dtXaHjJ9D1/2ZhogoF0ll13QOZyW5/9DZaI24v5RvupJ50vF/vOif/ oIsQkoLDdBGDjIrqq/t3i1AkHSFm3bhE0fuJE23IR/WcaDJO0waW0eQd9fx1zyPobK6t kORp7aC4PQZQXKA8hNTTpkRvA2R48eaX8XONDv3D3rBfPupJQKYvLpizXEynZGgaIVMA AFRKJcwnIwrygBL7/rJEKKO9NjQ57B/KROQ9YKo0wKqG2NDiau6sDHGj3iIiAJ35MjLo DBWQ== X-Forwarded-Encrypted: i=1; AKwUvBy/A6NByOKEcz+Wzmmjr6I+yHvtCyOzHr9omZ247LY9wO2ESwCh2RfLiZFkLX1pVJFGPXY6jnPSie5H/QI=@vger.kernel.org X-Gm-Message-State: AFuF++kxmubUy6eXLrcSkoHYiuG9l3bDK8A3kvtJy+HfTerilPS8Kzac xIMFolDPOrRpFYyeBRdVCwi86XMS9o7xKOtRtscR9MKIU/iKDz7joZopz0AALIdfwEF7rK4FWta 0AfF/pFJ04KqjbA== X-Received: from pga8.prod.google.com ([2002:a05:6a02:4f88:b0:cc7:66f5:c110]) (user=jmattson job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:a109:b0:3dd:85a9:55b3 with SMTP id adf61e73a8af0-3de0e7730e0mr661723637.30.1790209555098; Wed, 23 Sep 2026 17:25:55 -0700 (PDT) Date: Wed, 23 Sep 2026 17:25:42 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: Subject: [PATCH v2 1/4] KVM: x86: Advertise EFER_LMSLE_MBZ when KVM disallows EFER.LMSLE From: Jim Mattson To: seanjc@google.com, pbonzini@redhat.com Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, nikunj@amd.com, yosry@kernel.org, Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Set EFER_LMSLE_MBZ, CPUID.80000008H:EBX[bit 20], in KVM's supported CPUID whenever KVM refuses to set EFER.LMSLE, not just when hardware enumerates t= he defeature. KVM allows EFER.LMSLE if and only if nested SVM is supported, b= ut passes through hardware's EFER_LMSLE_MBZ as-is, i.e. KVM tells userspace th= at long mode segment limits are available on an Intel CPU, and on an AMD CPU w= ith kvm_amd.nested=3D0, and then rejects WRMSR(EFER) with EFER.LMSLE=3D1. The = bit means precisely "EFER.LMSLE must be zero." Note, this is purely an enumeration change. EFER.LMSLE is already gated on nested SVM: kvm_setup_efer_caps() adds EFER_LMSLE to supported_efer_bits on= ly if X86_FEATURE_SVM is supported, and svm_set_cpu_caps() sets X86_FEATURE_SVM if and only if "nested" is true. KVM's handling of EFER is unchanged; only what KVM tells userspace changes. Note #2, KVM now sets the defeature on Intel CPUs as well. That is no different than KVM setting NullSelectorClearsBase, another AMD-defined bit,= on CPUs that don't have the errata. And LMSLE's dependency on nested SVM is a historical artifact of commit eec4b140c924 ("KVM: SVM: Allow EFER.LMSLE to = be set with nested svm"), not an architectural requirement. Be aware that this changes KVM_GET_SUPPORTED_CPUID on all Intel hosts and on AMD hosts with kvm_amd.nested=3D0, i.e. userspace that reflects KVM's suppo= rted CPUID into the guest will start enumerating EFER_LMSLE_MBZ=3D1 to newly cre= ated guests. That's the correct enumeration, and it takes away only an EFER bit that KVM has never allowed to be set on such hosts, but it *is* guest visib= le. Deliberately omit a Fixes: tag: the misenumeration is benign in practice, as KVM's WRMSR(EFER) behavior is and always has been correct, and backporting a guest-visible CPUID change isn't worth the risk. Opportunistically key the EFER_LMSLE decision off kvm_cpu_cap_has() instead= of boot_cpu_has(), as PASSTHROUGH_F() sets EFER_LMSLE_MBZ based on *raw* CPUID, i.e. "clearcpuid=3D13:20" would otherwise get KVM to enumerate the defeatur= e and allow EFER.LMSLE=3D1, which is guaranteed to fail on VMRUN. Document KVM's enumeration of the defeature in api.rst. Assisted-by: LLM Signed-off-by: Jim Mattson --- Documentation/virt/kvm/api.rst | 11 +++++++++++ arch/x86/kvm/cpuid.c | 4 ++++ arch/x86/kvm/svm/svm.c | 10 ++++++++++ arch/x86/kvm/vmx/vmx.c | 7 +++++++ arch/x86/kvm/x86.c | 9 ++++++++- 5 files changed, 40 insertions(+), 1 deletion(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index e0430cc750c9..3c5bcdb0923f 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -9635,6 +9635,17 @@ On older versions of Linux, CPU[EAX=3D1]:ECX[24] (TS= C_DEADLINE) is not reported by is present and the kernel has enabled in-kernel emulation of the local API= C. On newer versions, ``KVM_GET_SUPPORTED_CPUID`` does report the bit as avai= lable. =20 +Long mode segment limits +~~~~~~~~~~~~~~~~~~~~~~~~ + +CPU[EAX=3D0x80000008]:EBX[20] (EFER_LMSLE_MBZ) is a "defeature" bit: it is= set +when the CPU does *not* support long mode segment limits, and so requires +EFER.LMSLE to be zero. KVM reports the bit via ``KVM_GET_SUPPORTED_CPUID`= ` if +and only if KVM refuses to set EFER.LMSLE, i.e. if the CPU doesn't support= long +mode segment limits, or if nested SVM is unsupported. KVM therefore repor= ts +the bit on all Intel hosts, as KVM allows EFER.LMSLE only when nested SVM = is +enabled. + CPU topology ~~~~~~~~~~~~ =20 diff --git a/arch/x86/kvm/cpuid.c b/arch/x86/kvm/cpuid.c index 851f151efb35..dbe20d5e6f80 100644 --- a/arch/x86/kvm/cpuid.c +++ b/arch/x86/kvm/cpuid.c @@ -1171,6 +1171,10 @@ void kvm_initialize_cpu_caps(void) F(AMD_STIBP), F(AMD_STIBP_ALWAYS_ON), F(AMD_IBRS_SAME_MODE), + /* + * Vendor code also sets EFER_LMSLE_MBZ if KVM itself + * can't support EFER.LMSLE, e.g. if nested SVM is disabled. + */ PASSTHROUGH_F(EFER_LMSLE_MBZ), F(AMD_PSFD), F(AMD_IBPB_RET), diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c index 7d59d301e1e5..58768e49505e 100644 --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ -5571,6 +5571,16 @@ static __init void svm_set_cpu_caps(void) boot_cpu_has(X86_FEATURE_AMD_SSBD)) kvm_cpu_cap_set(X86_FEATURE_VIRT_SSBD); =20 + /* + * Tell userspace that EFER.LMSLE must be zero if nested SVM is + * disabled, as KVM allows EFER.LMSLE if and only if nested SVM is + * supported (a historical artifact of commit eec4b140c924 ("KVM: SVM: + * Allow EFER.LMSLE to be set with nested svm"), not an architectural + * requirement). + */ + if (!nested) + kvm_cpu_cap_set(X86_FEATURE_EFER_LMSLE_MBZ); + if (enable_pmu) { /* * Enumerate support for PERFCTR_CORE if and only if KVM has diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c index 612ab07d4100..f144c1e1c63f 100644 --- a/arch/x86/kvm/vmx/vmx.c +++ b/arch/x86/kvm/vmx/vmx.c @@ -8137,6 +8137,13 @@ static __init void vmx_set_cpu_caps(void) kvm_cpu_cap_clear(X86_FEATURE_IBT); } =20 + /* + * CPUID 0x80000008. Tell userspace that EFER.LMSLE must be zero; KVM + * never allows EFER.LMSLE to be set on Intel CPUs, as KVM supports long + * mode segment limits only in conjunction with nested SVM. + */ + kvm_cpu_cap_set(X86_FEATURE_EFER_LMSLE_MBZ); + kvm_setup_xss_caps(); kvm_finalize_cpu_caps(); } diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 79468ddfe473..830cb9320d89 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -6930,7 +6930,14 @@ static void kvm_setup_efer_caps(void) =20 if (kvm_cpu_cap_has(X86_FEATURE_SVM)) { kvm_caps.supported_efer_bits |=3D EFER_SVME; - if (!boot_cpu_has(X86_FEATURE_EFER_LMSLE_MBZ)) + + /* + * Enumerating EFER_LMSLE_MBZ and allowing EFER.LMSLE=3D1 + * would be nonsensical. Note, vendor code sets the defeature + * if KVM can't support EFER.LMSLE for any reason, i.e. this + * needs to consult KVM's capabilities, not just raw CPUID. + */ + if (!kvm_cpu_cap_has(X86_FEATURE_EFER_LMSLE_MBZ)) kvm_caps.supported_efer_bits |=3D EFER_LMSLE; } } --=20 2.56.0.rc1.315.gc6ed9934b7-goog From nobody Thu Sep 24 13:36:57 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DAF6F2EC0B0 for ; Thu, 24 Sep 2026 00:25:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209558; cv=none; b=J6xI6dzCGrifYYiHypOOH1yO10N1CUXa3gFjE6vzM2yGPV4lsBJPexPlx0NfBJbR3PK36ee5bH2lOJfqfKQDoLkBYxdvFY9rhbeE72FzsAr9mJggoOao7zOmM7cd5hAjnNfvIS1BaAYgpKWn0X84xTkS79DkPiqwai+VJTHSXyQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209558; c=relaxed/simple; bh=YSrpIK2pVF/5FNDzNmbXXp4qlqZW/WMDSRuIs2fpCGI=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=HjgqNHHDa4UFjkIFhoWd9HuhCwDWqE/QWu4oEVDxfmyYX0FkmCQY3OqiEcwXM5ag3bg+3CfwJ+JHltwIqf+TgBMcYeqCoXdb9SHYe8sIj+n5vvBzCDFRPceRI2w3cu0/bGL5K8k8/p/I67ZKAWPL3Z5/g5fFB9MgL5ll4X2eYIs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=UkqBIMER; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="UkqBIMER" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-cc73895c7d5so1905302a12.1 for ; Wed, 23 Sep 2026 17:25:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790209556; x=1790814356; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=r+yU0a+PweSOLITCJc4zym8U6EzzEmx8B9YTHY0q+B0=; b=UkqBIMERU7OjqszGAva34vYyKMIQ3uLY0PK/2FrD8tkEGvGGkASx3Z6KoSbKpAs4WZ t66Qz0GEJ6J2WOcg7uS5tK05JKwSVkqcsDsGGFAJr1fCYQdmsH8dtKBksNUJZ4QaHVgE epaS4yW/K6o3Xk+uoq9Xpv7n8wt/E/sB9XBMP6tH6RdeFj/W+g4oKgRYB7ex2rQibfKB eTcU2KNsfbN+3MgbTmGffRaGInhBhZGoT94dJBOic29HCwsmHn4gmyr4fyy/g5DLQTqj JKSYK0JXz2Gx6zC7jrewUIcmxmrO+uRkdKtsgnRHhpGaADL/2JGKg9Yph27+HiAVWxZm EibA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790209556; x=1790814356; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=r+yU0a+PweSOLITCJc4zym8U6EzzEmx8B9YTHY0q+B0=; b=ngPIoUGBvlrM83/8Fxs4S/f3Vsd4PGhyhPHVQB85NT4tDGDTimPbJf81REZ7J7TcIr ddGJiT5wRDTMhRcxuvH7jJivDQaWzhAIFwzZpLeNYO7njmUnSx5PU3ekDvgiTxeGtEYd K0z7d73jqU1lAlgUzGYW4XnS44iJqtKVEZtAnk+WcJpBWjpUOlG9SQo+mpo9K11igNiZ wD+eiHxeZ16PWEpBCRxHRulqk8BZyVQIFToQWLQc4Zpmqad8vs9usT01QdSTULO4TxzA 3PvibDvhnbYcf2Fnyai/Alacct9NN5KQMtbfdQFBs9v/hqquk7azeXJhafawAVXSRW4v 0Qtw== X-Forwarded-Encrypted: i=1; AKwUvBzuGVv4XhWJdPUjBlDIO+bo5hxfrH/oi+/LvFBRCnqq3NTcTa9bov+wxY/vg7TVnkg2vfo/74W0K36RqBM=@vger.kernel.org X-Gm-Message-State: AFuF++lu8Ng4ykfT6v9DOn0FR0khrozINi/xuaALIFOW2QWYQpX/xBZD kLtDijsf7a8SBOYIUVxrOStragcjf6xqWFT++/6/6vUKFHmNoUFLysknk3nqCbKFh0II09MiSTK KiDKZ79UvrBDSeA== X-Received: from pga22.prod.google.com ([2002:a05:6a02:4f96:b0:cc4:33fd:a4f7]) (user=jmattson job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:4492:b0:3dd:85aa:452c with SMTP id adf61e73a8af0-3de0e770a3bmr694890637.30.1790209555955; Wed, 23 Sep 2026 17:25:55 -0700 (PDT) Date: Wed, 23 Sep 2026 17:25:43 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: <48e96e1a917a4bc543142014fa27a2357b114b60.1790208413.git.jmattson@google.com> Subject: [PATCH v2 2/4] KVM: x86: Honor the guest's EFER_LMSLE_MBZ From: Jim Mattson To: seanjc@google.com, pbonzini@redhat.com Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, nikunj@amd.com, yosry@kernel.org, Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Reject guest attempts to set EFER.LMSLE if userspace enumerates EFER_LMSLE_MBZ, CPUID.80000008H:EBX[bit 20], in guest CPUID, so that usersp= ace can hide long mode segment limits from the guest even on CPUs that support LMSLE, e.g. for migration compatibility between Rome and Milan. Enumerate the defeature as *partially* emulated, i.e. as supported by KVM b= ut never advertised, so that kvm_vcpu_after_set_cpuid() propagates userspace's= bit into vcpu->arch.cpu_caps. Checking guest_cpu_cap_has() in __kvm_valid_efer= () doesn't suffice on its own, because KVM clears the defeature in kvm_cpu_caps precisely on the CPUs where it needs to be emulated, i.e. cpu_caps would ne= ver hold the bit no matter what userspace puts in guest CPUID. Don't advertise= the defeature in KVM_GET_SUPPORTED_CPUID, as that would take EFER.LMSLE away fr= om userspace that blindly stuffs KVM's supported CPUID into KVM_SET_CPUID2. MONITOR/MWAIT gets the same treatment. Enforce the defeature only for guest-initiated writes, as it's a guest CPUID consistency check, not a host capability. See commit 11988499e62b ("KVM: x= 86: Skip EFER vs. guest CPUID checks for host-initiated writes"). Note, KVM_SET_SREGS does reject EFER.LMSLE, as it runs the full set of guest CPUID checks, as does nested VMRUN, i.e. L1 can't sneak EFER.LMSLE into L2 via vmcb12. Mask EFER_LMSLE out of the value consumed by efer_trap(), which handles SVM_EXIT_EFER_WRITE_TRAP. EFER writes are *trapped*, not intercepted, i.e. hardware has already committed the write by the time KVM gains control, and the trap is enabled if and only if the guest is SEV-ES, whose EFER lives in the encrypted VMSA and so can't be fixed up by KVM. Rejecting the write wo= uld inject a #GP *and* leave EFER.LMSLE set in the guest, which is strictly wor= se than honoring a write that hardware itself allowed. EFER_SVME is already masked out for a similar "KVM can't enforce this here" reason. Opportunistically fix the whitespace damage around __kvm_valid_efer(). Document the resulting two-level contract in api.rst, i.e. that userspace m= ay set the defeature even when KVM doesn't enumerate it. Suggested-by: Sean Christopherson Assisted-by: LLM Signed-off-by: Jim Mattson --- Documentation/virt/kvm/api.rst | 14 ++++++++++++++ arch/x86/kvm/cpuid.c | 16 ++++++++++++++++ arch/x86/kvm/msrs.c | 12 +++++++++++- arch/x86/kvm/svm/svm.c | 11 ++++++++++- 4 files changed, 51 insertions(+), 2 deletions(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index 3c5bcdb0923f..2852bf94828b 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -9646,6 +9646,20 @@ mode segment limits, or if nested SVM is unsupported= . KVM therefore reports the bit on all Intel hosts, as KVM allows EFER.LMSLE only when nested SVM = is enabled. =20 +KVM never reports the bit via ``KVM_GET_EMULATED_CPUID``, but userspace ma= y set +it via ``KVM_SET_CPUID2`` even on a host where KVM doesn't report it. KVM +honors the guest's enumeration and rejects EFER.LMSLE=3D1 accordingly. Th= at lets +userspace defeature a vCPU on a host that *does* support long mode segment +limits, so that the vCPU can later be migrated to a host that doesn't, e.g= . so +that a vCPU created on AMD Rome can be migrated to Milan and later, which +dropped support for long mode segment limits. The opposite direction need= s no +emulation, as a host that lacks long mode segment limits already enumerate= s the +defeature. + +Note, ``KVM_SET_MSRS`` is exempt from the check, as host-initiated MSR wri= tes +skip guest CPUID checks so that userspace can set MSRs before it sets guest +CPUID. ``KVM_SET_SREGS`` and nested VMRUN are not exempt. + CPU topology ~~~~~~~~~~~~ =20 diff --git a/arch/x86/kvm/cpuid.c b/arch/x86/kvm/cpuid.c index dbe20d5e6f80..53d205eda335 100644 --- a/arch/x86/kvm/cpuid.c +++ b/arch/x86/kvm/cpuid.c @@ -1419,6 +1419,22 @@ static int cpuid_func_emulated(struct kvm_cpuid_entr= y2 *entry, u32 func, u32 ind if (kvm_cpu_cap_has(X86_FEATURE_RDTSCP)) entry->ecx =3D feature_bit(RDPID); return 1; + case 0x80000008: + /* + * Honor the guest's EFER_LMSLE_MBZ even if the underlying CPU + * allows setting EFER.LMSLE, e.g. to allow migrating a vCPU + * between hosts with and without EFER.LMSLE support. To avoid + * breaking existing setups that reflect KVM's supported CPUID + * into the guest, KVM doesn't advertise EFER_LMSLE_MBZ unless + * KVM *can't* support EFER.LMSLE=3D1. + */ + if (include_partially_emulated && + !kvm_cpu_cap_has(X86_FEATURE_EFER_LMSLE_MBZ)) { + entry->ebx |=3D feature_bit(EFER_LMSLE_MBZ); + return 1; + } + /* Nothing in 0x80000008 is fully emulated, don't emit an entry. */ + return 0; default: return 0; } diff --git a/arch/x86/kvm/msrs.c b/arch/x86/kvm/msrs.c index dd3bb04878ca..b519fb90776e 100644 --- a/arch/x86/kvm/msrs.c +++ b/arch/x86/kvm/msrs.c @@ -598,9 +598,19 @@ static bool __kvm_valid_efer(struct kvm_vcpu *vcpu, u6= 4 efer) if (efer & EFER_NX && !guest_cpu_cap_has(vcpu, X86_FEATURE_NX)) return false; =20 - return true; + /* + * EFER_LMSLE_MBZ is a "defeature" bit, i.e. is set when the CPU does + * *not* support long mode segment limits, and so is the only EFER + * check whose polarity is inverted: EFER.LMSLE is legal if and only if + * the guest does *not* have the defeature. + */ + if (efer & EFER_LMSLE && + guest_cpu_cap_has(vcpu, X86_FEATURE_EFER_LMSLE_MBZ)) + return false; =20 + return true; } + bool kvm_valid_efer(struct kvm_vcpu *vcpu, u64 efer) { if (efer & ~kvm_caps.supported_efer_bits) diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c index 58768e49505e..90a80aad672f 100644 --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ -2759,10 +2759,19 @@ static int efer_trap(struct kvm_vcpu *vcpu) * bit in svm_set_efer(), but __kvm_valid_efer() checks it against * whether the guest has X86_FEATURE_SVM - this avoids a failure if * the guest doesn't have X86_FEATURE_SVM. + * + * Clear EFER_LMSLE for a related reason: EFER writes are *trapped*, + * not intercepted, i.e. hardware has already committed the write by + * the time KVM gains control, and the trap is enabled if and only if + * the guest is SEV-ES, whose EFER lives in the encrypted VMSA and so + * can't be fixed up by KVM. Rejecting EFER.LMSLE=3D1 would inject a #GP + * *and* leave EFER.LMSLE set in the guest, which is strictly worse + * than honoring a write that hardware itself allowed. */ msr_info.host_initiated =3D false; msr_info.index =3D MSR_EFER; - msr_info.data =3D to_svm(vcpu)->vmcb->control.exit_info_1 & ~EFER_SVME; + msr_info.data =3D to_svm(vcpu)->vmcb->control.exit_info_1 & + ~(EFER_SVME | EFER_LMSLE); ret =3D kvm_set_msr_common(vcpu, &msr_info); =20 return kvm_complete_insn_gp(vcpu, ret); --=20 2.56.0.rc1.315.gc6ed9934b7-goog From nobody Thu Sep 24 13:36:57 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BB9232F7F05 for ; Thu, 24 Sep 2026 00:25:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209559; cv=none; b=kVeICb+EOCIssQlulz0NrQX9PeKrRee6bMNj0w14wC2C+CAk7c9hIfvG2McXc+J8iDc0F0OzlfzPEeikdQNB45mkCbNkyulVij0sPUomtatE/G75OGM2uKd4Amj73OITRrclhxHtt6z7KCZENbOntHSEHuCcjDJNEtLzw4WrA9s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209559; c=relaxed/simple; bh=je0dnkfFNKk+uoQkxx/XvcJ1dcav2J5cIsKIcj+09fM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=pYkwh4EObO2RnB0ztowme8tiGSxImHhuKCtygKQKDZvpKxEPUtG2F0pu4Kgo4tH4niykslJoe7N6p16xCmwBUpy+x/H/sTUj4TJnMiQKCVfiaZVozJ5Gk9LfT6iOK8Uk7IfqWfQ0qE4K99mM1InKyV1gtHFtfPYtuu/INblfljg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=jfYUyMfK; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="jfYUyMfK" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2dd188e16a2so11708905ad.2 for ; Wed, 23 Sep 2026 17:25:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790209557; x=1790814357; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=rWvEr8HcLzBvZNb2ZaMPRoPvjV6bscZy0wkTIfmsztY=; b=jfYUyMfK+jYXpAVEdQvmv5EN1XahFjNUqqYd/3ncM3OYOky83e/GfH1D10ij31b/F6 FmVviRK4qfMjIEyu1ocphuZiB2TTyJhok9It8EweEUYzNGytpnltM0ktWsB7mdadGtzK 8Z2GxKD6gJzXlRUFSq+h6WE1djTVM/jcvVHNyQM3mrfALwcP9vOtFnYDsYc+8hYddlQ2 yxgDFu1vSFMZje9DprCnMRoCA8g7BIflNIpDpAOmHbuzWmtm3g25Zm2mP8DkwX2irwHx hOgEH+Vp5Z08RyPaKniNFX3TOmeUbBPNHyJPeMedExWgmUiTqZY2Aq7ZHeIa5I/EFZrA T+2g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790209557; x=1790814357; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rWvEr8HcLzBvZNb2ZaMPRoPvjV6bscZy0wkTIfmsztY=; b=xP/zqefM9JFNru3Do4vAHtLoOF33hd5rWqipYkK8I81dUfniGrrSbi6F3fiQm3YcBT 9r7icuvWffy/IXm5b3D07NN10tsL0PPDY3xCDz8be4bTQP8DC+3/mKr8XLPNxwfwtN/a GxneVgjavf7hjXucxo3GQPZOpV6vfOLQQIK01EiLtcCyrBInSOiEnRmw0CxLKloc/BJP AC70dubAOEN5nck1hRXBRMpBS6D0Zz12eYSGVIyP8KhWI1QlTHZ3ASzyIjmit7MfVcN9 pHSv1ebOlmN87dS7lLo9N4bMcXF24SfGLGor7jMjD8jkJlyI/y1+t36N0rCP3idlIQ9a e6Jw== X-Forwarded-Encrypted: i=1; AKwUvBw1n52Rhwdnkt0kr3Ncjr9csrP1f6wzeWoGpTVuc9kJKSpb7sW2hq0fEECq0fMd9IDwpZLV4+sYEmJ8k5s=@vger.kernel.org X-Gm-Message-State: AFuF++mrwdU/9rLq/wpBT3U2U5XJIaMg3Dz+ct3H/L8VJoxKWQojabQ8 Wdz7KOTi24UvN8g8dINIls1QHbVJuhaiFslsIJqkAq7LT5d1Fe1AK0tlOCauzztSK7S3E4u8uai Xz0orMA2dRdG/wQ== X-Received: from ples3.prod.google.com ([2002:a17:903:2143:b0:2df:527f:374a]) (user=jmattson job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:fda7:b0:2dd:c053:d73f with SMTP id d9443c01a7336-2df7dc37ebbmr6151175ad.38.1790209556802; Wed, 23 Sep 2026 17:25:56 -0700 (PDT) Date: Wed, 23 Sep 2026 17:25:44 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: Subject: [PATCH v2 3/4] KVM: selftests: Rename svm_nested_clear_efer_svme to svm_nested_efer_test From: Jim Mattson To: seanjc@google.com, pbonzini@redhat.com Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, nikunj@amd.com, yosry@kernel.org, Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Rename svm_nested_clear_efer_svme.c to svm_nested_efer_test.c so that the test can host coverage for other EFER bits whose behavior is tied to nested SVM, and so that the name matches the svm_nested__test convention us= ed by svm_nested_pat_test, svm_nested_shutdown_test, etc. No functional change intended. Assisted-by: LLM Signed-off-by: Jim Mattson --- tools/testing/selftests/kvm/Makefile.kvm | 2 +- .../{svm_nested_clear_efer_svme.c =3D> svm_nested_efer_test.c} | 0 2 files changed, 1 insertion(+), 1 deletion(-) rename tools/testing/selftests/kvm/x86/{svm_nested_clear_efer_svme.c =3D> = svm_nested_efer_test.c} (100%) diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selft= ests/kvm/Makefile.kvm index 96bab7002d39..2554464d8b96 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -117,7 +117,7 @@ TEST_GEN_PROGS_x86 +=3D x86/state_test TEST_GEN_PROGS_x86 +=3D x86/vmx_preemption_timer_test TEST_GEN_PROGS_x86 +=3D x86/svm_vmcall_test TEST_GEN_PROGS_x86 +=3D x86/svm_int_ctl_test -TEST_GEN_PROGS_x86 +=3D x86/svm_nested_clear_efer_svme +TEST_GEN_PROGS_x86 +=3D x86/svm_nested_efer_test TEST_GEN_PROGS_x86 +=3D x86/svm_nested_shutdown_test TEST_GEN_PROGS_x86 +=3D x86/svm_nested_soft_inject_test TEST_GEN_PROGS_x86 +=3D x86/svm_nested_vmcb12_gpa diff --git a/tools/testing/selftests/kvm/x86/svm_nested_clear_efer_svme.c b= /tools/testing/selftests/kvm/x86/svm_nested_efer_test.c similarity index 100% rename from tools/testing/selftests/kvm/x86/svm_nested_clear_efer_svme.c rename to tools/testing/selftests/kvm/x86/svm_nested_efer_test.c --=20 2.56.0.rc1.315.gc6ed9934b7-goog From nobody Thu Sep 24 13:36:57 2026 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C24B32EEE7C for ; Thu, 24 Sep 2026 00:25:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209564; cv=none; b=N4u86ITBWdlPs8GF6K1TjvHwQxV53F7w3NHzKF9FcdGeHwxn5IAiEoJPofrJz/HoITBX+7WjkGSQeM8qXSfyPCO1J0oWYklqlhQLPK3mc68WhSNuoRgbJwJ00NXjeZVV5FFvwAYazWpb5LeIK8gpDcSsNfLZ9pQnKO3UUGLnZKY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790209564; c=relaxed/simple; bh=W40w6Xz5WizzC2CsbKNXLo3vVFw6UxWPXABB0XtWee4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Sfk6YHxpZWOueN1Hrypdk4Xs2s+e8EyEldWLl1EhIEdxrbXWFmRIA4+7A2CI2XmBbnxXP1NTnEQbf+Zpf4RDk2QTxSE1G24/RSj1SoeT0lCuHFj9/q3d3D+4yFQmdHRuDGbCYj3QiSqMEntURtrfN0L0F7Od6KT7CJ/2foIlHRQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ASWHXTBK; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jmattson.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ASWHXTBK" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-cc21bc2923fso938704a12.2 for ; Wed, 23 Sep 2026 17:25:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790209558; x=1790814358; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=pso3DvjmD9/LlPFbVvpLmkrNSJZEBDMBWUOU1ztjzQc=; b=ASWHXTBKQJsEEc/Q0Rs2Z8oaRByBxo87L8Z564wkcGQltzfZyU67YiBlKSVBNp5C+h S26ktirIMM4P+L6j35/Dw7wfHPWn2gG2f9gx6AjLoIcaCLNl8l1BylXQ6BfpDtYgkII0 LVIcW6X8C9tPs8r8lMIri5plzU9NOiE6hRKR4Ku89Gqg5bn4EVGvLwOWS1BWSQAqbyz5 19arxFQY6nAxEtvxRCVDUGYjXtZkAdSTpLmAd4Boxbrqrmpq7WzSqG0xX5vwx1A1sKm7 p372VyQx6i4JVW4kTE0vXEa23vDopn/1TquMgFXTPooWFnW6TzHcuzCH93lo/p6LduFw umOQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790209558; x=1790814358; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pso3DvjmD9/LlPFbVvpLmkrNSJZEBDMBWUOU1ztjzQc=; b=it4CIf0Lo22TX6CgLjB3zhs/UxV3CSTV+swU2jngZizxNPgIAREasG+f5vENkB40fB X0qXBe0fj71y/QCIzBv1MqLbg+Oah3MyBF7/fcQXZp96CJG+L4C0poW2NJfL6mGx6eNG pcVnyN3UsB/rn62XKZakb6TbAdZZ0bcQQfthSHdmIkuo7y5Hceoy8vid8fERdPdTVwRR mFEdksQOnr8XlZkQ4x2hCB7N4+VOp+AwvWAxJGieGtRHpwFki71cAX2Rn/lqQoOaLvaQ sR/NQMyiyXkF6R4un9X7HZ9jqWLkEUgsT3Mp5EBT5G/OsfnHVHaB1kBQFIZt54Kmcad9 AiQQ== X-Forwarded-Encrypted: i=1; AKwUvBwO5HFNLKRD5+OIqppUNiIwvf3KVAbfH9+Bzn223jHx6IO6OwESBUDa5deDEOjO4JI1+Bg3p0a6c70obPM=@vger.kernel.org X-Gm-Message-State: AFuF++lKIO/Vb7SdlmSER8zpDdG34XhIOyZh1ZXASL7GushJ8dn8xm29 EvuzY9K40mvF3LjHugYlfOMdVxbC/vgpd4pBm6xHkpsDJ511qt7FajQ/QOSzCze/Q6aCYDTka9l JV/S4icv7Awrm7A== X-Received: from pgbfq8.prod.google.com ([2002:a05:6a02:2988:b0:c96:47b4:65c8]) (user=jmattson job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:2d07:b0:3dd:a195:dd5b with SMTP id adf61e73a8af0-3de0e93ffabmr703354637.61.1790209557724; Wed, 23 Sep 2026 17:25:57 -0700 (PDT) Date: Wed, 23 Sep 2026 17:25:45 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: Subject: [PATCH v2 4/4] KVM: selftests: Add coverage for the EFER_LMSLE_MBZ defeature From: Jim Mattson To: seanjc@google.com, pbonzini@redhat.com Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, nikunj@amd.com, yosry@kernel.org, Jim Mattson Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add coverage for KVM's virtualization of EFER_LMSLE_MBZ, CPUID.80000008H:EBX[bit 20], which is set when the CPU does *not* support l= ong mode segment limits. Verify that KVM's enumeration of the defeature matches the expectation derived from hardware plus kvm_amd's "nested" module param,= so that the assertion doesn't compare KVM's enumeration against itself, and th= at WRMSR(EFER), nested VMRUN, and KVM_SET_SREGS reject EFER.LMSLE=3D1 if and o= nly if guest CPUID enumerates the defeature. Also verify that host-initiated KVM_SET_MSRS is exempt from the check, as host-initiated writes skip guest CPUID checks so that userspace can set MSRs before it sets guest CPUID. Extend svm_nested_efer_test rather than add yet another test binary for a single CPUID bit; the test already has the L1/L2 harness needed to exercise EFER.LMSLE. Note, only a CPU that supports LMSLE, i.e. Rome and earlier, exercises the emulation path; elsewhere KVM enumerates the defeature straig= ht from hardware. Opportunistically drop the unnecessary #include of vmx.h. Assisted-by: LLM Signed-off-by: Jim Mattson --- .../selftests/kvm/include/x86/processor.h | 1 + .../selftests/kvm/x86/svm_nested_efer_test.c | 249 +++++++++++++++++- 2 files changed, 237 insertions(+), 13 deletions(-) diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/te= sting/selftests/kvm/include/x86/processor.h index 6e6f70035508..ba5d8b37edc1 100644 --- a/tools/testing/selftests/kvm/include/x86/processor.h +++ b/tools/testing/selftests/kvm/include/x86/processor.h @@ -216,6 +216,7 @@ struct kvm_x86_cpu_feature { #define X86_FEATURE_INVTSC KVM_X86_CPU_FEATURE(0x80000007, 0, EDX, 8) #define X86_FEATURE_RDPRU KVM_X86_CPU_FEATURE(0x80000008, 0, EBX, 4) #define X86_FEATURE_AMD_IBPB KVM_X86_CPU_FEATURE(0x80000008, 0, EBX, 12) +#define X86_FEATURE_EFER_LMSLE_MBZ KVM_X86_CPU_FEATURE(0x80000008, 0, EBX,= 20) #define X86_FEATURE_NPT KVM_X86_CPU_FEATURE(0x8000000A, 0, EDX, 0) #define X86_FEATURE_LBRV KVM_X86_CPU_FEATURE(0x8000000A, 0, EDX, 1) #define X86_FEATURE_NRIPS KVM_X86_CPU_FEATURE(0x8000000A, 0, EDX, 3) diff --git a/tools/testing/selftests/kvm/x86/svm_nested_efer_test.c b/tools= /testing/selftests/kvm/x86/svm_nested_efer_test.c index 6bc301207cbc..765476c79c96 100644 --- a/tools/testing/selftests/kvm/x86/svm_nested_efer_test.c +++ b/tools/testing/selftests/kvm/x86/svm_nested_efer_test.c @@ -1,16 +1,20 @@ // SPDX-License-Identifier: GPL-2.0-only /* + * Tests for KVM's handling of EFER bits whose behavior is tied to nested = SVM. + * * Copyright (C) 2026, Google LLC. */ +#include "test_util.h" #include "kvm_util.h" -#include "vmx.h" +#include "processor.h" #include "svm_util.h" #include "kselftest.h" =20 +static bool l2_ran; =20 -static void l2_guest_code(void) +static void l2_clear_efer_svme(void) { - unsigned long efer =3D rdmsr(MSR_EFER); + u64 efer =3D rdmsr(MSR_EFER); =20 /* generic_svm_setup() initializes EFER_SVME set for L2 */ GUEST_ASSERT(efer & EFER_SVME); @@ -20,31 +24,250 @@ static void l2_guest_code(void) GUEST_ASSERT(0); } =20 -static void l1_guest_code(struct svm_test_data *svm) +static void l1_clear_efer_svme(struct svm_test_data *svm) { - generic_svm_setup(svm, l2_guest_code); + generic_svm_setup(svm, l2_clear_efer_svme); run_guest(svm->vmcb, svm->vmcb_gpa); =20 /* Unreachable, L1 should be shutdown */ GUEST_ASSERT(0); } =20 -int main(int argc, char *argv[]) +static void l2_lmsle(void) +{ + GUEST_ASSERT(rdmsr(MSR_EFER) & EFER_LMSLE); + l2_ran =3D true; + vmmcall(); +} + +static void l1_lmsle(struct svm_test_data *svm) +{ + bool lmsle_mbz =3D this_cpu_has(X86_FEATURE_EFER_LMSLE_MBZ); + struct vmcb *vmcb =3D svm->vmcb; + u64 efer =3D rdmsr(MSR_EFER); + + /* + * Selftests' vCPUs are created with EFER.LMSLE clear; the sub-tests + * below need to start from a clean slate. + */ + GUEST_ASSERT(!(efer & EFER_LMSLE)); + GUEST_ASSERT(!l2_ran); + + /* + * Per AMD APM vol. 2, if CPUID.80000008H:EBX[bit 20] is set, "64-bit + * mode segment limit checking is not supported and attempting to set + * EFER.LMSLE =3D 1 causes a #GP exception". + */ + if (lmsle_mbz) { + GUEST_ASSERT_EQ(wrmsr_safe(MSR_EFER, efer | EFER_LMSLE), GP_VECTOR); + GUEST_ASSERT(!(rdmsr(MSR_EFER) & EFER_LMSLE)); + } else { + GUEST_ASSERT_EQ(wrmsr_safe(MSR_EFER, efer | EFER_LMSLE), 0); + GUEST_ASSERT(rdmsr(MSR_EFER) & EFER_LMSLE); + + /* + * Restore EFER so that generic_svm_setup() doesn't propagate + * EFER.LMSLE into vmcb12 on its own, i.e. so that the VMRUN + * sub-test actually tests what it thinks it's testing. + */ + wrmsr(MSR_EFER, efer); + } + + /* + * VMRUN's consistency checks reject "any MBZ bit of EFER", i.e. a + * vmcb12 with EFER.LMSLE set must generate VMEXIT_INVALID when the + * defeature is enumerated. + */ + generic_svm_setup(svm, l2_lmsle); + vmcb->save.efer |=3D EFER_LMSLE; + run_guest(vmcb, svm->vmcb_gpa); + + if (lmsle_mbz) { + GUEST_ASSERT_EQ(vmcb->control.exit_code, SVM_EXIT_ERR); + GUEST_ASSERT(!l2_ran); + } else { + GUEST_ASSERT_EQ(vmcb->control.exit_code, SVM_EXIT_VMMCALL); + GUEST_ASSERT(l2_ran); + GUEST_ASSERT(vmcb->save.efer & EFER_LMSLE); + } + + GUEST_DONE(); +} + +static struct kvm_vcpu *create_l1_vcpu(struct kvm_vm **vm, void *l1_guest_= code) { struct kvm_vcpu *vcpu; - struct kvm_vm *vm; - gva_t nested_gva =3D 0; + gva_t svm_gva; =20 - TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM)); + *vm =3D vm_create_with_one_vcpu(&vcpu, l1_guest_code); =20 - vm =3D vm_create_with_one_vcpu(&vcpu, l1_guest_code); + vcpu_alloc_svm(*vm, &svm_gva); + vcpu_args_set(vcpu, 1, svm_gva); + + return vcpu; +} + +static void test_enumeration(void) +{ + bool lmsle_mbz; =20 - vcpu_alloc_svm(vm, &nested_gva); - vcpu_args_set(vcpu, 1, nested_gva); + /* + * EFER_LMSLE_MBZ, CPUID.80000008H:EBX[bit 20], is a "defeature" bit, + * i.e. is set when the CPU does *not* support long mode segment + * limits. KVM enumerates the defeature if and only if KVM refuses to + * set EFER.LMSLE, i.e. if the CPU doesn't support LMSLE, or if KVM + * doesn't support nested SVM. Derive the expectation from raw CPUID + * and kvm_amd's "nested" module param rather than from + * kvm_cpu_has(X86_FEATURE_SVM), so that the assertion doesn't simply + * compare KVM's enumeration to itself. + * + * Note, kvm_amd's "nested" is an "int" module param, i.e. reads back + * as '1'/'0' and not as 'Y'/'N'. Query the param if and only if the + * CPU supports SVM, i.e. if and only if kvm_amd is the module in + * play. Checking the vendor string wouldn't suffice, e.g. Zhaoxin + * and Centaur CPUs are "GenuineIntel"-adjacent at best, but run VMX + * and thus load kvm_intel. + */ + lmsle_mbz =3D this_cpu_has(X86_FEATURE_EFER_LMSLE_MBZ) || + !this_cpu_has(X86_FEATURE_SVM) || + !get_kvm_amd_param_integer("nested"); + + TEST_ASSERT_EQ(kvm_cpu_has(X86_FEATURE_EFER_LMSLE_MBZ), lmsle_mbz); + + ksft_test_result_pass("KVM enumerates EFER_LMSLE_MBZ=3D%d\n", lmsle_mbz); +} + +static void test_clear_efer_svme(void) +{ + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + + vcpu =3D create_l1_vcpu(&vm, l1_clear_efer_svme); =20 vcpu_run(vcpu); TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_SHUTDOWN); =20 kvm_vm_free(vm); - return 0; + ksft_test_result_pass("L2 clearing EFER.SVME shuts down L1\n"); +} + +static void test_lmsle(bool lmsle_mbz) +{ + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + + vcpu =3D create_l1_vcpu(&vm, l1_lmsle); + + vcpu_set_or_clear_cpuid_feature(vcpu, X86_FEATURE_EFER_LMSLE_MBZ, + lmsle_mbz); + + vcpu_run(vcpu); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO); + + switch (get_ucall(vcpu, &uc)) { + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + case UCALL_DONE: + break; + default: + TEST_FAIL("Unexpected ucall: %lu", uc.cmd); + } + + kvm_vm_free(vm); + ksft_test_result_pass("Guest EFER_LMSLE_MBZ=3D%d\n", lmsle_mbz); +} + +static void test_host_initiated_lmsle(void) +{ + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + u64 efer; + + vm =3D vm_create_with_one_vcpu(&vcpu, NULL); + vcpu_set_cpuid_feature(vcpu, X86_FEATURE_EFER_LMSLE_MBZ); + + /* + * EFER_LMSLE_MBZ is a guest CPUID consistency check, not a host + * capability, i.e. must not be enforced against host-initiated writes, + * so that userspace can set MSRs before it sets guest CPUID. + */ + efer =3D vcpu_get_msr(vcpu, MSR_EFER); + TEST_ASSERT(!(efer & EFER_LMSLE), "EFER.LMSLE unexpectedly set"); + + vcpu_set_msr(vcpu, MSR_EFER, efer | EFER_LMSLE); + TEST_ASSERT_EQ(vcpu_get_msr(vcpu, MSR_EFER), efer | EFER_LMSLE); + + kvm_vm_free(vm); + ksft_test_result_pass("Host-initiated EFER.LMSLE=3D1 is allowed\n"); +} + +static void test_sregs_lmsle(void) +{ + struct kvm_vcpu *vcpu; + struct kvm_sregs sregs; + struct kvm_vm *vm; + int rc; + + vm =3D vm_create_with_one_vcpu(&vcpu, NULL); + vcpu_set_cpuid_feature(vcpu, X86_FEATURE_EFER_LMSLE_MBZ); + + /* + * Unlike KVM_SET_MSRS, KVM_SET_SREGS runs the full set of guest CPUID + * checks, i.e. rejects EFER.LMSLE even though it's host-initiated. + */ + vcpu_sregs_get(vcpu, &sregs); + TEST_ASSERT(!(sregs.efer & EFER_LMSLE), "EFER.LMSLE unexpectedly set"); + + sregs.efer |=3D EFER_LMSLE; + rc =3D _vcpu_sregs_set(vcpu, &sregs); + TEST_ASSERT(rc, "KVM allowed EFER.LMSLE with EFER_LMSLE_MBZ set"); + + kvm_vm_free(vm); + ksft_test_result_pass("KVM_SET_SREGS rejects EFER.LMSLE=3D1\n"); +} + +int main(int argc, char *argv[]) +{ + bool has_nested_svm, has_lmsle; + + ksft_print_header(); + ksft_set_plan(6); + + test_enumeration(); + + /* + * The sub-tests below need to actually run a nested guest, and the + * EFER.LMSLE sub-tests additionally need KVM to allow EFER.LMSLE. + * It's KVM's view of the world, not raw CPUID, that dictates whether + * EFER.LMSLE is allowed, i.e. whether the defeature is emulated. + */ + has_nested_svm =3D kvm_cpu_has(X86_FEATURE_SVM); + has_lmsle =3D has_nested_svm && + !kvm_cpu_has(X86_FEATURE_EFER_LMSLE_MBZ); + + if (!has_nested_svm) + ksft_print_msg("Nested SVM unsupported\n"); + else if (!has_lmsle) + ksft_print_msg("KVM doesn't support EFER.LMSLE\n"); + + if (has_nested_svm) { + test_clear_efer_svme(); + test_lmsle(true); + } else { + ksft_test_result_skip("L2 clearing EFER.SVME shuts down L1\n"); + ksft_test_result_skip("Guest EFER_LMSLE_MBZ=3D1\n"); + } + + if (has_lmsle) { + test_lmsle(false); + test_host_initiated_lmsle(); + test_sregs_lmsle(); + } else { + ksft_test_result_skip("Guest EFER_LMSLE_MBZ=3D0\n"); + ksft_test_result_skip("Host-initiated EFER.LMSLE=3D1 is allowed\n"); + ksft_test_result_skip("KVM_SET_SREGS rejects EFER.LMSLE=3D1\n"); + } + + ksft_finished(); } --=20 2.56.0.rc1.315.gc6ed9934b7-goog