From nobody Fri Sep 25 03:16:09 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B647D46D08C; Thu, 17 Sep 2026 07:21:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629695; cv=none; b=RS7552D5ymOCJHl21+3ISNB1ysuUWD7Z89F+h3X9OpgOpTc0G1fHwpQqqdauwozeHY5+Sj0LdwiWd6FDbUEVrv4xAq840zbQ96IoYpP78bMXGHlq3tBEuSCSxnazg+GmNAN5YjwFG74hkuWwlBXucnkQ70m/xwX8+/NYganQyMQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629695; c=relaxed/simple; bh=6X9vWFAJYhXEM3/Oimput5+0fsO5eVn90vKz882z0kM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lOeNlSw9LgErULwXFhGWdN3FXAv08iQEBKxyDb1GWjKKDhVZrmIfkhMAhhyNaL3SbIvQd4+WeIGzuu/uBzPMMsFXvHnJtAiejn5F2qIsfVSOA73ch57LpMQiQTD0BX4kZGnxE6+YkPc2ybfgH42rNwWVCdliQFZFkpTNg4mcWHw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=CATVygy2; arc=none smtp.client-ip=192.198.163.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="CATVygy2" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1789629691; x=1821165691; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=6X9vWFAJYhXEM3/Oimput5+0fsO5eVn90vKz882z0kM=; b=CATVygy2GjO2se4efCyesK7gg05JyjABZLInKTj/VN39HC6XJhWfQp7r baK9UZJ+69+dZSI6iAq/CICzZ37/0WM/DfOCJdLzzje+CMFCrOgFug+DH jxD4EWdRmjq3PcZ2wa+8Qmjj2DCtmC1Mxv6ApantMmfKh02m1+AbRZ0vM rmfApF8DvX83r8ZgSE/TY1lCewPWGLu2gflhQKRuoeKSqXpAx08vbcblQ y0JmLpSwr2PNBXUVZcuBMPqBXWeKekKskWfR8BvfJh22OQbvgR5kQ9Reg cwc2ZPlvGAD29piv6HaC30cuwQF6oq+mYNR3wrkyNKdvNgX2ypi+n/gRl g==; X-CSE-ConnectionGUID: fkP1ibI4TMqxbUVa0o/9vQ== X-CSE-MsgGUID: sUEVf9DhQ/W02hYOZjptpg== X-IronPort-AV: E=McAfee;i="6800,10657,11905"; a="531838" X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="531838" Received: from fmviesa012.fm.intel.com ([10.60.135.152]) by fmvoesa114.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:30 -0700 X-CSE-ConnectionGUID: iFcbTfoUToKAGGoiRP7Jiw== X-CSE-MsgGUID: 0N6tVA+FQ5++Tn3tQdKRdw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="1849257" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by smtpauth.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:27 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, tony.lindgren@linux.intel.com, kishen.maloor@intel.com, dedekind1@gmail.com, binbin.wu@linux.intel.com Subject: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Date: Thu, 17 Sep 2026 15:25:45 +0800 Message-ID: <20260917072548.2314491-2-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260917072548.2314491-1-binbin.wu@linux.intel.com> References: <20260917072548.2314491-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable CPUID feature bits that KVM supports, and build the masks during TDX hardware setup via tdx_initialize_cpu_cfg_caps(). The TDX module reports the CPUID bits that the VMM can directly configure for a TD, but KVM cannot blindly expose all reported bits to userspace. Certain features imply additional architectural state, e.g. one or more MSRs, that KVM must explicitly manage across host/guest transitions to prevent host state corruption. The existing hardcoded denylist cannot account for new host state clobbering features introduced by future TDX modules. Except for a few fixed-1 bits required for basic TDX support, host state clobbering features are either directly configurable or gated by TD ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly configurable bits to build an allowlist is therefore sufficient. Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can be built with similar feature-name based initializers. Directly configurable non-feature bits will be handled separately. The allowlist is prepared to be consumed by later patches to filter KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID feature bits stay hidden from userspace until KVM explicitly opts in. By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F(). Requiring support for non-TDX VMs avoids committing to TDX-specific behavior before general KVM support is established, and respects KVM's logic around disabling certain features, since the reasons for disabling them could apply to TDX as well. Allow exceptions through TDX_CFG_EXTRA_F() only with sufficient justification. Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM doesn't support for TDX yet. Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are not advertised in kvm_cpu_caps[]. The remaining directly configurable feature bits outside kvm_cpu_caps[] are left out of the allowlist: - Features forced to zero when #VE is reduced, or lacking KVM support for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A, RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES in a subsequent patch. - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE. - Features that can clobber host state and lack KVM support for TDX: FRED. - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel platform), and RAO_INT (defined only for future processors). Filtering KVM_TDX_CAPABILITIES in a subsequent patch will intentionally stop advertising the excluded bits as configurable, as the corresponding features are unsupported or cannot be properly virtualized. Signed-off-by: Binbin Wu --- v4: - Add XTPR and HT to the allowlist. - Explain in the changelog why some directly configurable bits are not added to the allowlist. - Improve comments/changelog. (Xiaoyao) - Fix the bug that AMX_COMPLEX should be added for CPUID_1E_1_EAX instead of CPUID_7_1_EDX. v3: - Drop the new data structure in v2 and only track feature bits by following the organization of kvm_cpu_caps[], handle non-feature bits separately. (Sean) - Use two versions of macros (TDX_CFG_F() vs. TDX_CFG_EXTRA_F()) to distinguish whether a supported TDX configurable CPUID bit should be checked against KVM's common cpu capabilities. - Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc. --- arch/x86/kvm/vmx/tdx.c | 164 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 164 insertions(+) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index 8a3f88b79b83d..2b51a85c998e8 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -52,6 +52,168 @@ __TDX_BUG_ON(__err, #__fn, __kvm, ", " #a1 " 0x%llx, " #a2 ", 0x%llx, " #= a3 " 0x%llx", \ a1, a2, a3) =20 +static u32 tdx_cpu_cfg_caps[NR_KVM_CPU_CAPS] __ro_after_init; +static_assert(ARRAY_SIZE(tdx_cpu_cfg_caps) =3D=3D ARRAY_SIZE(kvm_cpu_caps)= ); + +#define TDX_VALIDATE_CPU_CAP_USAGE(name) \ + BUILD_BUG_ON(__feature_leaf(X86_FEATURE_##name) !=3D \ + tdx_cpu_cap_init_in_progress) + +/* For a feature bit that needs to be cap'ed by kvm_cpu_caps[]. */ +#define TDX_CFG_F(name) \ +({ \ + TDX_VALIDATE_CPU_CAP_USAGE(name); \ + tdx_cfg_caps |=3D feature_bit(name); \ +}) + +/* + * For a feature bit that KVM allows for TDX guests even though it isn't + * advertised through kvm_cpu_caps[], e.g. MWAIT. Use this version only w= hen + * there is a justification. + */ +#define TDX_CFG_EXTRA_F(name) \ +({ \ + TDX_VALIDATE_CPU_CAP_USAGE(name); \ + tdx_cfg_caps_extra |=3D feature_bit(name); \ +}) + +#define tdx_cpu_cfg_cap_init(leaf, feature_initializers...) \ +do { \ + const u32 __maybe_unused tdx_cpu_cap_init_in_progress =3D leaf; \ + u32 tdx_cfg_caps_extra =3D 0; \ + u32 tdx_cfg_caps =3D 0; \ + \ + feature_initializers \ + tdx_cpu_cfg_caps[leaf] =3D (tdx_cfg_caps & kvm_cpu_caps[leaf]) | \ + tdx_cfg_caps_extra; \ +} while (0) + +/* + * Initialize tdx_cpu_cfg_caps[], the list of CPUID features that KVM + * supports for TDX guests. It covers only the directly configurable CPUID + * bits reported by the TDX module; features controlled by XFAM and + * ATTRIBUTES are handled separately. + */ +static void __init tdx_initialize_cpu_cfg_caps(void) +{ + tdx_cpu_cfg_cap_init(CPUID_1_ECX, + /* + * KVM allows userspace to enumerate MONITOR+MWAIT support to + * the guest, but the MWAIT feature flag is never advertised + * to userspace for non-TDX VMs. + */ + TDX_CFG_EXTRA_F(MWAIT), + /* + * XTPR can be exposed to a TD, but it never takes effect in + * the underlying hardware when the TD changes + * IA32_MISC_ENABLE[23]. + */ + TDX_CFG_EXTRA_F(XTPR), + TDX_CFG_F(TSC_DEADLINE_TIMER), + TDX_CFG_F(AVX), + TDX_CFG_F(F16C), + ); + + tdx_cpu_cfg_cap_init(CPUID_1_EDX, + TDX_CFG_F(MCE), + TDX_CFG_F(MTRR), + TDX_CFG_F(MCA), + TDX_CFG_F(SELFSNOOP), + /* + * HT is a topology enumeration bit that KVM doesn't care + * about, but userspace may want to expose it to the guests. + */ + TDX_CFG_EXTRA_F(HT), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_0_EBX, + TDX_CFG_F(BMI1), + /* HLE */ + TDX_CFG_F(BMI2), + TDX_CFG_F(ERMS), + /* RTM */ + TDX_CFG_F(AVX512F), + TDX_CFG_F(AVX512DQ), + TDX_CFG_F(ADX), + TDX_CFG_F(AVX512IFMA), + TDX_CFG_F(AVX512PF), + TDX_CFG_F(AVX512ER), + TDX_CFG_F(AVX512CD), + TDX_CFG_F(AVX512BW), + TDX_CFG_F(AVX512VL), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_ECX, + TDX_CFG_F(UMIP), + /* WAITPKG */ + TDX_CFG_F(AVX512_VBMI2), + TDX_CFG_F(GFNI), + TDX_CFG_F(VAES), + TDX_CFG_F(VPCLMULQDQ), + TDX_CFG_F(AVX512_VNNI), + TDX_CFG_F(AVX512_BITALG), + TDX_CFG_F(AVX512_VPOPCNTDQ), + TDX_CFG_F(LA57), + TDX_CFG_F(RDPID), + TDX_CFG_F(CLDEMOTE), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_EDX, + TDX_CFG_F(AVX512_4VNNIW), + TDX_CFG_F(AVX512_4FMAPS), + TDX_CFG_F(FSRM), + TDX_CFG_F(AVX512_VP2INTERSECT), + TDX_CFG_F(SERIALIZE), + TDX_CFG_F(TSXLDTRK), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_1_EAX, + TDX_CFG_F(SHA512), + TDX_CFG_F(SM3), + TDX_CFG_F(SM4), + TDX_CFG_F(AVX_VNNI), + TDX_CFG_F(AVX512_BF16), + TDX_CFG_F(CMPCCXADD), + TDX_CFG_F(FZRM), + TDX_CFG_F(FSRS), + TDX_CFG_F(FSRC), + /* FRED */ + TDX_CFG_F(LKGS), + TDX_CFG_F(WRMSRNS), + TDX_CFG_F(AMX_FP16), + TDX_CFG_F(AVX_IFMA), + TDX_CFG_F(LAM), + TDX_CFG_F(MOVRS), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_1_EDX, + TDX_CFG_F(AVX_VNNI_INT8), + TDX_CFG_F(AVX_NE_CONVERT), + TDX_CFG_F(AVX_VNNI_INT16), + TDX_CFG_F(PREFETCHITI), + TDX_CFG_F(AVX10), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_2_EDX, + TDX_CFG_F(DDPD_U), + TDX_CFG_F(MCDT_NO), + ); + + tdx_cpu_cfg_cap_init(CPUID_1E_1_EAX, + TDX_CFG_F(AMX_COMPLEX_ALIAS), + TDX_CFG_F(AMX_FP8), + TDX_CFG_F(AMX_TF32), + TDX_CFG_F(AMX_AVX512), + TDX_CFG_F(AMX_MOVRS), + ); + + tdx_cpu_cfg_cap_init(CPUID_8000_0008_EBX, + TDX_CFG_F(WBNOINVD), + ); +} + +#undef TDX_CFG_F +#undef TDX_CFG_EXTRA_F =20 bool enable_tdx __ro_after_init; module_param_named(tdx, enable_tdx, bool, 0444); @@ -3487,6 +3649,8 @@ int __init tdx_hardware_setup(void) return r; } =20 + tdx_initialize_cpu_cfg_caps(); + KVM_SANITY_CHECK_VM_STRUCT_SIZE(kvm_tdx); =20 vt_x86_ops.vm_size =3D max_t(unsigned int, vt_x86_ops.vm_size, sizeof(str= uct kvm_tdx)); --=20 2.46.0 From nobody Fri Sep 25 03:16:09 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 152B147604B; Thu, 17 Sep 2026 07:21:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629702; cv=none; b=ai5QawahQD/Faa5OI7hJeob6F4rKWXQtBb3LsNltRrfaeNZFBQOQHSymDWuhufB6r/oblPxpn0QStbEug5diCz9DC1hBCZm9sahZWlPEd1/GUcsfEa8VsdihTOjs5cy3iMkOaoCwcC1FjYNidd85+dGuE/QNal2zGNaUDV/nees= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629702; c=relaxed/simple; bh=oMsXHD3gcjjkL+qdh1k2ncXZr3LSlbqkqd60iPb3Py0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XzcAVnt2jQmekhoXPsELO3k3hWp0YRkB8TD86RDuyXL6+opNZdXAblxN+YLPXehayq9rCyqkCDxyazVaTvQfp+xitAkuMZJHtq9p/FY4TkbGe95trh94yDbna7r4im31FW/boSQYyPbvARFbP3r/O26Ub5CiwmThZokw8tfC9Vk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=PFwJkJiL; arc=none smtp.client-ip=192.198.163.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="PFwJkJiL" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1789629697; x=1821165697; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=oMsXHD3gcjjkL+qdh1k2ncXZr3LSlbqkqd60iPb3Py0=; b=PFwJkJiL05GqvPV9n4xRXB+JtsvcbFw+/2P9cteFFTsh6n4kfiNIMJH7 irWPIex8id8Fxcurd6vXsulTosZbntbuyjgozZLetIxy45k9kscGzKN1n 9Di6gxag6wU1ghQe0bo6bbIVPBa85N73p0ql6ntDmHdYbOCPU9p2JwxRw juakaIeaVtmAPzW2Qeo1/LTjJeonZoWUJaPNVewCnO8/K3Epdpdjs75Eg G5GVSNT5wSODK8Ah55LxI2yaLWg1xTaYi74JetzciKrm0R1H8mqGM03/X Vh4IJ+N4RDR7QYCsWGZ0f38kJ7cNMGmHRllcTeArhp9USo9iHEfxynbhN Q==; X-CSE-ConnectionGUID: Z7GARTPaRHSUQcofObSoIA== X-CSE-MsgGUID: vhGIV64pTxOLcMVc2OqbTw== X-IronPort-AV: E=McAfee;i="6800,10657,11905"; a="531856" X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="531856" Received: from fmviesa012.fm.intel.com ([10.60.135.152]) by fmvoesa114.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:33 -0700 X-CSE-ConnectionGUID: T5f5e0VtT6GNoCT2BC37zw== X-CSE-MsgGUID: 04p8HGb4Soa918/thMAeJg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="1849327" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by smtpauth.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:31 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, tony.lindgren@linux.intel.com, kishen.maloor@intel.com, dedekind1@gmail.com, binbin.wu@linux.intel.com Subject: [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Date: Thu, 17 Sep 2026 15:25:46 +0800 Message-ID: <20260917072548.2314491-3-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260917072548.2314491-1-binbin.wu@linux.intel.com> References: <20260917072548.2314491-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX directly configurable CPUID feature bits, even though KVM doesn't support MSR_IA32_CORE_CAPS for TDX guests, to accommodate the legacy TDX module definition and userspace's stale knowledge of it. Older TDX specifications define the CORE_CAPABILITIES CPUID bit as fixed-1, so userspace may expect the bit to be enabled for TDs. #VE reduction turns it into a directly configurable bit, so leaving it out of the allowlist would make the bit impossible to enable once KVM starts validating userspace's CPUID input, i.e. would be a surprising behavior change for such userspace. Reporting CORE_CAPABILITIES as directly configurable also lets userspace detect that the bit is no longer fixed-1, and thus correct its stale knowledge. Keep MSR_IA32_CORE_CAPS unsupported for TDs, as no existing TDX user needs guest access to the MSR. Note, CORE_CAPABILITIES is the only bit that is unsupported by KVM *and* changed from fixed-1 to directly configurable by #VE reduction, and no further #VE reductions are expected. Signed-off-by: Binbin Wu Reviewed-by: Tony Lindgren Reviewed-by: Xiaoyao Li --- v4: - Add #VE reduction related background to the changelog. (Kishen) - Add RB from Tony. v3: - Drop the code for MSR_IA32_CORE_CAPS access. --- arch/x86/kvm/vmx/tdx.c | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index 2b51a85c998e8..b34afc52b714e 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -165,6 +165,13 @@ static void __init tdx_initialize_cpu_cfg_caps(void) TDX_CFG_F(AVX512_VP2INTERSECT), TDX_CFG_F(SERIALIZE), TDX_CFG_F(TSXLDTRK), + /* + * KVM doesn't support MSR_IA32_CORE_CAPS, but older TDX specs + * define this bit as fixed-1. Report it as configurable to + * accommodate the legacy TDX module definition, and to let + * userspace detect that the bit is no longer fixed-1. + */ + TDX_CFG_EXTRA_F(CORE_CAPABILITIES), ); =20 tdx_cpu_cfg_cap_init(CPUID_7_1_EAX, --=20 2.46.0 From nobody Fri Sep 25 03:16:09 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9463847CC96; Thu, 17 Sep 2026 07:21:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629721; cv=none; b=PeKLKMkdeYXwUy/VwC2vTc5OesJmpevdxZkOUxA69ZyZI7lBSYZgpI7nTidlHxJONHTMYvypY4jsfbngvA8Jgc92MMaAL2s3D4kUSFZwAftsWLdUq4nmTpzNsHXnepyzLeu38zagPAdB5nbKBc/AxPWN8tKUkO2r5QT7fUJs2sI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629721; c=relaxed/simple; bh=RFo0GcFB8WvcT8YgsNJNcG8vvW7Ir69VueovFbcFSI4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UVDVpfZUe2DsYJ6KFBlsoXTPvoJ7yOQIeAN5M20qh4qGv1mwC0j2JKQQ9SeE6hbmYLD+ztfDe2EXL1Cg/rrcPXquuHYgpggpWGwo8CSth4Opx84V0T2XflgfRNQoS21NuWhSu/6htT+hwV6xw3aNYsmvKEN1B8nVO1Zug0nqY+c= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=YC/ATe5u; arc=none smtp.client-ip=192.198.163.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="YC/ATe5u" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1789629714; x=1821165714; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=RFo0GcFB8WvcT8YgsNJNcG8vvW7Ir69VueovFbcFSI4=; b=YC/ATe5uqW5mn029bDTpCkMW+ssVUSP3HcQsFg2pj30AU34FGM3+HG4P wBziJLOm/e+EYQO3FT5x8wQtCm3drZ1YHz6mDln6HLWbyiQbrKHTFGel3 EXOFvxQ0z7rz+QYqVEB9Cy92si4CYYnM4gSs74cCB8fG5CgLUBKlelicB M4uQSGfvRXFoCfbfC+r+iPqEB+Bitln2eU1feSXTbBWKgxlAHqS41q73k 4LRYfdGjU4fTGzFHoNBPdHaQfZutZGul4Nb9cIW7x2N8da86+PP0cCCZt mEhnskB3YXNteKdY2Xyl+OcExSExCpUG4tjd97W/DG+73/cMwHubj1shO w==; X-CSE-ConnectionGUID: jx44EVZmRiWVTwQEVusZwA== X-CSE-MsgGUID: B3SKfBBaRv2H2X++iMdc1Q== X-IronPort-AV: E=McAfee;i="6800,10657,11905"; a="531866" X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="531866" Received: from fmviesa012.fm.intel.com ([10.60.135.152]) by fmvoesa114.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:37 -0700 X-CSE-ConnectionGUID: xyBy01OFQV+qnAfd3Gw78w== X-CSE-MsgGUID: ZMkc8r9tS5uRlyEGO8j9Mg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="1849403" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by smtpauth.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:35 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, tony.lindgren@linux.intel.com, kishen.maloor@intel.com, dedekind1@gmail.com, binbin.wu@linux.intel.com Subject: [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Date: Thu, 17 Sep 2026 15:25:47 +0800 Message-ID: <20260917072548.2314491-4-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260917072548.2314491-1-binbin.wu@linux.intel.com> References: <20260917072548.2314491-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Filter the directly configurable CPUID bits reported through KVM_TDX_CAPABILITIES against KVM's TDX allowlist, and drop the hardcoded denylist based filtering. The TDX module reports directly configurable CPUID bits that it supports for a TD, but KVM must not expose bits that it doesn't support, as blindly exposing a host state clobbering feature can lead to host state corruption. The existing denylist, which clears only TSX and WAITPKG, is not fail-safe. Add tdx_get_cpuid_cfg_mask() to get the mask of directly configurable bits allowed by KVM for a given CPUID register, covering both feature bits, which come from tdx_cpu_cfg_caps[], and non-feature bits, which are enumerated at runtime. Apply the mask to every CPUID entry reported through KVM_TDX_CAPABILITIES. With the allowlist in place, newly introduced TDX directly configurable CPUID bits stay hidden from userspace until KVM explicitly opts in. Note, filtering KVM_TDX_CAPABILITIES intentionally stops advertising the directly configurable bits that aren't in the allowlist, as the corresponding features are unsupported or cannot be properly virtualized. Update the documentation to reflect the ABI change. Signed-off-by: Binbin Wu --- v4: - Rename functions. (Tony) tdx_cfg_non_feature_mask() -> tdx_get_cpuid_cfg_non_feature_mask() tdx_cfg_feature_mask() -> tdx_get_cpuid_cfg_feature_mask() tdx_get_allowed_cfg_cpuid_mask() -> tdx_get_cpuid_cfg_mask() - Update documentation. (Xiaoyao, Rick) v3: - Handle non-feature leafs at runtime. (Sean) - Add CPUID.0x24.0.EBX[7:0] into allowlist. There is a mismatch of the description about CPUID.0x24.0.EBX[7:0], which is listed as "XFAM & CPUID_Enabled & Native" but should be "XFAM & CPUID_Enabled & Configured & Native". --- Documentation/virt/kvm/x86/intel-tdx.rst | 8 ++ arch/x86/kvm/vmx/tdx.c | 95 ++++++++++++++++++------ 2 files changed, 80 insertions(+), 23 deletions(-) diff --git a/Documentation/virt/kvm/x86/intel-tdx.rst b/Documentation/virt/= kvm/x86/intel-tdx.rst index 6a222e9d09541..6beeb89d7f057 100644 --- a/Documentation/virt/kvm/x86/intel-tdx.rst +++ b/Documentation/virt/kvm/x86/intel-tdx.rst @@ -69,6 +69,14 @@ Return the TDX capabilities that current KVM supports wi= th the specific TDX module loaded in the system. It reports what features/capabilities are al= lowed to be configured to the TDX guest. =20 +Note, a CPUID feature bit that the TDX module reports as directly configur= able +is not necessarily reported as configurable by KVM. KVM omits features it +doesn't support. Generally speaking, KVM reports a feature as configurable +only if KVM supports the feature for both TDX and non-TDX VMs, though ther= e are +a handful of exceptions where KVM allows a feature for TDX guests that it = never +advertises for non-TDX VMs. Userspace must rely on KVM_TDX_CAPABILITIES to +determine what can be passed in cpuid to KVM_TDX_INIT_VM. + - id: KVM_TDX_CAPABILITIES - flags: must be 0 - data: pointer to struct kvm_tdx_capabilities diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index b34afc52b714e..ad7d70b36ffe5 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -294,36 +294,79 @@ static bool has_tsx(const struct kvm_cpuid_entry2 *en= try) (entry->ebx & TDX_FEATURE_TSX); } =20 -static void clear_tsx(struct kvm_cpuid_entry2 *entry) -{ - entry->ebx &=3D ~TDX_FEATURE_TSX; -} - static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry) { return entry->function =3D=3D 7 && entry->index =3D=3D 0 && (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG)); } =20 -static void clear_waitpkg(struct kvm_cpuid_entry2 *entry) -{ - entry->ecx &=3D ~__feature_bit(X86_FEATURE_WAITPKG); -} - -static void tdx_clear_unsupported_cpuid(struct kvm_cpuid_entry2 *entry) -{ - if (has_tsx(entry)) - clear_tsx(entry); - - if (has_waitpkg(entry)) - clear_waitpkg(entry); -} - static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry) { return has_tsx(entry) || has_waitpkg(entry); } =20 +#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0) + +static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int= reg) +{ + /* + * For a leaf/subleaf/register that will never be repurposed to hold + * feature bits, it's safe to return TDX_CPUID_ALL_ALLOWED_MASK, i.e. + * leave the TDX module's CPUID config mask intact. + */ + switch (function) { + case 1: + if (reg =3D=3D CPUID_EAX || reg =3D=3D CPUID_EBX) + return TDX_CPUID_ALL_ALLOWED_MASK; + return 0; + case 4: + case 0x18: + case 0x1f: + return TDX_CPUID_ALL_ALLOWED_MASK; + case 0x24: + if (index =3D=3D 0 && reg =3D=3D CPUID_EBX) + return GENMASK_U32(7, 0); + return 0; + case 0x80000008: + if (reg =3D=3D CPUID_EAX) + return TDX_CPUID_ALL_ALLOWED_MASK; + return 0; + default: + return 0; + } +} + +static u32 tdx_get_cpuid_cfg_feature_mask(u32 function, u32 index, int reg) +{ + for (int i =3D 0; i < NR_KVM_CPU_CAPS; i++) { + const struct cpuid_reg *cpuid =3D &reverse_cpuid[i]; + + if (!cpuid->function) + continue; + + if (cpuid->function =3D=3D function && cpuid->index =3D=3D index && + cpuid->reg =3D=3D reg) + return tdx_cpu_cfg_caps[i]; + } + + return 0; +} + +static u32 tdx_get_cpuid_cfg_mask(u32 function, u32 index, int reg) +{ + u32 non_feature_mask =3D tdx_get_cpuid_cfg_non_feature_mask(function, ind= ex, reg); + + /* Skip reverse_cpuid[] walk if it's already all allowed. */ + if (non_feature_mask =3D=3D TDX_CPUID_ALL_ALLOWED_MASK) + return TDX_CPUID_ALL_ALLOWED_MASK; + + /* + * It's possible that a CPUID register contains both feature and + * non-feature bits. + */ + return non_feature_mask | tdx_get_cpuid_cfg_feature_mask(function, index,= reg); +} + #define KVM_TDX_CPUID_NO_SUBLEAF ((__u32)-1) =20 static void td_init_cpuid_entry2(struct kvm_cpuid_entry2 *entry, unsigned = char idx) @@ -347,8 +390,6 @@ static void td_init_cpuid_entry2(struct kvm_cpuid_entry= 2 *entry, unsigned char i */ if (entry->function =3D=3D 0x80000008) entry->eax =3D tdx_set_guest_phys_addr_bits(entry->eax, 0xff); - - tdx_clear_unsupported_cpuid(entry); } =20 #define TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT BIT(1) @@ -371,8 +412,16 @@ static int init_kvm_tdx_caps(const struct tdx_sys_info= _td_conf *td_conf, caps->user_tdvmcallinfo_1_r11 =3D TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT; =20 - for (i =3D 0; i < td_conf->num_cpuid_config; i++) - td_init_cpuid_entry2(&caps->cpuid.entries[i], i); + for (i =3D 0; i < td_conf->num_cpuid_config; i++) { + struct kvm_cpuid_entry2 *e =3D &caps->cpuid.entries[i]; + + td_init_cpuid_entry2(e, i); + /* Only report the configurable bits allowed by KVM. */ + e->eax &=3D tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EAX); + e->ebx &=3D tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EBX); + e->ecx &=3D tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_ECX); + e->edx &=3D tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EDX); + } =20 return 0; } --=20 2.46.0 From nobody Fri Sep 25 03:16:09 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1D6C847ECDD; Thu, 17 Sep 2026 07:22:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629728; cv=none; b=F77otYUQLhndPcL7elQOKGm2d0w6RPMk13Q/5gBQSb3RecrRnZhnBLiT05u5daXuvv48aHljUw9sdNVQCy2gHgzJoJ2YXs0l/82APC30XqHeh8y5R7GPGbcxtvMJb+hGTxaXDFaRT/4/rK7mCTnuf/7RlJ/eFczdVSvDnns7HdU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789629728; c=relaxed/simple; bh=Bcx8AYHB1t/Ot/Sz8VFwmw35xbBw+Tf0ELwyt0yh56s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=APMlbYwQU5JvN9xCaaWDzAW9KMMzKXcPBhXzSTIP/mmddcTukyKN9KkU9XKOF3RTj65Fe8cYE4UaDhQnOxE6m28ZR07O9WpCJstUN6PaxDdwAq9Z0p4Yh4z95FBTO7lhv1ULwMz2RRsx4UOBKaPwgnoTZjf70qX265AL81qVMVY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=lT15/n2h; arc=none smtp.client-ip=192.198.163.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="lT15/n2h" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1789629722; x=1821165722; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=Bcx8AYHB1t/Ot/Sz8VFwmw35xbBw+Tf0ELwyt0yh56s=; b=lT15/n2hQFbgv1PpJr5aNg54SYhLk8TF62WFlw9UXHXd1AM378FnML0M QOeSj4Xg+SXj4XmUPNZRsgC9UAzhkd3ziMc1Qd2yolxWQbeo3mZI1ZUdo 9Ebouv2TTAIHRxG8B0x6nuWGoBCl6J1APdpQeJs50zamPUaBwkR9awAi0 6Nup2FHdkVE6vVmAp5hMUCxZEAdKCFSfe2q5M2zdihpkVaqN5qq9x3Udm Kq0eSYuwd3g4Lo3CmaHJoWCyqkBlm9jFehb/9hVGFSuIBYqFjBzHmP2Ws QeAU1TWt79nWFrDO//1zDVLAoBKojUeSlcmv+5W0hUlsPQer24Kcrs6Mu w==; X-CSE-ConnectionGUID: iIlwYpFARL2YLCk7dav3MQ== X-CSE-MsgGUID: Nal+Z+kxQtiLnFm+RxpIag== X-IronPort-AV: E=McAfee;i="6800,10657,11905"; a="531876" X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="531876" Received: from fmviesa012.fm.intel.com ([10.60.135.152]) by fmvoesa114.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:40 -0700 X-CSE-ConnectionGUID: 69pfWXumQeSHTVAE1YIgeA== X-CSE-MsgGUID: hsYhOe2OQRWXhmEbFdu8aQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,103,1787036400"; d="scan'208";a="1849453" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by smtpauth.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 Sep 2026 00:21:38 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, tony.lindgren@linux.intel.com, kishen.maloor@intel.com, dedekind1@gmail.com, binbin.wu@linux.intel.com Subject: [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Date: Thu, 17 Sep 2026 15:25:48 +0800 Message-ID: <20260917072548.2314491-5-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260917072548.2314491-1-binbin.wu@linux.intel.com> References: <20260917072548.2314491-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Validate the CPUID configuration provided by userspace through KVM_TDX_INIT_VM against KVM's TDX allowlist, and drop the hardcoded denylist based check. The TDX module lets the VMM configure certain CPUID features for a TD at initialization time, but KVM must strictly govern which of them userspace can actually enable, otherwise a host state clobbering feature could be enabled behind KVM's back. The existing check only rejects TSX and WAITPKG, i.e. it is not fail-safe, as any bit that a future TDX module makes configurable would be accepted even if KVM has no idea about the feature. Add tdx_has_unsupported_cpuid_cfg_bit() and reject KVM_TDX_INIT_VM if userspace sets any bit outside the mask returned by tdx_get_cpuid_cfg_mask(). There is no need to first mask the userspace input with the bits the TDX module reports as directly configurable, as anything outside that set is rejected by the TDX module itself. Also reject CPUID entries whose index differs from the value expected by the TDX module, as kvm_find_cpuid_entry2() ignores the index when KVM_CPUID_FLAG_SIGNIFCANT_INDEX is cleared, i.e. a mismatching entry could otherwise be applied to the wrong subleaf. Update the comments for KVM_TDX_INIT_VM in the uapi header and the TDX documentation accordingly. Signed-off-by: Binbin Wu Reviewed-by: Tony Lindgren --- v4: - Update the comments/documentation. - Collect RB tag from Tony. v3: - Check CPUID entry index mismatch b/t userspace input and TDX sysinfo configuration. (Sashiko) - No need to mask the userspace input with TDX module reported directly configurable bits first, since the userspace input should be subset of directly configurable bits. Otherwise, the input will be rejected by the TDX module. --- Documentation/virt/kvm/x86/intel-tdx.rst | 2 ++ arch/x86/include/uapi/asm/kvm.h | 2 ++ arch/x86/kvm/vmx/tdx.c | 41 ++++++++++++------------ 3 files changed, 25 insertions(+), 20 deletions(-) diff --git a/Documentation/virt/kvm/x86/intel-tdx.rst b/Documentation/virt/= kvm/x86/intel-tdx.rst index 6beeb89d7f057..36bd0fee2480d 100644 --- a/Documentation/virt/kvm/x86/intel-tdx.rst +++ b/Documentation/virt/kvm/x86/intel-tdx.rst @@ -135,6 +135,8 @@ KVM_CREATE_VM and before creating any VCPUs. /* * Call KVM_TDX_INIT_VM before vcpu creation, thus before * KVM_SET_CPUID2. + * KVM validates @cpuid, i.e. setting a bit that KVM_TDX_CAPABILIT= IES + * doesn't report as configurable fails the ioctl. * This configuration supersedes KVM_SET_CPUID2s for VCPUs because= the * TDX module directly virtualizes those CPUIDs without VMM. The = user * space VMM, e.g. qemu, should make KVM_SET_CPUID2 consistent with diff --git a/arch/x86/include/uapi/asm/kvm.h b/arch/x86/include/uapi/asm/kv= m.h index 1585ec8040666..784cb5cd7f664 100644 --- a/arch/x86/include/uapi/asm/kvm.h +++ b/arch/x86/include/uapi/asm/kvm.h @@ -1032,6 +1032,8 @@ struct kvm_tdx_init_vm { /* * Call KVM_TDX_INIT_VM before vcpu creation, thus before * KVM_SET_CPUID2. + * KVM validates @cpuid, i.e. setting a bit that KVM_TDX_CAPABILITIES + * doesn't report as configurable fails the ioctl. * This configuration supersedes KVM_SET_CPUID2s for VCPUs because the * TDX module directly virtualizes those CPUIDs without VMM. The user * space VMM, e.g. qemu, should make KVM_SET_CPUID2 consistent with diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index ad7d70b36ffe5..5426191905d0d 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -286,25 +286,6 @@ static u32 tdx_set_guest_phys_addr_bits(const u32 eax,= int addr_bits) return (eax & ~GENMASK(23, 16)) | (addr_bits & 0xff) << 16; } =20 -#define TDX_FEATURE_TSX (__feature_bit(X86_FEATURE_HLE) | __feature_bit(X8= 6_FEATURE_RTM)) - -static bool has_tsx(const struct kvm_cpuid_entry2 *entry) -{ - return entry->function =3D=3D 7 && entry->index =3D=3D 0 && - (entry->ebx & TDX_FEATURE_TSX); -} - -static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry) -{ - return entry->function =3D=3D 7 && entry->index =3D=3D 0 && - (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG)); -} - -static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry) -{ - return has_tsx(entry) || has_waitpkg(entry); -} - #define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0) =20 static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int= reg) @@ -2548,6 +2529,17 @@ static int setup_tdparams_eptp_controls(struct kvm_c= puid2 *cpuid, return 0; } =20 +static bool tdx_has_unsupported_cpuid_cfg_bit(const struct kvm_cpuid_entry= 2 *entry) +{ + u32 function =3D entry->function; + u32 index =3D entry->index; + + return (entry->eax & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EAX))= || + (entry->ebx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EBX))= || + (entry->ecx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_ECX))= || + (entry->edx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EDX)); +} + static int setup_tdparams_cpuids(struct kvm_cpuid2 *cpuid, struct td_params *td_params) { @@ -2571,7 +2563,16 @@ static int setup_tdparams_cpuids(struct kvm_cpuid2 *= cpuid, if (!entry) continue; =20 - if (tdx_unsupported_cpuid(entry)) + /* + * Reject entries whose index doesn't match the expected one. + * This catches userspace passing a CPUID entry with the + * KVM_CPUID_FLAG_SIGNIFCANT_INDEX flag cleared when the index + * is significant. + */ + if (entry->index !=3D tmp.index) + return -EINVAL; + + if (tdx_has_unsupported_cpuid_cfg_bit(entry)) return -EINVAL; =20 copy_cnt++; --=20 2.46.0