From nobody Mon Sep 28 03:41:40 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 28CC9385529; Thu, 27 Aug 2026 03:13:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800441; cv=none; b=IRIovqJakaB2C/SApARNaTfeFhUOaMb3dcX9rVXwTI3Rkpf9qJriEb2aHxYSfb1y52UwGbNSsSBkONxmL7AXOxY0ErOH0lpndn42skxKgHwF2JxJ5cL+NUyuvfm41GUdCfeva064OCHzCvRR99NhBSvrdH4h6gM4y/hc2A28i6Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800441; c=relaxed/simple; bh=PgeX2FlPPmtq8ItOCQvA8Gr2xrRj7YJ4vq9H3xkQ5y8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Y5vrZ/GAwBWA4XonuhbuZGQdlwrqiWSEFwCEZep7T4pwKPUl/NJHDkOOYERB6rgcz8VODjzeLgGX9mlsC/jN16ynl8YAykaoURjOT9jVK4wuDBwm1CJ2FYrasxGshfT4aQvUNQSwRdGCVFHa7PtDV7syYDTkTM/21nC0DTBztmM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=BURDQS2T; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="BURDQS2T" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1787800439; x=1819336439; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=PgeX2FlPPmtq8ItOCQvA8Gr2xrRj7YJ4vq9H3xkQ5y8=; b=BURDQS2T8Q0362BjiddE3iskqS+FlyIgLjiB9jD7ALYn13sAYq7nIomW VIyfqi5mkV0S0jFktf+NbQ3ZFkxpeEHwmbqR3PRrQTlBIktP3Tyd+sc5W SV/bxQTYbmdaAe8mmFiW+JW6AV1NWEOJThyk+OApTlP0hpw0BSN34OVhg lVPpBo2RBPW1dHJthfWJg01YsuQEZSg5JHZy2DCB+apSzNYxoQFimla5i emyuAQ4VgpPrUjYPG2472NAXh4YIVWC+QoqUBTxSM4h1uQffdnLN0Vk2s Yd2rGxuMKx6sQYu1VVtHi43IPL94qbcNNaC2XAo2YYC4aibyFcM6g2XW+ A==; X-CSE-ConnectionGUID: etIE1P2WRo6sR+U9GLiTVw== X-CSE-MsgGUID: BOIy311iTQ+ezwZQA1wZkA== X-IronPort-AV: E=McAfee;i="6800,10657,11887"; a="98964534" X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="98964534" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:13:59 -0700 X-CSE-ConnectionGUID: nr0eLgzsRzmr/VJx3vsRQw== X-CSE-MsgGUID: 1L598Ys9Th+E7gK6s9mqcg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="271264018" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by ORVIESA003-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:13:56 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, binbin.wu@linux.intel.com Subject: [PATCH v3 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Date: Thu, 27 Aug 2026 11:18:34 +0800 Message-ID: <20260827031837.2863609-2-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260827031837.2863609-1-binbin.wu@linux.intel.com> References: <20260827031837.2863609-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable CPUID feature bits that KVM supports, and build the masks during TDX hardware setup via tdx_initialize_cpu_cfg_caps(). The TDX module reports the CPUID bits that the VMM can directly configure for a TD, but KVM cannot blindly expose all reported bits to userspace. Certain features imply additional architectural state, e.g. one or more MSRs, that KVM must explicitly manage across host/guest transitions to prevent host state corruption. Today KVM relies on a hardcoded denylist, i.e. it clears a few known problematic bits, e.g. TSX and WAITPKG, and passes everything else through. A denylist is fundamentally fragile while an allowlist inverts the default, i.e. unknown configurable bits are hidden and not allowed to be enabled until KVM explicitly opts in. Except for a few fixed-1 bits required for basic TDX support, host state clobbering features are either directly configurable or gated by TD ATTRIBUTES/XFAM. Tracking only the directly configurable feature bits is therefore sufficient to serve the purpose while keeping the code footprint small. Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can be built with the similar feature-name based initializers. CPUID registers that hold directly configurable non-feature (multi-bit) fields are handled separately. The allowlist is consumed by later patches to filter KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID feature bits stay hidden from userspace until KVM explicitly opts in. Add comments as placeholders for HLE, RTM and WAITPKG, which KVM doesn't support for TDX yet. Signed-off-by: Binbin Wu Reviewed-by tags. --- v3: - Drop the new data structure in v2 and only track feature bits by following the organization of kvm_cpu_caps[], handle non-feature bits separately. (Sean) - Use two versions of macros (TDX_CFG_F() VS. TDX_CFG_EXTRA_F()) to distinguish whether a supported TDX configurable CPUID bit should be checked against KVM's common cpu capabilities. - Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc. --- arch/x86/kvm/vmx/tdx.c | 145 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 145 insertions(+) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index b272c20586a7..d4a3a42cfd9d 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -52,6 +52,149 @@ __TDX_BUG_ON(__err, #__fn, __kvm, ", " #a1 " 0x%llx, " #a2 ", 0x%llx, " #= a3 " 0x%llx", \ a1, a2, a3) =20 +static u32 tdx_cpu_cfg_caps[NR_KVM_CPU_CAPS] __ro_after_init; +static_assert(ARRAY_SIZE(tdx_cpu_cfg_caps) =3D=3D ARRAY_SIZE(kvm_cpu_caps)= ); + +#define TDX_VALIDATE_CPU_CAP_USAGE(name) \ + BUILD_BUG_ON(__feature_leaf(X86_FEATURE_##name) !=3D \ + tdx_cpu_cap_init_in_progress) + +/* For feature bit that KVM advertised through kvm_cpu_caps[]. */ +#define TDX_CFG_F(name) \ +({ \ + TDX_VALIDATE_CPU_CAP_USAGE(name); \ + tdx_cfg_caps |=3D feature_bit(name); \ +}) + +/* + * For feature bit KVM allows for TDX guests even though it is not adverti= sed + * through kvm_cpu_caps[], e.g. MWAIT. + */ +#define TDX_CFG_EXTRA_F(name) \ +({ \ + TDX_VALIDATE_CPU_CAP_USAGE(name); \ + tdx_cfg_extra_caps |=3D feature_bit(name); \ +}) + +#define tdx_cpu_cfg_cap_init(leaf, feature_initializers...) \ +do { \ + const u32 __maybe_unused tdx_cpu_cap_init_in_progress =3D leaf; \ + u32 tdx_cfg_extra_caps =3D 0; \ + u32 tdx_cfg_caps =3D 0; \ + \ + feature_initializers \ + tdx_cpu_cfg_caps[leaf] =3D (tdx_cfg_caps & kvm_cpu_caps[leaf]) | \ + tdx_cfg_extra_caps; \ +} while (0) + +/* + * Track only CPUID feature bits that are directly configurable by userspa= ce. + * Features controlled by XFAM or ATTRIBUTES are excluded; userspace cannot + * enable them until KVM adds support for the corresponding control. + */ +static void __init tdx_initialize_cpu_cfg_caps(void) +{ + tdx_cpu_cfg_cap_init(CPUID_1_ECX, + TDX_CFG_EXTRA_F(MWAIT), + TDX_CFG_F(TSC_DEADLINE_TIMER), + TDX_CFG_F(AVX), + TDX_CFG_F(F16C), + ); + + tdx_cpu_cfg_cap_init(CPUID_1_EDX, + TDX_CFG_F(MCE), + TDX_CFG_F(MTRR), + TDX_CFG_F(MCA), + TDX_CFG_F(SELFSNOOP), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_0_EBX, + TDX_CFG_F(BMI1), + /* HLE */ + TDX_CFG_F(BMI2), + TDX_CFG_F(ERMS), + /* RTM */ + TDX_CFG_F(AVX512F), + TDX_CFG_F(AVX512DQ), + TDX_CFG_F(ADX), + TDX_CFG_F(AVX512IFMA), + TDX_CFG_F(AVX512PF), + TDX_CFG_F(AVX512ER), + TDX_CFG_F(AVX512CD), + TDX_CFG_F(AVX512BW), + TDX_CFG_F(AVX512VL), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_ECX, + TDX_CFG_F(UMIP), + /* WAITPKG */ + TDX_CFG_F(AVX512_VBMI2), + TDX_CFG_F(GFNI), + TDX_CFG_F(VAES), + TDX_CFG_F(VPCLMULQDQ), + TDX_CFG_F(AVX512_VNNI), + TDX_CFG_F(AVX512_BITALG), + TDX_CFG_F(AVX512_VPOPCNTDQ), + TDX_CFG_F(LA57), + TDX_CFG_F(RDPID), + TDX_CFG_F(CLDEMOTE), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_EDX, + TDX_CFG_F(AVX512_4VNNIW), + TDX_CFG_F(AVX512_4FMAPS), + TDX_CFG_F(FSRM), + TDX_CFG_F(AVX512_VP2INTERSECT), + TDX_CFG_F(SERIALIZE), + TDX_CFG_F(TSXLDTRK), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_1_EAX, + TDX_CFG_F(SHA512), + TDX_CFG_F(SM3), + TDX_CFG_F(SM4), + TDX_CFG_F(AVX_VNNI), + TDX_CFG_F(AVX512_BF16), + TDX_CFG_F(CMPCCXADD), + TDX_CFG_F(FZRM), + TDX_CFG_F(FSRS), + TDX_CFG_F(FSRC), + TDX_CFG_F(LKGS), + TDX_CFG_F(WRMSRNS), + TDX_CFG_F(AMX_FP16), + TDX_CFG_F(AVX_IFMA), + TDX_CFG_F(LAM), + TDX_CFG_F(MOVRS), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_1_EDX, + TDX_CFG_F(AVX_VNNI_INT8), + TDX_CFG_F(AVX_NE_CONVERT), + TDX_CFG_F(AMX_COMPLEX), + TDX_CFG_F(AVX_VNNI_INT16), + TDX_CFG_F(PREFETCHITI), + TDX_CFG_F(AVX10), + ); + + tdx_cpu_cfg_cap_init(CPUID_7_2_EDX, + TDX_CFG_F(DDPD_U), + TDX_CFG_F(MCDT_NO), + ); + + tdx_cpu_cfg_cap_init(CPUID_1E_1_EAX, + TDX_CFG_F(AMX_FP8), + TDX_CFG_F(AMX_TF32), + TDX_CFG_F(AMX_AVX512), + TDX_CFG_F(AMX_MOVRS), + ); + + tdx_cpu_cfg_cap_init(CPUID_8000_0008_EBX, + TDX_CFG_F(WBNOINVD), + ); +} + +#undef TDX_CFG_F +#undef TDX_CFG_EXTRA_F =20 bool enable_tdx __ro_after_init; module_param_named(tdx, enable_tdx, bool, 0444); @@ -3481,6 +3624,8 @@ int __init tdx_hardware_setup(void) return r; } =20 + tdx_initialize_cpu_cfg_caps(); + KVM_SANITY_CHECK_VM_STRUCT_SIZE(kvm_tdx); =20 vt_x86_ops.vm_size =3D max_t(unsigned int, vt_x86_ops.vm_size, sizeof(str= uct kvm_tdx)); --=20 2.46.0 From nobody Mon Sep 28 03:41:40 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E421E38AC88; Thu, 27 Aug 2026 03:14:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800446; cv=none; b=R7jhpQLECLL1Lcp3fjWk/jdDsLkOmBH59aFvz/dot6J8AKjxOCIFw+nVEu8gVl4OC9N0pr8huv4YE2v81OsyUtuyTPBCgo54HK1y7OxwUMV/7MRVjhnn5DZ5fKVFUrQsDAGLXjVbC8JsHnQjJSFrWUc2FEFor7DWWhgrGvLs2O0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800446; c=relaxed/simple; bh=C7AG+FFJ4OjI97Ua5YzNJ5TTkUPgg7ctdaMgkiJfvGc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WV58zXPtPxJqnSU8C1h6d0dKxUIje3hLLA88ym+l0+1gPCocmvPQMQNOrKbrQ7wBzQG41Hu0O50nVSyMY5JBld6KfUF9aHrZHMOgIXG0+YSXyZazDfOFhLgyacPAIPpcW3nLprTC9Bfv7HXPBi/wMoajaZPUYOXYzjfa+C4x9uY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=fcTn263/; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="fcTn263/" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1787800442; x=1819336442; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=C7AG+FFJ4OjI97Ua5YzNJ5TTkUPgg7ctdaMgkiJfvGc=; b=fcTn263/b0g7qtQcxqgXciMyvn8JraRv6FCC9r3gUL5H6NmjnE7T7jxA 39EmJoZX4jnCNnQqEhapvg4eoQrLCxzlPoOLVW5rdfyWkJIEUi+MJ7cxv Lnrvt6U05c4ikIdIQkGG2BZqmP4r80wLXC4jr5268MDZ/QGNw0XDrp26F QVHytQ6qCGJdZIQovbYQUu8iyvm0quYErDtZzhI+PdhQaI+Izr5PFCymC Sc0Qbh1ugvHI4iswAQzn1y2wSB6Xwua2Swg7VrjZcuIRhc0n9hTV3hRo6 UAc6F7AIxs7BfO4li27858DPFYuFRCBlzk8XaHa3Nten+f/h27RDL+pHO g==; X-CSE-ConnectionGUID: HDyXIvinQluRq2qA7joKAw== X-CSE-MsgGUID: SVuFLbGvTpqNb+bAZRSjHQ== X-IronPort-AV: E=McAfee;i="6800,10657,11887"; a="98964543" X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="98964543" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:14:01 -0700 X-CSE-ConnectionGUID: RIN6N52eTdaHiAuIDTkFYw== X-CSE-MsgGUID: AV9OpNZ6TS+CHPeSuE8AwA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="271264035" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by ORVIESA003-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:13:59 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, binbin.wu@linux.intel.com Subject: [PATCH v3 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Date: Thu, 27 Aug 2026 11:18:35 +0800 Message-ID: <20260827031837.2863609-3-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260827031837.2863609-1-binbin.wu@linux.intel.com> References: <20260827031837.2863609-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX directly configurable CPUID feature bits, even though KVM doesn't support MSR_IA32_CORE_CAPS for TDX guests, to accommodate legacy TDX module behavior. Older TDX specs define the CORE_CAPABILITIES CPUID bit as fixed-1, so userspace may expect the bit to be enabled for TDs. If the bit becomes directly configurable in a newer TDX module but is not reported as such to userspace, userspace can no longer enable it once KVM starts validating the CPUID configuration input. Reporting CORE_CAPABILITIES as configurable keeps userspace able to enable the bit across the fixed-1 =3D> configurable transition, and lets userspace infer that the bit is no longer fixed-1 so it can adjust its expectations. Keep MSR_IA32_CORE_CAPS unsupported for TDX guests, as existing TDX users have not needed guest access to the MSR, and advertising the CPUID bit as configurable is enough for userspace to handle the legacy-module compatibility case. Signed-off-by: Binbin Wu Reviewed-by tags. Reviewed-by: Tony Lindgren --- v3: - Move this patch earlier in the series to avoid breaking userspace during bisection. (Xiaoyao) - Drop the code for MSR_IA32_CORE_CAPS access. --- arch/x86/kvm/vmx/tdx.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index d4a3a42cfd9d..b020518717ac 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -147,6 +147,12 @@ static void __init tdx_initialize_cpu_cfg_caps(void) TDX_CFG_F(AVX512_VP2INTERSECT), TDX_CFG_F(SERIALIZE), TDX_CFG_F(TSXLDTRK), + /* + * KVM does not support MSR_IA32_CORE_CAPS, but older TDX specs + * define this bit as fixed-1. Report it as configurable so + * userspace can know the feature is no longer a fixed-1 bit. + */ + TDX_CFG_EXTRA_F(CORE_CAPABILITIES), ); =20 tdx_cpu_cfg_cap_init(CPUID_7_1_EAX, --=20 2.46.0 From nobody Mon Sep 28 03:41:40 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A5142388E5A; Thu, 27 Aug 2026 03:14:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800447; cv=none; b=p7am6IEnzFgKlupUtWpWGQk7SQNmmfD+ThHKGOgWbUTJpzIeBB0sBcKUgmKO8PuBgNZTxBDjxzRnhdpa99WmhyQjsGtowcqGhDCPdAmRRC6OJF3XEi5pyFOlP2uFHX0C++Z+DgpgeNfzJNwLQmSBlP7yAvNv2JjBO1pRWvoW7gc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800447; c=relaxed/simple; bh=RRvUbCKMXF9daTfoWlWbm9gI2Ol0koeM0rWJjlqFPi4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=SetpxHVxHvxAqJu+2Mop+OA/GEWhU5+boY0Vl8cKuHACAM+ZXYDPmT4xVcqmJCtMXE25owIqTnvTR8rtJEr6mWuoH3O6T/V/VkIQUh/W8OOnDlZ3V5ynA6dGC/5k1AmpXa6nvZ5p0kUJYEPvKqGlEWhXhkoB/y7n4UN3ojk/tSo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=lnnnlh5f; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="lnnnlh5f" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1787800445; x=1819336445; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=RRvUbCKMXF9daTfoWlWbm9gI2Ol0koeM0rWJjlqFPi4=; b=lnnnlh5fwMafHErvcpN5oS1pP0vqgu5PHN5trXinuGUYFf3cTwAckdjz E06jNLggLi25BG0h4dBd+ggmArvCUlO5N7Jaif7voBk5S1WQTV45KCrzn DT/BomY89veDZYSzsIXmwmGY2C3sDI9J2Kv5SoiqNEeVdVUCRS+vq3iWy RCcIua3jQBJGYO0cycrWKBmw+nkt661tQmBsgE/BHjbRPgEvyaHsXKCzA cw7LP65fEmJntJzvp34I/ph3JJoDMhcG82SCeX036pS3JwZOMHI9XYQTy wFUPswuNol04pL7Heb5g62j5XlK/TSguKIZz2DlkcAomELlEcYMtEEwmD g==; X-CSE-ConnectionGUID: 3aDMmY1lTdWkoGXWuBjdKg== X-CSE-MsgGUID: Y/eSgohfRw2+6WvAJzrWwQ== X-IronPort-AV: E=McAfee;i="6800,10657,11887"; a="98964554" X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="98964554" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:14:04 -0700 X-CSE-ConnectionGUID: AVa4WG41TQeLLte1G3pa/w== X-CSE-MsgGUID: ccUGhdnuT9OXcFeoFoLmKw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="271264065" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by ORVIESA003-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:14:01 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, binbin.wu@linux.intel.com Subject: [PATCH v3 3/4] KVM: TDX: Filter configurable CPUID bits Date: Thu, 27 Aug 2026 11:18:36 +0800 Message-ID: <20260827031837.2863609-4-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260827031837.2863609-1-binbin.wu@linux.intel.com> References: <20260827031837.2863609-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Filter the directly configurable CPUID bits reported through KVM_TDX_CAPABILITIES against KVM's TDX allowlist, and drop the hardcoded denylist based filtering. The TDX module reports all directly configurable CPUID bits that it supports for a TD, but KVM must not expose bits that it doesn't support, as blindly exposing a host state clobbering feature can lead to host state corruption. The existing denylist, which clears only TSX and WAITPKG, is not fail-safe. Add tdx_get_allowed_cfg_cpuid_mask() to get the mask of directly configurable bits allowed by KVM for a given CPUID register, covering both feature bits, which come from tdx_cpu_cfg_caps[], and non-feature bits, which are enumerated at runtime. Apply the mask to every CPUID register reported through KVM_TDX_CAPABILITIES. With the allowlist in place, newly introduced TDX directly configurable CPUID bits stay hidden from userspace until KVM explicitly opts in. Signed-off-by: Binbin Wu Reviewed-by tags. --- v3: - Handle non-feature leafs at runtime. (Sean) - Add CPUID.0x24.0.EBX[7:0] into allow list. There is a mismatch of the description about CPUID.0x24.0.EBX[7:0], which is listed as "XFAM & CPUID_Enabled & Native" but should be "XFAM & CPUID_Enabled & Configured & Native". --- arch/x86/kvm/vmx/tdx.c | 84 +++++++++++++++++++++++++++++++++--------- 1 file changed, 66 insertions(+), 18 deletions(-) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index b020518717ac..e8951353de73 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -277,34 +277,76 @@ static bool has_tsx(const struct kvm_cpuid_entry2 *en= try) (entry->ebx & TDX_FEATURE_TSX); } =20 -static void clear_tsx(struct kvm_cpuid_entry2 *entry) -{ - entry->ebx &=3D ~TDX_FEATURE_TSX; -} - static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry) { return entry->function =3D=3D 7 && entry->index =3D=3D 0 && (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG)); } =20 -static void clear_waitpkg(struct kvm_cpuid_entry2 *entry) +static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry) { - entry->ecx &=3D ~__feature_bit(X86_FEATURE_WAITPKG); + return has_tsx(entry) || has_waitpkg(entry); } =20 -static void tdx_clear_unsupported_cpuid(struct kvm_cpuid_entry2 *entry) +#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0) + +static u32 tdx_cfg_non_feature_mask(u32 function, u32 index, int reg) { - if (has_tsx(entry)) - clear_tsx(entry); + /* + * For a leaf/subleaf/register that will never be repurposed to hold + * feature bits, it's safe to return TDX_CPUID_ALL_ALLOWED_MASK, i.e. + * leave the TDX module's CPUID config mask intact. + */ + switch (function) { + case 1: + if (reg =3D=3D CPUID_EAX || reg =3D=3D CPUID_EBX) + return TDX_CPUID_ALL_ALLOWED_MASK; + return 0; + case 4: + case 0x18: + case 0x1f: + return TDX_CPUID_ALL_ALLOWED_MASK; + case 0x24: + if (index =3D=3D 0 && reg =3D=3D CPUID_EBX) + return GENMASK_U32(7, 0); + return 0; + case 0x80000008: + if (reg =3D=3D CPUID_EAX) + return TDX_CPUID_ALL_ALLOWED_MASK; + return 0; + default: + return 0; + } +} =20 - if (has_waitpkg(entry)) - clear_waitpkg(entry); +static u32 tdx_cfg_feature_mask(u32 function, u32 index, int reg) +{ + for (int i =3D 0; i < NR_KVM_CPU_CAPS; i++) { + const struct cpuid_reg *cpuid =3D &reverse_cpuid[i]; + + if (!cpuid->function) + continue; + + if (cpuid->function =3D=3D function && cpuid->index =3D=3D index && + cpuid->reg =3D=3D reg) + return tdx_cpu_cfg_caps[i]; + } + + return 0; } =20 -static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry) +static u32 tdx_get_allowed_cfg_cpuid_mask(u32 function, u32 index, int reg) { - return has_tsx(entry) || has_waitpkg(entry); + u32 non_feature_mask =3D tdx_cfg_non_feature_mask(function, index, reg); + + if (non_feature_mask =3D=3D TDX_CPUID_ALL_ALLOWED_MASK) + return TDX_CPUID_ALL_ALLOWED_MASK; + + /* + * It's possible that a CPUID register contains both feature and + * non-feature bits. + */ + return non_feature_mask | tdx_cfg_feature_mask(function, index, reg); } =20 #define KVM_TDX_CPUID_NO_SUBLEAF ((__u32)-1) @@ -330,8 +372,6 @@ static void td_init_cpuid_entry2(struct kvm_cpuid_entry= 2 *entry, unsigned char i */ if (entry->function =3D=3D 0x80000008) entry->eax =3D tdx_set_guest_phys_addr_bits(entry->eax, 0xff); - - tdx_clear_unsupported_cpuid(entry); } =20 #define TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT BIT(1) @@ -354,8 +394,16 @@ static int init_kvm_tdx_caps(const struct tdx_sys_info= _td_conf *td_conf, caps->user_tdvmcallinfo_1_r11 =3D TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT; =20 - for (i =3D 0; i < td_conf->num_cpuid_config; i++) - td_init_cpuid_entry2(&caps->cpuid.entries[i], i); + for (i =3D 0; i < td_conf->num_cpuid_config; i++) { + struct kvm_cpuid_entry2 *e =3D &caps->cpuid.entries[i]; + + td_init_cpuid_entry2(e, i); + /* Only report the configurable bits allowed by KVM. */ + e->eax &=3D tdx_get_allowed_cfg_cpuid_mask(e->function, e->index, CPUID_= EAX); + e->ebx &=3D tdx_get_allowed_cfg_cpuid_mask(e->function, e->index, CPUID_= EBX); + e->ecx &=3D tdx_get_allowed_cfg_cpuid_mask(e->function, e->index, CPUID_= ECX); + e->edx &=3D tdx_get_allowed_cfg_cpuid_mask(e->function, e->index, CPUID_= EDX); + } =20 return 0; } --=20 2.46.0 From nobody Mon Sep 28 03:41:40 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DCCBA388E6B; Thu, 27 Aug 2026 03:14:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800448; cv=none; b=E9x+KY/TerjNc5AzkifLaw24V/sFjCOQCfGzw8ofAZyX0zHH8avXct2G6E13bEGUCkUl31h6m7ohkz1eIYK2v7fDu0U8iOOF5Z7h33M3AG2t/V1dzr1jLGL0gvbqrBROJJgtQHhiEjyPj/t5co4FfStbI4MA+Iv8i4ru9cSyRKI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787800448; c=relaxed/simple; bh=QOR5BllIEIxYwF69J+nzg+NHNYPor5+EuELD4Kql4dI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=k9g4bWSyDR7ldHBI9RlgpC5JtOjoPCTCGl1Botdnx62UO3H6SqDNN+m6gdZqJXdTLomHisH8Hqo/EwYY0mHdHwaf5G5GrbkonPsYr9Ga2E+5AWfo7fQpaAX5J7SoJCpsx/Sn6jJDvtRUA6RFezXu6MjvnvoEQnJUBXVM7qJM3rc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=ip1Txlix; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="ip1Txlix" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1787800447; x=1819336447; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=QOR5BllIEIxYwF69J+nzg+NHNYPor5+EuELD4Kql4dI=; b=ip1TxlixT05lPdgc7SGPlk1hcNyZYNNrNg1GQtjd8RvwYVu0kSxuWO3b K54fgEXp9LWo3naEQo515sJHL7D9LyROF7WdMcU83BUf1h6kAqdl5jWhr XDbqXh96wbvd7Kw19JzpG0lNcd+piXBtBNdpGvoEkhKtA0RWEGHCYau2C rblI2YDBKncdFDtyoKax+OBEmq+YA6LElWKchDRJlnqn2gxFmxigiUqBt XDbfJGPVsxsuMnLfdp26YO+CDoSbtB2ilopvuCIuH8Ri/3l+SP+ZDJ3jk RyJQeYfaTkWDU1lalezJS4BlxLTXvihuMxKiWwupjnO+ynnL+M5lje1O9 Q==; X-CSE-ConnectionGUID: 3oMH0IAiRnKaYy+RxN48SQ== X-CSE-MsgGUID: L+ahKcQlR1aMcRqERznzYg== X-IronPort-AV: E=McAfee;i="6800,10657,11887"; a="98964564" X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="98964564" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:14:07 -0700 X-CSE-ConnectionGUID: 4EEIuD7ZRh6PwztRZ2YW8w== X-CSE-MsgGUID: iiBgwFayQDS/D1laAu0m6Q== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,245,1779174000"; d="scan'208";a="271264125" Received: from litbin-desktop.sh.intel.com ([10.239.57.15]) by ORVIESA003-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 26 Aug 2026 20:14:04 -0700 From: Binbin Wu To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: seanjc@google.com, pbonzini@redhat.com, dave.hansen@linux.intel.com, andrew.cooper3@citrix.com, nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com, xiaoyao.li@intel.com, chao.gao@intel.com, binbin.wu@linux.intel.com Subject: [PATCH v3 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Date: Thu, 27 Aug 2026 11:18:37 +0800 Message-ID: <20260827031837.2863609-5-binbin.wu@linux.intel.com> X-Mailer: git-send-email 2.46.0 In-Reply-To: <20260827031837.2863609-1-binbin.wu@linux.intel.com> References: <20260827031837.2863609-1-binbin.wu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Validate the CPUID configuration provided by userspace through KVM_TDX_INIT_VM against KVM's TDX allowlist, and drop the hardcoded denylist based check. The TDX module lets the VMM configure certain CPUID features for a TD at initialization time, but KVM must strictly govern which of them userspace can actually enable, otherwise a host state clobbering feature could be enabled behind KVM's back. The existing check only rejects TSX and WAITPKG, i.e. it is not fail-safe, as any bit that a future TDX module makes configurable would be accepted even if KVM has no idea about the feature. Add tdx_has_unsupported_cfg_cpuid_bit() and reject KVM_TDX_INIT_VM if userspace sets any bit outside the mask returned by tdx_get_allowed_cfg_cpuid_mask(). There is no need to first mask the userspace input with the bits the TDX module reports as directly configurable, as anything outside that set is rejected by the TDX module itself. Also reject CPUID entries whose index differs from the value expected by the TDX module, as kvm_find_cpuid_entry2() ignores the index when KVM_CPUID_FLAG_SIGNIFCANT_INDEX is cleared, and so a mismatching entry could otherwise be applied to the wrong subleaf. Signed-off-by: Binbin Wu Reviewed-by tags. Reviewed-by: Tony Lindgren --- v3: - Check CPUID entry index mismatch b/t userspace input and TDX sysinfo configuration. (Sashiko) - No need to mask the userspace input with TDX module reported directly configurable bits first, since the userspace input should be subset of directly configurable bits. Otherwise, the input will be rejected by the TDX module. --- arch/x86/kvm/vmx/tdx.c | 41 +++++++++++++++++++++-------------------- 1 file changed, 21 insertions(+), 20 deletions(-) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index e8951353de73..12dea8775fd4 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -269,25 +269,6 @@ static u32 tdx_set_guest_phys_addr_bits(const u32 eax,= int addr_bits) return (eax & ~GENMASK(23, 16)) | (addr_bits & 0xff) << 16; } =20 -#define TDX_FEATURE_TSX (__feature_bit(X86_FEATURE_HLE) | __feature_bit(X8= 6_FEATURE_RTM)) - -static bool has_tsx(const struct kvm_cpuid_entry2 *entry) -{ - return entry->function =3D=3D 7 && entry->index =3D=3D 0 && - (entry->ebx & TDX_FEATURE_TSX); -} - -static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry) -{ - return entry->function =3D=3D 7 && entry->index =3D=3D 0 && - (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG)); -} - -static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry) -{ - return has_tsx(entry) || has_waitpkg(entry); -} - #define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0) =20 static u32 tdx_cfg_non_feature_mask(u32 function, u32 index, int reg) @@ -2533,6 +2514,17 @@ static int setup_tdparams_eptp_controls(struct kvm_c= puid2 *cpuid, return 0; } =20 +static bool tdx_has_unsupported_cfg_cpuid_bit(const struct kvm_cpuid_entry= 2 *entry) +{ + u32 function =3D entry->function; + u32 index =3D entry->index; + + return (entry->eax & ~tdx_get_allowed_cfg_cpuid_mask(function, index, CPU= ID_EAX)) || + (entry->ebx & ~tdx_get_allowed_cfg_cpuid_mask(function, index, CPU= ID_EBX)) || + (entry->ecx & ~tdx_get_allowed_cfg_cpuid_mask(function, index, CPU= ID_ECX)) || + (entry->edx & ~tdx_get_allowed_cfg_cpuid_mask(function, index, CPU= ID_EDX)); +} + static int setup_tdparams_cpuids(struct kvm_cpuid2 *cpuid, struct td_params *td_params) { @@ -2556,7 +2548,16 @@ static int setup_tdparams_cpuids(struct kvm_cpuid2 *= cpuid, if (!entry) continue; =20 - if (tdx_unsupported_cpuid(entry)) + /* + * Reject entries whose index does not match the expected one. + * This catches userspace passing a CPUID entry with the + * KVM_CPUID_FLAG_SIGNIFCANT_INDEX flag cleared when the index + * is significant. + */ + if (entry->index !=3D tmp.index) + return -EINVAL; + + if (tdx_has_unsupported_cfg_cpuid_bit(entry)) return -EINVAL; =20 copy_cnt++; --=20 2.46.0