arch/x86/kvm/vmx/tdx.c | 258 ++++++++++++++++++++++++++++++++++++----- 1 file changed, 229 insertions(+), 29 deletions(-)
Hi, The purpose of this patch series is to prevent userspace from enabling host state clobbering features that KVM does not support for TDX. A host state clobbering feature exposed on a new TDX module/platform can corrupt host state if KVM does not explicitly save and restore the related MSR(s) across host/guest transitions. If such a feature is blindly exposed to and used by a TD, the host will behave unexpectedly. Except for a few fixed-1 bits required for basic TDX support, host state clobbering features are either directly configurable or gated by TD ATTRIBUTES/XFAM. So an allowlist covering only the directly configurable CPUID bits, plus the corresponding filtering and validation, is sufficient to serve the purpose while keeping the code footprint small. Open question for Sean/Paolo ============================ This revision comes long after Sean confirmed the overall direction on RFC v2 [1], due to internal discussions about a potential backward compatibility issue: a CPUID field can change virtualization type to directly configurable. The TDX module can maintain backward compatibility in most cases; the problematic case is a single bit that was effectively "1" and non-configurable becoming directly configurable. The only scenario we can envision for that is an x86 feature deprecation. In that case the TDX module has to decide how to handle deprecation, especially with respect to TD migration compatibility. That pre-existing VMs with the deprecated feature enabled cannot migrate from an old platform to a new one is a problem shared by TDX and VMX, but the TDX module additionally has to decide: - Whether to allow TDs newly created on old platforms to migrate to new platforms. - Whether to allow TDs created on new platforms to migrate to old platforms. There are two options: 1. Do not offer the VMM the ability to configure the bit, i.e. simply report it as zero on platforms that no longer support the feature. This disallows both migration cases above. 2. Make the bit directly configurable so the VMM can clear it even on platforms that still support the feature, allowing both migration cases above. However, since the bit is not in KVM's allowlist, userspace would no longer be able to enable it on old platforms after a TDX module update. Adding all supported CPUID bits to the allowlist, regardless of their virtualization type in the TDX module, would avoid the compatibility issue for option 2, at the cost of a significantly larger series. Given that x86 feature deprecation is expected to be very rare, and that Intel will discuss such cases with the community beforehand, does the approach taken by this series look acceptable? If so, the goal for this revision is to collect detailed feedback and, where appropriate, Reviewed-by tags. Expected host state clobbering behavior for TDX =============================================== We also want to call for discussions about the expected host state clobbering behavior for TDX here for future features. For a normal VMX guest, VM entry/exit behavior for a given piece of CPU state is architecturally defined: state is either switched by hardware via VMCS host/guest fields, or left as the guest value on VM exit and managed by KVM in software. For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the TDX module does with a given piece of host state is defined by the TDX module ABI rather than by the x86 architecture. What we would like to align on is the expected baseline behavior of TDH.VP.ENTER for future features. The proposal is to have TDX simply match VMX behavior, i.e. on return from TDH.VP.ENTER, state that VMX would restore from the VMCS host fields is restored, and state that VMX would leave as the guest value is clobbered. That keeps a single model for VMM, and means enabling a new feature for TDs requires the same work flow as enabling it for VMX. FRED is a useful concrete example. Under VMX, the FRED host state in IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the TDX module is expected to clobber them on TDH.VP.ENTER return. The series ========== Starting with v2 [2], the series takes a simpler, TDX-contained approach instead of the comprehensive CPUID paranoid verification framework across VMX, SVM, and TDX proposed in v1 [3]. It validates only the TDX directly configurable CPUID bits, which are reported by the TDX module in CPUID_CONFIG fields that the VMM can configure for a TD. All filtering and validation logic is isolated within TDX code. This is sufficient to address the host clobbering issue because: - The TDX module will not introduce new host state clobbering features that are fixed-1. - Non-directly configurable feature bits (i.e. features controlled by XFAM or ATTRIBUTES) cannot be enabled by userspace without KVM's support. Specifically, this series builds a KVM-side allowlist of supported TDX directly configurable CPUID bits to: - Filter KVM_TDX_CAPABILITIES Replace the hardcoded denylist to only report configurable bits that KVM explicitly supports. - Validate KVM_TDX_INIT_VM Reject any configurable bit that the TDX module allows but KVM does not support, as well as CPUID entries with an unexpected subleaf. With this allowlist, newly added TDX configurable CPUID bits will not be exposed to userspace until KVM explicitly opts in after fulfilling the necessary virtualization requirements. The allowlist consists of two parts: - Feature CPUID bits, which are tracked in tdx_cpu_cfg_caps[] following the organization of kvm_cpu_caps[]. It holds KVM's supported masks for the TDX configurable CPUID feature bits. - Non-feature multi-bit fields, which are handled at runtime for the CPUID registers holding such fields. For compatibility with older TDX modules that report CORE_CAPABILITIES as fixed-1, report CORE_CAPABILITIES as configurable even though KVM does not support guest access to MSR_IA32_CORE_CAPS. This allows userspace to keep enabling CORE_CAPABILITIES when the bit changes from fixed-1 to directly configurable, and allows userspace to infer that the bit is no longer fixed-1 so it can adjust its expectations. Known limitation: - The series does not check KVM_SET_CPUID2 input for consistency with the CPUID configuration supplied through KVM_TDX_INIT_VM. This is not required for host safety today because KVM does not use its vCPU CPUID model to decide whether to manage host-clobbering state for TDs. A consistency check should be added if KVM ever starts relying on the vCPU CPUID model for such decisions. Changes from v2 [2]: - Drop the dedicated data structure introduced in v2; track only feature bits following the organization of kvm_cpu_caps[], and handle non-feature bits at runtime. (Sean) - Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc. - Add CPUID.0x24.0.EBX[7:0] into allow list. There is a mismatch of the description about CPUID.0x24.0.EBX[7:0], which is listed as "XFAM & CPUID_Enabled & Native" but should be "XFAM & CPUID_Enabled & Configured & Native". - Move the patch for CORE_CAPABILITIES earlier in the series to avoid breaking userspace during bisection. (Xiaoyao) - Report CORE_CAPABILITIES as a configurable bit to userspace for backward compatibility, but drop the MSR_IA32_CORE_CAPS access code. - Use two versions of macros to distinguish whether a supported TDX configurable CPUID bit should be checked against KVM's common CPU capabilities. - Validate KVM_TDX_INIT_VM input against the allowlist instead of masking it, and additionally reject entries whose index does not match the TDX sysinfo configuration. (Sashiko) Changes from v1 [3]: - Drop the overarching CPUID paranoid verification framework across VMX/SVM/TDX and the opt-in interface. (Sean) - Shift focus entirely to isolating and validating TDX directly configurable CPUID bits. The series follows the CSV version of the CPUID virtualization document from "Intel TDX Module ABI Definitions" [4], updated in June 2026. AI tools were used to: - Check the coverage of the directly configurable CPUID bits in the document against the allowlist in this series. - Polish the cover letter and commit messages. - Do code review. [1] https://lore.kernel.org/kvm/aj1fqZEcApyd4sxi@google.com [2] https://lore.kernel.org/kvm/20260604023314.3907511-1-binbin.wu@linux.intel.com [3] https://lore.kernel.org/kvm/20260417073610.3246316-1-binbin.wu@linux.intel.com [4] https://cdrdv2.intel.com/v1/dl/getContent/795381 Binbin Wu (4): KVM: TDX: Track configurable CPUID bits allowed by KVM KVM: TDX: Report CORE_CAPABILITIES as configurable KVM: TDX: Filter configurable CPUID bits KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM arch/x86/kvm/vmx/tdx.c | 258 ++++++++++++++++++++++++++++++++++++----- 1 file changed, 229 insertions(+), 29 deletions(-) base-commit: 76671054f9a1ff6abb976583cd8da37650acdc97 -- 2.46.0
On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote: > Hi, > > The purpose of this patch series is to prevent userspace from enabling > host state clobbering features that KVM does not support for TDX. A host > state clobbering feature exposed on a new TDX module/platform can corrupt > host state if KVM does not explicitly save and restore the related MSR(s) > across host/guest transitions. If such a feature is blindly exposed to > and used by a TD, the host will behave unexpectedly. > > Except for a few fixed-1 bits required for basic TDX support, host state > clobbering features are either directly configurable or gated by TD > ATTRIBUTES/XFAM. So an allowlist covering only the directly configurable > CPUID bits, plus the corresponding filtering and validation, is sufficient > to serve the purpose while keeping the code footprint small. > > Open question for Sean/Paolo > ============================ > This revision comes long after Sean confirmed the overall direction on > RFC v2 [1], due to internal discussions about a potential backward > compatibility issue: a CPUID field can change virtualization type to > directly configurable. The TDX module can maintain backward > compatibility in most cases; the problematic case is a single bit that > was effectively "1" and non-configurable becoming directly configurable. > The only scenario we can envision for that is an x86 feature deprecation. > > In that case the TDX module has to decide how to handle deprecation, > especially with respect to TD migration compatibility. That > pre-existing VMs with the deprecated feature enabled cannot migrate from > an old platform to a new one is a problem shared by TDX and VMX, but the > TDX module additionally has to decide: > - Whether to allow TDs newly created on old platforms to migrate to new > platforms. > - Whether to allow TDs created on new platforms to migrate to old > platforms. > > There are two options: > 1. Do not offer the VMM the ability to configure the bit, i.e. simply > report it as zero on platforms that no longer support the feature. > This disallows both migration cases above. > 2. Make the bit directly configurable so the VMM can clear it even on > platforms that still support the feature, allowing both migration > cases above. However, since the bit is not in KVM's allowlist, > userspace would no longer be able to enable it on old platforms after > a TDX module update. > > Adding all supported CPUID bits to the allowlist, regardless of their > virtualization type in the TDX module, would avoid the compatibility > issue for option 2, at the cost of a significantly larger series. I think we could only do this if no new fixed-1 bits are added. If they were, then a future fixes-1 bit could be added then deprecated. An old kernel would not know about it. But we will not see new fixed-1 bits? > > Given that x86 feature deprecation is expected to be very rare, and that > Intel will discuss such cases with the community beforehand, does the > approach taken by this series look acceptable? > In the event a feature bit is turned configurable to support creating TDs that can be migrated to new platforms without the feature, then the desired value for the bit on the old platform will be 0. Which is what it would end up being set as, if the bit was not in the KVM side allow list. So everything will work as desired for that use? But a user on the old platform that didn't care about migration and was using the feature, would suddenly lose it after TDX module upgrade that turned the bit configurable But it isn't a kernel bug, so I'd think to punt on this and let TDX module figure out a solution if/when the time comes. Probably various opt-in knobs to allow for working around it. If they are not worried, we don't need to worry and bloat the kernel to prepare for it. > If so, the goal for this > revision is to collect detailed feedback and, where appropriate, > Reviewed-by tags. > > Expected host state clobbering behavior for TDX > =============================================== > We also want to call for discussions about the expected host state > clobbering behavior for TDX here for future features. > > For a normal VMX guest, VM entry/exit behavior for a given piece of CPU > state is architecturally defined: state is either switched by hardware via > VMCS host/guest fields, or left as the guest value on VM exit and managed > by KVM in software. > > For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the > TDX module does with a given piece of host state is defined by the TDX > module ABI rather than by the x86 architecture. > > What we would like to align on is the expected baseline behavior of > TDH.VP.ENTER for future features. The proposal is to have TDX simply > match VMX behavior, i.e. on return from TDH.VP.ENTER, state that VMX would > restore from the VMCS host fields is restored, and state that VMX would > leave as the guest value is clobbered. That keeps a single model for VMM, > and means enabling a new feature for TDs requires the same work flow as > enabling it for VMX. Ideally the TDX save/restore would share code with normal VMs. On the other hand if we don't share enter/exit paths sufficiently, we may need to duplicate some save/restore in tdx code. Depending on how much of which category we have, it could be better for the kernel to have either one. And if TDX always saved/restored all state across VP.ENTER, then we don't need bit filtering? What are the problems then with exposing everything to userspace as we currently do? > > FRED is a useful concrete example. Under VMX, the FRED host state in > IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and > IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module > is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 > and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the > TDX module is expected to clobber them on TDH.VP.ENTER return. And so the ability to have a design like this for the feature depends on CPUID bit filtering to be in place, otherwise new modules on old kernels can clobber host state unexpectedly. Future features that follow this approach will need some VMM opt-in to say "VMM is filtering CPUID bit config", so TDX can safely expose new clobbering bits. Otherwise, features like FRED will be problems on old kernels. Do we want it? Where would it get plugged in? > > The series > ========== > > Starting with v2 [2], the series takes a simpler, TDX-contained approach > instead of the comprehensive CPUID paranoid verification framework across > VMX, SVM, and TDX proposed in v1 [3]. It validates only the TDX directly > configurable CPUID bits, which are reported by the TDX module in > CPUID_CONFIG fields that the VMM can configure for a TD. All filtering > and validation logic is isolated within TDX code. This is sufficient to > address the host clobbering issue because: > - The TDX module will not introduce new host state clobbering features > that are fixed-1. > - Non-directly configurable feature bits (i.e. features controlled by XFAM > or ATTRIBUTES) cannot be enabled by userspace without KVM's support. > > Specifically, this series builds a KVM-side allowlist of supported TDX > directly configurable CPUID bits to: > - Filter KVM_TDX_CAPABILITIES > Replace the hardcoded denylist to only report configurable bits that > KVM explicitly supports. > - Validate KVM_TDX_INIT_VM > Reject any configurable bit that the TDX module allows but KVM does > not support, as well as CPUID entries with an unexpected subleaf. > > With this allowlist, newly added TDX configurable CPUID bits will not be > exposed to userspace until KVM explicitly opts in after fulfilling the > necessary virtualization requirements. > > The allowlist consists of two parts: > - Feature CPUID bits, which are tracked in tdx_cpu_cfg_caps[] following > the organization of kvm_cpu_caps[]. It holds KVM's supported masks for > the TDX configurable CPUID feature bits. > - Non-feature multi-bit fields, which are handled at runtime for the CPUID > registers holding such fields. > > For compatibility with older TDX modules that report CORE_CAPABILITIES as > fixed-1, report CORE_CAPABILITIES as configurable even though KVM does not > support guest access to MSR_IA32_CORE_CAPS. This allows userspace to keep > enabling CORE_CAPABILITIES when the bit changes from fixed-1 to directly > configurable, and allows userspace to infer that the bit is no longer > fixed-1 so it can adjust its expectations. > > Known limitation: > - The series does not check KVM_SET_CPUID2 input for consistency with the > CPUID configuration supplied through KVM_TDX_INIT_VM. This is not > required for host safety today because KVM does not use its vCPU CPUID > model to decide whether to manage host-clobbering state for TDs. A > consistency check should be added if KVM ever starts relying on the vCPU > CPUID model for such decisions.
On 8/28/2026 3:33 AM, Edgecombe, Rick P wrote: > On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote: >> Hi, >> >> The purpose of this patch series is to prevent userspace from enabling >> host state clobbering features that KVM does not support for TDX. A host >> state clobbering feature exposed on a new TDX module/platform can corrupt >> host state if KVM does not explicitly save and restore the related MSR(s) >> across host/guest transitions. If such a feature is blindly exposed to >> and used by a TD, the host will behave unexpectedly. >> >> Except for a few fixed-1 bits required for basic TDX support, host state >> clobbering features are either directly configurable or gated by TD >> ATTRIBUTES/XFAM. So an allowlist covering only the directly configurable >> CPUID bits, plus the corresponding filtering and validation, is sufficient >> to serve the purpose while keeping the code footprint small. >> >> Open question for Sean/Paolo >> ============================ >> This revision comes long after Sean confirmed the overall direction on >> RFC v2 [1], due to internal discussions about a potential backward >> compatibility issue: a CPUID field can change virtualization type to >> directly configurable. The TDX module can maintain backward >> compatibility in most cases; the problematic case is a single bit that >> was effectively "1" and non-configurable becoming directly configurable. >> The only scenario we can envision for that is an x86 feature deprecation. >> >> In that case the TDX module has to decide how to handle deprecation, >> especially with respect to TD migration compatibility. That >> pre-existing VMs with the deprecated feature enabled cannot migrate from >> an old platform to a new one is a problem shared by TDX and VMX, but the >> TDX module additionally has to decide: >> - Whether to allow TDs newly created on old platforms to migrate to new >> platforms. >> - Whether to allow TDs created on new platforms to migrate to old >> platforms. >> >> There are two options: >> 1. Do not offer the VMM the ability to configure the bit, i.e. simply >> report it as zero on platforms that no longer support the feature. >> This disallows both migration cases above. >> 2. Make the bit directly configurable so the VMM can clear it even on >> platforms that still support the feature, allowing both migration >> cases above. However, since the bit is not in KVM's allowlist, >> userspace would no longer be able to enable it on old platforms after >> a TDX module update. >> >> Adding all supported CPUID bits to the allowlist, regardless of their >> virtualization type in the TDX module, would avoid the compatibility >> issue for option 2, at the cost of a significantly larger series. > > I think we could only do this if no new fixed-1 bits are added. If they were, > then a future fixes-1 bit could be added then deprecated. An old kernel would > not know about it. > > But we will not see new fixed-1 bits? The expectation is that no new fixed-1 bits will be added after the initial TDX basic support. A new fixed-1 bit means the x86 architecture itself leaves the TDX module no choice, so the TDX module will not add one unless some very unusual future architectural change forces it to (and it's hard to see such a design passing the design review in practice). If that ever does happen, it would imply a significant architectural change, and we can figure out how to handle it then. > >> >> Given that x86 feature deprecation is expected to be very rare, and that >> Intel will discuss such cases with the community beforehand, does the >> approach taken by this series look acceptable? >> > > In the event a feature bit is turned configurable to support creating TDs that > can be migrated to new platforms without the feature, then the desired value for > the bit on the old platform will be 0. Which is what it would end up being set > as, if the bit was not in the KVM side allow list. So everything will work as > desired for that use? Yes, it would work for migration cases. > > But a user on the old platform that didn't care about migration and was using > the feature, would suddenly lose it after TDX module upgrade that turned the bit > configurable Yes. This is exactly the backward compatibility issue for option 2. Though I am not sure how important it is when losing a feature that is being deprecated. > > But it isn't a kernel bug, so I'd think to punt on this and let TDX module > figure out a solution if/when the time comes. Probably various opt-in knobs to > allow for working around it. Do you mean on the old platforms, the according TDX module versions add opt-in knobs to allow the "effective 1 -> directly configurable" change? > If they are not worried, we don't need to worry and > bloat the kernel to prepare for it. > >> If so, the goal for this >> revision is to collect detailed feedback and, where appropriate, >> Reviewed-by tags. >> >> Expected host state clobbering behavior for TDX >> =============================================== >> We also want to call for discussions about the expected host state >> clobbering behavior for TDX here for future features. >> >> For a normal VMX guest, VM entry/exit behavior for a given piece of CPU >> state is architecturally defined: state is either switched by hardware via >> VMCS host/guest fields, or left as the guest value on VM exit and managed >> by KVM in software. >> >> For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the >> TDX module does with a given piece of host state is defined by the TDX >> module ABI rather than by the x86 architecture. >> >> What we would like to align on is the expected baseline behavior of >> TDH.VP.ENTER for future features. The proposal is to have TDX simply >> match VMX behavior, i.e. on return from TDH.VP.ENTER, state that VMX would >> restore from the VMCS host fields is restored, and state that VMX would >> leave as the guest value is clobbered. That keeps a single model for VMM, >> and means enabling a new feature for TDs requires the same work flow as >> enabling it for VMX. > > Ideally the TDX save/restore would share code with normal VMs. On the other hand > if we don't share enter/exit paths sufficiently, we may need to duplicate some > save/restore in tdx code. It probably needs some TDX specific handling, since the TDX module clobbers the MSRs (setting them to either their INIT values or some default values), whereas in the VMX case the guest values are left. Depending on how much of which category we have, it > could be better for the kernel to have either one. > > And if TDX always saved/restored all state across VP.ENTER, then we don't need > bit filtering? What are the problems then with exposing everything to userspace > as we currently do? One consideration is performance. - When The TDX module clobbers an MSR, there are 1 RDMSR + 2 WRMSR TDX WRMSR to restore guest value --- TDX RDMSR to save guest value TDX WRMSR to set to default value - When The TDX module save/restore an MSR, there are 2 RDMSR + 2 WRMSR TDX RDMSR to save host value TDX WRMSR to restore guest value --- TDX RDMSR to save guest value TDX WRMSR to restore host value In most cases where the TDX module chooses to clobber an MSR rather than restore the host value, that MSR isn't consumed in ring 0, so the Linux kernel/KVM simply restores the value before returning to userspace. That means for most TD entry/exit paths there is one extra MSR read per MSR. Previously in the PUCK meeting, we proposed that the TDX module preserve host state by default for any new feature, and provide an interface for the host VMM to opt in to "don't preserve" as an optimization. That proposal wasn't pursued, since Sean suggested letting KVM do the CPUID validation for TDX instead. > >> >> FRED is a useful concrete example. Under VMX, the FRED host state in >> IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and >> IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module >> is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 >> and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the >> TDX module is expected to clobber them on TDH.VP.ENTER return. > > And so the ability to have a design like this for the feature depends on CPUID > bit filtering to be in place, otherwise new modules on old kernels can clobber > host state unexpectedly. Yes. > Future features that follow this approach will need > some VMM opt-in to say "VMM is filtering CPUID bit config", so TDX can safely > expose new clobbering bits. Otherwise, features like FRED will be problems on > old kernels. If FRED support lands in KVM for non-TDX VMs before this filtering/invalidation is in place, adding FRED to the hardcoded denylist would work as a temporary solution. But that only holds for a well-behaved userspace VMM. QEMU, for example, only enables features supported by both KVM and TDX, so if FRED isn't supported in KVM for non-TDX VMs, QEMU won't enable it for TDX either. A malicious userspace VMM, however, can set a feature regardless of KVM's reported CPU capabilities, and could still cause trouble. > Do we want it? Where would it get plugged in? So if the approach in this series is taken, I think we still need such an opt-in interface to tell the TDX module that the VMM is now filtering the CPUID bits so that the TDX module knows that it's safe to report new host state clobbering features.
On 8/28/2026 11:19 AM, Binbin Wu wrote: > On 8/28/2026 3:33 AM, Edgecombe, Rick P wrote: >> On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote: <snip> >>> Expected host state clobbering behavior for TDX >>> =============================================== >>> We also want to call for discussions about the expected host state >>> clobbering behavior for TDX here for future features. >>> >>> For a normal VMX guest, VM entry/exit behavior for a given piece of CPU >>> state is architecturally defined: state is either switched by hardware via >>> VMCS host/guest fields, or left as the guest value on VM exit and managed >>> by KVM in software. >>> >>> For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the >>> TDX module does with a given piece of host state is defined by the TDX >>> module ABI rather than by the x86 architecture. >>> >>> What we would like to align on is the expected baseline behavior of >>> TDH.VP.ENTER for future features. The proposal is to have TDX simply >>> match VMX behavior, i.e. on return from TDH.VP.ENTER, state that VMX would >>> restore from the VMCS host fields is restored, and state that VMX would >>> leave as the guest value is clobbered. That keeps a single model for VMM, >>> and means enabling a new feature for TDs requires the same work flow as >>> enabling it for VMX. >> >> Ideally the TDX save/restore would share code with normal VMs. On the other hand >> if we don't share enter/exit paths sufficiently, we may need to duplicate some >> save/restore in tdx code. (Copy the FRED example here for reference) >>> >>> FRED is a useful concrete example. Under VMX, the FRED host state in >>> IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and >>> IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module >>> is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 >>> and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the >>> TDX module is expected to clobber them on TDH.VP.ENTER return. What I get, is not matching VMX behavior but matching the behavior KVM will perform for VMX. They are based on the assumption that KVM will always enable the save/restore VMCS fields for a new feature. But I don't think we can guarantee it. To me, "have TDX simply match VMX behavior" means: 1. if the VMX unconditionally save/restore a state, then TDX will do so. 2. if there are vm-entry/vm-exit load/save VMCS fields for a state, then provide the equivalent per-TD configurable interfaces which matches the VMCS fields. > It probably needs some TDX specific handling, since the TDX module clobbers the > MSRs (setting them to either their INIT values or some default values), whereas > in the VMX case the guest values are left. > >> Depending on how much of which category we have, it >> could be better for the kernel to have either one. >> >> And if TDX always saved/restored all state across VP.ENTER, then we don't need >> bit filtering? What are the problems then with exposing everything to userspace >> as we currently do? IMO, allowing userspace to expose/enable a new feature to a guest without KVM first evaluating it is always dangerous. It's not just about the state clobbering. We can know the implication for a new feature. For example, a feature consumes global per-socket resources. Allowing guest to use the feature might slowdown the host. Another example is a feature is used to catch bad behaviors, and in this case host would like to enforce the feature being forced on for the guest instead of allowing the guest to use (disable) the feature freely.
On Tue, 2026-09-01 at 17:38 +0800, Xiaoyao Li wrote: > > > Ideally the TDX save/restore would share code with normal VMs. On the > > > other hand if we don't share enter/exit paths sufficiently, we may need to > > > duplicate some save/restore in tdx code. > > (Copy the FRED example here for reference) > > >>> > >>> FRED is a useful concrete example. Under VMX, the FRED host state in > >>> IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and > >>> IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX > module > >>> is expected to restore these MSRs on TDH.VP.ENTER return. > IA32_FRED_RSP0 > >>> and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, > so the > >>> TDX module is expected to clobber them on TDH.VP.ENTER return. > > What I get, is not matching VMX behavior but matching the behavior KVM > will perform for VMX. They are based on the assumption that KVM will > always enable the save/restore VMCS fields for a new feature. But I > don't think we can guarantee it. > > To me, "have TDX simply match VMX behavior" means: > > 1. if the VMX unconditionally save/restore a state, then TDX will do so. > > 2. if there are vm-entry/vm-exit load/save VMCS fields for a state, then > provide the equivalent per-TD configurable interfaces which matches the > VMCS fields. What do you mean by this? Expose a TDX module interface to configure the clobber behavior for each feature with load/save configuration? That was similar to what we originally discussed, before pivoting to this solution. I was thinking if you configured a feature (for example shadow stack), it would automatically set the VMCS save/restore settings associated with that feature. (VM_EXIT_LOAD_CET_STATE/VM_ENTRY_LOAD_CET_STATE) This won't necessarily match KVM's behavior, because it could decide to not use the features. But we can probably get close with a simple rule that can make sense for all the VMMs. If later we want a host clobber interface on top of the bit filtering, in order to minimize TDX special handling, we can probably add it later for features we care about. I'd think we don't need it right now.
On 9/2/2026 1:41 AM, Edgecombe, Rick P wrote: > On Tue, 2026-09-01 at 17:38 +0800, Xiaoyao Li wrote: >>>> Ideally the TDX save/restore would share code with normal VMs. On the >>>> other hand if we don't share enter/exit paths sufficiently, we may need to >>>> duplicate some save/restore in tdx code. >> >> (Copy the FRED example here for reference) >> >> >>> >> >>> FRED is a useful concrete example. Under VMX, the FRED host state in >> >>> IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and >> >>> IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX >> module >> >>> is expected to restore these MSRs on TDH.VP.ENTER return. >> IA32_FRED_RSP0 >> >>> and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, >> so the >> >>> TDX module is expected to clobber them on TDH.VP.ENTER return. >> >> What I get, is not matching VMX behavior but matching the behavior KVM >> will perform for VMX. They are based on the assumption that KVM will >> always enable the save/restore VMCS fields for a new feature. But I >> don't think we can guarantee it. >> >> To me, "have TDX simply match VMX behavior" means: >> >> 1. if the VMX unconditionally save/restore a state, then TDX will do so. >> >> 2. if there are vm-entry/vm-exit load/save VMCS fields for a state, then >> provide the equivalent per-TD configurable interfaces which matches the >> VMCS fields. > > What do you mean by this? Expose a TDX module interface to configure the clobber > behavior for each feature with load/save configuration? That was similar to what > we originally discussed, before pivoting to this solution. yeah. This is what I meant. I was trying to show my literal understanding on "The proposal is to have TDX simply match VMX behavior". i.e., I don't think "have TDX simply match VMX behavior" is a good name/summary for what Binbin has proposed. > I was thinking if you configured a feature (for example shadow stack), it would > automatically set the VMCS save/restore settings associated with that feature. > (VM_EXIT_LOAD_CET_STATE/VM_ENTRY_LOAD_CET_STATE) > > This won't necessarily match KVM's behavior, because it could decide to not use > the features. But we can probably get close with a simple rule that can make > sense for all the VMMs. So the proposal is making TDX behave as if the relevant VMCS save/load controls (if any) are set around TDH.VP.ENTER. In fact, what matters for host vmm is just the VM_EXIT_LOAD_XXX control. So the proposal becomes "If there is VM_EXIT_LOAD_XXX control for a state, TDX needs to restore the host state after TDH.VP.ENTER. If no such contorl, TDX sets the state to INIT state after TDH.VP.ENTER". It's a fancy idea. And it provides a clear rule of how TDX handles states of a feature so that host VMM developers don't need to read the TDX module API to figure out what's the value of a state after TDH.VP.ENTER. It's helpful for host VMM developers, though I'm not sure on TDX module developers. > If later we want a host clobber interface on top of the bit filtering, in order > to minimize TDX special handling, we can probably add it later for features we > care about. I'd think we don't need it right now.
On Wed, 2026-09-02 at 18:29 +0800, Xiaoyao Li wrote: > So the proposal is making TDX behave as if the relevant VMCS save/load > controls (if any) are set around TDH.VP.ENTER. > > In fact, what matters for host vmm is just the VM_EXIT_LOAD_XXX control. > So the proposal becomes "If there is VM_EXIT_LOAD_XXX control for a > state, TDX needs to restore the host state after TDH.VP.ENTER. If no > such contorl, TDX sets the state to INIT state after TDH.VP.ENTER". > > It's a fancy idea. And it provides a clear rule of how TDX handles > states of a feature so that host VMM developers don't need to read the > TDX module API to figure out what's the value of a state after TDH.VP.ENTER. > > It's helpful for host VMM developers, though I'm not sure on TDX module > developers. I don't think they have agreed to it yet (Binbin?), but I think it mostly works this way already. Why do you think it is a burden on TDX module developers? Also, I think this can be the guideline. We could have exceptions.
On 9/2/2026 9:13 PM, Edgecombe, Rick P wrote: > On Wed, 2026-09-02 at 18:29 +0800, Xiaoyao Li wrote: >> So the proposal is making TDX behave as if the relevant VMCS save/load >> controls (if any) are set around TDH.VP.ENTER. >> >> In fact, what matters for host vmm is just the VM_EXIT_LOAD_XXX control. >> So the proposal becomes "If there is VM_EXIT_LOAD_XXX control for a >> state, TDX needs to restore the host state after TDH.VP.ENTER. If no >> such contorl, TDX sets the state to INIT state after TDH.VP.ENTER". >> >> It's a fancy idea. And it provides a clear rule of how TDX handles >> states of a feature so that host VMM developers don't need to read the >> TDX module API to figure out what's the value of a state after TDH.VP.ENTER. >> >> It's helpful for host VMM developers, though I'm not sure on TDX module >> developers. > > I don't think they have agreed to it yet (Binbin?), Not yet. We just brought the topic to the community for discussion first. > but I think it mostly works > this way already. Why do you think it is a burden on TDX module developers? > > Also, I think this can be the guideline. We could have exceptions.
On 9/2/2026 9:13 PM, Edgecombe, Rick P wrote: > On Wed, 2026-09-02 at 18:29 +0800, Xiaoyao Li wrote: >> So the proposal is making TDX behave as if the relevant VMCS save/load >> controls (if any) are set around TDH.VP.ENTER. >> >> In fact, what matters for host vmm is just the VM_EXIT_LOAD_XXX control. >> So the proposal becomes "If there is VM_EXIT_LOAD_XXX control for a >> state, TDX needs to restore the host state after TDH.VP.ENTER. If no >> such contorl, TDX sets the state to INIT state after TDH.VP.ENTER". >> >> It's a fancy idea. And it provides a clear rule of how TDX handles >> states of a feature so that host VMM developers don't need to read the >> TDX module API to figure out what's the value of a state after TDH.VP.ENTER. >> >> It's helpful for host VMM developers, though I'm not sure on TDX module >> developers. > > I don't think they have agreed to it yet (Binbin?), but I think it mostly works > this way already. Why do you think it is a burden on TDX module developers? Because it can bring confusion to TDX module developers. VM_EXIT_LOAD_XXX control is used by SEAM VMCS to load states for TDX module execution context. And I think for most features SEAM VMCS doesn't set it (I don't check it though). SEAMCALL is kind of a VM exit, and SEAMRET is kind of a VM entry. To automatically save and restore the host state, what TDX needs are VM_EXIT_SAVE_XXX, VM_ENTRY_LOAD_XXX, and vmcs guest state for XXX.
On Wed, 2026-09-02 at 21:39 +0800, Xiaoyao Li wrote: > > I don't think they have agreed to it yet (Binbin?), but I think it mostly > > works this way already. Why do you think it is a burden on TDX module > > developers? > > Because it can bring confusion to TDX module developers. VM_EXIT_LOAD_XXX > control is used by SEAM VMCS to load states for TDX module execution context. > And I think for most features SEAM VMCS doesn't set it (I don't check it > though). > > SEAMCALL is kind of a VM exit, and SEAMRET is kind of a VM entry. To > automatically save and restore the host state, what TDX needs are > VM_EXIT_SAVE_XXX, VM_ENTRY_LOAD_XXX, and vmcs guest state for XXX. I'd think we could avoid adding options for configuration that won't be used. If we did have a clobber control interface, matching VMX bits is an interesting idea. I'm a bit torn between wanting to fix the area once and for all with a full solution, and wanting to get this increasingly blocking CPUID bit fix in. I'm leaning towards just do the bit filtering and give the save/restore guidelines. Then we can do a save/restore control later if we find the guidelines are not sufficient.
On 9/2/2026 9:53 PM, Edgecombe, Rick P wrote: > On Wed, 2026-09-02 at 21:39 +0800, Xiaoyao Li wrote: >>> I don't think they have agreed to it yet (Binbin?), but I think it mostly >>> works this way already. Why do you think it is a burden on TDX module >>> developers? >> >> Because it can bring confusion to TDX module developers. VM_EXIT_LOAD_XXX >> control is used by SEAM VMCS to load states for TDX module execution context. >> And I think for most features SEAM VMCS doesn't set it (I don't check it >> though). >> >> SEAMCALL is kind of a VM exit, and SEAMRET is kind of a VM entry. To >> automatically save and restore the host state, what TDX needs are >> VM_EXIT_SAVE_XXX, VM_ENTRY_LOAD_XXX, and vmcs guest state for XXX. > > I'd think we could avoid adding options for configuration that won't be used. If > we did have a clobber control interface, matching VMX bits is an interesting > idea. > > I'm a bit torn between wanting to fix the area once and for all with a full > solution, and wanting to get this increasingly blocking CPUID bit fix in. > I'm > leaning towards just do the bit filtering and give the save/restore guidelines. I agree. I always believe the bit filtering introduced by this series makes KVM safer and it's anyway useful, while how TDX/TDX module save/restore host states can be another separate topic. > Then we can do a save/restore control later if we find the guidelines are not > sufficient.
On Fri, 2026-08-28 at 11:19 +0800, Binbin Wu wrote: > > > > But it isn't a kernel bug, so I'd think to punt on this and let TDX module > > figure out a solution if/when the time comes. Probably various opt-in knobs > > to allow for working around it. > > Do you mean on the old platforms, the according TDX module versions add opt-in > knobs to allow the "effective 1 -> directly configurable" change? Yea. Or the other way. Forcing the old fixed-1 bits to 1 automatically. Could be like a quirk like config thing. User/admin can decide if they want migration flexible design, or backwards compatibility for existing TDs. BUT, we could work out the details later if we think we have options. I think I've convinced myself we have options and can close this one. Agreed? <snip> > > > If FRED support lands in KVM for non-TDX VMs before this > filtering/invalidation is in place, adding FRED to the hardcoded denylist > would work as a temporary solution. This could easily happen. > > But that only holds for a well-behaved userspace VMM. QEMU, for example, only > enables features supported by both KVM and TDX, so if FRED isn't supported in > KVM for non-TDX VMs, QEMU won't enable it for TDX either. > > A malicious userspace VMM, however, can set a feature regardless of KVM's > reported CPU capabilities, and could still cause trouble. Right. It the problem I thought we would need an opt-in to avoid. > > > Do we want it? Where would it get plugged in? > > So if the approach in this series is taken, I think we still need such an opt- > in interface to tell the TDX module that the VMM is now filtering the CPUID > bits so that the TDX module knows that it's safe to report new host state > clobbering features. Yep. Can we think about what it would look like? Easiest would be a bit passed in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. Ok to me, for the simplicity. Could come with a nice comment.
On 8/29/2026 12:58 AM, Edgecombe, Rick P wrote: > On Fri, 2026-08-28 at 11:19 +0800, Binbin Wu wrote: >>> >>> But it isn't a kernel bug, so I'd think to punt on this and let TDX module >>> figure out a solution if/when the time comes. Probably various opt-in knobs >>> to allow for working around it. >> >> Do you mean on the old platforms, the according TDX module versions add opt-in >> knobs to allow the "effective 1 -> directly configurable" change? > > Yea. Or the other way. Forcing the old fixed-1 bits to 1 automatically. Could be > like a quirk like config thing. User/admin can decide if they want migration > flexible design, or backwards compatibility for existing TDs. It's all about the new created TDs on the old platforms. For existing TDs, the runtime TDX module update should not change the shape of the TDs. > BUT, we could work > out the details later if we think we have options. > > I think I've convinced myself we have options and can close this one. Agreed? Agree. > > <snip> >> >> >> If FRED support lands in KVM for non-TDX VMs before this >> filtering/invalidation is in place, adding FRED to the hardcoded denylist >> would work as a temporary solution. > > This could easily happen. > >> >> But that only holds for a well-behaved userspace VMM. QEMU, for example, only >> enables features supported by both KVM and TDX, so if FRED isn't supported in >> KVM for non-TDX VMs, QEMU won't enable it for TDX either. >> >> A malicious userspace VMM, however, can set a feature regardless of KVM's >> reported CPU capabilities, and could still cause trouble. > > Right. It the problem I thought we would need an opt-in to avoid. > >> >>> Do we want it? Where would it get plugged in? >> >> So if the approach in this series is taken, I think we still need such an opt- >> in interface to tell the TDX module that the VMM is now filtering the CPUID >> bits so that the TDX module knows that it's safe to report new host state >> clobbering features. > > Yep. Can we think about what it would look like? Easiest would be a bit passed > in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. Ok to me, > for the simplicity. Could come with a nice comment. The TDX module is initialized before KVM is loaded. I guess the upstream kernel doesn't support out of tree KVM code, so we can assume if the kernel has the code to opt-in the new host state clobbering features, KVM must have implemented the TDX CPUID filtering and validation? Also, do you think it's reasonable to backport this patch series to stable/LTS kernels as an alternative?
On 8/31/2026 1:01 PM, Binbin Wu wrote: >>> So if the approach in this series is taken, I think we still need such an opt- >>> in interface to tell the TDX module that the VMM is now filtering the CPUID >>> bits so that the TDX module knows that it's safe to report new host state >>> clobbering features. >> Yep. Can we think about what it would look like? Easiest would be a bit passed >> in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. Ok to me, >> for the simplicity. Could come with a nice comment. Can you just treat current KVM behavior of allowing userspace to enable any configurable bits as the bug of KVM and backport this series as Binbin suggested below? Instead of introducing more opt-in knobs. > The TDX module is initialized before KVM is loaded. I guess the upstream kernel > doesn't support out of tree KVM code, so we can assume if the kernel has the code to > opt-in the new host state clobbering features, KVM must have implemented the TDX > CPUID filtering and validation? > > Also, do you think it's reasonable to backport this patch series to stable/LTS > kernels as an alternative?
On Tue, 2026-09-01 at 17:42 +0800, Xiaoyao Li wrote: > On 8/31/2026 1:01 PM, Binbin Wu wrote: > > > > So if the approach in this series is taken, I think we still need such > > > > an opt- in interface to tell the TDX module that the VMM is now > > > > filtering the CPUID bits so that the TDX module knows that it's safe to > > > > report new host state clobbering features. > > > Yep. Can we think about what it would look like? Easiest would be a bit > > > passed in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. > > > Ok to me, for the simplicity. Could come with a nice comment. > > Can you just treat current KVM behavior of allowing userspace to enable > any configurable bits as the bug of KVM and backport this series as > Binbin suggested below? Instead of introducing more opt-in knobs. Ok, so if we are agreed on the other branch of the thread, the only big question is: Do we want an opt-in for future clobbering CPUID bits. I think either is ok. I don't love the precedent that we asserted that no new clobber bits could be added without opt-in, and then we would backport changes to allow this anyway. But on pure code, the backport would be simpler in the long term. If you guys are strongly in favor, I can agree. Are we sure no other VMM needs an opt-in, before finalizing it though? Binbin, can flag this to the TDX module team?
On 9/3/2026 12:09 AM, Edgecombe, Rick P wrote: > On Tue, 2026-09-01 at 17:42 +0800, Xiaoyao Li wrote: >> On 8/31/2026 1:01 PM, Binbin Wu wrote: >>>>> So if the approach in this series is taken, I think we still need such >>>>> an opt- in interface to tell the TDX module that the VMM is now >>>>> filtering the CPUID bits so that the TDX module knows that it's safe to >>>>> report new host state clobbering features. >>>> Yep. Can we think about what it would look like? Easiest would be a bit >>>> passed in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. >>>> Ok to me, for the simplicity. Could come with a nice comment. >> >> Can you just treat current KVM behavior of allowing userspace to enable >> any configurable bits as the bug of KVM and backport this series as >> Binbin suggested below? Instead of introducing more opt-in knobs. > > Ok, so if we are agreed on the other branch of the thread, the only big question > is: Do we want an opt-in for future clobbering CPUID bits. > > I think either is ok. I don't love the precedent that we asserted that no new > clobber bits could be added without opt-in, and then we would backport changes > to allow this anyway. But on pure code, the backport would be simpler in the > long term. If you guys are strongly in favor, I can agree. Had a discussion with Rick off list. We can defer the opt-in design or backport decision because the patches will still be correct and an improvement in any case.
On 9/3/2026 12:09 AM, Edgecombe, Rick P wrote: > On Tue, 2026-09-01 at 17:42 +0800, Xiaoyao Li wrote: >> On 8/31/2026 1:01 PM, Binbin Wu wrote: >>>>> So if the approach in this series is taken, I think we still need such >>>>> an opt- in interface to tell the TDX module that the VMM is now >>>>> filtering the CPUID bits so that the TDX module knows that it's safe to >>>>> report new host state clobbering features. >>>> Yep. Can we think about what it would look like? Easiest would be a bit >>>> passed in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. >>>> Ok to me, for the simplicity. Could come with a nice comment. >> >> Can you just treat current KVM behavior of allowing userspace to enable >> any configurable bits as the bug of KVM and backport this series as >> Binbin suggested below? Instead of introducing more opt-in knobs. > > Ok, so if we are agreed on the other branch of the thread, the only big question > is: Do we want an opt-in for future clobbering CPUID bits. > > I think either is ok. I don't love the precedent that we asserted that no new > clobber bits could be added without opt-in, and then we would backport changes > to allow this anyway. But on pure code, the backport would be simpler in the > long term. If you guys are strongly in favor, I can agree. > > Are we sure no other VMM needs an opt-in, before finalizing it though? Binbin, > can flag this to the TDX module team? AFAIK, no other VMM complained about the host clobbering issue. I will check it with the TDX module team.
On 9/1/2026 5:42 PM, Xiaoyao Li wrote: > On 8/31/2026 1:01 PM, Binbin Wu wrote: >>>> So if the approach in this series is taken, I think we still need >>>> such an opt- >>>> in interface to tell the TDX module that the VMM is now filtering >>>> the CPUID >>>> bits so that the TDX module knows that it's safe to report new host >>>> state >>>> clobbering features. >>> Yep. Can we think about what it would look like? Easiest would be a >>> bit passed >>> in TDH_SYS_CONFIG. But then arch/x86 is saying how KVM will behave. >>> Ok to me, >>> for the simplicity. Could come with a nice comment. > > Can you just treat current KVM behavior of allowing userspace to enable well, I meant "can we"... > any configurable bits as the bug of KVM and backport this series as > Binbin suggested below? Instead of introducing more opt-in knobs. > >> The TDX module is initialized before KVM is loaded. I guess the >> upstream kernel >> doesn't support out of tree KVM code, so we can assume if the kernel >> has the code to >> opt-in the new host state clobbering features, KVM must have >> implemented the TDX >> CPUID filtering and validation? >> >> Also, do you think it's reasonable to backport this patch series to >> stable/LTS >> kernels as an alternative? >
Hi Binbin, On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote: > > Specifically, this series builds a KVM-side allowlist of supported TDX > directly configurable CPUID bits to: > - Filter KVM_TDX_CAPABILITIES > Replace the hardcoded denylist to only report configurable bits that > KVM explicitly supports. > - Validate KVM_TDX_INIT_VM > Reject any configurable bit that the TDX module allows but KVM does > not support, as well as CPUID entries with an unexpected subleaf. Today's denylist only rejects TSX and WAITPKG. Everything else is allowed. Obviously, the TDX module allows directly setting virtual CPUID values only for a subset of CPUID leaves, not all of them. So "everything else" above is that subset minus TSX and WAITPKG. My question is: is there a feature in that "everything else" that causes host state clobbering today? In other words, does this patch only build the infrastructure for addressing future clobbering issues, or does it also fix a specific bug? Thanks!
On 9/8/2026 5:42 PM, Artem Bityutskiy wrote: > Hi Binbin, > > On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote: >> >> Specifically, this series builds a KVM-side allowlist of supported TDX >> directly configurable CPUID bits to: >> - Filter KVM_TDX_CAPABILITIES >> Replace the hardcoded denylist to only report configurable bits that >> KVM explicitly supports. >> - Validate KVM_TDX_INIT_VM >> Reject any configurable bit that the TDX module allows but KVM does >> not support, as well as CPUID entries with an unexpected subleaf. > > Today's denylist only rejects TSX and WAITPKG. Everything else is allowed. > > Obviously, the TDX module allows directly setting virtual CPUID values > only for a subset of CPUID leaves, not all of them. So "everything else" > above is that subset minus TSX and WAITPKG. > > My question is: is there a feature in that "everything else" that > causes host state clobbering today? > > In other words, does this patch only build the infrastructure for > addressing future clobbering issues, or does it also fix a specific > bug? The Denylist works fine today. But the TDX module evolves. There are new host state clobbering features coming, e.g. FRED. The Denylist based solution is not a clean solution: - It couples the feature enabling for normal VMs with TDX tightly. - Each time a new feature is added in the denylist, it needs to be backported to old KVM versions > > Thanks!
**Disclaimer**: I am new to KVM and TDX, still learning, let me know if
some of my comments are off.
It took me several days digging through docs and code to understand what is
going on here, so let me summarize my understanding below and please correct
me where I am wrong.
On Thu, 2026-08-27 at 11:18 +0800, Binbin Wu wrote:
> Hi,
>
> The purpose of this patch series is to prevent userspace from enabling
> host state clobbering features that KVM does not support for TDX. A host
> state clobbering feature exposed on a new TDX module/platform can corrupt
> host state if KVM does not explicitly save and restore the related MSR(s)
> across host/guest transitions. If such a feature is blindly exposed to
> and used by a TD, the host will behave unexpectedly.
>
> Except for a few fixed-1 bits required for basic TDX support, host state
> clobbering features are either directly configurable or gated by TD
> ATTRIBUTES/XFAM. So an allowlist covering only the directly configurable
> CPUID bits, plus the corresponding filtering and validation, is sufficient
> to serve the purpose while keeping the code footprint small.
So the general principle is:
- KVM is responsible for protecting the host state from being clobbered
by the TD. Different approaches can be taken here.
- Saving and restoring the host state within KVM itself.
- For a subset of registers, relying on the TDX module to save and
restore the host state.
- Disallowing certain TD features altogether if they pose a risk to
host state.
- TDX module is responsible for saving and restoring TD state.
We have 2 boundaries: VMM <-> TDX module and TDX module <-> TD. Only the
first one is relevant to host state clobbering.
Approach
========
Your patch-set adds an explicit allowlist to KVM: a TD may only use a
feature if every register it needs is either saved/restored by KVM
itself, or is preserved across TDH.VP.ENTER by the TDX module or HW.
I think this approach is safe and sound: each feature is checked before
it is added to the allowlist. No surprises.
A CPU feature is associated with a group of registers, e.g. WAITPKG is
the IA32_UMWAIT_CONTROL MSR. So instead of trying to control individual
register access by TD, KVM controls which features are exposed to the
TD.
CPUID is the actual mechanism for turning CPU features on/off for a TD.
This patch works because the TDX module does not let a TD access the
registers behind a CPU feature unless that feature is exposed via the
virtual CPUID. An attempt to do so results in #UD or #GP.
Not every configurable CPUID field gates some registers, though. Some
are pure enumeration, e.g. family/model/stepping and cache parameters.
So the approach is for KVM to look at what features are enabled in the
virtualized CPUID, and reject the unsafe ones.
Details
=======
TDX module enables TD features only at TD build time, via `TD_PARAMS`
structure of the `TDH.MNG.INIT` seamcall. This happens in the
KVM_TDX_INIT_VM ioctl, after which the feature set cannot be widened.
Migration cannot change them either: the `TDH.IMPORT.STATE.IMMUTABLE`
seamcall imports the source TD's configuration state as-is.
The virtual CPUID values are built based on what the TDX module
supports, what the host supports, and what userspace passed in via the
KVM_TDX_INIT_VM ioctl.
The ioctl carries 'ATTRIBUTES', 'XFAM', and 'CPUID_CONFIG', all fields
of 'TD_PARAMS'. ATTRIBUTES and XFAM turn features on/off and shape
CPUID, but KVM only allows bits it already knows about, so there is no
clobbering risk there.
CPUID_CONFIG is different: it lets userspace set CPUID leaf values
directly (but not for any CPUID, only for pieces of it the TDX module
allows, and KVM enumerates them via the KVM_TDX_CAPABILITIES ioctl).
That is the problematic part. KVM does not check its contents, beyond a
denylist that only filters out TSX and WAITPKG features. Everything
else is allowed.
So your patch-set basically kicks out the current denylist mechanism
and replaces it with an allowlist of directly configurable CPUID leaves
and bits.
There is a long list of CPUID pieces you allowed. I wanted to
acknowledge that it must have been quite an effort to go through all of
them and determine which ones are safe to allow.
Did I understand your work correctly?
> Expected host state clobbering behavior for TDX
> ===============================================
> We also want to call for discussions about the expected host state
> clobbering behavior for TDX here for future features.
>
> For a normal VMX guest, VM entry/exit behavior for a given piece of CPU
> state is architecturally defined: state is either switched by hardware via
> VMCS host/guest fields, or left as the guest value on VM exit and managed
> by KVM in software.
>
> For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the
> TDX module does with a given piece of host state is defined by the TDX
> module ABI rather than by the x86 architecture.
I think I found this contract: the TDH.VP.ENTER definition in the ABI
spec, section "CPU State Preservation Following a Successful TD Entry
and a TD Exit". It refers to a table that lists the MSRs whose value
may not be preserved across TD entry and exit, with the condition for
each, e.g.
IA32_PL0_SSP Init(XFAM[11] | XFAM[12])
IA32_UMWAIT_CONTROL Init(virt. CPUID(7,0).ECX[5])
This is from an older version of the TDX ABI specification. I could not
find 'msr_preservation.pdf' published. But I assume it is published.
Anyway, seems to be a clear contract to me.
>
> What we would like to align on is the expected baseline behavior of
> TDH.VP.ENTER for future features. The proposal is to have TDX simply
> match VMX behavior, i.e. on return from TDH.VP.ENTER, state that VMX would
> restore from the VMCS host fields is restored, and state that VMX would
> leave as the guest value is clobbered. That keeps a single model for VMM,
> and means enabling a new feature for TDs requires the same work flow as
> enabling it for VMX.
Let me restate this to check I follow.
For a VMX guest, it is the VMCS host-state area that is used for
restoring host state. KVM writes it in advance, hardware restores it
upon VM exit.
Most of it is loaded unconditionally: control registers, RSP, RIP,
SYSENTER MSRs and some more. But there are some MSRs that
are restored only if KVM configures the corresponding control bits.
Example: IA32_PAT, IA32_EFER. If the control is clear, the MSR keeps
the guest value upon VM exit.
In KVM some of those controls are set once, e.g. for IA32_PAT. Others
are toggled at run time, e.g. "load IA32_EFER" and "load
IA32_PERF_GLOBAL_CTRL". So there is dynamicity there.
So the proposal is that TDH.VP.ENTER should draw the same line: whatever
VMX would reload from the host-state area is preserved, whatever VMX
leaves as the guest value is clobbered and KVM handles it. Is that
right?
So with TDX module there is a twist. First of all, my understanding is
that the SEAM VMCS (which controls VMM <-> TDX module state
save/restore) is configured when the TDX module is loaded and stays
fixed after that. So no dynamicity there.
Second, the SDM says SEAMCALL behaves like an SMM VM exit and SEAMRET
like a VM entry returning from SMM (SDM 35.1.1, 35.1.2), and SMM VM
exits save state into the guest-state area (SDM 34.15.2.2).
For reference: An SMM VM exit is a VM exit that begins outside SMM
and that ends in SMM.
So if I understand correctly, for VMM <-> TDX module switches, the
hardware uses the VMCS guest-state area, not the host-state area.
Not like in VMX guest case. On TDH.VP.ENTER it saves the VMM state
to the guest-state area, and on return it restores the VMM state.
IOW, looks like there may be significant differences between the VMM
<-> TDX module state handling and VMX VM entries and exits.
So maybe the way to go is to just follow 'msr_preservation.pdf' and
adjust the allowlist? I find this approach safe and acceptable.
Artem.
On Tue, 2026-09-08 at 23:30 +0300, Artem Bityutskiy wrote: > **Disclaimer**: I am new to KVM and TDX, still learning, let me know if > some of my comments are off. > > ... > > So maybe the way to go is to just follow 'msr_preservation.pdf' and > adjust the allowlist? I find this approach safe and acceptable. We are kind of discussing what recommendations we should give about how msr_preservation.pdf should be defined for new CPUID bit based features. So saying to follow msr_preservation.pdf is self referential. Again, please do not treat the TDX specs as something to be handed down and "followed". I think this is something to get used to for TDX. I mean, upstream never wants to adapt to platform arch that fits awkwardly, but it's even tougher to swallow when the arch is mostly SW defined. And further, the people working on the TDX arch want to hear such requirements from upstream. Here, the thing to discuss is how TDX should define new features that will clobber host state (e.g. bits that would appear in msr_preservation.pdf as not being preserved). There have been several PUCK discussions on the problem this series is tackling, and actually several attempts to solve the problem during the base series. More recently a host clobbering control was proposed that attempted to make it safe, but it was not accepted. That proposal brought up the topic of whether having select states clobbered was actually an unproven optimization. Now that we are moving back to an allow list type solution, what guidance should we give on this other surfaced topic. Since TDX shares some save/restore logic with normal VMs, we should have it work well with that code. So forgetting about the performance optimization question, how to have it work in a sensible way with the shared code paths.
On Tue, 2026-09-08 at 22:31 +0000, Edgecombe, Rick P wrote:
> We are kind of discussing what recommendations we should give about how
> msr_preservation.pdf should be defined for new CPUID bit based features. So
> saying to follow msr_preservation.pdf is self referential.
OK, thanks for elaborating. The e-mail was vague about that.
> Again, please do not treat the TDX specs as something to be handed down and
> "followed". I think this is something to get used to for TDX. I mean, upstream
> never wants to adapt to platform arch that fits awkwardly, but it's even tougher
> to swallow when the arch is mostly SW defined. And further, the people working
> on the TDX arch want to hear such requirements from upstream.
Agreed, that matches my understanding. Reminders are useful in general,
but this was not that case.
> Here, the thing to discuss is how TDX should define new features that will
> clobber host state (e.g. bits that would appear in msr_preservation.pdf as not
> being preserved).
>
> There have been several PUCK discussions on the problem this series is tackling,
> and actually several attempts to solve the problem during the base series. More
> recently a host clobbering control was proposed that attempted to make it safe,
> but it was not accepted. That proposal brought up the topic of whether having
> select states clobbered was actually an unproven optimization.
>
> Now that we are moving back to an allow list type solution, what guidance should
> we give on this other surfaced topic. Since TDX shares some save/restore logic
> with normal VMs, we should have it work well with that code. So forgetting about
> the performance optimization question, how to have it work in a sensible way
> with the shared code paths.
Rick,
In general, the TDX case and the VMX case behaving the same way is
best. I thought that consistency is always obviously a good thing,
but let me acknowledge it explicitly.
I was trying to dig deeper into the proposal and analyze it.
<cite>
The proposal is to have TDX simply match VMX behavior, i.e. on return
from TDH.VP.ENTER, state that VMX would restore from the VMCS host
fields is restored, and state that VMX would leave as the guest value
is clobbered. That keeps a single model for VMM, and means enabling a
new feature for TDs requires the same work flow as enabling it for VMX.
</cite>
My point was that hardware behaves differently for VMX guests and for
the VMM<->TDX. This is not about following "boss specs", it is what the
SDM describes.
For VMX, the VMCS host-state area is loaded by hardware on VM exit
(SDM 27.5). For SEAMCALL and SEAMRET, the SDM seems to say they operate
like an SMM VM exit and a VM entry returning from SMM (SDM 35.1), and
those save state into the guest-state area of the transfer VMCS
(SDM 34.15.2.2).
VMX:
VM entry = load guest state from guest-state area
VM exit = save guest state into guest-state area,
load host state from host-state area
TDX:
SEAMCALL = save host state into SEAM VMCS guest-state area,
load module state from SEAM VMCS host-state area
SEAMRET = restore host state from SEAM VMCS guest-state area
I may be reading the SDM wrong, let me know.
So there are differences, and I was hoping to:
- Be corrected if I misinterpret the SDM and how things work.
- Get comments on whether the proposal took this into account.
- Get comments on how this affects, or does not affect, the proposal.
Thanks, Artem.
On 9/9/2026 2:52 PM, Artem Bityutskiy wrote: > On Tue, 2026-09-08 at 22:31 +0000, Edgecombe, Rick P wrote: >> We are kind of discussing what recommendations we should give about how >> msr_preservation.pdf should be defined for new CPUID bit based features. So >> saying to follow msr_preservation.pdf is self referential. > > OK, thanks for elaborating. The e-mail was vague about that. > >> Again, please do not treat the TDX specs as something to be handed down and >> "followed". I think this is something to get used to for TDX. I mean, upstream >> never wants to adapt to platform arch that fits awkwardly, but it's even tougher >> to swallow when the arch is mostly SW defined. And further, the people working >> on the TDX arch want to hear such requirements from upstream. > > Agreed, that matches my understanding. Reminders are useful in general, > but this was not that case. > >> Here, the thing to discuss is how TDX should define new features that will >> clobber host state (e.g. bits that would appear in msr_preservation.pdf as not >> being preserved). >> >> There have been several PUCK discussions on the problem this series is tackling, >> and actually several attempts to solve the problem during the base series. More >> recently a host clobbering control was proposed that attempted to make it safe, >> but it was not accepted. That proposal brought up the topic of whether having >> select states clobbered was actually an unproven optimization. >> >> Now that we are moving back to an allow list type solution, what guidance should >> we give on this other surfaced topic. Since TDX shares some save/restore logic >> with normal VMs, we should have it work well with that code. So forgetting about >> the performance optimization question, how to have it work in a sensible way >> with the shared code paths. > > Rick, > > In general, the TDX case and the VMX case behaving the same way is > best. I thought that consistency is always obviously a good thing, > but let me acknowledge it explicitly. > > I was trying to dig deeper into the proposal and analyze it. > > <cite> > The proposal is to have TDX simply match VMX behavior, i.e. on return > from TDH.VP.ENTER, state that VMX would restore from the VMCS host > fields is restored, and state that VMX would leave as the guest value > is clobbered. That keeps a single model for VMM, and means enabling a > new feature for TDs requires the same work flow as enabling it for VMX. > </cite> > > My point was that hardware behaves differently for VMX guests and for > the VMM<->TDX. This is not about following "boss specs", it is what the > SDM describes. > > For VMX, the VMCS host-state area is loaded by hardware on VM exit > (SDM 27.5). For SEAMCALL and SEAMRET, the SDM seems to say they operate > like an SMM VM exit and a VM entry returning from SMM (SDM 35.1), and > those save state into the guest-state area of the transfer VMCS > (SDM 34.15.2.2). > > VMX: > VM entry = load guest state from guest-state area > VM exit = save guest state into guest-state area, > load host state from host-state area > > TDX: > SEAMCALL = save host state into SEAM VMCS guest-state area, > load module state from SEAM VMCS host-state area > SEAMRET = restore host state from SEAM VMCS guest-state area > > I may be reading the SDM wrong, let me know. That's my understanding too. > > So there are differences, and I was hoping to: > - Be corrected if I misinterpret the SDM and how things work. > - Get comments on whether the proposal took this into account. > - Get comments on how this affects, or does not affect, the proposal. For host state clobbering behavior, we cares about the values of the host (VMX root mode) after SEAMRET. When there is a control/field for "load host state from host-state area", I think there are two cases: - If there is the corresponding control/field for "load guest state from guest-state area", the TDX module could leverage it. - If there is no such corresponding control/field for "load guest state from guest-state area", the TDX module could do it in software way to mimic it. So from the view of the VMM, it can have the aligned behavior on host state clobbering behavior.
On Wed, 2026-09-09 at 16:48 +0800, Binbin Wu wrote: > > VMX: > > VM entry = load guest state from guest-state area > > VM exit = save guest state into guest-state area, > > load host state from host-state area > > > > TDX: > > SEAMCALL = save host state into SEAM VMCS guest-state area, > > load module state from SEAM VMCS host-state area > > SEAMRET = restore host state from SEAM VMCS guest-state area > > > > I may be reading the SDM wrong, let me know. > > That's my understanding too. Good, thanks for confirming. > > > > So there are differences, and I was hoping to: > > - Be corrected if I misinterpret the SDM and how things work. > > - Get comments on whether the proposal took this into account. > > - Get comments on how this affects, or does not affect, the proposal. > > For host state clobbering behavior, we cares about the values of the host (VMX > root mode) after SEAMRET. > > When there is a control/field for "load host state from host-state area", I > think there are two cases: > - If there is the corresponding control/field for "load guest state from > guest-state area", the TDX module could leverage it. > - If there is no such corresponding control/field for "load guest state from > guest-state area", the TDX module could do it in software way to mimic it. > > So from the view of the VMM, it can have the aligned behavior on host state > clobbering behavior. Now I see what you mean: make msr_preservation.pdf follow the same rule as the VMX host-state restore, and let the TDX module help where HW behaves differently (call this SW restore vs HW restore via VMCS). That sounds good to me. My only doubt is whether it can be guaranteed in every case. A HW restore happens after a SW restore. E.g., IA32_DEBUGCTL - HW clears it on VM exit (SDM 30.5.1), so whatever TDX module puts there on the exit path, will be overwritten. Not that this is an issue today, just using this as an example. But I'd guess there would be only few problematic cases (if any). Then you wrote this: <cite> FRED is a useful concrete example. Under VMX, the FRED host state in IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the TDX module is expected to clobber them on TDH.VP.ENTER return. </cite> That one reads as obviously right to me. If VMX and TDX differed in how IA32_FRED_RSP0 and IA32_PL0_SSP are handled, that would be a red flag. Did you go through all the MSRs and check that the VMX and TDX behavior matches today? Thanks!
On 9/9/2026 7:20 PM, Artem Bityutskiy wrote: > On Wed, 2026-09-09 at 16:48 +0800, Binbin Wu wrote: >>> VMX: >>> VM entry = load guest state from guest-state area >>> VM exit = save guest state into guest-state area, >>> load host state from host-state area >>> >>> TDX: >>> SEAMCALL = save host state into SEAM VMCS guest-state area, >>> load module state from SEAM VMCS host-state area >>> SEAMRET = restore host state from SEAM VMCS guest-state area >>> >>> I may be reading the SDM wrong, let me know. >> >> That's my understanding too. > > Good, thanks for confirming. > >>> >>> So there are differences, and I was hoping to: >>> - Be corrected if I misinterpret the SDM and how things work. >>> - Get comments on whether the proposal took this into account. >>> - Get comments on how this affects, or does not affect, the proposal. >> >> For host state clobbering behavior, we cares about the values of the host (VMX >> root mode) after SEAMRET. >> >> When there is a control/field for "load host state from host-state area", I >> think there are two cases: >> - If there is the corresponding control/field for "load guest state from >> guest-state area", the TDX module could leverage it. >> - If there is no such corresponding control/field for "load guest state from >> guest-state area", the TDX module could do it in software way to mimic it. >> >> So from the view of the VMM, it can have the aligned behavior on host state >> clobbering behavior. > > Now I see what you mean: make msr_preservation.pdf follow the same rule > as the VMX host-state restore, and let the TDX module help where HW behaves > differently (call this SW restore vs HW restore via VMCS). > > That sounds good to me. > > My only doubt is whether it can be guaranteed in every case. A HW restore > happens after a SW restore. > > E.g., IA32_DEBUGCTL - HW clears it on VM exit (SDM 30.5.1), so whatever > TDX module puts there on the exit path, will be overwritten. Not that this > is an issue today, just using this as an example. But SEAMRET is actually a VM Entry, I think these special cases are during VM Exit. For a VM entry, the TDX module should be able to set whatever valid values for host. > > But I'd guess there would be only few problematic cases (if any). > > Then you wrote this: > > <cite> > FRED is a useful concrete example. Under VMX, the FRED host state in > IA32_FRED_CONFIG, IA32_FRED_STKLVLS, IA32_FRED_RSP1-3 and > IA32_FRED_SSP1-3 is covered by the VMCS host-state area, so the TDX module > is expected to restore these MSRs on TDH.VP.ENTER return. IA32_FRED_RSP0 > and IA32_PL0_SSP (a.k.a. IA32_FRED_SSP0) are handled by software, so the > TDX module is expected to clobber them on TDH.VP.ENTER return. > </cite> > > That one reads as obviously right to me. If VMX and TDX differed in how > IA32_FRED_RSP0 and IA32_PL0_SSP are handled, that would be a red flag. > > Did you go through all the MSRs and check that the VMX and TDX behavior > matches today? Not yet. This topic is put in the cover letter for discussions. And the patch series itself doesn't depend on conclusion of the discussions. > > Thanks!
On 9/9/2026 4:30 AM, Artem Bityutskiy wrote: > **Disclaimer**: I am new to KVM and TDX, still learning, let me know if > some of my comments are off. > > It took me several days digging through docs and code to understand what is > going on here, Thanks! [...] > > CPUID_CONFIG is different: it lets userspace set CPUID leaf values > directly (but not for any CPUID, only for pieces of it the TDX module > allows, and KVM enumerates them via the KVM_TDX_CAPABILITIES ioctl). > > That is the problematic part. KVM does not check its contents, beyond a > denylist that only filters out TSX and WAITPKG features. Everything > else is allowed. > > So your patch-set basically kicks out the current denylist mechanism > and replaces it with an allowlist of directly configurable CPUID leaves > and bits. > [...] > > Did I understand your work correctly? Yes, you understand it correctly. > >> Expected host state clobbering behavior for TDX >> =============================================== >> We also want to call for discussions about the expected host state >> clobbering behavior for TDX here for future features. >> >> For a normal VMX guest, VM entry/exit behavior for a given piece of CPU >> state is architecturally defined: state is either switched by hardware via >> VMCS host/guest fields, or left as the guest value on VM exit and managed >> by KVM in software. >> >> For TDs, the host/guest transition goes through TDH.VP.ENTER, and what the >> TDX module does with a given piece of host state is defined by the TDX >> module ABI rather than by the x86 architecture. > > I think I found this contract: the TDH.VP.ENTER definition in the ABI > spec, section "CPU State Preservation Following a Successful TD Entry > and a TD Exit". It refers to a table that lists the MSRs whose value > may not be preserved across TD entry and exit, with the condition for > each, e.g. > > IA32_PL0_SSP Init(XFAM[11] | XFAM[12]) > IA32_UMWAIT_CONTROL Init(virt. CPUID(7,0).ECX[5]) > > This is from an older version of the TDX ABI specification. I could not > find 'msr_preservation.pdf' published. But I assume it is published. You can find these tables from the "Intel TDX Module ABI Definitions" part on https://www.intel.com/content/www/us/en/developer/tools/trust-domain-extensions/documentation.html [...] > So maybe the way to go is to just follow 'msr_preservation.pdf' and > adjust the allowlist? I find this approach safe and acceptable. > Rick explained the long history: https://lore.kernel.org/all/58c185c82658819454a9950f37c6424226a098bb.camel@intel.com/ The Denylist based solution is not a clean solution: - It couples the feature enabling for normal VMs with TDX tightly. - Each time a new feature is added in the denylist, it needs to be backported to old KVM versions > Artem.
On 9/9/2026 7:54 AM, Binbin Wu wrote: > >> So maybe the way to go is to just follow 'msr_preservation.pdf' and >> adjust the allowlist? I find this approach safe and acceptable. >> > > Rick explained the long history: > https://lore.kernel.org/all/58c185c82658819454a9950f37c6424226a098bb.camel@intel.com/ > > The Denylist based solution is not a clean solution: > - It couples the feature enabling for normal VMs with TDX tightly. > - Each time a new feature is added in the denylist, it needs to be backported to old KVM versions > Please ignore this part since I replied to the wrong thread. > >> Artem. >
© 2016 - 2026 Red Hat, Inc.