From nobody Fri Jul 24 21:53:01 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7E3C742BC22 for ; Thu, 23 Jul 2026 09:44:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784799869; cv=none; b=cZUXx+7mu8l8Ma8nj1g0Gf1YE1v5erSSvLzzIQnXIv+MNsRvfGClvfO66BYXUC1Jp8CIwbsVrvY5dujbv3C739NJ65VqElSTu7KLLX3I0aosKJ1keqPwzw0qBn8oP6LMRZ8aSloqgfR7ySWuQLqgB/vev4g/u1JWwW6e/wVUhpE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784799869; c=relaxed/simple; bh=8LcfHWlPiRLeBmOzU9vGGv4EbjPyHeH3olbOBNC2Kx8=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=PTHsAzH1t5UlqcNny4cvf8WJDnZJbDNp763hrpF23R4yzzJ/irk6nGA0B16VitCbKnYTsVnlQwqJvPJ15PnW8Nzjy6DODIStjzLc8L+VSrGapxCZ0ucpq+pysuO93l9RZxodlB75rYHHrVnwKhNEWcwjRbFMS/q+psZ0U2PT3qk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=P5zBVM/1; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=k7Z3m8nb; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="P5zBVM/1"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="k7Z3m8nb" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784799865; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding; bh=tD9r13MXqpkctapd3kVKw58/yj0JDaoTRqy5Ai0sP+c=; b=P5zBVM/1gHG1vSXMKrAp6TzivQ+k3yKVIdhyF5ikeag0nt/NXWVDMb0D6UeWtyK3ZByoZW 7CNI1SDuSnuKs4az6uHYmhPYuhoLbL+Vaz3Ll2gbZxpo2gTyJxobOVqBLKL5N6+SOECXym hjSTdkEAcli5prS32JlBBcHu0X6szYU= Received: from mail-wr1-f71.google.com (mail-wr1-f71.google.com [209.85.221.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-662-VID7wlDAOcGoyGyzYE-OKg-1; Thu, 23 Jul 2026 05:44:23 -0400 X-MC-Unique: VID7wlDAOcGoyGyzYE-OKg-1 X-Mimecast-MFC-AGG-ID: VID7wlDAOcGoyGyzYE-OKg_1784799862 Received: by mail-wr1-f71.google.com with SMTP id ffacd0b85a97d-47f83416551so450786f8f.3 for ; Thu, 23 Jul 2026 02:44:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1784799862; x=1785404662; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=tD9r13MXqpkctapd3kVKw58/yj0JDaoTRqy5Ai0sP+c=; b=k7Z3m8nbcXdWjG785v86xAH7y5kNFPqozIM2mJ38d6xO5UOr9jZooIwYTFMf2Ut3x2 2cM4WDQW/s6r5mEeEo4BgaY1LiMhooQccc8DjAtr6Wqrm2nO+kXdWhyLsXx/bNPVa6jg maIs0/KveURDZPr1F/LV/zDhKEqSU/DKUH1ydNkb1TCJkUZ3r5y0ZfKEschXgORDIaHL y843Qa4UEzvbSP26QfiMtaVu3HclekmeBBzSsr2R5M1LPRHRrdhNrXDGwMm3P0E5X3Gb f253X7hW6qYUhoc8RPvSsywX0P+5k/OtisvALAleQo06CvlAvFzFzyDbKH/EVSND+Fbv O3VA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784799862; x=1785404662; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tD9r13MXqpkctapd3kVKw58/yj0JDaoTRqy5Ai0sP+c=; b=PArHq+y+jOCS5CfY88GkHXd20BXy9NG7wfYG2/8rXv29TySdtQFFjJV8onyNQjs/VL yN77vVj0sT8SEQ6mP08sPk1CKT0zMDyDS8zovNF7nweiKDscvCI8P7Xj5gLOVfYY+jNc bVCffqBsV+nn1LfEhUzdB8x6oSJE8+5Hb5U4VUUvJdNDPgPzNqeUrIhQ9uKjUHDRvEdp 5jR9aUMXcf8fVTYobhkEMAzi+CdwTIf0qqmLlCWaNId3Obeof+XwdNwT2YkP1g/DeyAO MS9JckFmWoFKK++9+HDlaayvl9+iV61nIKp2a4xiiBt9FNUGKnXu1FtpAvRHqyewhaX/ TYtg== X-Gm-Message-State: AOJu0Yz5I4ljJVqKHe4d1gK+Vv6FXJsdGj26HMAPeIhnuOgWn2deoMqU eDrGHbDsKk5hZtp3Sn/j6TGyjOcoQIlTnkaqDJG8kC0LPh3v3wFkxSu0qhaj6IfIn9mYkbxXzIJ S1GogrQIXxuY3CGtrd2PSFMUEs+vZXw9gob53EpT4j8B8y0hjrRGeUgknP4T8rpyA16W7NEL0xR 0092/YExqyD4fMi7OK5eAR38uuDjE3Nw0zOORNvIWF49B74ikF0w== X-Gm-Gg: AR+sD10Q4g1obEpWnRpN2PqdR4aUOs1OJ+hm5zUMtXhXXfOAqf08cA7iSYSDuz9IjRW SHXAvvnE9QDLMqbNewRnnlFd2tZKi3Y6olmcL9vMjIT+AUcxI7eX9whyRxAcjPdmgM+bi3C9alb 12+b5bZe2DO8ESc6n88ScHdL0qTPtRhOD8QOprt5tAg6pWgJAyh0iInlmdasr+l8c6f4TMTYvND URFhJpm41ZFyw4ossNqgiZeArQa3HjHbr7OMFpKAna9LgJRQuErPgVWqzGLASsF6H+F0Eyh0tFt cbpR1rD4800ALysLCJQUx2Ez9VkweW27mu5Ujzf/6nDjDEowxp8fJobDTVi7musivBrYZhEJnuc JdXcm7mvq8Wr9kw+Go9Mw+oeqVQhRpQ3g8BWSx5hofgnDjAk1bkd5kjP5RfETP9C35/uTeCfino 3jfqAL X-Received: by 2002:a05:6000:27d3:b0:47f:9266:9bde with SMTP id ffacd0b85a97d-47f9266a0b3mr754374f8f.4.1784799861966; Thu, 23 Jul 2026 02:44:21 -0700 (PDT) X-Received: by 2002:a05:6000:27d3:b0:47f:9266:9bde with SMTP id ffacd0b85a97d-47f9266a0b3mr754332f8f.4.1784799861338; Thu, 23 Jul 2026 02:44:21 -0700 (PDT) Received: from [192.168.10.48] ([151.49.94.110]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-47f85bdac28sm13099297f8f.16.2026.07.23.02.44.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 23 Jul 2026 02:44:20 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Vitaly Kuznetsov , Alexander Lougovski Subject: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled Date: Thu, 23 Jul 2026 11:44:19 +0200 Message-ID: <20260723094419.630204-1-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Red Hat is seeing multiple reports of Windows memory corruptions (and consequent BSODs) with hv-tlbflush=3Don, on AMD processors only. The crashes, while extremely rare, happen even with a stock configuration, but with Driver Verifier enabled they can be detected after approximately 200 VM hours. In particular, Alexander Lougovski measured the following: - on AMD Turin, 15 crashes in 3300 VM hours - on AMD Milan, 2 crashes in 500 VM hours (there are fewer hours here due to the host being smaller) - on Intel Sapphire Rapids, 0 crashes in 8000 VM hours - on AMD Turin with full TLB flush (not exactly this patch but similar), no crashes in ~2 weeks of run time which should also be ~7000 VM hours For Turin, the microcode version was 0x0b002162, which (assuming this is the same issue) should not be affected by the problem listed in https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-mach= ines-running-on-amd.html; on the other hand that problem should not apply to earlier processors. AMD has not provided any information or analysis yet, and when we asked we didn't know yet that it reproduced on Milan as well. As to the workload, Alexander threw more or less everything at the same time at the VM: - a full Windows Defender scan every 30 minutes - a disk I/O job - a loop doing repeated mmap of system files (mostly to hope that it triggers some consistency check in the Windows memory manager) - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries (75% write/25% read) via sqlcmd Driver Verifier is able to detect BSODs more or less at the same time as the pages are freed. They mostly happen in the Windows Defender filter driver, but occasionally also in the networking stack (e.g., afd.sys) or elsewhere in the filesystem stack (e.g., fltmgr.sys). The flush is issued from kvm_hv_vcpu_flush_tlb(), which receives the cross-CPU requests from the Hyper-V TLB flush hypercalls via a kfifo and is invoked by the KVM_REQ_HV_TLB_FLUSH request. The mechanism is the same for both Intel and AMD, and the handler for both vendors is a simple INVVPID(ADDR)/INVLPGA instruction. Because the request is handled on the destination CPU, there is a question of what happens if the VM is migrated across physical CPUs. In that case, the INVLPGA instruction would use a stale svm->vmcb->control.asid; but if anything that might do an *unnecessary* flush (on an asid that's being used for another VM) and then pre_svm_run() would force a full TLB rebuild. So, for lack of better ideas, this patch forces a full ASID bump in svm_flush_tlb_gva(). To avoid paying the price on Intel and also to avoid unnecessary loops on AMD, the flush_tlb_gva op now returns whether it did a full flush or not; kvm_hv_vcpu_flush_tlb() takes note and exits its loops immediately. While there is an obvious performance impact, about half of the benefit from Hyper-V tlbflush is preserved (10% vs. 20% on the SQL Server workload). kvm_mmu_invalidate_addr() is the only other caller of the flush_tlb_gva op. The change would have a performance impact on every intercepted INVLPG and, for nested SVM, on every L1 INVLPGA. For INVLPGA specifically, this covers the same suspected issue but for nested hypervisors, so it is correct to apply the workaround; for INVLPG on shadow paging, instead, the impact would be stronger and, due to lack of data, for now the use of INVLPGA is left in place in svm_flush_tlb_gva(). Analyzed-by: Vitaly Kuznetsov Analyzed-by: Alexander Lougovski Signed-off-by: Paolo Bonzini --- arch/x86/include/asm/kvm_host.h | 2 +- arch/x86/kvm/hyperv.c | 7 ++++--- arch/x86/kvm/mmu/mmu.c | 2 +- arch/x86/kvm/svm/svm.c | 27 ++++++++++++++++++++------- arch/x86/kvm/vmx/main.c | 4 ++-- arch/x86/kvm/vmx/vmx.c | 2 +- arch/x86/kvm/vmx/x86_ops.h | 2 +- 7 files changed, 30 insertions(+), 16 deletions(-) diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_hos= t.h index b517257a6315..eca04d4b974e 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -1751,7 +1751,7 @@ struct kvm_x86_ops { * Can potentially get non-canonical addresses through INVLPGs, which * the implementation may choose to ignore if appropriate. */ - void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr); + void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr, bool *full); =20 /* * Flush any TLB entries created by the guest. Like tlb_flush_gva(), diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c index 1ee0d23f8949..5ec1bf28a195 100644 --- a/arch/x86/kvm/hyperv.c +++ b/arch/x86/kvm/hyperv.c @@ -1974,6 +1974,7 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu) u64 entries[KVM_HV_TLB_FLUSH_FIFO_SIZE]; int i, j, count; gva_t gva; + bool full =3D false; =20 if (!tdp_enabled || !hv_vcpu) return -EINVAL; @@ -1982,7 +1983,7 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu) =20 count =3D kfifo_out(&tlb_flush_fifo->entries, entries, KVM_HV_TLB_FLUSH_F= IFO_SIZE); =20 - for (i =3D 0; i < count; i++) { + for (i =3D 0; i < count && !full; i++) { if (entries[i] =3D=3D KVM_HV_TLB_FLUSHALL_ENTRY) goto out_flush_all; =20 @@ -1991,11 +1992,11 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu) * pages to flush. */ gva =3D entries[i] & PAGE_MASK; - for (j =3D 0; j < (entries[i] & ~PAGE_MASK) + 1; j++) { + for (j =3D 0; j < (entries[i] & ~PAGE_MASK) + 1 && !full; j++) { if (is_noncanonical_invlpg_address(gva + j * PAGE_SIZE, vcpu)) continue; =20 - kvm_x86_call(flush_tlb_gva)(vcpu, gva + j * PAGE_SIZE); + kvm_x86_call(flush_tlb_gva)(vcpu, gva + j * PAGE_SIZE, &full); } =20 ++vcpu->stat.tlb_flush; diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 6c13da942bfc..bea5499fc9f1 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -6672,7 +6672,7 @@ void kvm_mmu_invalidate_addr(struct kvm_vcpu *vcpu, s= truct kvm_pagewalk *w, if (is_noncanonical_invlpg_address(addr, vcpu)) return; =20 - kvm_x86_call(flush_tlb_gva)(vcpu, addr); + kvm_x86_call(flush_tlb_gva)(vcpu, addr, NULL); =20 if (tdp_enabled) return; diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c index ef69a51ab27f..0bd63970305b 100644 --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ -4222,13 +4222,6 @@ static void svm_flush_tlb_all(struct kvm_vcpu *vcpu) svm_flush_tlb_asid(vcpu); } =20 -static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva) -{ - struct vcpu_svm *svm =3D to_svm(vcpu); - - invlpga(gva, svm->vmcb->control.asid); -} - static void svm_flush_tlb_guest(struct kvm_vcpu *vcpu) { kvm_register_mark_dirty(vcpu, VCPU_REG_ERAPS); @@ -4236,6 +4229,26 @@ static void svm_flush_tlb_guest(struct kvm_vcpu *vcp= u) svm_flush_tlb_asid(vcpu); } =20 +static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva, bool *full) +{ + struct vcpu_svm *svm =3D to_svm(vcpu); + + /* + * INVLPGA has had errata on Genoa and Turin, and even on older + * generations there were reports of Windows BSODs if INVLPGA + * was used for Hyper-V tlbflush. Use it only for shadow paging + * where it seems to be okay. + */ + if (!npt_enabled) { + invlpga(gva, svm->vmcb->control.asid); + return; + } + + svm_flush_tlb_guest(vcpu); + if (full) + *full =3D true; +} + static inline void sync_cr8_to_lapic(struct kvm_vcpu *vcpu) { struct vcpu_svm *svm =3D to_svm(vcpu); diff --git a/arch/x86/kvm/vmx/main.c b/arch/x86/kvm/vmx/main.c index 83d9921277ea..f204a0fc0a57 100644 --- a/arch/x86/kvm/vmx/main.c +++ b/arch/x86/kvm/vmx/main.c @@ -535,12 +535,12 @@ static void vt_flush_tlb_current(struct kvm_vcpu *vcp= u) vmx_flush_tlb_current(vcpu); } =20 -static void vt_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr) +static void vt_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full) { if (is_td_vcpu(vcpu)) return; =20 - vmx_flush_tlb_gva(vcpu, addr); + vmx_flush_tlb_gva(vcpu, addr, full); } =20 static void vt_flush_tlb_guest(struct kvm_vcpu *vcpu) diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c index 3681d565f177..e27084d30a5e 100644 --- a/arch/x86/kvm/vmx/vmx.c +++ b/arch/x86/kvm/vmx/vmx.c @@ -3374,7 +3374,7 @@ void vmx_flush_tlb_current(struct kvm_vcpu *vcpu) vpid_sync_context(vmx_get_current_vpid(vcpu)); } =20 -void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr) +void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full) { /* * vpid_sync_vcpu_addr() is a nop if vpid=3D=3D0, see the comment in diff --git a/arch/x86/kvm/vmx/x86_ops.h b/arch/x86/kvm/vmx/x86_ops.h index 409858074246..17595d52985c 100644 --- a/arch/x86/kvm/vmx/x86_ops.h +++ b/arch/x86/kvm/vmx/x86_ops.h @@ -82,7 +82,7 @@ void vmx_set_rflags(struct kvm_vcpu *vcpu, unsigned long = rflags); bool vmx_get_if_flag(struct kvm_vcpu *vcpu); void vmx_flush_tlb_all(struct kvm_vcpu *vcpu); void vmx_flush_tlb_current(struct kvm_vcpu *vcpu); -void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr); +void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full); void vmx_flush_tlb_guest(struct kvm_vcpu *vcpu); void vmx_set_interrupt_shadow(struct kvm_vcpu *vcpu, int mask); u32 vmx_get_interrupt_shadow(struct kvm_vcpu *vcpu); --=20 2.55.0