From nobody Tue Sep 29 02:34:37 2026 Received: from mail-pg1-f170.google.com (mail-pg1-f170.google.com [209.85.215.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1538D3368AA for ; Thu, 13 Aug 2026 04:39:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786595983; cv=none; b=WemiMwiGeoxZkq4EaLfXvehiPWsQ2DL40BjlvBI3WMX/Gl6xwgEdQWLAF2/K9bqqHy3xcRosX3tou2szmD2WE3AdMHIkWcwVqlM8F8/carkwgV43PhbO/6HuzqXm6gh2iUWWltq/L/z1s/FGUrqEOBz5YiXNn6yOWbKnBXKo4fk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786595983; c=relaxed/simple; bh=MrA1eUn49JHHtkGSWBVP0LTaU/MxZzymHsQV1rsechk=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=GDLmbgxzrrA6c4u7x7x/vvR0EeSEdCvFmm4ehVVlSdS7yDMCjMZriat8KkCt2zeYJjN+FiFWfS0YAKenHfkPiPuqWjcYOIcA5N2DxbA9kAyN4eEaXOh6M77CnZf2d1lA/f55hDG5HLBuVkyHmD3GVa9DvaX1xFiTqg9V7IX7uoc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=oyXlXK49; arc=none smtp.client-ip=209.85.215.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="oyXlXK49" Received: by mail-pg1-f170.google.com with SMTP id 41be03b00d2f7-cbedb88aa34so301167a12.3 for ; Wed, 12 Aug 2026 21:39:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786595981; x=1787200781; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=QY0PhIaQe303xAQrTg2FMWpE3pS6R1npxD02nSQKkGo=; b=oyXlXK494Z/C6uht5qpC3+QJ3ej6LcpMvgO00eTR9vcmcgZfySI5GuzQy0YwPYr+Wb OtmHz8tYVWAUZkXyXedaIkBxXhWWTzY/iw6ocV1LE/oTpKZNtpeTiVE213Hgt2kpUHrb wO51ZcKRW+4wZh9wMozFYId0u1Uky4SfugYyi9c6sPV1sCZq0VXvrss+N4N1ool1n3PW OwMtma1nSXrDCsHRoQZv8Z3vw0CUAd9dqN1WX6GugWm9WUf6cHiB03131Ikkm259sLxv vmVcLQNCXfAAViPADi7fcI2bOHze88PWIXymoKCgLU6Wiznhf7WTaYXGevNY7coB+dwk LK0A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786595981; x=1787200781; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QY0PhIaQe303xAQrTg2FMWpE3pS6R1npxD02nSQKkGo=; b=VWTGFUYqHKdmmYiXeMXBq6fCs/ouGILJa7P19tDj/9KIAtpNrXysqRcD7gpKcQ9263 xhBUecNfN2Aehgsnqlcs8WKrsYGb73ML2hyiktxthutR7axLimyUikFmYLDkWYrk0Ham HyYPfV8WidGx6FXlScg6YipJKn1gYjLO5i4tvdKxe/R1iEAvS8O+1+G8/miEsWHp4kp7 nshXfD9hKLikYZE+N0MYUCjRE4mxc3peW+xRpPAZn8T+aNhin9YO6rAK//rDOzDUuBl+ Fw1uNBiZWEj8i9+HN/8ujZ2wjvwOywf+53/eDgmjloZBpw86pi8Ma1GKQXy74TmD5/VW 493A== X-Forwarded-Encrypted: i=1; AHgh+RraImzgQBJQ78AVYWnPIy6eu2fuDUTh+pYh5kvFdXZQMrwrnEtJS5nN2quAxhnODwp8guCR38TTTWQlenc=@vger.kernel.org X-Gm-Message-State: AOJu0Yx+Boo2s6Iv/eqPoF61W/dPR3JbMjlAbPr6Km+MOPjeAQLsOYcM yyE9g3CHcSr4bGKMttV//HDZIFii7N6RkFbhNwNK/P8U4RhMqRfJV8Qk X-Gm-Gg: AR+sD12MDgqijMhFjeSnSCGZRWfcvDS+C/w2r6ANxXssrST2ajC5vMcx6oDJdKRYYlG IsorXjqbCRs6AXCHqkSP0ZnUDb+8E6y6+1OhAt5HvqRE2OELH3QiaVqzoaoLa+9yoRvVbUFbcai stQwNfOQqcFh5culVX+P8qZNrhqqGhkIz6x9Q9QlH5B3TrJitJD9shmfPVvCmav8rLg/ysxhbt6 zGHkh4ifTFAzsl3gll4bJKj1F9NdAh8cIlVmMRnIS8mGklfN44Pxm2E2QSOtQGYbkZUJvo7jrDj yNpqIfjjLXZYk48tVQzKcqXvvLzqC7KCKdJsRZlXV2KVaGFEe5t04PvFfAt2Aw4YHPHrgHbIBoU iM7ItnfFa1UfQ7ovap/zTpn0qDrVqQJDZdO91PZXluPHqaE8f4ZVJPjX3EUmMry7+jRIGXJv57E ox+iCrCxCQ8MlGcOpAR46AlD8fJNZtzt/2RyR4u3xDYD6GqZIYlkgkUETcRlfBUj2PbKmZqyFpt u7cRT5IvPv2GOCrTRARkz9HDic= X-Received: by 2002:a05:6a20:549c:b0:3c3:a9ad:a747 with SMTP id adf61e73a8af0-3cc553d8f30mr4537871637.26.1786595981038; Wed, 12 Aug 2026 21:39:41 -0700 (PDT) Received: from CT104.localdomain ([118.223.79.56]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-cbef710b05csm417325a12.15.2026.08.12.21.39.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 12 Aug 2026 21:39:40 -0700 (PDT) From: Jinwoo Lee To: seanjc@google.com, pbonzini@redhat.com Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Jinwoo Lee , stable@vger.kernel.org Subject: [PATCH] KVM: nVMX: Re-arm the vmcs12 pages request if mapping the pages fails Date: Thu, 13 Aug 2026 04:39:32 +0000 Message-ID: <20260813043932.3214460-1-rkskek9254@gmail.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Re-arm KVM_REQ_GET_NESTED_STATE_PAGES when nested_get_vmcs12_pages() fails, so that KVM retries the mapping on the next KVM_RUN instead of resuming L2 with a stale vmcs02. On failure KVM exits to userspace with KVM_EXIT_INTERNAL_ERROR but leaves the vCPU in guest mode with vmcs02 loaded. The request has already been consumed by kvm_check_request() in vcpu_enter_guest(), and nothing re-arms it, so a subsequent KVM_RUN goes straight to VM-Enter. vmcs02's APIC_ACCESS_ADDR, VIRTUAL_APIC_PAGE_ADDR and POSTED_INTR_DESC_ADDR still hold the host physical addresses that were mapped for the previous nested VM-Enter. Those pages have already been unmapped and unpinned by nested_put_vmcs12_pages(), which runs after vmx_switch_vmcs() to vmcs01 and therefore cannot clear the vmcs02 fields, and prepare_vmcs02_early() re-arms the controls that consume them without rewriting the address fields. Hardware then accesses pages that KVM no longer holds a reference to. The SECONDARY_EXEC_VIRTUALIZE_APIC_ACCESSES branch makes this worse by returning before the CPU_BASED_TPR_SHADOW and posted interrupt fallbacks, which would otherwise write INVALID_GPA to VIRTUAL_APIC_PAGE_ADDR and clear PIN_BASED_POSTED_INTR. Commit 671ddc700fd0 ("KVM: nVMX: Don't leak L1 MMIO regions to L2") replaced the "clear the control" fallback with an error return. The intent is right, but it left vmcs02 in a usable state. Re-arming the request closes that; if the mapping keeps failing KVM keeps exiting to userspace, which is noisy but safe. Note this relies on KVM_REQ_GET_NESTED_STATE_PAGES being cleared on nested VM-Exit, which "KVM: nVMX: Ensure KVM_REQ_GET_NESTED_STATE_PAGES is cleared on VM-Exit" makes unconditional. Fixes: 671ddc700fd0 ("KVM: nVMX: Don't leak L1 MMIO regions to L2") Cc: stable@vger.kernel.org Signed-off-by: Jinwoo Lee --- Notes for reviewers, not intended for the commit log. Affected versions: v5.4-rc5 (671ddc700fd0) through v7.2-rc6. Verified that nested_get_vmcs12_pages() and the KVM_REQ_GET_NESTED_STATE_PAGES consumer in vcpu_enter_guest() are unchanged in kvm-x86/next as of 2026-08-13. Disclosure: this was found with AI-assisted code review, so per Documentation/process/security-bugs.rst it is being reported publicly rather than to security@kernel.org. What I verified empirically, on the RSM path with the load_pdptrs() abort: - KVM consumes KVM_REQ_GET_NESTED_STATE_PAGES, the mapping fails, KVM_RUN returns 0 with run->exit_reason left at KVM_EXIT_UNKNOWN, and the vCPU = is still in guest mode. - The request is not re-armed, and the next KVM_RUN VM-Enters L2, which t= hen executes with vmcs02 still naming the previously mapped pages. Confirm= ed deterministically (4/4) with a selftest, plus a control run showing that the same sequence without the poison maps successfully. - With the patch applied the code compiles clean, but I have not been abl= e to boot a patched kernel, so the fix itself is not runtime tested. The selftest fails on an unpatched kernel as expected. What I did not verify: - The APIC-access branch end to end. That is the interesting one, becaus= e it is reachable by L1 alone (point vmcs12->apic_access_addr at an unbacked GPA) and it returns before the CPU_BASED_TPR_SHADOW and posted interrupt fallbacks. This host does not expose SECONDARY_EXEC_VIRTUALIZE_APIC_ACCESSES or PIN_BASED_POSTED_INTR to L1,= so I could only reach the load_pdptrs() abort, which needs userspace to po= ison the PDPTEs through a KVM_GUESTDBG_SINGLESTEP window and is therefore not guest-triggerable on its own. - Whether a stale page is actually reused by the host. The selftest keeps every page allocated as its own guest RAM for the whole run. On whether userspace resumes: QEMU's kvm_cpu_exec() treats KVM_INTERNAL_ERROR_EMULATION as recoverable and returns EXCP_INTERRUPT when kvm_arch_stop_on_emulation_error() is false, which for x86 is the case when= the guest is in protected mode at CPL 3 (target/i386/kvm/kvm.c). It does not re-push nested state on that path. This is from reading qemu.git at 055952c0aa91; I have not run it. A selftest is available. I have not included it here per the reproducer guidance in security-bugs.rst; happy to post it if you want it. arch/x86/kvm/vmx/nested.c | 18 ++++++++++++++++-- 1 file changed, 16 insertions(+), 2 deletions(-) diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c index ddf6df7bee93..9e9bd6c541ba 100644 --- a/arch/x86/kvm/vmx/nested.c +++ b/arch/x86/kvm/vmx/nested.c @@ -3453,7 +3453,7 @@ static bool nested_get_vmcs12_pages(struct kvm_vcpu *= vcpu) * state which can lead to a load of wrong PDPTRs. */ if (CC(!load_pdptrs(vcpu, vcpu->arch.cr3))) - return false; + goto fail; } @@ -3469,7 +3469,7 @@ static bool nested_get_vmcs12_pages(struct kvm_vcpu *= vcpu) vcpu->run->internal.suberror =3D KVM_INTERNAL_ERROR_EMULATION; vcpu->run->internal.ndata =3D 0; - return false; + goto fail; } } @@ -3525,6 +3525,20 @@ static bool nested_get_vmcs12_pages(struct kvm_vcpu = *vcpu) exec_controls_clearbit(vmx, CPU_BASED_USE_MSR_BITMAPS); return true; + +fail: + /* + * Re-arm the request so that KVM retries the mapping instead of running + * L2 with a stale vmcs02. Bailing here leaves the vCPU in guest mode + * with vmcs02 loaded and its APIC-access, virtual-APIC and posted + * interrupt descriptor addresses still pointing at the host pages that + * were mapped for the *previous* nested VM-Enter, which have since been + * unmapped and unpinned by nested_put_vmcs12_pages(). KVM returns to + * userspace without leaving guest mode, so if userspace resumes the + * vCPU, VM-Enter succeeds and hardware accesses those stale HPAs. + */ + kvm_make_request(KVM_REQ_GET_NESTED_STATE_PAGES, vcpu); + return false; } static bool vmx_get_nested_state_pages(struct kvm_vcpu *vcpu) -- 2.43.0