From nobody Mon Sep 28 04:49:02 2026 Received: from mail-pg1-f197.google.com (mail-pg1-f197.google.com [209.85.215.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 199A0479883 for ; Wed, 26 Aug 2026 16:42:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762571; cv=none; b=Eze/A2LIrWbfS0vC9fw2BgMt1Zfj+Ucc8W0O6DD2d2/7zKOkgpFvrddtNOqLmYtgx1rYXKz/RzhPFnConllpc3jx/0CzInTJM/XhIirxriC4v+y06qiIOcyOL/ncml8uMxDk20PH+fPWkbrHmFeS+sMuNATB7bPUaMS4FVa+hzA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762571; c=relaxed/simple; bh=i578dgYavrIdohpu+SpEj4jJ+sVyZzrBwA5orlX4/TM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=C+myOhUB4Hd/zozmJhcoqAUV+SXGnkeGY7QnafphnskaUW5yCbuTPK9as4qfWmrePRVuotXJx2g+sGb6wLWcO/yUmRIRhFTXqT6vhMBVMlmFbujgbEebo2drJ3QyKJEhgvxkwJTCsuZDy0/ugnu8OvRVoZ4iZlq9AcFA1vJCWDo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=tSUIZZ0y; arc=none smtp.client-ip=209.85.215.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="tSUIZZ0y" Received: by mail-pg1-f197.google.com with SMTP id 41be03b00d2f7-cbedb8673ceso1145281a12.0 for ; Wed, 26 Aug 2026 09:42:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787762538; x=1788367338; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=t27G+87/xldPf9EXxl/42NH6liZLtdScRqOAABQHwp0=; b=tSUIZZ0y4SjXHgxEnyp4zLNI5D3UOVcK+Bn9+5T8D2Zo4qXvCucomKd8DEL/6w3H8t cMZG948x2CaiyVn/yZOQixWIe1lDynVzttkCjkR5zxhXF5uHocnaD+hqtUj3f5pnc/xA tmjOfZvI1S6e+hkWCqg3iMhywnYbLL503JBAVgDwMyXQo7RuSMLZvvaWUAHTDF8kd0fa qSp4laKfkZKUZ+dSAUipo1jlL0rwaNGMjn7dCfe5B8sncmmzX7I/UrGlJGBRVjxCoyNC XtdsK1HPHOI1xe0dimhify2pUEzwOWjG2Ic954iXIO/J/b7X35Y/VcUbCI3F3UEoJBu8 WAdQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787762538; x=1788367338; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=t27G+87/xldPf9EXxl/42NH6liZLtdScRqOAABQHwp0=; b=KgkseHCKcXBXoyk+KTgXcsiuag3fe+jylduLGKCS8i8ZJ1ulwm5+Pb2kdIWFr2sMqb CMncskDqwl8MJkJ27DXPleqO/RJaVHu/Ut8R1cGIa+ZFuOe+L4mbo9m0q8fa9Z7Y+u+p 1YD8VQATJbjwVJjwHNmL1v7K3g4zVUUdEKRC8Gp8xJ5VUP3vqAZUoNeVIb9wLj2VY6FY W7DiKFTc9hFaIHOyLC4OLpE4oE8CmgDnOe1u5/85MxA5qD31eemAFtk+cm9Ssj+KYLyV IU2dTnjEFp8v/ntpy8FkgIfhdIIex/G6CxryZqdh9so5z9236VOcRZBtCvcDCltmNmA3 eSmQ== X-Forwarded-Encrypted: i=1; AHgh+RoLtStVoF9ZylAJzPnynP6ne399lqBQYU5EfZDYOkqGW0NfvTlvPSHg+2w2QfJaI09EI1NyB2gAuSSEIoI=@vger.kernel.org X-Gm-Message-State: AFuF++lNSndh7P60p/8OHIwTZTGkxtHvdHAG9W17Jo8q485BREEN/0gc VC5skgwxSAvV/yp/0GMcit07VT7G0svEEx3bCqwmKeaaALK9n9k25Lyg8l5ZnndW6CPqPK4S88H VanQRsQ== X-Received: from pghp20.prod.google.com ([2002:a63:fe14:0:b0:c9a:c533:8316]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:4f92:b0:848:2e3c:9955 with SMTP id d2e1a72fcca58-85371fa599amr14451999b3a.4.1787762537296; Wed, 26 Aug 2026 09:42:17 -0700 (PDT) Reply-To: Sean Christopherson Date: Wed, 26 Aug 2026 09:42:11 -0700 In-Reply-To: <20260826164214.756512-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826164214.756512-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.860.g4b6b3295ed-goog Message-ID: <20260826164214.756512-2-seanjc@google.com> Subject: [PATCH v2 1/4] KVM: x86/mmu: Reload MMU on *every* page pre-fault attempt/iteration From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Rick Edgecombe , Kai Huang , Yan Zhao , Sashiko Bot Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Reload the MMU (which is a nop if the MMU doesn't need to be reloaded) on every attempt to pre-fault a guest page, i.e. when the page fault path signals that the caller should retry. If the synchronize_srcu_expedited() in kvm_invalidate_memslot() completes before kvm_vcpu_pre_fault_memory() grabs SRCU, but kvm_mmu_reload() in the pre-fault path completes before kvm_invalidate_memslot() triggers x86's "fast zap all", then the pre-fault task will reach kvm_tdp_page_prefault() with an invalid root. Attempting to fault-in memory with an invalid root ultimately puts kvm_tdp_page_prefault() into an infinite (breakable) retry loop, which manifests most obviously as a hang in the pre_fault_memory_test selftest, but also eventually causes RCU (SRCU?) to complain. INFO: rcu_tasks detected stalls on tasks: 000000000cda47bd: .. nvcsw: 6/6 holdout: 1 idle_cpu: -1/25 task:pre_fault_memor state:R running task stack:12696 pid:95588 tgid:95588 ppid:95584 task_flags:0x400000 flags:0x00080801 Call Trace: lock_release+0x4e/0x320 __get_user_pages+0x546/0xcd0 up_read+0x1b/0x30 get_user_pages_unlocked+0xee/0x350 hva_to_pfn+0xd3/0x3d0 [kvm] lock_release+0x4e/0x320 xa_load+0x5c/0x170 xa_load+0x14c/0x170 __kvm_faultin_pfn+0xd9/0x130 [kvm] lock_acquire+0x65/0x2b0 lock_release+0x4e/0x320 kvm_mmu_faultin_pfn+0x1e1/0x690 [kvm] gup_fast_fallback+0x63e/0xdf0 kvm_tdp_page_fault+0xeb/0x140 [kvm] kvm_mmu_do_page_fault+0x12e/0x200 [kvm] kvm_arch_vcpu_pre_fault_memory+0x16e/0x200 [kvm] kvm_vcpu_pre_fault_memory+0xc1/0x1f0 [kvm] kvm_vcpu_pre_fault_memory+0x116/0x1f0 [kvm] kvm_vcpu_ioctl+0x3a4/0x6b0 [kvm] clockevents_program_event+0x5d/0x170 __se_sys_ioctl+0x6d/0xb0 entry_SYSCALL_64_after_hwframe+0x4b/0x53 do_syscall_64+0x10a/0x480 __irq_exit_rcu+0x8e/0x140 entry_SYSCALL_64_after_hwframe+0x4b/0x53 Fixes: 6e01b7601dfe ("KVM: x86: Implement kvm_arch_vcpu_pre_fault_memory()") Reviewed-by: Rick Edgecombe Reviewed-by: Kai Huang Signed-off-by: Sean Christopherson --- arch/x86/kvm/mmu/mmu.c | 12 ++++-------- 1 file changed, 4 insertions(+), 8 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 064ecc33b926..3220f05387b5 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -5061,6 +5061,10 @@ static int kvm_tdp_page_prefault(struct kvm_vcpu *vc= pu, gpa_t gpa, if (kvm_check_request(KVM_REQ_VM_DEAD, vcpu)) return -EIO; =20 + r =3D kvm_mmu_reload(vcpu); + if (r) + return r; + cond_resched(); r =3D kvm_mmu_do_page_fault(vcpu, gpa, error_code, true, NULL, level); } while (r =3D=3D RET_PF_RETRY); @@ -5101,14 +5105,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu = *vcpu, if (kvm_is_gfn_alias(vcpu->kvm, gpa_to_gfn(range->gpa))) return -EINVAL; =20 - /* - * reload is efficient when called repeatedly, so we can do it on - * every iteration. - */ - r =3D kvm_mmu_reload(vcpu); - if (r) - return r; - direct_bits =3D 0; if (kvm_arch_has_private_mem(vcpu->kvm) && kvm_mem_is_private(vcpu->kvm, gpa_to_gfn(range->gpa))) --=20 2.55.0.860.g4b6b3295ed-goog From nobody Mon Sep 28 04:49:02 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1A26F479890 for ; Wed, 26 Aug 2026 16:42:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762590; cv=none; b=NgIAm5kM4Xik3bXP5IZfvhUwD82ngjnI63Cy/j8p1LM5pJNE6SY5rFyDKxctSJ7D7+kLXII1WAYCDLAp/yCJhR3vRnUfJ4BCkchqIKlqQz/P7suSQV7qvOQEnpObzwRLWYoN9WwDEC9PqvxuqxoHjXAduvD5+SDp/B00/cA2Hc8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762590; c=relaxed/simple; bh=BEhe2r+8iU4EkJKKeUDkD+GqT9e+NS4i1MkXdRfLIgA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=asNs/fMSKXw6hAA6ZBXkHnW92vGPUuxzhq01EEOwFFrfxyBOtTsST+AQzUqDlJbGIaq8G+XODpeMiZgMfAO6S9Y0HqKwr08cuiDcN02dvbMdpnqaEKpVTxm4bXZEUm5+qSYWNTM9Y3J8fh96KsQDJJiI5aOnnfHq38pqpeqOY9k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=rd1EuFxO; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="rd1EuFxO" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2cfca8558d2so16871025ad.2 for ; Wed, 26 Aug 2026 09:42:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787762539; x=1788367339; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=Ji0mUaDNfVUZXzbmrEerFV372UGViNyfEX1CgVwEdcw=; b=rd1EuFxOE7fMErYv4lq06VO33Ck9hx3jzPxEOU4a4g1wpaCn/h3V73xoT4FzV+PMpP JT2o0Vs2RtD3koE9JX/XUqqby8I7fmHOAXVj1TCwGq9qK3F+lnlcktZv+3llkJIdZ92h 2CXp92hb1WOsBwUitGcl0LYdnV65oklSH8GrsVttaT/V5+UqSX4rjLAk8HE3ehoWeFmL j5fwgqeTv113BEOsmVZhTiul5Yqtcra7ieAe/i56gyVTN3fEENxvcpFhvXrcjQv3ISdv K1PB8fQV16r/6I8zrmgVnycc7XEyiLl/svCQGe8VLwdeQC2SrT2mSXJ8V+CeHkEI7BY8 heUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787762539; x=1788367339; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=Ji0mUaDNfVUZXzbmrEerFV372UGViNyfEX1CgVwEdcw=; b=UZIcRoNs47kxlixBFpfSTQYK+v0ONcKF+Yi4QCzFP1GBwzvFiEk7oqpVw2ozA9ABOU t1Do7JQkfA60Sow8dbAbZ80nbuRQC20asdDMdOqVeI+b28FpIr+uoKqsAX23q3P2OplN 4fEAG5mMr1HmQiuEPgZaGn3iY/DYOGJ3fwrIhH1uctKiTbpy0ZFUrb4tli3bDrHYfCna HxzgwfqKHtEaOirhPYuxu6d3PF+OWYUzoLprNLnc7p5SswODCgJctEnV6bbAdJHlZtld h1xTcXowzn465SnLOtRE3GRs96OKflycNlelIA42oUC4Q57kcwGhuRfh2xmXepbLfuKs v/cQ== X-Forwarded-Encrypted: i=1; AHgh+Ro+HMlgPiEQHbGxbCfHNjx5hTTmquXDIdvsbdhO0p/86N+6tqbVFCxWNiZxFFt8G1pjoqDoQZ8w0GMuVPM=@vger.kernel.org X-Gm-Message-State: AFuF++m3MRnRDxJumNLztIV6nLlrddOdSHirrMpqE1QJyEI4NdLwVvD8 SyDo0/0gz1Y16lI9pS9IRDhkacGBPMh71T1Lthei0EHmfcH0lpFepz4UMjJKFKOO0yKFTolRedt NYsWQKw== X-Received: from plef20.prod.google.com ([2002:a17:902:f394:b0:2cc:79e2:e717]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:ef50:b0:2d7:1aeb:739 with SMTP id d9443c01a7336-2d71aeb1484mr67465065ad.13.1787762538425; Wed, 26 Aug 2026 09:42:18 -0700 (PDT) Reply-To: Sean Christopherson Date: Wed, 26 Aug 2026 09:42:12 -0700 In-Reply-To: <20260826164214.756512-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826164214.756512-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.860.g4b6b3295ed-goog Message-ID: <20260826164214.756512-3-seanjc@google.com> Subject: [PATCH v2 2/4] KVM: x86/mmu: Harden "map private PFN" against unexpected root invalidation From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Rick Edgecombe , Kai Huang , Yan Zhao , Sashiko Bot Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Move kvm_tdp_mmu_map_private_pfn()'s reload of the MMU into its tight loop so that an unexpected root invalidation has a better chance of being handled gracefully, even though it should be impossible for the vCPU's root to be invalidated after the initial reload. As is, encountering an invalid root is *guaranteed* to put the task into an infinite loop (albeit a breakable loop that honors NEED_RESCHED). Add a WARN to try and detect bugs that break KVM's expectations, along with a comment to explain why it should be impossible for the root to be invalidated. Note, the loop in question doesn't actually check for a stale page fault, i.e. likely won't detect an invalid loop in the first place. That bug will be addressed shortly. Cc: Kai Huang Cc: Yan Zhao Cc: Rick Edgecombe Signed-off-by: Sean Christopherson --- arch/x86/kvm/mmu/mmu.c | 16 ++++++++++++---- 1 file changed, 12 insertions(+), 4 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 3220f05387b5..1969c26861e5 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -5209,10 +5209,6 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcp= u, gfn_t gfn, kvm_pfn_t pfn) if (kvm_gfn_is_write_tracked(kvm, fault.slot, fault.gfn)) return -EPERM; =20 - r =3D kvm_mmu_reload(vcpu); - if (r) - return r; - r =3D mmu_topup_memory_caches(vcpu, false); if (r) return r; @@ -5224,10 +5220,22 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vc= pu, gfn_t gfn, kvm_pfn_t pfn) if (kvm_test_request(KVM_REQ_VM_DEAD, vcpu)) return -EIO; =20 + r =3D kvm_mmu_reload(vcpu); + if (r) + return r; + cond_resched(); =20 guard(read_lock)(&kvm->mmu_lock); =20 + /* + * Because slots_lock is held, it should be impossible for *any* + * roots to be invalidated after the initial MMU reload. WARN, + * but continue on; the above MMU reload will do the right thing + * if the current root is actually invalid. + */ + WARN_ON_ONCE(kvm_test_request(KVM_REQ_MMU_FREE_OBSOLETE_ROOTS, vcpu)); + r =3D kvm_tdp_mmu_map(vcpu, &fault); } while (r =3D=3D RET_PF_RETRY); =20 --=20 2.55.0.860.g4b6b3295ed-goog From nobody Mon Sep 28 04:49:02 2026 Received: from mail-pg1-f200.google.com (mail-pg1-f200.google.com [209.85.215.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1A13147988C for ; Wed, 26 Aug 2026 16:42:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762554; cv=none; b=IuxTJ0n9mAJC6UHhx9Y3axiiFxDV4pGdfaQ4CHScBBMr1urrlZQbdodhCcWKak9mggweM8bBzuVwzfcsoIG9PvcLa4bUdMIhR9N642/wyOzAxQS3iSX4cxz2Z4wYZsVz1cYoZtFUCpIEQxlFdkgF0t91oaZeiSKHikzRZzME/4Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762554; c=relaxed/simple; bh=N3KhGuurOfNAdflb4lT6TlTE27zh2L+SziWwZzTsYoY=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=WO8aBOXWhr720z4CCnxfhxsyiHx73VAKZSqZKAjRlR+h17WwNKLKSbREdxHS+j3urACuJbp09NzjYYoYE4+WKgU/h4hwM7KLAIap93U0nGuYGv+SZHdbFGhPAbhC3YV2RtaiGKK8PCnpGmWqGHofVAX2BFGtUDqRfpSjCHqy8hM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=hV4FhYtr; arc=none smtp.client-ip=209.85.215.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="hV4FhYtr" Received: by mail-pg1-f200.google.com with SMTP id 41be03b00d2f7-cbb20f82a0eso856667a12.0 for ; Wed, 26 Aug 2026 09:42:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787762540; x=1788367340; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=T9ggCE2s4lQUcVzu7FyWkmkWQene3zbgM9NYxUP0czg=; b=hV4FhYtrVsqpYlqspNSJ66iqkGEbPoQyqTF4n6VQx9xW490UO/LLkay7j2Gt7x5hzy 7emkWgHDgGAJkupB0FNNMWSrJ9JQ37CkDZnnuO85308h2E7wD5SkBh8a71XMwWKVqbgs txTDshkMUgywBxKQqo4j1qg/ldAtXtFK60ZleQxBLlw6uhG/L8rP32a1NxdkDLCcfCd5 1l6XcOxOn5esGWTblmW70lZb+xDILNTNRBhjrjc8gC5Nhj+AJXoHDftawRJrdPeT3+IN RXi6prRs+15dCTXYCIRPzocc5dQCiuM2nHQy8uGoYkUA8XKbMPx42lwgv0eceAE8Uodd cL5Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787762540; x=1788367340; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=T9ggCE2s4lQUcVzu7FyWkmkWQene3zbgM9NYxUP0czg=; b=bh1rAP2S/MK2lC7mM5wmXEjOyQwsxPVtAUFhXht0VZcPqgpA8Zz8hO+4LBNR+Dy1dh Tkh1TVyiYYxHBWC/S0hOSdXrGOGkKWvInZWGmRNlcSk+++5cJDIRlwCkoGpiyhvfTxlv GuP3BKw11JZzf0tYp+dyXPmqj6OafGEQYjRPc1y48oMrjRyv+fnxAY3oMAheXSumOK7K W+gPhxSAoPtJ+TGwq6WiSiYnAy47+0a8UurxArK1lkJw1licehZz305fDfpbg9lNs9Tm zf2pQ6Pz4hx85ese7fGlvl4HIdVppA9dnUdGvboy38nxmU7Fg82tZR+IYrda4vv0Z0fs sslg== X-Forwarded-Encrypted: i=1; AHgh+RpMEJpeHmUXnmbHiVbYwhSOrp7zTsDRmaZ8dm9N55IisjiDSAi0sfXQk1NCSwdV38AhB5QXqD+YOc8x+lk=@vger.kernel.org X-Gm-Message-State: AFuF++lQWPfD3JmA8Ehl3NHqvyiqTRDxFfTGEwyQc+dvNFm2hvOFT4bA FJkA1D9iasaQ838Askt+Wqb7j8W++kJ1QxLi8GWDWDOvnIonq+dCTavCoCaQWzAinklsRdyg6H0 l8gzThA== X-Received: from pgbct7.prod.google.com ([2002:a05:6a02:2107:b0:cc1:d48e:146c]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:e210:b0:3bf:63af:855 with SMTP id adf61e73a8af0-3cf762865fcmr14724358637.1.1787762539550; Wed, 26 Aug 2026 09:42:19 -0700 (PDT) Reply-To: Sean Christopherson Date: Wed, 26 Aug 2026 09:42:13 -0700 In-Reply-To: <20260826164214.756512-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826164214.756512-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.860.g4b6b3295ed-goog Message-ID: <20260826164214.756512-4-seanjc@google.com> Subject: [PATCH v2 3/4] KVM: x86/mmu: Top-up memory caches when retrying "map private PFN" From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Rick Edgecombe , Kai Huang , Yan Zhao , Sashiko Bot Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When mapping a private PFN in TDX's post-populate callback, top-up the memory caches on every attempt to map the PFN to harden against bugs in the map flow that could consume cache entries even if mapping ultimately fails. E.g. as pointed out by Sashiko, the in-progress Dynamic PAMT support could consume PAMT cache entries on TDX-Module lock contention. Harden KVM even though consuming an entry on failure is considered a KVM bug. Retry should only be encountered if KVM is buggy (the locks held by the sole call path will prevent retries from being needed due to TDX-specific details, and memory can be faulted in only once the VM is TD_STATE_RUNNABLE, and KVM_TDX_INIT_MEM_REGION is only usable if the VM is *not* TD_STATE_RUNNABLE), top-up is "free" if there's no work to be done, and populating a TDX guest's memory is a slow path, i.e. there's no meaningful downside to the hardening. Reported-by: Sashiko Bot Closes: https://lore.kernel.org/all/20260718061050.E17B01F000E9@smtp.kernel= .org Reviewed-by: Rick Edgecombe Signed-off-by: Sean Christopherson --- arch/x86/kvm/mmu/mmu.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 1969c26861e5..19a501029f08 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -5209,10 +5209,6 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcp= u, gfn_t gfn, kvm_pfn_t pfn) if (kvm_gfn_is_write_tracked(kvm, fault.slot, fault.gfn)) return -EPERM; =20 - r =3D mmu_topup_memory_caches(vcpu, false); - if (r) - return r; - do { if (signal_pending(current)) return -EINTR; @@ -5224,6 +5220,10 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcp= u, gfn_t gfn, kvm_pfn_t pfn) if (r) return r; =20 + r =3D mmu_topup_memory_caches(vcpu, false); + if (r) + return r; + cond_resched(); =20 guard(read_lock)(&kvm->mmu_lock); --=20 2.55.0.860.g4b6b3295ed-goog From nobody Mon Sep 28 04:49:02 2026 Received: from mail-pg1-f200.google.com (mail-pg1-f200.google.com [209.85.215.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA2AF470E8E for ; Wed, 26 Aug 2026 16:42:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762570; cv=none; b=LvOTidgupfAhjvzv07affX78gaD4UEV5B9ShizvgD1zkKxi2IMDI0jQn2V27fUzJmDGyJyrqbkSFAuPsLrLcuUxLHXx3HbnoxVpUX1MqImT8anwIcQ3QGudxhb2IZ4KQnueGOx9C2M68b7XgTpCDOSn3CY56WB6/pW9I/owS1D0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787762570; c=relaxed/simple; bh=4eSK8pU1dmh/XamK6Y3En78UgfUaTPDadfBbBp/fvUI=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=axsas0nW7wg5HovQ2oeij8Q30t5BL5u5qmVxqg1G10B27Eyxsderg63O0Cim6FvJDc6TeiEtDSBBKn2NbdfhgCvQur3AU+KOrLk6rZxO0Lz2CAOoXuBShWiB7SCuri/Yk4tQykcoEdtWCQhPdPC3OVAbBCY7JhhrYLvf7vwQ2xA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=pgz8r9Z5; arc=none smtp.client-ip=209.85.215.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="pgz8r9Z5" Received: by mail-pg1-f200.google.com with SMTP id 41be03b00d2f7-cc1a439db36so1344605a12.2 for ; Wed, 26 Aug 2026 09:42:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787762541; x=1788367341; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:from:to:cc:subject:date:message-id :reply-to:content-type; bh=BBZMmsivox8ZLj9npWSJpItnIK5TSzdBCqDSLHP2rcA=; b=pgz8r9Z5JP95qPHy6H0jr/Xbn1mMDMUNS1QQi7Hcm9iaPofbvxn43EgW8mJB8jFj+D wu8evYH3lgODrq2vl4s5CY379N+lrNq+Xux7jDC6C6Ue0R2gr1b1W1JXjbVIgbskEFOv 2jqyRc+tz06QUDXUSX/0xam5Q008ZMt4DH6w5U7k8zWQve1rWxFvwiFLwyzBKFyri1TG qM3b4GGZg+mHKUZ3qyOLyu7O04X+z5s7lr59agwP1AdGGmoNIGNH1z0V4lfSDBJ045L0 4bMwylC9l23Dqxk18rB+mMLPJ/oaPLsQ5xGiRTON7CB6PcllcZD9WS2M0gJE2tEkbNxE ZPzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787762541; x=1788367341; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:reply-to:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=BBZMmsivox8ZLj9npWSJpItnIK5TSzdBCqDSLHP2rcA=; b=Y1jiOkufp13vywEdM9Vs15ZzQM+y4BXxjAN/gnFpN6+w9YiiMBdOsC2Rl8gul+gOHZ DI0Atlq0nkHGHKbvDsqXPOWjQqq7tokcgHhOxztZwjpe5EV1YbcWrfYeT5bZDpbB9KKE aijixW5M66N3owUmlgGKG67lF/DRwlgbwpSEZGP1jz7UKROtbtuQRNjxTY/Ry37VYKRa tTE2/esqPhoW3+ipSgOzeuk/OlXiXHVf0bmXAyrbD+ZvceOaYnyGD7+hwfyfk9U+ctih NsZJeX7zLKtmhPiNm0NUWP8oNBt5HJOQA8qYur/IyyGPMAYnkJH1D75+/6R8W6KgNEPT +VPg== X-Forwarded-Encrypted: i=1; AHgh+RpKDtt9QCOF0bwzSLoFqh2Cx6Jb86FHwsTdqTOfooWhtvafzp8BkRKGIwcbEoEjgic5qIOCiZTQLK2bASE=@vger.kernel.org X-Gm-Message-State: AFuF++linzIZGWoKXqFoqgNsLRNusSUOFSmFgfOzyKyOgchm+gK0HZmY B8xyqZtCUxDrvfUNv8gvVyhcIVSzBrfpJY/3iu39pV87X+PHNfjSbAtoNKJdsZGp1Ob949NeokA aCr6Khg== X-Received: from pgdq18.prod.google.com ([2002:a63:9812:0:b0:cc1:bf16:dc3c]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:8087:b0:853:2ca8:c42d with SMTP id d2e1a72fcca58-85371fa59a6mr17184684b3a.2.1787762540664; Wed, 26 Aug 2026 09:42:20 -0700 (PDT) Reply-To: Sean Christopherson Date: Wed, 26 Aug 2026 09:42:14 -0700 In-Reply-To: <20260826164214.756512-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826164214.756512-1-seanjc@google.com> X-Mailer: git-send-email 2.55.0.860.g4b6b3295ed-goog Message-ID: <20260826164214.756512-5-seanjc@google.com> Subject: [PATCH v2 4/4] KVM: x86/mmu: Add sanity check to detect stale page faults in "map private PFN" From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Rick Edgecombe , Kai Huang , Yan Zhao , Sashiko Bot Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Harden the "map private PFN" flow against potentially-fatal bugs or future KVM changes by checking for a stale "fault" prior to actually mapping the PFN into the guest. While it should be impossible for the "page fault" to become stale, the sanity check is cheap, whereas a broken assumption would have a high probability of leading to a guest-expoitable use-after-free. Snapshot the invalidation sequence after acquiring mmu_lock to avoid false positives, even though doing so completely voids anys and all protection against unexpected invalidations. Pretty much the entire point of kvm_tdp_mmu_map_private_pfn() is that it allows mapping a PFN that was gifted by the caller, i.e. the caller would have to mess up its one and only responsibility. Signed-off-by: Sean Christopherson --- arch/x86/kvm/mmu/mmu.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 19a501029f08..79c450d677b4 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -5236,6 +5236,18 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcp= u, gfn_t gfn, kvm_pfn_t pfn) */ WARN_ON_ONCE(kvm_test_request(KVM_REQ_MMU_FREE_OBSOLETE_ROOTS, vcpu)); =20 + /* + * Snapshot the invalidation sequence counter after acquiring + * mmu_lock, as guest_memfd guarantees the validity of the pfn, + * i.e. any concurrent invalidations are guaranteed to be + * irrelevant. + */ + fault.mmu_seq =3D vcpu->kvm->mmu_invalidate_seq; + if (is_page_fault_stale(vcpu, &fault)) { + r =3D RET_PF_RETRY; + continue; + } + r =3D kvm_tdp_mmu_map(vcpu, &fault); } while (r =3D=3D RET_PF_RETRY); =20 --=20 2.55.0.860.g4b6b3295ed-goog