From nobody Tue Sep 29 08:25:36 2026 Received: from mail-qk1-f178.google.com (mail-qk1-f178.google.com [209.85.222.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B77BF40B6F2 for ; Mon, 10 Aug 2026 15:27:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786375655; cv=none; b=G7D2cOg6SPo129WIWGFXovmSi7xjVS87gvAdC981FofrBpTQL0aVZiDwWlzk80Ej12LRIm9TCJIDHKWFUFAdCEzzUcrvA6TC2iAQ6iG8CcbiATsXCfi+reUMqNF2I3qunJoZmDrphGK4612UeYYz5NDIUlrY1LiyWoPlJjRhDLw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786375655; c=relaxed/simple; bh=/+GXtYNIQEoanbyO74E2Z/1ZvUJU9TFAq2W8M3Yf/NA=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=mUxyrriryuJQZatLNLcTCL/F+QwwwKyM4ndCVdMNJhFOUnMlKyujnfKaYtgLFx2i6/PG0st5E6fqw5xwaWQAU0CGWnv8RJRrt481+/qWyPvyk5oFcM3KPWJsnxW9Q3Sqcm7ZAWsOpRGYaijZFKNdad321Ubylt0JRJDEQmIVYcw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=trailofbits.com; spf=pass smtp.mailfrom=trailofbits.com; dkim=pass (2048-bit key) header.d=trailofbits.com header.i=@trailofbits.com header.b=Hx5o9khX; arc=none smtp.client-ip=209.85.222.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=trailofbits.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=trailofbits.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=trailofbits.com header.i=@trailofbits.com header.b="Hx5o9khX" Received: by mail-qk1-f178.google.com with SMTP id af79cd13be357-92e6391b114so148010285a.3 for ; Mon, 10 Aug 2026 08:27:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=trailofbits.com; s=google; t=1786375651; x=1786980451; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=JeSEH2kZC7O/8qyc+dhZdHeAalMWmSlfY9/FgFd5LWI=; b=Hx5o9khXTSdMRo8B6NjbuZIghCjC/jKrSGbmooI/KZgmZkK5lXF2wHEUut9Pn6Nw9r KE9M0V8jnXdUDjFEQLO9HxqCLBQmCXv4ugykJct8oN/bJli5F9DHRKIUa5mYP6aPOuyx EMpCOBpnV5CJ568/e+pMkwWzRPPjCcg11Gif+XmDebRsL40CC/UEzbi+4vDKGlFZYWSC rdXWI4La+97l8cA//krwLYGcQjW19SekV9LZsw/pXW6zDoTnwz90U8Q1VD+8dWU4+tPa me+7ZMm5PwiGqwBIJ13jOst20JXcvVGy1yRo93ySY/RDVMTa8q0q277jPReM12S6H7ra 5R5w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786375651; x=1786980451; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JeSEH2kZC7O/8qyc+dhZdHeAalMWmSlfY9/FgFd5LWI=; b=RXRo57KWCqR+NxNWBW/oni7sIYqnsmuB0PCorS7bsxLO8bhtAvA4cG5tcMHeyFCFLI jzf+hTSLrgcTAkdybBrSQEi4pl8iF8Afn1cQBPVFp6DTOQTd+fT6NUzL1vwcUf/lJHvF pwa1y6yBRERrzEaE4KQXTkrjR9hL7KvglpaxfWKInNNo4PDfZulG4OZnPIDWk+Rmpuyz GkltFFh96ww3IxEHnHkmr+dyNSXXbhUjGnQhQcMqs465pyJ+QnEGaBAa5sWY1MhuI/jT Y8gLw5AuNaSlcg2kwvoSSto+LZ5KiELTAuTmrIndCDpt1xwMTCtOu/cbrAPY6KcSHv75 9PXg== X-Forwarded-Encrypted: i=1; AHgh+RpN/eWe7yhMMci3UnHUUTeZyf7Qh+3HKjCet9zhI03EhZ7Yc34lbhu0966c0/BnIv7Jvlb+Dr9EMNwkwDs=@vger.kernel.org X-Gm-Message-State: AOJu0Yw8ESaTJgY8r5d7pwYn/aiWXomy/cf8AvxxwNjVL7YCHY9fn4zk 5qfsw2d6WUFqaxhEWP4FoNofiq6QiAerPN2HudHW6ST6jlcwNOHcUilAp85NaG9DBRo= X-Gm-Gg: AR+sD12XdSQ/96FrNCCLSN2+nd/gC5lkF2Nsktvm8qiJCvrQQIPtsdI98PTe+qC670x B0u6fvjxEYWTNHAAMSnmMOsfpCmAU0dEP099IX+h8rvCGadvjmUUaAtNrFKJCOuuYk1SdXcCep3 e98Ri2rF9S0SjgbyJTRiFB4XtBwnd6UY/+aLeBaOzbHG7cxggRghWwr+5z8HPAsIZcOW76vL3Om tkB+uq3sPwqNj3YUOUNV9js+XcE31WUmtKFDrQoXcERQpb4dambZzjhQ1QpQEgvpndjqQiXjhHe d1PABue7RG9U4aeHGl6WWGDJ9JMExknMYPSHVMbIztWI2V69oYgDn9xje/h5QccUaq1eWBZOA5c teFVYRkZ8gUvU8vV+oYFlryqUdfqqSLd4EeDfsA1D8rnvQ6e2y/0/m4No1UcEPRYAZDXfVWs9SX 12EiotfXXQU/fbXVRtpEjV5+l4Mb+9oITtV3yq7/1yPBhF6tss3fIPWA7qsJM5ISYHhg== X-Received: by 2002:a05:620a:2b85:b0:916:10f6:780f with SMTP id af79cd13be357-9366655a73bmr3212596685a.27.1786375651355; Mon, 10 Aug 2026 08:27:31 -0700 (PDT) Received: from localhost ([146.190.222.192]) by smtp.gmail.com with UTF8SMTPSA id af79cd13be357-9366dc51b43sm808858685a.0.2026.08.10.08.27.30 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Mon, 10 Aug 2026 08:27:31 -0700 (PDT) From: David Lee To: pbonzini@redhat.com Cc: Kyle Zeng , Dominik 'Disconnect3d' Czarnota , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org, David Lee Subject: [PATCH] KVM: pfncache: track all MMU notifier invalidations Date: Mon, 10 Aug 2026 15:27:29 +0000 Message-ID: <20260810152730.841260-1-david.lee@trailofbits.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kyle Zeng There is a race condition in KVM's gfn-to-pfn cache refresh and MMU notifier handling. An HVA-backed cache can publish a stale PFN and kernel virtual address after the corresponding userspace mapping has been invalidated. The Xen shared-info HVA interface immediately reads and writes through that stale address, resulting in a host-kernel use-after-free. The cache refresh path in virt/kvm/pfncache.c drops gpc->lock while resolving and mapping an HVA. It uses mn_active_invalidate_count and mmu_invalidate_seq to detect an MMU notifier interval that overlaps this unlocked window. However, mmu_invalidate_seq is advanced only when the invalidated HVA overlaps a KVM memslot. HVA-backed caches are explicitly allowed to refer to memory outside all memslots. If such an invalidation starts and finishes while gpc->valid is false, the active count returns to zero without a sequence change and the refresh accepts a stale PFN. An unprivileged process with access to /dev/kvm can reach this path with KVM_XEN_ATTR_TYPE_SHARED_INFO_HVA. KASAN-detected use-after-free in kvm_xen_shared_info_init(). The affected function reads and writes Xen wall-clock fields through the stale mapping, so the issue can cause a host-kernel crash and memory corruption. The attached KASAN output confirms: BUG: KASAN: use-after-free in kvm_xen_shared_info_init+0x344/0x3d0 [kvm] Read of size 4 at addr ffff888046000900 by task poc/1266 Add a notifier-specific sequence that advances for every completed invalidate interval before mn_active_invalidate_count is decremented, and use that sequence for pfncache retry. The existing barrier pairing then guarantees refresh observes either an active invalidation or a sequence change. Fixes: 721f5b0dda78 ("KVM: pfncache: allow a cache to be activated with a f= ixed (userspace) HVA") Cc: stable@vger.kernel.org # 6.9+ Assisted-by: Codex:gpt-5.6-sol Codex:gpt-5.5-cyber Signed-off-by: Kyle Zeng Co-developed-by: David Lee Signed-off-by: David Lee --- Bug found and triaged by OpenAI Security Research and validated by Trail of Bits. Trail of Bits has a reproducer for this bug that triggers a KASAN use-after-free and can share if needed. include/linux/kvm_host.h | 1 + virt/kvm/kvm_main.c | 9 ++++++++- virt/kvm/pfncache.c | 18 +++++++++--------- 3 files changed, 18 insertions(+), 10 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index ab8cfaec8..0ac382cd9 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -800,6 +800,7 @@ struct kvm { /* Used to wait for completion of MMU notifiers. */ spinlock_t mn_invalidate_lock; unsigned long mn_active_invalidate_count; + unsigned long mn_invalidate_seq; struct rcuwait mn_memslots_update_rcuwait; =20 /* For management / invalidation of gfn_to_pfn_caches */ diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 45e784462..5e43dd63c 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -812,8 +812,15 @@ static void kvm_mmu_notifier_invalidate_range_end(stru= ct mmu_notifier *mn, =20 /* Pairs with the increment in range_start(). */ spin_lock(&kvm->mn_invalidate_lock); - if (!WARN_ON_ONCE(!kvm->mn_active_invalidate_count)) + if (!WARN_ON_ONCE(!kvm->mn_active_invalidate_count)) { + kvm->mn_invalidate_seq++; + /* + * Publish the sequence update before dropping the active count + * so that pfncache refreshes observe one or the other. + */ + smp_wmb(); --kvm->mn_active_invalidate_count; + } wake =3D !kvm->mn_active_invalidate_count; spin_unlock(&kvm->mn_invalidate_lock); =20 diff --git a/virt/kvm/pfncache.c b/virt/kvm/pfncache.c index 728d2c1b4..d360f1eda 100644 --- a/virt/kvm/pfncache.c +++ b/virt/kvm/pfncache.c @@ -124,7 +124,7 @@ static void gpc_unmap(kvm_pfn_t pfn, void *khva) #endif } =20 -static inline bool mmu_notifier_retry_cache(struct kvm *kvm, unsigned long= mmu_seq) +static inline bool mmu_notifier_retry_cache(struct kvm *kvm, unsigned long= mn_seq) { /* * mn_active_invalidate_count acts for all intents and purposes @@ -136,20 +136,20 @@ static inline bool mmu_notifier_retry_cache(struct kv= m *kvm, unsigned long mmu_s * Note, it does not matter that mn_active_invalidate_count * is not protected by gpc->lock. It is guaranteed to * be elevated before the mmu_notifier acquires gpc->lock, and - * isn't dropped until after mmu_invalidate_seq is updated. + * isn't dropped until after mn_invalidate_seq is updated. */ - if (kvm->mn_active_invalidate_count) + if (READ_ONCE(kvm->mn_active_invalidate_count)) return true; =20 /* * Ensure mn_active_invalidate_count is read before - * mmu_invalidate_seq. This pairs with the smp_wmb() in + * mn_invalidate_seq. This pairs with the smp_wmb() in * mmu_notifier_invalidate_range_end() to guarantee either the * old (non-zero) value of mn_active_invalidate_count or the - * new (incremented) value of mmu_invalidate_seq is observed. + * new (incremented) value of mn_invalidate_seq is observed. */ smp_rmb(); - return kvm->mmu_invalidate_seq !=3D mmu_seq; + return READ_ONCE(kvm->mn_invalidate_seq) !=3D mn_seq; } =20 static kvm_pfn_t hva_to_pfn_retry(struct gfn_to_pfn_cache *gpc) @@ -158,7 +158,7 @@ static kvm_pfn_t hva_to_pfn_retry(struct gfn_to_pfn_cac= he *gpc) void *old_khva =3D (void *)PAGE_ALIGN_DOWN((uintptr_t)gpc->khva); kvm_pfn_t new_pfn =3D KVM_PFN_ERR_FAULT; void *new_khva =3D NULL; - unsigned long mmu_seq; + unsigned long mn_seq; struct page *page; =20 struct kvm_follow_pfn kfp =3D { @@ -181,7 +181,7 @@ static kvm_pfn_t hva_to_pfn_retry(struct gfn_to_pfn_cac= he *gpc) gpc->valid =3D false; =20 do { - mmu_seq =3D gpc->kvm->mmu_invalidate_seq; + mn_seq =3D READ_ONCE(gpc->kvm->mn_invalidate_seq); smp_rmb(); =20 write_unlock_irq(&gpc->lock); @@ -232,7 +232,7 @@ static kvm_pfn_t hva_to_pfn_retry(struct gfn_to_pfn_cac= he *gpc) * attempting to refresh. */ WARN_ON_ONCE(gpc->valid); - } while (mmu_notifier_retry_cache(gpc->kvm, mmu_seq)); + } while (mmu_notifier_retry_cache(gpc->kvm, mn_seq)); =20 gpc->valid =3D true; gpc->pfn =3D new_pfn; --=20 2.53.0