From nobody Fri Oct 2 06:58:31 2026 Received: from mail-pl1-f170.google.com (mail-pl1-f170.google.com [209.85.214.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CA4634570C7 for ; Tue, 4 Aug 2026 10:58:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785841091; cv=none; b=HcXFVMWI8ubwClnI4zwAlGt+NW0auBEe80hQgoCUdpywiQAN1guhTYS3t2lPZPEgIHZA4nli4FGRgXOYfT62bjg+ZqLPKjL2NY+BS4ESKGnYI9iXMZgh9J+YluVwkXk7XqyLqpr9V4h5/xPtMrOLr1YmqwcCWDkX/B+yNPuzq60= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785841091; c=relaxed/simple; bh=qNsGDjSXmrDO4E8OltKgVJGP7btPSiJzKaauWKqcujE=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=a0VcirePom9lm5Vnm2T2zRPTl9lHzhopq0a6eGxLw0LGeYk3b0cIed1YOZ6Mzirc6hnjcV5DEWIU0uU5j4LEU6VNJ3W+Dm+FSaosrfLwehUA2O8hdLIj85Qi8x9Dh6l7hW75+sIYKFRX0jeHM5ncDIWLNDqR26YgXG7T0unQz9I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jW2HLtQ6; arc=none smtp.client-ip=209.85.214.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jW2HLtQ6" Received: by mail-pl1-f170.google.com with SMTP id d9443c01a7336-2cfff5f88dbso51807575ad.3 for ; Tue, 04 Aug 2026 03:58:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785841082; x=1786445882; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=7gmvruGvhObjRkfatzPup5KdvMlhQ+iXfk/s6L7SaZ4=; b=jW2HLtQ60QQFt+SjWe9vngmz0t8kJKeFS10Fxm6bv03XD/UFVbQ0JcmmXbZClfN8L0 okzAlUj+Z5dsm4lCkGAc0PQw6YB90LPqcoIDAgmpk+lZirFfAeRy38LA4YdEYxxRLkLe X4NlA/hSdlKwkqXom7iCZaXhbRlyvnUyeoFRD4y3EBcTR3gu1qdBzbbONTkCeP+SxuoE WhN602rJZXk2sO4FvJNuH5bcxrEfzMCgQV0TXvlV6Y3nOx3e0i39vlktP62arco9m4X4 kSuC5k6BvCopopt1MHZN2GqjDNR3ZXE10MfeFl4grVhmKuvQ55jw7qjNVDX0DqVgdwDG cOsQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785841082; x=1786445882; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7gmvruGvhObjRkfatzPup5KdvMlhQ+iXfk/s6L7SaZ4=; b=dOXeMk0wlcQzpnSY4lRyX3OAJgu3IlNWNASiz4T9yebU5lg9My4yQXOZ6QmAva6ZNB id6kiZytbKikwmTInuTHaLRVjBCMib+yGAFTmW2R295JBS0w94HmaF7NSeTipGsc56Uq djMuB+7YN4b3r5+iiVJCdlExDVqQruNxqFWQYtms+cDzr1BIzsAXiruzTHhBimCVBqke 26jg9waFKhloDPCdB6aLtxAogHxBCkhiE72wGim8BOwKsSaWcA0aWnHyOb1JdPGOkvZZ ktj+CuE7Upz9F1fxje1j2JtP4KoyOvD0I59Uj6yZ9GKvLQWMaBY25FwjmmAhw++5fGOn twyA== X-Forwarded-Encrypted: i=1; AHgh+RoDUhPxACgOalibzBeKEdXjx4XDcnwhmXspzhtSU7/ffsU2QNkaM8TBsLEt4Fx62goI/7CZLZvmulzfuVE=@vger.kernel.org X-Gm-Message-State: AOJu0YzwA0RgBt2OW6MkvCTPzMk1tLZ986YRdmpTEGk9OAAVy+05cvBp geRPfGV+qj/BWe0yk1gA3tfPJBpzLniEqrakKeBeDUQYuKLAayheJDtR X-Gm-Gg: AR+sD10tthOAXfWq5z1v8PYfU+0A4E6uxG1Twe8M8iB2iqtR6BAI0P/q30ZuW94AHiC u+cUIwy/+qHnraAP8gd1r7n1i1HuudvLYEGByAjB78OMLV0gQZyC8c/iHeDPtDCSAUaOw1EUdve dbgekUWDIipxraJZyBg1aSYB7nV4DOMwlACNF7ub7NKfNkPQdAXdnjVl1me7ak3/ftC4fqrwPPl wPX7Jm9sbZexeDZFn3gPQPacv67u1ZkKUd9RcPtwVRSfJCzC9NrfRYHR9yPJWmAlcYqC0aUqJWb vwS5cWM8Qb4k3KYoKyMiuvhy36CRi1RTpYT6qQHwgU+ob8z3LFvIUm2Z+TGSFC2owoO0cJ10PAT VFJYij6ZXeVnEQSUWCHHiL05dOi7tnA9z0GZEJjlq4774/1j51lSnKai7pJMLPI1bRfh4Lnx4pz jpSdbyw8r9SrDaBgvrePYr9mdzU1FxQIXgcC/grjopEdycefxy+ypcnKFsCkDMk3dD85uNGiFBi wdIa86Ax2JMkQ== X-Received: by 2002:a17:903:22d1:b0:2ca:619f:9733 with SMTP id d9443c01a7336-2d0522233d3mr134076635ad.17.1785841081905; Tue, 04 Aug 2026 03:58:01 -0700 (PDT) Received: from DESKTOP-58PM2JF.localdomain ([116.84.110.160]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d0b6fbf2c6sm102515ad.34.2026.08.04.03.58.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 03:58:01 -0700 (PDT) From: Jinu Kim To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, x86@kernel.org Subject: [PATCH v2] KVM: x86/mmu: Write-protect tracked GFNs in all address spaces Date: Tue, 4 Aug 2026 19:57:55 +0900 Message-ID: <20260804105755.276646-1-kimjw04271234@gmail.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" KVM relies on write tracking to fault all subsequent guest CPU writes to a GFN that backs a shadow page. The write-protection installed when tracking starts is currently restricted to the supplied memslot. With SMM, the same backing page can be mapped through both x86 address spaces. If the peer address space already has a writable SPTE, a guest write through that mapping bypasses page tracking and leaves KVM's shadow state stale. On current mainline, changing a nested EPT PDE through the surviving SMM mapping leaves L2 using the old translation even after a valid INVEPT. Performing the same change through the tracked address space faults and updates the translation. Write-protect existing mappings in both x86 address spaces whenever KVM registers or synchronizes a tracked GFN. Revoking existing SPTEs is not sufficient. A later fault on another GFN in the same 2 MiB or 1 GiB region can recreate a writable huge SPTE over the tracked GFN. When selecting a mapping level, consult dynamic large-page restrictions in every peer memslot that overlaps the candidate huge-page range. Keep slot-layout and memory-attribute restrictions local to the address space that owns the mapping, and keep all counters in their owning memslots. This restores the invariant that a tracked GFN cannot remain, or become, CPU-writable through another x86 address space. Fixes: 699023e23965 ("KVM: x86: add SMM to the MMU role, support SMRAM addr= ess space") Cc: stable@vger.kernel.org Assisted-by: Codex:GPT-5 Signed-off-by: Jinu Kim --- Changes in v2: - Cover writable huge-SPTE recreation in addition to existing SPTEs. - Check every peer memslot that overlaps the candidate huge-page range. - Propagate only dynamic peer restrictions and keep per-slot accounting and intrinsic restrictions local. Testing: - The existing-SPTE reproducer produced before=3DA hidden=3DA hidden_invept=3DA control=3DB on stock v7.2-rc6, and before=3DA hidden=3DB hidden_invept=3DB control=3DB with this patch. - A huge-SPTE re-arm reproducer placed the tracked GFN one 4 KiB page into a 2 MiB mapping, then faulted a sibling GFN before writing the tracked GFN. An existing-SPTE-only fix left hidden=3DA hidden_invept=3DA; this patch produced hidden=3DB hidden_invept=3DB. - Re-ran the existing-SPTE reproducer with tdp_mmu=3DN; this patch produced before=3DA hidden=3DB hidden_invept=3DB control=3DB. - Ran all 103 tests in the default x86 KVM selftests collection. 65 passed, 35 skipped, and the three failures reproduced with the parent commit in the same nested test environment. - Ran all 87 tests in the default x86 kvm-unit-tests suite. 53 passed, 30 skipped, and the four failures reproduced with the parent commit in the same nested test environment. - Built and booted a matching bzImage, kvm.ko, kvm-intel.ko, and irqbypass.ko from commit e38548d14bfb. - Passed scripts/checkpatch.pl --strict and a W=3D1 build of kvm.ko and kvm-intel.ko. arch/x86/kvm/mmu.h | 11 +++++ arch/x86/kvm/mmu/mmu.c | 77 +++++++++++++++++++++++++++------ arch/x86/kvm/mmu/mmu_internal.h | 3 ++ arch/x86/kvm/mmu/page_track.c | 2 +- arch/x86/kvm/x86.c | 8 ++-- 5 files changed, 84 insertions(+), 17 deletions(-) diff --git a/arch/x86/kvm/mmu.h b/arch/x86/kvm/mmu.h index e1bb663ebbd58..2dd89fcba0aea 100644 --- a/arch/x86/kvm/mmu.h +++ b/arch/x86/kvm/mmu.h @@ -274,6 +274,17 @@ static inline bool kvm_memslots_have_rmaps(struct kvm = *kvm) return !tdp_mmu_enabled || kvm_shadow_root_allocated(kvm); } =20 +/* + * The upper bits of disallow_lpage describe restrictions that are intrins= ic + * to the memslot or its memory attributes. The lower bits refcount dynam= ic + * restrictions, e.g. shadowed or externally write-tracked GFNs. + */ +#define KVM_LPAGE_MIXED_FLAG BIT(31) +#define KVM_LPAGE_SLOT_DISALLOW_FLAG BIT(30) +#define KVM_LPAGE_DISALLOW_FLAGS (KVM_LPAGE_MIXED_FLAG | \ + KVM_LPAGE_SLOT_DISALLOW_FLAG) +#define KVM_LPAGE_DYNAMIC_DISALLOW_MASK GENMASK(29, 0) + static inline gfn_t gfn_to_index(gfn_t gfn, gfn_t base_gfn, int level) { /* KVM_HPAGE_GFN_SHIFT(PG_LEVEL_4K) must be 0. */ diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 66e69d2a41b3c..69e33140723c9 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -742,13 +742,43 @@ static bool kvm_gfn_is_lpage_allowed(struct kvm *kvm, return true; } =20 -/* - * The most significant bit in disallow_lpage tracks whether or not memory - * attributes are mixed, i.e. not identical for all gfns at the current le= vel. - * The lower order bits are used to refcount other cases where a hugepage = is - * disallowed, e.g. if KVM has shadow a page table at the gfn. - */ -#define KVM_LPAGE_MIXED_FLAG BIT(31) +static bool kvm_gfn_is_lpage_allowed_for_mapping(struct kvm *kvm, + const struct kvm_memory_slot *slot, + gfn_t gfn, int level) +{ + const struct kvm_memory_slot *other_slot; + struct kvm_memslot_iter iter; + struct kvm_memslots *slots; + gfn_t start, end; + + if (lpage_info_slot(gfn, slot, level)->disallow_lpage) + return false; + + if (kvm_arch_nr_memslot_as_ids(kvm) =3D=3D 1) + return true; + + start =3D gfn_round_for_level(gfn, level); + end =3D start + KVM_PAGES_PER_HPAGE(level); + slots =3D __kvm_memslots(kvm, slot->as_id ^ 1); + + if (kvm_memslots_empty(slots)) + return true; + + other_slot =3D __gfn_to_memslot(slots, start); + if (other_slot && other_slot->base_gfn + other_slot->npages >=3D end) + return !(lpage_info_slot(start, other_slot, level)->disallow_lpage & + KVM_LPAGE_DYNAMIC_DISALLOW_MASK); + + kvm_for_each_memslot_in_gfn_range(&iter, slots, start, end) { + gfn_t slot_gfn =3D max(start, iter.slot->base_gfn); + + if (lpage_info_slot(slot_gfn, iter.slot, level)->disallow_lpage & + KVM_LPAGE_DYNAMIC_DISALLOW_MASK) + return false; + } + + return true; +} =20 static void update_gfn_disallow_lpage_count(const struct kvm_memory_slot *= slot, gfn_t gfn, int count) @@ -761,7 +791,8 @@ static void update_gfn_disallow_lpage_count(const struc= t kvm_memory_slot *slot, =20 old =3D linfo->disallow_lpage; linfo->disallow_lpage +=3D count; - WARN_ON_ONCE((old ^ linfo->disallow_lpage) & KVM_LPAGE_MIXED_FLAG); + WARN_ON_ONCE((old ^ linfo->disallow_lpage) & + KVM_LPAGE_DISALLOW_FLAGS); } } =20 @@ -801,7 +832,7 @@ static void account_shadowed(struct kvm *kvm, struct kv= m_mmu_page *sp) =20 kvm_mmu_gfn_disallow_lpage(slot, gfn); =20 - if (kvm_mmu_slot_gfn_write_protect(kvm, slot, gfn, PG_LEVEL_4K)) + if (kvm_mmu_gfn_write_protect(kvm, slot, gfn, PG_LEVEL_4K)) kvm_flush_remote_tlbs_gfn(kvm, gfn, PG_LEVEL_4K); } =20 @@ -1510,12 +1541,33 @@ bool kvm_mmu_slot_gfn_write_protect(struct kvm *kvm, return write_protected; } =20 +bool kvm_mmu_gfn_write_protect(struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + int min_level) +{ + struct kvm_memory_slot *other_slot; + bool write_protected; + + BUILD_BUG_ON(KVM_MAX_NR_ADDRESS_SPACES > 2); + + write_protected =3D kvm_mmu_slot_gfn_write_protect(kvm, slot, gfn, min_le= vel); + if (kvm_arch_nr_memslot_as_ids(kvm) > 1) { + other_slot =3D __gfn_to_memslot(__kvm_memslots(kvm, slot->as_id ^ 1), gf= n); + if (other_slot) + write_protected |=3D kvm_mmu_slot_gfn_write_protect(kvm, + other_slot, + gfn, min_level); + } + + return write_protected; +} + static bool kvm_vcpu_write_protect_gfn(struct kvm_vcpu *vcpu, u64 gfn) { struct kvm_memory_slot *slot; =20 slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); - return kvm_mmu_slot_gfn_write_protect(vcpu->kvm, slot, gfn, PG_LEVEL_4K); + return kvm_mmu_gfn_write_protect(vcpu->kvm, slot, gfn, PG_LEVEL_4K); } =20 static bool kvm_zap_rmap(struct kvm *kvm, struct kvm_rmap_head *rmap_head, @@ -3396,7 +3448,6 @@ static u8 kvm_gmem_max_mapping_level(struct kvm *kvm,= struct kvm_page_fault *fau int kvm_mmu_max_mapping_level(struct kvm *kvm, struct kvm_page_fault *faul= t, const struct kvm_memory_slot *slot, gfn_t gfn) { - struct kvm_lpage_info *linfo; int host_level, max_level; bool is_private; =20 @@ -3412,8 +3463,8 @@ int kvm_mmu_max_mapping_level(struct kvm *kvm, struct= kvm_page_fault *fault, =20 max_level =3D min(max_level, max_huge_page_level); for ( ; max_level > PG_LEVEL_4K; max_level--) { - linfo =3D lpage_info_slot(gfn, slot, max_level); - if (!linfo->disallow_lpage) + if (kvm_gfn_is_lpage_allowed_for_mapping(kvm, slot, gfn, + max_level)) break; } =20 diff --git a/arch/x86/kvm/mmu/mmu_internal.h b/arch/x86/kvm/mmu/mmu_interna= l.h index 73cdcbccc89e8..424ff75433582 100644 --- a/arch/x86/kvm/mmu/mmu_internal.h +++ b/arch/x86/kvm/mmu/mmu_internal.h @@ -207,6 +207,9 @@ void kvm_mmu_gfn_allow_lpage(const struct kvm_memory_sl= ot *slot, gfn_t gfn); bool kvm_mmu_slot_gfn_write_protect(struct kvm *kvm, struct kvm_memory_slot *slot, u64 gfn, int min_level); +bool kvm_mmu_gfn_write_protect(struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + int min_level); =20 /* Flush the given page (huge or not) of guest memory. */ static inline void kvm_flush_remote_tlbs_gfn(struct kvm *kvm, gfn_t gfn, i= nt level) diff --git a/arch/x86/kvm/mmu/page_track.c b/arch/x86/kvm/mmu/page_track.c index 7e8195a311bb0..f32caafcca8db 100644 --- a/arch/x86/kvm/mmu/page_track.c +++ b/arch/x86/kvm/mmu/page_track.c @@ -106,7 +106,7 @@ void __kvm_write_track_add_gfn(struct kvm *kvm, struct = kvm_memory_slot *slot, */ kvm_mmu_gfn_disallow_lpage(slot, gfn); =20 - if (kvm_mmu_slot_gfn_write_protect(kvm, slot, gfn, PG_LEVEL_4K)) + if (kvm_mmu_gfn_write_protect(kvm, slot, gfn, PG_LEVEL_4K)) kvm_flush_remote_tlbs(kvm); } =20 diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 47cb9eba113b1..bfe350567cea3 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -13557,9 +13557,10 @@ static int kvm_alloc_memslot_metadata(struct kvm *= kvm, slot->arch.lpage_info[i - 1] =3D linfo; =20 if (slot->base_gfn & (KVM_PAGES_PER_HPAGE(level) - 1)) - linfo[0].disallow_lpage =3D 1; + linfo[0].disallow_lpage =3D KVM_LPAGE_SLOT_DISALLOW_FLAG; if ((slot->base_gfn + npages) & (KVM_PAGES_PER_HPAGE(level) - 1)) - linfo[lpages - 1].disallow_lpage =3D 1; + linfo[lpages - 1].disallow_lpage =3D + KVM_LPAGE_SLOT_DISALLOW_FLAG; ugfn =3D slot->userspace_addr >> PAGE_SHIFT; /* * If the gfn and userspace address are not aligned wrt each @@ -13569,7 +13570,8 @@ static int kvm_alloc_memslot_metadata(struct kvm *k= vm, unsigned long j; =20 for (j =3D 0; j < lpages; ++j) - linfo[j].disallow_lpage =3D 1; + linfo[j].disallow_lpage =3D + KVM_LPAGE_SLOT_DISALLOW_FLAG; } } =20 base-commit: 848acc8ffe1b7cd5f1bf427b93069becfebc2c9d --=20 2.43.0