From nobody Thu Sep 24 13:37:22 2026 Received: from mail-wm1-f45.google.com (mail-wm1-f45.google.com [209.85.128.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DC5E2382287 for ; Fri, 12 Jun 2026 16:24:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281458; cv=none; b=VvQfWvdqUWO5Rj7KDZ8mcMoTLBqT0p+WvLEnIdSfh9yaFBdkJMu6j0USB+Iy8Q+Y4Kt7B8t2cwZ7yw481APJ3HWH0zvtwHwvwWO1WY2xvR8rVJifndGb3vHokZuDZmor+cpdIhaMPVu+j7gVvaARuW7SjQUf1FEoaxJxRnZdoWE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281458; c=relaxed/simple; bh=Ek5sg8ia7dHrpdcCzHmFxOvsiqSRhQeTCIZ4Fbont5Q=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=i4VCR9pNKeoTUKBlpGa9kIJTpBByNQYGSYaR2/nuABcsUoIhFbXhuy8hxXu4R7OK1IYF6jRJmK1RtiDYD5xlEuTptc/cKjGCw3mjJTfsvYKBrWSgkJTX6WimdrCUrtIyvvvrVdquxEnQAEaHDUgqmINc09U1s2FKtEFMB9RvJE0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=io+CayzF; arc=none smtp.client-ip=209.85.128.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="io+CayzF" Received: by mail-wm1-f45.google.com with SMTP id 5b1f17b1804b1-490bb83a3f6so9578855e9.0 for ; Fri, 12 Jun 2026 09:24:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781281454; x=1781886254; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=tdYyfsvx1DQaEQW7SfjjNGQD4l8zjElTcwX+OTVhZvA=; b=io+CayzFz4Uk2FNGC95HpAvXh/oJq7US3dD/KCDftdTO/QmdQj9yAudNKO1NfMCRAv qcics5FowaDmK3SpF7CS3UI0vuLK9wuQa//jYtfMhFz380Q5MX4aOrAR/vbShXHc8QYc uSxy+gKK3DVpit+m/+sV5UPkoRqbgNQXXzV/w87643eLyyR9dQtxRZprlFCKlzb7p9JN 0NiLKPN0+xiCyqBCcEuaEaq3kmBXJGnHhfHDGJBAErwxj8RoBD5q/XmUjzPIOFUAeEKF 6xqJ46NxzV23u1A2m5R1OvLI2MlVzSl/vauAvo99XPEhpTJI9InkjjMl6aAQwl/eKO/m wM7Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781281454; x=1781886254; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=tdYyfsvx1DQaEQW7SfjjNGQD4l8zjElTcwX+OTVhZvA=; b=EXUMg0R4i3kPKEeaGv+8TJeaNysxcj3GYpgTlDJNYQwnOJ7GwMN9EtMRs0QInAHCk1 j9ug5yF5yO5TyI0sMRUUWtzuYnPqmvvDnDodnjLfDS2GL5m9Uznl+Dnu1FnFPLuZT53K CYb09kQRk5vKjHgMM0GcqfoWxpzg6Ag1pwikwbzVbWdCFnx2VjOMCAWvpdUV9nWXFVKw xuBnbTknoIln/itOCn8C8XT9/Msg6fE3eyOtQY64V3RB2a0/KQidkLlYkIfQESBWbeRY GhdIoDlA9AnmPERj7tLL66xfbjLH9f3+bvCt60Rz4EdRYD9atRUWn4w8QsVoHLJ1IEhu R6uA== X-Forwarded-Encrypted: i=1; AFNElJ95fijJGh5AF0rTzMAjk1iSxxpD0S/zRD+pedSlR0iOJ8Hw7jetVi+kCKRcPB7Erv32vQ0t2G4vvfb/zYo=@vger.kernel.org X-Gm-Message-State: AOJu0YxLjh2m9eylUYsiCdud8E2E484l9B6XTk9N3GxueUErvN17vgy5 GoXdpAjhZcpL4hRd3F32edoNBeF9OKpsS+rWVej9My5ncD78GC/XPq/D X-Gm-Gg: Acq92OGe1rCE5buPKqT0jTcXneSahY2KZ4CxUpP9gJ5gxz1j7aD/GMSS1pfNjJ+CIlO uxBLAPmqV7+l+xOfLg/pvGEYV4zB6vJupnxk7DMmwl2dkz3zTYYqjyhrQiYY+JYYY//ikKN57v2 ngNabayh6L3gWLAtc/aCwVfsTqKKucMbhHsr/kMfc5kkzppCAx0NCkWiVTACTGkp20jcMi+MQz+ 5Uxsh7LgxApFS/XqQxq6YgzgVsSoJUZYWr1cMLT2K8nugXJSNwWS6JcMRvgWIQVWemmY7Z0z0rs 9tV9d0hVEthP7FM+HGf7TdnOBuflvqjFc3m6CJiFvNKrxZoMzVaXrSdgE0iFzw52blj/DYGGD9E hmtEHsmgjqKHlbVenV+pB5sdGJVpmiPoxA60MBm6jCCGNfZ5jihq0zCxEnP+S8+Ep9y34jQ/kD9 gLOIZUsMvxkg1/ihRHAO34pySuVQdR4q/vozNiTyEQMwxPBs0xgq58xv2+nMgxhw== X-Received: by 2002:a05:600c:3548:b0:490:ea8a:32d0 with SMTP id 5b1f17b1804b1-490ec501917mr49869245e9.20.1781281454334; Fri, 12 Jun 2026 09:24:14 -0700 (PDT) Received: from f4d4888f22f2.ant.amazon.com.com ([15.248.2.31]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490ea95c51dsm57620935e9.1.2026.06.12.09.24.13 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 12 Jun 2026 09:24:13 -0700 (PDT) From: Jack Thomson To: maz@kernel.org, oupton@kernel.org, pbonzini@redhat.com Cc: joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, shuah@kernel.org, corbet@lwn.net, vladimir.murzin@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, isaku.yamahata@intel.com, Jack Thomson Subject: [PATCH v5 1/5] KVM: arm64: Pass walk flags to kvm_pgtable_get_leaf() Date: Fri, 12 Jun 2026 17:23:49 +0100 Message-ID: <20260612162354.73378-2-jackabt.amazon@gmail.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260612162354.73378-1-jackabt.amazon@gmail.com> References: <20260612162354.73378-1-jackabt.amazon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jack Thomson Allow callers of kvm_pgtable_get_leaf() to specify the page-table walk flags, in preparation for performing walks under the MMU read lock. Reading a stage-2 leaf while only holding the read lock requires KVM_PGTABLE_WALK_SHARED: parallel faults (which also only hold the read lock) can unlink table pages and free them via RCU, so the walker must be inside an RCU read-side critical section, which the shared walk flag provides via kvm_pgtable_walk_begin(). All existing callers either hold the write lock, walk with interrupts disabled, or run at hyp where shared walks are rejected; they keep the current behaviour by passing no flags. No functional change intended. Signed-off-by: Jack Thomson --- arch/arm64/include/asm/kvm_pgtable.h | 5 ++++- arch/arm64/kvm/hyp/nvhe/mem_protect.c | 10 +++++----- arch/arm64/kvm/hyp/pgtable.c | 5 +++-- arch/arm64/kvm/mmu.c | 2 +- arch/arm64/kvm/nested.c | 2 +- 5 files changed, 14 insertions(+), 10 deletions(-) diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/= kvm_pgtable.h index 41a8687938eb..d0167f7dfbee 100644 --- a/arch/arm64/include/asm/kvm_pgtable.h +++ b/arch/arm64/include/asm/kvm_pgtable.h @@ -859,6 +859,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr,= u64 size, * @addr: Input address for the start of the walk. * @ptep: Pointer to storage for the retrieved PTE. * @level: Pointer to storage for the level of the retrieved PTE. + * @flags: Flags to control the page-table walk + * (see struct kvm_pgtable_visit_ctx). * * The offset of @addr within a page is ignored. * @@ -869,7 +871,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr,= u64 size, * Return: 0 on success, negative error code on failure. */ int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr, - kvm_pte_t *ptep, s8 *level); + kvm_pte_t *ptep, s8 *level, + enum kvm_pgtable_walk_flags flags); =20 /** * kvm_pgtable_stage2_pte_prot() - Retrieve the protection attributes of a diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvh= e/mem_protect.c index 25f04629014e..3b765c9ff7e8 100644 --- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c +++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c @@ -522,7 +522,7 @@ static int host_stage2_adjust_range(u64 addr, struct kv= m_mem_range *range) int ret; =20 hyp_assert_lock_held(&host_mmu.lock); - ret =3D kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level, 0); if (ret) return ret; =20 @@ -890,7 +890,7 @@ static int get_valid_guest_pte(struct pkvm_hyp_vm *vm, = u64 ipa, kvm_pte_t *ptep, s8 level; int ret; =20 - ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0); if (ret) return ret; if (guest_pte_is_poisoned(pte)) @@ -939,7 +939,7 @@ int __pkvm_vcpu_in_poison_fault(struct pkvm_hyp_vcpu *h= yp_vcpu) ipa |=3D FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(&hyp_vcpu->vcpu)); =20 guest_lock_component(vm); - ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0); if (ret) goto unlock; =20 @@ -1293,7 +1293,7 @@ static int host_stage2_get_guest_info(phys_addr_t phy= s, struct pkvm_hyp_vm **vm, return -EPERM; } =20 - ret =3D kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level, 0); if (ret) return ret; =20 @@ -1522,7 +1522,7 @@ static int __check_host_shared_guest(struct pkvm_hyp_= vm *vm, u64 *__phys, u64 ip s8 level; int ret; =20 - ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0); if (ret) return ret; if (!kvm_pte_valid(pte)) diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c index 0c1defa5fb0f..6a839a32e246 100644 --- a/arch/arm64/kvm/hyp/pgtable.c +++ b/arch/arm64/kvm/hyp/pgtable.c @@ -298,12 +298,13 @@ static int leaf_walker(const struct kvm_pgtable_visit= _ctx *ctx, } =20 int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr, - kvm_pte_t *ptep, s8 *level) + kvm_pte_t *ptep, s8 *level, + enum kvm_pgtable_walk_flags flags) { struct leaf_walk_data data; struct kvm_pgtable_walker walker =3D { .cb =3D leaf_walker, - .flags =3D KVM_PGTABLE_WALK_LEAF, + .flags =3D flags | KVM_PGTABLE_WALK_LEAF, .arg =3D &data, }; int ret; diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c index 4da9281312eb..c720f07cb82e 100644 --- a/arch/arm64/kvm/mmu.c +++ b/arch/arm64/kvm/mmu.c @@ -839,7 +839,7 @@ static int get_user_mapping_size(struct kvm *kvm, u64 a= ddr) * IPI-ing threads). */ local_irq_save(flags); - ret =3D kvm_pgtable_get_leaf(&pgt, addr, &pte, &level); + ret =3D kvm_pgtable_get_leaf(&pgt, addr, &pte, &level, 0); local_irq_restore(flags); =20 if (ret) diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c index 38f672e94087..e45aed6d9e65 100644 --- a/arch/arm64/kvm/nested.c +++ b/arch/arm64/kvm/nested.c @@ -559,7 +559,7 @@ static u8 get_guest_mapping_ttl(struct kvm_s2_mmu *mmu,= u64 addr) return 0; =20 tmp &=3D ~(sz - 1); - if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL)) + if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL, 0)) goto again; if (!(pte & PTE_VALID)) goto again; --=20 2.43.0 From nobody Thu Sep 24 13:37:22 2026 Received: from mail-wm1-f53.google.com (mail-wm1-f53.google.com [209.85.128.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6DE0C3D5657 for ; Fri, 12 Jun 2026 16:24:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281460; cv=none; b=TiBHHMKJsNQEazB6WxJHn+PGWjIXZhcnZMr+yboI0q20xeqWsYKBTE3YDK5ivU+GQmPbdc86O7kESgfOR4X2LRQCZ6ntRbVzNfA4Mhx0piTMgaJQgQyuJiD2BZBg59VumJcnzVsnMCqqrXPwT5JMXNJrcpxXIuoXEqqgG0v5J38= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281460; c=relaxed/simple; bh=QRDRXTCX50amxKD5gLh7D6OBaXCfaBfzgoB4N+6bbzg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=QdLntKI/SOXfxcHiLCmmr8+fRy+pBJ8sb08rnTjNC7o0JlAHmnKJEQpV3EmodEQreaSbC9Zwy+gH/itRVxjqrW2PNce9oZDKsKddLSihK5srHTs1gomw1WJCHfhwCpG8BDA6oN07VwpzabE4WUzfE+ccj9mm1/SdmCk+685xfS8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NoBrRehd; arc=none smtp.client-ip=209.85.128.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NoBrRehd" Received: by mail-wm1-f53.google.com with SMTP id 5b1f17b1804b1-490ac10e337so7787675e9.3 for ; Fri, 12 Jun 2026 09:24:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781281456; x=1781886256; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=mJahK7oX1jpIBSr/85+UyDd//PxGmbpfN024eZe1Ebc=; b=NoBrRehdDwJVQdKa0C37zfYA0hO1eLZxOeIoJgnQZPq8ucrodATeQo4X4YpaNNQ0Md osOLlTZeLA1VzYGwGsJ1wpA7PP5QDgn+Nrsv8nw9J1odCsDJegjQPp3SMJ08I2uXX1F6 QYz4/D/vh4dg9d3gkQGcPd/glSiOyVYmDIE8MKfJskbXNXTJei7HetJpnJCe+uAYMUkF pHlPu51NPntaMF7PAHlWoxozVaeoxn9wFg5js6aWzEvv/Yb0bmsgxnd0iab5CZjazxI6 UzVV+9zw/Zs35zUxaIgihk+CTgZEjHE08QS1KiRmzBNfSsgPUFrpL9Gx8qCf3XFxTkPi dlWA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781281456; x=1781886256; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=mJahK7oX1jpIBSr/85+UyDd//PxGmbpfN024eZe1Ebc=; b=Nb1RyK3mnjZm34pNPR7DSX/D0pCSuBt8aZ0z4Sa3jMEoymlgVNLw5A0tAILLuAkqCo VfDC48JRgbhMpBeO3t9hU5AqmDBh58QDAaiXtVxzY9Nd6bFR1DeMTaM3pGINRZF7bUNT hyOIncjduv9YUi/eiB8M8o69Z5p6nKk82bsIt3ivzpDHkg5drMjnaLDMbs0v8vY8MJnS 333Svz7C9sgDReVgbhTb6twIQklZkhL2ulBcbmhE9JkDIPmd5dkZ345q5KnW+XF69/oj fiBKT/KdMJ94tNg7E5FwBHHD2CZpa0yoVXFb1mg97txI7XTk5j7pYsZO4uG2TVq9/egl OLzw== X-Forwarded-Encrypted: i=1; AFNElJ8X+7Lrhhl5MzJUU0XLCUGxcXQFtfFH82nPpD24iKUCFisLDh6g+cTWqwH3DYKvS24ppYJBISDgmTDZg+Q=@vger.kernel.org X-Gm-Message-State: AOJu0YyhqqBfURhY4MR0AjhCBJHDLtBniqg3AQDPKUqiMujtB+7GNkaG fOWYxXjYzYhO6VwJDiIsicsyEDLF+7SD0neC8yhdR1hslUhfMJosPkPc X-Gm-Gg: Acq92OE5I1qJJet/d4jdg6FX1lImbW2n5G+vk70PsZ4M+717djR6QKwvXNAs5dIk4Ez LVA7jaaHWPW9geoxbP9fnoVzAdSwCYnFKOPNQyWSjY/bi3QqhXr73q0CMuWV4O5lJAk9SXjB/Tv DyVHy7lxBJ7qodLBAZguc6CjBMbQAPoCXe/SanCJCEbNik0av7aXAm9wQ9km/MivQwcuTawOKUP oHz8i1lk3si0FBSdC1e2VYVpekzDzz3ptaqwSeBwYqrdq2K+vDDgbEx0pVOfSQi3d1b425RHglN 6WwVau3qrf5zcay5zITIPhGjqhvPSjr1P7msy2BXCBkGLXcgweCWfNpVTkCe8sEeQsdHLo+FkSa 0pKQViOAs7bQ9b3w4vmQCQkxpo390YQkv50gpvpdZ7oKcNI/ZGQoYyXOffn7SuRuusr4RekN1Wo LffOjV4CNJsTW7WueFqbP+oqnT59Am5uJvcCvssXKMc0qTj7aEuQ3ibRmL5iB2FQ== X-Received: by 2002:a05:600c:4745:b0:490:bd66:e523 with SMTP id 5b1f17b1804b1-492200c04a9mr1802785e9.20.1781281455606; Fri, 12 Jun 2026 09:24:15 -0700 (PDT) Received: from f4d4888f22f2.ant.amazon.com.com ([15.248.2.31]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490ea95c51dsm57620935e9.1.2026.06.12.09.24.14 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 12 Jun 2026 09:24:15 -0700 (PDT) From: Jack Thomson To: maz@kernel.org, oupton@kernel.org, pbonzini@redhat.com Cc: joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, shuah@kernel.org, corbet@lwn.net, vladimir.murzin@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, isaku.yamahata@intel.com, Jack Thomson Subject: [PATCH v5 2/5] KVM: arm64: Add pre_fault_memory implementation Date: Fri, 12 Jun 2026 17:23:50 +0100 Message-ID: <20260612162354.73378-3-jackabt.amazon@gmail.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260612162354.73378-1-jackabt.amazon@gmail.com> References: <20260612162354.73378-1-jackabt.amazon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jack Thomson Add arm64 support for KVM_PRE_FAULT_MEMORY by synthesizing a read data abort and routing it through the existing stage-2 fault handlers. Treat the requested GPA as an IPA in the userspace-owned VM's memslot space and always target the canonical stage-2, even if the vCPU last ran with a nested/shadow MMU selected. If the vCPU last ran in a nested context, switch to the canonical stage-2 with the vCPU put/load helpers so VMID, VNCR and shadow-MMU refcount state stay consistent. Leave the switch in place for the ioctl; vcpu_put() at ioctl exit drops the hw_mmu and the next vcpu_load() reselects the correct MMU from vCPU state. Check existing mappings with a shared page-table walk under the MMU read lock, and use the resulting walk level when constructing the synthetic fault. Report poisoned pages through the ioctl return path with -EHWPOISON instead of also queueing SIGBUS, and use the installed mapping size to advance the prefault range. Advertise KVM_CAP_PRE_FAULT_MEMORY on arm64. Protected VMs remain unsupported: pKVM filters the capability, and the ioctl returns -EOPNOTSUPP if invoked anyway. Signed-off-by: Jack Thomson --- Documentation/virt/kvm/api.rst | 18 +++- arch/arm64/kvm/Kconfig | 1 + arch/arm64/kvm/arm.c | 1 + arch/arm64/kvm/mmu.c | 162 +++++++++++++++++++++++++++++++++ 4 files changed, 178 insertions(+), 4 deletions(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index 52bbbb553ce1..657e05656fa6 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -6462,7 +6462,7 @@ See KVM_SET_USER_MEMORY_REGION2 for additional detail= s. --------------------------- =20 :Capability: KVM_CAP_PRE_FAULT_MEMORY -:Architectures: none +:Architectures: x86, arm64 :Type: vcpu ioctl :Parameters: struct kvm_pre_fault_memory (in/out) :Returns: 0 if at least one page is processed, < 0 on error @@ -6470,11 +6470,14 @@ See KVM_SET_USER_MEMORY_REGION2 for additional deta= ils. Errors: =20 =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + EAGAIN A memslot update raced with the ioctl before any page was + processed. EINVAL The specified `gpa` and `size` were invalid (e.g. not page aligned, causes an overflow, or size is zero). ENOENT The specified `gpa` is outside defined memslots. EINTR An unmasked signal is pending and no page was processed. EFAULT The parameter address was invalid. + EHWPOISON A poisoned host page was encountered. EOPNOTSUPP Mapping memory for a GPA is unsupported by the hypervisor, and/or for the current vCPU state/mode. EIO unexpected error conditions (also causes a WARN) @@ -6494,7 +6497,14 @@ Errors: KVM_PRE_FAULT_MEMORY populates KVM's stage-2 page tables used to map memory for the current vCPU state. KVM maps memory as if the vCPU generated a stage-2 read page fault, e.g. faults in memory as needed, but doesn't break -CoW. However, KVM does not mark any newly created stage-2 PTE as Accessed. +CoW. However, on x86, KVM does not mark any newly created stage-2 PTE as +Accessed. On arm64, newly created stage-2 PTEs are marked Accessed. + +On arm64, `gpa` is interpreted as an IPA in the userspace-owned VM's +memslot address space. If the vCPU most recently ran a nested guest, KVM +still targets the VM's canonical stage-2, and does not interpret `gpa` as +a nested guest IPA or target the nested/shadow stage-2 selected by the +vCPU's last run state. =20 In the case of confidential VM types where there is an initial set up of private guest memory before the guest is 'finalized'/measured, this ioctl @@ -6507,9 +6517,9 @@ case, the ioctl can be called in parallel. =20 When the ioctl returns, the input values are updated to point to the remaining range. If `size` > 0 on return, the caller can just issue -the ioctl again with the same `struct kvm_map_memory` argument. +the ioctl again with the same `struct kvm_pre_fault_memory` argument. =20 -Shadow page tables cannot support this ioctl because they +On x86, shadow page tables cannot support this ioctl because they are indexed by virtual address or nested guest physical address. Calling this ioctl when the guest is using shadow page tables (for example because it is running a nested guest with nested page tables) diff --git a/arch/arm64/kvm/Kconfig b/arch/arm64/kvm/Kconfig index 449154f9a485..6b89262e8ba7 100644 --- a/arch/arm64/kvm/Kconfig +++ b/arch/arm64/kvm/Kconfig @@ -24,6 +24,7 @@ menuconfig KVM select HAVE_KVM_CPU_RELAX_INTERCEPT select KVM_MMIO select KVM_GENERIC_DIRTYLOG_READ_PROTECT + select KVM_GENERIC_PRE_FAULT_MEMORY select VIRT_XFER_TO_GUEST_WORK select KVM_VFIO select HAVE_KVM_DIRTY_RING_ACQ_REL diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 9453321ef8c6..dcb92bee13af 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -392,6 +392,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long = ext) case KVM_CAP_COUNTER_OFFSET: case KVM_CAP_ARM_WRITABLE_IMP_ID_REGS: case KVM_CAP_ARM_SEA_TO_USER: + case KVM_CAP_PRE_FAULT_MEMORY: r =3D 1; break; case KVM_CAP_SET_GUEST_DEBUG2: diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c index c720f07cb82e..4bf048bbcf8b 100644 --- a/arch/arm64/kvm/mmu.c +++ b/arch/arm64/kvm/mmu.c @@ -1571,6 +1571,8 @@ struct kvm_s2_fault_desc { struct kvm_s2_trans *nested; struct kvm_memory_slot *memslot; unsigned long hva; + unsigned long *page_size; + bool prefault; }; =20 static int gmem_abort(const struct kvm_s2_fault_desc *s2fd) @@ -1882,6 +1884,13 @@ static int kvm_s2_fault_pin_pfn(const struct kvm_s2_= fault_desc *s2fd, &s2vi->map_writable, &s2vi->page); if (unlikely(is_error_noslot_pfn(s2vi->pfn))) { if (s2vi->pfn =3D=3D KVM_PFN_ERR_HWPOISON) { + /* + * When prefaulting, report the poison via -EHWPOISON + * only; don't also queue a SIGBUS as the run path + * does for the faulting vCPU thread. + */ + if (s2fd->prefault) + return -EHWPOISON; kvm_send_hwpoison_signal(s2fd->hva, __ffs(s2vi->vma_pagesize)); return 0; } @@ -2053,6 +2062,9 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault= _desc *s2fd, kvm_release_faultin_page(kvm, s2vi->page, !!ret, writable); kvm_fault_unlock(kvm); =20 + if (s2fd->page_size && !ret) + *s2fd->page_size =3D mapping_size; + /* * Mark the page dirty only if the fault is handled successfully, * making sure we adjust the canonical IPA if the mapping size has @@ -2757,3 +2769,153 @@ void kvm_toggle_cache(struct kvm_vcpu *vcpu, bool w= as_enabled) =20 trace_kvm_toggle_cache(*vcpu_pc(vcpu), was_enabled, now_enabled); } + +/* + * Prefaulting always targets the canonical stage-2. If the vCPU last ran + * in a nested context, swap in the canonical MMU via the vCPU put/load + * helpers so that preemption, VMID, VNCR fixmap and shadow-MMU refcount + * state stay consistent. + * + * The swap is deliberately not undone: nothing runs in between the + * per-page invocations of kvm_arch_vcpu_pre_fault_memory() except the + * generic prefault loop, and the vcpu_put() at ioctl exit discards + * vcpu->arch.hw_mmu anyway (see kvm_vcpu_put_hw_mmu()), so the next + * vcpu_load() re-derives the correct MMU from the vCPU's context. If the + * prefault task is preempted in the meantime, kvm_vcpu_put_hw_mmu() + * keeps the canonical MMU in place for the reload. Leaving the swap in + * place also bounds the cost to at most one put/load pair per ioctl, + * rather than two pairs per prefaulted page. + */ +static void kvm_pre_fault_load_canonical_mmu(struct kvm_vcpu *vcpu) +{ + if (!vcpu_has_nv(vcpu) || vcpu->arch.hw_mmu =3D=3D &vcpu->kvm->arch.mmu) + return; + + preempt_disable(); + kvm_arch_vcpu_put(vcpu); + vcpu->arch.hw_mmu =3D &vcpu->kvm->arch.mmu; + kvm_arch_vcpu_load(vcpu, smp_processor_id()); + preempt_enable(); +} + +long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, + struct kvm_pre_fault_memory *range) +{ + struct kvm_vcpu_fault_info *fault_info =3D &vcpu->arch.fault; + struct kvm_vcpu_fault_info fault_backup =3D *fault_info; + s8 walk_level =3D KVM_PGTABLE_LAST_LEVEL; + unsigned long page_size =3D PAGE_SIZE; + struct kvm_memory_slot *memslot; + phys_addr_t gpa =3D range->gpa; + struct kvm_pgtable *pgt; + phys_addr_t end; + kvm_pte_t pte; + hva_t hva; + gfn_t gfn; + long ret; + + if (vcpu_is_protected(vcpu)) + return -EOPNOTSUPP; + + /* + * Interpret range->gpa in the userspace-owned VM's IPA space, not in + * any nested guest IPA space that may have been active on the vCPU's + * last run. Always target the canonical stage-2. + */ + kvm_pre_fault_load_canonical_mmu(vcpu); + + if (gpa >=3D kvm_phys_size(vcpu->arch.hw_mmu)) { + ret =3D -ENOENT; + goto out; + } + + gfn =3D gpa_to_gfn(gpa); + memslot =3D gfn_to_memslot(vcpu->kvm, gfn); + if (!memslot) { + ret =3D -ENOENT; + goto out; + } + + /* + * A racing memslot deletion or move installs an invalid slot before + * zapping stage-2. Ask userspace to retry once the update settles. + */ + if (memslot->flags & KVM_MEMSLOT_INVALID) { + ret =3D -EAGAIN; + goto out; + } + + /* + * pKVM stage-2 mappings aren't directly walkable from the host; let + * the fault path handle both new and existing mappings. + */ + if (!is_protected_kvm_enabled()) { + pgt =3D vcpu->arch.hw_mmu->pgt; + scoped_guard(read_lock, &vcpu->kvm->mmu_lock) { + ret =3D kvm_pgtable_get_leaf(pgt, gpa, &pte, &walk_level, + KVM_PGTABLE_WALK_SHARED); + } + if (ret) + goto out; + + if (kvm_pte_valid(pte)) { + page_size =3D kvm_granule_size(walk_level); + if (!(pte & KVM_PTE_LEAF_ATTR_LO_S2_AF)) + handle_access_fault(vcpu, gpa); + goto out_success; + } + } + + /* + * Synthesize a read translation fault for the canonical IPA, at the + * level where the stage-2 walk currently ends (the last level under + * pKVM, where stage-2 isn't walkable from the host). + */ + fault_info->esr_el2 =3D (ESR_ELx_EC_DABT_LOW << ESR_ELx_EC_SHIFT) | + ESR_ELx_IL | ESR_ELx_FSC_FAULT_L(walk_level); + fault_info->hpfar_el2 =3D HPFAR_EL2_NS | + FIELD_PREP(HPFAR_EL2_FIPA, gpa >> 12); + + struct kvm_s2_fault_desc s2fd =3D { + .vcpu =3D vcpu, + .fault_ipa =3D gpa, + .nested =3D NULL, + .memslot =3D memslot, + .page_size =3D &page_size, + .prefault =3D true, + }; + + /* + * As in the run path, -EAGAIN from the abort handlers is treated as + * progress: either a parallel fault installed the mapping, or a racing + * invalidation is in flight and the next access will refault. + */ + if (kvm_slot_has_gmem(memslot)) { + ret =3D gmem_abort(&s2fd); + } else { + hva =3D gfn_to_hva_memslot_prot(memslot, gfn, NULL); + if (kvm_is_error_hva(hva)) { + ret =3D -EFAULT; + goto out; + } + + s2fd.hva =3D hva; + ret =3D user_mem_abort(&s2fd); + } + + if (ret < 0) + goto out; + +out_success: + end =3D ALIGN_DOWN(gpa, page_size) + page_size; + ret =3D min_t(u64, range->size, end - gpa); +out: + /* + * Restore the synthetic fault state so a subsequent KVM_RUN does not + * observe it. kvm_handle_mmio_return() runs before guest entry can + * refresh fault.esr_el2 from hardware, so leaving the synthetic ESR + * in place would corrupt the completion of a pending MMIO exit. + */ + *fault_info =3D fault_backup; + return ret; +} --=20 2.43.0 From nobody Thu Sep 24 13:37:22 2026 Received: from mail-wm1-f42.google.com (mail-wm1-f42.google.com [209.85.128.42]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B03FF3859F5 for ; Fri, 12 Jun 2026 16:24:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.42 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281462; cv=none; b=qzQQro3A3+shgY3/s3CRj8PRoYz0xjT+VSnONomhKuJTTYi9f1E0TvZSNytv1l0vDuZ6+XGQeQXumGeKB8Ao+rSwL47xEFjpEEdVriSha3zwc1MbGX4xJtA9ufmtuo5NHqsBUddvGtTyfCuy2tEO7kD+UF/HLj2T5iK3/kwvjz8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281462; c=relaxed/simple; bh=H2Kd07qElqnot6gVepkdwLKpMe9/R/K8xbz744caVbU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TRntTo8loQApPYVazILNXLL/Ip1UIbSucjIwhU5JAdY2x1xhvoU0TjbCXZqe005SiIRX5xG5A32o3M494+UdQCWiD9VupN7X2Oj/sKhpqCRD5syig7wPnFkXYkAcjs/cViywt5Z3CoXwPLSOqRIqwrFuRijkxZ9v9NfPUvvNXhY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=V+ornFK4; arc=none smtp.client-ip=209.85.128.42 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="V+ornFK4" Received: by mail-wm1-f42.google.com with SMTP id 5b1f17b1804b1-490afc47455so5503005e9.2 for ; Fri, 12 Jun 2026 09:24:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781281457; x=1781886257; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=tydmeyYrLPvT4qet9l6Pm8OKfFhD4dz0649Lu1ejs0k=; b=V+ornFK40pMnEL6osF/kpsRP7kftKkRR63KmY1fDtEtdEOoetHZAGyxb1o4GnysgeI Ee8owvBrO8/x/LpcFl9PkhkGbtC0JbeljkyubOaSDwYzppWW7G9IqPJJgRND+r6CcsrW QzX1oCVEMCkvk7dSKC1kjdkL1HtxCAMGjzU1u57aUynFdumMcefgoU+30Ve3PFR0v9Zw RBeYpNlibmbd4iuO9f5AcPau8waqIVJdNv7HTls7xKaStoO6+S+pOagnlWedR4HfBw9f Svdcj6rQn2pyk4YGLKx1zJiVpRESMrMdugi6aaZfzwBmZSMNlaWP/PzVekM6jRd/cD9M +Wxw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781281457; x=1781886257; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=tydmeyYrLPvT4qet9l6Pm8OKfFhD4dz0649Lu1ejs0k=; b=jSSJzsGnLFbExgbnFrhhOWAsC3R/nY1ob9zfpbzmxAHOcfAVC5SKUFEsmt5IqmNFpc a1yUOVuOxko3dYe27UKi61ZM64vXNrIjPfyzDzN3m75oWOVnKs4x0fb965azKK8YeBMR UZFtaYKq0/ooIhAFHdaM6iTXJYMD+lGvRAyIBQgiIMCUV4fNdDkmIGIq/q6TDlg3FDuF s7asVsSZd3vy+ROens6KvXm6G2pzk9lPMRY8ZCcrL+A/Qh97eWQTR42/11Ms/L7CohLf 8KcrTzb+81h2j12W2qGtuo/kW5Q73UeTSFMgv6gByuNz6XztvXHDluyFrlay/hr/jt6a emhg== X-Forwarded-Encrypted: i=1; AFNElJ/EZpRIEQe9PSdX7JVtIm3/WmqALzqoBLwdKzEip5C0XPQLxtYOHGj7bZfZ1IDTIM11V5uBsNUTHiI9diY=@vger.kernel.org X-Gm-Message-State: AOJu0YyCK1ED1nlVm8ZKlVLM/vLQFyBPgHzh0kMOEs/porKQTOl//QNW PTGHsgpVJlMX9svW1JmqreJfd8WCJPGOqzh6gChOnLEoXiheMMFpyiajd19bxQ== X-Gm-Gg: Acq92OGp0ByxKluiKIJ4qjUK+g1p5Wd6i5rb+8Rw9w7MBX4ZHmzcitVQ4raSlK2DyAO woegA34CwGxSPNccpomcnJmlyfqWF/lRvTv/KK6BpuhkxA0WbsOg39ele2vQ4F602f8EaJyJ39/ kwVbHlz3T3JAa/2LGQj6d4r8PlbxSAJAtVRmtg31Cvn2L5zBWccNGqSlvjbOdQPC4qKThIJ/bpR +J2WDLcduUuYIuwaxQEg0piDiwlYW1C3xXguE4aBudVfFOJosvUgaAZwCTiYaKW1QEHkCD1apdH +LhiMrj7AvGWK1YKeWk5vnx/VftA6bn1NthtMcwRibywuS7mI5P7A0iaTf8gKL616X+XAnGG9ds 3xxT8gtuGqQ+aou2WvEgYCn9DVsA3Ah1V4bHxH1yduCxikCg5ICFxoUlkYdGOAJPF+IS2rzW7U7 9aweuZXNy+z1rXACXP5mXDDGrSsN9t2DvOv9p/Sc5GjpW0qanjev8zW1teu1UFGg== X-Received: by 2002:a05:600c:c0d1:10b0:490:b9c3:6c69 with SMTP id 5b1f17b1804b1-490ec50f80cmr36665975e9.30.1781281456986; Fri, 12 Jun 2026 09:24:16 -0700 (PDT) Received: from f4d4888f22f2.ant.amazon.com.com ([15.248.2.31]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490ea95c51dsm57620935e9.1.2026.06.12.09.24.15 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 12 Jun 2026 09:24:16 -0700 (PDT) From: Jack Thomson To: maz@kernel.org, oupton@kernel.org, pbonzini@redhat.com Cc: joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, shuah@kernel.org, corbet@lwn.net, vladimir.murzin@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, isaku.yamahata@intel.com, Jack Thomson Subject: [PATCH v5 3/5] KVM: selftests: Enable pre_fault_memory_test for arm64 Date: Fri, 12 Jun 2026 17:23:51 +0100 Message-ID: <20260612162354.73378-4-jackabt.amazon@gmail.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260612162354.73378-1-jackabt.amazon@gmail.com> References: <20260612162354.73378-1-jackabt.amazon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jack Thomson Enable the pre_fault_memory_test to run on arm64 by making it work with different guest page sizes and testing multiple guest configurations. Update the test_assert to compare against the UCALL_EXIT_REASON, for portability, as arm64 exits with KVM_EXIT_MMIO while x86 uses KVM_EXIT_IO. Signed-off-by: Jack Thomson --- tools/testing/selftests/kvm/Makefile.kvm | 1 + .../selftests/kvm/pre_fault_memory_test.c | 115 ++++++++++++++---- 2 files changed, 92 insertions(+), 24 deletions(-) diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selft= ests/kvm/Makefile.kvm index 9118a5a51b89..4609d8f23e38 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -194,6 +194,7 @@ TEST_GEN_PROGS_arm64 +=3D guest_memfd_test TEST_GEN_PROGS_arm64 +=3D mmu_stress_test TEST_GEN_PROGS_arm64 +=3D rseq_test TEST_GEN_PROGS_arm64 +=3D steal_time +TEST_GEN_PROGS_arm64 +=3D pre_fault_memory_test =20 TEST_GEN_PROGS_s390 =3D $(TEST_GEN_PROGS_COMMON) TEST_GEN_PROGS_s390 +=3D s390/memop diff --git a/tools/testing/selftests/kvm/pre_fault_memory_test.c b/tools/te= sting/selftests/kvm/pre_fault_memory_test.c index fcb57fd034e6..9f5f0d1a5db1 100644 --- a/tools/testing/selftests/kvm/pre_fault_memory_test.c +++ b/tools/testing/selftests/kvm/pre_fault_memory_test.c @@ -11,19 +11,29 @@ #include #include #include +#include =20 /* Arbitrarily chosen values */ -#define TEST_SIZE (SZ_2M + PAGE_SIZE) -#define TEST_NPAGES (TEST_SIZE / PAGE_SIZE) +#define TEST_BASE_SIZE SZ_2M #define TEST_SLOT 10 =20 +/* Storage of test info to share with guest code */ +struct test_config { + u64 page_size; + u64 test_size; + u64 test_num_pages; +}; + +static struct test_config test_config; + static void guest_code(u64 base_gva) { volatile u64 val __used; + struct test_config *config =3D &test_config; int i; =20 - for (i =3D 0; i < TEST_NPAGES; i++) { - u64 *src =3D (u64 *)(base_gva + i * PAGE_SIZE); + for (i =3D 0; i < config->test_num_pages; i++) { + u64 *src =3D (u64 *)(base_gva + i * config->page_size); =20 val =3D *src; } @@ -56,7 +66,7 @@ static void *delete_slot_worker(void *__data) cpu_relax(); =20 vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, data->gpa, - TEST_SLOT, TEST_NPAGES, data->flags); + TEST_SLOT, test_config.test_num_pages, data->flags); =20 return NULL; } @@ -149,8 +159,8 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64= base_gpa, u64 offset, /* * Assert success if prefaulting the entire range should succeed, i.e. * complete with no bytes remaining. Otherwise prefaulting should have - * failed due to ENOENT (due to RET_PF_EMULATE for emulated MMIO when - * no memslot exists). + * failed due to ENOENT (no memslot exists for the GPA; on x86 this + * surfaces via RET_PF_EMULATE). */ if (!expected_left) TEST_ASSERT_VM_VCPU_IOCTL(!ret, KVM_PRE_FAULT_MEMORY, ret, vcpu->vm); @@ -159,43 +169,70 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u= 64 base_gpa, u64 offset, KVM_PRE_FAULT_MEMORY, ret, vcpu->vm); } =20 -static void __test_pre_fault_memory(unsigned long vm_type, bool private) +struct test_params { + unsigned long vm_type; + bool private; +}; + +static void __test_pre_fault_memory(enum vm_guest_mode guest_mode, void *a= rg) { - gpa_t gpa, gva, alignment, guest_page_size; + gpa_t gpa, gva, alignment, guest_page_size, host_page_size; + struct test_params *p =3D arg; const struct vm_shape shape =3D { - .mode =3D VM_MODE_DEFAULT, - .type =3D vm_type, + .mode =3D guest_mode, + .type =3D p->vm_type, }; struct kvm_vcpu *vcpu; struct kvm_run *run; struct kvm_vm *vm; struct ucall uc; =20 + pr_info("Testing guest mode: %s\n", vm_guest_mode_string(guest_mode)); + vm =3D vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); =20 - alignment =3D guest_page_size =3D vm_guest_mode_params[VM_MODE_DEFAULT].p= age_size; - gpa =3D (vm->max_gfn - TEST_NPAGES) * guest_page_size; + guest_page_size =3D vm_guest_mode_params[guest_mode].page_size; + host_page_size =3D getpagesize(); + + test_config.page_size =3D guest_page_size; + test_config.test_size =3D align_up(TEST_BASE_SIZE + test_config.page_size, + host_page_size); + test_config.test_num_pages =3D vm_calc_num_guest_pages(vm->mode, test_con= fig.test_size); + + gpa =3D (vm->max_gfn - test_config.test_num_pages) * test_config.page_siz= e; alignment =3D SZ_2M; + alignment =3D max(alignment, host_page_size); gpa =3D align_down(gpa, alignment); gva =3D gpa & ((1ULL << (vm->va_bits - 1)) - 1); =20 - vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa, TEST_SLOT, - TEST_NPAGES, private ? KVM_MEM_GUEST_MEMFD : 0); - virt_map(vm, gva, gpa, TEST_NPAGES); + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, + gpa, TEST_SLOT, test_config.test_num_pages, + p->private ? KVM_MEM_GUEST_MEMFD : 0); + virt_map(vm, gva, gpa, test_config.test_num_pages); =20 - if (private) - vm_mem_set_private(vm, gpa, TEST_SIZE); + if (p->private) + vm_mem_set_private(vm, gpa, test_config.test_size); =20 - pre_fault_memory(vcpu, gpa, 0, SZ_2M, 0, private); - pre_fault_memory(vcpu, gpa, SZ_2M, PAGE_SIZE * 2, PAGE_SIZE, private); - pre_fault_memory(vcpu, gpa, TEST_SIZE, PAGE_SIZE, PAGE_SIZE, private); + pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, p->private); + /* Retry the same range after the first prefault attempt. */ + pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, p->private); + pre_fault_memory(vcpu, gpa, + test_config.test_size - host_page_size, + host_page_size * 2, host_page_size, p->private); + pre_fault_memory(vcpu, gpa, test_config.test_size, + host_page_size, host_page_size, p->private); =20 vcpu_args_set(vcpu, 1, gva); + + /* Export the shared variables to the guest. */ + sync_global_to_guest(vm, test_config); + vcpu_run(vcpu); =20 run =3D vcpu->run; - TEST_ASSERT(run->exit_reason =3D=3D KVM_EXIT_IO, - "Wanted KVM_EXIT_IO, got exit reason: %u (%s)", + TEST_ASSERT(run->exit_reason =3D=3D UCALL_EXIT_REASON, + "Wanted %s, got exit reason: %u (%s)", + exit_reason_str(UCALL_EXIT_REASON), run->exit_reason, exit_reason_str(run->exit_reason)); =20 switch (get_ucall(vcpu, &uc)) { @@ -214,16 +251,46 @@ static void __test_pre_fault_memory(unsigned long vm_= type, bool private) =20 static void test_pre_fault_memory(unsigned long vm_type, bool private) { + struct test_params p =3D { + .vm_type =3D vm_type, + .private =3D private, + }; + if (vm_type && !(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(vm_type))) { pr_info("Skipping tests for vm_type 0x%lx\n", vm_type); return; } =20 - __test_pre_fault_memory(vm_type, private); + for_each_guest_mode(__test_pre_fault_memory, &p); +} + +static void help(char *name) +{ + puts(""); + printf("usage: %s [-h] [-m mode]\n", name); + puts(""); + guest_modes_help(); + puts(""); } =20 int main(int argc, char *argv[]) { + int opt; + + guest_modes_append_default(); + + while ((opt =3D getopt(argc, argv, "hm:")) !=3D -1) { + switch (opt) { + case 'm': + guest_modes_cmdline(optarg); + break; + case 'h': + default: + help(argv[0]); + exit(0); + } + } + TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY)); =20 test_pre_fault_memory(0, false); --=20 2.43.0 From nobody Thu Sep 24 13:37:22 2026 Received: from mail-wm1-f45.google.com (mail-wm1-f45.google.com [209.85.128.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3925B369D56 for ; Fri, 12 Jun 2026 16:24:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281463; cv=none; b=byGlhupeyEJ8jpTYetZdXetdx5aM1unNkchTYJrPQK0nRJhGGWExjfWb+qsIpQVRjkWb3JvqHFZiUhWGRYrSI0SANUJGgZkRq63KeOXNOtwFPLDSnE0BqxCxs8SzM5HMAgq9r2ohJB+63Dw5KuBywTbTAva43lZCKoa/07ANxSU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281463; c=relaxed/simple; bh=KmMepa6SILU+p5gEjwRcDM4ZL8kxCOFddSZl35hjD5I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B5GuupvR1cMZx9wsH8idRgS1d2JHclJFUGHv3ycZS9Lt2Rhfq7eROYp0/5VyV1rxG739Xd0riFXLwlh0dHMf7w0YW8fRBJnGTxUsn+W+v7fY+3kv9o8Jz86vz0VrtNdO7UoDOc+Zb1miPxA2RwlWwBUkfgZNWx2GxPciRzKRPkM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=H2euK0jM; arc=none smtp.client-ip=209.85.128.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="H2euK0jM" Received: by mail-wm1-f45.google.com with SMTP id 5b1f17b1804b1-491b390f9e9so4948825e9.0 for ; Fri, 12 Jun 2026 09:24:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781281458; x=1781886258; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=Tuh44YqrjzSlLd6mm7aQwU0y2sKDDLs9iqICVctwKS8=; b=H2euK0jMcKLvCfRWGX6UYiKL+KhzjvVA8yan1E4HoPE+w71Gcj60T3H2zeWSEjq7Ih /+pz/E9rBXSdeLUGyVonD8d1Iz3AlpEkVEO3WZS5rnD7y3CG4fdRBlrIB7uhx024Ev7Q 9ei1n9SAkSUzzTcNvIgr7PmPfsXFkGDYVTT0IEGTUCUCPl2W10r79P+t9++NH/LCzkuV qqNX2eEmkQ+3YWPQz4bBMb+wOSkeX8iOgUvy0CiFATlk5yHbRZbNGvJQcLeDrlsfrrxG jecJ7Vq42HhV4HJjPvg2qQNk8otVAjkEG0OELV7L5wx5RKWfrA8UsJNZvArDtI1f38gq Ux/Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781281458; x=1781886258; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=Tuh44YqrjzSlLd6mm7aQwU0y2sKDDLs9iqICVctwKS8=; b=Ps6wo3NegyhKuAqwZB5m52Q7U+RBt67s9t24aQ5K7WTGrK6XAGvY/4kcodXdyM5BjV SvQlkymxKY32kFtTXQQZNyNG4xX7k1OOcgnirblV+tfq3GITBYZlIyyswMoPwRId+DC7 0Td101U1xuK+J9Qd1Ev2R2/YgbJR5OCJziGGXPBfZ2oOC8uX5y81LOLbdoTvGMezH9PC Y4s60vtf7h/ud8bGTcyMdwlaIEK2RrJMuPkeMk4th1AhKiakQyYrHaffj9gsO5rm1ngP 2Ff0xmvQAXqvpaDVy8Yi7M54nzMcUDNcyKMHZBRZj6fZaRe85rGT1dBOeSQjddqy6UeQ xAWw== X-Forwarded-Encrypted: i=1; AFNElJ/R6IbkEjt8bzyyep2Q7ej+Zy1t1BNW0Xs0E6qqNjxIYufadaqyDOi5GK8oP0mdzuIWqRcryjlMhAmrTTI=@vger.kernel.org X-Gm-Message-State: AOJu0YwgehBGo+Og2OI9qKMJ0S4qgebtJVq1AgbHJmkqSQk3ZMO++cT/ c84rqFETD1mf5pyNs51Wo+MmWmP+ztpOsOSWBb3Cp5mU2nbM5HI3qMMx X-Gm-Gg: Acq92OFqbgt5f9IioMo4HRQ8u4xW0lCTuAohcESxzJ5/+CtXzdgyywatP+o12fgBByS GsGJAZMm1FwUtQiufhezmaBwjfm9YHkGYXqk8J8YSC5seW06TPoWYk2eV+D1u3lvxnwBM7Ntces JbEYhjAy7EJIjPImS86oFg7rW3E7EZPk5YXn/GLGjx6MVtQHuO0Fs7L9r+UActbCA4qBCHBxUD4 fAV2C/H7r7U8rbK1aN4otlqEaPhzLOAL5Aze/lzZB7jv7mTQxaAMQAVXimtTKfsVJhA3PcvHOk0 a2xH/5DXvlsLatvYJBd+2eg3QAyD5/fAQOvNnqEdOwHdK3ZQ0/30LcCeKw3GyXVyKZ90bdPzbc+ msAOMuFjk5TpanyOaETVtR1VJ6A9VdUlTT0EKkh6o8OvoyagLkoRz8TNCwIR0Da9WA9/LwOxMYk 4+DYEtomWs3zB/umgmG9dazc5+69Wa+LYmMFKIL3hQNlZybvLQ9bGyVHjI3fpBbw== X-Received: by 2002:a05:600c:468d:b0:490:9588:bdae with SMTP id 5b1f17b1804b1-490ec4ee664mr53299705e9.18.1781281458289; Fri, 12 Jun 2026 09:24:18 -0700 (PDT) Received: from f4d4888f22f2.ant.amazon.com.com ([15.248.2.31]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490ea95c51dsm57620935e9.1.2026.06.12.09.24.17 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 12 Jun 2026 09:24:18 -0700 (PDT) From: Jack Thomson To: maz@kernel.org, oupton@kernel.org, pbonzini@redhat.com Cc: joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, shuah@kernel.org, corbet@lwn.net, vladimir.murzin@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, isaku.yamahata@intel.com, Jack Thomson Subject: [PATCH v5 4/5] KVM: selftests: Add option for different backing in pre-fault tests Date: Fri, 12 Jun 2026 17:23:52 +0100 Message-ID: <20260612162354.73378-5-jackabt.amazon@gmail.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260612162354.73378-1-jackabt.amazon@gmail.com> References: <20260612162354.73378-1-jackabt.amazon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jack Thomson Add a -s option to specify different memory backing types for the pre-fault tests (e.g. anonymous, hugetlb), allowing testing of the pre-fault functionality across different memory configurations. Signed-off-by: Jack Thomson --- .../selftests/kvm/pre_fault_memory_test.c | 51 +++++++++++++------ 1 file changed, 36 insertions(+), 15 deletions(-) diff --git a/tools/testing/selftests/kvm/pre_fault_memory_test.c b/tools/te= sting/selftests/kvm/pre_fault_memory_test.c index 9f5f0d1a5db1..c850cf28e86a 100644 --- a/tools/testing/selftests/kvm/pre_fault_memory_test.c +++ b/tools/testing/selftests/kvm/pre_fault_memory_test.c @@ -45,6 +45,7 @@ struct slot_worker_data { struct kvm_vm *vm; gpa_t gpa; u32 flags; + enum vm_mem_backing_src_type mem_backing_src; bool worker_ready; bool prefault_ready; bool recreate_slot; @@ -65,14 +66,16 @@ static void *delete_slot_worker(void *__data) while (!READ_ONCE(data->recreate_slot)) cpu_relax(); =20 - vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, data->gpa, + vm_userspace_mem_region_add(vm, data->mem_backing_src, data->gpa, TEST_SLOT, test_config.test_num_pages, data->flags); =20 return NULL; } =20 static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offs= et, - u64 size, u64 expected_left, bool private) + u64 size, u64 expected_left, + enum vm_mem_backing_src_type mem_backing_src, + bool private) { struct kvm_pre_fault_memory range =3D { .gpa =3D base_gpa + offset, @@ -83,6 +86,7 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 b= ase_gpa, u64 offset, .vm =3D vcpu->vm, .gpa =3D base_gpa, .flags =3D private ? KVM_MEM_GUEST_MEMFD : 0, + .mem_backing_src =3D mem_backing_src, }; bool slot_recreated =3D false; pthread_t slot_worker; @@ -172,11 +176,13 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u= 64 base_gpa, u64 offset, struct test_params { unsigned long vm_type; bool private; + enum vm_mem_backing_src_type mem_backing_src; }; =20 static void __test_pre_fault_memory(enum vm_guest_mode guest_mode, void *a= rg) { gpa_t gpa, gva, alignment, guest_page_size, host_page_size; + gpa_t backing_src_pagesz, mem_page_size; struct test_params *p =3D arg; const struct vm_shape shape =3D { .mode =3D guest_mode, @@ -188,24 +194,28 @@ static void __test_pre_fault_memory(enum vm_guest_mod= e guest_mode, void *arg) struct ucall uc; =20 pr_info("Testing guest mode: %s\n", vm_guest_mode_string(guest_mode)); + pr_info("Testing memory backing src type: %s\n", + vm_mem_backing_src_alias(p->mem_backing_src)->name); =20 vm =3D vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); =20 guest_page_size =3D vm_guest_mode_params[guest_mode].page_size; host_page_size =3D getpagesize(); + backing_src_pagesz =3D get_backing_src_pagesz(p->mem_backing_src); + mem_page_size =3D max(host_page_size, backing_src_pagesz); =20 test_config.page_size =3D guest_page_size; test_config.test_size =3D align_up(TEST_BASE_SIZE + test_config.page_size, - host_page_size); + mem_page_size); test_config.test_num_pages =3D vm_calc_num_guest_pages(vm->mode, test_con= fig.test_size); =20 gpa =3D (vm->max_gfn - test_config.test_num_pages) * test_config.page_siz= e; alignment =3D SZ_2M; - alignment =3D max(alignment, host_page_size); + alignment =3D max(alignment, mem_page_size); gpa =3D align_down(gpa, alignment); gva =3D gpa & ((1ULL << (vm->va_bits - 1)) - 1); =20 - vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, + vm_userspace_mem_region_add(vm, p->mem_backing_src, gpa, TEST_SLOT, test_config.test_num_pages, p->private ? KVM_MEM_GUEST_MEMFD : 0); virt_map(vm, gva, gpa, test_config.test_num_pages); @@ -213,14 +223,18 @@ static void __test_pre_fault_memory(enum vm_guest_mod= e guest_mode, void *arg) if (p->private) vm_mem_set_private(vm, gpa, test_config.test_size); =20 - pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, p->private); + pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, + p->mem_backing_src, p->private); /* Retry the same range after the first prefault attempt. */ - pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, p->private); + pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0, + p->mem_backing_src, p->private); pre_fault_memory(vcpu, gpa, test_config.test_size - host_page_size, - host_page_size * 2, host_page_size, p->private); + host_page_size * 2, host_page_size, + p->mem_backing_src, p->private); pre_fault_memory(vcpu, gpa, test_config.test_size, - host_page_size, host_page_size, p->private); + host_page_size, host_page_size, + p->mem_backing_src, p->private); =20 vcpu_args_set(vcpu, 1, gva); =20 @@ -249,11 +263,13 @@ static void __test_pre_fault_memory(enum vm_guest_mod= e guest_mode, void *arg) kvm_vm_free(vm); } =20 -static void test_pre_fault_memory(unsigned long vm_type, bool private) +static void test_pre_fault_memory(unsigned long vm_type, enum vm_mem_backi= ng_src_type backing_src, + bool private) { struct test_params p =3D { .vm_type =3D vm_type, .private =3D private, + .mem_backing_src =3D backing_src, }; =20 if (vm_type && !(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(vm_type))) { @@ -267,23 +283,28 @@ static void test_pre_fault_memory(unsigned long vm_ty= pe, bool private) static void help(char *name) { puts(""); - printf("usage: %s [-h] [-m mode]\n", name); + printf("usage: %s [-h] [-m mode] [-s mem-type]\n", name); puts(""); guest_modes_help(); + backing_src_help("-s"); puts(""); } =20 int main(int argc, char *argv[]) { + enum vm_mem_backing_src_type backing =3D DEFAULT_VM_MEM_SRC; int opt; =20 guest_modes_append_default(); =20 - while ((opt =3D getopt(argc, argv, "hm:")) !=3D -1) { + while ((opt =3D getopt(argc, argv, "hm:s:")) !=3D -1) { switch (opt) { case 'm': guest_modes_cmdline(optarg); break; + case 's': + backing =3D parse_backing_src_type(optarg); + break; case 'h': default: help(argv[0]); @@ -293,10 +314,10 @@ int main(int argc, char *argv[]) =20 TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY)); =20 - test_pre_fault_memory(0, false); + test_pre_fault_memory(0, backing, false); #ifdef __x86_64__ - test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, false); - test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, true); + test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, false); + test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, true); #endif return 0; } --=20 2.43.0 From nobody Thu Sep 24 13:37:22 2026 Received: from mail-wm1-f45.google.com (mail-wm1-f45.google.com [209.85.128.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 244303E6383 for ; Fri, 12 Jun 2026 16:24:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281464; cv=none; b=VWO/Z8L1LWGXWXm/F4M9lDt9l2YGXjcSgoaWKI1+dxDjDffvdW1H/StjvPEPyqN6zizjvsoz6A/PDhejMF5Kx4k0TawXUgIpwbJLTsJMx5aWsYeCqzDox3ekXlmeqGAdGPqPCvnqy3LfLIYUVj4OxTi9EyPKe18qndnQdgH4924= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781281464; c=relaxed/simple; bh=vAClADCgCQBCt/qYeMJkO5SXTU35Ou+fa4NCteEDdtc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ZURwhv6vFMeO38cAFebx8zQZXhX6yJRHl2fWLAs3LPP7n2t5RyhkjH/txB+ZIPK1trcHItIgTjgDEp+SEnUCn/Kg9TfUPKItbngv3mMXM+C4UDdUcKzmIfXJvkrTU2nxb8LijewvsX5g8IkFjkLz6lGp8mXFgBvZzpvdqtt+/SE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=S2iWDs44; arc=none smtp.client-ip=209.85.128.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="S2iWDs44" Received: by mail-wm1-f45.google.com with SMTP id 5b1f17b1804b1-490ace40f4bso10873965e9.3 for ; Fri, 12 Jun 2026 09:24:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781281459; x=1781886259; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=ow6Lcgf/RFgv5jAm1y7NuVlWFYYd0/JrOjI+zfS6dAY=; b=S2iWDs44ogcjVukGLcI7YlzsV4PaID2rpGPZr1MqyAbWhW1rmgRyhQ1hTUKaqenRqz aiebOIt8CE+ZtAVoRgmZw4j+hH9vIDuMOiejjqw1fMFu1fSEfTpz3DwipdEDjhBOAqXC pbiSoMoPKidY02V8gTbdg/nHQdSWYLE3WhmDNSo/Rm8l6sm3Pp27dKRM8QI9ZgzlbHK0 YrMF4bKe8/uZG15etrB+wAyszZB2L8RaP9M2RdeKxJdlmMuprAHZZv8U6KgwRKm5tLcf lrCHtDWTGnlgVR43bxJyXay4agd4Rx0YLC1hFIFVThEjtpbo8OLl9BeVhVMNbx0oJkkf RvAA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781281459; x=1781886259; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=ow6Lcgf/RFgv5jAm1y7NuVlWFYYd0/JrOjI+zfS6dAY=; b=cUkJJe5II3Hglh390AdJ442cv8makhujrMnSOiulEoB5vUDRUtm4ir/QQoIiRLmgvJ Zay15sOb28ZBpgQyON+P0t9YDI5w6OhZBMQjbfjiI8kMa+jdlaCBKGa+89HKy3fei6xt GGD+oNxK0CH3uUHm5lgCwrBakF7ZGVi3DgLHdLAn9cha4bGcAYHJSGMQwJzK6zXenui2 UZ9+a14WcS/pz992O4dba/qNw1uRU6aaU+cxBcScJmRM0ObX5MQSV7RX3nVqOiM0TgbV KsyHanHk+Ued28VCg6uxK3jWHQiOAMxXN/COCxLAg/9FTpVuewYGxuWVeHwSh+KhJPzx Rrxw== X-Forwarded-Encrypted: i=1; AFNElJ9OpY8cS1g9EouZ1ogJZUnQxoK8gQdQwqVCDBZtVYrZtJocnvzv2lpqG3PEeRjFfB58UyvXiDhb/4Fr+Dk=@vger.kernel.org X-Gm-Message-State: AOJu0Yy3PaHfrfAVPjzcAVge/uEf5weIe1XB4Ghj4LgFQmHyB8d6hh+w CsTab8nPMfHOIJfSoTqn8oCliXNJt0WLzBOX7mR5whFE0R9uL6TsbMdv X-Gm-Gg: Acq92OELOuFFJuFjmxPC1S2gn46Yk0Mp9iKEOO4mMbAH0UmXboG/jJf1rsed3NnTJT/ Ip1bxTOn0CwB4C3wixVov7BiHMgxaDM2vuyB2oP3MVHg+3c5nN3BJrP14vnPx/D9RGL7LAJFr+k fM1gca6FQPI8y9boQGVFkzVw+3ycA+3eHWy6LVwp3rcsVhDF0ZV896BHwuJR2e/JNZ7Jkub8KMe EJIfJ/T2nSuJCO0BoeqzbR7NTm5lHoUwjZB+kQNV8MFYh81aeiYcP0HUkw4F+hhYTLmbJNE55iD Zo6taWulALuxgp4jTOKHX7Q7FW1sm2GDMLgcl15Q5BlzsV7zPP5EPrWGKxCm2fH3PZJxWxXyrAj KEy8X3pTpXtXq6oKAOpOrtmr0BI5N84KH6fvW2kjZFpW0OwjVz+77y49Qsdic0JXEei7QMxr2Sa mk0939uJfykOoXG63ao/nNaj7zys2kQjMzYbKLieCkTVfgkUHoeJ5ToN0MMzDo3g== X-Received: by 2002:a05:600c:4e48:b0:490:e60b:6860 with SMTP id 5b1f17b1804b1-4922005e235mr2153615e9.7.1781281459529; Fri, 12 Jun 2026 09:24:19 -0700 (PDT) Received: from f4d4888f22f2.ant.amazon.com.com ([15.248.2.31]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-490ea95c51dsm57620935e9.1.2026.06.12.09.24.18 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 12 Jun 2026 09:24:19 -0700 (PDT) From: Jack Thomson To: maz@kernel.org, oupton@kernel.org, pbonzini@redhat.com Cc: joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, shuah@kernel.org, corbet@lwn.net, vladimir.murzin@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, isaku.yamahata@intel.com, Jack Thomson Subject: [PATCH v5 5/5] KVM: selftests: Add nested pre-fault test for arm64 Date: Fri, 12 Jun 2026 17:23:53 +0100 Message-ID: <20260612162354.73378-6-jackabt.amazon@gmail.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260612162354.73378-1-jackabt.amazon@gmail.com> References: <20260612162354.73378-1-jackabt.amazon@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jack Thomson Add an arm64 nested-virt selftest for KVM_PRE_FAULT_MEMORY. The guest enters vEL1 and exits to userspace with a nested/shadow stage-2 MMU as the vCPU's last-run context. Before prefaulting, userspace enables HCR_EL2.VM and points VTTBR_EL2 at an empty nested stage-2 root. A prefault implementation that incorrectly treats the userspace GPA as an L2 IPA will fail the ioctl; the correct path swaps to the canonical stage-2 and succeeds. Restore the original nested state before resuming the guest, then touch the prefaulted range to check that vEL1 still runs correctly. Signed-off-by: Jack Thomson --- tools/testing/selftests/kvm/Makefile.kvm | 1 + .../kvm/arm64/nv_pre_fault_memory_test.c | 200 ++++++++++++++++++ 2 files changed, 201 insertions(+) create mode 100644 tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_t= est.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selft= ests/kvm/Makefile.kvm index 4609d8f23e38..63d79245b47d 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -170,6 +170,7 @@ TEST_GEN_PROGS_arm64 +=3D arm64/debug-exceptions TEST_GEN_PROGS_arm64 +=3D arm64/hello_el2 TEST_GEN_PROGS_arm64 +=3D arm64/host_sve TEST_GEN_PROGS_arm64 +=3D arm64/hypercalls +TEST_GEN_PROGS_arm64 +=3D arm64/nv_pre_fault_memory_test TEST_GEN_PROGS_arm64 +=3D arm64/external_aborts TEST_GEN_PROGS_arm64 +=3D arm64/page_fault_test TEST_GEN_PROGS_arm64 +=3D arm64/psci_test diff --git a/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c b= /tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c new file mode 100644 index 000000000000..2bbd5540599c --- /dev/null +++ b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c @@ -0,0 +1,200 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * nv_pre_fault_memory_test - Test KVM_PRE_FAULT_MEMORY on a vCPU whose + * last-run context is nested. + * + * The guest starts at vEL2, mirrors its EL2 translation regime into the + * real EL1 registers, drops HCR_EL2.TGE and ERETs to vEL1, then exits to + * userspace from vEL1 so that the vCPU's last-run context selects a + * shadow stage-2 MMU. Userspace then enables an empty nested stage-2 + * before prefaulting. Prefaulting must target the canonical stage-2, + * regardless of the vCPU's nested state. + */ +#include "kvm_util.h" +#include "processor.h" +#include "test_util.h" +#include "ucall.h" + +#include +#include + +#define TEST_MEM_SLOT 10 +#define NESTED_S2_ROOT_SLOT 11 +#define TEST_MEM_SIZE SZ_2M +#define TEST_MEM_GPA SZ_1G +#define NESTED_S2_ROOT_GPA (TEST_MEM_GPA + TEST_MEM_SIZE) + +struct nested_s2_state { + u64 hcr_el2; + u64 vttbr_el2; +}; + +static void guest_el1_code(void) +{ + u64 offset; + + GUEST_ASSERT_EQ(get_current_el(), 1); + + /* Exit to userspace with the vEL1 (nested) context live. */ + GUEST_SYNC(1); + + /* + * Touch the prefaulted range. vstage-2 is disabled, so the shadow + * stage-2 is a 1:1 view of the canonical IPA space. + */ + for (offset =3D 0; offset < TEST_MEM_SIZE; offset +=3D SZ_4K) + READ_ONCE(*(u64 *)(TEST_MEM_GPA + offset)); + + GUEST_DONE(); +} + +static void guest_code(void) +{ + u64 sp; + + GUEST_ASSERT_EQ(get_current_el(), 2); + + /* + * Mirror the EL2 translation regime into the real EL1 registers so + * that vEL1 runs on the test's stage-1 page tables. With E2H=3D1, the + * _EL1 accessors read the EL2 registers, and the _EL12 accessors + * write the real EL1 registers. + */ + write_sysreg_s(read_sysreg(sctlr_el1), SYS_SCTLR_EL12); + write_sysreg_s(read_sysreg(tcr_el1), SYS_TCR_EL12); + write_sysreg_s(read_sysreg(ttbr0_el1), SYS_TTBR0_EL12); + write_sysreg_s(read_sysreg(mair_el1), SYS_MAIR_EL12); + write_sysreg_s(read_sysreg(cpacr_el1), SYS_CPACR_EL12); + + /* Run vEL1 on the same stack. */ + asm volatile("mov %0, sp" : "=3Dr"(sp)); + write_sysreg(sp, sp_el1); + + /* + * Drop TGE so that vEL1 is a nested context rather than host EL0. + * KVM backs it with a shadow stage-2 MMU even though vstage-2 is + * disabled (HCR_EL2.VM=3D0). + */ + write_sysreg(read_sysreg(hcr_el2) & ~HCR_EL2_TGE, hcr_el2); + isb(); + + write_sysreg(PSR_MODE_EL1h | PSR_F_BIT | PSR_I_BIT | PSR_A_BIT | + PSR_D_BIT, spsr_el2); + write_sysreg((u64)guest_el1_code, elr_el2); + asm volatile("eret"); + + GUEST_ASSERT(false); +} + +static void pre_fault(struct kvm_vcpu *vcpu, u64 gpa, u64 size) +{ + struct kvm_pre_fault_memory range =3D { + .gpa =3D gpa, + .size =3D size, + }; + int ret; + + do { + ret =3D __vcpu_ioctl(vcpu, KVM_PRE_FAULT_MEMORY, &range); + } while (ret < 0 && errno =3D=3D EINTR); + + TEST_ASSERT(!ret, "KVM_PRE_FAULT_MEMORY failed, ret: %d errno: %d", + ret, errno); + TEST_ASSERT_EQ(range.size, 0); +} + +static struct nested_s2_state enable_empty_nested_s2(struct kvm_vcpu *vcpu) +{ + struct nested_s2_state state =3D { + .hcr_el2 =3D vcpu_get_reg(vcpu, KVM_ARM64_SYS_REG(SYS_HCR_EL2)), + .vttbr_el2 =3D vcpu_get_reg(vcpu, + KVM_ARM64_SYS_REG(SYS_VTTBR_EL2)), + }; + + TEST_ASSERT(!(state.hcr_el2 & HCR_EL2_TGE), + "vCPU should be in nested/vEL1 context"); + + vcpu_set_reg(vcpu, KVM_ARM64_SYS_REG(SYS_VTTBR_EL2), + NESTED_S2_ROOT_GPA); + vcpu_set_reg(vcpu, KVM_ARM64_SYS_REG(SYS_HCR_EL2), + state.hcr_el2 | HCR_EL2_VM); + + return state; +} + +static void restore_nested_s2(struct kvm_vcpu *vcpu, + struct nested_s2_state *state) +{ + vcpu_set_reg(vcpu, KVM_ARM64_SYS_REG(SYS_HCR_EL2), state->hcr_el2); + vcpu_set_reg(vcpu, KVM_ARM64_SYS_REG(SYS_VTTBR_EL2), + state->vttbr_el2); +} + +int main(void) +{ + struct nested_s2_state s2; + struct kvm_vcpu_init init; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + u64 npages; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_ARM_EL2)); + TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY)); + + vm =3D vm_create(1); + + kvm_get_default_vcpu_target(vm, &init); + init.features[0] |=3D BIT(KVM_ARM_VCPU_HAS_EL2); + vcpu =3D aarch64_vcpu_add(vm, 0, &init, guest_code); + kvm_arch_vm_finalize_vcpus(vm); + + npages =3D TEST_MEM_SIZE / vm->page_size; + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, TEST_MEM_GPA, + TEST_MEM_SLOT, npages, 0); + virt_map(vm, TEST_MEM_GPA, TEST_MEM_GPA, npages); + + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, + NESTED_S2_ROOT_GPA, NESTED_S2_ROOT_SLOT, + 1, 0); + + /* Run the guest until it has ERET'd from vEL2 to vEL1. */ + vcpu_run(vcpu); + switch (get_ucall(vcpu, &uc)) { + case UCALL_SYNC: + TEST_ASSERT_EQ(uc.args[1], 1); + break; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + default: + TEST_FAIL("Unhandled ucall: %ld", uc.cmd); + } + + /* + * The vCPU's last-run context is vEL1, backed by a shadow stage-2 + * MMU. Enable nested stage-2 with an empty root so that the ioctl + * fails if it tries to interpret the userspace GPA as an L2 IPA. + * Prefault in two halves so that the second ioctl exercises a + * repeated shadow-MMU attach and canonical stage-2 swap. + */ + s2 =3D enable_empty_nested_s2(vcpu); + pre_fault(vcpu, TEST_MEM_GPA, TEST_MEM_SIZE / 2); + pre_fault(vcpu, TEST_MEM_GPA + TEST_MEM_SIZE / 2, TEST_MEM_SIZE / 2); + restore_nested_s2(vcpu, &s2); + + /* Resume at vEL1 and touch the prefaulted range. */ + vcpu_run(vcpu); + switch (get_ucall(vcpu, &uc)) { + case UCALL_DONE: + break; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + default: + TEST_FAIL("Unhandled ucall: %ld", uc.cmd); + } + + kvm_vm_free(vm); + return 0; +} --=20 2.43.0