From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B9F533A7587 for ; Thu, 16 Jul 2026 18:15:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225708; cv=none; b=TYFodwo0C4OdmXFdUAMRgc185jjKU7rgHSY+e3rFRQNfl0kMPxk9Cu3qQ9monuAjy3p0qghFGx/Az/sUZ+U3VEV3cHqzefWwgCwuuf7Hljjw9QYDjzG1NdoaReeWcgmsQuLEtnNkdJta/9B6dkXHE/W4/NeY0u1Xi5mWCgJp++s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225708; c=relaxed/simple; bh=Gmk3fsCPJ9/fpzOw/uDpZP5Yi/5iPCaqU5WO1EXGx4E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=b3IZ0CeZZNSnmccVikKRzOKxioWLGXhr6ZSpjtpN7sA9/G+WanDdZTDe1J2FsfUatPrXUD0MhzztSv02Aa3r5pJVeJ68XyaCegnr/wAdtEeiiZHQvOd5tZcob66S+FJrG50sDXd2nNtUUvn/Hlb9NSqBeJzO/tA6RJIEEWpIiW4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=QH+48PvY; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="QH+48PvY" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225701; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=A0nQvlR4hNkHzkahJsScWQc9QTREYmqYVvHYGclEvv0=; b=QH+48PvYvPHeCN3II5x/n/vrK0EfDFAf1G1MfDdpslZqx7UdbQLyMRzfpLyp3c1vhIyze8 exHOlKiB6SwUZqCNyEYwEnizivECNO6Jj2f14cht8oOwHWPW8akxV1v8s4PgczKPZ0LSlZ wQxTTlA3QVFbd+Eb0xiZrESu+Iyh5/I= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-301-2m4bTt_wOu2MyZ7F5nBSSg-1; Thu, 16 Jul 2026 14:14:59 -0400 X-MC-Unique: 2m4bTt_wOu2MyZ7F5nBSSg-1 X-Mimecast-MFC-AGG-ID: 2m4bTt_wOu2MyZ7F5nBSSg_1784225698 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 816751955F04; Thu, 16 Jul 2026 18:14:58 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 10737195604E; Thu, 16 Jul 2026 18:14:57 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 01/24] KVM: selftests: Take into account mixed memory fault flags Date: Thu, 16 Jul 2026 14:14:33 -0400 Message-ID: <20260716181456.402786-2-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Update x86_64/private_mem_kvm_exits_test to take into account memory fault flags might contain multiple bits set while remaining valid---for example KVM might return KVM_MEMORY_EXIT_FLAG_READ in addition to KVM_MEMORY_EXIT_FLAG_PRIVATE. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- .../testing/selftests/kvm/x86/private_mem_kvm_exits_test.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c b= /tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c index 10db9fe6d906..70ed3d591ab6 100644 --- a/tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c +++ b/tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c @@ -75,7 +75,8 @@ static void test_private_access_memslot_deleted(void) exit_reason =3D (u32)(u64)thread_return; =20 TEST_ASSERT_EQ(exit_reason, KVM_EXIT_MEMORY_FAULT); - TEST_ASSERT_EQ(vcpu->run->memory_fault.flags, KVM_MEMORY_EXIT_FLAG_PRIVAT= E); + TEST_ASSERT(vcpu->run->memory_fault.flags & KVM_MEMORY_EXIT_FLAG_PRIVATE, + "Memory fault didn't occur on a private memory access"); TEST_ASSERT_EQ(vcpu->run->memory_fault.gpa, EXITS_TEST_GPA); TEST_ASSERT_EQ(vcpu->run->memory_fault.size, EXITS_TEST_SIZE); =20 @@ -104,7 +105,8 @@ static void test_private_access_memslot_not_private(voi= d) exit_reason =3D run_vcpu_get_exit_reason(vcpu); =20 TEST_ASSERT_EQ(exit_reason, KVM_EXIT_MEMORY_FAULT); - TEST_ASSERT_EQ(vcpu->run->memory_fault.flags, KVM_MEMORY_EXIT_FLAG_PRIVAT= E); + TEST_ASSERT(vcpu->run->memory_fault.flags & KVM_MEMORY_EXIT_FLAG_PRIVATE, + "Memory fault didn't occur on a private memory access"); TEST_ASSERT_EQ(vcpu->run->memory_fault.gpa, EXITS_TEST_GPA); TEST_ASSERT_EQ(vcpu->run->memory_fault.size, EXITS_TEST_SIZE); =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5725342F717 for ; Thu, 16 Jul 2026 18:15:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225708; cv=none; b=YoymSFz3Rk0Si8i0o9CHzo77dLy8cAjrYeSOceECQw996LLdnz+Fa6J2mcBqbBDBHdy4pqE5LtQ1e0wwxoFjFJaXZsXuWLZlEwTVlE2FZy0dliDE1XQzi6oAP/CNCR8yQIMG7IMkpuaUu0ZHFpp76pj+8hW3mo0Yrlvl570b5nY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225708; c=relaxed/simple; bh=T+BmnNJAay5ftJzxLRL/6cpRSh2IaCrhRltBeUXgn+o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=iSx8ePPZA4XU5ECQXZImnfFGMEiuYlsqTNhXSuyHnZCs8xUknvzXcq8EvGHT/QfZSDFFwCD41CJL6WYOns9G31yDNu8yVdiBjQeT1hvqk4ywquqgx24F+9IS33lhHF8tGa0XYZzgDvUbXeHZxAOnBTKxVuMT0iLNX+lKH9hD/mQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=XsQB4Cc8; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="XsQB4Cc8" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225701; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=bF3tu8c+2PZb25EXkIFhdBMK4Xxc6B2MkNju/pWEoKo=; b=XsQB4Cc894yURCf3yQ5ejEhZQ8htUA7ULaMvdRfwtyeRkaDwjP0f9hUOpBawDTxqp696vD UQMjVggY4F0xeups1yorzq8su8eu72Jg3eU4IOzRLl/fmxKhcm5BNr5zy2uaamUbHCq6wH 5Rw4AxMc1Zujp1zijCigao6uXDuVu1g= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-292-vpcU5VLEOVa8lUfxCJkTgg-1; Thu, 16 Jul 2026 14:15:00 -0400 X-MC-Unique: vpcU5VLEOVa8lUfxCJkTgg-1 X-Mimecast-MFC-AGG-ID: vpcU5VLEOVa8lUfxCJkTgg_1784225699 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 581CB180074E; Thu, 16 Jul 2026 18:14:59 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id A77D2195604E; Thu, 16 Jul 2026 18:14:58 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com, Anish Moorthy , Sean Christopherson Subject: [PATCH 02/24] KVM: Define and communicate KVM_EXIT_MEMORY_FAULT RWX flags to userspace Date: Thu, 16 Jul 2026 14:14:34 -0400 Message-ID: <20260716181456.402786-3-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Anish Moorthy kvm_prepare_memory_fault_exit() already takes parameters describing the RWX-ness of the relevant access but doesn't actually do anything with them. Define and use the flags necessary to pass this information on to userspace. Suggested-by: Sean Christopherson Link: https://lore.kernel.org/kvm/ZR4N8cwzTMDanPUY@google.com/ Signed-off-by: Anish Moorthy Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- Documentation/virt/kvm/api.rst | 5 +++++ include/linux/kvm_host.h | 9 ++++++++- include/uapi/linux/kvm.h | 3 +++ 3 files changed, 16 insertions(+), 1 deletion(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index a5f9ee92f43e..1be1bc0de5d2 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -7259,6 +7259,9 @@ spec refer, https://github.com/riscv/riscv-sbi-doc. =20 /* KVM_EXIT_MEMORY_FAULT */ struct { + #define KVM_MEMORY_EXIT_FLAG_READ (1ULL << 0) + #define KVM_MEMORY_EXIT_FLAG_WRITE (1ULL << 1) + #define KVM_MEMORY_EXIT_FLAG_EXEC (1ULL << 2) #define KVM_MEMORY_EXIT_FLAG_PRIVATE (1ULL << 3) __u64 flags; __u64 gpa; @@ -7270,6 +7273,8 @@ could not be resolved by KVM. The 'gpa' and 'size' (= in bytes) describe the guest physical address range [gpa, gpa + size) of the fault. The 'flags' = field describes properties of the faulting access that are likely pertinent: =20 + - KVM_MEMORY_EXIT_FLAG_READ/WRITE/EXEC - When set, indicates that the mem= ory + fault occurred on a read/write/exec access respectively. - KVM_MEMORY_EXIT_FLAG_PRIVATE - When set, indicates the memory fault occ= urred on a private memory access. When clear, indicates the fault occurred o= n a shared access. diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index ab8cfaec82d3..2278b17f2f28 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2519,8 +2519,15 @@ static inline void kvm_prepare_memory_fault_exit(str= uct kvm_vcpu *vcpu, vcpu->run->memory_fault.gpa =3D gpa; vcpu->run->memory_fault.size =3D size; =20 - /* RWX flags are not (yet) defined or communicated to userspace. */ vcpu->run->memory_fault.flags =3D 0; + + if (is_write) + vcpu->run->memory_fault.flags |=3D KVM_MEMORY_EXIT_FLAG_WRITE; + else if (is_exec) + vcpu->run->memory_fault.flags |=3D KVM_MEMORY_EXIT_FLAG_EXEC; + else + vcpu->run->memory_fault.flags |=3D KVM_MEMORY_EXIT_FLAG_READ; + if (is_private) vcpu->run->memory_fault.flags |=3D KVM_MEMORY_EXIT_FLAG_PRIVATE; } diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 419011097fa8..720b1ffe880b 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -457,6 +457,9 @@ struct kvm_run { } notify; /* KVM_EXIT_MEMORY_FAULT */ struct { +#define KVM_MEMORY_EXIT_FLAG_READ (1ULL << 0) +#define KVM_MEMORY_EXIT_FLAG_WRITE (1ULL << 1) +#define KVM_MEMORY_EXIT_FLAG_EXEC (1ULL << 2) #define KVM_MEMORY_EXIT_FLAG_PRIVATE (1ULL << 3) __u64 flags; __u64 gpa; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D9B723AF678 for ; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225711; cv=none; b=qXNWN27oCeB8nMDqhCszmCgvR8+0WKjqeoXVoE/vDgEfX7jnXRFO0ERLCucUtnN5Eqi8RZRYfQxM6KAzk8FgYzFPrZxc+6I91lDthB6y9wyZ/CA7s/TOA6wD/qnwdMHiQ41HVnYQ/yyJ4Hr+cRIWLpt4kpGGejnqtR8cUbHqGLc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225711; c=relaxed/simple; bh=a8PF7SbM0fd04w0iq2kodW9LEQN4/g9z3k0clpm7csA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=OYmH+cjoq2cQY5oRyyvBQg+Sf9FNZFrd1ojx2hKesMyyEzYUaEdmEquLcvwj0VQ11ecMkui2+OOdvp1lPf6ck4Ss2rxmKkNFlD5lCRSTySbQuqXtRd/oruvq8QJ3JmdYaoZl11Tc0xUHQtUB8XAreRSNzjUa4tcTTYp9QETgiYk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=RYSkxA7W; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="RYSkxA7W" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225704; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=wNzIu7vEv4oBxoHkiuC1yFCUUcwdFOHiiEk7gNbeV+M=; b=RYSkxA7WeTKq9iX7d7cRrgKLvT3SldPQ7T+A1A7KKfg4kZYH7Nq2Lrqs9n2Ri2/G3hRMdx GI+bAfwCi6za0ASnnoxQQ6UvdG3qda3frF6ymbw/gsBsA3xr9POl6JG7KdGh+DtEl4vMYq hbg7ePJgLnQxpDt5ZwjXoYfvw693m2g= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-689-tNiJotxWN2iLopizr2b6qA-1; Thu, 16 Jul 2026 14:15:01 -0400 X-MC-Unique: tNiJotxWN2iLopizr2b6qA-1 X-Mimecast-MFC-AGG-ID: tNiJotxWN2iLopizr2b6qA_1784225700 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 128FF1800667; Thu, 16 Jul 2026 18:15:00 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 7E068195604E; Thu, 16 Jul 2026 18:14:59 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 03/24] KVM: x86: hyperv: Introduce memory fault on hcalls with bad ingpas Date: Thu, 16 Jul 2026 14:14:35 -0400 Message-ID: <20260716181456.402786-4-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/include/asm/kvm_host.h | 1 + arch/x86/kvm/hyperv.c | 32 ++++++++++++++++++++++++++++++++ arch/x86/kvm/x86.c | 7 +++++++ include/uapi/linux/kvm.h | 1 + 4 files changed, 41 insertions(+) diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_hos= t.h index b517257a6315..d0c1d0ab03c2 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -1274,6 +1274,7 @@ struct kvm_hv { struct kvm_hv_syndbg hv_syndbg; =20 bool xsaves_xsavec_checked; + bool hcall_fault_exit; }; #endif =20 diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c index 1ee0d23f8949..d688759df081 100644 --- a/arch/x86/kvm/hyperv.c +++ b/arch/x86/kvm/hyperv.c @@ -2531,11 +2531,30 @@ static bool hv_check_hypercall_access(struct kvm_vc= pu_hv *hv_vcpu, u16 code) return true; } =20 +static unsigned int kvm_hv_hypercall_mem_access(u16 code) +{ + switch (code) { + case HVCALL_SIGNAL_EVENT: + case HVCALL_FLUSH_VIRTUAL_ADDRESS_LIST: + case HVCALL_FLUSH_VIRTUAL_ADDRESS_LIST_EX: + case HVCALL_FLUSH_VIRTUAL_ADDRESS_SPACE: + case HVCALL_FLUSH_VIRTUAL_ADDRESS_SPACE_EX: + case HVCALL_SEND_IPI: + case HVCALL_SEND_IPI_EX: + return KVM_MEMORY_EXIT_FLAG_READ; + } + + return KVM_MEMORY_EXIT_FLAG_WRITE; +} + int kvm_hv_hypercall(struct kvm_vcpu *vcpu) { struct kvm_vcpu_hv *hv_vcpu =3D to_hv_vcpu(vcpu); + struct kvm *kvm =3D vcpu->kvm; struct kvm_hv_hcall hc; u64 ret =3D HV_STATUS_SUCCESS; + unsigned int access; + unsigned long addr; =20 /* * hypercall generates UD from non zero cpl and real mode @@ -2590,6 +2609,19 @@ int kvm_hv_hypercall(struct kvm_vcpu *vcpu) kvm_hv_hypercall_read_xmm(&hc); } =20 + if (!hc.fast && kvm->arch.hyperv.hcall_fault_exit) { + bool writable =3D true; + access =3D kvm_hv_hypercall_mem_access(hc.code); + addr =3D kvm_vcpu_gfn_to_hva_prot(vcpu, gpa_to_gfn(hc.ingpa), &writable); + if (addr =3D=3D KVM_HVA_ERR_BAD || + (access =3D=3D KVM_MEMORY_EXIT_FLAG_WRITE && !writable)) { + kvm_prepare_memory_fault_exit(vcpu, hc.ingpa, PAGE_SIZE, + access =3D=3D KVM_MEMORY_EXIT_FLAG_WRITE, + false, false); + return -EFAULT; + } + } + switch (hc.code) { case HVCALL_NOTIFY_LONG_SPIN_WAIT: if (unlikely(hc.rep || hc.var_cnt)) { diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 0626e835e9eb..352d289076f6 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -2224,6 +2224,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, lon= g ext) case KVM_CAP_HYPERV_CPUID: case KVM_CAP_HYPERV_ENFORCE_CPUID: case KVM_CAP_SYS_HYPERV_CPUID: + case KVM_CAP_HYPERV_HCALL_FAULT_EXIT: #endif case KVM_CAP_PCI_SEGMENT: case KVM_CAP_DEBUGREGS: @@ -4190,6 +4191,12 @@ int kvm_vm_ioctl_enable_cap(struct kvm *kvm, mutex_unlock(&kvm->lock); break; } +#ifdef CONFIG_KVM_HYPERV + case KVM_CAP_HYPERV_HCALL_FAULT_EXIT: + kvm->arch.hyperv.hcall_fault_exit =3D cap->args[0]; + r =3D 0; + break; +#endif default: r =3D -EINVAL; break; diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 720b1ffe880b..3b235e953923 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1000,6 +1000,7 @@ struct kvm_enable_cap { #define KVM_CAP_S390_KEYOP 247 #define KVM_CAP_S390_VSIE_ESAMODE 248 #define KVM_CAP_S390_HPAGE_2G 249 +#define KVM_CAP_HYPERV_HCALL_FAULT_EXIT 250 =20 struct kvm_irq_routing_irqchip { __u32 irqchip; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A70E748034F for ; Thu, 16 Jul 2026 18:16:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225778; cv=none; b=XD1fFffSAfOIZDAFNeJCWgb9yLiUmMrYac5AzvhSZcWc27+oBSoud9ykPx5poeQGmy5WHY3o6wqith8t4GRXsytHRrq6tz2vqHuVxOIVmC+rsYqI67PQYwfQ9f+J7ZdFYVabmMqY2GE0+OirsvlvqDg7rH2Ygjoeinz5f0/6ViQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225778; c=relaxed/simple; bh=4Dv48tDQIx40WR6Ur6ghchIGY+vgtRefoSNxFqnLb2Q=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=qoVfTyVjqg8DUzTbh0DAPLzoMLgX4hmQwHJm5t+pGnuqY/kHM492S8S+wgiMLRAEW8lWZW5/oGLFomBsM8WcBYbXOT1dJSnYG6nyYvefyGmdLu4QvFk2/u6Bdv/WFKh4Onrpub+/nqeWVGK918hrfIXLHddkPnhA1aCneOFw7js= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=LRIED+GQ; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="LRIED+GQ" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225772; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=yv9YPvkjkDF2qJ7dc9+Bs6eFu6QdBTMu8Ml4KSx/0K8=; b=LRIED+GQRgAhS0ZbqkZdtkXNKAg7O/JKLMLimTA52wwrGOhQaPW0SU3sMjiGTXOPuuCo9A h+zo6nd85LnOv2fu2xIQbD1995KPQm38R7qiNAbq9YOxjmeKxTM0rRceb5THN9cyrvwATE Fsx+9l683nq8qtUvizwuNonNeQFhtbM= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-435-HDruLpmEPVai2wWFknyzsA-1; Thu, 16 Jul 2026 14:15:01 -0400 X-MC-Unique: HDruLpmEPVai2wWFknyzsA-1 X-Mimecast-MFC-AGG-ID: HDruLpmEPVai2wWFknyzsA_1784225700 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 938EF1955F7E; Thu, 16 Jul 2026 18:15:00 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 20FD31955D86; Thu, 16 Jul 2026 18:15:00 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 04/24] KVM: x86/mmu: intersect writability from __kvm_faultin_pfn with fault->map_writable Date: Thu, 16 Jul 2026 14:14:36 -0400 Message-ID: <20260716181456.402786-5-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" fault->map_writable is currently a pure output of __kvm_faultin_pfn(), which is the only thing that restricts it. This will no longer hold once memory protections derived from memory attributes are applied: those compute their own access permissions that combine with those from __kvm_faultin_pfn(). Applying them *before* faulting in the pfn lets a fault that violates the attributes exit to userspace without the cost of gup and/or an async #PF; but it means that permissions will then be restricted in two independent steps, first by memory attributes and then by __kvm_faultin_pfn= (). Switch fault->map_writable to that model by letting kvm_mmu_faultin_pfn() only clear bits rather than assign them. No functional change intended: nothing writes fault->map_writable between the initializer and __kvm_mmu_faultin_pfn() yet, so the AND is equivalent to the assignment it replaces. Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 6c13da942bfc..cf6a409b76ae 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -4603,7 +4603,7 @@ static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *= vcpu, return r; } =20 - fault->map_writable =3D !(fault->slot->flags & KVM_MEM_READONLY); + fault->map_writable &=3D !(fault->slot->flags & KVM_MEM_READONLY); fault->max_level =3D kvm_max_level_for_order(max_order); =20 return RET_PF_CONTINUE; @@ -4613,13 +4613,14 @@ static int __kvm_mmu_faultin_pfn(struct kvm_vcpu *v= cpu, struct kvm_page_fault *fault) { unsigned int foll =3D fault->write ? FOLL_WRITE : 0; + bool writable; =20 if (fault->is_private || kvm_memslot_is_gmem_only(fault->slot)) return kvm_mmu_faultin_pfn_gmem(vcpu, fault); =20 foll |=3D FOLL_NOWAIT; fault->pfn =3D __kvm_faultin_pfn(fault->slot, fault->gfn, foll, - &fault->map_writable, &fault->refcounted_page); + &writable, &fault->refcounted_page); =20 /* * If resolving the page failed because I/O is needed to fault-in the @@ -4628,7 +4629,7 @@ static int __kvm_mmu_faultin_pfn(struct kvm_vcpu *vcp= u, * other failures are terminal, i.e. retrying won't help. */ if (fault->pfn !=3D KVM_PFN_ERR_NEEDS_IO) - return RET_PF_CONTINUE; + goto out_pf_continue; =20 if (!fault->prefetch && kvm_can_do_async_pf(vcpu)) { trace_kvm_try_async_get_page(fault->addr, fault->gfn); @@ -4649,8 +4650,10 @@ static int __kvm_mmu_faultin_pfn(struct kvm_vcpu *vc= pu, foll |=3D FOLL_INTERRUPTIBLE; foll &=3D ~FOLL_NOWAIT; fault->pfn =3D __kvm_faultin_pfn(fault->slot, fault->gfn, foll, - &fault->map_writable, &fault->refcounted_page); + &writable, &fault->refcounted_page); =20 +out_pf_continue: + fault->map_writable &=3D writable; return RET_PF_CONTINUE; } =20 @@ -4968,6 +4971,7 @@ static int kvm_mmu_do_page_fault(struct kvm_vcpu *vcp= u, gpa_t cr2_or_gpa, .is_private =3D err & PFERR_PRIVATE_ACCESS, =20 .pfn =3D KVM_PFN_ERR_FAULT, + .map_writable =3D true, }; int r; =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 87791444704 for ; Thu, 16 Jul 2026 18:16:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225778; cv=none; b=JDuEY99KYzVtcxrN7IKn8KKzY0u0xlG8rLRSkEGlb/v8Yu8vEOIot0GAWj1O+ykS5xoL40SCnotkkz9ryCnNmvi+lKYv3YrSB6rTTaif5L9b9O5MhxeC0gPcyUc/wlMe/LiQNx40nQwT4tKUe7vhCGfzOnhYJMNqJ0JtlebPv/Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225778; c=relaxed/simple; bh=sJl61TpXYYzOmR4Cot6bFVcB9aleftR4740JKOenP4A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=JY/ayS5PeGo4iE0W52fhMy8N2LeSy1cbW6JpYtJoZCGeP64aWIbejvH+WV0c8ZebQZiguK/wEOa/P9phHe/jzaSXTj98BVxwO+bTBUlfnitI2R3uKqIaQb0groK3FnDqw9JE47ymDC9ekvdlR6HqpUYwZ29pm/8C++Cqc9h0NQw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=EE/vTt8b; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="EE/vTt8b" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225772; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=1gnRsNsviw69Fi8JHkPSX5eQmaBMo6EzZm6CD4BcwSA=; b=EE/vTt8bCbPSz1eZeavy3DE9jVJOOTL4JpKyLkwHKAEiK9Tjo/ypGpgDoFt0DOeqUGnm23 hOZriaZQjG8b6fPz4gEoq98sSz3/CI6i4nsjsQYCKgigjcekSB9TbuzVBkwY+7GhKMukoy 5njF2JLpXsIpf7OGPJJxbQgCWqAVbjQ= Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-681-s8H3r7T_PFK0uyHS77QVfw-1; Thu, 16 Jul 2026 14:15:02 -0400 X-MC-Unique: s8H3r7T_PFK0uyHS77QVfw-1 X-Mimecast-MFC-AGG-ID: s8H3r7T_PFK0uyHS77QVfw_1784225701 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 34EB21955D59; Thu, 16 Jul 2026 18:15:01 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id B83AD195604E; Thu, 16 Jul 2026 18:15:00 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 05/24] KVM: x86/mmu: Extend map_writable to a full ACC_* mask Date: Thu, 16 Jul 2026 14:14:37 -0400 Message-ID: <20260716181456.402786-6-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" Support for memory protection attributes opens the door to installing non-executable mappings. Instead of introducing yet another member in struct kvm_page_fault and another argument to make_spte(), make the existing member map_writable a mask of ACC_* bits. This also avoids the need for make_spte() to map a single bool to either the NX bit or the XS/XU bits together. Unlike for mappings that are not writable because the fault did not request write premission, it is not not necessary to track executability for these SPTEs; the gfn is always available and it will be possible to access the attributes directly in FNAME(sync_spte). Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 21 ++++++++++++--------- arch/x86/kvm/mmu/mmu_internal.h | 2 +- arch/x86/kvm/mmu/paging_tmpl.h | 8 +++++--- arch/x86/kvm/mmu/spte.c | 12 ++++++------ arch/x86/kvm/mmu/spte.h | 2 +- arch/x86/kvm/mmu/tdp_mmu.c | 2 +- 6 files changed, 26 insertions(+), 21 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index cf6a409b76ae..424f4e113682 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -3074,7 +3074,7 @@ static int mmu_set_spte(struct kvm_vcpu *vcpu, struct= kvm_memory_slot *slot, u64 spte; =20 /* Prefetching always gets a writable pfn. */ - bool host_writable =3D !fault || fault->map_writable; + unsigned host_access =3D fault ? fault->host_access : ACC_ALL; bool prefetch =3D !fault || fault->prefetch; bool write_fault =3D fault && fault->write; =20 @@ -3111,7 +3111,7 @@ static int mmu_set_spte(struct kvm_vcpu *vcpu, struct= kvm_memory_slot *slot, } =20 wrprot =3D make_spte(vcpu, sp, slot, pte_access, gfn, pfn, *sptep, prefet= ch, - false, host_writable, &spte); + false, host_access, &spte); =20 if (*sptep =3D=3D spte) { ret =3D RET_PF_SPURIOUS; @@ -3558,7 +3558,7 @@ static int kvm_handle_noslot_fault(struct kvm_vcpu *v= cpu, =20 fault->slot =3D NULL; fault->pfn =3D KVM_PFN_NOSLOT; - fault->map_writable =3D false; + fault->host_access =3D 0; =20 /* * If MMIO caching is disabled, emulate immediately without @@ -4583,7 +4583,8 @@ static void kvm_mmu_finish_page_fault(struct kvm_vcpu= *vcpu, struct kvm_page_fault *fault, int r) { kvm_release_faultin_page(vcpu->kvm, fault->refcounted_page, - r =3D=3D RET_PF_RETRY, fault->map_writable); + r =3D=3D RET_PF_RETRY, + !!(fault->host_access & ACC_WRITE_MASK)); } =20 static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *vcpu, @@ -4603,9 +4604,10 @@ static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu = *vcpu, return r; } =20 - fault->map_writable &=3D !(fault->slot->flags & KVM_MEM_READONLY); - fault->max_level =3D kvm_max_level_for_order(max_order); + if (fault->slot->flags & KVM_MEM_READONLY) + fault->host_access &=3D ~ACC_WRITE_MASK; =20 + fault->max_level =3D kvm_max_level_for_order(max_order); return RET_PF_CONTINUE; } =20 @@ -4653,7 +4655,8 @@ static int __kvm_mmu_faultin_pfn(struct kvm_vcpu *vcp= u, &writable, &fault->refcounted_page); =20 out_pf_continue: - fault->map_writable &=3D writable; + if (!writable) + fault->host_access &=3D ~ACC_WRITE_MASK; return RET_PF_CONTINUE; } =20 @@ -4971,7 +4974,7 @@ static int kvm_mmu_do_page_fault(struct kvm_vcpu *vcp= u, gpa_t cr2_or_gpa, .is_private =3D err & PFERR_PRIVATE_ACCESS, =20 .pfn =3D KVM_PFN_ERR_FAULT, - .map_writable =3D true, + .host_access =3D ACC_ALL, }; int r; =20 @@ -5165,7 +5168,7 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcpu= , gfn_t gfn, kvm_pfn_t pfn) .gfn =3D gfn, .slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn), .pfn =3D pfn, - .map_writable =3D true, + .host_access =3D ACC_ALL, }; struct kvm *kvm =3D vcpu->kvm; int r; diff --git a/arch/x86/kvm/mmu/mmu_internal.h b/arch/x86/kvm/mmu/mmu_interna= l.h index c29002c60126..00215b9f309f 100644 --- a/arch/x86/kvm/mmu/mmu_internal.h +++ b/arch/x86/kvm/mmu/mmu_internal.h @@ -280,7 +280,7 @@ struct kvm_page_fault { unsigned long mmu_seq; kvm_pfn_t pfn; struct page *refcounted_page; - bool map_writable; + u8 host_access; =20 /* * Indicates the guest is trying to write a gfn that contains one or diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index e73fc09ec4db..1871d334fed7 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -933,7 +933,7 @@ static gpa_t FNAME(gva_to_gpa)(struct kvm_vcpu *vcpu, s= truct kvm_pagewalk *w, */ static int FNAME(sync_spte)(struct kvm_vcpu *vcpu, struct kvm_mmu_page *sp= , int i) { - bool host_writable; + u8 host_access; gpa_t first_pte_gpa; u64 *sptep, spte; struct kvm_memory_slot *slot; @@ -990,11 +990,13 @@ static int FNAME(sync_spte)(struct kvm_vcpu *vcpu, st= ruct kvm_mmu_page *sp, int =20 sptep =3D &sp->spt[i]; spte =3D *sptep; - host_writable =3D spte & shadow_host_writable_mask; + host_access =3D ACC_ALL; + if (!(spte & shadow_host_writable_mask)) + host_access &=3D ~ACC_WRITE_MASK; slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); make_spte(vcpu, sp, slot, pte_access, gfn, spte_to_pfn(spte), spte, true, true, - host_writable, &spte); + host_access, &spte); =20 /* * There is no need to mark the pfn dirty, as the new protections must diff --git a/arch/x86/kvm/mmu/spte.c b/arch/x86/kvm/mmu/spte.c index bdf72a98c19c..bc6d32c4c854 100644 --- a/arch/x86/kvm/mmu/spte.c +++ b/arch/x86/kvm/mmu/spte.c @@ -188,7 +188,7 @@ bool make_spte(struct kvm_vcpu *vcpu, struct kvm_mmu_pa= ge *sp, const struct kvm_memory_slot *slot, unsigned int pte_access, gfn_t gfn, kvm_pfn_t pfn, u64 old_spte, bool prefetch, bool synchronizing, - bool host_writable, u64 *new_spte) + unsigned int host_access, u64 *new_spte) { int level =3D sp->role.level; u64 spte =3D SPTE_MMU_PRESENT_MASK; @@ -206,6 +206,11 @@ bool make_spte(struct kvm_vcpu *vcpu, struct kvm_mmu_p= age *sp, if (!prefetch || synchronizing) spte |=3D shadow_accessed_mask; =20 + if (host_access & ACC_WRITE_MASK) + spte |=3D shadow_host_writable_mask; + + pte_access &=3D host_access; + /* * For simplicity, enforce the NX huge page mitigation even if not * strictly necessary. KVM could ignore the mitigation if paging is @@ -245,11 +250,6 @@ bool make_spte(struct kvm_vcpu *vcpu, struct kvm_mmu_p= age *sp, if (kvm_x86_ops.get_mt_mask) spte |=3D kvm_x86_call(get_mt_mask)(vcpu, gfn, kvm_is_mmio_pfn(pfn, &is_host_mmio)); - if (host_writable) - spte |=3D shadow_host_writable_mask; - else - pte_access &=3D ~ACC_WRITE_MASK; - if (shadow_me_value && !kvm_is_mmio_pfn(pfn, &is_host_mmio)) spte |=3D shadow_me_value; =20 diff --git a/arch/x86/kvm/mmu/spte.h b/arch/x86/kvm/mmu/spte.h index e730717824b3..589f3954633e 100644 --- a/arch/x86/kvm/mmu/spte.h +++ b/arch/x86/kvm/mmu/spte.h @@ -563,7 +563,7 @@ bool make_spte(struct kvm_vcpu *vcpu, struct kvm_mmu_pa= ge *sp, const struct kvm_memory_slot *slot, unsigned int pte_access, gfn_t gfn, kvm_pfn_t pfn, u64 old_spte, bool prefetch, bool synchronizing, - bool host_writable, u64 *new_spte); + unsigned int host_access, u64 *new_spte); u64 make_small_spte(struct kvm *kvm, u64 huge_spte, union kvm_mmu_page_role role, int index); u64 make_huge_spte(struct kvm *kvm, u64 small_spte, int level); diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c index ce3f2efadb05..4d9e7b6868e2 100644 --- a/arch/x86/kvm/mmu/tdp_mmu.c +++ b/arch/x86/kvm/mmu/tdp_mmu.c @@ -1143,7 +1143,7 @@ static int tdp_mmu_map_handle_target_level(struct kvm= _vcpu *vcpu, else wrprot =3D make_spte(vcpu, sp, fault->slot, sp->role.access, iter->gfn, fault->pfn, iter->old_spte, fault->prefetch, - false, fault->map_writable, &new_spte); + false, fault->host_access, &new_spte); =20 if (new_spte =3D=3D iter->old_spte) ret =3D RET_PF_SPURIOUS; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9EB00443E3E for ; Thu, 16 Jul 2026 18:15:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225714; cv=none; b=b7rWBJU2hNyjDv+3t36ykY2jbD3t8mD0jyi+48GPYJc/GzLNuLXNzfnraJZimoprRsZEA5dQQhqPJz9wcFmmMpUP81nfWuZGDmr4CkwC4R9p+QL9ZEj7wpmlcDvruPJHduvNxN96t/fh1D7uVE8DP/jZ4TzInqhPxKGZGOnDpmg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225714; c=relaxed/simple; bh=Yn6hB0cTunTTJYa9SwdDquNDODpP5QDIpcgYE69Rs74=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=swh8gZd/VPfjyI4Bv6uzFFDABMya14amk+IXUXhgE8ERhGcPYItm+Uchy1Cu7ossCJ4AE6CvmogJ98kwjOyUTI9W0cBx+6Xk4t35uvoe8ZNM/hyHuqVpp4lzjcA89C9OJ9K+6axiVdv8H6m0Y+wtCesnQDmLVgellriPEwrWjL0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=ObtiSO0E; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="ObtiSO0E" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225706; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=XG6Tb7rfF0+kImzRSQfE5PFeJcO4058fC8goQhkRW6Q=; b=ObtiSO0EjfI1frR8k8x6c5mkA7r7ZvuM0z/02Fo6NWRrr5vI3xSnvgmPriYC4onHwGeaMi 59YgYlBJp5eD3NSEbSjWARb1HKXIp+Cemty3EaXdhP9qhkKOuabGHETyRhjWDATGZEQCQ/ M+0UBDCiEcEmfhUxCYjt5oEApp3xnTk= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-210-gt5qS8mbNFaXx-LPygfNyg-1; Thu, 16 Jul 2026 14:15:02 -0400 X-MC-Unique: gt5qS8mbNFaXx-LPygfNyg-1 X-Mimecast-MFC-AGG-ID: gt5qS8mbNFaXx-LPygfNyg_1784225701 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id CC3171955EA2; Thu, 16 Jul 2026 18:15:01 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 5B5B4195604E; Thu, 16 Jul 2026 18:15:01 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 06/24] KVM: x86: Avoid warning when installing non-private memory attributes Date: Thu, 16 Jul 2026 14:14:38 -0400 Message-ID: <20260716181456.402786-7-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne In preparation to introducing RWX memory attributes. Make sure user-space is attempting to install a memory attribute with KVM_MEMORY_ATTRIBUTE_PRIVATE before throwing a warning on systems with no private memory support. The WARN is really a duplicate of the kvm_supported_mem_attributes() test in kvm_vm_ioctl_set_mem_attributes(), but then all of them should be redundant... Keep it as a special check for the private-memory attribute, to avoid that kvm_mmu_page_fault() sets PFERR_PRIVATE_ACCESS; that bit would send KVM down kvm_mmu_faultin_pfn_gmem(), which is such a wrong path that it's worth catching it early. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 20 ++++++++++++-------- 1 file changed, 12 insertions(+), 8 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 424f4e113682..055e0b45a8ee 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -8078,9 +8078,15 @@ static void hugepage_set_mixed(struct kvm_memory_slo= t *slot, gfn_t gfn, bool kvm_arch_pre_set_memory_attributes(struct kvm *kvm, struct kvm_gfn_range *range) { + unsigned long attrs =3D range->arg.attributes; struct kvm_memory_slot *slot =3D range->slot; int level; =20 + if (!kvm_arch_has_private_mem(kvm)) { + WARN_ON(attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE); + attrs &=3D ~KVM_MEMORY_ATTRIBUTE_PRIVATE; + } + /* * Zap SPTEs even if the slot can't be mapped PRIVATE. KVM x86 only * supports KVM_MEMORY_ATTRIBUTE_PRIVATE, and so it *seems* like KVM @@ -8092,9 +8098,6 @@ bool kvm_arch_pre_set_memory_attributes(struct kvm *k= vm, * Zapping SPTEs in this case ensures KVM will reassess whether or not * a hugepage can be used for affected ranges. */ - if (WARN_ON_ONCE(!kvm_arch_has_private_mem(kvm))) - return false; - if (WARN_ON_ONCE(range->end <=3D range->start)) return false; =20 @@ -8165,16 +8168,17 @@ bool kvm_arch_post_set_memory_attributes(struct kvm= *kvm, lockdep_assert_held_write(&kvm->mmu_lock); lockdep_assert_held(&kvm->slots_lock); =20 + if (!kvm_arch_has_private_mem(kvm)) { + WARN_ON(attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE); + attrs &=3D ~KVM_MEMORY_ATTRIBUTE_PRIVATE; + } + /* * Calculate which ranges can be mapped with hugepages even if the slot * can't map memory PRIVATE. KVM mustn't create a SHARED hugepage over * a range that has PRIVATE GFNs, and conversely converting a range to * SHARED may now allow hugepages. - */ - if (WARN_ON_ONCE(!kvm_arch_has_private_mem(kvm))) - return false; - - /* + * * The sequence matters here: upper levels consume the result of lower * level's scanning. */ --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 583F7433E71 for ; Thu, 16 Jul 2026 18:15:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225710; cv=none; b=Q10oqvxo5ogLE+32IH6qkJu7cDWrYHjUVaWHp8saVzfESlIeipfxy3Z7TikpFk1320Cia3f+h+h1KjgtcruCuP0tzTMqtzk562LZz/7/X1+9EuobrwJCqvDOhTXO4/VeNBS9YCicPbc9l0acSzND+Kc3aSUPmOmGOotcbw+AdkQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225710; c=relaxed/simple; bh=2rVrADV0daX86CQLSPFvHaFYidtq/pYjVHvQAhS/5po=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=ZhAdCxcfHCMpiDU4Qbp58OSEZHcmSSS4tmcUkCKNr+AVhB+sEsFyTImp+IIe3lMWqDUrNGDhUS5U+vFN7QzJa/9sxQKihnNZ48Q2ZEaRY12oPhSzFwS0e3BCQU8i0E8BXjplTGX602iy0L0eBYvav9mmpY92bvn9ZI1eMuKFz28= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=HepXAIIf; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="HepXAIIf" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225705; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6jGh0+icnGcvjYFFwFz54AZGxyxD4J+S55ZiS6S4zxs=; b=HepXAIIfr/aceOgEL+6eLtWyLA4yJ0qnyjOizFAV63/+s0tvQqsOAwunUrECEOe8DA8aZH 8mnFcXBE/6dtzk05PEmgWg6l74Z/c/WnCjWFmXCWOrmqd4esBK8DxD44cqCA7zwRSU6wiq EiV6wnUtIaqkFh5soPRz5y21DK5egCs= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-516-tysaZ_vxOimQNrTmmLQbVQ-1; Thu, 16 Jul 2026 14:15:03 -0400 X-MC-Unique: tysaZ_vxOimQNrTmmLQbVQ-1 X-Mimecast-MFC-AGG-ID: tysaZ_vxOimQNrTmmLQbVQ_1784225702 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 6FD0B1800578; Thu, 16 Jul 2026 18:15:02 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id F287D195604E; Thu, 16 Jul 2026 18:15:01 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 07/24] KVM: x86/mmu: Init memslot hugepage information for non-private_mem VMs too Date: Thu, 16 Jul 2026 14:14:39 -0400 Message-ID: <20260716181456.402786-8-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne The list of supported memory attributes is about to grow, and they will be available on any VM type (as opposed to only ones targeted at confidential computing). As such, update the check in kvm_mmu_init_memslot_memory_attributes() to initialize huge page information if any kind of memory attributes is available for the VM. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 2 +- include/linux/kvm_host.h | 5 +++++ virt/kvm/kvm_main.c | 2 +- 3 files changed, 7 insertions(+), 2 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 055e0b45a8ee..80f581afb6a0 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -8231,7 +8231,7 @@ void kvm_mmu_init_memslot_memory_attributes(struct kv= m *kvm, { int level; =20 - if (!kvm_arch_has_private_mem(kvm)) + if (!kvm_supported_mem_attributes(kvm)) return; =20 for (level =3D PG_LEVEL_2M; level <=3D KVM_MAX_HUGEPAGE_LEVEL; level++) { diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 2278b17f2f28..4731e8a9e6db 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2552,12 +2552,17 @@ bool kvm_arch_pre_set_memory_attributes(struct kvm = *kvm, struct kvm_gfn_range *range); bool kvm_arch_post_set_memory_attributes(struct kvm *kvm, struct kvm_gfn_range *range); +u64 kvm_supported_mem_attributes(struct kvm *kvm); =20 static inline bool kvm_mem_is_private(struct kvm *kvm, gfn_t gfn) { return kvm_get_memory_attributes(kvm, gfn) & KVM_MEMORY_ATTRIBUTE_PRIVATE; } #else +static inline u64 kvm_supported_mem_attributes(struct kvm *kvm) +{ + return 0; +} static inline bool kvm_mem_is_private(struct kvm *kvm, gfn_t gfn) { return false; diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index e44c20c04961..4ecc3579d163 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2419,7 +2419,7 @@ static int kvm_vm_ioctl_clear_dirty_log(struct kvm *k= vm, #endif /* CONFIG_KVM_GENERIC_DIRTYLOG_READ_PROTECT */ =20 #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES -static u64 kvm_supported_mem_attributes(struct kvm *kvm) +u64 kvm_supported_mem_attributes(struct kvm *kvm) { if (!kvm || kvm_arch_has_private_mem(kvm)) return KVM_MEMORY_ATTRIBUTE_PRIVATE; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 65CDF2D0606 for ; Thu, 16 Jul 2026 18:15:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225714; cv=none; b=YegFPBtJc8rVG6AYbtlQV9W4Uo2YTzzp2rLObR4rcM9VnhZDNmeuwWuj8IVctgylJ/KXMriDIN+scen5Bj3bqGxkO26VTg1inEoWwbaRZHfBAJ0nFw1aT4jM+CzmQo9gCeNf6zmkV6AFBnJMgA3lWRpKG4Ohp0cq++zEAdVGNzs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225714; c=relaxed/simple; bh=ymQ1dL8MxCQZw1HZ3sOtC6Cl0rDVVCH/oyQIhxYOd84=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Q39RoNh9txQou577rXW2dYf/Eq0xjBrWKOrhwOslKIsxtTmTUf/2cPPmMwz1pT8crrG21WSmXUKsI6lQcc0eG6/EoeM8J2k2iquMB59UzG43peEwhaM74HDr8eywU2T3Ot9xS5Sd/zN27rF965GVnoIE3RlxKjm7I2W10LdI2E4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=RF2d81tv; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="RF2d81tv" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225707; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=h40LU/DLcCjOqTvxof1lSNIVRlTqzpnwdFbpy6VN+1Y=; b=RF2d81tv5dE2rdXWX2cyV6tt/epsSGueuyH82zLCgkK1GiVKB94CjL40cEuDRIqM5ahsjw 2xRlLeDcJnk2vh7ymJlvxbPi7GpuMUN2mU9wDP3wIcZ/EFupunDoE3wbFz0G2mlHV1XYp4 DJyaUx+uOM52HGt4CLja6YQMXa1iUVE= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-410-e269NZb_Mlmuy1RwAhf36A-1; Thu, 16 Jul 2026 14:15:04 -0400 X-MC-Unique: e269NZb_Mlmuy1RwAhf36A-1 X-Mimecast-MFC-AGG-ID: e269NZb_Mlmuy1RwAhf36A_1784225703 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 2A36718001F3; Thu, 16 Jul 2026 18:15:03 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 95632195604E; Thu, 16 Jul 2026 18:15:02 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 08/24] KVM: pass kvm == NULL case to kvm_arch_has_private_mem Date: Thu, 16 Jul 2026 14:14:40 -0400 Message-ID: <20260716181456.402786-9-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" Allow the architecture-specific code to enable CONFIG_KVM_GENERIC_MEMORY_AT= TRIBUTES without exposing KVM_MEMORY_ATTRIBUTE_PRIVATE. This is mostly for consistency after introducing memory protection attributes; an architecture like Arm might want to add support for KVM_MEMORY_ATTRIBUTE_PRIVATE without having the protection attributes show up in KVM_CHECK_EXTENSION. Signed-off-by: Paolo Bonzini --- arch/x86/include/asm/kvm_host.h | 2 +- virt/kvm/kvm_main.c | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_hos= t.h index d0c1d0ab03c2..5d329fbb8acb 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2005,7 +2005,7 @@ enum kvm_intr_type { (!!in_nmi() =3D=3D ((vcpu)->arch.handling_intr_from_guest =3D=3D KVM_HAN= DLING_NMI))) =20 #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES -#define kvm_arch_has_private_mem(kvm) ((kvm)->arch.has_private_mem) +#define kvm_arch_has_private_mem(kvm) (!(kvm) || (kvm)->arch.has_private_m= em) #endif =20 #define kvm_arch_has_readonly_mem(kvm) (!(kvm)->arch.has_protected_state) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 4ecc3579d163..fa4473d7c920 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2421,7 +2421,7 @@ static int kvm_vm_ioctl_clear_dirty_log(struct kvm *k= vm, #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES u64 kvm_supported_mem_attributes(struct kvm *kvm) { - if (!kvm || kvm_arch_has_private_mem(kvm)) + if (kvm_arch_has_private_mem(kvm)) return KVM_MEMORY_ATTRIBUTE_PRIVATE; =20 return 0; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 45011324B24 for ; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225713; cv=none; b=aewza/3DyzPUJ/On5cyezuSMAZxfiLAkfcdkkmkG5tPiyHrHw3taL9IocbFDyfw/uLOaR+hiHpjYv1Mq8ODxeytp3SZ+9jLBE4UYU0uWlWHKCR74GEES1suxEUkc2mLp1535uF7piR2EOGsIWTPf4B5EzFZEoSmQcRvh9a2zx54= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225713; c=relaxed/simple; bh=eEza9SAclNDjEeSbNvKxGGfBlR++oqUO6WMBMW+L2pk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=ZrBFwVEU5Wmow48O2M/3w7/RSfeM7Iu21l+hQ8ZZBxM9cEPFrW6YFgUD0zvvC1u2GmLN276ouHbobnpWo6qFVKar3EsVpZEbPVQChh0Nmu4FyqM39JAM0APNZCTuB0zH2XkQmQljDeJfj+u4Z1gCj53ZbYfQ/JUhy/8JWUJqD0o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=JxGGx1ZO; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="JxGGx1ZO" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225706; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=8ywHgIy2wj+S8WnAk9F643i7jhZecJqHJ9G3mfZTAtw=; b=JxGGx1ZO21HEwXv0NVvh9vRghv3Gb3zC+PUdM2rBRP2oYI6j9P9WTg7oXlOW0eWiQci8Lz 6hz7sHSCL885noUs1O1yaBiXiGw46UAMe85bk7jx5nFsNHi5ZWYv9KGv13IbdpabriOYf3 YG63Evlr7rVMWXhmo2XUWHpT+gknMHA= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-345-QEAVmtG_OeCjROLMAosSxQ-1; Thu, 16 Jul 2026 14:15:04 -0400 X-MC-Unique: QEAVmtG_OeCjROLMAosSxQ-1 X-Mimecast-MFC-AGG-ID: QEAVmtG_OeCjROLMAosSxQ_1784225703 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id A92591800677; Thu, 16 Jul 2026 18:15:03 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 38872195604E; Thu, 16 Jul 2026 18:15:03 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 09/24] KVM: Introduce NR/NW/NX memory attributes Date: Thu, 16 Jul 2026 14:14:41 -0400 Message-ID: <20260716181456.402786-10-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Introduce memory attributes to map guest physical memory regions as non-readable, non-writable, and/or non-executable. Only a subset of flag combinations is supported. Notably write and exec permissions require read permission, and memory protection attributes are incompatible with private memory. As mentioned in 5a475554db1e ("KVM: Introduce per-page memory attributes", 2023-11-13), bits 0-2 of the memory attributes were reserved for RWX protection; they are negated to support current memory attribute users which use 0 to indicate no special treatment. Since 0 is not available, a non-negated version of the flags would need an extra bit to express no-access (R=3D0/W=3D0/X=3D0) mappings. Unfortunately this precaution did not age too well; KVM now supports MBEC/GMET and adding mode-based memory protections will require a non-contiguous bit. But that's something left for later. Different architectures may have different limitations on the set of valid protections, for example execution-only and XU=3D0 mappings are supported by Intel but not AMD processors[1]. So, add an architecture-spec= ific callback and add a basic implementation for x86. [1] When adding support for MBEC/GMET, since NX would remain to mean no execution at all, it is possible to use either an NXS bit or two separate NXS/NXU bits in addition to NX. The former would only support permissions that are available with either MBEC or GMET. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- Documentation/virt/kvm/api.rst | 14 +++++++--- arch/x86/include/asm/kvm_host.h | 1 + arch/x86/kvm/Kconfig | 4 +-- arch/x86/kvm/mmu/mmu.c | 47 +++++++++++++++++++++++++-------- arch/x86/kvm/x86.c | 2 -- include/linux/kvm_host.h | 28 +++++++++++++++++++- include/uapi/linux/kvm.h | 3 +++ virt/kvm/kvm_main.c | 32 +++++++++++++++++++--- 8 files changed, 107 insertions(+), 24 deletions(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index 1be1bc0de5d2..d11142c5eca1 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -6394,15 +6394,23 @@ of guest physical memory. __u64 flags; }; =20 + #define KVM_MEMORY_ATTRIBUTE_NR (1ULL << 0) + #define KVM_MEMORY_ATTRIBUTE_NW (1ULL << 1) + #define KVM_MEMORY_ATTRIBUTE_NX (1ULL << 2) #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3) =20 The address and size must be page aligned. The supported attributes can be retrieved via ioctl(KVM_CHECK_EXTENSION) on KVM_CAP_MEMORY_ATTRIBUTES. If executed on a VM, KVM_CAP_MEMORY_ATTRIBUTES precisely returns the attribut= es supported by that VM. If executed at system scope, KVM_CAP_MEMORY_ATTRIBU= TES -returns all attributes supported by KVM. The only attribute defined at th= is -time is KVM_MEMORY_ATTRIBUTE_PRIVATE, which marks the associated gfn as be= ing -guest private memory. +returns all attributes supported by KVM. The attribute defined at this +time are: + + - KVM_MEMORY_ATTRIBUTE_NR/NW/NX - Respectively marks the memory region as + non-read, non-write and/or non-exec. Note that write-only, exec-only a= nd + write-exec mappings are not supported. + - KVM_MEMORY_ATTRIBUTE_PRIVATE - Which marks the associated gfn as being = guest + private memory. =20 Note, there is no "get" API. Userspace is responsible for explicitly trac= king the state of a gfn/page as needed. diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_hos= t.h index 5d329fbb8acb..268e98a36643 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2005,6 +2005,7 @@ enum kvm_intr_type { (!!in_nmi() =3D=3D ((vcpu)->arch.handling_intr_from_guest =3D=3D KVM_HAN= DLING_NMI))) =20 #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES +#define kvm_arch_has_memory_protection_attributes(kvm) (!(kvm) || !(kvm)->= arch.has_private_mem) #define kvm_arch_has_private_mem(kvm) (!(kvm) || (kvm)->arch.has_private_m= em) #endif =20 diff --git a/arch/x86/kvm/Kconfig b/arch/x86/kvm/Kconfig index 801bf9e520db..d9c5bd71e1c5 100644 --- a/arch/x86/kvm/Kconfig +++ b/arch/x86/kvm/Kconfig @@ -48,6 +48,7 @@ config KVM_X86 select KVM_GENERIC_PRE_FAULT_MEMORY select KVM_WERROR if WERROR select KVM_GUEST_MEMFD if X86_64 + select KVM_GENERIC_MEMORY_ATTRIBUTES =20 config KVM tristate "Kernel-based Virtual Machine (KVM) support" @@ -84,7 +85,6 @@ config KVM_SW_PROTECTED_VM bool "Enable support for KVM software-protected VMs" depends on EXPERT depends on KVM_X86 && X86_64 - select KVM_GENERIC_MEMORY_ATTRIBUTES help Enable support for KVM software-protected VMs. Currently, software- protected VMs are purely a development and testing vehicle for @@ -135,7 +135,6 @@ config KVM_INTEL_TDX bool "Intel Trust Domain Extensions (TDX) support" default y depends on INTEL_TDX_HOST - select KVM_GENERIC_MEMORY_ATTRIBUTES select HAVE_KVM_ARCH_GMEM_POPULATE help Provides support for launching Intel Trust Domain Extensions (TDX) @@ -159,7 +158,6 @@ config KVM_AMD_SEV depends on KVM_AMD && X86_64 depends on CRYPTO_DEV_SP_PSP && !(KVM_AMD=3Dy && CRYPTO_DEV_CCP_DD=3Dm) select ARCH_HAS_CC_PLATFORM - select KVM_GENERIC_MEMORY_ATTRIBUTES select HAVE_KVM_ARCH_GMEM_PREPARE select HAVE_KVM_ARCH_GMEM_INVALIDATE select HAVE_KVM_ARCH_GMEM_POPULATE diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 80f581afb6a0..b6463b0b0b6d 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -8056,7 +8056,6 @@ void kvm_mmu_pre_destroy_vm(struct kvm *kvm) vhost_task_stop(kvm->arch.nx_huge_page_recovery_thread); } =20 -#ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES static bool hugepage_test_mixed(struct kvm_memory_slot *slot, gfn_t gfn, int level) { @@ -8088,16 +8087,23 @@ bool kvm_arch_pre_set_memory_attributes(struct kvm = *kvm, } =20 /* - * Zap SPTEs even if the slot can't be mapped PRIVATE. KVM x86 only - * supports KVM_MEMORY_ATTRIBUTE_PRIVATE, and so it *seems* like KVM - * can simply ignore such slots. But if userspace is making memory - * PRIVATE, then KVM must prevent the guest from accessing the memory - * as shared. And if userspace is making memory SHARED and this point - * is reached, then at least one page within the range was previously - * PRIVATE, i.e. the slot's possible hugepage ranges are changing. - * Zapping SPTEs in this case ensures KVM will reassess whether or not - * a hugepage can be used for affected ranges. + * For KVM_MEMORY_ATTRIBUTE_PRIVATE: + * Zap SPTEs even if the slot can't be mapped PRIVATE. KVM x86 only + * supports KVM_MEMORY_ATTRIBUTE_PRIVATE, and so it *seems* like KVM + * can simply ignore such slots. But if userspace is making memory + * PRIVATE, then KVM must prevent the guest from accessing the memory + * as shared. And if userspace is making memory SHARED and this point + * is reached, then at least one page within the range was previously + * PRIVATE, i.e. the slot's possible hugepage ranges are changing. + * Zapping SPTEs in this case ensures KVM will reassess whether or not + * a hugepage can be used for affected ranges. + * + * For KVM_MEMORY_ATTRIBUTE_NR/NW/NX: + * Zap even when loosening restrictions R=3D>RW, which is nost strictly + * necessary, but will allow KVM to reasses whether a hugepage can be + * used for the affected pages. */ + if (WARN_ON_ONCE(range->end <=3D range->start)) return false; =20 @@ -8262,4 +8268,23 @@ void kvm_mmu_init_memslot_memory_attributes(struct k= vm *kvm, } } } -#endif + +/* The bits are flipped but remain in the same position. */ +#define KVM_PROT_READ KVM_MEMORY_ATTRIBUTE_NR +#define KVM_PROT_WRITE KVM_MEMORY_ATTRIBUTE_NW +#define KVM_PROT_EXEC KVM_MEMORY_ATTRIBUTE_NX + +bool kvm_arch_mem_attributes_supported_prot(struct kvm *kvm, unsigned long= attrs) +{ + unsigned long prot =3D (attrs & KVM_MEMORY_ATTRIBUTE_PROT) ^ KVM_MEMORY_A= TTRIBUTE_PROT; + + /* Private memory and access permissions are incompatible */ + if (attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE) + return false; + + /* For now do now support exec-only, even though EPT can handle it. */ + if (prot && !(prot & KVM_PROT_READ)) + return false; + + return true; +} diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 352d289076f6..e2240f817c22 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -10084,9 +10084,7 @@ static int kvm_alloc_memslot_metadata(struct kvm *k= vm, } } =20 -#ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES kvm_mmu_init_memslot_memory_attributes(kvm, slot); -#endif =20 if (kvm_page_track_create_memslot(kvm, slot, npages)) goto out_free; diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 4731e8a9e6db..2dd1b65799e4 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -722,7 +722,9 @@ static inline int kvm_arch_vcpu_memslots_id(struct kvm_= vcpu *vcpu) } #endif =20 -#ifndef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES +#ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES +bool kvm_arch_mem_attributes_supported_prot(struct kvm *kvm, unsigned long= attrs); +#else static inline bool kvm_arch_has_private_mem(struct kvm *kvm) { return false; @@ -2540,7 +2542,25 @@ static inline bool kvm_memslot_is_gmem_only(const st= ruct kvm_memory_slot *slot) return slot->flags & KVM_MEMSLOT_GMEM_ONLY; } =20 +static inline bool kvm_mem_attributes_may_read(u64 attrs) +{ + return !(attrs & KVM_MEMORY_ATTRIBUTE_NR); +} + +static inline bool kvm_mem_attributes_may_write(u64 attrs) +{ + return !(attrs & KVM_MEMORY_ATTRIBUTE_NW); +} + +static inline bool kvm_mem_attributes_may_exec(u64 attrs) +{ + return !(attrs & KVM_MEMORY_ATTRIBUTE_NX); +} + #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES +#define KVM_MEMORY_ATTRIBUTE_PROT \ + (KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW | KVM_MEMORY_ATTRIBUTE= _NX) + static inline unsigned long kvm_get_memory_attributes(struct kvm *kvm, gfn= _t gfn) { return xa_to_value(xa_load(&kvm->mem_attr_array, gfn)); @@ -2552,6 +2572,7 @@ bool kvm_arch_pre_set_memory_attributes(struct kvm *k= vm, struct kvm_gfn_range *range); bool kvm_arch_post_set_memory_attributes(struct kvm *kvm, struct kvm_gfn_range *range); +bool kvm_mem_attributes_valid(struct kvm *kvm, unsigned long attrs); u64 kvm_supported_mem_attributes(struct kvm *kvm); =20 static inline bool kvm_mem_is_private(struct kvm *kvm, gfn_t gfn) @@ -2559,6 +2580,11 @@ static inline bool kvm_mem_is_private(struct kvm *kv= m, gfn_t gfn) return kvm_get_memory_attributes(kvm, gfn) & KVM_MEMORY_ATTRIBUTE_PRIVATE; } #else +static inline bool kvm_mem_attributes_valid(struct kvm *kvm, + unsigned long attrs) +{ + return false; +} static inline u64 kvm_supported_mem_attributes(struct kvm *kvm) { return 0; diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 3b235e953923..6a96eb3218dc 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1653,6 +1653,9 @@ struct kvm_memory_attributes { __u64 flags; }; =20 +#define KVM_MEMORY_ATTRIBUTE_NR (1ULL << 0) +#define KVM_MEMORY_ATTRIBUTE_NW (1ULL << 1) +#define KVM_MEMORY_ATTRIBUTE_NX (1ULL << 2) #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3) =20 #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest= _memfd) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index fa4473d7c920..a47b62c2c9ce 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2421,10 +2421,15 @@ static int kvm_vm_ioctl_clear_dirty_log(struct kvm = *kvm, #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES u64 kvm_supported_mem_attributes(struct kvm *kvm) { - if (kvm_arch_has_private_mem(kvm)) - return KVM_MEMORY_ATTRIBUTE_PRIVATE; + u64 supported_attrs =3D 0; =20 - return 0; + if (kvm_arch_has_memory_protection_attributes(kvm)) + supported_attrs |=3D KVM_MEMORY_ATTRIBUTE_PROT; + + if (kvm_arch_has_private_mem(kvm)) + supported_attrs |=3D KVM_MEMORY_ATTRIBUTE_PRIVATE; + + return supported_attrs; } =20 /* @@ -2596,6 +2601,25 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm= , gfn_t start, gfn_t end, =20 return r; } + +bool __weak kvm_arch_mem_attributes_supported_prot(struct kvm *kvm, unsign= ed long attrs) +{ + WARN_ON_ONCE("KVM_MEMORY_ATTRIBUTE_PROT requires kvm_arch_mem_attributes_= supported_prot()"); + return false; +} + +bool kvm_mem_attributes_valid(struct kvm *kvm, unsigned long attrs) +{ + if (attrs & ~kvm_supported_mem_attributes(kvm)) + return false; + + if ((attrs & KVM_MEMORY_ATTRIBUTE_PROT) && + !kvm_arch_mem_attributes_supported_prot(kvm, attrs)) + return false; + + return true; +} + static int kvm_vm_ioctl_set_mem_attributes(struct kvm *kvm, struct kvm_memory_attributes *attrs) { @@ -2604,7 +2628,7 @@ static int kvm_vm_ioctl_set_mem_attributes(struct kvm= *kvm, /* flags is currently not used. */ if (attrs->flags) return -EINVAL; - if (attrs->attributes & ~kvm_supported_mem_attributes(kvm)) + if (!kvm_mem_attributes_valid(kvm, attrs->attributes)) return -EINVAL; if (attrs->size =3D=3D 0 || attrs->address + attrs->size < attrs->address) return -EINVAL; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1D3C446042 for ; Thu, 16 Jul 2026 18:15:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225722; cv=none; b=jEY6e8vJ+Xy2OMlUzi9qFuwM9B7tz3PY7ONYFPWUfw+sn/2UZnL1xZQ5wZOR9GQb/X06hgU3pLrK4RbcK2F9b9KELM3oEOy3EhpiNp1YzgyptFTGGN6tqgtAPiVYqcct8a8En88JdT29qhhdRKOYWAQodbW3SijcekBrLrPM5yE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225722; c=relaxed/simple; bh=RbB7+rV+Unnkg8P9NDzl3Ie4WCg1IS/jJMOdDDxS3Gg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=MUbnOw4aYA09SHeymvazZWoBjrAXQoC9FyUa+GCB7HEFPl6KAUA0YHGc+BUDUtObof/JpU5mUIhUZLRAQmLzs4b8Lw1GvKt7/20F22YKRd0bitBF8/dRTZ2mmeag2wX3P6LE5yUvwmW5JXF7PLBa+NnTV/003K+oWq1zVGIeZv0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=is+PyTG3; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="is+PyTG3" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225712; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Tw1WPOSN+5h2YK/GVlX7f0rfZxbYC8aN2XeVHPaNYxU=; b=is+PyTG3OzkBnsxacjXT6pxUhxYIqNKAuCOs0Z0Blq8v9CGBcXl5wvOzBnghhFelrDb4O8 8r5yF2+/BzS4X7p1aucHefnDziha3/NFGfNL8+0Em8gsSAZBvP4HblwdXgwcM3ZpfDBYub VdO+dOcOejT2/8+qunrSPXhkElEFBqQ= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-584-Z8Gc3n8-M_iYNGzGbcFgEg-1; Thu, 16 Jul 2026 14:15:06 -0400 X-MC-Unique: Z8Gc3n8-M_iYNGzGbcFgEg-1 X-Mimecast-MFC-AGG-ID: Z8Gc3n8-M_iYNGzGbcFgEg_1784225704 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id B64181800672; Thu, 16 Jul 2026 18:15:04 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 481AB1800480; Thu, 16 Jul 2026 18:15:04 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 10/24] KVM: Include memory protections in result of gfn->hva conversion Date: Thu, 16 Jul 2026 14:14:42 -0400 Message-ID: <20260716181456.402786-11-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne All paths that lead to guest memory accesses now need to check whether memory attributes allow that access. Users of gfn_to_hva, kvm_vcpu_gfn_to_hva and their *_prot variant can get it more or less for free via erroneous return values, so do this first. Note however that this is not true of the variants that take a cached kvm_memslots pointer. These include caches (gfn-to-hva and gfn-to-pfn) and page faults, both of which will need specific changes; but the more optimized functions in virt/kvm/kvm_main.c such as kvm_read_guest() and kvm_write_guest() also retrieve the memslot high in the call chain, and therefore they will need changes in __kvm_read/write_guest_page(). Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- include/linux/kvm_host.h | 18 +++++++++++++ virt/kvm/kvm_main.c | 55 +++++++++++++++++++++++++++++++++++++--- 2 files changed, 69 insertions(+), 4 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 2dd1b65799e4..f10ab70b99dc 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2593,8 +2593,26 @@ static inline bool kvm_mem_is_private(struct kvm *kv= m, gfn_t gfn) { return false; } +static inline unsigned long kvm_get_memory_attributes(struct kvm *kvm, gfn= _t gfn) +{ + return 0; +} #endif /* CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES */ =20 +static inline int kvm_mem_attributes_may_read_gfn(struct kvm *kvm, gfn_t g= fn) +{ + unsigned long attrs =3D kvm_get_memory_attributes(kvm, gfn); + + return kvm_mem_attributes_may_read(attrs); +} + +static inline int kvm_mem_attributes_may_write_gfn(struct kvm *kvm, gfn_t = gfn) +{ + unsigned long attrs =3D kvm_get_memory_attributes(kvm, gfn); + + return kvm_mem_attributes_may_write(attrs); +} + #ifdef CONFIG_KVM_GUEST_MEMFD int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index a47b62c2c9ce..56016aab0aad 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2764,13 +2764,38 @@ EXPORT_SYMBOL_FOR_KVM_INTERNAL(gfn_to_hva_memslot); =20 unsigned long gfn_to_hva(struct kvm *kvm, gfn_t gfn) { - return gfn_to_hva_many(gfn_to_memslot(kvm, gfn), gfn, NULL); + unsigned long addr; + + addr =3D gfn_to_hva_many(gfn_to_memslot(kvm, gfn), gfn, NULL); + if (kvm_is_error_hva(addr)) + return addr; + + if (!kvm_mem_attributes_may_read_gfn(kvm, gfn)) + return KVM_HVA_ERR_BAD; + + if (!kvm_mem_attributes_may_write_gfn(kvm, gfn)) + return KVM_HVA_ERR_RO_BAD; + + return addr; } EXPORT_SYMBOL_FOR_KVM_INTERNAL(gfn_to_hva); =20 unsigned long kvm_vcpu_gfn_to_hva(struct kvm_vcpu *vcpu, gfn_t gfn) { - return gfn_to_hva_many(kvm_vcpu_gfn_to_memslot(vcpu, gfn), gfn, NULL); + struct kvm *kvm =3D vcpu->kvm; + unsigned long addr; + + addr =3D gfn_to_hva_many(kvm_vcpu_gfn_to_memslot(vcpu, gfn), gfn, NULL); + if (kvm_is_error_hva(addr)) + return addr; + + if (!kvm_mem_attributes_may_read_gfn(kvm, gfn)) + return KVM_HVA_ERR_BAD; + + if (!kvm_mem_attributes_may_write_gfn(kvm, gfn)) + return KVM_HVA_ERR_RO_BAD; + + return addr; } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_gfn_to_hva); =20 @@ -2796,15 +2821,37 @@ unsigned long gfn_to_hva_memslot_prot(struct kvm_me= mory_slot *slot, unsigned long gfn_to_hva_prot(struct kvm *kvm, gfn_t gfn, bool *writable) { struct kvm_memory_slot *slot =3D gfn_to_memslot(kvm, gfn); + unsigned long addr; =20 - return gfn_to_hva_memslot_prot(slot, gfn, writable); + addr =3D gfn_to_hva_memslot_prot(slot, gfn, writable); + if (kvm_is_error_hva(addr)) + return addr; + + if (!kvm_mem_attributes_may_read_gfn(kvm, gfn)) + return KVM_HVA_ERR_BAD; + + if (!kvm_mem_attributes_may_write_gfn(kvm, gfn)) + *writable =3D false; + + return addr; } =20 unsigned long kvm_vcpu_gfn_to_hva_prot(struct kvm_vcpu *vcpu, gfn_t gfn, b= ool *writable) { struct kvm_memory_slot *slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); + unsigned long addr; =20 - return gfn_to_hva_memslot_prot(slot, gfn, writable); + addr =3D gfn_to_hva_memslot_prot(slot, gfn, writable); + if (kvm_is_error_hva(addr)) + return addr; + + if (!kvm_mem_attributes_may_read_gfn(vcpu->kvm, gfn)) + return KVM_HVA_ERR_BAD; + + if (!kvm_mem_attributes_may_write_gfn(vcpu->kvm, gfn)) + *writable =3D false; + + return addr; } =20 static bool kvm_is_ad_tracked_page(struct page *page) --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 68A0E4446E2 for ; Thu, 16 Jul 2026 18:15:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225716; cv=none; b=NR+i1lUbBhrFgj1bYpx6+hSH4hDBmUUjZbjgRW5uQ0UHyCcHFWs8V5R0ljXl8VVy0bGjHTuurHCv5pAapxwzmCKjNKmbQ5Ucvo8BsOw2dIpgJeqPFW0N8jzDHl84kKBMcPnZ8v/PuSNt6IzgDB5ouqfrC2CozU5TK4bGVqmLZtg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225716; c=relaxed/simple; bh=2e+U27xduGBRxfBgJ+9hvS/HFbgVzEXO1NYF3Pm2hJE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=rX+lZYVMfq2VOLroINvcrm8bGXVjuymuJKbd4Ryk81NSx2d+iHo4YW/0vREcw3FigLrQkMpBnqIha0dO6zpU+afUz9h+2CGr0ELmo7Wtizs5Ec0jihO3huiWBcGc+qP3VFyLycLF6ahYh/1PzJy7iEf3pkKvSVyKU5z12cUr/Ls= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=GFp5M2TK; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="GFp5M2TK" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225708; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=TY1b9TTussYYzB0mCcw/1CQOTncA91YZzylI01X+vNc=; b=GFp5M2TKwJHGaTWl05mfjOUZcLWbE62nXWAN8GlarycxNTHnSG+NWMdz2Uj5pYO4NB/RnD 3mO0/1MLLh06xH3VXB35NX5TJoae5z3N2pVENF8m3KjQAAh9eLInzuBEnBOzkfKV7Ds6SM KufLYAxgnl7HqhLCg4WMqhihXsgSa6k= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-191-XacLad_MPfuT7VG3OegCKg-1; Thu, 16 Jul 2026 14:15:06 -0400 X-MC-Unique: XacLad_MPfuT7VG3OegCKg-1 X-Mimecast-MFC-AGG-ID: XacLad_MPfuT7VG3OegCKg_1784225705 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 593691955F7E; Thu, 16 Jul 2026 18:15:05 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id DDB1D1800480; Thu, 16 Jul 2026 18:15:04 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 11/24] KVM: Take memory protections into account in kvm_read/write_guest() Date: Thu, 16 Jul 2026 14:14:43 -0400 Message-ID: <20260716181456.402786-12-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Take into account memory attributes when accessing guest memory through the kvm_read/write_guest() family of functions. All of these pass a struct kvm_memory_slot pointer to the actual workhorse functions, in order to share code between the VM-wide and vCPU-specific version of the functions (the latter of which handles the multi-address-space case). For this reason they need specific changes and do not work even though the gfn_to_hva() path has been taught already about memory protection attributes. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- virt/kvm/kvm_main.c | 26 +++++++++++++++++++------- 1 file changed, 19 insertions(+), 7 deletions(-) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 56016aab0aad..d920ce5b6739 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -3249,8 +3249,8 @@ static int next_segment(unsigned long len, int offset) } =20 /* Copy @len bytes from guest memory at '(@gfn * PAGE_SIZE) + @offset' to = @data */ -static int __kvm_read_guest_page(struct kvm_memory_slot *slot, gfn_t gfn, - void *data, int offset, int len) +static int __kvm_read_guest_page(struct kvm *kvm, struct kvm_memory_slot *= slot, + gfn_t gfn, void *data, int offset, int len) { int r; unsigned long addr; @@ -3261,6 +3261,10 @@ static int __kvm_read_guest_page(struct kvm_memory_s= lot *slot, gfn_t gfn, addr =3D gfn_to_hva_memslot_prot(slot, gfn, NULL); if (kvm_is_error_hva(addr)) return -EFAULT; + + if (!kvm_mem_attributes_may_read_gfn(kvm, gfn)) + return -EFAULT; + r =3D __copy_from_user(data, (void __user *)addr + offset, len); if (r) return -EFAULT; @@ -3272,7 +3276,7 @@ int kvm_read_guest_page(struct kvm *kvm, gfn_t gfn, v= oid *data, int offset, { struct kvm_memory_slot *slot =3D gfn_to_memslot(kvm, gfn); =20 - return __kvm_read_guest_page(slot, gfn, data, offset, len); + return __kvm_read_guest_page(kvm, slot, gfn, data, offset, len); } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_read_guest_page); =20 @@ -3281,7 +3285,7 @@ int kvm_vcpu_read_guest_page(struct kvm_vcpu *vcpu, g= fn_t gfn, void *data, { struct kvm_memory_slot *slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); =20 - return __kvm_read_guest_page(slot, gfn, data, offset, len); + return __kvm_read_guest_page(vcpu->kvm, slot, gfn, data, offset, len); } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest_page); =20 @@ -3325,8 +3329,9 @@ int kvm_vcpu_read_guest(struct kvm_vcpu *vcpu, gpa_t = gpa, void *data, unsigned l } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest); =20 -static int __kvm_read_guest_atomic(struct kvm_memory_slot *slot, gfn_t gfn, - void *data, int offset, unsigned long len) +static int __kvm_read_guest_atomic(struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + void *data, int offset, unsigned long len) { int r; unsigned long addr; @@ -3334,6 +3339,9 @@ static int __kvm_read_guest_atomic(struct kvm_memory_= slot *slot, gfn_t gfn, if (WARN_ON_ONCE(offset + len > PAGE_SIZE)) return -EFAULT; =20 + if (!kvm_mem_attributes_may_read_gfn(kvm, gfn)) + return -EFAULT; + addr =3D gfn_to_hva_memslot_prot(slot, gfn, NULL); if (kvm_is_error_hva(addr)) return -EFAULT; @@ -3352,7 +3360,7 @@ int kvm_vcpu_read_guest_atomic(struct kvm_vcpu *vcpu,= gpa_t gpa, struct kvm_memory_slot *slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); int offset =3D offset_in_page(gpa); =20 - return __kvm_read_guest_atomic(slot, gfn, data, offset, len); + return __kvm_read_guest_atomic(vcpu->kvm, slot, gfn, data, offset, len); } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest_atomic); =20 @@ -3370,6 +3378,10 @@ static int __kvm_write_guest_page(struct kvm *kvm, addr =3D gfn_to_hva_memslot(memslot, gfn); if (kvm_is_error_hva(addr)) return -EFAULT; + + if (!kvm_mem_attributes_may_write_gfn(kvm, gfn)) + return -EFAULT; + r =3D __copy_to_user((void __user *)addr + offset, data, len); if (r) return -EFAULT; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 61FFB443E5F for ; Thu, 16 Jul 2026 18:15:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225717; cv=none; b=Okc3VQrSBC9ARDVYWgcQ5jbkRCNwizosmnsWV1hjyocP1wzT738FPfi9jAXNtBDGD+B4V54dPhtm+NysrC+fp1ayQ545DqLZrt0gioCG9xH7t/jI4u1diAvwwfLb1txoM48TCZtrw3kwWzr4GCS9QAyErdbZQBTWtzMGH4hw/H4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225717; c=relaxed/simple; bh=xDYSoUPrA5BnWA99MVHEfmO3r+n0ypDkoYPnNjV4p/0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=WercM+PmjvOEehMFrB+YBLKwtcwAtcDsqBZpidNpLL3wnAK3DKRrk7bkaa8qSmARLQJPFbjt3HN1cYGAnggzyNFIARCRL7sK58lk/gTqgKOOxN2PRvoqKsZQwKv0JV7OQ6/COXytITOOgxaf2ZZJsKn+siMdekH6MJOPYEwAdys= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Gd8IOwmn; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Gd8IOwmn" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225708; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=9Bf3DEj+Rm6QGcYhGWBRVx8rQIjklskExWGYebe9dcE=; b=Gd8IOwmnvE6ZhLltkDGrIhnOaNvYzFvXFG8Nz3cwi32AJMjMzhOqyuXeYLpBuq+OX76kGo Eu1e44otvi8IEiwCtjnpOkz7FAsCi5vO0cIwaGpz0hzW16JF8DRk9ub85Qz2HzRdMq6MTx d00/zHKLuSPR+P3YSltl7zgvEU4WN4U= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-370-YhdexhexMWi-kTv2A0JlMg-1; Thu, 16 Jul 2026 14:15:07 -0400 X-MC-Unique: YhdexhexMWi-kTv2A0JlMg-1 X-Mimecast-MFC-AGG-ID: YhdexhexMWi-kTv2A0JlMg_1784225706 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id EEFD81955F30; Thu, 16 Jul 2026 18:15:05 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 7F4641800480; Thu, 16 Jul 2026 18:15:05 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 12/24] KVM: Encapsulate memattrs array into anonymous struct Date: Thu, 16 Jul 2026 14:14:44 -0400 Message-ID: <20260716181456.402786-13-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne The metadata surrounding memory attributes is about to grow, so encapsulate the memory attributes array within an anonymous struct to provide namespacing. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- include/linux/kvm_host.h | 8 +++++--- virt/kvm/kvm_main.c | 10 +++++----- 2 files changed, 10 insertions(+), 8 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index f10ab70b99dc..341f2e97f3cb 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -874,8 +874,10 @@ struct kvm { struct notifier_block pm_notifier; #endif #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES - /* Protected by slots_lock (for writes) and RCU (for reads) */ - struct xarray mem_attr_array; + struct { + /* Protected by slots_lock (for writes) and RCU (for reads) */ + struct xarray array; + } mem_attrs; #endif char stats_id[KVM_STATS_NAME_SIZE]; }; @@ -2563,7 +2565,7 @@ static inline bool kvm_mem_attributes_may_exec(u64 at= trs) =20 static inline unsigned long kvm_get_memory_attributes(struct kvm *kvm, gfn= _t gfn) { - return xa_to_value(xa_load(&kvm->mem_attr_array, gfn)); + return xa_to_value(xa_load(&kvm->mem_attrs.array, gfn)); } =20 bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t e= nd, diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index d920ce5b6739..20113069562e 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -1116,7 +1116,7 @@ static struct kvm *kvm_create_vm(unsigned long type, = const char *fdname) rcuwait_init(&kvm->mn_memslots_update_rcuwait); xa_init(&kvm->vcpu_array); #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES - xa_init(&kvm->mem_attr_array); + xa_init(&kvm->mem_attrs.array); #endif =20 INIT_LIST_HEAD(&kvm->gpc_list); @@ -1301,7 +1301,7 @@ static void kvm_destroy_vm(struct kvm *kvm) srcu_barrier(&kvm->srcu); cleanup_srcu_struct(&kvm->srcu); #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES - xa_destroy(&kvm->mem_attr_array); + xa_destroy(&kvm->mem_attrs.array); #endif kvm_arch_free_vm(kvm); preempt_notifier_dec(); @@ -2439,7 +2439,7 @@ u64 kvm_supported_mem_attributes(struct kvm *kvm) bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t e= nd, unsigned long mask, unsigned long attrs) { - XA_STATE(xas, &kvm->mem_attr_array, start); + XA_STATE(xas, &kvm->mem_attrs.array, start); unsigned long index; void *entry; =20 @@ -2578,7 +2578,7 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm,= gfn_t start, gfn_t end, * partway through setting the new attributes. */ for (i =3D start; i < end; i++) { - r =3D xa_reserve(&kvm->mem_attr_array, i, GFP_KERNEL_ACCOUNT); + r =3D xa_reserve(&kvm->mem_attrs.array, i, GFP_KERNEL_ACCOUNT); if (r) goto out_unlock; =20 @@ -2588,7 +2588,7 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm,= gfn_t start, gfn_t end, kvm_handle_gfn_range(kvm, &pre_set_range); =20 for (i =3D start; i < end; i++) { - r =3D xa_err(xa_store(&kvm->mem_attr_array, i, entry, + r =3D xa_err(xa_store(&kvm->mem_attrs.array, i, entry, GFP_KERNEL_ACCOUNT)); KVM_BUG_ON(r, kvm); cond_resched(); --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7E8534446E7 for ; Thu, 16 Jul 2026 18:15:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225716; cv=none; b=QA/TH1wenJiwRCIjVKlM9Fck/Njh3N9+d8CE5HKeEfk48fv4VwgM/ZFs2auGEspwa+3htlRiFSNPS/b3V0KBqoUAWXGV1kMCUYk9BwI/xW6rzCUTCS4PeSqfP1fjxmlu/a4UQzqpLZMegKGIrO1LsS3BYmQJWoaiYDVLXQHm0yY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225716; c=relaxed/simple; bh=FmJzLb+n+bJ8qEVB8q4rI758/WTEOrbdo3jqTnz2IHg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=iBf4PxVIXifJQKCqi4/5N7UwtC1hDQ/FWMiHo6NHDU3cBlzCqWwOuiaSMmSxBz0UiOvS+heJ8bPcToJuwlVt7WGQGcH5gymcfmfErmmnvG7pOCO/VsiCGnANsZIqzIP2W1ipUUJEmKcA/V6Fil+17uYziq2d9elgmSvuaJAHbyg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=KPMX2bLC; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="KPMX2bLC" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225709; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=u0+M/PykJDVqAjN7ZHI2LjwPbsl16uAh/kMFqVLU5gQ=; b=KPMX2bLCKzhHG5WujiYpCPTfd5GgB5DamrU5zDZcetRtRLX1NPyS/XW4Zb3EC2sS0sS4x9 ZsMXxw+gT4i0sKmpmGxZhx/yvsDBazP4+XS0PjXp+Lf0rtzoxIBpXAPh3ujAYfI1LbgHsQ 3LNsgBUrVigIv6XiRkKMZKhA5C2z8zk= Received: from mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-507-HyLXb4zROBqgkmG2Z3V9kA-1; Thu, 16 Jul 2026 14:15:07 -0400 X-MC-Unique: HyLXb4zROBqgkmG2Z3V9kA-1 X-Mimecast-MFC-AGG-ID: HyLXb4zROBqgkmG2Z3V9kA_1784225706 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 90755180062C; Thu, 16 Jul 2026 18:15:06 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 20C401800480; Thu, 16 Jul 2026 18:15:06 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 13/24] KVM: Introduce kvm_check_gen()/kvm_memslots_check_gen() Date: Thu, 16 Jul 2026 14:14:45 -0400 Message-ID: <20260716181456.402786-14-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" In many cases, retrieving kvm_memslots is followed by a check on the generation of the slots. Introduce a helper function that either does the check alone, or compounds it with returning the struct kvm_memslots* to the caller. Signed-off-by: Paolo Bonzini --- arch/x86/kvm/x86.c | 10 ++-------- include/linux/kvm_host.h | 12 ++++++++++++ virt/kvm/kvm_main.c | 8 ++++---- virt/kvm/pfncache.c | 10 ++++------ 4 files changed, 22 insertions(+), 18 deletions(-) diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index e2240f817c22..58f544df47e1 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -2032,7 +2032,6 @@ static void record_steal_time(struct kvm_vcpu *vcpu) { struct gfn_to_hva_cache *ghc =3D &vcpu->arch.st.cache; struct kvm_steal_time __user *st; - struct kvm_memslots *slots; gpa_t gpa =3D vcpu->arch.st.msr_val & KVM_STEAL_VALID_BITS; u64 steal; u32 version; @@ -2048,9 +2047,7 @@ static void record_steal_time(struct kvm_vcpu *vcpu) if (WARN_ON_ONCE(current->mm !=3D vcpu->kvm->mm)) return; =20 - slots =3D kvm_memslots(vcpu->kvm); - - if (unlikely(slots->generation !=3D ghc->generation || + if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->generation) || gpa !=3D ghc->gpa || kvm_is_error_hva(ghc->hva) || !ghc->memslot)) { /* We rely on the fact that it fits in a single page. */ @@ -2601,7 +2598,6 @@ static void kvm_steal_time_set_preempted(struct kvm_v= cpu *vcpu) { struct gfn_to_hva_cache *ghc =3D &vcpu->arch.st.cache; struct kvm_steal_time __user *st; - struct kvm_memslots *slots; static const u8 preempted =3D KVM_VCPU_PREEMPTED; gpa_t gpa =3D vcpu->arch.st.msr_val & KVM_STEAL_VALID_BITS; =20 @@ -2628,9 +2624,7 @@ static void kvm_steal_time_set_preempted(struct kvm_v= cpu *vcpu) if (unlikely(current->mm !=3D vcpu->kvm->mm)) return; =20 - slots =3D kvm_memslots(vcpu->kvm); - - if (unlikely(slots->generation !=3D ghc->generation || + if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->generation) || gpa !=3D ghc->gpa || kvm_is_error_hva(ghc->hva) || !ghc->memslot)) return; diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 341f2e97f3cb..122af87d6f9a 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2559,6 +2559,18 @@ static inline bool kvm_mem_attributes_may_exec(u64 a= ttrs) return !(attrs & KVM_MEMORY_ATTRIBUTE_NX); } =20 +static inline bool kvm_memslots_check_gen(struct kvm *kvm, u64 slots_gener= ation, struct kvm_memslots **p_slots) +{ + struct kvm_memslots *slots =3D *p_slots =3D kvm_memslots(kvm); + return slots->generation =3D=3D slots_generation; +} + +static inline bool kvm_check_gen(struct kvm *kvm, u64 slots_generation) +{ + struct kvm_memslots *slots; + return kvm_memslots_check_gen(kvm, slots_generation, &slots); +} + #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES #define KVM_MEMORY_ATTRIBUTE_PROT \ (KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW | KVM_MEMORY_ATTRIBUTE= _NX) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 20113069562e..1afaec56ecb3 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -3502,14 +3502,14 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, = struct gfn_to_hva_cache *ghc, void *data, unsigned int offset, unsigned long len) { - struct kvm_memslots *slots =3D kvm_memslots(kvm); + struct kvm_memslots *slots; int r; gpa_t gpa =3D ghc->gpa + offset; =20 if (WARN_ON_ONCE(len + offset > ghc->len)) return -EINVAL; =20 - if (slots->generation !=3D ghc->generation) { + if (unlikely(!kvm_memslots_check_gen(kvm, ghc->generation, &slots))) { if (__kvm_gfn_to_hva_cache_init(slots, ghc, ghc->gpa, ghc->len)) return -EFAULT; } @@ -3540,14 +3540,14 @@ int kvm_read_guest_offset_cached(struct kvm *kvm, s= truct gfn_to_hva_cache *ghc, void *data, unsigned int offset, unsigned long len) { - struct kvm_memslots *slots =3D kvm_memslots(kvm); + struct kvm_memslots *slots; int r; gpa_t gpa =3D ghc->gpa + offset; =20 if (WARN_ON_ONCE(len + offset > ghc->len)) return -EINVAL; =20 - if (slots->generation !=3D ghc->generation) { + if (unlikely(!kvm_memslots_check_gen(kvm, ghc->generation, &slots))) { if (__kvm_gfn_to_hva_cache_init(slots, ghc, ghc->gpa, ghc->len)) return -EFAULT; } diff --git a/virt/kvm/pfncache.c b/virt/kvm/pfncache.c index 728d2c1b488a..e09703d249bb 100644 --- a/virt/kvm/pfncache.c +++ b/virt/kvm/pfncache.c @@ -72,8 +72,6 @@ static bool kvm_gpc_is_valid_len(gpa_t gpa, unsigned long= uhva, =20 bool kvm_gpc_check(struct gfn_to_pfn_cache *gpc, unsigned long len) { - struct kvm_memslots *slots =3D kvm_memslots(gpc->kvm); - if (!gpc->active) return false; =20 @@ -81,7 +79,7 @@ bool kvm_gpc_check(struct gfn_to_pfn_cache *gpc, unsigned= long len) * If the page was cached from a memslot, make sure the memslots have * not been re-configured. */ - if (!kvm_is_error_gpa(gpc->gpa) && gpc->generation !=3D slots->generation) + if (!kvm_is_error_gpa(gpc->gpa) && !kvm_check_gen(gpc->kvm, gpc->generati= on)) return false; =20 if (kvm_is_error_hva(gpc->uhva)) @@ -290,12 +288,12 @@ static int __kvm_gpc_refresh(struct gfn_to_pfn_cache = *gpc, gpa_t gpa, unsigned l if (gpc->uhva !=3D old_uhva) hva_change =3D true; } else { - struct kvm_memslots *slots =3D kvm_memslots(gpc->kvm); + struct kvm_memslots *slots; =20 page_offset =3D offset_in_page(gpa); =20 - if (gpc->gpa !=3D gpa || gpc->generation !=3D slots->generation || - kvm_is_error_hva(gpc->uhva)) { + if (!kvm_memslots_check_gen(gpc->kvm, gpc->generation, &slots) || + gpc->gpa !=3D gpa || kvm_is_error_hva(gpc->uhva)) { gfn_t gfn =3D gpa_to_gfn(gpa); =20 gpc->gpa =3D gpa; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B622437869 for ; Thu, 16 Jul 2026 18:15:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225719; cv=none; b=MOa3fIhOmcRVfriu2HwNqPa4mNe9Z8WBqTdEoAd9w7VWXWrjc1C9EXFQuNEZO/RzKqQrrXvbGYqUe5K4caAxKM87e92rlg0LjHpb3brI62/HTCkWvgcLaAFBIhmYYRu1rrjarL+mtVv6qKHN71IoOperR4Xc+YVXDfrd71szgkk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225719; c=relaxed/simple; bh=Kj/39F27i6nm/FzzqZdRX1SWP/oRHWxE/TjvNdNFZWA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Fj/gdOq0N/6P0srXaEEpy6mrJcrFnVNPxVNbbQ54NpxNQO7wk6Os8HKFCpMEeFOF3H0AiOr4ncdQxlwMKlyhRMAnbzGEnWO8orMjG8L4YBlXd/Faw5YxRs1x2+51a1ECfQGtZMrVGp7i5SH9DZ13ebFGCRaLhqShjurhVlofDGw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=QixbCBwB; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="QixbCBwB" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225711; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=NkFNzpBnvWvuQQmK9EC1837UEMgr2MUOQKRAKWrsHRc=; b=QixbCBwB2JLVWAcuPhE4nFC7EJuoACvTQiBSp/1pINlmuFoGcY04IPaomSgFbKPjwSMwQ0 cnInGtNCgCe7xEOwPdMRkXAJibyNap1G8plyuVO1zrw/0pzl4pL9cnobFEJGtzIfSchWiC bZHePQ/joDX9Cntn2HhMBbQYzbCiy8s= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-494-VxWoXOBUPO2GkHcm5aj0wA-1; Thu, 16 Jul 2026 14:15:08 -0400 X-MC-Unique: VxWoXOBUPO2GkHcm5aj0wA-1 X-Mimecast-MFC-AGG-ID: VxWoXOBUPO2GkHcm5aj0wA_1784225707 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 323AE1955EBC; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id B6D671800480; Thu, 16 Jul 2026 18:15:06 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 14/24] KVM: Introduce a generation number for memory attributes Date: Thu, 16 Jul 2026 14:14:46 -0400 Message-ID: <20260716181456.402786-15-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Introduce a generation number to track memory attribute modifications. This will allow KVM components to invalidate any assumptions they might have about guest's physical addresses and their permissions when memory attributes change. Like with memory slot updates, it's mandatory for components that access guest memory based on cached information to do so within a KVM SRCU read-side critical section, and that they validate the generation number before accessing memory. This, in combination with the synchronize_srcu() call within the memory attributes ioctl handler, ensures the following: - A memory attribute modification operation only returns after all users of outdated GPA data are done running. - Any component accessing cached data after the memory attribute modification returned will see the updated generation number. Additionally, loads/stores of the generation number have acquire/release semantics; which ensures all attribute writes are visible before updating the generation, and loads from attributes happen after having read the current generation number. Ultimately, since synchronize_srcu_expedited() is an expensive operation, only perform it when absolutely necessary. Do so if the introduced memory attribute is known to require synchronization or if the attribute being cleared contained a memory attribute that required synchronization. There shouldn't be any performance loss for memory attributes that don't require synchronization (and in general for any VMM that does not apply memory protections), because the attributes generation will always remain unchanged. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/x86.c | 4 +- include/linux/kvm_host.h | 33 ++++++++++++++-- include/linux/kvm_types.h | 6 ++- include/trace/events/kvm.h | 14 +++++-- virt/kvm/kvm_main.c | 80 +++++++++++++++++++++++++++++++++----- virt/kvm/pfncache.c | 13 ++++--- 6 files changed, 123 insertions(+), 27 deletions(-) diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 58f544df47e1..5229f40d074c 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -2047,7 +2047,7 @@ static void record_steal_time(struct kvm_vcpu *vcpu) if (WARN_ON_ONCE(current->mm !=3D vcpu->kvm->mm)) return; =20 - if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->generation) || + if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->slots_generation, ghc->attrs_= generation) || gpa !=3D ghc->gpa || kvm_is_error_hva(ghc->hva) || !ghc->memslot)) { /* We rely on the fact that it fits in a single page. */ @@ -2624,7 +2624,7 @@ static void kvm_steal_time_set_preempted(struct kvm_v= cpu *vcpu) if (unlikely(current->mm !=3D vcpu->kvm->mm)) return; =20 - if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->generation) || + if (unlikely(!kvm_check_gen(vcpu->kvm, ghc->slots_generation, ghc->attrs_= generation) || gpa !=3D ghc->gpa || kvm_is_error_hva(ghc->hva) || !ghc->memslot)) return; diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 122af87d6f9a..1fad3eb303c1 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -877,6 +877,7 @@ struct kvm { struct { /* Protected by slots_lock (for writes) and RCU (for reads) */ struct xarray array; + u64 generation; } mem_attrs; #endif char stats_id[KVM_STATS_NAME_SIZE]; @@ -2559,29 +2560,49 @@ static inline bool kvm_mem_attributes_may_exec(u64 = attrs) return !(attrs & KVM_MEMORY_ATTRIBUTE_NX); } =20 -static inline bool kvm_memslots_check_gen(struct kvm *kvm, u64 slots_gener= ation, struct kvm_memslots **p_slots) +static inline bool kvm_memslots_check_gen(struct kvm *kvm, u64 slots_gener= ation, + u64 attrs_generation, struct kvm_memslots **p_slots) { struct kvm_memslots *slots =3D *p_slots =3D kvm_memslots(kvm); - return slots->generation =3D=3D slots_generation; + return slots->generation =3D=3D slots_generation && kvm->mem_attrs.genera= tion =3D=3D attrs_generation; } =20 -static inline bool kvm_check_gen(struct kvm *kvm, u64 slots_generation) +static inline bool kvm_check_gen(struct kvm *kvm, u64 slots_generation, + u64 attrs_generation) { struct kvm_memslots *slots; - return kvm_memslots_check_gen(kvm, slots_generation, &slots); + return kvm_memslots_check_gen(kvm, slots_generation, attrs_generation, &s= lots); } =20 #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES #define KVM_MEMORY_ATTRIBUTE_PROT \ (KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW | KVM_MEMORY_ATTRIBUTE= _NX) =20 +#define KVM_MEMORY_ATTRIBUTE_NEEDS_SYNC_MASK KVM_MEMORY_ATTRIBUTE_PROT + static inline unsigned long kvm_get_memory_attributes(struct kvm *kvm, gfn= _t gfn) { return xa_to_value(xa_load(&kvm->mem_attrs.array, gfn)); } =20 +static inline u64 kvm_mem_attributes_generation(struct kvm *kvm) +{ + RCU_LOCKDEP_WARN(!lockdep_is_held(&kvm->slots_lock) && + !srcu_read_lock_held(&kvm->srcu), + "Suspicious memory attribute generation usage\n"); + + /* + * The acquire pairs with the release in kvm_vm_set_mem_attributes(). + * Memory attributes should only be queried _after_ storing the + * generation number. + */ + return smp_load_acquire(&kvm->mem_attrs.generation); +} + bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t e= nd, unsigned long mask, unsigned long attrs); +bool kvm_range_has_any_memory_attributes(struct kvm *kvm, gfn_t start, gfn= _t end, + unsigned long mask); bool kvm_arch_pre_set_memory_attributes(struct kvm *kvm, struct kvm_gfn_range *range); bool kvm_arch_post_set_memory_attributes(struct kvm *kvm, @@ -2611,6 +2632,10 @@ static inline unsigned long kvm_get_memory_attribute= s(struct kvm *kvm, gfn_t gfn { return 0; } +static inline u64 kvm_mem_attributes_generation(struct kvm *kvm) +{ + return 0; +} #endif /* CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES */ =20 static inline int kvm_mem_attributes_may_read_gfn(struct kvm *kvm, gfn_t g= fn) diff --git a/include/linux/kvm_types.h b/include/linux/kvm_types.h index a568d8e6f4e8..7d911220e00d 100644 --- a/include/linux/kvm_types.h +++ b/include/linux/kvm_types.h @@ -74,7 +74,8 @@ typedef u64 hfn_t; typedef hfn_t kvm_pfn_t; =20 struct gfn_to_hva_cache { - u64 generation; + u64 slots_generation; + u64 attrs_generation; gpa_t gpa; unsigned long hva; unsigned long len; @@ -82,7 +83,8 @@ struct gfn_to_hva_cache { }; =20 struct gfn_to_pfn_cache { - u64 generation; + u64 slots_generation; + u64 attrs_generation; gpa_t gpa; unsigned long uhva; struct kvm_memory_slot *memslot; diff --git a/include/trace/events/kvm.h b/include/trace/events/kvm.h index b282e3a86769..a620131e9010 100644 --- a/include/trace/events/kvm.h +++ b/include/trace/events/kvm.h @@ -365,23 +365,29 @@ TRACE_EVENT(kvm_dirty_ring_exit, * @attr: The value of the attribute being set. */ TRACE_EVENT(kvm_vm_set_mem_attributes, - TP_PROTO(gfn_t start, gfn_t end, unsigned long attr), - TP_ARGS(start, end, attr), + TP_PROTO(gfn_t start, gfn_t end, unsigned long attr, bool sync, u64 gener= ation), + TP_ARGS(start, end, attr, sync, generation), =20 TP_STRUCT__entry( __field(gfn_t, start) __field(gfn_t, end) __field(unsigned long, attr) + __field(bool, sync) + __field(u64, generation) ), =20 TP_fast_assign( __entry->start =3D start; __entry->end =3D end; __entry->attr =3D attr; + __entry->sync =3D sync; + __entry->generation =3D generation; ), =20 - TP_printk("%#016llx -- %#016llx [0x%lx]", - __entry->start, __entry->end, __entry->attr) + TP_printk("%#016llx -- %#016llx [0x%lx], sync %d gen %llu", + __entry->start, __entry->end, __entry->attr, + __entry->sync, __entry->generation) + ); #endif /* CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES */ =20 diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 1afaec56ecb3..2fe4087319ad 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -1117,6 +1117,7 @@ static struct kvm *kvm_create_vm(unsigned long type, = const char *fdname) xa_init(&kvm->vcpu_array); #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES xa_init(&kvm->mem_attrs.array); + kvm->mem_attrs.generation =3D 0; #endif =20 INIT_LIST_HEAD(&kvm->gpc_list); @@ -2467,6 +2468,51 @@ bool kvm_range_has_memory_attributes(struct kvm *kvm= , gfn_t start, gfn_t end, return true; } =20 +/* + * Returns true if _any_ gfns in the range [@start, @end) have attributes = that + * match _any_ bit in @mask. + */ +bool kvm_range_has_any_memory_attributes(struct kvm *kvm, gfn_t start, gfn= _t end, + unsigned long mask) +{ + XA_STATE(xas, &kvm->mem_attrs.array, start); + void *entry; + + mask &=3D kvm_supported_mem_attributes(kvm); + if (!mask) + return false; + + if (end =3D=3D start + 1) + return !!(kvm_get_memory_attributes(kvm, start) & mask); + + guard(rcu)(); + for (;;) { + do { + entry =3D xas_next(&xas); + } while (xas_retry(&xas, entry)); + + if (xas.xa_index >=3D end) + break; + + if (xa_to_value(entry) & mask) + return true; + } + + return false; +} + +static bool kvm_range_memory_attributes_need_sync(struct kvm *kvm, + gfn_t start, gfn_t end, + unsigned long attributes) +{ + u64 mask =3D KVM_MEMORY_ATTRIBUTE_NEEDS_SYNC_MASK; + + if (attributes & mask) + return true; + + return kvm_range_has_any_memory_attributes(kvm, start, end, mask); +} + static __always_inline void kvm_handle_gfn_range(struct kvm *kvm, struct kvm_mmu_notifier_range *range) { @@ -2559,20 +2605,25 @@ static int kvm_vm_set_mem_attributes(struct kvm *kv= m, gfn_t start, gfn_t end, .on_lock =3D kvm_mmu_invalidate_end, .may_block =3D true, }; + bool sync =3D false; unsigned long i; void *entry; int r =3D 0; =20 entry =3D attributes ? xa_mk_value(attributes) : NULL; =20 - trace_kvm_vm_set_mem_attributes(start, end, attributes); - mutex_lock(&kvm->slots_lock); =20 /* Nothing to do if the entire range has the desired attributes. */ if (kvm_range_has_memory_attributes(kvm, start, end, ~0, attributes)) goto out_unlock; =20 + sync =3D kvm_range_memory_attributes_need_sync(kvm, start, end, + attributes); + + trace_kvm_vm_set_mem_attributes(start, end, attributes, sync, + kvm->mem_attrs.generation + 1); + /* * Reserve memory ahead of time to avoid having to deal with failures * partway through setting the new attributes. @@ -2594,10 +2645,16 @@ static int kvm_vm_set_mem_attributes(struct kvm *kv= m, gfn_t start, gfn_t end, cond_resched(); } =20 + /* Pairs with acquire in kvm_mem_attributes_generation() */ + smp_store_release(&kvm->mem_attrs.generation, + kvm->mem_attrs.generation + 1); + kvm_handle_gfn_range(kvm, &post_set_range); =20 out_unlock: mutex_unlock(&kvm->slots_lock); + if (sync) + synchronize_srcu_expedited(&kvm->srcu); =20 return r; } @@ -3449,7 +3506,8 @@ int kvm_vcpu_write_guest(struct kvm_vcpu *vcpu, gpa_t= gpa, const void *data, } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_write_guest); =20 -static int __kvm_gfn_to_hva_cache_init(struct kvm_memslots *slots, +static int __kvm_gfn_to_hva_cache_init(struct kvm *kvm, + struct kvm_memslots *slots, struct gfn_to_hva_cache *ghc, gpa_t gpa, unsigned long len) { @@ -3459,8 +3517,8 @@ static int __kvm_gfn_to_hva_cache_init(struct kvm_mem= slots *slots, gfn_t nr_pages_needed =3D end_gfn - start_gfn + 1; gfn_t nr_pages_avail; =20 - /* Update ghc->generation before performing any error checks. */ - ghc->generation =3D slots->generation; + /* Update ghc->slots_generation before performing any error checks. */ + ghc->slots_generation =3D slots->generation; =20 if (start_gfn > end_gfn) { ghc->hva =3D KVM_HVA_ERR_BAD; @@ -3479,6 +3537,8 @@ static int __kvm_gfn_to_hva_cache_init(struct kvm_mem= slots *slots, return -EFAULT; } =20 + ghc->attrs_generation =3D kvm_mem_attributes_generation(kvm); + /* Use the slow path for cross page reads and writes. */ if (nr_pages_needed =3D=3D 1) ghc->hva +=3D offset; @@ -3494,7 +3554,7 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct= gfn_to_hva_cache *ghc, gpa_t gpa, unsigned long len) { struct kvm_memslots *slots =3D kvm_memslots(kvm); - return __kvm_gfn_to_hva_cache_init(slots, ghc, gpa, len); + return __kvm_gfn_to_hva_cache_init(kvm, slots, ghc, gpa, len); } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gfn_to_hva_cache_init); =20 @@ -3509,8 +3569,8 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, st= ruct gfn_to_hva_cache *ghc, if (WARN_ON_ONCE(len + offset > ghc->len)) return -EINVAL; =20 - if (unlikely(!kvm_memslots_check_gen(kvm, ghc->generation, &slots))) { - if (__kvm_gfn_to_hva_cache_init(slots, ghc, ghc->gpa, ghc->len)) + if (unlikely(!kvm_memslots_check_gen(kvm, ghc->slots_generation, ghc->att= rs_generation, &slots))) { + if (__kvm_gfn_to_hva_cache_init(kvm, slots, ghc, ghc->gpa, ghc->len)) return -EFAULT; } =20 @@ -3547,8 +3607,8 @@ int kvm_read_guest_offset_cached(struct kvm *kvm, str= uct gfn_to_hva_cache *ghc, if (WARN_ON_ONCE(len + offset > ghc->len)) return -EINVAL; =20 - if (unlikely(!kvm_memslots_check_gen(kvm, ghc->generation, &slots))) { - if (__kvm_gfn_to_hva_cache_init(slots, ghc, ghc->gpa, ghc->len)) + if (unlikely(!kvm_memslots_check_gen(kvm, ghc->slots_generation, ghc->att= rs_generation, &slots))) { + if (__kvm_gfn_to_hva_cache_init(kvm, slots, ghc, ghc->gpa, ghc->len)) return -EFAULT; } =20 diff --git a/virt/kvm/pfncache.c b/virt/kvm/pfncache.c index e09703d249bb..46ffae69fe77 100644 --- a/virt/kvm/pfncache.c +++ b/virt/kvm/pfncache.c @@ -76,10 +76,11 @@ bool kvm_gpc_check(struct gfn_to_pfn_cache *gpc, unsign= ed long len) return false; =20 /* - * If the page was cached from a memslot, make sure the memslots have - * not been re-configured. + * If the page was cached from a memslot, make sure the memslots nor + * memory attributes have not been re-configured. */ - if (!kvm_is_error_gpa(gpc->gpa) && !kvm_check_gen(gpc->kvm, gpc->generati= on)) + if (!kvm_is_error_gpa(gpc->gpa) && + !kvm_check_gen(gpc->kvm, gpc->slots_generation, gpc->attrs_generation= )) return false; =20 if (kvm_is_error_hva(gpc->uhva)) @@ -253,6 +254,7 @@ static kvm_pfn_t hva_to_pfn_retry(struct gfn_to_pfn_cac= he *gpc) =20 static int __kvm_gpc_refresh(struct gfn_to_pfn_cache *gpc, gpa_t gpa, unsi= gned long uhva) { + struct kvm *kvm =3D gpc->kvm; unsigned long page_offset; bool unmap_old =3D false; unsigned long old_uhva; @@ -292,12 +294,13 @@ static int __kvm_gpc_refresh(struct gfn_to_pfn_cache = *gpc, gpa_t gpa, unsigned l =20 page_offset =3D offset_in_page(gpa); =20 - if (!kvm_memslots_check_gen(gpc->kvm, gpc->generation, &slots) || + if (!kvm_memslots_check_gen(gpc->kvm, gpc->slots_generation, gpc->attrs_= generation, &slots) || gpc->gpa !=3D gpa || kvm_is_error_hva(gpc->uhva)) { gfn_t gfn =3D gpa_to_gfn(gpa); =20 + gpc->attrs_generation =3D kvm_mem_attributes_generation(kvm); gpc->gpa =3D gpa; - gpc->generation =3D slots->generation; + gpc->slots_generation =3D slots->generation; gpc->memslot =3D __gfn_to_memslot(slots, gfn); gpc->uhva =3D gfn_to_hva_memslot(gpc->memslot, gfn); =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7C6A1444700 for ; Thu, 16 Jul 2026 18:15:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225717; cv=none; b=rfGbjxB3E8WZi2mFLcavLWbysLqrd6qhUzluFwV4PjXzD+0W/hK8HnXPfkXB0sCp0FGQMTUWojJrVnOx1drC7FyggVeg6kxEhRsjvsVlbqB2mmLGdZpuKgfwJ8iYMYFLdl3vgNSx7p5kWVG6VMko4VF6xaVxr3BKukRO0BU23uk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225717; c=relaxed/simple; bh=4gx87GDfWBQArS2ZGfXt/bxpjrLyyn9ZT0k59nikohw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=VL/mzjqw/ee7VQYk/moNXQ0eV921XWjHgHdNeqRR+LVIdZCKqafQZca5pd4gDpZ7/+K8dBaBzYhjSvNhhKANayzW6yn4EkuDvz/TjsmjLF9HBS1vpnaPuXq1RGqUlQz3QAIGnD1/bLeBV17KbcYXQ8uI+G2k+o2ZdhgwRmQV0OQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Fsu9KTtF; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Fsu9KTtF" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225710; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=5YNOczln1vCwhatoD+84vy5EZHAV+CFehLYCLKBd8B4=; b=Fsu9KTtFUaB83HcOQ4aJAUOW7SWuXQJA0PAVix7SuBEehoplMOV+7VWGiqJlXYhhMfwoN8 TDaNPEpeq0Bfx44D9r0qAwWYwJDwOoCo4pMjJF9M2Ds/xOdDOAUP7c8jfiTwKYq3A+4XF7 6o+QpERsm/Qu9RzKvPuyI+ukJrRdjQs= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-209-uNaYeMmGMp6fg3Xgn7jveQ-1; Thu, 16 Jul 2026 14:15:08 -0400 X-MC-Unique: uNaYeMmGMp6fg3Xgn7jveQ-1 X-Mimecast-MFC-AGG-ID: uNaYeMmGMp6fg3Xgn7jveQ_1784225707 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id C85B41955F08; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 58A341800480; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 15/24] KVM: Take memory protections into account for accesses with cached gfn->hva Date: Thu, 16 Jul 2026 14:14:47 -0400 Message-ID: <20260716181456.402786-16-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Account for memory attributes when accessing guest memory through kvm_get/put_guest(). This requires tracking the memory attributes generation as part of gfn_to_hva_cache's data, invalidate the cached information if the generation changes, and failing to refresh the cache if restrictive memory attributes are found within the GPA range. Similar to how gfn_to_hva_cache disallows caching gfns mapped within read-only memory slots, gfns marked as read-only by memory attributes will also fail to initialize. Unsurprisingly, the same behaviour applies to gfns mapped as non-accessible (NR/NW), while non-executable mappings are okay. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- include/linux/kvm_host.h | 19 +++++++++++++++++-- virt/kvm/kvm_main.c | 15 +++++++++++---- 2 files changed, 28 insertions(+), 6 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 1fad3eb303c1..71ab2cbecbd1 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -1351,7 +1351,8 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct= gfn_to_hva_cache *ghc, typeof(v) __user *__uaddr =3D (typeof(__uaddr))(__addr + offset); \ int __ret =3D -EFAULT; \ \ - if (!kvm_is_error_hva(__addr)) \ + if (!kvm_is_error_hva(__addr) && \ + kvm_mem_attributes_may_read_gfn(kvm, gfn)) \ __ret =3D get_user(v, __uaddr); \ __ret; \ }) @@ -1371,7 +1372,8 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct= gfn_to_hva_cache *ghc, typeof(v) __user *__uaddr =3D (typeof(__uaddr))(__addr + offset); \ int __ret =3D -EFAULT; \ \ - if (!kvm_is_error_hva(__addr)) \ + if (!kvm_is_error_hva(__addr) && \ + kvm_mem_attributes_may_write_gfn(kvm, gfn)) \ __ret =3D put_user(v, __uaddr); \ if (!__ret) \ mark_page_dirty(kvm, gfn); \ @@ -2632,6 +2634,12 @@ static inline unsigned long kvm_get_memory_attribute= s(struct kvm *kvm, gfn_t gfn { return 0; } +static inline bool kvm_range_has_any_memory_attributes(struct kvm *kvm, + gfn_t start, gfn_t end, + unsigned long mask) +{ + return false; +} static inline u64 kvm_mem_attributes_generation(struct kvm *kvm) { return 0; @@ -2652,6 +2660,13 @@ static inline int kvm_mem_attributes_may_write_gfn(s= truct kvm *kvm, gfn_t gfn) return kvm_mem_attributes_may_write(attrs); } =20 +static inline bool kvm_range_has_rw_memory_protections(struct kvm *kvm, + gfn_t start, gfn_t end) +{ + return kvm_range_has_any_memory_attributes(kvm, start, end, + KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW); +} + #ifdef CONFIG_KVM_GUEST_MEMFD int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 2fe4087319ad..a6fd17851c2f 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -3529,15 +3529,22 @@ static int __kvm_gfn_to_hva_cache_init(struct kvm *= kvm, * If the requested region crosses two memslots, we still * verify that the entire region is valid here. */ - for ( ; start_gfn <=3D end_gfn; start_gfn +=3D nr_pages_avail) { - ghc->memslot =3D __gfn_to_memslot(slots, start_gfn); - ghc->hva =3D gfn_to_hva_many(ghc->memslot, start_gfn, - &nr_pages_avail); + for (gfn_t gfn =3D start_gfn ; gfn <=3D end_gfn; gfn +=3D nr_pages_avail)= { + ghc->memslot =3D __gfn_to_memslot(slots, gfn); + ghc->hva =3D gfn_to_hva_many(ghc->memslot, gfn, &nr_pages_avail); if (kvm_is_error_hva(ghc->hva)) return -EFAULT; } =20 + /* + * RW memory attributes are incompatible with GHC. The RW protection + * check has to happen after storing the generation number. + */ ghc->attrs_generation =3D kvm_mem_attributes_generation(kvm); + if (kvm_range_has_rw_memory_protections(kvm, start_gfn, end_gfn + 1)) { + ghc->hva =3D KVM_HVA_ERR_BAD; + return -EFAULT; + } =20 /* Use the slow path for cross page reads and writes. */ if (nr_pages_needed =3D=3D 1) --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 29CEE446050 for ; Thu, 16 Jul 2026 18:15:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; cv=none; b=mk8OhECCURygvbsb3/uzfioGZ1EL2Seb4yVZIHuyj/wx5Wf07ZfIllYfdQF9kLnpbKSym9ktOHJ2VU6w0IeToHD8HhwEK6rVU4IYNcvoQeu/e4djRg3cuBoP3DCLTAirjKER6sV5s5jw1zlmTaZDH1eDbo37Uv35wopwy2Qrx+k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; c=relaxed/simple; bh=w7MCxn6ioQ4dt8pPbZ0HMRdHRJUHlmZ0rQY/csJg0AA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=F/3X0ApqrwP6O7Gl2z4NHyohjTnOsmhExqH0p0dw1v1D1pgnH08s4d5F3VAyvI+mSnsYHRlG0LAI+heTEd1E3lk661BkwqN5zn/8/CFJxOKhX/zIsUyl2GKhxVLzsA/kTm2zkKyd7/iCYx8cw5YkrUqjlGNacbvtRRKGIUA04dk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=bwBDbCau; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="bwBDbCau" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225714; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=fgIXjluZig1qTYu6rzzcNeV9+C4wTNcZU/cG7DYvGAA=; b=bwBDbCaurmZNVBKHKE7E/yhdzRmbMavx7nT8PL34AciwI7+PqeF0LRTTtgwzUNaOSTEkMe rVfKB/UArZR5gJkkF5UTAwvpRSL8ZqJSojdPttN1gTofXKMFOmb00aoBgrWvXE8e1jpAl0 bAMTQfqbmuvbUkzffCEAUxVBARvO5iw= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-656-uEvr9yjKM7mg2uMsjz18zQ-1; Thu, 16 Jul 2026 14:15:09 -0400 X-MC-Unique: uEvr9yjKM7mg2uMsjz18zQ-1 X-Mimecast-MFC-AGG-ID: uEvr9yjKM7mg2uMsjz18zQ_1784225708 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 7F4B31955E87; Thu, 16 Jul 2026 18:15:08 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id EE8B51800480; Thu, 16 Jul 2026 18:15:07 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 16/24] KVM: pfncache: Fail to refresh if it contains memory protections Date: Thu, 16 Jul 2026 14:14:48 -0400 Message-ID: <20260716181456.402786-17-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Track the memory attributes generation as part of pfncache's data, invalidate the cached information if the generation changes, and fail to refresh the cache if restrictive memory attributes are found within the GPA range. Similar to how pfncache disallows caching gfns mapped within read-only memory slots, gfns marked as read-only by memory attributes will also fail to initialize. Unsurprisingly, the same behaviour applies to gfns mapped as non-accessible (NR/NW), while non-executable mappings are fine. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- virt/kvm/pfncache.c | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/virt/kvm/pfncache.c b/virt/kvm/pfncache.c index 46ffae69fe77..39935740136d 100644 --- a/virt/kvm/pfncache.c +++ b/virt/kvm/pfncache.c @@ -298,7 +298,18 @@ static int __kvm_gpc_refresh(struct gfn_to_pfn_cache *= gpc, gpa_t gpa, unsigned l gpc->gpa !=3D gpa || kvm_is_error_hva(gpc->uhva)) { gfn_t gfn =3D gpa_to_gfn(gpa); =20 + /* + * RW memory attributes are incompatible with GPC. The + * RW protection check has to happen after storing the + * generation number. + */ gpc->attrs_generation =3D kvm_mem_attributes_generation(kvm); + if (kvm_range_has_rw_memory_protections(kvm, gfn, gfn + 1)) { + gpc->uhva =3D KVM_HVA_ERR_BAD; + ret =3D -EFAULT; + goto out; + } + gpc->gpa =3D gpa; gpc->slots_generation =3D slots->generation; gpc->memslot =3D __gfn_to_memslot(slots, gfn); --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 16C93443E2E for ; Thu, 16 Jul 2026 18:15:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225719; cv=none; b=fYwSu9fgZUPjSIPUZs5Cc4hobiXpfCN++vRhb81HnUmwCGqWD/921tlvxUFpP39rMh+1+M3SSLoGTodr20Jfh5H5axmlIrIc6IUr8pdyt0Ni9WDqzi76a0xubQdYbhXQOr2nSIo/JPN5HqD++ZOeC6Rd127EcuS4yYpKRuGnIAQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225719; c=relaxed/simple; bh=0pNoM3dY4hAyzeTYjBVMfHxYcmOfZy5uy6XvHHy/87g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=MNru+K/G7eRha4sSdWiq84L6Sc/w8GYboIot2TD//JoII2ZjicpAVPoE0BvWtTj5BGaAh/EmWhXdVGozNpmJyH9lpDSK4HonO/n23+ztmy2171y83T2CZdrp0VDlVG+mcl3bMFYYYo2GOl8DQiyZwX+/IKcTGDx74fyzRSLGMmw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=iI1rS2dX; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="iI1rS2dX" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225711; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=mlpzWgYUP8YYr8Qsq+Jl34UUqNPGRwsPg+yx3ZDZOiA=; b=iI1rS2dXQGAfkLZsg1bgXQmUQ3OiVjf+ap4P68RDr74Ki8ID2hEGFfk2XkW84HVC7ng7Jh ogIhJgeuFm89R0N95gkg5f9bWAUBQu5Nu11Z0du3aLonxL0dACvbdiJGqUHlO7HltJ2EeY cqwfkmbl593cmuQYqf/i55afFbQI1i4= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-499-3O3E3_2NPdCDhcEOzaO0GA-1; Thu, 16 Jul 2026 14:15:09 -0400 X-MC-Unique: 3O3E3_2NPdCDhcEOzaO0GA-1 X-Mimecast-MFC-AGG-ID: 3O3E3_2NPdCDhcEOzaO0GA_1784225709 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 0C4811955EAC; Thu, 16 Jul 2026 18:15:09 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 9026D1800480; Thu, 16 Jul 2026 18:15:08 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 17/24] KVM: x86/mmu: Take memory protection attributes into account during faults Date: Thu, 16 Jul 2026 14:14:49 -0400 Message-ID: <20260716181456.402786-18-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Take memory protection attributes when faulting guest memory. Prohibited memory accesses will cause a user-space -EFAULT exit just like private memory accesses. Userspace will either enable the access or bump it to the guest as some kind of exception (e.g. a VTL return). Since the struct kvm_page_fault already has the access type in PFERR_* format, the check is done via the kvm_page_format permissions table. This means that it supports naturally all page table format variants, and it can even handle mode-based memory protection when the host uses MBEC/GMET. The only thing that needs some care is to build the restricted ACC_* mask with the root page's own access mask as a base (and not ACC_ALL). Otherwise, supervisor mode execution would be handled incorrectly on AMD processors with GMET. To avoid spamming the trace buffer too much, the new trace event only kicks in if memory protection attributes are present for the faulted gfn. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 49 +++++++++++++++++++++++++++++++++ arch/x86/kvm/mmu/mmu_internal.h | 19 +++++++++++++ arch/x86/kvm/mmu/mmutrace.h | 29 +++++++++++++++++++ arch/x86/kvm/mmu/paging_tmpl.h | 2 +- arch/x86/kvm/mmu/spte.h | 11 ++------ 5 files changed, 100 insertions(+), 10 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index b6463b0b0b6d..59c2a04f5648 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -4611,6 +4611,50 @@ static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu = *vcpu, return RET_PF_CONTINUE; } =20 +static inline unsigned kvm_get_gfn_protections(struct kvm_vcpu *vcpu, gfn_= t gfn) +{ + struct kvm *kvm =3D vcpu->kvm; + unsigned int access =3D vcpu->arch.mmu->root_role.access; + unsigned long attrs =3D kvm_get_memory_attributes(kvm, gfn); + if (!attrs) + return access; + + WARN_ON_ONCE(!kvm_mem_attributes_valid(kvm, attrs)); + + if (!kvm_mem_attributes_may_read(attrs)) + access &=3D ~ACC_READ_MASK; + if (!kvm_mem_attributes_may_write(attrs)) + access &=3D ~ACC_WRITE_MASK; + if (!kvm_mem_attributes_may_exec(attrs)) { + access &=3D ~ACC_EXEC_MASK; + if (shadow_xu_mask) + access &=3D ~ACC_USER_EXEC_MASK; + } + + return access; +} + +static int kvm_faultin_memory_protections(struct kvm_vcpu *vcpu, + struct kvm_page_fault *fault) +{ + unsigned access; + + /* Memory attributes don't apply to MMIO regions */ + if (unlikely(!fault->slot)) + return RET_PF_CONTINUE; + + access =3D kvm_get_gfn_protections(vcpu, fault->gfn); + if (access =3D=3D ACC_ALL) + return RET_PF_CONTINUE; + + trace_kvm_faultin_memory_protections(vcpu, fault, access); + if (__permission_fault(vcpu->arch.mmu, access, fault)) + return -EFAULT; + + fault->host_access &=3D access; + return RET_PF_CONTINUE; +} + static int __kvm_mmu_faultin_pfn(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) { @@ -4691,6 +4735,11 @@ static int kvm_mmu_faultin_pfn(struct kvm_vcpu *vcpu, if (unlikely(!slot)) return kvm_handle_noslot_fault(vcpu, fault, access); =20 + if (kvm_faultin_memory_protections(vcpu, fault)) { + kvm_mmu_prepare_memory_fault_exit(vcpu, fault); + return -EFAULT; + } + /* * Retry the page fault if the gfn hit a memslot that is being deleted * or moved. This ensures any existing SPTEs for the old memslot will diff --git a/arch/x86/kvm/mmu/mmu_internal.h b/arch/x86/kvm/mmu/mmu_interna= l.h index 00215b9f309f..0bad60c5aab6 100644 --- a/arch/x86/kvm/mmu/mmu_internal.h +++ b/arch/x86/kvm/mmu/mmu_internal.h @@ -290,6 +290,25 @@ struct kvm_page_fault { bool write_fault_to_shadow_pgtable; }; =20 +/* + * Returns true if the access indicated by @fault is forbidden by the exis= ting + * SPTE protections. + */ +static inline bool __permission_fault(struct kvm_mmu *mmu, unsigned access, + struct kvm_page_fault *fault) +{ + unsigned pfec; + + /* + * RSVD is handled elsewhere, and is used for SMAP in the context + * of accessing fmt.permissions[]. SPTEs never use PK or SS, as + * they are not supported for shadow paging and irrelevant for TDP. + */ + pfec =3D fault->error_code & ( + PFERR_WRITE_MASK | PFERR_USER_MASK | PFERR_FETCH_MASK); + return (mmu->fmt.permissions[pfec >> 1] >> access) & 1; +} + /* * Return values of handle_mmio_page_fault(), mmu.page_fault(), fast_page_= fault(), * and of course kvm_mmu_do_page_fault(). diff --git a/arch/x86/kvm/mmu/mmutrace.h b/arch/x86/kvm/mmu/mmutrace.h index 8354d9f39777..6bbd66a827b6 100644 --- a/arch/x86/kvm/mmu/mmutrace.h +++ b/arch/x86/kvm/mmu/mmutrace.h @@ -447,6 +447,35 @@ TRACE_EVENT( __entry->gfn, __entry->spte, __entry->level, __entry->errno) ); =20 +TRACE_EVENT(kvm_faultin_memory_protections, + TP_PROTO(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault, + unsigned access), + TP_ARGS(vcpu, fault, access), + + TP_STRUCT__entry( + __field(unsigned int, vcpu_id) + __field(unsigned long, guest_rip) + __field(u64, fault_address) + __field(bool, write) + __field(bool, exec) + __field(unsigned, access) + ), + + TP_fast_assign( + __entry->vcpu_id =3D vcpu->vcpu_id; + __entry->guest_rip =3D kvm_rip_read(vcpu); + __entry->fault_address =3D fault->gfn; + __entry->write =3D fault->write; + __entry->exec =3D fault->exec; + __entry->access =3D access; + ), + + TP_printk("vcpu %d rip 0x%lx gfn 0x%016llx access %s protections 0x%x", + __entry->vcpu_id, __entry->guest_rip, __entry->fault_address, + __entry->exec ? "X" : (__entry->write ? "W" : "R"), + __entry->access) +); + #endif /* _TRACE_KVMMMU_H */ =20 #undef TRACE_INCLUDE_PATH diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index 1871d334fed7..cdf05cd76d63 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -990,7 +990,7 @@ static int FNAME(sync_spte)(struct kvm_vcpu *vcpu, stru= ct kvm_mmu_page *sp, int =20 sptep =3D &sp->spt[i]; spte =3D *sptep; - host_access =3D ACC_ALL; + host_access =3D kvm_get_gfn_protections(vcpu, gfn); if (!(spte & shadow_host_writable_mask)) host_access &=3D ~ACC_WRITE_MASK; slot =3D kvm_vcpu_gfn_to_memslot(vcpu, gfn); diff --git a/arch/x86/kvm/mmu/spte.h b/arch/x86/kvm/mmu/spte.h index 589f3954633e..14f322944671 100644 --- a/arch/x86/kvm/mmu/spte.h +++ b/arch/x86/kvm/mmu/spte.h @@ -491,7 +491,7 @@ static inline bool is_mmu_writable_spte(u64 spte) static inline bool spte_permission_fault(struct kvm_mmu *mmu, u64 spte, struct kvm_page_fault *fault) { - unsigned pfec, pte_access; + unsigned pte_access; =20 if (!is_shadow_present_pte(spte)) return true; @@ -511,14 +511,7 @@ static inline bool spte_permission_fault(struct kvm_mm= u *mmu, u64 spte, pte_access |=3D spte & shadow_xu_mask ? ACC_USER_EXEC_MASK : 0; } =20 - /* - * RSVD is handled elsewhere, and is used for SMAP in the context - * of accessing fmt.permissions[]. SPTEs never use PK or SS, as - * they are not supported for shadow paging and irrelevant for TDP. - */ - pfec =3D fault->error_code & ( - PFERR_WRITE_MASK | PFERR_USER_MASK | PFERR_FETCH_MASK); - return (mmu->fmt.permissions[pfec >> 1] >> pte_access) & 1; + return __permission_fault(mmu, pte_access, fault); } =20 /* --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2D27444605B for ; Thu, 16 Jul 2026 18:15:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225726; cv=none; b=aO3H8ic3Ply8OQ9IedDXKlgr6pNhLPHXEoQlKSU6SyqpGfcG+seRFfdzZmMDdoUNd7ZLSHn1b0QXyfGpszr9+Y3dWDBlda2RuNnZR0OvKOZw4QK4P4NBE7iMVYaaL2nMHRFxNtXHU/VycFR3yj1G/g7LZRPRCDIevhN2MPNkS2M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225726; c=relaxed/simple; bh=7LsMRzOtY0JTrNv8gK7CEE2+mee4K8ZPe2OjPuTcAWk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=RpAk+s6IUD+3zM1V4uJb89rQmxLiMIl6eyMQpUNzYnaY7WnzFuCC1dD6gbd6ApnPevmEBtgQO1ELk4l6lsH9GFCSzzX+pSg5R7dLrMcHVZga9GLyNtLUSH47WXAsi7QDEDK4bIdt1YEO9tDTb+DM/MXOVW5Xw/8n9a6y9ObdVeo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Z5gstPTb; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Z5gstPTb" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225714; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=37r7DMJnDYdCZNxuuPu+eWT3OhHL3Y8s3njCFCoQsG0=; b=Z5gstPTb74vbhQwpeKeEw3um/I2dQzct/HVcRQ6U50+MnxA6mFflAN2lkdHO61cD6O2Jbf 1VQV0SUXkeAux6IRORxBBHa3PuZtsTf3t4SHt17IoYU0oRt8dwgQDS/5H71Vi6Xgkj42rE HdiaNz5fXNoQBye+VkfN1aAaIu+kMLw= Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-350--FllGWkTP8S1PZZoi-pEww-1; Thu, 16 Jul 2026 14:15:10 -0400 X-MC-Unique: -FllGWkTP8S1PZZoi-pEww-1 X-Mimecast-MFC-AGG-ID: -FllGWkTP8S1PZZoi-pEww_1784225709 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id A1C151956044; Thu, 16 Jul 2026 18:15:09 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 31CE71800480; Thu, 16 Jul 2026 18:15:09 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 18/24] KVM: x86/mmu: Issue memory fault exit if walk failed due to memory attribute Date: Thu, 16 Jul 2026 14:14:50 -0400 Message-ID: <20260716181456.402786-19-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne During execution, page table entries are subject to memory protection simply because their GPA is obtained with an EPT walk. During emulation, the checks need to be done by hand and result in an -EFAULT userspace exit. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/paging_tmpl.h | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index cdf05cd76d63..23d7a7d9769f 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -89,6 +89,7 @@ struct guest_walker { unsigned int pt_access[PT_MAX_FULL_LEVELS]; unsigned int pte_access; gfn_t gfn; + bool memory_attributes_fault; struct x86_exception fault; }; =20 @@ -342,6 +343,7 @@ static int FNAME(walk_addr_generic)(struct guest_walker= *walker, =20 trace_kvm_mmu_pagetable_walk(addr, access); retry_walk: + walker->memory_attributes_fault =3D false; walker->level =3D w->cpu_role.base.level; pte =3D kvm_mmu_get_guest_pgd(vcpu, w); have_ad =3D PT_HAVE_ACCESSED_DIRTY(w); @@ -411,6 +413,12 @@ static int FNAME(walk_addr_generic)(struct guest_walke= r *walker, if (unlikely(kvm_is_error_hva(host_addr))) goto error; =20 + if (!kvm_mem_attributes_may_read_gfn(vcpu->kvm, gpa_to_gfn(real_gpa))) { + walker->memory_attributes_fault =3D true; + walker->gfn =3D gpa_to_gfn(real_gpa); + goto error; + } + ptep_user =3D (pt_element_t __user *)((void *)host_addr + offset); if (unlikely(get_user(pte, ptep_user))) goto error; @@ -820,6 +828,12 @@ static int FNAME(page_fault)(struct kvm_vcpu *vcpu, st= ruct kvm_page_fault *fault * The page is not mapped by the guest. Let the guest handle it. */ if (!r) { + if (walker.memory_attributes_fault) { + fault->gfn =3D walker.gfn; + kvm_mmu_prepare_memory_fault_exit(vcpu, fault); + return -EFAULT; + } + if (!fault->prefetch) __kvm_inject_emulated_page_fault(vcpu, &walker.fault, true); =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 26C01444702 for ; Thu, 16 Jul 2026 18:15:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; cv=none; b=qP07ycJwkPcaIft4pmM252Uqws4MIZSTnqe6Q65KEDhpdjfz+kLFTeR0L3wptEHtFmiFQOx3sQoenlFtGNy1R68phXngJZdI+LZ3MrLAykdSnlMzHVrmWj6NthL/AlloeW+SmGzRIKPU22WcAX1mEV4i4tOaEBQ+BAmP5/XtZwI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; c=relaxed/simple; bh=3GJ14NznGThun6XqlqpKmqJR+Xni8XyTM17YQYecO3o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=AGr0bWCOYCr7FMykza1fEZCjw3ccsBX2zzcnMxHPw8OpjimCAQfcZJE44gKptK/ofCUDVW1+3WJGULjoMDTGMebe+Na1CF6DMOGXVLuWQKjbPrUMg9r3acg88CqUt04gpDU+h63h1MXbtpS1LSDLFD45AxCm7z82jXTDQVgqk+I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=a0uyRHFh; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="a0uyRHFh" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225715; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=+YDPV/jgMTVU5QltBcVl6pkFbi8mbShuNaZfaELV+ck=; b=a0uyRHFhkNT8ogtLBV5CMqm5Kq2CEdJybCo+N6DZHLoxuaMd5hj+1tFCyqgP5Z78P3fPn3 V1nT+iusQc1MOO3yeKE0lKmbmVBBqUF2PpkWNB79YxGWuiVAqoq6xEaLb3s71IMbQZJ9g5 PlTx0Q9Mos/6kocxyKtjHytb5hxMxvo= Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-622-oBwcCW2mOACbpJjPSZ0qUA-1; Thu, 16 Jul 2026 14:15:11 -0400 X-MC-Unique: oBwcCW2mOACbpJjPSZ0qUA-1 X-Mimecast-MFC-AGG-ID: oBwcCW2mOACbpJjPSZ0qUA_1784225710 Received: from mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.93]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 4CD861955E8C; Thu, 16 Jul 2026 18:15:10 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id C7FD21800480; Thu, 16 Jul 2026 18:15:09 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 19/24] KVM: x86/mmu: Do not update accessed/dirty if guest PTE is read-only Date: Thu, 16 Jul 2026 14:14:51 -0400 Message-ID: <20260716181456.402786-20-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.93 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne In this case do not go all the way out to userspace [RFC] Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/paging_tmpl.h | 3 +++ 1 file changed, 3 insertions(+) diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index 23d7a7d9769f..ee5c23d1b77a 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -419,6 +419,9 @@ static int FNAME(walk_addr_generic)(struct guest_walker= *walker, goto error; } =20 + if (!kvm_mem_attributes_may_write_gfn(vcpu->kvm, gpa_to_gfn(real_gpa))) + walker->pte_writable[walker->level - 1] =3D false; + ptep_user =3D (pt_element_t __user *)((void *)host_addr + offset); if (unlikely(get_user(pte, ptep_user))) goto error; --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A3DF144606C for ; Thu, 16 Jul 2026 18:15:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225724; cv=none; b=qAM2EI6Kjnijqbt1mXH+RmtFD1WFoWvytGVatEVi/VHja/jQXu6epTASp7DfQwxEms+SgsG/5k1+nlJ/ufonaUNap0fO5rVIJq34mtAQfqrwRNStpB91t1Uluilnv9oOK0PjpKAIbp63u0KsKcfyk4RqXRbqiOiSHN6rf0uzqXw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225724; c=relaxed/simple; bh=6utOy21KcM4jcYjjQ5iL5Jv1gjyFvDWujTytWlx+NcY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=jzqJiPDavszl0NjoB74e5AKT9G7osyxBkxwC6afjUafvQRQoHvftqGXntzupR8L/j24gGMC6zuGHWV0cf5NBSkzdI7UMv8qMG6Rfsr26sYQXENphts6fcrwwgBMGAZkUjIsNhY1jBdKHaFWuE/UEMc6Ue+7zZY4TBpE85XdAxfI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Z5oGLjGF; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Z5oGLjGF" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225715; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=26RlGOwVuSPPPr5KjC7HxaW648Zwf/4T1lFh+L2G5ZI=; b=Z5oGLjGFlnclQLyidfWH0Wf54b5OnS4G7zrp+2EWDBQcrnNO+NPefu387bWHpXL/x/zZ6T MgqfoWrNL0xsfNyyQwksi8xUpL95/QuVRtS7yC6e3rHL0xEop67hRPTw88c6WiQXktdy1V qQiNQye5UhcXl88kjLzEpgxrF6kVAqw= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-680-phfHM5JiONWll_ixbyeHhQ-1; Thu, 16 Jul 2026 14:15:12 -0400 X-MC-Unique: phfHM5JiONWll_ixbyeHhQ-1 X-Mimecast-MFC-AGG-ID: phfHM5JiONWll_ixbyeHhQ_1784225711 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 7C9B11955E90; Thu, 16 Jul 2026 18:15:11 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 0063E195604E; Thu, 16 Jul 2026 18:15:10 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 20/24] KVM: x86/mmu: Do not prefetch sptes on gfns backed by memory attributes Date: Thu, 16 Jul 2026 14:14:52 -0400 Message-ID: <20260716181456.402786-21-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Prefetched SPTEs are always given full access. Do not prefetch GFNs that have memory protections applied. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 59c2a04f5648..8e85f672a08d 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -3156,6 +3156,9 @@ static bool kvm_mmu_prefetch_sptes(struct kvm_vcpu *v= cpu, gfn_t gfn, u64 *sptep, return false; =20 for (i =3D 0; i < nr_pages; i++, gfn++, sptep++) { + if (kvm_get_memory_attributes(vcpu->kvm, gfn)) + continue; + mmu_set_spte(vcpu, slot, sptep, access, gfn, page_to_pfn(pages[i]), NULL); =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B8EF446054 for ; Thu, 16 Jul 2026 18:15:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; cv=none; b=P3zhqW75Mo0FZDuLU/o7ijU1mPkg2JnTePhAptMqkw79oR2gxNpm+993lq41T5fcKcG6pULdWOytzhHqpu+pD0dwUutOEBkM8EggScqkhSITmL3COomenm3mDUyOYukpWdZVgANfh6dvkKf9+ApTC1ImieImwndA3zaT91l1CwU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225725; c=relaxed/simple; bh=TamKDFoXwMXvt9o99JKWhsElHOQQJNHR4/PZAeRDtDQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gQvVJA6M0f8DblyrVE9wbcNnbWJNJbJwa/NPH4V/hFY0xDSyv4A6XPW3krrJnWu4R5r1l5u5m2KQHN3ST6+X7E7NQVcsHEMTvS/vTJafTgRNK6TPgGVkU2NylKPFHX52eoUdH0OZdtA1eM7z/HhVAwhgazkpqQgjChwsFVfES9Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=RmtKEhG7; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="RmtKEhG7" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225714; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=QoRxtY/U4Rt7RD5JrPiCkn7z8UituYgwYp9NiSV3kS4=; b=RmtKEhG7DPV7zUkx1N6Q6keUONrCBNsDq2diSm0EyCgv6uMkbo5JN3DCWpdENCuvn47elf bht+XI62fDWZDgyMEV/r/GjTNWgwwveHjv5G1I5pSxvjUzfgN8JWW6osax78Y7tEKAiVap aNFZtiYnKDTBIxEYtwb6H0ELQiJa7aI= Received: from mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-489-sfaeIMkvMH6THvWiCKuv5Q-1; Thu, 16 Jul 2026 14:15:13 -0400 X-MC-Unique: sfaeIMkvMH6THvWiCKuv5Q-1 X-Mimecast-MFC-AGG-ID: sfaeIMkvMH6THvWiCKuv5Q_1784225712 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 20782180066E; Thu, 16 Jul 2026 18:15:12 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id A1C0D195604E; Thu, 16 Jul 2026 18:15:11 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 21/24] KVM: x86/mmu: Obsolete all roots if memattr contains gPTEs Date: Thu, 16 Jul 2026 14:14:53 -0400 Message-ID: <20260716181456.402786-22-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne see comments inside - this patch is incomplete. Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- arch/x86/kvm/mmu/mmu.c | 47 ++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 45 insertions(+), 2 deletions(-) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 8e85f672a08d..676bc6b61f03 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -6989,11 +6989,11 @@ static void kvm_zap_obsolete_pages(struct kvm *kvm) * not use any resource of the being-deleted slot or all slots * after calling the function. */ -static void kvm_mmu_zap_all_fast(struct kvm *kvm) +static void __kvm_mmu_zap_all_fast(struct kvm *kvm) { lockdep_assert_held(&kvm->slots_lock); + lockdep_assert_held_write(&kvm->mmu_lock); =20 - write_lock(&kvm->mmu_lock); trace_kvm_mmu_zap_all_fast(kvm); =20 /* @@ -7030,7 +7030,21 @@ static void kvm_mmu_zap_all_fast(struct kvm *kvm) kvm_make_all_cpus_request(kvm, KVM_REQ_MMU_FREE_OBSOLETE_ROOTS); =20 kvm_zap_obsolete_pages(kvm); +} =20 +/* + * Fast invalidate all shadow pages and use lock-break technique + * to zap obsolete pages. + * + * It's required when memslot is being deleted or VM is being + * destroyed, in these cases, we should ensure that KVM MMU does + * not use any resource of the being-deleted slot or all slots + * after calling the function. + */ +static void kvm_mmu_zap_all_fast(struct kvm *kvm) +{ + write_lock(&kvm->mmu_lock); + __kvm_mmu_zap_all_fast(kvm); write_unlock(&kvm->mmu_lock); =20 /* @@ -8221,6 +8235,8 @@ bool kvm_arch_post_set_memory_attributes(struct kvm *= kvm, { unsigned long attrs =3D range->arg.attributes; struct kvm_memory_slot *slot =3D range->slot; + bool gen_update =3D false; + struct kvm_mmu_page *sp; int level; =20 lockdep_assert_held_write(&kvm->mmu_lock); @@ -8281,6 +8297,33 @@ bool kvm_arch_post_set_memory_attributes(struct kvm = *kvm, hugepage_set_mixed(slot, gfn, level); } } + + /* + * There are special considerations when applying an memory protection + * attibute against a GPTE page. If set read-only, access/dirty bits + * within that page shouldn't be updated. If set non-accesible, + * accessing a virtual address that requires traversing that GPTE page + * should fault. + * + * On TDP enabled guests, the CPU faults on the GPTE address upon + * detecting such a situation. + * + * On non-TDP, upon detecting this situation, and based on the fact it + * should be a rare occasion, invalidate all the mmu roots. + */ + for (gfn_t gfn =3D range->start; gfn < range->end; gfn++) { + for_each_gfn_valid_sp_with_gptes(kvm, sp, gfn) { + gen_update =3D true; + trace_printk("needs gen update! %llx\n", gfn); + goto exit_loop; + } + } +exit_loop: + if (gen_update) + __kvm_mmu_zap_all_fast(kvm); + + // todo: what to do with kvm_tdp_mmu_zap_invalidated_roots()? + // add kvm_arch_post_set_memory_attributes_unlocked? return false; } =20 --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5A398446840 for ; Thu, 16 Jul 2026 18:15:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225726; cv=none; b=GyyKQgsrg/7QjvAdpAmzc3MvvzE/7kL8t8HHTd7BJVMRUNMiqq8itJ77EFDzQGxL1JMQmHJihb3WX8VohT8paMTjSpro80e4BnQtvXUG14N04jUDwikmRNRrYu/SudncDfBkkUfKxGjYu6cHO7bTArB8O+FpHu3jnatIJdlS4z8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225726; c=relaxed/simple; bh=zSf5J4TagZlI2eYRClP6nKX6VceYrLsbJd1z4PhFFN8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=WP4PbPKli1pd8DI4ap+HDkOcvvZiUGviv+AvZODncEw2I7EcaWCCfAjIwyuSnoPvzOHtLbrw9P+zqmnLEf0TKBgxGYCxmCkA2cbgKAgpLK05stzHCgSbzclvnq4Va7YFGbcgZ+bPQuLOaYEvNuEuvpfXsLEn0itUDOhYcz2AmIY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=JeNlzGA9; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="JeNlzGA9" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225717; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=N9BuHBT5zfcEzvmuoZO/X3B0wpsVcYOl+5SVDoPWrFo=; b=JeNlzGA9SGiStJYs1EZ1ZgtvubZbf8dLdb+oAdbSGgfWby2UMMkNFJYVQD+9xZGi0RYnS/ vtttQmF+b8RLVRe4EbQWVPBwnDpOe+ZH9Dregls76n2h+jG/4fTLhOUj1tfXoIApL8wIkI 6RLIX1EuUSNKyL+3Kgqn2feR2CsdVUQ= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-523-gTLeAb-SNWyD-IjGpOpOGA-1; Thu, 16 Jul 2026 14:15:13 -0400 X-MC-Unique: gTLeAb-SNWyD-IjGpOpOGA-1 X-Mimecast-MFC-AGG-ID: gTLeAb-SNWyD-IjGpOpOGA_1784225712 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id B6A041955F26; Thu, 16 Jul 2026 18:15:12 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 447B3195604E; Thu, 16 Jul 2026 18:15:12 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 22/24] KVM: x86: selftests: Introduce memory attributes test Date: Thu, 16 Jul 2026 14:14:54 -0400 Message-ID: <20260716181456.402786-23-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- tools/include/uapi/linux/kvm.h | 3 + tools/testing/selftests/kvm/Makefile.kvm | 1 + .../testing/selftests/kvm/include/kvm_util.h | 23 +- .../testing/selftests/kvm/memory_attributes.c | 362 ++++++++++++++++++ .../selftests/kvm/x86/memory_attributes.c | 43 +++ 5 files changed, 422 insertions(+), 10 deletions(-) create mode 100644 tools/testing/selftests/kvm/memory_attributes.c create mode 100644 tools/testing/selftests/kvm/x86/memory_attributes.c diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h index d0c0c8605976..fe2f500c932d 100644 --- a/tools/include/uapi/linux/kvm.h +++ b/tools/include/uapi/linux/kvm.h @@ -1639,6 +1639,9 @@ struct kvm_memory_attributes { __u64 flags; }; =20 +#define KVM_MEMORY_ATTRIBUTE_NR (1ULL << 0) +#define KVM_MEMORY_ATTRIBUTE_NW (1ULL << 1) +#define KVM_MEMORY_ATTRIBUTE_NX (1ULL << 2) #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3) =20 #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest= _memfd) diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selft= ests/kvm/Makefile.kvm index 4ace12606e93..3e2a6e4398eb 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -162,6 +162,7 @@ TEST_GEN_PROGS_x86 +=3D rseq_test TEST_GEN_PROGS_x86 +=3D steal_time TEST_GEN_PROGS_x86 +=3D system_counter_offset_test TEST_GEN_PROGS_x86 +=3D pre_fault_memory_test +TEST_GEN_PROGS_x86 +=3D memory_attributes =20 # Compiled outputs used by test targets TEST_GEN_PROGS_EXTENDED_x86 +=3D x86/nx_huge_pages_test diff --git a/tools/testing/selftests/kvm/include/kvm_util.h b/tools/testing= /selftests/kvm/include/kvm_util.h index 04a910164a29..54d73ea46e11 100644 --- a/tools/testing/selftests/kvm/include/kvm_util.h +++ b/tools/testing/selftests/kvm/include/kvm_util.h @@ -418,24 +418,27 @@ static inline void vm_enable_cap(struct kvm_vm *vm, u= 32 cap, u64 arg0) vm_ioctl(vm, KVM_ENABLE_CAP, &enable_cap); } =20 -static inline void vm_set_memory_attributes(struct kvm_vm *vm, gpa_t gpa, - u64 size, u64 attributes) +static inline int __vm_set_memory_attributes(struct kvm_vm *vm, u64 gpa, + u64 size, u64 attributes, + u64 flags) { struct kvm_memory_attributes attr =3D { .attributes =3D attributes, .address =3D gpa, .size =3D size, - .flags =3D 0, + .flags =3D flags, }; =20 - /* - * KVM_SET_MEMORY_ATTRIBUTES overwrites _all_ attributes. These flows - * need significant enhancements to support multiple attributes. - */ - TEST_ASSERT(!attributes || attributes =3D=3D KVM_MEMORY_ATTRIBUTE_PRIVATE, - "Update me to support multiple attributes!"); + return __vm_ioctl(vm, KVM_SET_MEMORY_ATTRIBUTES, &attr); +} =20 - vm_ioctl(vm, KVM_SET_MEMORY_ATTRIBUTES, &attr); +static inline void vm_set_memory_attributes(struct kvm_vm *vm, gpa_t gpa, + u64 size, u64 attributes) +{ + int rc; + + rc =3D __vm_set_memory_attributes(vm, gpa, size, attributes, 0); + TEST_ASSERT_VM_VCPU_IOCTL(!rc, KVM_SET_MEMORY_ATTRIBUTES, rc, vm); } =20 =20 diff --git a/tools/testing/selftests/kvm/memory_attributes.c b/tools/testin= g/selftests/kvm/memory_attributes.c new file mode 100644 index 000000000000..7066ae791b99 --- /dev/null +++ b/tools/testing/selftests/kvm/memory_attributes.c @@ -0,0 +1,362 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (C) 2024, Amazon.com, Inc. or its affiliates. All Rights Rese= rved + * + * Test for KVM_MEMORY_ATTRIBUTES + */ +#include +#include + +#include "test_util.h" +#include "ucall_common.h" +#include "kvm_util.h" +#include "processor.h" +#include "hyperv.h" +#include "apic.h" +#include "asm/pvclock-abi.h" + +#define KVM_MEMORY_ATTRIBUTE_NO_ACCESS \ + (KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW | \ + KVM_MEMORY_ATTRIBUTE_NX) + +#define MMIO_GPA 0x700000000 +#define MMIO_GVA MMIO_GPA + +enum { + TEST_OP_NOP, + TEST_OP_READ, + TEST_OP_WRITE, + TEST_OP_EXEC, + TEST_OP_EXIT, +}; + +const char *test_op_names[] =3D +{ + [TEST_OP_READ] =3D "Read", + [TEST_OP_WRITE] =3D "Write", + [TEST_OP_EXEC] =3D "Exec", + [TEST_OP_EXIT] =3D "Exit", +}; + +struct test_data { + uint8_t op; + int stage; + gva_t vaddr; + + struct kvm_vcpu *vcpu; +}; + +static struct test_data *test_data; + +static uint64_t arch_controlled_read(gva_t addr); +static void arch_controlled_write(gva_t addr, uint64_t val); +static void arch_controlled_exec(gva_t addr); +static void arch_write_return_insn(struct kvm_vm *vm, gpa_t vaddr); + +static void guest_code(void *data) +{ + struct test_data *test_data =3D data; + int stage =3D 1; + + while (true) { + gva_t vaddr =3D READ_ONCE(test_data->vaddr); + + switch(READ_ONCE(test_data->op)) { + case TEST_OP_READ: + (void) arch_controlled_read(vaddr); + GUEST_SYNC(stage++); + break; + case TEST_OP_WRITE: + arch_controlled_write(vaddr, 1); + GUEST_SYNC(stage++); + break; + case TEST_OP_EXEC: + arch_controlled_exec(vaddr); + GUEST_SYNC(stage++); + break; + default: + goto exit; + }; + } + +exit: + GUEST_DONE(); +} + +static void vcpu_run_and_inc_stage(struct kvm_vcpu *vcpu) +{ + struct ucall uc; + + vcpu_run(vcpu); + + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO); + switch (get_ucall(vcpu, &uc)) { + case UCALL_SYNC: + TEST_ASSERT(uc.args[1] =3D=3D test_data->stage, + "Unexpected stage: %ld (%d expected)", + uc.args[1], test_data->stage); + break; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + /* NOT REACHED */ + default: + TEST_FAIL("Unknown ucall %lu", uc.cmd); + } + + test_data->stage++; +} + +static void test_page_restricted(struct kvm_vcpu *vcpu, int op, + gva_t vaddr, gpa_t fault_paddr, + uint64_t fault_reason) +{ + struct kvm_vm *vm =3D vcpu->vm; + int rc; + + test_data->op =3D op; + test_data->vaddr =3D vaddr; + + rc =3D _vcpu_run(vcpu); + TEST_ASSERT(rc =3D=3D -1 && errno =3D=3D EFAULT, + "KVM_RUN IOCTL didn't return EFAULT on %s, rc %d, errno %d", + test_op_names[op], rc, errno); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_MEMORY_FAULT); + TEST_ASSERT_EQ(vcpu->run->memory_fault.gpa, fault_paddr); + TEST_ASSERT_EQ(vcpu->run->memory_fault.flags, fault_reason); + TEST_ASSERT_EQ(vcpu->run->memory_fault.size, vm->page_size); +} + +static void test_page_accessible(struct kvm_vcpu *vcpu, int op, gva_t vadd= r) +{ + test_data->op =3D op; + test_data->vaddr =3D vaddr; + vcpu_run_and_inc_stage(vcpu); +} + +#include "x86/memory_attributes.c" + +/* + * We want to test the following cases: + * - Sucessful access to GPAs backed by memory attributes (for ex. read ac= cess + * on an read-only page). + * - First fault after setting memory attributes, with unpopulated SPTEs/E= PTS. + * - Fault caused by an SPTE/EPT reflecting the memory attributes. + * + * The list of ops below tests the 3 situations for each memory attribute + * combination. + */ +const struct memory_access { + const char *name; + uint64_t attrs; + int ops[5]; +} access_array[] =3D { + { "all allowed", 0, { TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC } }, + { "no write (unmapped)", KVM_MEMORY_ATTRIBUTE_NW, + { TEST_OP_WRITE, TEST_OP_READ, TEST_OP_EXEC, TEST_OP_WRITE } }, + { "no write (mapped)", KVM_MEMORY_ATTRIBUTE_NW, + { TEST_OP_READ, TEST_OP_WRITE, TEST_OP_READ, TEST_OP_EXEC, TEST_OP_WRIT= E } }, + { "no exec (unmapped)", KVM_MEMORY_ATTRIBUTE_NX, + { TEST_OP_EXEC, TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC } }, + { "no exec (mapped)", KVM_MEMORY_ATTRIBUTE_NX, + { TEST_OP_READ, TEST_OP_EXEC, TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC= } }, + { "read only", KVM_MEMORY_ATTRIBUTE_NW | KVM_MEMORY_ATTRIBUTE_NX, + { TEST_OP_EXEC, TEST_OP_WRITE, TEST_OP_READ, TEST_OP_WRITE, + TEST_OP_EXEC } }, + { "no access (map on read)", KVM_MEMORY_ATTRIBUTE_NO_ACCESS, + { TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC } }, + { "no access (map on write)", KVM_MEMORY_ATTRIBUTE_NO_ACCESS, + { TEST_OP_WRITE, TEST_OP_READ, TEST_OP_EXEC } }, + { "no access (map on exec)", KVM_MEMORY_ATTRIBUTE_NO_ACCESS, + { TEST_OP_EXEC, TEST_OP_READ, TEST_OP_WRITE } }, + /* Verify everything is back to normal */ + { "all allowed (2)", 0, { TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC } }, +}; + +static void test_page_access(struct kvm_vcpu *vcpu, gva_t vaddr, + uint64_t attrs, const int ops[]) +{ + struct kvm_vm *vm =3D vcpu->vm; + gpa_t paddr =3D addr_gva2gpa(vm, vaddr); + + for (int i =3D 0; i < ARRAY_SIZE(access_array[0].ops); i++) { + int op =3D ops[i]; + + if (op =3D=3D TEST_OP_NOP) + continue; + + /* + * We're about to have the guest jump into 'vaddr', make it a + * 'ret' instruction so it returns right away. + */ + if (op =3D=3D TEST_OP_EXEC) + arch_write_return_insn(vm, paddr); + + vm_set_memory_attributes(vm, paddr, vm->page_size, attrs); + + /* + * Attributes are negated, a match means the operation should + * fail. + */ + if (attrs & BIT_ULL(op - 1)) { + test_page_restricted(vcpu, op, vaddr, paddr, BIT(op - 1)); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + } + + test_page_accessible(vcpu, op, vaddr); + } + +} + +static void test_memory_access(struct kvm_vcpu *vcpu, gva_t test_vm_vaddr, + size_t size) +{ + struct kvm_vm *vm =3D vcpu->vm; + gpa_t test_vm_paddr =3D addr_gva2gpa(vm, test_vm_vaddr); + + printf("gva %lx gpa %lx\n", test_vm_vaddr, test_vm_paddr); + for (size_t i =3D 0; i < ARRAY_SIZE(access_array); i++) { + uint64_t attrs =3D access_array[i].attrs; + + printf("starting %s test...\n", access_array[i].name); + vm_set_memory_attributes(vm, test_vm_paddr, size, attrs); + + for (gva_t vaddr =3D test_vm_vaddr; + vaddr < test_vm_vaddr + size; vaddr +=3D PAGE_SIZE) { + test_page_access(vcpu, vaddr, attrs, access_array[i].ops); + } + } +} + +static void test_memattrs_ignore_mmio(struct kvm_vcpu *vcpu) +{ + struct kvm_vm *vm =3D vcpu->vm; + + vm_set_memory_attributes(vm, MMIO_GPA, vm->page_size, + KVM_MEMORY_ATTRIBUTE_NO_ACCESS); + + test_data->op =3D TEST_OP_READ; + test_data->vaddr =3D MMIO_GVA; + vcpu_run(vcpu); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_MMIO); + TEST_ASSERT_EQ(vcpu->run->mmio.phys_addr, MMIO_GPA); + TEST_ASSERT_EQ(vcpu->run->mmio.is_write, 0); + TEST_ASSERT_EQ(vcpu->run->mmio.len, 8); + vcpu_run_and_inc_stage(vcpu); + + test_data->op =3D TEST_OP_WRITE; + test_data->vaddr =3D MMIO_GVA; + vcpu_run(vcpu); + TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_MMIO); + TEST_ASSERT_EQ(vcpu->run->mmio.phys_addr, MMIO_GPA); + TEST_ASSERT_EQ(vcpu->run->mmio.is_write, 1); + TEST_ASSERT_EQ(vcpu->run->mmio.len, 8); + vcpu_run_and_inc_stage(vcpu); + + vm_set_memory_attributes(vm, MMIO_GPA, vm->page_size, 0); +} + +static void test_input_validation(struct kvm_vm *vm) +{ + uint64_t flags, gpa =3D 0, size =3D 0, attrs =3D 0; + int rc; + + /* 'flags' is unsupported */ + flags =3D BIT(0); + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* 'size' can't be 0 */ + flags =3D 0; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* 'gpa' shouldn't overflow */ + gpa =3D 0ULL - vm->page_size; + size =3D vm->page_size; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* 'gpa' should be page aligned */ + gpa =3D 1; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* 'size' should be page aligned */ + gpa =3D 0; + size =3D 1; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* exec mappings require read access */ + size =3D vm->page_size; + attrs =3D KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* write mappings require read access */ + size =3D vm->page_size; + attrs =3D KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NX; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); + + /* private mappings are incompatible with access restrictions */ + attrs =3D KVM_MEMORY_ATTRIBUTE_NW | KVM_MEMORY_ATTRIBUTE_PRIVATE; + rc =3D __vm_set_memory_attributes(vm, gpa, size, attrs, flags); + TEST_ASSERT_VM_VCPU_IOCTL(rc =3D=3D -1 && errno =3D=3D EINVAL, + KVM_SET_MEMORY_ATTRIBUTES, rc, vm); +} + +static void test_finalize(struct kvm_vcpu *vcpu) +{ + test_data->op =3D TEST_OP_EXIT; + vcpu_run(vcpu); + TEST_ASSERT_EQ(get_ucall(vcpu, NULL), UCALL_DONE); +} + +static struct test_data *init_test_data(struct kvm_vcpu *vcpu) +{ + struct kvm_vm *vm =3D vcpu->vm; + gva_t test_data_vm_vaddr; + + test_data_vm_vaddr =3D vm_alloc_page(vm); + vcpu_args_set(vcpu, 1, test_data_vm_vaddr); + + test_data =3D addr_gva2hva(vm, test_data_vm_vaddr); + test_data->stage =3D 1; + test_data->vcpu =3D vcpu; + + return test_data; +} + +int main(int argc, char *argv[]) +{ + uint32_t guest_page_size =3D vm_guest_mode_params[VM_MODE_DEFAULT].page_s= ize; + unsigned int ptes_per_page =3D guest_page_size / 8; + size_t size =3D guest_page_size * ptes_per_page * 2; /* 2 huge-pages */ + struct kvm_vcpu *vcpu; + gva_t test_mem; + struct kvm_vm *vm; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_MEMORY_ATTRIBUTES) & + KVM_MEMORY_ATTRIBUTE_NO_ACCESS); + + vm =3D __vm_create_with_one_vcpu(&vcpu, size, guest_code); + test_mem =3D vm_alloc(vm, size, KVM_UTIL_MIN_VADDR); + virt_map(vcpu->vm, MMIO_GVA, MMIO_GPA, 1); + test_data =3D init_test_data(vcpu); + + test_input_validation(vm); + test_memory_access(vcpu, test_mem, size); + test_memattrs_ignore_mmio(vcpu); + test_finalize(vcpu); + + kvm_vm_free(vm); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/memory_attributes.c b/tools/te= sting/selftests/kvm/x86/memory_attributes.c new file mode 100644 index 000000000000..2e1148f5146d --- /dev/null +++ b/tools/testing/selftests/kvm/x86/memory_attributes.c @@ -0,0 +1,43 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (C) 2024, Amazon.com, Inc. or its affiliates. All Rights Rese= rved + * + * Test for KVM_MEMORY_ATTRIBUTES + */ +#include "kvm_util.h" +#include "apic.h" + +uint64_t arch_controlled_read(gva_t addr) +{ + uint64_t val; + + asm volatile("mov %[addr], %%rax \n\r" + "mov (%%rax), %[val] \n\r" + : [val] "=3Dr" (val) + : [addr] "m"(addr) + : "memory", "rax"); + + return val; +} + +void arch_controlled_write(gva_t addr, uint64_t val) +{ + asm volatile("mov %[addr], %%rax \n\r" + "mov %[val], %%rbx \n\r" + "mov %%rbx, (%%rax) \n\r" + :: [addr] "m" (addr), [val] "m" (val) + : "memory", "rax", "rbx"); +} + +void arch_controlled_exec(gva_t addr) +{ + asm volatile("mov %[addr], %%rax \n\r" + "call *%%rax \n\t" + :: [addr] "m"(addr) + : "memory", "rax"); +} + +void arch_write_return_insn(struct kvm_vm *vm, gpa_t vaddr) +{ + memset(addr_gpa2hva(vm, vaddr), 0xc3, 1); +} --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BB6CB446826 for ; Thu, 16 Jul 2026 18:15:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225728; cv=none; b=a2G8LsCqSpZ1u1nyP5mV3iJTLcQv2475bpVVChodIJ+4zl6Li/frBFhOa/Riaze+fssKp6piC2CBGaxZ55VwveHVKE1czKK0DDegRANuwthGIZIsUftB/sK3q2/zQk3Mb+kH7dAsS5CT3hRwopZuNPEL/bWLfZT0263+S8WR/Yo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225728; c=relaxed/simple; bh=X2Mp5tqWG/hlLXxMeFeNp6AJ24Jm10L1p4PmCQnTdF0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=k6Sj5cmR09ktznbQ1vG5IZ7YNi5w7mdXLVCPYXZ0WupZycwrX1BRJybo79HzSdgbQdlLM9+Pi0Dfm0joy56s1TPcp1CRp3QEe7wdPVtrs9vGDDosRHRWaB88VkhXWuFQwCl2GeQ/AlpUdh0gSpSCozYco1TZJgCBgn1lFXf83/M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=eAxR705V; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="eAxR705V" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225717; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=bf+r9m213uUwsm4gz+90/tHgCey5uMsxmGmIjXRW02Q=; b=eAxR705VTyDdZkd89m2oI1NgkdRjHW8yuMiThJ//ebov46tvr87GHm566ZOv90C9kWPSKy 4phxTTO53mxB7rjbVxqFzN0/AGUCkDAYnwn5/v4H1JC44rkktRjOzdy12eqdml8WlElRER vogOTl0saDebM9lKOPdXgQc4Udp04Uk= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-606-oXL4oLieOqeyirX-Tcubsw-1; Thu, 16 Jul 2026 14:15:14 -0400 X-MC-Unique: oXL4oLieOqeyirX-Tcubsw-1 X-Mimecast-MFC-AGG-ID: oXL4oLieOqeyirX-Tcubsw_1784225713 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 597B31955E75; Thu, 16 Jul 2026 18:15:13 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id DB946195604E; Thu, 16 Jul 2026 18:15:12 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 23/24] KVM: x86: selftests: Introduce memory attributes PTE test Date: Thu, 16 Jul 2026 14:14:55 -0400 Message-ID: <20260716181456.402786-24-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Introduce the memory attributes PTE tests. This test confirms the behaviour of memory attributes access restrictions when installed on a page that holds guest page table entries. Notably two cases are taken into account: - The page is made non-accesible. In such case the next access to a virtual memory address translated by that paging structure should fault. - The page is made read-only. In such case the next access to a virtual memory address translated by that paging structure should either fault or succeed yet not perform any writes into the PTE (accessed and dirty bits). Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- .../testing/selftests/kvm/memory_attributes.c | 38 +++++- .../selftests/kvm/x86/memory_attributes.c | 126 ++++++++++++++++++ 2 files changed, 162 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/kvm/memory_attributes.c b/tools/testin= g/selftests/kvm/memory_attributes.c index 7066ae791b99..c66d5d085a9c 100644 --- a/tools/testing/selftests/kvm/memory_attributes.c +++ b/tools/testing/selftests/kvm/memory_attributes.c @@ -22,11 +22,16 @@ #define MMIO_GPA 0x700000000 #define MMIO_GVA MMIO_GPA =20 +#define PT_WRITABLE_MASK BIT_ULL(1) +#define PT_ACCESSED_MASK BIT_ULL(5) +#define PTE_VADDR 0x1000000000 + enum { TEST_OP_NOP, TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC, + TEST_OP_INVPLG, TEST_OP_EXIT, }; =20 @@ -35,6 +40,7 @@ const char *test_op_names[] =3D [TEST_OP_READ] =3D "Read", [TEST_OP_WRITE] =3D "Write", [TEST_OP_EXEC] =3D "Exec", + [TEST_OP_INVPLG] =3D "Invplg", [TEST_OP_EXIT] =3D "Exit", }; =20 @@ -42,6 +48,7 @@ struct test_data { uint8_t op; int stage; gva_t vaddr; + uint64_t expected_val; =20 struct kvm_vcpu *vcpu; }; @@ -59,6 +66,7 @@ static void guest_code(void *data) int stage =3D 1; =20 while (true) { + uint64_t expected_val =3D READ_ONCE(test_data->expected_val); gva_t vaddr =3D READ_ONCE(test_data->vaddr); =20 switch(READ_ONCE(test_data->op)) { @@ -67,13 +75,20 @@ static void guest_code(void *data) GUEST_SYNC(stage++); break; case TEST_OP_WRITE: - arch_controlled_write(vaddr, 1); + arch_controlled_write(vaddr, expected_val); GUEST_SYNC(stage++); break; case TEST_OP_EXEC: arch_controlled_exec(vaddr); GUEST_SYNC(stage++); break; +#ifdef __x86_64__ + case TEST_OP_INVPLG: + asm volatile("invlpg (%0)" + :: "b" (vaddr): "memory"); + GUEST_SYNC(stage++); + break; +#endif default: goto exit; }; @@ -106,6 +121,21 @@ static void vcpu_run_and_inc_stage(struct kvm_vcpu *vc= pu) test_data->stage++; } =20 +static int test_page(struct kvm_vcpu *vcpu, int op, gva_t vaddr) +{ + int rc; + + test_data->op =3D op; + test_data->vaddr =3D vaddr; + + rc =3D _vcpu_run(vcpu); + + if (rc >=3D 0) + test_data->stage++; + + return rc < 0 ? -errno : rc; +} + static void test_page_restricted(struct kvm_vcpu *vcpu, int op, gva_t vaddr, gpa_t fault_paddr, uint64_t fault_reason) @@ -122,7 +152,8 @@ static void test_page_restricted(struct kvm_vcpu *vcpu,= int op, test_op_names[op], rc, errno); TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_MEMORY_FAULT); TEST_ASSERT_EQ(vcpu->run->memory_fault.gpa, fault_paddr); - TEST_ASSERT_EQ(vcpu->run->memory_fault.flags, fault_reason); + if (fault_reason) + TEST_ASSERT_EQ(vcpu->run->memory_fault.flags, fault_reason); TEST_ASSERT_EQ(vcpu->run->memory_fault.size, vm->page_size); } =20 @@ -355,6 +386,9 @@ int main(int argc, char *argv[]) test_input_validation(vm); test_memory_access(vcpu, test_mem, size); test_memattrs_ignore_mmio(vcpu); +#ifdef __x86_64__ + arch_test_memory_access_pte(vcpu, test_mem); +#endif test_finalize(vcpu); =20 kvm_vm_free(vm); diff --git a/tools/testing/selftests/kvm/x86/memory_attributes.c b/tools/te= sting/selftests/kvm/x86/memory_attributes.c index 2e1148f5146d..3d2bcfe71884 100644 --- a/tools/testing/selftests/kvm/x86/memory_attributes.c +++ b/tools/testing/selftests/kvm/x86/memory_attributes.c @@ -41,3 +41,129 @@ void arch_write_return_insn(struct kvm_vm *vm, gpa_t va= ddr) { memset(addr_gpa2hva(vm, vaddr), 0xc3, 1); } + +/* + * This test validates that, during a page walk, if the page a PTE is plac= ed in + * is read-only the accesss and dirty bits will not be written. Note There= 's a + * slight variation in behaviour between TDP and non-TDP VMs: + * - With TDP enabled, KVM issues a fault exit upon observing the non-wri= table + * page. + * - With non-TDP, the access bit is not set, but the walk succeeds. + * + * This is aligned with read-only memslots' behaviour. + */ +static void test_memory_access_pte_ro(struct kvm_vcpu *vcpu, gva_t vaddr) +{ + struct kvm_vm *vm =3D vcpu->vm; + gpa_t paddr; + u64 *pte; + const u64 accessed_mask =3D PTE_ACCESSED_MASK(&vm->mmu); + + pte =3D vm_get_pte(vm, vaddr); + paddr =3D addr_hva2gpa(vm, pte) & GENMASK(61, vm->page_shift); + + *pte &=3D ~accessed_mask; + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= W); + if (test_page(vcpu, TEST_OP_READ, vaddr) < 0) { + test_page_restricted(vcpu, TEST_OP_READ, vaddr, paddr, + /* write PTE's accessed bit */ + KVM_MEMORY_EXIT_FLAG_WRITE); + + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + test_page_accessible(vcpu, TEST_OP_READ, vaddr); + TEST_ASSERT_EQ(*pte & accessed_mask, accessed_mask); + + /* Re-run the test, now vaddr is backed by an EPT. */ + *pte &=3D ~accessed_mask; + vm_set_memory_attributes(vm, paddr, vm->page_size, + KVM_MEMORY_ATTRIBUTE_NW); + test_page_restricted(vcpu, TEST_OP_READ, vaddr, paddr, + /* write PTE's accessed bit */ + KVM_MEMORY_EXIT_FLAG_WRITE); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + test_page_accessible(vcpu, TEST_OP_READ, vaddr); + } else { + TEST_ASSERT_EQ(*pte & accessed_mask, 0); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + } +} + +/* + * This test validates that, during a page walk, if the page a PTE is plac= ed in + * is maked as non-accesible, KVM issues a fault exit. + */ +static void test_memory_access_pte_nr(struct kvm_vcpu *vcpu, gva_t vaddr) +{ + struct kvm_vm *vm =3D vcpu->vm; + gpa_t paddr; + uint64_t *pte; + + pte =3D vm_get_pte(vm, vaddr); + paddr =3D addr_hva2gpa(vm, pte) & GENMASK(61, vm->page_shift); + + vm_set_memory_attributes(vm, paddr, vm->page_size, + KVM_MEMORY_ATTRIBUTE_NO_ACCESS); + + test_page_restricted(vcpu, TEST_OP_READ, vaddr, paddr, 0); + + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + test_page_accessible(vcpu, TEST_OP_READ, vaddr); + + /* Re-run the test, now vaddr is backed by an SPTE. */ + vm_set_memory_attributes(vm, paddr, vm->page_size, + KVM_MEMORY_ATTRIBUTE_NO_ACCESS); + test_page_restricted(vcpu, TEST_OP_READ, vaddr, paddr, 0); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + test_page_accessible(vcpu, TEST_OP_READ, vaddr); +} + +static void test_memory_access_sync_spte(struct kvm_vcpu *vcpu, gva_t vadd= r) +{ + struct kvm_vm *vm =3D vcpu->vm; + gpa_t paddr =3D addr_gva2gpa(vm, vaddr); + uint64_t *pte, old_pte, new_pte; + + pte =3D vm_get_pte(vm, vaddr); + gpa_t pte_paddr =3D addr_hva2gpa(vm, pte); + gpa_t pte_page_paddr =3D pte_paddr & GENMASK(61, vm->page_shift); + int pte_offset =3D pte_paddr - pte_page_paddr; + virt_pg_map(vm, PTE_VADDR, pte_page_paddr); + old_pte =3D *pte; + + /* Set vmaddr as non-executable */ + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= X); + + /* + * Make sure SPTEs are populated as previous op might have destroyed + * them. We new have a non-executable SPTE. + */ + test_page_accessible(vcpu, TEST_OP_READ, vaddr); + + /* + * Update PTE, make it non-writable and flush TLBs to make sure we go + * through the sync_spte path. This should update the SPTE and make it + * read-only. + */ + new_pte =3D (old_pte & ~PT_WRITABLE_MASK) | PT_ACCESSED_MASK; + test_data->expected_val =3D new_pte; + test_page_accessible(vcpu, TEST_OP_WRITE, PTE_VADDR + pte_offset); + TEST_ASSERT_EQ(*pte, new_pte); + test_page_accessible(vcpu, TEST_OP_INVPLG, vaddr); + + /* The not executable attrs remain valid */ + arch_write_return_insn(vm, vaddr); + test_page_restricted(vcpu, TEST_OP_EXEC, vaddr, paddr, + KVM_MEMORY_EXIT_FLAG_EXEC); + + /* Cleanup */ + *pte =3D old_pte; + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + test_page_accessible(vcpu, TEST_OP_EXEC, vaddr); +} + +static void arch_test_memory_access_pte(struct kvm_vcpu *vcpu, gva_t vaddr) +{ + test_memory_access_pte_nr(vcpu, vaddr); + test_memory_access_pte_ro(vcpu, vaddr); + test_memory_access_sync_spte(vcpu, vaddr); +} --=20 2.52.0 From nobody Mon Jul 27 18:59:23 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4D8FD445AFC for ; Thu, 16 Jul 2026 18:15:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225728; cv=none; b=JiiKDVxKnnwngvA48uU+qN1b1VOZ3tREIkfwhDOzZCj0kljnf3fE+gPTl9uC22l+nXEtjTn+R5HEbgeur9IWm/TgLFGVpHhGqgr54hwJJF1vAST7KJczstmkKmwIqd8TealVdM6nibXXSFMAf/tiKriCALM8SdiGqHtZgXK6iqE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784225728; c=relaxed/simple; bh=eP2HFC3tD3/Odtv7gB+vAzMwhRUmOmY+r+l4Pc96sS4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=qSZhEgjvBwapVaAnsQpEi9BY+K98aoTBYQCL313rWD47wlos0UM0nPJVo9ylvzdyEFPfiRIPg3GBC+xS/tsLg77d+ThPXT7QsGPf5MnXfQOqKcrBWyIhNRguoUy2LMr/LhWvBbD3wLB+QrrE9NTFp1emeAjnMWY2hZJ1y7EYToY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Q3B82AL5; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Q3B82AL5" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784225718; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=QMR1jZbfx2MG/9n6DX7JRLieA7qEPCzB1A94SEYRI2Y=; b=Q3B82AL5tukrst1Z1JP96mhjNg0cr5rvAkUGHs4YahVhJz7+Ta9J4tnwSkJvg9j3xczOFT 9YQJT8yNbtDTSkwVxDVOyy7SK3GgBUtQvnDKJQAK+hFFkv6gI0E/KxTebisFWUVk9rbNxq anbf89oEkc0rwWgRhLKuHA7trhodhJ4= Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-120-tlhDSg_pOsm1wyj0LIgNCg-1; Thu, 16 Jul 2026 14:15:14 -0400 X-MC-Unique: tlhDSg_pOsm1wyj0LIgNCg-1 X-Mimecast-MFC-AGG-ID: tlhDSg_pOsm1wyj0LIgNCg_1784225714 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id EF5441800671; Thu, 16 Jul 2026 18:15:13 +0000 (UTC) Received: from virtlab1023.lab.eng.rdu2.redhat.lab.eng.rdu2.redhat.com (virtlab1023.lab.eng.rdu2.redhat.com [10.8.1.187]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 7E569195604E; Thu, 16 Jul 2026 18:15:13 +0000 (UTC) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: nsaenz@amazon.com Subject: [PATCH 24/24] KVM: x86: selftests: Introduce memory attributes side-channel tests Date: Thu, 16 Jul 2026 14:14:56 -0400 Message-ID: <20260716181456.402786-25-pbonzini@redhat.com> In-Reply-To: <20260716181456.402786-1-pbonzini@redhat.com> References: <20260716181456.402786-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Content-Type: text/plain; charset="utf-8" From: Nicolas Saenz Julienne Introduce memory attributes selftests to catch any vulnerable side-channels. Memory attributes access restrictions are vulnerable to side-channel attacks. This means that any KVM operation initiated by the guest that requires guest memory access (which is the case for most para-virtualised interfaces) needs to consider memory attributes. The tests confirm this requirement is upheld for a variety of use-cases, and were selected to exercise specific approaches to guest memory access, including: - kvm_read/write_guest() - gfn_to_hva_cache - gfn_to_pfn_cache - Guest page walker (which uses __try_cmpxchg_user()) Signed-off-by: Nicolas Saenz Julienne Signed-off-by: Paolo Bonzini --- .../testing/selftests/kvm/include/kvm_util.h | 10 + .../selftests/kvm/include/x86/processor.h | 1 + .../testing/selftests/kvm/lib/x86/processor.c | 5 + .../testing/selftests/kvm/memory_attributes.c | 63 ++++- .../selftests/kvm/x86/memory_attributes.c | 216 ++++++++++++++++++ 5 files changed, 294 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/kvm/include/kvm_util.h b/tools/testing= /selftests/kvm/include/kvm_util.h index 54d73ea46e11..3160a25966a1 100644 --- a/tools/testing/selftests/kvm/include/kvm_util.h +++ b/tools/testing/selftests/kvm/include/kvm_util.h @@ -883,6 +883,16 @@ static inline int vcpu_get_stats_fd(struct kvm_vcpu *v= cpu) return fd; } =20 +static inline struct kvm_translation vcpu_translate(struct kvm_vcpu *vcpu, + u64 gva) +{ + struct kvm_translation tr; + + tr.linear_address =3D gva; + vcpu_ioctl(vcpu, KVM_TRANSLATE, &tr); + return tr; +} + int __kvm_has_device_attr(int dev_fd, u32 group, u64 attr); =20 static inline void kvm_has_device_attr(int dev_fd, u32 group, u64 attr) diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/te= sting/selftests/kvm/include/x86/processor.h index 7d3a27bc0d84..d9876377968f 100644 --- a/tools/testing/selftests/kvm/include/x86/processor.h +++ b/tools/testing/selftests/kvm/include/x86/processor.h @@ -1422,6 +1422,7 @@ static inline bool kvm_is_lbrv_enabled(void) return !!get_kvm_amd_param_integer("lbrv"); } =20 +u64 *vm_get_pte_level(struct kvm_vm *vm, gva_t gva, int *level); u64 *vm_get_pte(struct kvm_vm *vm, gva_t gva); =20 u64 kvm_hypercall(u64 nr, u64 a0, u64 a1, u64 a2, u64 a3); diff --git a/tools/testing/selftests/kvm/lib/x86/processor.c b/tools/testin= g/selftests/kvm/lib/x86/processor.c index ef56dcefe011..8b990016e764 100644 --- a/tools/testing/selftests/kvm/lib/x86/processor.c +++ b/tools/testing/selftests/kvm/lib/x86/processor.c @@ -397,6 +397,11 @@ u64 *tdp_get_pte(struct kvm_vm *vm, u64 l2_gpa) return __vm_get_page_table_entry(vm, &vm->stage2_mmu, l2_gpa, &level); } =20 +u64 *vm_get_pte_level(struct kvm_vm *vm, gva_t gva, int *level) +{ + return __vm_get_page_table_entry(vm, &vm->mmu, gva, level); +} + u64 *vm_get_pte(struct kvm_vm *vm, gva_t gva) { int level =3D PG_LEVEL_4K; diff --git a/tools/testing/selftests/kvm/memory_attributes.c b/tools/testin= g/selftests/kvm/memory_attributes.c index c66d5d085a9c..843fafcf4d96 100644 --- a/tools/testing/selftests/kvm/memory_attributes.c +++ b/tools/testing/selftests/kvm/memory_attributes.c @@ -19,6 +19,8 @@ (KVM_MEMORY_ATTRIBUTE_NR | KVM_MEMORY_ATTRIBUTE_NW | \ KVM_MEMORY_ATTRIBUTE_NX) =20 +#define HV_STATUS_INVALID_HYPERCALL_INPUT 3 + #define MMIO_GPA 0x700000000 #define MMIO_GVA MMIO_GPA =20 @@ -26,12 +28,35 @@ #define PT_ACCESSED_MASK BIT_ULL(5) #define PTE_VADDR 0x1000000000 =20 +static volatile uint64_t ipis_rcvd; + +static pthread_t vcpu_thread; + +struct hv_vpset { + u64 format; + u64 valid_bank_mask; + u64 bank_contents[2]; +}; + +enum HV_GENERIC_SET_FORMAT { + HV_GENERIC_SET_SPARSE_4K, + HV_GENERIC_SET_ALL, +}; + +struct hv_send_ipi_ex { + u32 vector; + u32 reserved; + struct hv_vpset vp_set; +}; + enum { TEST_OP_NOP, TEST_OP_READ, TEST_OP_WRITE, TEST_OP_EXEC, TEST_OP_INVPLG, + TEST_OP_HYPERV_HYPERCALL_INPUT, + TEST_OP_MONITOR_ADDRESS, TEST_OP_EXIT, }; =20 @@ -41,6 +66,8 @@ const char *test_op_names[] =3D [TEST_OP_WRITE] =3D "Write", [TEST_OP_EXEC] =3D "Exec", [TEST_OP_INVPLG] =3D "Invplg", + [TEST_OP_HYPERV_HYPERCALL_INPUT] =3D "HvHcall input", + [TEST_OP_MONITOR_ADDRESS] =3D "Monitor address", [TEST_OP_EXIT] =3D "Exit", }; =20 @@ -48,6 +75,8 @@ struct test_data { uint8_t op; int stage; gva_t vaddr; + gpa_t paddr; + uint8_t confirm_read; uint64_t expected_val; =20 struct kvm_vcpu *vcpu; @@ -59,19 +88,27 @@ static uint64_t arch_controlled_read(gva_t addr); static void arch_controlled_write(gva_t addr, uint64_t val); static void arch_controlled_exec(gva_t addr); static void arch_write_return_insn(struct kvm_vm *vm, gpa_t vaddr); +static void arch_guest_init(void); =20 static void guest_code(void *data) { struct test_data *test_data =3D data; int stage =3D 1; =20 + arch_guest_init(); + while (true) { uint64_t expected_val =3D READ_ONCE(test_data->expected_val); + bool confirm_read =3D READ_ONCE(test_data->confirm_read); gva_t vaddr =3D READ_ONCE(test_data->vaddr); + gpa_t paddr =3D READ_ONCE(test_data->paddr); + u64 val; =20 switch(READ_ONCE(test_data->op)) { case TEST_OP_READ: - (void) arch_controlled_read(vaddr); + val =3D arch_controlled_read(vaddr); + if (confirm_read) + GUEST_ASSERT_EQ(expected_val, val); GUEST_SYNC(stage++); break; case TEST_OP_WRITE: @@ -88,6 +125,25 @@ static void guest_code(void *data) :: "b" (vaddr): "memory"); GUEST_SYNC(stage++); break; + case TEST_OP_HYPERV_HYPERCALL_INPUT: { + hyperv_hypercall(HVCALL_SEND_IPI_EX, paddr, 0); + asm volatile ("sti; hlt; cli;"); + GUEST_ASSERT_EQ(ipis_rcvd, 1); + GUEST_SYNC(stage++); + break; + } + case TEST_OP_MONITOR_ADDRESS: { + uint64_t *pval =3D (uint64_t *)vaddr; + uint64_t val =3D READ_ONCE(*pval); + + while (READ_ONCE(*pval) =3D=3D val && + /* So host can force the op out of the loop */ + READ_ONCE(test_data->op) =3D=3D TEST_OP_MONITOR_ADDRESS) + asm volatile("nop"); + + GUEST_SYNC(stage++); + break; + } #endif default: goto exit; @@ -145,6 +201,7 @@ static void test_page_restricted(struct kvm_vcpu *vcpu,= int op, =20 test_data->op =3D op; test_data->vaddr =3D vaddr; + test_data->paddr =3D fault_paddr; =20 rc =3D _vcpu_run(vcpu); TEST_ASSERT(rc =3D=3D -1 && errno =3D=3D EFAULT, @@ -382,12 +439,16 @@ int main(int argc, char *argv[]) test_mem =3D vm_alloc(vm, size, KVM_UTIL_MIN_VADDR); virt_map(vcpu->vm, MMIO_GVA, MMIO_GPA, 1); test_data =3D init_test_data(vcpu); +#ifdef __x86_64__ + vcpu_set_hv_cpuid(vcpu); +#endif =20 test_input_validation(vm); test_memory_access(vcpu, test_mem, size); test_memattrs_ignore_mmio(vcpu); #ifdef __x86_64__ arch_test_memory_access_pte(vcpu, test_mem); + arch_test_side_channels(vcpu, test_mem, size); #endif test_finalize(vcpu); =20 diff --git a/tools/testing/selftests/kvm/x86/memory_attributes.c b/tools/te= sting/selftests/kvm/x86/memory_attributes.c index 3d2bcfe71884..5203d284d2d3 100644 --- a/tools/testing/selftests/kvm/x86/memory_attributes.c +++ b/tools/testing/selftests/kvm/x86/memory_attributes.c @@ -42,6 +42,11 @@ void arch_write_return_insn(struct kvm_vm *vm, gpa_t vad= dr) memset(addr_gpa2hva(vm, vaddr), 0xc3, 1); } =20 +void arch_guest_init(void) +{ + x2apic_enable(); +} + /* * This test validates that, during a page walk, if the page a PTE is plac= ed in * is read-only the accesss and dirty bits will not be written. Note There= 's a @@ -167,3 +172,214 @@ static void arch_test_memory_access_pte(struct kvm_vc= pu *vcpu, gva_t vaddr) test_memory_access_pte_ro(vcpu, vaddr); test_memory_access_sync_spte(vcpu, vaddr); } + +#define IPI_VECTOR 0xfe + +static void guest_ipi_handler_hv(struct ex_regs *regs) +{ + ipis_rcvd++; + wrmsr(HV_X64_MSR_EOI, 1); +} + +/* + * This test verifies that the Hyper-V hypercall exit handler takes memory + * attributes into account before accessing input data. It coordinates wit= h the + * guest through the 'TEST_OP_HYPERV_HYPERCALL_INPUT' operation and instru= cts + * the guest to issue two PV IPIs. The first PV IPI fails because the input + * data is held in read-protected memory. Subsequently, the memory protect= ion + * is lifted, and the second PV IPI succeeds. + */ +static void test_side_channel_hyperv_hypercall_inputs(struct kvm_vcpu *vcp= u, + gva_t vaddr, + size_t size) +{ + struct kvm_vm *vm =3D vcpu->vm; + struct hv_send_ipi_ex *ipi_ex =3D addr_gva2hva(vm, vaddr); + gpa_t paddr =3D addr_gva2gpa(vm, test_data->vaddr); + + if (!kvm_has_cap(KVM_CAP_HYPERV_SEND_IPI) || + !kvm_has_cap(KVM_CAP_HYPERV_HCALL_FAULT_EXIT)) + return; + + vm_enable_cap(vcpu->vm, KVM_CAP_HYPERV_HCALL_FAULT_EXIT, 1); + + ipis_rcvd =3D 0; + vm_install_exception_handler(vm, IPI_VECTOR, guest_ipi_handler_hv); + vcpu_set_msr(vcpu, HV_X64_MSR_GUEST_OS_ID, HYPERV_LINUX_OS_ID); + + ipi_ex->vector =3D IPI_VECTOR; + ipi_ex->vp_set.format =3D HV_GENERIC_SET_ALL; + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= O_ACCESS); + test_page_restricted(vcpu, TEST_OP_HYPERV_HYPERCALL_INPUT, vaddr, paddr, + KVM_MEMORY_EXIT_FLAG_READ); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + vcpu_run_and_inc_stage(vcpu); +} + +/* + * Verifies that the guest page table walker fails the walk if it encounte= rs a + * page table entry address read-protected by a memory attribute. + */ +static void test_side_channel_emul_page_walks(struct kvm_vcpu *vcpu, + gva_t test_vm_addr, + size_t size) +{ + const uint64_t pte_addr_mask =3D GENMASK(51, 12); + struct kvm_vm *vm =3D vcpu->vm; + struct kvm_translation tr; + int level =3D PG_LEVEL_1G; + gpa_t paddr; + uint64_t *pte; + + pte =3D vm_get_pte_level(vm, test_vm_addr, &level); + TEST_ASSERT_EQ(level, PG_LEVEL_1G); + paddr =3D *pte & pte_addr_mask; + + tr =3D vcpu_translate(vcpu, test_vm_addr); + TEST_ASSERT_EQ(tr.valid, true); + TEST_ASSERT_EQ(tr.physical_address, addr_gva2gpa(vm, test_vm_addr)); + + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= O_ACCESS); + tr =3D vcpu_translate(vcpu, test_vm_addr); + TEST_ASSERT_EQ(tr.valid, false); + + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); +} + +static void vm_set_vapic_addr(struct kvm_vcpu *vcpu, uint64_t addr) +{ + struct kvm_vapic_addr va; + + va.vapic_addr =3D addr; + vcpu_ioctl(vcpu, KVM_SET_VAPIC_ADDR, &va); +} + +/* + * Perform a dummy regs update to issue a KVM_REQ_EVENT. This forces + * vapic to be synced before entering the guest. + */ +static void vcpu_force_vapic_update(struct kvm_vcpu *vcpu) +{ + struct kvm_regs regs; + + vcpu_regs_get(vcpu, ®s); + vcpu_regs_set(vcpu, ®s); +} + +/* + * Setup the vapic address on a GPA that is write-protected. Force an vapic + * update and validate its contents were not changes. Then, lift the write + * restriction and validate the page's contents are updated. + */ +static void test_side_channel_vapic_addr(struct kvm_vcpu *vcpu) +{ + struct kvm_vm *vm =3D vcpu->vm; + gva_t vaddr =3D vm_alloc_page(vm); + gpa_t paddr =3D addr_gva2gpa(vm, vaddr); + + vm_set_vapic_addr(vcpu, paddr); + test_data->op =3D TEST_OP_READ; + test_data->vaddr =3D vaddr; + test_data->confirm_read =3D 1; + test_data->expected_val =3D ~0ULL >> 32; + memset(addr_gva2hva(vm, vaddr), 0xff, sizeof(uint32_t)); + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= W); + vcpu_run_and_inc_stage(vcpu); + + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + vcpu_force_vapic_update(vcpu); + test_data->expected_val =3D 0ULL; + vcpu_run_and_inc_stage(vcpu); + + vm_set_vapic_addr(vcpu, 0); + test_data->confirm_read =3D 0; +} + +static void *vcpu_worker(void *data) +{ + struct test_data *test_data =3D data; + struct kvm_vcpu *vcpu =3D test_data->vcpu; + + vcpu_run_and_inc_stage(vcpu); + + return NULL; +} + +/* + * Set up the pvclock page and validate that KVM periodically updates the + * 'tsc_timestamp' field. Subsequently, make the pvclock page non-writable= and + * verify that the 'tsc_timestamp' field is no longer updated. + */ +static void test_side_channel_pvclock(struct kvm_vcpu *vcpu) +{ + struct kvm_vm *vm =3D vcpu->vm; + gva_t vaddr =3D vm_alloc_page(vm); + gpa_t paddr =3D addr_gva2gpa(vm, vaddr); + vcpu_set_msr(vcpu, MSR_KVM_SYSTEM_TIME_NEW, paddr | 0x1); + + test_data->op =3D TEST_OP_MONITOR_ADDRESS; + test_data->vaddr =3D vaddr + offsetof(struct pvclock_vcpu_time_info, tsc_= timestamp); + + pthread_create(&vcpu_thread, NULL, vcpu_worker, test_data); + usleep(msecs_to_usecs(1000)); + TEST_ASSERT_EQ(pthread_tryjoin_np(vcpu_thread, NULL), 0); + + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= W); + pthread_create(&vcpu_thread, NULL, vcpu_worker, test_data); + usleep(msecs_to_usecs(1000)); + TEST_ASSERT_EQ(pthread_tryjoin_np(vcpu_thread, NULL), EBUSY); + + /* Force the 'monitor_address' guest operation to finish */ + test_data->op =3D TEST_OP_NOP; + TEST_ASSERT_EQ(pthread_join(vcpu_thread, NULL), 0); + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); + vcpu_set_msr(vcpu, MSR_KVM_SYSTEM_TIME_NEW, 0); +} + +/* + * Write to MSR_KVM_WALL_CLOCK_NEW and verify that the struct's 'version' = field + * is updated. Subsequently, make the target guest physical address + * non-writable, and verify the 'version' field isn't updated anymore. + */ +static void test_side_channel_wallclock(struct kvm_vcpu *vcpu) +{ + struct kvm_vm *vm =3D vcpu->vm; + gva_t vaddr =3D vm_alloc_page(vm); + gpa_t paddr =3D addr_gva2gpa(vm, vaddr); + struct pvclock_wall_clock *wc =3D addr_gva2hva(vm, vaddr); + + wc->version =3D 0x0; + vcpu_set_msr(vcpu, MSR_KVM_WALL_CLOCK_NEW, paddr); + TEST_ASSERT_EQ(wc->version, 2); + + vm_set_memory_attributes(vm, paddr, vm->page_size, KVM_MEMORY_ATTRIBUTE_N= W); + vcpu_set_msr(vcpu, MSR_KVM_WALL_CLOCK_NEW, paddr); + TEST_ASSERT_EQ(wc->version, 2); + + vm_set_memory_attributes(vm, paddr, vm->page_size, 0); +} + +/* + * Memory attributes are vulnerable to side-channel attacks. This means th= at + * any KVM operation initiated by the guest that requires guest memory acc= ess + * (which is the case for most pv-interfaces) needs to consider memory + * attributes. + * + * The following tests validate that this requirement is upheld for a vari= ety + * of use-cases. These test cases were selected to exercise specific appro= aches + * to accessing guest memory, including: + * + * - kvm_read/write_guest() + * - gfn_to_hva_cache + * - gfn_to_pfn_cache + * - Guest page walker + */ +static void arch_test_side_channels(struct kvm_vcpu *vcpu, gva_t test_vm_a= ddr, + size_t size) +{ + test_side_channel_hyperv_hypercall_inputs(vcpu, test_vm_addr, size); + test_side_channel_emul_page_walks(vcpu, test_vm_addr, size); + test_side_channel_vapic_addr(vcpu); + test_side_channel_wallclock(vcpu); + test_side_channel_pvclock(vcpu); +} --=20 2.52.0