From nobody Thu Sep 24 01:08:34 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=linux.microsoft.com ARC-Seal: i=1; a=rsa-sha256; t=1785927932; cv=none; d=zohomail.com; s=zohoarc; b=VOWIaFRB3pYsK0wZUCbkMmPZgGj6t9t8KZskLwss45mOKVcYN7iVMUxuAvgfWKtY4a3ypZoOvnIfRYItzZ70mA7egLfWCNMGnybwH1fkeUBtmw5qyALWA/grihI+Ydh1+oS/SklotueRFkg8cDw8iU6An9E1iroxZkRTPTBvlfY= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785927932; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=cKpLSPc5Ft5FlEL2TPYsuyNO1mO7zlGBjKjNjZCumb0=; b=YoV7ABm2Y5Zl9B5O3k2/I9MnB42LPXDwX8UAPrt0NpLVgVPvYSjFSlIVUL7jAgC1osxa8v/H7Bwgs03yW86Osy3QyBQ6xlhtZBobdf6JyDdBW/ODyIC24PYtNWUI3GKca9rcGKa6+e3sqriB8lIyb+3rHUEAm1gIMUW9GYuOTgs= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785927932691299.201444274354; Wed, 5 Aug 2026 04:05:32 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wrZQR-0000mJ-7i; Wed, 05 Aug 2026 07:04:43 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wrZQP-0000kO-E6 for qemu-devel@nongnu.org; Wed, 05 Aug 2026 07:04:41 -0400 Received: from linux.microsoft.com ([13.77.154.182]) by eggs.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wrZQL-0000BO-5y for qemu-devel@nongnu.org; Wed, 05 Aug 2026 07:04:41 -0400 Received: from fedora.hsd1.wa.comcast.net (unknown [52.148.140.42]) by linux.microsoft.com (Postfix) with ESMTPSA id 8681D20B716A; Wed, 5 Aug 2026 04:04:13 -0700 (PDT) DKIM-Filter: OpenDKIM Filter v2.11.0 linux.microsoft.com 8681D20B716A DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.microsoft.com; s=default; t=1785927853; bh=cKpLSPc5Ft5FlEL2TPYsuyNO1mO7zlGBjKjNjZCumb0=; h=From:To:Cc:Subject:Date:In-Reply-To:References:From; b=jeJe8BEZS6SaED7DkDupduivx9FuLihWlULQySkpXZdklJW1abcHTC3S57j4UNwDV X8nGxMlcTK+8ySS9WnKbVSuZ5Xj8LVBfBlRhjoQgnu4IGJ2wYl9U+rViLbgRuyUhEf HXZEnN/Q0znekiHMYYzYeBASwBuyJ6l9Y/xT+M/Q= From: Sriram Nambakam To: qemu-devel@nongnu.org Cc: kvm@vger.kernel.org Subject: [RFC PATCH v1 1/5] kvm: add userspace handlers for VM planes and VBS VTL calls Date: Wed, 5 Aug 2026 04:04:28 -0700 Message-ID: <20260805110432.25167-2-snambakam@linux.microsoft.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260805110432.25167-1-snambakam@linux.microsoft.com> References: <20260805110432.25167-1-snambakam@linux.microsoft.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=13.77.154.182; envelope-from=snambakam@linux.microsoft.com; helo=linux.microsoft.com X-Spam_score_int: -19 X-Spam_score: -2.0 X-Spam_bar: -- X-Spam_report: (-2.0 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @linux.microsoft.com) X-ZM-MESSAGEID: 1785927934179158500 (cherry picked from commit 47d0da9ac50640f96564997edc4b79f44e70cffa) --- accel/kvm/kvm-all.c | 18 + include/standard-headers/linux/kvm_para.h | 3 + include/system/kvm_int.h | 17 + linux-headers/linux/kvm.h | 3 + target/i386/kvm/kvm.c | 776 ++++++++++++++++++++++ 5 files changed, 817 insertions(+) diff --git a/accel/kvm/kvm-all.c b/accel/kvm/kvm-all.c index 46a14ac0f4..80a9791078 100644 --- a/accel/kvm/kvm-all.c +++ b/accel/kvm/kvm-all.c @@ -822,6 +822,24 @@ void kvm_close(void) close(kvm_get_plane_fd(kvm_state, plane_id)); kvm_set_plane_fd(kvm_state, plane_id, -1); } while (plane_id !=3D 0); + if (kvm_state->vm_planes) { + unsigned int i, j; + for (i =3D 1; i < kvm_state->vm_plane_count; i++) { + struct kvm_vm_plane_state *ps =3D &kvm_state->vm_planes[i]; + if (ps->vcpu_fds) { + for (j =3D 0; j < ps->vcpu_count; j++) { + if (ps->vcpu_fds[j] >=3D 0) { + close(ps->vcpu_fds[j]); + } + } + g_free(ps->vcpu_fds); + ps->vcpu_fds =3D NULL; + } + } + g_free(kvm_state->vm_planes); + kvm_state->vm_planes =3D NULL; + kvm_state->vm_plane_count =3D 0; + } close(kvm_state->fd); kvm_state->fd =3D -1; } diff --git a/include/standard-headers/linux/kvm_para.h b/include/standard-h= eaders/linux/kvm_para.h index 015c166302..ad19aac9d0 100644 --- a/include/standard-headers/linux/kvm_para.h +++ b/include/standard-headers/linux/kvm_para.h @@ -30,6 +30,9 @@ #define KVM_HC_SEND_IPI 10 #define KVM_HC_SCHED_YIELD 11 #define KVM_HC_MAP_GPA_RANGE 12 +#define KVM_HC_VM_PLANES_CONFIG 13 +#define KVM_HC_VM_PLANES_ACTIVATE 14 +#define KVM_HC_VBS_VTL_CALL 15 =20 /* * hypercalls use architecture specific diff --git a/include/system/kvm_int.h b/include/system/kvm_int.h index 70b381f1ba..e7c9dc95a2 100644 --- a/include/system/kvm_int.h +++ b/include/system/kvm_int.h @@ -109,6 +109,19 @@ struct KVMPlane { bool vcpu_dirty; }; =20 +/* Per-plane VM state managed by LVBS hypercall handlers. + * The plane fd itself is owned by the accel layer and accessed via + * kvm_get_plane_fd(s, plane_id); only LVBS-specific state lives here. */ +struct kvm_vm_plane_state { + int *vcpu_fds; + unsigned int vcpu_count; + uint64_t load_offset; + uint64_t memory_size; + uint64_t entry_point; + void *host_addr; /* host pointer to plane RAM (cleared after la= unch) */ + char cmdline[512]; +}; + struct KVMState { AccelState parent_obj; @@ -176,6 +189,10 @@ struct KVMState uint16_t xen_evtchn_max_pirq; char *device; OnOffAuto honor_guest_pat; + /* VM planes state (populated by LVBS hypercall handlers) */ + struct kvm_vm_plane_state *vm_planes; + unsigned int vm_plane_count; + unsigned int vm_planes_max; }; =20 static inline void kvm_set_plane_fd(KVMState *s, unsigned plane, int fd) diff --git a/linux-headers/linux/kvm.h b/linux-headers/linux/kvm.h index 8caa3eccce..d66ce6d272 100644 --- a/linux-headers/linux/kvm.h +++ b/linux-headers/linux/kvm.h @@ -1654,6 +1654,9 @@ struct kvm_memory_attributes { =20 #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3) =20 +#define KVM_MEMORY_ATTRIBUTE_NO_WRITE (1ULL << 4) +#define KVM_MEMORY_ATTRIBUTE_NO_EXEC (1ULL << 5) + #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest= _memfd) #define GUEST_MEMFD_FLAG_MMAP (1ULL << 0) #define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1) diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c index 13cfa60071..9d89e5cac6 100644 --- a/target/i386/kvm/kvm.c +++ b/target/i386/kvm/kvm.c @@ -71,6 +71,13 @@ #include "exec/memattrs.h" #include "exec/target_page.h" #include "trace.h" +#include "system/address-spaces.h" +#include "system/memory.h" +#include "qemu/thread.h" +#include "qemu/timer.h" +#include +#include +#include =20 #include CONFIG_DEVICES =20 @@ -3588,6 +3595,13 @@ int kvm_arch_init(MachineState *ms, KVMState *s) kvm_vmfd_add_change_notifier(&kvm_vmfd_change_notifier); } =20 + /* Enable userspace exit for VM planes and VBS hypercalls (LVBS). */ + if (!kvm_enable_hypercall(BIT_ULL(KVM_HC_VM_PLANES_CONFIG) | + BIT_ULL(KVM_HC_VM_PLANES_ACTIVATE) | + BIT_ULL(KVM_HC_VBS_VTL_CALL))) { + warn_report("kvm: failed to enable VM planes / VBS hypercall exit"= ); + } + /* * Most x86 CPUs in current use have self-snoop, so honoring guest PAT= is * preferable. As well, the bochs video driver bug which motivated ma= king @@ -6519,10 +6533,772 @@ static int kvm_handle_hc_map_gpa_range(X86CPU *cpu= , struct kvm_run *run) return 0; } =20 +/* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * LVBS =E2=80=94 VM planes + VBS VTL hypercall handlers + * + * Guest-side struct vm_plane_config layout (672 bytes, see Linux + * include/linux/vm_planes.h): + * off 0 u64 load_offset + * off 8 u64 memory_size + * off 16 u64 entry_point + * off 24 u32 vcpu_count + * off 28 u32 kernel_format + * off 32 char kernel[128] + * off 160 char cmdline[512] + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D */ + +#define VM_PLANE_CFG_STRIDE 672 + +#define VBS_CALL_INIT 0x0001 +#define VBS_CALL_SHUTDOWN 0x0002 +#define VBS_CALL_PROTECT_MEMORY 0x0100 +#define VBS_CALL_SEAL_KERNEL 0x0101 +#define VBS_CALL_VALIDATE_MODULE 0x0200 +#define VBS_CALL_SET_MODULE_PERMS 0x0201 +#define VBS_CALL_UNLOAD_MODULE 0x0202 +#define VBS_CALL_ADD_KEY 0x0300 +#define VBS_CALL_REVOKE_KEY 0x0301 +#define VBS_CALL_SEND_CERTS 0x0302 +#define VBS_CALL_KEXEC_VALIDATE 0x0400 +#define VBS_CALL_KEXEC_INVALIDATE 0x0401 + +#define CA_OFF_CALL_ID 4 +#define CA_OFF_STATUS 8 +#define CA_OFF_ARG_SIZE 12 +#define CA_OFF_RESP_SIZE 16 +#define CA_OFF_BUFFER 20 + +struct vm_plane_boot_ctx { + int vcpu_fd; + int vcpu_mmap_size; + unsigned int vcpu_idx; + uint64_t plane_id; + unsigned int *halted_count; + QemuMutex *mutex; + QemuCond *cond; + int result; + int log_fd; + QemuMutex wake_mutex; + QemuCond wake_cond; + bool halted; + bool kick; + bool stopped; +}; + +static int kvm_handle_hc_vm_planes_config(X86CPU *cpu, struct kvm_run *run) +{ + uint64_t gpa =3D run->hypercall.args[0]; + uint64_t plane_count =3D run->hypercall.args[1]; + uint64_t plane_id; + unsigned int plane0_vcpu_count; + unsigned int *plane0_vcpu_ids; + CPUState *cs; + KVMState *s =3D kvm_state; + + if (!gpa || !plane_count) { + run->hypercall.ret =3D -EINVAL; + return 0; + } + + if (!s->vm_planes_max) { + int max =3D kvm_vm_ioctl(s, KVM_CHECK_EXTENSION, KVM_CAP_PLANES); + if (max <=3D 0) { + error_report("vm_planes: KVM does not support planes"); + run->hypercall.ret =3D -ENOTSUP; + return 0; + } + s->vm_planes_max =3D max; + } + + if (plane_count > s->vm_planes_max) { + error_report("vm_planes: requested %" PRIu64 " planes but KVM " + "supports %u", plane_count, s->vm_planes_max); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + if (s->vm_planes) { + info_report("vm_planes: already configured, ignoring"); + run->hypercall.ret =3D 0; + return 0; + } + + s->vm_planes =3D g_new0(struct kvm_vm_plane_state, plane_count); + s->vm_plane_count =3D plane_count; + + plane0_vcpu_count =3D 0; + CPU_FOREACH(cs) { + plane0_vcpu_count++; + } + if (!plane0_vcpu_count) { + error_report("vm_planes: no plane0 vCPUs available"); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + plane0_vcpu_ids =3D g_new(unsigned int, plane0_vcpu_count); + plane0_vcpu_count =3D 0; + CPU_FOREACH(cs) { + plane0_vcpu_ids[plane0_vcpu_count++] =3D cs->cpu_index; + } + + for (plane_id =3D 1; plane_id < plane_count; plane_id++) { + uint64_t plane_gpa =3D gpa + (plane_id * VM_PLANE_CFG_STRIDE); + uint64_t load_offset =3D 0, memory_size =3D 0, entry_point =3D 0; + uint32_t vcpu_count =3D 0; + struct kvm_vm_plane_state *ps =3D &s->vm_planes[plane_id]; + char cmdline_buf[512]; + MemoryRegionSection section; + int plane_fd; + unsigned int i; + + cpu_physical_memory_read(plane_gpa + 0, &load_offset, 8); + cpu_physical_memory_read(plane_gpa + 8, &memory_size, 8); + cpu_physical_memory_read(plane_gpa + 16, &entry_point, 8); + cpu_physical_memory_read(plane_gpa + 24, &vcpu_count, 4); + + if (!memory_size || !vcpu_count) { + error_report("vm_planes: plane %" PRIu64 " invalid " + "(load_offset=3D0x%" PRIx64 " size=3D0x%" PRIx64 + " vcpus=3D%u)", + plane_id, load_offset, memory_size, vcpu_count); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + if (vcpu_count > plane0_vcpu_count) { + error_report("vm_planes: plane %" PRIu64 " requests %u vCPUs, " + "but plane0 has %u", plane_id, vcpu_count, + plane0_vcpu_count); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + memset(cmdline_buf, 0, sizeof(cmdline_buf)); + cpu_physical_memory_read(plane_gpa + 160, cmdline_buf, + sizeof(cmdline_buf)); + cmdline_buf[sizeof(cmdline_buf) - 1] =3D '\0'; + memcpy(ps->cmdline, cmdline_buf, sizeof(ps->cmdline)); + + plane_fd =3D kvm_vm_ioctl(s, KVM_CREATE_PLANE, (int)plane_id); + if (plane_fd < 0) { + error_report("vm_planes: KVM_CREATE_PLANE plane %" PRIu64 + " failed: %s", plane_id, strerror(errno)); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -errno; + return 0; + } + /* The kvm_vm_plane_ioctl path looks up the plane fd via + * kvm_get_plane_fd / kvm_set_plane_fd, so register the fd. */ + kvm_set_plane_fd(s, plane_id, plane_fd); + + section =3D memory_region_find(get_system_memory(), + load_offset, memory_size); + if (!section.mr || !memory_region_is_ram(section.mr)) { + error_report("vm_planes: plane %" PRIu64 " GPA 0x%" PRIx64 + " size 0x%" PRIx64 " is not RAM", + plane_id, load_offset, memory_size); + if (section.mr) { + memory_region_unref(section.mr); + } + close(plane_fd); + kvm_set_plane_fd(s, plane_id, -1); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -ENOMEM; + return 0; + } + ps->host_addr =3D memory_region_get_ram_ptr(section.mr) + + section.offset_within_region; + memory_region_unref(section.mr); + + ps->vcpu_fds =3D g_new0(int, vcpu_count); + for (i =3D 0; i < vcpu_count; i++) { + int vcpu_fd; + unsigned int vcpu_id =3D plane0_vcpu_ids[i]; + + vcpu_fd =3D kvm_vm_plane_ioctl(s, plane_id, KVM_CREATE_VCPU, + (void *)(uintptr_t)vcpu_id); + if (vcpu_fd < 0) { + error_report("vm_planes: KVM_CREATE_VCPU plane %" PRIu64 + " vcpu %u failed: %s", + plane_id, vcpu_id, strerror(errno)); + close(plane_fd); + kvm_set_plane_fd(s, plane_id, -1); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -errno; + return 0; + } + + ps->vcpu_fds[i] =3D vcpu_fd; + } + + ps->vcpu_count =3D vcpu_count; + ps->load_offset =3D load_offset; + ps->memory_size =3D memory_size; + ps->entry_point =3D entry_point; + + info_report("vm_planes: plane %" PRIu64 " ready =E2=80=94 GPA 0x%"= PRIx64 + " size 0x%" PRIx64 " entry 0x%" PRIx64 " vcpus %u", + plane_id, load_offset, memory_size, entry_point, + vcpu_count); + } + + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D 0; + return 0; +} + +static int kvm_init_plane_vcpu(int vcpu_fd, uint64_t entry_addr, + uint64_t stack_addr, uint64_t zero_page_gpa, + uint64_t page_table_gpa, uint64_t gdt_gpa, + bool is_bsp) +{ + struct kvm_regs regs =3D {}; + struct kvm_sregs sregs =3D {}; + struct kvm_mp_state mp =3D { + .mp_state =3D is_bsp ? KVM_MP_STATE_RUNNABLE + : KVM_MP_STATE_INIT_RECEIVED, + }; + int ret; + + sregs.cs.base =3D 0; sregs.cs.limit =3D 0xffffffff; sregs.cs.selector = =3D 0x10; + sregs.cs.type =3D 0xb; sregs.cs.present =3D 1; sregs.cs.dpl =3D 0; + sregs.cs.db =3D 0; sregs.cs.s =3D 1; sregs.cs.l =3D 1; sregs.cs.g =3D = 1; + + sregs.ds.base =3D 0; sregs.ds.limit =3D 0xffffffff; sregs.ds.selector = =3D 0x18; + sregs.ds.type =3D 0x3; sregs.ds.present =3D 1; sregs.ds.dpl =3D 0; + sregs.ds.db =3D 1; sregs.ds.s =3D 1; sregs.ds.g =3D 1; + sregs.es =3D sregs.ds; + sregs.ss =3D sregs.ds; + sregs.fs =3D sregs.ds; sregs.fs.selector =3D 0; + sregs.gs =3D sregs.fs; + + sregs.gdt.base =3D gdt_gpa; sregs.gdt.limit =3D 0x2f; + sregs.idt.base =3D 0; sregs.idt.limit =3D 0xffff; + sregs.tr.base =3D 0; sregs.tr.limit =3D 0x67; sregs.tr.selector =3D 0x= 28; + sregs.tr.type =3D 0xb; sregs.tr.present =3D 1; sregs.tr.dpl =3D 0; sre= gs.tr.s =3D 0; + sregs.ldt.unusable =3D 1; + + sregs.cr3 =3D page_table_gpa; + sregs.cr4 =3D (1u << 5); /* PAE */ + sregs.cr0 =3D (1u << 0) | (1u << 4) | (1u << 5) | (1u << 16) | (1u << = 31); + sregs.efer =3D (1u << 0) | (1u << 8) | (1u << 10) | (1u << 11); + sregs.apic_base =3D 0xfee00000 | (1u << 11); + if (is_bsp) { + sregs.apic_base |=3D (1u << 8); + } + + ret =3D ioctl(vcpu_fd, KVM_SET_SREGS, &sregs); + if (ret < 0) { + error_report("vm_planes: KVM_SET_SREGS: %s", strerror(errno)); + return -errno; + } + + regs.rip =3D entry_addr; + regs.rsp =3D stack_addr; + regs.rsi =3D zero_page_gpa; + regs.rflags =3D 0x2; + ret =3D ioctl(vcpu_fd, KVM_SET_REGS, ®s); + if (ret < 0) { + error_report("vm_planes: KVM_SET_REGS: %s", strerror(errno)); + return -errno; + } + + ret =3D ioctl(vcpu_fd, KVM_SET_MP_STATE, &mp); + if (ret < 0) { + error_report("vm_planes: KVM_SET_MP_STATE: %s", strerror(errno)); + return -errno; + } + return 0; +} + +static void *vm_plane_vcpu_thread(void *arg) +{ + struct vm_plane_boot_ctx *ctx =3D arg; + struct kvm_run *kvm_run; + int ret; + bool boot_signaled =3D false; + + kvm_run =3D mmap(NULL, ctx->vcpu_mmap_size, PROT_READ | PROT_WRITE, + MAP_SHARED, ctx->vcpu_fd, 0); + if (kvm_run =3D=3D MAP_FAILED) { + error_report("vm_planes: plane %" PRIu64 " vcpu %u: mmap failed: %= s", + ctx->plane_id, ctx->vcpu_idx, strerror(errno)); + ctx->result =3D -errno; + qemu_mutex_lock(ctx->mutex); + (*ctx->halted_count)++; + qemu_cond_signal(ctx->cond); + qemu_mutex_unlock(ctx->mutex); + return NULL; + } + + for (;;) { + ret =3D ioctl(ctx->vcpu_fd, KVM_RUN, 0); + if (ret < 0) { + if (errno =3D=3D EINTR || errno =3D=3D EAGAIN) { + continue; + } + error_report("vm_planes: plane %" PRIu64 " vcpu %u: KVM_RUN: %= s", + ctx->plane_id, ctx->vcpu_idx, strerror(errno)); + ctx->result =3D -errno; + break; + } + + switch (kvm_run->exit_reason) { + case KVM_EXIT_HLT: + if (!boot_signaled) { + boot_signaled =3D true; + qemu_mutex_lock(ctx->mutex); + (*ctx->halted_count)++; + qemu_cond_signal(ctx->cond); + qemu_mutex_unlock(ctx->mutex); + } + qemu_mutex_lock(&ctx->wake_mutex); + ctx->halted =3D true; + while (!ctx->kick && !ctx->stopped) { + qemu_cond_wait(&ctx->wake_cond, &ctx->wake_mutex); + } + ctx->halted =3D false; + ctx->kick =3D false; + if (ctx->stopped) { + qemu_mutex_unlock(&ctx->wake_mutex); + goto done; + } + qemu_mutex_unlock(&ctx->wake_mutex); + break; + + case KVM_EXIT_IO: { + uint8_t *io_data =3D (uint8_t *)kvm_run + kvm_run->io.data_off= set; + size_t io_size =3D kvm_run->io.size * kvm_run->io.count; + uint16_t port =3D kvm_run->io.port; + + if (kvm_run->io.direction =3D=3D KVM_EXIT_IO_OUT) { + if (port =3D=3D 0x3f8 && ctx->log_fd >=3D 0) { + ssize_t w =3D write(ctx->log_fd, io_data, io_size); + (void)w; + } + } else { + memset(io_data, 0, io_size); + switch (port) { + case 0x3fa: memset(io_data, 0xc1, io_size); break; + case 0x3fb: memset(io_data, 0x03, io_size); break; + case 0x3fc: memset(io_data, 0x08, io_size); break; + case 0x3fd: memset(io_data, 0x60, io_size); break; + case 0x3fe: memset(io_data, 0xb0, io_size); break; + default: break; + } + } + usleep(100); + break; + } + + case KVM_EXIT_MMIO: + usleep(100); + break; + + case KVM_EXIT_SHUTDOWN: { + struct kvm_regs dbg =3D {}; + ioctl(ctx->vcpu_fd, KVM_GET_REGS, &dbg); + error_report("vm_planes: plane %" PRIu64 " vcpu %u: shutdown " + "RIP=3D0x%" PRIx64 " RSP=3D0x%" PRIx64, + ctx->plane_id, ctx->vcpu_idx, + (uint64_t)dbg.rip, (uint64_t)dbg.rsp); + ctx->result =3D -EFAULT; + goto done; + } + + case KVM_EXIT_FAIL_ENTRY: + error_report("vm_planes: plane %" PRIu64 " vcpu %u: entry fail= ure " + "0x%" PRIx64, ctx->plane_id, ctx->vcpu_idx, + (uint64_t)kvm_run->fail_entry.hardware_entry_fail= ure_reason); + ctx->result =3D -EFAULT; + goto done; + + case KVM_EXIT_INTERNAL_ERROR: + error_report("vm_planes: plane %" PRIu64 " vcpu %u: internal " + "error %u", ctx->plane_id, ctx->vcpu_idx, + kvm_run->internal.suberror); + ctx->result =3D -EFAULT; + goto done; + + default: + error_report("vm_planes: plane %" PRIu64 " vcpu %u: unexpected= " + "exit %u", ctx->plane_id, ctx->vcpu_idx, + kvm_run->exit_reason); + ctx->result =3D -EFAULT; + goto done; + } + } + +done: + if (ctx->log_fd >=3D 0 && ctx->vcpu_idx =3D=3D 0) { + close(ctx->log_fd); + } + munmap(kvm_run, ctx->vcpu_mmap_size); + qemu_mutex_lock(ctx->mutex); + if (!boot_signaled) { + (*ctx->halted_count)++; + qemu_cond_signal(ctx->cond); + } + qemu_mutex_unlock(ctx->mutex); + return NULL; +} + +static int kvm_handle_hc_vm_planes_activate(X86CPU *cpu, struct kvm_run *r= un) +{ + uint64_t gpa =3D run->hypercall.args[0]; + uint64_t plane_count =3D run->hypercall.args[1]; + uint64_t plane_id; + KVMState *s =3D kvm_state; + int vcpu_mmap_size; + + if (!gpa || !plane_count || !s->vm_planes || + plane_count !=3D s->vm_plane_count) { + run->hypercall.ret =3D -EINVAL; + return 0; + } + + vcpu_mmap_size =3D kvm_ioctl(s, KVM_GET_VCPU_MMAP_SIZE, 0); + if (vcpu_mmap_size <=3D 0) { + error_report("vm_planes: KVM_GET_VCPU_MMAP_SIZE failed"); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + for (plane_id =3D 1; plane_id < plane_count; plane_id++) { + struct kvm_vm_plane_state *ps =3D &s->vm_planes[plane_id]; + uint64_t stack_addr; + uint64_t entry_point =3D 0; + uint64_t cmdline_gpa, zero_page_gpa; + uint64_t pt_base, pml4_gpa, pdpt_gpa, pd_base, gdt_gpa_val; + struct vm_plane_boot_ctx *ctxs; + QemuThread *threads; + QemuMutex mutex; + QemuCond cond; + unsigned int halted_count =3D 0; + unsigned int i; + int plane_log_fd =3D -1; + + if (kvm_get_plane_fd(s, plane_id) < 0 || !ps->vcpu_count || + !ps->host_addr) { + error_report("vm_planes: plane %" PRIu64 " not configured", + plane_id); + run->hypercall.ret =3D -EINVAL; + return 0; + } + + cpu_physical_memory_read(gpa + (plane_id * VM_PLANE_CFG_STRIDE) + = 16, + &entry_point, 8); + if (!entry_point) { + error_report("vm_planes: plane %" PRIu64 " bad entry_point", + plane_id); + run->hypercall.ret =3D -EIO; + return 0; + } + ps->entry_point =3D entry_point; + + stack_addr =3D ps->load_offset + ps->memory_size; + cmdline_gpa =3D stack_addr - 0x1000; + zero_page_gpa =3D stack_addr - 0x2000; + pt_base =3D ps->load_offset + ps->memory_size - 0x10000; + pml4_gpa =3D pt_base; + pdpt_gpa =3D pt_base + 0x1000; + pd_base =3D pt_base + 0x2000; + gdt_gpa_val =3D pt_base + 0x6000; + +#define PLANE_HOST(g) ((uint8_t *)ps->host_addr + ((g) - ps->load_offset)) + + /* cmdline */ + { + size_t cl =3D strlen(ps->cmdline) + 1; + memcpy(PLANE_HOST(cmdline_gpa), ps->cmdline, cl); + } + + /* boot_params zero page */ + { + uint8_t zp[4096] =3D {}; + uint32_t cl_ptr =3D (uint32_t)(cmdline_gpa & 0xffffffff); + uint32_t cl_hi =3D (uint32_t)(cmdline_gpa >> 32); + struct { + uint64_t addr; + uint64_t size; + uint32_t type; + } QEMU_PACKED e820 =3D { + ps->load_offset, ps->memory_size, 1, + }; + + zp[0x1fe] =3D 0x55; zp[0x1ff] =3D 0xAA; + zp[0x202] =3D 'H'; zp[0x203] =3D 'd'; + zp[0x204] =3D 'r'; zp[0x205] =3D 'S'; + zp[0x206] =3D 0x0f; zp[0x207] =3D 0x02; + zp[0x210] =3D 0xff; + memcpy(&zp[0x228], &cl_ptr, 4); + memcpy(&zp[0x0c8], &cl_hi, 4); + zp[0x1e8] =3D 1; + memcpy(&zp[0x2d0], &e820, 20); + memcpy(PLANE_HOST(zero_page_gpa), zp, sizeof(zp)); + } + + /* Identity-mapped page tables (PML4 =E2=86=92 PDPT =E2=86=92 4=C3= =97PD with 2MB pages) */ + { + uint8_t page[4096]; + uint64_t *entries; + int pd_idx; + + memset(page, 0, sizeof(page)); + entries =3D (uint64_t *)page; + entries[0] =3D pdpt_gpa | 0x3; + memcpy(PLANE_HOST(pml4_gpa), page, 4096); + + memset(page, 0, sizeof(page)); + entries =3D (uint64_t *)page; + for (pd_idx =3D 0; pd_idx < 4; pd_idx++) { + entries[pd_idx] =3D (pd_base + pd_idx * 0x1000) | 0x3; + } + memcpy(PLANE_HOST(pdpt_gpa), page, 4096); + + for (pd_idx =3D 0; pd_idx < 4; pd_idx++) { + int j; + memset(page, 0, sizeof(page)); + entries =3D (uint64_t *)page; + for (j =3D 0; j < 512; j++) { + uint64_t phys =3D ((uint64_t)pd_idx << 30) | + ((uint64_t)j << 21); + entries[j] =3D phys | 0x83; + } + memcpy(PLANE_HOST(pd_base + pd_idx * 0x1000), page, 4096); + } + } + + /* Minimal GDT */ + { + uint8_t gdt[48] =3D {}; + uint64_t *gdt64 =3D (uint64_t *)gdt; + + gdt64[0] =3D 0; + gdt64[1] =3D 0; + gdt64[2] =3D 0x00af9a000000ffffULL; + gdt64[3] =3D 0x00cf92000000ffffULL; + gdt64[4] =3D 0; + gdt64[5] =3D 0x0000890000000067ULL; + memcpy(PLANE_HOST(gdt_gpa_val), gdt, sizeof(gdt)); + } + + /* Initialize all plane vCPUs */ + for (i =3D 0; i < ps->vcpu_count; i++) { + int ret =3D kvm_init_plane_vcpu(ps->vcpu_fds[i], entry_point, + stack_addr, zero_page_gpa, pml4_= gpa, + gdt_gpa_val, i =3D=3D 0); + if (ret) { + error_report("vm_planes: init plane %" PRIu64 " vcpu %u " + "failed", plane_id, i); + run->hypercall.ret =3D ret; + return 0; + } + } +#undef PLANE_HOST + + /* Serial log */ + { + char lp[256]; + snprintf(lp, sizeof(lp), "/tmp/plane%" PRIu64 "-serial.log", + plane_id); + plane_log_fd =3D open(lp, O_CREAT | O_WRONLY | O_TRUNC, 0644); + } + + /* Spawn vCPU threads */ + qemu_mutex_init(&mutex); + qemu_cond_init(&cond); + + ctxs =3D g_new0(struct vm_plane_boot_ctx, ps->vcpu_count); + threads =3D g_new0(QemuThread, ps->vcpu_count); + + for (i =3D 0; i < ps->vcpu_count; i++) { + char name[48]; + + ctxs[i].vcpu_fd =3D ps->vcpu_fds[i]; + ctxs[i].vcpu_mmap_size =3D vcpu_mmap_size; + ctxs[i].vcpu_idx =3D i; + ctxs[i].plane_id =3D plane_id; + ctxs[i].halted_count =3D &halted_count; + ctxs[i].mutex =3D &mutex; + ctxs[i].cond =3D &cond; + ctxs[i].result =3D 0; + ctxs[i].log_fd =3D (i =3D=3D 0) ? plane_log_fd : -1; + qemu_mutex_init(&ctxs[i].wake_mutex); + qemu_cond_init(&ctxs[i].wake_cond); + ctxs[i].halted =3D false; + ctxs[i].kick =3D false; + ctxs[i].stopped =3D false; + + snprintf(name, sizeof(name), "plane%" PRIu64 "-vcpu%u", + plane_id, i); + qemu_thread_create(&threads[i], name, vm_plane_vcpu_thread, + &ctxs[i], QEMU_THREAD_JOINABLE); + } + + /* Wait up to 5s for plane to reach first HLT */ + { + int64_t dl =3D qemu_clock_get_ns(QEMU_CLOCK_REALTIME) + + 5LL * 1000000000LL; + qemu_mutex_lock(&mutex); + while (halted_count < ps->vcpu_count) { + int64_t now =3D qemu_clock_get_ns(QEMU_CLOCK_REALTIME); + if (now >=3D dl) { + info_report("vm_planes: plane %" PRIu64 " boot timeout= " + "(%u/%u halted)", plane_id, halted_count, + ps->vcpu_count); + break; + } + qemu_cond_timedwait(&cond, &mutex, 1000); + } + qemu_mutex_unlock(&mutex); + } + + /* Note: threads keep running for the plane's lifetime; we + * intentionally leak ctxs/threads =E2=80=94 they outlive this cal= l. */ + + /* Seal plane memory via KVM_SET_MEMORY_ATTRIBUTES(NO_WRITE|NO_EXE= C) */ + { + struct kvm_memory_attributes ma =3D { + .address =3D ps->load_offset, + .size =3D ps->memory_size, + .attributes =3D KVM_MEMORY_ATTRIBUTE_NO_WRITE | + KVM_MEMORY_ATTRIBUTE_NO_EXEC, + .flags =3D 0, + }; + int pr =3D kvm_vm_ioctl(s, KVM_SET_MEMORY_ATTRIBUTES, &ma); + if (pr < 0) { + warn_report("vm_planes: plane %" PRIu64 " seal failed (%d)= ", + plane_id, pr); + } else { + info_report("vm_planes: plane %" PRIu64 " memory sealed", + plane_id); + } + } + ps->host_addr =3D NULL; + + info_report("vm_planes: plane %" PRIu64 " launched =E2=80=94 entry= 0x%" PRIx64 + " vcpus %u", plane_id, entry_point, ps->vcpu_count); + } + + run->hypercall.ret =3D 0; + return 0; +} + +static int vbs_apply_protection(KVMState *s, uint64_t gpa, uint64_t size, + uint32_t perms) +{ + uint64_t attrs =3D 0; + struct kvm_memory_attributes ma; + + if (!(perms & 2)) { + attrs |=3D KVM_MEMORY_ATTRIBUTE_NO_WRITE; + } + if (!(perms & 4)) { + attrs |=3D KVM_MEMORY_ATTRIBUTE_NO_EXEC; + } + if (!attrs) { + return 0; + } + + ma.address =3D gpa; + ma.size =3D size; + ma.attributes =3D attrs; + ma.flags =3D 0; + return kvm_vm_ioctl(s, KVM_SET_MEMORY_ATTRIBUTES, &ma); +} + +static int32_t vbs_handle_protect_memory(KVMState *s, uint64_t ca_gpa) +{ + uint64_t gpa, size; + uint32_t perms, arg_size; + + cpu_physical_memory_read(ca_gpa + CA_OFF_ARG_SIZE, &arg_size, 4); + if (arg_size < 24) { + return -22; + } + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 0, &gpa, 8); + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 8, &size, 8); + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 16, &perms, 4); + return vbs_apply_protection(s, gpa, size, perms); +} + +static int32_t vbs_handle_seal_kernel(KVMState *s, uint64_t ca_gpa) +{ + uint64_t text_gpa, text_size, rodata_gpa, rodata_size; + uint32_t arg_size; + int ret; + + cpu_physical_memory_read(ca_gpa + CA_OFF_ARG_SIZE, &arg_size, 4); + if (arg_size < 40) { + return -22; + } + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 0, &text_gpa, 8); + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 8, &text_size, 8); + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 16, &rodata_gpa, 8); + cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 24, &rodata_size, 8); + + ret =3D vbs_apply_protection(s, text_gpa, text_size, 1 | 4); + if (ret < 0) { + return ret; + } + return vbs_apply_protection(s, rodata_gpa, rodata_size, 1); +} + +static int kvm_handle_hc_vbs_vtl_call(X86CPU *cpu, struct kvm_run *run) +{ + uint64_t ca_gpa =3D run->hypercall.args[0]; + uint32_t call_id; + int32_t status; + KVMState *s =3D kvm_state; + + cpu_physical_memory_read(ca_gpa + CA_OFF_CALL_ID, &call_id, 4); + + switch (call_id) { + case VBS_CALL_INIT: + case VBS_CALL_SHUTDOWN: + status =3D 0; + break; + case VBS_CALL_PROTECT_MEMORY: + status =3D vbs_handle_protect_memory(s, ca_gpa); + break; + case VBS_CALL_SEAL_KERNEL: + status =3D vbs_handle_seal_kernel(s, ca_gpa); + break; + case VBS_CALL_VALIDATE_MODULE: + case VBS_CALL_SET_MODULE_PERMS: + case VBS_CALL_UNLOAD_MODULE: + case VBS_CALL_ADD_KEY: + case VBS_CALL_REVOKE_KEY: + case VBS_CALL_SEND_CERTS: + case VBS_CALL_KEXEC_VALIDATE: + case VBS_CALL_KEXEC_INVALIDATE: + status =3D 0; /* acknowledge */ + break; + default: + warn_report("vbs_vtl_call: unknown call_id 0x%04x", call_id); + status =3D -38; + break; + } + + cpu_physical_memory_write(ca_gpa + CA_OFF_STATUS, &status, 4); + run->hypercall.ret =3D 0; + return 0; +} + static int kvm_handle_hypercall(X86CPU *cpu, struct kvm_run *run) { if (run->hypercall.nr =3D=3D KVM_HC_MAP_GPA_RANGE) return kvm_handle_hc_map_gpa_range(cpu, run); + if (run->hypercall.nr =3D=3D KVM_HC_VM_PLANES_CONFIG) + return kvm_handle_hc_vm_planes_config(cpu, run); + if (run->hypercall.nr =3D=3D KVM_HC_VM_PLANES_ACTIVATE) + return kvm_handle_hc_vm_planes_activate(cpu, run); + if (run->hypercall.nr =3D=3D KVM_HC_VBS_VTL_CALL) + return kvm_handle_hc_vbs_vtl_call(cpu, run); =20 return -EINVAL; } --=20 2.55.0