From nobody Mon Sep 21 02:45:37 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=linux.microsoft.com ARC-Seal: i=1; a=rsa-sha256; t=1785927952; cv=none; d=zohomail.com; s=zohoarc; b=J5zfERnBDczrhDy3eMMskaTwd1PCzCgh6z+17F6udK6HyliS9gaRUaMpfM4UVTHX58YPs2Owuqdaf8ga5JqY8GfLb8+vZXWvMw3qiMyEoKExKcNMG7iMpIfAHvf7fjPiypmmTQQEbauN2vfJfb4XvBLZouMgfFAItl8poVSowkQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785927952; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=yjbu25v6uYx/tC7572WVZCaIO5Pyb1mNt2gKgiiUzwo=; b=RY101OYP508qMK2v4eeqrhYCBYmv/IOLzW2JzZEf/19eiM4FhsppFZKGDuctJuizkZoetV3m5moYxiMCLDGY8PFQ23v8j44ox7avhbtOEOIJVYrz8pvS7nz2yOJkaMdsPOmYwWBC3yYrws/3NM2deCmAqD5H4RkOEofLjAYr7AI= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785927952529552.4619189264348; Wed, 5 Aug 2026 04:05:52 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wrZQT-0000mu-CX; Wed, 05 Aug 2026 07:04:45 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wrZQQ-0000m9-Uo for qemu-devel@nongnu.org; Wed, 05 Aug 2026 07:04:43 -0400 Received: from linux.microsoft.com ([13.77.154.182]) by eggs.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wrZQN-0000DV-6Q for qemu-devel@nongnu.org; Wed, 05 Aug 2026 07:04:42 -0400 Received: from fedora.hsd1.wa.comcast.net (unknown [52.148.140.42]) by linux.microsoft.com (Postfix) with ESMTPSA id 9046920B716D; Wed, 5 Aug 2026 04:04:16 -0700 (PDT) DKIM-Filter: OpenDKIM Filter v2.11.0 linux.microsoft.com 9046920B716D DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.microsoft.com; s=default; t=1785927856; bh=yjbu25v6uYx/tC7572WVZCaIO5Pyb1mNt2gKgiiUzwo=; h=From:To:Cc:Subject:Date:In-Reply-To:References:From; b=sbFYykHAZKwQXyxWBPI0FoBVuLeeMLplg1WhwbCEb+VeSDKbfCGu4ufDU6bZ6amKK JkOIhOM3ijItBvFOviFBGuNLTSxKpENfJNrjMEt2vjqFYEW21xpykxAxlU6hLK5MZu E9jIH1iDv9WtFn39jaSumxVZdrxHaZ71Vj1shxDw= From: Sriram Nambakam To: qemu-devel@nongnu.org Cc: kvm@vger.kernel.org Subject: [RFC PATCH v1 4/5] target/i386/kvm: run the secure plane in-kernel (Option B) Date: Wed, 5 Aug 2026 04:04:31 -0700 Message-ID: <20260805110432.25167-5-snambakam@linux.microsoft.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260805110432.25167-1-snambakam@linux.microsoft.com> References: <20260805110432.25167-1-snambakam@linux.microsoft.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=13.77.154.182; envelope-from=snambakam@linux.microsoft.com; helo=linux.microsoft.com X-Spam_score_int: -19 X-Spam_score: -2.0 X-Spam_bar: -- X-Spam_report: (-2.0 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @linux.microsoft.com) X-ZM-MESSAGEID: 1785927954270158500 Move secure-plane execution into the kernel and out of QEMU. Plane vCPUs are now created and initialised on their owning plane-0 CPU thread via run_on_cpu() (plane_vcpu_create_cb / plane_vcpu_init_cb), and the activate path leaves the secure vCPU stopped: it boots lazily inside KVM on the first KVM_HC_VBS_VTL_CALL, which KVM now services itself. Drop the userspace secure-world emulation that this replaces: the dedicated vm_plane_vcpu_thread (with its 8250 emulation and per-plane serial capture), vbs_apply_protection, vbs_handle_protect_memory, vbs_handle_seal_kernel and the kvm_handle_hc_vbs_vtl_call dispatch. Track each plane vCPU's owning CPU index (vcpu_cpu_index) so the run_on_cpu callbacks target the right thread. Signed-off-by: Sriram Nambakam (cherry picked from commit cbecc9eda1514d200fefe2bb89f330be4f1ab1be) --- include/system/kvm_int.h | 1 + target/i386/kvm/kvm.c | 536 ++++++++++++--------------------------- 2 files changed, 157 insertions(+), 380 deletions(-) diff --git a/include/system/kvm_int.h b/include/system/kvm_int.h index e7c9dc95a2..cde0a148e1 100644 --- a/include/system/kvm_int.h +++ b/include/system/kvm_int.h @@ -114,6 +114,7 @@ struct KVMPlane { * kvm_get_plane_fd(s, plane_id); only LVBS-specific state lives here. */ struct kvm_vm_plane_state { int *vcpu_fds; + unsigned int *vcpu_cpu_index; /* owning plane-0 CPU index per plane v= CPU */ unsigned int vcpu_count; uint64_t load_offset; uint64_t memory_size; diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c index 9d89e5cac6..5ac7962e61 100644 --- a/target/i386/kvm/kvm.c +++ b/target/i386/kvm/kvm.c @@ -6541,10 +6541,9 @@ static int kvm_handle_hc_map_gpa_range(X86CPU *cpu, = struct kvm_run *run) * off 0 u64 load_offset * off 8 u64 memory_size * off 16 u64 entry_point - * off 24 u32 vcpu_count - * off 28 u32 kernel_format - * off 32 char kernel[128] - * off 160 char cmdline[512] + * off 24 u32 kernel_format + * off 28 char kernel[128] + * off 156 char cmdline[512] * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D */ =20 #define VM_PLANE_CFG_STRIDE 672 @@ -6568,23 +6567,37 @@ static int kvm_handle_hc_map_gpa_range(X86CPU *cpu,= struct kvm_run *run) #define CA_OFF_RESP_SIZE 16 #define CA_OFF_BUFFER 20 =20 -struct vm_plane_boot_ctx { - int vcpu_fd; - int vcpu_mmap_size; - unsigned int vcpu_idx; - uint64_t plane_id; - unsigned int *halted_count; - QemuMutex *mutex; - QemuCond *cond; - int result; - int log_fd; - QemuMutex wake_mutex; - QemuCond wake_cond; - bool halted; - bool kick; - bool stopped; +/* + * Plane>0 vCPU creation and initialization must run on the OWNING plane-0 + * vCPU's thread. A plane>0 vCPU shares its plane-0 sibling's + * struct kvm_vcpu_common, including the single embedded preempt_notifier. + * KVM_CREATE_VCPU and KVM_SET_{S,}REGS / KVM_SET_MP_STATE all vcpu_load() + * the plane vCPU, which registers that shared notifier on the *calling* + * task. If issued from the config/activate-hypercall vCPU's thread while= a + * sibling is concurrently in KVM_RUN on another host CPU, the same + * hlist_node would be linked onto two tasks' preempt-notifier lists -> + * list corruption -> host hard lockup. run_on_cpu() forces the sibling o= ut + * of KVM_RUN (KVM_RUN returns, vcpu_put() unregisters the notifier) and + * runs the work on that sibling's own thread, so the notifier is owned by + * exactly one task. For the current vCPU, run_on_cpu() runs inline. + */ +struct plane_vcpu_create_ctx { + KVMState *s; + unsigned int plane_id; + unsigned int vcpu_id; + int fd; /* out: vcpu fd, or -errno on failure */ }; =20 +static void plane_vcpu_create_cb(CPUState *cs, run_on_cpu_data data) +{ + struct plane_vcpu_create_ctx *ctx =3D data.host_ptr; + int fd; + + fd =3D kvm_vm_plane_ioctl(ctx->s, ctx->plane_id, KVM_CREATE_VCPU, + (void *)(uintptr_t)ctx->vcpu_id); + ctx->fd =3D (fd < 0) ? -errno : fd; +} + static int kvm_handle_hc_vm_planes_config(X86CPU *cpu, struct kvm_run *run) { uint64_t gpa =3D run->hypercall.args[0]; @@ -6645,7 +6658,7 @@ static int kvm_handle_hc_vm_planes_config(X86CPU *cpu= , struct kvm_run *run) for (plane_id =3D 1; plane_id < plane_count; plane_id++) { uint64_t plane_gpa =3D gpa + (plane_id * VM_PLANE_CFG_STRIDE); uint64_t load_offset =3D 0, memory_size =3D 0, entry_point =3D 0; - uint32_t vcpu_count =3D 0; + uint32_t vcpu_count =3D plane0_vcpu_count; struct kvm_vm_plane_state *ps =3D &s->vm_planes[plane_id]; char cmdline_buf[512]; MemoryRegionSection section; @@ -6655,29 +6668,24 @@ static int kvm_handle_hc_vm_planes_config(X86CPU *c= pu, struct kvm_run *run) cpu_physical_memory_read(plane_gpa + 0, &load_offset, 8); cpu_physical_memory_read(plane_gpa + 8, &memory_size, 8); cpu_physical_memory_read(plane_gpa + 16, &entry_point, 8); - cpu_physical_memory_read(plane_gpa + 24, &vcpu_count, 4); =20 - if (!memory_size || !vcpu_count) { + /* + * joergroedel plane model: a plane has exactly one vCPU per + * plane-0 vCPU =E2=80=94 each is the sibling of a plane-0 vCPU sh= aring the + * same logical CPU. The guest does not configure a count; QEMU + * mirrors the plane-0 vCPU set. + */ + if (!memory_size) { error_report("vm_planes: plane %" PRIu64 " invalid " - "(load_offset=3D0x%" PRIx64 " size=3D0x%" PRIx64 - " vcpus=3D%u)", - plane_id, load_offset, memory_size, vcpu_count); - g_free(plane0_vcpu_ids); - run->hypercall.ret =3D -EINVAL; - return 0; - } - - if (vcpu_count > plane0_vcpu_count) { - error_report("vm_planes: plane %" PRIu64 " requests %u vCPUs, " - "but plane0 has %u", plane_id, vcpu_count, - plane0_vcpu_count); + "(load_offset=3D0x%" PRIx64 " size=3D0x%" PRIx64 = ")", + plane_id, load_offset, memory_size); g_free(plane0_vcpu_ids); run->hypercall.ret =3D -EINVAL; return 0; } =20 memset(cmdline_buf, 0, sizeof(cmdline_buf)); - cpu_physical_memory_read(plane_gpa + 160, cmdline_buf, + cpu_physical_memory_read(plane_gpa + 156, cmdline_buf, sizeof(cmdline_buf)); cmdline_buf[sizeof(cmdline_buf) - 1] =3D '\0'; memcpy(ps->cmdline, cmdline_buf, sizeof(ps->cmdline)); @@ -6714,24 +6722,46 @@ static int kvm_handle_hc_vm_planes_config(X86CPU *c= pu, struct kvm_run *run) memory_region_unref(section.mr); =20 ps->vcpu_fds =3D g_new0(int, vcpu_count); + ps->vcpu_cpu_index =3D g_new0(unsigned int, vcpu_count); for (i =3D 0; i < vcpu_count; i++) { - int vcpu_fd; unsigned int vcpu_id =3D plane0_vcpu_ids[i]; + CPUState *target =3D qemu_get_cpu(vcpu_id); + struct plane_vcpu_create_ctx cctx =3D { + .s =3D s, + .plane_id =3D plane_id, + .vcpu_id =3D vcpu_id, + .fd =3D -EINVAL, + }; + + if (!target) { + error_report("vm_planes: plane %" PRIu64 " no CPU for id= =3D%u", + plane_id, vcpu_id); + close(plane_fd); + kvm_set_plane_fd(s, plane_id, -1); + g_free(plane0_vcpu_ids); + run->hypercall.ret =3D -EINVAL; + return 0; + } =20 - vcpu_fd =3D kvm_vm_plane_ioctl(s, plane_id, KVM_CREATE_VCPU, - (void *)(uintptr_t)vcpu_id); - if (vcpu_fd < 0) { + /* Create on the owning CPU's thread; see plane_vcpu_create_cb= . */ + bql_lock(); + run_on_cpu(target, plane_vcpu_create_cb, + RUN_ON_CPU_HOST_PTR(&cctx)); + bql_unlock(); + + if (cctx.fd < 0) { error_report("vm_planes: KVM_CREATE_VCPU plane %" PRIu64 " vcpu %u failed: %s", - plane_id, vcpu_id, strerror(errno)); + plane_id, vcpu_id, strerror(-cctx.fd)); close(plane_fd); kvm_set_plane_fd(s, plane_id, -1); g_free(plane0_vcpu_ids); - run->hypercall.ret =3D -errno; + run->hypercall.ret =3D cctx.fd; return 0; } =20 - ps->vcpu_fds[i] =3D vcpu_fd; + ps->vcpu_fds[i] =3D cctx.fd; + ps->vcpu_cpu_index[i] =3D vcpu_id; } =20 ps->vcpu_count =3D vcpu_count; @@ -6757,10 +6787,6 @@ static int kvm_init_plane_vcpu(int vcpu_fd, uint64_t= entry_addr, { struct kvm_regs regs =3D {}; struct kvm_sregs sregs =3D {}; - struct kvm_mp_state mp =3D { - .mp_state =3D is_bsp ? KVM_MP_STATE_RUNNABLE - : KVM_MP_STATE_INIT_RECEIVED, - }; int ret; =20 sregs.cs.base =3D 0; sregs.cs.limit =3D 0xffffffff; sregs.cs.selector = =3D 0x10; @@ -6806,144 +6832,42 @@ static int kvm_init_plane_vcpu(int vcpu_fd, uint64= _t entry_addr, return -errno; } =20 - ret =3D ioctl(vcpu_fd, KVM_SET_MP_STATE, &mp); - if (ret < 0) { - error_report("vm_planes: KVM_SET_MP_STATE: %s", strerror(errno)); - return -errno; - } + /* + * Do NOT issue KVM_SET_MP_STATE on a plane (plane_level > 0) vCPU fd: + * the VM-planes kernel only whitelists a subset of vCPU ioctls for + * plane vCPUs (KVM_SET_*REGS/SREGS/FPU/LAPIC/...), and KVM_SET_MP_STA= TE + * is intentionally excluded -> it returns -EINVAL. The kernel already + * establishes the correct initial MP state at create time: the plane + * BSP (vcpu_id =3D=3D bsp_vcpu_id) is left RUNNABLE, and APs are left + * UNINITIALIZED (wait-for-INIT), which is the correct power-on state. + * The plane guest's own SMP bringup INIT/SIPIs its APs via the + * in-kernel LAPIC, so no userspace MP-state poke is needed. + */ return 0; } =20 -static void *vm_plane_vcpu_thread(void *arg) -{ - struct vm_plane_boot_ctx *ctx =3D arg; - struct kvm_run *kvm_run; - int ret; - bool boot_signaled =3D false; - - kvm_run =3D mmap(NULL, ctx->vcpu_mmap_size, PROT_READ | PROT_WRITE, - MAP_SHARED, ctx->vcpu_fd, 0); - if (kvm_run =3D=3D MAP_FAILED) { - error_report("vm_planes: plane %" PRIu64 " vcpu %u: mmap failed: %= s", - ctx->plane_id, ctx->vcpu_idx, strerror(errno)); - ctx->result =3D -errno; - qemu_mutex_lock(ctx->mutex); - (*ctx->halted_count)++; - qemu_cond_signal(ctx->cond); - qemu_mutex_unlock(ctx->mutex); - return NULL; - } - - for (;;) { - ret =3D ioctl(ctx->vcpu_fd, KVM_RUN, 0); - if (ret < 0) { - if (errno =3D=3D EINTR || errno =3D=3D EAGAIN) { - continue; - } - error_report("vm_planes: plane %" PRIu64 " vcpu %u: KVM_RUN: %= s", - ctx->plane_id, ctx->vcpu_idx, strerror(errno)); - ctx->result =3D -errno; - break; - } - - switch (kvm_run->exit_reason) { - case KVM_EXIT_HLT: - if (!boot_signaled) { - boot_signaled =3D true; - qemu_mutex_lock(ctx->mutex); - (*ctx->halted_count)++; - qemu_cond_signal(ctx->cond); - qemu_mutex_unlock(ctx->mutex); - } - qemu_mutex_lock(&ctx->wake_mutex); - ctx->halted =3D true; - while (!ctx->kick && !ctx->stopped) { - qemu_cond_wait(&ctx->wake_cond, &ctx->wake_mutex); - } - ctx->halted =3D false; - ctx->kick =3D false; - if (ctx->stopped) { - qemu_mutex_unlock(&ctx->wake_mutex); - goto done; - } - qemu_mutex_unlock(&ctx->wake_mutex); - break; - - case KVM_EXIT_IO: { - uint8_t *io_data =3D (uint8_t *)kvm_run + kvm_run->io.data_off= set; - size_t io_size =3D kvm_run->io.size * kvm_run->io.count; - uint16_t port =3D kvm_run->io.port; - - if (kvm_run->io.direction =3D=3D KVM_EXIT_IO_OUT) { - if (port =3D=3D 0x3f8 && ctx->log_fd >=3D 0) { - ssize_t w =3D write(ctx->log_fd, io_data, io_size); - (void)w; - } - } else { - memset(io_data, 0, io_size); - switch (port) { - case 0x3fa: memset(io_data, 0xc1, io_size); break; - case 0x3fb: memset(io_data, 0x03, io_size); break; - case 0x3fc: memset(io_data, 0x08, io_size); break; - case 0x3fd: memset(io_data, 0x60, io_size); break; - case 0x3fe: memset(io_data, 0xb0, io_size); break; - default: break; - } - } - usleep(100); - break; - } - - case KVM_EXIT_MMIO: - usleep(100); - break; - - case KVM_EXIT_SHUTDOWN: { - struct kvm_regs dbg =3D {}; - ioctl(ctx->vcpu_fd, KVM_GET_REGS, &dbg); - error_report("vm_planes: plane %" PRIu64 " vcpu %u: shutdown " - "RIP=3D0x%" PRIx64 " RSP=3D0x%" PRIx64, - ctx->plane_id, ctx->vcpu_idx, - (uint64_t)dbg.rip, (uint64_t)dbg.rsp); - ctx->result =3D -EFAULT; - goto done; - } - - case KVM_EXIT_FAIL_ENTRY: - error_report("vm_planes: plane %" PRIu64 " vcpu %u: entry fail= ure " - "0x%" PRIx64, ctx->plane_id, ctx->vcpu_idx, - (uint64_t)kvm_run->fail_entry.hardware_entry_fail= ure_reason); - ctx->result =3D -EFAULT; - goto done; - - case KVM_EXIT_INTERNAL_ERROR: - error_report("vm_planes: plane %" PRIu64 " vcpu %u: internal " - "error %u", ctx->plane_id, ctx->vcpu_idx, - kvm_run->internal.suberror); - ctx->result =3D -EFAULT; - goto done; +/* Initialize a plane vCPU on its owning CPU's thread (see + * plane_vcpu_create_cb for why KVM_SET_*REGS / KVM_SET_MP_STATE, which + * vcpu_load() the shared common, must not race the running sibling). */ +struct plane_vcpu_init_ctx { + int vcpu_fd; + uint64_t entry_addr; + uint64_t stack_addr; + uint64_t zero_page_gpa; + uint64_t page_table_gpa; + uint64_t gdt_gpa; + bool is_bsp; + int ret; /* out: 0 on success, -errno on failure */ +}; =20 - default: - error_report("vm_planes: plane %" PRIu64 " vcpu %u: unexpected= " - "exit %u", ctx->plane_id, ctx->vcpu_idx, - kvm_run->exit_reason); - ctx->result =3D -EFAULT; - goto done; - } - } +static void plane_vcpu_init_cb(CPUState *cs, run_on_cpu_data data) +{ + struct plane_vcpu_init_ctx *ctx =3D data.host_ptr; =20 -done: - if (ctx->log_fd >=3D 0 && ctx->vcpu_idx =3D=3D 0) { - close(ctx->log_fd); - } - munmap(kvm_run, ctx->vcpu_mmap_size); - qemu_mutex_lock(ctx->mutex); - if (!boot_signaled) { - (*ctx->halted_count)++; - qemu_cond_signal(ctx->cond); - } - qemu_mutex_unlock(ctx->mutex); - return NULL; + ctx->ret =3D kvm_init_plane_vcpu(ctx->vcpu_fd, ctx->entry_addr, + ctx->stack_addr, ctx->zero_page_gpa, + ctx->page_table_gpa, ctx->gdt_gpa, + ctx->is_bsp); } =20 static int kvm_handle_hc_vm_planes_activate(X86CPU *cpu, struct kvm_run *r= un) @@ -6952,7 +6876,6 @@ static int kvm_handle_hc_vm_planes_activate(X86CPU *c= pu, struct kvm_run *run) uint64_t plane_count =3D run->hypercall.args[1]; uint64_t plane_id; KVMState *s =3D kvm_state; - int vcpu_mmap_size; =20 if (!gpa || !plane_count || !s->vm_planes || plane_count !=3D s->vm_plane_count) { @@ -6960,26 +6883,13 @@ static int kvm_handle_hc_vm_planes_activate(X86CPU = *cpu, struct kvm_run *run) return 0; } =20 - vcpu_mmap_size =3D kvm_ioctl(s, KVM_GET_VCPU_MMAP_SIZE, 0); - if (vcpu_mmap_size <=3D 0) { - error_report("vm_planes: KVM_GET_VCPU_MMAP_SIZE failed"); - run->hypercall.ret =3D -EINVAL; - return 0; - } - for (plane_id =3D 1; plane_id < plane_count; plane_id++) { struct kvm_vm_plane_state *ps =3D &s->vm_planes[plane_id]; uint64_t stack_addr; uint64_t entry_point =3D 0; uint64_t cmdline_gpa, zero_page_gpa; uint64_t pt_base, pml4_gpa, pdpt_gpa, pd_base, gdt_gpa_val; - struct vm_plane_boot_ctx *ctxs; - QemuThread *threads; - QemuMutex mutex; - QemuCond cond; - unsigned int halted_count =3D 0; unsigned int i; - int plane_log_fd =3D -1; =20 if (kvm_get_plane_fd(s, plane_id) < 0 || !ps->vcpu_count || !ps->host_addr) { @@ -7086,98 +6996,60 @@ static int kvm_handle_hc_vm_planes_activate(X86CPU = *cpu, struct kvm_run *run) memcpy(PLANE_HOST(gdt_gpa_val), gdt, sizeof(gdt)); } =20 - /* Initialize all plane vCPUs */ + /* Initialize all plane vCPUs on their owning CPU threads. */ for (i =3D 0; i < ps->vcpu_count; i++) { - int ret =3D kvm_init_plane_vcpu(ps->vcpu_fds[i], entry_point, - stack_addr, zero_page_gpa, pml4_= gpa, - gdt_gpa_val, i =3D=3D 0); - if (ret) { - error_report("vm_planes: init plane %" PRIu64 " vcpu %u " - "failed", plane_id, i); - run->hypercall.ret =3D ret; + CPUState *target =3D qemu_get_cpu(ps->vcpu_cpu_index[i]); + struct plane_vcpu_init_ctx ictx =3D { + .vcpu_fd =3D ps->vcpu_fds[i], + .entry_addr =3D entry_point, + .stack_addr =3D stack_addr, + .zero_page_gpa =3D zero_page_gpa, + .page_table_gpa =3D pml4_gpa, + .gdt_gpa =3D gdt_gpa_val, + .is_bsp =3D (i =3D=3D 0), + .ret =3D -EINVAL, + }; + + if (!target) { + error_report("vm_planes: plane %" PRIu64 " no CPU for vcpu= %u", + plane_id, i); + run->hypercall.ret =3D -EINVAL; return 0; } - } -#undef PLANE_HOST - - /* Serial log */ - { - char lp[256]; - snprintf(lp, sizeof(lp), "/tmp/plane%" PRIu64 "-serial.log", - plane_id); - plane_log_fd =3D open(lp, O_CREAT | O_WRONLY | O_TRUNC, 0644); - } =20 - /* Spawn vCPU threads */ - qemu_mutex_init(&mutex); - qemu_cond_init(&cond); - - ctxs =3D g_new0(struct vm_plane_boot_ctx, ps->vcpu_count); - threads =3D g_new0(QemuThread, ps->vcpu_count); + /* See plane_vcpu_init_cb: must run on the sibling's own threa= d. */ + bql_lock(); + run_on_cpu(target, plane_vcpu_init_cb, + RUN_ON_CPU_HOST_PTR(&ictx)); + bql_unlock(); =20 - for (i =3D 0; i < ps->vcpu_count; i++) { - char name[48]; - - ctxs[i].vcpu_fd =3D ps->vcpu_fds[i]; - ctxs[i].vcpu_mmap_size =3D vcpu_mmap_size; - ctxs[i].vcpu_idx =3D i; - ctxs[i].plane_id =3D plane_id; - ctxs[i].halted_count =3D &halted_count; - ctxs[i].mutex =3D &mutex; - ctxs[i].cond =3D &cond; - ctxs[i].result =3D 0; - ctxs[i].log_fd =3D (i =3D=3D 0) ? plane_log_fd : -1; - qemu_mutex_init(&ctxs[i].wake_mutex); - qemu_cond_init(&ctxs[i].wake_cond); - ctxs[i].halted =3D false; - ctxs[i].kick =3D false; - ctxs[i].stopped =3D false; - - snprintf(name, sizeof(name), "plane%" PRIu64 "-vcpu%u", - plane_id, i); - qemu_thread_create(&threads[i], name, vm_plane_vcpu_thread, - &ctxs[i], QEMU_THREAD_JOINABLE); - } - - /* Wait up to 5s for plane to reach first HLT */ - { - int64_t dl =3D qemu_clock_get_ns(QEMU_CLOCK_REALTIME) + - 5LL * 1000000000LL; - qemu_mutex_lock(&mutex); - while (halted_count < ps->vcpu_count) { - int64_t now =3D qemu_clock_get_ns(QEMU_CLOCK_REALTIME); - if (now >=3D dl) { - info_report("vm_planes: plane %" PRIu64 " boot timeout= " - "(%u/%u halted)", plane_id, halted_count, - ps->vcpu_count); - break; - } - qemu_cond_timedwait(&cond, &mutex, 1000); + if (ictx.ret) { + error_report("vm_planes: init plane %" PRIu64 " vcpu %u " + "failed", plane_id, i); + run->hypercall.ret =3D ictx.ret; + return 0; } - qemu_mutex_unlock(&mutex); } +#undef PLANE_HOST =20 - /* Note: threads keep running for the plane's lifetime; we - * intentionally leak ctxs/threads =E2=80=94 they outlive this cal= l. */ - - /* Seal plane memory via KVM_SET_MEMORY_ATTRIBUTES(NO_WRITE|NO_EXE= C) */ - { - struct kvm_memory_attributes ma =3D { - .address =3D ps->load_offset, - .size =3D ps->memory_size, - .attributes =3D KVM_MEMORY_ATTRIBUTE_NO_WRITE | - KVM_MEMORY_ATTRIBUTE_NO_EXEC, - .flags =3D 0, - }; - int pr =3D kvm_vm_ioctl(s, KVM_SET_MEMORY_ATTRIBUTES, &ma); - if (pr < 0) { - warn_report("vm_planes: plane %" PRIu64 " seal failed (%d)= ", - plane_id, pr); - } else { - info_report("vm_planes: plane %" PRIu64 " memory sealed", - plane_id); - } - } + /* + * Full Option B: do NOT run the secure plane from a dedicated + * userspace thread. The plane's vCPU has been initialized (entry + * point, page tables, MP state) but is left STOPPED. It boots + * lazily and in-kernel the first time the normal plane issues a + * VBS/VTL call: KVM switches to the secure plane within the normal + * plane's KVM_RUN, the secure kernel boots to its dispatch loop a= nd + * parks itself via KVM_HC_VBS_VTL_RETURN. This removes the racy + * per-plane thread that concurrently drove the shared vcpu->run + * page and caused host lockups. + * + * NB: the secure plane's own memory is intentionally NOT sealed + * here. Because memory attributes are currently applied to every + * plane's EPT alike, sealing the secure plane's region NO_EXEC + * would prevent the secure kernel from executing its own dispatch + * loop. Protecting the secure plane's memory from the normal pla= ne + * requires plane-aware EPT attributes and is left as a follow-up. + */ ps->host_addr =3D NULL; =20 info_report("vm_planes: plane %" PRIu64 " launched =E2=80=94 entry= 0x%" PRIx64 @@ -7188,107 +7060,13 @@ static int kvm_handle_hc_vm_planes_activate(X86CPU= *cpu, struct kvm_run *run) return 0; } =20 -static int vbs_apply_protection(KVMState *s, uint64_t gpa, uint64_t size, - uint32_t perms) -{ - uint64_t attrs =3D 0; - struct kvm_memory_attributes ma; - - if (!(perms & 2)) { - attrs |=3D KVM_MEMORY_ATTRIBUTE_NO_WRITE; - } - if (!(perms & 4)) { - attrs |=3D KVM_MEMORY_ATTRIBUTE_NO_EXEC; - } - if (!attrs) { - return 0; - } - - ma.address =3D gpa; - ma.size =3D size; - ma.attributes =3D attrs; - ma.flags =3D 0; - return kvm_vm_ioctl(s, KVM_SET_MEMORY_ATTRIBUTES, &ma); -} - -static int32_t vbs_handle_protect_memory(KVMState *s, uint64_t ca_gpa) -{ - uint64_t gpa, size; - uint32_t perms, arg_size; - - cpu_physical_memory_read(ca_gpa + CA_OFF_ARG_SIZE, &arg_size, 4); - if (arg_size < 24) { - return -22; - } - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 0, &gpa, 8); - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 8, &size, 8); - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 16, &perms, 4); - return vbs_apply_protection(s, gpa, size, perms); -} - -static int32_t vbs_handle_seal_kernel(KVMState *s, uint64_t ca_gpa) -{ - uint64_t text_gpa, text_size, rodata_gpa, rodata_size; - uint32_t arg_size; - int ret; - - cpu_physical_memory_read(ca_gpa + CA_OFF_ARG_SIZE, &arg_size, 4); - if (arg_size < 40) { - return -22; - } - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 0, &text_gpa, 8); - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 8, &text_size, 8); - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 16, &rodata_gpa, 8); - cpu_physical_memory_read(ca_gpa + CA_OFF_BUFFER + 24, &rodata_size, 8); - - ret =3D vbs_apply_protection(s, text_gpa, text_size, 1 | 4); - if (ret < 0) { - return ret; - } - return vbs_apply_protection(s, rodata_gpa, rodata_size, 1); -} - -static int kvm_handle_hc_vbs_vtl_call(X86CPU *cpu, struct kvm_run *run) -{ - uint64_t ca_gpa =3D run->hypercall.args[0]; - uint32_t call_id; - int32_t status; - KVMState *s =3D kvm_state; - - cpu_physical_memory_read(ca_gpa + CA_OFF_CALL_ID, &call_id, 4); - - switch (call_id) { - case VBS_CALL_INIT: - case VBS_CALL_SHUTDOWN: - status =3D 0; - break; - case VBS_CALL_PROTECT_MEMORY: - status =3D vbs_handle_protect_memory(s, ca_gpa); - break; - case VBS_CALL_SEAL_KERNEL: - status =3D vbs_handle_seal_kernel(s, ca_gpa); - break; - case VBS_CALL_VALIDATE_MODULE: - case VBS_CALL_SET_MODULE_PERMS: - case VBS_CALL_UNLOAD_MODULE: - case VBS_CALL_ADD_KEY: - case VBS_CALL_REVOKE_KEY: - case VBS_CALL_SEND_CERTS: - case VBS_CALL_KEXEC_VALIDATE: - case VBS_CALL_KEXEC_INVALIDATE: - status =3D 0; /* acknowledge */ - break; - default: - warn_report("vbs_vtl_call: unknown call_id 0x%04x", call_id); - status =3D -38; - break; - } - - cpu_physical_memory_write(ca_gpa + CA_OFF_STATUS, &status, 4); - run->hypercall.ret =3D 0; - return 0; -} - +/* + * Full Option B: VBS/VTL calls (KVM_HC_VBS_VTL_CALL / _RETURN) are now + * serviced entirely in-kernel by switching to the secure plane, which runs + * a real in-guest dispatcher. QEMU no longer emulates the secure world, = so + * those hypercalls never exit to userspace here. Only the plane + * configuration/activation hypercalls are handled by QEMU. + */ static int kvm_handle_hypercall(X86CPU *cpu, struct kvm_run *run) { if (run->hypercall.nr =3D=3D KVM_HC_MAP_GPA_RANGE) @@ -7297,8 +7075,6 @@ static int kvm_handle_hypercall(X86CPU *cpu, struct k= vm_run *run) return kvm_handle_hc_vm_planes_config(cpu, run); if (run->hypercall.nr =3D=3D KVM_HC_VM_PLANES_ACTIVATE) return kvm_handle_hc_vm_planes_activate(cpu, run); - if (run->hypercall.nr =3D=3D KVM_HC_VBS_VTL_CALL) - return kvm_handle_hc_vbs_vtl_call(cpu, run); =20 return -EINVAL; } --=20 2.55.0