From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425742; cv=none; d=zohomail.com; s=zohoarc; b=jaOhG3ariNaVN0aVzgiKz5+nzjXnVT8o5gnLBtwUTPyV8z3e8JdH4extvmZ0ntJg5AKbI4McMLNaZA0ke/pbr8ZsclDsa1bqX5G9fEnWo74JOm1FNpqG4Tz86tMHaAkM+FjFHUXx9fu8SX3VaK1INSYXfJYrCAK0K1ivLu3mteU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425742; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=Il7VWeSEPUOBiY8Vlk0cJzi/A/8E9vL/OigwqXSjBD0=; b=XgHSaOgGP+yfG5J3opguuOqR0er9yjcbsDXmuEYP3wJKFLEBWfo2VuHhmAtwWj33b7wd1xBoN3FkZMtj4WLkCsVVWhBtKSLZ1zy9w7JMs+uq/cVVdrt9ICjM6UBFDfoEqYdIhlbzX2WVnt/8RVhpIu3jmaLXP1hLn6IN+fOcaZU= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787425742448282.75027705599996; Sat, 22 Aug 2026 12:09:02 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr4v-000411-Eq; Sat, 22 Aug 2026 15:08:29 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr4u-00040j-F1 for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:28 -0400 Received: from mail-yx1-xb132.google.com ([2607:f8b0:4864:20::b132]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr4s-0005kQ-BO for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:28 -0400 Received: by mail-yx1-xb132.google.com with SMTP id 956f58d0204a3-66c9995ca60so5321313d50.1 for ; Sat, 22 Aug 2026 12:08:26 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66cf4665056sm1205778d50.8.2026.08.22.12.08.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425705; x=1788030505; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Il7VWeSEPUOBiY8Vlk0cJzi/A/8E9vL/OigwqXSjBD0=; b=diziOCGZdnUEJU3CPHGf5H9feSrTXY6Z73Mw7Cd8g6Hfkh/broC2H7v2vMcGPFfdiX bncpJftL3LCdREVPQBTprbr3muh3AQdreMzj/miYOELTH8UYaEsxBOEFjQZRbdkj1oFS gmFY2gC6iOl4zOqoVEnPyXiziaHO92GFcEwR6wEFlp/FcrvTso7Ia3YFf3aLvyfvW9dJ zPcB99vRD0ruu4fNfsuqfLPEyS3keDxu5FogF0dt4T5pp+O9QFxgsmRqWmVK74lZl6hi vqopyIYj8V2a8fUUYFqPXECLTPXTd3S7HPmOOLvgp6zdApui8MA9tuECrbAsv8ah97N6 SWsw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425705; x=1788030505; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Il7VWeSEPUOBiY8Vlk0cJzi/A/8E9vL/OigwqXSjBD0=; b=YhOIHBZictWqskOTpLFoA2pwiE+Ct9MuoC1cLqxaC9cR5snxb5aadzb5jUTHZ64QTE hFU44Py1ONJehNqBNZ2WNcjE1ynvopHlJIox67yvp6SXLVIVHW7S7nwtQSAxyH7NWABl XHpMwCdQRaaviNyT4iXcZL+39AmqQwFzG9Kz3K/2L7Z5Xdj6BvbEx6b10hWwYP3sLx0d We9ki5JbRuMIvMyYvuJUy5XZkpTV5kSYzURoS0Xgi6WGv6u4rTezwEca9JNmZLwbSS+N vG98JggQEzRfkUB56/Dkks3JteeIwKGiYiQEiAmxiC82dnFS98Ka4MuHZoS4omIhhfZb XmnA== X-Gm-Message-State: AFuF++lD0+n/D5DjJGD7OWz22NdTQNI+D7Y4OUq5d668dINY9BvBdCWO fqOD52NBIm7y+d+TsumpR4bgeHs5RYOoBsI7KMq/Z3QGcHjBWXrgoC1jylWJ1/Sx0d0= X-Gm-Gg: AR+sD12lTgRMM0TiGpf7/q9SIgWDnZkO3ZSsDlJXDy5t694rn2RJ3A4+Goo5xHiezda 45QrjCKDcASnQ6zTzvmgx3Gz+CONfgVInxAxsiK0QO4javxDAUqGH5lVafsQs3maWxoL0lUlH/D 7A/BVqEGLLO+86suH4aWTD5fboHv9MEoJ/1NlAKX933l+fzfyff5hNC1VQ1osccnBubEN2ZbyET wB9uYoSkzKthqbHgu61uScxIWDuLgcFLlLolHBm/3SLN16sOviAvO4QzkW8XwwWMkM/OsL4Bl30 Tnm2fyBxZ+GS51UX3U+G07PoxGMRm5sWqkZsADdPC4OLgIILU6v+mmv+8AtFucI7jY5vJpoOzer ykt0IP0fIuMqO43OpqeYFRpeoB8kDjMELpWSoxdgtveLQL3Ae0jc57mNjfmktWYiHXbgbwVilnv VwRvWkQ/I/VxciTszkz6xXjleQQbUu/GHvDBeInEIJrwj96USR4AhZOXHukDqGyw9jhImB7GCOv 4fHi2IRunSrg95xk8VvHi6uTUA9dy7vUXpndUbH X-Received: by 2002:a53:c505:0:b0:666:c4c5:7525 with SMTP id 956f58d0204a3-66ce5b83b4dmr3438472d50.24.1787425704951; Sat, 22 Aug 2026 12:08:24 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 1/7] accel/tcg: fold the dynamic cflags into CPUState::tcg_cflags Date: Sat, 22 Aug 2026 15:08:12 -0400 Message-ID: <20260822190818.1829249-2-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b132; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb132.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425743024158500 Content-Type: text/plain; charset="utf-8" curr_cflags() is called once per TB dispatch, from helper_lookup_tb_ptr() and from the cpu_exec() loop. It recomputes the same value every time: uint32_t cflags =3D cpu->tcg_cflags; if (unlikely(cpu_single_stepping(cpu))) { ... } else if (qatomic_read(&one_insn_per_tb)) { ... } else if (qemu_loglevel_mask(CPU_LOG_TB_NOCHAIN)) { ... } That is three loads and three branches on the hottest path in the interpreter, for state that changes only when gdb enables single-step, when one-insn-per-tb is toggled, or when the log mask changes. None of the three has to be sampled at dispatch time. Fold each into CPUState::tcg_cflags where it changes and curr_cflags() becomes a single load of a field that TB lookup has to read anyway. The derived bits -- CF_COUNT_MASK, CF_NO_GOTO_TB, CF_NO_GOTO_PTR and CF_SINGLE_STEP -- are never set by tcg_cflags_set(), so tcg_update_cflags() can recompute them in place without disturbing the rest, and conversely tcg_cflags_set() ORs in its bits without disturbing them. There are three places to call it: - tcg_exec_realizefn(), so that a CPU created after the command line has been parsed starts out with the right value. This covers user-only, where tcg_cpu_init_cflags() is not reached. linux-user's cpu_copy() copies tcg_cflags wholesale, so a cloned thread inherits it. - cpu_single_step(), which changes one CPU and runs either on that CPU's thread or with it stopped. - tcg_set_one_insn_per_tb() and qemu_set_log_internal(), which change every CPU. Both can be reached from the monitor while the vCPUs are running -- 'one-insn-per-tb on' and 'log nochain' -- so the update is queued with async_safe_run_on_cpu() and each CPU writes its own cflags with the others halted. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, in a build configured with --enable-lto: before: 1,646,994,254,249 instructions after: 1,562,204,796,597 instructions -5.15% That workload issues 8.4 billion dispatches, so the per-call saving is small but the aggregate is not. The emulated compiler produces byte-identical output before and after. Wall clock does not move: 133.19s to 132.58s, a 0.46% difference against a run-to-run spread larger than that. The removed work is a few predictable loads and branches that the host executes largely in parallel with the surrounding dispatch, so this patch is worth taking for the instruction count and for what it enables, not for a time saving that can be measured on its own. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec-common.c | 33 ++++++++++++++++++++++++++++++--- accel/tcg/cpu-exec.c | 3 +++ accel/tcg/internal-common.h | 11 +++++++++-- accel/tcg/tcg-all.c | 1 + cpu-target.c | 3 +++ include/system/tcg.h | 12 ++++++++++++ stubs/meson.build | 1 + stubs/tcg-cflags.c | 16 ++++++++++++++++ util/log.c | 4 ++++ 9 files changed, 79 insertions(+), 5 deletions(-) create mode 100644 stubs/tcg-cflags.c diff --git ./accel/tcg/cpu-exec-common.c ./accel/tcg/cpu-exec-common.c index 44e84344f3..dd2be475e2 100644 --- ./accel/tcg/cpu-exec-common.c +++ ./accel/tcg/cpu-exec-common.c @@ -36,9 +36,16 @@ void tcg_cflags_set(CPUState *cpu, uint32_t flags) cpu->tcg_cflags |=3D flags; } =20 -uint32_t curr_cflags(CPUState *cpu) +/* + * The bits of CPUState::tcg_cflags that tcg_cflags_set() never sets, beca= use + * they are derived from gdb single-step, one-insn-per-tb and -d nochain. + */ +#define CF_DERIVED (CF_COUNT_MASK | CF_NO_GOTO_TB | CF_NO_GOTO_PTR | \ + CF_SINGLE_STEP) + +void tcg_update_cflags(CPUState *cpu) { - uint32_t cflags =3D cpu->tcg_cflags; + uint32_t cflags =3D cpu->tcg_cflags & ~CF_DERIVED; =20 /* * Record gdb single-step. We should be exiting the TB by raising @@ -55,7 +62,27 @@ uint32_t curr_cflags(CPUState *cpu) cflags |=3D CF_NO_GOTO_TB; } =20 - return cflags; + cpu->tcg_cflags =3D cflags; +} + +static void tcg_update_cflags_work(CPUState *cpu, run_on_cpu_data data) +{ + tcg_update_cflags(cpu); +} + +void tcg_update_all_cflags(void) +{ + CPUState *cpu; + + /* + * one-insn-per-tb and -d nochain can both be changed from the monitor + * while the vCPUs are running. Have each CPU update its own cflags + * with the others halted, so that no dispatch can read a value that + * another thread is in the middle of writing. + */ + CPU_FOREACH(cpu) { + async_safe_run_on_cpu(cpu, tcg_update_cflags_work, RUN_ON_CPU_NULL= ); + } } =20 /* exit the current TB, but without causing any exception to be raised */ diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 257211235d..148e0f583e 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -1068,6 +1068,9 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_target_initialized =3D true; } =20 + /* Pick up one-insn-per-tb and -d nochain from the command line. */ + tcg_update_cflags(cpu); + cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); tlb_init(cpu); #ifndef CONFIG_USER_ONLY diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 9e7be2d78d..853d1b51ee 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -69,8 +69,15 @@ void tlb_destroy(CPUState *cpu); bool tcg_exec_realizefn(CPUState *cpu, Error **errp); void tcg_exec_unrealizefn(CPUState *cpu); =20 -/* current cflags for hashing/comparison */ -uint32_t curr_cflags(CPUState *cpu); +/* + * Current cflags for hashing/comparison. Everything that feeds into the + * value is folded into CPUState::tcg_cflags when it changes, by + * tcg_update_cflags(), so that TB dispatch only has to load it. + */ +static inline uint32_t curr_cflags(CPUState *cpu) +{ + return cpu->tcg_cflags; +} =20 void tb_check_watchpoint(CPUState *cpu, uintptr_t retaddr); =20 diff --git ./accel/tcg/tcg-all.c ./accel/tcg/tcg-all.c index 7186c10cf0..c9874a286a 100644 --- ./accel/tcg/tcg-all.c +++ ./accel/tcg/tcg-all.c @@ -254,6 +254,7 @@ static void tcg_set_one_insn_per_tb(Object *obj, bool v= alue, Error **errp) s->one_insn_per_tb =3D value; /* Set the global also: this changes the behaviour */ qatomic_set(&one_insn_per_tb, value); + tcg_update_all_cflags(); } =20 static void tcg_accel_class_init(ObjectClass *oc, const void *data) diff --git ./cpu-target.c ./cpu-target.c index 4783845c9b..50be591acf 100644 --- ./cpu-target.c +++ ./cpu-target.c @@ -24,6 +24,7 @@ #include "exec/replay-core.h" #include "exec/log.h" #include "hw/core/cpu.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 /* enable or disable single step mode. EXCP_DEBUG is returned by the @@ -35,6 +36,8 @@ void cpu_single_step(CPUState *cpu, unsigned flags) cpu->singlestep_flags, flags); cpu->singlestep_flags =3D flags; =20 + tcg_update_cflags(cpu); + #if !defined(CONFIG_USER_ONLY) const AccelOpsClass *ops =3D cpus_get_accel(); if (ops->update_guest_debug) { diff --git ./include/system/tcg.h ./include/system/tcg.h index 7622dcea30..2c2dbc753b 100644 --- ./include/system/tcg.h +++ ./include/system/tcg.h @@ -17,6 +17,18 @@ extern bool tcg_allowed; #define tcg_enabled() 0 #endif =20 +/* + * Recompute the parts of CPUState::tcg_cflags that TB dispatch consumes b= ut + * tcg_cflags_set() does not provide: gdb single-step, one-insn-per-tb and + * the CPU_LOG_TB_NOCHAIN log flag. Call whenever one of those changes. + * + * tcg_update_cflags() updates one CPU and must be called from that CPU's + * thread, or with it stopped. tcg_update_all_cflags() updates every CPU + * and is safe to call from the monitor while the vCPUs run. + */ +void tcg_update_cflags(CPUState *cpu); +void tcg_update_all_cflags(void); + /** * qemu_tcg_mttcg_enabled: * Check whether we are running MultiThread TCG or not. diff --git ./stubs/meson.build ./stubs/meson.build index 3b2f2680b1..0025e79226 100644 --- ./stubs/meson.build +++ ./stubs/meson.build @@ -3,6 +3,7 @@ # below, so that it is clear who needs the stubbed functionality. =20 stub_ss.add(files('cpu-get-clock.c')) +stub_ss.add(files('tcg-cflags.c')) stub_ss.add(files('fdset.c')) stub_ss.add(files('iothread-lock.c')) stub_ss.add(files('is-daemonized.c')) diff --git ./stubs/tcg-cflags.c ./stubs/tcg-cflags.c new file mode 100644 index 0000000000..cb278e94aa --- /dev/null +++ ./stubs/tcg-cflags.c @@ -0,0 +1,16 @@ +/* + * Stub for tcg_update_all_cflags(), for binaries that link util/log.c + * or cpu-target.c but not TCG. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include "qemu/osdep.h" +#include "system/tcg.h" + +void tcg_update_cflags(CPUState *cpu) +{ +} + +void tcg_update_all_cflags(void) +{ +} diff --git ./util/log.c ./util/log.c index 7cffbc1bf8..3fa46a67fa 100644 --- ./util/log.c +++ ./util/log.c @@ -27,6 +27,7 @@ #include "qemu/thread.h" #include "qemu/lockable.h" #include "qemu/rcu.h" +#include "system/tcg.h" #ifdef CONFIG_LINUX #include #endif @@ -301,6 +302,9 @@ static bool qemu_set_log_internal(const char *filename,= bool changed_name, #endif qemu_loglevel =3D log_flags; =20 + /* CPU_LOG_TB_NOCHAIN feeds into the per-CPU cflags. */ + tcg_update_all_cflags(); + daemonized =3D is_daemonized(); need_to_open_file =3D false; if (!daemonized) { --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787806998; cv=none; d=zohomail.com; s=zohoarc; b=LcgagdL6SE/EH8JF04s3VlFCTQepP7q0NsHftENsmHn5iu0qOODqACXxQD//s6w0p+SwhIHVdU1mHWEr8DRCC265YKp61lM5B+1pA56ols+ZLM4uuANfVUKCcVIaC7BhvFqB7YLwnrB4O5nCVp9ZipBo5OnvrdT5Uagzb00ApsI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787806998; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=rWVNU5X35JNMpH6UmX6plkLu9+bpHKzoChMIjMge2go=; b=C7NU9Yxc+7YJJZ1pZw0BOp50mIOHXeH1aIZ4qw2hcFtwdFtEqk9B3PijGOiYIlsIhbnk0Zhb8WbIPCcObWA0TOAOngbICswhRvh5KerNpkGJLUSgA8PBoqqVf73rjhoo5hZcaZZjULcr2Vf/I7Q7lfYvvVJms4D3Nl48SL9q14M= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787806998687364.0772347817748; Wed, 26 Aug 2026 22:03:18 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGU-0005og-AV; Thu, 27 Aug 2026 01:03:02 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGQ-0005oO-Sm for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:02:58 -0400 Received: from mail-yx1-xb134.google.com ([2607:f8b0:4864:20::b134]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGO-0002Hm-OB for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:02:58 -0400 Received: by mail-yx1-xb134.google.com with SMTP id 956f58d0204a3-66dbce0b639so916503d50.1 for ; Wed, 26 Aug 2026 22:02:56 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66d24721e47sm2409680d50.13.2026.08.26.22.02.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:02:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806975; x=1788411775; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rWVNU5X35JNMpH6UmX6plkLu9+bpHKzoChMIjMge2go=; b=KeMZzPqaFeOgrTN7MVQdXfpP477xa/EcRWur/+N+jwyvLquA/6dURNiz2y0fiws2Cw pVgJWNf2Qf/4R0dw6zkewipcoU2x3JQ847kjSD9zSnmr+NE4oEo7flUPVLmTS6sAhvrm TJZ/NDVL7vcB3Uha7pHmDos1Q1RKgBMqzZI9cdGeVRaRfOdgr6YATpcmCCscM/VYBvFf ZGudcgUm7fA1jC/SM8dQpNZZJDXDO6N5fbT/07sn+ytmiE5Myfmmnp+Fc2zBY8N376zp QNd9g94XYQHQJdvzJtJhAde3t8y6afF9PwHbocJDCQt8gi8IDTVpR9wnTfEDW/khVvX3 vMIA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806975; x=1788411775; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=rWVNU5X35JNMpH6UmX6plkLu9+bpHKzoChMIjMge2go=; b=WkQY3ZgW5pdQR3cAYonC4cEqFEwzHnpV/EiOFij23+rfk4UKAnLTZB99OuDXulWxWY t8+JOs3R7+FKOqBSuYY35RtxEZwgnPpq5mLllLIjRQovwDnjlREXB6q1uv0Y+tzg1DQv a6QFn1xPglYD/jj7cMh7fTH1YYYeShogxkfRABiAvy8694t/9J6AxzwuQffsFVOHGa/H 0PRtWwQGzq32A1FjkPygEbF9ox1WXPVsJ9E9eA7YRVS8FxtaXNofxTIh8w+63FRqDHWr sZ7boMzGIHbn3YwD0ciBs99VYAq65DbUjP0PtJT8wptQl2lbxnMUOjcSS0iHtZb452rQ WZFg== X-Gm-Message-State: AFuF++k3Dnznnko3ooK7erYZ47fiW76nRweN94eoqeCLINYhkQ0bqPYy 5x/Qf4QzZpKUlKPnr3WOUNnDYxiunxu3Nurp4izlUxDkUoc9BZ7Il1nl2M4eeZxe X-Gm-Gg: AR+sD11QdwYpEymealNmwQmuwwCmkqQAmUgMMAibhRl7aXHVw9uf8zOI+wBwGvu6G7M jW3EQXPrrf5I85tskPCqtFTXwMj7+AQSHBq6PO1z4oWbRdXFDaxHKPf6OH+V/v2sCxk1mQW9O6n dsqWTSDq/UcBlNaNv1TYk0v6pF11y2ncW8F0wCCHyKx5/nZgbCF20w6fjxY3pvUmfdfkJBJJHjV kMEaRRM3yF6flQblUOtnoZZ+E4oz61NG5zSiFZacPfTQb+wKphpzJluMR4TzsJJzTG79CR9Tdlo P/3tPXQlUAUAT1Sv5gI9XX9pTORPRMb1d5LAiQYGy42y7o3Mj6oMdNN870KQSBe/aGWaxVlW9hG R1dbFm4L0ulVoTHiuOObWEBw7Unnw87q9BfcYoOrwjh0B4R2lJRQtP8EO6JRaTRhUE3+LMu156/ j3RGEEkies9iSie0yUvgxTxK4PM8d05FXBEmsgbaeYe9A8yNSsqB2F85ErvW3ow+u8vPtrggx24 VM5mzDmYasW3FRV+5Nup6Z1m1/30F9HmMdlI1D7 X-Received: by 2002:a53:ad09:0:b0:66c:4efb:6518 with SMTP id 956f58d0204a3-66d25700958mr3876985d50.21.1787806975347; Wed, 26 Aug 2026 22:02:55 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 1/9] accel/tcg: fold the dynamic cflags into CPUState::tcg_cflags Date: Thu, 27 Aug 2026 01:02:33 -0400 Message-ID: <20260827050241.3713332-2-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b134; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb134.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807000174158500 Content-Type: text/plain; charset="utf-8" curr_cflags() is called once per TB dispatch, from helper_lookup_tb_ptr() and from the cpu_exec() loop. It recomputes the same value every time: uint32_t cflags =3D cpu->tcg_cflags; if (unlikely(cpu_single_stepping(cpu))) { ... } else if (qatomic_read(&one_insn_per_tb)) { ... } else if (qemu_loglevel_mask(CPU_LOG_TB_NOCHAIN)) { ... } That is three loads and three branches on the hottest path in the interpreter, for state that changes only when gdb enables single-step, when one-insn-per-tb is toggled, or when the log mask changes. None of the three has to be sampled at dispatch time. Fold each into CPUState::tcg_cflags where it changes and curr_cflags() becomes a single load of a field that TB lookup has to read anyway. The derived bits -- CF_COUNT_MASK, CF_NO_GOTO_TB, CF_NO_GOTO_PTR and CF_SINGLE_STEP -- are never set by tcg_cflags_set(), so tcg_update_cflags() can recompute them in place without disturbing the rest, and conversely tcg_cflags_set() ORs in its bits without disturbing them. There are three places to call it: - tcg_exec_realizefn(), so that a CPU created after the command line has been parsed starts out with the right value. This covers user-only, where tcg_cpu_init_cflags() is not reached. linux-user's cpu_copy() copies tcg_cflags wholesale, so a cloned thread inherits it. - cpu_single_step(), which changes one CPU. gdb is the only caller that matters; in system mode it runs with the vCPUs stopped, and in user mode gdb_continue_partial() can reach a thread that is still running, because gdb_handlesig() stops only the thread that trapped. That is exactly the plain cross-thread store to another CPU's CPUState that cpu->singlestep_flags already was, read back by that CPU through cpu_single_stepping() in curr_cflags(). This patch changes which field carries it, not who writes it or how. - hmp_one_insn_per_tb() and hmp_log(), which change every CPU while the vCPUs are running, so the update is queued with async_run_on_cpu() and each CPU writes its own cflags from its own thread. The command line spellings of those two settings need nothing: they are parsed before any CPU is realized, so tcg_exec_realizefn() picks them up. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, in a build configured with --enable-lto: before: 1,646,994,254,249 instructions after: 1,562,204,796,597 instructions -5.15% That workload issues 8.4 billion dispatches, so the per-call saving is small but the aggregate is not. The emulated compiler produces byte-identical output before and after. Wall clock does not move: 133.19s to 132.58s, a 0.46% difference against a run-to-run spread larger than that. The removed work is a few predictable loads and branches that the host executes largely in parallel with the surrounding dispatch, so this patch is worth taking for the instruction count and for what it enables, not for a time saving that can be measured on its own. v4: Update the cflags from the HMP handlers for 'log' and 'one-insn-per-tb' rather than from qemu_set_log_internal() and the accelerator property setter. Those are the paths that reach a running vCPU, and the monitor is the only thing that does. Suggested by Richard Henderson. v4: Queue the per-CPU update with async_run_on_cpu() rather than async_safe_run_on_cpu(). Halting the other vCPUs buys nothing: the queued work already runs on the owning CPU's own thread. Suggested by Alex Bennee, who also asked whether there are cross-vCPU updates of tcg_cflags at all. With this change the monitor path has none: the only remaining writer from another thread is cpu_single_step(), above, which is neither new nor made worse here. v4: Move the stub to accel/stubs/, which is where the other accelerator stubs live. Signed-off-by: Matt Turner Reviewed-by: Richard Henderson --- accel/stubs/meson.build | 1 + accel/stubs/tcg-stub.c | 16 ++++++++++++++++ accel/tcg/cpu-exec-common.c | 33 ++++++++++++++++++++++++++++++--- accel/tcg/cpu-exec.c | 3 +++ accel/tcg/internal-common.h | 11 +++++++++-- cpu-target.c | 3 +++ include/system/tcg.h | 12 ++++++++++++ monitor/hmp-cmds.c | 5 +++++ system/runstate-hmp-cmds.c | 4 ++++ 9 files changed, 83 insertions(+), 5 deletions(-) create mode 100644 accel/stubs/tcg-stub.c diff --git ./accel/stubs/meson.build ./accel/stubs/meson.build index 7c6d7ad943..ccad583e64 100644 --- ./accel/stubs/meson.build +++ ./accel/stubs/meson.build @@ -4,6 +4,7 @@ stub_ss.add(files( 'nitro-stub.c', 'mshv-stub.c', 'nvmm-stub.c', + 'tcg-stub.c', 'whpx-stub.c', 'xen-stub.c', )) diff --git ./accel/stubs/tcg-stub.c ./accel/stubs/tcg-stub.c new file mode 100644 index 0000000000..f9e1bd22d6 --- /dev/null +++ ./accel/stubs/tcg-stub.c @@ -0,0 +1,16 @@ +/* + * Stubs for the TCG entry points in system/tcg.h, for binaries that link + * cpu-target.c or the HMP command handlers but not TCG. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include "qemu/osdep.h" +#include "system/tcg.h" + +void tcg_update_cflags(CPUState *cpu) +{ +} + +void tcg_update_all_cflags(void) +{ +} diff --git ./accel/tcg/cpu-exec-common.c ./accel/tcg/cpu-exec-common.c index 44e84344f3..9f3517f36b 100644 --- ./accel/tcg/cpu-exec-common.c +++ ./accel/tcg/cpu-exec-common.c @@ -36,9 +36,16 @@ void tcg_cflags_set(CPUState *cpu, uint32_t flags) cpu->tcg_cflags |=3D flags; } =20 -uint32_t curr_cflags(CPUState *cpu) +/* + * The bits of CPUState::tcg_cflags that tcg_cflags_set() never sets, beca= use + * they are derived from gdb single-step, one-insn-per-tb and -d nochain. + */ +#define CF_DERIVED (CF_COUNT_MASK | CF_NO_GOTO_TB | CF_NO_GOTO_PTR | \ + CF_SINGLE_STEP) + +void tcg_update_cflags(CPUState *cpu) { - uint32_t cflags =3D cpu->tcg_cflags; + uint32_t cflags =3D cpu->tcg_cflags & ~CF_DERIVED; =20 /* * Record gdb single-step. We should be exiting the TB by raising @@ -55,7 +62,27 @@ uint32_t curr_cflags(CPUState *cpu) cflags |=3D CF_NO_GOTO_TB; } =20 - return cflags; + cpu->tcg_cflags =3D cflags; +} + +static void tcg_update_cflags_work(CPUState *cpu, run_on_cpu_data data) +{ + tcg_update_cflags(cpu); +} + +void tcg_update_all_cflags(void) +{ + CPUState *cpu; + + /* + * one-insn-per-tb and -d nochain can both be changed from the monitor + * while the vCPUs are running. Queue the update onto each CPU rather + * than writing tcg_cflags from here, so that the field is only ever + * written by the CPU that owns it. + */ + CPU_FOREACH(cpu) { + async_run_on_cpu(cpu, tcg_update_cflags_work, RUN_ON_CPU_NULL); + } } =20 /* exit the current TB, but without causing any exception to be raised */ diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 257211235d..148e0f583e 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -1068,6 +1068,9 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_target_initialized =3D true; } =20 + /* Pick up one-insn-per-tb and -d nochain from the command line. */ + tcg_update_cflags(cpu); + cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); tlb_init(cpu); #ifndef CONFIG_USER_ONLY diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 9e7be2d78d..853d1b51ee 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -69,8 +69,15 @@ void tlb_destroy(CPUState *cpu); bool tcg_exec_realizefn(CPUState *cpu, Error **errp); void tcg_exec_unrealizefn(CPUState *cpu); =20 -/* current cflags for hashing/comparison */ -uint32_t curr_cflags(CPUState *cpu); +/* + * Current cflags for hashing/comparison. Everything that feeds into the + * value is folded into CPUState::tcg_cflags when it changes, by + * tcg_update_cflags(), so that TB dispatch only has to load it. + */ +static inline uint32_t curr_cflags(CPUState *cpu) +{ + return cpu->tcg_cflags; +} =20 void tb_check_watchpoint(CPUState *cpu, uintptr_t retaddr); =20 diff --git ./cpu-target.c ./cpu-target.c index 4783845c9b..50be591acf 100644 --- ./cpu-target.c +++ ./cpu-target.c @@ -24,6 +24,7 @@ #include "exec/replay-core.h" #include "exec/log.h" #include "hw/core/cpu.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 /* enable or disable single step mode. EXCP_DEBUG is returned by the @@ -35,6 +36,8 @@ void cpu_single_step(CPUState *cpu, unsigned flags) cpu->singlestep_flags, flags); cpu->singlestep_flags =3D flags; =20 + tcg_update_cflags(cpu); + #if !defined(CONFIG_USER_ONLY) const AccelOpsClass *ops =3D cpus_get_accel(); if (ops->update_guest_debug) { diff --git ./include/system/tcg.h ./include/system/tcg.h index 7622dcea30..2c2dbc753b 100644 --- ./include/system/tcg.h +++ ./include/system/tcg.h @@ -17,6 +17,18 @@ extern bool tcg_allowed; #define tcg_enabled() 0 #endif =20 +/* + * Recompute the parts of CPUState::tcg_cflags that TB dispatch consumes b= ut + * tcg_cflags_set() does not provide: gdb single-step, one-insn-per-tb and + * the CPU_LOG_TB_NOCHAIN log flag. Call whenever one of those changes. + * + * tcg_update_cflags() updates one CPU and must be called from that CPU's + * thread, or with it stopped. tcg_update_all_cflags() updates every CPU + * and is safe to call from the monitor while the vCPUs run. + */ +void tcg_update_cflags(CPUState *cpu); +void tcg_update_all_cflags(void); + /** * qemu_tcg_mttcg_enabled: * Check whether we are running MultiThread TCG or not. diff --git ./monitor/hmp-cmds.c ./monitor/hmp-cmds.c index 4e8d996dbb..b83551ea54 100644 --- ./monitor/hmp-cmds.c +++ ./monitor/hmp-cmds.c @@ -39,6 +39,7 @@ #include "system/hw_accel.h" #include "system/memory.h" #include "system/system.h" +#include "system/tcg.h" #include "disas/disas.h" =20 /* Please update hmp-commands.hx when adding or changing commands */ @@ -335,7 +336,11 @@ void hmp_log(Monitor *mon, const QDict *qdict) =20 if (!qemu_set_log(mask, &err)) { error_report_err(err); + return; } + + /* CPU_LOG_TB_NOCHAIN feeds into the per-CPU cflags. */ + tcg_update_all_cflags(); } =20 void hmp_gdbserver(Monitor *mon, const QDict *qdict) diff --git ./system/runstate-hmp-cmds.c ./system/runstate-hmp-cmds.c index 02d1d42bf3..86754a37f8 100644 --- ./system/runstate-hmp-cmds.c +++ ./system/runstate-hmp-cmds.c @@ -22,6 +22,7 @@ #include "qapi/qapi-commands-run-state.h" #include "qobject/qdict.h" #include "qemu/accel.h" +#include "system/tcg.h" =20 void hmp_info_status(Monitor *mon, const QDict *qdict) { @@ -64,6 +65,9 @@ void hmp_one_insn_per_tb(Monitor *mon, const QDict *qdict) /* If the property exists then setting it can never fail */ object_property_set_bool(OBJECT(accel), "one-insn-per-tb", newval, &error_abort); + + /* one-insn-per-tb feeds into the per-CPU cflags. */ + tcg_update_all_cflags(); } =20 void hmp_watchdog_action(Monitor *mon, const QDict *qdict) --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807054; cv=none; d=zohomail.com; s=zohoarc; b=VpFSijAk7mDLZysc17mDqOCGtE+8GC/fTgdtVdgGYY0lyOjK8sH/wkzDAMH/o3PlKf5BL4J55Z+5jmW1AFIEyOXE5X7OithHRSfD6rfTI4/N+UGOzMjghtov0AfOl7G/hUyuAiDrXEAzO4aFvSbcs40w9UU4AvqABXOKYcQLMyU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807054; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=ZLy+Cn9PnVzddXZFMa8JH1IMrm1W9LnlWo+fFcdWXluXxx7nj/UygFBOKCFz39D8eQRNRtOa7NOLzCeo161xeTq+LcZaEnMkuvmq2xhnqsaykW1MD1Jpo4gH6GiGwRRemu2TaIpTXemqVBoJbnCC3goMD30U+qrCFCRMiw1jf48= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 178780705497014.258746700118422; Wed, 26 Aug 2026 22:04:14 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGX-0005pj-NR; Thu, 27 Aug 2026 01:03:05 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGS-0005od-Us for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:00 -0400 Received: from mail-yx1-xb131.google.com ([2607:f8b0:4864:20::b131]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGR-0002I8-6D for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:00 -0400 Received: by mail-yx1-xb131.google.com with SMTP id 956f58d0204a3-66b32bb75beso2652676d50.2 for ; Wed, 26 Aug 2026 22:02:58 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b5bc34f86sm4317827b3.8.2026.08.26.22.02.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:02:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806978; x=1788411778; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=FoH7vIoRz3/R4+Nu0rNV1SaR/A2/BXmU0tsDUCHXFoOnofXlumrXNx4DjvTUssn2lt tmKs/nSOgFbOscSc6LabJlhu1xop7FvUhOLO4XzzZRQSsJ1hFggihvfKYgAmKONJlIOs T8FpympBl0DVyAjdJKlwZdnF5yKeEPMt9uPYXjhbymzmASwJobPORyHnfoWvl6Vu0Y1g 3Pz8favWZOOCjg3TFN0tLMGnviRHoqGk/VcNTFNIzAxPhGOeRpKhLz5Fu88jO+CYYw9Q xR4/f4B59E/pvpSm7Vi2W9SuJyxxX02VQtbDxgD81ANnVEq5KEeQezXku2TUDseVd6Gs nQeQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806978; x=1788411778; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=FMsgdHjcAKLAQijnunxqj//+RjHxvXxFpsJD23eYSRk0PjTAQraZpZA1ksdPhzAeFh DklsMP++Q98bQ6jSb/lsbJ5CenJD2GwfuKOwa9QB4mDvl0fG9ojvMmbNmYYevGreFwVL WqLBxIEwOhNLi/WEY4MIBqsGKL2dpsMVzVnYN+nf+3Y/X75z9rHcwE07ur1F+hszokqn MgdvRwMly+H0fqsCEHj8i+On03B6xgHQDrggdhFKRSpOL9zdOGHBpzJHimHjH6tytq8m CiSvJo67q63pxR0W51NcoD1mA4ilIPBdej5JhGtLdTr5ppG+BNovQVx1v2i6N8jgBJl0 vILg== X-Gm-Message-State: AFuF++mjOuM76yaOZdD3JzMc6NdZKouSIjgonzB98Nc+EehQzaVFuBm1 NRwkTURTwKVO6ijKba1+4+nFnpdTgeCs74W5VT/6DT3AmmMiWQ434tO6iX0PlaUC X-Gm-Gg: AR+sD13ULDM/qSnjUV6t5abdBbuBzt6fJMr7FDPPX/sijRajOF7LKtP2SRgo01Kk7N9 iUiF7E67RkmeDP1woUNE4P1iHZOe5cP2t2g8QBpZhoqMMgt1PH+PRjo81pGidTkoW+FmpmxmjZg 6iGxSiO5XGs+uqRmh2OgdQoLFvhEdSN9sNYbcYlhXMB0vy96fFg9GsRaAx0XHywaPMVkM8dHikg eTvNtgsVtypn/SslVhUN2+zWmLIiKPhQbIV/oKR0zuOOuF8O5j8RUkC+9eoOwZwvmHatmWp3e0S RQSyUtZwKvjEb6EixGfQdpAIPSmbbOGRUdZgdczNmorY1KKk5/TWvfpDRqo/2PsxhuEVjlYnwzN byJVQ6W3NXk7oFZ/fUOZ0NFm+KkOZ1gb6Aw6TSEl261uGQq3jT0fIMmFGqLoBTFih5IZDeyeG2W 0C8Mjtfm1Oq2PydokLEKLv2Jr8fsD1oxU3wK/db45AoGkFcIWRK11ILieanRtPsB2Xzyufo6Fkw 6wwv1oEcnXG6ZXoXzZtD+b60ArJ9JyHSPNXv1OD X-Received: by 2002:a53:d5ca:0:b0:667:f400:e822 with SMTP id 956f58d0204a3-66d256acceemr4053961d50.3.1787806977905; Wed, 26 Aug 2026 22:02:57 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 2/9] accel/tcg: enlarge the TB jump cache to 64K entries Date: Thu, 27 Aug 2026 01:02:34 -0400 Message-ID: <20260827050241.3713332-3-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b131; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb131.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807056214158500 Content-Type: text/plain; charset="utf-8" The per-CPU TB jump cache has held 4096 entries since it was introduced. That is too small for guests running large programs: an emulated compiler misses often enough that the fallback qht lookup shows up prominently in a profile. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host. The compile performs 34.2 billion TB executions, of which 8.4 billion take the indirect dispatch path. Sizing curve, on top of the preceding patch, instructions retired and wall clock: 12 bits ( 64 KiB): 1,562,204,796,597 132.58s 14 bits ( 256 KiB): 1,493,318,515,396 -4.41% 124.67s -5.97% 16 bits ( 1 MiB): 1,469,772,951,575 -5.92% 121.04s -8.71% 18 bits ( 4 MiB): 1,462,309,832,762 -6.39% 119.82s -9.62% 16 bits is the knee. 18 buys another 0.47% of instructions for four times the memory. It does show a further 1.01% of wall clock, which is outside the 0.70% run-to-run spread at 16 bits, so the effect is probably real -- but paying four times the memory for it is a poor trade, and instructions retired does not account for the data cache pressure of a 4 MiB table. In a perf profile the mechanism is visible directly: tb_htable_lookup(), which is where qht_lookup_custom() lands once it is inlined in an LTO build, falls from 6.10% of samples to 1.66%. The cost is memory: the cache grows from 64 KiB to 1 MiB, once per CPUState. In linux-user that is per guest thread rather than per process, so a threaded guest pays it as many times as it has threads, exactly as system emulation pays it per vCPU. The allocation is g_new0(), so the pages are faulted in as the cache is touched and a thread that runs a small amount of code touches a small part of it, but the address space is committed either way. So this may still want to be tunable, or scaled from the number of CPUs, rather than raised unconditionally. I do not have a threaded workload where the smaller cache is the better trade, and would welcome one. v4: Fix the claim that a linux-user process is a single vCPU. The cache is per CPUState, and linux-user creates one per guest thread. Pointed out by Richard Henderson. Signed-off-by: Matt Turner --- accel/tcg/tb-jmp-cache.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git ./accel/tcg/tb-jmp-cache.h ./accel/tcg/tb-jmp-cache.h index c3a505e394..268dacd7ba 100644 --- ./accel/tcg/tb-jmp-cache.h +++ ./accel/tcg/tb-jmp-cache.h @@ -12,7 +12,7 @@ #include "qemu/rcu.h" #include "exec/cpu-common.h" =20 -#define TB_JMP_CACHE_BITS 12 +#define TB_JMP_CACHE_BITS 16 #define TB_JMP_CACHE_SIZE (1 << TB_JMP_CACHE_BITS) =20 /* --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425773; cv=none; d=zohomail.com; s=zohoarc; b=D5P2uLdOouz7UvRt2Ij6OUk8zqJSjFdHpdZWWnc0rpA3kPdNbXISEHMENil5TAZxalHApMOOoHD50oiTD8dC7Fjl4wkMA9sXVavaIwi50qTg/zX3RlAc4ErsbPUH7FX3aEitIQG1UoT8Opy05NpLE5Hj7YFzQ6Ir6y2ZlCOZBIo= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425773; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=j3mECKD7SIh7xN5kiEZe0QA6kTy9Z31/NtbGK372goo=; b=Fh7L+S0Hko53L6pvSWG88WbPWH7RA4Due+g+xY5m+dcrg8HcEbdmj34R580XLO+pPIl6BzDS5Gw23jLGJQpQ5DGdrSpGw+fahSB6Be1BB42A4uTwdQqtFNyMuKYpLnGW0KefBsrnGAYae/2UshTHkhVpPiaymusamRl8+9WYrU4= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787425773265287.1260754787179; Sat, 22 Aug 2026 12:09:33 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr4y-000425-5D; Sat, 22 Aug 2026 15:08:32 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr4w-00041c-Uw for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:30 -0400 Received: from mail-yw1-x112c.google.com ([2607:f8b0:4864:20::112c]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr4v-0005kv-Cf for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:30 -0400 Received: by mail-yw1-x112c.google.com with SMTP id 00721157ae682-836eac5682bso39449467b3.0 for ; Sat, 22 Aug 2026 12:08:29 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84ca62f6ad5sm13393687b3.18.2026.08.22.12.08.25 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425708; x=1788030508; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=j3mECKD7SIh7xN5kiEZe0QA6kTy9Z31/NtbGK372goo=; b=tOQVW8y1FAkWWd6b4de2d6psUUWZBfyFJN7iiGPoGG3AulobRvjqILhhJFWp3Uijl8 YG567vdg6OufGaCfjGaYnc0ceyVfJUYBbgi5F+ocUlf0HB3j/RJ429UkprHcioXZ9p8e PXm7BlJcOhR11MrQ7nTR1mTAhWm4bHOf+/EQpKwBGvyzoQCILzIWrHsoSv086zEUiaW5 NwwqNBNz00KFUfeK8vWhbt+Aum5wVaHDMs/aECGgmSLBMZJvus+YbbVmwjfZ7oDJ+rMh 45RTlBVe4Xw5/GAOOMAaA8tm9Onl5vCW1cdk88PqSbCHyL/OzBgPdDSDGREObh0eMLbv Hsuw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425708; x=1788030508; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=j3mECKD7SIh7xN5kiEZe0QA6kTy9Z31/NtbGK372goo=; b=n+wSuXcIBB0GoQM85jpw1DC8SihiqOM3n+fjrUGmEvOr6dW8/2apPTOL0kvapSAh6A thQzbwre1JEqOcTkVH8xdHcROI+2Oc6R4JaecPWkuJzUERNLvs1KFJNoqD4nMfT57XSt BdU8vIDmAhL9r7IEA3ZLiEAdLTrGjBBu7aKu4GLtGibIPgShDHxAIqytKtxl/tvlxdua /9p89MN2K7Hd5Xwl4XzCvz217ZjWuF5uLIzSs8PI0PdISFRgKHozTu2x08Ii1u/4KAwB NUUCUeG3UQTLqCQLf2znnwVwWYY1coRIv9iFsRYAAYPvYYrzJ22UdBl0n2wWAlJyTK9h KR2A== X-Gm-Message-State: AFuF++l9VJ1JSw4+sY+x0dth+4Imne/3CvZN0NreDeMkTDy051qFzujY fJjPdHCpArFDUScSmX8Ogj1Tfnfob4ELaoYfEgKZF5+ekOLf48CsljXD8+6JL3/zZq4= X-Gm-Gg: AR+sD126tFkm+S3T9b6TFje5MeVxuNtwWDh0TukKQyYxpq7RyY48+LuZ8BrtH4I0XoR imAG2pWh+9wkDVWelPPrpuOYEjpTfDuDrT9vervtHQK4Ljy6f303K0BDJtcpXkSwhYb9MBEKMLk rQ/pt2jB7JUWjBPVBROpg3RwkkKc61PIMNYkRAokF/u3FH+6t05s+i38BEwsgGs+dhvpzrA83Tw jlG1C72MP5mQCLAE9fg1/kumWWCk7ZQ8pM4QPoi/iwuZzr4qhIhkyvv2uz2xR/6j2ZQikw2P23d oIITsgZm3Mv5VUaTtWYJ42Gz2t3cpK8aIrB+RbNf2AZKAGFaB9OerQIzziRGyKbR9HvPdbtfEPw Z5IflO/tsynOHw/TBJXFToQI8Ed0y+TF+rg+nAZFO92NRuQQpHijyh33dHytE1RVB/W9f6/ZmuK oz7X/kLH10VgL3nmMYshAR9NEIf7576/r9XPC9/Br3IL4xOlX+90gIy3coyJzelq3JfUZh9xnOz eYGkRe+oECgdba7IWtBU2xjlebAe4Aj7qjS3zD6+4n4wQ== X-Received: by 2002:a05:690c:e142:10b0:7ef:9fd9:db07 with SMTP id 00721157ae682-84a199d646bmr39517897b3.12.1787425708068; Sat, 22 Aug 2026 12:08:28 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 2/7] accel/tcg: enlarge the TB jump cache to 64K entries Date: Sat, 22 Aug 2026 15:08:13 -0400 Message-ID: <20260822190818.1829249-3-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112c; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112c.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425774880158500 Content-Type: text/plain; charset="utf-8" The per-CPU TB jump cache has held 4096 entries since it was introduced. That is too small for guests running large programs: an emulated compiler misses often enough that the fallback qht lookup shows up prominently in a profile. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host. The compile performs 34.2 billion TB executions, of which 8.4 billion take the indirect dispatch path. Sizing curve, on top of the preceding patch, instructions retired and wall clock: 12 bits ( 64 KiB): 1,562,204,796,597 132.58s 14 bits ( 256 KiB): 1,493,318,515,396 -4.41% 124.67s -5.97% 16 bits ( 1 MiB): 1,469,772,951,575 -5.92% 121.04s -8.71% 18 bits ( 4 MiB): 1,462,309,832,762 -6.39% 119.82s -9.62% 16 bits is the knee. 18 buys another 0.47% of instructions for four times the memory. It does show a further 1.01% of wall clock, which is outside the 0.70% run-to-run spread at 16 bits, so the effect is probably real -- but paying four times the memory for it is a poor trade, and instructions retired does not account for the data cache pressure of a 4 MiB table. In a perf profile the mechanism is visible directly: tb_htable_lookup(), which is where qht_lookup_custom() lands once it is inlined in an LTO build, falls from 6.10% of samples to 1.66%. The cost is memory: the cache grows from 64 KiB to 1 MiB per vCPU. That is easy to justify for a single-vCPU linux-user process and less obvious for system emulation with many vCPUs, so this may want to be sized by target or made tunable rather than raised unconditionally. Signed-off-by: Matt Turner --- accel/tcg/tb-jmp-cache.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git ./accel/tcg/tb-jmp-cache.h ./accel/tcg/tb-jmp-cache.h index c3a505e394..268dacd7ba 100644 --- ./accel/tcg/tb-jmp-cache.h +++ ./accel/tcg/tb-jmp-cache.h @@ -12,7 +12,7 @@ #include "qemu/rcu.h" #include "exec/cpu-common.h" =20 -#define TB_JMP_CACHE_BITS 12 +#define TB_JMP_CACHE_BITS 16 #define TB_JMP_CACHE_SIZE (1 << TB_JMP_CACHE_BITS) =20 /* --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807048; cv=none; d=zohomail.com; s=zohoarc; b=jMMJT+q76A38pf/LoOix0mxcCIySISbFWHdOH0zcAyYow7KUnp1Gp6zBxTtOW+R0KGACpriwloOp/K1GffjhUHJJhVEhBC2K9kPfnMBlyYC9S3VMB9qCSTd5s0L/+JXFWXIPIfmdtgpISfJZbiT0liubslXzq7Rff+aBVTORRdI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807048; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=A04acGkxXJevk4ozrOtszJaOKo97LvcuoAyqcn0iCpgzgK1UcJeGm+KCREVZJ5r0c+6ruMFAZBfXAkv+DABxo4auGJrABgedLMQQopAVPyeMn+xpdZqKJIzqD6TGJve/bUrBTgPOjDjVKPlROhCCU+2C38RVjuNeytOVUFu5xL8= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807048989603.3306974565028; Wed, 26 Aug 2026 22:04:08 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGX-0005pn-Ne; Thu, 27 Aug 2026 01:03:05 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGV-0005pG-G7 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:03 -0400 Received: from mail-yw1-x1130.google.com ([2607:f8b0:4864:20::1130]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGT-0002Ij-RG for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:03 -0400 Received: by mail-yw1-x1130.google.com with SMTP id 00721157ae682-8585c16ab5aso4602327b3.2 for ; Wed, 26 Aug 2026 22:03:01 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b5bc34efasm4267567b3.7.2026.08.26.22.02.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:02:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806981; x=1788411781; darn=nongnu.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=qV9zbJMhgOHzlGynQt2Lh+gHUtOPxB63IvJAqGgKPlYVCX3wvV6WvnVAQH4yLx638+ qqrBUMchMZDJWT/8DIltdDGx5STx0wauGe1JxYB5rVRAfed+hXo4PFp3lmWXhihpWDku qarREQ2FypHFgsVlcIT3VQKeNyGo632LXUR1oAzPqlxqcu3XnpHoPuywe7GVNkxYWAB1 QNrMSbWv8RUKblgwKaFJnBie7Hv92/C3bNRPUBB0XNGda1TyEfScgMgHsE4OEGQQlgeQ YdMJ/0t/KmHBw60/ZM8SsC8DWPPzlv4aCzx31WS9vFiVBlKTeu4hU5ML2e3I8t49Rbk4 QGxg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806981; x=1788411781; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=SO53SjwCKwFSjYa+TzZm+qAn1Xj0R7zmPlqYq4ZiE/7aRi3cQqAQXXcTaAds8cKuGV 6QZ+P2iblgWLEBBEc+eE2KdIKui3IeGJEwrvkt74w43hsU9W4IURvHDaZ50RWIiImXKI cL0eHpRf7vCoOz7r2L0BccnwYz1+8c5pxgFgS1MxXjVxRrx5CCulwF3rgoBnOsZxVMsD BzWwsc56daqt9ctZz6zc84TtSJLQwzpcwCX4nUSEOksYJV7hBwd+6upMZLr75L9FHuV+ avVSbxoG36ksB5zg19MMnAoIiJljDKkF2M7sWpFmZcx69DndsCRx688YABR4GCt0jrFB XX+Q== X-Gm-Message-State: AFuF++nfHfdRcsF4pUQmyneIW6IYJ3Q2v4m4Y4W0kzwlsnGhyVPWPfA9 DUM3QeDlFcR1mYBuGlfbe2Ck9RlCEht7Hz2/6acA8+Z/olS51GwRX6j9hrk6AGlc X-Gm-Gg: AR+sD10SrQNevDf9OPgJ0bV9wOoO/ar7kFxMuv21ho2ZPFJc9QdvTyALzNBBILLkvtU ynMhhM3g3H/+6xXv0RXU+IC7NFvWV2ojdvcmCkrkzJouszdw3CwBofc/8g9b5faoNFuJnUSiNDV bdL8Y97HSwsRPmQOCAnRLOMehZ+edAm62nbMGDIrv+HRwjyvejwroHvuAu7ttcR8YxZr7Y4MgY4 obfEloEgcvMW9iZaYGC1oDd+OrTcmbXlVCk95/NRqMSYj+Wy29rOAy40XnQsNcV9eaKs0Bkb6KK GKug0gQTD9CKSRAa8B5v1wV6MNa4QzNfcVqcMzhEyjuIPV5Q8/9PToMeFDB/TSrY/JM0iyNJFb0 BgTQOJlrcvH6Tmg3JfbRRC3ONV3EyrTatrm5rScNY+U3VsCoj3pqrPldyfD3pF6qJqVbgAPoNwJ QakdGfIH5TBGDWgIuViJtJzG2h11ADYtpkno6eIv6kycMhuORzNxhgl6d1xf/3yw1SCuSeGjPNl sVsdomtdtZ7pqRuFMy9lHO1PdRVxE2t24aH6wGu X-Received: by 2002:a05:690c:399:b0:841:e9e4:8ab9 with SMTP id 00721157ae682-857402a4134mr53042687b3.25.1787806980573; Wed, 26 Aug 2026 22:03:00 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 3/9] accel/tcg: skip the can_do_io stores in user-only builds Date: Thu, 27 Aug 2026 01:02:35 -0400 Message-ID: <20260827050241.3713332-4-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1130; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1130.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807050009158500 Every translation block stores to cpu->neg.can_do_io twice: false before the first instruction, true before the last one. Nothing reads it in a user-only build. There is no memory-mapped I/O in linux-user, and every reader is in system_ss: cputlb.c, watchpoint.c, icount-common.c and tcg-accel-ops-icount.c. Two stores per TB is not much on its own, but TBs are short. An emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) executes 34.2 billion TBs at 6.04 guest instructions each, so this is 68 billion stores for nothing. Measured on an x86-64 host, LTO build, on top of the preceding two patches: before: 1,469,772,951,575 instructions after: 1,402,667,803,616 instructions -4.57% before: 121.04s wall clock after: 115.56s wall clock -4.53% The emulated compiler produces byte-identical output. v3: Use #ifndef CONFIG_USER_ONLY again rather than if (IS_ENABLED(CONFIG_USER_ONLY)). QEMU's IS_ENABLED() is IS_EMPTY(), which is only true for a symbol Meson defines empty; CONFIG_USER_ONLY is defined as 1, so the test was always false and v2 emitted the two stores after all. The measurements above are from the working form. Reviewed-by: Richard Henderson Reviewed-by: Philippe Mathieu-Daud=C3=A9 Signed-off-by: Matt Turner --- accel/tcg/translator.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 57daded60f..6c8fcd7a20 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -21,12 +21,14 @@ #include "disas/disas.h" #include "tb-internal.h" =20 +#ifndef CONFIG_USER_ONLY static void set_can_do_io(DisasContextBase *db, bool val) { QEMU_BUILD_BUG_ON(sizeof_field(CPUState, neg.can_do_io) !=3D 1); tcg_gen_st8_i32(tcg_constant_i32(val), tcg_env, offsetof(CPUState, neg.can_do_io) - sizeof(CPUState)); } +#endif =20 bool translator_io_start(DisasContextBase *db) { @@ -210,17 +212,25 @@ void translator_loop(CPUState *cpu, TranslationBlock = *tb, int *max_insns, /* * Manage can_do_io for the translation block: set to false before * the first insn and set to true before the last insn. + * + * Nothing reads can_do_io in user-only builds. There is no MMIO + * there, and every reader (cputlb.c, watchpoint.c, icount) is in + * system_ss, so skip the two stores per TB entirely. */ if (db->num_insns =3D=3D 1) { tcg_debug_assert(first_insn_start =3D=3D db->insn_start); } else { tcg_debug_assert(first_insn_start !=3D db->insn_start); +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D first_insn_start; set_can_do_io(db, false); +#endif } +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D db->insn_start; set_can_do_io(db, true); tcg_ctx->emit_before_op =3D NULL; +#endif =20 /* May be used by disas_log or plugin callbacks. */ tb->size =3D db->pc_next - db->pc_first; --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425745; cv=none; d=zohomail.com; s=zohoarc; b=TZuPmQzDR7pVwb9TgxzjxPHE2EwIMjJ1raMDlnE8MdY6ShngCl+rO8IJnqzRcQQHoCi9aoVNlCqr2tzTYPhCdlVjGcCadtixXIh9GcJn154vjRUxWYTtkgRiZYcOgXUz4kcTXuqnQSXvD/KMxIqtVTI32IorzOKCwLkK9S6z0AQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425745; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=Qm9lS7Om247BNcUvVpk+qh0av0gdYX84kx3RG6dwG6XU7XvN5bw6oJlyxsS5ovNlHDP9KtE36/PXEf3yrej8PVwjy4H8yqgofkFfXDGKoBTFc/QN4ciEnMcFStiigNONrI22xhkODsbKE4+arfx83NmB41RCkPlLLCHLiJIMNjw= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787425745285655.006863298901; Sat, 22 Aug 2026 12:09:05 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr51-00047D-VA; Sat, 22 Aug 2026 15:08:36 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr50-00044c-1g for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:34 -0400 Received: from mail-yw1-x112d.google.com ([2607:f8b0:4864:20::112d]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr4y-0005lk-CO for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:33 -0400 Received: by mail-yw1-x112d.google.com with SMTP id 00721157ae682-81f64e8dfbcso32052717b3.2 for ; Sat, 22 Aug 2026 12:08:32 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84cab854483sm13222017b3.33.2026.08.22.12.08.28 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425711; x=1788030511; darn=nongnu.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=L6IgcZQk+sAoLDyzk9uG8nbvrDJ4yGicKQxlZg8wiaOin9ABIPTXTxm3ksoLk4AdJL cmD/FnURoLhCkrCf6QoeQxDCuxM5gXIqZgLHXbBZGReje/FyvivsMi90oofNhVfSd34f iSdWk7BP8rW0hMh5J7FmgIjGRir3DzJWhV9Txw7k1dO0sCFwftytpTKayJA/EnUcr/nk HO04sISyslhCZCRkNO611NEpYJDUx71N5MzsXtG/opSrLYso/cCZVjFTV5tVsgyiIQ0y nNVh4JIwUUz6UN3jCkVtqJcdnrE6gBNXORSkYVfD3W8Fq3tTKKMK/iQPES7h6C29AIl2 gywQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425711; x=1788030511; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=TiFw1+cLMc9KdbbEqRaoCoqjo5a8XonLHNmVsy3+eyoq73yhLb13bCgjTAaX/FFRr4 MbLg87woJjCcTl9C768x+TuBcrYZMGaDck2o4B/t5jjNC4S848NEQ19XfuwQd2loU4UW iPB/kECQu83ECt2ZwAbhwzhBl5IhxunXZhkdEQjNMuFNsjma/jmZdUHvr1zSQfCNe7pI mBds2GfYGnHfqsbVcOCxhvbomzBqdtZWS4uatvcs/wNL+WdmquxCCT1X+TJ6cNCHWcKG nsIDsW108YEIEPwv4w/Xo2E+GdrF/Nsil7PWkZJcEudNU6hKWHTpOl0/dep8LafeeLKO c/XA== X-Gm-Message-State: AFuF++lnL6qAG40c2ccba0RewS0+W4QEv0K2m/2ag4v02EWnJseTYo7X St5fAnnzF294/ZdxMASRcAGyDoOxQKQ+tOVpdC3R5/uHfI3rLd8CN/dJ4x4ZQ7mPOmY= X-Gm-Gg: AR+sD13lXUN89kG4rsi1lGEfkgV6P0pFuZc3qOcDQ0B+yju4NhsZAW9u41xQNBpK0/B 8w6pYJBpyoIgBC4bqnObNFM38laVmCxDBFQOHrKfkLamxxlRWwW5OJooteNvRlmjPtXR0SjuM+Z nDnWxhGKDhcu3YZXgGtXEp9sqD/ZEz5FoIKeumXD4y3HCqIvm980ffWi8G6PAM2aLY4+bYiAL9g 3NUsZ/Rsb9WuYifNJlgYxVWKKiD4LRsjx4G/028FOEUIxjFme+E/nUoGhGaXXSZJ5Xi/RNV/pzP RiRnBxI8Ku2pIeuUNQogZHC5CkXsbms84rddPNt+AOisodt/g/vopKJvn35tUMZMIIgGfSAoT5/ 6cRoXe1DvBQAOP6jkCLCzkyKiljviGUw2XAMqn2azltySNvxAM73bGARryVQevamSF/Ytl2+cb6 O5C2MRFk95vNNhVSY99uYDb2I8Yo/WiVrzZgg0gxn54F0DF/lY8ydqTE8yjPM8Ekz+WAPJ1nbKX p8OPpzuQBnI2ybkJPr15ERWg0cBcKclHHSfThBg X-Received: by 2002:a05:690c:e1c4:10b0:820:1281:8deb with SMTP id 00721157ae682-849f08113c4mr51716717b3.8.1787425711249; Sat, 22 Aug 2026 12:08:31 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 3/7] accel/tcg: skip the can_do_io stores in user-only builds Date: Sat, 22 Aug 2026 15:08:14 -0400 Message-ID: <20260822190818.1829249-4-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112d; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112d.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425746709158500 Every translation block stores to cpu->neg.can_do_io twice: false before the first instruction, true before the last one. Nothing reads it in a user-only build. There is no memory-mapped I/O in linux-user, and every reader is in system_ss: cputlb.c, watchpoint.c, icount-common.c and tcg-accel-ops-icount.c. Two stores per TB is not much on its own, but TBs are short. An emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) executes 34.2 billion TBs at 6.04 guest instructions each, so this is 68 billion stores for nothing. Measured on an x86-64 host, LTO build, on top of the preceding two patches: before: 1,469,772,951,575 instructions after: 1,402,667,803,616 instructions -4.57% before: 121.04s wall clock after: 115.56s wall clock -4.53% The emulated compiler produces byte-identical output. v3: Use #ifndef CONFIG_USER_ONLY again rather than if (IS_ENABLED(CONFIG_USER_ONLY)). QEMU's IS_ENABLED() is IS_EMPTY(), which is only true for a symbol Meson defines empty; CONFIG_USER_ONLY is defined as 1, so the test was always false and v2 emitted the two stores after all. The measurements above are from the working form. Reviewed-by: Richard Henderson Reviewed-by: Philippe Mathieu-Daud=C3=A9 Signed-off-by: Matt Turner --- accel/tcg/translator.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 57daded60f..6c8fcd7a20 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -21,12 +21,14 @@ #include "disas/disas.h" #include "tb-internal.h" =20 +#ifndef CONFIG_USER_ONLY static void set_can_do_io(DisasContextBase *db, bool val) { QEMU_BUILD_BUG_ON(sizeof_field(CPUState, neg.can_do_io) !=3D 1); tcg_gen_st8_i32(tcg_constant_i32(val), tcg_env, offsetof(CPUState, neg.can_do_io) - sizeof(CPUState)); } +#endif =20 bool translator_io_start(DisasContextBase *db) { @@ -210,17 +212,25 @@ void translator_loop(CPUState *cpu, TranslationBlock = *tb, int *max_insns, /* * Manage can_do_io for the translation block: set to false before * the first insn and set to true before the last insn. + * + * Nothing reads can_do_io in user-only builds. There is no MMIO + * there, and every reader (cputlb.c, watchpoint.c, icount) is in + * system_ss, so skip the two stores per TB entirely. */ if (db->num_insns =3D=3D 1) { tcg_debug_assert(first_insn_start =3D=3D db->insn_start); } else { tcg_debug_assert(first_insn_start !=3D db->insn_start); +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D first_insn_start; set_can_do_io(db, false); +#endif } +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D db->insn_start; set_can_do_io(db, true); tcg_ctx->emit_before_op =3D NULL; +#endif =20 /* May be used by disas_log or plugin callbacks. */ tb->size =3D db->pc_next - db->pc_first; --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807043; cv=none; d=zohomail.com; s=zohoarc; b=kBOe0pcPhA79c1AH7Lx58hmb9JnZ2AN4uOa15MYOwlU+tS9cHY0cAmqhN+2ENvEwBPkWFXnH4k8T/GiDp6mlDGtTl8H3tZujCIOsDSi0XT6yHKgWBdYWrLsvV7d59PWir/y76P9lVrzAqFUyOE2Vn8ooqxTc/2SNTWEQKbZIpsQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807043; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=PtS9C6criMikEeHhjOAa2GbahJameLBPB99rYyPv+nQ=; b=N2rFygv8STPx26Vnxpj/+MRfeKtVrJgpMzctvWF8H6zyiDqF+Qdg3+JgMKchS9fOXbOLX6czBbZkPyv2Vnr2SKNR65LL8PFEjK9BsfMtjS5amnZXD8de8clffFFg13AVI2SlT+RsWja38UwEWG77Ed596T+4FDAM3gTXHnRSnLs= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807043583742.5871719325893; Wed, 26 Aug 2026 22:04:03 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGa-0005qd-Co; Thu, 27 Aug 2026 01:03:08 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGZ-0005qQ-HP for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:07 -0400 Received: from mail-yx1-xb129.google.com ([2607:f8b0:4864:20::b129]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGW-0002JD-Je for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:07 -0400 Received: by mail-yx1-xb129.google.com with SMTP id 956f58d0204a3-66d1442b24eso2242609d50.1 for ; Wed, 26 Aug 2026 22:03:04 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b6143fe16sm4110607b3.29.2026.08.26.22.03.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806983; x=1788411783; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PtS9C6criMikEeHhjOAa2GbahJameLBPB99rYyPv+nQ=; b=rAFEoS/zoyLMld5tHBRcRobuqAteYFK6e6ZNydxJP9g57HOSGfo3+tccGoUWNq8RaG VpP1DtxK3yBJNMIp73KsMQigPd926kthyfHTyScY+OuPE9aVOguSNuSbjg58F/6pHYs4 FVhBKziJy+fVsck2QM7EfC12Y3HItlXRlG4jZ7+Ris14h3h3qqvpR6cuFx+KetGNYpde mzffb3Q64LE75X1bxlTdH6R1RFuOMfI/eXVOIWnvPNdb2ugiEFKBtHBcn075Fo2EyYoh fn7iyGURJ9nLj7VdKe3f2f81kF9aEF/EhrkUn61LktECFa8ZP/vQ/vsie8S6ei63H0G0 LRUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806983; x=1788411783; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PtS9C6criMikEeHhjOAa2GbahJameLBPB99rYyPv+nQ=; b=ARAmaGcEUuzg2y2BVStEitdGseDWNmwXI+QMRMBzcmaHQm7UYIreuCnS5kQbKe42D2 HvSqAjQVfBnjoDvtLydngS/5VNf7dLluFl9OEl/vN8VkY6wY+zmcxPbshDLRDKYir8L4 PHVYshw+Q78yUe8NHyOIg0pFOYGvwpRRDrS+3CqeTdDzKavCyMmWjYbfS65E+O48Fa+R Q5UTMexjaJnBLzJbWerABWKGETj/3Xr5PcawGx7wuezKIA8QdYDVmEmf17mCD+VlaljL YfmidwjTWgRWtp6VMc/L4dYDxEO9BOTCArst66VWUqZW1dqzGPdplP1E9HXZdeJxLZL8 Ki6w== X-Gm-Message-State: AFuF++m1azWTHcery0hNH4BJIhvo8/Z36JUmXw+gVEztP2NPd1R4yGE1 hmSlwmSRCphaBbJbqm8wuAXEe1ZY0KuhVdc4JdC2vD0pKUEXki0qgbnQ9HuwHp/2 X-Gm-Gg: AR+sD10r8vjGs3GsynDEaDztxh+vAAtW2fRp7lrSnA7mmqz3OcI6EF4eJxTOEby9Gn0 pjY5njb6b+rajdDX9FCVN//BxA7ruTejYhVwhlPAyQDykeiXsqks7dAq/uxKi96NmOZslLkhIpn de86mLIUNkKXwvOcboZLpi4eeJQagaWvro74heaJil7okfde62zLv8+kkQHswOWzEQtQerpKHKK ZxfFxZ0CnBpU6+OwJtWeXh3wifdFfinSO2F1/Z7flsk5/FwRdIIks2aagoxxFzcsoXftI0WLFkg 17091ItzisBezg/T/KP5bDzHQSwG9Dp4X81Q0YAgD9OWXp0nEdcQMeP960ZSa5AsS3O2psxV6Rt O3dCq5ecjHWHXoHAm6Z3b7Wot5ExhL9YlNq9RDznXJBfkr/XUGUlShM2sjfEekzngBD0zME+xAZ TFD4mdgkEmLJIxe4Rs/7t/tAFHVgW3fiwCZX9EwhDVBD2GNIVKbjMKKNndEWfisbOmtp/2hu3/v BZuX55n7CziSvQzoV2WbIEwu2PdOGpNW/KGAITtiwZ0qeaCprg= X-Received: by 2002:a53:ac9c:0:b0:668:9fa5:b9cc with SMTP id 956f58d0204a3-66d2551013fmr4446592d50.0.1787806983083; Wed, 26 Aug 2026 22:03:03 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 4/9] tcg: pass the destination to tcg_gen_lookup_and_goto_ptr() Date: Thu, 27 Aug 2026 01:02:36 -0400 Message-ID: <20260827050241.3713332-5-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b129; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb129.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807046044158500 Content-Type: text/plain; charset="utf-8" tcg_gen_lookup_and_goto_ptr() takes no arguments and emits a call to helper_lookup_tb_ptr(), which recovers the destination PC from env by calling back into the target through TCGCPUOps::get_tb_cpu_state(). At translation time the caller already has the destination PC in a temp, and knows the flags, cflags and cs_base any destination it may reach has to match, because they are the ones the block being generated was translated with. Pass both, so that a later patch can use them to look the destination up inline. Nothing reads them yet and the generated code does not change. The contract on @pc is the whole of the interface: it must hold exactly what get_tb_cpu_state() reports as the pc for the destination block. Five targets keep their PC in a temp whose value is that pc by construction and so can pass it: alpha, loongarch, mips, ppc and s390x. Everything else passes NULL and keeps today's behavior. For six of those the TB pc is derived and passing the PC temp would be wrong: avr's TB pc is the word address doubled, i386's is eip before segmentation, riscv masks it to 32 bits when xl is MXL_RV32, hppa derives it from the IAQ, hexagon adjusts it inside a hardware loop, and sparc puts npc in cs_base. The remaining seven -- arm, m68k, microblaze, or1k, rx, sh4 and tricore -- look like they could pass it, but I have not convinced myself of the contract for them and have nothing to test them with. Each is a one-line change for whoever wants it. The common entry point takes a TCGTemp rather than a TCGv and reads the width from it, because the translators that are built for both values of TARGET_LONG_BITS -- arm, s390x, microblaze -- cannot include tcg-op.h. tcg-op.h wraps it for everyone else. This is the same split as tcg_gen_qemu_ld_*_chk(). v4: Split out of "tcg: probe the TB jump cache inline instead of calling a helper", which did the API change and the inline probe in one patch. Requested by Richard Henderson. Signed-off-by: Matt Turner --- include/tcg/tcg-op-common.h | 15 ++++++++++++--- include/tcg/tcg-op.h | 12 ++++++++++++ target/alpha/translate.c | 4 ++-- target/arm/tcg/translate-a64.c | 4 ++-- target/arm/tcg/translate.c | 10 +++++----- target/avr/translate.c | 4 ++-- target/hexagon/translate.c | 4 ++-- target/hppa/translate.c | 6 +++--- target/i386/tcg/translate.c | 2 +- .../loongarch/tcg/insn_trans/trans_branch.c.inc | 2 +- target/loongarch/tcg/translate.c | 4 ++-- target/m68k/translate.c | 2 +- target/microblaze/translate.c | 4 ++-- target/mips/tcg/nanomips_translate.c.inc | 2 +- target/mips/tcg/translate.c | 6 +++--- target/or1k/translate.c | 4 ++-- target/ppc/translate.c | 4 ++-- target/riscv/tcg/insn_trans/trans_rvzce.c.inc | 4 ++-- target/riscv/tcg/translate.c | 2 +- target/rx/translate.c | 4 ++-- target/s390x/tcg/translate.c | 5 +++-- target/sh4/translate.c | 4 ++-- target/sparc/translate.c | 4 ++-- target/tricore/translate.c | 4 ++-- tcg/tcg-op.c | 3 ++- 25 files changed, 71 insertions(+), 48 deletions(-) diff --git ./include/tcg/tcg-op-common.h ./include/tcg/tcg-op-common.h index 9b321f959c..34102b3b7a 100644 --- ./include/tcg/tcg-op-common.h +++ ./include/tcg/tcg-op-common.h @@ -75,15 +75,24 @@ void tcg_gen_exit_tb(const TranslationBlock *tb, unsign= ed idx); void tcg_gen_goto_tb(unsigned idx); =20 /** - * tcg_gen_lookup_and_goto_ptr() - look up the current TB, jump to it if v= alid - * @addr: Guest address of the target TB + * tcg_gen_lookup_and_goto_ptr() - look up the destination TB, jump to it + * @pc: temp holding the destination guest PC, or NULL + * @tb: the translation block being generated * * If the TB is not valid, jump to the epilogue. * + * The lookup is a call to helper_lookup_tb_ptr(). @pc and @tb describe t= he + * destination for a faster lookup that a later patch adds, and neither is + * used yet. When @pc is non-NULL it must hold exactly the value + * get_tb_cpu_state() reports as the pc for the destination, and the + * destination must match @tb's flags, cflags and cs_base. A target whose + * pc is derived rather than being that key -- avr's word address, i386's + * eip before segmentation -- must pass NULL. + * * This operation is optional. If the TCG backend does not implement goto_= ptr, * this op is equivalent to calling tcg_gen_exit_tb() with 0 as the argume= nt. */ -void tcg_gen_lookup_and_goto_ptr(void); +void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, const TranslationBlock *= tb); =20 void tcg_gen_plugin_cb(unsigned from); void tcg_gen_plugin_mem_cb(TCGv_i64 addr, unsigned meminfo); diff --git ./include/tcg/tcg-op.h ./include/tcg/tcg-op.h index 3721164236..24d567bd2f 100644 --- ./include/tcg/tcg-op.h +++ ./include/tcg/tcg-op.h @@ -49,6 +49,18 @@ typedef TCGv_i64 TCGv; #error Unhandled TARGET_LONG_BITS value #endif =20 +/* + * See tcg_gen_lookup_and_goto_ptr_tmp(). @pc may be NULL, for a target + * whose guest PC is not directly the key a destination block is found by. + * A translator that is built for more than one value of TARGET_LONG_BITS, + * and so cannot include this header, calls the _tmp() form directly. + */ +static inline void +tcg_gen_lookup_and_goto_ptr(TCGv pc, const TranslationBlock *tb) +{ + tcg_gen_lookup_and_goto_ptr_tmp(pc ? tcgv_tl_temp(pc) : NULL, tb); +} + #if TARGET_LONG_BITS =3D=3D 64 #define tcg_gen_movi_tl tcg_gen_movi_i64 #define tcg_gen_mov_tl tcg_gen_mov_i64 diff --git ./target/alpha/translate.c ./target/alpha/translate.c index c66e3f9c14..822f5cc120 100644 --- ./target/alpha/translate.c +++ ./target/alpha/translate.c @@ -449,7 +449,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, int32_t disp) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_pc_disp(ctx, cpu_pc, disp); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); } } =20 @@ -2917,7 +2917,7 @@ static void alpha_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cpu) gen_pc_disp(ctx, cpu_pc, 0); /* FALLTHRU */ case DISAS_PC_UPDATED: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); break; case DISAS_PC_UPDATED_NOCHAIN: tcg_gen_exit_tb(NULL, 0); diff --git ./target/arm/tcg/translate-a64.c ./target/arm/tcg/translate-a64.c index 4f9a93950b..d1dd33a1af 100644 --- ./target/arm/tcg/translate-a64.c +++ ./target/arm/tcg/translate-a64.c @@ -562,7 +562,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_sl= ot_idx, int64_t diff) if (s->ss_active) { gen_step_complete_exception(s); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); s->base.is_jmp =3D DISAS_NORETURN; } } @@ -11250,7 +11250,7 @@ static void aarch64_tr_tb_stop(DisasContextBase *dc= base, CPUState *cpu) gen_a64_update_pc(dc, 4); /* fall through */ case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; case DISAS_NORETURN: case DISAS_SWI: diff --git ./target/arm/tcg/translate.c ./target/arm/tcg/translate.c index c866148383..ca701b9cbc 100644 --- ./target/arm/tcg/translate.c +++ ./target/arm/tcg/translate.c @@ -1306,9 +1306,9 @@ void write_neon_element64(TCGv_i64 src, int reg, int = ele, MemOp memop) } } =20 -static void gen_goto_ptr(void) +static void gen_goto_ptr(DisasContext *s) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, s->base.tb); } =20 /* This will end the TB but doesn't guarantee we'll return to @@ -1336,7 +1336,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_= slot_idx, int64_t diff) tcg_gen_exit_tb(s->base.tb, tb_slot_idx); } else { gen_update_pc(s, diff); - gen_goto_ptr(); + gen_goto_ptr(s); } s->base.is_jmp =3D DISAS_NORETURN; } @@ -1373,7 +1373,7 @@ static void gen_jmp_tb(DisasContext *s, int64_t diff,= int tbno) * and don't chain to another TB. */ gen_update_pc(s, diff); - gen_goto_ptr(); + gen_goto_ptr(s); s->base.is_jmp =3D DISAS_NORETURN; break; default: @@ -6858,7 +6858,7 @@ static void arm_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cpu) gen_update_pc(dc, curr_insn_len(dc)); /* fall through */ case DISAS_JUMP: - gen_goto_ptr(); + gen_goto_ptr(dc); break; case DISAS_UPDATE_EXIT: gen_update_pc(dc, curr_insn_len(dc)); diff --git ./target/avr/translate.c ./target/avr/translate.c index 3c57606097..8f2e0baa67 100644 --- ./target/avr/translate.c +++ ./target/avr/translate.c @@ -992,7 +992,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, tcg_gen_exit_tb(tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } @@ -2778,7 +2778,7 @@ static void avr_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cs) /* fall through */ case DISAS_LOOKUP: if (!force_exit) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; } /* fall through */ diff --git ./target/hexagon/translate.c ./target/hexagon/translate.c index 06a8159d28..cc230b08d1 100644 --- ./target/hexagon/translate.c +++ ./target/hexagon/translate.c @@ -181,7 +181,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, if (move_to_pc) { tcg_gen_movi_tl(hex_gpr[HEX_REG_PC], dest); } - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } =20 @@ -218,7 +218,7 @@ static void gen_end_tb(DisasContext *ctx) gen_set_label(skip); gen_goto_tb(ctx, 1, ctx->next_PC, false); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } =20 ctx->base.is_jmp =3D DISAS_NORETURN; diff --git ./target/hppa/translate.c ./target/hppa/translate.c index 002189ddfb..cf8f1a2c13 100644 --- ./target/hppa/translate.c +++ ./target/hppa/translate.c @@ -816,7 +816,7 @@ static void gen_goto_tb(DisasContext *ctx, int which, tcg_gen_goto_tb(which); tcg_gen_exit_tb(ctx->base.tb, which); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } =20 @@ -2027,7 +2027,7 @@ static bool do_ibranch(DisasContext *ctx, unsigned li= nk, store_psw_xb(ctx, PSW_B); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return nullify_end(ctx); } @@ -4838,7 +4838,7 @@ static void hppa_tr_tb_stop(DisasContextBase *dcbase,= CPUState *cs) } /* FALLTHRU */ case DISAS_IAQ_N_UPDATED: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_EXIT: tcg_gen_exit_tb(NULL, 0); diff --git ./target/i386/tcg/translate.c ./target/i386/tcg/translate.c index 2115c5cd24..66a0ee3cdf 100644 --- ./target/i386/tcg/translate.c +++ ./target/i386/tcg/translate.c @@ -2005,7 +2005,7 @@ gen_eob(DisasContext *s, int mode) } else if (mode =3D=3D DISAS_JUMP && /* give irqs a chance to happen */ !inhibit_reset) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); } else { tcg_gen_exit_tb(NULL, 0); } diff --git ./target/loongarch/tcg/insn_trans/trans_branch.c.inc ./target/lo= ongarch/tcg/insn_trans/trans_branch.c.inc index da07778658..57d9d47353 100644 --- ./target/loongarch/tcg/insn_trans/trans_branch.c.inc +++ ./target/loongarch/tcg/insn_trans/trans_branch.c.inc @@ -27,7 +27,7 @@ static bool trans_jirl(DisasContext *ctx, arg_jirl *a) tcg_gen_mov_tl(cpu_pc, addr); tcg_gen_movi_tl(dest, make_address_pc(ctx, ctx->base.pc_next + 4)); gen_set_gpr(a->rd, dest, EXT_NONE); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return true; } diff --git ./target/loongarch/tcg/translate.c ./target/loongarch/tcg/transl= ate.c index 124dce6269..a45a51852a 100644 --- ./target/loongarch/tcg/translate.c +++ ./target/loongarch/tcg/translate.c @@ -111,7 +111,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, vaddr dest) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { tcg_gen_movi_tl(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); } } =20 @@ -311,7 +311,7 @@ static void loongarch_tr_tb_stop(DisasContextBase *dcba= se, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: tcg_gen_movi_tl(cpu_pc, ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); break; case DISAS_TOO_MANY: gen_goto_tb(ctx, 0, ctx->base.pc_next); diff --git ./target/m68k/translate.c ./target/m68k/translate.c index 138c89d3e5..73691bc0d1 100644 --- ./target/m68k/translate.c +++ ./target/m68k/translate.c @@ -6095,7 +6095,7 @@ static void m68k_tr_tb_stop(DisasContextBase *dcbase,= CPUState *cpu) if (dc->ss_active) { gen_raise_exception_format2(dc, EXCP_TRACE, dc->pc_prev); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } break; case DISAS_EXIT: diff --git ./target/microblaze/translate.c ./target/microblaze/translate.c index 8b219afb5d..851b372f8f 100644 --- ./target/microblaze/translate.c +++ ./target/microblaze/translate.c @@ -127,7 +127,7 @@ static void gen_goto_tb(DisasContext *dc, unsigned tb_s= lot_idx, vaddr dest) tcg_gen_exit_tb(dc->base.tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, dc->base.tb); } dc->base.is_jmp =3D DISAS_NORETURN; } @@ -1764,7 +1764,7 @@ static void mb_tr_tb_stop(DisasContextBase *dcb, CPUS= tate *cs) /* Indirect jump (or direct jump w/ goto_tb disabled) */ tcg_gen_mov_i32(cpu_pc, cpu_btarget); tcg_gen_discard_i32(cpu_btarget); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, dc->base.tb); return; =20 default: diff --git ./target/mips/tcg/nanomips_translate.c.inc ./target/mips/tcg/nan= omips_translate.c.inc index 4b0b01ba37..007e29f9ac 100644 --- ./target/mips/tcg/nanomips_translate.c.inc +++ ./target/mips/tcg/nanomips_translate.c.inc @@ -2406,7 +2406,7 @@ static void gen_compute_nanomips_pbalrsc_branch(Disas= Context *ctx, int rs, =20 /* unconditional branch to register */ tcg_gen_mov_tl(cpu_PC, btarget); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); } =20 /* nanoMIPS Branches */ diff --git ./target/mips/tcg/translate.c ./target/mips/tcg/translate.c index e3467d1525..73abfbb5d4 100644 --- ./target/mips/tcg/translate.c +++ ./target/mips/tcg/translate.c @@ -4374,7 +4374,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned t= b_slot_idx, tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_save_pc(dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); } } =20 @@ -11014,7 +11014,7 @@ static void gen_branch(DisasContext *ctx, int insn_= bytes) } else { tcg_gen_mov_tl(cpu_PC, btarget); } - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); break; default: LOG_DISAS("unknown branch 0x%x\n", proc_hflags); @@ -15244,7 +15244,7 @@ static void mips_tr_tb_stop(DisasContextBase *dcbas= e, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: gen_save_pc(ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); break; case DISAS_NEXT: case DISAS_TOO_MANY: diff --git ./target/or1k/translate.c ./target/or1k/translate.c index eb4485312f..4907284a6d 100644 --- ./target/or1k/translate.c +++ ./target/or1k/translate.c @@ -1605,7 +1605,7 @@ static void openrisc_tr_tb_stop(DisasContextBase *dcb= ase, CPUState *cs) /* The jump destination is indirect/computed; use jmp_pc. */ tcg_gen_mov_i32(cpu_pc, jmp_pc); tcg_gen_discard_i32(jmp_pc); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; } /* The jump destination is direct; use jmp_pc_imm. @@ -1622,7 +1622,7 @@ static void openrisc_tr_tb_stop(DisasContextBase *dcb= ase, CPUState *cs) break; } tcg_gen_movi_i32(cpu_pc, jmp_dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; =20 case DISAS_EXIT: diff --git ./target/ppc/translate.c ./target/ppc/translate.c index 06ed2adf10..42924281b0 100644 --- ./target/ppc/translate.c +++ ./target/ppc/translate.c @@ -3664,7 +3664,7 @@ static void gen_lookup_and_goto_ptr(DisasContext *ctx) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_nip, ctx->base.tb); } } =20 @@ -6690,7 +6690,7 @@ static void ppc_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cs) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_nip, ctx->base.tb); break; =20 case DISAS_EXIT_UPDATE: diff --git ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc ./target/riscv/t= cg/insn_trans/trans_rvzce.c.inc index 71b4ca5473..3f1e7c039e 100644 --- ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc +++ ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc @@ -213,7 +213,7 @@ static bool gen_pop(DisasContext *ctx, arg_cmpp *a, boo= l ret, bool ret_val) } #endif tcg_gen_mov_tl(cpu_pc, ret_addr); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; } =20 @@ -334,7 +334,7 @@ static bool trans_cm_jalt(DisasContext *ctx, arg_cm_jal= t *a) =20 tcg_gen_mov_tl(cpu_pc, addr); =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return true; } diff --git ./target/riscv/tcg/translate.c ./target/riscv/tcg/translate.c index 9684dbe752..8475ab43b4 100644 --- ./target/riscv/tcg/translate.c +++ ./target/riscv/tcg/translate.c @@ -287,7 +287,7 @@ static void lookup_and_goto_ptr(DisasContext *ctx) gen_helper_itrigger_match(tcg_env); } #endif - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } =20 static void exit_tb(DisasContext *ctx) diff --git ./target/rx/translate.c ./target/rx/translate.c index 132d495710..e5a9783d84 100644 --- ./target/rx/translate.c +++ ./target/rx/translate.c @@ -161,7 +161,7 @@ static void gen_goto_tb(DisasContext *dc, unsigned tb_s= lot_idx, vaddr dest) tcg_gen_exit_tb(dc->base.tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } dc->base.is_jmp =3D DISAS_NORETURN; } @@ -2242,7 +2242,7 @@ static void rx_tr_tb_stop(DisasContextBase *dcbase, C= PUState *cs) gen_goto_tb(ctx, 0, dcbase->pc_next); break; case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_UPDATE: tcg_gen_movi_i32(cpu_pc, ctx->base.pc_next); diff --git ./target/s390x/tcg/translate.c ./target/s390x/tcg/translate.c index 1b6023168b..607c039419 100644 --- ./target/s390x/tcg/translate.c +++ ./target/s390x/tcg/translate.c @@ -1162,7 +1162,7 @@ static DisasJumpType help_branch(DisasContext *s, Dis= asCompare *c, tcg_gen_goto_tb(0); tcg_gen_exit_tb(s->base.tb, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(tcgv_i64_temp(psw_addr), s->base.t= b); } =20 gen_set_label(lab); @@ -6477,7 +6477,8 @@ static void s390x_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cs) if (dc->exit_to_mainloop) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(tcgv_i64_temp(psw_addr), + dc->base.tb); } break; default: diff --git ./target/sh4/translate.c ./target/sh4/translate.c index 373950fd66..a4be456bd9 100644 --- ./target/sh4/translate.c +++ ./target/sh4/translate.c @@ -242,7 +242,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, vaddr dest) if (use_exit_tb(ctx)) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } ctx->base.is_jmp =3D DISAS_NORETURN; @@ -258,7 +258,7 @@ static void gen_jump(DisasContext * ctx) if (use_exit_tb(ctx)) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } else { diff --git ./target/sparc/translate.c ./target/sparc/translate.c index 3156be6a94..2ae0a02c44 100644 --- ./target/sparc/translate.c +++ ./target/sparc/translate.c @@ -376,7 +376,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_sl= ot_idx, /* jump to another page: we can use an indirect jump */ tcg_gen_movi_tl(cpu_pc, pc); tcg_gen_movi_tl(cpu_npc, npc); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); } } =20 @@ -5807,7 +5807,7 @@ static void sparc_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cs) tcg_gen_movi_tl(cpu_npc, dc->npc); } if (may_lookup) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } else { tcg_gen_exit_tb(NULL, 0); } diff --git ./target/tricore/translate.c ./target/tricore/translate.c index 8cd6b58f66..1d7f54f6df 100644 --- ./target/tricore/translate.c +++ ./target/tricore/translate.c @@ -2857,7 +2857,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned t= b_slot_index, vaddr dest) tcg_gen_exit_tb(ctx->base.tb, tb_slot_index); } else { gen_save_pc(dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } @@ -8478,7 +8478,7 @@ static void tricore_tr_tb_stop(DisasContextBase *dcba= se, CPUState *cpu) tcg_gen_exit_tb(NULL, 0); break; case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_NORETURN: break; diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index 28d3b2a847..2fda6e5c07 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -2715,7 +2715,7 @@ void tcg_gen_goto_tb(unsigned idx) tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } =20 -void tcg_gen_lookup_and_goto_ptr(void) +void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, const TranslationBlock *= tb) { TCGv_ptr ptr; =20 @@ -2725,6 +2725,7 @@ void tcg_gen_lookup_and_goto_ptr(void) } =20 plugin_gen_disable_mem_helpers(); + ptr =3D tcg_temp_ebb_new_ptr(); gen_helper_lookup_tb_ptr(ptr, tcg_env); tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425764; cv=none; d=zohomail.com; s=zohoarc; b=Yi+rsC5xFOUWsi1M3dByP93Z2g3wn38Ugb/3yf7K3N/TL5uiMGwrSGXqLKGCp3bPJcBhBcOVw1FVIqb6cln/cCm9o13fgMAs8enrOOudQuZAxLavi574LhYfX28x5Qnw/TTmx4jd6xu/XqjsSfj5G7qzX6j4g44jm236nih6UrA= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425764; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=z9P/7RakXdfn5VCCXUOm88mF4ZhBdXuP/baYPcdZBcI=; b=WgtccU4zsOIEkv0R93ggnmHYVkJeD7qqvNxTEKKrN9ncI8zJ4mmgd+V3z0tDumK+HgKOeTxejRXkbO6jeukMWb/BL7jjIX7fsimNT+npac03aiFk/dmWEvEMJ9msR0S4o9Joe8zL8nic4jsroFOtNI6drb9A0AkOUDbJRgg+tTk= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 17874257648806.079138111611201; Sat, 22 Aug 2026 12:09:24 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr57-00048w-MF; Sat, 22 Aug 2026 15:08:41 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr55-00048a-RQ for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:39 -0400 Received: from mail-yx1-xb12c.google.com ([2607:f8b0:4864:20::b12c]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr51-0005lz-LL for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:39 -0400 Received: by mail-yx1-xb12c.google.com with SMTP id 956f58d0204a3-66c82b32121so3202486d50.1 for ; Sat, 22 Aug 2026 12:08:35 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66cf4665056sm1205893d50.8.2026.08.22.12.08.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425714; x=1788030514; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=z9P/7RakXdfn5VCCXUOm88mF4ZhBdXuP/baYPcdZBcI=; b=BbzipP0pRttYvhKD9o/bX07UAKwzmYy6x+7G4AjJUowEdh21dNrJ1putjZ1l53cDS8 dLZ5Ew3M2M9EvBRZzXX/GlAlVFuauV7eu56GHPAv0+5eeppkc7hnKrpoIaClONceGHsf q539UzeK9noURaokXm64CFRnwIDYmaAUBAKtxK1NBfmvCP2Yi2nmITW6A35sEIoinwhD hcIgc7JsqPfSnzXFee5X9n2Fe0Ke042/H8WcV1pUhz16j5tu7i7W8C6IXz/QQNW5Bzsb pu6EsX+Oj4gUHF2ea0b01SIRHkQ12/JkIINspTTkEZpNj++N1rMO7UnrWsUEo/i64+WH rB+A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425714; x=1788030514; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=z9P/7RakXdfn5VCCXUOm88mF4ZhBdXuP/baYPcdZBcI=; b=lVk0JybZvDHSdiupW2E+ZewjYXBlX8N8r2m240Ed23TpEvtNie5/Np0xtPJTHmdM3f SSiEc6dKxiw4hAZyk09/7EpIz4C2uRUOM4uJOd3Z3oJdtYvZjY47HDxMy4Xl9XuS7I6T nWLLRHgjQjvFlIOUXs320iYI3d6OOSd0rMav8QNth6oJXAOc7bVoqdXwcIAGyy54kgYJ YQK13erwyQ3s6WymL9sECXeI+GdKpaO+btGYmbL5wUlKjfXHXzL8xp6jjONgIxm9dMOZ 8MWemfSjluAYjbagUUS04HEsWj1ZIyUuXeNWC+ff62ZAgKgbVSk2OVXrWwYbwQEuy52S 822w== X-Gm-Message-State: AFuF++nKAkYTOwR9DAwlI4kfYk3FbZQNBJVDYDqynjl8ZitBS68KY3LO BSqL982VMbyyS5OYajv1BsWY4fVUBVtSZb+7ABDYNfILrb7C8a0bt7bXP94Gn/RJP8s= X-Gm-Gg: AR+sD11Rak2hOL28UqeObd1sw2jctq9rahAmq1lOOkgmg3HQOyfnJIBNOfvSYu0dQx4 0k+tZcLBYhoIdUUzAN59pUVLdEvA0a29IUVRKCgNe/egNH1xbJqORT7oKR5TSD9Qa/Az73kvi32 kRaMnNoOBsVFdp4S1GfUnYPWm8m8SWevScQXRrWLPOaRBf9ExKeTtbT14GgMa8lDLCSU3Hc7DjN 8OJcK39ejV5GIsaQSqeA9Afmrt1L0syB4ocp6M99oeuLqSS+s4zs+kz/RvoepVP9rsZXgs/IY8m 6QLfQVzQ4F5TRsVQRHOWds+3rAD2fhU7aQljzbp4oIbrkHkob2H53+Nv39doMX0Uca1rbHScFYy iftN9TKnohzxdc3GbieNzZra37xK3DBa6r3b7TU5RGQ0i/mdhwTwV7Xu70mlfJRmmhp3xxDupGt 11PQEV10czIUfoIP4xq5LlSNWtOdP6jDVGU00aaAag98VOA1WUK/yJ7CsMXVhZFMCHOuzE33+cP uOEGq4BPGGHRvEV1dJKHL4z+g96xhQuu050ydg5H0dTB2r4vBg= X-Received: by 2002:a05:690e:450c:10b0:66c:3483:8261 with SMTP id 956f58d0204a3-66cf21bc221mr1560116d50.34.1787425713935; Sat, 22 Aug 2026 12:08:33 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 4/7] RFC: tcg: probe the TB jump cache inline instead of calling a helper Date: Sat, 22 Aug 2026 15:08:15 -0400 Message-ID: <20260822190818.1829249-5-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b12c; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb12c.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425766883158500 Content-Type: text/plain; charset="utf-8" Every indirect branch that cannot use goto_tb ends in tcg_gen_lookup_and_goto_ptr(), which calls helper_lookup_tb_ptr(). For an emulated compiler that is 8.4 billion helper calls in a single translation unit: 24.6% of all TB exits take this path, because jsr/ret/jmp have a register destination and because goto_tb is restricted to same-page targets. The helper itself is already tight, but each call pays for a call frame, the can_do_io store, the get_tb_cpu_state() indirect call through TCGCPUOps, curr_cflags(), and a breakpoint check, before it gets to the jump cache probe that almost always hits (95.8% for this workload). Emit the probe inline instead. The destination PC is already in a TCG temp, and the flags, cflags and cs_base the destination must match are constants at translation time, so the fast path is a hash, four guarded loads and a goto_ptr. Only a miss calls the helper, which still owns filling the cache. tcg_gen_lookup_and_goto_ptr() therefore takes the destination PC and the TB being generated, and decides for itself whether to emit the probe or the old helper call; there is no second entry point for targets that opt in. A target that cannot name its destination in a single temp passes NULL and gets the helper. Since the probe hashes and compares the PC as one 64-bit value, a 32-bit guest PC also falls back. The PC a target passes must be exactly what get_tb_cpu_state() reports for the destination, which is the whole of the contract. alpha, loongarch, mips, ppc and s390x pass their PC register, whose value is that pc by construction. The rest pass NULL for now: avr's TB pc is the word address doubled, i386's is eip before segmentation, riscv masks it to 32 bits when xl is MXL_RV32, hppa derives it from the IAQ, hexagon adjusts it inside a hardware loop, and sparc puts npc in cs_base so the guard could not hit anyway. Each of those is a one-line change for whoever wants to measure it. Two details matter for the generated code. The flags and cflags guards are folded into a single aligned 64-bit load and compare, since the fields are adjacent. And each path emits its own goto_ptr rather than branching to a shared one: a temp live across the label is spilled and reloaded on every dispatch, which cost 6.3% on its own. The probe cannot check everything the helper checks, and the one that matters is breakpoints. check_for_breakpoints() raises EXCP_DEBUG on an exact pc match and selects CF_BP_PAGE cflags for the rest of the page, and setting a breakpoint deliberately invalidates no TB, so a block translated before the breakpoint was set is still sitting in the jump cache. Rather than pay for a breakpoint test on the fast path, give the probe its own base pointer, tb_jmp_cache_probe, that nothing else reads, and point it at a page of zeroes while any breakpoint is set. Every entry the probe finds then has a NULL tb, so every dispatch misses into the helper and the old behaviour is restored exactly. cpu_breakpoint_insert() poisons the pointer, so the poison takes effect at the next dispatch rather than whenever that vCPU next reaches its main loop, which matters because a vCPU chaining indirectly need never reach it. The main loop puts the pointer back once the last breakpoint is gone; that is a load and a compare per block dispatched from the main loop, and nothing at all in generated code. The flags and cflags constants are safe against the other things that can change them. CF_PARALLEL is only ever set by begin_parallel_context(), which flushes first, so no block predating it survives to dispatch. gdb single-step is only turned on with the CPU stopped, and a block translated without CF_SINGLE_STEP can only be re-entered through tb_lookup(), which from then on demands the new cflags -- so a stale-cflags block is never the one running. What is left is one_insn_per_tb and -d nochain, which the monitor can toggle under a running vCPU without a flush; see below. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding three patches: before: 1,402,667,803,616 instructions after: 916,415,123,244 instructions -34.67% before: 115.56s wall clock after: 85.59s wall clock -25.94% The gap between the two is the point at which this stops being a straight-line win: the helper call was highly predictable work that the host pipelined well, so removing it retires far fewer instructions than it saves time. IPC falls from 2.48 to 2.17 across this patch for that reason. Despite emitting more code, this also reduces instruction cache pressure, because a dispatch no longer jumps into qemu's .text and evicts translated code: before: 11,735,141,703 L1-icache-load-misses after: 7,154,863,292 L1-icache-load-misses -39.0% The mechanism is visible directly in a profile: helper_lookup_tb_ptr() falls from 31.01% of samples to 0.35%, and qemu's own .text falls from 38.8% to 5.3%, with the balance moving into generated code. Combined with the three preceding patches, against an unmodified LTO build, 1,646,994,254,249 instructions fall to 916,415,123,244, or -44.36%. The emulated compiler produces byte-identical output throughout. Open issues, hence RFC: - one_insn_per_tb and CPU_LOG_TB_NOCHAIN can be toggled from the monitor while a vCPU is inside a block that was translated without them. The block keeps dispatching inline with the old cflags until it exits for some other reason. Poisoning the probe from tcg_update_all_curr_cflags() would close it. - The jump cache entry is read without qatomic_read(); entries are invalidated concurrently by setting tb to NULL. - Only alpha has been measured. The other four targets that pass a PC are built and boot-tested only. v3: Fold the fast path into tcg_gen_lookup_and_goto_ptr() instead of adding tcg_gen_lookup_and_goto_ptr_inline() beside it (Richard). It now takes the destination PC and the TB unconditionally, from all 38 call sites, and picks the probe or the helper itself. Translators built for both values of TARGET_LONG_BITS -- arm, s390x, microblaze -- cannot include tcg-op.h, so the common entry point takes a TCGTemp and reads the width from it, and tcg-op.h wraps that for everyone else; this is the same split as tcg_gen_qemu_ld_*_chk(). Compare cs_base too. v2 listed this as an open issue, and closing it is what lets the choice be made generically rather than per target: a target that uses cs_base would otherwise have been enabled silently by a decision keyed on PC width alone. It costs a load and a compare on the fast path, and the numbers above were measured with it in place. Audited which targets may pass a real PC, the contract being that it is exactly what get_tb_cpu_state() reports for the destination. Five do; the rest pass NULL and keep the helper call, sparc among them because it puts npc in cs_base and so could essentially never hit. Poison the probe from cpu_breakpoint_insert() rather than only from the poisoned CPU's own main loop. gdb inserts a breakpoint into every CPU (tcg_insert_gdbstub_breakpoint()), and a thread already inside generated code, dispatching indirectly, need never return to the main loop -- so it would keep dispatching inline and run past a breakpoint another thread had just set. Upstream has no such window: its helper_lookup_tb_ptr() sees the new breakpoint at the next indirect branch. The un-poison in tcg_cpu_sync_jmp_cache() now re-checks after its store, with a barrier, so that it loses the race with a concurrent insert in the safe direction. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec.c | 102 ++++++++++++++++++ accel/tcg/internal-common.h | 2 + cpu-common.c | 11 ++ include/hw/core/cpu.h | 9 ++ include/system/tcg.h | 9 ++ include/tcg/tcg-op-common.h | 16 ++- include/tcg/tcg-op.h | 12 +++ stubs/tcg-cflags.c | 8 +- target/alpha/translate.c | 4 +- target/arm/tcg/translate-a64.c | 4 +- target/arm/tcg/translate.c | 10 +- target/avr/translate.c | 4 +- target/hexagon/translate.c | 4 +- target/hppa/translate.c | 6 +- target/i386/tcg/translate.c | 2 +- .../tcg/insn_trans/trans_branch.c.inc | 2 +- target/loongarch/tcg/translate.c | 4 +- target/m68k/translate.c | 2 +- target/microblaze/translate.c | 4 +- target/mips/tcg/nanomips_translate.c.inc | 2 +- target/mips/tcg/translate.c | 6 +- target/or1k/translate.c | 4 +- target/ppc/translate.c | 4 +- target/riscv/tcg/insn_trans/trans_rvzce.c.inc | 4 +- target/riscv/tcg/translate.c | 2 +- target/rx/translate.c | 4 +- target/s390x/tcg/translate.c | 5 +- target/sh4/translate.c | 4 +- target/sparc/translate.c | 4 +- target/tricore/translate.c | 4 +- tcg/tcg-op.c | 92 +++++++++++++++- 31 files changed, 300 insertions(+), 50 deletions(-) diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 148e0f583e..e546f717e8 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -752,6 +752,99 @@ static inline bool cpu_handle_exception(CPUState *cpu,= int *ret) return false; } =20 +/* + * The inline jump cache probe reads cpu->tb_jmp_cache_probe and takes the + * slow path when the entry it finds has a NULL tb. Pointing the probe at= a + * region that is all zeroes therefore forces every indirect dispatch into + * helper_lookup_tb_ptr(), which does the full lookup the inline probe only + * approximates. The real jump cache is untouched, so no contents are lost + * and recovery is a single store. + * + * Only ever read from, and only the tb field of one entry per dispatch, so + * one shared zero-filled cache is enough for every CPU. + */ +static const CPUJumpCache *tb_jmp_cache_poison(void) +{ + static CPUJumpCache *poison; + + if (unlikely(poison =3D=3D NULL)) { + /* Raced allocations are harmless: both are all zeroes. */ + qatomic_cmpxchg(&poison, NULL, g_new0(CPUJumpCache, 1)); + } + return poison; +} + +/* + * Whether the generated code may dispatch to the next block by itself. + * + * The inline probe matches on the destination pc and on the flags and + * cflags the dispatching block was translated with. It does not consult + * cpu->breakpoints, so it must not run while one is set: setting a + * breakpoint deliberately invalidates nothing, and check_for_breakpoints() + * both raises EXCP_DEBUG on an exact match and picks CF_BP_PAGE cflags for + * the rest of the page. A block translated before the breakpoint was set= is + * therefore still in the jump cache, and dispatching to it inline would s= tep + * straight over the breakpoint. + */ +static bool tcg_cpu_may_dispatch(CPUState *cpu) +{ + return QTAILQ_EMPTY(&cpu->breakpoints); +} + +/* + * Poison @cpu's probe, from any thread. Called when a breakpoint is + * inserted, which is what makes the poison take effect at the dispatch + * after the insert rather than whenever @cpu next reaches its main loop: + * a vCPU chaining indirectly need never reach it, and would run past a + * breakpoint another thread had just set. + * + * A plain store is enough. The value only ever costs a slow path that is + * correct on its own, and the generated code re-reads the base on every + * dispatch. Un-poisoning is tcg_cpu_sync_jmp_cache()'s job. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu) +{ + if (qatomic_read(&cpu->tb_jmp_cache_probe) !=3D NULL) { + qatomic_set(&cpu->tb_jmp_cache_probe, + (CPUJumpCache *)tb_jmp_cache_poison()); + } +} + +/* + * Called from the main loop, which is the only context that can establish + * that no reason to be poisoned is left. Cheap enough to call every time + * round: the common case is a load, a compare and no store at all. + */ +void tcg_cpu_sync_jmp_cache(CPUState *cpu) +{ + CPUJumpCache *want; + + if (qatomic_read(&cpu->tb_jmp_cache_probe) =3D=3D NULL) { + return; /* not realized, or already unrealized */ + } + + want =3D tcg_cpu_may_dispatch(cpu) + ? cpu->tb_jmp_cache + : (CPUJumpCache *)tb_jmp_cache_poison(); + + if (qatomic_read(&cpu->tb_jmp_cache_probe) !=3D want) { + qatomic_set(&cpu->tb_jmp_cache_probe, want); + + /* + * Un-poisoning races a concurrent tcg_cpu_poison_jmp_cache(): the + * reason may have appeared after tcg_cpu_may_dispatch() read it a= nd + * the poison may have landed before the store above. Order the + * store against a re-read, and lose the race in the safe directio= n. + */ + if (want =3D=3D cpu->tb_jmp_cache) { + smp_mb(); + if (!tcg_cpu_may_dispatch(cpu)) { + tcg_cpu_poison_jmp_cache(cpu); + } + } + } +} + void tcg_kick_vcpu_thread(CPUState *cpu) { /* @@ -964,6 +1057,13 @@ cpu_exec_loop(CPUState *cpu, SyncClocks *sc) break; } =20 + /* + * Reaching here means the main loop has just re-evaluated + * everything the inline probe assumes, so this is where the + * probe is allowed to come back after a poison. + */ + tcg_cpu_sync_jmp_cache(cpu); + tb =3D tb_lookup(cpu, s); if (tb =3D=3D NULL) { CPUJumpCache *jc; @@ -1072,6 +1172,7 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_update_cflags(cpu); =20 cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); + qatomic_set(&cpu->tb_jmp_cache_probe, cpu->tb_jmp_cache); tlb_init(cpu); #ifndef CONFIG_USER_ONLY tcg_iommu_init_notifier_list(cpu); @@ -1089,5 +1190,6 @@ void tcg_exec_unrealizefn(CPUState *cpu) #endif /* !CONFIG_USER_ONLY */ =20 tlb_destroy(cpu); + qatomic_set(&cpu->tb_jmp_cache_probe, NULL); g_free_rcu(cpu->tb_jmp_cache, rcu); } diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 853d1b51ee..9d1f6712d6 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -144,6 +144,8 @@ void page_table_config_init(void); G_NORETURN void cpu_io_recompile(CPUState *cpu, uintptr_t retaddr); #endif /* CONFIG_USER_ONLY */ =20 +void tcg_cpu_sync_jmp_cache(CPUState *cpu); + void tb_phys_invalidate(TranslationBlock *tb, tb_page_addr_t page_addr); void tb_set_jmp_target(TranslationBlock *tb, int n, uintptr_t addr); =20 diff --git ./cpu-common.c ./cpu-common.c index adb76b3a78..3aed0156e6 100644 --- ./cpu-common.c +++ ./cpu-common.c @@ -22,6 +22,7 @@ #include "exec/cpu-common.h" #include "hw/core/cpu.h" #include "qemu/lockable.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 QemuMutex qemu_cpu_list_lock; @@ -429,6 +430,16 @@ int cpu_breakpoint_insert(CPUState *cpu, vaddr pc, int= flags, *breakpoint =3D bp; } =20 + /* + * Nothing is invalidated here, so blocks translated before this point + * are still live and still dispatch to each other without consulting + * cpu->breakpoints. Stop the ones that can: a TCG vCPU dispatching + * inline reads a base pointer that this poisons, so the next dispatch + * takes the slow path and sees the new breakpoint. @cpu may be anoth= er + * thread, and may be running. + */ + tcg_cpu_poison_jmp_cache(cpu); + trace_breakpoint_insert(cpu->cpu_index, pc, flags); return 0; } diff --git ./include/hw/core/cpu.h ./include/hw/core/cpu.h index 81af7b9ee1..bd2cdd2a0b 100644 --- ./include/hw/core/cpu.h +++ ./include/hw/core/cpu.h @@ -519,6 +519,15 @@ struct CPUState { MemoryRegion *memory; =20 struct CPUJumpCache *tb_jmp_cache; + /* + * @tb_jmp_cache_probe: base the inline jump cache probe reads. + * + * Normally @tb_jmp_cache. Pointed at a shared page of zeroes to force + * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (); + * see tcg_cpu_sync_jmp_cache(). NULL before tcg_exec_realizefn() and + * after tcg_exec_unrealizefn(). + */ + struct CPUJumpCache *tb_jmp_cache_probe; =20 GArray *gdb_regs; int gdb_num_regs; diff --git ./include/system/tcg.h ./include/system/tcg.h index 2c2dbc753b..bf05db1329 100644 --- ./include/system/tcg.h +++ ./include/system/tcg.h @@ -29,6 +29,15 @@ extern bool tcg_allowed; void tcg_update_cflags(CPUState *cpu); void tcg_update_all_cflags(void); =20 +/* + * Force @cpu's generated code back into the slow dispatch path, which + * re-checks everything the inline jump cache probe assumes. Safe to call + * from any thread, and a no-op for a CPU that is not running TCG. Call + * whenever something the probe cannot see changes under a running vCPU; + * the main loop undoes it once the reason is gone. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu); + /** * qemu_tcg_mttcg_enabled: * Check whether we are running MultiThread TCG or not. diff --git ./include/tcg/tcg-op-common.h ./include/tcg/tcg-op-common.h index 9b321f959c..ba580c7fb7 100644 --- ./include/tcg/tcg-op-common.h +++ ./include/tcg/tcg-op-common.h @@ -75,15 +75,25 @@ void tcg_gen_exit_tb(const TranslationBlock *tb, unsign= ed idx); void tcg_gen_goto_tb(unsigned idx); =20 /** - * tcg_gen_lookup_and_goto_ptr() - look up the current TB, jump to it if v= alid - * @addr: Guest address of the target TB + * tcg_gen_lookup_and_goto_ptr() - look up the destination TB, jump to it + * @pc: temp holding the destination guest PC, or NULL + * @tb: the translation block being generated * * If the TB is not valid, jump to the epilogue. * + * The lookup is normally a call to helper_lookup_tb_ptr(). If @pc is + * non-NULL and the destination can be keyed on it directly, the jump cache + * is probed inline instead and only a miss reaches the helper. @pc must + * then hold exactly the value get_tb_cpu_state() reports as the pc for the + * destination; a target whose pc is derived (avr's word address, i386's + * eip before segmentation) must pass NULL. The destination is required to + * match @tb's flags, cflags and cs_base, which is what makes them + * constants in the probe. + * * This operation is optional. If the TCG backend does not implement goto_= ptr, * this op is equivalent to calling tcg_gen_exit_tb() with 0 as the argume= nt. */ -void tcg_gen_lookup_and_goto_ptr(void); +void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, const TranslationBlock *= tb); =20 void tcg_gen_plugin_cb(unsigned from); void tcg_gen_plugin_mem_cb(TCGv_i64 addr, unsigned meminfo); diff --git ./include/tcg/tcg-op.h ./include/tcg/tcg-op.h index 3721164236..b6c7c6fea2 100644 --- ./include/tcg/tcg-op.h +++ ./include/tcg/tcg-op.h @@ -49,6 +49,18 @@ typedef TCGv_i64 TCGv; #error Unhandled TARGET_LONG_BITS value #endif =20 +/* + * See tcg_gen_lookup_and_goto_ptr_tmp(). @pc may be NULL, for a target + * whose guest PC is not directly the key the jump cache is indexed by. + * A translator that is built for more than one value of TARGET_LONG_BITS, + * and so cannot include this header, calls the _tmp() form directly. + */ +static inline void +tcg_gen_lookup_and_goto_ptr(TCGv pc, const TranslationBlock *tb) +{ + tcg_gen_lookup_and_goto_ptr_tmp(pc ? tcgv_tl_temp(pc) : NULL, tb); +} + #if TARGET_LONG_BITS =3D=3D 64 #define tcg_gen_movi_tl tcg_gen_movi_i64 #define tcg_gen_mov_tl tcg_gen_mov_i64 diff --git ./stubs/tcg-cflags.c ./stubs/tcg-cflags.c index cb278e94aa..4ac8a82e92 100644 --- ./stubs/tcg-cflags.c +++ ./stubs/tcg-cflags.c @@ -1,6 +1,6 @@ /* - * Stub for tcg_update_all_cflags(), for binaries that link util/log.c - * or cpu-target.c but not TCG. + * Stubs for the TCG entry points in system/tcg.h, for binaries that link + * util/log.c, cpu-target.c or cpu-common.c but not TCG. * * SPDX-License-Identifier: GPL-2.0-or-later */ @@ -14,3 +14,7 @@ void tcg_update_cflags(CPUState *cpu) void tcg_update_all_cflags(void) { } + +void tcg_cpu_poison_jmp_cache(CPUState *cpu) +{ +} diff --git ./target/alpha/translate.c ./target/alpha/translate.c index c66e3f9c14..822f5cc120 100644 --- ./target/alpha/translate.c +++ ./target/alpha/translate.c @@ -449,7 +449,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, int32_t disp) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_pc_disp(ctx, cpu_pc, disp); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); } } =20 @@ -2917,7 +2917,7 @@ static void alpha_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cpu) gen_pc_disp(ctx, cpu_pc, 0); /* FALLTHRU */ case DISAS_PC_UPDATED: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); break; case DISAS_PC_UPDATED_NOCHAIN: tcg_gen_exit_tb(NULL, 0); diff --git ./target/arm/tcg/translate-a64.c ./target/arm/tcg/translate-a64.c index 4f9a93950b..d1dd33a1af 100644 --- ./target/arm/tcg/translate-a64.c +++ ./target/arm/tcg/translate-a64.c @@ -562,7 +562,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_sl= ot_idx, int64_t diff) if (s->ss_active) { gen_step_complete_exception(s); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); s->base.is_jmp =3D DISAS_NORETURN; } } @@ -11250,7 +11250,7 @@ static void aarch64_tr_tb_stop(DisasContextBase *dc= base, CPUState *cpu) gen_a64_update_pc(dc, 4); /* fall through */ case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; case DISAS_NORETURN: case DISAS_SWI: diff --git ./target/arm/tcg/translate.c ./target/arm/tcg/translate.c index c866148383..ca701b9cbc 100644 --- ./target/arm/tcg/translate.c +++ ./target/arm/tcg/translate.c @@ -1306,9 +1306,9 @@ void write_neon_element64(TCGv_i64 src, int reg, int = ele, MemOp memop) } } =20 -static void gen_goto_ptr(void) +static void gen_goto_ptr(DisasContext *s) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, s->base.tb); } =20 /* This will end the TB but doesn't guarantee we'll return to @@ -1336,7 +1336,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_= slot_idx, int64_t diff) tcg_gen_exit_tb(s->base.tb, tb_slot_idx); } else { gen_update_pc(s, diff); - gen_goto_ptr(); + gen_goto_ptr(s); } s->base.is_jmp =3D DISAS_NORETURN; } @@ -1373,7 +1373,7 @@ static void gen_jmp_tb(DisasContext *s, int64_t diff,= int tbno) * and don't chain to another TB. */ gen_update_pc(s, diff); - gen_goto_ptr(); + gen_goto_ptr(s); s->base.is_jmp =3D DISAS_NORETURN; break; default: @@ -6858,7 +6858,7 @@ static void arm_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cpu) gen_update_pc(dc, curr_insn_len(dc)); /* fall through */ case DISAS_JUMP: - gen_goto_ptr(); + gen_goto_ptr(dc); break; case DISAS_UPDATE_EXIT: gen_update_pc(dc, curr_insn_len(dc)); diff --git ./target/avr/translate.c ./target/avr/translate.c index 3c57606097..8f2e0baa67 100644 --- ./target/avr/translate.c +++ ./target/avr/translate.c @@ -992,7 +992,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, tcg_gen_exit_tb(tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } @@ -2778,7 +2778,7 @@ static void avr_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cs) /* fall through */ case DISAS_LOOKUP: if (!force_exit) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; } /* fall through */ diff --git ./target/hexagon/translate.c ./target/hexagon/translate.c index 06a8159d28..cc230b08d1 100644 --- ./target/hexagon/translate.c +++ ./target/hexagon/translate.c @@ -181,7 +181,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, if (move_to_pc) { tcg_gen_movi_tl(hex_gpr[HEX_REG_PC], dest); } - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } =20 @@ -218,7 +218,7 @@ static void gen_end_tb(DisasContext *ctx) gen_set_label(skip); gen_goto_tb(ctx, 1, ctx->next_PC, false); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } =20 ctx->base.is_jmp =3D DISAS_NORETURN; diff --git ./target/hppa/translate.c ./target/hppa/translate.c index 002189ddfb..cf8f1a2c13 100644 --- ./target/hppa/translate.c +++ ./target/hppa/translate.c @@ -816,7 +816,7 @@ static void gen_goto_tb(DisasContext *ctx, int which, tcg_gen_goto_tb(which); tcg_gen_exit_tb(ctx->base.tb, which); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } =20 @@ -2027,7 +2027,7 @@ static bool do_ibranch(DisasContext *ctx, unsigned li= nk, store_psw_xb(ctx, PSW_B); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return nullify_end(ctx); } @@ -4838,7 +4838,7 @@ static void hppa_tr_tb_stop(DisasContextBase *dcbase,= CPUState *cs) } /* FALLTHRU */ case DISAS_IAQ_N_UPDATED: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_EXIT: tcg_gen_exit_tb(NULL, 0); diff --git ./target/i386/tcg/translate.c ./target/i386/tcg/translate.c index 2115c5cd24..66a0ee3cdf 100644 --- ./target/i386/tcg/translate.c +++ ./target/i386/tcg/translate.c @@ -2005,7 +2005,7 @@ gen_eob(DisasContext *s, int mode) } else if (mode =3D=3D DISAS_JUMP && /* give irqs a chance to happen */ !inhibit_reset) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); } else { tcg_gen_exit_tb(NULL, 0); } diff --git ./target/loongarch/tcg/insn_trans/trans_branch.c.inc ./target/lo= ongarch/tcg/insn_trans/trans_branch.c.inc index da07778658..57d9d47353 100644 --- ./target/loongarch/tcg/insn_trans/trans_branch.c.inc +++ ./target/loongarch/tcg/insn_trans/trans_branch.c.inc @@ -27,7 +27,7 @@ static bool trans_jirl(DisasContext *ctx, arg_jirl *a) tcg_gen_mov_tl(cpu_pc, addr); tcg_gen_movi_tl(dest, make_address_pc(ctx, ctx->base.pc_next + 4)); gen_set_gpr(a->rd, dest, EXT_NONE); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return true; } diff --git ./target/loongarch/tcg/translate.c ./target/loongarch/tcg/transl= ate.c index 124dce6269..a45a51852a 100644 --- ./target/loongarch/tcg/translate.c +++ ./target/loongarch/tcg/translate.c @@ -111,7 +111,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, vaddr dest) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { tcg_gen_movi_tl(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); } } =20 @@ -311,7 +311,7 @@ static void loongarch_tr_tb_stop(DisasContextBase *dcba= se, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: tcg_gen_movi_tl(cpu_pc, ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_pc, ctx->base.tb); break; case DISAS_TOO_MANY: gen_goto_tb(ctx, 0, ctx->base.pc_next); diff --git ./target/m68k/translate.c ./target/m68k/translate.c index 138c89d3e5..73691bc0d1 100644 --- ./target/m68k/translate.c +++ ./target/m68k/translate.c @@ -6095,7 +6095,7 @@ static void m68k_tr_tb_stop(DisasContextBase *dcbase,= CPUState *cpu) if (dc->ss_active) { gen_raise_exception_format2(dc, EXCP_TRACE, dc->pc_prev); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } break; case DISAS_EXIT: diff --git ./target/microblaze/translate.c ./target/microblaze/translate.c index 8b219afb5d..851b372f8f 100644 --- ./target/microblaze/translate.c +++ ./target/microblaze/translate.c @@ -127,7 +127,7 @@ static void gen_goto_tb(DisasContext *dc, unsigned tb_s= lot_idx, vaddr dest) tcg_gen_exit_tb(dc->base.tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, dc->base.tb); } dc->base.is_jmp =3D DISAS_NORETURN; } @@ -1764,7 +1764,7 @@ static void mb_tr_tb_stop(DisasContextBase *dcb, CPUS= tate *cs) /* Indirect jump (or direct jump w/ goto_tb disabled) */ tcg_gen_mov_i32(cpu_pc, cpu_btarget); tcg_gen_discard_i32(cpu_btarget); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(NULL, dc->base.tb); return; =20 default: diff --git ./target/mips/tcg/nanomips_translate.c.inc ./target/mips/tcg/nan= omips_translate.c.inc index 4b0b01ba37..007e29f9ac 100644 --- ./target/mips/tcg/nanomips_translate.c.inc +++ ./target/mips/tcg/nanomips_translate.c.inc @@ -2406,7 +2406,7 @@ static void gen_compute_nanomips_pbalrsc_branch(Disas= Context *ctx, int rs, =20 /* unconditional branch to register */ tcg_gen_mov_tl(cpu_PC, btarget); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); } =20 /* nanoMIPS Branches */ diff --git ./target/mips/tcg/translate.c ./target/mips/tcg/translate.c index e3467d1525..73abfbb5d4 100644 --- ./target/mips/tcg/translate.c +++ ./target/mips/tcg/translate.c @@ -4374,7 +4374,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned t= b_slot_idx, tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_save_pc(dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); } } =20 @@ -11014,7 +11014,7 @@ static void gen_branch(DisasContext *ctx, int insn_= bytes) } else { tcg_gen_mov_tl(cpu_PC, btarget); } - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); break; default: LOG_DISAS("unknown branch 0x%x\n", proc_hflags); @@ -15244,7 +15244,7 @@ static void mips_tr_tb_stop(DisasContextBase *dcbas= e, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: gen_save_pc(ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_PC, ctx->base.tb); break; case DISAS_NEXT: case DISAS_TOO_MANY: diff --git ./target/or1k/translate.c ./target/or1k/translate.c index eb4485312f..4907284a6d 100644 --- ./target/or1k/translate.c +++ ./target/or1k/translate.c @@ -1605,7 +1605,7 @@ static void openrisc_tr_tb_stop(DisasContextBase *dcb= ase, CPUState *cs) /* The jump destination is indirect/computed; use jmp_pc. */ tcg_gen_mov_i32(cpu_pc, jmp_pc); tcg_gen_discard_i32(jmp_pc); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; } /* The jump destination is direct; use jmp_pc_imm. @@ -1622,7 +1622,7 @@ static void openrisc_tr_tb_stop(DisasContextBase *dcb= ase, CPUState *cs) break; } tcg_gen_movi_i32(cpu_pc, jmp_dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); break; =20 case DISAS_EXIT: diff --git ./target/ppc/translate.c ./target/ppc/translate.c index 06ed2adf10..42924281b0 100644 --- ./target/ppc/translate.c +++ ./target/ppc/translate.c @@ -3664,7 +3664,7 @@ static void gen_lookup_and_goto_ptr(DisasContext *ctx) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_nip, ctx->base.tb); } } =20 @@ -6690,7 +6690,7 @@ static void ppc_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cs) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(cpu_nip, ctx->base.tb); break; =20 case DISAS_EXIT_UPDATE: diff --git ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc ./target/riscv/t= cg/insn_trans/trans_rvzce.c.inc index 71b4ca5473..3f1e7c039e 100644 --- ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc +++ ./target/riscv/tcg/insn_trans/trans_rvzce.c.inc @@ -213,7 +213,7 @@ static bool gen_pop(DisasContext *ctx, arg_cmpp *a, boo= l ret, bool ret_val) } #endif tcg_gen_mov_tl(cpu_pc, ret_addr); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; } =20 @@ -334,7 +334,7 @@ static bool trans_cm_jalt(DisasContext *ctx, arg_cm_jal= t *a) =20 tcg_gen_mov_tl(cpu_pc, addr); =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); ctx->base.is_jmp =3D DISAS_NORETURN; return true; } diff --git ./target/riscv/tcg/translate.c ./target/riscv/tcg/translate.c index 9684dbe752..8475ab43b4 100644 --- ./target/riscv/tcg/translate.c +++ ./target/riscv/tcg/translate.c @@ -287,7 +287,7 @@ static void lookup_and_goto_ptr(DisasContext *ctx) gen_helper_itrigger_match(tcg_env); } #endif - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } =20 static void exit_tb(DisasContext *ctx) diff --git ./target/rx/translate.c ./target/rx/translate.c index 132d495710..e5a9783d84 100644 --- ./target/rx/translate.c +++ ./target/rx/translate.c @@ -161,7 +161,7 @@ static void gen_goto_tb(DisasContext *dc, unsigned tb_s= lot_idx, vaddr dest) tcg_gen_exit_tb(dc->base.tb, tb_slot_idx); } else { tcg_gen_movi_i32(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } dc->base.is_jmp =3D DISAS_NORETURN; } @@ -2242,7 +2242,7 @@ static void rx_tr_tb_stop(DisasContextBase *dcbase, C= PUState *cs) gen_goto_tb(ctx, 0, dcbase->pc_next); break; case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_UPDATE: tcg_gen_movi_i32(cpu_pc, ctx->base.pc_next); diff --git ./target/s390x/tcg/translate.c ./target/s390x/tcg/translate.c index 1b6023168b..607c039419 100644 --- ./target/s390x/tcg/translate.c +++ ./target/s390x/tcg/translate.c @@ -1162,7 +1162,7 @@ static DisasJumpType help_branch(DisasContext *s, Dis= asCompare *c, tcg_gen_goto_tb(0); tcg_gen_exit_tb(s->base.tb, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(tcgv_i64_temp(psw_addr), s->base.t= b); } =20 gen_set_label(lab); @@ -6477,7 +6477,8 @@ static void s390x_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cs) if (dc->exit_to_mainloop) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr_tmp(tcgv_i64_temp(psw_addr), + dc->base.tb); } break; default: diff --git ./target/sh4/translate.c ./target/sh4/translate.c index 373950fd66..a4be456bd9 100644 --- ./target/sh4/translate.c +++ ./target/sh4/translate.c @@ -242,7 +242,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, vaddr dest) if (use_exit_tb(ctx)) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } } ctx->base.is_jmp =3D DISAS_NORETURN; @@ -258,7 +258,7 @@ static void gen_jump(DisasContext * ctx) if (use_exit_tb(ctx)) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } else { diff --git ./target/sparc/translate.c ./target/sparc/translate.c index 3156be6a94..2ae0a02c44 100644 --- ./target/sparc/translate.c +++ ./target/sparc/translate.c @@ -376,7 +376,7 @@ static void gen_goto_tb(DisasContext *s, unsigned tb_sl= ot_idx, /* jump to another page: we can use an indirect jump */ tcg_gen_movi_tl(cpu_pc, pc); tcg_gen_movi_tl(cpu_npc, npc); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, s->base.tb); } } =20 @@ -5807,7 +5807,7 @@ static void sparc_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cs) tcg_gen_movi_tl(cpu_npc, dc->npc); } if (may_lookup) { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, dc->base.tb); } else { tcg_gen_exit_tb(NULL, 0); } diff --git ./target/tricore/translate.c ./target/tricore/translate.c index 8cd6b58f66..1d7f54f6df 100644 --- ./target/tricore/translate.c +++ ./target/tricore/translate.c @@ -2857,7 +2857,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned t= b_slot_index, vaddr dest) tcg_gen_exit_tb(ctx->base.tb, tb_slot_index); } else { gen_save_pc(dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); } ctx->base.is_jmp =3D DISAS_NORETURN; } @@ -8478,7 +8478,7 @@ static void tricore_tr_tb_stop(DisasContextBase *dcba= se, CPUState *cpu) tcg_gen_exit_tb(NULL, 0); break; case DISAS_JUMP: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_lookup_and_goto_ptr(NULL, ctx->base.tb); break; case DISAS_NORETURN: break; diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index 28d3b2a847..ce77541eab 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -28,6 +28,8 @@ #include "tcg/tcg-op-common.h" #include "exec/translation-block.h" #include "exec/plugin-gen.h" +#include "hw/core/cpu.h" +#include "../accel/tcg/tb-jmp-cache.h" #include "tcg-internal.h" #include "tcg-has.h" =20 @@ -2715,7 +2717,81 @@ void tcg_gen_goto_tb(unsigned idx) tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } =20 -void tcg_gen_lookup_and_goto_ptr(void) +static void gen_jmp_cache_probe(TCGv_i64 pc, const TranslationBlock *tb) +{ + TCGv_ptr jc, ent, tbp, ptr; + TCGv_i64 h, tmp; + TCGLabel *slow; + uint64_t fpair; + + QEMU_BUILD_BUG_ON(sizeof(((CPUJumpCache *)0)->array[0]) !=3D 16); + QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, cflags) !=3D + offsetof(TranslationBlock, flags) + 4); + + jc =3D tcg_temp_ebb_new_ptr(); + ent =3D tcg_temp_ebb_new_ptr(); + tbp =3D tcg_temp_ebb_new_ptr(); + ptr =3D tcg_temp_ebb_new_ptr(); + h =3D tcg_temp_ebb_new_i64(); + tmp =3D tcg_temp_ebb_new_i64(); + slow =3D gen_new_label(); + + /* h =3D tb_jmp_cache_hash_func(pc) * sizeof(array[0]) */ + tcg_gen_shri_i64(h, pc, TB_JMP_CACHE_BITS); + tcg_gen_xor_i64(h, h, pc); + tcg_gen_andi_i64(h, h, TB_JMP_CACHE_SIZE - 1); + tcg_gen_shli_i64(h, h, 4); + + /* + * Not cpu->tb_jmp_cache: the probe reads its own base so that the main + * loop can poison it, which is how conditions the probe cannot test f= or + * itself force every dispatch back into the helper. See + * tcg_cpu_sync_jmp_cache(). + */ + tcg_gen_ld_ptr(jc, tcg_env, + offsetof(CPUState, tb_jmp_cache_probe) - sizeof(CPUStat= e)); + tcg_gen_trunc_i64_ptr(ent, h); + tcg_gen_add_ptr(ent, jc, ent); + + tcg_gen_ld_ptr(tbp, ent, offsetof(CPUJumpCache, array[0].tb)); + tcg_gen_brcondi_ptr(TCG_COND_EQ, tbp, 0, slow); + + tcg_gen_ld_i64(tmp, ent, offsetof(CPUJumpCache, array[0].pc)); + tcg_gen_brcond_i64(TCG_COND_NE, tmp, pc, slow); + + /* + * flags and cflags are adjacent uint32_t, so one aligned 64-bit load + * and compare covers both. + */ +#if HOST_BIG_ENDIAN + fpair =3D ((uint64_t)tb->flags << 32) | tb->cflags; +#else + fpair =3D ((uint64_t)tb->cflags << 32) | tb->flags; +#endif + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, flags)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, fpair, slow); + + /* + * The destination must have been translated with the same cs_base, wh= ich + * the pc alone does not imply on a target that uses it. + */ + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, cs_base)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, tb->cs_base, slow); + + tcg_gen_ld_ptr(ptr, tbp, offsetof(TranslationBlock, tc.ptr)); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); + + /* + * Emit a second goto_ptr rather than branching to a shared one: a temp + * live across the label would be spilled and reloaded on every dispat= ch. + */ + gen_set_label(slow); + ptr =3D tcg_temp_ebb_new_ptr(); + gen_helper_lookup_tb_ptr(ptr, tcg_env); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); +} + +void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, const TranslationBlock *= tb) { TCGv_ptr ptr; =20 @@ -2724,7 +2800,21 @@ void tcg_gen_lookup_and_goto_ptr(void) return; } =20 + /* + * No icount_decr poll is needed for this exit: the helper is called on + * every dispatch and returns to the main loop while an exit is pendin= g. + */ plugin_gen_disable_mem_helpers(); + + /* + * The inline probe hashes and compares the pc as a single 64-bit valu= e. + * A target with a 32-bit guest PC keeps the helper call. + */ + if (pc && pc->type =3D=3D TCG_TYPE_I64) { + gen_jmp_cache_probe(temp_tcgv_i64(pc), tb); + return; + } + ptr =3D tcg_temp_ebb_new_ptr(); gen_helper_lookup_tb_ptr(ptr, tcg_env); tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425768; cv=none; d=zohomail.com; s=zohoarc; b=Wbg2mNiDFvKmUiaHJ2MmaQwnyztz+SR1/8wgKWXJyrIFEDMQIuGNLrZLQ5ofn2CH0o/ipyWEMm9mqY2ZV+/Ggq44C/e/ldTzpssdcA+RnqwDdWI1q/2XS7VbUTGIvo9PucMatmCHM8hJy5mAALsYS+axQwPCe5UiT0cAKb6JyD0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425768; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=QCeGKvgZvzLRmULF1kf11XJuTp71B39WYQckVmi9p4Y=; b=FE+VOVX7LROna0xdZLvTFpZDxkCN++WST7PU4/Ik5VW2JlDH1yBMrICjodHrX2yEZ2vPjmjbx2OYCGP+EBb2ezc0bI1Tn6aKCBG7YGxza97BphkCN6iqRrmzwuwnqZYTMRA+lZadk/oqwp4GvVq7oIUwNYuzccM8kqQYoENTZvM= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787425768340418.5852021366836; Sat, 22 Aug 2026 12:09:28 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr58-0004A0-L6; Sat, 22 Aug 2026 15:08:42 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr57-00048o-2J for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:41 -0400 Received: from mail-yw1-x112e.google.com ([2607:f8b0:4864:20::112e]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr54-0005m7-Dr for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:40 -0400 Received: by mail-yw1-x112e.google.com with SMTP id 00721157ae682-81f36179d72so35163097b3.2 for ; Sat, 22 Aug 2026 12:08:38 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84ca5187583sm13931197b3.4.2026.08.22.12.08.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425717; x=1788030517; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QCeGKvgZvzLRmULF1kf11XJuTp71B39WYQckVmi9p4Y=; b=gXE4DdpqNKaPdRhdglhKQ14xJdEbxl39iy5hWDPaTVwVCTI//ED0IxW87rXbxWQ7/z DSSO+eR+GyjDOrAdOz70cc69dgdHeeIkEjpThGh8gYndn+6WvXhsm/HbGvx/je1DQQSa KW0sA5NaAka4uvafKuVZt7aM5tWqURqjikBRbCInjwE+/iWrKrB4mF3t/hZEwVEJAriM kVPL6r/8mFkLif5IpTKOtnGfBeTZqGl/3depLqLhk3lIDlC14jn9mgCnIRh+gFoMfZYE Uk19SD4X4lYBHRate4L6qg9NmJgV8oFEXpyHf9YGtsiu3djLgUXI5X1LuUXgCuoDMNer h8FA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425717; x=1788030517; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=QCeGKvgZvzLRmULF1kf11XJuTp71B39WYQckVmi9p4Y=; b=HEblkvOoyCiuA2xJn3CB/ODD2fvLf0NWzccMamr/1FUayH+/AMRLqe38K2niKliFsW +LKmuqItP9qIXUERNbmNCwtthLeWEQs9WGY/C3BHucdrxzw10JWo4mVRNnxQvzg20Ztc VfZaOCKIrrgMR6Zn77O7ZK967OB+VwoZ5qd67GU7u5l6k6iTpdkr4lXiOZn1O6Ju/g7K kjcAWQmt+ID43TL7neNVn2K3x/sz/zbRHbXD9yRd0CybLZ/pAJeQ3ZCohZp2J2qaQcFT FyekLZL9lLDrhekThVR0n+t3E/rP7LlQPFhHbX1N2jxVa2TNv2zCov17Ijqqc1obo+Sc hvlQ== X-Gm-Message-State: AFuF++nK2l9BrMhUvmzCxIakNq/frfATXxq1yW7WqaBCxLSdfkhz/lt3 kybholWX7bW1rEcsZCQcqaK07XWvH56fveYlO7S2sFy64iQYxyLw/7K80/LCd+swsCg= X-Gm-Gg: AR+sD13oC8mtfL46jaDI0DBbkimjn+HDENRFHCYrArha7e7xnW4P+1Aav5P4r5Td7bL bTiTENbcUhs9tU1HmWP8P0OqbDtgiRVw3sYn0kBBb5in3WIOj0AGG6nVWTIyV8IlPfhf0ax0woR Iv5gcqaRKFmRJuVanOIjeq6AP5Bujbf+VVW2AVk07qB0wJnmm6cEEo22xsmYNJ9SANszDF5fG2R pyg+HwkYY2gSPxxkVtr2GJrda9+r7Pr0SU8fYN/dZNs/L7GHSPknlKDwX+pZT6DpzTQi7SW1XN/ qhs9v/FV5MGXWjzx4HfWN+jaDhPSbJx5NYlCM9WvfdyaD4WHvn1YcWLhpibrm/SFr8U4MEfAsnX zQEfpF2prm22jf9hv+J2yI3krcUK5CrbGz65eNCWipO4J9tcc9rJqSKqdlA8LkSN8RAQd2bTihr jPH0+S1kCoYAKZW4lTe2qjQUMifM2h8mMUVKoRv/5n8A+qz7dBRoTSdGbIorArHY7uHKPbDgWy3 yQy/xD4QRFtNH5JdADrbhkpPSXAHI5ygVql9cyP X-Received: by 2002:a05:690c:dc4:b0:80e:38c5:bc17 with SMTP id 00721157ae682-84c9366089emr25466277b3.7.1787425717172; Sat, 22 Aug 2026 12:08:37 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 5/7] RFC: accel/tcg: allow cross-page goto_tb chaining in user-only builds Date: Sat, 22 Aug 2026 15:08:16 -0400 Message-ID: <20260822190818.1829249-6-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112e; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112e.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425768894158500 Content-Type: text/plain; charset="utf-8" translator_use_goto_tb() refuses to chain unless the destination is on the same page as the start of the TB. For guests whose text is much larger than a page this is expensive: an emulated alpha gcc compiling a 255k line translation unit takes the indirect dispatch path for 8.4 billion of its 34.2 billion TB exits, and a large share of those are ordinary direct branches that simply crossed an 8 KiB page boundary. The restriction was made unconditional by d3a2a1d803 ("accel/tcg: Introduce translator_use_goto_tb"), whose rationale was: Various targets avoid the page crossing test for CONFIG_USER_ONLY, but that is wrong: mmap and mprotect can change page permissions. That is true, but in user-only builds the invalidation path already covers it. There are no page tables: every mmap, mprotect and munmap reaches page_set_flags(), which calls tb_invalidate_phys_range() whenever the flags actually change, and tb_phys_invalidate() calls tb_jmp_unlink() to reset incoming jumps. A chained cross-page jump is therefore broken whenever the destination page's permissions change. This is not true in system mode, where TBs are keyed by physical address and a page table change invalidates nothing, so the restriction is kept there. The rule protects one more thing, which the original rationale does not mention: it guarantees that execution cannot enter a page without a TB lookup, and so without check_for_breakpoints(). That is what makes a breakpoint set after a block was translated take effect, since insertion deliberately invalidates nothing. A link established before the breakpoint was set would jump straight over it. So the chaining is only enabled for a run that can never acquire a breakpoint. In user-only mode every breakpoint comes from gdb -- BP_CPU is g_assert_not_reached() there, and the guest cannot ask for one -- and gdb has to be requested with -g before the first block is translated, even though with suspend=3Dn it may connect later. gdb_may_set_breakpoints() reports whether it was, and is fixed for the lifetime of the process. Add tests/tcg/alpha/test-xpage-chain.c to cover both hazards directly. It places a direct branch near the end of one page targeting the next page, runs it 200000 times so the chain is established, then checks that mprotect(PROT_NONE) makes the next call fault, and that remapping the page with different code runs the new code rather than a stale translation. The test detects the hazard it is meant to detect: with the tb_invalidate_phys_range() call in page_set_flags() commented out, it fails both phases, executing page B after PROT_NONE and returning the stale result. Run with -b, the same binary stops once the chain is established and lets tests/tcg/alpha/gdbstub/xpage-bp.py set a breakpoint on the far side of it, which the next call has to stop on. With gdb_may_set_breakpoints() forced to false so that the chaining stays on under gdb, that breakpoint is missed and the test fails, which is what makes it a test of the gate rather than of gdb. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation on an x86-64 host, LTO build, on top of the preceding patches: before: 916,415,123,244 instructions after: 891,254,240,071 instructions -2.75% before: 85.59s wall clock after: 81.45s wall clock -4.84% Note that this is worth more in time than in instructions, the reverse of the preceding patch: a chained jump replaces a cache probe whose loads can miss, so the instructions it removes are more expensive than average. Measured before the inline jump cache probe, when a missed chain cost a helper call rather than an inline probe, the same change was worth -7.9%. RFC because this reverses a deliberate decision and the reasoning above wants review from someone who knows the invalidation paths better than I do. v3: Only take the shortcut when no gdbstub was requested. The same-page rule also forces a lookup, and so a breakpoint check, on entry to every page; without that, a chain established before a breakpoint was set runs past it. Reported by Richard Henderson. v3: Change translator_use_goto_tb() rather than translator_is_same_page(). i386, riscv and s390x call translator_is_same_page() for something else -- enforcing that only a single-insn TB may cross a page -- and v2 changed their TB boundaries in user-only mode as a side effect. alpha does not call it, so the numbers above are unaffected. v3: Add the gdbstub half of the test. Signed-off-by: Matt Turner --- accel/tcg/translator.c | 33 ++++++- gdbstub/user.c | 14 +++ include/gdbstub/user.h | 11 +++ tests/tcg/alpha/Makefile.target | 17 +++- tests/tcg/alpha/gdbstub/xpage-bp.py | 34 +++++++ tests/tcg/alpha/test-xpage-chain.c | 144 ++++++++++++++++++++++++++++ 6 files changed, 251 insertions(+), 2 deletions(-) create mode 100644 tests/tcg/alpha/gdbstub/xpage-bp.py create mode 100644 tests/tcg/alpha/test-xpage-chain.c diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 6c8fcd7a20..8879cd626f 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -15,6 +15,9 @@ #include "accel/tcg/cpu-mmu-index.h" #include "exec/target_page.h" #include "exec/translator.h" +#ifdef CONFIG_USER_ONLY +#include "gdbstub/user.h" +#endif #include "exec/plugin-gen.h" #include "tcg/tcg-op-common.h" #include "internal-common.h" @@ -110,6 +113,34 @@ bool translator_is_same_page(const DisasContextBase *d= b, vaddr addr) return ((addr ^ db->pc_first) & TARGET_PAGE_MASK) =3D=3D 0; } =20 +/* + * Whether a direct jump may be chained to a destination outside the page + * the TB started in. + * + * In user-only mode there are no page tables. Every mmap, mprotect and + * munmap goes through page_set_flags(), which calls tb_invalidate_phys_ra= nge() + * whenever the flags actually change, and tb_phys_invalidate() unlinks + * incoming jumps. A cross-page link is therefore broken whenever the + * destination page's permissions change. + * + * What the same-page rule also provides is that execution cannot enter a = page + * without a TB lookup, and so without check_for_breakpoints(), which is w= hat + * makes a breakpoint set after a block was translated take effect. Nothi= ng + * invalidates on breakpoint insertion, so a link established beforehand w= ould + * jump straight over it. In user-only mode breakpoints only ever come fr= om + * gdb -- BP_CPU is g_assert_not_reached() there and the guest has no way = to + * ask for one -- and gdb has to be requested with -g before the first blo= ck + * is translated, so a run that has no gdbstub can never acquire a breakpo= int. + */ +static bool use_cross_page_goto_tb(void) +{ +#ifdef CONFIG_USER_ONLY + return !gdb_may_set_breakpoints(); +#else + return false; +#endif +} + bool translator_use_goto_tb(DisasContextBase *db, vaddr dest) { /* Suppress goto_tb if requested. */ @@ -118,7 +149,7 @@ bool translator_use_goto_tb(DisasContextBase *db, vaddr= dest) } =20 /* Check for the dest on the same page as the start of the TB. */ - return translator_is_same_page(db, dest); + return use_cross_page_goto_tb() || translator_is_same_page(db, dest); } =20 void translator_loop(CPUState *cpu, TranslationBlock *tb, int *max_insns, diff --git ./gdbstub/user.c ./gdbstub/user.c index 9e6f9a6f37..d810f0f38c 100644 --- ./gdbstub/user.c +++ ./gdbstub/user.c @@ -470,6 +470,18 @@ static void *gdbserver_accept_thread(void *arg) =20 #define USAGE "\nUsage: -g {port|path}[,suspend=3D{y|n}]" =20 +/* + * Set before the guest runs and never cleared, so that code translated at + * any point can rely on it: with suspend=3Dn gdb may connect long after + * startup, and once connected it can insert a breakpoint at any time. + */ +static bool gdbserver_requested; + +bool gdb_may_set_breakpoints(void) +{ + return gdbserver_requested; +} + bool gdbserver_start(const char *args, Error **errp) { g_auto(GStrv) argv =3D g_strsplit(args, ",", 0); @@ -513,6 +525,8 @@ bool gdbserver_start(const char *args, Error **errp) return false; } =20 + gdbserver_requested =3D true; + if (suspend) { if (gdbserver_accept(port, gdb_fd, port_or_path)) { gdb_handlesig(first_cpu, 0, NULL, NULL, 0); diff --git ./include/gdbstub/user.h ./include/gdbstub/user.h index 654986d483..c091cd9758 100644 --- ./include/gdbstub/user.h +++ ./include/gdbstub/user.h @@ -11,6 +11,17 @@ =20 #define MAX_SIGINFO_LENGTH 128 =20 +/** + * gdb_may_set_breakpoints() - whether a breakpoint can ever be inserted + * + * In user-only mode every breakpoint comes from gdb, and gdb is only ever + * reachable if -g was given at startup, before the guest ran a single + * instruction. A run that has no gdbstub can therefore never acquire a + * breakpoint, which lets translation take shortcuts that a breakpoint + * would invalidate. Stays true once true, even if gdb detaches. + */ +bool gdb_may_set_breakpoints(void); + /** * gdb_handlesig() - yield control to gdb * @cpu: CPU diff --git ./tests/tcg/alpha/Makefile.target ./tests/tcg/alpha/Makefile.tar= get index 36d8ed1eae..1a3f541bec 100644 --- ./tests/tcg/alpha/Makefile.target +++ ./tests/tcg/alpha/Makefile.target @@ -5,7 +5,7 @@ ALPHA_SRC=3D$(SRC_PATH)/tests/tcg/alpha VPATH+=3D$(ALPHA_SRC) =20 -ALPHA_TESTS=3Dhello-alpha test-cond test-cmov test-ovf test-cvttq +ALPHA_TESTS=3Dhello-alpha test-cond test-cmov test-ovf test-cvttq test-xpa= ge-chain TESTS+=3D$(ALPHA_TESTS) =20 test-cmov: EXTRA_CFLAGS=3D-DTEST_CMOV @@ -16,3 +16,18 @@ test-cmov: test-cond.c test-plugin-mem-access: CFLAGS+=3D-mbwx =20 run-test-cmov: test-cmov + +ifneq ($(GDB),) +GDB_SCRIPT=3D$(SRC_PATH)/tests/guest-debug/run-test.py + +# The chaining this exercises is only enabled when no gdbstub was requeste= d, +# so what is under test here is that requesting one turns it back off. +run-gdbstub-xpage-bp: test-xpage-chain + $(call run-test, $@, $(GDB_SCRIPT) \ + --gdb $(GDB) \ + --qemu $(QEMU) --qargs "$(QEMU_OPTS)" \ + --bin "$< -b" --test $(ALPHA_SRC)/gdbstub/xpage-bp.py, \ + breakpoint behind an established cross-page chain) + +EXTRA_RUNS +=3D run-gdbstub-xpage-bp +endif diff --git ./tests/tcg/alpha/gdbstub/xpage-bp.py ./tests/tcg/alpha/gdbstub/= xpage-bp.py new file mode 100644 index 0000000000..f0ec14cdec --- /dev/null +++ ./tests/tcg/alpha/gdbstub/xpage-bp.py @@ -0,0 +1,34 @@ +"""Test that a breakpoint set after a cross-page chain is established is h= it. + +translator_use_goto_tb() lets a direct branch chain to another page in +user-only builds, which is only safe because a run with no gdbstub can nev= er +acquire a breakpoint. This runs with one, so the chaining must be off and +the breakpoint must still be reached. + +This runs as a sourced script (via -x, via run-test.py). + +SPDX-License-Identifier: GPL-2.0-or-later +""" +from test_gdbstub import main, report + + +def run_test(): + """Run through the tests one by one""" + gdb.Breakpoint("break_here") + gdb.execute("continue") + + # The chain exists by now; put a breakpoint on the far side of it. + target =3D int(gdb.parse_and_eval("(unsigned long)page_b_entry")) + gdb.execute("break *{}".format(target)) + gdb.execute("continue") + + pc =3D int(gdb.parse_and_eval("(unsigned long)$pc")) + report(pc =3D=3D target, "stopped at {:#x}, expected {:#x}".format(pc,= target)) + + gdb.execute("delete") + gdb.execute("continue") + exitcode =3D int(gdb.parse_and_eval("$_exitcode")) + report(exitcode =3D=3D 0, "{} =3D=3D 0".format(exitcode)) + + +main(run_test) diff --git ./tests/tcg/alpha/test-xpage-chain.c ./tests/tcg/alpha/test-xpag= e-chain.c new file mode 100644 index 0000000000..23b916ffe1 --- /dev/null +++ ./tests/tcg/alpha/test-xpage-chain.c @@ -0,0 +1,144 @@ +/* + * Cross-page TB chaining hazard test. + * + * Phase 1: a direct branch (br) near the end of page A targets page B. + * Run it enough times that QEMU chains TB_A -> TB_B. + * Phase 2: mprotect page B away. Re-running must fault. + * Phase 3: remap page B with different code. Re-running must execute the + * NEW code, not a stale chained translation of the old code. + * + * With -b, phases 2 and 3 are replaced by a stop at break_here(), where t= he + * gdbstub test sets a breakpoint on page B -- after the chain exists -- a= nd + * checks that re-running the chain still stops on it. See + * tests/tcg/alpha/gdbstub/xpage-bp.py. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include +#include +#include +#include + +#define PS 8192 + +static sigjmp_buf jb; +/* + * Written by the SIGSEGV handler and read by main(), so it must not be + * cached in a register across the faulting call. + */ +static volatile sig_atomic_t caught; + +/* Where the branch lands, for the gdbstub test to set a breakpoint on. */ +unsigned int *page_b_entry; + +/* Somewhere for the gdbstub test to stop once the chain is established. */ +void __attribute__((noinline)) break_here(void) +{ + asm volatile (""); +} + +static void segv(int sig) +{ + caught =3D 1; + siglongjmp(jb, 1); +} + +/* lda $0, imm($31) -> v0 =3D imm */ +static unsigned int lda_v0(int imm) +{ + return 0x201F0000u | (unsigned short)imm; +} + +int main(int argc, char **argv) +{ + bool bp_mode =3D argc > 1 && strcmp(argv[1], "-b") =3D=3D 0; + struct sigaction sa; + int rc =3D 0; + unsigned char *m =3D mmap(NULL, 2 * PS, PROT_READ | PROT_WRITE | PROT_= EXEC, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (m =3D=3D MAP_FAILED) { + perror("mmap"); + return 2; + } + + unsigned char *pa =3D m, *pb =3D m + PS; + unsigned int *entry =3D (unsigned int *)(pa + PS - 64); + unsigned int *tgt =3D (unsigned int *)(pb + 16); + + page_b_entry =3D tgt; + + entry[0] =3D lda_v0(1); + long disp =3D ((long)tgt - ((long)&entry[1] + 4)) / 4; + entry[1] =3D 0xC3E00000u | (unsigned int)(disp & 0x1FFFFF); /* br $31= ,tgt */ + tgt[0] =3D 0x6BFA8001u; /* ret = */ + __builtin___clear_cache((char *)m, (char *)m + 2 * PS); + + long (*fn)(void) =3D (long (*)(void))entry; + + for (int i =3D 0; i < 200000; i++) { + if (fn() !=3D 1) { + printf("FAIL: phase 1 wrong result\n"); + return 1; + } + } + printf("phase 1 ok (chained)\n"); + + if (bp_mode) { + /* + * The chain from page A to page B now exists. gdb puts a breakpo= int + * on page_b_entry here; the call below has to stop on it rather t= han + * jump over it. + */ + break_here(); + if (fn() !=3D 1) { + printf("FAIL: bp phase wrong result\n"); + return 1; + } + printf("bp phase ok\n"); + return 0; + } + + memset(&sa, 0, sizeof(sa)); + sa.sa_handler =3D segv; + sigemptyset(&sa.sa_mask); + if (sigaction(SIGSEGV, &sa, NULL) !=3D 0) { + perror("sigaction"); + return 2; + } + if (mprotect(pb, PS, PROT_NONE) !=3D 0) { + perror("mprotect"); + return 2; + } + if (sigsetjmp(jb, 1) =3D=3D 0) { + fn(); + printf("FAIL: phase 2 executed page B after mprotect(PROT_NONE)\n"= ); + rc =3D 1; + } else if (!caught) { + printf("FAIL: phase 2 longjmp without entering the handler\n"); + rc =3D 1; + } else { + printf("phase 2 ok (faulted)\n"); + } + + /* Phase 3: remap with different code, expect the new code to run. */ + if (mprotect(pb, PS, PROT_READ | PROT_WRITE | PROT_EXEC) !=3D 0) { + perror("mprotect back"); + return 2; + } + tgt[0] =3D lda_v0(2); + tgt[1] =3D 0x6BFA8001u; + __builtin___clear_cache((char *)pb, (char *)pb + PS); + + long r =3D fn(); + if (r !=3D 2) { + printf("FAIL: phase 3 returned %ld, expected 2 (stale chain)\n", r= ); + rc =3D 1; + } else { + printf("phase 3 ok (new code ran)\n"); + } + return rc; +} --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807040; cv=none; d=zohomail.com; s=zohoarc; b=L7f5zLYoOYRus3azuyjmXXLeTRGg6Q2V09f0R7TtYi1w24ncqD7LH4WSgmkFHbNGtRUpw6WbR0NmH1tpLCMes3XJcuIWwhfs1MFMkYg11PrHiiUn9aX5jfP43w7grfqDl4Spa0zQvHYqgYUh+ANTG5XUw2Z2W/IWo1mmatpAodg= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807040; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=ZnObpIoTaCD5yxi6xstyDb0vWidqbX29THsSMuAvPIs=; b=lVQN1yJny5pvhRYpMk/RnKtjlpMNUikHIdNwnbg2XoTBEt0lEdAQ65X4xrgzpyFZpv4TiUWN5vGN4+dDyMwUFqY5xDTez6Tu8nuuupjFYNMnD3wOmT7GqcN9Dx+Uv93RWJ4FiPO0Z7ycVI+8dbvdqyiHJqrcAGSp3mw/PfIn/UY= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 178780704063795.51452282400237; Wed, 26 Aug 2026 22:04:00 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGg-0005rV-LL; Thu, 27 Aug 2026 01:03:14 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGb-0005ql-69 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:09 -0400 Received: from mail-yw1-x1136.google.com ([2607:f8b0:4864:20::1136]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGZ-0002JX-11 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:08 -0400 Received: by mail-yw1-x1136.google.com with SMTP id 00721157ae682-857d1184d29so19194417b3.2 for ; Wed, 26 Aug 2026 22:03:06 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b6143e97bsm4133407b3.26.2026.08.26.22.03.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806986; x=1788411786; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ZnObpIoTaCD5yxi6xstyDb0vWidqbX29THsSMuAvPIs=; b=AOz6dFf8tzj8bh+BDVOuG/Dde6fh0y+RGXF31KsX3VvxhykDbjuAJQyMhF7kQCQUoO RgwkinHNER0u79Pn3MUf+D5woLvwKbQfwSXlMK0qOi11SIsY0P3GawCClddpUcsqgMSB JGIB2gyIdUBaTtzfFEWdGOyWOQzQvENEiKBOdVPAMjMmaa3mjFkLNZXqQcUWji5ybsjV kBhssrPQ9yYTIC/R+zUC84irPt3QlCvUSWdxhtVyoCbKRVFYuVkZeqT2ame2xk2f7U2f QrU5L2ee64b8xhwqb11IcugxEQ83NxQkp0i+425FSNmS11jJvUB2XyYuagy8lle4spWc dBNw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806986; x=1788411786; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ZnObpIoTaCD5yxi6xstyDb0vWidqbX29THsSMuAvPIs=; b=puDUsHMJ3GPaS8WXooTMPnSOlWryI8w/UYZ6rTU1FV8fs/Ia7xdma2zxJKmnOBWreC e5WfSHSCVlsYcifb8b+8EtYiaY8odludBvFZrNiZgPSQpn7kemSsJcHzBrETTvalGk60 Py9t+F3Y01xW4Jv1myu50rdGck2LXeG0w/i68gH29P7snQ45Jf9cqF5iPdWBzI3sr1UG mNQYpBSYiUNlRnpJifqQgmzmThtV0+ji3MuiFJVW7Bp6ly+6TQ4RopcC18JAgxcgL7yz KQK8Omyh9lWtZLkuFiSce368pA7NTeo1yUfPDMVI7g3340fXbddJZE7qMiRuYqpJ3GHL 6Vbw== X-Gm-Message-State: AFuF++lF6guNoMg5o8FocqHEizdxabD/HDS1t49JlwUB+2/FSvZYoQxY GlGtkXCe9atzjqU7zygJtr4pVBb9EYwYJfUukf9m+o2DRdq7IJPzzXzasnFOCieg X-Gm-Gg: AR+sD10ooAdnp5b+XTJX6kZyQM9xbsN94ERe+BmdVDCDnRb57s6OIf4YU/ZVFO6ftur 7+rsgjmYwjSWU/ls3NKQdxJri4Sz4FH+vegKjNkhAp1ISXVlXPzgNr59KYvEV+vI5KLGl0u1Wtk Ur7UulieNRM5Mhw+RFaeLCoejX8cly/qBgpdbgIsThw8+ODCShzSSuyJS6/6fMTW71MIlQE8ZdO FDLHjCy6DmqgLXXs8+0KBKc/1J5UQUOmoyrY6JjtfGn7mYXjxfbFm8bzK4VTSF12EE4RhU9hrAL Vxh6FBdqwyJae11MX6LpEyjRinJOIpQ4n8I1JbQLmbXBAFt1QdVY8e8vMnmbz58W11bAcLgbhLU RRSDeYKQqO0q0wWAfh3DhOXm3miHxeQm+eem2slLD7ZMQoZFvqm4N2mLWVtKY2l+BGoLvwtOZ3W nkarwg2ssUEF+pXpD8Cf9Rc0iEH2j/DjSA9hbBGNVmAaYOAegNIXMGJqjlic90SZpW6FDRqLi7/ ruU+LTagzAAi5P+MI8VkgYXOziSZioa1zrnldpT X-Received: by 2002:a05:690c:c507:b0:825:c3fc:e63f with SMTP id 00721157ae682-8573c9cf6c7mr48415177b3.8.1787806985778; Wed, 26 Aug 2026 22:03:05 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 5/9] accel/tcg: give the TB jump cache a second base pointer for generated code Date: Thu, 27 Aug 2026 01:02:37 -0400 Message-ID: <20260827050241.3713332-6-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1136; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1136.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807042108158500 Content-Type: text/plain; charset="utf-8" Add CPUState::tb_jmp_cache_probe, a base pointer that only generated code will read. Normally it is cpu->tb_jmp_cache. Pointing it instead at a shared, permanently zero-filled CPUJumpCache makes every entry generated code finds have a NULL tb, so every lookup done through it misses. The real jump cache is not touched, so nothing is lost and recovery is a single store. Nothing reads it yet. The next patch probes the jump cache from generated code, and that probe cannot check everything helper_lookup_tb_ptr() checks; poisoning this pointer is how the conditions it cannot check force it back into the helper. The one that matters here is breakpoints. check_for_breakpoints() raises EXCP_DEBUG on an exact pc match and selects CF_BP_PAGE cflags for the rest of the page, and inserting a breakpoint deliberately invalidates no TB, so a block translated before the breakpoint was set is still sitting in the jump cache and would be dispatched to directly. cpu_breakpoint_insert() poisons the target CPU, rather than leaving it to that CPU's own main loop, because gdb inserts a breakpoint into every CPU (tcg_insert_gdbstub_breakpoint()) and a thread already inside generated code dispatching to itself need never return to its main loop. The main loop puts the pointer back once the last breakpoint is gone, re-checking after the store so that it loses a race with a concurrent insert in the safe direction. The poison cache is a plain static rather than a const one so that it lands in .bss: a megabyte of const zeroes would be a megabyte of .rodata in every emulator binary, whereas .bss costs nothing on disk and faults in only the handful of pages a poisoned run happens to probe. v4: Split out of "tcg: probe the TB jump cache inline instead of calling a helper". Requested by Richard Henderson. v4: Make the poison a static object rather than allocating one on first use. Suggested by Richard Henderson, who asked for const; see above for why it is not. Signed-off-by: Matt Turner --- accel/stubs/tcg-stub.c | 6 ++- accel/tcg/cpu-exec.c | 96 +++++++++++++++++++++++++++++++++++++ accel/tcg/internal-common.h | 2 + cpu-common.c | 11 +++++ include/hw/core/cpu.h | 9 ++++ include/system/tcg.h | 9 ++++ 6 files changed, 132 insertions(+), 1 deletion(-) diff --git ./accel/stubs/tcg-stub.c ./accel/stubs/tcg-stub.c index f9e1bd22d6..8298e4a1f5 100644 --- ./accel/stubs/tcg-stub.c +++ ./accel/stubs/tcg-stub.c @@ -1,6 +1,6 @@ /* * Stubs for the TCG entry points in system/tcg.h, for binaries that link - * cpu-target.c or the HMP command handlers but not TCG. + * cpu-target.c, cpu-common.c or the HMP command handlers but not TCG. * * SPDX-License-Identifier: GPL-2.0-or-later */ @@ -14,3 +14,7 @@ void tcg_update_cflags(CPUState *cpu) void tcg_update_all_cflags(void) { } + +void tcg_cpu_poison_jmp_cache(CPUState *cpu) +{ +} diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 148e0f583e..c2a9679cd7 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -752,6 +752,93 @@ static inline bool cpu_handle_exception(CPUState *cpu,= int *ret) return false; } =20 +/* + * The inline jump cache probe reads cpu->tb_jmp_cache_probe and takes the + * slow path when the entry it finds has a NULL tb. Pointing the probe at= a + * region that is all zeroes therefore forces every indirect dispatch into + * helper_lookup_tb_ptr(), which does the full lookup the inline probe only + * approximates. The real jump cache is untouched, so no contents are lost + * and recovery is a single store. + * + * Only ever read from, and only one entry per dispatch, so one shared + * zero-filled cache is enough for every CPU. Not const: that would put a + * megabyte of zeroes in .rodata and so in the binary, where .bss costs + * nothing on disk and only faults in the handful of pages a poisoned run + * happens to probe. + */ +static CPUJumpCache tb_jmp_cache_poison; + +/* + * Whether the generated code may dispatch to the next block by itself. + * + * The inline probe matches on the destination pc and on the flags and + * cflags the dispatching block was translated with. It does not consult + * cpu->breakpoints, so it must not run while one is set: setting a + * breakpoint deliberately invalidates nothing, and check_for_breakpoints() + * both raises EXCP_DEBUG on an exact match and picks CF_BP_PAGE cflags for + * the rest of the page. A block translated before the breakpoint was set= is + * therefore still in the jump cache, and dispatching to it inline would s= tep + * straight over the breakpoint. + */ +static bool tcg_cpu_may_dispatch(CPUState *cpu) +{ + return QTAILQ_EMPTY(&cpu->breakpoints); +} + +/* + * Poison @cpu's probe, from any thread. Called when a breakpoint is + * inserted, which is what makes the poison take effect at the dispatch + * after the insert rather than whenever @cpu next reaches its main loop: + * a vCPU chaining indirectly need never reach it, and would run past a + * breakpoint another thread had just set. + * + * A plain store is enough. The value only ever costs a slow path that is + * correct on its own, and the generated code re-reads the base on every + * dispatch. Un-poisoning is tcg_cpu_sync_jmp_cache()'s job. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu) +{ + if (qatomic_read(&cpu->tb_jmp_cache_probe) !=3D NULL) { + qatomic_set(&cpu->tb_jmp_cache_probe, &tb_jmp_cache_poison); + } +} + +/* + * Called from the main loop, which is the only context that can establish + * that no reason to be poisoned is left. Cheap enough to call every time + * round: the common case is a load, a compare and no store at all. + */ +void tcg_cpu_sync_jmp_cache(CPUState *cpu) +{ + CPUJumpCache *want; + + if (qatomic_read(&cpu->tb_jmp_cache_probe) =3D=3D NULL) { + return; /* not realized, or already unrealized */ + } + + want =3D tcg_cpu_may_dispatch(cpu) + ? cpu->tb_jmp_cache + : &tb_jmp_cache_poison; + + if (qatomic_read(&cpu->tb_jmp_cache_probe) !=3D want) { + qatomic_set(&cpu->tb_jmp_cache_probe, want); + + if (want =3D=3D cpu->tb_jmp_cache) { + /* + * Un-poisoning races a concurrent tcg_cpu_poison_jmp_cache(): + * the reason may have appeared after tcg_cpu_may_dispatch() r= ead + * it, and the poison may have landed before the store above. + * Order that store against the re-read below, so that the race + * is lost in the safe direction. + */ + smp_mb(); + if (!tcg_cpu_may_dispatch(cpu)) { + tcg_cpu_poison_jmp_cache(cpu); + } + } + } +} + void tcg_kick_vcpu_thread(CPUState *cpu) { /* @@ -964,6 +1051,13 @@ cpu_exec_loop(CPUState *cpu, SyncClocks *sc) break; } =20 + /* + * Reaching here means the main loop has just re-evaluated + * everything the inline probe assumes, so this is where the + * probe is allowed to come back after a poison. + */ + tcg_cpu_sync_jmp_cache(cpu); + tb =3D tb_lookup(cpu, s); if (tb =3D=3D NULL) { CPUJumpCache *jc; @@ -1072,6 +1166,7 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_update_cflags(cpu); =20 cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); + qatomic_set(&cpu->tb_jmp_cache_probe, cpu->tb_jmp_cache); tlb_init(cpu); #ifndef CONFIG_USER_ONLY tcg_iommu_init_notifier_list(cpu); @@ -1089,5 +1184,6 @@ void tcg_exec_unrealizefn(CPUState *cpu) #endif /* !CONFIG_USER_ONLY */ =20 tlb_destroy(cpu); + qatomic_set(&cpu->tb_jmp_cache_probe, NULL); g_free_rcu(cpu->tb_jmp_cache, rcu); } diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 853d1b51ee..9d1f6712d6 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -144,6 +144,8 @@ void page_table_config_init(void); G_NORETURN void cpu_io_recompile(CPUState *cpu, uintptr_t retaddr); #endif /* CONFIG_USER_ONLY */ =20 +void tcg_cpu_sync_jmp_cache(CPUState *cpu); + void tb_phys_invalidate(TranslationBlock *tb, tb_page_addr_t page_addr); void tb_set_jmp_target(TranslationBlock *tb, int n, uintptr_t addr); =20 diff --git ./cpu-common.c ./cpu-common.c index adb76b3a78..3aed0156e6 100644 --- ./cpu-common.c +++ ./cpu-common.c @@ -22,6 +22,7 @@ #include "exec/cpu-common.h" #include "hw/core/cpu.h" #include "qemu/lockable.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 QemuMutex qemu_cpu_list_lock; @@ -429,6 +430,16 @@ int cpu_breakpoint_insert(CPUState *cpu, vaddr pc, int= flags, *breakpoint =3D bp; } =20 + /* + * Nothing is invalidated here, so blocks translated before this point + * are still live and still dispatch to each other without consulting + * cpu->breakpoints. Stop the ones that can: a TCG vCPU dispatching + * inline reads a base pointer that this poisons, so the next dispatch + * takes the slow path and sees the new breakpoint. @cpu may be anoth= er + * thread, and may be running. + */ + tcg_cpu_poison_jmp_cache(cpu); + trace_breakpoint_insert(cpu->cpu_index, pc, flags); return 0; } diff --git ./include/hw/core/cpu.h ./include/hw/core/cpu.h index 81af7b9ee1..bd2cdd2a0b 100644 --- ./include/hw/core/cpu.h +++ ./include/hw/core/cpu.h @@ -519,6 +519,15 @@ struct CPUState { MemoryRegion *memory; =20 struct CPUJumpCache *tb_jmp_cache; + /* + * @tb_jmp_cache_probe: base the inline jump cache probe reads. + * + * Normally @tb_jmp_cache. Pointed at a shared page of zeroes to force + * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (); + * see tcg_cpu_sync_jmp_cache(). NULL before tcg_exec_realizefn() and + * after tcg_exec_unrealizefn(). + */ + struct CPUJumpCache *tb_jmp_cache_probe; =20 GArray *gdb_regs; int gdb_num_regs; diff --git ./include/system/tcg.h ./include/system/tcg.h index 2c2dbc753b..bf05db1329 100644 --- ./include/system/tcg.h +++ ./include/system/tcg.h @@ -29,6 +29,15 @@ extern bool tcg_allowed; void tcg_update_cflags(CPUState *cpu); void tcg_update_all_cflags(void); =20 +/* + * Force @cpu's generated code back into the slow dispatch path, which + * re-checks everything the inline jump cache probe assumes. Safe to call + * from any thread, and a no-op for a CPU that is not running TCG. Call + * whenever something the probe cannot see changes under a running vCPU; + * the main loop undoes it once the reason is gone. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu); + /** * qemu_tcg_mttcg_enabled: * Check whether we are running MultiThread TCG or not. --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425753; cv=none; d=zohomail.com; s=zohoarc; b=nTW9eK6+/y9PnU5WrI5YJDP3MHdIOlVhqlMNJrEDKksjkbVexenzCm3IkyJBJi5Auf+C4mGHPQ5Nz7MXXTe25tTIpC33iaMQoZETVvF7ErkGJ6evo9o1nW7J7gxLD/JVnSB2/bqdrY528gIEGZngt4KFQqJj+MaqLK7Q6kHrze0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425753; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=SeRu4vnI69HRB4IvNiphxPhFUxC0jRHtQT+PelOqh98=; b=MGlvTJl6kTn8MXFUb6D5RIE6F5pvx3vHEcF9K+YEZuYBDw5B23av6OQfNu4qCAivEJ3poVGWCAMDtBi7wBwuPAJDihbMt0tM6WOtjuBGWazJupb2S78uu53OAA1rmyq1n9HG7wzDqX/cRXLE0pRHiWUr8hRLzEa4sktaFVGy4PU= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787425753576566.106264038262; Sat, 22 Aug 2026 12:09:13 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr5B-0004BR-4H; Sat, 22 Aug 2026 15:08:45 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr59-0004Af-L3 for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:43 -0400 Received: from mail-yx1-xb135.google.com ([2607:f8b0:4864:20::b135]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr57-0005mM-2c for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:43 -0400 Received: by mail-yx1-xb135.google.com with SMTP id 956f58d0204a3-66c7127a73dso2388227d50.2 for ; Sat, 22 Aug 2026 12:08:40 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84ca527e9basm13558207b3.5.2026.08.22.12.08.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425720; x=1788030520; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SeRu4vnI69HRB4IvNiphxPhFUxC0jRHtQT+PelOqh98=; b=ogkpBPFgfIFBfLy1ad3UDNJNmluJJV986Vt7X2pdKxTUEkqgQcdUMErWfcs0KtJKPU r1GjLPg8Rw2x8pVmmYghXp6+ZGPuS8ynbp0quZbH7g/RQO6aKRBryyN68h2AnthYdrXF hQlrubKA7aySrW0xW1zBCtV7bJEKaWalyKZGpXhf1lw/bVWBsGZQvCdacpLOoxssMQ6Y clmtudodDExSwR7BxUbcASYF31H5+PO7EljvxluIFz6I1APMULuzpRPocFvSfLFBo3fK 8viVOITkNVVoVaWTPEW6VKR0enVTzqDcE+iwVD042L5xaAqMsGxlssg0//gQlJpEz2dL WMJw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425720; x=1788030520; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=SeRu4vnI69HRB4IvNiphxPhFUxC0jRHtQT+PelOqh98=; b=bqXmQN17+EmKRm4HQwQWQza7m3SrfAP3hZwRdT+OkGnhxo7guVTeoAS9VcDh1TZGRJ MKKpMN5bP7nDRasr13+9iVRskFB9+AzhzH0RDn/f0vTh/Yp5Iw6G3PnWA3VkEoSGX1ad Un7MuXqRDLJDqOjsWsuXBLBDJAYfMabZ0sG7N4F3y7mCa1ynjGozPcLY/+wGpHI4VNbp 8KU7fIoq2hbtFyjPfqb7Y0wrqj3KdF9HWs2bTep4+7fgrkCQ79AHAENH5yLKXVIaceN5 y3qVDLAjEi8l87xywMhkoisgx70jqPDLY+RVrPbJzkYo8pBKaqJBmya4+HxHrEmNzmOh UIpQ== X-Gm-Message-State: AFuF++nd0uJzpg/kgQagGxYbIvAopmefakRmM0brMlh72M8ofPvGAPUA UuHDAPAS0iUt74cxJywsqnZw2oKDil5RqQDL7kIdD0GS/5rAhpevwJOANaPXS3YXF6A= X-Gm-Gg: AR+sD11BpoyDrL7rmPbY3crXL/K5PC516XgXtJd/oaSzTLw50Kcxcims72QFb1SSmvV TYOizJvH6jY1QDeCsD1/5bkK6Dt9qmc/uMDOeH5OzJ9a/DdTiDtSGTqrGvO28F2/6xWz9CFnhpj 8HdsikREaqHsmgRsCZxuLBTIU2JPe+WV0vofzMBI288h+N1Dz7+2JkFMww25vDMeJ1nTTzkn13M IavLSHZ5+xpMhw/2hgrsVmQE1xH0p/lPganouIfu7rjnB4rjZZOIh3VICZ8g880FFObizuJY1gk yVssCdeJW85RYIseh0EKzHFG2FXJY0mfvhIpuDvmFilefEXMv/6uVpBc9sWqo7dhNyQZtQRZrUD 1wM4nHAz9GyymP48pAjLHs/D9s5GOa9tdtCkJRkj86Uf70yhoQroPl3Puf5MmnhrbPNyYtCy7tC 4bfZMv9fBvqPikH7lp5JCKEGBfB7dJvWrcbTQO/l43xMEnU8oaDK0j1YJiAOrp5XGdjKHDJEpBC kVnQQW5bW7s93qhZulgSoCUPhTFqO1mwi5HwjGX X-Received: by 2002:a53:b3c3:0:b0:66c:486c:aa07 with SMTP id 956f58d0204a3-66cf21bd42amr1757095d50.35.1787425719582; Sat, 22 Aug 2026 12:08:39 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 6/7] RFC: accel/tcg: poison the jump cache instead of polling for indirect exits Date: Sat, 22 Aug 2026 15:08:17 -0400 Message-ID: <20260822190818.1829249-7-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b135; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb135.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425754893158500 Content-Type: text/plain; charset="utf-8" Every translation block begins by loading cpu->neg.icount_decr.u32, testing it and branching to the exit path. That is three host instructions at the t= op of every TB, and blocks are short: an emulated alpha gcc 16.2.0 compiling a 255k line translation unit executes 34.2 billion of them at 6.04 guest instructions each. A block does not need to poll if every way out of it already reaches a chec= k. A goto_tb does not: it chains straight into its destination, with nothing in between that looks at icount_decr, so the destination has to poll on entry. An indirect exit does. The out-of-line path calls helper_lookup_tb_ptr() every time, so it only needs the helper to return the epilogue while an exit is pending. The inline probe is already covered by the machinery the breakpoint patch added: it reads its base pointer from cpu->tb_jmp_cache_probe and takes the slow path when the entry it finds has= a NULL tb, so pointing that base at a page of zeroes turns every indirect dispatch into a miss, and a miss lands in the same helper. So a pending exit becomes one more reason for tcg_cpu_may_dispatch() to say no. The two places that set icount_decr.u16.high poison the probe; the main loop puts it back once the flag is clear, on the same pass that already re-evaluates the breakpoint state. The real tb_jmp_cache is untouched throughout, so no cache contents are lost, and the fast path pays nothing: the base was a load from CPUState either way. The poll is therefore emitted only in blocks that emit a goto_tb. Whether a block does is not known until its last exit has been generated, so the decision is deferred and the load and branch are emitted retroactively at t= he head of the block in gen_tb_end(), using the same emit_before_op mechanism the can_do_io stores use. icount opts out and keeps the counter unconditionally. Interrupt latency is bounded at one block, as before. It does not depend on the shape of the guest's control flow graph: a block either polls on entry = or is checked on the way out, and no run of blocks can avoid both. What changes is where the check sits, not how often one happens. tests/tcg/alpha/test-indirect-irq.c is added for this: a loop whose only ba= ck edge is an indirect branch, under alarm(1). That loop's block emits no goto_tb, so it no longer polls, and the test passes only because the dispat= ch notices instead -- it hangs if the poison is removed, which is what makes i= t a test of the new mechanism rather than of the old poll. The other alpha tests still pass and the emulated compiler still produces byte-identical output. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, = on top of the preceding patches: before: 891,254,240,071 instructions after: 868,811,832,620 instructions -2.52% before: 81.45s wall clock after: 79.85s wall clock -1.96% The emulated compiler produces byte-identical output. RFC because: - The un-poison in the main loop races a concurrent poison from another thread. The existing barrier around icount_decr.u16.high covers it -- a poison that lands after the sync also re-set the flag, and exit_request w= as stored before it -- but this deserves more eyes than the single-threaded user-mode testing I have given it. - Only the inline probe needs the poison, and only alpha uses the inline probe today. Targets on the out-of-line path are covered by the helper check alone, but that has not been measured. - The shared zero-filled CPUJumpCache is a 1MB allocation that is never written. A read-only mapping would express that better. v3: Rebased onto the removal of "only poll for interrupts in blocks that can close a cycle", which v2 sat on top of and which is dropped: it let a straight-line run of arbitrary length go unchecked, since a block with = no backward edge polled nowhere (Richard). The rule is now that a block polls iff it emits a goto_tb, rather than iff it can close a control flow cycle. That keeps the bound at one block without any analysis of the guest's control flow graph, so the objection to the dropped patch does not carry over. The deferred-emission machine= ry it needs moves here from that patch; DisasContextBase::needs_exit_check and the hook in translator_use_goto_tb() are gone with it, and the flag is now set by tcg_gen_goto_tb() rather than by goto_ptr emission. All of v2's measurements were dropped: they were taken with the cycle-analysis patch underneath, which changes both the baseline and what is left to remove, so none of them described this patch. The numbers above are a fresh measurement of the series as it now stands. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec.c | 23 ++++++++++-- accel/tcg/tcg-accel-ops.c | 1 + accel/tcg/translator.c | 51 ++++++++++++++++++++++++-- include/hw/core/cpu.h | 8 +++-- include/tcg/tcg.h | 2 ++ tcg/tcg-op.c | 13 +++++-- tests/tcg/alpha/Makefile.target | 3 +- tests/tcg/alpha/test-indirect-irq.c | 55 +++++++++++++++++++++++++++++ 8 files changed, 144 insertions(+), 12 deletions(-) create mode 100644 tests/tcg/alpha/test-indirect-irq.c diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index e546f717e8..62f6984f7a 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -388,6 +388,16 @@ const void *HELPER(lookup_tb_ptr)(CPUArchState *env) */ cpu->neg.can_do_io =3D true; =20 + /* + * A block that dispatches indirectly does not emit the icount_decr po= ll, + * so this is where a pending exit is noticed for that path: either the + * probe was poisoned and every dispatch arrives here, or the target u= ses + * the out-of-line lookup and always did. + */ + if (unlikely(cpu_loop_exit_requested(cpu))) { + return tcg_code_gen_epilogue; + } + TCGTBCPUState s =3D cpu->cc->tcg_ops->get_tb_cpu_state(cpu); s.cflags =3D curr_cflags(cpu); =20 @@ -757,8 +767,9 @@ static inline bool cpu_handle_exception(CPUState *cpu, = int *ret) * slow path when the entry it finds has a NULL tb. Pointing the probe at= a * region that is all zeroes therefore forces every indirect dispatch into * helper_lookup_tb_ptr(), which does the full lookup the inline probe only - * approximates. The real jump cache is untouched, so no contents are lost - * and recovery is a single store. + * approximates and returns to the main loop while an exit is pending. The + * real jump cache is untouched, so no contents are lost and recovery is a + * single store. * * Only ever read from, and only the tb field of one entry per dispatch, so * one shared zero-filled cache is enough for every CPU. @@ -785,10 +796,13 @@ static const CPUJumpCache *tb_jmp_cache_poison(void) * the rest of the page. A block translated before the breakpoint was set= is * therefore still in the jump cache, and dispatching to it inline would s= tep * straight over the breakpoint. + * + * A block that dispatches indirectly also does not emit the icount_decr + * poll, so the dispatch is where a pending exit has to be noticed. */ static bool tcg_cpu_may_dispatch(CPUState *cpu) { - return QTAILQ_EMPTY(&cpu->breakpoints); + return QTAILQ_EMPTY(&cpu->breakpoints) && !cpu_loop_exit_requested(cpu= ); } =20 /* @@ -857,6 +871,9 @@ void tcg_kick_vcpu_thread(CPUState *cpu) =20 /* Ensure cpu_exec will see the exit request after TCG has exited. */ qatomic_store_release(&cpu->neg.icount_decr.u16.high, -1); + + /* Blocks that only dispatch indirectly do not poll; stop them chainin= g. */ + tcg_cpu_poison_jmp_cache(cpu); } =20 static inline bool icount_exit_request(CPUState *cpu) diff --git ./accel/tcg/tcg-accel-ops.c ./accel/tcg/tcg-accel-ops.c index 560fe2554b..9eb9e861ac 100644 --- ./accel/tcg/tcg-accel-ops.c +++ ./accel/tcg/tcg-accel-ops.c @@ -106,6 +106,7 @@ void tcg_handle_interrupt(CPUState *cpu, int mask) qemu_cpu_kick(cpu); } else { qatomic_set(&cpu->neg.icount_decr.u16.high, -1); + tcg_cpu_poison_jmp_cache(cpu); } } =20 diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 8879cd626f..89d255bd04 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -45,12 +45,35 @@ bool translator_io_start(DisasContextBase *db) return true; } =20 +/* + * A block that ends in a goto_tb chains straight to its destination: noth= ing + * between the two looks at icount_decr, so the destination has to poll on + * entry. A block whose exits are all indirect does not, because the disp= atch + * itself notices -- a pending exit poisons tb_jmp_cache_probe, so the pro= be + * misses into helper_lookup_tb_ptr(), which returns the epilogue. Every = block + * therefore either polls on entry or is checked as it leaves, which bounds + * interrupt latency at one block without looking at the shape of the gues= t's + * control flow graph. + * + * Which kind a block is is not known until its last exit has been emitted= , so + * defer the decision to gen_tb_end() and emit the poll retroactively. + * + * icount needs the counter unconditionally, so it opts out. + */ +static bool defer_exit_check(uint32_t cflags) +{ + return !(cflags & CF_USE_ICOUNT); +} + static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t cflags) { TCGv_i32 count =3D NULL; TCGOp *icount_start_insn =3D NULL; =20 - if ((cflags & CF_USE_ICOUNT) || !(cflags & CF_NOIRQ)) { + tcg_ctx->exit_check_needed =3D false; + + if ((cflags & CF_USE_ICOUNT) || + (!(cflags & CF_NOIRQ) && !defer_exit_check(cflags))) { count =3D tcg_temp_new_i32(); tcg_gen_ld_i32(count, tcg_env, offsetof(CPUState, neg.icount_decr.u32) - @@ -76,6 +99,9 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t= cflags) */ if (cflags & CF_NOIRQ) { tcg_ctx->exitreq_label =3D NULL; + } else if (defer_exit_check(cflags)) { + /* Emitted retroactively by gen_tb_end(), if this TB emits a goto_= tb. */ + tcg_ctx->exitreq_label =3D gen_new_label(); } else { tcg_ctx->exitreq_label =3D gen_new_label(); tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_label); @@ -91,7 +117,8 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_= t cflags) } =20 static void gen_tb_end(const TranslationBlock *tb, uint32_t cflags, - TCGOp *icount_start_insn, int num_insns) + TCGOp *icount_start_insn, int num_insns, + TCGOp *first_insn_start) { if (cflags & CF_USE_ICOUNT) { /* @@ -102,6 +129,23 @@ static void gen_tb_end(const TranslationBlock *tb, uin= t32_t cflags, tcgv_i32_arg(tcg_constant_i32(num_insns))); } =20 + if (tcg_ctx->exitreq_label && defer_exit_check(cflags) && + !(cflags & CF_NOIRQ)) { + if (tcg_ctx->exit_check_needed) { + TCGv_i32 count =3D tcg_temp_new_i32(); + TCGOp *save =3D tcg_ctx->emit_before_op; + + tcg_ctx->emit_before_op =3D first_insn_start; + tcg_gen_ld_i32(count, tcg_env, + offsetof(CPUState, neg.icount_decr.u32) - + sizeof(CPUState)); + tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_la= bel); + tcg_ctx->emit_before_op =3D save; + } else { + tcg_ctx->exitreq_label =3D NULL; + } + } + if (tcg_ctx->exitreq_label) { gen_set_label(tcg_ctx->exitreq_label); tcg_gen_exit_tb(tb, TB_EXIT_REQUESTED); @@ -238,7 +282,8 @@ void translator_loop(CPUState *cpu, TranslationBlock *t= b, int *max_insns, =20 /* Emit code to exit the TB, as indicated by db->is_jmp. */ ops->tb_stop(db, cpu); - gen_tb_end(tb, cflags, icount_start_insn, db->num_insns); + gen_tb_end(tb, cflags, icount_start_insn, db->num_insns, + first_insn_start); =20 /* * Manage can_do_io for the translation block: set to false before diff --git ./include/hw/core/cpu.h ./include/hw/core/cpu.h index bd2cdd2a0b..4272740303 100644 --- ./include/hw/core/cpu.h +++ ./include/hw/core/cpu.h @@ -523,9 +523,11 @@ struct CPUState { * @tb_jmp_cache_probe: base the inline jump cache probe reads. * * Normally @tb_jmp_cache. Pointed at a shared page of zeroes to force - * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (); - * see tcg_cpu_sync_jmp_cache(). NULL before tcg_exec_realizefn() and - * after tcg_exec_unrealizefn(). + * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (), + * either because a breakpoint is set or because an exit is pending; s= ee + * tcg_cpu_sync_jmp_cache(). Only generated code and the accessors in + * cpu-exec.c may touch it. NULL before tcg_exec_realizefn() and after + * tcg_exec_unrealizefn(). */ struct CPUJumpCache *tb_jmp_cache_probe; =20 diff --git ./include/tcg/tcg.h ./include/tcg/tcg.h index 7669dc1c2d..df08c10544 100644 --- ./include/tcg/tcg.h +++ ./include/tcg/tcg.h @@ -389,6 +389,8 @@ struct TCGContext { struct TCGLabelPoolData *pool_labels; =20 TCGLabel *exitreq_label; + /* Set by goto_tb emission: this TB chains without reaching a check. */ + bool exit_check_needed; =20 #ifdef CONFIG_PLUGIN /* diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index ce77541eab..9ed04d50f3 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -2713,6 +2713,13 @@ void tcg_gen_goto_tb(unsigned idx) tcg_debug_assert((tcg_ctx->goto_tb_issue_mask & (1 << idx)) =3D=3D 0); tcg_ctx->goto_tb_issue_mask |=3D 1 << idx; #endif + /* + * A goto_tb chains straight into the destination, with nothing in bet= ween + * that looks at icount_decr, so this TB has to poll on entry. See + * defer_exit_check(). + */ + tcg_ctx->exit_check_needed =3D true; + plugin_gen_disable_mem_helpers(); tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } @@ -2801,8 +2808,10 @@ void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, co= nst TranslationBlock *tb) } =20 /* - * No icount_decr poll is needed for this exit: the helper is called on - * every dispatch and returns to the main loop while an exit is pendin= g. + * No icount_decr poll is needed for this exit. The helper returns to= the + * main loop while an exit is pending, and a pending exit poisons + * tb_jmp_cache_probe, so the inline path below finds a NULL tb and fa= lls + * into that same helper. */ plugin_gen_disable_mem_helpers(); =20 diff --git ./tests/tcg/alpha/Makefile.target ./tests/tcg/alpha/Makefile.tar= get index 1a3f541bec..334a088848 100644 --- ./tests/tcg/alpha/Makefile.target +++ ./tests/tcg/alpha/Makefile.target @@ -5,7 +5,8 @@ ALPHA_SRC=3D$(SRC_PATH)/tests/tcg/alpha VPATH+=3D$(ALPHA_SRC) =20 -ALPHA_TESTS=3Dhello-alpha test-cond test-cmov test-ovf test-cvttq test-xpa= ge-chain +ALPHA_TESTS=3Dhello-alpha test-cond test-cmov test-ovf test-cvttq test-xpa= ge-chain \ + test-indirect-irq TESTS+=3D$(ALPHA_TESTS) =20 test-cmov: EXTRA_CFLAGS=3D-DTEST_CMOV diff --git ./tests/tcg/alpha/test-indirect-irq.c ./tests/tcg/alpha/test-ind= irect-irq.c new file mode 100644 index 0000000000..df8eaed6e3 --- /dev/null +++ ./tests/tcg/alpha/test-indirect-irq.c @@ -0,0 +1,55 @@ +/* + * A loop whose only back edge is an indirect branch must still be + * interruptible. + * + * Blocks that dispatch indirectly do not emit the icount_decr poll; a pen= ding + * exit instead poisons the inline jump cache probe so that the dispatch f= alls + * into helper_lookup_tb_ptr(), which returns to the main loop. If that + * mechanism breaks, this program never leaves the loop and the test times + * out rather than failing an assertion. + * + * A computed goto is used deliberately: a plain while(1) would end the bl= ock + * with a direct backward branch, that is a goto_tb, and a block that emit= s a + * goto_tb still polls -- so it would not exercise the path under test. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include + +static volatile sig_atomic_t fired; +static volatile unsigned long iterations; + +static void handler(int sig) +{ + fired =3D 1; +} + +int main(void) +{ + /* + * Indexing a table with a volatile index, rather than jumping through= a + * volatile pointer: gcc happily proves a single-valued pointer consta= nt + * and emits a direct branch, which is the case this test is not about. + */ + void *target[2]; + volatile int idx =3D 0; + + assert(signal(SIGALRM, handler) !=3D SIG_ERR); + alarm(1); + + target[0] =3D &&spin; + target[1] =3D &&out; +spin: + iterations++; + if (!fired) { + goto *target[idx]; + } +out: + + printf("interrupted after %lu iterations\n", iterations); + return 0; +} --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807022; cv=none; d=zohomail.com; s=zohoarc; b=DfFgwCBCO8fzEeeDEYDpGGmiX+AxJ9U/MgzQ+AlXbv35rbabiVDS4q6v45KoEQLy6O6AcD7f/oDQ2P/Si0H5XPQS0igfFND57AXCLE6vOr2h6K83HpwBrh4r/ru4eXsMu6H8kL8qvHdcMN/F1ahJ83ZqdJ/FxmwDuph157JbxYQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807022; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=QYGTe5Dtc8jpyUVbafTGVnqH+rZmDr8bZimPftm6jiM=; b=MbtXWQ0Kdt750Rk/TUE+EpWOtbi1AV8lWANm2ro6vxGl6VGLJzfY+z1EHwmOTbmd5SxuHUSs3Q4GXIcuyFbTe65f3cafikYUQ6yKEB7xlv5JOa6rTVJXHNGhVZeVfXXZibw8jcQWi85U9kWNt04RyDG1spTbooAZr4GBBn4bRQw= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807022697618.9823436707264; Wed, 26 Aug 2026 22:03:42 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGj-0005s3-5A; Thu, 27 Aug 2026 01:03:17 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGd-0005rI-H6 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:12 -0400 Received: from mail-yw1-x1129.google.com ([2607:f8b0:4864:20::1129]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGb-0002Ju-6A for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:11 -0400 Received: by mail-yw1-x1129.google.com with SMTP id 00721157ae682-85a50f6a7f7so15272317b3.2 for ; Wed, 26 Aug 2026 22:03:08 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b6143ff05sm4120317b3.32.2026.08.26.22.03.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806988; x=1788411788; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QYGTe5Dtc8jpyUVbafTGVnqH+rZmDr8bZimPftm6jiM=; b=Wgi3CDrLKN+MYNEmWQCQj9Zkiaz9WVweKdfwUu7qWZviuUu4WwZdCifGHt7/wABedW Gji42kborOuKDjjzFORmCWqnhs5Qen8/0IoHb1KYl2qmOk0UcBABm76dVrSIeV1w59Lj Z5AeD7Cc+VdfjNFkzFB7hJroaoJhmHp+prWFPZsJPXwd3PTi3qWc40VXm/33vsfoesfC awnHVtLNbQ1eA1MwdQNBwYEeoviaD18yORuiZjzqo5xGidibih41OwRIt3JKERn37Ub1 KvDFLIZKABOtgHrLkzL3JT5HIBOBrhkR8GWwAwYbPOFLifzCP8w4eN61cTOyJcwp4ZJd 7pCQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806988; x=1788411788; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=QYGTe5Dtc8jpyUVbafTGVnqH+rZmDr8bZimPftm6jiM=; b=hUbdSQbapAbYwdNih6q5TgTuBQ2eNY84FWM1tMHk2BWR1vUqMmuSA34Sp3shEJ/Cq3 YIYdqTHX2E8+hgfr30Z3R/uY6TTuhS2p/we2dmv9EPW51jU0D92ZeV4qHp8ANS1YpP+T zZ1ve0JOOLv8Qpo5qOQ9OUzSqkSzyXa0c+OX0NzoMysGzH4Jou5VU6H81IgAKOaoxnuF QYrWnO1hcb/piBQ0xCxHSJWnHHDc96f4pOg4RE1dmZahlyd4eBORsRSc9K5+XWUWyYst Ce7Ftw7oKRjlmIsSAJQxyWLFEBecZF181lhwspvOO4cW4KS+B860QbFnCUzAMkyo/G0y UG2A== X-Gm-Message-State: AFuF++nlf3TFdyIOZovgSGP2hwsGPEEEcOF6Q8PNFJZ7hTGG5uAYdV1Z yBAMFxN7DCkcJ77lr4QdAmG+b/2CwbgrbGvEryycYkM870vvYKOfEquVhrN5Oxbc X-Gm-Gg: AR+sD10134yIECcza7FHzo+SKpFBaXII3Lqw6d5W/vUqoWGbPruhMjsdJmsc4GDPiZ+ 4a1hU8ybTHk0Y2v2S/FBahlJUugo/2Z6UDZ8+4f8YVYWuajE/RQ0REWHzfb0s4nCCyLBi3aDU3J cncRLFoqsKCW/f+vzXIvE2aw/3/hLquLv84gj5CDTXrPGUgQ2sD6jpZ0zfptwQxe9gt9udj5FEz jjFavYBV6TpN9K5xq/eEer1qjq46bE1DQZ6Dm0kQfXTXkefUTGeCQ7nvrFaPjx4UxfLt8tC5ACZ krT5FOnL14CWj7Us5vuMKsgccGRNtw0rkKpn/ieCDG1zd4x31M5IvwdqiX49sn8l12zqmPWuiGG Gnlbme18Y+ukk89WQmPkWeaXMaa+ZWWNTovDvXZ9vksp5xGxwUoZalKG9KcT0/rvnKh6+1/pFpS hgwYJ6ZUUsnTNWbDif9/v95rqD5Hmdn2Atrr2MXMyNvmfk37e3AAmpkvLSD0apF/5ENh7u94hxy DAOEnqP3vsYQ5yW6yYVf8fjvozsbgOg7bKizzG8 X-Received: by 2002:a05:690c:4b11:b0:81e:cc9f:3dfe with SMTP id 00721157ae682-8573c008d0bmr54581107b3.1.1787806987967; Wed, 26 Aug 2026 22:03:07 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 6/9] RFC: tcg: probe the TB jump cache inline instead of calling a helper Date: Thu, 27 Aug 2026 01:02:38 -0400 Message-ID: <20260827050241.3713332-7-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1129; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1129.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807024116158500 Content-Type: text/plain; charset="utf-8" Every indirect branch that cannot use goto_tb ends in tcg_gen_lookup_and_goto_ptr(), which calls helper_lookup_tb_ptr(). For an emulated compiler that is 8.4 billion helper calls in a single translation unit: 24.6% of all TB exits take this path, because jsr/ret/jmp have a register destination and because goto_tb is restricted to same-page targets. The helper itself is already tight, but each call pays for a call frame, the can_do_io store, the get_tb_cpu_state() indirect call through TCGCPUOps, curr_cflags(), and a breakpoint check, before it gets to the jump cache probe that almost always hits (95.8% for this workload). Emit the probe inline instead. The two preceding patches supply what it needs: the destination PC is in a TCG temp, the flags, cflags and cs_base the destination must match are constants at translation time, and tb_jmp_cache_probe is a base pointer the main loop can poison. The fast path is therefore a hash, four guarded loads and a goto_ptr. Only a miss calls the helper, which still owns filling the cache. Two details matter for the generated code. The flags and cflags guards are folded into a single aligned 64-bit load and compare, since the fields are adjacent. And each path emits its own goto_ptr rather than branching to a shared one: a temp live across the label is spilled and reloaded on every dispatch, which cost 6.3% on its own. The flags and cflags constants are safe against the other things that can change them. CF_PARALLEL is only ever set by begin_parallel_context(), which flushes first, so no block predating it survives to dispatch. gdb single-step is only turned on with the CPU stopped, and a block translated without CF_SINGLE_STEP can only be re-entered through tb_lookup(), which from then on demands the new cflags -- so a stale-cflags block is never the one running. What is left is one_insn_per_tb and -d nochain, which the monitor can toggle under a running vCPU without a flush; see below. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding patches: before: 1,402,667,803,616 instructions after: 916,415,123,244 instructions -34.67% before: 115.56s wall clock after: 85.59s wall clock -25.94% The gap between the two is the point at which this stops being a straight-line win: the helper call was highly predictable work that the host pipelined well, so removing it retires far fewer instructions than it saves time. IPC falls from 2.48 to 2.17 across this patch for that reason. Despite emitting more code, this also reduces instruction cache pressure, because a dispatch no longer jumps into qemu's .text and evicts translated code: before: 11,735,141,703 L1-icache-load-misses after: 7,154,863,292 L1-icache-load-misses -39.0% The mechanism is visible directly in a profile: helper_lookup_tb_ptr() falls from 31.01% of samples to 0.35%, and qemu's own .text falls from 38.8% to 5.3%, with the balance moving into generated code. Combined with the preceding patches, against an unmodified LTO build, 1,646,994,254,249 instructions fall to 916,415,123,244, or -44.36%. The emulated compiler produces byte-identical output throughout. Open issues, hence RFC: - one_insn_per_tb and CPU_LOG_TB_NOCHAIN can be toggled from the monitor while a vCPU is inside a block that was translated without them. The block keeps dispatching inline with the old cflags until it exits for some other reason. Poisoning the probe from tcg_update_all_cflags() would close it. - The jump cache entry is read without qatomic_read(); entries are invalidated concurrently by setting tb to NULL. - Only alpha has been measured. The other four targets that pass a PC are built and boot-tested only. v4: Split out of the patch that also changed the tcg_gen_lookup_and_goto_ptr() API and introduced tb_jmp_cache_probe, which are now the two preceding patches. Requested by Richard Henderson. v4: Emit the softmmu form of tb_jmp_cache_hash_func() under CONFIG_SOFTMMU rather than the user-only form everywhere. v3 emitted the user-only hash unconditionally, which was wrong for system mode and was only not a correctness bug because a wrong index simply misses. Caught by Richard Henderson. tcg-op.c is compiled once per build rather than once per target, but CONFIG_SOFTMMU is set for it, and TARGET_PAGE_BITS -- a load from target_page here -- is fixed long before any translation happens. v4: Compare the pc before testing tb for NULL. On a hash miss the pc is the field most likely to differ, and an unused entry has a zero pc that only pc 0 can match, so the tb test buys nothing ahead of it. Suggested by Richard Henderson. v4: Assert that offsetof(TranslationBlock, flags) is 8-byte aligned, since folding the flags and cflags guards into one 64-bit load relies on it and nothing else does. Requested by Richard Henderson. v4: Zero-extend a 32-bit guest PC instead of falling back to the helper. Suggested by Richard Henderson. The high half then folds to a compare against zero. v4: Describe cs_base in the probe as a second word of target-specific flags rather than by name. Suggested by Richard Henderson. Signed-off-by: Matt Turner --- include/tcg/tcg-op-common.h | 14 ++-- tcg/tcg-op.c | 126 ++++++++++++++++++++++++++++++++++++ 2 files changed, 133 insertions(+), 7 deletions(-) diff --git ./include/tcg/tcg-op-common.h ./include/tcg/tcg-op-common.h index 34102b3b7a..b8399229c9 100644 --- ./include/tcg/tcg-op-common.h +++ ./include/tcg/tcg-op-common.h @@ -81,13 +81,13 @@ void tcg_gen_goto_tb(unsigned idx); * * If the TB is not valid, jump to the epilogue. * - * The lookup is a call to helper_lookup_tb_ptr(). @pc and @tb describe t= he - * destination for a faster lookup that a later patch adds, and neither is - * used yet. When @pc is non-NULL it must hold exactly the value - * get_tb_cpu_state() reports as the pc for the destination, and the - * destination must match @tb's flags, cflags and cs_base. A target whose - * pc is derived rather than being that key -- avr's word address, i386's - * eip before segmentation -- must pass NULL. + * The lookup is normally a call to helper_lookup_tb_ptr(). If @pc is + * non-NULL the jump cache is probed inline instead, and only a miss reach= es + * the helper. @pc must then hold exactly the value get_tb_cpu_state() + * reports as the pc for the destination; a target whose pc is derived + * (avr's word address, i386's eip before segmentation) must pass NULL. T= he + * destination is required to match @tb's flags, cflags and cs_base, which + * is what makes them constants in the probe. * * This operation is optional. If the TCG backend does not implement goto_= ptr, * this op is equivalent to calling tcg_gen_exit_tb() with 0 as the argume= nt. diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index 2fda6e5c07..cf7b6882d8 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -28,6 +28,8 @@ #include "tcg/tcg-op-common.h" #include "exec/translation-block.h" #include "exec/plugin-gen.h" +#include "hw/core/cpu.h" +#include "../accel/tcg/tb-hash.h" #include "tcg-internal.h" #include "tcg-has.h" =20 @@ -2715,6 +2717,112 @@ void tcg_gen_goto_tb(unsigned idx) tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } =20 +static void gen_jmp_cache_hash(TCGv_i64 h, TCGv_i64 pc) +{ +#ifdef CONFIG_SOFTMMU + /* + * tb_jmp_cache_hash_func(), softmmu form. TARGET_PAGE_BITS is a load + * from target_page in this translation unit, but it is decided long + * before any translation happens, so it is a constant here. + */ + int shift =3D TARGET_PAGE_BITS - TB_JMP_PAGE_BITS; + TCGv_i64 tmp =3D tcg_temp_ebb_new_i64(); + + tcg_gen_shri_i64(tmp, pc, shift); + tcg_gen_xor_i64(tmp, tmp, pc); + tcg_gen_shri_i64(h, tmp, shift); + tcg_gen_andi_i64(h, h, TB_JMP_PAGE_MASK); + tcg_gen_andi_i64(tmp, tmp, TB_JMP_ADDR_MASK); + tcg_gen_or_i64(h, h, tmp); + tcg_temp_free_i64(tmp); +#else + /* tb_jmp_cache_hash_func(), user-only form. */ + tcg_gen_shri_i64(h, pc, TB_JMP_CACHE_BITS); + tcg_gen_xor_i64(h, h, pc); + tcg_gen_andi_i64(h, h, TB_JMP_CACHE_SIZE - 1); +#endif +} + +static void gen_jmp_cache_probe(TCGv_i64 pc, const TranslationBlock *tb) +{ + TCGv_ptr jc, ent, tbp, ptr; + TCGv_i64 h, tmp; + TCGLabel *slow; + uint64_t fpair; + + QEMU_BUILD_BUG_ON(sizeof(((CPUJumpCache *)0)->array[0]) !=3D 16); + QEMU_BUILD_BUG_ON(offsetof(CPUJumpCache, array[0].pc) % 8 !=3D 0); + /* One 64-bit load has to cover both, so they must be adjacent... */ + QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, cflags) !=3D + offsetof(TranslationBlock, flags) + 4); + /* ...and aligned, which nothing else currently relies on. */ + QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, flags) % 8 !=3D 0); + + jc =3D tcg_temp_ebb_new_ptr(); + ent =3D tcg_temp_ebb_new_ptr(); + tbp =3D tcg_temp_ebb_new_ptr(); + ptr =3D tcg_temp_ebb_new_ptr(); + h =3D tcg_temp_ebb_new_i64(); + tmp =3D tcg_temp_ebb_new_i64(); + slow =3D gen_new_label(); + + /* ent =3D &jc->array[tb_jmp_cache_hash_func(pc)] */ + gen_jmp_cache_hash(h, pc); + tcg_gen_shli_i64(h, h, 4); + + /* + * Not cpu->tb_jmp_cache: the probe reads its own base so that the main + * loop can poison it, which is how conditions the probe cannot test f= or + * itself force every dispatch back into the helper. See + * tcg_cpu_sync_jmp_cache(). + */ + tcg_gen_ld_ptr(jc, tcg_env, + offsetof(CPUState, tb_jmp_cache_probe) - sizeof(CPUStat= e)); + tcg_gen_trunc_i64_ptr(ent, h); + tcg_gen_add_ptr(ent, jc, ent); + + /* + * The pc first: on a hash miss it is the field most likely to differ, + * and an entry whose tb is NULL has a zero pc that only pc 0 matches. + */ + tcg_gen_ld_i64(tmp, ent, offsetof(CPUJumpCache, array[0].pc)); + tcg_gen_brcond_i64(TCG_COND_NE, tmp, pc, slow); + + tcg_gen_ld_ptr(tbp, ent, offsetof(CPUJumpCache, array[0].tb)); + tcg_gen_brcondi_ptr(TCG_COND_EQ, tbp, 0, slow); + + /* + * flags and cflags are adjacent uint32_t, so one aligned 64-bit load + * and compare covers both. + */ +#if HOST_BIG_ENDIAN + fpair =3D ((uint64_t)tb->flags << 32) | tb->cflags; +#else + fpair =3D ((uint64_t)tb->cflags << 32) | tb->flags; +#endif + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, flags)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, fpair, slow); + + /* + * cs_base is a second word of target-specific flags despite the name, + * and the pc alone does not imply it on a target that uses it. + */ + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, cs_base)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, tb->cs_base, slow); + + tcg_gen_ld_ptr(ptr, tbp, offsetof(TranslationBlock, tc.ptr)); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); + + /* + * Emit a second goto_ptr rather than branching to a shared one: a temp + * live across the label would be spilled and reloaded on every dispat= ch. + */ + gen_set_label(slow); + ptr =3D tcg_temp_ebb_new_ptr(); + gen_helper_lookup_tb_ptr(ptr, tcg_env); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); +} + void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, const TranslationBlock *= tb) { TCGv_ptr ptr; @@ -2726,6 +2834,24 @@ void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, co= nst TranslationBlock *tb) =20 plugin_gen_disable_mem_helpers(); =20 + if (pc) { + TCGv_i64 pc64; + + /* + * The jump cache is keyed on a vaddr, so a 32-bit guest PC is + * compared as its zero-extension. The high half then folds to a + * constant compare against zero. + */ + if (pc->type =3D=3D TCG_TYPE_I32) { + pc64 =3D tcg_temp_ebb_new_i64(); + tcg_gen_extu_i32_i64(pc64, temp_tcgv_i32(pc)); + } else { + pc64 =3D temp_tcgv_i64(pc); + } + gen_jmp_cache_probe(pc64, tb); + return; + } + ptr =3D tcg_temp_ebb_new_ptr(); gen_helper_lookup_tb_ptr(ptr, tcg_env); tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787425765; cv=none; d=zohomail.com; s=zohoarc; b=Jzv4rYeAz0DNq9c+4Dbxu6qzzGk0Hh09oayb1gUWiP10DxnaTDciDMvFKD4TUgokmj2ItdgSRdF5yl8XmnfIY92dhwMn7uKXqqRQMJTLfdciygnm64iYxuVDkxai5QKadsBki2/cQXPBOR9+iVTITYwy/g8rYfePnv54MqnPkCA= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787425765; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=OHjPdAZ7tK8mubjCMhyER0kSir+YM3BIpc18K2DMUtc=; b=PXq8OckTpL/KFLrezv0EsjzB2hJyTsrZMo58E3Iit3rGaBe9FcAagcvhOYZr1dGNMszq4DKhZrVfNCa45OfZ/QHTFWvqGkK2M89NosN8gKOPdQTIeMzimJp9xH6h3bouYGnHFAYAEUzq25mVGTnlvOUQ4jLBf88bb4RELtSHn6k= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 17874257659521021.1529953666047; Sat, 22 Aug 2026 12:09:25 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wxr5D-0004C1-49; Sat, 22 Aug 2026 15:08:47 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wxr5C-0004Bp-6Z for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:46 -0400 Received: from mail-yw1-x1130.google.com ([2607:f8b0:4864:20::1130]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wxr59-0005mZ-Qb for qemu-devel@nongnu.org; Sat, 22 Aug 2026 15:08:45 -0400 Received: by mail-yw1-x1130.google.com with SMTP id 00721157ae682-81dfdbd86d1so29722787b3.1 for ; Sat, 22 Aug 2026 12:08:43 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84ca5188982sm13908967b3.10.2026.08.22.12.08.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 22 Aug 2026 12:08:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787425723; x=1788030523; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OHjPdAZ7tK8mubjCMhyER0kSir+YM3BIpc18K2DMUtc=; b=N1YP0U/XlH2kL751uZDPw/uDAI7K8jKcSalmDr51bWxNaMVmt2ZYIlsHz6PgUs1bhC UVDIeWY8MzT1/CPVaWQvY/pFCyEDVgdjYiB8HLoBz0MN9lFeFv9FPLlE9yZ8I4DqQawh dXXEKYoyFEL8bpz13bNZvZNbxcYiVfwNFgHLZ+AnATb1P6FVgMSI0u3h4AQrl7iRcr3p Wfhuam0kdG03CemWEUXUF2yp3xhh17q0o+YpNDiyNAPSmVi0DKsDshWEcQ/GFGc+sPj8 ZwxAvCp47gYK9NPRFQwaO6v3wtL2exLer4Jf3T7XxIjvQMEBKGqh3jpXieALzs2+m9tc b1lA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787425723; x=1788030523; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=OHjPdAZ7tK8mubjCMhyER0kSir+YM3BIpc18K2DMUtc=; b=Njm/Jh19REiW9SWlpOSyHC+eP5v9qLzELn/dwusLY4h2aEZEWp+SVXs/dXdJzTrzE+ wLzwnf69NaLQA+gm2HrHd4JdD4rZa0re+t0GDPddfe7yyqgABGCjh/8c2YSK1IL9omIW coe0qNYfFtuiCkxeerqHz2GRow8t5NzWc3bqKyx1NZVjlkoSk9dUcS9UEmVbmpuXcAnH DKas5fZ7fpf2Io2na50Jtn20vRNCJMuM09rnt4Sdc9iOyiu8Zv99eluJH6GNwSCKLe1+ OcwXEugFI5RZ8GmZLdVR+WChv0/4IOSVMD7sFmOhhijhpPSmQQ5XJL4jXMJOqLSHTDla hiXg== X-Gm-Message-State: AFuF++njUQ5LUa/DnnmEKUQr6O/OXDFrqFUrqmYM+FkA315F2pjJgb4X UOaxA0mPmw4hfi/b3QUQZ8K7y+XLHF67mZW2Q5T2FL8yNwK2M4a/f65DjkLxCps9oNA= X-Gm-Gg: AR+sD12pcxHk0TqD+bQ/kdvrVj292ZxauZNGUQZ6I1zRwd7oK0VKXPBAAie+CouGTK1 mHz4+whNlM6aeLseiMV4ZoPOadH6wRnV4zQn+RACa+XijU7Xh77QRYc1ABfonkppAhi+QQUrmhI DU7+/XU5PCKVoUZ6Bb61WyxSZzRsNhGlprIZmL6osYmsOdPrNjlzUZU+m9WdR87gFPHtF1pjk3a vi5dfsxj0hZ8CyC+gw+OEocryY4fMo0W3NBO5RJDG+auP85o/ro4P6znCWJdol6IGZR7Qz8V9vW nya10gQo/w5w/Z4DuwP9lQ22Bm0N/gQgaoX/JbKbgExIyZOVwi265udIGE3SRlZHgOVpCIOmst6 dnSz+Kbjz3hLsTHXTuhhoJSiTQmb2gkkCXn4fpjkAIHxPx+bj4a/BFWeptOoJu4k84nSqfJDUnN Nb38VwRIzLDGggqmdyJTtEWIHLGWVu6g0PjoRcYoSwBZAB4TMv6w2OedumS4CIWkJXZFfNwXKoy jb1mQ8vybSD+Ck4H5MG6WjQfcgHAJHxBOA0uzPk X-Received: by 2002:a05:690c:63ca:b0:80e:5236:b944 with SMTP id 00721157ae682-849f2a39745mr53932557b3.10.1787425722386; Sat, 22 Aug 2026 12:08:42 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v3 7/7] RFC: tcg: fold a guest displacement into the host addressing mode Date: Sat, 22 Aug 2026 15:08:18 -0400 Message-ID: <20260822190818.1829249-8-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1130; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1130.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787425766922158500 Content-Type: text/plain; charset="utf-8" Nothing in the TCG frontend interface can express a based memory access. tcg_gen_qemu_ld/st take an address and nothing else, so a target with a displacement in its load and store encodings -- which is most of them -- has to materialise the address first: ldq a1,8(a0) -> mov 0x80(%rbp),%rbx reload a0 lea 0x8(%rbx),%r12 address mov (%r12),%r12 the load mov %r12,0x88(%rbp) spill a1 The lea is pure loss on a host whose addressing mode has a displacement field sitting empty. It also needs a register, at the point in a block where pressure is highest. Fold it. After optimisation, look for an add of a constant immediately before a guest access, defining that access's address operand, and move the constant into a new second constant argument on the op. The add is left for liveness to remove, so nothing breaks if its result has another use. Only the immediately preceding op is examined: that is what the frontends emit, and a window of one op means the pass does not have to reason about what could have happened in between. The one thing it does check is that the add did not clobber the base it read, since the access now reads that base directly. Targets opt in with TCG_TARGET_HAS_ldst_disp and an out_disp member on TCGOutOpQemuLdSt. Without it the pass does not run, the displacement stays zero and the existing out member is called exactly as before, so no other backend changes behaviour or needs touching. For x86_64 the displacement goes in the disp32 that prepare_host_addr() already fills in for guest_base. The fold is refused unless there is no slow path at all, which means user-only -- softmmu compares the unadjusted address against the TLB -- and an access needing no alignment test, since the slow path hands addr_reg to the helper and that register no longer holds the full guest address. It is also refused if guest_base plus the displacement leaves disp32. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding patches, against a control measured in the same session: before: 868,811,832,620 instructions, 79.85s after: 819,262,147,022 instructions, 77.30s -5.70% instructions, -3.20% wall Emitted code shrinks from 50.55MB to 48.80MB over the run, 167.4 to 161.6 bytes per block. Per Alpha opcode, the host bytes emitted for an access fall as expected and nothing else moves: ldq 18.3 -> 15.4 ldah 20.9 -> 20.9 ldl 16.6 -> 14.1 lda 12.9 -> 12.9 stq 12.8 -> 9.7 mov 9.8 -> 9.8 The emulated compiler produces byte-identical output and the alpha tests still pass, including with a non-zero guest_base forced via -B. RFC because: - Only wired up for x86_64, and only for qemu_ld and qemu_st; the i128 qemu_ld2 and qemu_st2 pairs are left alone. - Requiring that no slow path exists is stricter than necessary. Recording the displacement in TCGLabelQemuLdst and emitting one lea on the slow path would cover alignment-checked accesses too, at no fast path cost. - Softmmu wants the displacement folded into the TLB comparison as well, which is a bigger change than this one. - A one op window catches everything the frontends emit today but is trivially defeated by anything scheduled in between. Signed-off-by: Matt Turner --- include/tcg/tcg-opc.h | 9 +++- tcg/tcg-op-ldst.c | 3 +- tcg/tcg.c | 86 ++++++++++++++++++++++++++++++++++++- tcg/x86_64/tcg-target.c.inc | 61 ++++++++++++++++++++++++++ tcg/x86_64/tcg-target.h | 3 ++ 5 files changed, 158 insertions(+), 4 deletions(-) diff --git ./include/tcg/tcg-opc.h ./include/tcg/tcg-opc.h index f3a81d5d7f..92fd34d3e3 100644 --- ./include/tcg/tcg-opc.h +++ ./include/tcg/tcg-opc.h @@ -125,8 +125,13 @@ DEF(goto_ptr, 0, 1, 0, TCG_OPF_BB_EXIT | TCG_OPF_BB_EN= D) DEF(plugin_cb, 0, 0, 1, TCG_OPF_NOT_PRESENT) DEF(plugin_mem_cb, 0, 1, 1, TCG_OPF_NOT_PRESENT) =20 -DEF(qemu_ld, 1, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) -DEF(qemu_st, 0, 2, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +/* + * The second constant argument is a displacement to add to the address, + * zero unless a target advertises TCG_TARGET_HAS_ldst_disp and the fold in + * fold_ldst_disp() applied. + */ +DEF(qemu_ld, 1, 1, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +DEF(qemu_st, 0, 2, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) DEF(qemu_ld2, 2, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) DEF(qemu_st2, 0, 3, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) =20 diff --git ./tcg/tcg-op-ldst.c ./tcg/tcg-op-ldst.c index 22211ccb45..ffc5e651a6 100644 --- ./tcg/tcg-op-ldst.c +++ ./tcg/tcg-op-ldst.c @@ -92,7 +92,8 @@ static MemOp tcg_canonicalize_memop(MemOp op, bool is64, = bool st) static void gen_ldst1(TCGOpcode opc, TCGType type, TCGTemp *v, TCGTemp *addr, MemOpIdx oi) { - TCGOp *op =3D tcg_gen_op3(opc, type, temp_arg(v), temp_arg(addr), oi); + /* The trailing zero is the address displacement; see fold_ldst_disp()= . */ + TCGOp *op =3D tcg_gen_op4(opc, type, temp_arg(v), temp_arg(addr), oi, = 0); TCGOP_FLAGS(op) =3D get_memop(oi) & MO_SIZE; } =20 diff --git ./tcg/tcg.c ./tcg/tcg.c index 489df0e738..9e6da41887 100644 --- ./tcg/tcg.c +++ ./tcg/tcg.c @@ -1058,6 +1058,13 @@ typedef struct TCGOutOpQemuLdSt { TCGOutOp base; void (*out)(TCGContext *s, TCGType type, TCGReg dest, TCGReg addr, MemOpIdx oi); + /* + * As out(), for an access at addr + disp. Only required of targets th= at + * define TCG_TARGET_HAS_ldst_disp; for everyone else fold_ldst_disp() + * never runs and the displacement is always zero. + */ + void (*out_disp)(TCGContext *s, TCGType type, TCGReg dest, + TCGReg addr, MemOpIdx oi, int32_t disp); } TCGOutOpQemuLdSt; =20 typedef struct TCGOutOpQemuLdSt2 { @@ -3574,6 +3581,77 @@ static void move_label_uses(TCGLabel *to, TCGLabel *= from) QSIMPLEQ_CONCAT(&to->branches, &from->branches); } =20 +#ifndef TCG_TARGET_HAS_ldst_disp +#define TCG_TARGET_HAS_ldst_disp 0 +static bool tcg_target_ldst_disp_ok(TCGContext *s, MemOpIdx oi, int64_t di= sp) +{ + return false; +} +#endif + +/* + * Fold "add addr, base, $disp" into the guest access that follows it, so + * that the displacement becomes part of the host addressing mode instead = of + * a separate instruction. Frontends have no way to express this: there is + * no displacement operand on tcg_gen_qemu_ld/st, so a based access always + * costs an extra add, and an extra register to hold its result. + * + * Only an add in the op immediately before the access is recognised. That + * is what the frontends emit, and a window of one op means no analysis is + * needed of what might have happened in between. The add is left in place; + * liveness removes it if its result has no other use. + */ +static void __attribute__((noinline)) +fold_ldst_disp(TCGContext *s) +{ + TCGOp *op; + + if (!TCG_TARGET_HAS_ldst_disp) { + return; + } + + QTAILQ_FOREACH(op, &s->ops, link) { + TCGOp *prev; + TCGTemp *cts; + int64_t disp; + + switch (op->opc) { + case INDEX_op_qemu_ld: + case INDEX_op_qemu_st: + break; + default: + continue; + } + + prev =3D QTAILQ_PREV(op, link); + if (prev =3D=3D NULL || prev->opc !=3D INDEX_op_add || + TCGOP_TYPE(prev) !=3D s->addr_type) { + continue; + } + + /* + * The add must define the address operand, and must not have + * clobbered the base it read: after the fold the access reads the + * base directly, so the base has to still hold its original value. + */ + if (prev->args[0] !=3D op->args[1] || prev->args[0] =3D=3D prev->a= rgs[1]) { + continue; + } + + cts =3D arg_temp(prev->args[2]); + if (cts->kind !=3D TEMP_CONST) { + continue; + } + disp =3D cts->val; + if (disp =3D=3D 0 || !tcg_target_ldst_disp_ok(s, op->args[2], disp= )) { + continue; + } + + op->args[1] =3D prev->args[1]; + op->args[3] =3D disp; + } +} + /* Reachable analysis : remove unreachable code. */ static void __attribute__((noinline)) reachable_code_pass(TCGContext *s) @@ -5728,7 +5806,12 @@ static void tcg_reg_alloc_op(TCGContext *s, const TC= GOp *op) const TCGOutOpQemuLdSt *out =3D container_of(all_outop[op->opc], TCGOutOpQemuLdSt, base); =20 - out->out(s, type, new_args[0], new_args[1], new_args[2]); + if (new_args[3]) { + out->out_disp(s, type, new_args[0], new_args[1], + new_args[2], new_args[3]); + } else { + out->out(s, type, new_args[0], new_args[1], new_args[2]); + } } break; =20 @@ -6611,6 +6694,7 @@ int tcg_gen_code(TCGContext *s, TranslationBlock *tb,= uint64_t pc_start) tcg_temp_ebb_reset_freed(s); =20 tcg_optimize(s); + fold_ldst_disp(s); =20 reachable_code_pass(s); liveness_pass_0(s); diff --git ./tcg/x86_64/tcg-target.c.inc ./tcg/x86_64/tcg-target.c.inc index 2c8f1f3e58..d72c3db13d 100644 --- ./tcg/x86_64/tcg-target.c.inc +++ ./tcg/x86_64/tcg-target.c.inc @@ -2027,6 +2027,39 @@ static TCGLabelQemuLdst *prepare_host_addr(TCGContex= t *s, HostAddress *h, return ldst; } =20 +/* + * Whether the displacement of a guest access can be folded into the host + * addressing mode rather than materialised by a separate lea. + * + * Folding rewrites the access to use base + disp, so nothing may need a + * register holding the complete guest address. The softmmu TLB comparison + * does, and so does any slow path, which hands addr_reg to the helper. In + * user-only mode prepare_host_addr() creates a slow path only for an + * alignment test, so requiring that none is needed rules it out. What is + * left to check is guest_base, which shares the disp32 field. + */ +static bool tcg_target_ldst_disp_ok(TCGContext *s, MemOpIdx oi, int64_t di= sp) +{ +#ifdef CONFIG_USER_ONLY + MemOp opc =3D get_memop(oi); + TCGAtomAlign aa; + int64_t ofs; + + if (tcg_use_softmmu || s->addr_type !=3D TCG_TYPE_I64) { + return false; + } + aa =3D atom_and_align_for_opc(s, opc, MO_ATOM_WITHIN16, + (opc & MO_SIZE) =3D=3D MO_128); + if (aa.align) { + return false; + } + ofs =3D (int64_t)x86_guest_base.ofs + disp; + return ofs =3D=3D (int32_t)ofs; +#else + return false; +#endif +} + static void tcg_out_qemu_ld_direct(TCGContext *s, TCGReg datalo, TCGReg da= tahi, HostAddress h, TCGType type, MemOp memo= p) { @@ -2183,9 +2216,23 @@ static void tgen_qemu_ld(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_ld_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, true); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_ld_direct(s, data, -1, h, type, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_ld =3D { .base.static_constraint =3D C_O1_I1(r, L), .out =3D tgen_qemu_ld, + .out_disp =3D tgen_qemu_ld_disp, }; =20 static void tgen_qemu_ld2(TCGContext *s, TCGType type, TCGReg datalo, @@ -2321,9 +2368,23 @@ static void tgen_qemu_st(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_st_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, false); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_st_direct(s, data, -1, h, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_st =3D { .base.static_constraint =3D C_O0_I2(L, L), .out =3D tgen_qemu_st, + .out_disp =3D tgen_qemu_st_disp, }; =20 static void tgen_qemu_st2(TCGContext *s, TCGType type, TCGReg datalo, diff --git ./tcg/x86_64/tcg-target.h ./tcg/x86_64/tcg-target.h index 7ebae56a7d..8f2315c15e 100644 --- ./tcg/x86_64/tcg-target.h +++ ./tcg/x86_64/tcg-target.h @@ -30,6 +30,9 @@ #define TCG_TARGET_NB_REGS 32 #define MAX_CODE_GEN_BUFFER_SIZE (2 * GiB) =20 +/* A guest displacement can go in the disp32 of the addressing mode. */ +#define TCG_TARGET_HAS_ldst_disp 1 + typedef enum { TCG_REG_EAX =3D 0, TCG_REG_ECX, --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807055; cv=none; d=zohomail.com; s=zohoarc; b=O42w0jV8IssVUQpRhjrCQSQoo4pRpY3WCrc2m/P8b7TrXRQETrLoQyaHkkX+v7q9OZ4Tlo2zISEIvaasijtk0wVIGQQcySHZzXfdQHTF8ZywoXdmKOlR840fP2JoI+3U7UJ3l7hNPINBJorqVYjIRTZwZgh4GBLgjmA8UYEPvU0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807055; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=Lra74KjfPiopd29u8wDtZNL7yuLaPJ5gVF0OQzSZUSHaKp20xmEaklVxgKHREqBUrZdHbOnaMe7x22p0O7Zg8Zr7NaJgEmCS0dJI5kgbAtbXOWJVx94jDFx3y0bw9h3eOWz+uzPsp2CLOUNwJRcLEjy+aPTA/vyVSc9AeWBr2os= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807055595884.0655465993174; Wed, 26 Aug 2026 22:04:15 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGo-0005xk-2H; Thu, 27 Aug 2026 01:03:22 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGj-0005ta-Sn for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:17 -0400 Received: from mail-yw1-x112a.google.com ([2607:f8b0:4864:20::112a]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGf-0002KY-P4 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:16 -0400 Received: by mail-yw1-x112a.google.com with SMTP id 00721157ae682-85aeb506b78so7446527b3.3 for ; Wed, 26 Aug 2026 22:03:12 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b61f249f2sm4170677b3.41.2026.08.26.22.03.08 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806991; x=1788411791; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=lpt85zE8lJgotVhsn1N65g0wpJ+XlTGNz9QbMpQ/ywuSYfx1L9cyyGvHzOyiIBCWM9 +UYQSjEvXk/NT3cuBoUrBicIkY2lWpi4w797hAxplZyBthtqGod4cPt6swJxk2SxdkWt JqsUWXzeTbkZD06cS7f5Vn7FU+QpTXTpbgSJTOPER4+ZalbRS8GmLFnEtwdTsRjGwslC yJQMkG480a8cI4XmyaJK8UyQGGBHPSLyRvCCX2VjPSFUz+XQOATlM3jvGBLCc+g1dyoO 6/1IbvRasGgpSQWyBDGWlZ0mHQpgQryNpfcWfup2akvN+Pbs6giG+L1aWlUrdD4Ac63t TTyA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806991; x=1788411791; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=FT3aGi2B+WuitiywZPl7xOZfwiVRkayDxqhuGkA4wujBKxB5wmyTrWzeoc7ncfJA5N KU1NlWg+WYgTS58AE9CHhtAxEUYk9l2mlCwuymQTil89Qz+09rpoA6k1/Up2Nlm6+TCT b3uuQ3DSUcizC6DyFfNxGsxcXovTxmT5zK+rVoqskc7ogGPYf7dXXwLUrPx2nIO26r0b KlpmaU2YvC8VGjFlOKGZf+QAWwZh10ui+c4Hs+XDljkZN8xM9Knfonn0rI9INMTo1xNE wsovCdr2nA+Vl3BhdQ2tDQ5/4X9eVbJceFH6HdoL3SZ3w1COJYu83Ekpl2b9IBxiQ3Tj 62OA== X-Gm-Message-State: AFuF++lsN2CnuqTRqimAhPZxG38/k+NoFiH7ytW5hvZ5+p1OfiJvyXxu dF5zOevt5freJfH0fNJjBBxDrNNtvNP0q/ZG0d52eU9nMFgkcmMIc/iSyiWVqTEx X-Gm-Gg: AR+sD13brbrboUJcjIpBtxEbihxfkXuYFJTun5XJxOYhV86qjEF0Zk3AaSk4fI2bw// 9FhGIuDsUlRfzZUQyS/2SgHidHdLvnlkxF9Qz4xFl0r4FEEcroskPFWXP9Ghm2hMqB7F2xIfFZO nAxW0UBKi9ylRmpfNM4NIgFMp+o7kZPUWNiN5m0WnChk9oYeu5gLdZFctx2Mam5T5Lr3lUNuNWR XSDf6Mu92HUeo7+tCr+TbNZKQymJMcogtXqgPnB6dLIwjVaibKo85IVh2m8WGC1BZb8dtAiq1uW xP8Z4IeTfC7ETUkZuTh96+uFSGucSzGFDRXBQsIu+r5yb7CdUoqaBIzIGa2ZTYnxtxXumzGMF0c IYkQw0K+n+bZlVr+1BamzAXeip9y5GL5AKbyDSCi+YxNprhw3rwfFmLSjSAMMSS4MTeiBwPJQeE COC2nD4f429eAd7rOLOkze8SGWiQuV/lh+VgKyV1gPV0PKgyHPD740DrQh1yFJATJAVoPWZe28g nBYIx6XpQ1F1HPNvS5E9QON2f3XmTSPeT1CG/PQ X-Received: by 2002:a05:690c:4026:b0:81e:c17d:7b8f with SMTP id 00721157ae682-85741ef0e6bmr48337037b3.31.1787806991297; Wed, 26 Aug 2026 22:03:11 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 7/9] RFC: accel/tcg: allow cross-page goto_tb chaining in user-only builds Date: Thu, 27 Aug 2026 01:02:39 -0400 Message-ID: <20260827050241.3713332-8-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112a; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112a.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807056252158500 Content-Type: text/plain; charset="utf-8" translator_use_goto_tb() refuses to chain unless the destination is on the same page as the start of the TB. For guests whose text is much larger than a page this is expensive: an emulated alpha gcc compiling a 255k line translation unit takes the indirect dispatch path for 8.4 billion of its 34.2 billion TB exits, and a large share of those are ordinary direct branches that simply crossed an 8 KiB page boundary. The restriction was made unconditional by d3a2a1d803 ("accel/tcg: Introduce translator_use_goto_tb"), whose rationale was: Various targets avoid the page crossing test for CONFIG_USER_ONLY, but that is wrong: mmap and mprotect can change page permissions. That is true, but in user-only builds the invalidation path already covers it. There are no page tables: every mmap, mprotect and munmap reaches page_set_flags(), which calls tb_invalidate_phys_range() whenever the flags actually change, and tb_phys_invalidate() calls tb_jmp_unlink() to reset incoming jumps. A chained cross-page jump is therefore broken whenever the destination page's permissions change. This is not true in system mode, where TBs are keyed by physical address and a page table change invalidates nothing, so the restriction is kept there. The rule protects one more thing, which the original rationale does not mention: it guarantees that execution cannot enter a page without a TB lookup, and so without check_for_breakpoints(). That is what makes a breakpoint set after a block was translated take effect, since insertion deliberately invalidates nothing. A link established before the breakpoint was set would jump straight over it. So the chaining is only enabled for a run that can never acquire a breakpoint. In user-only mode every breakpoint comes from gdb -- BP_CPU is g_assert_not_reached() there, and the guest cannot ask for one -- and gdb has to be requested with -g before the first block is translated, even though with suspend=3Dn it may connect later. gdb_may_set_breakpoints() reports whether it was, and is fixed for the lifetime of the process. Add tests/tcg/multiarch/test-xpage-chain.c to cover both hazards directly. It writes the last instruction of one page and the first of the next, so that the fall-through between them is a cross-page goto_tb, runs it 200000 times so the chain is established, then checks that mprotect(PROT_NONE) makes the next call fault, and that different code written into the page once it is mapped back runs rather than a stale translation. The two instructions -- set the return value register, and return -- are all the architecture specific code there is; thirteen architectures supply them and the rest skip. The test detects the hazard it is meant to detect: with the tb_invalidate_phys_range() call in page_set_flags() commented out, it fails both phases, executing page B after PROT_NONE and returning the stale result. Run with -b, the same binary stops once the chain is established and lets tests/tcg/multiarch/gdbstub/xpage-bp.py set a breakpoint on the far side of= it, which the next call has to stop on. With gdb_may_set_breakpoints() forced to false so that the chaining stays on under gdb, that breakpoint is missed and the test fails, which is what makes it a test of the gate rather than of gdb. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation on an x86-64 host, LTO build, on top of the preceding patches: before: 916,415,123,244 instructions after: 891,254,240,071 instructions -2.75% before: 85.59s wall clock after: 81.45s wall clock -4.84% Note that this is worth more in time than in instructions, the reverse of the preceding patch: a chained jump replaces a cache probe whose loads can miss, so the instructions it removes are more expensive than average. Measured before the inline jump cache probe, when a missed chain cost a helper call rather than an inline probe, the same change was worth -7.9%. RFC because this reverses a deliberate decision and the reasoning above wants review from someone who knows the invalidation paths better than I do. v3: Only take the shortcut when no gdbstub was requested. The same-page rule also forces a lookup, and so a breakpoint check, on entry to every page; without that, a chain established before a breakpoint was set runs past it. Reported by Richard Henderson. v3: Change translator_use_goto_tb() rather than translator_is_same_page(). i386, riscv and s390x call translator_is_same_page() for something else -- enforcing that only a single-insn TB may cross a page -- and v2 changed their TB boundaries in user-only mode as a side effect. alpha does not call it, so the numbers above are unaffected. v3: Add the gdbstub half of the test. v4: Move the test to tests/tcg/multiarch so that every *-user target runs it, rather than only alpha. Requested by Alex Bennee. The direct branch is gone with it: a fall-through off the end of a page is a cross-page goto_tb just the same, and needs no per-architecture branch encoding or displacement arithmetic, only "set the return value" and "return". Built and run under qemu-user on aarch64, alpha, arm, hppa, loongarch64, m68k, mips, ppc, ppc64le, riscv64, s390x, sh4, sparc64 and x86_64; ppc64 ELFv1 skips, because a function pointer there is a descriptor rather than a code address. Signed-off-by: Matt Turner --- accel/tcg/translator.c | 33 ++- gdbstub/user.c | 14 + include/gdbstub/user.h | 11 + tests/tcg/multiarch/Makefile.target | 12 +- tests/tcg/multiarch/gdbstub/xpage-bp.py | 37 +++ tests/tcg/multiarch/test-xpage-chain.c | 336 ++++++++++++++++++++++++ 6 files changed, 441 insertions(+), 2 deletions(-) create mode 100644 tests/tcg/multiarch/gdbstub/xpage-bp.py create mode 100644 tests/tcg/multiarch/test-xpage-chain.c diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 6c8fcd7a20..8879cd626f 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -15,6 +15,9 @@ #include "accel/tcg/cpu-mmu-index.h" #include "exec/target_page.h" #include "exec/translator.h" +#ifdef CONFIG_USER_ONLY +#include "gdbstub/user.h" +#endif #include "exec/plugin-gen.h" #include "tcg/tcg-op-common.h" #include "internal-common.h" @@ -110,6 +113,34 @@ bool translator_is_same_page(const DisasContextBase *d= b, vaddr addr) return ((addr ^ db->pc_first) & TARGET_PAGE_MASK) =3D=3D 0; } =20 +/* + * Whether a direct jump may be chained to a destination outside the page + * the TB started in. + * + * In user-only mode there are no page tables. Every mmap, mprotect and + * munmap goes through page_set_flags(), which calls tb_invalidate_phys_ra= nge() + * whenever the flags actually change, and tb_phys_invalidate() unlinks + * incoming jumps. A cross-page link is therefore broken whenever the + * destination page's permissions change. + * + * What the same-page rule also provides is that execution cannot enter a = page + * without a TB lookup, and so without check_for_breakpoints(), which is w= hat + * makes a breakpoint set after a block was translated take effect. Nothi= ng + * invalidates on breakpoint insertion, so a link established beforehand w= ould + * jump straight over it. In user-only mode breakpoints only ever come fr= om + * gdb -- BP_CPU is g_assert_not_reached() there and the guest has no way = to + * ask for one -- and gdb has to be requested with -g before the first blo= ck + * is translated, so a run that has no gdbstub can never acquire a breakpo= int. + */ +static bool use_cross_page_goto_tb(void) +{ +#ifdef CONFIG_USER_ONLY + return !gdb_may_set_breakpoints(); +#else + return false; +#endif +} + bool translator_use_goto_tb(DisasContextBase *db, vaddr dest) { /* Suppress goto_tb if requested. */ @@ -118,7 +149,7 @@ bool translator_use_goto_tb(DisasContextBase *db, vaddr= dest) } =20 /* Check for the dest on the same page as the start of the TB. */ - return translator_is_same_page(db, dest); + return use_cross_page_goto_tb() || translator_is_same_page(db, dest); } =20 void translator_loop(CPUState *cpu, TranslationBlock *tb, int *max_insns, diff --git ./gdbstub/user.c ./gdbstub/user.c index 9e6f9a6f37..d810f0f38c 100644 --- ./gdbstub/user.c +++ ./gdbstub/user.c @@ -470,6 +470,18 @@ static void *gdbserver_accept_thread(void *arg) =20 #define USAGE "\nUsage: -g {port|path}[,suspend=3D{y|n}]" =20 +/* + * Set before the guest runs and never cleared, so that code translated at + * any point can rely on it: with suspend=3Dn gdb may connect long after + * startup, and once connected it can insert a breakpoint at any time. + */ +static bool gdbserver_requested; + +bool gdb_may_set_breakpoints(void) +{ + return gdbserver_requested; +} + bool gdbserver_start(const char *args, Error **errp) { g_auto(GStrv) argv =3D g_strsplit(args, ",", 0); @@ -513,6 +525,8 @@ bool gdbserver_start(const char *args, Error **errp) return false; } =20 + gdbserver_requested =3D true; + if (suspend) { if (gdbserver_accept(port, gdb_fd, port_or_path)) { gdb_handlesig(first_cpu, 0, NULL, NULL, 0); diff --git ./include/gdbstub/user.h ./include/gdbstub/user.h index 654986d483..c091cd9758 100644 --- ./include/gdbstub/user.h +++ ./include/gdbstub/user.h @@ -11,6 +11,17 @@ =20 #define MAX_SIGINFO_LENGTH 128 =20 +/** + * gdb_may_set_breakpoints() - whether a breakpoint can ever be inserted + * + * In user-only mode every breakpoint comes from gdb, and gdb is only ever + * reachable if -g was given at startup, before the guest ran a single + * instruction. A run that has no gdbstub can therefore never acquire a + * breakpoint, which lets translation take shortcuts that a breakpoint + * would invalidate. Stays true once true, even if gdb detaches. + */ +bool gdb_may_set_breakpoints(void); + /** * gdb_handlesig() - yield control to gdb * @cpu: CPU diff --git ./tests/tcg/multiarch/Makefile.target ./tests/tcg/multiarch/Make= file.target index ab4bf9c5d5..f8a91fed2c 100644 --- ./tests/tcg/multiarch/Makefile.target +++ ./tests/tcg/multiarch/Makefile.target @@ -143,6 +143,15 @@ run-gdbstub-follow-fork-mode-parent: follow-fork-mode --bin $< --test $(MULTIARCH_SRC)/gdbstub/follow-fork-mode-parent.py, \ following parents on fork) =20 +# The chaining this exercises is only enabled when no gdbstub was requeste= d, +# so what is under test here is that requesting one turns it back off. +run-gdbstub-xpage-bp: test-xpage-chain + $(call run-test, $@, $(GDB_SCRIPT) \ + --gdb $(GDB) \ + --qemu $(QEMU) --qargs "$(QEMU_OPTS)" \ + --bin "$< -b" --test $(MULTIARCH_SRC)/gdbstub/xpage-bp.py, \ + breakpoint behind an established cross-page chain) + run-gdbstub-late-attach: late-attach $(call run-test, $@, env LATE_ATTACH_PY=3D1 $(GDB_SCRIPT) \ --gdb $(GDB) \ @@ -159,7 +168,8 @@ EXTRA_RUNS +=3D run-gdbstub-sha1 run-gdbstub-qxfer-auxv= -read \ run-gdbstub-registers run-gdbstub-prot-none \ run-gdbstub-catch-syscalls run-gdbstub-follow-fork-mode-child \ run-gdbstub-follow-fork-mode-parent \ - run-gdbstub-qxfer-siginfo-read run-gdbstub-late-attach + run-gdbstub-qxfer-siginfo-read run-gdbstub-late-attach \ + run-gdbstub-xpage-bp =20 # ARM Compatible Semi Hosting Tests # diff --git ./tests/tcg/multiarch/gdbstub/xpage-bp.py ./tests/tcg/multiarch/= gdbstub/xpage-bp.py new file mode 100644 index 0000000000..f40024f16d --- /dev/null +++ ./tests/tcg/multiarch/gdbstub/xpage-bp.py @@ -0,0 +1,37 @@ +"""Test that a breakpoint set after a cross-page chain is established is h= it. + +translator_use_goto_tb() lets a direct branch chain to another page in +user-only builds, which is only safe because a run with no gdbstub can nev= er +acquire a breakpoint. This runs with one, so the chaining must be off and +the breakpoint must still be reached. + +This runs as a sourced script (via -x, via run-test.py). + +SPDX-License-Identifier: GPL-2.0-or-later +""" +from test_gdbstub import main, report + + +def run_test(): + """Run through the tests one by one""" + gdb.Breakpoint("break_here") + gdb.execute("continue") + + # The chain exists by now; put a breakpoint on the far side of it. + target =3D int(gdb.parse_and_eval("(unsigned long)page_b_entry")) + if target =3D=3D 0: + report(True, "no code emitters for this architecture, skipped") + return + gdb.execute("break *{}".format(target)) + gdb.execute("continue") + + pc =3D int(gdb.parse_and_eval("(unsigned long)$pc")) + report(pc =3D=3D target, "stopped at {:#x}, expected {:#x}".format(pc,= target)) + + gdb.execute("delete") + gdb.execute("continue") + exitcode =3D int(gdb.parse_and_eval("$_exitcode")) + report(exitcode =3D=3D 0, "{} =3D=3D 0".format(exitcode)) + + +main(run_test) diff --git ./tests/tcg/multiarch/test-xpage-chain.c ./tests/tcg/multiarch/t= est-xpage-chain.c new file mode 100644 index 0000000000..a4e34149e7 --- /dev/null +++ ./tests/tcg/multiarch/test-xpage-chain.c @@ -0,0 +1,336 @@ +/* + * Cross-page TB chaining hazard test. + * + * Two adjacent pages of hand-written code. The last instruction of page A + * sets the return value and falls through into page B, which returns; a TB + * always ends at a page boundary, so page A reaches page B through a + * cross-page goto_tb. + * + * Phase 1: run it enough times that QEMU chains TB_A -> TB_B. + * Phase 2: mprotect page B away. Re-running must fault. + * Phase 3: map it back and write different code into it. Re-running must + * execute the NEW code, not a stale chained translation. + * + * With -b, phases 2 and 3 are replaced by a stop at break_here(), where t= he + * gdbstub test sets a breakpoint on page B -- after the chain exists -- a= nd + * checks that re-running the chain still stops on it. See + * tests/tcg/multiarch/gdbstub/xpage-bp.py. + * + * The code the two pages hold is architecture specific, so each + * architecture supplies two emitters: + * + * emit_set_ret(p, val) - set the integer return value register to val + * emit_ret(p) - return to the caller + * + * both writing at @p and returning the number of bytes written. Neither + * may contain a branch: the fall-through from page A into page B is the + * whole point, and a delay slot must not straddle the boundary. An + * architecture that supplies neither skips the test. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include +#include +#include +#include +#include + +static inline size_t put32(void *p, uint32_t insn) +{ + memcpy(p, &insn, sizeof(insn)); + return sizeof(insn); +} + +static inline size_t put16(void *p, uint16_t insn) +{ + memcpy(p, &insn, sizeof(insn)); + return sizeof(insn); +} + +#if defined(__aarch64__) +#define HAVE_EMITTERS +/* movz w0, #val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x52800000u | ((uint32_t)val << 5)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0xd65f03c0u); /* ret */ +} +#elif defined(__alpha__) +#define HAVE_EMITTERS +/* lda $0, val($31) */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x201f0000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x6bfa8001u); /* ret */ +} +#elif defined(__arm__) +#define HAVE_EMITTERS +/* mov r0, #val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0xe3a00000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0xe12fff1eu); /* bx lr */ +} +#elif defined(__hppa__) +#define HAVE_EMITTERS +/* ldi val, %ret0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x341c0000u | ((uint32_t)val << 1)); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0xe840c000u); /* bv %r0(%rp) */ + return n + put32((char *)p + n, 0x08000240u); /* nop (delay slot= ) */ +} +#elif defined(__i386__) || defined(__x86_64__) +#define HAVE_EMITTERS +/* mov $val, %eax */ +static size_t emit_set_ret(void *p, int val) +{ + uint32_t imm =3D val; + *(unsigned char *)p =3D 0xb8; + return 1 + put32((char *)p + 1, imm); +} +static size_t emit_ret(void *p) +{ + *(unsigned char *)p =3D 0xc3; /* ret */ + return 1; +} +#elif defined(__loongarch64) +#define HAVE_EMITTERS +/* ori $a0, $zero, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x03800004u | ((uint32_t)val << 10)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x4c000020u); /* jr $ra */ +} +#elif defined(__m68k__) +#define HAVE_EMITTERS +/* moveq #val, %d0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put16(p, 0x7000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + return put16(p, 0x4e75u); /* rts */ +} +#elif defined(__mips__) +#define HAVE_EMITTERS +/* li $v0, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x24020000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0x03e00008u); /* jr $ra */ + return n + put32((char *)p + n, 0x00000000u); /* nop (delay slot= ) */ +} +/* + * ELFv1 function pointers are descriptors rather than code addresses, so + * there is nothing to call the raw code through. + */ +#elif defined(__powerpc__) && \ + (!defined(__powerpc64__) || (defined(_CALL_ELF) && _CALL_ELF =3D=3D = 2)) +#define HAVE_EMITTERS +/* li r3, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x38600000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x4e800020u); /* blr */ +} +#elif defined(__riscv) +#define HAVE_EMITTERS +/* addi a0, zero, val -- the 4 byte form, never c.li */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x00000513u | ((uint32_t)val << 20)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x00008067u); /* jalr zero, 0(ra= ) */ +} +#elif defined(__s390x__) +#define HAVE_EMITTERS +/* lghi %r2, val */ +static size_t emit_set_ret(void *p, int val) +{ + size_t n =3D put16(p, 0xa729u); + return n + put16((char *)p + n, (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put16(p, 0x07feu); /* br %r14 */ +} +#elif defined(__sh__) +#define HAVE_EMITTERS +/* mov #val, r0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put16(p, 0xe000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put16(p, 0x000bu); /* rts */ + return n + put16((char *)p + n, 0x0009u); /* nop (delay slot= ) */ +} +#elif defined(__sparc__) +#define HAVE_EMITTERS +/* mov val, %o0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x90102000u | (uint32_t)(val & 0x1fff)); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0x81c3e008u); /* retl */ + return n + put32((char *)p + n, 0x01000000u); /* nop (delay slot= ) */ +} +#endif + +/* Where the fall-through lands, for the gdbstub test to breakpoint on. */ +void *page_b_entry; + +/* Somewhere for the gdbstub test to stop once the chain is established. */ +void __attribute__((noinline)) break_here(void) +{ + asm volatile (""); +} + +#ifdef HAVE_EMITTERS +static sigjmp_buf jb; +/* + * Written by the SIGSEGV handler and read by main(), so it must not be + * cached in a register across the faulting call. + */ +static volatile sig_atomic_t caught; + +static void segv(int sig) +{ + caught =3D 1; + siglongjmp(jb, 1); +} +#endif + +int main(int argc, char **argv) +{ + bool bp_mode =3D argc > 1 && strcmp(argv[1], "-b") =3D=3D 0; +#ifndef HAVE_EMITTERS + printf("SKIP: no code emitters for this architecture\n"); + if (bp_mode) { + break_here(); + } + return 0; +#else + unsigned char tmp[16]; + struct sigaction sa; + long (*fn)(void); + size_t setlen, n; + long ps =3D sysconf(_SC_PAGESIZE); + int rc =3D 0; + unsigned char *m =3D mmap(NULL, 2 * ps, PROT_READ | PROT_WRITE | PROT_= EXEC, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (m =3D=3D MAP_FAILED) { + perror("mmap"); + return 2; + } + + unsigned char *pb =3D m + ps; + + /* + * Page A ends with the store to the return value register, so that the + * next instruction executed is the first one on page B. + */ + setlen =3D emit_set_ret(tmp, 1); + memcpy(pb - setlen, tmp, setlen); + emit_ret(pb); + __builtin___clear_cache((char *)m, (char *)m + 2 * ps); + + page_b_entry =3D pb; + fn =3D (long (*)(void))(pb - setlen); + + for (int i =3D 0; i < 200000; i++) { + if (fn() !=3D 1) { + printf("FAIL: phase 1 wrong result\n"); + return 1; + } + } + printf("phase 1 ok (chained)\n"); + + if (bp_mode) { + /* + * The chain from page A to page B now exists. gdb puts a breakpo= int + * on page_b_entry here; the call below has to stop on it rather t= han + * jump over it. + */ + break_here(); + if (fn() !=3D 1) { + printf("FAIL: bp phase wrong result\n"); + return 1; + } + printf("bp phase ok\n"); + return 0; + } + + memset(&sa, 0, sizeof(sa)); + sa.sa_handler =3D segv; + sigemptyset(&sa.sa_mask); + if (sigaction(SIGSEGV, &sa, NULL) !=3D 0) { + perror("sigaction"); + return 2; + } + if (mprotect(pb, ps, PROT_NONE) !=3D 0) { + perror("mprotect"); + return 2; + } + if (sigsetjmp(jb, 1) =3D=3D 0) { + fn(); + printf("FAIL: phase 2 executed page B after mprotect(PROT_NONE)\n"= ); + rc =3D 1; + } else if (!caught) { + printf("FAIL: phase 2 longjmp without entering the handler\n"); + rc =3D 1; + } else { + printf("phase 2 ok (faulted)\n"); + } + + /* Phase 3: map back, overwrite, expect the new code to run. */ + if (mprotect(pb, ps, PROT_READ | PROT_WRITE | PROT_EXEC) !=3D 0) { + perror("mprotect back"); + return 2; + } + n =3D emit_set_ret(pb, 2); + emit_ret(pb + n); + __builtin___clear_cache((char *)pb, (char *)pb + ps); + + long r =3D fn(); + if (r !=3D 2) { + printf("FAIL: phase 3 returned %ld, expected 2 (stale chain)\n", r= ); + rc =3D 1; + } else { + printf("phase 3 ok (new code ran)\n"); + } + return rc; +#endif +} --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807013; cv=none; d=zohomail.com; s=zohoarc; b=mh/UpVSL8hoI9U6rYx8fFll7Z/wQRIvPqksB97xjise9F13WwvlTTVeqL7TNPha19+2Jfo8jEYOQsG7pK7L4gRUPPoXjEabYAzStwtidBnEvoiqiMRGwEn45y8igCzWV6+Z21yfu6KDOY0gqybhvU0iW+wKJupW1hrd7JdR7104= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807013; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=IieFYtDc0DxvX0y8JT1W+rF6uG/bZgRcpKzBTVDticE=; b=h0a72mVApJQFab+NEwGDpSbLYHUkfztfoXm9SIHLB/d5WzJmFsmgrth6LXEhxcRhukuY24TXovWaNfXV5KgWtnoz0Sa03nGebCU3/4XLaCHs2jbZxd0yZI4bEt0QojpuZ9SXZ1wRX6zkDnLYxlDv62n8O1h5W4uZwe/CWr4agcU= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807013347682.6232844819771; Wed, 26 Aug 2026 22:03:33 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGn-0005wh-Kl; Thu, 27 Aug 2026 01:03:21 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGj-0005tb-T7 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:17 -0400 Received: from mail-yw1-x1134.google.com ([2607:f8b0:4864:20::1134]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGg-0002Kp-Os for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:17 -0400 Received: by mail-yw1-x1134.google.com with SMTP id 00721157ae682-81ee6b2da98so26786967b3.3 for ; Wed, 26 Aug 2026 22:03:14 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b5bc331c5sm4272317b3.3.2026.08.26.22.03.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806993; x=1788411793; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=IieFYtDc0DxvX0y8JT1W+rF6uG/bZgRcpKzBTVDticE=; b=OuKOufnnA9nfCwp3o3CQG3hB+BxnRhqEx8fY3iiLPWZfTerf9lmW4ljMGCsL2+p/Gj 9wgKeU06yDk24cG3p+4itJlZw5yicseCczaHxETK+PYo0LUdKqJ+DNGcCYo9hwe6QgL0 6c/xtcOrpfo2icPbAqQqCR0LX2oEZSJZF0bmNqmiHj6+S2LvOsqAkE8W9fEdnerq0FZ2 VftDRQk80ZTyKQe2k+XrN+qsEunfG9YN86egHQm0pOLAJOp0RrMrA3N0LvAhgj6dualJ Zb0n8RWNJd7MbzqzCYAAtuQUe3b1JgJ3C9+sbD6SgXmbgQzAzDvldOfWCXPG/O1tBJef f5pg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806993; x=1788411793; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=IieFYtDc0DxvX0y8JT1W+rF6uG/bZgRcpKzBTVDticE=; b=Kd1kFWFaltLqpTVDwJhNXlXpSin2ELaT93N3mCYoCk/1VZP0leu5gvZgyYuNNeu0ed rkjkadeXmv0MoY/3+PbqyQhk8ebrY8UkKUUMDUlk84Y2dehVFKm/wq9rfo6n1c0yNSi7 wg45hFrGLiIv+v+P9VagUO7p44e6hycHQuk3yVKZD5HEaiYJsUUcFc0OGgQbU1JtFvzt IMKOMJUJn5gUYUCS0GW9hz4Qxz6B2gHPEAFaFG6MQYfolGqXf89th6qk8hEZsx2JxRsS lG8gZ2eFKS0jqFlgm5dHc9HYgwxeuKQdD6wg9qB5KupEzeeUkshAzi0ztgl0WqjG35tk OEfg== X-Gm-Message-State: AFuF++kTNzVAVomDIKWnJUlwp+X7NDyQ1onqSjeS2OWZO1hh070xCOM9 /WQ1G3l7EZP8hxMzPZ0TSd2+xDzJMDW4OE5+/pn/LLmSLtQM1P70r9bXRpXJzsmU X-Gm-Gg: AR+sD11uJ8VX1IGKf+HgBFdwRMJyW1zQEeIWZJc4SDo6IXR084asn/JAo06jl5SG93e E0FCLZ/mksAUsIBXYFNx9Yyt/6972TJKFf/Y0qF5brsU0WyhK3P4YzXAv34CnXgM+Q6PksChCsl DwG1V3nrJSLA+3mNJD9BZOJi9sRPpqFbn6ARI0C4A1wOzHRNsBMyjgS9syg5GEbBWfScEe7Unz7 5/ZiPdLq10zTD9rByzHx423tPtB/dpgzbzc7M7/Am5Mt6phEaDbbLrw5XHWEsod6nwkbt7YMCrM HQxWd7v3Ko/5k6HP4NEuRiuxSbe5b8EMF3uOtetAnpSd4RTFbGDclRpr2vYSVatOmeBvmwcJV9Q mN2Q1wGSOTVe8fnjEFrcuKiCjswx63fbtogeHbP3oMttjVzurTj279g+q31UKNW7VE7KFCrYKtw g0Ed1EpG9p+cMpfqr5I8/Ij+E7tWV6cb0oxNE24kuLorgVTIYqeb//aGDM2mwkmUg6E98fFjW7R j1j7hy0qSFScv6m5/jnlHOTUiDQ3iGVprYhtoo1 X-Received: by 2002:a05:690c:3388:b0:81e:89b2:91ae with SMTP id 00721157ae682-8573f8d462amr59125217b3.21.1787806993191; Wed, 26 Aug 2026 22:03:13 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 8/9] RFC: accel/tcg: poison the jump cache instead of polling for indirect exits Date: Thu, 27 Aug 2026 01:02:40 -0400 Message-ID: <20260827050241.3713332-9-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1134; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1134.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807013994158500 Content-Type: text/plain; charset="utf-8" Every translation block begins by loading cpu->neg.icount_decr.u32, testing it and branching to the exit path. That is three host instructions at the t= op of every TB, and blocks are short: an emulated alpha gcc 16.2.0 compiling a 255k line translation unit executes 34.2 billion of them at 6.04 guest instructions each. A block does not need to poll if every way out of it already reaches a chec= k. A goto_tb does not: it chains straight into its destination, with nothing in between that looks at icount_decr, so the destination has to poll on entry. An indirect exit does. The out-of-line path calls helper_lookup_tb_ptr() every time, so it only needs the helper to return the epilogue while an exit is pending. The inline probe is already covered by the machinery the breakpoint patch added: it reads its base pointer from cpu->tb_jmp_cache_probe and takes the slow path when the entry it finds has= a NULL tb, so pointing that base at a page of zeroes turns every indirect dispatch into a miss, and a miss lands in the same helper. So a pending exit becomes one more reason for tcg_cpu_may_dispatch() to say no. The two places that set icount_decr.u16.high poison the probe; the main loop puts it back once the flag is clear, on the same pass that already re-evaluates the breakpoint state. The real tb_jmp_cache is untouched throughout, so no cache contents are lost, and the fast path pays nothing: the base was a load from CPUState either way. The poll is therefore emitted only in blocks that emit a goto_tb. Whether a block does is not known until its last exit has been generated, so the decision is deferred and the load and branch are emitted retroactively at t= he head of the block in gen_tb_end(), using the same emit_before_op mechanism the can_do_io stores use. icount opts out and keeps the counter unconditionally. Interrupt latency is bounded at one block, as before. It does not depend on the shape of the guest's control flow graph: a block either polls on entry = or is checked on the way out, and no run of blocks can avoid both. What changes is where the check sits, not how often one happens. tests/tcg/multiarch/test-indirect-irq.c is added for this: a loop whose only back edge is an indirect branch, under alarm(1). That loop's block emits no goto_tb, so it no longer polls, and the test passes only because the dispat= ch notices instead -- it hangs if the poison is removed, which is what makes i= t a test of the new mechanism rather than of the old poll. Nothing in it is architecture specific: the loop is a computed goto, which every target's compiler supports, so it covers whichever targets go on to use the inline probe. The other alpha tests still pass and the emulated compiler still produces byte-identical output. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, = on top of the preceding patches: before: 891,254,240,071 instructions after: 868,811,832,620 instructions -2.52% before: 81.45s wall clock after: 79.85s wall clock -1.96% The emulated compiler produces byte-identical output. RFC because: - The un-poison in the main loop races a concurrent poison from another thread. The existing barrier around icount_decr.u16.high covers it -- a poison that lands after the sync also re-set the flag, and exit_request w= as stored before it -- but this deserves more eyes than the single-threaded user-mode testing I have given it. - Only the inline probe needs the poison, and only alpha uses the inline probe today. Targets on the out-of-line path are covered by the helper check alone, but that has not been measured. - The shared zero-filled CPUJumpCache is a 1MB allocation that is never written. A read-only mapping would express that better. v3: Rebased onto the removal of "only poll for interrupts in blocks that can close a cycle", which v2 sat on top of and which is dropped: it let a straight-line run of arbitrary length go unchecked, since a block with = no backward edge polled nowhere (Richard). The rule is now that a block polls iff it emits a goto_tb, rather than iff it can close a control flow cycle. That keeps the bound at one block without any analysis of the guest's control flow graph, so the objection to the dropped patch does not carry over. The deferred-emission machine= ry it needs moves here from that patch; DisasContextBase::needs_exit_check and the hook in translator_use_goto_tb() are gone with it, and the flag is now set by tcg_gen_goto_tb() rather than by goto_ptr emission. All of v2's measurements were dropped: they were taken with the cycle-analysis patch underneath, which changes both the baseline and what is left to remove, so none of them described this patch. The numbers above are a fresh measurement of the series as it now stands. v4: Moved the test from tests/tcg/alpha/ to tests/tcg/multiarch/: the mechanism is generic and nothing in the test is alpha specific (Alex). The performance numbers above are the v3 measurements, not re-run: the machine they were taken on is busy. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec.c | 23 +++++++-- accel/tcg/tcg-accel-ops.c | 1 + accel/tcg/translator.c | 51 ++++++++++++++++++-- include/hw/core/cpu.h | 8 ++-- include/tcg/tcg.h | 2 + tcg/tcg-op.c | 13 ++++++ tests/tcg/multiarch/test-indirect-irq.c | 62 +++++++++++++++++++++++++ 7 files changed, 151 insertions(+), 9 deletions(-) create mode 100644 tests/tcg/multiarch/test-indirect-irq.c diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index c2a9679cd7..9ca5894108 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -388,6 +388,16 @@ const void *HELPER(lookup_tb_ptr)(CPUArchState *env) */ cpu->neg.can_do_io =3D true; =20 + /* + * A block that dispatches indirectly does not emit the icount_decr po= ll, + * so this is where a pending exit is noticed for that path: either the + * probe was poisoned and every dispatch arrives here, or the target u= ses + * the out-of-line lookup and always did. + */ + if (unlikely(cpu_loop_exit_requested(cpu))) { + return tcg_code_gen_epilogue; + } + TCGTBCPUState s =3D cpu->cc->tcg_ops->get_tb_cpu_state(cpu); s.cflags =3D curr_cflags(cpu); =20 @@ -757,8 +767,9 @@ static inline bool cpu_handle_exception(CPUState *cpu, = int *ret) * slow path when the entry it finds has a NULL tb. Pointing the probe at= a * region that is all zeroes therefore forces every indirect dispatch into * helper_lookup_tb_ptr(), which does the full lookup the inline probe only - * approximates. The real jump cache is untouched, so no contents are lost - * and recovery is a single store. + * approximates and returns to the main loop while an exit is pending. The + * real jump cache is untouched, so no contents are lost and recovery is a + * single store. * * Only ever read from, and only one entry per dispatch, so one shared * zero-filled cache is enough for every CPU. Not const: that would put a @@ -779,10 +790,13 @@ static CPUJumpCache tb_jmp_cache_poison; * the rest of the page. A block translated before the breakpoint was set= is * therefore still in the jump cache, and dispatching to it inline would s= tep * straight over the breakpoint. + * + * A block that dispatches indirectly also does not emit the icount_decr + * poll, so the dispatch is where a pending exit has to be noticed. */ static bool tcg_cpu_may_dispatch(CPUState *cpu) { - return QTAILQ_EMPTY(&cpu->breakpoints); + return QTAILQ_EMPTY(&cpu->breakpoints) && !cpu_loop_exit_requested(cpu= ); } =20 /* @@ -851,6 +865,9 @@ void tcg_kick_vcpu_thread(CPUState *cpu) =20 /* Ensure cpu_exec will see the exit request after TCG has exited. */ qatomic_store_release(&cpu->neg.icount_decr.u16.high, -1); + + /* Blocks that only dispatch indirectly do not poll; stop them chainin= g. */ + tcg_cpu_poison_jmp_cache(cpu); } =20 static inline bool icount_exit_request(CPUState *cpu) diff --git ./accel/tcg/tcg-accel-ops.c ./accel/tcg/tcg-accel-ops.c index 560fe2554b..9eb9e861ac 100644 --- ./accel/tcg/tcg-accel-ops.c +++ ./accel/tcg/tcg-accel-ops.c @@ -106,6 +106,7 @@ void tcg_handle_interrupt(CPUState *cpu, int mask) qemu_cpu_kick(cpu); } else { qatomic_set(&cpu->neg.icount_decr.u16.high, -1); + tcg_cpu_poison_jmp_cache(cpu); } } =20 diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 8879cd626f..89d255bd04 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -45,12 +45,35 @@ bool translator_io_start(DisasContextBase *db) return true; } =20 +/* + * A block that ends in a goto_tb chains straight to its destination: noth= ing + * between the two looks at icount_decr, so the destination has to poll on + * entry. A block whose exits are all indirect does not, because the disp= atch + * itself notices -- a pending exit poisons tb_jmp_cache_probe, so the pro= be + * misses into helper_lookup_tb_ptr(), which returns the epilogue. Every = block + * therefore either polls on entry or is checked as it leaves, which bounds + * interrupt latency at one block without looking at the shape of the gues= t's + * control flow graph. + * + * Which kind a block is is not known until its last exit has been emitted= , so + * defer the decision to gen_tb_end() and emit the poll retroactively. + * + * icount needs the counter unconditionally, so it opts out. + */ +static bool defer_exit_check(uint32_t cflags) +{ + return !(cflags & CF_USE_ICOUNT); +} + static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t cflags) { TCGv_i32 count =3D NULL; TCGOp *icount_start_insn =3D NULL; =20 - if ((cflags & CF_USE_ICOUNT) || !(cflags & CF_NOIRQ)) { + tcg_ctx->exit_check_needed =3D false; + + if ((cflags & CF_USE_ICOUNT) || + (!(cflags & CF_NOIRQ) && !defer_exit_check(cflags))) { count =3D tcg_temp_new_i32(); tcg_gen_ld_i32(count, tcg_env, offsetof(CPUState, neg.icount_decr.u32) - @@ -76,6 +99,9 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t= cflags) */ if (cflags & CF_NOIRQ) { tcg_ctx->exitreq_label =3D NULL; + } else if (defer_exit_check(cflags)) { + /* Emitted retroactively by gen_tb_end(), if this TB emits a goto_= tb. */ + tcg_ctx->exitreq_label =3D gen_new_label(); } else { tcg_ctx->exitreq_label =3D gen_new_label(); tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_label); @@ -91,7 +117,8 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_= t cflags) } =20 static void gen_tb_end(const TranslationBlock *tb, uint32_t cflags, - TCGOp *icount_start_insn, int num_insns) + TCGOp *icount_start_insn, int num_insns, + TCGOp *first_insn_start) { if (cflags & CF_USE_ICOUNT) { /* @@ -102,6 +129,23 @@ static void gen_tb_end(const TranslationBlock *tb, uin= t32_t cflags, tcgv_i32_arg(tcg_constant_i32(num_insns))); } =20 + if (tcg_ctx->exitreq_label && defer_exit_check(cflags) && + !(cflags & CF_NOIRQ)) { + if (tcg_ctx->exit_check_needed) { + TCGv_i32 count =3D tcg_temp_new_i32(); + TCGOp *save =3D tcg_ctx->emit_before_op; + + tcg_ctx->emit_before_op =3D first_insn_start; + tcg_gen_ld_i32(count, tcg_env, + offsetof(CPUState, neg.icount_decr.u32) - + sizeof(CPUState)); + tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_la= bel); + tcg_ctx->emit_before_op =3D save; + } else { + tcg_ctx->exitreq_label =3D NULL; + } + } + if (tcg_ctx->exitreq_label) { gen_set_label(tcg_ctx->exitreq_label); tcg_gen_exit_tb(tb, TB_EXIT_REQUESTED); @@ -238,7 +282,8 @@ void translator_loop(CPUState *cpu, TranslationBlock *t= b, int *max_insns, =20 /* Emit code to exit the TB, as indicated by db->is_jmp. */ ops->tb_stop(db, cpu); - gen_tb_end(tb, cflags, icount_start_insn, db->num_insns); + gen_tb_end(tb, cflags, icount_start_insn, db->num_insns, + first_insn_start); =20 /* * Manage can_do_io for the translation block: set to false before diff --git ./include/hw/core/cpu.h ./include/hw/core/cpu.h index bd2cdd2a0b..4272740303 100644 --- ./include/hw/core/cpu.h +++ ./include/hw/core/cpu.h @@ -523,9 +523,11 @@ struct CPUState { * @tb_jmp_cache_probe: base the inline jump cache probe reads. * * Normally @tb_jmp_cache. Pointed at a shared page of zeroes to force - * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (); - * see tcg_cpu_sync_jmp_cache(). NULL before tcg_exec_realizefn() and - * after tcg_exec_unrealizefn(). + * every inline dispatch to miss and fall back to helper_lookup_tb_ptr= (), + * either because a breakpoint is set or because an exit is pending; s= ee + * tcg_cpu_sync_jmp_cache(). Only generated code and the accessors in + * cpu-exec.c may touch it. NULL before tcg_exec_realizefn() and after + * tcg_exec_unrealizefn(). */ struct CPUJumpCache *tb_jmp_cache_probe; =20 diff --git ./include/tcg/tcg.h ./include/tcg/tcg.h index 7669dc1c2d..df08c10544 100644 --- ./include/tcg/tcg.h +++ ./include/tcg/tcg.h @@ -389,6 +389,8 @@ struct TCGContext { struct TCGLabelPoolData *pool_labels; =20 TCGLabel *exitreq_label; + /* Set by goto_tb emission: this TB chains without reaching a check. */ + bool exit_check_needed; =20 #ifdef CONFIG_PLUGIN /* diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index cf7b6882d8..d384325a2e 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -2713,6 +2713,13 @@ void tcg_gen_goto_tb(unsigned idx) tcg_debug_assert((tcg_ctx->goto_tb_issue_mask & (1 << idx)) =3D=3D 0); tcg_ctx->goto_tb_issue_mask |=3D 1 << idx; #endif + /* + * A goto_tb chains straight into the destination, with nothing in bet= ween + * that looks at icount_decr, so this TB has to poll on entry. See + * defer_exit_check(). + */ + tcg_ctx->exit_check_needed =3D true; + plugin_gen_disable_mem_helpers(); tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } @@ -2834,6 +2841,12 @@ void tcg_gen_lookup_and_goto_ptr_tmp(TCGTemp *pc, co= nst TranslationBlock *tb) =20 plugin_gen_disable_mem_helpers(); =20 + /* + * Neither path below needs an icount_decr poll. The helper returns to + * the main loop while an exit is pending, and a pending exit poisons + * tb_jmp_cache_probe, so the inline probe finds a NULL tb and falls i= nto + * that same helper. + */ if (pc) { TCGv_i64 pc64; =20 diff --git ./tests/tcg/multiarch/test-indirect-irq.c ./tests/tcg/multiarch/= test-indirect-irq.c new file mode 100644 index 0000000000..a672faf641 --- /dev/null +++ ./tests/tcg/multiarch/test-indirect-irq.c @@ -0,0 +1,62 @@ +/* + * A loop whose only back edge is an indirect branch must still be + * interruptible. + * + * Blocks that dispatch indirectly do not emit the icount_decr poll; a pen= ding + * exit instead poisons the inline jump cache probe so that the dispatch f= alls + * into helper_lookup_tb_ptr(), which returns to the main loop. If that + * mechanism breaks, this program never leaves the loop and the test times + * out rather than failing an assertion. + * + * A computed goto is used deliberately: a plain while(1) would end the bl= ock + * with a direct backward branch, that is a goto_tb, and a block that emit= s a + * goto_tb still polls -- so it would not exercise the path under test. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include +#include + +/* Written by the handler, read by the loop, so it must not be cached. */ +static volatile sig_atomic_t fired; +/* Read after the loop, so the loop must not optimize the increment away. = */ +static volatile unsigned long iterations; + +static void handler(int sig) +{ + fired =3D 1; +} + +int main(void) +{ + /* + * Indexing a table with a volatile index, rather than jumping through= a + * volatile pointer: gcc happily proves a single-valued pointer consta= nt + * and emits a direct branch, which is the case this test is not about. + */ + volatile int idx =3D 0; + void *target[2]; + struct sigaction sa; + + memset(&sa, 0, sizeof(sa)); + sa.sa_handler =3D handler; + sigemptyset(&sa.sa_mask); + assert(sigaction(SIGALRM, &sa, NULL) =3D=3D 0); + alarm(1); + + target[0] =3D &&spin; + target[1] =3D &&out; +spin: + iterations++; + if (!fired) { + goto *target[idx]; + } +out: + + printf("interrupted after %lu iterations\n", iterations); + return 0; +} --=20 2.54.0 From nobody Mon Sep 28 00:53:22 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1787807028; cv=none; d=zohomail.com; s=zohoarc; b=Go/ohwEquybcjJSK4MM9KdbuIKQXYatp5zp5XGCp6SM/fpghAqm04JEhUBE8oXYwPoHS5ljCaKxrOR/w6n96WhiFdhBbsgxyZESEj9vwma21nw8A0JnEjSkYX/argx1Vp+HewZ56gf0fgmUgOqIFVUMR10SFDIs9SXjRXgCU8h8= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1787807028; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=GhFbOi2D7TSDX/5NYHihYJQXa2xvjlerNJOOPcgiuis0Fvsm3eSpztiBzcdayLrkBRFbXTkwRwCVc6bVhqE7LduYl31AsyGpWfmZE+PLwxIH9Dv1zT1zZh85fk8GvRQhz8G8oO9c1kbynPaMZ129Pa2WqfTaDqpGUUPe6vVjg7k= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1787807028186512.7501214016534; Wed, 26 Aug 2026 22:03:48 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wzSGp-0005y2-Og; Thu, 27 Aug 2026 01:03:23 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wzSGm-0005uR-32 for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:20 -0400 Received: from mail-yw1-x112f.google.com ([2607:f8b0:4864:20::112f]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wzSGj-0002LA-Lj for qemu-devel@nongnu.org; Thu, 27 Aug 2026 01:03:19 -0400 Received: by mail-yw1-x112f.google.com with SMTP id 00721157ae682-81ee6b2da98so26787437b3.3 for ; Wed, 26 Aug 2026 22:03:16 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85b5c7f3d67sm4296127b3.15.2026.08.26.22.03.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 22:03:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787806996; x=1788411796; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=jpGMC0EUURyGoCiFMsT59gJZt47tbTVj6eAN/+P/tI4REmmn2Sps+8pLEam0i83p6n bdnYrtdXjWmbY7A+LZUrz4omMqRBQkEnySzS/tfNRopEt/NowUuY7Za+ADNk/KvS3bne LnahPT8BTbY7g+MUflPLlT1NZSOSS9+e3TkMesA/ahYPccCsxgqIjRFA9VsXjoPHBVrR JPAoSpNoV69nnS7mzGO+lLq1ewM3gJAcB0U5wDUMkyCPaf61j0wMq2WB3b8bdW9PiTyH 0+oHUwFT2jJJK9Z87QIJqSLsMScbfg1yy/Z/k49WO5IrtIgtegRnAElnQopKvIIE66Gu VwPw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787806996; x=1788411796; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=dMtEHQYXLnsxqhMGg/2vSbb0NuESXDPunNyoGfi0LOp3esc9G2k6BfCOQ/vUpCyqDS fTisJr1xdkCGszjLoVL+C+Nq2uPt2FBec3BK47kxyoD/2rhuicbTWASAh1ardQX2/uoM wIue4c8K4TujYNMTruzyWvTPN+oWDHFU1DRo4qXpXptcKMxBLwjstWxJF8mVnhkPthm1 7IjTyN/dJe/fiLPLK0KsqoamRvY/LVgT55qYhsmgmu3gA3AW2zPHEmpeJB3EvhK8/uyI OPD3zI1g38TEBunitSxqjJBFyfngez1dA69lX22n1/28JiubcExuVegrGCG19osFNyH9 8ZNg== X-Gm-Message-State: AFuF++mFgwQaNsSePCkDIUx0cA7lEQpuRc+AAsL5+lhb+a1zMaXaeZyw mtzcWpM0l3GNGgLT4Y+IGk3tVvBHCT6o1EyIDHHQek1WQUL3qjyB/YtB110JfZjT X-Gm-Gg: AR+sD108scyb6Xp7VmehSa/J83KPJX7aQ+qTqO+ImeDDCfseCEcpAQ5Zv5irJhgi8GX oKRIoYEKClS7AQu+0lyHyJXr3HZrSAQVxFInWe282HEiBYOWbpVP3sGW5ixhRaY8ZpgnH3nF1or Vn7MmpRz6+k5tcz2Cg/POD/ynk3Soc+/MpZMAUvydTNaxrd3UrlWSGNtmhWgmm/HJafcQN2uMLk iCnU6AdY227qzOKY0H3PHiuysdM5Bl/llp0r24h9G7Jll2/3bSCWeb/bjCgziNtXsquOLq8klYR +9ezBdmJb7mk3DfKk3I1gHAOVNqglI1wWTIuUo9cbxAxlbosshKl6imACZS8I9picT9D7ULpoJL QKB/xwObcpi4sGpOKk6Y047DhaLBXyO9QfEY95OcccpVp91k5inRsRh8M4xTAjyWDmNcMep0Mqw l0SxlLJGpE/HDg0LYzQv7aerI9ewfWhUyJJEjg4uS9fCLiIzFnPqfCIhYup8Jil1wT8/PATF93e tZg/qu5dMzKFSMdQpWdmdWjm3JSQUarUw0vwo9KDw== X-Received: by 2002:a05:690c:60c5:b0:854:5db3:4ed0 with SMTP id 00721157ae682-8574153a296mr54240007b3.31.1787806995751; Wed, 26 Aug 2026 22:03:15 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v4 9/9] RFC: tcg: fold a guest displacement into the host addressing mode Date: Thu, 27 Aug 2026 01:02:41 -0400 Message-ID: <20260827050241.3713332-10-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260822190818.1829249-1-mattst88@gmail.com> References: <20260822190818.1829249-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112f; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112f.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1787807030021158500 Content-Type: text/plain; charset="utf-8" Nothing in the TCG frontend interface can express a based memory access. tcg_gen_qemu_ld/st take an address and nothing else, so a target with a displacement in its load and store encodings -- which is most of them -- has to materialize the address first: ldq a1,8(a0) -> mov 0x80(%rbp),%rbx reload a0 lea 0x8(%rbx),%r12 address mov (%r12),%r12 the load mov %r12,0x88(%rbp) spill a1 The lea is pure loss on a host whose addressing mode has a displacement field sitting empty. It also needs a register, at the point in a block where pressure is highest. Fold it. After optimization, look for an add of a constant immediately before a guest access, defining that access's address operand, and move the constant into a new second constant argument on the op. The add is left for liveness to remove, so nothing breaks if its result has another use. Only the immediately preceding op is examined: that is what the frontends emit, and a window of one op means the pass does not have to reason about what could have happened in between. The one thing it does check is that the add did not clobber the base it read, since the access now reads that base directly. Targets opt in with TCG_TARGET_HAS_ldst_disp and an out_disp member on TCGOutOpQemuLdSt. Without it the pass does not run, the displacement stays zero and the existing out member is called exactly as before, so no other backend changes behavior or needs touching. The fold is refused unless the access has no slow path at all, since the slow path hands addr_reg to the helper and that register no longer holds the full guest address. That is decided generically: user-only, because softmmu compares the unadjusted address against the TLB; a 64-bit address type, because a 32-bit one wraps where a host displacement would not; and no alignment test on the access. For x86_64 the displacement goes in the disp32 that prepare_host_addr() already fills in for guest_base, so all the backend has left to check is that guest_base plus the displacement still fits there. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding patches, against a control measured in the same session: before: 868,811,832,620 instructions, 79.85s after: 819,262,147,022 instructions, 77.30s -5.70% instructions, -3.20% wall Emitted code shrinks from 50.55MB to 48.80MB over the run, 167.4 to 161.6 bytes per block. Per Alpha opcode, the host bytes emitted for an access fall as expected and nothing else moves: ldq 18.3 -> 15.4 ldah 20.9 -> 20.9 ldl 16.6 -> 14.1 lda 12.9 -> 12.9 stq 12.8 -> 9.7 mov 9.8 -> 9.8 The emulated compiler produces byte-identical output and the alpha tests still pass, including with a non-zero guest_base forced via -B. RFC because: - Only wired up for x86_64, and only for qemu_ld and qemu_st; the i128 qemu_ld2 and qemu_st2 pairs are left alone. - Requiring that no slow path exists is stricter than necessary. The fast path test can stay on the base register as long as the displacement is itself a multiple of the required alignment, which it is for anything a frontend emits for a struct or stack access. Recording the displacement in TCGLabelQemuLdst and emitting one lea on the slow path would then cover alignment-checked accesses too, at no fast path cost. - Softmmu wants the displacement folded into the TLB comparison as well, which is a bigger change than this one. - A one op window catches everything the frontends emit today but is trivially defeated by anything scheduled in between. v4: - Hoisted the compilation mode tests -- tcg_use_softmmu and the 64-bit address type -- out of the backend hook and into fold_ldst_disp(), next to the TCG_TARGET_HAS_ldst_disp test, so the loop is not entered at all when the mode rules the fold out. - Pass MemOp rather than MemOpIdx to the backend hook; nothing about the mmu_idx is relevant to it. - Moved the alignment test into generic code as ldst_disp_needs_align(), so a backend does not have to repeat the atom_and_align_for_opc() call. The exact answer depends on the host's atomicity capabilities, which the generic pass does not know, so it answers for the most restrictive host. That is the same answer for everything the frontends actually emit -- MO_ATOM_IFALIGN is the default -- and conservative for the handful of MO_ATOM_WITHIN16 and MO_ATOM_SUBALIGN accesses, which lose the fold on a host that could have taken it. - What is left of the x86_64 hook is the guest_base test, so it now lives beside x86_guest_base under the CONFIG_USER_ONLY that declares it. - Refuse a displacement that does not fit in an int32_t, which is what out_disp() takes. Not reachable with any real guest_base, but the pass should not offer the backend something the interface cannot carry. - The numbers above are unchanged from v3: they have not been re-measured on the restructured patch, which is not expected to move them since the accesses in this workload are all MO_ATOM_IFALIGN. Signed-off-by: Matt Turner --- include/tcg/tcg-opc.h | 9 ++- tcg/tcg-op-ldst.c | 3 +- tcg/tcg.c | 132 +++++++++++++++++++++++++++++++++++- tcg/x86_64/tcg-target.c.inc | 41 +++++++++++ tcg/x86_64/tcg-target.h | 3 + 5 files changed, 184 insertions(+), 4 deletions(-) diff --git ./include/tcg/tcg-opc.h ./include/tcg/tcg-opc.h index f3a81d5d7f..92fd34d3e3 100644 --- ./include/tcg/tcg-opc.h +++ ./include/tcg/tcg-opc.h @@ -125,8 +125,13 @@ DEF(goto_ptr, 0, 1, 0, TCG_OPF_BB_EXIT | TCG_OPF_BB_EN= D) DEF(plugin_cb, 0, 0, 1, TCG_OPF_NOT_PRESENT) DEF(plugin_mem_cb, 0, 1, 1, TCG_OPF_NOT_PRESENT) =20 -DEF(qemu_ld, 1, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) -DEF(qemu_st, 0, 2, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +/* + * The second constant argument is a displacement to add to the address, + * zero unless a target advertises TCG_TARGET_HAS_ldst_disp and the fold in + * fold_ldst_disp() applied. + */ +DEF(qemu_ld, 1, 1, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +DEF(qemu_st, 0, 2, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) DEF(qemu_ld2, 2, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) DEF(qemu_st2, 0, 3, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) =20 diff --git ./tcg/tcg-op-ldst.c ./tcg/tcg-op-ldst.c index 22211ccb45..ffc5e651a6 100644 --- ./tcg/tcg-op-ldst.c +++ ./tcg/tcg-op-ldst.c @@ -92,7 +92,8 @@ static MemOp tcg_canonicalize_memop(MemOp op, bool is64, = bool st) static void gen_ldst1(TCGOpcode opc, TCGType type, TCGTemp *v, TCGTemp *addr, MemOpIdx oi) { - TCGOp *op =3D tcg_gen_op3(opc, type, temp_arg(v), temp_arg(addr), oi); + /* The trailing zero is the address displacement; see fold_ldst_disp()= . */ + TCGOp *op =3D tcg_gen_op4(opc, type, temp_arg(v), temp_arg(addr), oi, = 0); TCGOP_FLAGS(op) =3D get_memop(oi) & MO_SIZE; } =20 diff --git ./tcg/tcg.c ./tcg/tcg.c index 489df0e738..466604eb97 100644 --- ./tcg/tcg.c +++ ./tcg/tcg.c @@ -1058,6 +1058,13 @@ typedef struct TCGOutOpQemuLdSt { TCGOutOp base; void (*out)(TCGContext *s, TCGType type, TCGReg dest, TCGReg addr, MemOpIdx oi); + /* + * As out(), for an access at addr + disp. Only required of targets th= at + * define TCG_TARGET_HAS_ldst_disp; for everyone else fold_ldst_disp() + * never runs and the displacement is always zero. + */ + void (*out_disp)(TCGContext *s, TCGType type, TCGReg dest, + TCGReg addr, MemOpIdx oi, int32_t disp); } TCGOutOpQemuLdSt; =20 typedef struct TCGOutOpQemuLdSt2 { @@ -3574,6 +3581,123 @@ static void move_label_uses(TCGLabel *to, TCGLabel = *from) QSIMPLEQ_CONCAT(&to->branches, &from->branches); } =20 +#ifndef TCG_TARGET_HAS_ldst_disp +#define TCG_TARGET_HAS_ldst_disp 0 +#define tcg_target_ldst_disp_ok(s, opc, disp) false +#endif + +/* + * Return true if @opc needs an alignment test in the fast path. + * + * atom_and_align_for_opc() gives the exact answer, but only once the host= 's + * atomicity capabilities are known, and those belong to the backend. Answ= er + * instead for the most restrictive host, which is valid for all of them. + */ +static bool ldst_disp_needs_align(MemOp opc) +{ + MemOp size =3D opc & MO_SIZE; + + if (memop_alignment_bits(opc)) { + return true; + } + switch (opc & MO_ATOM_MASK) { + case MO_ATOM_NONE: + case MO_ATOM_IFALIGN: + case MO_ATOM_IFALIGN_PAIR: + return false; + case MO_ATOM_WITHIN16: + /* Misalignment implies !within16, and therefore no atomicity. */ + return size !=3D MO_128; + case MO_ATOM_WITHIN16_PAIR: + case MO_ATOM_SUBALIGN: + return size !=3D MO_8; + default: + g_assert_not_reached(); + } +} + +/* + * Fold "add addr, base, $disp" into the guest access that follows it, so + * that the displacement becomes part of the host addressing mode instead = of + * a separate instruction. Frontends have no way to express this: there is + * no displacement operand on tcg_gen_qemu_ld/st, so a based access always + * costs an extra add, and an extra register to hold its result. + * + * Only an add in the op immediately before the access is recognized. That + * is what the frontends emit, and a window of one op means no analysis is + * needed of what might have happened in between. The add is left in place; + * liveness removes it if its result has no other use. + */ +static void __attribute__((noinline)) +fold_ldst_disp(TCGContext *s) +{ + TCGOp *op; + + /* + * The fold requires that the access have no slow path, because the sl= ow + * path hands the address operand to the helper and that register no + * longer holds the complete guest address. That means user-only, since + * softmmu compares the unadjusted address against the TLB. It also + * requires a 64-bit address type: for a 32-bit one the add wraps and a + * host displacement would not. + */ + if (!TCG_TARGET_HAS_ldst_disp || tcg_use_softmmu || + s->addr_type !=3D TCG_TYPE_I64) { + return; + } + + QTAILQ_FOREACH(op, &s->ops, link) { + TCGOp *prev; + TCGTemp *cts; + int64_t disp; + MemOp opc; + + switch (op->opc) { + case INDEX_op_qemu_ld: + case INDEX_op_qemu_st: + break; + default: + continue; + } + + opc =3D get_memop(op->args[2]); + if (ldst_disp_needs_align(opc)) { + continue; + } + + prev =3D QTAILQ_PREV(op, link); + if (prev =3D=3D NULL || prev->opc !=3D INDEX_op_add || + TCGOP_TYPE(prev) !=3D s->addr_type) { + continue; + } + + /* + * The add must define the address operand, and must not have + * clobbered the base it read: after the fold the access reads the + * base directly, so the base has to still hold its original value. + */ + if (prev->args[0] !=3D op->args[1] || prev->args[0] =3D=3D prev->a= rgs[1]) { + continue; + } + + cts =3D arg_temp(prev->args[2]); + if (cts->kind !=3D TEMP_CONST) { + continue; + } + /* out_disp() takes an int32_t, so anything wider cannot be passed= . */ + disp =3D cts->val; + if (disp !=3D (int32_t)disp) { + continue; + } + if (disp =3D=3D 0 || !tcg_target_ldst_disp_ok(s, opc, disp)) { + continue; + } + + op->args[1] =3D prev->args[1]; + op->args[3] =3D disp; + } +} + /* Reachable analysis : remove unreachable code. */ static void __attribute__((noinline)) reachable_code_pass(TCGContext *s) @@ -5728,7 +5852,12 @@ static void tcg_reg_alloc_op(TCGContext *s, const TC= GOp *op) const TCGOutOpQemuLdSt *out =3D container_of(all_outop[op->opc], TCGOutOpQemuLdSt, base); =20 - out->out(s, type, new_args[0], new_args[1], new_args[2]); + if (new_args[3]) { + out->out_disp(s, type, new_args[0], new_args[1], + new_args[2], new_args[3]); + } else { + out->out(s, type, new_args[0], new_args[1], new_args[2]); + } } break; =20 @@ -6611,6 +6740,7 @@ int tcg_gen_code(TCGContext *s, TranslationBlock *tb,= uint64_t pc_start) tcg_temp_ebb_reset_freed(s); =20 tcg_optimize(s); + fold_ldst_disp(s); =20 reachable_code_pass(s); liveness_pass_0(s); diff --git ./tcg/x86_64/tcg-target.c.inc ./tcg/x86_64/tcg-target.c.inc index 2c8f1f3e58..9b177d3475 100644 --- ./tcg/x86_64/tcg-target.c.inc +++ ./tcg/x86_64/tcg-target.c.inc @@ -1892,6 +1892,18 @@ static HostAddress x86_guest_base =3D { .index =3D -1 }; =20 +/* + * Whether the displacement of a guest access can be folded into the host + * addressing mode rather than materialized by a separate lea. The generic + * pass has already established that the access has no slow path, so all + * that is left is guest_base, which shares the disp32 field. + */ +static bool tcg_target_ldst_disp_ok(TCGContext *s, MemOp opc, int32_t disp) +{ + int64_t ofs =3D (int64_t)x86_guest_base.ofs + disp; + return ofs =3D=3D (int32_t)ofs; +} + #if defined(__linux__) # include # include @@ -1917,6 +1929,7 @@ static inline int setup_guest_base_seg(void) #endif #else # define x86_guest_base (*(HostAddress *)({ qemu_build_not_reached(); NULL= ; })) +# define tcg_target_ldst_disp_ok(s, opc, disp) false #endif /* CONFIG_USER_ONLY */ #ifndef setup_guest_base_seg # define setup_guest_base_seg() 0 @@ -2183,9 +2196,23 @@ static void tgen_qemu_ld(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_ld_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, true); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_ld_direct(s, data, -1, h, type, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_ld =3D { .base.static_constraint =3D C_O1_I1(r, L), .out =3D tgen_qemu_ld, + .out_disp =3D tgen_qemu_ld_disp, }; =20 static void tgen_qemu_ld2(TCGContext *s, TCGType type, TCGReg datalo, @@ -2321,9 +2348,23 @@ static void tgen_qemu_st(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_st_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, false); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_st_direct(s, data, -1, h, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_st =3D { .base.static_constraint =3D C_O0_I2(L, L), .out =3D tgen_qemu_st, + .out_disp =3D tgen_qemu_st_disp, }; =20 static void tgen_qemu_st2(TCGContext *s, TCGType type, TCGReg datalo, diff --git ./tcg/x86_64/tcg-target.h ./tcg/x86_64/tcg-target.h index 7ebae56a7d..8f2315c15e 100644 --- ./tcg/x86_64/tcg-target.h +++ ./tcg/x86_64/tcg-target.h @@ -30,6 +30,9 @@ #define TCG_TARGET_NB_REGS 32 #define MAX_CODE_GEN_BUFFER_SIZE (2 * GiB) =20 +/* A guest displacement can go in the disp32 of the addressing mode. */ +#define TCG_TARGET_HAS_ldst_disp 1 + typedef enum { TCG_REG_EAX =3D 0, TCG_REG_ECX, --=20 2.54.0