From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234650; cv=none; d=zohomail.com; s=zohoarc; b=mR1dMAE02Ynt/WMPMnd/PnbPxJsgUY5Ty1JacEr9/2lUMumX+DP66YQADAUJnMtSwsGHIhWlD2UG/gzwZ9sEVIZnkN9fI7ti9fMFBWntvVVJ5MZXq5soDmPJNFCNGFfD67uZvfMXLCV1ONNd+t4BCFvE/FDL1uztQT2nZoweHnc= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234650; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=w7dxb1OJU0LPOmKMjovKdzBr/gV3Lga/pfQME9xMvpA=; b=m4DwMCSHxabXb+cPkaFYfULjiksMPwyIraR00Zb01DX1eBNJzk0vvuw07pWV+XFDpX6Q4yXVDIQk0dMBHIPUSkDWIYnoKlV1Cd/W5imbf7Qe2lZeOuvRq0B4j3c455yDLWUJcCqnti/MC0ACpB4OY6sm2xCVzjuzwoXp5k4ch1k= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234650419340.5352923475572; Mon, 31 Aug 2026 20:50:50 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUR-0000aS-87; Mon, 31 Aug 2026 23:48:51 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUP-0000Zv-1m for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:49 -0400 Received: from mail-yw1-x112b.google.com ([2607:f8b0:4864:20::112b]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUM-0004Kt-8V for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:48 -0400 Received: by mail-yw1-x112b.google.com with SMTP id 00721157ae682-857ff9fef54so5066407b3.1 for ; Mon, 31 Aug 2026 20:48:45 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-86326649752sm41728887b3.3.2026.08.31.20.48.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234525; x=1788839325; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=w7dxb1OJU0LPOmKMjovKdzBr/gV3Lga/pfQME9xMvpA=; b=So9zQronuqkg9pZrgHRcfUK6jy0VW8/Au5oMH/lN8YQ/0JMtgHu/zHJy4p7nEWgthP GrHUGNGnudZ7Mx7BlF7p9d6NuoODdzkHGfLLIdIWr0CpDFLWUIWUmMtuDvEq04gPerPs N0kPA31ZGIzqYIxN6krCQUOgb5et4mwbdKSD7qhCvNRXWakAO5FT7yJhcMtFHsmXtwA7 VAEnTiQeRqqVQ7BHIxfURnv7PlC2eZPbX6uh1GaGcPvH/8JyRfw8qLOVO69WZFwssIJ3 v/J2nTaemEszjgmR13AdGIVkFJF0IfiRWehFZPZT4CifzBRizFVTRbUYkH9He/oxGgX/ +Obg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234525; x=1788839325; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=w7dxb1OJU0LPOmKMjovKdzBr/gV3Lga/pfQME9xMvpA=; b=lH9+MogKZ3x/IjELSG/rZl5o2Mt9+LjKW8X0UOPo4AZQ1yVW9sRVKpLWc3DSAs07rp 4zXPtYAUl2OSxxZLslxfGfWeqXP+XWuZMTwohfFO/fEsEAfHpLUGAaVvXcY7rLvhycJy 96mWe7Zy4Qs1ZSC4ByGenVANAUfo4p6O8rxVsNDpHeyBSHsR3S/sTZbA+d9BgK3DQFkf bx33ht7iq3kGlV3A4ZcQXO5dqOVgVQwzhBVINPyi0Kk+SvtU4ck7jIFckv+n77X1Wq3E CZzLC+3ueI4nPh7pseEymZWf7YChyvBi8/0mICgZSQdOy/tc42dZEr9E+wJiqQKh7rL1 xhmw== X-Gm-Message-State: AFuF++mY+LesjP2BnVPsSaVxmuJBlmV3uhbiyd+MJylK+XUbKS4Fject QNQRSoUnMek8r6KdzWdGnDUKfgnlnb/CqpyhrgQahAczUHTGF7DqHTbskBXgUg== X-Gm-Gg: AYBFou2dFt4WLRNStJpTzRYODPV8Nm5COEl4at/n9+O+UBAnToe13htkgdRWD+XNuGg phpdqIGS/hjisaZNE8g4SpocgiQf1X7N2bmHhMSHt3C0ap575N6LtGbdTUOtE5BmMXab8xhmbap P9H/iq+Yg8z6WsHMlBj5NvJFlmtqGeFsWfF5pGwh/YNVnKAFxTluh1ux8aIVO/P4Bb9+xdUq/kM cQ3XBFWEBgpcHS5DeDDllsOrQWe9Nt61Cfi2r0e4qMocpJLkgGIBno+ckzH4jqASpc21vwW38sO TsNtNO1TTgrxZfhlPY1TC1/k/JSRKmkZRiA5sh3ENpOXHYfDF0vOTF+hz6Jt8DOH0PzTNHMVabm TAt4te+FlTf9oceOlZvmT5S+r6egvIsRhFX8Z1U/N8WomHPS0y0nu6yIyn4tLHGoHJEQpfAAFfD BimvagkGUmDzz8U8RL+c4P5m0BIaEnpNbAljam/0aCZRKVwFwzgZqj5pCpeqbvlngFQn71kMjoZ dfs/W6sPwAJmnWPiXjH2GPwRDu72GWKFvE16+1NNQ== X-Received: by 2002:a05:690c:309:b0:864:a8f4:faf0 with SMTP id 00721157ae682-864a8f521femr59318437b3.10.1788234524422; Mon, 31 Aug 2026 20:48:44 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 1/9] accel/tcg: fold the dynamic cflags into CPUState::tcg_cflags Date: Mon, 31 Aug 2026 23:48:00 -0400 Message-ID: <20260901034808.3524945-2-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112b; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112b.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234652593154100 Content-Type: text/plain; charset="utf-8" curr_cflags() is called once per TB dispatch, from helper_lookup_tb_ptr() and from the cpu_exec() loop. It recomputes the same value every time: uint32_t cflags =3D cpu->tcg_cflags; if (unlikely(cpu_single_stepping(cpu))) { ... } else if (qatomic_read(&one_insn_per_tb)) { ... } else if (qemu_loglevel_mask(CPU_LOG_TB_NOCHAIN)) { ... } That is three loads and three branches on the hottest path in the interpreter, for state that changes only when gdb enables single-step, when one-insn-per-tb is toggled, or when the log mask changes. None of the three has to be sampled at dispatch time. Fold each into CPUState::tcg_cflags where it changes and curr_cflags() becomes a single load of a field that TB lookup has to read anyway. The derived bits -- CF_COUNT_MASK, CF_NO_GOTO_TB, CF_NO_GOTO_PTR and CF_SINGLE_STEP -- are never set by tcg_cflags_set(), so tcg_update_cflags() can recompute them in place without disturbing the rest, and conversely tcg_cflags_set() ORs in its bits without disturbing them. There are three places to call it: - tcg_exec_realizefn(), so that a CPU created after the command line has been parsed starts out with the right value. This covers user-only, where tcg_cpu_init_cflags() is not reached. linux-user's cpu_copy() copies tcg_cflags wholesale, so a cloned thread inherits it. - cpu_single_step(), which changes one CPU. gdb is the only caller that matters; in system mode it runs with the vCPUs stopped, and in user mode gdb_continue_partial() can reach a thread that is still running, because gdb_handlesig() stops only the thread that trapped. That is exactly the plain cross-thread store to another CPU's CPUState that cpu->singlestep_flags already was, read back by that CPU through cpu_single_stepping() in curr_cflags(). This patch changes which field carries it, not who writes it or how. - hmp_one_insn_per_tb() and hmp_log(), which change every CPU while the vCPUs are running, so the update is queued with async_run_on_cpu() and each CPU writes its own cflags from its own thread. The command line spellings of those two settings need nothing: they are parsed before any CPU is realized, so tcg_exec_realizefn() picks them up. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, in a build configured with --enable-lto: before: 1,646,994,254,249 instructions after: 1,562,204,796,597 instructions -5.15% That workload issues 8.4 billion dispatches, so the per-call saving is small but the aggregate is not. The emulated compiler produces byte-identical output before and after. Wall clock does not move: 133.19s to 132.58s, a 0.46% difference against a run-to-run spread larger than that. The removed work is a few predictable loads and branches that the host executes largely in parallel with the surrounding dispatch, so this patch is worth taking for the instruction count and for what it enables, not for a time saving that can be measured on its own. v4: Update the cflags from the HMP handlers for 'log' and 'one-insn-per-tb' rather than from qemu_set_log_internal() and the accelerator property setter. Those are the paths that reach a running vCPU, and the monitor is the only thing that does. Suggested by Richard Henderson. v4: Queue the per-CPU update with async_run_on_cpu() rather than async_safe_run_on_cpu(). Halting the other vCPUs buys nothing: the queued work already runs on the owning CPU's own thread. Suggested by Alex Bennee, who also asked whether there are cross-vCPU updates of tcg_cflags at all. With this change the monitor path has none: the only remaining writer from another thread is cpu_single_step(), above, which is neither new nor made worse here. v4: Move the stub to accel/stubs/, which is where the other accelerator stubs live. Signed-off-by: Matt Turner Reviewed-by: Richard Henderson --- accel/stubs/meson.build | 1 + accel/stubs/tcg-stub.c | 16 ++++++++++++++++ accel/tcg/cpu-exec-common.c | 33 ++++++++++++++++++++++++++++++--- accel/tcg/cpu-exec.c | 3 +++ accel/tcg/internal-common.h | 11 +++++++++-- cpu-target.c | 3 +++ include/system/tcg.h | 12 ++++++++++++ monitor/hmp-cmds.c | 5 +++++ system/runstate-hmp-cmds.c | 4 ++++ 9 files changed, 83 insertions(+), 5 deletions(-) create mode 100644 accel/stubs/tcg-stub.c diff --git ./accel/stubs/meson.build ./accel/stubs/meson.build index 7c6d7ad943..ccad583e64 100644 --- ./accel/stubs/meson.build +++ ./accel/stubs/meson.build @@ -4,6 +4,7 @@ stub_ss.add(files( 'nitro-stub.c', 'mshv-stub.c', 'nvmm-stub.c', + 'tcg-stub.c', 'whpx-stub.c', 'xen-stub.c', )) diff --git ./accel/stubs/tcg-stub.c ./accel/stubs/tcg-stub.c new file mode 100644 index 0000000000..f9e1bd22d6 --- /dev/null +++ ./accel/stubs/tcg-stub.c @@ -0,0 +1,16 @@ +/* + * Stubs for the TCG entry points in system/tcg.h, for binaries that link + * cpu-target.c or the HMP command handlers but not TCG. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include "qemu/osdep.h" +#include "system/tcg.h" + +void tcg_update_cflags(CPUState *cpu) +{ +} + +void tcg_update_all_cflags(void) +{ +} diff --git ./accel/tcg/cpu-exec-common.c ./accel/tcg/cpu-exec-common.c index 44e84344f3..9f3517f36b 100644 --- ./accel/tcg/cpu-exec-common.c +++ ./accel/tcg/cpu-exec-common.c @@ -36,9 +36,16 @@ void tcg_cflags_set(CPUState *cpu, uint32_t flags) cpu->tcg_cflags |=3D flags; } =20 -uint32_t curr_cflags(CPUState *cpu) +/* + * The bits of CPUState::tcg_cflags that tcg_cflags_set() never sets, beca= use + * they are derived from gdb single-step, one-insn-per-tb and -d nochain. + */ +#define CF_DERIVED (CF_COUNT_MASK | CF_NO_GOTO_TB | CF_NO_GOTO_PTR | \ + CF_SINGLE_STEP) + +void tcg_update_cflags(CPUState *cpu) { - uint32_t cflags =3D cpu->tcg_cflags; + uint32_t cflags =3D cpu->tcg_cflags & ~CF_DERIVED; =20 /* * Record gdb single-step. We should be exiting the TB by raising @@ -55,7 +62,27 @@ uint32_t curr_cflags(CPUState *cpu) cflags |=3D CF_NO_GOTO_TB; } =20 - return cflags; + cpu->tcg_cflags =3D cflags; +} + +static void tcg_update_cflags_work(CPUState *cpu, run_on_cpu_data data) +{ + tcg_update_cflags(cpu); +} + +void tcg_update_all_cflags(void) +{ + CPUState *cpu; + + /* + * one-insn-per-tb and -d nochain can both be changed from the monitor + * while the vCPUs are running. Queue the update onto each CPU rather + * than writing tcg_cflags from here, so that the field is only ever + * written by the CPU that owns it. + */ + CPU_FOREACH(cpu) { + async_run_on_cpu(cpu, tcg_update_cflags_work, RUN_ON_CPU_NULL); + } } =20 /* exit the current TB, but without causing any exception to be raised */ diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 257211235d..148e0f583e 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -1068,6 +1068,9 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_target_initialized =3D true; } =20 + /* Pick up one-insn-per-tb and -d nochain from the command line. */ + tcg_update_cflags(cpu); + cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); tlb_init(cpu); #ifndef CONFIG_USER_ONLY diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 9e7be2d78d..853d1b51ee 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -69,8 +69,15 @@ void tlb_destroy(CPUState *cpu); bool tcg_exec_realizefn(CPUState *cpu, Error **errp); void tcg_exec_unrealizefn(CPUState *cpu); =20 -/* current cflags for hashing/comparison */ -uint32_t curr_cflags(CPUState *cpu); +/* + * Current cflags for hashing/comparison. Everything that feeds into the + * value is folded into CPUState::tcg_cflags when it changes, by + * tcg_update_cflags(), so that TB dispatch only has to load it. + */ +static inline uint32_t curr_cflags(CPUState *cpu) +{ + return cpu->tcg_cflags; +} =20 void tb_check_watchpoint(CPUState *cpu, uintptr_t retaddr); =20 diff --git ./cpu-target.c ./cpu-target.c index 4783845c9b..50be591acf 100644 --- ./cpu-target.c +++ ./cpu-target.c @@ -24,6 +24,7 @@ #include "exec/replay-core.h" #include "exec/log.h" #include "hw/core/cpu.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 /* enable or disable single step mode. EXCP_DEBUG is returned by the @@ -35,6 +36,8 @@ void cpu_single_step(CPUState *cpu, unsigned flags) cpu->singlestep_flags, flags); cpu->singlestep_flags =3D flags; =20 + tcg_update_cflags(cpu); + #if !defined(CONFIG_USER_ONLY) const AccelOpsClass *ops =3D cpus_get_accel(); if (ops->update_guest_debug) { diff --git ./include/system/tcg.h ./include/system/tcg.h index 7622dcea30..2c2dbc753b 100644 --- ./include/system/tcg.h +++ ./include/system/tcg.h @@ -17,6 +17,18 @@ extern bool tcg_allowed; #define tcg_enabled() 0 #endif =20 +/* + * Recompute the parts of CPUState::tcg_cflags that TB dispatch consumes b= ut + * tcg_cflags_set() does not provide: gdb single-step, one-insn-per-tb and + * the CPU_LOG_TB_NOCHAIN log flag. Call whenever one of those changes. + * + * tcg_update_cflags() updates one CPU and must be called from that CPU's + * thread, or with it stopped. tcg_update_all_cflags() updates every CPU + * and is safe to call from the monitor while the vCPUs run. + */ +void tcg_update_cflags(CPUState *cpu); +void tcg_update_all_cflags(void); + /** * qemu_tcg_mttcg_enabled: * Check whether we are running MultiThread TCG or not. diff --git ./monitor/hmp-cmds.c ./monitor/hmp-cmds.c index 4e8d996dbb..b83551ea54 100644 --- ./monitor/hmp-cmds.c +++ ./monitor/hmp-cmds.c @@ -39,6 +39,7 @@ #include "system/hw_accel.h" #include "system/memory.h" #include "system/system.h" +#include "system/tcg.h" #include "disas/disas.h" =20 /* Please update hmp-commands.hx when adding or changing commands */ @@ -335,7 +336,11 @@ void hmp_log(Monitor *mon, const QDict *qdict) =20 if (!qemu_set_log(mask, &err)) { error_report_err(err); + return; } + + /* CPU_LOG_TB_NOCHAIN feeds into the per-CPU cflags. */ + tcg_update_all_cflags(); } =20 void hmp_gdbserver(Monitor *mon, const QDict *qdict) diff --git ./system/runstate-hmp-cmds.c ./system/runstate-hmp-cmds.c index 02d1d42bf3..86754a37f8 100644 --- ./system/runstate-hmp-cmds.c +++ ./system/runstate-hmp-cmds.c @@ -22,6 +22,7 @@ #include "qapi/qapi-commands-run-state.h" #include "qobject/qdict.h" #include "qemu/accel.h" +#include "system/tcg.h" =20 void hmp_info_status(Monitor *mon, const QDict *qdict) { @@ -64,6 +65,9 @@ void hmp_one_insn_per_tb(Monitor *mon, const QDict *qdict) /* If the property exists then setting it can never fail */ object_property_set_bool(OBJECT(accel), "one-insn-per-tb", newval, &error_abort); + + /* one-insn-per-tb feeds into the per-CPU cflags. */ + tcg_update_all_cflags(); } =20 void hmp_watchdog_action(Monitor *mon, const QDict *qdict) --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234599; cv=none; d=zohomail.com; s=zohoarc; b=klQorjM4UcaG3ndJA/mEvVhwWZfCy/JTytzSCLhWPtyZHIQJveluBMwMgqu4TkQrP1MDDJ90JZm3KPtLBhVZnK96Xxx54wlKE1tlkpgi5axY236qCo3IYBjZS+Tqvbgzw7t2xy68qOKp6rzNng9xS2ur58ktS0xg/I7ztcLkbls= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234599; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=KhVxgMvDHM0ZH0+WGPYXRDF+WZiLdAMJdvaMWa7aV9F/FSQASH5/lhiuleKfo30xQ+yHkCfuiIYpnmScXAL370G3vw1oCTZoZSCejGAxgFgMC9OV9C8oUf2q/U2gSE8mHdhjQCAoWYu1ubi28x0c/Vm6UNScLFh6+sGr7Mb26sQ= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234599349551.3633795464174; Mon, 31 Aug 2026 20:49:59 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUS-0000bK-En; Mon, 31 Aug 2026 23:48:52 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUP-0000aA-GV for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:49 -0400 Received: from mail-yw1-x112e.google.com ([2607:f8b0:4864:20::112e]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUN-0004L8-TZ for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:49 -0400 Received: by mail-yw1-x112e.google.com with SMTP id 00721157ae682-866e57f63a3so4904617b3.3 for ; Mon, 31 Aug 2026 20:48:47 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4eb56b6dsm7390775d50.7.2026.08.31.20.48.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234526; x=1788839326; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=UQTiKLNu3EYDxM8fq74sUilbEAYkji9QDYicDIEKTuI5N9f1SttlowcTskrCt99ZD+ ryNzn9xuAny5CjZVtRLE1xMRw8bimkBga4OKDj/cK7kZ7LV8zWng5gK4577gRrDhGNd9 iPWXE+cPGQTGfog/TCzQJ/FQeACRq4vhzQr5jWF0AtdO1kazbbmaoh3PkAPQtGWJy6Wc /Qg1yUjm6XN5dmOzmd6JHQz2X681pRyPFiB3VTSB5/1nUvxzELOPZeql+tHsfgPSUCZc WQp5+8VNxSn7oGKj84KbXO2DF2jsQFD0ZRlz+JNh9XEj1gK0jiWCyBsKlx5L4PmLAZMN dHdA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234526; x=1788839326; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=dddZb2vZrpIrHqwA49O0MST0IY6y/BPuLvI5U4ufjcg=; b=V6NjaiQKn2IdgBhsu/BqkJsjAHsbKaRo/U7cD3Vy3gAN2IB/8JNyxcT/8vQhlpFfcZ OWGWnf2eY+YvKAFd4W6Hyr7WyOJLOjmdogeV0KHBjKGmu4KV0raK0Pb7BxTuD5Hb0AcX Q069QWfaqKN0StlWxdF+ft/B6OAJkXdUiLUID52+TkIeDrnGT3kEhfiWxBXuiqWQGtzr WLdrwg8zX/rXueelAtPz5yKqop/khkClqqf9fmYqi4UqxLljoR1sFhQ1+Y+g3Az/WDZ/ 4PyozSDVjIBOZyvsDYsvkGmVsyUo9904UNVSBIIJXJ/+s8BPsVwDlovwNVMJ9aEfAvZs UuyA== X-Gm-Message-State: AFuF++ms6ybVe6Ul5W8Gw5p2j9AiMoxDBAxivETYm1MFJ7FKs+ZXj0H0 v1eG+KlHLSHXXro9Y7Rp8nJatmu9J2sDzioj4Pwi6Z80RnojNZcAqQb6ECbKbg== X-Gm-Gg: AYBFou3aN9ilq3b1fiRvKd0xK45TvLDsE4iSQ0syDTu1BgKFvK4gwzZHknkGmIbuTAS ck/2u4fzXxiveI7KqeL1dsc5szCEvHDeU/L3FTpG4JhyEbRYFM2rYrCx+ogonKWtPP1fV6Fg3Mb Z06DInUOdbsAjm3HuEP9VfMZ6pvYDqPTJdREdrRq2DbRuFdiX3Vx2krgeyMH/SAS8wKKDRRgu3/ HvlN6drE1oV2WrgjNUDFqpsqkOn2ycUKD0e+RqAuNo6tKK6FyZBj3x/Y242Rvxf9an7szHMgIoJ L/4HLqUXS/Sly8algt2Bc9v0fywjdWCU83OQiQV9JePDEzvpMAYS7jzu5G6s9B5Rk1QSqj5Zn02 ZZ2Rrm6qkSDePqlC2mrs+LfVIj5YGj4ftFg72yQ9R65avOalk86C4A4tu4UnNbTVKilTaI1vJCc VLMj6mugpUIXCL8U2WCCUeZjTM0sDj8UKMMIskOcRDxhEzlZ+rHq/Np1pTUGWWkR3MNmfXEuitE WHVjKSAzFKuTrmQg4SMGY9heNrK5y8momCdWaoXFKI3T9Pgu0I= X-Received: by 2002:a05:690e:801:20b0:66f:812e:5c03 with SMTP id 956f58d0204a3-66f812e649dmr2264372d50.2.1788234526648; Mon, 31 Aug 2026 20:48:46 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 2/9] accel/tcg: enlarge the TB jump cache to 64K entries Date: Mon, 31 Aug 2026 23:48:01 -0400 Message-ID: <20260901034808.3524945-3-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112e; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112e.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234602238154100 Content-Type: text/plain; charset="utf-8" The per-CPU TB jump cache has held 4096 entries since it was introduced. That is too small for guests running large programs: an emulated compiler misses often enough that the fallback qht lookup shows up prominently in a profile. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host. The compile performs 34.2 billion TB executions, of which 8.4 billion take the indirect dispatch path. Sizing curve, on top of the preceding patch, instructions retired and wall clock: 12 bits ( 64 KiB): 1,562,204,796,597 132.58s 14 bits ( 256 KiB): 1,493,318,515,396 -4.41% 124.67s -5.97% 16 bits ( 1 MiB): 1,469,772,951,575 -5.92% 121.04s -8.71% 18 bits ( 4 MiB): 1,462,309,832,762 -6.39% 119.82s -9.62% 16 bits is the knee. 18 buys another 0.47% of instructions for four times the memory. It does show a further 1.01% of wall clock, which is outside the 0.70% run-to-run spread at 16 bits, so the effect is probably real -- but paying four times the memory for it is a poor trade, and instructions retired does not account for the data cache pressure of a 4 MiB table. In a perf profile the mechanism is visible directly: tb_htable_lookup(), which is where qht_lookup_custom() lands once it is inlined in an LTO build, falls from 6.10% of samples to 1.66%. The cost is memory: the cache grows from 64 KiB to 1 MiB, once per CPUState. In linux-user that is per guest thread rather than per process, so a threaded guest pays it as many times as it has threads, exactly as system emulation pays it per vCPU. The allocation is g_new0(), so the pages are faulted in as the cache is touched and a thread that runs a small amount of code touches a small part of it, but the address space is committed either way. So this may still want to be tunable, or scaled from the number of CPUs, rather than raised unconditionally. I do not have a threaded workload where the smaller cache is the better trade, and would welcome one. v4: Fix the claim that a linux-user process is a single vCPU. The cache is per CPUState, and linux-user creates one per guest thread. Pointed out by Richard Henderson. Signed-off-by: Matt Turner --- accel/tcg/tb-jmp-cache.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git ./accel/tcg/tb-jmp-cache.h ./accel/tcg/tb-jmp-cache.h index c3a505e394..268dacd7ba 100644 --- ./accel/tcg/tb-jmp-cache.h +++ ./accel/tcg/tb-jmp-cache.h @@ -12,7 +12,7 @@ #include "qemu/rcu.h" #include "exec/cpu-common.h" =20 -#define TB_JMP_CACHE_BITS 12 +#define TB_JMP_CACHE_BITS 16 #define TB_JMP_CACHE_SIZE (1 << TB_JMP_CACHE_BITS) =20 /* --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234632; cv=none; d=zohomail.com; s=zohoarc; b=SSwmqg0k1s+vN3FGNKQaFCN3pri8LGEV7QBOt1FK5hhTJc+D4fWEBAGMHiyg23H+y7VqJCWp4G8micmSX7d9lP8ukdJ+3hcIZQ4GZxuy7rAxgRGhXZFsxcEr6+ngtV8df8c5YCtibVjdrQsS6ptG5sGeouqiRvMMqwzz6sAUJGk= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234632; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=EZ0mUt6zMR1Qj3TCNQcNILV0Gq61IJOwfusx8T+mGxIHF50LQTUBX3RelS3JcgvyEBa7aFTVB88+Lh8Os+vCtiVLZQjQsrB7nWl0vRnkFSPAmtx3fQsNAf7hejqStLLYxszeIWizSTv2HWLM0caJ3araq0Mi4pHrVH1Y54vAo/U= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234632712236.57890198691325; Mon, 31 Aug 2026 20:50:32 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUS-0000bL-Tu; Mon, 31 Aug 2026 23:48:52 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUR-0000aR-4u for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:51 -0400 Received: from mail-yw1-x112a.google.com ([2607:f8b0:4864:20::112a]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUP-0004LN-Hy for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:50 -0400 Received: by mail-yw1-x112a.google.com with SMTP id 00721157ae682-81f36179d72so49828077b3.2 for ; Mon, 31 Aug 2026 20:48:49 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4ea2d6cbsm7335283d50.0.2026.08.31.20.48.47 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234528; x=1788839328; darn=nongnu.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=qPmzS/BdgUMH8cYcvO4X1aHPnxlwtTv1uikyST8hsRu4h4IEkQCheoysDXZeN989Fp xAOl4aLbljABk3Z2700Yz8LQqmQbYwe34i48h6Mxf1M4SNmaUfvpPwBuapXUcLM+rlJ0 DiOiJLH/fDYjmQBjQlsKJmOmm6KeM+iC8T08SwEKuqyf96TdsAde3c+rCgr5rhoQ8oDq dbQyj8xhTV4wW+jpQrhwyt2rYTgFeswZruZ9GW9ufDGxWnSNXVegNOtUv7Cj8LJu9VMk YLsvfQ7uHu2lmqg2TJth4xKoisJb2RwsLDLr9ubBn/xbYkXKVkRiC4i7SlRMVCIZU9qo xT/A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234528; x=1788839328; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=qoJCBK/eJXzgVER7xSpcF+/HlWQEqa8pDFs252ZZUaI=; b=Yf69vSSe4o42NXox8Lu9RC+Y+GX5w85TTTDrtUJFRY2pXk6Oeo6JkylYoR7yJShshN LtszKF607pAV3rV3vV2CdGZXkz0MtX2g3SOxvqqmi7i02o35HErNUQaPYsFlGhdzDiIQ AHSsNEBt+mB1R5JgN52yLGPBoeDX9YjpJ0Ma7IvGziFW1IA+ARN01Y8m/umY9Q9t9z0M wRdEONdUL9JCIeDO8ObGz7VEzEjJNDO5d09fk5TwX8/G8uw/knhv8FfxL8xvXWKPInen sIGvqBQz75K27yHhW7SG4HeNrJpKIiX5mxUbv6s+dz0TpJ5EQaWWdkKzEQ7A/pOHfCE5 xUrw== X-Gm-Message-State: AFuF++nMWm/X/W3UhSfDrytIJkLDnA3J0NT90nRod0aV3QmVko1aUOpF +tZ0BAkB4eiQWnvbulm4ZgtVfmxtLj/3yHO7WYLvFnQXK5yQOZhiLW3VgLXwBA== X-Gm-Gg: AYBFou0HReXiXDAhRdJFinjtnY2j+TrsLZDLCDPciKOMQrAoNu6zRn0D9brZ4ktvsy9 rZ8S8uJO9/ugVKGoe5Rb+DQ6ia+el8QAVrWAJexH9CeJLOQGsjUOutBhIb9hA/ifCd6rsykkDXf tCPSfQD12eG/1bwh9WEBF0RHJW0N6QpXXl/TeK1HoOOZcYOgEr7DWe/b/Mo8szjnzol5v7+z5ia /kV7ovstprauy9TloBru43kGYsKsmjWcHMBi+pH7gjJfoIfWzGIqipRKy77D6e00UiUH8X1fRyg jhGkCnQ/1HRwUS+nOF05L472b6AvQFkjxQjQ4fSopwuoil+A6+M/5XvUIidFbZRS3TKCo22Go9x zq/YimRDcCi3/0Ffu6tA2eUrOoX4N9kssVpH34V8P//NYXE2yT69olW6FHiee9F9q7ThIsG9clO opvf8u5Fl4gA/uqrlDizwaFMn53TZav8uEJKUaaVJO5mqWtqiGw7+F9l4nrEZmC/ytEeHPO6wue 2kFIanEUOUrjdNH3Xh18XRxLZ4z6AW2i3Aj3gyw X-Received: by 2002:a05:690e:1443:b0:66e:6089:b3ab with SMTP id 956f58d0204a3-66e6089bff8mr5321106d50.47.1788234528345; Mon, 31 Aug 2026 20:48:48 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 3/9] accel/tcg: skip the can_do_io stores in user-only builds Date: Mon, 31 Aug 2026 23:48:02 -0400 Message-ID: <20260901034808.3524945-4-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112a; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112a.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234634350158500 Every translation block stores to cpu->neg.can_do_io twice: false before the first instruction, true before the last one. Nothing reads it in a user-only build. There is no memory-mapped I/O in linux-user, and every reader is in system_ss: cputlb.c, watchpoint.c, icount-common.c and tcg-accel-ops-icount.c. Two stores per TB is not much on its own, but TBs are short. An emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) executes 34.2 billion TBs at 6.04 guest instructions each, so this is 68 billion stores for nothing. Measured on an x86-64 host, LTO build, on top of the preceding two patches: before: 1,469,772,951,575 instructions after: 1,402,667,803,616 instructions -4.57% before: 121.04s wall clock after: 115.56s wall clock -4.53% The emulated compiler produces byte-identical output. v3: Use #ifndef CONFIG_USER_ONLY again rather than if (IS_ENABLED(CONFIG_USER_ONLY)). QEMU's IS_ENABLED() is IS_EMPTY(), which is only true for a symbol Meson defines empty; CONFIG_USER_ONLY is defined as 1, so the test was always false and v2 emitted the two stores after all. The measurements above are from the working form. Reviewed-by: Richard Henderson Reviewed-by: Philippe Mathieu-Daud=C3=A9 Signed-off-by: Matt Turner --- accel/tcg/translator.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 57daded60f..6c8fcd7a20 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -21,12 +21,14 @@ #include "disas/disas.h" #include "tb-internal.h" =20 +#ifndef CONFIG_USER_ONLY static void set_can_do_io(DisasContextBase *db, bool val) { QEMU_BUILD_BUG_ON(sizeof_field(CPUState, neg.can_do_io) !=3D 1); tcg_gen_st8_i32(tcg_constant_i32(val), tcg_env, offsetof(CPUState, neg.can_do_io) - sizeof(CPUState)); } +#endif =20 bool translator_io_start(DisasContextBase *db) { @@ -210,17 +212,25 @@ void translator_loop(CPUState *cpu, TranslationBlock = *tb, int *max_insns, /* * Manage can_do_io for the translation block: set to false before * the first insn and set to true before the last insn. + * + * Nothing reads can_do_io in user-only builds. There is no MMIO + * there, and every reader (cputlb.c, watchpoint.c, icount) is in + * system_ss, so skip the two stores per TB entirely. */ if (db->num_insns =3D=3D 1) { tcg_debug_assert(first_insn_start =3D=3D db->insn_start); } else { tcg_debug_assert(first_insn_start !=3D db->insn_start); +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D first_insn_start; set_can_do_io(db, false); +#endif } +#ifndef CONFIG_USER_ONLY tcg_ctx->emit_before_op =3D db->insn_start; set_can_do_io(db, true); tcg_ctx->emit_before_op =3D NULL; +#endif =20 /* May be used by disas_log or plugin callbacks. */ tb->size =3D db->pc_next - db->pc_first; --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234623; cv=none; d=zohomail.com; s=zohoarc; b=ZbCuDQNAXRtPZz8NxDYmo4EX671vmLkke+nsIXYwg+H7v3kVmimI19ReNtqzKVnB759NeUQxnDXU24nX++uU5rPFxnpHSEZCOj0Lei+7tYaf8Qs4Oks5p2fbSQFP5Ma8hPO5QPwmo5W4ku32wNCK4dg3sQ6m0hL7+HyJdmFz24c= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234623; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=HOjRzfQpe3kE34Y2l5KyKrR2biv2SpLqm7CIXewxRtU=; b=CTTqEQr5udPdOWPVZD4vKubv/tY83MlxH1F9wAlM5DrtWA04bY6qwvD75Zm2nyRluwleuk7erZ65Q1TegRskR56fM+EdInbp8GqZoIDnyG+QwZX6fn/rhs1XhTgDJqoV1ZC7FtsH7b/t58knnMz+BSJpeqJ1FmOZxil4Tysy1v8= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234623125862.8494409826422; Mon, 31 Aug 2026 20:50:23 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUj-0000eL-ED; Mon, 31 Aug 2026 23:49:09 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUa-0000cb-Qe for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:00 -0400 Received: from mail-yw1-x112e.google.com ([2607:f8b0:4864:20::112e]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUX-0004M0-NP for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:00 -0400 Received: by mail-yw1-x112e.google.com with SMTP id 00721157ae682-856114a8247so4741047b3.2 for ; Mon, 31 Aug 2026 20:48:57 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85e66f9a5d8sm66385787b3.41.2026.08.31.20.48.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:49 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234537; x=1788839337; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HOjRzfQpe3kE34Y2l5KyKrR2biv2SpLqm7CIXewxRtU=; b=s3WaxJXVR4vZx+CvsTI7De9GD0E//SP/jRfU13QYcpm4X5yuwIGTzFrQ/DXmkQa1MI OWEH2eU4dPdg5c9UkMsIsGtFCGSioLG9N006CsIcBkW34J3K6cFzNdXEeKIBuo0p7kEA FY7biKsnpdwMe+SpG6+Wh4fnnTu+HybHB1i9Q5jSYF2QRDOAdUklQ0+hCET8Mb7PUTzk 8CT/1N2ThloWqQmQynKGkKgNFDjj+jeFp3FlUcwq1ox733eNLhpPJ/YDDJngdJh+pweH Q07IEZ1daBctrMYQPKrkccYn8MDfa1TwUG22OBMiCC8o65B6AqoG6V5wgYmrIm/vxJ66 aOoQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234537; x=1788839337; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=HOjRzfQpe3kE34Y2l5KyKrR2biv2SpLqm7CIXewxRtU=; b=cchsepQ/8DpQDYZwp9QPhGvz+ECyVmypwOk4UoFZd17QEHikVdQ5u2uOuY4cXy4Vgv ojEGxHDaVhEddpoL0XC1K54/C/0rpii0bxE53rjS/ZNpH4fLG50y6GUwOUuP903AbEzt 9ZPoA6t9KYLNqO2eOvI8OTDE6/I8kt/SUnhaR2SHc0IdCdwVURJkWBhfaP9HDJvNCtEO /o1bHdCK8HnQ5LgQTyvkmFi0vaDfI4Ea9WKJMEJ0Mp6Jt2trpZxr6BeARBDyd1o2WtUw 9wNuo/LP8A+7qgBGSOMSuO+KZJ6dXm+ga/hHcr3kDTbYnkTDDG2CgpiUsJV/Oii3k6Tr ga9w== X-Gm-Message-State: AFuF++lWc9H5AJ0kKW5yPSFDYoc3TatB55FL9zTlCDTgfmD9jl8FUMf/ 8H6YhYPOazhR+V1flM3IvC3c/4yUo4xfrsL/MONGHp13eInyBV0wM1o3Lo+Ogg== X-Gm-Gg: AYBFou2CWdcSely8u4O7I/vDmy+Su/ThhCTE43LDPAE5RJVjh86f6SqHVbtYayB2VSh IJ1ya/3bNg3gHBc7i3KhvGnAniXxpNHlTB+Ys1C4gucYaAuqhhgmmuEDfGKGGNhZ2y93fr4JqLq iNSSNvmpP8j0XpW2lzRJdd5jUbgO2JtH9WOqG+jHVyYP90YZb82zI19OGBguIFeJ/lWRYpURmdf 8I8P02GiUY79Z4PybPrsU1pS69jJS6+6F4QvUqfUYdbiiyMGgLQUxa3jXeROvc4FMYR3Y28TZRj 99w1sBQVvVIl7ZJbYTW5YI8B31pTaEiMu+/GYtVjsuyBZZVQsI3gW++B3cOFV2bbIV2NXy8uR1A 5g+y0Js7GwDA7r3/5xWfvvWRZkGdvzNk4a0WgeZ2rF6vLdLl3FuC0rT0FyGqL03dXHqIup0/Fzo h4M/F9DQJKyQ5ZTKE/s1FPxqq4xlWUbq35YfAJrvX6Cg5x0/tAb6IsrSrQsrxhs2Se7FFZMo9cZ Qpv2FumenRWMCCQ+zuajtCkRvLiuPJPq7XMN2XG X-Received: by 2002:a05:690c:386:b0:855:2d6d:2db5 with SMTP id 00721157ae682-85d69e03e9cmr118514477b3.12.1788234531581; Mon, 31 Aug 2026 20:48:51 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 4/9] tcg: add tcg_gen_goto_jc_{i32,i64,tl}() Date: Mon, 31 Aug 2026 23:48:03 -0400 Message-ID: <20260901034808.3524945-5-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112e; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112e.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234626339158500 Content-Type: text/plain; charset="utf-8" tcg_gen_lookup_and_goto_ptr() takes no arguments and emits a call to helper_lookup_tb_ptr(), which recovers the destination PC from env by calling back into the target through TCGCPUOps::get_tb_cpu_state(). At translation time the caller often already has the destination PC in a temp, and knows the flags, cflags and cs_base any destination it may reach has to match, because they are the ones the block being generated was translated with. A later patch uses that to look the destination up inline. But it is not something every caller of tcg_gen_lookup_and_goto_ptr() can promise, and the promise is subtle. target/arm has case DISAS_UPDATE_NOCHAIN: gen_update_pc(dc, curr_insn_len(dc)); /* fall through */ case DISAS_JUMP: gen_goto_ptr(); break; where DISAS_JUMP could make the promise and DISAS_UPDATE_NOCHAIN could not, because it is there precisely because the state changed. Two call sites, one line apart, on opposite sides of the contract. So add a second entry point rather than growing an argument on the first. tcg_gen_lookup_and_goto_ptr() keeps today's meaning and today's signature: dispatch, and let the helper work out where. tcg_gen_goto_jc_*() means dispatch to the destination that env already describes, and takes the pc as proof that the caller knows which one that is. Targets migrate one call site at a time, and a call site that cannot promise simply does not move. The contract is: - @pc holds exactly what get_tb_cpu_state() reports as the destination pc. - The flags and cs_base it reports are the ones this block was translated with, which is what lets them be constants in the generated code. --enable-debug-tcg checks all three against get_tb_cpu_state() at run time, via a new helper_goto_jc_check(). That turns a mistake into an assertion at the offending call site instead of a block that runs with someone else's flags. Five targets have a call site whose pc temp is that key by construction, and are migrated here: alpha, loongarch, mips, ppc and s390x. Nothing else changes; the generated code does not change either, since goto_jc still emits the same helper call for now. For six targets the TB pc is derived and passing the pc temp would be wrong: avr's TB pc is the word address doubled, i386's is eip before segmentation, riscv masks it to 32 bits when xl is MXL_RV32, hppa derives it from the IAQ, hexagon adjusts it inside a hardware loop, and sparc puts npc in cs_base. The remaining seven -- arm, m68k, microblaze, or1k, rx, sh4 and tricore -- have call sites that look like they could move, but I have not convinced myself of the contract for them and have nothing to test them with. Each is a one-line change for whoever wants it, and debug-tcg will say if it is wrong. The i32 and i64 forms are separate functions, with a _tl alias in tcg-op.h, as for most everything else. A translator built for more than one value of TARGET_LONG_BITS cannot include tcg-op.h and calls the sized form directly, which is what s390x does here. Neither form takes the TranslationBlock: tcg_ctx->gen_tb is the block being generated, the same one tcg_gen_goto_tb() and tcg_gen_lookup_and_goto_ptr() already read, so there is no way for a caller to pass the wrong one. v4: Split out of "tcg: probe the TB jump cache inline instead of calling a helper", which did the API change and the inline probe in one patch. Requested by Richard Henderson. v5: Add a new interface rather than growing an argument on tcg_gen_lookup_and_goto_ptr(), and check the contract under --enable-debug-tcg. Requested by Richard Henderson, who named it tcg_gen_goto_jc_*(); the DISAS_UPDATE_NOCHAIN example above is his. v5: Define _i32 and _i64 entry points with a _tl alias in tcg-op.h, rather than one entry point taking a TCGTemp. Requested by Richard Henderson: as targets migrate to single-binary, code is built once and stops relying on TARGET_LONG_BITS, so the TCGTemp split was the wrong shape. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec.c | 27 +++++++++ accel/tcg/tcg-runtime.h | 4 ++ include/tcg/tcg-op-common.h | 20 +++++++ include/tcg/tcg-op.h | 2 + target/alpha/translate.c | 4 +- .../tcg/insn_trans/trans_branch.c.inc | 2 +- target/loongarch/tcg/translate.c | 4 +- target/mips/tcg/nanomips_translate.c.inc | 2 +- target/mips/tcg/translate.c | 6 +- target/ppc/translate.c | 4 +- target/s390x/tcg/translate.c | 4 +- tcg/tcg-op.c | 59 +++++++++++++++++-- 12 files changed, 119 insertions(+), 19 deletions(-) diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index 148e0f583e..ca90a77a7b 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -407,6 +407,33 @@ const void *HELPER(lookup_tb_ptr)(CPUArchState *env) return tb->tc.ptr; } =20 +#ifdef CONFIG_DEBUG_TCG +/** + * helper_goto_jc_check: check the contract of tcg_gen_goto_jc_*() + * @env: current cpu state + * @pc: the destination pc the caller passed at translation time + * @flags: the flags the dispatching block was translated with + * @cs_base: the cs_base the dispatching block was translated with + * + * A goto_jc looks the destination up on the caller's @pc with the flags a= nd + * cs_base of the block doing the dispatching, so all three have to be what + * get_tb_cpu_state() reports by the time the dispatch runs. That is a + * property of the translator, not of the generated code, so check it here + * rather than leaving a target that gets it wrong to be debugged as a blo= ck + * running with someone else's flags. + */ +void HELPER(goto_jc_check)(CPUArchState *env, uint64_t pc, uint64_t flags, + uint64_t cs_base) +{ + CPUState *cpu =3D env_cpu(env); + TCGTBCPUState s =3D cpu->cc->tcg_ops->get_tb_cpu_state(cpu); + + assert(s.pc =3D=3D pc); + assert(s.flags =3D=3D flags); + assert(s.cs_base =3D=3D cs_base); +} +#endif + /* Return the current PC from CPU, which may be cached in TB. */ static vaddr log_pc(CPUState *cpu, const TranslationBlock *tb) { diff --git ./accel/tcg/tcg-runtime.h ./accel/tcg/tcg-runtime.h index 0b832176b3..ec99170698 100644 --- ./accel/tcg/tcg-runtime.h +++ ./accel/tcg/tcg-runtime.h @@ -22,6 +22,10 @@ DEF_HELPER_FLAGS_1(ctpop_i64, TCG_CALL_NO_RWG_SE, i64, i= 64) =20 DEF_HELPER_FLAGS_1(lookup_tb_ptr, TCG_CALL_NO_WG_SE, cptr, env) =20 +#ifdef CONFIG_DEBUG_TCG +DEF_HELPER_FLAGS_4(goto_jc_check, TCG_CALL_NO_WG_SE, void, env, i64, i64, = i64) +#endif + DEF_HELPER_FLAGS_1(exit_atomic, TCG_CALL_NO_WG, noreturn, env) =20 #ifndef IN_HELPER_PROTO diff --git ./include/tcg/tcg-op-common.h ./include/tcg/tcg-op-common.h index 9b321f959c..4f334faaaa 100644 --- ./include/tcg/tcg-op-common.h +++ ./include/tcg/tcg-op-common.h @@ -85,6 +85,26 @@ void tcg_gen_goto_tb(unsigned idx); */ void tcg_gen_lookup_and_goto_ptr(void); =20 +/** + * tcg_gen_goto_jc_i32() - dispatch to the destination TB via the jump cac= he + * tcg_gen_goto_jc_i64() - dispatch to the destination TB via the jump cac= he + * @pc: temp holding the destination guest PC + * + * As tcg_gen_lookup_and_goto_ptr(), but the caller states where the + * dispatch is going, which allows the lookup to be done inline. + * + * The contract is that when this runs, the CPU state must already be + * exactly the destination's: @pc must hold what get_tb_cpu_state() would + * report as the destination pc, and the flags and cs_base it would report + * must be the ones the block being generated was translated with. A + * translator that has not finished updating the state, or whose pc is + * derived rather than being the lookup key -- avr's word address, i386's + * eip before segmentation -- must use tcg_gen_lookup_and_goto_ptr() + * instead. --enable-debug-tcg checks the contract at runtime. + */ +void tcg_gen_goto_jc_i32(TCGv_i32 pc); +void tcg_gen_goto_jc_i64(TCGv_i64 pc); + void tcg_gen_plugin_cb(unsigned from); void tcg_gen_plugin_mem_cb(TCGv_i64 addr, unsigned meminfo); =20 diff --git ./include/tcg/tcg-op.h ./include/tcg/tcg-op.h index 3721164236..cd4794d745 100644 --- ./include/tcg/tcg-op.h +++ ./include/tcg/tcg-op.h @@ -38,6 +38,7 @@ typedef TCGv_i32 TCGv; #define tcgv_tl_temp tcgv_i32_temp #define tcg_gen_qemu_ld_tl tcg_gen_qemu_ld_i32 #define tcg_gen_qemu_st_tl tcg_gen_qemu_st_i32 +#define tcg_gen_goto_jc_tl tcg_gen_goto_jc_i32 #elif TARGET_LONG_BITS =3D=3D 64 typedef TCGv_i64 TCGv; #define tcg_temp_new() tcg_temp_new_i64() @@ -45,6 +46,7 @@ typedef TCGv_i64 TCGv; #define tcgv_tl_temp tcgv_i64_temp #define tcg_gen_qemu_ld_tl tcg_gen_qemu_ld_i64 #define tcg_gen_qemu_st_tl tcg_gen_qemu_st_i64 +#define tcg_gen_goto_jc_tl tcg_gen_goto_jc_i64 #else #error Unhandled TARGET_LONG_BITS value #endif diff --git ./target/alpha/translate.c ./target/alpha/translate.c index c66e3f9c14..8318487cd8 100644 --- ./target/alpha/translate.c +++ ./target/alpha/translate.c @@ -449,7 +449,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, int32_t disp) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_pc_disp(ctx, cpu_pc, disp); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_pc); } } =20 @@ -2917,7 +2917,7 @@ static void alpha_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cpu) gen_pc_disp(ctx, cpu_pc, 0); /* FALLTHRU */ case DISAS_PC_UPDATED: - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_pc); break; case DISAS_PC_UPDATED_NOCHAIN: tcg_gen_exit_tb(NULL, 0); diff --git ./target/loongarch/tcg/insn_trans/trans_branch.c.inc ./target/lo= ongarch/tcg/insn_trans/trans_branch.c.inc index da07778658..d4318dfa43 100644 --- ./target/loongarch/tcg/insn_trans/trans_branch.c.inc +++ ./target/loongarch/tcg/insn_trans/trans_branch.c.inc @@ -27,7 +27,7 @@ static bool trans_jirl(DisasContext *ctx, arg_jirl *a) tcg_gen_mov_tl(cpu_pc, addr); tcg_gen_movi_tl(dest, make_address_pc(ctx, ctx->base.pc_next + 4)); gen_set_gpr(a->rd, dest, EXT_NONE); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_pc); ctx->base.is_jmp =3D DISAS_NORETURN; return true; } diff --git ./target/loongarch/tcg/translate.c ./target/loongarch/tcg/transl= ate.c index 124dce6269..6ac0c3773a 100644 --- ./target/loongarch/tcg/translate.c +++ ./target/loongarch/tcg/translate.c @@ -111,7 +111,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned tb_= slot_idx, vaddr dest) tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { tcg_gen_movi_tl(cpu_pc, dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_pc); } } =20 @@ -311,7 +311,7 @@ static void loongarch_tr_tb_stop(DisasContextBase *dcba= se, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: tcg_gen_movi_tl(cpu_pc, ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_pc); break; case DISAS_TOO_MANY: gen_goto_tb(ctx, 0, ctx->base.pc_next); diff --git ./target/mips/tcg/nanomips_translate.c.inc ./target/mips/tcg/nan= omips_translate.c.inc index 4b0b01ba37..106f49990d 100644 --- ./target/mips/tcg/nanomips_translate.c.inc +++ ./target/mips/tcg/nanomips_translate.c.inc @@ -2406,7 +2406,7 @@ static void gen_compute_nanomips_pbalrsc_branch(Disas= Context *ctx, int rs, =20 /* unconditional branch to register */ tcg_gen_mov_tl(cpu_PC, btarget); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_PC); } =20 /* nanoMIPS Branches */ diff --git ./target/mips/tcg/translate.c ./target/mips/tcg/translate.c index e3467d1525..dea1ba4c1e 100644 --- ./target/mips/tcg/translate.c +++ ./target/mips/tcg/translate.c @@ -4374,7 +4374,7 @@ static void gen_goto_tb(DisasContext *ctx, unsigned t= b_slot_idx, tcg_gen_exit_tb(ctx->base.tb, tb_slot_idx); } else { gen_save_pc(dest); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_PC); } } =20 @@ -11014,7 +11014,7 @@ static void gen_branch(DisasContext *ctx, int insn_= bytes) } else { tcg_gen_mov_tl(cpu_PC, btarget); } - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_PC); break; default: LOG_DISAS("unknown branch 0x%x\n", proc_hflags); @@ -15244,7 +15244,7 @@ static void mips_tr_tb_stop(DisasContextBase *dcbas= e, CPUState *cs) switch (ctx->base.is_jmp) { case DISAS_STOP: gen_save_pc(ctx->base.pc_next); - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_PC); break; case DISAS_NEXT: case DISAS_TOO_MANY: diff --git ./target/ppc/translate.c ./target/ppc/translate.c index 06ed2adf10..21e21102fc 100644 --- ./target/ppc/translate.c +++ ./target/ppc/translate.c @@ -3664,7 +3664,7 @@ static void gen_lookup_and_goto_ptr(DisasContext *ctx) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_nip); } } =20 @@ -6690,7 +6690,7 @@ static void ppc_tr_tb_stop(DisasContextBase *dcbase, = CPUState *cs) pmu_count_insns(ctx); } =20 - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_tl(cpu_nip); break; =20 case DISAS_EXIT_UPDATE: diff --git ./target/s390x/tcg/translate.c ./target/s390x/tcg/translate.c index 1b6023168b..906951c7df 100644 --- ./target/s390x/tcg/translate.c +++ ./target/s390x/tcg/translate.c @@ -1162,7 +1162,7 @@ static DisasJumpType help_branch(DisasContext *s, Dis= asCompare *c, tcg_gen_goto_tb(0); tcg_gen_exit_tb(s->base.tb, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_i64(psw_addr); } =20 gen_set_label(lab); @@ -6477,7 +6477,7 @@ static void s390x_tr_tb_stop(DisasContextBase *dcbase= , CPUState *cs) if (dc->exit_to_mainloop) { tcg_gen_exit_tb(NULL, 0); } else { - tcg_gen_lookup_and_goto_ptr(); + tcg_gen_goto_jc_i64(psw_addr); } break; default: diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index 28d3b2a847..a2f35359fe 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -2715,18 +2715,65 @@ void tcg_gen_goto_tb(unsigned idx) tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } =20 +static void gen_lookup_tb_ptr_and_goto(void) +{ + TCGv_ptr ptr =3D tcg_temp_ebb_new_ptr(); + + gen_helper_lookup_tb_ptr(ptr, tcg_env); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); + tcg_temp_free_ptr(ptr); +} + void tcg_gen_lookup_and_goto_ptr(void) { - TCGv_ptr ptr; - if (tcg_ctx->gen_tb->cflags & CF_NO_GOTO_PTR) { tcg_gen_exit_tb(NULL, 0); return; } =20 plugin_gen_disable_mem_helpers(); - ptr =3D tcg_temp_ebb_new_ptr(); - gen_helper_lookup_tb_ptr(ptr, tcg_env); - tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); - tcg_temp_free_ptr(ptr); + gen_lookup_tb_ptr_and_goto(); +} + +/* + * The common half of tcg_gen_goto_jc_i32() and tcg_gen_goto_jc_i64(). @pc + * is widened to i64 because the jump cache is keyed on a vaddr; for a + * 32-bit guest PC that is its zero extension. + */ +static void gen_goto_jc(TCGv_i64 pc) +{ + const TranslationBlock *tb =3D tcg_ctx->gen_tb; + + if (tb->cflags & CF_NO_GOTO_PTR) { + tcg_gen_exit_tb(NULL, 0); + return; + } + + plugin_gen_disable_mem_helpers(); + +#ifdef CONFIG_DEBUG_TCG + /* + * The caller has asserted that env already describes the destination. + * Check it, rather than leaving a target that gets it wrong to be + * debugged as a block that runs with someone else's flags. + */ + gen_helper_goto_jc_check(tcg_env, pc, tcg_constant_i64(tb->flags), + tcg_constant_i64(tb->cs_base)); +#endif + + gen_lookup_tb_ptr_and_goto(); +} + +void tcg_gen_goto_jc_i64(TCGv_i64 pc) +{ + gen_goto_jc(pc); +} + +void tcg_gen_goto_jc_i32(TCGv_i32 pc) +{ + TCGv_i64 pc64 =3D tcg_temp_ebb_new_i64(); + + tcg_gen_extu_i32_i64(pc64, pc); + gen_goto_jc(pc64); + tcg_temp_free_i64(pc64); } --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234637; cv=none; d=zohomail.com; s=zohoarc; b=G9T9A8pJuJwctpEJPX6g9yikAyklo35koMi8TYGhieOJShKdLynEgbNqINSMsE3R0xGxYsKNvJzmop++fPmfYcZxBBv4UMo7lrD56mByUlfwZbKlwmyhlPNi6UCjgXuEjVWZaTMI6uUfoT75KaVknxHS8VDNDV4jEwbcDMlNY3w= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234637; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=bwP7cQaRP/GqhipJ0CESBAL8FySLg3GbdoTs9nGIJiA=; b=hhNvWrEtZ5CIlMgokkLPZ90C8LGMCoMl63odOUTrjRtU1pABts57HzNsFdFmcPD9UNUXXTQzp8Z7Xc/AJePhSuQV08dseA1SnadwUcwkP3WNs9Sye4d36ulbzEJg4rz3TzGCuJZGV9R+wL4BfnpUhVAxjkHTQbwCCpGjwkfwC1w= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234637585121.87583237653689; Mon, 31 Aug 2026 20:50:37 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUa-0000cZ-M1; Mon, 31 Aug 2026 23:49:00 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUY-0000c3-Of for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:58 -0400 Received: from mail-yx1-xb12c.google.com ([2607:f8b0:4864:20::b12c]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUW-0004Lw-W1 for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:58 -0400 Received: by mail-yx1-xb12c.google.com with SMTP id 956f58d0204a3-66d23c88af0so2044488d50.1 for ; Mon, 31 Aug 2026 20:48:56 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4ed1fbebsm7511499d50.17.2026.08.31.20.48.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234536; x=1788839336; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=bwP7cQaRP/GqhipJ0CESBAL8FySLg3GbdoTs9nGIJiA=; b=CVI6oL1SkYwShF9iInuPNBIjNk5Y0wPycpB1H9AdxQXjJqNRGe9VXlgHqWlAmmYwpl 8loKCFCycF/ZmIByjTYHt8TkqS+xOcTiZKNkWP7HZ/M1xU3hs9Ln6RPTjpS2I2ipjMUS b72U8Y7pDk9dGK2a61iMazH/0c2ldMd4mm84ipi1+B+Gfra6oIfpDzMr5wCAlqi3Xp92 togRk1upKXVKUc+RA0pfn4aMSER4ou3AX2y4j1x8KMEksH3GBIvQclTuSdNW1+inlFKz HsUjk0z1O7Skvg3XGNNAd/U8gIDskOdwRlfIuT2Qjps0WfMG3ZZj66fUNkyJT6444Xt5 7KlA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234536; x=1788839336; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=bwP7cQaRP/GqhipJ0CESBAL8FySLg3GbdoTs9nGIJiA=; b=T1Qj8vabKn1tKgNhX87fnTnZW/mX/Cb3T+/summRjWiyNgz/IjOF73XHOZEV+ghrxh EmHikbzeUFQESIlZNds2qEuXTdv25fk593gBh0bcyq+Bi2niw6lEQ6YpmIoh7Po5rJfj +6H9M8itsT9VhErk9LTIEXdXOf0LupTTjzX2SL64pX6zaje+dN40wHXW+Let3+fzdt4T YDY8SkYq21x8MnAV36yFDg4FXLiDb34bi4uihUCPQbVZi1za/HbQWuJUyJ093yyIeqz6 CdEJ5OeVS/nXypW6SDGI6OFixn4+Q47W68KaOZ7eqjBITVQZ7Lgr+/q6tKxdWLFBiRVA 63uA== X-Gm-Message-State: AFuF++lGO+cGI8sS3b1sR/LIT9jyq4/whsxAKxU5fdtuGDFkY6q5+vjW /k49oOl60Z9bBMEftBft1krjUsa/qKfghppB1Cfx14CKW/twKKWd2seo/xv2YA== X-Gm-Gg: AYBFou2hQacsaVEvM56p0D1R9YYkO/AYWiWDWzYpRcayU9e/LDWBnodglvp+UBqWUom FqER55SLUjwzNLf61a0389aljSxGEiE28rnpTtizM4lQeaoRAHicas9EiURw5aWri0E0cUUpsxM KA/L8S5V1rnQiHFUGbpRtE0Rmw1nngGrAn/H5tnj80DEALSYydTjjJkhPy+2qeIMXkGO0e1LPsL y7UBVusBOA7DI2AcKInVzL081UWO6uYuF0S9o06l6/Cj/arpOutlWHtBwQXgNJoAqW8yn/KpZxb OKFenbqtLb60BMRXFpipkTPQ1vjm2p8lwchKbIUwKhaMoY+nPJL1It22uxFteD/ZGgZzluqWSfl mU6/VSt/z+WYiWyBEecj1q5ckYYa4fDZCdYDJjfL3iaL11cEyUCiOtXDcmUfFWo+qsd97Cm/70O gu9JdS754aljDvJ31eWlYNmeZy9IpP1fcxtc2yq12c9kya02sgOZ3ZK7vcAaDfImUYSS+GqV1X0 W7DX22pyeg8exCV23BWmwg3v3zbL+6A9zL7BOnn X-Received: by 2002:a05:690e:1386:b0:66c:51c5:61df with SMTP id 956f58d0204a3-66f875fc4a2mr1869738d50.44.1788234533325; Mon, 31 Aug 2026 20:48:53 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 5/9] accel/tcg: add CF_NO_GOTO_JC, set while a breakpoint is present Date: Mon, 31 Aug 2026 23:48:04 -0400 Message-ID: <20260901034808.3524945-6-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b12c; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb12c.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234640303154100 Content-Type: text/plain; charset="utf-8" The next patch dispatches a goto_jc by probing the TB jump cache from generated code. That probe cannot check everything helper_lookup_tb_ptr() checks, and the one that matters is breakpoints: check_for_breakpoints() raises EXCP_DEBUG on an exact pc match and selects CF_BP_PAGE cflags for the rest of the page, and inserting a breakpoint deliberately invalidates no TB. The probe does compare the destination's cflags against the cflags of the block doing the dispatching, and only takes the destination when they are equal. So a cflag is all that is needed. Add CF_NO_GOTO_JC, set it in CPUState::tcg_cflags while cpu->breakpoints is non-empty, and blocks translated from then on both decline to dispatch inline themselves -- the next patch makes them emit the plain helper call -- and are unreachable from blocks that do, because their cflags no longer match. The two ends of the flag are cpu_breakpoint_insert() and cpu_breakpoint_remove_by_ref(), which are the only places the list changes. Both already run either on the CPU's own thread or with it stopped, or reach another CPU exactly as cpu_single_step() does, which is where the previous patch put the same kind of update. That leaves blocks translated before the breakpoint was inserted, which are still live and still chain to each other. They do so on the old cflags, so inline dispatch among them keeps working until the vCPU reaches its main loop, which then looks up with the new cflags and translates afresh. In system mode gdb inserts breakpoints with the vCPUs stopped, so there is no window at all. In user mode the window is the one goto_tb chaining already has: a chained direct jump consults nothing either, and is not broken by inserting a breakpoint. Nothing reads CF_NO_GOTO_JC yet; the next patch does. v5: New patch, replacing "accel/tcg: give the TB jump cache a second base pointer for generated code", which forced the same fallback by pointing generated code at a zero-filled jump cache when a breakpoint was inserted, and needed a cross-thread poison and an un-poison race to do it. Richard Henderson suggested a cflag instead, and pointed out that the previous patch had already shown how to update tcg_cflags from gdbstub. The base pointer comes back later in the series, for pending exits, which a cflag cannot express. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec-common.c | 13 ++++++++++++- cpu-common.c | 7 +++++++ include/exec/translation-block.h | 1 + 3 files changed, 20 insertions(+), 1 deletion(-) diff --git ./accel/tcg/cpu-exec-common.c ./accel/tcg/cpu-exec-common.c index 9f3517f36b..a3148bbf8f 100644 --- ./accel/tcg/cpu-exec-common.c +++ ./accel/tcg/cpu-exec-common.c @@ -41,7 +41,7 @@ void tcg_cflags_set(CPUState *cpu, uint32_t flags) * they are derived from gdb single-step, one-insn-per-tb and -d nochain. */ #define CF_DERIVED (CF_COUNT_MASK | CF_NO_GOTO_TB | CF_NO_GOTO_PTR | \ - CF_SINGLE_STEP) + CF_SINGLE_STEP | CF_NO_GOTO_JC) =20 void tcg_update_cflags(CPUState *cpu) { @@ -62,6 +62,17 @@ void tcg_update_cflags(CPUState *cpu) cflags |=3D CF_NO_GOTO_TB; } =20 + /* + * A block that dispatches through the jump cache inline does not cons= ult + * cpu->breakpoints, and inserting a breakpoint deliberately invalidat= es + * nothing. Give blocks translated while one is set a distinct cflags= , so + * that they neither dispatch inline themselves nor are reached by a b= lock + * that does, and check_for_breakpoints() gets to run on every dispatc= h. + */ + if (unlikely(!QTAILQ_EMPTY(&cpu->breakpoints))) { + cflags |=3D CF_NO_GOTO_JC; + } + cpu->tcg_cflags =3D cflags; } =20 diff --git ./cpu-common.c ./cpu-common.c index adb76b3a78..3178601987 100644 --- ./cpu-common.c +++ ./cpu-common.c @@ -22,6 +22,7 @@ #include "exec/cpu-common.h" #include "hw/core/cpu.h" #include "qemu/lockable.h" +#include "system/tcg.h" #include "trace/trace-root.h" =20 QemuMutex qemu_cpu_list_lock; @@ -429,6 +430,9 @@ int cpu_breakpoint_insert(CPUState *cpu, vaddr pc, int = flags, *breakpoint =3D bp; } =20 + /* The first breakpoint takes the CPU off the inline dispatch path. */ + tcg_update_cflags(cpu); + trace_breakpoint_insert(cpu->cpu_index, pc, flags); return 0; } @@ -456,6 +460,9 @@ void cpu_breakpoint_remove_by_ref(CPUState *cpu, CPUBre= akpoint *bp) { QTAILQ_REMOVE(&cpu->breakpoints, bp, entry); =20 + /* The last breakpoint puts the CPU back on it. */ + tcg_update_cflags(cpu); + trace_breakpoint_remove(cpu->cpu_index, bp->pc, bp->flags); g_free(bp); } diff --git ./include/exec/translation-block.h ./include/exec/translation-bl= ock.h index 40cc699031..8c4778c681 100644 --- ./include/exec/translation-block.h +++ ./include/exec/translation-block.h @@ -84,6 +84,7 @@ struct TranslationBlock { #define CF_NOIRQ 0x00010000 /* Generate an uninterruptible TB */ #define CF_PCREL 0x00020000 /* Opcodes in TB are PC-relative */ #define CF_BP_PAGE 0x00040000 /* Breakpoint present in code page */ +#define CF_NO_GOTO_JC 0x00080000 /* Do not dispatch via the inline prob= e */ #define CF_CLUSTER_MASK 0xff000000 /* Top 8 bits are cluster ID */ #define CF_CLUSTER_SHIFT 24 =20 --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234558; cv=none; d=zohomail.com; s=zohoarc; b=XFU0AJBzPz8mLUdKpcfx8EWqjnBdHZFB1co2ySxeQ+IcyHNAuXlXi1pnBq+BDTEIOYddzZ3yNoHZiPfkX8sHA19EeiDzeOKeOgkXCnp5Q5OLeN4F29Fk+vhigqqUAbkROiDlwqS6305v+MA1rROsxVrEipL0h6d0+DOIcze5QsQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234558; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=hU417OCLlkLOavVulyBQtQTYsid7DrkibHmsCfBSBuo=; b=PwYTf7CBM2jfu2ig4jKZPdZ5Irj+sKiumhG0AXhWMxYevdGzwGA0AmAjju6o27K9dAYvsz7Q9Z7MpE+2UcFSUqiTVyFD/ysHT2lcI1AHAEIfD9SrLYXoTHgWHTuQ7EceNIplovyDcOGlkvaRM8/UdBC8ihPHqK/is59IRR0de/Y= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234558782888.9458126948261; Mon, 31 Aug 2026 20:49:18 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUa-0000cX-3g; Mon, 31 Aug 2026 23:49:00 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUY-0000bw-FD for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:58 -0400 Received: from mail-yx1-xb130.google.com ([2607:f8b0:4864:20::b130]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUW-0004Lr-85 for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:48:58 -0400 Received: by mail-yx1-xb130.google.com with SMTP id 956f58d0204a3-66dee7a0c09so3083649d50.0 for ; Mon, 31 Aug 2026 20:48:55 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4ed2a78bsm7444135d50.20.2026.08.31.20.48.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234535; x=1788839335; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=hU417OCLlkLOavVulyBQtQTYsid7DrkibHmsCfBSBuo=; b=QfQ/gVTwJlYcrQrkbZ9r805kweEjnTCsUtZfnNBvtxAGLFlStP6h4y1SgThhHWLm6K 0x+l6qspvSNd3BjFp2Jfy826AokF2ks2jIqWNxFHWhL5PIlGzSxEeL72YZZUOBTKRMB7 bD5uau8ZV985RmSmRZyzn/TX92ADlmVjwhtl+J1cLkVBGMUgE4i3srYfj9e9CX4UfNIZ xoH16AOZ5F0U2e099IHfgqxcO2JTzsE/arZDubxw08tsn900RWR9J0AOWp2WVRr65f+N WGQrNtD46//WAepOVx48QCODSTvjbI/IqpTJmQj5bLcQw8eAyjy99WabhiZPGbAVnSPt gyKA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234535; x=1788839335; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=hU417OCLlkLOavVulyBQtQTYsid7DrkibHmsCfBSBuo=; b=DYimtegxmxqqN2FI5CKhW6zXrKxzqP5eCEBTit8ywopFeYArma43jkUNcTLwOvqiTl xyexGPqMVXgP104J6yxwcOiPAZrq/fFuD364CuAbsVBpkX0Ks1hW3LXy+8nUZHnPXWnP D9V8nkM2T0KCFtqoVOc72eSY3dNjTHqoQsrnacCYZp6FupSM3sg7g3CqWT2pnqwzYcWm yMhqtd0zCOwHpsJHa4dAEA2RnICIj2G4b2EC0uLVg4Ci/g8dVdoX+WK69b/S/L5csZNS s2T/HtBNm2x/G+QcwzJmBsNax8BS4ZwVcd9zvjty9P/AZWDO06ZkwtY39ooUkYXf/iqP A6AA== X-Gm-Message-State: AFuF++mySejQDEIfMH3Du6wUXfNgD4QB42V8E2eOL9wocp1GpmVVJh4Y AraAj+2xH4E4LAU8O+N+Fjg96r9rghclQX36bFGnxXZtTFBaZobUQCWNLA5h9A== X-Gm-Gg: AYBFou28rhRHv444XNOuMWryA30S5P8hbfvTQ3pBsAeAi2LsiRJl1+lxZZ9k0S2WkfA PZtOiZUl49n5YdPOdKSbERLw3uH2NFALvEXWhjbOQOtrTMxghxJyvjrdZtxN1g7hXFxvEK8oZiy XB9loH0kDoG4SKD6Hp4MJV7L61pdtpbl6z6YuMbTrycyl/4b9KVAIoVRmCJhyz9y+3ryyjfS7Rw GVyWHFrr9ZuCcuNot/W6SkOaIH0AQJOsrGDCrOw87m9tp+KOedInBHGWiS4TkmxsEic/H2BTzKB z9VyFVd7Ipp6fOZ77zPMzn5u8EzMPl+IYfG4p2emnDkYtYe1SqMDc4tHtqa7SpgtiNb5TXhhq2t 3AOSOoNW+nu45pOUNpNDCrGgbWne2ouRxh1ZTeyxL9GOJRY7MjzXWHbcLV/kixGrir9L6ofQIUj lopbN/kpNujIxKgE2nFrUdpAuLPo0VSOU0MHhh/C+kTY5E+WyffJhksxAILaYz2mUDmBm+O/0as RpCQOIFcyLJuScoOTsVtHd4sBzWwqbiggPMjnAR X-Received: by 2002:a05:690e:1589:10b0:66c:bd06:e89b with SMTP id 956f58d0204a3-66e4c746d23mr7671577d50.32.1788234534859; Mon, 31 Aug 2026 20:48:54 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 6/9] RFC: tcg: probe the TB jump cache inline instead of calling a helper Date: Mon, 31 Aug 2026 23:48:05 -0400 Message-ID: <20260901034808.3524945-7-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b130; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb130.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234561397154100 Content-Type: text/plain; charset="utf-8" Every indirect branch that cannot use goto_tb ends in a dispatch that calls helper_lookup_tb_ptr(). For an emulated compiler that is 8.4 billion helper calls in a single translation unit: 24.6% of all TB exits take this path, because jsr/ret/jmp have a register destination and because goto_tb is restricted to same-page targets. The helper itself is already tight, but each call pays for a call frame, the can_do_io store, the get_tb_cpu_state() indirect call through TCGCPUOps, curr_cflags(), and a breakpoint check, before it gets to the jump cache probe that almost always hits (95.8% for this workload). Emit the probe inline instead, for the callers that have migrated to tcg_gen_goto_jc_*(). Those supply what it needs: the destination PC is in a TCG temp, and the flags, cflags and cs_base the destination must match are constants at translation time. The fast path is therefore a hash, four guarded loads and a goto_ptr. Only a miss calls the helper, which still owns filling the cache. Two details matter for the generated code. The flags and cflags guards are folded into a single aligned 64-bit load and compare, since the fields are adjacent. And each path emits its own goto_ptr rather than branching to a shared one: a temp live across the label is spilled and reloaded on every dispatch, which cost 6.3% on its own. The flags and cflags constants are safe against the other things that can change them. CF_PARALLEL is only ever set by begin_parallel_context(), which flushes first, so no block predating it survives to dispatch. gdb single-step is only turned on with the CPU stopped, and a block translated without CF_SINGLE_STEP can only be re-entered through tb_lookup(), which from then on demands the new cflags -- so a stale-cflags block is never the one running. Breakpoints are handled by CF_NO_GOTO_JC, added by the previous patch: while one is set, blocks are translated with a cflags that both keeps them off the inline path and keeps them unreachable from blocks already on it. What is left is one_insn_per_tb and -d nochain; see below. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding patches: before: 1,402,667,803,616 instructions after: 916,415,123,244 instructions -34.67% before: 115.56s wall clock after: 85.59s wall clock -25.94% The gap between the two is the point at which this stops being a straight-line win: the helper call was highly predictable work that the host pipelined well, so removing it retires far fewer instructions than it saves time. IPC falls from 2.48 to 2.17 across this patch for that reason. Despite emitting more code, this also reduces instruction cache pressure, because a dispatch no longer jumps into qemu's .text and evicts translated code: before: 11,735,141,703 L1-icache-load-misses after: 7,154,863,292 L1-icache-load-misses -39.0% The mechanism is visible directly in a profile: helper_lookup_tb_ptr() falls from 31.01% of samples to 0.35%, and qemu's own .text falls from 38.8% to 5.3%, with the balance moving into generated code. Combined with the preceding patches, against an unmodified LTO build, 1,646,994,254,249 instructions fall to 916,415,123,244, or -44.36%. The emulated compiler produces byte-identical output throughout. A follow-up worth having: the probe is emitted entirely out of generic TCG ops, and several backends can do much better than the result. x86_64 and s390x have memory-operand comparisons; aarch64 can form env + off + h * 16 with a shift-add, load (tb, pc) and (cs_base, flags) with two ldp, and halve the branches with ccmp. That wants a backend expansion of a dedicated opcode, which is a separate series. Open issues, hence RFC: - one_insn_per_tb and CPU_LOG_TB_NOCHAIN can be toggled from the monitor while a vCPU is inside a block that was translated without them. The block keeps dispatching inline on the old cflags until it exits for some other reason. This is the same window goto_tb chaining already has, since a chained direct jump consults nothing either, but it is worth saying out loud. - The jump cache entry is read without qatomic_read(); entries are invalidated concurrently by setting tb to NULL. - Only alpha has been measured. The other four targets that use goto_jc are built and boot-tested only. v4: Split out of the patch that also changed the tcg_gen_lookup_and_goto_ptr() API and introduced tb_jmp_cache_probe, which are now the two preceding patches. Requested by Richard Henderson. v4: Emit the softmmu form of tb_jmp_cache_hash_func() under CONFIG_SOFTMMU rather than the user-only form everywhere. v3 emitted the user-only hash unconditionally, which was wrong for system mode and was only not a correctness bug because a wrong index simply misses. Caught by Richard Henderson. tcg-op.c is compiled once per build rather than once per target, but CONFIG_SOFTMMU is set for it, and TARGET_PAGE_BITS -- a load from target_page here -- is fixed long before any translation happens. v4: Compare the pc before testing tb for NULL. On a hash miss the pc is the field most likely to differ, and an unused entry has a zero pc that only pc 0 can match, so the tb test buys nothing ahead of it. Suggested by Richard Henderson. v4: Assert that offsetof(TranslationBlock, flags) is 8-byte aligned, since folding the flags and cflags guards into one 64-bit load relies on it and nothing else does. Requested by Richard Henderson. v4: Zero-extend a 32-bit guest PC instead of falling back to the helper. Suggested by Richard Henderson. The high half then folds to a compare against zero. v4: Describe cs_base in the probe as a second word of target-specific flags rather than by name. Suggested by Richard Henderson. v5: Build the folded flags/cflags constant with deposit64() rather than under #if HOST_BIG_ENDIAN, so both arms compile on every host. Requested by Richard Henderson. v5: Read cpu->tb_jmp_cache directly, and honor CF_NO_GOTO_JC rather than a poisoned base pointer, which is no longer how breakpoints are handled. A separate base pointer comes back later in the series for pending exits. v5: Note the backend expansion this wants as a follow-up. Suggested by Richard Henderson, whose list it is. Signed-off-by: Matt Turner --- include/tcg/tcg-op-common.h | 3 +- tcg/tcg-op.c | 110 +++++++++++++++++++++++++++++++++++- 2 files changed, 111 insertions(+), 2 deletions(-) diff --git ./include/tcg/tcg-op-common.h ./include/tcg/tcg-op-common.h index 4f334faaaa..f41f3ee58f 100644 --- ./include/tcg/tcg-op-common.h +++ ./include/tcg/tcg-op-common.h @@ -91,7 +91,8 @@ void tcg_gen_lookup_and_goto_ptr(void); * @pc: temp holding the destination guest PC * * As tcg_gen_lookup_and_goto_ptr(), but the caller states where the - * dispatch is going, which allows the lookup to be done inline. + * dispatch is going, so the TB jump cache is probed inline and only a miss + * reaches helper_lookup_tb_ptr(). * * The contract is that when this runs, the CPU state must already be * exactly the destination's: @pc must hold what get_tb_cpu_state() would diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index a2f35359fe..b10b2d66d5 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -28,6 +28,8 @@ #include "tcg/tcg-op-common.h" #include "exec/translation-block.h" #include "exec/plugin-gen.h" +#include "hw/core/cpu.h" +#include "../accel/tcg/tb-hash.h" #include "tcg-internal.h" #include "tcg-has.h" =20 @@ -2735,6 +2737,102 @@ void tcg_gen_lookup_and_goto_ptr(void) gen_lookup_tb_ptr_and_goto(); } =20 +static void gen_jmp_cache_hash(TCGv_i64 h, TCGv_i64 pc) +{ +#ifdef CONFIG_SOFTMMU + /* + * tb_jmp_cache_hash_func(), softmmu form. TARGET_PAGE_BITS is a load + * from target_page in this translation unit, but it is decided long + * before any translation happens, so it is a constant here. + */ + int shift =3D TARGET_PAGE_BITS - TB_JMP_PAGE_BITS; + TCGv_i64 tmp =3D tcg_temp_ebb_new_i64(); + + tcg_gen_shri_i64(tmp, pc, shift); + tcg_gen_xor_i64(tmp, tmp, pc); + tcg_gen_shri_i64(h, tmp, shift); + tcg_gen_andi_i64(h, h, TB_JMP_PAGE_MASK); + tcg_gen_andi_i64(tmp, tmp, TB_JMP_ADDR_MASK); + tcg_gen_or_i64(h, h, tmp); + tcg_temp_free_i64(tmp); +#else + /* tb_jmp_cache_hash_func(), user-only form. */ + tcg_gen_shri_i64(h, pc, TB_JMP_CACHE_BITS); + tcg_gen_xor_i64(h, h, pc); + tcg_gen_andi_i64(h, h, TB_JMP_CACHE_SIZE - 1); +#endif +} + +static void gen_jmp_cache_probe(TCGv_i64 pc, const TranslationBlock *tb) +{ + TCGv_ptr jc, ent, tbp, ptr; + TCGv_i64 h, tmp; + TCGLabel *slow; + uint64_t fpair; + + QEMU_BUILD_BUG_ON(sizeof(((CPUJumpCache *)0)->array[0]) !=3D 16); + QEMU_BUILD_BUG_ON(offsetof(CPUJumpCache, array[0].pc) % 8 !=3D 0); + /* One 64-bit load has to cover both, so they must be adjacent... */ + QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, cflags) !=3D + offsetof(TranslationBlock, flags) + 4); + /* ...and aligned, which nothing else currently relies on. */ + QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, flags) % 8 !=3D 0); + + jc =3D tcg_temp_ebb_new_ptr(); + ent =3D tcg_temp_ebb_new_ptr(); + tbp =3D tcg_temp_ebb_new_ptr(); + ptr =3D tcg_temp_ebb_new_ptr(); + h =3D tcg_temp_ebb_new_i64(); + tmp =3D tcg_temp_ebb_new_i64(); + slow =3D gen_new_label(); + + /* ent =3D &jc->array[tb_jmp_cache_hash_func(pc)] */ + gen_jmp_cache_hash(h, pc); + tcg_gen_shli_i64(h, h, 4); + + tcg_gen_ld_ptr(jc, tcg_env, + offsetof(CPUState, tb_jmp_cache) - sizeof(CPUState)); + tcg_gen_trunc_i64_ptr(ent, h); + tcg_gen_add_ptr(ent, jc, ent); + + /* + * The pc first: on a hash miss it is the field most likely to differ, + * and an entry whose tb is NULL has a zero pc that only pc 0 matches. + */ + tcg_gen_ld_i64(tmp, ent, offsetof(CPUJumpCache, array[0].pc)); + tcg_gen_brcond_i64(TCG_COND_NE, tmp, pc, slow); + + tcg_gen_ld_ptr(tbp, ent, offsetof(CPUJumpCache, array[0].tb)); + tcg_gen_brcondi_ptr(TCG_COND_EQ, tbp, 0, slow); + + /* + * flags and cflags are adjacent uint32_t, so one aligned 64-bit load + * and compare covers both. + */ + fpair =3D (HOST_BIG_ENDIAN + ? deposit64(tb->cflags, 32, 32, tb->flags) + : deposit64(tb->flags, 32, 32, tb->cflags)); + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, flags)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, fpair, slow); + + /* + * cs_base is a second word of target-specific flags despite the name, + * and the pc alone does not imply it on a target that uses it. + */ + tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, cs_base)); + tcg_gen_brcondi_i64(TCG_COND_NE, tmp, tb->cs_base, slow); + + tcg_gen_ld_ptr(ptr, tbp, offsetof(TranslationBlock, tc.ptr)); + tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr)); + + /* + * Emit a second goto_ptr rather than branching to a shared one: a temp + * live across the label would be spilled and reloaded on every dispat= ch. + */ + gen_set_label(slow); + gen_lookup_tb_ptr_and_goto(); +} + /* * The common half of tcg_gen_goto_jc_i32() and tcg_gen_goto_jc_i64(). @pc * is widened to i64 because the jump cache is keyed on a vaddr; for a @@ -2761,7 +2859,17 @@ static void gen_goto_jc(TCGv_i64 pc) tcg_constant_i64(tb->cs_base)); #endif =20 - gen_lookup_tb_ptr_and_goto(); + /* + * A breakpoint is the one thing the probe cannot check for itself, so + * while one is set the flag is set too and every dispatch takes the + * helper, which does check. See tcg_update_cflags(). + */ + if (tb->cflags & CF_NO_GOTO_JC) { + gen_lookup_tb_ptr_and_goto(); + return; + } + + gen_jmp_cache_probe(pc, tb); } =20 void tcg_gen_goto_jc_i64(TCGv_i64 pc) --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234629; cv=none; d=zohomail.com; s=zohoarc; b=fpTIY2qLm8N/5+kkOB1WS22aJ5MUTEtHSWReVACriBGr2gRCdSXC+nrzN7eREpXKXWt6MfOgOf2v8L8104sXAcyaW1+/0x4I9Y5VEhbFOG2syQeTXW0s+PhS8VBWTXzHS+g6Zy37/EPgx9vOuxpsWFT9OG/8g13bD4kGCVEwxRg= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234629; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=fbnESuNHhFjy/7oDq+9x064cCTF6lHATYsYiIMQH+oRUFBb6fWW9sxNjkkA8CSaIVklfGbfzxgsShGXAWtXEfqt/8k6nuQH32jyq6mJSzJbEaycPHBiE/UJa+mRcWMX51FwexpUSgBDvQEXzpUBsP9JZ6oaD78Gx1UOtnaSj95I= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234629863629.6388452048573; Mon, 31 Aug 2026 20:50:29 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUn-0000g4-Mp; Mon, 31 Aug 2026 23:49:13 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUc-0000cx-Ki for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:04 -0400 Received: from mail-yw1-x1134.google.com ([2607:f8b0:4864:20::1134]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUZ-0004MB-2e for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:02 -0400 Received: by mail-yw1-x1134.google.com with SMTP id 00721157ae682-8200b55dc47so40814327b3.3 for ; Mon, 31 Aug 2026 20:48:58 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85e5ed1dc88sm66934857b3.19.2026.08.31.20.48.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234538; x=1788839338; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=JqnoARXK3tCv/uKcoHuX/492ZqF7WpaNOjDEYR55U/BeDri1hW8HmXUQFTTbuHSv8X qpL3pdCOG96+EYgVZpzu0LBK6x9tIl/psPGorbLfEv8U9qMmhcLpzwfEZ+uU44Fa4D4g bt5xPgUSTYjSf4yALhOx9xohdRKlZXWBC7D8wE779Yeq2ThPqXYX2EEAcwI5R7SCcRPs w6ZU8IHqiTkgk2Q0OD5OHW3Mj0MTi/cbYf+VzpfIbKTN7GRiRW1sKzCwdyaZYz6BJoOA GB7o5IN6ucqiGndr2f9A8v3Nl+gmlc6Cq9r3SM2kmuTsYX8ZIyd2VRlV5quSKvLhA6DB IATw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234538; x=1788839338; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=GTyVp0vlfc04HWA2atZi7IJalMnDmcZPoaeKa+M3cjU=; b=DL1HFeuQNmX2lQ6EalVrb0g8rJFuY+MByccYZ0Ul56wQbE/MVgR78y4MuU/smFPBAZ Yay5pPA4kypkb04KmQ3D0LczvXY6OTFaSh+B6wDWhK4o3HQLMy/JJOoXmOf7OYIyqwFi mdClz6LbxDJn2TdQq6HTRSqz8XShm6+KlOsin1AkI2iGULiB7KdiuwvvyRnOJzzzf0g+ Z1exHsQ1DteuvasNfbVfVBg3B/DaNVYhqWpJYihv5sL82iTWYC62ppnxLiNT9PiSIUTj XzqIOtFsRE4jIbkaXNJ7PZJ+KxkjS18LyRf75YOYpkMpV9ImwGVJ+8V//At82HtDqwIJ hOag== X-Gm-Message-State: AFuF++mHaLUNWwn5CSwMA98Y5+W4eIFHtAnANYd9mwrqUn+5w3YoOY2f kuQiDdnwk6TWCVFFT6kIqEd5MTdZXwdYz66JWKdCKr6tk3YwLDP4xdpjpa0J0g== X-Gm-Gg: AYBFou0lSItuo6HKPgB1ZQpnuOq7hW84uvuV344hg/zAKBxYB8nnL+kVIQ3b78vVyNL cGG5Q0ZUa3z3mQY0FnMhAzmrMsexmjeAVpz/arFfBbaG4A6p00BfMlRU0E+MZSVMbgpNKXDO9rF SrkSWNlH8KjfjlmYQFw7JfeD6N7ooG2mEe9fOPd1M2m6dnwGENWiLJFADiGBFNbFoAgkgfYWH2D XKI2J8Jkq+3gTDO0lQBNVj3RNYeywD/UUzh/K9DQgU6ONxxiLkPBllPBvptAiLfr9rvJP9rfTNB sscyIHWonPfrpJefJVAlwFQV6HvqODzjZx/Fp8kvtccz+xvW3J/omSIt3vmVqKJRSryxFPRjWGE cDPmu9jw3IBkqNlPwiGrhbDtdkjBFwLd6KkTVz6bF9SMlOJf8QBv/e1FRuH01WjTMRRTujhfB44 xQEZx2YGAe7cOSSbGYDQtbWoZoelbi+lbdJtlE0M6e52B/2VtciV5PPWyKT0L8wA6DyUDuixFwQ 0y+qTYuR0UpTzuim5OLKW7eHBhqsdYPtB4mpaWa X-Received: by 2002:a05:690c:660b:b0:856:b8a6:5e79 with SMTP id 00721157ae682-85d6ec1c18amr108548797b3.23.1788234537695; Mon, 31 Aug 2026 20:48:57 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 7/9] RFC: accel/tcg: allow cross-page goto_tb chaining in user-only builds Date: Mon, 31 Aug 2026 23:48:06 -0400 Message-ID: <20260901034808.3524945-8-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1134; envelope-from=mattst88@gmail.com; helo=mail-yw1-x1134.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234634342154100 Content-Type: text/plain; charset="utf-8" translator_use_goto_tb() refuses to chain unless the destination is on the same page as the start of the TB. For guests whose text is much larger than a page this is expensive: an emulated alpha gcc compiling a 255k line translation unit takes the indirect dispatch path for 8.4 billion of its 34.2 billion TB exits, and a large share of those are ordinary direct branches that simply crossed an 8 KiB page boundary. The restriction was made unconditional by d3a2a1d803 ("accel/tcg: Introduce translator_use_goto_tb"), whose rationale was: Various targets avoid the page crossing test for CONFIG_USER_ONLY, but that is wrong: mmap and mprotect can change page permissions. That is true, but in user-only builds the invalidation path already covers it. There are no page tables: every mmap, mprotect and munmap reaches page_set_flags(), which calls tb_invalidate_phys_range() whenever the flags actually change, and tb_phys_invalidate() calls tb_jmp_unlink() to reset incoming jumps. A chained cross-page jump is therefore broken whenever the destination page's permissions change. This is not true in system mode, where TBs are keyed by physical address and a page table change invalidates nothing, so the restriction is kept there. The rule protects one more thing, which the original rationale does not mention: it guarantees that execution cannot enter a page without a TB lookup, and so without check_for_breakpoints(). That is what makes a breakpoint set after a block was translated take effect, since insertion deliberately invalidates nothing. A link established before the breakpoint was set would jump straight over it. So the chaining is only enabled for a run that can never acquire a breakpoint. In user-only mode every breakpoint comes from gdb -- BP_CPU is g_assert_not_reached() there, and the guest cannot ask for one -- and gdb has to be requested with -g before the first block is translated, even though with suspend=3Dn it may connect later. gdb_may_set_breakpoints() reports whether it was, and is fixed for the lifetime of the process. Add tests/tcg/multiarch/test-xpage-chain.c to cover both hazards directly. It writes the last instruction of one page and the first of the next, so that the fall-through between them is a cross-page goto_tb, runs it 200000 times so the chain is established, then checks that mprotect(PROT_NONE) makes the next call fault, and that different code written into the page once it is mapped back runs rather than a stale translation. The two instructions -- set the return value register, and return -- are all the architecture specific code there is; thirteen architectures supply them and the rest skip. The test detects the hazard it is meant to detect: with the tb_invalidate_phys_range() call in page_set_flags() commented out, it fails both phases, executing page B after PROT_NONE and returning the stale result. Run with -b, the same binary stops once the chain is established and lets tests/tcg/multiarch/gdbstub/xpage-bp.py set a breakpoint on the far side of= it, which the next call has to stop on. With gdb_may_set_breakpoints() forced to false so that the chaining stays on under gdb, that breakpoint is missed and the test fails, which is what makes it a test of the gate rather than of gdb. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation on an x86-64 host, LTO build, on top of the preceding patches: before: 916,415,123,244 instructions after: 891,254,240,071 instructions -2.75% before: 85.59s wall clock after: 81.45s wall clock -4.84% Note that this is worth more in time than in instructions, the reverse of the preceding patch: a chained jump replaces a cache probe whose loads can miss, so the instructions it removes are more expensive than average. Measured before the inline jump cache probe, when a missed chain cost a helper call rather than an inline probe, the same change was worth -7.9%. RFC because this reverses a deliberate decision and the reasoning above wants review from someone who knows the invalidation paths better than I do. v3: Only take the shortcut when no gdbstub was requested. The same-page rule also forces a lookup, and so a breakpoint check, on entry to every page; without that, a chain established before a breakpoint was set runs past it. Reported by Richard Henderson. v3: Change translator_use_goto_tb() rather than translator_is_same_page(). i386, riscv and s390x call translator_is_same_page() for something else -- enforcing that only a single-insn TB may cross a page -- and v2 changed their TB boundaries in user-only mode as a side effect. alpha does not call it, so the numbers above are unaffected. v3: Add the gdbstub half of the test. v4: Move the test to tests/tcg/multiarch so that every *-user target runs it, rather than only alpha. Requested by Alex Bennee. The direct branch is gone with it: a fall-through off the end of a page is a cross-page goto_tb just the same, and needs no per-architecture branch encoding or displacement arithmetic, only "set the return value" and "return". Built and run under qemu-user on aarch64, alpha, arm, hppa, loongarch64, m68k, mips, ppc, ppc64le, riscv64, s390x, sh4, sparc64 and x86_64; ppc64 ELFv1 skips, because a function pointer there is a descriptor rather than a code address. Signed-off-by: Matt Turner --- accel/tcg/translator.c | 33 ++- gdbstub/user.c | 14 + include/gdbstub/user.h | 11 + tests/tcg/multiarch/Makefile.target | 12 +- tests/tcg/multiarch/gdbstub/xpage-bp.py | 37 +++ tests/tcg/multiarch/test-xpage-chain.c | 336 ++++++++++++++++++++++++ 6 files changed, 441 insertions(+), 2 deletions(-) create mode 100644 tests/tcg/multiarch/gdbstub/xpage-bp.py create mode 100644 tests/tcg/multiarch/test-xpage-chain.c diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 6c8fcd7a20..8879cd626f 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -15,6 +15,9 @@ #include "accel/tcg/cpu-mmu-index.h" #include "exec/target_page.h" #include "exec/translator.h" +#ifdef CONFIG_USER_ONLY +#include "gdbstub/user.h" +#endif #include "exec/plugin-gen.h" #include "tcg/tcg-op-common.h" #include "internal-common.h" @@ -110,6 +113,34 @@ bool translator_is_same_page(const DisasContextBase *d= b, vaddr addr) return ((addr ^ db->pc_first) & TARGET_PAGE_MASK) =3D=3D 0; } =20 +/* + * Whether a direct jump may be chained to a destination outside the page + * the TB started in. + * + * In user-only mode there are no page tables. Every mmap, mprotect and + * munmap goes through page_set_flags(), which calls tb_invalidate_phys_ra= nge() + * whenever the flags actually change, and tb_phys_invalidate() unlinks + * incoming jumps. A cross-page link is therefore broken whenever the + * destination page's permissions change. + * + * What the same-page rule also provides is that execution cannot enter a = page + * without a TB lookup, and so without check_for_breakpoints(), which is w= hat + * makes a breakpoint set after a block was translated take effect. Nothi= ng + * invalidates on breakpoint insertion, so a link established beforehand w= ould + * jump straight over it. In user-only mode breakpoints only ever come fr= om + * gdb -- BP_CPU is g_assert_not_reached() there and the guest has no way = to + * ask for one -- and gdb has to be requested with -g before the first blo= ck + * is translated, so a run that has no gdbstub can never acquire a breakpo= int. + */ +static bool use_cross_page_goto_tb(void) +{ +#ifdef CONFIG_USER_ONLY + return !gdb_may_set_breakpoints(); +#else + return false; +#endif +} + bool translator_use_goto_tb(DisasContextBase *db, vaddr dest) { /* Suppress goto_tb if requested. */ @@ -118,7 +149,7 @@ bool translator_use_goto_tb(DisasContextBase *db, vaddr= dest) } =20 /* Check for the dest on the same page as the start of the TB. */ - return translator_is_same_page(db, dest); + return use_cross_page_goto_tb() || translator_is_same_page(db, dest); } =20 void translator_loop(CPUState *cpu, TranslationBlock *tb, int *max_insns, diff --git ./gdbstub/user.c ./gdbstub/user.c index 9e6f9a6f37..d810f0f38c 100644 --- ./gdbstub/user.c +++ ./gdbstub/user.c @@ -470,6 +470,18 @@ static void *gdbserver_accept_thread(void *arg) =20 #define USAGE "\nUsage: -g {port|path}[,suspend=3D{y|n}]" =20 +/* + * Set before the guest runs and never cleared, so that code translated at + * any point can rely on it: with suspend=3Dn gdb may connect long after + * startup, and once connected it can insert a breakpoint at any time. + */ +static bool gdbserver_requested; + +bool gdb_may_set_breakpoints(void) +{ + return gdbserver_requested; +} + bool gdbserver_start(const char *args, Error **errp) { g_auto(GStrv) argv =3D g_strsplit(args, ",", 0); @@ -513,6 +525,8 @@ bool gdbserver_start(const char *args, Error **errp) return false; } =20 + gdbserver_requested =3D true; + if (suspend) { if (gdbserver_accept(port, gdb_fd, port_or_path)) { gdb_handlesig(first_cpu, 0, NULL, NULL, 0); diff --git ./include/gdbstub/user.h ./include/gdbstub/user.h index 654986d483..c091cd9758 100644 --- ./include/gdbstub/user.h +++ ./include/gdbstub/user.h @@ -11,6 +11,17 @@ =20 #define MAX_SIGINFO_LENGTH 128 =20 +/** + * gdb_may_set_breakpoints() - whether a breakpoint can ever be inserted + * + * In user-only mode every breakpoint comes from gdb, and gdb is only ever + * reachable if -g was given at startup, before the guest ran a single + * instruction. A run that has no gdbstub can therefore never acquire a + * breakpoint, which lets translation take shortcuts that a breakpoint + * would invalidate. Stays true once true, even if gdb detaches. + */ +bool gdb_may_set_breakpoints(void); + /** * gdb_handlesig() - yield control to gdb * @cpu: CPU diff --git ./tests/tcg/multiarch/Makefile.target ./tests/tcg/multiarch/Make= file.target index ab4bf9c5d5..f8a91fed2c 100644 --- ./tests/tcg/multiarch/Makefile.target +++ ./tests/tcg/multiarch/Makefile.target @@ -143,6 +143,15 @@ run-gdbstub-follow-fork-mode-parent: follow-fork-mode --bin $< --test $(MULTIARCH_SRC)/gdbstub/follow-fork-mode-parent.py, \ following parents on fork) =20 +# The chaining this exercises is only enabled when no gdbstub was requeste= d, +# so what is under test here is that requesting one turns it back off. +run-gdbstub-xpage-bp: test-xpage-chain + $(call run-test, $@, $(GDB_SCRIPT) \ + --gdb $(GDB) \ + --qemu $(QEMU) --qargs "$(QEMU_OPTS)" \ + --bin "$< -b" --test $(MULTIARCH_SRC)/gdbstub/xpage-bp.py, \ + breakpoint behind an established cross-page chain) + run-gdbstub-late-attach: late-attach $(call run-test, $@, env LATE_ATTACH_PY=3D1 $(GDB_SCRIPT) \ --gdb $(GDB) \ @@ -159,7 +168,8 @@ EXTRA_RUNS +=3D run-gdbstub-sha1 run-gdbstub-qxfer-auxv= -read \ run-gdbstub-registers run-gdbstub-prot-none \ run-gdbstub-catch-syscalls run-gdbstub-follow-fork-mode-child \ run-gdbstub-follow-fork-mode-parent \ - run-gdbstub-qxfer-siginfo-read run-gdbstub-late-attach + run-gdbstub-qxfer-siginfo-read run-gdbstub-late-attach \ + run-gdbstub-xpage-bp =20 # ARM Compatible Semi Hosting Tests # diff --git ./tests/tcg/multiarch/gdbstub/xpage-bp.py ./tests/tcg/multiarch/= gdbstub/xpage-bp.py new file mode 100644 index 0000000000..f40024f16d --- /dev/null +++ ./tests/tcg/multiarch/gdbstub/xpage-bp.py @@ -0,0 +1,37 @@ +"""Test that a breakpoint set after a cross-page chain is established is h= it. + +translator_use_goto_tb() lets a direct branch chain to another page in +user-only builds, which is only safe because a run with no gdbstub can nev= er +acquire a breakpoint. This runs with one, so the chaining must be off and +the breakpoint must still be reached. + +This runs as a sourced script (via -x, via run-test.py). + +SPDX-License-Identifier: GPL-2.0-or-later +""" +from test_gdbstub import main, report + + +def run_test(): + """Run through the tests one by one""" + gdb.Breakpoint("break_here") + gdb.execute("continue") + + # The chain exists by now; put a breakpoint on the far side of it. + target =3D int(gdb.parse_and_eval("(unsigned long)page_b_entry")) + if target =3D=3D 0: + report(True, "no code emitters for this architecture, skipped") + return + gdb.execute("break *{}".format(target)) + gdb.execute("continue") + + pc =3D int(gdb.parse_and_eval("(unsigned long)$pc")) + report(pc =3D=3D target, "stopped at {:#x}, expected {:#x}".format(pc,= target)) + + gdb.execute("delete") + gdb.execute("continue") + exitcode =3D int(gdb.parse_and_eval("$_exitcode")) + report(exitcode =3D=3D 0, "{} =3D=3D 0".format(exitcode)) + + +main(run_test) diff --git ./tests/tcg/multiarch/test-xpage-chain.c ./tests/tcg/multiarch/t= est-xpage-chain.c new file mode 100644 index 0000000000..a4e34149e7 --- /dev/null +++ ./tests/tcg/multiarch/test-xpage-chain.c @@ -0,0 +1,336 @@ +/* + * Cross-page TB chaining hazard test. + * + * Two adjacent pages of hand-written code. The last instruction of page A + * sets the return value and falls through into page B, which returns; a TB + * always ends at a page boundary, so page A reaches page B through a + * cross-page goto_tb. + * + * Phase 1: run it enough times that QEMU chains TB_A -> TB_B. + * Phase 2: mprotect page B away. Re-running must fault. + * Phase 3: map it back and write different code into it. Re-running must + * execute the NEW code, not a stale chained translation. + * + * With -b, phases 2 and 3 are replaced by a stop at break_here(), where t= he + * gdbstub test sets a breakpoint on page B -- after the chain exists -- a= nd + * checks that re-running the chain still stops on it. See + * tests/tcg/multiarch/gdbstub/xpage-bp.py. + * + * The code the two pages hold is architecture specific, so each + * architecture supplies two emitters: + * + * emit_set_ret(p, val) - set the integer return value register to val + * emit_ret(p) - return to the caller + * + * both writing at @p and returning the number of bytes written. Neither + * may contain a branch: the fall-through from page A into page B is the + * whole point, and a delay slot must not straddle the boundary. An + * architecture that supplies neither skips the test. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include +#include +#include +#include +#include + +static inline size_t put32(void *p, uint32_t insn) +{ + memcpy(p, &insn, sizeof(insn)); + return sizeof(insn); +} + +static inline size_t put16(void *p, uint16_t insn) +{ + memcpy(p, &insn, sizeof(insn)); + return sizeof(insn); +} + +#if defined(__aarch64__) +#define HAVE_EMITTERS +/* movz w0, #val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x52800000u | ((uint32_t)val << 5)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0xd65f03c0u); /* ret */ +} +#elif defined(__alpha__) +#define HAVE_EMITTERS +/* lda $0, val($31) */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x201f0000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x6bfa8001u); /* ret */ +} +#elif defined(__arm__) +#define HAVE_EMITTERS +/* mov r0, #val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0xe3a00000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0xe12fff1eu); /* bx lr */ +} +#elif defined(__hppa__) +#define HAVE_EMITTERS +/* ldi val, %ret0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x341c0000u | ((uint32_t)val << 1)); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0xe840c000u); /* bv %r0(%rp) */ + return n + put32((char *)p + n, 0x08000240u); /* nop (delay slot= ) */ +} +#elif defined(__i386__) || defined(__x86_64__) +#define HAVE_EMITTERS +/* mov $val, %eax */ +static size_t emit_set_ret(void *p, int val) +{ + uint32_t imm =3D val; + *(unsigned char *)p =3D 0xb8; + return 1 + put32((char *)p + 1, imm); +} +static size_t emit_ret(void *p) +{ + *(unsigned char *)p =3D 0xc3; /* ret */ + return 1; +} +#elif defined(__loongarch64) +#define HAVE_EMITTERS +/* ori $a0, $zero, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x03800004u | ((uint32_t)val << 10)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x4c000020u); /* jr $ra */ +} +#elif defined(__m68k__) +#define HAVE_EMITTERS +/* moveq #val, %d0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put16(p, 0x7000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + return put16(p, 0x4e75u); /* rts */ +} +#elif defined(__mips__) +#define HAVE_EMITTERS +/* li $v0, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x24020000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0x03e00008u); /* jr $ra */ + return n + put32((char *)p + n, 0x00000000u); /* nop (delay slot= ) */ +} +/* + * ELFv1 function pointers are descriptors rather than code addresses, so + * there is nothing to call the raw code through. + */ +#elif defined(__powerpc__) && \ + (!defined(__powerpc64__) || (defined(_CALL_ELF) && _CALL_ELF =3D=3D = 2)) +#define HAVE_EMITTERS +/* li r3, val */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x38600000u | (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x4e800020u); /* blr */ +} +#elif defined(__riscv) +#define HAVE_EMITTERS +/* addi a0, zero, val -- the 4 byte form, never c.li */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x00000513u | ((uint32_t)val << 20)); +} +static size_t emit_ret(void *p) +{ + return put32(p, 0x00008067u); /* jalr zero, 0(ra= ) */ +} +#elif defined(__s390x__) +#define HAVE_EMITTERS +/* lghi %r2, val */ +static size_t emit_set_ret(void *p, int val) +{ + size_t n =3D put16(p, 0xa729u); + return n + put16((char *)p + n, (uint16_t)val); +} +static size_t emit_ret(void *p) +{ + return put16(p, 0x07feu); /* br %r14 */ +} +#elif defined(__sh__) +#define HAVE_EMITTERS +/* mov #val, r0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put16(p, 0xe000u | (uint8_t)val); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put16(p, 0x000bu); /* rts */ + return n + put16((char *)p + n, 0x0009u); /* nop (delay slot= ) */ +} +#elif defined(__sparc__) +#define HAVE_EMITTERS +/* mov val, %o0 */ +static size_t emit_set_ret(void *p, int val) +{ + return put32(p, 0x90102000u | (uint32_t)(val & 0x1fff)); +} +static size_t emit_ret(void *p) +{ + size_t n =3D put32(p, 0x81c3e008u); /* retl */ + return n + put32((char *)p + n, 0x01000000u); /* nop (delay slot= ) */ +} +#endif + +/* Where the fall-through lands, for the gdbstub test to breakpoint on. */ +void *page_b_entry; + +/* Somewhere for the gdbstub test to stop once the chain is established. */ +void __attribute__((noinline)) break_here(void) +{ + asm volatile (""); +} + +#ifdef HAVE_EMITTERS +static sigjmp_buf jb; +/* + * Written by the SIGSEGV handler and read by main(), so it must not be + * cached in a register across the faulting call. + */ +static volatile sig_atomic_t caught; + +static void segv(int sig) +{ + caught =3D 1; + siglongjmp(jb, 1); +} +#endif + +int main(int argc, char **argv) +{ + bool bp_mode =3D argc > 1 && strcmp(argv[1], "-b") =3D=3D 0; +#ifndef HAVE_EMITTERS + printf("SKIP: no code emitters for this architecture\n"); + if (bp_mode) { + break_here(); + } + return 0; +#else + unsigned char tmp[16]; + struct sigaction sa; + long (*fn)(void); + size_t setlen, n; + long ps =3D sysconf(_SC_PAGESIZE); + int rc =3D 0; + unsigned char *m =3D mmap(NULL, 2 * ps, PROT_READ | PROT_WRITE | PROT_= EXEC, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (m =3D=3D MAP_FAILED) { + perror("mmap"); + return 2; + } + + unsigned char *pb =3D m + ps; + + /* + * Page A ends with the store to the return value register, so that the + * next instruction executed is the first one on page B. + */ + setlen =3D emit_set_ret(tmp, 1); + memcpy(pb - setlen, tmp, setlen); + emit_ret(pb); + __builtin___clear_cache((char *)m, (char *)m + 2 * ps); + + page_b_entry =3D pb; + fn =3D (long (*)(void))(pb - setlen); + + for (int i =3D 0; i < 200000; i++) { + if (fn() !=3D 1) { + printf("FAIL: phase 1 wrong result\n"); + return 1; + } + } + printf("phase 1 ok (chained)\n"); + + if (bp_mode) { + /* + * The chain from page A to page B now exists. gdb puts a breakpo= int + * on page_b_entry here; the call below has to stop on it rather t= han + * jump over it. + */ + break_here(); + if (fn() !=3D 1) { + printf("FAIL: bp phase wrong result\n"); + return 1; + } + printf("bp phase ok\n"); + return 0; + } + + memset(&sa, 0, sizeof(sa)); + sa.sa_handler =3D segv; + sigemptyset(&sa.sa_mask); + if (sigaction(SIGSEGV, &sa, NULL) !=3D 0) { + perror("sigaction"); + return 2; + } + if (mprotect(pb, ps, PROT_NONE) !=3D 0) { + perror("mprotect"); + return 2; + } + if (sigsetjmp(jb, 1) =3D=3D 0) { + fn(); + printf("FAIL: phase 2 executed page B after mprotect(PROT_NONE)\n"= ); + rc =3D 1; + } else if (!caught) { + printf("FAIL: phase 2 longjmp without entering the handler\n"); + rc =3D 1; + } else { + printf("phase 2 ok (faulted)\n"); + } + + /* Phase 3: map back, overwrite, expect the new code to run. */ + if (mprotect(pb, ps, PROT_READ | PROT_WRITE | PROT_EXEC) !=3D 0) { + perror("mprotect back"); + return 2; + } + n =3D emit_set_ret(pb, 2); + emit_ret(pb + n); + __builtin___clear_cache((char *)pb, (char *)pb + ps); + + long r =3D fn(); + if (r !=3D 2) { + printf("FAIL: phase 3 returned %ld, expected 2 (stale chain)\n", r= ); + rc =3D 1; + } else { + printf("phase 3 ok (new code ran)\n"); + } + return rc; +#endif +} --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234633; cv=none; d=zohomail.com; s=zohoarc; b=c4QeP6MzvqT0VsttfFDymWkXdBoMzyPL5j3vPvki1TVXUBmzL0Vz53V0FMjKbYT3+qNAUvm7krHl2XlBYwSclV8c1mVZ31jjptSmCmt/0yizdNcY0PDg6o91MRQyNg0UuU43ph6COB0yRF4ihXrzP6iVD584qU7QDlZiBtQnXZY= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234633; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=2aNNw+LRwkAiqFyZ0GoG7IJ1ns7+1hc/AivPTPBuYsM=; b=Yl8Y5kcc306oZneATWneOwSrcAWrOGpgxX2jaMPLtsPM1idXAAzCCOyy7yKRBptuw+ykqdVaJaRaEYYkbbIg3D09tgSayehFa00V2Ft6Q7TpuckwW+vklWLYne/2fPVQM94qO3upzOQlVnKv3x/A4C9hREBNnwlV+bv4u5iP3eA= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234633405808.1982090446119; Mon, 31 Aug 2026 20:50:33 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUt-0000wF-Kp; Mon, 31 Aug 2026 23:49:19 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUf-0000eH-PW for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:09 -0400 Received: from mail-yx1-xb133.google.com ([2607:f8b0:4864:20::b133]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUc-0004Me-Dq for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:05 -0400 Received: by mail-yx1-xb133.google.com with SMTP id 956f58d0204a3-66b32bb75beso3903687d50.2 for ; Mon, 31 Aug 2026 20:49:01 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4ed2986csm7466463d50.18.2026.08.31.20.48.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:48:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234541; x=1788839341; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=2aNNw+LRwkAiqFyZ0GoG7IJ1ns7+1hc/AivPTPBuYsM=; b=iqwGQlJBdXwG6wT79b18Q8NmMPLIdyoV3EZb+lIZUZhKtmjEV9U/62EihvGdg70vOV UI37JJpdCohh+S1MeoB2iYPSscNSMZ1beMsIwnLCw36osGpzoCIq4/LHfMvRqEYaPsGc okAg4AUTgPqXyWZS5r5sg8gblCSPdypgX4QIoRc79Tw0DbKMZ2I3kVF1lFkds1WJqYfm gAoy15DhlUG+CmTmti/bSDFlk0idC6Nt/b82SuDlvYvOKeWp5Lrx8Be+GfIvq8drs6q9 TJnmUHrnaoCIKI0819YkmxSAkSTC09+Ft30WeTHcukh/IMcKNNasqWbZYUVlzjnfx3nY h8cg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234541; x=1788839341; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=2aNNw+LRwkAiqFyZ0GoG7IJ1ns7+1hc/AivPTPBuYsM=; b=LGcJX5I+td7iBxflOwPCA9w2Ox0VdXOT01RAKHuvhfH0zXV7wb4TL03Gk2Nn4ShYVG KqRiP/yRqubnljssDq5CHGjTA20yOPFdQFPwQm/8cyzW2EFZAGGxGQPs853iZ2sU0bO0 6Lg9s3Y/c/RsiCQM847dpkuLSyQqe2kAo1DNkkyiHYnWe18ST8kvZa5gNAI1pXsUhTzZ iARZEzlwGgBmy/adj+lv7A4AaUVLzbMDjci7aZDVix+8IX0/V0EIUV15Ka+ScVcNcfhM umsq1XgfI3lGGvEVlXLUgvLJ5GvzVjhQOampqoz9mVn2kAi2hqfvFRnCGMs7Dlvu4Y3T 9NLQ== X-Gm-Message-State: AFuF++l7XArTzn8GVTDBc9NIOrdbsmDyZuOhNsxpi6CnEzkQ3ohb90Q2 ZHsIsB6l846VY79JKozkRY9GBt2Ui8z3vxCrd479bjZqi97dRW0Ivhyq/+sk3A== X-Gm-Gg: AYBFou2qfCmLXvmTZ5HCPArvNyyiiOu8JiZiiK3J4Bxm9zecPmkMqPMgXZWj1m70LpE q2sxIDO6nB6D485xuI7jqYnYsfqaXTfJjLYEjtW5KcTSuQChknE7JXK31MVQFkvISet/C021R+u LcskpGIfjgHSZ3NJAuT7HaMHubQm/aYPRSQcqH+vcPT1A0pQatZrAUGQ526eM6BtbQTYcdLQ0EE DXdaceuoJzFeMPbW8zcGlFDCaF/SsJyUy/86EXYJfUJ9coQ6B4gDZLTC5+qRDsf0xeWldypj2q7 PPI2JndEvn1EkXTn3oR5IY7wGXYzHgUqkYMPAn+k5UVW0CDLHFiAlIEFgSVfQnPAe/yMU2VyqL2 MT/McMmH50SDx0nwm8z9baS6F8mt0VRLtgDqS/l7vd2GhtQPoM15JQaETDDCQKE4tWTHGqDDGYI Dq8Plg23Q/AkRlCefgJ/SsSBps1czeXsgsFGt1iMlrTvkyY6lav1b481hTnkubyZ6f9L8H4reaI Wn/mHv392m+nMtIy3f1rew3dnAZ+AbY9RC7gCQc X-Received: by 2002:a05:690e:4501:10b0:66e:54a0:cdd3 with SMTP id 956f58d0204a3-66e54a0e97amr5804793d50.42.1788234540725; Mon, 31 Aug 2026 20:49:00 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 8/9] RFC: accel/tcg: poison the jump cache instead of polling for indirect exits Date: Mon, 31 Aug 2026 23:48:07 -0400 Message-ID: <20260901034808.3524945-9-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::b133; envelope-from=mattst88@gmail.com; helo=mail-yx1-xb133.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234634481158500 Content-Type: text/plain; charset="utf-8" Every translation block begins by loading cpu->neg.icount_decr.u32, testing it and branching to the exit path. That is three host instructions at the t= op of every TB, and blocks are short: an emulated alpha gcc 16.2.0 compiling a 255k line translation unit executes 34.2 billion of them at 6.04 guest instructions each. A block does not need to poll if every way out of it already reaches a chec= k. A goto_tb does not: it chains straight into its destination, with nothing in between that looks at icount_decr, so the destination has to poll on entry. An indirect exit does. The out-of-line path calls helper_lookup_tb_ptr() every time, so it only needs the helper to return the epilogue while an exit is pending. The inline probe needs a way to be told, so give it one: CPUState::tb_jmp_cache_probe, the base pointer it reads. Normally that is cpu->tb_jmp_cache; pointed at a shared page of zeroes instead, every entry the probe finds has a NULL tb, every dispatch misses, and a miss lands in t= he same helper. The real jump cache is untouched, so no cache contents are los= t, and the fast path pays nothing: the base was a load from CPUState either wa= y. The two places that set icount_decr.u16.high poison the probe; the main loop puts it back once cpu_handle_interrupt() has cleared the reason. The poison is a single read-only mapping shared by every CPU, because nothing may ever write to it and a stray store into a page every vCPU dispatches through is worth trapping rather than debugging. The poll is therefore emitted only in blocks that emit a goto_tb. Whether a block does is not known until its last exit has been generated, so the decision is deferred and the load and branch are emitted retroactively at t= he head of the block in gen_tb_end(), using the same emit_before_op mechanism the can_do_io stores use. icount opts out and keeps the counter unconditionally. Interrupt latency is bounded at one block, as before. It does not depend on the shape of the guest's control flow graph: a block either polls on entry = or is checked on the way out, and no run of blocks can avoid both. What changes is where the check sits, not how often one happens. tests/tcg/multiarch/test-indirect-irq.c is added for this: a loop whose only back edge is an indirect branch, under alarm(1). That loop's block emits no goto_tb, so it no longer polls, and the test passes only because the dispat= ch notices instead -- it hangs if the poison is removed, which is what makes i= t a test of the new mechanism rather than of the old poll. Nothing in it is architecture specific: the loop is a computed goto, which every target's compiler supports, so it covers whichever targets go on to use the inline probe. The other alpha tests still pass and the emulated compiler still produces byte-identical output. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, = on top of the preceding patches: before: 891,254,240,071 instructions after: 868,811,832,620 instructions -2.52% before: 81.45s wall clock after: 79.85s wall clock -1.96% The emulated compiler produces byte-identical output. RFC because: - The un-poison in the main loop races a concurrent poison from another thread. The existing barrier around icount_decr.u16.high covers it -- a poison that lands after the sync also re-set the flag, and exit_request w= as stored before it -- but this deserves more eyes than the single-threaded user-mode testing I have given it. - Only the inline probe needs the poison, and only alpha uses the inline probe today. Targets on the out-of-line path are covered by the helper check alone, but that has not been measured. v3: Rebased onto the removal of "only poll for interrupts in blocks that can close a cycle", which v2 sat on top of and which is dropped: it let a straight-line run of arbitrary length go unchecked, since a block with = no backward edge polled nowhere (Richard). The rule is now that a block polls iff it emits a goto_tb, rather than iff it can close a control flow cycle. That keeps the bound at one block without any analysis of the guest's control flow graph, so the objection to the dropped patch does not carry over. The deferred-emission machine= ry it needs moves here from that patch; DisasContextBase::needs_exit_check and the hook in translator_use_goto_tb() are gone with it, and the flag is now set by tcg_gen_goto_tb() rather than by goto_ptr emission. All of v2's measurements were dropped: they were taken with the cycle-analysis patch underneath, which changes both the baseline and what is left to remove, so none of them described this patch. The numbers above are a fresh measurement of the series as it now stands. v4: Moved the test from tests/tcg/alpha/ to tests/tcg/multiarch/: the mechanism is generic and nothing in the test is alpha specific (Alex). The performance numbers above are the v3 measurements, not re-run: the machine they were taken on is busy. v5: CPUState::tb_jmp_cache_probe moves here from what was patch 5, which used it for breakpoints too. Breakpoints are now a cflag, so a pending exit is the only reason left to poison, and the machinery shrinks to match: no NULL states to handle, no cross-thread poison from cpu_breakpoint_insert(), and one condition rather than two. v5: Map the poison read-only rather than leaving it a writable .bss object. Requested by Richard Henderson. It costs a page-aligned 1MB allocation at startup instead of nothing on disk, which the enforcement is worth. qemu_mprotect_ro() is added for it, alongside the _rw, _rwx and _none forms already there. v5: Drop the NULL checks in the poison and sync helpers. Requested by Richard Henderson: the sync is only ever called by the main loop, so it cannot see an unrealized CPU, and unrealize now leaves the probe pointi= ng at the poison rather than at NULL, so neither has an unrealized state to consider. Signed-off-by: Matt Turner --- accel/tcg/cpu-exec.c | 91 +++++++++++++++++++++++++ accel/tcg/internal-common.h | 9 +++ accel/tcg/tcg-accel-ops.c | 2 + accel/tcg/translator.c | 51 +++++++++++++- include/hw/core/cpu.h | 9 +++ include/qemu/mprotect.h | 1 + include/tcg/tcg.h | 2 + tcg/tcg-op.c | 21 +++++- tests/tcg/multiarch/test-indirect-irq.c | 62 +++++++++++++++++ util/osdep.c | 9 +++ 10 files changed, 253 insertions(+), 4 deletions(-) create mode 100644 tests/tcg/multiarch/test-indirect-irq.c diff --git ./accel/tcg/cpu-exec.c ./accel/tcg/cpu-exec.c index ca90a77a7b..7a371e6928 100644 --- ./accel/tcg/cpu-exec.c +++ ./accel/tcg/cpu-exec.c @@ -19,6 +19,9 @@ =20 #include "qemu/osdep.h" #include "qemu/qemu-print.h" +#include "qemu/error-report.h" +#include "qemu/memalign.h" +#include "qemu/mprotect.h" #include "qapi/error.h" #include "qapi/type-helpers.h" #include "hw/core/cpu.h" @@ -388,6 +391,16 @@ const void *HELPER(lookup_tb_ptr)(CPUArchState *env) */ cpu->neg.can_do_io =3D true; =20 + /* + * A block that dispatches indirectly does not emit the icount_decr po= ll, + * so this is where a pending exit is noticed for that path: either the + * probe was poisoned and every dispatch arrives here, or the target u= ses + * the out-of-line lookup and always did. + */ + if (unlikely(cpu_loop_exit_requested(cpu))) { + return tcg_code_gen_epilogue; + } + TCGTBCPUState s =3D cpu->cc->tcg_ops->get_tb_cpu_state(cpu); s.cflags =3D curr_cflags(cpu); =20 @@ -779,6 +792,70 @@ static inline bool cpu_handle_exception(CPUState *cpu,= int *ret) return false; } =20 +/* + * The inline jump cache probe reads cpu->tb_jmp_cache_probe and takes the + * slow path when the entry it finds has a NULL tb. Pointing the probe at= a + * region that is all zeroes therefore forces every indirect dispatch into + * helper_lookup_tb_ptr(), which returns the epilogue while an exit is + * pending. The real jump cache is untouched, so no contents are lost and + * recovery is a single store. + * + * Only ever read from, and only one entry per dispatch, so one shared + * zero-filled cache is enough for every CPU. Mapped read-only, since + * nothing may write to it and a stray store into a shared page every vCPU + * dispatches through is worth trapping rather than debugging. + */ +static CPUJumpCache *tb_jmp_cache_poison; + +static void tb_jmp_cache_poison_init(void) +{ + size_t align =3D qemu_real_host_page_size(); + size_t size =3D ROUND_UP(sizeof(CPUJumpCache), align); + void *p =3D qemu_memalign(align, size); + + memset(p, 0, size); + if (qemu_mprotect_ro(p, size) < 0) { + /* Only the enforcement is lost; the zeroes are what matter. */ + warn_report("could not write-protect the jump cache poison"); + } + tb_jmp_cache_poison =3D p; +} + +/* + * Poison @cpu's probe, from any thread. A plain store is enough: the val= ue + * only ever costs a slow path that is correct on its own, and the generat= ed + * code re-reads the base on every dispatch. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu) +{ + qatomic_set(&cpu->tb_jmp_cache_probe, tb_jmp_cache_poison); +} + +/* + * Called from @cpu's own main loop, which is the only context that can + * establish that no reason to be poisoned is left. + */ +void tcg_cpu_sync_jmp_cache(CPUState *cpu) +{ + if (qatomic_read(&cpu->tb_jmp_cache_probe) =3D=3D cpu->tb_jmp_cache) { + return; + } + + qatomic_set(&cpu->tb_jmp_cache_probe, cpu->tb_jmp_cache); + + /* + * Another thread may have set icount_decr.u16.high after the caller + * decided no exit was pending, and its poison may have landed before + * the store above. Order that store against the re-read, so the race + * is lost in the safe direction: an exit that is still pending here + * poisons again, and the dispatch after it returns to the main loop. + */ + smp_mb(); + if (unlikely(cpu_loop_exit_requested(cpu))) { + tcg_cpu_poison_jmp_cache(cpu); + } +} + void tcg_kick_vcpu_thread(CPUState *cpu) { /* @@ -791,6 +868,9 @@ void tcg_kick_vcpu_thread(CPUState *cpu) =20 /* Ensure cpu_exec will see the exit request after TCG has exited. */ qatomic_store_release(&cpu->neg.icount_decr.u16.high, -1); + + /* Blocks that only dispatch indirectly do not poll; stop them chainin= g. */ + tcg_cpu_poison_jmp_cache(cpu); } =20 static inline bool icount_exit_request(CPUState *cpu) @@ -991,6 +1071,13 @@ cpu_exec_loop(CPUState *cpu, SyncClocks *sc) break; } =20 + /* + * cpu_handle_interrupt() has just cleared everything that wou= ld + * make a dispatch have to come back here, so this is where the + * probe is allowed to return after a poison. + */ + tcg_cpu_sync_jmp_cache(cpu); + tb =3D tb_lookup(cpu, s); if (tb =3D=3D NULL) { CPUJumpCache *jc; @@ -1092,6 +1179,7 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) assert(tcg_ops->get_tb_cpu_state); assert(tcg_ops->mmu_index); tcg_ops->initialize(); + tb_jmp_cache_poison_init(); tcg_target_initialized =3D true; } =20 @@ -1099,6 +1187,7 @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp) tcg_update_cflags(cpu); =20 cpu->tb_jmp_cache =3D g_new0(CPUJumpCache, 1); + qatomic_set(&cpu->tb_jmp_cache_probe, cpu->tb_jmp_cache); tlb_init(cpu); #ifndef CONFIG_USER_ONLY tcg_iommu_init_notifier_list(cpu); @@ -1116,5 +1205,7 @@ void tcg_exec_unrealizefn(CPUState *cpu) #endif /* !CONFIG_USER_ONLY */ =20 tlb_destroy(cpu); + /* Not NULL: nothing then has to special-case an unrealized CPU. */ + tcg_cpu_poison_jmp_cache(cpu); g_free_rcu(cpu->tb_jmp_cache, rcu); } diff --git ./accel/tcg/internal-common.h ./accel/tcg/internal-common.h index 853d1b51ee..6faa039850 100644 --- ./accel/tcg/internal-common.h +++ ./accel/tcg/internal-common.h @@ -144,6 +144,15 @@ void page_table_config_init(void); G_NORETURN void cpu_io_recompile(CPUState *cpu, uintptr_t retaddr); #endif /* CONFIG_USER_ONLY */ =20 +/* + * Force @cpu's generated code back into helper_lookup_tb_ptr(), which + * re-checks everything the inline jump cache probe cannot. Safe to call + * from any thread. tcg_cpu_sync_jmp_cache() undoes it, and is for the + * owning CPU's main loop only. + */ +void tcg_cpu_poison_jmp_cache(CPUState *cpu); +void tcg_cpu_sync_jmp_cache(CPUState *cpu); + void tb_phys_invalidate(TranslationBlock *tb, tb_page_addr_t page_addr); void tb_set_jmp_target(TranslationBlock *tb, int n, uintptr_t addr); =20 diff --git ./accel/tcg/tcg-accel-ops.c ./accel/tcg/tcg-accel-ops.c index 560fe2554b..63a15f1689 100644 --- ./accel/tcg/tcg-accel-ops.c +++ ./accel/tcg/tcg-accel-ops.c @@ -38,6 +38,7 @@ #include "exec/cputlb.h" #include "exec/hwaddr.h" #include "exec/tb-flush.h" +#include "internal-common.h" #include "exec/translation-block.h" #include "exec/watchpoint.h" #include "gdbstub/enums.h" @@ -106,6 +107,7 @@ void tcg_handle_interrupt(CPUState *cpu, int mask) qemu_cpu_kick(cpu); } else { qatomic_set(&cpu->neg.icount_decr.u16.high, -1); + tcg_cpu_poison_jmp_cache(cpu); } } =20 diff --git ./accel/tcg/translator.c ./accel/tcg/translator.c index 8879cd626f..89d255bd04 100644 --- ./accel/tcg/translator.c +++ ./accel/tcg/translator.c @@ -45,12 +45,35 @@ bool translator_io_start(DisasContextBase *db) return true; } =20 +/* + * A block that ends in a goto_tb chains straight to its destination: noth= ing + * between the two looks at icount_decr, so the destination has to poll on + * entry. A block whose exits are all indirect does not, because the disp= atch + * itself notices -- a pending exit poisons tb_jmp_cache_probe, so the pro= be + * misses into helper_lookup_tb_ptr(), which returns the epilogue. Every = block + * therefore either polls on entry or is checked as it leaves, which bounds + * interrupt latency at one block without looking at the shape of the gues= t's + * control flow graph. + * + * Which kind a block is is not known until its last exit has been emitted= , so + * defer the decision to gen_tb_end() and emit the poll retroactively. + * + * icount needs the counter unconditionally, so it opts out. + */ +static bool defer_exit_check(uint32_t cflags) +{ + return !(cflags & CF_USE_ICOUNT); +} + static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t cflags) { TCGv_i32 count =3D NULL; TCGOp *icount_start_insn =3D NULL; =20 - if ((cflags & CF_USE_ICOUNT) || !(cflags & CF_NOIRQ)) { + tcg_ctx->exit_check_needed =3D false; + + if ((cflags & CF_USE_ICOUNT) || + (!(cflags & CF_NOIRQ) && !defer_exit_check(cflags))) { count =3D tcg_temp_new_i32(); tcg_gen_ld_i32(count, tcg_env, offsetof(CPUState, neg.icount_decr.u32) - @@ -76,6 +99,9 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_t= cflags) */ if (cflags & CF_NOIRQ) { tcg_ctx->exitreq_label =3D NULL; + } else if (defer_exit_check(cflags)) { + /* Emitted retroactively by gen_tb_end(), if this TB emits a goto_= tb. */ + tcg_ctx->exitreq_label =3D gen_new_label(); } else { tcg_ctx->exitreq_label =3D gen_new_label(); tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_label); @@ -91,7 +117,8 @@ static TCGOp *gen_tb_start(DisasContextBase *db, uint32_= t cflags) } =20 static void gen_tb_end(const TranslationBlock *tb, uint32_t cflags, - TCGOp *icount_start_insn, int num_insns) + TCGOp *icount_start_insn, int num_insns, + TCGOp *first_insn_start) { if (cflags & CF_USE_ICOUNT) { /* @@ -102,6 +129,23 @@ static void gen_tb_end(const TranslationBlock *tb, uin= t32_t cflags, tcgv_i32_arg(tcg_constant_i32(num_insns))); } =20 + if (tcg_ctx->exitreq_label && defer_exit_check(cflags) && + !(cflags & CF_NOIRQ)) { + if (tcg_ctx->exit_check_needed) { + TCGv_i32 count =3D tcg_temp_new_i32(); + TCGOp *save =3D tcg_ctx->emit_before_op; + + tcg_ctx->emit_before_op =3D first_insn_start; + tcg_gen_ld_i32(count, tcg_env, + offsetof(CPUState, neg.icount_decr.u32) - + sizeof(CPUState)); + tcg_gen_brcondi_i32(TCG_COND_LT, count, 0, tcg_ctx->exitreq_la= bel); + tcg_ctx->emit_before_op =3D save; + } else { + tcg_ctx->exitreq_label =3D NULL; + } + } + if (tcg_ctx->exitreq_label) { gen_set_label(tcg_ctx->exitreq_label); tcg_gen_exit_tb(tb, TB_EXIT_REQUESTED); @@ -238,7 +282,8 @@ void translator_loop(CPUState *cpu, TranslationBlock *t= b, int *max_insns, =20 /* Emit code to exit the TB, as indicated by db->is_jmp. */ ops->tb_stop(db, cpu); - gen_tb_end(tb, cflags, icount_start_insn, db->num_insns); + gen_tb_end(tb, cflags, icount_start_insn, db->num_insns, + first_insn_start); =20 /* * Manage can_do_io for the translation block: set to false before diff --git ./include/hw/core/cpu.h ./include/hw/core/cpu.h index 81af7b9ee1..c8669f2cad 100644 --- ./include/hw/core/cpu.h +++ ./include/hw/core/cpu.h @@ -519,6 +519,15 @@ struct CPUState { MemoryRegion *memory; =20 struct CPUJumpCache *tb_jmp_cache; + /* + * @tb_jmp_cache_probe: the base the inline jump cache probe reads. + * + * Normally @tb_jmp_cache. Pointed at a shared read-only page of zero= es + * while an exit is pending, so that every inline dispatch misses and + * falls back to helper_lookup_tb_ptr(), which returns to the main loo= p. + * Only generated code and the accessors in cpu-exec.c may touch it. + */ + struct CPUJumpCache *tb_jmp_cache_probe; =20 GArray *gdb_regs; int gdb_num_regs; diff --git ./include/qemu/mprotect.h ./include/qemu/mprotect.h index 1e83d1433e..4fc13d79f6 100644 --- ./include/qemu/mprotect.h +++ ./include/qemu/mprotect.h @@ -8,6 +8,7 @@ #define QEMU_MPROTECT_H =20 int qemu_mprotect_rw(void *addr, size_t size); +int qemu_mprotect_ro(void *addr, size_t size); int qemu_mprotect_rwx(void *addr, size_t size); int qemu_mprotect_none(void *addr, size_t size); =20 diff --git ./include/tcg/tcg.h ./include/tcg/tcg.h index 7669dc1c2d..df08c10544 100644 --- ./include/tcg/tcg.h +++ ./include/tcg/tcg.h @@ -389,6 +389,8 @@ struct TCGContext { struct TCGLabelPoolData *pool_labels; =20 TCGLabel *exitreq_label; + /* Set by goto_tb emission: this TB chains without reaching a check. */ + bool exit_check_needed; =20 #ifdef CONFIG_PLUGIN /* diff --git ./tcg/tcg-op.c ./tcg/tcg-op.c index b10b2d66d5..1c5c5ec1d3 100644 --- ./tcg/tcg-op.c +++ ./tcg/tcg-op.c @@ -2713,6 +2713,13 @@ void tcg_gen_goto_tb(unsigned idx) tcg_debug_assert((tcg_ctx->goto_tb_issue_mask & (1 << idx)) =3D=3D 0); tcg_ctx->goto_tb_issue_mask |=3D 1 << idx; #endif + /* + * A goto_tb chains straight into the destination, with nothing in bet= ween + * that looks at icount_decr, so this TB has to poll on entry. See + * defer_exit_check(). + */ + tcg_ctx->exit_check_needed =3D true; + plugin_gen_disable_mem_helpers(); tcg_gen_op1i(INDEX_op_goto_tb, 0, idx); } @@ -2790,8 +2797,13 @@ static void gen_jmp_cache_probe(TCGv_i64 pc, const T= ranslationBlock *tb) gen_jmp_cache_hash(h, pc); tcg_gen_shli_i64(h, h, 4); =20 + /* + * Not cpu->tb_jmp_cache: the probe reads its own base so that the main + * loop can poison it, which is how a pending exit forces every dispat= ch + * back into the helper. See tcg_cpu_sync_jmp_cache(). + */ tcg_gen_ld_ptr(jc, tcg_env, - offsetof(CPUState, tb_jmp_cache) - sizeof(CPUState)); + offsetof(CPUState, tb_jmp_cache_probe) - sizeof(CPUStat= e)); tcg_gen_trunc_i64_ptr(ent, h); tcg_gen_add_ptr(ent, jc, ent); =20 @@ -2849,6 +2861,13 @@ static void gen_goto_jc(TCGv_i64 pc) =20 plugin_gen_disable_mem_helpers(); =20 + /* + * Neither path below needs an icount_decr poll. The helper returns to + * the main loop while an exit is pending, and a pending exit poisons + * tb_jmp_cache_probe, so the inline probe finds a NULL tb and falls i= nto + * that same helper. + */ + #ifdef CONFIG_DEBUG_TCG /* * The caller has asserted that env already describes the destination. diff --git ./tests/tcg/multiarch/test-indirect-irq.c ./tests/tcg/multiarch/= test-indirect-irq.c new file mode 100644 index 0000000000..a672faf641 --- /dev/null +++ ./tests/tcg/multiarch/test-indirect-irq.c @@ -0,0 +1,62 @@ +/* + * A loop whose only back edge is an indirect branch must still be + * interruptible. + * + * Blocks that dispatch indirectly do not emit the icount_decr poll; a pen= ding + * exit instead poisons the inline jump cache probe so that the dispatch f= alls + * into helper_lookup_tb_ptr(), which returns to the main loop. If that + * mechanism breaks, this program never leaves the loop and the test times + * out rather than failing an assertion. + * + * A computed goto is used deliberately: a plain while(1) would end the bl= ock + * with a direct backward branch, that is a goto_tb, and a block that emit= s a + * goto_tb still polls -- so it would not exercise the path under test. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ +#include +#include +#include +#include +#include +#include + +/* Written by the handler, read by the loop, so it must not be cached. */ +static volatile sig_atomic_t fired; +/* Read after the loop, so the loop must not optimize the increment away. = */ +static volatile unsigned long iterations; + +static void handler(int sig) +{ + fired =3D 1; +} + +int main(void) +{ + /* + * Indexing a table with a volatile index, rather than jumping through= a + * volatile pointer: gcc happily proves a single-valued pointer consta= nt + * and emits a direct branch, which is the case this test is not about. + */ + volatile int idx =3D 0; + void *target[2]; + struct sigaction sa; + + memset(&sa, 0, sizeof(sa)); + sa.sa_handler =3D handler; + sigemptyset(&sa.sa_mask); + assert(sigaction(SIGALRM, &sa, NULL) =3D=3D 0); + alarm(1); + + target[0] =3D &&spin; + target[1] =3D &&out; +spin: + iterations++; + if (!fired) { + goto *target[idx]; + } +out: + + printf("interrupted after %lu iterations\n", iterations); + return 0; +} diff --git ./util/osdep.c ./util/osdep.c index 4a8b8b5a90..9c72a42b4f 100644 --- ./util/osdep.c +++ ./util/osdep.c @@ -99,6 +99,15 @@ int qemu_mprotect_rw(void *addr, size_t size) #endif } =20 +int qemu_mprotect_ro(void *addr, size_t size) +{ +#ifdef _WIN32 + return qemu_mprotect__osdep(addr, size, PAGE_READONLY); +#else + return qemu_mprotect__osdep(addr, size, PROT_READ); +#endif +} + int qemu_mprotect_rwx(void *addr, size_t size) { #ifdef _WIN32 --=20 2.54.0 From nobody Sat Sep 26 20:50:25 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1788234568; cv=none; d=zohomail.com; s=zohoarc; b=jHGU+RYNe6eY/LpXJGSc2b5NV1TkRQKqLM8P9qpLSTKEWDOTimNQQR1VdTs7BlCjSUeW7NFPkCn08GqbgnwefTSDo/ywnNl2SMCHS/eGnC369z9RQtjrCdaNXmtJRa3UQY8FE0nIjxytAOuUpteGBMG39nvTH+XyRXsHurgi9j0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1788234568; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=DkbGiT0O49wAsgajIdIH3GDtMj4n9sAaajbSP09NK3UFNEt6I/1r3+zWXeeVO0f2HFcanBIwa7n/cfs/+9fxG/vA7mTFHT8Jw1XGTQEux+W7CY3wvCAwgh+zDE82TL8OUfczB5VXFU1ZQIy68l5bLVCozQUVc6hNBrnt1pk2Hmw= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1788234568225637.8132958414839; Mon, 31 Aug 2026 20:49:28 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x1FUr-0000h0-TW; Mon, 31 Aug 2026 23:49:17 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x1FUj-0000fG-JI for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:09 -0400 Received: from mail-yw1-x112b.google.com ([2607:f8b0:4864:20::112b]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1x1FUf-0004ND-D3 for qemu-devel@nongnu.org; Mon, 31 Aug 2026 23:49:07 -0400 Received: by mail-yw1-x112b.google.com with SMTP id 00721157ae682-836cb2fa1bcso40429687b3.1 for ; Mon, 31 Aug 2026 20:49:05 -0700 (PDT) Received: from localhost (107-220-129-194.lightspeed.chrlnc.sbcglobal.net. [107.220.129.194]) by smtp.gmail.com with ESMTPSA id 00721157ae682-85e5ed18556sm66125667b3.11.2026.08.31.20.49.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 31 Aug 2026 20:49:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788234544; x=1788839344; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=h0Zbi1Xhmejw3sfocQlwLAXI2bDUAx30CmvNNJKtBrU2cFwbVpejAjoo9MbTcTIBjL W1/eEYCzX0x+7GBHcysFqb9mxBhdif5R4E8RJf6znOCrvEZOmoi+UJjn1ujdjXTgHRau +6wBuM5P5joVWYPw85tgxpYg3bZK7WvvhTasTdwwX8GLBaoxeqbHwKj+DL8DaRwUHB3M 9uWpLp2RCZqL+F43g/CO4yAsryz+92vpAeUrm70AlxRfG0h6Xf7Wfmu0XoIc6Lnj7bKs Y6H6eLyyzRuUqqugZ1vvS2ED8TUHF6SbnylPORqNc9M+cik3vWnVVpnmGKxWlN8glciG CA5A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788234544; x=1788839344; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=abTvdvVoaUPAt7wEkDm5qW6FwkZ8iBbbJSMuWBq+oyc=; b=mh9XqAAavEJl9APo4U1IFrtXPfeRJoz3teS0O4ImpxIx94fjK1pgT25vrQs3xfuBS4 0BBfLQYB/o4R8MqXd1hJfmf0aCROlwh0pw6vsqPh+WcaSd+aAt1Y38c9AU1B1PRI+VCU NPOQNx06hYDy1snGAIENTu0dj2BbPjAE/SYsCHG1qOBaWnvEsiVPF97KWb+xYQfXihl/ hAR7SVmcgRPKjV9zz+F0rlc/kEq4X0e8B2l7rxLFTdPMaf0JWRHXu9tkSiXhol6j9FH0 GbaubcKa79zdxDxwGT14lzQQiQu3zXvfBhz+o+T/r7+c1nlXGmAEu/ENS2zTbXw+n6TC p9XA== X-Gm-Message-State: AFuF++ks4Af4TUlSCoSXMquRYHE1YDKgHQn/vCIwMVaNAm78Czxa1aAW 3rja4/vJ2BKiK8CZUxWdkFLbdFW6UnOeyXdwJcPyZJm49ybnCwrZp1Qgg4hAIw== X-Gm-Gg: AYBFou0dmlpzwwMVBFN3+tX/mvrnaA4o5GYDwNgYsESuYIUuYOOvk49b1jaNSsOgwkb FgVIR+rx89IixyeO/wokV8xyRBdgj2ixT2NMYsPU4OVHMRn+tXU8gPN8NK9PkkuWPJCsC9CrfuJ RpgyaoIo2bdIP20u23CppO74YyBltS9Bakub+VSN0YZCoEnS0jiT78Z1oOGl8TbtyD/rgiR38QQ uysXqNQ5CD1T2ePgHXVW3GuDJGvB5Qpn83aseGRWSLoDzvgw/pj9hKBZEJtlOcBsD51iNDSAhLg PL1lt5TPgn47Ff3l7Fpz22KpxzMt9EIQM/woxURIM+t36ipFiP1zfi4vNxpPEkOyNuR5uSMk4Ag mxYSQuMaveJBly0stCCYxBzeLD98ZEU4Z4yBJKITFyb1oMQAg8IcYPP+zVKX4hi5DG5ZIvtKn4T OF+bFnvPGFVgGcfcsv32JjlJxrqFYow3izQjpz+lrhMlXLjsAdcLF+/Hrnt9DmmxCYuXck6V9dB zHRAEFhQiZE8iH+SrPwmU0Y/jdLQJYn2tYMEXZm X-Received: by 2002:a05:690c:c623:b0:809:9422:8c47 with SMTP id 00721157ae682-85d6d889183mr99833927b3.22.1788234544140; Mon, 31 Aug 2026 20:49:04 -0700 (PDT) From: Matt Turner To: qemu-devel@nongnu.org Cc: richard.henderson@linaro.org, pbonzini@redhat.com, philmd@oss.qualcomm.com, alex.bennee@linaro.org, zhao1.liu@intel.com, Matt Turner Subject: [PATCH v5 9/9] RFC: tcg: fold a guest displacement into the host addressing mode Date: Mon, 31 Aug 2026 23:48:08 -0400 Message-ID: <20260901034808.3524945-10-mattst88@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260901034808.3524945-1-mattst88@gmail.com> References: <20260827050241.3713332-1-mattst88@gmail.com> <20260901034808.3524945-1-mattst88@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::112b; envelope-from=mattst88@gmail.com; helo=mail-yw1-x112b.google.com X-Spam_score_int: -17 X-Spam_score: -1.8 X-Spam_bar: - X-Spam_report: (-1.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_ENVFROM_END_DIGIT=0.25, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1788234570092158500 Content-Type: text/plain; charset="utf-8" Nothing in the TCG frontend interface can express a based memory access. tcg_gen_qemu_ld/st take an address and nothing else, so a target with a displacement in its load and store encodings -- which is most of them -- has to materialize the address first: ldq a1,8(a0) -> mov 0x80(%rbp),%rbx reload a0 lea 0x8(%rbx),%r12 address mov (%r12),%r12 the load mov %r12,0x88(%rbp) spill a1 The lea is pure loss on a host whose addressing mode has a displacement field sitting empty. It also needs a register, at the point in a block where pressure is highest. Fold it. After optimization, look for an add of a constant immediately before a guest access, defining that access's address operand, and move the constant into a new second constant argument on the op. The add is left for liveness to remove, so nothing breaks if its result has another use. Only the immediately preceding op is examined: that is what the frontends emit, and a window of one op means the pass does not have to reason about what could have happened in between. The one thing it does check is that the add did not clobber the base it read, since the access now reads that base directly. Targets opt in with TCG_TARGET_HAS_ldst_disp and an out_disp member on TCGOutOpQemuLdSt. Without it the pass does not run, the displacement stays zero and the existing out member is called exactly as before, so no other backend changes behavior or needs touching. The fold is refused unless the access has no slow path at all, since the slow path hands addr_reg to the helper and that register no longer holds the full guest address. That is decided generically: user-only, because softmmu compares the unadjusted address against the TLB; a 64-bit address type, because a 32-bit one wraps where a host displacement would not; and no alignment test on the access. For x86_64 the displacement goes in the disp32 that prepare_host_addr() already fills in for guest_base, so all the backend has left to check is that guest_base plus the displacement still fits there. Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host, LTO build, on top of the preceding patches, against a control measured in the same session: before: 868,811,832,620 instructions, 79.85s after: 819,262,147,022 instructions, 77.30s -5.70% instructions, -3.20% wall Emitted code shrinks from 50.55MB to 48.80MB over the run, 167.4 to 161.6 bytes per block. Per Alpha opcode, the host bytes emitted for an access fall as expected and nothing else moves: ldq 18.3 -> 15.4 ldah 20.9 -> 20.9 ldl 16.6 -> 14.1 lda 12.9 -> 12.9 stq 12.8 -> 9.7 mov 9.8 -> 9.8 The emulated compiler produces byte-identical output and the alpha tests still pass, including with a non-zero guest_base forced via -B. RFC because: - Only wired up for x86_64, and only for qemu_ld and qemu_st; the i128 qemu_ld2 and qemu_st2 pairs are left alone. - Requiring that no slow path exists is stricter than necessary. The fast path test can stay on the base register as long as the displacement is itself a multiple of the required alignment, which it is for anything a frontend emits for a struct or stack access. Recording the displacement in TCGLabelQemuLdst and emitting one lea on the slow path would then cover alignment-checked accesses too, at no fast path cost. - Softmmu wants the displacement folded into the TLB comparison as well, which is a bigger change than this one. - A one op window catches everything the frontends emit today but is trivially defeated by anything scheduled in between. v4: - Hoisted the compilation mode tests -- tcg_use_softmmu and the 64-bit address type -- out of the backend hook and into fold_ldst_disp(), next to the TCG_TARGET_HAS_ldst_disp test, so the loop is not entered at all when the mode rules the fold out. - Pass MemOp rather than MemOpIdx to the backend hook; nothing about the mmu_idx is relevant to it. - Moved the alignment test into generic code as ldst_disp_needs_align(), so a backend does not have to repeat the atom_and_align_for_opc() call. The exact answer depends on the host's atomicity capabilities, which the generic pass does not know, so it answers for the most restrictive host. That is the same answer for everything the frontends actually emit -- MO_ATOM_IFALIGN is the default -- and conservative for the handful of MO_ATOM_WITHIN16 and MO_ATOM_SUBALIGN accesses, which lose the fold on a host that could have taken it. - What is left of the x86_64 hook is the guest_base test, so it now lives beside x86_guest_base under the CONFIG_USER_ONLY that declares it. - Refuse a displacement that does not fit in an int32_t, which is what out_disp() takes. Not reachable with any real guest_base, but the pass should not offer the backend something the interface cannot carry. - The numbers above are unchanged from v3: they have not been re-measured on the restructured patch, which is not expected to move them since the accesses in this workload are all MO_ATOM_IFALIGN. Signed-off-by: Matt Turner --- include/tcg/tcg-opc.h | 9 ++- tcg/tcg-op-ldst.c | 3 +- tcg/tcg.c | 132 +++++++++++++++++++++++++++++++++++- tcg/x86_64/tcg-target.c.inc | 41 +++++++++++ tcg/x86_64/tcg-target.h | 3 + 5 files changed, 184 insertions(+), 4 deletions(-) diff --git ./include/tcg/tcg-opc.h ./include/tcg/tcg-opc.h index f3a81d5d7f..92fd34d3e3 100644 --- ./include/tcg/tcg-opc.h +++ ./include/tcg/tcg-opc.h @@ -125,8 +125,13 @@ DEF(goto_ptr, 0, 1, 0, TCG_OPF_BB_EXIT | TCG_OPF_BB_EN= D) DEF(plugin_cb, 0, 0, 1, TCG_OPF_NOT_PRESENT) DEF(plugin_mem_cb, 0, 1, 1, TCG_OPF_NOT_PRESENT) =20 -DEF(qemu_ld, 1, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) -DEF(qemu_st, 0, 2, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +/* + * The second constant argument is a displacement to add to the address, + * zero unless a target advertises TCG_TARGET_HAS_ldst_disp and the fold in + * fold_ldst_disp() applied. + */ +DEF(qemu_ld, 1, 1, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) +DEF(qemu_st, 0, 2, 2, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_OP= F_INT) DEF(qemu_ld2, 2, 1, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) DEF(qemu_st2, 0, 3, 1, TCG_OPF_CALL_CLOBBER | TCG_OPF_SIDE_EFFECTS | TCG_O= PF_INT) =20 diff --git ./tcg/tcg-op-ldst.c ./tcg/tcg-op-ldst.c index 22211ccb45..ffc5e651a6 100644 --- ./tcg/tcg-op-ldst.c +++ ./tcg/tcg-op-ldst.c @@ -92,7 +92,8 @@ static MemOp tcg_canonicalize_memop(MemOp op, bool is64, = bool st) static void gen_ldst1(TCGOpcode opc, TCGType type, TCGTemp *v, TCGTemp *addr, MemOpIdx oi) { - TCGOp *op =3D tcg_gen_op3(opc, type, temp_arg(v), temp_arg(addr), oi); + /* The trailing zero is the address displacement; see fold_ldst_disp()= . */ + TCGOp *op =3D tcg_gen_op4(opc, type, temp_arg(v), temp_arg(addr), oi, = 0); TCGOP_FLAGS(op) =3D get_memop(oi) & MO_SIZE; } =20 diff --git ./tcg/tcg.c ./tcg/tcg.c index 489df0e738..466604eb97 100644 --- ./tcg/tcg.c +++ ./tcg/tcg.c @@ -1058,6 +1058,13 @@ typedef struct TCGOutOpQemuLdSt { TCGOutOp base; void (*out)(TCGContext *s, TCGType type, TCGReg dest, TCGReg addr, MemOpIdx oi); + /* + * As out(), for an access at addr + disp. Only required of targets th= at + * define TCG_TARGET_HAS_ldst_disp; for everyone else fold_ldst_disp() + * never runs and the displacement is always zero. + */ + void (*out_disp)(TCGContext *s, TCGType type, TCGReg dest, + TCGReg addr, MemOpIdx oi, int32_t disp); } TCGOutOpQemuLdSt; =20 typedef struct TCGOutOpQemuLdSt2 { @@ -3574,6 +3581,123 @@ static void move_label_uses(TCGLabel *to, TCGLabel = *from) QSIMPLEQ_CONCAT(&to->branches, &from->branches); } =20 +#ifndef TCG_TARGET_HAS_ldst_disp +#define TCG_TARGET_HAS_ldst_disp 0 +#define tcg_target_ldst_disp_ok(s, opc, disp) false +#endif + +/* + * Return true if @opc needs an alignment test in the fast path. + * + * atom_and_align_for_opc() gives the exact answer, but only once the host= 's + * atomicity capabilities are known, and those belong to the backend. Answ= er + * instead for the most restrictive host, which is valid for all of them. + */ +static bool ldst_disp_needs_align(MemOp opc) +{ + MemOp size =3D opc & MO_SIZE; + + if (memop_alignment_bits(opc)) { + return true; + } + switch (opc & MO_ATOM_MASK) { + case MO_ATOM_NONE: + case MO_ATOM_IFALIGN: + case MO_ATOM_IFALIGN_PAIR: + return false; + case MO_ATOM_WITHIN16: + /* Misalignment implies !within16, and therefore no atomicity. */ + return size !=3D MO_128; + case MO_ATOM_WITHIN16_PAIR: + case MO_ATOM_SUBALIGN: + return size !=3D MO_8; + default: + g_assert_not_reached(); + } +} + +/* + * Fold "add addr, base, $disp" into the guest access that follows it, so + * that the displacement becomes part of the host addressing mode instead = of + * a separate instruction. Frontends have no way to express this: there is + * no displacement operand on tcg_gen_qemu_ld/st, so a based access always + * costs an extra add, and an extra register to hold its result. + * + * Only an add in the op immediately before the access is recognized. That + * is what the frontends emit, and a window of one op means no analysis is + * needed of what might have happened in between. The add is left in place; + * liveness removes it if its result has no other use. + */ +static void __attribute__((noinline)) +fold_ldst_disp(TCGContext *s) +{ + TCGOp *op; + + /* + * The fold requires that the access have no slow path, because the sl= ow + * path hands the address operand to the helper and that register no + * longer holds the complete guest address. That means user-only, since + * softmmu compares the unadjusted address against the TLB. It also + * requires a 64-bit address type: for a 32-bit one the add wraps and a + * host displacement would not. + */ + if (!TCG_TARGET_HAS_ldst_disp || tcg_use_softmmu || + s->addr_type !=3D TCG_TYPE_I64) { + return; + } + + QTAILQ_FOREACH(op, &s->ops, link) { + TCGOp *prev; + TCGTemp *cts; + int64_t disp; + MemOp opc; + + switch (op->opc) { + case INDEX_op_qemu_ld: + case INDEX_op_qemu_st: + break; + default: + continue; + } + + opc =3D get_memop(op->args[2]); + if (ldst_disp_needs_align(opc)) { + continue; + } + + prev =3D QTAILQ_PREV(op, link); + if (prev =3D=3D NULL || prev->opc !=3D INDEX_op_add || + TCGOP_TYPE(prev) !=3D s->addr_type) { + continue; + } + + /* + * The add must define the address operand, and must not have + * clobbered the base it read: after the fold the access reads the + * base directly, so the base has to still hold its original value. + */ + if (prev->args[0] !=3D op->args[1] || prev->args[0] =3D=3D prev->a= rgs[1]) { + continue; + } + + cts =3D arg_temp(prev->args[2]); + if (cts->kind !=3D TEMP_CONST) { + continue; + } + /* out_disp() takes an int32_t, so anything wider cannot be passed= . */ + disp =3D cts->val; + if (disp !=3D (int32_t)disp) { + continue; + } + if (disp =3D=3D 0 || !tcg_target_ldst_disp_ok(s, opc, disp)) { + continue; + } + + op->args[1] =3D prev->args[1]; + op->args[3] =3D disp; + } +} + /* Reachable analysis : remove unreachable code. */ static void __attribute__((noinline)) reachable_code_pass(TCGContext *s) @@ -5728,7 +5852,12 @@ static void tcg_reg_alloc_op(TCGContext *s, const TC= GOp *op) const TCGOutOpQemuLdSt *out =3D container_of(all_outop[op->opc], TCGOutOpQemuLdSt, base); =20 - out->out(s, type, new_args[0], new_args[1], new_args[2]); + if (new_args[3]) { + out->out_disp(s, type, new_args[0], new_args[1], + new_args[2], new_args[3]); + } else { + out->out(s, type, new_args[0], new_args[1], new_args[2]); + } } break; =20 @@ -6611,6 +6740,7 @@ int tcg_gen_code(TCGContext *s, TranslationBlock *tb,= uint64_t pc_start) tcg_temp_ebb_reset_freed(s); =20 tcg_optimize(s); + fold_ldst_disp(s); =20 reachable_code_pass(s); liveness_pass_0(s); diff --git ./tcg/x86_64/tcg-target.c.inc ./tcg/x86_64/tcg-target.c.inc index 2c8f1f3e58..9b177d3475 100644 --- ./tcg/x86_64/tcg-target.c.inc +++ ./tcg/x86_64/tcg-target.c.inc @@ -1892,6 +1892,18 @@ static HostAddress x86_guest_base =3D { .index =3D -1 }; =20 +/* + * Whether the displacement of a guest access can be folded into the host + * addressing mode rather than materialized by a separate lea. The generic + * pass has already established that the access has no slow path, so all + * that is left is guest_base, which shares the disp32 field. + */ +static bool tcg_target_ldst_disp_ok(TCGContext *s, MemOp opc, int32_t disp) +{ + int64_t ofs =3D (int64_t)x86_guest_base.ofs + disp; + return ofs =3D=3D (int32_t)ofs; +} + #if defined(__linux__) # include # include @@ -1917,6 +1929,7 @@ static inline int setup_guest_base_seg(void) #endif #else # define x86_guest_base (*(HostAddress *)({ qemu_build_not_reached(); NULL= ; })) +# define tcg_target_ldst_disp_ok(s, opc, disp) false #endif /* CONFIG_USER_ONLY */ #ifndef setup_guest_base_seg # define setup_guest_base_seg() 0 @@ -2183,9 +2196,23 @@ static void tgen_qemu_ld(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_ld_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, true); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_ld_direct(s, data, -1, h, type, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_ld =3D { .base.static_constraint =3D C_O1_I1(r, L), .out =3D tgen_qemu_ld, + .out_disp =3D tgen_qemu_ld_disp, }; =20 static void tgen_qemu_ld2(TCGContext *s, TCGType type, TCGReg datalo, @@ -2321,9 +2348,23 @@ static void tgen_qemu_st(TCGContext *s, TCGType type= , TCGReg data, } } =20 +static void tgen_qemu_st_disp(TCGContext *s, TCGType type, TCGReg data, + TCGReg addr, MemOpIdx oi, int32_t disp) +{ + TCGLabelQemuLdst *ldst; + HostAddress h; + + ldst =3D prepare_host_addr(s, &h, addr, oi, false); + /* tcg_target_ldst_disp_ok() has ruled out every slow path. */ + tcg_debug_assert(ldst =3D=3D NULL); + h.ofs +=3D disp; + tcg_out_qemu_st_direct(s, data, -1, h, get_memop(oi)); +} + static const TCGOutOpQemuLdSt outop_qemu_st =3D { .base.static_constraint =3D C_O0_I2(L, L), .out =3D tgen_qemu_st, + .out_disp =3D tgen_qemu_st_disp, }; =20 static void tgen_qemu_st2(TCGContext *s, TCGType type, TCGReg datalo, diff --git ./tcg/x86_64/tcg-target.h ./tcg/x86_64/tcg-target.h index 7ebae56a7d..8f2315c15e 100644 --- ./tcg/x86_64/tcg-target.h +++ ./tcg/x86_64/tcg-target.h @@ -30,6 +30,9 @@ #define TCG_TARGET_NB_REGS 32 #define MAX_CODE_GEN_BUFFER_SIZE (2 * GiB) =20 +/* A guest displacement can go in the disp32 of the addressing mode. */ +#define TCG_TARGET_HAS_ldst_disp 1 + typedef enum { TCG_REG_EAX =3D 0, TCG_REG_ECX, --=20 2.54.0