Series comparison

-[PULL 00/26] target-arm queue
+[PULL 00/72] target-arm queue
-Hi; hopefully this is the last arm pullreq before softfreeze.
+First arm pullreq of the cycle; this is mostly my softfloat NaN
-There's a handful of miscellaneous bug fixes here, but the
+handling series. (Lots more in my to-review queue, but I don't
-bulk of the pullreq is Mostafa's implementation of 2-stage
+like pullreqs growing too close to a hundred patches at a time :-))
 translation in the SMMUv3.
 thanks
 -- PMM
-The following changes since commit d74ec4d7dda6322bcc51d1b13ccbd993d3574795:
+The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:
-  Merge tag 'pull-trivial-patches' of https://gitlab.com/mjt0k/qemu into staging (2024-07-18 10:07:23 +1000)
+  Open 10.0 development tree (2024-12-10 17:41:17 +0000)
 are available in the Git repository at:
-  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20240718
+  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211
-for you to fetch changes up to 30a1690f2402e6c1582d5b3ebcf7940bfe2fad4b:
+for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:
-  hvf: arm: Do not advance PC when raising an exception (2024-07-18 13:49:30 +0100)
+  MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)
 ----------------------------------------------------------------
 target-arm queue:
- * Fix handling of LDAPR/STLR with negative offset
+ * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
- * LDAPR should honour SCTLR_ELx.nAA
+ * fpu: Make muladd NaN handling runtime-selected, not compile-time
- * Use float_status copy in sme_fmopa_s
+ * fpu: Make default NaN pattern runtime-selected, not compile-time
- * hw/display/bcm2835_fb: fix fb_use_offsets condition
+ * fpu: Minor NaN-related cleanups
- * hw/arm/smmuv3: Support and advertise nesting
+ * MAINTAINERS: email address updates
  * Use FPST_F16 for SME FMOPA (widening)
  * tests/arm-cpu-features: Do not assume PMU availability
  * hvf: arm: Do not advance PC when raising an exception
 ----------------------------------------------------------------
-Akihiko Odaki (2):
+Bernhard Beschow (5):
-      tests/arm-cpu-features: Do not assume PMU availability
+      hw/net/lan9118: Extract lan9118_phy
-      hvf: arm: Do not advance PC when raising an exception
+      hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
       hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
       hw/net/lan9118_phy: Reuse MII constants
       hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
-Daniyal Khan (2):
+Leif Lindholm (1):
-      target/arm: Use float_status copy in sme_fmopa_s
+      MAINTAINERS: update email address for Leif Lindholm
       tests/tcg/aarch64: Add test cases for SME FMOPA (widening)
-Mostafa Saleh (18):
+Peter Maydell (54):
-      hw/arm/smmu-common: Add missing size check for stage-1
+      fpu: handle raising Invalid for infzero in pick_nan_muladd
-      hw/arm/smmu: Fix IPA for stage-2 events
+      fpu: Check for default_nan_mode before calling pickNaNMulAdd
-      hw/arm/smmuv3: Fix encoding of CLASS in events
+      softfloat: Allow runtime choice of inf * 0 + NaN result
-      hw/arm/smmu: Use enum for SMMU stage
+      tests/fp: Explicitly set inf-zero-nan rule
-      hw/arm/smmu: Split smmuv3_translate()
+      target/arm: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu: Consolidate ASID and VMID types
+      target/s390: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu: Introduce CACHED_ENTRY_TO_ADDR
+      target/ppc: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmuv3: Translate CD and TT using stage-2 table
+      target/mips: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu-common: Rework TLB lookup for nesting
+      target/sparc: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu-common: Add support for nested TLB
+      target/xtensa: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu-common: Support nested translation
+      target/x86: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu: Support nesting in smmuv3_range_inval()
+      target/loongarch: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu: Introduce smmu_iotlb_inv_asid_vmid
+      target/hppa: Set FloatInfZeroNaNRule explicitly
-      hw/arm/smmu: Support nesting in the rest of commands
+      softfloat: Pass have_snan to pickNaNMulAdd
-      hw/arm/smmuv3: Support nested SMMUs in smmuv3_notify_iova()
+      softfloat: Allow runtime choice of NaN propagation for muladd
-      hw/arm/smmuv3: Handle translation faults according to SMMUPTWEventInfo
+      tests/fp: Explicitly set 3-NaN propagation rule
-      hw/arm/smmuv3: Support and advertise nesting
+      target/arm: Set Float3NaNPropRule explicitly
-      hw/arm/smmu: Refactor SMMU OAS
+      target/loongarch: Set Float3NaNPropRule explicitly
       target/ppc: Set Float3NaNPropRule explicitly
       target/s390x: Set Float3NaNPropRule explicitly
       target/sparc: Set Float3NaNPropRule explicitly
       target/mips: Set Float3NaNPropRule explicitly
       target/xtensa: Set Float3NaNPropRule explicitly
       target/i386: Set Float3NaNPropRule explicitly
       target/hppa: Set Float3NaNPropRule explicitly
       fpu: Remove use_first_nan field from float_status
       target/m68k: Don't pass NULL float_status to floatx80_default_nan()
       softfloat: Create floatx80 default NaN from parts64_default_nan
       target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
       target/m68k: In frem helper, initialize local float_status from env->fp_status
       target/m68k: Init local float_status from env fp_status in gdb get/set reg
       target/sparc: Initialize local scratch float_status from env->fp_status
       target/ppc: Use env->fp_status in helper_compute_fprf functions
       fpu: Allow runtime choice of default NaN value
       tests/fp: Set default NaN pattern explicitly
       target/microblaze: Set default NaN pattern explicitly
       target/i386: Set default NaN pattern explicitly
       target/hppa: Set default NaN pattern explicitly
       target/alpha: Set default NaN pattern explicitly
       target/arm: Set default NaN pattern explicitly
       target/loongarch: Set default NaN pattern explicitly
       target/m68k: Set default NaN pattern explicitly
       target/mips: Set default NaN pattern explicitly
       target/openrisc: Set default NaN pattern explicitly
       target/ppc: Set default NaN pattern explicitly
       target/sh4: Set default NaN pattern explicitly
       target/rx: Set default NaN pattern explicitly
       target/s390x: Set default NaN pattern explicitly
       target/sparc: Set default NaN pattern explicitly
       target/xtensa: Set default NaN pattern explicitly
       target/hexagon: Set default NaN pattern explicitly
       target/riscv: Set default NaN pattern explicitly
       target/tricore: Set default NaN pattern explicitly
       fpu: Remove default handling for dnan_pattern
-Peter Maydell (2):
+Richard Henderson (11):
-      target/arm: Fix handling of LDAPR/STLR with negative offset
+      target/arm: Copy entire float_status in is_ebf
-      target/arm: LDAPR should honour SCTLR_ELx.nAA
+      softfloat: Inline pickNaNMulAdd
       softfloat: Use goto for default nan case in pick_nan_muladd
       softfloat: Remove which from parts_pick_nan_muladd
       softfloat: Pad array size in pick_nan_muladd
       softfloat: Move propagateFloatx80NaN to softfloat.c
       softfloat: Use parts_pick_nan in propagateFloatx80NaN
       softfloat: Inline pickNaN
       softfloat: Share code between parts_pick_nan cases
       softfloat: Sink frac_cmp in parts_pick_nan until needed
       softfloat: Replace WHICH with RET in parts_pick_nan
-Richard Henderson (1):
+Vikram Garhwal (1):
-      target/arm: Use FPST_F16 for SME FMOPA (widening)
+      MAINTAINERS: Add correct email address for Vikram Garhwal
-SamJakob (1):
+ MAINTAINERS                       |   4 +-
-      hw/display/bcm2835_fb: fix fb_use_offsets condition
+ include/fpu/softfloat-helpers.h   |  38 +++-
+ include/fpu/softfloat-types.h     |  89 +++++++-
- hw/arm/smmuv3-internal.h          |  19 +-
+ include/hw/net/imx_fec.h          |   9 +-
- include/hw/arm/smmu-common.h      |  46 +++-
+ include/hw/net/lan9118_phy.h      |  37 ++++
- target/arm/tcg/a64.decode         |   2 +-
+ include/hw/net/mii.h              |   6 +
- hw/arm/smmu-common.c              | 312 ++++++++++++++++++++++---
+ target/mips/fpu_helper.h          |  20 ++
- hw/arm/smmuv3.c                   | 467 +++++++++++++++++++++++++-------------
+ target/sparc/helper.h             |   4 +-
- hw/display/bcm2835_fb.c           |   2 +-
+ fpu/softfloat.c                   |  19 ++
- target/arm/hvf/hvf.c              |   1 +
+ hw/net/imx_fec.c                  | 146 ++------------
- target/arm/tcg/sme_helper.c       |   2 +-
+ hw/net/lan9118.c                  | 137 ++-----------
- target/arm/tcg/translate-a64.c    |   2 +-
+ hw/net/lan9118_phy.c              | 222 ++++++++++++++++++++
- target/arm/tcg/translate-sme.c    |  12 +-
+ linux-user/arm/nwfpe/fpa11.c      |   5 +
- tests/qtest/arm-cpu-features.c    |  13 +-
+ target/alpha/cpu.c                |   2 +
- tests/tcg/aarch64/sme-fmopa-1.c   |  63 +++++
+ target/arm/cpu.c                  |  10 +
- tests/tcg/aarch64/sme-fmopa-2.c   |  56 +++++
+ target/arm/tcg/vec_helper.c       |  20 +-
- tests/tcg/aarch64/sme-fmopa-3.c   |  63 +++++
+ target/hexagon/cpu.c              |   2 +
- hw/arm/trace-events               |  26 ++-
+ target/hppa/fpu_helper.c          |  12 ++
- tests/tcg/aarch64/Makefile.target |   5 +-
+ target/i386/tcg/fpu_helper.c      |  12 ++
-files changed, 846 insertions(+), 245 deletions(-)
+ target/loongarch/tcg/fpu_helper.c |  14 +-
- create mode 100644 tests/tcg/aarch64/sme-fmopa-1.c
+ target/m68k/cpu.c                 |  14 +-
- create mode 100644 tests/tcg/aarch64/sme-fmopa-2.c
+ target/m68k/fpu_helper.c          |   6 +-
- create mode 100644 tests/tcg/aarch64/sme-fmopa-3.c
+ target/m68k/helper.c              |   6 +-
  target/microblaze/cpu.c           |   2 +
  target/mips/msa.c                 |  10 +
  target/openrisc/cpu.c             |   2 +
  target/ppc/cpu_init.c             |  19 ++
  target/ppc/fpu_helper.c           |   3 +-
  target/riscv/cpu.c                |   2 +
  target/rx/cpu.c                   |   2 +
  target/s390x/cpu.c                |   5 +
  target/sh4/cpu.c                  |   2 +
  target/sparc/cpu.c                |   6 +
  target/sparc/fop_helper.c         |   8 +-
  target/sparc/translate.c          |   4 +-
  target/tricore/helper.c           |   2 +
  target/xtensa/cpu.c               |   4 +
  target/xtensa/fpu_helper.c        |   3 +-
  tests/fp/fp-bench.c               |   7 +
  tests/fp/fp-test-log2.c           |   1 +
  tests/fp/fp-test.c                |   7 +
  fpu/softfloat-parts.c.inc         | 152 +++++++++++---
  fpu/softfloat-specialize.c.inc    | 412 ++------------------------------------
  .mailmap                          |   5 +-
  hw/net/Kconfig                    |   5 +
  hw/net/meson.build                |   1 +
  hw/net/trace-events               |  10 +-
 files changed, 778 insertions(+), 730 deletions(-)
  create mode 100644 include/hw/net/lan9118_phy.h
  create mode 100644 hw/net/lan9118_phy.c

-[PULL 24/26] tests/tcg/aarch64: Add test cases for SME FMOPA (widening)
+[PULL 01/72] hw/net/lan9118: Extract lan9118_phy
-From: Daniyal Khan <danikhan632@gmail.com>
+From: Bernhard Beschow <shentey@gmail.com>
-Signed-off-by: Daniyal Khan <danikhan632@gmail.com>
+A very similar implementation of the same device exists in imx_fec. Prepare for
-Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+a common implementation by extracting a device model into its own files.
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
-Message-id: 20240717060149.204788-4-richard.henderson@linaro.org
+Some migration state has been moved into the new device model which breaks
-Message-Id: 172090222034.13953.16888708708822922098-1@git.sr.ht
+migration compatibility for the following machines:
-[rth: Split test from a larger patch, tidy assembly]
+* smdkc210
-Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+* realview-*
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+* vexpress-*
 * kzm
 * mps2-*
 While breaking migration ABI, fix the size of the MII registers to be 16 bit,
 as defined by IEEE 802.3u.
 Signed-off-by: Bernhard Beschow <shentey@gmail.com>
 Tested-by: Guenter Roeck <linux@roeck-us.net>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20241102125724.532843-2-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- tests/tcg/aarch64/sme-fmopa-1.c   | 63 +++++++++++++++++++++++++++++++
+ include/hw/net/lan9118_phy.h |  37 ++++++++
- tests/tcg/aarch64/sme-fmopa-2.c   | 56 +++++++++++++++++++++++++++
+ hw/net/lan9118.c             | 137 +++++-----------------------
- tests/tcg/aarch64/sme-fmopa-3.c   | 63 +++++++++++++++++++++++++++++++
+ hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
- tests/tcg/aarch64/Makefile.target |  5 ++-
+ hw/net/Kconfig               |   4 +
-files changed, 185 insertions(+), 2 deletions(-)
+ hw/net/meson.build           |   1 +
- create mode 100644 tests/tcg/aarch64/sme-fmopa-1.c
+files changed, 233 insertions(+), 115 deletions(-)
- create mode 100644 tests/tcg/aarch64/sme-fmopa-2.c
+ create mode 100644 include/hw/net/lan9118_phy.h
- create mode 100644 tests/tcg/aarch64/sme-fmopa-3.c
+ create mode 100644 hw/net/lan9118_phy.c
-diff --git a/tests/tcg/aarch64/sme-fmopa-1.c b/tests/tcg/aarch64/sme-fmopa-1.c
+diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
 new file mode 100644
 index XXXXXXX..XXXXXXX
 --- /dev/null
-+++ b/tests/tcg/aarch64/sme-fmopa-1.c
++++ b/include/hw/net/lan9118_phy.h
 @@ -XXX,XX +XXX,XX @@
 +/*
-+ * SME outer product, 1 x 1.
++ * SMSC LAN9118 PHY emulation
-+ * SPDX-License-Identifier: GPL-2.0-or-later
++ *
 + * Copyright (c) 2009 CodeSourcery, LLC.
 + * Written by Paul Brook
 + *
 + * This work is licensed under the terms of the GNU GPL, version 2 or later.
 + * See the COPYING file in the top-level directory.
 + */
 +
-+#include <stdio.h>
++#ifndef HW_NET_LAN9118_PHY_H
-+
++#define HW_NET_LAN9118_PHY_H
-+static void foo(float *dst)
++
-+{
++#include "qom/object.h"
-+    asm(".arch_extension sme\n\t"
++#include "hw/sysbus.h"
-+        "smstart\n\t"
++
-+        "ptrue p0.s, vl4\n\t"
++#define TYPE_LAN9118_PHY "lan9118-phy"
-+        "fmov z0.s, #1.0\n\t"
++OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
-+        /*
++
-+         * An outer product of a vector of 1.0 by itself should be a matrix of 1.0.
++typedef struct Lan9118PhyState {
-+         * Note that we are using tile 1 here (za1.s) rather than tile 0.
++    SysBusDevice parent_obj;
-+         */
++
-+        "zero {za}\n\t"
++    uint16_t status;
-+        "fmopa za1.s, p0/m, p0/m, z0.s, z0.s\n\t"
++    uint16_t control;
-+        /*
++    uint16_t advertise;
-+         * Read the first 4x4 sub-matrix of elements from tile 1:
++    uint16_t ints;
-+         * Note that za1h should be interchangeable here.
++    uint16_t int_mask;
-+         */
++    qemu_irq irq;
-+        "mov w12, #0\n\t"
++    bool link_down;
-+        "mova z0.s, p0/m, za1v.s[w12, #0]\n\t"
++} Lan9118PhyState;
-+        "mova z1.s, p0/m, za1v.s[w12, #1]\n\t"
++
-+        "mova z2.s, p0/m, za1v.s[w12, #2]\n\t"
++void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
-+        "mova z3.s, p0/m, za1v.s[w12, #3]\n\t"
++void lan9118_phy_reset(Lan9118PhyState *s);
-+        /*
++uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
-+         * And store them to the input pointer (dst in the C code):
++void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
-+         */
++
-+        "st1w {z0.s}, p0, [%0]\n\t"
++#endif
-+        "add x0, x0, #16\n\t"
+diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
-+        "st1w {z1.s}, p0, [x0]\n\t"
+index XXXXXXX..XXXXXXX 100644
-+        "add x0, x0, #16\n\t"
+--- a/hw/net/lan9118.c
-+        "st1w {z2.s}, p0, [x0]\n\t"
++++ b/hw/net/lan9118.c
-+        "add x0, x0, #16\n\t"
+@@ -XXX,XX +XXX,XX @@
-+        "st1w {z3.s}, p0, [x0]\n\t"
+ #include "net/net.h"
-+        "smstop"
+ #include "net/eth.h"
-+        : : "r"(dst)
+ #include "hw/irq.h"
-+        : "x12", "d0", "d1", "d2", "d3", "memory");
++#include "hw/net/lan9118_phy.h"
-+}
+ #include "hw/net/lan9118.h"
-+
+ #include "hw/ptimer.h"
-+int main()
+ #include "hw/qdev-properties.h"
-+{
+@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
-+    float dst[16] = { };
+ #define MAC_CR_RXEN     0x00000004
-+
+ #define MAC_CR_RESERVED 0x7f404213
-+    foo(dst);
-+
+-#define PHY_INT_ENERGYON            0x80
-+    for (int i = 0; i < 16; i++) {
+-#define PHY_INT_AUTONEG_COMPLETE    0x40
-+        if (dst[i] != 1.0f) {
+-#define PHY_INT_FAULT               0x20
-+            goto failure;
+-#define PHY_INT_DOWN                0x10
-+        }
+-#define PHY_INT_AUTONEG_LP          0x08
 -#define PHY_INT_PARFAULT            0x04
 -#define PHY_INT_AUTONEG_PAGE        0x02
 -
  #define GPT_TIMER_EN    0x20000000
  /*
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
      uint32_t mac_mii_data;
      uint32_t mac_flow;
 -    uint32_t phy_status;
 -    uint32_t phy_control;
 -    uint32_t phy_advertise;
 -    uint32_t phy_int;
 -    uint32_t phy_int_mask;
 +    Lan9118PhyState mii;
 +    IRQState mii_irq;
      int32_t eeprom_writable;
      uint8_t eeprom[128];
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
  static const VMStateDescription vmstate_lan9118 = {
      .name = "lan9118",
 -    .version_id = 2,
 -    .minimum_version_id = 1,
 +    .version_id = 3,
 +    .minimum_version_id = 3,
      .fields = (const VMStateField[]) {
          VMSTATE_PTIMER(timer, lan9118_state),
          VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
          VMSTATE_UINT32(mac_mii_acc, lan9118_state),
          VMSTATE_UINT32(mac_mii_data, lan9118_state),
          VMSTATE_UINT32(mac_flow, lan9118_state),
 -        VMSTATE_UINT32(phy_status, lan9118_state),
 -        VMSTATE_UINT32(phy_control, lan9118_state),
 -        VMSTATE_UINT32(phy_advertise, lan9118_state),
 -        VMSTATE_UINT32(phy_int, lan9118_state),
 -        VMSTATE_UINT32(phy_int_mask, lan9118_state),
          VMSTATE_INT32(eeprom_writable, lan9118_state),
          VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
          VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
      lan9118_mac_changed(s);
  }
 -static void phy_update_irq(lan9118_state *s)
 +static void lan9118_update_irq(void *opaque, int n, int level)
  {
 -    if (s->phy_int & s->phy_int_mask) {
 +    lan9118_state *s = opaque;
 +
 +    if (level) {
          s->int_sts |= PHY_INT;
      } else {
          s->int_sts &= ~PHY_INT;
@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
      lan9118_update(s);
  }
 -static void phy_update_link(lan9118_state *s)
 -{
 -    /* Autonegotiation status mirrors link status.  */
 -    if (qemu_get_queue(s->nic)->link_down) {
 -        s->phy_status &= ~0x0024;
 -        s->phy_int |= PHY_INT_DOWN;
 -    } else {
 -        s->phy_status |= 0x0024;
 -        s->phy_int |= PHY_INT_ENERGYON;
 -        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
 -    }
 -    phy_update_irq(s);
 -}
 -
  static void lan9118_set_link(NetClientState *nc)
  {
 -    phy_update_link(qemu_get_nic_opaque(nc));
 -}
 -
 -static void phy_reset(lan9118_state *s)
 -{
 -    s->phy_status = 0x7809;
 -    s->phy_control = 0x3000;
 -    s->phy_advertise = 0x01e1;
 -    s->phy_int_mask = 0;
 -    s->phy_int = 0;
 -    phy_update_link(s);
 +    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
 +                            nc->link_down);
  }
  static void lan9118_reset(DeviceState *d)
@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
      s->read_word_n = 0;
      s->write_word_n = 0;
 -    phy_reset(s);
 -
      s->eeprom_writable = 0;
      lan9118_reload_eeprom(s);
  }
@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
      uint32_t status;
      /* FIXME: Honor TX disable, and allow queueing of packets.  */
 -    if (s->phy_control & 0x4000)  {
 +    if (s->mii.control & 0x4000) {
          /* This assumes the receive routine doesn't touch the VLANClient.  */
          qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
      } else {
@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
      }
  }
 -static uint32_t do_phy_read(lan9118_state *s, int reg)
 -{
 -    uint32_t val;
 -
 -    switch (reg) {
 -    case 0: /* Basic Control */
 -        return s->phy_control;
 -    case 1: /* Basic Status */
 -        return s->phy_status;
 -    case 2: /* ID1 */
 -        return 0x0007;
 -    case 3: /* ID2 */
 -        return 0xc0d1;
 -    case 4: /* Auto-neg advertisement */
 -        return s->phy_advertise;
 -    case 5: /* Auto-neg Link Partner Ability */
 -        return 0x0f71;
 -    case 6: /* Auto-neg Expansion */
 -        return 1;
 -        /* TODO 17, 18, 27, 29, 30, 31 */
 -    case 29: /* Interrupt source.  */
 -        val = s->phy_int;
 -        s->phy_int = 0;
 -        phy_update_irq(s);
 -        return val;
 -    case 30: /* Interrupt mask */
 -        return s->phy_int_mask;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_read: PHY read reg %d\n", reg);
 -        return 0;
 -    }
 -}
 -
 -static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
 -{
 -    switch (reg) {
 -    case 0: /* Basic Control */
 -        if (val & 0x8000) {
 -            phy_reset(s);
 -            break;
 -        }
 -        s->phy_control = val & 0x7980;
 -        /* Complete autonegotiation immediately.  */
 -        if (val & 0x1000) {
 -            s->phy_status |= 0x0020;
 -        }
 -        break;
 -    case 4: /* Auto-neg advertisement */
 -        s->phy_advertise = (val & 0x2d7f) | 0x80;
 -        break;
 -        /* TODO 17, 18, 27, 31 */
 -    case 30: /* Interrupt mask */
 -        s->phy_int_mask = val & 0xff;
 -        phy_update_irq(s);
 -        break;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 -    }
 -}
 -
  static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
  {
      switch (reg) {
@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
          if (val & 2) {
              DPRINTF("PHY write %d = 0x%04x\n",
                      (val >> 6) & 0x1f, s->mac_mii_data);
 -            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
 +            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
          } else {
 -            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
 +            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
              DPRINTF("PHY read %d = 0x%04x\n",
                      (val >> 6) & 0x1f, s->mac_mii_data);
          }
@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
          break;
      case CSR_PMT_CTRL:
          if (val & 0x400) {
 -            phy_reset(s);
 +            lan9118_phy_reset(&s->mii);
          }
          s->pmt_ctrl &= ~0x34e;
          s->pmt_ctrl |= (val & 0x34e);
@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
      const MemoryRegionOps *mem_ops =
              s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
 +    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
 +    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
 +    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
 +        return;
 +    }
-+    /* success */
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
-+    return 0;
++
-+
+     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
-+ failure:
+                           "lan9118-mmio", 0x100);
-+    for (int i = 0; i < 16; i++) {
+     sysbus_init_mmio(sbd, &s->mmio);
-+        printf("%f%c", dst[i], i % 4 == 3 ? '\n' : ' ');
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
 +    }
 +    return 1;
 +}
 diff --git a/tests/tcg/aarch64/sme-fmopa-2.c b/tests/tcg/aarch64/sme-fmopa-2.c
 new file mode 100644
 index XXXXXXX..XXXXXXX
 --- /dev/null
-+++ b/tests/tcg/aarch64/sme-fmopa-2.c
++++ b/hw/net/lan9118_phy.c
 @@ -XXX,XX +XXX,XX @@
 +/*
-+ * SME outer product, FZ vs FZ16
++ * SMSC LAN9118 PHY emulation
-+ * SPDX-License-Identifier: GPL-2.0-or-later
++ *
 + * Copyright (c) 2009 CodeSourcery, LLC.
 + * Written by Paul Brook
 + *
 + * This code is licensed under the GNU GPL v2
 + *
 + * Contributions after 2012-01-13 are licensed under the terms of the
 + * GNU GPL, version 2 or (at your option) any later version.
 + */
 +
-+#include <stdint.h>
++#include "qemu/osdep.h"
-+#include <stdio.h>
++#include "hw/net/lan9118_phy.h"
-+
++#include "hw/irq.h"
-+static void test_fmopa(uint32_t *result)
++#include "hw/resettable.h"
-+{
++#include "migration/vmstate.h"
-+    asm(".arch_extension sme\n\t"
++#include "qemu/log.h"
-+        "smstart\n\t"               /* Z*, P* and ZArray cleared */
++
-+        "ptrue p2.b, vl16\n\t"      /* Limit vector length to 16 */
++#define PHY_INT_ENERGYON            (1 << 7)
-+        "ptrue p5.b, vl16\n\t"
++#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
-+        "movi d0, #0x00ff\n\t"      /* fp16 denormal */
++#define PHY_INT_FAULT               (1 << 5)
-+        "movi d16, #0x00ff\n\t"
++#define PHY_INT_DOWN                (1 << 4)
-+        "mov w15, #0x0001000000\n\t" /* FZ=1, FZ16=0 */
++#define PHY_INT_AUTONEG_LP          (1 << 3)
-+        "msr fpcr, x15\n\t"
++#define PHY_INT_PARFAULT            (1 << 2)
-+        "fmopa za3.s, p2/m, p5/m, z16.h, z0.h\n\t"
++#define PHY_INT_AUTONEG_PAGE        (1 << 1)
-+        "mov w15, #0\n\t"
++
-+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
++static void lan9118_phy_update_irq(Lan9118PhyState *s)
-+        "add %0, %0, #16\n\t"
++{
-+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
++    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
-+        "mov w15, #2\n\t"
++}
-+        "add %0, %0, #16\n\t"
++
-+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
++uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
-+        "add %0, %0, #16\n\t"
++{
-+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
++    uint16_t val;
-+        "smstop"
++
-+        : "+r"(result) :
++    switch (reg) {
-+        : "x15", "x16", "p2", "p5", "d0", "d16", "memory");
++    case 0: /* Basic Control */
-+}
++        return s->control;
-+
++    case 1: /* Basic Status */
-+int main(void)
++        return s->status;
-+{
++    case 2: /* ID1 */
-+    uint32_t result[4 * 4] = { };
++        return 0x0007;
-+
++    case 3: /* ID2 */
-+    test_fmopa(result);
++        return 0xc0d1;
-+
++    case 4: /* Auto-neg advertisement */
-+    if (result[0] != 0x2f7e0100) {
++        return s->advertise;
-+        printf("Test failed: Incorrect output in first 4 bytes\n"
++    case 5: /* Auto-neg Link Partner Ability */
-+               "Expected: %08x\n"
++        return 0x0f71;
-+               "Got:      %08x\n",
++    case 6: /* Auto-neg Expansion */
 +               0x2f7e0100, result[0]);
 +        return 1;
++        /* TODO 17, 18, 27, 29, 30, 31 */
++    case 29: /* Interrupt source. */
++        val = s->ints;
++        s->ints = 0;
++        lan9118_phy_update_irq(s);
++        return val;
++    case 30: /* Interrupt mask */
++        return s->int_mask;
++    default:
++        qemu_log_mask(LOG_GUEST_ERROR,
++                      "lan9118_phy_read: PHY read reg %d\n", reg);
++        return 0;
 +    }
-+
++}
-+    for (int i = 1; i < 16; ++i) {
++
-+        if (result[i] != 0) {
++void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
-+            printf("Test failed: Non-zero word at position %d\n", i);
++{
-+            return 1;
++    switch (reg) {
 +    case 0: /* Basic Control */
 +        if (val & 0x8000) {
 +            lan9118_phy_reset(s);
 +            break;
 +        }
++        s->control = val & 0x7980;
++        /* Complete autonegotiation immediately. */
++        if (val & 0x1000) {
++            s->status |= 0x0020;
++        }
++        break;
++    case 4: /* Auto-neg advertisement */
++        s->advertise = (val & 0x2d7f) | 0x80;
++        break;
++        /* TODO 17, 18, 27, 31 */
++    case 30: /* Interrupt mask */
++        s->int_mask = val & 0xff;
++        lan9118_phy_update_irq(s);
++        break;
++    default:
++        qemu_log_mask(LOG_GUEST_ERROR,
++                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 +    }
-+
++}
-+    return 0;
++
-+}
++void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
-diff --git a/tests/tcg/aarch64/sme-fmopa-3.c b/tests/tcg/aarch64/sme-fmopa-3.c
++{
-new file mode 100644
++    s->link_down = link_down;
-index XXXXXXX..XXXXXXX
++
---- /dev/null
++    /* Autonegotiation status mirrors link status. */
-+++ b/tests/tcg/aarch64/sme-fmopa-3.c
++    if (link_down) {
-@@ -XXX,XX +XXX,XX @@
++        s->status &= ~0x0024;
-+/*
++        s->ints |= PHY_INT_DOWN;
-+ * SME outer product, [ 1 2 3 4 ] squared
++    } else {
-+ * SPDX-License-Identifier: GPL-2.0-or-later
++        s->status |= 0x0024;
-+ */
++        s->ints |= PHY_INT_ENERGYON;
-+
++        s->ints |= PHY_INT_AUTONEG_COMPLETE;
-+#include <stdio.h>
++    }
-+#include <stdint.h>
++    lan9118_phy_update_irq(s);
-+#include <string.h>
++}
-+#include <math.h>
++
-+
++void lan9118_phy_reset(Lan9118PhyState *s)
-+static const float i_1234[4] = {
++{
-+    1.0f, 2.0f, 3.0f, 4.0f
++    s->control = 0x3000;
 +    s->status = 0x7809;
 +    s->advertise = 0x01e1;
 +    s->int_mask = 0;
 +    s->ints = 0;
 +    lan9118_phy_update_link(s, s->link_down);
 +}
 +
 +static void lan9118_phy_reset_hold(Object *obj, ResetType type)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    lan9118_phy_reset(s);
 +}
 +
 +static void lan9118_phy_init(Object *obj)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
 +}
 +
 +static const VMStateDescription vmstate_lan9118_phy = {
 +    .name = "lan9118-phy",
 +    .version_id = 1,
 +    .minimum_version_id = 1,
 +    .fields = (const VMStateField[]) {
 +        VMSTATE_UINT16(control, Lan9118PhyState),
 +        VMSTATE_UINT16(status, Lan9118PhyState),
 +        VMSTATE_UINT16(advertise, Lan9118PhyState),
 +        VMSTATE_UINT16(ints, Lan9118PhyState),
 +        VMSTATE_UINT16(int_mask, Lan9118PhyState),
 +        VMSTATE_BOOL(link_down, Lan9118PhyState),
 +        VMSTATE_END_OF_LIST()
 +    }
 +};
 +
-+static const float expected[4] = {
++static void lan9118_phy_class_init(ObjectClass *klass, void *data)
-+    4.515625f, 5.750000f, 6.984375f, 8.218750f
++{
 +    ResettableClass *rc = RESETTABLE_CLASS(klass);
 +    DeviceClass *dc = DEVICE_CLASS(klass);
 +
 +    rc->phases.hold = lan9118_phy_reset_hold;
 +    dc->vmsd = &vmstate_lan9118_phy;
 +}
 +
 +static const TypeInfo types[] = {
 +    {
 +        .name          = TYPE_LAN9118_PHY,
 +        .parent        = TYPE_SYS_BUS_DEVICE,
 +        .instance_size = sizeof(Lan9118PhyState),
 +        .instance_init = lan9118_phy_init,
 +        .class_init    = lan9118_phy_class_init,
 +    }
 +};
 +
-+static void test_fmopa(float *result)
++DEFINE_TYPES(types)
-+{
+diff --git a/hw/net/Kconfig b/hw/net/Kconfig
 +    asm(".arch_extension sme\n\t"
 +        "smstart\n\t"               /* ZArray cleared */
 +        "ptrue p2.b, vl16\n\t"      /* Limit vector length to 16 */
 +        "ld1w {z0.s}, p2/z, [%1]\n\t"
 +        "mov w15, #0\n\t"
 +        "mov za3h.s[w15, 0], p2/m, z0.s\n\t"
 +        "mov za3h.s[w15, 1], p2/m, z0.s\n\t"
 +        "mov w15, #2\n\t"
 +        "mov za3h.s[w15, 0], p2/m, z0.s\n\t"
 +        "mov za3h.s[w15, 1], p2/m, z0.s\n\t"
 +        "msr fpcr, xzr\n\t"
 +        "fmopa za3.s, p2/m, p2/m, z0.h, z0.h\n\t"
 +        "mov w15, #0\n\t"
 +        "st1w {za3h.s[w15, 0]}, p2, [%0]\n"
 +        "add %0, %0, #16\n\t"
 +        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
 +        "mov w15, #2\n\t"
 +        "add %0, %0, #16\n\t"
 +        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
 +        "add %0, %0, #16\n\t"
 +        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
 +        "smstop"
 +        : "+r"(result) : "r"(i_1234)
 +        : "x15", "x16", "p2", "d0", "memory");
 +}
 +
 +int main(void)
 +{
 +    float result[4 * 4] = { };
 +    int ret = 0;
 +
 +    test_fmopa(result);
 +
 +    for (int i = 0; i < 4; i++) {
 +        float actual = result[i];
 +        if (fabsf(actual - expected[i]) > 0.001f) {
 +            printf("Test failed at element %d: Expected %f, got %f\n",
 +                   i, expected[i], actual);
 +            ret = 1;
 +        }
 +    }
 +    return ret;
 +}
 diff --git a/tests/tcg/aarch64/Makefile.target b/tests/tcg/aarch64/Makefile.target
 index XXXXXXX..XXXXXXX 100644
---- a/tests/tcg/aarch64/Makefile.target
+--- a/hw/net/Kconfig
-+++ b/tests/tcg/aarch64/Makefile.target
++++ b/hw/net/Kconfig
-@@ -XXX,XX +XXX,XX @@ endif
+@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
+ config SMC91C111
- # SME Tests
+     bool
- ifneq ($(CROSS_AS_HAS_ARMV9_SME),)
--AARCH64_TESTS += sme-outprod1 sme-smopa-1 sme-smopa-2
++config LAN9118_PHY
--sme-outprod1 sme-smopa-1 sme-smopa-2: CFLAGS += $(CROSS_AS_HAS_ARMV9_SME)
++    bool
-+SME_TESTS = sme-outprod1 sme-smopa-1 sme-smopa-2 sme-fmopa-1 sme-fmopa-2 sme-fmopa-3
++
-+AARCH64_TESTS += $(SME_TESTS)
+ config LAN9118
-+$(SME_TESTS): CFLAGS += $(CROSS_AS_HAS_ARMV9_SME)
+     bool
- endif
++    select LAN9118_PHY
+     select PTIMER
- # System Registers Tests
  config NE2000_ISA
 diff --git a/hw/net/meson.build b/hw/net/meson.build
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/meson.build
 +++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
  system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
  system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
 +system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
  system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
  system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
  system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
 --
 .34.1

-New patch
+[PULL 02/72] hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
+From: Bernhard Beschow <shentey@gmail.com>
+imx_fec models the same PHY as lan9118_phy. The code is almost the same with
+imx_fec having more logging and tracing. Merge these improvements into
+lan9118_phy and reuse in imx_fec to fix the code duplication.
+Some migration state how resides in the new device model which breaks migration
+compatibility for the following machines:
+* imx25-pdk
+* sabrelite
+* mcimx7d-sabre
+* mcimx6ul-evk
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241102125724.532843-3-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ include/hw/net/imx_fec.h |   9 ++-
+ hw/net/imx_fec.c         | 146 ++++-----------------------------------
+ hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
+ hw/net/Kconfig           |   1 +
+ hw/net/trace-events      |  10 +--
+files changed, 85 insertions(+), 163 deletions(-)
+diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/hw/net/imx_fec.h
++++ b/include/hw/net/imx_fec.h
+@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
+ #define TYPE_IMX_ENET "imx.enet"
+ #include "hw/sysbus.h"
++#include "hw/net/lan9118_phy.h"
++#include "hw/irq.h"
+ #include "net/net.h"
+ #define ENET_EIR               1
+@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
+     uint32_t tx_descriptor[ENET_TX_RING_NUM];
+     uint32_t tx_ring_num;
+-    uint32_t phy_status;
+-    uint32_t phy_control;
+-    uint32_t phy_advertise;
+-    uint32_t phy_int;
+-    uint32_t phy_int_mask;
++    Lan9118PhyState mii;
++    IRQState mii_irq;
+     uint32_t phy_num;
+     bool phy_connected;
+     struct IMXFECState *phy_consumer;
+diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/imx_fec.c
++++ b/hw/net/imx_fec.c
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
+ static const VMStateDescription vmstate_imx_eth = {
+     .name = TYPE_IMX_FEC,
+-    .version_id = 2,
+-    .minimum_version_id = 2,
++    .version_id = 3,
++    .minimum_version_id = 3,
+     .fields = (const VMStateField[]) {
+         VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
+         VMSTATE_UINT32(rx_descriptor, IMXFECState),
+         VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
+-        VMSTATE_UINT32(phy_status, IMXFECState),
+-        VMSTATE_UINT32(phy_control, IMXFECState),
+-        VMSTATE_UINT32(phy_advertise, IMXFECState),
+-        VMSTATE_UINT32(phy_int, IMXFECState),
+-        VMSTATE_UINT32(phy_int_mask, IMXFECState),
+         VMSTATE_END_OF_LIST()
+     },
+     .subsections = (const VMStateDescription * const []) {
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
+     },
+ };
+-#define PHY_INT_ENERGYON            (1 << 7)
+-#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+-#define PHY_INT_FAULT               (1 << 5)
+-#define PHY_INT_DOWN                (1 << 4)
+-#define PHY_INT_AUTONEG_LP          (1 << 3)
+-#define PHY_INT_PARFAULT            (1 << 2)
+-#define PHY_INT_AUTONEG_PAGE        (1 << 1)
+-
+ static void imx_eth_update(IMXFECState *s);
+ /*
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
+  * For now we don't handle any GPIO/interrupt line, so the OS will
+  * have to poll for the PHY status.
+  */
+-static void imx_phy_update_irq(IMXFECState *s)
++static void imx_phy_update_irq(void *opaque, int n, int level)
+ {
+-    imx_eth_update(s);
+-}
+-
+-static void imx_phy_update_link(IMXFECState *s)
+-{
+-    /* Autonegotiation status mirrors link status.  */
+-    if (qemu_get_queue(s->nic)->link_down) {
+-        trace_imx_phy_update_link("down");
+-        s->phy_status &= ~0x0024;
+-        s->phy_int |= PHY_INT_DOWN;
+-    } else {
+-        trace_imx_phy_update_link("up");
+-        s->phy_status |= 0x0024;
+-        s->phy_int |= PHY_INT_ENERGYON;
+-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
+-    }
+-    imx_phy_update_irq(s);
++    imx_eth_update(opaque);
+ }
+ static void imx_eth_set_link(NetClientState *nc)
+ {
+-    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
+-}
+-
+-static void imx_phy_reset(IMXFECState *s)
+-{
+-    trace_imx_phy_reset();
+-
+-    s->phy_status = 0x7809;
+-    s->phy_control = 0x3000;
+-    s->phy_advertise = 0x01e1;
+-    s->phy_int_mask = 0;
+-    s->phy_int = 0;
+-    imx_phy_update_link(s);
++    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
++                            nc->link_down);
+ }
+ static uint32_t imx_phy_read(IMXFECState *s, int reg)
+ {
+-    uint32_t val;
+     uint32_t phy = reg / 32;
+     if (!s->phy_connected) {
+@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
+     reg %= 32;
+-    switch (reg) {
+-    case 0:     /* Basic Control */
+-        val = s->phy_control;
+-        break;
+-    case 1:     /* Basic Status */
+-        val = s->phy_status;
+-        break;
+-    case 2:     /* ID1 */
+-        val = 0x0007;
+-        break;
+-    case 3:     /* ID2 */
+-        val = 0xc0d1;
+-        break;
+-    case 4:     /* Auto-neg advertisement */
+-        val = s->phy_advertise;
+-        break;
+-    case 5:     /* Auto-neg Link Partner Ability */
+-        val = 0x0f71;
+-        break;
+-    case 6:     /* Auto-neg Expansion */
+-        val = 1;
+-        break;
+-    case 29:    /* Interrupt source.  */
+-        val = s->phy_int;
+-        s->phy_int = 0;
+-        imx_phy_update_irq(s);
+-        break;
+-    case 30:    /* Interrupt mask */
+-        val = s->phy_int_mask;
+-        break;
+-    case 17:
+-    case 18:
+-    case 27:
+-    case 31:
+-        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        val = 0;
+-        break;
+-    default:
+-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        val = 0;
+-        break;
+-    }
+-
+-    trace_imx_phy_read(val, phy, reg);
+-
+-    return val;
++    return lan9118_phy_read(&s->mii, reg);
+ }
+ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
+@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
+     reg %= 32;
+-    trace_imx_phy_write(val, phy, reg);
+-
+-    switch (reg) {
+-    case 0:     /* Basic Control */
+-        if (val & 0x8000) {
+-            imx_phy_reset(s);
+-        } else {
+-            s->phy_control = val & 0x7980;
+-            /* Complete autonegotiation immediately.  */
+-            if (val & 0x1000) {
+-                s->phy_status |= 0x0020;
+-            }
+-        }
+-        break;
+-    case 4:     /* Auto-neg advertisement */
+-        s->phy_advertise = (val & 0x2d7f) | 0x80;
+-        break;
+-    case 30:    /* Interrupt mask */
+-        s->phy_int_mask = val & 0xff;
+-        imx_phy_update_irq(s);
+-        break;
+-    case 17:
+-    case 18:
+-    case 27:
+-    case 31:
+-        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        break;
+-    default:
+-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        break;
+-    }
++    lan9118_phy_write(&s->mii, reg, val);
+ }
+ static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
+     s->rx_descriptor = 0;
+     memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
+-
+-    /* We also reset the PHY */
+-    imx_phy_reset(s);
+ }
+ static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
+     sysbus_init_irq(sbd, &s->irq[0]);
+     sysbus_init_irq(sbd, &s->irq[1]);
++    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
++    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
++    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
++        return;
++    }
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
++
+     qemu_macaddr_default_if_unset(&s->conf.macaddr);
+     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@
+  * Copyright (c) 2009 CodeSourcery, LLC.
+  * Written by Paul Brook
+  *
++ * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
++ *
+  * This code is licensed under the GNU GPL v2
+  *
+  * Contributions after 2012-01-13 are licensed under the terms of the
+@@ -XXX,XX +XXX,XX @@
+ #include "hw/resettable.h"
+ #include "migration/vmstate.h"
+ #include "qemu/log.h"
++#include "trace.h"
+ #define PHY_INT_ENERGYON            (1 << 7)
+ #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+     switch (reg) {
+     case 0: /* Basic Control */
+-        return s->control;
++        val = s->control;
++        break;
+     case 1: /* Basic Status */
+-        return s->status;
++        val = s->status;
++        break;
+     case 2: /* ID1 */
+-        return 0x0007;
++        val = 0x0007;
++        break;
+     case 3: /* ID2 */
+-        return 0xc0d1;
++        val = 0xc0d1;
++        break;
+     case 4: /* Auto-neg advertisement */
+-        return s->advertise;
++        val = s->advertise;
++        break;
+     case 5: /* Auto-neg Link Partner Ability */
+-        return 0x0f71;
++        val = 0x0f71;
++        break;
+     case 6: /* Auto-neg Expansion */
+-        return 1;
+-        /* TODO 17, 18, 27, 29, 30, 31 */
++        val = 1;
++        break;
+     case 29: /* Interrupt source. */
+         val = s->ints;
+         s->ints = 0;
+         lan9118_phy_update_irq(s);
+-        return val;
++        break;
+     case 30: /* Interrupt mask */
+-        return s->int_mask;
++        val = s->int_mask;
++        break;
++    case 17:
++    case 18:
++    case 27:
++    case 31:
++        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
++                      __func__, reg);
++        val = 0;
++        break;
+     default:
+-        qemu_log_mask(LOG_GUEST_ERROR,
+-                      "lan9118_phy_read: PHY read reg %d\n", reg);
+-        return 0;
++        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
++                      __func__, reg);
++        val = 0;
++        break;
+     }
++
++    trace_lan9118_phy_read(val, reg);
++
++    return val;
+ }
+ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+ {
++    trace_lan9118_phy_write(val, reg);
++
+     switch (reg) {
+     case 0: /* Basic Control */
+         if (val & 0x8000) {
+             lan9118_phy_reset(s);
+-            break;
+-        }
+-        s->control = val & 0x7980;
+-        /* Complete autonegotiation immediately. */
+-        if (val & 0x1000) {
+-            s->status |= 0x0020;
++        } else {
++            s->control = val & 0x7980;
++            /* Complete autonegotiation immediately. */
++            if (val & 0x1000) {
++                s->status |= 0x0020;
++            }
+         }
+         break;
+     case 4: /* Auto-neg advertisement */
+         s->advertise = (val & 0x2d7f) | 0x80;
+         break;
+-        /* TODO 17, 18, 27, 31 */
+     case 30: /* Interrupt mask */
+         s->int_mask = val & 0xff;
+         lan9118_phy_update_irq(s);
+         break;
++    case 17:
++    case 18:
++    case 27:
++    case 31:
++        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
++                      __func__, reg);
++        break;
+     default:
+-        qemu_log_mask(LOG_GUEST_ERROR,
+-                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
++        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
++                      __func__, reg);
++        break;
+     }
+ }
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+     /* Autonegotiation status mirrors link status. */
+     if (link_down) {
++        trace_lan9118_phy_update_link("down");
+         s->status &= ~0x0024;
+         s->ints |= PHY_INT_DOWN;
+     } else {
++        trace_lan9118_phy_update_link("up");
+         s->status |= 0x0024;
+         s->ints |= PHY_INT_ENERGYON;
+         s->ints |= PHY_INT_AUTONEG_COMPLETE;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+ void lan9118_phy_reset(Lan9118PhyState *s)
+ {
++    trace_lan9118_phy_reset();
++
+     s->control = 0x3000;
+     s->status = 0x7809;
+     s->advertise = 0x01e1;
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
+     .version_id = 1,
+     .minimum_version_id = 1,
+     .fields = (const VMStateField[]) {
+-        VMSTATE_UINT16(control, Lan9118PhyState),
+         VMSTATE_UINT16(status, Lan9118PhyState),
++        VMSTATE_UINT16(control, Lan9118PhyState),
+         VMSTATE_UINT16(advertise, Lan9118PhyState),
+         VMSTATE_UINT16(ints, Lan9118PhyState),
+         VMSTATE_UINT16(int_mask, Lan9118PhyState),
+diff --git a/hw/net/Kconfig b/hw/net/Kconfig
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/Kconfig
++++ b/hw/net/Kconfig
+@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
+ config IMX_FEC
+     bool
++    select LAN9118_PHY
+ config CADENCE
+     bool
+diff --git a/hw/net/trace-events b/hw/net/trace-events
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/trace-events
++++ b/hw/net/trace-events
+@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
+ allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
+ allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
++# lan9118_phy.c
++lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
++lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
++lan9118_phy_update_link(const char *s) "%s"
++lan9118_phy_reset(void) ""
++
+ # lance.c
+ lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
+ lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
+@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
+ i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
+ # imx_fec.c
+-imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
+ imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
+-imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
+ imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
+-imx_phy_update_link(const char *s) "%s"
+-imx_phy_reset(void) ""
+ imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
+ imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
+ imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
+--
+.34.1

-[PULL 01/26] target/arm: Fix handling of LDAPR/STLR with negative offset
+[PULL 03/72] hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
-When we converted the LDAPR/STLR instructions to decodetree we
+From: Bernhard Beschow <shentey@gmail.com>
 accidentally introduced a regression where the offset is negative.
 The 9-bit immediate field is signed, and the old hand decoder
 correctly used sextract32() to get it out of the insn word,
 but the ldapr_stlr_i pattern in the decode file used "imm:9"
 instead of "imm:s9", so it treated the field as unsigned.
-Fix the pattern to treat the field as a signed immediate.
+Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
 fixes the MSB of selector field to be zero, as specified in the datasheet.
-Cc: qemu-stable@nongnu.org
+Fixes: 2a424990170b "LAN9118 emulation"
-Fixes: 2521b6073b7 ("target/arm: Convert LDAPR/STLR (imm) to decodetree")
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
-Resolves: https://gitlab.com/qemu-project/qemu/-/issues/2419
+Tested-by: Guenter Roeck <linux@roeck-us.net>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20241102125724.532843-4-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240709134504.3500007-2-peter.maydell@linaro.org
 ---
- target/arm/tcg/a64.decode | 2 +-
+ hw/net/lan9118_phy.c | 2 +-
 file changed, 1 insertion(+), 1 deletion(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/hw/net/lan9118_phy.c
-+++ b/target/arm/tcg/a64.decode
++++ b/hw/net/lan9118_phy.c
-@@ -XXX,XX +XXX,XX @@ LDAPR           sz:2 111 0 00 1 0 1 11111 1100 00 rn:5 rt:5
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
- LDRA            11 111 0 00 m:1 . 1 ......... w:1 1 rn:5 rt:5 imm=%ldra_imm
+         val = s->advertise;
+         break;
- &ldapr_stlr_i   rn rt imm sz sign ext
+     case 5: /* Auto-neg Link Partner Ability */
--@ldapr_stlr_i   .. ...... .. . imm:9 .. rn:5 rt:5 &ldapr_stlr_i
+-        val = 0x0f71;
-+@ldapr_stlr_i   .. ...... .. . imm:s9 .. rn:5 rt:5 &ldapr_stlr_i
++        val = 0x0fe1;
- STLR_i          sz:2 011001 00 0 ......... 00 ..... ..... @ldapr_stlr_i sign=0 ext=0
+         break;
- LDAPR_i         sz:2 011001 01 0 ......... 00 ..... ..... @ldapr_stlr_i sign=0 ext=0
+     case 6: /* Auto-neg Expansion */
- LDAPR_i         00 011001 10 0 ......... 00 ..... ..... @ldapr_stlr_i sign=1 ext=0 sz=0
+         val = 1;
 --
 .34.1

-[PULL 26/26] hvf: arm: Do not advance PC when raising an exception
+[PULL 04/72] hw/net/lan9118_phy: Reuse MII constants
-From: Akihiko Odaki <akihiko.odaki@daynix.com>
+From: Bernhard Beschow <shentey@gmail.com>
-hvf did not advance PC when raising an exception for most unhandled
+Prefer named constants over magic values for better readability.
 system registers, but it mistakenly advanced PC when raising an
 exception for GICv3 registers.
-Cc: qemu-stable@nongnu.org
-Fixes: a2260983c655 ("hvf: arm: Add support for GICv3")
-Signed-off-by: Akihiko Odaki <akihiko.odaki@daynix.com>
-Message-id: 20240716-pmu-v3-4-8c7c1858a227@daynix.com
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Message-id: 20241102125724.532843-5-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/hvf/hvf.c | 1 +
+ include/hw/net/mii.h |  6 +++++
-file changed, 1 insertion(+)
+ hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
 files changed, 46 insertions(+), 23 deletions(-)
-diff --git a/target/arm/hvf/hvf.c b/target/arm/hvf/hvf.c
+diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/hvf/hvf.c
+--- a/include/hw/net/mii.h
-+++ b/target/arm/hvf/hvf.c
++++ b/include/hw/net/mii.h
-@@ -XXX,XX +XXX,XX @@ static int hvf_sysreg_read(CPUState *cpu, uint32_t reg, uint32_t rt)
+@@ -XXX,XX +XXX,XX @@
-         /* Call the TCG sysreg handler. This is only safe for GICv3 regs. */
+ #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
-         if (!hvf_sysreg_read_cp(cpu, reg, &val)) {
+ #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
-             hvf_raise_exception(cpu, EXCP_UDEF, syn_uncategorized());
-+            return 1;
++#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
  #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
  #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
  #define MII_ANAR_TXFD       (1 << 8)
@@ -XXX,XX +XXX,XX @@
  #define MII_ANAR_10FD       (1 << 6)
  #define MII_ANAR_10         (1 << 5)
  #define MII_ANAR_CSMACD     (1 << 0)
 +#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
  #define MII_ANLPAR_ACK      (1 << 14)
  #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
@@ -XXX,XX +XXX,XX @@
  #define RTL8201CP_PHYID1    0x0000
  #define RTL8201CP_PHYID2    0x8201
 +/* SMSC LAN9118 */
 +#define SMSCLAN9118_PHYID1  0x0007
 +#define SMSCLAN9118_PHYID2  0xc0d1
 +
  /* RealTek 8211E */
  #define RTL8211E_PHYID1     0x001c
  #define RTL8211E_PHYID2     0xc915
 diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/lan9118_phy.c
 +++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
  #include "qemu/osdep.h"
  #include "hw/net/lan9118_phy.h"
 +#include "hw/net/mii.h"
  #include "hw/irq.h"
  #include "hw/resettable.h"
  #include "migration/vmstate.h"
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
      uint16_t val;
      switch (reg) {
 -    case 0: /* Basic Control */
 +    case MII_BMCR:
          val = s->control;
          break;
 -    case 1: /* Basic Status */
 +    case MII_BMSR:
          val = s->status;
          break;
 -    case 2: /* ID1 */
 -        val = 0x0007;
 +    case MII_PHYID1:
 +        val = SMSCLAN9118_PHYID1;
          break;
 -    case 3: /* ID2 */
 -        val = 0xc0d1;
 +    case MII_PHYID2:
 +        val = SMSCLAN9118_PHYID2;
          break;
 -    case 4: /* Auto-neg advertisement */
 +    case MII_ANAR:
          val = s->advertise;
          break;
 -    case 5: /* Auto-neg Link Partner Ability */
 -        val = 0x0fe1;
 +    case MII_ANLPAR:
 +        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
 +              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
 +              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
          break;
 -    case 6: /* Auto-neg Expansion */
 -        val = 1;
 +    case MII_ANER:
 +        val = MII_ANER_NWAY;
          break;
      case 29: /* Interrupt source. */
          val = s->ints;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
      trace_lan9118_phy_write(val, reg);
      switch (reg) {
 -    case 0: /* Basic Control */
 -        if (val & 0x8000) {
 +    case MII_BMCR:
 +        if (val & MII_BMCR_RESET) {
              lan9118_phy_reset(s);
          } else {
 -            s->control = val & 0x7980;
 +            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
 +                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
 +                                MII_BMCR_CTST);
              /* Complete autonegotiation immediately. */
 -            if (val & 0x1000) {
 -                s->status |= 0x0020;
 +            if (val & MII_BMCR_AUTOEN) {
 +                s->status |= MII_BMSR_AN_COMP;
              }
          }
          break;
-     case SYSREG_DBGBVR0_EL1:
+-    case 4: /* Auto-neg advertisement */
 -        s->advertise = (val & 0x2d7f) | 0x80;
 +    case MII_ANAR:
 +        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
 +                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
 +                               MII_ANAR_SELECT))
 +                     | MII_ANAR_TX;
          break;
      case 30: /* Interrupt mask */
          s->int_mask = val & 0xff;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
      /* Autonegotiation status mirrors link status. */
      if (link_down) {
          trace_lan9118_phy_update_link("down");
 -        s->status &= ~0x0024;
 +        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
          s->ints |= PHY_INT_DOWN;
      } else {
          trace_lan9118_phy_update_link("up");
 -        s->status |= 0x0024;
 +        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
          s->ints |= PHY_INT_ENERGYON;
          s->ints |= PHY_INT_AUTONEG_COMPLETE;
      }
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
  {
      trace_lan9118_phy_reset();
 -    s->control = 0x3000;
 -    s->status = 0x7809;
 -    s->advertise = 0x01e1;
 +    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
 +    s->status = MII_BMSR_100TX_FD
 +                | MII_BMSR_100TX_HD
 +                | MII_BMSR_10T_FD
 +                | MII_BMSR_10T_HD
 +                | MII_BMSR_AUTONEG
 +                | MII_BMSR_EXTCAP;
 +    s->advertise = MII_ANAR_TXFD
 +                   | MII_ANAR_TX
 +                   | MII_ANAR_10FD
 +                   | MII_ANAR_10
 +                   | MII_ANAR_CSMACD;
      s->int_mask = 0;
      s->ints = 0;
      lan9118_phy_update_link(s, s->link_down);
 --
 .34.1

-[PULL 25/26] tests/arm-cpu-features: Do not assume PMU availability
+[PULL 05/72] hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
-From: Akihiko Odaki <akihiko.odaki@daynix.com>
+From: Bernhard Beschow <shentey@gmail.com>
-Asahi Linux supports KVM but lacks PMU support.
+The real device advertises this mode and the device model already advertises
 mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
 make the model more realistic.
-Signed-off-by: Akihiko Odaki <akihiko.odaki@daynix.com>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
-Message-id: 20240716-pmu-v3-1-8c7c1858a227@daynix.com
+Tested-by: Guenter Roeck <linux@roeck-us.net>
 Message-id: 20241102125724.532843-6-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- tests/qtest/arm-cpu-features.c | 13 ++++++++-----
+ hw/net/lan9118_phy.c | 4 ++--
-file changed, 8 insertions(+), 5 deletions(-)
+file changed, 2 insertions(+), 2 deletions(-)
-diff --git a/tests/qtest/arm-cpu-features.c b/tests/qtest/arm-cpu-features.c
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
 index XXXXXXX..XXXXXXX 100644
---- a/tests/qtest/arm-cpu-features.c
+--- a/hw/net/lan9118_phy.c
-+++ b/tests/qtest/arm-cpu-features.c
++++ b/hw/net/lan9118_phy.c
-@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
-     assert_set_feature(qts, "host", "kvm-no-adjvtime", false);
+         break;
+     case MII_ANAR:
-     if (g_str_equal(qtest_get_arch(), "aarch64")) {
+         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
-+        bool kvm_supports_pmu;
+-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
-         bool kvm_supports_steal_time;
+-                               MII_ANAR_SELECT))
-         bool kvm_supports_sve;
++                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
-         char max_name[8], name[8];
++                               MII_ANAR_10 | MII_ANAR_SELECT))
-@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
+                      | MII_ANAR_TX;
+         break;
-         assert_has_feature_enabled(qts, "host", "aarch64");
+     case 30: /* Interrupt mask */
 -        /* Enabling and disabling pmu should always work. */
 -        assert_has_feature_enabled(qts, "host", "pmu");
 -        assert_set_feature(qts, "host", "pmu", false);
 -        assert_set_feature(qts, "host", "pmu", true);
 -
          /*
           * Some features would be enabled by default, but they're disabled
           * because this instance of KVM doesn't support them. Test that the
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
          assert_has_feature(qts, "host", "sve");
          resp = do_query_no_props(qts, "host");
 +        kvm_supports_pmu = resp_get_feature(resp, "pmu");
          kvm_supports_steal_time = resp_get_feature(resp, "kvm-steal-time");
          kvm_supports_sve = resp_get_feature(resp, "sve");
          vls = resp_get_sve_vls(resp);
          qobject_unref(resp);
 +        if (kvm_supports_pmu) {
 +            /* If we have pmu then we should be able to toggle it. */
 +            assert_set_feature(qts, "host", "pmu", false);
 +            assert_set_feature(qts, "host", "pmu", true);
 +        }
 +
          if (kvm_supports_steal_time) {
              /* If we have steal-time then we should be able to toggle it. */
              assert_set_feature(qts, "host", "kvm-steal-time", false);
 --
 .34.1

-[PULL 14/26] hw/arm/smmu-common: Support nested translation
+[PULL 06/72] fpu: handle raising Invalid for infzero in pick_nan_muladd
-From: Mostafa Saleh <smostafa@google.com>
+For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
 Invalid for the multiplication of 0 by infinity.  Currently we handle
 this in the per-architecture ifdef ladder in pickNaNMulAdd().
 However, since this isn't really architecture specific we can hoist
 it up to the generic code.
-When nested translation is requested, do the following:
+For the cases where the infzero test in pickNaNMulAdd was
-- Translate stage-1 table address IPA into PA through stage-2.
+returning 2, we can delete the check entirely and allow the
-- Translate stage-1 table walk output (IPA) through stage-2.
+code to fall into the normal pick-a-NaN handling, because this
-- Create a single TLB entry from stage-1 and stage-2 translations
+will return 2 anyway (input 'c' being the only NaN in this case).
-  using logic introduced before.
+For the cases where infzero was returning 3 to indicate "return
 the default NaN", we must retain that "return 3".
-smmu_ptw() has a new argument SMMUState which include the TLB as
+For Arm, this looks like it might be a behaviour change because we
-stage-1 table address can be cached in there.
+used to set float_flag_invalid | float_flag_invalid_imz only if C is
 a quiet NaN.  However, it is not, because Arm target code never looks
 at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
 already raised float_flag_invalid via the "abc_mask &
 float_cmask_snan" check in pick_nan_muladd.
-Also in smmu_ptw(), a separate path used for nesting to simplify the
+For any target architecture using the "default implementation" at the
-code, although some logic can be combined.
+bottom of the ifdef, this is a behaviour change but will be fixing a
 bug (where we failed to raise the Invalid exception for (0 * inf +
 QNaN).  The architectures using the default case are:
  * hppa
  * i386
  * sh4
  * tricore
-With nested translation class of translation fault can be different,
+The x86, Tricore and SH4 CPU architecture manuals are clear that this
-from the class of the translation, as faults from translating stage-1
+should have raised Invalid; HPPA is a bit vaguer but still seems
-tables are considered as CLASS_TT and not CLASS_IN, a new member
+clear enough.
 "is_ipa_descriptor" added to "SMMUPTWEventInfo" to differ faults
 from walking stage 1 translation table and faults from translating
 an IPA for a transaction.
-Signed-off-by: Mostafa Saleh <smostafa@google.com>
-Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
-Reviewed-by: Eric Auger <eric.auger@redhat.com>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
-Message-id: 20240715084519.1189624-12-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
 ---
- include/hw/arm/smmu-common.h |  7 ++--
+ fpu/softfloat-parts.c.inc      | 13 +++++++------
- hw/arm/smmu-common.c         | 74 +++++++++++++++++++++++++++++++-----
+ fpu/softfloat-specialize.c.inc | 29 +----------------------------
- hw/arm/smmuv3.c              | 14 +++++++
+files changed, 8 insertions(+), 34 deletions(-)
 files changed, 82 insertions(+), 13 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPTWEventInfo {
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-     SMMUStage stage;
+                                             int ab_mask, int abc_mask)
-     SMMUPTWEventType type;
+ {
-     dma_addr_t addr; /* fetched address that induced an abort, if any */
+     int which;
-+    bool is_ipa_descriptor; /* src for fault in nested translation. */
++    bool infzero = (ab_mask == float_cmask_infzero);
- } SMMUPTWEventInfo;
+     if (unlikely(abc_mask & float_cmask_snan)) {
- typedef struct SMMUTransTableInfo {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
+     }
-  * smmu_ptw - Perform the page table walk for a given iova / access flags
-  * pair, according to @cfg translation config
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
-  */
+-                          ab_mask == float_cmask_infzero, s);
--int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
++    if (infzero) {
--             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
++        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
--
++        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
 +int smmu_ptw(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t iova,
 +             IOMMUAccessFlags perm, SMMUTLBEntry *tlbe,
 +             SMMUPTWEventInfo *info);
  /*
   * smmu_translate - Look for a translation in TLB, if not, do a PTW.
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
      return NULL;
  }
 +/* Translate stage-1 table address using stage-2 page table. */
 +static inline int translate_table_addr_ipa(SMMUState *bs,
 +                                           dma_addr_t *table_addr,
 +                                           SMMUTransCfg *cfg,
 +                                           SMMUPTWEventInfo *info)
 +{
 +    dma_addr_t addr = *table_addr;
 +    SMMUTLBEntry *cached_entry;
 +    int asid;
 +
 +    /*
 +     * The translation table walks performed from TTB0 or TTB1 are always
 +     * performed in IPA space if stage 2 translations are enabled.
 +     */
 +    asid = cfg->asid;
 +    cfg->stage = SMMU_STAGE_2;
 +    cfg->asid = -1;
 +    cached_entry = smmu_translate(bs, cfg, addr, IOMMU_RO, info);
 +    cfg->asid = asid;
 +    cfg->stage = SMMU_NESTED;
 +
 +    if (cached_entry) {
 +        *table_addr = CACHED_ENTRY_TO_ADDR(cached_entry, addr);
 +        return 0;
 +    }
 +
-+    info->stage = SMMU_STAGE_2;
++    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
-+    info->addr = addr;
-+    info->is_ipa_descriptor = true;
+     if (s->default_nan_mode || which == 3) {
-+    return -EINVAL;
+-        /*
-+}
+-         * Note that this check is after pickNaNMulAdd so that function
 -         * has an opportunity to set the Invalid flag for infzero.
 -         */
          parts_default_nan(a, s);
          return a;
      }
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
       * the default NaN
       */
      if (infzero && is_qnan(c_cls)) {
 -        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
          return 3;
      }
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
           * case sets InvalidOp and returns the default NaN
           */
          if (infzero) {
 -            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
              return 3;
          }
          /* Prefer sNaN over qNaN, in the a, b, c order. */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
           * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
           * case sets InvalidOp and returns the input value 'c'
           */
 -        if (infzero) {
 -            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
 -            return 2;
 -        }
          /* Prefer sNaN over qNaN, in the c, a, b order. */
          if (is_snan(c_cls)) {
              return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
       * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
       * case sets InvalidOp and returns the input value 'c'
       */
 -    if (infzero) {
 -        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
 -        return 2;
 -    }
 +
- /**
+     /* Prefer sNaN over qNaN, in the c, a, b order. */
-  * smmu_ptw_64_s1 - VMSAv8-64 Walk of the page tables for a given IOVA
+     if (is_snan(c_cls)) {
-+ * @bs: smmu state which includes TLB instance
+         return 2;
-  * @cfg: translation config
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-  * @iova: iova to translate
+      * to return an input NaN if we have one (ie c) rather than generating
-  * @perm: access type
+      * a default NaN
-@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
+      */
-  * Upon success, @tlbe is filled with translated_addr and entry
+-    if (infzero) {
-  * permission rights.
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-  */
+-        return 2;
--static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
+-    }
-+static int smmu_ptw_64_s1(SMMUState *bs, SMMUTransCfg *cfg,
-                           dma_addr_t iova, IOMMUAccessFlags perm,
+     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-                           SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
+      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
- {
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
+         return 1;
                  goto error;
              }
              baseaddr = get_table_pte_address(pte, granule_sz);
 +            if (cfg->stage == SMMU_NESTED) {
 +                if (translate_table_addr_ipa(bs, &baseaddr, cfg, info)) {
 +                    goto error;
 +                }
 +            }
              level++;
              continue;
          } else if (is_page_pte(pte, level)) {
@@ -XXX,XX +XXX,XX @@ error:
   * combine S1 and S2 TLB entries into a single entry.
   * As a result the S1 entry is overriden with combined data.
   */
 -static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
 -                                                SMMUTLBEntry *tlbe_s2,
 -                                                dma_addr_t iova,
 -                                                SMMUTransCfg *cfg)
 +static void combine_tlb(SMMUTLBEntry *tlbe, SMMUTLBEntry *tlbe_s2,
 +                        dma_addr_t iova, SMMUTransCfg *cfg)
  {
      if (tlbe_s2->entry.addr_mask < tlbe->entry.addr_mask) {
          tlbe->entry.addr_mask = tlbe_s2->entry.addr_mask;
@@ -XXX,XX +XXX,XX @@ static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
  /**
   * smmu_ptw - Walk the page tables for an IOVA, according to @cfg
   *
 + * @bs: smmu state which includes TLB instance
   * @cfg: translation configuration
   * @iova: iova to translate
   * @perm: tentative access type
@@ -XXX,XX +XXX,XX @@ static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
   *
   * return 0 on success
   */
 -int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
 -             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 +int smmu_ptw(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t iova,
 +             IOMMUAccessFlags perm, SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
  {
 +    int ret;
 +    SMMUTLBEntry tlbe_s2;
 +    dma_addr_t ipa;
 +
      if (cfg->stage == SMMU_STAGE_1) {
 -        return smmu_ptw_64_s1(cfg, iova, perm, tlbe, info);
 +        return smmu_ptw_64_s1(bs, cfg, iova, perm, tlbe, info);
      } else if (cfg->stage == SMMU_STAGE_2) {
          /*
           * If bypassing stage 1(or unimplemented), the input address is passed
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
          return smmu_ptw_64_s2(cfg, iova, perm, tlbe, info);
      }
+ #elif defined(TARGET_RISCV)
--    g_assert_not_reached();
+-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
-+    /* SMMU_NESTED. */
+-    if (infzero) {
-+    ret = smmu_ptw_64_s1(bs, cfg, iova, perm, tlbe, info);
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-+    if (ret) {
+-    }
-+        return ret;
+     return 3; /* default NaN */
-+    }
+ #elif defined(TARGET_S390X)
-+
+     if (infzero) {
-+    ipa = CACHED_ENTRY_TO_ADDR(tlbe, iova);
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-+    ret = smmu_ptw_64_s2(cfg, ipa, perm, &tlbe_s2, info);
+         return 3;
 +    if (ret) {
 +        return ret;
 +    }
 +
 +    combine_tlb(tlbe, &tlbe_s2, iova, cfg);
 +    return 0;
  }
  SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
      }
-     cached_entry = g_new0(SMMUTLBEntry, 1);
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
--    status = smmu_ptw(cfg, addr, flag, cached_entry, info);
+         return 2;
-+    status = smmu_ptw(bs, cfg, addr, flag, cached_entry, info);
+     }
-     if (status) {
+ #elif defined(TARGET_SPARC)
-             g_free(cached_entry);
+-    /* For (inf,0,nan) return c. */
-             return NULL;
+-    if (infzero) {
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-index XXXXXXX..XXXXXXX 100644
+-        return 2;
---- a/hw/arm/smmuv3.c
+-    }
-+++ b/hw/arm/smmuv3.c
+     /* Prefer SNaN over QNaN, order C, B, A. */
-@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
+     if (is_snan(c_cls)) {
-     if (!cached_entry) {
+         return 2;
-         /* All faults from PTW has S2 field. */
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-         event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
+      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
-+        /*
+      * an input NaN if we have one (ie c).
-+         * Fault class is set as follows based on "class" input to
+      */
-+         * the function and to "ptw_info" from "smmu_translate()"
+-    if (infzero) {
-+         * For stage-1:
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-+         *   - EABT => CLASS_TT (hardcoded)
+-        return 2;
-+         *   - other events => CLASS_IN (input to function)
+-    }
-+         * For stage-2 => CLASS_IN (input to function)
+     if (status->use_first_nan) {
-+         * For nested, for all events:
+         if (is_nan(a_cls)) {
-+         *  - CD fetch => CLASS_CD (input to function)
+             return 0;
 +         *  - walking stage 1 translation table  => CLASS_TT (from
 +         *    is_ipa_descriptor or input in case of TTBx)
 +         *  - s2 translation => CLASS_IN (input to function)
 +         */
 +        class = ptw_info.is_ipa_descriptor ? SMMU_CLASS_TT : class;
          switch (ptw_info.type) {
          case SMMU_PTW_ERR_WALK_EABT:
              event->type = SMMU_EVT_F_WALK_EABT;
 --
 .34.1

-New patch
+[PULL 07/72] fpu: Check for default_nan_mode before calling pickNaNMulAdd
+If the target sets default_nan_mode then we're always going to return
+the default NaN, and pickNaNMulAdd() no longer has any side effects.
+For consistency with pickNaN(), check for default_nan_mode before
+calling pickNaNMulAdd().
+When we convert pickNaNMulAdd() to allow runtime selection of the NaN
+propagation rule, this means we won't have to make the targets which
+use default_nan_mode also set a propagation rule.
+Since RiscV always uses default_nan_mode, this allows us to remove
+its ifdef case from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 8 ++++++--
+ fpu/softfloat-specialize.c.inc | 9 +++++++--
+files changed, 13 insertions(+), 4 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+     }
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    if (s->default_nan_mode) {
++        which = 3;
++    } else {
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    }
+-    if (s->default_nan_mode || which == 3) {
++    if (which == 3) {
+         parts_default_nan(a, s);
+         return a;
+     }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
++    /*
++     * We guarantee not to require the target to tell us how to
++     * pick a NaN if we're always returning the default NaN.
++     * But if we're not in default-NaN mode then the target must
++     * specify.
++     */
++    assert(!status->default_nan_mode);
+ #if defined(TARGET_ARM)
+     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
+      * the default NaN
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     } else {
+         return 1;
+     }
+-#elif defined(TARGET_RISCV)
+-    return 3; /* default NaN */
+ #elif defined(TARGET_S390X)
+     if (infzero) {
+         return 3;
+--
+.34.1

-[PULL 10/26] hw/arm/smmu: Introduce CACHED_ENTRY_TO_ADDR
+[PULL 08/72] softfloat: Allow runtime choice of inf * 0 + NaN result
-From: Mostafa Saleh <smostafa@google.com>
+IEEE 758 does not define a fixed rule for what NaN to return in
+the case of a fused multiply-add of inf * 0 + NaN. Different
-Soon, smmuv3_do_translate() will be used to translate the CD and the
+architectures thus do different things:
-TTBx, instead of re-writting the same logic to convert the returned
+ * some return the default NaN
-cached entry to an address, add a new macro CACHED_ENTRY_TO_ADDR.
+ * some return the input NaN
+ * Arm returns the default NaN if the input NaN is quiet,
-Reviewed-by: Eric Auger <eric.auger@redhat.com>
+   and the input NaN if it is signalling
-Signed-off-by: Mostafa Saleh <smostafa@google.com>
-Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
+We want to make this logic be runtime selected rather than
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+hardcoded into the binary, because:
-Message-id: 20240715084519.1189624-8-smostafa@google.com
+ * this will let us have multiple targets in one QEMU binary
  * the Arm FEAT_AFP architectural feature includes letting
    the guest select a NaN propagation rule at runtime
 In this commit we add an enum for the propagation rule, the field in
 float_status, and the corresponding getters and setters.  We change
 pickNaNMulAdd to honour this, but because all targets still leave
 this field at its default 0 value, the fallback logic will pick the
 rule type with the old ifdef ladder.
 Note that four architectures both use the muladd softfloat functions
 and did not have a branch of the ifdef ladder to specify their
 behaviour (and so were ending up with the "default" case, probably
 wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
 default_nan_mode, and so will never get into pickNaNMulAdd().  For
 HPPA and i386 we retain the same behaviour as the old default-case,
 which is to not ever return the default NaN.  This might not be
 correct but it is not a behaviour change.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
 ---
- include/hw/arm/smmu-common.h | 3 +++
+ include/fpu/softfloat-helpers.h | 11 ++++
- hw/arm/smmuv3.c              | 3 +--
+ include/fpu/softfloat-types.h   | 23 +++++++++
-files changed, 4 insertions(+), 2 deletions(-)
+ fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
+files changed, 95 insertions(+), 30 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/include/fpu/softfloat-helpers.h
-+++ b/include/hw/arm/smmu-common.h
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
- #define VMSA_IDXMSK(isz, strd, lvl)         ((1ULL << \
+     status->float_2nan_prop_rule = rule;
-                                              VMSA_BIT_LVL(isz, strd, lvl)) - 1)
+ }
-+#define CACHED_ENTRY_TO_ADDR(ent, addr)      ((ent)->entry.translated_addr + \
++static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
-+                                             ((addr) & (ent)->entry.addr_mask))
++                                             float_status *status)
 +{
 +    status->float_infzeronan_rule = rule;
 +}
 +
  static inline void set_flush_to_zero(bool val, float_status *status)
  {
      status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
      return status->float_2nan_prop_rule;
  }
 +static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 +{
 +    return status->float_infzeronan_rule;
 +}
 +
  static inline bool get_flush_to_zero(float_status *status)
  {
      return status->flush_to_zero;
 diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-types.h
 +++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
      float_2nan_prop_x87,
  } Float2NaNPropRule;
 +/*
 + * Rule for result of fused multiply-add 0 * Inf + NaN.
 + * This must be a NaN, but implementations differ on whether this
 + * is the input NaN or the default NaN.
 + *
 + * You don't need to set this if default_nan_mode is enabled.
 + * When not in default-NaN mode, it is an error for the target
 + * not to set the rule in float_status if it uses muladd, and we
 + * will assert if we need to handle an input NaN and no rule was
 + * selected.
 + */
 +typedef enum __attribute__((__packed__)) {
 +    /* No propagation rule specified */
 +    float_infzeronan_none = 0,
 +    /* Result is never the default NaN (so always the input NaN) */
 +    float_infzeronan_dnan_never,
 +    /* Result is always the default NaN */
 +    float_infzeronan_dnan_always,
 +    /* Result is the default NaN if the input NaN is quiet */
 +    float_infzeronan_dnan_if_qnan,
 +} FloatInfZeroNaNRule;
 +
  /*
-  * Page table walk error types
+  * Floating Point Status. Individual architectures may maintain
-  */
+  * several versions of float_status for different functions. The
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
      bool flush_to_zero;
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/hw/arm/smmuv3.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ epilogue:
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-     switch (status) {
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-     case SMMU_TRANS_SUCCESS:
+                          bool infzero, float_status *status)
-         entry.perm = cached_entry->entry.perm;
+ {
--        entry.translated_addr = cached_entry->entry.translated_addr +
++    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
--                                    (addr & cached_entry->entry.addr_mask);
++
-+        entry.translated_addr = CACHED_ENTRY_TO_ADDR(cached_entry, addr);
+     /*
-         entry.addr_mask = cached_entry->entry.addr_mask;
+      * We guarantee not to require the target to tell us how to
-         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
+      * pick a NaN if we're always returning the default NaN.
-                                        entry.translated_addr, entry.perm,
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
       * specify.
       */
      assert(!status->default_nan_mode);
 +
 +    if (rule == float_infzeronan_none) {
 +        /*
 +         * Temporarily fall back to ifdef ladder
 +         */
  #if defined(TARGET_ARM)
 -    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
 -     * the default NaN
 -     */
 -    if (infzero && is_qnan(c_cls)) {
 -        return 3;
 +        /*
 +         * For ARM, the (inf,zero,qnan) case returns the default NaN,
 +         * but (inf,zero,snan) returns the input NaN.
 +         */
 +        rule = float_infzeronan_dnan_if_qnan;
 +#elif defined(TARGET_MIPS)
 +        if (snan_bit_is_one(status)) {
 +            /*
 +             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the default NaN
 +             */
 +            rule = float_infzeronan_dnan_always;
 +        } else {
 +            /*
 +             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the input value 'c'
 +             */
 +            rule = float_infzeronan_dnan_never;
 +        }
 +#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
 +    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
 +    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
 +        /*
 +         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 +         * case sets InvalidOp and returns the input value 'c'
 +         */
 +        /*
 +         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
 +         * to return an input NaN if we have one (ie c) rather than generating
 +         * a default NaN
 +         */
 +        rule = float_infzeronan_dnan_never;
 +#elif defined(TARGET_S390X)
 +        rule = float_infzeronan_dnan_always;
 +#endif
      }
 +    if (infzero) {
 +        /*
 +         * Inf * 0 + NaN -- some implementations return the default NaN here,
 +         * and some return the input NaN.
 +         */
 +        switch (rule) {
 +        case float_infzeronan_dnan_never:
 +            return 2;
 +        case float_infzeronan_dnan_always:
 +            return 3;
 +        case float_infzeronan_dnan_if_qnan:
 +            return is_qnan(c_cls) ? 3 : 2;
 +        default:
 +            g_assert_not_reached();
 +        }
 +    }
 +
 +#if defined(TARGET_ARM)
 +
      /* This looks different from the ARM ARM pseudocode, because the ARM ARM
       * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
       */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      }
  #elif defined(TARGET_MIPS)
      if (snan_bit_is_one(status)) {
 -        /*
 -         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the default NaN
 -         */
 -        if (infzero) {
 -            return 3;
 -        }
          /* Prefer sNaN over qNaN, in the a, b, c order. */
          if (is_snan(a_cls)) {
              return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              return 2;
          }
      } else {
 -        /*
 -         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the input value 'c'
 -         */
          /* Prefer sNaN over qNaN, in the c, a, b order. */
          if (is_snan(c_cls)) {
              return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
  #elif defined(TARGET_LOONGARCH64)
 -    /*
 -     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 -     * case sets InvalidOp and returns the input value 'c'
 -     */
 -
      /* Prefer sNaN over qNaN, in the c, a, b order. */
      if (is_snan(c_cls)) {
          return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          return 1;
      }
  #elif defined(TARGET_PPC)
 -    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
 -     * to return an input NaN if we have one (ie c) rather than generating
 -     * a default NaN
 -     */
 -
      /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
       * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
       */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          return 1;
      }
  #elif defined(TARGET_S390X)
 -    if (infzero) {
 -        return 3;
 -    }
 -
      if (is_snan(a_cls)) {
          return 0;
      } else if (is_snan(b_cls)) {
 --
 .34.1

-New patch
+[PULL 09/72] tests/fp: Explicitly set inf-zero-nan rule
+Explicitly set a rule in the softfloat tests for the inf-zero-nan
+muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
+and so we should select here the Arm rule of
+float_infzeronan_dnan_if_qnan.
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 5 +++++
+ tests/fp/fp-test.c  | 5 +++++
+files changed, 10 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+ {
+     bench_func_t f;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+ {
+     unsigned int i;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+     verCases_maxErrorCount = n_max_errors;
+--
+.34.1

-New patch
+[PULL 10/72] target/arm: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the Arm target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 3 +++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 4 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
++ *    and the input NaN if it is signalling
+  */
+ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_ARM)
+-        /*
+-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+-         * but (inf,zero,snan) returns the input NaN.
+-         */
+-        rule = float_infzeronan_dnan_if_qnan;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             /*
+              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 11/72] target/s390: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for s390, so we
+can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_infzeronan_rule(float_infzeronan_dnan_always,
++                                  &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * a default NaN
+          */
+         rule = float_infzeronan_dnan_never;
+-#elif defined(TARGET_S390X)
+-        rule = float_infzeronan_dnan_always;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 12/72] target/ppc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the PPC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 7 +------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
++     * to return an input NaN if we have one (ie c) rather than generating
++     * a default NaN
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+              */
+             rule = float_infzeronan_dnan_never;
+         }
+-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
++#elif defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+          */
+-        /*
+-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+-         * to return an input NaN if we have one (ie c) rather than generating
+-         * a default NaN
+-         */
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 13/72] target/mips: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the MIPS target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h       |  9 +++++++++
+ target/mips/msa.c              |  4 ++++
+ fpu/softfloat-specialize.c.inc | 16 +---------------
+files changed, 14 insertions(+), 15 deletions(-)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
+ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+ {
+     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
++    FloatInfZeroNaNRule izn_rule;
+     /*
+      * With nan2008, SNaNs are silenced in the usual way.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+      */
+     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
+     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
++    /*
++     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
++     * case sets InvalidOp and returns the default NaN.
++     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'.
++     */
++    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
++    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+ }
+ static inline void restore_fp_status(CPUMIPSState *env)
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* set proper signanling bit meaning ("1" means "quiet") */
+     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
++
++    /* Inf * 0 + NaN returns the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never,
++                              &env->active_tc.msa_fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_MIPS)
+-        if (snan_bit_is_one(status)) {
+-            /*
+-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the default NaN
+-             */
+-            rule = float_infzeronan_dnan_always;
+-        } else {
+-            /*
+-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the input value 'c'
+-             */
+-            rule = float_infzeronan_dnan_never;
+-        }
+-#elif defined(TARGET_SPARC) || \
++#if defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+--
+.34.1

-New patch
+[PULL 14/72] target/sparc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the SPARC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 +--
+files changed, 3 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_SPARC) || \
+-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 15/72] target/xtensa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the xtensa target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
+---
+ target/xtensa/cpu.c            | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/xtensa/cpu.c
++++ b/target/xtensa/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
+     reset_mmu(env);
+     cs->halted = env->runstall;
+ #endif
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_no_signaling_nans(!dfpu, &env->fp_status);
+     xtensa_use_first_nan(env, !dfpu);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 16/72] target/x86: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the x86 target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 8 insertions(+), 1 deletion(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
++    /*
++     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
++     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
++     * specified -- for 0 * inf + NaN the input NaN is selected, and if
++     * there are multiple input NaNs they are selected in the order a, b, c.
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * Temporarily fall back to ifdef ladder
+          */
+ #if defined(TARGET_HPPA) || \
+-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
++    defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+--
+.34.1

-New patch
+[PULL 17/72] target/loongarch: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the loongarch target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 5 +++++
+ fpu/softfloat-specialize.c.inc    | 7 +------
+files changed, 6 insertions(+), 6 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+                             &env->fp_status);
+     set_flush_to_zero(0, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_HPPA) || \
+-    defined(TARGET_LOONGARCH)
+-        /*
+-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+-         * case sets InvalidOp and returns the input value 'c'
+-         */
++#if defined(TARGET_HPPA)
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 18/72] target/hppa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the HPPA target,
+so we can remove the ifdef from pickNaNMulAdd().
+As this is the last target to be converted to explicitly setting
+the rule, we can remove the fallback code in pickNaNMulAdd()
+entirely.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       |  2 ++
+ fpu/softfloat-specialize.c.inc | 13 +------------
+files changed, 3 insertions(+), 12 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
+-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+-
+     /*
+      * We guarantee not to require the target to tell us how to
+      * pick a NaN if we're always returning the default NaN.
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      */
+     assert(!status->default_nan_mode);
+-    if (rule == float_infzeronan_none) {
+-        /*
+-         * Temporarily fall back to ifdef ladder
+-         */
+-#if defined(TARGET_HPPA)
+-        rule = float_infzeronan_dnan_never;
+-#endif
+-    }
+-
+     if (infzero) {
+         /*
+          * Inf * 0 + NaN -- some implementations return the default NaN here,
+          * and some return the input NaN.
+          */
+-        switch (rule) {
++        switch (status->float_infzeronan_rule) {
+         case float_infzeronan_dnan_never:
+             return 2;
+         case float_infzeronan_dnan_always:
+--
+.34.1

-New patch
+[PULL 19/72] softfloat: Pass have_snan to pickNaNMulAdd
+The new implementation of pickNaNMulAdd() will find it convenient
+to know whether at least one of the three arguments to the muladd
+was a signaling NaN. We already calculate that in the caller,
+so pass it in as a new bool have_snan.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 5 +++--
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 4 insertions(+), 3 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+ {
+     int which;
+     bool infzero = (ab_mask == float_cmask_infzero);
++    bool have_snan = (abc_mask & float_cmask_snan);
+-    if (unlikely(abc_mask & float_cmask_snan)) {
++    if (unlikely(have_snan)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
+     }
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+     if (s->default_nan_mode) {
+         which = 3;
+     } else {
+-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
+     }
+     if (which == 3) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
+ *----------------------------------------------------------------------------*/
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+-                         bool infzero, float_status *status)
++                         bool infzero, bool have_snan, float_status *status)
+ {
+     /*
+      * We guarantee not to require the target to tell us how to
+--
+.34.1

-[PULL 11/26] hw/arm/smmuv3: Translate CD and TT using stage-2 table
+[PULL 20/72] softfloat: Allow runtime choice of NaN propagation for muladd
-From: Mostafa Saleh <smostafa@google.com>
+IEEE 758 does not define a fixed rule for which NaN to pick as the
+result if both operands of a 3-operand fused multiply-add operation
-According to ARM SMMU architecture specification (ARM IHI 0070 F.b),
+are NaNs.  As a result different architectures have ended up with
-In "5.2 Stream Table Entry":
+different rules for propagating NaNs.
- [51:6] S1ContextPtr
- If Config[1] == 1 (stage 2 enabled), this pointer is an IPA translated by
+QEMU currently hardcodes the NaN propagation logic into the binary
- stage 2 and the programmed value must be within the range of the IAS.
+because pickNaNMulAdd() has an ifdef ladder for different targets.
+We want to make the propagation rule instead be selectable at
-In "5.4.1 CD notes":
+runtime, because:
- The translation table walks performed from TTB0 or TTB1 are always performed
+ * this will let us have multiple targets in one QEMU binary
- in IPA space if stage 2 translations are enabled.
+ * the Arm FEAT_AFP architectural feature includes letting
+   the guest select a NaN propagation rule at runtime
-This patch implements translation of the S1 context descriptor pointer and
-TTBx base addresses through the S2 stage (IPA -> PA)
+In this commit we add an enum for the propagation rule, the field in
+float_status, and the corresponding getters and setters.  We change
-smmuv3_do_translate() is updated to have one arg which is translation
+pickNaNMulAdd to honour this, but because all targets still leave
-class, this is useful to:
+this field at its default 0 value, the fallback logic will pick the
- - Decide wether a translation is stage-2 only or use the STE config.
+rule type with the old ifdef ladder.
- - Populate the class in case of faults, WALK_EABT is left unchanged
-   for stage-1 as it is always IN, while stage-2 would match the
+It's valid not to set a propagation rule if default_nan_mode is
-   used class (TT, IN, CD), this will change slightly when the ptw
+enabled, because in that case there's no need to pick a NaN; all the
-   supports nested translation as it can also issue TT event with
+callers of pickNaNMulAdd() catch this case and skip calling it.
-   class IN.
 In case for stage-2 only translation, used in the context of nested
 translation, the stage and asid are saved and restored before and
 after calling smmu_translate().
 Translating CD or TTBx can fail for the following reasons:
 ) Large address size: This is described in
    (3.4.3 Address sizes of SMMU-originated accesses)
    - For CD ptr larger than IAS, for SMMUv3.1, it can trigger either
      C_BAD_STE or Translation fault, we implement the latter as it
      requires no extra code.
    - For TTBx, if larger than the effective stage 1 output address size, it
      triggers C_BAD_CD.
 ) Faults from PTWs (7.3 Event records)
    - F_ADDR_SIZE: large address size after first level causes stage 2 Address
      Size fault (Also in 3.4.3 Address sizes of SMMU-originated accesses)
    - F_PERMISSION: Same as an address translation. However, when
      CLASS == CD, the access is implicitly Data and a read.
    - F_ACCESS: Same as an address translation.
    - F_TRANSLATION: Same as an address translation.
    - F_WALK_EABT: Same as an address translation.
   These are already implemented in the PTW logic, so no extra handling
   required.
 As in CD and TTBx translation context, the iova is not known, setting
 the InputAddr was removed from "smmuv3_do_translate" and set after
 from "smmuv3_translate" with the new function "smmuv3_fixup_event"
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-9-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
 ---
- hw/arm/smmuv3.c | 120 +++++++++++++++++++++++++++++++++++++++++-------
+ include/fpu/softfloat-helpers.h |  11 +++
-file changed, 103 insertions(+), 17 deletions(-)
+ include/fpu/softfloat-types.h   |  55 +++++++++++
+ fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+files changed, 107 insertions(+), 126 deletions(-)
 diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/include/fpu/softfloat-helpers.h
-+++ b/hw/arm/smmuv3.c
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@ static int smmu_get_ste(SMMUv3State *s, dma_addr_t addr, STE *buf,
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
+     status->float_2nan_prop_rule = rule;
  }
-+static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
++static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
-+                                                 SMMUTransCfg *cfg,
++                                            float_status *status)
-+                                                 SMMUEventInfo *event,
++{
-+                                                 IOMMUAccessFlags flag,
++    status->float_3nan_prop_rule = rule;
-+                                                 SMMUTLBEntry **out_entry,
++}
-+                                                 SMMUTranslationClass class);
++
- /* @ssid > 0 not supported yet */
+ static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
--static int smmu_get_cd(SMMUv3State *s, STE *ste, uint32_t ssid,
+                                              float_status *status)
 -                       CD *buf, SMMUEventInfo *event)
 +static int smmu_get_cd(SMMUv3State *s, STE *ste, SMMUTransCfg *cfg,
 +                       uint32_t ssid, CD *buf, SMMUEventInfo *event)
  {
-     dma_addr_t addr = STE_CTXPTR(ste);
+@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
-     int ret, i;
+     return status->float_2nan_prop_rule;
-+    SMMUTranslationStatus status;
+ }
-+    SMMUTLBEntry *entry;
++static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
-     trace_smmuv3_get_cd(addr);
++{
-+
++    return status->float_3nan_prop_rule;
-+    if (cfg->stage == SMMU_NESTED) {
++}
-+        status = smmuv3_do_translate(s, addr, cfg, event,
++
-+                                     IOMMU_RO, &entry, SMMU_CLASS_CD);
+ static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
-+
+ {
-+        /* Same PTW faults are reported but with CLASS = CD. */
+     return status->float_infzeronan_rule;
-+        if (status != SMMU_TRANS_SUCCESS) {
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
-+            return -EINVAL;
+index XXXXXXX..XXXXXXX 100644
-+        }
+--- a/include/fpu/softfloat-types.h
-+
++++ b/include/fpu/softfloat-types.h
-+        addr = CACHED_ENTRY_TO_ADDR(entry, addr);
+@@ -XXX,XX +XXX,XX @@ this code that are retained.
  #ifndef SOFTFLOAT_TYPES_H
  #define SOFTFLOAT_TYPES_H
 +#include "hw/registerfields.h"
 +
  /*
   * Software IEC/IEEE floating-point types.
   */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
      float_2nan_prop_x87,
  } Float2NaNPropRule;
 +/*
 + * 3-input NaN propagation rule, for fused multiply-add. Individual
 + * architectures have different rules for which input NaN is
 + * propagated to the output when there is more than one NaN on the
 + * input.
 + *
 + * If default_nan_mode is enabled then it is valid not to set a NaN
 + * propagation rule, because the softfloat code guarantees not to try
 + * to pick a NaN to propagate in default NaN mode.  When not in
 + * default-NaN mode, it is an error for the target not to set the rule
 + * in float_status if it uses a muladd, and we will assert if we need
 + * to handle an input NaN and no rule was selected.
 + *
 + * The naming scheme for Float3NaNPropRule values is:
 + *  float_3nan_prop_s_abc:
 + *    = "Prefer SNaN over QNaN, then operand A over B over C"
 + *  float_3nan_prop_abc:
 + *    = "Prefer A over B over C regardless of SNaN vs QNAN"
 + *
 + * For QEMU, the multiply-add operation is A * B + C.
 + */
 +
 +/*
 + * We set the Float3NaNPropRule enum values up so we can select the
 + * right value in pickNaNMulAdd in a data driven way.
 + */
 +FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
 +FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
 +FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
 +FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
 +
 +#define PROPRULE(X, Y, Z) \
 +    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
 +
 +typedef enum __attribute__((__packed__)) {
 +    float_3nan_prop_none = 0,     /* No propagation rule specified */
 +    float_3nan_prop_abc = PROPRULE(0, 1, 2),
 +    float_3nan_prop_acb = PROPRULE(0, 2, 1),
 +    float_3nan_prop_bac = PROPRULE(1, 0, 2),
 +    float_3nan_prop_bca = PROPRULE(1, 2, 0),
 +    float_3nan_prop_cab = PROPRULE(2, 0, 1),
 +    float_3nan_prop_cba = PROPRULE(2, 1, 0),
 +    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
 +} Float3NaNPropRule;
 +
 +#undef PROPRULE
 +
  /*
   * Rule for result of fused multiply-add 0 * Inf + NaN.
   * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    Float3NaNPropRule float_3nan_prop_rule;
      FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
  static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                           bool infzero, bool have_snan, float_status *status)
  {
 +    FloatClass cls[3] = { a_cls, b_cls, c_cls };
 +    Float3NaNPropRule rule = status->float_3nan_prop_rule;
 +    int which;
 +
      /*
       * We guarantee not to require the target to tell us how to
       * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
 +    if (rule == float_3nan_prop_none) {
  #if defined(TARGET_ARM)
 -
 -    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
 -     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
 -     */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else {
 -        return 1;
 -    }
 +        /*
 +         * This looks different from the ARM ARM pseudocode, because the ARM ARM
 +         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
 +         */
 +        rule = float_3nan_prop_s_cab;
  #elif defined(TARGET_MIPS)
 -    if (snan_bit_is_one(status)) {
 -        /* Prefer sNaN over qNaN, in the a, b, c order. */
 -        if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_snan(c_cls)) {
 -            return 2;
 -        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 +        if (snan_bit_is_one(status)) {
 +            rule = float_3nan_prop_s_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_s_cab;
          }
 -    } else {
 -        /* Prefer sNaN over qNaN, in the c, a, b order. */
 -        if (is_snan(c_cls)) {
 -            return 2;
 -        } else if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_qnan(c_cls)) {
 -            return 2;
 -        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -    }
  #elif defined(TARGET_LOONGARCH64)
 -    /* Prefer sNaN over qNaN, in the c, a, b order. */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else {
 -        return 1;
 -    }
 +        rule = float_3nan_prop_s_cab;
  #elif defined(TARGET_PPC)
 -    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 -     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 -     */
 -    if (is_nan(a_cls)) {
 -        return 0;
 -    } else if (is_nan(c_cls)) {
 -        return 2;
 -    } else {
 -        return 1;
 -    }
 +        /*
 +         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 +         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 +         */
 +        rule = float_3nan_prop_acb;
  #elif defined(TARGET_S390X)
 -    if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 2;
 -    }
 +        rule = float_3nan_prop_s_abc;
  #elif defined(TARGET_SPARC)
 -    /* Prefer SNaN over QNaN, order C, B, A. */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 0;
 -    }
 +        rule = float_3nan_prop_s_cba;
  #elif defined(TARGET_XTENSA)
 -    /*
 -     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
 -     * an input NaN if we have one (ie c).
 -     */
 -    if (status->use_first_nan) {
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else if (is_nan(b_cls)) {
 -            return 1;
 +        if (status->use_first_nan) {
 +            rule = float_3nan_prop_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_cba;
          }
 -    } else {
 -        if (is_nan(c_cls)) {
 -            return 2;
 -        } else if (is_nan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -    }
  #else
 -    /* A default implementation: prefer a to b to c.
 -     * This is unlikely to actually match any real implementation.
 -     */
 -    if (is_nan(a_cls)) {
 -        return 0;
 -    } else if (is_nan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 2;
 -    }
 +        rule = float_3nan_prop_abc;
  #endif
 +    }
 +
-     /* TODO: guarantee 64-bit single-copy atomicity */
++    assert(rule != float_3nan_prop_none);
-     ret = dma_memory_read(&address_space_memory, addr, buf, sizeof(*buf),
++    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-                           MEMTXATTRS_UNSPECIFIED);
++        /* We have at least one SNaN input and should prefer it */
-@@ -XXX,XX +XXX,XX @@ static int smmu_find_ste(SMMUv3State *s, uint32_t sid, STE *ste,
++        do {
-     return 0;
++            which = rule & R_3NAN_1ST_MASK;
 +            rule >>= R_3NAN_1ST_LENGTH;
 +        } while (!is_snan(cls[which]));
 +    } else {
 +        do {
 +            which = rule & R_3NAN_1ST_MASK;
 +            rule >>= R_3NAN_1ST_LENGTH;
 +        } while (!is_nan(cls[which]));
 +    }
 +    return which;
  }
--static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
+ /*----------------------------------------------------------------------------
 +static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
 +                     CD *cd, SMMUEventInfo *event)
  {
      int ret = -EINVAL;
      int i;
 +    SMMUTranslationStatus status;
 +    SMMUTLBEntry *entry;
      if (!CD_VALID(cd) || !CD_AARCH64(cd)) {
          goto bad_cd;
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
          tt->tsz = tsz;
          tt->ttb = CD_TTB(cd, i);
 +
          if (tt->ttb & ~(MAKE_64BIT_MASK(0, cfg->oas))) {
              goto bad_cd;
          }
 +
 +        /* Translate the TTBx, from IPA to PA if nesting is enabled. */
 +        if (cfg->stage == SMMU_NESTED) {
 +            status = smmuv3_do_translate(s, tt->ttb, cfg, event, IOMMU_RO,
 +                                         &entry, SMMU_CLASS_TT);
 +            /*
 +             * Same PTW faults are reported but with CLASS = TT.
 +             * If TTBx is larger than the effective stage 1 output addres
 +             * size, it reports C_BAD_CD, which is handled by the above case.
 +             */
 +            if (status != SMMU_TRANS_SUCCESS) {
 +                return -EINVAL;
 +            }
 +            tt->ttb = CACHED_ENTRY_TO_ADDR(entry, tt->ttb);
 +        }
 +
          tt->had = CD_HAD(cd, i);
          trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz, tt->had);
      }
@@ -XXX,XX +XXX,XX @@ static int smmuv3_decode_config(IOMMUMemoryRegion *mr, SMMUTransCfg *cfg,
          return 0;
      }
 -    ret = smmu_get_cd(s, &ste, 0 /* ssid */, &cd, event);
 +    ret = smmu_get_cd(s, &ste, cfg, 0 /* ssid */, &cd, event);
      if (ret) {
          return ret;
      }
 -    return decode_cd(cfg, &cd, event);
 +    return decode_cd(s, cfg, &cd, event);
  }
  /**
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
                                                   SMMUTransCfg *cfg,
                                                   SMMUEventInfo *event,
                                                   IOMMUAccessFlags flag,
 -                                                 SMMUTLBEntry **out_entry)
 +                                                 SMMUTLBEntry **out_entry,
 +                                                 SMMUTranslationClass class)
  {
      SMMUPTWEventInfo ptw_info = {};
      SMMUState *bs = ARM_SMMU(s);
      SMMUTLBEntry *cached_entry = NULL;
 +    int asid, stage;
 +    bool desc_s2_translation = class != SMMU_CLASS_IN;
 +
 +    /*
 +     * The function uses the argument class to identify which stage is used:
 +     * - CLASS = IN: Means an input translation, determine the stage from STE.
 +     * - CLASS = CD: Means the addr is an IPA of the CD, and it would be
 +     *   translated using the stage-2.
 +     * - CLASS = TT: Means the addr is an IPA of the stage-1 translation table
 +     *   and it would be translated using the stage-2.
 +     * For the last 2 cases instead of having intrusive changes in the common
 +     * logic, we modify the cfg to be a stage-2 translation only in case of
 +     * nested, and then restore it after.
 +     */
 +    if (desc_s2_translation) {
 +        asid = cfg->asid;
 +        stage = cfg->stage;
 +        cfg->asid = -1;
 +        cfg->stage = SMMU_STAGE_2;
 +    }
      cached_entry = smmu_translate(bs, cfg, addr, flag, &ptw_info);
 +
 +    if (desc_s2_translation) {
 +        cfg->asid = asid;
 +        cfg->stage = stage;
 +    }
 +
      if (!cached_entry) {
          /* All faults from PTW has S2 field. */
          event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
          switch (ptw_info.type) {
          case SMMU_PTW_ERR_WALK_EABT:
              event->type = SMMU_EVT_F_WALK_EABT;
 -            event->u.f_walk_eabt.addr = addr;
              event->u.f_walk_eabt.rnw = flag & 0x1;
              event->u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
 -                                          SMMU_CLASS_IN : SMMU_CLASS_TT;
 +                                          class : SMMU_CLASS_TT;
              event->u.f_walk_eabt.addr2 = ptw_info.addr;
              break;
          case SMMU_PTW_ERR_TRANSLATION:
              if (PTW_RECORD_FAULT(cfg)) {
                  event->type = SMMU_EVT_F_TRANSLATION;
 -                event->u.f_translation.addr = addr;
                  event->u.f_translation.addr2 = ptw_info.addr;
 -                event->u.f_translation.class = SMMU_CLASS_IN;
 +                event->u.f_translation.class = class;
                  event->u.f_translation.rnw = flag & 0x1;
              }
              break;
          case SMMU_PTW_ERR_ADDR_SIZE:
              if (PTW_RECORD_FAULT(cfg)) {
                  event->type = SMMU_EVT_F_ADDR_SIZE;
 -                event->u.f_addr_size.addr = addr;
                  event->u.f_addr_size.addr2 = ptw_info.addr;
 -                event->u.f_addr_size.class = SMMU_CLASS_IN;
 +                event->u.f_addr_size.class = class;
                  event->u.f_addr_size.rnw = flag & 0x1;
              }
              break;
          case SMMU_PTW_ERR_ACCESS:
              if (PTW_RECORD_FAULT(cfg)) {
                  event->type = SMMU_EVT_F_ACCESS;
 -                event->u.f_access.addr = addr;
                  event->u.f_access.addr2 = ptw_info.addr;
 -                event->u.f_access.class = SMMU_CLASS_IN;
 +                event->u.f_access.class = class;
                  event->u.f_access.rnw = flag & 0x1;
              }
              break;
          case SMMU_PTW_ERR_PERMISSION:
              if (PTW_RECORD_FAULT(cfg)) {
                  event->type = SMMU_EVT_F_PERMISSION;
 -                event->u.f_permission.addr = addr;
                  event->u.f_permission.addr2 = ptw_info.addr;
 -                event->u.f_permission.class = SMMU_CLASS_IN;
 +                event->u.f_permission.class = class;
                  event->u.f_permission.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
      return SMMU_TRANS_SUCCESS;
  }
 +/*
 + * Sets the InputAddr for an SMMU_TRANS_ERROR, as it can't be
 + * set from all contexts, as smmuv3_get_config() can return
 + * translation faults in case of nested translation (for CD
 + * and TTBx). But in that case the iova is not known.
 + */
 +static void smmuv3_fixup_event(SMMUEventInfo *event, hwaddr iova)
 +{
 +    switch (event->type) {
 +    case SMMU_EVT_F_WALK_EABT:
 +    case SMMU_EVT_F_TRANSLATION:
 +    case SMMU_EVT_F_ADDR_SIZE:
 +    case SMMU_EVT_F_ACCESS:
 +    case SMMU_EVT_F_PERMISSION:
 +        event->u.f_walk_eabt.addr = iova;
 +        break;
 +    default:
 +        break;
 +    }
 +}
 +
  /* Entry point to SMMU, does everything. */
  static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                                        IOMMUAccessFlags flag, int iommu_idx)
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
          goto epilogue;
      }
 -    status = smmuv3_do_translate(s, addr, cfg, &event, flag, &cached_entry);
 +    status = smmuv3_do_translate(s, addr, cfg, &event, flag,
 +                                 &cached_entry, SMMU_CLASS_IN);
  epilogue:
      qemu_mutex_unlock(&s->mutex);
@@ -XXX,XX +XXX,XX @@ epilogue:
                                       entry.perm);
          break;
      case SMMU_TRANS_ERROR:
 +        smmuv3_fixup_event(&event, addr);
          qemu_log_mask(LOG_GUEST_ERROR,
                        "%s translation failed for iova=0x%"PRIx64" (%s)\n",
                        mr->parent_obj.name, addr, smmu_event_string(event.type));
 --
 .34.1

-New patch
+[PULL 21/72] tests/fp: Explicitly set 3-NaN propagation rule
+Explicitly set a rule in the softfloat tests for propagating NaNs in
+the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
+so we should select here the Arm rule of float_3nan_prop_s_cab.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 1 +
+ tests/fp/fp-test.c  | 1 +
+files changed, 2 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-New patch
+[PULL 22/72] target/arm: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for Arm, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 5 +++++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 6 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
++ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
++ *    but note that for QEMU muladd is a * b + c, whereas for
++ *    the pseudocode function the arguments are in the order c, a, b.
+  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+  *    and the input NaN if it is signalling
+  */
+@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     }
+     if (rule == float_3nan_prop_none) {
+-#if defined(TARGET_ARM)
+-        /*
+-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+-         */
+-        rule = float_3nan_prop_s_cab;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             rule = float_3nan_prop_s_abc;
+         } else {
+--
+.34.1

-New patch
+[PULL 23/72] target/loongarch: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for loongarch, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 1 +
+ fpu/softfloat-specialize.c.inc    | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      * case sets InvalidOp and returns the input value 'c'
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_LOONGARCH64)
+-        rule = float_3nan_prop_s_cab;
+ #elif defined(TARGET_PPC)
+         /*
+          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+--
+.34.1

-New patch
+[PULL 24/72] target/ppc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for PPC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 6 ------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * NaN propagation for fused multiply-add:
++     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
++     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
++     * whereas QEMU labels the operands as (a * b) + c.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
+     /*
+      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+      * to return an input NaN if we have one (ie c) rather than generating
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_PPC)
+-        /*
+-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+-         */
+-        rule = float_3nan_prop_acb;
+ #elif defined(TARGET_S390X)
+         rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+--
+.34.1

-New patch
+[PULL 25/72] target/s390x: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for s390x, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 1 +
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
+        /* fall through */
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_S390X)
+-        rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+         rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+--
+.34.1

-New patch
+[PULL 26/72] target/sparc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for SPARC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
++    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_SPARC)
+-        rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+         if (status->use_first_nan) {
+             rule = float_3nan_prop_abc;
+--
+.34.1

-New patch
+[PULL 27/72] target/mips: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for Arm, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h       | 4 ++++
+ target/mips/msa.c              | 3 +++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 8 insertions(+), 7 deletions(-)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+ {
+     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
+     FloatInfZeroNaNRule izn_rule;
++    Float3NaNPropRule nan3_rule;
+     /*
+      * With nan2008, SNaNs are silenced in the usual way.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+      */
+     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
++    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
++    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
++
+ }
+ static inline void restore_fp_status(CPUMIPSState *env)
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab,
+                              &env->active_tc.msa_fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
++                             &env->active_tc.msa_fp_status);
++
+     /* clear float_status exception flags */
+     set_float_exception_flags(0, &env->active_tc.msa_fp_status);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     }
+     if (rule == float_3nan_prop_none) {
+-#if defined(TARGET_MIPS)
+-        if (snan_bit_is_one(status)) {
+-            rule = float_3nan_prop_s_abc;
+-        } else {
+-            rule = float_3nan_prop_s_cab;
+-        }
+-#elif defined(TARGET_XTENSA)
++#if defined(TARGET_XTENSA)
+         if (status->use_first_nan) {
+             rule = float_3nan_prop_abc;
+         } else {
+--
+.34.1

-[PULL 12/26] hw/arm/smmu-common: Rework TLB lookup for nesting
+[PULL 28/72] target/xtensa: Set Float3NaNPropRule explicitly
-From: Mostafa Saleh <smostafa@google.com>
+Set the Float3NaNPropRule explicitly for xtensa, and remove the
 ifdef from pickNaNMulAdd().
-In the next patch, combine_tlb() will be added which combines 2 TLB
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-entries into one for nested translations, which chooses the granule
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-and level from the smallest entry.
+Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
 ---
  target/xtensa/fpu_helper.c     | 2 ++
  fpu/softfloat-specialize.c.inc | 8 --------
 files changed, 2 insertions(+), 8 deletions(-)
-This means that with nested translation, an entry can be cached with
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
 the granule of stage-2 and not stage-1.
 However, currently, the lookup for an IOVA is done with input stage
 granule, which is stage-1 for nested configuration, which will not
 work with the above logic.
 This patch reworks lookup in that case, so it falls back to stage-2
 granule if no entry is found using stage-1 granule.
 Also, drop aligning the iova to avoid over-aligning in case the iova
 is cached with a smaller granule, the TLB lookup will align the iova
 anyway for each granule and level, and the page table walker doesn't
 consider the page offset bits.
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-10-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  hw/arm/smmu-common.c | 64 +++++++++++++++++++++++++++++---------------
 file changed, 43 insertions(+), 21 deletions(-)
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/target/xtensa/fpu_helper.c
-+++ b/hw/arm/smmu-common.c
++++ b/target/xtensa/fpu_helper.c
-@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
+@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
-     return key;
+     set_use_first_nan(use_first, &env->fp_status);
      set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                               &env->fp_status);
 +    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
 +                             &env->fp_status);
  }
--SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+ void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
--                                SMMUTransTableInfo *tt, hwaddr iova)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-+static SMMUTLBEntry *smmu_iotlb_lookup_all_levels(SMMUState *bs,
+index XXXXXXX..XXXXXXX 100644
-+                                                  SMMUTransCfg *cfg,
+--- a/fpu/softfloat-specialize.c.inc
-+                                                  SMMUTransTableInfo *tt,
++++ b/fpu/softfloat-specialize.c.inc
-+                                                  hwaddr iova)
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
  {
      uint8_t tg = (tt->granule_sz - 10) / 2;
      uint8_t inputsize = 64 - tt->tsz;
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
          }
          level++;
      }
-+    return entry;
-+}
+     if (rule == float_3nan_prop_none) {
-+
+-#if defined(TARGET_XTENSA)
-+/**
+-        if (status->use_first_nan) {
-+ * smmu_iotlb_lookup - Look up for a TLB entry.
+-            rule = float_3nan_prop_abc;
-+ * @bs: SMMU state which includes the TLB instance
+-        } else {
-+ * @cfg: Configuration of the translation
+-            rule = float_3nan_prop_cba;
-+ * @tt: Translation table info (granule and tsz)
+-        }
-+ * @iova: IOVA address to lookup
+-#else
-+ *
+         rule = float_3nan_prop_abc;
-+ * returns a valid entry on success, otherwise NULL.
+-#endif
 + * In case of nested translation, tt can be updated to include
 + * the granule of the found entry as it might different from
 + * the IOVA granule.
 + */
 +SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
 +                                SMMUTransTableInfo *tt, hwaddr iova)
 +{
 +    SMMUTLBEntry *entry = NULL;
 +
 +    entry = smmu_iotlb_lookup_all_levels(bs, cfg, tt, iova);
 +    /*
 +     * For nested translation also try the s2 granule, as the TLB will insert
 +     * it if the size of s2 tlb entry was smaller.
 +     */
 +    if (!entry && (cfg->stage == SMMU_NESTED) &&
 +        (cfg->s2cfg.granule_sz != tt->granule_sz)) {
 +        tt->granule_sz = cfg->s2cfg.granule_sz;
 +        entry = smmu_iotlb_lookup_all_levels(bs, cfg, tt, iova);
 +    }
      if (entry) {
          cfg->iotlb_hits++;
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
  SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
                               IOMMUAccessFlags flag, SMMUPTWEventInfo *info)
  {
 -    uint64_t page_mask, aligned_addr;
      SMMUTLBEntry *cached_entry = NULL;
      SMMUTransTableInfo *tt;
      int status;
      /*
 -     * Combined attributes used for TLB lookup, as only one stage is supported,
 -     * it will hold attributes based on the enabled stage.
 +     * Combined attributes used for TLB lookup, holds the attributes for
 +     * the input stage.
       */
      SMMUTransTableInfo tt_combined;
 -    if (cfg->stage == SMMU_STAGE_1) {
 +    if (cfg->stage == SMMU_STAGE_2) {
 +        /* Stage2. */
 +        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
 +        tt_combined.tsz = cfg->s2cfg.tsz;
 +    } else {
          /* Select stage1 translation table. */
          tt = select_tt(cfg, addr);
          if (!tt) {
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
          }
          tt_combined.granule_sz = tt->granule_sz;
          tt_combined.tsz = tt->tsz;
 -
 -    } else {
 -        /* Stage2. */
 -        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
 -        tt_combined.tsz = cfg->s2cfg.tsz;
      }
--    /*
+     assert(rule != float_3nan_prop_none);
 -     * TLB lookup looks for granule and input size for a translation stage,
 -     * as only one stage is supported right now, choose the right values
 -     * from the configuration.
 -     */
 -    page_mask = (1ULL << tt_combined.granule_sz) - 1;
 -    aligned_addr = addr & ~page_mask;
 -
 -    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
 +    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, addr);
      if (cached_entry) {
          if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
              info->type = SMMU_PTW_ERR_PERMISSION;
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
      }
      cached_entry = g_new0(SMMUTLBEntry, 1);
 -    status = smmu_ptw(cfg, aligned_addr, flag, cached_entry, info);
 +    status = smmu_ptw(cfg, addr, flag, cached_entry, info);
      if (status) {
              g_free(cached_entry);
              return NULL;
 --
 .34.1

-New patch
+[PULL 29/72] target/i386: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for i386.  We had no
+i386-specific behaviour in the old ifdef ladder, so we were using the
+default "prefer a then b then c" fallback; this is actually the
+correct per-the-spec handling for i386.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c | 1 +
+file changed, 1 insertion(+)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      * there are multiple input NaNs they are selected in the order a, b, c.
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+--
+.34.1

-New patch
+[PULL 30/72] target/hppa: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for HPPA, and remove the
+ifdef from pickNaNMulAdd().
+HPPA is the only target that was using the default branch of the
+ifdef ladder (other targets either do not use muladd or set
+default_nan_mode), so we can remove the ifdef fallback entirely now
+(allowing the "rule not set" case to fall into the default of the
+switch statement and assert).
+We add a TODO note that the HPPA rule is probably wrong; this is
+not a behavioural change for this refactoring.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 4 ----
+files changed, 8 insertions(+), 4 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * TODO: The HPPA architecture reference only documents its NaN
++     * propagation rule for 2-operand operations. Testing on real hardware
++     * might be necessary to confirm whether this order for muladd is correct.
++     * Not preferring the SNaN is almost certainly incorrect as it diverges
++     * from the documented rules for 2-operand operations.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         }
+     }
+-    if (rule == float_3nan_prop_none) {
+-        rule = float_3nan_prop_abc;
+-    }
+-
+     assert(rule != float_3nan_prop_none);
+     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+         /* We have at least one SNaN input and should prefer it */
+--
+.34.1

-New patch
+[PULL 31/72] fpu: Remove use_first_nan field from float_status
+The use_first_nan field in float_status was an xtensa-specific way to
+select at runtime from two different NaN propagation rules.  Now that
+xtensa is using the target-agnostic NaN propagation rule selection
+that we've just added, we can remove use_first_nan, because there is
+no longer any code that reads it.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
+---
+ include/fpu/softfloat-helpers.h | 5 -----
+ include/fpu/softfloat-types.h   | 1 -
+ target/xtensa/fpu_helper.c      | 1 -
+files changed, 7 deletions(-)
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/fpu/softfloat-helpers.h
++++ b/include/fpu/softfloat-helpers.h
+@@ -XXX,XX +XXX,XX @@ static inline void set_snan_bit_is_one(bool val, float_status *status)
+     status->snan_bit_is_one = val;
+ }
+-static inline void set_use_first_nan(bool val, float_status *status)
+-{
+-    status->use_first_nan = val;
+-}
+-
+ static inline void set_no_signaling_nans(bool val, float_status *status)
+ {
+     status->no_signaling_nans = val;
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/fpu/softfloat-types.h
++++ b/include/fpu/softfloat-types.h
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
+      * softfloat-specialize.inc.c)
+      */
+     bool snan_bit_is_one;
+-    bool use_first_nan;
+     bool no_signaling_nans;
+     /* should overflowed results subtract re_bias to its exponent? */
+     bool rebias_overflow;
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/xtensa/fpu_helper.c
++++ b/target/xtensa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ static const struct {
+ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
+ {
+-    set_use_first_nan(use_first, &env->fp_status);
+     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
+                              &env->fp_status);
+     set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
+--
+.34.1

-New patch
+[PULL 32/72] target/m68k: Don't pass NULL float_status to floatx80_default_nan()
+Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
+to get the NaN bit pattern to reset the FPU registers. This
+works because it happens that our implementation of
+floatx80_default_nan() doesn't actually look at the float_status
+pointer except for TARGET_MIPS. However, this isn't guaranteed,
+and to be able to remove the ifdef in floatx80_default_nan()
+we're going to need a real float_status here.
+Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
+earlier, and thus can pass it to floatx80_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c | 12 +++++++-----
+file changed, 7 insertions(+), 5 deletions(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+     CPUState *cs = CPU(obj);
+     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
+     CPUM68KState *env = cpu_env(cs);
+-    floatx80 nan = floatx80_default_nan(NULL);
++    floatx80 nan;
+     int i;
+     if (mcc->parent_phases.hold) {
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+ #else
+     cpu_m68k_set_sr(env, SR_S | SR_I);
+ #endif
+-    for (i = 0; i < 8; i++) {
+-        env->fregs[i].d = nan;
+-    }
+-    cpu_m68k_set_fpcr(env, 0);
+     /*
+      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
+      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++
++    nan = floatx80_default_nan(&env->fp_status);
++    for (i = 0; i < 8; i++) {
++        env->fregs[i].d = nan;
++    }
++    cpu_m68k_set_fpcr(env, 0);
+     env->fpsr = 0;
+     /* TODO: We should set PC from the interrupt vector.  */
+--
+.34.1

-New patch
+[PULL 33/72] softfloat: Create floatx80 default NaN from parts64_default_nan
+We create our 128-bit default NaN by calling parts64_default_nan()
+and then adjusting the result.  We can do the same trick for creating
+the floatx80 default NaN, which lets us drop a target ifdef.
+floatx80 is used only by:
+ i386
+ m68k
+ arm nwfpe old floating-point emulation emulation support
+    (which is essentially dead, especially the parts involving floatx80)
+ PPC (only in the xsrqpxp instruction, which just rounds an input
+    value by converting to floatx80 and back, so will never generate
+    the default NaN)
+The floatx80 default NaN as currently implemented is:
+ m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
+ i386: sign = 1, exp = 1...1, int = 1, frac = 10...0
+These are the same as the parts64_default_nan for these architectures.
+This is technically a possible behaviour change for arm linux-user
+nwfpe emulation emulation, because the default NaN will now have the
+sign bit clear.  But we were already generating a different floatx80
+default NaN from the real kernel emulation we are supposedly
+following, which appears to use an all-bits-1 value:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267
+This won't affect the only "real" use of the nwfpe emulation, which
+is ancient binaries that used it as part of the old floating point
+calling convention; that only uses loads and stores of 32 and 64 bit
+floats, not any of the floatx80 behaviour the original hardware had.
+We also get the nwfpe float64 default NaN value wrong:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
+so if we ever cared about this obscure corner the right fix would be
+to correct that so nwfpe used its own default-NaN setting rather
+than the Arm VFP one.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
+---
+ fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
+file changed, 10 insertions(+), 10 deletions(-)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
+ floatx80 floatx80_default_nan(float_status *status)
+ {
+     floatx80 r;
++    /*
++     * Extrapolate from the choices made by parts64_default_nan to fill
++     * in the floatx80 format. We assume that floatx80's explicit
++     * integer bit is always set (this is true for i386 and m68k,
++     * which are the only real users of this format).
++     */
++    FloatParts64 p64;
++    parts64_default_nan(&p64, status);
+-    /* None of the targets that have snan_bit_is_one use floatx80.  */
+-    assert(!snan_bit_is_one(status));
+-#if defined(TARGET_M68K)
+-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
+-    r.high = 0x7FFF;
+-#else
+-    /* X86 */
+-    r.low = UINT64_C(0xC000000000000000);
+-    r.high = 0xFFFF;
+-#endif
++    r.high = 0x7FFF | (p64.sign << 15);
++    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
+     return r;
+ }
+--
+.34.1

-New patch
+[PULL 34/72] target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
+In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
+a zero-initialized float_status struct to float32_is_quiet_nan() and
+float64_is_quiet_nan(), with the cryptic comment "for
+snan_bit_is_one".
+This pattern appears to have been copied from target/riscv, where it
+is used because the functions there do not have ready access to the
+CPU state struct. The comment presumably refers to the fact that the
+main reason the is_quiet_nan() functions want the float_state is
+because they want to know about the snan_bit_is_one config.
+In the loongarch helpers, though, we have the CPU state struct
+to hand. Use the usual env->fp_status here. This avoids our needing
+to track that we need to update the initializer of the local
+float_status structs when the core softfloat code adds new
+options for targets to configure their behaviour.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
+     } else if (float32_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float32_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
+     } else if (float64_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float64_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+--
+.34.1

-New patch
+[PULL 35/72] target/m68k: In frem helper, initialize local float_status from env->fp_status
+In the frem helper, we have a local float_status because we want to
+execute the floatx80_div() with a custom rounding mode.  Instead of
+zero-initializing the local float_status and then having to set it up
+with the m68k standard behaviour (including the NaN propagation rule
+and copying the rounding precision from env->fp_status), initialize
+it as a complete copy of env->fp_status. This will avoid our having
+to add new code in this function for every new config knob we add
+to fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
+---
+ target/m68k/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/fpu_helper.c
++++ b/target/m68k/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
+     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
+     if (!floatx80_is_any_nan(fp_rem)) {
+-        float_status fp_status = { };
++        /* Use local temporary fp_status to set different rounding mode */
++        float_status fp_status = env->fp_status;
+         uint32_t quotient;
+         int sign;
+         /* Calculate quotient directly using round to nearest mode */
+-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
+         set_float_rounding_mode(float_round_nearest_even, &fp_status);
+-        set_floatx80_rounding_precision(
+-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
+         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
+         sign = extractFloatx80Sign(fp_quot.d);
+--
+.34.1

-New patch
+[PULL 36/72] target/m68k: Init local float_status from env fp_status in gdb get/set reg
+In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
+from float64 to floatx80 using a scratch float_status, because we
+don't want the conversion to affect the CPU's floating point exception
+status. Currently we use a zero-initialized float_status. This will
+get steadily more awkward as we add config knobs to float_status
+that the target must initialize. Avoid having to add any of that
+configuration here by instead initializing our local float_status
+from the env->fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
+---
+ target/m68k/helper.c | 6 ++++--
+file changed, 4 insertions(+), 2 deletions(-)
+diff --git a/target/m68k/helper.c b/target/m68k/helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/helper.c
++++ b/target/m68k/helper.c
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
+     }
+     switch (n) {
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
+         return 8;
+     }
+--
+.34.1

-[PULL 16/26] hw/arm/smmu: Introduce smmu_iotlb_inv_asid_vmid
+[PULL 37/72] target/sparc: Initialize local scratch float_status from env->fp_status
-From: Mostafa Saleh <smostafa@google.com>
+In the helper functions flcmps and flcmpd we use a scratch float_status
 so that we don't change the CPU state if the comparison raises any
 floating point exception flags. Instead of zero-initializing this
 scratch float_status, initialize it as a copy of env->fp_status. This
 avoids the need to explicitly initialize settings like the NaN
 propagation rule or others we might add to softfloat in future.
-Soon, Instead of doing TLB invalidation by ASID only, VMID will be
+To do this we need to pass the CPU env pointer in to the helper.
 also required.
 Add smmu_iotlb_inv_asid_vmid() which invalidates by both ASID and VMID.
-However, at the moment this function is only used in SMMU_CMD_TLBI_NH_ASID
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-which is a stage-1 command, so passing VMID = -1 keeps the original
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-behaviour.
+Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
 ---
  target/sparc/helper.h     | 4 ++--
  target/sparc/fop_helper.c | 8 ++++----
  target/sparc/translate.c  | 4 ++--
 files changed, 8 insertions(+), 8 deletions(-)
-Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
+diff --git a/target/sparc/helper.h b/target/sparc/helper.h
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-14-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  include/hw/arm/smmu-common.h |  2 +-
  hw/arm/smmu-common.c         | 20 +++++++++++++-------
  hw/arm/smmuv3.c              |  2 +-
  hw/arm/trace-events          |  2 +-
 files changed, 16 insertions(+), 10 deletions(-)
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/target/sparc/helper.h
-+++ b/include/hw/arm/smmu-common.h
++++ b/target/sparc/helper.h
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
+@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
- SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
+ DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
-                                 uint8_t tg, uint8_t level);
+ DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
- void smmu_iotlb_inv_all(SMMUState *s);
+ DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
--void smmu_iotlb_inv_asid(SMMUState *s, int asid);
+-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
-+void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid);
+-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
- void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
++DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
- void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
++DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
-                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
+ DEF_HELPER_2(raise_exception, noreturn, env, int)
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
  DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
 diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/target/sparc/fop_helper.c
-+++ b/hw/arm/smmu-common.c
++++ b/target/sparc/fop_helper.c
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_all(SMMUState *s)
+@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
-     g_hash_table_remove_all(s->iotlb);
+     return finish_fcmp(env, r, GETPC());
  }
--static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
+-uint32_t helper_flcmps(float32 src1, float32 src2)
--                                         gpointer user_data)
++uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
 +static gboolean smmu_hash_remove_by_asid_vmid(gpointer key, gpointer value,
 +                                              gpointer user_data)
  {
--    int asid = *(int *)user_data;
+     /*
-+    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
+      * FLCMP never raises an exception nor modifies any FSR fields.
-     SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
+      * Perform the comparison with a dummy fp environment.
+      */
--    return SMMU_IOTLB_ASID(*iotlb_key) == asid;
+-    float_status discard = { };
-+    return (SMMU_IOTLB_ASID(*iotlb_key) == info->asid) &&
++    float_status discard = env->fp_status;
-+           (SMMU_IOTLB_VMID(*iotlb_key) == info->vmid);
+     FloatRelation r;
      set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
      g_assert_not_reached();
  }
- static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
+-uint32_t helper_flcmpd(float64 src1, float64 src2)
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
++uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
-                                 &info);
+ {
 -    float_status discard = { };
 +    float_status discard = env->fp_status;
      FloatRelation r;
      set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
 diff --git a/target/sparc/translate.c b/target/sparc/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/sparc/translate.c
 +++ b/target/sparc/translate.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
      src1 = gen_load_fpr_F(dc, a->rs1);
      src2 = gen_load_fpr_F(dc, a->rs2);
 -    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
 +    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
      return advance_pc(dc);
  }
--void smmu_iotlb_inv_asid(SMMUState *s, int asid)
+@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
-+void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid)
- {
+     src1 = gen_load_fpr_D(dc, a->rs1);
--    trace_smmu_iotlb_inv_asid(asid);
+     src2 = gen_load_fpr_D(dc, a->rs2);
--    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid, &asid);
+-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
-+    SMMUIOTLBPageInvInfo info = {
++    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
-+        .asid = asid,
+     return advance_pc(dc);
 +        .vmid = vmid,
 +    };
 +
 +    trace_smmu_iotlb_inv_asid_vmid(asid, vmid);
 +    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_vmid, &info);
  }
- void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
-index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
-+++ b/hw/arm/smmuv3.c
-@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
-             trace_smmuv3_cmdq_tlbi_nh_asid(asid);
-             smmu_inv_notifiers_all(&s->smmu_state);
--            smmu_iotlb_inv_asid(bs, asid);
-+            smmu_iotlb_inv_asid_vmid(bs, asid, -1);
-             break;
-         }
-         case SMMU_CMD_TLBI_NH_ALL:
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
-index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/trace-events
-+++ b/hw/arm/trace-events
-@@ -XXX,XX +XXX,XX @@ smmu_ptw_page_pte(int stage, int level,  uint64_t iova, uint64_t baseaddr, uint6
- smmu_ptw_block_pte(int stage, int level, uint64_t baseaddr, uint64_t pteaddr, uint64_t pte, uint64_t iova, uint64_t gpa, int bsize_mb) "stage=%d level=%d base@=0x%"PRIx64" pte@=0x%"PRIx64" pte=0x%"PRIx64" iova=0x%"PRIx64" block address = 0x%"PRIx64" block size = %d MiB"
- smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "baseaddr=0x%"PRIx64" index=0x%x, pteaddr=0x%"PRIx64", pte=0x%"PRIx64
- smmu_iotlb_inv_all(void) "IOTLB invalidate all"
--smmu_iotlb_inv_asid(int asid) "IOTLB invalidate asid=%d"
-+smmu_iotlb_inv_asid_vmid(int asid, int vmid) "IOTLB invalidate asid=%d vmid=%d"
- smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
- smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
- smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
 --
 .34.1

-New patch
+[PULL 38/72] target/ppc: Use env->fp_status in helper_compute_fprf functions
+In the helper_compute_fprf functions, we pass a dummy float_status
+in to the is_signaling_nan() function. This is unnecessary, because
+we have convenient access to the CPU env pointer here and that
+is already set up with the correct values for the snan_bit_is_one
+and no_signaling_nans config settings. is_signaling_nan() doesn't
+ever update the fp_status with any exception flags, so there is
+no reason not to use env->fp_status here.
+Use env->fp_status instead of the dummy fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
+---
+ target/ppc/fpu_helper.c | 3 +--
+file changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/fpu_helper.c
++++ b/target/ppc/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
+     } else if (tp##_is_infinity(arg)) {                           \
+         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
+     } else {                                                      \
+-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
+-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
++        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
+             fprf = 0x00 << FPSCR_FPRF;                            \
+         } else {                                                  \
+             fprf = 0x11 << FPSCR_FPRF;                            \
+--
+.34.1

-[PULL 23/26] target/arm: Use FPST_F16 for SME FMOPA (widening)
+[PULL 39/72] target/arm: Copy entire float_status in is_ebf
 From: Richard Henderson <richard.henderson@linaro.org>
-This operation has float16 inputs and thus must use
+Now that float_status has a bunch of fp parameters,
-the FZ16 control not the FZ control.
+it is easier to copy an existing structure than create
 one from scratch.  Begin by copying the structure that
 corresponds to the FPSR and make only the adjustments
 required for BFloat16 semantics.
-Cc: qemu-stable@nongnu.org
-Fixes: 3916841ac75 ("target/arm: Implement FMOPA, FMOPS (widening)")
-Reported-by: Daniyal Khan <danikhan632@gmail.com>
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240717060149.204788-3-richard.henderson@linaro.org
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Resolves: https://gitlab.com/qemu-project/qemu/-/issues/2374
+Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/translate-sme.c | 12 ++++++++----
+ target/arm/tcg/vec_helper.c | 20 +++++++-------------
-file changed, 8 insertions(+), 4 deletions(-)
+file changed, 7 insertions(+), 13 deletions(-)
-diff --git a/target/arm/tcg/translate-sme.c b/target/arm/tcg/translate-sme.c
+diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-sme.c
+--- a/target/arm/tcg/vec_helper.c
-+++ b/target/arm/tcg/translate-sme.c
++++ b/target/arm/tcg/vec_helper.c
-@@ -XXX,XX +XXX,XX @@ static bool do_outprod(DisasContext *s, arg_op *a, MemOp esz,
+@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
       * no effect on AArch32 instructions.
       */
      bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
 -    *statusp = (float_status){
 -        .tininess_before_rounding = float_tininess_before_rounding,
 -        .float_rounding_mode = float_round_to_odd_inf,
 -        .flush_to_zero = true,
 -        .flush_inputs_to_zero = true,
 -        .default_nan_mode = true,
 -    };
 +
 +    *statusp = env->vfp.fp_status;
 +    set_default_nan_mode(true, statusp);
      if (ebf) {
 -        float_status *fpst = &env->vfp.fp_status;
 -        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
 -        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
 -        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
 -
          /* EBF=1 needs to do a step with round-to-odd semantics */
          *oddstatusp = *statusp;
          set_float_rounding_mode(float_round_to_odd, oddstatusp);
 +    } else {
 +        set_flush_to_zero(true, statusp);
 +        set_flush_inputs_to_zero(true, statusp);
 +        set_float_rounding_mode(float_round_to_odd_inf, statusp);
      }
 -
      return ebf;
  }
- static bool do_outprod_fpst(DisasContext *s, arg_op *a, MemOp esz,
-+                            ARMFPStatusFlavour e_fpst,
-                             gen_helper_gvec_5_ptr *fn)
- {
-     int svl = streaming_vec_reg_size(s);
-@@ -XXX,XX +XXX,XX @@ static bool do_outprod_fpst(DisasContext *s, arg_op *a, MemOp esz,
-     zm = vec_full_reg_ptr(s, a->zm);
-     pn = pred_full_reg_ptr(s, a->pn);
-     pm = pred_full_reg_ptr(s, a->pm);
--    fpst = fpstatus_ptr(FPST_FPCR);
-+    fpst = fpstatus_ptr(e_fpst);
-     fn(za, zn, zm, pn, pm, fpst, tcg_constant_i32(desc));
-     return true;
- }
--TRANS_FEAT(FMOPA_h, aa64_sme, do_outprod_fpst, a, MO_32, gen_helper_sme_fmopa_h)
--TRANS_FEAT(FMOPA_s, aa64_sme, do_outprod_fpst, a, MO_32, gen_helper_sme_fmopa_s)
--TRANS_FEAT(FMOPA_d, aa64_sme_f64f64, do_outprod_fpst, a, MO_64, gen_helper_sme_fmopa_d)
-+TRANS_FEAT(FMOPA_h, aa64_sme, do_outprod_fpst, a,
-+           MO_32, FPST_FPCR_F16, gen_helper_sme_fmopa_h)
-+TRANS_FEAT(FMOPA_s, aa64_sme, do_outprod_fpst, a,
-+           MO_32, FPST_FPCR, gen_helper_sme_fmopa_s)
-+TRANS_FEAT(FMOPA_d, aa64_sme_f64f64, do_outprod_fpst, a,
-+           MO_64, FPST_FPCR, gen_helper_sme_fmopa_d)
- /* TODO: FEAT_EBF16 */
- TRANS_FEAT(BFMOPA, aa64_sme, do_outprod, a, MO_32, gen_helper_sme_bfmopa)
 --
 .34.1

-[PULL 15/26] hw/arm/smmu: Support nesting in smmuv3_range_inval()
+[PULL 40/72] fpu: Allow runtime choice of default NaN value
-From: Mostafa Saleh <smostafa@google.com>
+Currently we hardcode the default NaN value in parts64_default_nan()
 using a compile-time ifdef ladder. This is awkward for two cases:
  * for single-QEMU-binary we can't hard-code target-specifics like this
  * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
    (specifically the sign bit is different)
-With nesting, we would need to invalidate IPAs without
+Add a field to float_status to specify the default NaN value; fall
-over-invalidating stage-1 IOVAs. This can be done by
+back to the old ifdef behaviour if these are not set.
 distinguishing IPAs in the TLBs by having ASID=-1.
 To achieve that, rework the invalidation for IPAs to have a
 separate function, while for IOVA invalidation ASID=-1 means
 invalidate for all ASIDs.
-Reviewed-by: Eric Auger <eric.auger@redhat.com>
+The default NaN value is specified by setting a uint8_t to a
-Signed-off-by: Mostafa Saleh <smostafa@google.com>
+pattern corresponding to the sign and upper fraction parts of
-Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
+the NaN; the lower bits of the fraction are set from bit 0 of
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+the pattern.
-Message-id: 20240715084519.1189624-13-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
 ---
- include/hw/arm/smmu-common.h |  3 ++-
+ include/fpu/softfloat-helpers.h | 11 +++++++
- hw/arm/smmu-common.c         | 47 ++++++++++++++++++++++++++++++++++++
+ include/fpu/softfloat-types.h   | 10 ++++++
- hw/arm/smmuv3.c              | 23 ++++++++++++------
+ fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
- hw/arm/trace-events          |  2 +-
+files changed, 54 insertions(+), 22 deletions(-)
 files changed, 66 insertions(+), 9 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/include/fpu/softfloat-helpers.h
-+++ b/include/hw/arm/smmu-common.h
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_asid(SMMUState *s, int asid);
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
- void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
+     status->float_infzeronan_rule = rule;
  void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                           uint8_t tg, uint64_t num_pages, uint8_t ttl);
 -
 +void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
 +                        uint64_t num_pages, uint8_t ttl);
  /* Unmap the range of all the notifiers registered to any IOMMU mr */
  void smmu_inv_notifiers_all(SMMUState *s);
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid_vmid_iova(gpointer key, gpointer value,
             ((entry->iova & ~info->mask) == info->iova);
  }
-+static gboolean smmu_hash_remove_by_vmid_ipa(gpointer key, gpointer value,
++static inline void set_float_default_nan_pattern(uint8_t dnan_pattern,
-+                                             gpointer user_data)
++                                                 float_status *status)
 +{
-+    SMMUTLBEntry *iter = (SMMUTLBEntry *)value;
++    status->default_nan_pattern = dnan_pattern;
 +    IOMMUTLBEntry *entry = &iter->entry;
 +    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
 +    SMMUIOTLBKey iotlb_key = *(SMMUIOTLBKey *)key;
 +
 +    if (SMMU_IOTLB_ASID(iotlb_key) >= 0) {
 +        /* This is a stage-1 address. */
 +        return false;
 +    }
 +    if (info->vmid != SMMU_IOTLB_VMID(iotlb_key)) {
 +        return false;
 +    }
 +    return ((info->iova & ~entry->addr_mask) == entry->iova) ||
 +           ((entry->iova & ~info->mask) == info->iova);
 +}
 +
- void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
+ static inline void set_flush_to_zero(bool val, float_status *status)
                           uint8_t tg, uint64_t num_pages, uint8_t ttl)
  {
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
+     status->flush_to_zero = val;
-                                 &info);
+@@ -XXX,XX +XXX,XX @@ static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status
      return status->float_infzeronan_rule;
  }
-+/*
++static inline uint8_t get_float_default_nan_pattern(float_status *status)
 + * Similar to smmu_iotlb_inv_iova(), but for Stage-2, ASID is always -1,
 + * in Stage-1 invalidation ASID = -1, means don't care.
 + */
 +void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
 +                        uint64_t num_pages, uint8_t ttl)
 +{
-+    uint8_t granule = tg ? tg * 2 + 10 : 12;
++    return status->default_nan_pattern;
 +    int asid = -1;
 +
 +   if (ttl && (num_pages == 1)) {
 +        SMMUIOTLBKey key = smmu_get_iotlb_key(asid, vmid, ipa, tg, ttl);
 +
 +        if (g_hash_table_remove(s->iotlb, &key)) {
 +            return;
 +        }
 +    }
 +
 +    SMMUIOTLBPageInvInfo info = {
 +        .iova = ipa,
 +        .vmid = vmid,
 +        .mask = (num_pages << granule) - 1};
 +
 +    g_hash_table_foreach_remove(s->iotlb,
 +                                smmu_hash_remove_by_vmid_ipa,
 +                                &info);
 +}
 +
- void smmu_iotlb_inv_asid(SMMUState *s, int asid)
+ static inline bool get_flush_to_zero(float_status *status)
  {
-     trace_smmu_iotlb_inv_asid(asid);
+     return status->flush_to_zero;
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/include/fpu/softfloat-types.h
-+++ b/hw/arm/smmuv3.c
++++ b/include/fpu/softfloat-types.h
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
-     }
+     /* should denormalised inputs go to zero and set the input_denormal flag? */
- }
+     bool flush_inputs_to_zero;
+     bool default_nan_mode;
--static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
++    /*
-+static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
++     * The pattern to use for the default NaN. Here the high bit specifies
 +     * the default NaN's sign bit, and bits 6..0 specify the high bits of the
 +     * fractional part. The low bits of the fractional part are copies of bit 0.
 +     * The exponent of the default NaN is (as for any NaN) always all 1s.
 +     * Note that a value of 0 here is not a valid NaN. The target must set
 +     * this to the correct non-zero value, or we will assert when trying to
 +     * create a default NaN.
 +     */
 +    uint8_t default_nan_pattern;
      /*
       * The flags below are not used on all specializations and may
       * constant fold away (see snan_bit_is_one()/no_signalling_nans() in
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
  {
-     dma_addr_t end, addr = CMD_ADDR(cmd);
+     bool sign = 0;
-     uint8_t type = CMD_TYPE(cmd);
+     uint64_t frac;
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
++    uint8_t dnan_pattern = status->default_nan_pattern;
-     }
++    if (dnan_pattern == 0) {
-     if (!tg) {
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
--        trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf);
+-    /* !snan_bit_is_one, set all bits */
-+        trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf, stage);
+-    frac = (1ULL << DECOMPOSED_BINARY_POINT) - 1;
-         smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1);
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64) \
--        smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
++        /* Sign bit clear, all frac bits set */
-+        if (stage == SMMU_STAGE_1) {
++        dnan_pattern = 0b01111111;
-+            smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
++#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
      || defined(TARGET_MICROBLAZE)
 -    /* !snan_bit_is_one, set sign and msb */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    sign = 1;
 +        /* Sign bit set, most significant frac bit set */
 +        dnan_pattern = 0b11000000;
  #elif defined(TARGET_HPPA)
 -    /* snan_bit_is_one, set msb-1.  */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 2);
 +        /* Sign bit clear, msb-1 frac bit set */
 +        dnan_pattern = 0b00100000;
  #elif defined(TARGET_HEXAGON)
 -    sign = 1;
 -    frac = ~0ULL;
 +        /* Sign bit set, all frac bits set. */
 +        dnan_pattern = 0b11111111;
  #else
 -    /*
 -     * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 -     * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 -     * do not have floating-point.
 -     */
 -    if (snan_bit_is_one(status)) {
 -        /* set all bits other than msb */
 -        frac = (1ULL << (DECOMPOSED_BINARY_POINT - 1)) - 1;
 -    } else {
 -        /* set msb */
 -        frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    }
 +        /*
 +         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 +         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 +         * do not have floating-point.
 +         */
 +        if (snan_bit_is_one(status)) {
 +            /* sign bit clear, set all frac bits other than msb */
 +            dnan_pattern = 0b00111111;
 +        } else {
-+            smmu_iotlb_inv_ipa(s, vmid, addr, tg, 1, ttl);
++            /* sign bit clear, set frac msb */
 +            dnan_pattern = 0b01000000;
 +        }
-         return;
+ #endif
-     }
++    }
++    assert(dnan_pattern != 0);
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
++
-         uint64_t mask = dma_aligned_pow2_mask(addr, end, 64);
++    sign = dnan_pattern >> 7;
++    /*
-         num_pages = (mask + 1) >> granule;
++     * Place default_nan_pattern [6:0] into bits [62:56],
--        trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages, ttl, leaf);
++     * and replecate bit [0] down into [55:0]
-+        trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages,
++     */
-+                                 ttl, leaf, stage);
++    frac = deposit64(0, DECOMPOSED_BINARY_POINT - 7, 7, dnan_pattern);
-         smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages);
++    frac = deposit64(frac, 0, DECOMPOSED_BINARY_POINT - 7, -(dnan_pattern & 1));
--        smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
-+        if (stage == SMMU_STAGE_1) {
+     *p = (FloatParts64) {
-+            smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
+         .cls = float_class_qnan,
 +        } else {
 +            smmu_iotlb_inv_ipa(s, vmid, addr, tg, num_pages, ttl);
 +        }
          addr += mask + 1;
      }
  }
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
                  cmd_error = SMMU_CERROR_ILL;
                  break;
              }
 -            smmuv3_range_inval(bs, &cmd);
 +            smmuv3_range_inval(bs, &cmd, SMMU_STAGE_1);
              break;
          case SMMU_CMD_TLBI_S12_VMALL:
          {
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
               * As currently only either s1 or s2 are supported
               * we can reuse same function for s2.
               */
 -            smmuv3_range_inval(bs, &cmd);
 +            smmuv3_range_inval(bs, &cmd, SMMU_STAGE_2);
              break;
          case SMMU_CMD_TLBI_EL3_ALL:
          case SMMU_CMD_TLBI_EL3_VA:
 diff --git a/hw/arm/trace-events b/hw/arm/trace-events
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/trace-events
 +++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%x - end=0x%x"
  smmuv3_cmdq_cfgi_cd(uint32_t sid) "sid=0x%x"
  smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
  smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
 -smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
 +smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf, int stage) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d stage=%d"
  smmuv3_cmdq_tlbi_nh(void) ""
  smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
  smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
 --
 .34.1

-New patch
+[PULL 41/72] tests/fp: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the tests/fp code.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c     | 1 +
+ tests/fp/fp-test-log2.c | 1 +
+ tests/fp/fp-test.c      | 1 +
+files changed, 3 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
++    set_float_default_nan_pattern(0b01000000, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test-log2.c
++++ b/tests/fp/fp-test-log2.c
+@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
+     int i;
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_rounding_mode(float_round_nearest_even, &qsf);
+     test.d = 0.0;
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-New patch
+[PULL 42/72] target/microblaze: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
+---
+ target/microblaze/cpu.c        | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 +--
+files changed, 3 insertions(+), 2 deletions(-)
+diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/microblaze/cpu.c
++++ b/target/microblaze/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
+      * this architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN: sign bit set, most significant frac bit set */
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
+ #if defined(CONFIG_USER_ONLY)
+     /* start in user mode with interrupts enabled.  */
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
+-    || defined(TARGET_MICROBLAZE)
++#elif defined(TARGET_I386) || defined(TARGET_X86_64)
+         /* Sign bit set, most significant frac bit set */
+         dnan_pattern = 0b11000000;
+ #elif defined(TARGET_HPPA)
+--
+.34.1

-New patch
+[PULL 43/72] target/i386: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 4 ++++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 4 insertions(+), 3 deletions(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
++    /* Default NaN: sign bit set, most significant frac bit set */
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
++    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
++    set_float_default_nan_pattern(0b11000000, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
+-        /* Sign bit set, most significant frac bit set */
+-        dnan_pattern = 0b11000000;
+ #elif defined(TARGET_HPPA)
+         /* Sign bit clear, msb-1 frac bit set */
+         dnan_pattern = 0b00100000;
+--
+.34.1

-New patch
+[PULL 44/72] target/hppa: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 2 insertions(+), 3 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN: sign bit clear, msb-1 frac bit set */
++    set_float_default_nan_pattern(0b00100000, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HPPA)
+-        /* Sign bit clear, msb-1 frac bit set */
+-        dnan_pattern = 0b00100000;
+ #elif defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+--
+.34.1

-New patch
+[PULL 45/72] target/alpha: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the alpha target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-40-peter.maydell@linaro.org
+---
+ target/alpha/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/alpha/cpu.c b/target/alpha/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/alpha/cpu.c
++++ b/target/alpha/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void alpha_cpu_initfn(Object *obj)
+      * operand in Fa. That is float_2nan_prop_ba.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ #if defined(CONFIG_USER_ONLY)
+     env->flags = ENV_FLAG_PS_USER | ENV_FLAG_FEN;
+     cpu_alpha_store_fpcr(env, (uint64_t)(FPCR_INVD | FPCR_DZED | FPCR_OVFD
+--
+.34.1

-New patch
+[PULL 46/72] target/arm: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the arm target.
+This includes setting it for the old linux-user nwfpe emulation.
+For nwfpe, our default doesn't match the real kernel, but we
+avoid making a behaviour change in this commit.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
+---
+ linux-user/arm/nwfpe/fpa11.c | 5 +++++
+ target/arm/cpu.c             | 2 ++
+files changed, 7 insertions(+)
+diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
+index XXXXXXX..XXXXXXX 100644
+--- a/linux-user/arm/nwfpe/fpa11.c
++++ b/linux-user/arm/nwfpe/fpa11.c
+@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
+    * this late date.
+    */
+   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
++  /*
++   * Use the same default NaN value as Arm VFP. This doesn't match
++   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
++   */
++  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
+ }
+ void SetRoundingMode(const unsigned int opcode)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *    the pseudocode function the arguments are in the order c, a, b.
+  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+  *    and the input NaN if it is signalling
++ *  * Default NaN has sign bit clear, msb frac bit set
+  */
+ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
++    set_float_default_nan_pattern(0b01000000, s);
+ }
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
+--
+.34.1

-New patch
+[PULL 47/72] target/loongarch: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for loongarch.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-42-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+--
+.34.1

-New patch
+[PULL 48/72] target/m68k: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for m68k.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c              | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++    /* Default NaN: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     nan = floatx80_default_nan(&env->fp_status);
+     for (i = 0; i < 8; i++) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
++#if defined(TARGET_SPARC)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+ #elif defined(TARGET_HEXAGON)
+--
+.34.1

-New patch
+[PULL 49/72] target/mips: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for MIPS. Note that this
+is our only target which currently changes the default NaN
+at runtime (which it was previously doing indirectly when it
+changed the snan_bit_is_one setting).
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h | 7 +++++++
+ target/mips/msa.c        | 3 +++
+files changed, 10 insertions(+)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
++    /*
++     * With nan2008, the default NaN value has the sign bit clear and the
++     * frac msb set; with the older mode, the sign bit is clear, and all
++     * frac bits except the msb are set.
++     */
++    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
++                                  &env->active_fpu.fp_status);
+ }
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* Inf * 0 + NaN returns the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                               &env->active_tc.msa_fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000,
++                                  &env->active_tc.msa_fp_status);
+ }
+--
+.34.1

-New patch
+[PULL 50/72] target/openrisc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for openrisc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-45-peter.maydell@linaro.org
+---
+ target/openrisc/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/openrisc/cpu.c b/target/openrisc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/openrisc/cpu.c
++++ b/target/openrisc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void openrisc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &cpu->env.fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000, &cpu->env.fp_status);
+ #ifndef CONFIG_USER_ONLY
+     cpu->env.picmr = 0x00000000;
+--
+.34.1

-New patch
+[PULL 51/72] target/ppc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for ppc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-46-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c | 4 ++++
+file changed, 4 insertions(+)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
++    /* Default NaN: sign bit clear, set frac msb */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
++    set_float_default_nan_pattern(0b01000000, &env->vec_status);
++
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+--
+.34.1

-New patch
+[PULL 52/72] target/sh4: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for sh4. Note that sh4
+is one of the only three targets (the others being HPPA and
+sometimes MIPS) that has snan_bit_is_one set.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-47-peter.maydell@linaro.org
+---
+ target/sh4/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/sh4/cpu.c b/target/sh4/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sh4/cpu.c
++++ b/target/sh4/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void superh_cpu_reset_hold(Object *obj, ResetType type)
+     set_flush_to_zero(1, &env->fp_status);
+ #endif
+     set_default_nan_mode(1, &env->fp_status);
++    /* sign bit clear, set all frac bits other than msb */
++    set_float_default_nan_pattern(0b00111111, &env->fp_status);
+ }
+ static void superh_cpu_disas_set_info(CPUState *cpu, disassemble_info *info)
+--
+.34.1

-New patch
+[PULL 53/72] target/rx: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for rx.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-48-peter.maydell@linaro.org
+---
+ target/rx/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/rx/cpu.c b/target/rx/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/rx/cpu.c
++++ b/target/rx/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void rx_cpu_reset_hold(Object *obj, ResetType type)
+      * then prefer dest over source", which is float_2nan_prop_s_ab.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN value: sign bit clear, set frac msb */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ }
+ static ObjectClass *rx_cpu_class_by_name(const char *cpu_model)
+--
+.34.1

-New patch
+[PULL 54/72] target/s390x: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for s390x.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-49-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
++        /* Default NaN value: sign bit clear, frac msb set */
++        set_float_default_nan_pattern(0b01000000, &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+--
+.34.1

-New patch
+[PULL 55/72] target/sparc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for SPARC, and remove
+the ifdef from parts64_default_nan.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 5 +----
+files changed, 3 insertions(+), 4 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN value: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC)
+-        /* Sign bit clear, all frac bits set */
+-        dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HEXAGON)
++#if defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+ #else
+--
+.34.1

-[PULL 03/26] hw/display/bcm2835_fb: fix fb_use_offsets condition
+[PULL 56/72] target/xtensa: Set default NaN pattern explicitly
-From: SamJakob <me@samjakob.com>
+Set the default NaN pattern explicitly for xtensa.
-It is common practice when implementing double-buffering on VideoCore
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-to do so by multiplying the height of the virtual buffer by the
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-number of virtual screens desired (i.e., two - in the case of
+Message-id: 20241202131347.498124-51-peter.maydell@linaro.org
-double-bufferring).
+---
  target/xtensa/cpu.c | 2 ++
 file changed, 2 insertions(+)
-At present, this won't work in QEMU because the logic in
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
 fb_use_offsets require that both the virtual width and height exceed
 their physical counterparts.
 This appears to be unintentional/a typo and indeed the comment
 states; "Experimentally, the hardware seems to do this only if the
 viewport size is larger than the physical screen".  The
 viewport/virtual size would be larger than the physical size if
 either virtual dimension were larger than their physical counterparts
 and not necessarily both.
 Signed-off-by: SamJakob <me@samjakob.com>
 Message-id: 20240713160353.62410-1-me@samjakob.com
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  hw/display/bcm2835_fb.c | 2 +-
 file changed, 1 insertion(+), 1 deletion(-)
 diff --git a/hw/display/bcm2835_fb.c b/hw/display/bcm2835_fb.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/display/bcm2835_fb.c
+--- a/target/xtensa/cpu.c
-+++ b/hw/display/bcm2835_fb.c
++++ b/target/xtensa/cpu.c
-@@ -XXX,XX +XXX,XX @@ static bool fb_use_offsets(BCM2835FBConfig *config)
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
-      * viewport size is larger than the physical screen. (It doesn't
+     /* For inf * 0 + NaN, return the input NaN */
-      * prevent the guest setting this silly viewport setting, though...)
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
-      */
+     set_no_signaling_nans(!dfpu, &env->fp_status);
--    return config->xres_virtual > config->xres &&
++    /* Default NaN value: sign bit clear, set frac msb */
-+    return config->xres_virtual > config->xres ||
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
-         config->yres_virtual > config->yres;
+     xtensa_use_first_nan(env, !dfpu);
  }
 --
 .34.1

-New patch
+[PULL 57/72] target/hexagon: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for hexagon.
+Remove the ifdef from parts64_default_nan(); the only
+remaining unconverted targets all use the default case.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
+---
+ target/hexagon/cpu.c           | 2 ++
+ fpu/softfloat-specialize.c.inc | 5 -----
+files changed, 2 insertions(+), 5 deletions(-)
+diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hexagon/cpu.c
++++ b/target/hexagon/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
+     set_default_nan_mode(1, &env->fp_status);
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
++    /* Default NaN value: sign bit set, all frac bits set */
++    set_float_default_nan_pattern(0b11111111, &env->fp_status);
+ }
+ static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_HEXAGON)
+-        /* Sign bit set, all frac bits set. */
+-        dnan_pattern = 0b11111111;
+-#else
+         /*
+          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
+          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+             /* sign bit clear, set frac msb */
+             dnan_pattern = 0b01000000;
+         }
+-#endif
+     }
+     assert(dnan_pattern != 0);
+--
+.34.1

-New patch
+[PULL 58/72] target/riscv: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for riscv.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-53-peter.maydell@linaro.org
+---
+ target/riscv/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/riscv/cpu.c
++++ b/target/riscv/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void riscv_cpu_reset_hold(Object *obj, ResetType type)
+     cs->exception_index = RISCV_EXCP_NONE;
+     env->load_res = -1;
+     set_default_nan_mode(1, &env->fp_status);
++    /* Default NaN value: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+     env->vill = true;
+ #ifndef CONFIG_USER_ONLY
+--
+.34.1

-[PULL 13/26] hw/arm/smmu-common: Add support for nested TLB
+[PULL 59/72] target/tricore: Set default NaN pattern explicitly
-From: Mostafa Saleh <smostafa@google.com>
+Set the default NaN pattern explicitly for tricore.
-This patch adds support for nested (combined) TLB entries.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-The main function combine_tlb() is not used here but in the next
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-patches, but to simplify the patches it is introduced first.
+Message-id: 20241202131347.498124-54-peter.maydell@linaro.org
 ---
  target/tricore/helper.c | 2 ++
 file changed, 2 insertions(+)
-Main changes:
+diff --git a/target/tricore/helper.c b/target/tricore/helper.c
 ) New field added in the SMMUTLBEntry struct: parent_perm, for
    nested TLB, holds the stage-2 permission, this can be used to know
    the origin of a permission fault from a cached entry as caching
    the “and” of the permissions loses this information.
    SMMUPTWEventInfo is used to hold information about PTW faults so
    the event can be populated, the value of stage used to be set
    based on the current stage for TLB permission faults, however
    with the parent_perm, it is now set based on which perm has
    the missing permission
    When nesting is not enabled it has the same value as perm which
    doesn't change the logic.
 ) As combined TLB implementation is used, the combination logic
    chooses:
    - tg and level from the entry which has the smallest addr_mask.
    - Based on that the iova that would be cached is recalculated.
    - Translated_addr is chosen from stage-2.
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-11-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  include/hw/arm/smmu-common.h |  1 +
  hw/arm/smmu-common.c         | 37 ++++++++++++++++++++++++++++++++----
 files changed, 34 insertions(+), 4 deletions(-)
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/target/tricore/helper.c
-+++ b/include/hw/arm/smmu-common.h
++++ b/target/tricore/helper.c
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTLBEntry {
+@@ -XXX,XX +XXX,XX @@ void fpu_set_state(CPUTriCoreState *env)
-     IOMMUTLBEntry entry;
+     set_flush_to_zero(1, &env->fp_status);
-     uint8_t level;
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
-     uint8_t granule;
+     set_default_nan_mode(1, &env->fp_status);
-+    IOMMUAccessFlags parent_perm;
++    /* Default NaN pattern: sign bit clear, frac msb set */
- } SMMUTLBEntry;
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
  /* Stage-2 configuration. */
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
          tlbe->entry.translated_addr = gpa;
          tlbe->entry.iova = iova & ~mask;
          tlbe->entry.addr_mask = mask;
 -        tlbe->entry.perm = PTE_AP_TO_PERM(ap);
 +        tlbe->parent_perm = PTE_AP_TO_PERM(ap);
 +        tlbe->entry.perm = tlbe->parent_perm;
          tlbe->level = level;
          tlbe->granule = granule_sz;
          return 0;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
          tlbe->entry.translated_addr = gpa;
          tlbe->entry.iova = ipa & ~mask;
          tlbe->entry.addr_mask = mask;
 -        tlbe->entry.perm = s2ap;
 +        tlbe->parent_perm = s2ap;
 +        tlbe->entry.perm = tlbe->parent_perm;
          tlbe->level = level;
          tlbe->granule = granule_sz;
          return 0;
@@ -XXX,XX +XXX,XX @@ error:
      return -EINVAL;
  }
-+/*
+ uint32_t psw_read(CPUTriCoreState *env)
 + * combine S1 and S2 TLB entries into a single entry.
 + * As a result the S1 entry is overriden with combined data.
 + */
 +static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
 +                                                SMMUTLBEntry *tlbe_s2,
 +                                                dma_addr_t iova,
 +                                                SMMUTransCfg *cfg)
 +{
 +    if (tlbe_s2->entry.addr_mask < tlbe->entry.addr_mask) {
 +        tlbe->entry.addr_mask = tlbe_s2->entry.addr_mask;
 +        tlbe->granule = tlbe_s2->granule;
 +        tlbe->level = tlbe_s2->level;
 +    }
 +
 +    tlbe->entry.translated_addr = CACHED_ENTRY_TO_ADDR(tlbe_s2,
 +                                    tlbe->entry.translated_addr);
 +
 +    tlbe->entry.iova = iova & ~tlbe->entry.addr_mask;
 +    /* parent_perm has s2 perm while perm keeps s1 perm. */
 +    tlbe->parent_perm = tlbe_s2->entry.perm;
 +    return;
 +}
 +
  /**
   * smmu_ptw - Walk the page tables for an IOVA, according to @cfg
   *
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
      cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, addr);
      if (cached_entry) {
 -        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
 +        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm &
 +            cached_entry->parent_perm & IOMMU_WO)) {
              info->type = SMMU_PTW_ERR_PERMISSION;
 -            info->stage = cfg->stage;
 +            info->stage = !(cached_entry->entry.perm & IOMMU_WO) ?
 +                          SMMU_STAGE_1 :
 +                          SMMU_STAGE_2;
              return NULL;
          }
          return cached_entry;
 --
 .34.1

-[PULL 02/26] target/arm: LDAPR should honour SCTLR_ELx.nAA
+[PULL 60/72] fpu: Remove default handling for dnan_pattern
-In commit c1a1f80518d360b when we added the FEAT_LSE2 relaxations to
+Now that all our targets have bene converted to explicitly specify
-the alignment requirements for atomic and ordered loads and stores,
+their pattern for the default NaN value we can remove the remaining
-we didn't quite get it right for LDAPR/LDAPRH/LDAPRB with no
+fallback code in parts64_default_nan().
 immediate offset.  These instructions were handled in the old decoder
 as part of disas_ldst_atomic(), but unlike all the other insns that
 function decoded (LDADD, LDCLR, etc) these insns are "ordered", not
 "atomic", so they should be using check_ordered_align() rather than
 check_atomic_align().  Commit c1a1f80518d360b used
 check_atomic_align() regardless for everything in
 disas_ldst_atomic().  We then carried that incorrect check over in
 the decodetree conversion, where LDAPR/LDAPRH/LDAPRB are now handled
 by trans_LDAPR().
-The effect is that when FEAT_LSE2 is implemented, these instructions
-don't honour the SCTLR_ELx.nAA bit and will generate alignment
-faults when they should not.
-(The LDAPR insns with an immediate offset were in disas_ldst_ldapr_stlr()
-and then in trans_LDAPR_i() and trans_STLR_i(), and have always used
-the correct check_ordered_align().)
-Use check_ordered_align() in trans_LDAPR().
-Cc: qemu-stable@nongnu.org
-Fixes: c1a1f80518d360b ("target/arm: Relax ordered/atomic alignment checks for LSE2")
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240709134504.3500007-3-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
 ---
- target/arm/tcg/translate-a64.c | 2 +-
+ fpu/softfloat-specialize.c.inc | 14 --------------
-file changed, 1 insertion(+), 1 deletion(-)
+file changed, 14 deletions(-)
-diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-a64.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/tcg/translate-a64.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static bool trans_LDAPR(DisasContext *s, arg_LDAPR *a)
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
-     if (a->rn == 31) {
+     uint64_t frac;
-         gen_check_sp_alignment(s);
+     uint8_t dnan_pattern = status->default_nan_pattern;
-     }
--    mop = check_atomic_align(s, a->rn, a->sz);
+-    if (dnan_pattern == 0) {
-+    mop = check_ordered_align(s, a->rn, 0, false, a->sz);
+-        /*
-     clean_addr = gen_mte_check1(s, cpu_reg_sp(s, a->rn), false,
+-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-                                 a->rn != 31, mop);
+-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
-     /*
+-         * do not have floating-point.
 -         */
 -        if (snan_bit_is_one(status)) {
 -            /* sign bit clear, set all frac bits other than msb */
 -            dnan_pattern = 0b00111111;
 -        } else {
 -            /* sign bit clear, set frac msb */
 -            dnan_pattern = 0b01000000;
 -        }
 -    }
      assert(dnan_pattern != 0);
      sign = dnan_pattern >> 7;
 --
 .34.1

-[PULL 08/26] hw/arm/smmu: Split smmuv3_translate()
+[PULL 61/72] softfloat: Inline pickNaNMulAdd
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-smmuv3_translate() does everything from STE/CD parsing to TLB lookup
+Inline pickNaNMulAdd into its only caller.  This makes
-and PTW.
+one assert redundant with the immediately preceding IF.
-Soon, when nesting is supported, stage-1 data (tt, CD) needs to be
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-translated using stage-2.
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
-Split smmuv3_translate() to 3 functions:
+[PMM: keep comment from old code in new location]
 - smmu_translate(): in smmu-common.c, which does the TLB lookup, PTW,
   TLB insertion, all the functions are already there, this just puts
   them together.
   This also simplifies the code as it consolidates event generation
   in case of TLB lookup permission failure or in TT selection.
 - smmuv3_do_translate(): in smmuv3.c, Calls smmu_translate() and does
   the event population in case of errors.
 - smmuv3_translate(), now calls smmuv3_do_translate() for
   translation while the rest is the same.
 Also, add stage in trace_smmuv3_translate_success()
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-6-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/hw/arm/smmu-common.h |   8 ++
+ fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
- hw/arm/smmu-common.c         |  59 +++++++++++
+ fpu/softfloat-specialize.c.inc | 54 ----------------------------------
- hw/arm/smmuv3.c              | 194 +++++++++++++----------------------
+files changed, 40 insertions(+), 55 deletions(-)
  hw/arm/trace-events          |   2 +-
 files changed, 142 insertions(+), 121 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
- int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
+     }
-              SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
+     if (s->default_nan_mode) {
-+
++        /*
-+/*
++         * We guarantee not to require the target to tell us how to
-+ * smmu_translate - Look for a translation in TLB, if not, do a PTW.
++         * pick a NaN if we're always returning the default NaN.
-+ * Returns NULL on PTW error or incase of TLB permission errors.
++         * But if we're not in default-NaN mode then the target must
-+ */
++         * specify.
-+SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
++         */
-+                             IOMMUAccessFlags flag, SMMUPTWEventInfo *info);
+         which = 3;
-+
++    } else if (infzero) {
- /**
++        /*
-  * select_tt - compute which translation table shall be used according to
++         * Inf * 0 + NaN -- some implementations return the
-  * the input iova and translation config and return the TT specific info
++         * default NaN here, and some return the input NaN.
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
++         */
-index XXXXXXX..XXXXXXX 100644
++        switch (s->float_infzeronan_rule) {
---- a/hw/arm/smmu-common.c
++        case float_infzeronan_dnan_never:
-+++ b/hw/arm/smmu-common.c
++            which = 2;
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
      g_assert_not_reached();
  }
 +SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
 +                             IOMMUAccessFlags flag, SMMUPTWEventInfo *info)
 +{
 +    uint64_t page_mask, aligned_addr;
 +    SMMUTLBEntry *cached_entry = NULL;
 +    SMMUTransTableInfo *tt;
 +    int status;
 +
 +    /*
 +     * Combined attributes used for TLB lookup, as only one stage is supported,
 +     * it will hold attributes based on the enabled stage.
 +     */
 +    SMMUTransTableInfo tt_combined;
 +
 +    if (cfg->stage == SMMU_STAGE_1) {
 +        /* Select stage1 translation table. */
 +        tt = select_tt(cfg, addr);
 +        if (!tt) {
 +            info->type = SMMU_PTW_ERR_TRANSLATION;
 +            info->stage = SMMU_STAGE_1;
 +            return NULL;
 +        }
 +        tt_combined.granule_sz = tt->granule_sz;
 +        tt_combined.tsz = tt->tsz;
 +
 +    } else {
 +        /* Stage2. */
 +        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
 +        tt_combined.tsz = cfg->s2cfg.tsz;
 +    }
 +
 +    /*
 +     * TLB lookup looks for granule and input size for a translation stage,
 +     * as only one stage is supported right now, choose the right values
 +     * from the configuration.
 +     */
 +    page_mask = (1ULL << tt_combined.granule_sz) - 1;
 +    aligned_addr = addr & ~page_mask;
 +
 +    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
 +    if (cached_entry) {
 +        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
 +            info->type = SMMU_PTW_ERR_PERMISSION;
 +            info->stage = cfg->stage;
 +            return NULL;
 +        }
 +        return cached_entry;
 +    }
 +
 +    cached_entry = g_new0(SMMUTLBEntry, 1);
 +    status = smmu_ptw(cfg, aligned_addr, flag, cached_entry, info);
 +    if (status) {
 +            g_free(cached_entry);
 +            return NULL;
 +    }
 +    smmu_iotlb_insert(bs, cfg, cached_entry);
 +    return cached_entry;
 +}
 +
  /**
   * The bus number is used for lookup when SID based invalidation occurs.
   * In that case we lazily populate the SMMUPciBus array from the bus hash
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_flush_config(SMMUDevice *sdev)
      g_hash_table_remove(bc->configs, sdev);
  }
 +/* Do translation with TLB lookup. */
 +static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
 +                                                 SMMUTransCfg *cfg,
 +                                                 SMMUEventInfo *event,
 +                                                 IOMMUAccessFlags flag,
 +                                                 SMMUTLBEntry **out_entry)
 +{
 +    SMMUPTWEventInfo ptw_info = {};
 +    SMMUState *bs = ARM_SMMU(s);
 +    SMMUTLBEntry *cached_entry = NULL;
 +
 +    cached_entry = smmu_translate(bs, cfg, addr, flag, &ptw_info);
 +    if (!cached_entry) {
 +        /* All faults from PTW has S2 field. */
 +        event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
 +        switch (ptw_info.type) {
 +        case SMMU_PTW_ERR_WALK_EABT:
 +            event->type = SMMU_EVT_F_WALK_EABT;
 +            event->u.f_walk_eabt.addr = addr;
 +            event->u.f_walk_eabt.rnw = flag & 0x1;
 +            event->u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
 +                                          SMMU_CLASS_IN : SMMU_CLASS_TT;
 +            event->u.f_walk_eabt.addr2 = ptw_info.addr;
 +            break;
-+        case SMMU_PTW_ERR_TRANSLATION:
++        case float_infzeronan_dnan_always:
-+            if (PTW_RECORD_FAULT(cfg)) {
++            which = 3;
 +                event->type = SMMU_EVT_F_TRANSLATION;
 +                event->u.f_translation.addr = addr;
 +                event->u.f_translation.addr2 = ptw_info.addr;
 +                event->u.f_translation.class = SMMU_CLASS_IN;
 +                event->u.f_translation.rnw = flag & 0x1;
 +            }
 +            break;
-+        case SMMU_PTW_ERR_ADDR_SIZE:
++        case float_infzeronan_dnan_if_qnan:
-+            if (PTW_RECORD_FAULT(cfg)) {
++            which = is_qnan(c->cls) ? 3 : 2;
 +                event->type = SMMU_EVT_F_ADDR_SIZE;
 +                event->u.f_addr_size.addr = addr;
 +                event->u.f_addr_size.addr2 = ptw_info.addr;
 +                event->u.f_addr_size.class = SMMU_CLASS_IN;
 +                event->u.f_addr_size.rnw = flag & 0x1;
 +            }
 +            break;
 +        case SMMU_PTW_ERR_ACCESS:
 +            if (PTW_RECORD_FAULT(cfg)) {
 +                event->type = SMMU_EVT_F_ACCESS;
 +                event->u.f_access.addr = addr;
 +                event->u.f_access.addr2 = ptw_info.addr;
 +                event->u.f_access.class = SMMU_CLASS_IN;
 +                event->u.f_access.rnw = flag & 0x1;
 +            }
 +            break;
 +        case SMMU_PTW_ERR_PERMISSION:
 +            if (PTW_RECORD_FAULT(cfg)) {
 +                event->type = SMMU_EVT_F_PERMISSION;
 +                event->u.f_permission.addr = addr;
 +                event->u.f_permission.addr2 = ptw_info.addr;
 +                event->u.f_permission.class = SMMU_CLASS_IN;
 +                event->u.f_permission.rnw = flag & 0x1;
 +            }
 +            break;
 +        default:
 +            g_assert_not_reached();
 +        }
-+        return SMMU_TRANS_ERROR;
+     } else {
-+    }
+-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
-+    *out_entry = cached_entry;
++        FloatClass cls[3] = { a->cls, b->cls, c->cls };
-+    return SMMU_TRANS_SUCCESS;
++        Float3NaNPropRule rule = s->float_3nan_prop_rule;
 +}
 +
-+/* Entry point to SMMU, does everything. */
++        assert(rule != float_3nan_prop_none);
- static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
++        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-                                       IOMMUAccessFlags flag, int iommu_idx)
++            /* We have at least one SNaN input and should prefer it */
- {
++            do {
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
++                which = rule & R_3NAN_1ST_MASK;
-     SMMUEventInfo event = {.type = SMMU_EVT_NONE,
++                rule >>= R_3NAN_1ST_LENGTH;
-                            .sid = sid,
++            } while (!is_snan(cls[which]));
-                            .inval_ste_allowed = false};
++        } else {
--    SMMUPTWEventInfo ptw_info = {};
++            do {
-     SMMUTranslationStatus status;
++                which = rule & R_3NAN_1ST_MASK;
--    SMMUState *bs = ARM_SMMU(s);
++                rule >>= R_3NAN_1ST_LENGTH;
--    uint64_t page_mask, aligned_addr;
++            } while (!is_nan(cls[which]));
--    SMMUTLBEntry *cached_entry = NULL;
++        }
--    SMMUTransTableInfo *tt;
+     }
-     SMMUTransCfg *cfg = NULL;
-     IOMMUTLBEntry entry = {
+     if (which == 3) {
-         .target_as = &address_space_memory,
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+index XXXXXXX..XXXXXXX 100644
-         .addr_mask = ~(hwaddr)0,
+--- a/fpu/softfloat-specialize.c.inc
-         .perm = IOMMU_NONE,
++++ b/fpu/softfloat-specialize.c.inc
-     };
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
      }
  }
 -/*----------------------------------------------------------------------------
 -| Select which NaN to propagate for a three-input operation.
 -| For the moment we assume that no CPU needs the 'larger significand'
 -| information.
 -| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
 -*----------------------------------------------------------------------------*/
 -static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
 -                         bool infzero, bool have_snan, float_status *status)
 -{
 -    FloatClass cls[3] = { a_cls, b_cls, c_cls };
 -    Float3NaNPropRule rule = status->float_3nan_prop_rule;
 -    int which;
 -
 -    /*
--     * Combined attributes used for TLB lookup, as only one stage is supported,
+-     * We guarantee not to require the target to tell us how to
--     * it will hold attributes based on the enabled stage.
+-     * pick a NaN if we're always returning the default NaN.
 -     * But if we're not in default-NaN mode then the target must
 -     * specify.
 -     */
--    SMMUTransTableInfo tt_combined;
+-    assert(!status->default_nan_mode);
 +    SMMUTLBEntry *cached_entry = NULL;
      qemu_mutex_lock(&s->mutex);
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
          goto epilogue;
      }
 -    if (cfg->stage == SMMU_STAGE_1) {
 -        /* Select stage1 translation table. */
 -        tt = select_tt(cfg, addr);
 -        if (!tt) {
 -            if (cfg->record_faults) {
 -                event.type = SMMU_EVT_F_TRANSLATION;
 -                event.u.f_translation.addr = addr;
 -                event.u.f_translation.rnw = flag & 0x1;
 -            }
 -            status = SMMU_TRANS_ERROR;
 -            goto epilogue;
 -        }
 -        tt_combined.granule_sz = tt->granule_sz;
 -        tt_combined.tsz = tt->tsz;
 -
--    } else {
+-    if (infzero) {
--        /* Stage2. */
+-        /*
--        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
+-         * Inf * 0 + NaN -- some implementations return the default NaN here,
--        tt_combined.tsz = cfg->s2cfg.tsz;
+-         * and some return the input NaN.
--    }
+-         */
--    /*
+-        switch (status->float_infzeronan_rule) {
--     * TLB lookup looks for granule and input size for a translation stage,
+-        case float_infzeronan_dnan_never:
--     * as only one stage is supported right now, choose the right values
+-            return 2;
--     * from the configuration.
+-        case float_infzeronan_dnan_always:
--     */
+-            return 3;
--    page_mask = (1ULL << tt_combined.granule_sz) - 1;
+-        case float_infzeronan_dnan_if_qnan:
--    aligned_addr = addr & ~page_mask;
+-            return is_qnan(c_cls) ? 3 : 2;
 -
 -    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
 -    if (cached_entry) {
 -        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
 -            status = SMMU_TRANS_ERROR;
 -            /*
 -             * We know that the TLB only contains either stage-1 or stage-2 as
 -             * nesting is not supported. So it is sufficient to check the
 -             * translation stage to know the TLB stage for now.
 -             */
 -            event.u.f_walk_eabt.s2 = (cfg->stage == SMMU_STAGE_2);
 -            if (PTW_RECORD_FAULT(cfg)) {
 -                event.type = SMMU_EVT_F_PERMISSION;
 -                event.u.f_permission.addr = addr;
 -                event.u.f_permission.rnw = flag & 0x1;
 -            }
 -        } else {
 -            status = SMMU_TRANS_SUCCESS;
 -        }
 -        goto epilogue;
 -    }
 -
 -    cached_entry = g_new0(SMMUTLBEntry, 1);
 -
 -    if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
 -        /* All faults from PTW has S2 field. */
 -        event.u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
 -        g_free(cached_entry);
 -        switch (ptw_info.type) {
 -        case SMMU_PTW_ERR_WALK_EABT:
 -            event.type = SMMU_EVT_F_WALK_EABT;
 -            event.u.f_walk_eabt.addr = addr;
 -            event.u.f_walk_eabt.rnw = flag & 0x1;
 -            /* Stage-2 (only) is class IN while stage-1 is class TT */
 -            event.u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
 -                                         SMMU_CLASS_IN : SMMU_CLASS_TT;
 -            event.u.f_walk_eabt.addr2 = ptw_info.addr;
 -            break;
 -        case SMMU_PTW_ERR_TRANSLATION:
 -            if (PTW_RECORD_FAULT(cfg)) {
 -                event.type = SMMU_EVT_F_TRANSLATION;
 -                event.u.f_translation.addr = addr;
 -                event.u.f_translation.addr2 = ptw_info.addr;
 -                event.u.f_translation.class = SMMU_CLASS_IN;
 -                event.u.f_translation.rnw = flag & 0x1;
 -            }
 -            break;
 -        case SMMU_PTW_ERR_ADDR_SIZE:
 -            if (PTW_RECORD_FAULT(cfg)) {
 -                event.type = SMMU_EVT_F_ADDR_SIZE;
 -                event.u.f_addr_size.addr = addr;
 -                event.u.f_addr_size.addr2 = ptw_info.addr;
 -                event.u.f_translation.class = SMMU_CLASS_IN;
 -                event.u.f_addr_size.rnw = flag & 0x1;
 -            }
 -            break;
 -        case SMMU_PTW_ERR_ACCESS:
 -            if (PTW_RECORD_FAULT(cfg)) {
 -                event.type = SMMU_EVT_F_ACCESS;
 -                event.u.f_access.addr = addr;
 -                event.u.f_access.addr2 = ptw_info.addr;
 -                event.u.f_translation.class = SMMU_CLASS_IN;
 -                event.u.f_access.rnw = flag & 0x1;
 -            }
 -            break;
 -        case SMMU_PTW_ERR_PERMISSION:
 -            if (PTW_RECORD_FAULT(cfg)) {
 -                event.type = SMMU_EVT_F_PERMISSION;
 -                event.u.f_permission.addr = addr;
 -                event.u.f_permission.addr2 = ptw_info.addr;
 -                event.u.f_translation.class = SMMU_CLASS_IN;
 -                event.u.f_permission.rnw = flag & 0x1;
 -            }
 -            break;
 -        default:
 -            g_assert_not_reached();
 -        }
--        status = SMMU_TRANS_ERROR;
+-    }
 -
 -    assert(rule != float_3nan_prop_none);
 -    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
 -        /* We have at least one SNaN input and should prefer it */
 -        do {
 -            which = rule & R_3NAN_1ST_MASK;
 -            rule >>= R_3NAN_1ST_LENGTH;
 -        } while (!is_snan(cls[which]));
 -    } else {
--        smmu_iotlb_insert(bs, cfg, cached_entry);
+-        do {
--        status = SMMU_TRANS_SUCCESS;
+-            which = rule & R_3NAN_1ST_MASK;
 -            rule >>= R_3NAN_1ST_LENGTH;
 -        } while (!is_nan(cls[which]));
 -    }
-+    status = smmuv3_do_translate(s, addr, cfg, &event, flag, &cached_entry);
+-    return which;
+-}
- epilogue:
+-
-     qemu_mutex_unlock(&s->mutex);
+ /*----------------------------------------------------------------------------
-@@ -XXX,XX +XXX,XX @@ epilogue:
+ | Returns 1 if the double-precision floating-point value `a' is a quiet
-                                     (addr & cached_entry->entry.addr_mask);
+ | NaN; otherwise returns 0.
          entry.addr_mask = cached_entry->entry.addr_mask;
          trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
 -                                       entry.translated_addr, entry.perm);
 +                                       entry.translated_addr, entry.perm,
 +                                       cfg->stage);
          break;
      case SMMU_TRANS_DISABLE:
          entry.perm = flag;
 diff --git a/hw/arm/trace-events b/hw/arm/trace-events
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/trace-events
 +++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_get_ste(uint64_t addr) "STE addr: 0x%"PRIx64
  smmuv3_translate_disable(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x bypass (smmu disabled) iova:0x%"PRIx64" is_write=%d"
  smmuv3_translate_bypass(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x STE bypass iova:0x%"PRIx64" is_write=%d"
  smmuv3_translate_abort(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x abort on iova:0x%"PRIx64" is_write=%d"
 -smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm) "%s sid=0x%x iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x"
 +smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm, int stage) "%s sid=0x%x iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x stage=%d"
  smmuv3_get_cd(uint64_t addr) "CD addr: 0x%"PRIx64
  smmuv3_decode_cd(uint32_t oas) "oas=%d"
  smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz, bool had) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d had:%d"
 --
 .34.1

-[PULL 21/26] hw/arm/smmu: Refactor SMMU OAS
+[PULL 62/72] softfloat: Use goto for default nan case in pick_nan_muladd
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-SMMUv3 OAS is currently hardcoded in the code to 44 bits, for nested
+Remove "3" as a special case for which and simply
-configurations that can be a problem, as stage-2 might be shared with
+branch to return the desired value.
 the CPU which might have different PARANGE, and according to SMMU manual
 ARM IHI 0070F.b:
 .3.6 SMMU_IDR5, OAS must match the system physical address size.
-This patch doesn't change the SMMU OAS, but refactors the code to
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-make it easier to do that:
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-- Rely everywhere on IDR5 for reading OAS instead of using the
+Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
   SMMU_IDR5_OAS macro, so, it is easier just to change IDR5 and
   it propagages correctly.
 - Add additional checks when OAS is greater than 48bits.
 - Remove unused functions/macros: pa_range/MAX_PA.
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-19-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h | 13 -------------
+ fpu/softfloat-parts.c.inc | 20 ++++++++++----------
- hw/arm/smmu-common.c     |  7 ++++---
+file changed, 10 insertions(+), 10 deletions(-)
  hw/arm/smmuv3.c          | 35 ++++++++++++++++++++++++++++-------
 files changed, 32 insertions(+), 23 deletions(-)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3-internal.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static inline int oas2bits(int oas_field)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-     return -1;
+          * But if we're not in default-NaN mode then the target must
- }
+          * specify.
+          */
--static inline int pa_range(STE *ste)
+-        which = 3;
--{
++        goto default_nan;
--    int oas_field = MIN(STE_S2PS(ste), SMMU_IDR5_OAS);
+     } else if (infzero) {
--
+         /*
--    if (!STE_S2AA64(ste)) {
+          * Inf * 0 + NaN -- some implementations return the
--        return 40;
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
           */
          switch (s->float_infzeronan_rule) {
          case float_infzeronan_dnan_never:
 -            which = 2;
              break;
          case float_infzeronan_dnan_always:
 -            which = 3;
 -            break;
 +            goto default_nan;
          case float_infzeronan_dnan_if_qnan:
 -            which = is_qnan(c->cls) ? 3 : 2;
 +            if (is_qnan(c->cls)) {
 +                goto default_nan;
 +            }
              break;
          default:
              g_assert_not_reached();
          }
 +        which = 2;
      } else {
          FloatClass cls[3] = { a->cls, b->cls, c->cls };
          Float3NaNPropRule rule = s->float_3nan_prop_rule;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          }
      }
 -    if (which == 3) {
 -        parts_default_nan(a, s);
 -        return a;
 -    }
 -
--    return oas2bits(oas_field);
+     switch (which) {
--}
+     case 0:
--
+         break;
--#define MAX_PA(ste) ((1 << pa_range(ste)) - 1)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
--
+         parts_silence_nan(a, s);
- /* CD fields */
+     }
+     return a;
  #define CD_VALID(x)   extract32((x)->word[0], 31, 1)
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUState *bs, SMMUTransCfg *cfg,
      inputsize = 64 - tt->tsz;
      level = 4 - (inputsize - 4) / stride;
      indexmask = VMSA_IDXMSK(inputsize, stride, level);
 -    baseaddr = extract64(tt->ttb, 0, 48);
 +
-+    baseaddr = extract64(tt->ttb, 0, cfg->oas);
++ default_nan:
-     baseaddr &= ~indexmask;
++    parts_default_nan(a, s);
++    return a;
      while (level < VMSA_LEVELS) {
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
       * Get the ttb from concatenated structure.
       * The offset is the idx * size of each ttb(number of ptes * (sizeof(pte))
       */
 -    uint64_t baseaddr = extract64(cfg->s2cfg.vttb, 0, 48) + (1 << stride) *
 -                                  idx * sizeof(uint64_t);
 +    uint64_t baseaddr = extract64(cfg->s2cfg.vttb, 0, cfg->s2cfg.eff_ps) +
 +                                  (1 << stride) * idx * sizeof(uint64_t);
      dma_addr_t indexmask = VMSA_IDXMSK(inputsize, stride, level);
      baseaddr &= ~indexmask;
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static bool s2t0sz_valid(SMMUTransCfg *cfg)
      }
      if (cfg->s2cfg.granule_sz == 16) {
 -        return (cfg->s2cfg.tsz >= 64 - oas2bits(SMMU_IDR5_OAS));
 +        return (cfg->s2cfg.tsz >= 64 - cfg->s2cfg.eff_ps);
      }
 -    return (cfg->s2cfg.tsz >= MAX(64 - oas2bits(SMMU_IDR5_OAS), 16));
 +    return (cfg->s2cfg.tsz >= MAX(64 - cfg->s2cfg.eff_ps, 16));
  }
  /*
-@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
-     return nr_concat <= VMSA_MAX_S2_CONCAT;
- }
--static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
-+static int decode_ste_s2_cfg(SMMUv3State *s, SMMUTransCfg *cfg,
-+                             STE *ste)
- {
-+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
-+
-     if (STE_S2AA64(ste) == 0x0) {
-         qemu_log_mask(LOG_UNIMP,
-                       "SMMUv3 AArch32 tables not supported\n");
-@@ -XXX,XX +XXX,XX @@ static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
-     }
-     /* For AA64, The effective S2PS size is capped to the OAS. */
--    cfg->s2cfg.eff_ps = oas2bits(MIN(STE_S2PS(ste), SMMU_IDR5_OAS));
-+    cfg->s2cfg.eff_ps = oas2bits(MIN(STE_S2PS(ste), oas));
-+    /*
-+     * For SMMUv3.1 and later, when OAS == IAS == 52, the stage 2 input
-+     * range is further limited to 48 bits unless STE.S2TG indicates a
-+     * 64KB granule.
-+     */
-+    if (cfg->s2cfg.granule_sz != 16) {
-+        cfg->s2cfg.eff_ps = MIN(cfg->s2cfg.eff_ps, 48);
-+    }
-     /*
-      * It is ILLEGAL for the address in S2TTB to be outside the range
-      * described by the effective S2PS value.
-@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
-                       STE *ste, SMMUEventInfo *event)
- {
-     uint32_t config;
-+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
-     int ret;
-     if (!STE_VALID(ste)) {
-@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
-          * Stage-1 OAS defaults to OAS even if not enabled as it would be used
-          * in input address check for stage-2.
-          */
--        cfg->oas = oas2bits(SMMU_IDR5_OAS);
--        ret = decode_ste_s2_cfg(cfg, ste);
-+        cfg->oas = oas2bits(oas);
-+        ret = decode_ste_s2_cfg(s, cfg, ste);
-         if (ret) {
-             goto bad_ste;
-         }
-@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
-     int i;
-     SMMUTranslationStatus status;
-     SMMUTLBEntry *entry;
-+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
-     if (!CD_VALID(cd) || !CD_AARCH64(cd)) {
-         goto bad_cd;
-@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
-     cfg->aa64 = true;
-     cfg->oas = oas2bits(CD_IPS(cd));
--    cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
-+    cfg->oas = MIN(oas2bits(oas), cfg->oas);
-     cfg->tbi = CD_TBI(cd);
-     cfg->asid = CD_ASID(cd);
-     cfg->affd = CD_AFFD(cd);
-@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
-             goto bad_cd;
-         }
-+        /*
-+         * An address greater than 48 bits in size can only be output from a
-+         * TTD when, in SMMUv3.1 and later, the effective IPS is 52 and a 64KB
-+         * granule is in use for that translation table
-+         */
-+        if (tt->granule_sz != 16) {
-+            cfg->oas = MIN(cfg->oas, 48);
-+        }
-         tt->tsz = tsz;
-         tt->ttb = CD_TTB(cd, i);
 --
 .34.1

-[PULL 07/26] hw/arm/smmu: Use enum for SMMU stage
+[PULL 63/72] softfloat: Remove which from parts_pick_nan_muladd
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Currently, translation stage is represented as an int, where 1 is stage-1 and
+Assign the pointer return value to 'a' directly,
-is stage-2, when nested is added, 3 would be confusing to represent nesting,
+rather than going through an intermediary index.
 so we use an enum instead.
-While keeping the same values, this is useful for:
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
- - Doing tricks with bit masks, where BIT(0) is stage-1 and BIT(1) is
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-   stage-2 and both is nested.
+Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
  - Tracing, as stage is printed as int.
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Message-id: 20240715084519.1189624-5-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/hw/arm/smmu-common.h | 11 +++++++++--
+ fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
- hw/arm/smmu-common.c         | 14 +++++++-------
+file changed, 10 insertions(+), 22 deletions(-)
  hw/arm/smmuv3.c              | 17 +++++++++--------
 files changed, 25 insertions(+), 17 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ typedef enum {
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-     SMMU_PTW_ERR_PERMISSION,  /* Permission fault */
+                                             FloatPartsN *c, float_status *s,
- } SMMUPTWEventType;
+                                             int ab_mask, int abc_mask)
 +/* SMMU Stage */
 +typedef enum {
 +    SMMU_STAGE_1 = 1,
 +    SMMU_STAGE_2,
 +    SMMU_NESTED,
 +} SMMUStage;
 +
  typedef struct SMMUPTWEventInfo {
 -    int stage;
 +    SMMUStage stage;
      SMMUPTWEventType type;
      dma_addr_t addr; /* fetched address that induced an abort, if any */
  } SMMUPTWEventInfo;
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUS2Cfg {
   */
  typedef struct SMMUTransCfg {
      /* Shared fields between stage-1 and stage-2. */
 -    int stage;                 /* translation stage */
 +    SMMUStage stage;           /* translation stage */
      bool disabled;             /* smmu is disabled */
      bool bypassed;             /* translation is bypassed */
      bool aborted;              /* translation is aborted */
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
                            SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
  {
-     dma_addr_t baseaddr, indexmask;
+-    int which;
--    int stage = cfg->stage;
+     bool infzero = (ab_mask == float_cmask_infzero);
-+    SMMUStage stage = cfg->stage;
+     bool have_snan = (abc_mask & float_cmask_snan);
-     SMMUTransTableInfo *tt = select_tt(cfg, iova);
++    FloatPartsN *ret;
-     uint8_t level, granule_sz, inputsize, stride;
+     if (unlikely(have_snan)) {
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-     info->type = SMMU_PTW_ERR_TRANSLATION;
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         default:
- error:
+             g_assert_not_reached();
 -    info->stage = 1;
 +    info->stage = SMMU_STAGE_1;
      tlbe->entry.perm = IOMMU_NONE;
      return -EINVAL;
  }
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
                            dma_addr_t ipa, IOMMUAccessFlags perm,
                            SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
  {
 -    const int stage = 2;
 +    const SMMUStage stage = SMMU_STAGE_2;
      int granule_sz = cfg->s2cfg.granule_sz;
      /* ARM DDI0487I.a: Table D8-7. */
      int inputsize = 64 - cfg->s2cfg.tsz;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
  error_ipa:
      info->addr = ipa;
  error:
 -    info->stage = 2;
 +    info->stage = SMMU_STAGE_2;
      tlbe->entry.perm = IOMMU_NONE;
      return -EINVAL;
  }
@@ -XXX,XX +XXX,XX @@ error:
  int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
               SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
  {
 -    if (cfg->stage == 1) {
 +    if (cfg->stage == SMMU_STAGE_1) {
          return smmu_ptw_64_s1(cfg, iova, perm, tlbe, info);
 -    } else if (cfg->stage == 2) {
 +    } else if (cfg->stage == SMMU_STAGE_2) {
          /*
           * If bypassing stage 1(or unimplemented), the input address is passed
           * directly to stage 2 as IPA. If the input address of a transaction
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
           */
          if (iova >= (1ULL << cfg->oas)) {
              info->type = SMMU_PTW_ERR_ADDR_SIZE;
 -            info->stage = 1;
 +            info->stage = SMMU_STAGE_1;
              tlbe->entry.perm = IOMMU_NONE;
              return -EINVAL;
          }
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+-        which = 2;
-index XXXXXXX..XXXXXXX 100644
++        ret = c;
---- a/hw/arm/smmuv3.c
+     } else {
-+++ b/hw/arm/smmuv3.c
+-        FloatClass cls[3] = { a->cls, b->cls, c->cls };
-@@ -XXX,XX +XXX,XX @@
++        FloatPartsN *val[3] = { a, b, c };
- #include "smmuv3-internal.h"
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
- #include "smmu-internal.h"
+         assert(rule != float_3nan_prop_none);
--#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == 1) ? (cfg)->record_faults : \
+         if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-+#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == SMMU_STAGE_1) ? \
+             /* We have at least one SNaN input and should prefer it */
-+                                 (cfg)->record_faults : \
+             do {
-                                  (cfg)->s2cfg.record_faults)
+-                which = rule & R_3NAN_1ST_MASK;
++                ret = val[rule & R_3NAN_1ST_MASK];
- /**
+                 rule >>= R_3NAN_1ST_LENGTH;
-@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
+-            } while (!is_snan(cls[which]));
++            } while (!is_snan(ret->cls));
- static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
+         } else {
- {
+             do {
--    cfg->stage = 2;
+-                which = rule & R_3NAN_1ST_MASK;
-+    cfg->stage = SMMU_STAGE_2;
++                ret = val[rule & R_3NAN_1ST_MASK];
+                 rule >>= R_3NAN_1ST_LENGTH;
-     if (STE_S2AA64(ste) == 0x0) {
+-            } while (!is_nan(cls[which]));
-         qemu_log_mask(LOG_UNIMP,
++            } while (!is_nan(ret->cls));
-@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
+         }
      /* we support only those at the moment */
      cfg->aa64 = true;
 -    cfg->stage = 1;
 +    cfg->stage = SMMU_STAGE_1;
      cfg->oas = oas2bits(CD_IPS(cd));
      cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
@@ -XXX,XX +XXX,XX @@ static int smmuv3_decode_config(IOMMUMemoryRegion *mr, SMMUTransCfg *cfg,
          return ret;
      }
--    if (cfg->aborted || cfg->bypassed || (cfg->stage == 2)) {
+-    switch (which) {
-+    if (cfg->aborted || cfg->bypassed || (cfg->stage == SMMU_STAGE_2)) {
+-    case 0:
-         return 0;
+-        break;
 -    case 1:
 -        a = b;
 -        break;
 -    case 2:
 -        a = c;
 -        break;
 -    default:
 -        g_assert_not_reached();
 +    if (is_snan(ret->cls)) {
 +        parts_silence_nan(ret, s);
      }
+-    if (is_snan(a->cls)) {
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+-        parts_silence_nan(a, s);
-         goto epilogue;
+-    }
-     }
+-    return a;
++    return ret;
--    if (cfg->stage == 1) {
-+    if (cfg->stage == SMMU_STAGE_1) {
+  default_nan:
-         /* Select stage1 translation table. */
+     parts_default_nan(a, s);
          tt = select_tt(cfg, addr);
          if (!tt) {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
               * nesting is not supported. So it is sufficient to check the
               * translation stage to know the TLB stage for now.
               */
 -            event.u.f_walk_eabt.s2 = (cfg->stage == 2);
 +            event.u.f_walk_eabt.s2 = (cfg->stage == SMMU_STAGE_2);
              if (PTW_RECORD_FAULT(cfg)) {
                  event.type = SMMU_EVT_F_PERMISSION;
                  event.u.f_permission.addr = addr;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
      if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
          /* All faults from PTW has S2 field. */
 -        event.u.f_walk_eabt.s2 = (ptw_info.stage == 2);
 +        event.u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
          g_free(cached_entry);
          switch (ptw_info.type) {
          case SMMU_PTW_ERR_WALK_EABT:
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              event.u.f_walk_eabt.addr = addr;
              event.u.f_walk_eabt.rnw = flag & 0x1;
              /* Stage-2 (only) is class IN while stage-1 is class TT */
 -            event.u.f_walk_eabt.class = (ptw_info.stage == 2) ?
 +            event.u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
                                           SMMU_CLASS_IN : SMMU_CLASS_TT;
              event.u.f_walk_eabt.addr2 = ptw_info.addr;
              break;
 --
 .34.1

-[PULL 22/26] target/arm: Use float_status copy in sme_fmopa_s
+[PULL 64/72] softfloat: Pad array size in pick_nan_muladd
-From: Daniyal Khan <danikhan632@gmail.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-We made a copy above because the fp exception flags
+While all indices into val[] should be in [0-2], the mask
-are not propagated back to the FPST register, but
+applied is two bits.  To help static analysis see there is
-then failed to use the copy.
+no possibility of read beyond the end of the array, pad the
 array to 4 entries, with the final being (implicitly) NULL.
-Cc: qemu-stable@nongnu.org
-Fixes: 558e956c719 ("target/arm: Implement FMOPA, FMOPS (non-widening)")
-Signed-off-by: Daniyal Khan <danikhan632@gmail.com>
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
 Message-id: 20240717060149.204788-2-richard.henderson@linaro.org
 [rth: Split from a larger patch]
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/sme_helper.c | 2 +-
+ fpu/softfloat-parts.c.inc | 2 +-
 file changed, 1 insertion(+), 1 deletion(-)
-diff --git a/target/arm/tcg/sme_helper.c b/target/arm/tcg/sme_helper.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/sme_helper.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/tcg/sme_helper.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ void HELPER(sme_fmopa_s)(void *vza, void *vzn, void *vzm, void *vpn,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-                         if (pb & 1) {
+         }
-                             uint32_t *a = vza_row + H1_4(col);
+         ret = c;
-                             uint32_t *m = vzm + H1_4(col);
+     } else {
--                            *a = float32_muladd(n, *m, *a, 0, vst);
+-        FloatPartsN *val[3] = { a, b, c };
-+                            *a = float32_muladd(n, *m, *a, 0, &fpst);
++        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
-                         }
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
-                         col += 4;
-                         pb >>= 4;
+         assert(rule != float_3nan_prop_none);
 --
 .34.1

-[PULL 20/26] hw/arm/smmuv3: Support and advertise nesting
+[PULL 65/72] softfloat: Move propagateFloatx80NaN to softfloat.c
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Everything is in place, consolidate parsing of STE cfg and setting
+This function is part of the public interface and
-translation stage.
+is not "specialized" to any target in any way.
-Advertise nesting if stage requested is "nested".
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
+Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-18-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3.c | 35 ++++++++++++++++++++++++++---------
+ fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
-file changed, 26 insertions(+), 9 deletions(-)
+ fpu/softfloat-specialize.c.inc | 52 ----------------------------------
 files changed, 52 insertions(+), 52 deletions(-)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/fpu/softfloat.c
-+++ b/hw/arm/smmuv3.c
++++ b/fpu/softfloat.c
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
-     /* Based on sys property, the stages supported in smmu will be advertised.*/
+     *zExpPtr = 1 - shiftCount;
      if (s->stage && !strcmp("2", s->stage)) {
          s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S2P, 1);
 +    } else if (s->stage && !strcmp("nested", s->stage)) {
 +        s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S1P, 1);
 +        s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S2P, 1);
      } else {
          s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S1P, 1);
      }
@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
  static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
  {
 -    cfg->stage = SMMU_STAGE_2;
 -
      if (STE_S2AA64(ste) == 0x0) {
          qemu_log_mask(LOG_UNIMP,
                        "SMMUv3 AArch32 tables not supported\n");
@@ -XXX,XX +XXX,XX @@ bad_ste:
      return -EINVAL;
  }
-+static void decode_ste_config(SMMUTransCfg *cfg, uint32_t config)
++/*----------------------------------------------------------------------------
 +| Takes two extended double-precision floating-point values `a' and `b', one
 +| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
 +| `b' is a signaling NaN, the invalid exception is raised.
 +*----------------------------------------------------------------------------*/
 +
 +floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 +{
++    bool aIsLargerSignificand;
++    FloatClass a_cls, b_cls;
 +
-+    if (STE_CFG_ABORT(config)) {
++    /* This is not complete, but is good enough for pickNaN.  */
-+        cfg->aborted = true;
++    a_cls = (!floatx80_is_any_nan(a)
-+        return;
++             ? float_class_normal
-+    }
++             : floatx80_is_signaling_nan(a, status)
-+    if (STE_CFG_BYPASS(config)) {
++             ? float_class_snan
-+        cfg->bypassed = true;
++             : float_class_qnan);
-+        return;
++    b_cls = (!floatx80_is_any_nan(b)
 +             ? float_class_normal
 +             : floatx80_is_signaling_nan(b, status)
 +             ? float_class_snan
 +             : float_class_qnan);
 +
 +    if (is_snan(a_cls) || is_snan(b_cls)) {
 +        float_raise(float_flag_invalid, status);
 +    }
 +
-+    if (STE_CFG_S1_ENABLED(config)) {
++    if (status->default_nan_mode) {
-+        cfg->stage = SMMU_STAGE_1;
++        return floatx80_default_nan(status);
 +    }
 +
-+    if (STE_CFG_S2_ENABLED(config)) {
++    if (a.low < b.low) {
-+        cfg->stage |= SMMU_STAGE_2;
++        aIsLargerSignificand = 0;
 +    } else if (b.low < a.low) {
 +        aIsLargerSignificand = 1;
 +    } else {
 +        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
 +    }
 +
 +    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
 +        if (is_snan(b_cls)) {
 +            return floatx80_silence_nan(b, status);
 +        }
 +        return b;
 +    } else {
 +        if (is_snan(a_cls)) {
 +            return floatx80_silence_nan(a, status);
 +        }
 +        return a;
 +    }
 +}
 +
- /* Returns < 0 in case of invalid STE, 0 otherwise */
+ /*----------------------------------------------------------------------------
- static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
+ | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
-                       STE *ste, SMMUEventInfo *event)
+ | and extended significand formed by the concatenation of `zSig0' and `zSig1',
-@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
-     config = STE_CONFIG(ste);
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
--    if (STE_CFG_ABORT(config)) {
+@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
--        cfg->aborted = true;
+     return a;
--        return 0;
+ }
 -/*----------------------------------------------------------------------------
 -| Takes two extended double-precision floating-point values `a' and `b', one
 -| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
 -| `b' is a signaling NaN, the invalid exception is raised.
 -*----------------------------------------------------------------------------*/
 -
 -floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 -{
 -    bool aIsLargerSignificand;
 -    FloatClass a_cls, b_cls;
 -
 -    /* This is not complete, but is good enough for pickNaN.  */
 -    a_cls = (!floatx80_is_any_nan(a)
 -             ? float_class_normal
 -             : floatx80_is_signaling_nan(a, status)
 -             ? float_class_snan
 -             : float_class_qnan);
 -    b_cls = (!floatx80_is_any_nan(b)
 -             ? float_class_normal
 -             : floatx80_is_signaling_nan(b, status)
 -             ? float_class_snan
 -             : float_class_qnan);
 -
 -    if (is_snan(a_cls) || is_snan(b_cls)) {
 -        float_raise(float_flag_invalid, status);
 -    }
-+    decode_ste_config(cfg, config);
+-
+-    if (status->default_nan_mode) {
--    if (STE_CFG_BYPASS(config)) {
+-        return floatx80_default_nan(status);
--        cfg->bypassed = true;
+-    }
-+    if (cfg->aborted || cfg->bypassed) {
+-
-         return 0;
+-    if (a.low < b.low) {
-     }
+-        aIsLargerSignificand = 0;
+-    } else if (b.low < a.low) {
-@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
+-        aIsLargerSignificand = 1;
+-    } else {
-     /* we support only those at the moment */
+-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-     cfg->aa64 = true;
+-    }
--    cfg->stage = SMMU_STAGE_1;
+-
+-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-     cfg->oas = oas2bits(CD_IPS(cd));
+-        if (is_snan(b_cls)) {
-     cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
+-            return floatx80_silence_nan(b, status);
 -        }
 -        return b;
 -    } else {
 -        if (is_snan(a_cls)) {
 -            return floatx80_silence_nan(a, status);
 -        }
 -        return a;
 -    }
 -}
 -
  /*----------------------------------------------------------------------------
  | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
  | NaN; otherwise returns 0.
 --
 .34.1

-[PULL 09/26] hw/arm/smmu: Consolidate ASID and VMID types
+[PULL 66/72] softfloat: Use parts_pick_nan in propagateFloatx80NaN
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-ASID and VMID used to be uint16_t in the translation config, however,
+Unpacking and repacking the parts may be slightly more work
-in other contexts they can be int as -1 in case of TLB invalidation,
+than we did before, but we get to reuse more code.  For a
-to represent all (don’t care).
+code path handling exceptional values, this is an improvement.
 When stage-2 was added asid was set to -1 in stage-2 and vmid to -1
 in stage-1 configs. However, that meant they were set as (65536),
 this was not an issue as nesting was not supported and no
 commands/lookup uses both.
-With nesting, it’s critical to get this right as translation must be
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-tagged correctly with ASID/VMID, and with ASID=-1 meaning stage-2.
+Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
-Represent ASID/VMID everywhere as int.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  fpu/softfloat.c | 43 +++++--------------------------------------
 file changed, 5 insertions(+), 38 deletions(-)
-Reviewed-by: Eric Auger <eric.auger@redhat.com>
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-7-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  include/hw/arm/smmu-common.h | 14 +++++++-------
  hw/arm/smmu-common.c         | 10 +++++-----
  hw/arm/smmuv3.c              |  4 ++--
  hw/arm/trace-events          | 18 +++++++++---------
 files changed, 23 insertions(+), 23 deletions(-)
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat.c
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat.c
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUS2Cfg {
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
-     bool record_faults;     /* Record fault events (S2R) */
-     uint8_t granule_sz;     /* Granule page shift (based on S2TG) */
+ floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-     uint8_t eff_ps;         /* Effective PA output range (based on S2PS) */
+ {
--    uint16_t vmid;          /* Virtual Machine ID (S2VMID) */
+-    bool aIsLargerSignificand;
-+    int vmid;               /* Virtual Machine ID (S2VMID) */
+-    FloatClass a_cls, b_cls;
-     uint64_t vttb;          /* Address of translation table base (S2TTB) */
++    FloatParts128 pa, pb, *pr;
- } SMMUS2Cfg;
+-    /* This is not complete, but is good enough for pickNaN.  */
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransCfg {
+-    a_cls = (!floatx80_is_any_nan(a)
-     uint64_t ttb;              /* TT base address */
+-             ? float_class_normal
-     uint8_t oas;               /* output address width */
+-             : floatx80_is_signaling_nan(a, status)
-     uint8_t tbi;               /* Top Byte Ignore */
+-             ? float_class_snan
--    uint16_t asid;
+-             : float_class_qnan);
-+    int asid;
+-    b_cls = (!floatx80_is_any_nan(b)
-     SMMUTransTableInfo tt[2];
+-             ? float_class_normal
-     /* Used by stage-2 only. */
+-             : floatx80_is_signaling_nan(b, status)
-     struct SMMUS2Cfg s2cfg;
+-             ? float_class_snan
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPciBus {
+-             : float_class_qnan);
+-
- typedef struct SMMUIOTLBKey {
+-    if (is_snan(a_cls) || is_snan(b_cls)) {
-     uint64_t iova;
+-        float_raise(float_flag_invalid, status);
--    uint16_t asid;
+-    }
--    uint16_t vmid;
+-
-+    int asid;
+-    if (status->default_nan_mode) {
-+    int vmid;
++    if (!floatx80_unpack_canonical(&pa, a, status) ||
-     uint8_t tg;
++        !floatx80_unpack_canonical(&pb, b, status)) {
-     uint8_t level;
+         return floatx80_default_nan(status);
- } SMMUIOTLBKey;
+     }
-@@ -XXX,XX +XXX,XX @@ SMMUDevice *smmu_find_sdev(SMMUState *s, uint32_t sid);
- SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+-    if (a.low < b.low) {
-                                 SMMUTransTableInfo *tt, hwaddr iova);
+-        aIsLargerSignificand = 0;
- void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
+-    } else if (b.low < a.low) {
--SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint16_t vmid, uint64_t iova,
+-        aIsLargerSignificand = 1;
-+SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
+-    } else {
-                                 uint8_t tg, uint8_t level);
+-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
- void smmu_iotlb_inv_all(SMMUState *s);
+-    }
--void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
+-
--void smmu_iotlb_inv_vmid(SMMUState *s, uint16_t vmid);
+-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-+void smmu_iotlb_inv_asid(SMMUState *s, int asid);
+-        if (is_snan(b_cls)) {
-+void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
+-            return floatx80_silence_nan(b, status);
- void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
+-        }
-                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
+-        return b;
+-    } else {
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
+-        if (is_snan(a_cls)) {
-index XXXXXXX..XXXXXXX 100644
+-            return floatx80_silence_nan(a, status);
---- a/hw/arm/smmu-common.c
+-        }
-+++ b/hw/arm/smmu-common.c
+-        return a;
-@@ -XXX,XX +XXX,XX @@ static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
+-    }
-            (k1->vmid == k2->vmid);
++    pr = parts_pick_nan(&pa, &pb, status);
 +    return floatx80_round_pack_canonical(pr, status);
  }
--SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint16_t vmid, uint64_t iova,
+ /*----------------------------------------------------------------------------
 +SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
                                  uint8_t tg, uint8_t level)
  {
      SMMUIOTLBKey key = {.asid = asid, .vmid = vmid, .iova = iova,
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_all(SMMUState *s)
  static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
                                           gpointer user_data)
  {
 -    uint16_t asid = *(uint16_t *)user_data;
 +    int asid = *(int *)user_data;
      SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
      return SMMU_IOTLB_ASID(*iotlb_key) == asid;
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
  static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
                                           gpointer user_data)
  {
 -    uint16_t vmid = *(uint16_t *)user_data;
 +    int vmid = *(int *)user_data;
      SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
      return SMMU_IOTLB_VMID(*iotlb_key) == vmid;
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                                  &info);
  }
 -void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
 +void smmu_iotlb_inv_asid(SMMUState *s, int asid)
  {
      trace_smmu_iotlb_inv_asid(asid);
      g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid, &asid);
  }
 -void smmu_iotlb_inv_vmid(SMMUState *s, uint16_t vmid)
 +void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
  {
      trace_smmu_iotlb_inv_vmid(vmid);
      g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid, &vmid);
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
          }
          case SMMU_CMD_TLBI_NH_ASID:
          {
 -            uint16_t asid = CMD_ASID(&cmd);
 +            int asid = CMD_ASID(&cmd);
              if (!STAGE1_SUPPORTED(s)) {
                  cmd_error = SMMU_CERROR_ILL;
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
              break;
          case SMMU_CMD_TLBI_S12_VMALL:
          {
 -            uint16_t vmid = CMD_VMID(&cmd);
 +            int vmid = CMD_VMID(&cmd);
              if (!STAGE2_SUPPORTED(s)) {
                  cmd_error = SMMU_CERROR_ILL;
 diff --git a/hw/arm/trace-events b/hw/arm/trace-events
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/trace-events
 +++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_ptw_page_pte(int stage, int level,  uint64_t iova, uint64_t baseaddr, uint6
  smmu_ptw_block_pte(int stage, int level, uint64_t baseaddr, uint64_t pteaddr, uint64_t pte, uint64_t iova, uint64_t gpa, int bsize_mb) "stage=%d level=%d base@=0x%"PRIx64" pte@=0x%"PRIx64" pte=0x%"PRIx64" iova=0x%"PRIx64" block address = 0x%"PRIx64" block size = %d MiB"
  smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "baseaddr=0x%"PRIx64" index=0x%x, pteaddr=0x%"PRIx64", pte=0x%"PRIx64
  smmu_iotlb_inv_all(void) "IOTLB invalidate all"
 -smmu_iotlb_inv_asid(uint16_t asid) "IOTLB invalidate asid=%d"
 -smmu_iotlb_inv_vmid(uint16_t vmid) "IOTLB invalidate vmid=%d"
 -smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
 +smmu_iotlb_inv_asid(int asid) "IOTLB invalidate asid=%d"
 +smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
 +smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
  smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
 -smmu_iotlb_lookup_hit(uint16_t asid, uint16_t vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 -smmu_iotlb_lookup_miss(uint16_t asid, uint16_t vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 -smmu_iotlb_insert(uint16_t asid, uint16_t vmid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d vmid=%d addr=0x%"PRIx64" tg=%d level=%d"
 +smmu_iotlb_lookup_hit(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 +smmu_iotlb_lookup_miss(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 +smmu_iotlb_insert(int asid, int vmid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d vmid=%d addr=0x%"PRIx64" tg=%d level=%d"
  # smmuv3.c
  smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
@@ -XXX,XX +XXX,XX @@ smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t p
  smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
  smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
  smmuv3_cmdq_tlbi_nh(void) ""
 -smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
 -smmuv3_cmdq_tlbi_s12_vmid(uint16_t vmid) "vmid=%d"
 +smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
 +smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
  smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
  smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
  smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
 -smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint16_t vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
 +smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
  # strongarm.c
  strongarm_uart_update_parameters(const char *label, int speed, char parity, int data_bits, int stop_bits) "%s speed=%d parity=%c data=%d stop=%d"
 --
 .34.1

-[PULL 18/26] hw/arm/smmuv3: Support nested SMMUs in smmuv3_notify_iova()
+[PULL 67/72] softfloat: Inline pickNaN
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-IOMMUTLBEvent only understands IOVA, for stage-1 or stage-2
+Inline pickNaN into its only caller.  This makes one assert
-SMMU instances we consider the input address as the IOVA, but when
+redundant with the immediately preceding IF.
-nesting is used, we can't mix stage-1 and stage-2 addresses, so for
-nesting only stage-1 is considered the IOVA and would be notified.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Signed-off-by: Mostafa Saleh <smostafa@google.com>
+Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-16-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3.c     | 39 +++++++++++++++++++++++++--------------
+ fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
- hw/arm/trace-events |  2 +-
+ fpu/softfloat-specialize.c.inc | 96 ----------------------------------
-files changed, 26 insertions(+), 15 deletions(-)
+files changed, 73 insertions(+), 105 deletions(-)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ epilogue:
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
-  * @iova: iova
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-  * @tg: translation granule (if communicated through range invalidation)
+                                      float_status *s)
   * @num_pages: number of @granule sized pages (if tg != 0), otherwise 1
 + * @stage: Which stage(1 or 2) is used
   */
  static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
                                 IOMMUNotifier *n,
                                 int asid, int vmid,
                                 dma_addr_t iova, uint8_t tg,
 -                               uint64_t num_pages)
 +                               uint64_t num_pages, int stage)
  {
-     SMMUDevice *sdev = container_of(mr, SMMUDevice, iommu);
++    int cmp, which;
 +    SMMUEventInfo eventinfo = {.inval_ste_allowed = true};
 +    SMMUTransCfg *cfg = smmuv3_get_config(sdev, &eventinfo);
      IOMMUTLBEvent event;
      uint8_t granule;
 -    SMMUv3State *s = sdev->smmu;
 +
-+    if (!cfg) {
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+        return;
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
      }
      if (s->default_nan_mode) {
          parts_default_nan(a, s);
 -    } else {
 -        int cmp = frac_cmp(a, b);
 -        if (cmp == 0) {
 -            cmp = a->sign < b->sign;
 -        }
 +        return a;
 +    }
 -        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
 -            a = b;
 -        }
 +    cmp = frac_cmp(a, b);
 +    if (cmp == 0) {
 +        cmp = a->sign < b->sign;
 +    }
 +
-+    /*
++    switch (s->float_2nan_prop_rule) {
-+     * stage is passed from TLB invalidation commands which can be either
++    case float_2nan_prop_s_ab:
-+     * stage-1 or stage-2.
+         if (is_snan(a->cls)) {
-+     * However, IOMMUTLBEvent only understands IOVA, for stage-1 or stage-2
+-            parts_silence_nan(a, s);
-+     * SMMU instances we consider the input address as the IOVA, but when
++            which = 0;
-+     * nesting is used, we can't mix stage-1 and stage-2 addresses, so for
++        } else if (is_snan(b->cls)) {
-+     * nesting only stage-1 is considered the IOVA and would be notified.
++            which = 1;
-+     */
++        } else if (is_qnan(a->cls)) {
-+    if ((stage == SMMU_STAGE_2) && (cfg->stage == SMMU_NESTED))
++            which = 0;
-+        return;
++        } else {
++            which = 1;
      if (!tg) {
 -        SMMUEventInfo eventinfo = {.inval_ste_allowed = true};
 -        SMMUTransCfg *cfg = smmuv3_get_config(sdev, &eventinfo);
          SMMUTransTableInfo *tt;
 -        if (!cfg) {
 -            return;
 -        }
 -
          if (asid >= 0 && cfg->asid != asid) {
              return;
          }
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
++        break;
-             return;
++    case float_2nan_prop_s_ba:
-         }
++        if (is_snan(b->cls)) {
++            which = 1;
--        if (STAGE1_SUPPORTED(s)) {
++        } else if (is_snan(a->cls)) {
-+        if (stage == SMMU_STAGE_1) {
++            which = 0;
-             tt = select_tt(cfg, iova);
++        } else if (is_qnan(b->cls)) {
-             if (!tt) {
++            which = 1;
-                 return;
++        } else {
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
++            which = 0;
- /* invalidate an asid/vmid/iova range tuple in all mr's */
++        }
- static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
++        break;
-                                       dma_addr_t iova, uint8_t tg,
++    case float_2nan_prop_ab:
--                                      uint64_t num_pages)
++        which = is_nan(a->cls) ? 0 : 1;
-+                                      uint64_t num_pages, int stage)
++        break;
- {
++    case float_2nan_prop_ba:
-     SMMUDevice *sdev;
++        which = is_nan(b->cls) ? 1 : 0;
++        break;
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
++    case float_2nan_prop_x87:
-         IOMMUNotifier *n;
++        /*
++         * This implements x87 NaN propagation rules:
-         trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, vmid,
++         * SNaN + QNaN => return the QNaN
--                                        iova, tg, num_pages);
++         * two SNaNs => return the one with the larger significand, silenced
-+                                        iova, tg, num_pages, stage);
++         * two QNaNs => return the one with the larger significand
++         * SNaN and a non-NaN => return the SNaN, silenced
-         IOMMU_NOTIFIER_FOREACH(n, mr) {
++         * QNaN and a non-NaN => return the QNaN
--            smmuv3_notify_iova(mr, n, asid, vmid, iova, tg, num_pages);
++         *
-+            smmuv3_notify_iova(mr, n, asid, vmid, iova, tg, num_pages, stage);
++         * If we get down to comparing significands and they are the same,
-         }
++         * return the NaN with the positive sign bit (if any).
 +         */
 +        if (is_snan(a->cls)) {
 +            if (is_snan(b->cls)) {
 +                which = cmp > 0 ? 0 : 1;
 +            } else {
 +                which = is_qnan(b->cls) ? 1 : 0;
 +            }
 +        } else if (is_qnan(a->cls)) {
 +            if (is_snan(b->cls) || !is_qnan(b->cls)) {
 +                which = 0;
 +            } else {
 +                which = cmp > 0 ? 0 : 1;
 +            }
 +        } else {
 +            which = 1;
 +        }
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +
 +    if (which) {
 +        a = b;
 +    }
 +    if (is_snan(a->cls)) {
 +        parts_silence_nan(a, s);
      }
      return a;
  }
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
      }
  }
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
+-/*----------------------------------------------------------------------------
-     if (!tg) {
+-| Select which NaN to propagate for a two-input operation.
-         trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf, stage);
+-| IEEE754 doesn't specify all the details of this, so the
--        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1);
+-| algorithm is target-specific.
-+        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1, stage);
+-| The routine is passed various bits of information about the
-         if (stage == SMMU_STAGE_1) {
+-| two NaNs and should return 0 to select NaN a and 1 for NaN b.
-             smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
+-| Note that signalling NaNs are always squashed to quiet NaNs
-         } else {
+-| by the caller, by calling floatXX_silence_nan() before
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
+-| returning them.
-         num_pages = (mask + 1) >> granule;
+-|
-         trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages,
+-| aIsLargerSignificand is only valid if both a and b are NaNs
-                                  ttl, leaf, stage);
+-| of some kind, and is true if a has the larger significand,
--        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages);
+-| or if both a and b have the same significand but a is
-+        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages, stage);
+-| positive but b is negative. It is only needed for the x87
-         if (stage == SMMU_STAGE_1) {
+-| tie-break rule.
-             smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
+-*----------------------------------------------------------------------------*/
-         } else {
+-
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
+-static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-index XXXXXXX..XXXXXXX 100644
+-                   bool aIsLargerSignificand, float_status *status)
---- a/hw/arm/trace-events
+-{
-+++ b/hw/arm/trace-events
+-    /*
-@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
+-     * We guarantee not to require the target to tell us how to
- smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
+-     * pick a NaN if we're always returning the default NaN.
- smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
+-     * But if we're not in default-NaN mode then the target must
- smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
+-     * specify via set_float_2nan_prop_rule().
--smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
+-     */
-+smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages, int stage) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" stage=%d"
+-    assert(!status->default_nan_mode);
+-
- # strongarm.c
+-    switch (status->float_2nan_prop_rule) {
- strongarm_uart_update_parameters(const char *label, int speed, char parity, int data_bits, int stop_bits) "%s speed=%d parity=%c data=%d stop=%d"
+-    case float_2nan_prop_s_ab:
 -        if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_s_ba:
 -        if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_ab:
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_ba:
 -        if (is_nan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_x87:
 -        /*
 -         * This implements x87 NaN propagation rules:
 -         * SNaN + QNaN => return the QNaN
 -         * two SNaNs => return the one with the larger significand, silenced
 -         * two QNaNs => return the one with the larger significand
 -         * SNaN and a non-NaN => return the SNaN, silenced
 -         * QNaN and a non-NaN => return the QNaN
 -         *
 -         * If we get down to comparing significands and they are the same,
 -         * return the NaN with the positive sign bit (if any).
 -         */
 -        if (is_snan(a_cls)) {
 -            if (is_snan(b_cls)) {
 -                return aIsLargerSignificand ? 0 : 1;
 -            }
 -            return is_qnan(b_cls) ? 1 : 0;
 -        } else if (is_qnan(a_cls)) {
 -            if (is_snan(b_cls) || !is_qnan(b_cls)) {
 -                return 0;
 -            } else {
 -                return aIsLargerSignificand ? 0 : 1;
 -            }
 -        } else {
 -            return 1;
 -        }
 -    default:
 -        g_assert_not_reached();
 -    }
 -}
 -
  /*----------------------------------------------------------------------------
  | Returns 1 if the double-precision floating-point value `a' is a quiet
  | NaN; otherwise returns 0.
 --
 .34.1

-[PULL 19/26] hw/arm/smmuv3: Handle translation faults according to SMMUPTWEventInfo
+[PULL 68/72] softfloat: Share code between parts_pick_nan cases
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Previously, to check if faults are enabled, it was sufficient to check
+Remember if there was an SNaN, and use that to simplify
-the current stage of translation and check the corresponding
+float_2nan_prop_s_{ab,ba} to only the snan component.
-record_faults flag.
+Then, fall through to the corresponding
 float_2nan_prop_{ab,ba} case to handle any remaining
 nans, which must be quiet.
-However, with nesting, it is possible for stage-1 (nested) translation
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-to trigger a stage-2 fault, so we check SMMUPTWEventInfo as it would
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-have the correct stage set from the page table walk.
+Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-17-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3.c | 15 ++++++++-------
+ fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
-file changed, 8 insertions(+), 7 deletions(-)
+file changed, 12 insertions(+), 20 deletions(-)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
- #include "smmuv3-internal.h"
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- #include "smmu-internal.h"
+                                      float_status *s)
+ {
--#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == SMMU_STAGE_1) ? \
++    bool have_snan = false;
--                                 (cfg)->record_faults : \
+     int cmp, which;
--                                 (cfg)->s2cfg.record_faults)
-+#define PTW_RECORD_FAULT(ptw_info, cfg) (((ptw_info).stage == SMMU_STAGE_1 && \
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+                                        (cfg)->record_faults) || \
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-+                                        ((ptw_info).stage == SMMU_STAGE_2 && \
++        have_snan = true;
-+                                        (cfg)->s2cfg.record_faults))
+     }
- /**
+     if (s->default_nan_mode) {
-  * smmuv3_trigger_irq - pulse @irq if enabled and update
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
-             event->u.f_walk_eabt.addr2 = ptw_info.addr;
+     switch (s->float_2nan_prop_rule) {
-             break;
+     case float_2nan_prop_s_ab:
-         case SMMU_PTW_ERR_TRANSLATION:
+-        if (is_snan(a->cls)) {
--            if (PTW_RECORD_FAULT(cfg)) {
+-            which = 0;
-+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
+-        } else if (is_snan(b->cls)) {
-                 event->type = SMMU_EVT_F_TRANSLATION;
+-            which = 1;
-                 event->u.f_translation.addr2 = ptw_info.addr;
+-        } else if (is_qnan(a->cls)) {
-                 event->u.f_translation.class = class;
+-            which = 0;
-@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
+-        } else {
-             }
+-            which = 1;
-             break;
++        if (have_snan) {
-         case SMMU_PTW_ERR_ADDR_SIZE:
++            which = is_snan(a->cls) ? 0 : 1;
--            if (PTW_RECORD_FAULT(cfg)) {
++            break;
-+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
+         }
-                 event->type = SMMU_EVT_F_ADDR_SIZE;
+-        break;
-                 event->u.f_addr_size.addr2 = ptw_info.addr;
+-    case float_2nan_prop_s_ba:
-                 event->u.f_addr_size.class = class;
+-        if (is_snan(b->cls)) {
-@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
+-            which = 1;
-             }
+-        } else if (is_snan(a->cls)) {
-             break;
+-            which = 0;
-         case SMMU_PTW_ERR_ACCESS:
+-        } else if (is_qnan(b->cls)) {
--            if (PTW_RECORD_FAULT(cfg)) {
+-            which = 1;
-+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
+-        } else {
-                 event->type = SMMU_EVT_F_ACCESS;
+-            which = 0;
-                 event->u.f_access.addr2 = ptw_info.addr;
+-        }
-                 event->u.f_access.class = class;
+-        break;
-@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
++        /* fall through */
-             }
+     case float_2nan_prop_ab:
-             break;
+         which = is_nan(a->cls) ? 0 : 1;
-         case SMMU_PTW_ERR_PERMISSION:
+         break;
--            if (PTW_RECORD_FAULT(cfg)) {
++    case float_2nan_prop_s_ba:
-+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
++        if (have_snan) {
-                 event->type = SMMU_EVT_F_PERMISSION;
++            which = is_snan(b->cls) ? 1 : 0;
-                 event->u.f_permission.addr2 = ptw_info.addr;
++            break;
-                 event->u.f_permission.class = class;
++        }
 +        /* fall through */
      case float_2nan_prop_ba:
          which = is_nan(b->cls) ? 1 : 0;
          break;
 --
 .34.1

-[PULL 05/26] hw/arm/smmu: Fix IPA for stage-2 events
+[PULL 69/72] softfloat: Sink frac_cmp in parts_pick_nan until needed
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-For the following events (ARM IHI 0070 F.b - 7.3 Event records):
+Move the fractional comparison to the end of the
-- F_TRANSLATION
+float_2nan_prop_x87 case.  This is not required for
-- F_ACCESS
+any other 2nan propagation rule.  Reorganize the
-- F_PERMISSION
+x87 case itself to break out of the switch when the
-- F_ADDR_SIZE
+fractional comparison is not required.
-If fault occurs at stage 2, S2 == 1 and:
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-  - If translating an IPA for a transaction (whether by input to
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-    stage 2-only configuration, or after successful stage 1 translation),
+Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
     CLASS == IN, and IPA is provided.
 At the moment only CLASS == IN is used which indicates input
 translation.
 However, this was not implemented correctly, as for stage 2, the code
 only sets the  S2 bit but not the IPA.
 This field has the same bits as FetchAddr in F_WALK_EABT which is
 populated correctly, so we don’t change that.
 The setting of this field should be done from the walker as the IPA address
 wouldn't be known in case of nesting.
 For stage 1, the spec says:
   If fault occurs at stage 1, S2 == 0 and:
   CLASS == IN, IPA is UNKNOWN.
 So, no need to set it to for stage 1, as ptw_info is initialised by zero in
 smmuv3_translate().
 Fixes: e703f7076a “hw/arm/smmuv3: Add page table walk for stage-2”
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Message-id: 20240715084519.1189624-3-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmu-common.c | 10 ++++++----
+ fpu/softfloat-parts.c.inc | 19 +++++++++----------
- hw/arm/smmuv3.c      |  4 ++++
+file changed, 9 insertions(+), 10 deletions(-)
 files changed, 10 insertions(+), 4 deletions(-)
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmu-common.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-      */
+         return a;
      if (ipa >= (1ULL << inputsize)) {
          info->type = SMMU_PTW_ERR_TRANSLATION;
 -        goto error;
 +        goto error_ipa;
      }
-     while (level < VMSA_LEVELS) {
+-    cmp = frac_cmp(a, b);
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
+-    if (cmp == 0) {
 -        cmp = a->sign < b->sign;
 -    }
 -
      switch (s->float_2nan_prop_rule) {
      case float_2nan_prop_s_ab:
          if (have_snan) {
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
           * return the NaN with the positive sign bit (if any).
           */
-         if (!PTE_AF(pte) && !cfg->s2cfg.affd) {
+         if (is_snan(a->cls)) {
-             info->type = SMMU_PTW_ERR_ACCESS;
+-            if (is_snan(b->cls)) {
--            goto error;
+-                which = cmp > 0 ? 0 : 1;
-+            goto error_ipa;
+-            } else {
 +            if (!is_snan(b->cls)) {
                  which = is_qnan(b->cls) ? 1 : 0;
 +                break;
              }
          } else if (is_qnan(a->cls)) {
              if (is_snan(b->cls) || !is_qnan(b->cls)) {
                  which = 0;
 -            } else {
 -                which = cmp > 0 ? 0 : 1;
 +                break;
              }
          } else {
              which = 1;
 +            break;
          }
++        cmp = frac_cmp(a, b);
-         s2ap = PTE_AP(pte);
++        if (cmp == 0) {
-         if (is_permission_fault_s2(s2ap, perm)) {
++            cmp = a->sign < b->sign;
-             info->type = SMMU_PTW_ERR_PERMISSION;
++        }
--            goto error;
++        which = cmp > 0 ? 0 : 1;
-+            goto error_ipa;
+         break;
-         }
+     default:
+         g_assert_not_reached();
          /*
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
           */
          if (gpa >= (1ULL << cfg->s2cfg.eff_ps)) {
              info->type = SMMU_PTW_ERR_ADDR_SIZE;
 -            goto error;
 +            goto error_ipa;
          }
          tlbe->entry.translated_addr = gpa;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
      }
      info->type = SMMU_PTW_ERR_TRANSLATION;
 +error_ipa:
 +    info->addr = ipa;
  error:
      info->stage = 2;
      tlbe->entry.perm = IOMMU_NONE;
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              if (PTW_RECORD_FAULT(cfg)) {
                  event.type = SMMU_EVT_F_TRANSLATION;
                  event.u.f_translation.addr = addr;
 +                event.u.f_translation.addr2 = ptw_info.addr;
                  event.u.f_translation.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              if (PTW_RECORD_FAULT(cfg)) {
                  event.type = SMMU_EVT_F_ADDR_SIZE;
                  event.u.f_addr_size.addr = addr;
 +                event.u.f_addr_size.addr2 = ptw_info.addr;
                  event.u.f_addr_size.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              if (PTW_RECORD_FAULT(cfg)) {
                  event.type = SMMU_EVT_F_ACCESS;
                  event.u.f_access.addr = addr;
 +                event.u.f_access.addr2 = ptw_info.addr;
                  event.u.f_access.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              if (PTW_RECORD_FAULT(cfg)) {
                  event.type = SMMU_EVT_F_PERMISSION;
                  event.u.f_permission.addr = addr;
 +                event.u.f_permission.addr2 = ptw_info.addr;
                  event.u.f_permission.rnw = flag & 0x1;
              }
              break;
 --
 .34.1

-[PULL 17/26] hw/arm/smmu: Support nesting in the rest of commands
+[PULL 70/72] softfloat: Replace WHICH with RET in parts_pick_nan
-From: Mostafa Saleh <smostafa@google.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Some commands need rework for nesting, as they used to assume S1
+Replace the "index" selecting between A and B with a result variable
-and S2 are mutually exclusive:
+of the proper type.  This improves clarity within the function.
-- CMD_TLBI_NH_ASID: Consider VMID if stage-2 is supported
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-- CMD_TLBI_NH_ALL: Consider VMID if stage-2 is supported, otherwise
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-  invalidate everything, this required a new vmid invalidation
+Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
   function for stage-1 only (ASID >= 0)
 Also, rework trace events to reflect the new implementation.
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Message-id: 20240715084519.1189624-15-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/hw/arm/smmu-common.h |  1 +
+ fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
- hw/arm/smmu-common.c         | 16 ++++++++++++++++
+file changed, 13 insertions(+), 15 deletions(-)
  hw/arm/smmuv3.c              | 28 ++++++++++++++++++++++++++--
  hw/arm/trace-events          |  4 +++-
 files changed, 46 insertions(+), 3 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- void smmu_iotlb_inv_all(SMMUState *s);
+                                      float_status *s)
  void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid);
  void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
 +void smmu_iotlb_inv_vmid_s1(SMMUState *s, int vmid);
  void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                           uint8_t tg, uint64_t num_pages, uint8_t ttl);
  void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
      return SMMU_IOTLB_VMID(*iotlb_key) == vmid;
  }
 +static gboolean smmu_hash_remove_by_vmid_s1(gpointer key, gpointer value,
 +                                            gpointer user_data)
 +{
 +    int vmid = *(int *)user_data;
 +    SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
 +
 +    return (SMMU_IOTLB_VMID(*iotlb_key) == vmid) &&
 +           (SMMU_IOTLB_ASID(*iotlb_key) >= 0);
 +}
 +
  static gboolean smmu_hash_remove_by_asid_vmid_iova(gpointer key, gpointer value,
                                                gpointer user_data)
  {
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
+     bool have_snan = false;
-     g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid, &vmid);
+-    int cmp, which;
- }
++    FloatPartsN *ret;
++    int cmp;
-+inline void smmu_iotlb_inv_vmid_s1(SMMUState *s, int vmid)
-+{
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+    trace_smmu_iotlb_inv_vmid_s1(vmid);
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-+    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid_s1, &vmid);
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-+}
+     switch (s->float_2nan_prop_rule) {
-+
+     case float_2nan_prop_s_ab:
- /* VMSAv8-64 Translation */
+         if (have_snan) {
+-            which = is_snan(a->cls) ? 0 : 1;
- /**
++            ret = is_snan(a->cls) ? a : b;
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+             break;
-index XXXXXXX..XXXXXXX 100644
+         }
---- a/hw/arm/smmuv3.c
+         /* fall through */
-+++ b/hw/arm/smmuv3.c
+     case float_2nan_prop_ab:
-@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
+-        which = is_nan(a->cls) ? 0 : 1;
-         case SMMU_CMD_TLBI_NH_ASID:
++        ret = is_nan(a->cls) ? a : b;
-         {
+         break;
-             int asid = CMD_ASID(&cmd);
+     case float_2nan_prop_s_ba:
-+            int vmid = -1;
+         if (have_snan) {
+-            which = is_snan(b->cls) ? 1 : 0;
-             if (!STAGE1_SUPPORTED(s)) {
++            ret = is_snan(b->cls) ? b : a;
-                 cmd_error = SMMU_CERROR_ILL;
+             break;
          }
          /* fall through */
      case float_2nan_prop_ba:
 -        which = is_nan(b->cls) ? 1 : 0;
 +        ret = is_nan(b->cls) ? b : a;
          break;
      case float_2nan_prop_x87:
          /*
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
           */
          if (is_snan(a->cls)) {
              if (!is_snan(b->cls)) {
 -                which = is_qnan(b->cls) ? 1 : 0;
 +                ret = is_qnan(b->cls) ? b : a;
                  break;
              }
+         } else if (is_qnan(a->cls)) {
-+            /*
+             if (is_snan(b->cls) || !is_qnan(b->cls)) {
-+             * VMID is only matched when stage 2 is supported, otherwise set it
+-                which = 0;
-+             * to -1 as the value used for stage-1 only VMIDs.
++                ret = a;
-+             */
+                 break;
-+            if (STAGE2_SUPPORTED(s)) {
+             }
-+                vmid = CMD_VMID(&cmd);
+         } else {
-+            }
+-            which = 1;
-+
++            ret = b;
              trace_smmuv3_cmdq_tlbi_nh_asid(asid);
              smmu_inv_notifiers_all(&s->smmu_state);
 -            smmu_iotlb_inv_asid_vmid(bs, asid, -1);
 +            smmu_iotlb_inv_asid_vmid(bs, asid, vmid);
              break;
          }
-         case SMMU_CMD_TLBI_NH_ALL:
+         cmp = frac_cmp(a, b);
-+        {
+         if (cmp == 0) {
-+            int vmid = -1;
+             cmp = a->sign < b->sign;
-+
+         }
-             if (!STAGE1_SUPPORTED(s)) {
+-        which = cmp > 0 ? 0 : 1;
-                 cmd_error = SMMU_CERROR_ILL;
++        ret = cmp > 0 ? a : b;
-                 break;
+         break;
-             }
+     default:
-+
+         g_assert_not_reached();
-+            /*
+     }
-+             * If stage-2 is supported, invalidate for this VMID only, otherwise
-+             * invalidate the whole thing.
+-    if (which) {
-+             */
+-        a = b;
-+            if (STAGE2_SUPPORTED(s)) {
++    if (is_snan(ret->cls)) {
-+                vmid = CMD_VMID(&cmd);
++        parts_silence_nan(ret, s);
-+                trace_smmuv3_cmdq_tlbi_nh(vmid);
+     }
-+                smmu_iotlb_inv_vmid_s1(bs, vmid);
+-    if (is_snan(a->cls)) {
-+                break;
+-        parts_silence_nan(a, s);
-+            }
+-    }
-             QEMU_FALLTHROUGH;
+-    return a;
-+        }
++    return ret;
-         case SMMU_CMD_TLBI_NSNH_ALL:
+ }
--            trace_smmuv3_cmdq_tlbi_nh();
-+            trace_smmuv3_cmdq_tlbi_nsnh();
+ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
              smmu_inv_notifiers_all(&s->smmu_state);
              smmu_iotlb_inv_all(bs);
              break;
 diff --git a/hw/arm/trace-events b/hw/arm/trace-events
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/trace-events
 +++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "base
  smmu_iotlb_inv_all(void) "IOTLB invalidate all"
  smmu_iotlb_inv_asid_vmid(int asid, int vmid) "IOTLB invalidate asid=%d vmid=%d"
  smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
 +smmu_iotlb_inv_vmid_s1(int vmid) "IOTLB invalidate vmid=%d"
  smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
  smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
  smmu_iotlb_lookup_hit(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_cd(uint32_t sid) "sid=0x%x"
  smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
  smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
  smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf, int stage) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d stage=%d"
 -smmuv3_cmdq_tlbi_nh(void) ""
 +smmuv3_cmdq_tlbi_nh(int vmid) "vmid=%d"
 +smmuv3_cmdq_tlbi_nsnh(void) ""
  smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
  smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
  smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
 --
 .34.1

-[PULL 06/26] hw/arm/smmuv3: Fix encoding of CLASS in events
+[PULL 71/72] MAINTAINERS: update email address for Leif Lindholm
-From: Mostafa Saleh <smostafa@google.com>
+From: Leif Lindholm <quic_llindhol@quicinc.com>
-The SMMUv3 spec (ARM IHI 0070 F.b - 7.3 Event records) defines the
+I'm migrating to Qualcomm's new open source email infrastructure, so
-class of events faults as:
+update my email address, and update the mailmap to match.
-CLASS: The class of the operation that caused the fault:
+Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
-- 0b00: CD, CD fetch.
+Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
-- 0b01: TTD, Stage 1 translation table fetch.
+Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
-- 0b10: IN, Input address
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-However, this value was not set and left as 0 which means CD and not
+Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
 IN (0b10).
 Another problem was that stage-2 class is considered IN not TT for
 EABT, according to the spec:
     Translation of an IPA after successful stage 1 translation (or,
     in stage 2-only configuration, an input IPA)
     - S2 == 1 (stage 2), CLASS == IN (Input to stage)
 This would change soon when nested translations are supported.
 While at it, add an enum for class as it would be used for nesting.
 However, at the moment stage-1 and stage-2 use the same class values,
 except for EABT.
 Fixes: 9bde7f0674 “hw/arm/smmuv3: Implement translate callback”
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 20240715084519.1189624-4-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h | 6 ++++++
+ MAINTAINERS | 2 +-
- hw/arm/smmuv3.c          | 8 +++++++-
+ .mailmap    | 5 +++--
-files changed, 13 insertions(+), 1 deletion(-)
+files changed, 4 insertions(+), 3 deletions(-)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/MAINTAINERS
-+++ b/hw/arm/smmuv3-internal.h
++++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ typedef enum SMMUTranslationStatus {
+@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
-     SMMU_TRANS_SUCCESS,
+ SBSA-REF
- } SMMUTranslationStatus;
+ M: Radoslaw Biernacki <rad@semihalf.com>
+ M: Peter Maydell <peter.maydell@linaro.org>
-+typedef enum SMMUTranslationClass {
+-R: Leif Lindholm <quic_llindhol@quicinc.com>
-+    SMMU_CLASS_CD,
++R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
-+    SMMU_CLASS_TT,
+ R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
-+    SMMU_CLASS_IN,
+ L: qemu-arm@nongnu.org
-+} SMMUTranslationClass;
+ S: Maintained
-+
+diff --git a/.mailmap b/.mailmap
  /* MMIO Registers */
  REG32(IDR0,                0x0)
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/.mailmap
-+++ b/hw/arm/smmuv3.c
++++ b/.mailmap
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
-             event.type = SMMU_EVT_F_WALK_EABT;
+ Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
-             event.u.f_walk_eabt.addr = addr;
+ James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
-             event.u.f_walk_eabt.rnw = flag & 0x1;
+ Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
--            event.u.f_walk_eabt.class = 0x1;
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-+            /* Stage-2 (only) is class IN while stage-1 is class TT */
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
-+            event.u.f_walk_eabt.class = (ptw_info.stage == 2) ?
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
-+                                         SMMU_CLASS_IN : SMMU_CLASS_TT;
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
-             event.u.f_walk_eabt.addr2 = ptw_info.addr;
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
-             break;
+ Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
-         case SMMU_PTW_ERR_TRANSLATION:
+ Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+ Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
                  event.type = SMMU_EVT_F_TRANSLATION;
                  event.u.f_translation.addr = addr;
                  event.u.f_translation.addr2 = ptw_info.addr;
 +                event.u.f_translation.class = SMMU_CLASS_IN;
                  event.u.f_translation.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                  event.type = SMMU_EVT_F_ADDR_SIZE;
                  event.u.f_addr_size.addr = addr;
                  event.u.f_addr_size.addr2 = ptw_info.addr;
 +                event.u.f_translation.class = SMMU_CLASS_IN;
                  event.u.f_addr_size.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                  event.type = SMMU_EVT_F_ACCESS;
                  event.u.f_access.addr = addr;
                  event.u.f_access.addr2 = ptw_info.addr;
 +                event.u.f_translation.class = SMMU_CLASS_IN;
                  event.u.f_access.rnw = flag & 0x1;
              }
              break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                  event.type = SMMU_EVT_F_PERMISSION;
                  event.u.f_permission.addr = addr;
                  event.u.f_permission.addr2 = ptw_info.addr;
 +                event.u.f_translation.class = SMMU_CLASS_IN;
                  event.u.f_permission.rnw = flag & 0x1;
              }
              break;
 --
 .34.1

-[PULL 04/26] hw/arm/smmu-common: Add missing size check for stage-1
+[PULL 72/72] MAINTAINERS: Add correct email address for Vikram Garhwal
-From: Mostafa Saleh <smostafa@google.com>
+From: Vikram Garhwal <vikram.garhwal@bytedance.com>
-According to the SMMU architecture specification (ARM IHI 0070 F.b),
+Previously, maintainer role was paused due to inactive email id. Commit id:
-in “3.4 Address sizes”
+c009d715721861984c4987bcc78b7ee183e86d75.
     The address output from the translation causes a stage 1 Address Size
     fault if it exceeds the range of the effective IPA size for the given CD.
-However, this check was missing.
+Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
+Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
-There is already a similar check for stage-2 against effective PA.
+Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
 Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Mostafa Saleh <smostafa@google.com>
 Message-id: 20240715084519.1189624-2-smostafa@google.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmu-common.c | 10 ++++++++++
+ MAINTAINERS | 2 ++
-file changed, 10 insertions(+)
+file changed, 2 insertions(+)
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
+diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/MAINTAINERS
-+++ b/hw/arm/smmu-common.c
++++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
+@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
-             goto error;
-         }
+ Xilinx CAN
+ M: Francisco Iglesias <francisco.iglesias@amd.com>
-+        /*
++M: Vikram Garhwal <vikram.garhwal@bytedance.com>
-+         * The address output from the translation causes a stage 1 Address
+ S: Maintained
-+         * Size fault if it exceeds the range of the effective IPA size for
+ F: hw/net/can/xlnx-*
-+         * the given CD.
+ F: include/hw/net/xlnx-*
-+         */
+@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
-+        if (gpa >= (1ULL << cfg->oas)) {
+ CAN bus subsystem and hardware
-+            info->type = SMMU_PTW_ERR_ADDR_SIZE;
+ M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
-+            goto error;
+ M: Francisco Iglesias <francisco.iglesias@amd.com>
-+        }
++M: Vikram Garhwal <vikram.garhwal@bytedance.com>
-+
+ S: Maintained
-         tlbe->entry.translated_addr = gpa;
+ W: https://canbus.pages.fel.cvut.cz/
-         tlbe->entry.iova = iova & ~mask;
+ F: net/can/*
          tlbe->entry.addr_mask = mask;
 --
 .34.1

Hi; hopefully this is the last arm pullreq before softfreeze.
There's a handful of miscellaneous bug fixes here, but the
bulk of the pullreq is Mostafa's implementation of 2-stage
translation in the SMMUv3.

thanks
-- PMM

The following changes since commit d74ec4d7dda6322bcc51d1b13ccbd993d3574795:

Merge tag 'pull-trivial-patches' of https://gitlab.com/mjt0k/qemu into staging (2024-07-18 10:07:23 +1000)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20240718

for you to fetch changes up to 30a1690f2402e6c1582d5b3ebcf7940bfe2fad4b:

hvf: arm: Do not advance PC when raising an exception (2024-07-18 13:49:30 +0100)

----------------------------------------------------------------
target-arm queue:
 * Fix handling of LDAPR/STLR with negative offset
 * LDAPR should honour SCTLR_ELx.nAA
 * Use float_status copy in sme_fmopa_s
 * hw/display/bcm2835_fb: fix fb_use_offsets condition
 * hw/arm/smmuv3: Support and advertise nesting
 * Use FPST_F16 for SME FMOPA (widening)
 * tests/arm-cpu-features: Do not assume PMU availability
 * hvf: arm: Do not advance PC when raising an exception

----------------------------------------------------------------
Akihiko Odaki (2):
      tests/arm-cpu-features: Do not assume PMU availability
      hvf: arm: Do not advance PC when raising an exception

Daniyal Khan (2):
      target/arm: Use float_status copy in sme_fmopa_s
      tests/tcg/aarch64: Add test cases for SME FMOPA (widening)

Mostafa Saleh (18):
      hw/arm/smmu-common: Add missing size check for stage-1
      hw/arm/smmu: Fix IPA for stage-2 events
      hw/arm/smmuv3: Fix encoding of CLASS in events
      hw/arm/smmu: Use enum for SMMU stage
      hw/arm/smmu: Split smmuv3_translate()
      hw/arm/smmu: Consolidate ASID and VMID types
      hw/arm/smmu: Introduce CACHED_ENTRY_TO_ADDR
      hw/arm/smmuv3: Translate CD and TT using stage-2 table
      hw/arm/smmu-common: Rework TLB lookup for nesting
      hw/arm/smmu-common: Add support for nested TLB
      hw/arm/smmu-common: Support nested translation
      hw/arm/smmu: Support nesting in smmuv3_range_inval()
      hw/arm/smmu: Introduce smmu_iotlb_inv_asid_vmid
      hw/arm/smmu: Support nesting in the rest of commands
      hw/arm/smmuv3: Support nested SMMUs in smmuv3_notify_iova()
      hw/arm/smmuv3: Handle translation faults according to SMMUPTWEventInfo
      hw/arm/smmuv3: Support and advertise nesting
      hw/arm/smmu: Refactor SMMU OAS

Peter Maydell (2):
      target/arm: Fix handling of LDAPR/STLR with negative offset
      target/arm: LDAPR should honour SCTLR_ELx.nAA

Richard Henderson (1):
      target/arm: Use FPST_F16 for SME FMOPA (widening)

SamJakob (1):
      hw/display/bcm2835_fb: fix fb_use_offsets condition

hw/arm/smmuv3-internal.h          |  19 +-
 include/hw/arm/smmu-common.h      |  46 +++-
 target/arm/tcg/a64.decode         |   2 +-
 hw/arm/smmu-common.c              | 312 ++++++++++++++++++++++---
 hw/arm/smmuv3.c                   | 467 +++++++++++++++++++++++++-------------
 hw/display/bcm2835_fb.c           |   2 +-
 target/arm/hvf/hvf.c              |   1 +
 target/arm/tcg/sme_helper.c       |   2 +-
 target/arm/tcg/translate-a64.c    |   2 +-
 target/arm/tcg/translate-sme.c    |  12 +-
 tests/qtest/arm-cpu-features.c    |  13 +-
 tests/tcg/aarch64/sme-fmopa-1.c   |  63 +++++
 tests/tcg/aarch64/sme-fmopa-2.c   |  56 +++++
 tests/tcg/aarch64/sme-fmopa-3.c   |  63 +++++
 hw/arm/trace-events               |  26 ++-
 tests/tcg/aarch64/Makefile.target |   5 +-
 16 files changed, 846 insertions(+), 245 deletions(-)
 create mode 100644 tests/tcg/aarch64/sme-fmopa-1.c
 create mode 100644 tests/tcg/aarch64/sme-fmopa-2.c
 create mode 100644 tests/tcg/aarch64/sme-fmopa-3.c

When we converted the LDAPR/STLR instructions to decodetree we
accidentally introduced a regression where the offset is negative.
The 9-bit immediate field is signed, and the old hand decoder
correctly used sextract32() to get it out of the insn word,
but the ldapr_stlr_i pattern in the decode file used "imm:9"
instead of "imm:s9", so it treated the field as unsigned.

Fix the pattern to treat the field as a signed immediate.

Cc: qemu-stable@nongnu.org
Fixes: 2521b6073b7 ("target/arm: Convert LDAPR/STLR (imm) to decodetree")
Resolves: https://gitlab.com/qemu-project/qemu/-/issues/2419
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240709134504.3500007-2-peter.maydell@linaro.org
---
 target/arm/tcg/a64.decode | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/a64.decode
+++ b/target/arm/tcg/a64.decode
@@ -XXX,XX +XXX,XX @@ LDAPR           sz:2 111 0 00 1 0 1 11111 1100 00 rn:5 rt:5
 LDRA            11 111 0 00 m:1 . 1 ......... w:1 1 rn:5 rt:5 imm=%ldra_imm
 
 &ldapr_stlr_i   rn rt imm sz sign ext
-@ldapr_stlr_i   .. ...... .. . imm:9 .. rn:5 rt:5 &ldapr_stlr_i
+@ldapr_stlr_i   .. ...... .. . imm:s9 .. rn:5 rt:5 &ldapr_stlr_i
 STLR_i          sz:2 011001 00 0 ......... 00 ..... ..... @ldapr_stlr_i sign=0 ext=0
 LDAPR_i         sz:2 011001 01 0 ......... 00 ..... ..... @ldapr_stlr_i sign=0 ext=0
 LDAPR_i         00 011001 10 0 ......... 00 ..... ..... @ldapr_stlr_i sign=1 ext=0 sz=0
-- 
2.34.1

In commit c1a1f80518d360b when we added the FEAT_LSE2 relaxations to
the alignment requirements for atomic and ordered loads and stores,
we didn't quite get it right for LDAPR/LDAPRH/LDAPRB with no
immediate offset.  These instructions were handled in the old decoder
as part of disas_ldst_atomic(), but unlike all the other insns that
function decoded (LDADD, LDCLR, etc) these insns are "ordered", not
"atomic", so they should be using check_ordered_align() rather than
check_atomic_align().  Commit c1a1f80518d360b used
check_atomic_align() regardless for everything in
disas_ldst_atomic().  We then carried that incorrect check over in
the decodetree conversion, where LDAPR/LDAPRH/LDAPRB are now handled
by trans_LDAPR().

The effect is that when FEAT_LSE2 is implemented, these instructions
don't honour the SCTLR_ELx.nAA bit and will generate alignment
faults when they should not.

(The LDAPR insns with an immediate offset were in disas_ldst_ldapr_stlr()
and then in trans_LDAPR_i() and trans_STLR_i(), and have always used
the correct check_ordered_align().)

Use check_ordered_align() in trans_LDAPR().

Cc: qemu-stable@nongnu.org
Fixes: c1a1f80518d360b ("target/arm: Relax ordered/atomic alignment checks for LSE2")
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240709134504.3500007-3-peter.maydell@linaro.org
---
 target/arm/tcg/translate-a64.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-a64.c
+++ b/target/arm/tcg/translate-a64.c
@@ -XXX,XX +XXX,XX @@ static bool trans_LDAPR(DisasContext *s, arg_LDAPR *a)
     if (a->rn == 31) {
         gen_check_sp_alignment(s);
     }
-    mop = check_atomic_align(s, a->rn, a->sz);
+    mop = check_ordered_align(s, a->rn, 0, false, a->sz);
     clean_addr = gen_mte_check1(s, cpu_reg_sp(s, a->rn), false,
                                 a->rn != 31, mop);
     /*
-- 
2.34.1

From: SamJakob <me@samjakob.com>

It is common practice when implementing double-buffering on VideoCore
to do so by multiplying the height of the virtual buffer by the
number of virtual screens desired (i.e., two - in the case of
double-bufferring).

At present, this won't work in QEMU because the logic in
fb_use_offsets require that both the virtual width and height exceed
their physical counterparts.

This appears to be unintentional/a typo and indeed the comment
states; "Experimentally, the hardware seems to do this only if the
viewport size is larger than the physical screen".  The
viewport/virtual size would be larger than the physical size if
either virtual dimension were larger than their physical counterparts
and not necessarily both.

Signed-off-by: SamJakob <me@samjakob.com>
Message-id: 20240713160353.62410-1-me@samjakob.com
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/display/bcm2835_fb.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/hw/display/bcm2835_fb.c b/hw/display/bcm2835_fb.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/display/bcm2835_fb.c
+++ b/hw/display/bcm2835_fb.c
@@ -XXX,XX +XXX,XX @@ static bool fb_use_offsets(BCM2835FBConfig *config)
      * viewport size is larger than the physical screen. (It doesn't
      * prevent the guest setting this silly viewport setting, though...)
      */
-    return config->xres_virtual > config->xres &&
+    return config->xres_virtual > config->xres ||
         config->yres_virtual > config->yres;
 }
 
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

According to the SMMU architecture specification (ARM IHI 0070 F.b),
in “3.4 Address sizes”
    The address output from the translation causes a stage 1 Address Size
    fault if it exceeds the range of the effective IPA size for the given CD.

However, this check was missing.

There is already a similar check for stage-2 against effective PA.

Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Message-id: 20240715084519.1189624-2-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-common.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
             goto error;
         }
 
+        /*
+         * The address output from the translation causes a stage 1 Address
+         * Size fault if it exceeds the range of the effective IPA size for
+         * the given CD.
+         */
+        if (gpa >= (1ULL << cfg->oas)) {
+            info->type = SMMU_PTW_ERR_ADDR_SIZE;
+            goto error;
+        }
+
         tlbe->entry.translated_addr = gpa;
         tlbe->entry.iova = iova & ~mask;
         tlbe->entry.addr_mask = mask;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

For the following events (ARM IHI 0070 F.b - 7.3 Event records):
- F_TRANSLATION
- F_ACCESS
- F_PERMISSION
- F_ADDR_SIZE

If fault occurs at stage 2, S2 == 1 and:
  - If translating an IPA for a transaction (whether by input to
    stage 2-only configuration, or after successful stage 1 translation),
    CLASS == IN, and IPA is provided.

At the moment only CLASS == IN is used which indicates input
translation.

However, this was not implemented correctly, as for stage 2, the code
only sets the  S2 bit but not the IPA.

This field has the same bits as FetchAddr in F_WALK_EABT which is
populated correctly, so we don’t change that.
The setting of this field should be done from the walker as the IPA address
wouldn't be known in case of nesting.

For stage 1, the spec says:
  If fault occurs at stage 1, S2 == 0 and:
  CLASS == IN, IPA is UNKNOWN.

So, no need to set it to for stage 1, as ptw_info is initialised by zero in
smmuv3_translate().

Fixes: e703f7076a “hw/arm/smmuv3: Add page table walk for stage-2”
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Message-id: 20240715084519.1189624-3-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-common.c | 10 ++++++----
 hw/arm/smmuv3.c      |  4 ++++
 2 files changed, 10 insertions(+), 4 deletions(-)

diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
      */
     if (ipa >= (1ULL << inputsize)) {
         info->type = SMMU_PTW_ERR_TRANSLATION;
-        goto error;
+        goto error_ipa;
     }
 
     while (level < VMSA_LEVELS) {
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
          */
         if (!PTE_AF(pte) && !cfg->s2cfg.affd) {
             info->type = SMMU_PTW_ERR_ACCESS;
-            goto error;
+            goto error_ipa;
         }
 
         s2ap = PTE_AP(pte);
         if (is_permission_fault_s2(s2ap, perm)) {
             info->type = SMMU_PTW_ERR_PERMISSION;
-            goto error;
+            goto error_ipa;
         }
 
         /*
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
          */
         if (gpa >= (1ULL << cfg->s2cfg.eff_ps)) {
             info->type = SMMU_PTW_ERR_ADDR_SIZE;
-            goto error;
+            goto error_ipa;
         }
 
         tlbe->entry.translated_addr = gpa;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
     }
     info->type = SMMU_PTW_ERR_TRANSLATION;
 
+error_ipa:
+    info->addr = ipa;
 error:
     info->stage = 2;
     tlbe->entry.perm = IOMMU_NONE;
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             if (PTW_RECORD_FAULT(cfg)) {
                 event.type = SMMU_EVT_F_TRANSLATION;
                 event.u.f_translation.addr = addr;
+                event.u.f_translation.addr2 = ptw_info.addr;
                 event.u.f_translation.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             if (PTW_RECORD_FAULT(cfg)) {
                 event.type = SMMU_EVT_F_ADDR_SIZE;
                 event.u.f_addr_size.addr = addr;
+                event.u.f_addr_size.addr2 = ptw_info.addr;
                 event.u.f_addr_size.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             if (PTW_RECORD_FAULT(cfg)) {
                 event.type = SMMU_EVT_F_ACCESS;
                 event.u.f_access.addr = addr;
+                event.u.f_access.addr2 = ptw_info.addr;
                 event.u.f_access.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             if (PTW_RECORD_FAULT(cfg)) {
                 event.type = SMMU_EVT_F_PERMISSION;
                 event.u.f_permission.addr = addr;
+                event.u.f_permission.addr2 = ptw_info.addr;
                 event.u.f_permission.rnw = flag & 0x1;
             }
             break;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

The SMMUv3 spec (ARM IHI 0070 F.b - 7.3 Event records) defines the
class of events faults as:

CLASS: The class of the operation that caused the fault:
- 0b00: CD, CD fetch.
- 0b01: TTD, Stage 1 translation table fetch.
- 0b10: IN, Input address

However, this value was not set and left as 0 which means CD and not
IN (0b10).

Another problem was that stage-2 class is considered IN not TT for
EABT, according to the spec:
    Translation of an IPA after successful stage 1 translation (or,
    in stage 2-only configuration, an input IPA)
    - S2 == 1 (stage 2), CLASS == IN (Input to stage)

This would change soon when nested translations are supported.

While at it, add an enum for class as it would be used for nesting.
However, at the moment stage-1 and stage-2 use the same class values,
except for EABT.

Fixes: 9bde7f0674 “hw/arm/smmuv3: Implement translate callback”
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Message-id: 20240715084519.1189624-4-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h | 6 ++++++
 hw/arm/smmuv3.c          | 8 +++++++-
 2 files changed, 13 insertions(+), 1 deletion(-)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ typedef enum SMMUTranslationStatus {
     SMMU_TRANS_SUCCESS,
 } SMMUTranslationStatus;
 
+typedef enum SMMUTranslationClass {
+    SMMU_CLASS_CD,
+    SMMU_CLASS_TT,
+    SMMU_CLASS_IN,
+} SMMUTranslationClass;
+
 /* MMIO Registers */
 
 REG32(IDR0,                0x0)
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             event.type = SMMU_EVT_F_WALK_EABT;
             event.u.f_walk_eabt.addr = addr;
             event.u.f_walk_eabt.rnw = flag & 0x1;
-            event.u.f_walk_eabt.class = 0x1;
+            /* Stage-2 (only) is class IN while stage-1 is class TT */
+            event.u.f_walk_eabt.class = (ptw_info.stage == 2) ?
+                                         SMMU_CLASS_IN : SMMU_CLASS_TT;
             event.u.f_walk_eabt.addr2 = ptw_info.addr;
             break;
         case SMMU_PTW_ERR_TRANSLATION:
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                 event.type = SMMU_EVT_F_TRANSLATION;
                 event.u.f_translation.addr = addr;
                 event.u.f_translation.addr2 = ptw_info.addr;
+                event.u.f_translation.class = SMMU_CLASS_IN;
                 event.u.f_translation.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                 event.type = SMMU_EVT_F_ADDR_SIZE;
                 event.u.f_addr_size.addr = addr;
                 event.u.f_addr_size.addr2 = ptw_info.addr;
+                event.u.f_translation.class = SMMU_CLASS_IN;
                 event.u.f_addr_size.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                 event.type = SMMU_EVT_F_ACCESS;
                 event.u.f_access.addr = addr;
                 event.u.f_access.addr2 = ptw_info.addr;
+                event.u.f_translation.class = SMMU_CLASS_IN;
                 event.u.f_access.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                 event.type = SMMU_EVT_F_PERMISSION;
                 event.u.f_permission.addr = addr;
                 event.u.f_permission.addr2 = ptw_info.addr;
+                event.u.f_translation.class = SMMU_CLASS_IN;
                 event.u.f_permission.rnw = flag & 0x1;
             }
             break;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Currently, translation stage is represented as an int, where 1 is stage-1 and
2 is stage-2, when nested is added, 3 would be confusing to represent nesting,
so we use an enum instead.

While keeping the same values, this is useful for:
 - Doing tricks with bit masks, where BIT(0) is stage-1 and BIT(1) is
   stage-2 and both is nested.
 - Tracing, as stage is printed as int.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Message-id: 20240715084519.1189624-5-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h | 11 +++++++++--
 hw/arm/smmu-common.c         | 14 +++++++-------
 hw/arm/smmuv3.c              | 17 +++++++++--------
 3 files changed, 25 insertions(+), 17 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef enum {
     SMMU_PTW_ERR_PERMISSION,  /* Permission fault */
 } SMMUPTWEventType;
 
+/* SMMU Stage */
+typedef enum {
+    SMMU_STAGE_1 = 1,
+    SMMU_STAGE_2,
+    SMMU_NESTED,
+} SMMUStage;
+
 typedef struct SMMUPTWEventInfo {
-    int stage;
+    SMMUStage stage;
     SMMUPTWEventType type;
     dma_addr_t addr; /* fetched address that induced an abort, if any */
 } SMMUPTWEventInfo;
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUS2Cfg {
  */
 typedef struct SMMUTransCfg {
     /* Shared fields between stage-1 and stage-2. */
-    int stage;                 /* translation stage */
+    SMMUStage stage;           /* translation stage */
     bool disabled;             /* smmu is disabled */
     bool bypassed;             /* translation is bypassed */
     bool aborted;              /* translation is aborted */
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
                           SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
     dma_addr_t baseaddr, indexmask;
-    int stage = cfg->stage;
+    SMMUStage stage = cfg->stage;
     SMMUTransTableInfo *tt = select_tt(cfg, iova);
     uint8_t level, granule_sz, inputsize, stride;
 
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
     info->type = SMMU_PTW_ERR_TRANSLATION;
 
 error:
-    info->stage = 1;
+    info->stage = SMMU_STAGE_1;
     tlbe->entry.perm = IOMMU_NONE;
     return -EINVAL;
 }
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
                           dma_addr_t ipa, IOMMUAccessFlags perm,
                           SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
-    const int stage = 2;
+    const SMMUStage stage = SMMU_STAGE_2;
     int granule_sz = cfg->s2cfg.granule_sz;
     /* ARM DDI0487I.a: Table D8-7. */
     int inputsize = 64 - cfg->s2cfg.tsz;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
 error_ipa:
     info->addr = ipa;
 error:
-    info->stage = 2;
+    info->stage = SMMU_STAGE_2;
     tlbe->entry.perm = IOMMU_NONE;
     return -EINVAL;
 }
@@ -XXX,XX +XXX,XX @@ error:
 int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
              SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
-    if (cfg->stage == 1) {
+    if (cfg->stage == SMMU_STAGE_1) {
         return smmu_ptw_64_s1(cfg, iova, perm, tlbe, info);
-    } else if (cfg->stage == 2) {
+    } else if (cfg->stage == SMMU_STAGE_2) {
         /*
          * If bypassing stage 1(or unimplemented), the input address is passed
          * directly to stage 2 as IPA. If the input address of a transaction
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
          */
         if (iova >= (1ULL << cfg->oas)) {
             info->type = SMMU_PTW_ERR_ADDR_SIZE;
-            info->stage = 1;
+            info->stage = SMMU_STAGE_1;
             tlbe->entry.perm = IOMMU_NONE;
             return -EINVAL;
         }
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@
 #include "smmuv3-internal.h"
 #include "smmu-internal.h"
 
-#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == 1) ? (cfg)->record_faults : \
+#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == SMMU_STAGE_1) ? \
+                                 (cfg)->record_faults : \
                                  (cfg)->s2cfg.record_faults)
 
 /**
@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
 
 static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
 {
-    cfg->stage = 2;
+    cfg->stage = SMMU_STAGE_2;
 
     if (STE_S2AA64(ste) == 0x0) {
         qemu_log_mask(LOG_UNIMP,
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
 
     /* we support only those at the moment */
     cfg->aa64 = true;
-    cfg->stage = 1;
+    cfg->stage = SMMU_STAGE_1;
 
     cfg->oas = oas2bits(CD_IPS(cd));
     cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
@@ -XXX,XX +XXX,XX @@ static int smmuv3_decode_config(IOMMUMemoryRegion *mr, SMMUTransCfg *cfg,
         return ret;
     }
 
-    if (cfg->aborted || cfg->bypassed || (cfg->stage == 2)) {
+    if (cfg->aborted || cfg->bypassed || (cfg->stage == SMMU_STAGE_2)) {
         return 0;
     }
 
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         goto epilogue;
     }
 
-    if (cfg->stage == 1) {
+    if (cfg->stage == SMMU_STAGE_1) {
         /* Select stage1 translation table. */
         tt = select_tt(cfg, addr);
         if (!tt) {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
              * nesting is not supported. So it is sufficient to check the
              * translation stage to know the TLB stage for now.
              */
-            event.u.f_walk_eabt.s2 = (cfg->stage == 2);
+            event.u.f_walk_eabt.s2 = (cfg->stage == SMMU_STAGE_2);
             if (PTW_RECORD_FAULT(cfg)) {
                 event.type = SMMU_EVT_F_PERMISSION;
                 event.u.f_permission.addr = addr;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
 
     if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
         /* All faults from PTW has S2 field. */
-        event.u.f_walk_eabt.s2 = (ptw_info.stage == 2);
+        event.u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
         g_free(cached_entry);
         switch (ptw_info.type) {
         case SMMU_PTW_ERR_WALK_EABT:
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
             event.u.f_walk_eabt.addr = addr;
             event.u.f_walk_eabt.rnw = flag & 0x1;
             /* Stage-2 (only) is class IN while stage-1 is class TT */
-            event.u.f_walk_eabt.class = (ptw_info.stage == 2) ?
+            event.u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
                                          SMMU_CLASS_IN : SMMU_CLASS_TT;
             event.u.f_walk_eabt.addr2 = ptw_info.addr;
             break;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

smmuv3_translate() does everything from STE/CD parsing to TLB lookup
and PTW.

Soon, when nesting is supported, stage-1 data (tt, CD) needs to be
translated using stage-2.

Split smmuv3_translate() to 3 functions:

- smmu_translate(): in smmu-common.c, which does the TLB lookup, PTW,
  TLB insertion, all the functions are already there, this just puts
  them together.
  This also simplifies the code as it consolidates event generation
  in case of TLB lookup permission failure or in TT selection.

- smmuv3_do_translate(): in smmuv3.c, Calls smmu_translate() and does
  the event population in case of errors.

- smmuv3_translate(), now calls smmuv3_do_translate() for
  translation while the rest is the same.

Also, add stage in trace_smmuv3_translate_success()

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-6-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |   8 ++
 hw/arm/smmu-common.c         |  59 +++++++++++
 hw/arm/smmuv3.c              | 194 +++++++++++++----------------------
 hw/arm/trace-events          |   2 +-
 4 files changed, 142 insertions(+), 121 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
 int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
              SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
 
+
+/*
+ * smmu_translate - Look for a translation in TLB, if not, do a PTW.
+ * Returns NULL on PTW error or incase of TLB permission errors.
+ */
+SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
+                             IOMMUAccessFlags flag, SMMUPTWEventInfo *info);
+
 /**
  * select_tt - compute which translation table shall be used according to
  * the input iova and translation config and return the TT specific info
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
     g_assert_not_reached();
 }
 
+SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
+                             IOMMUAccessFlags flag, SMMUPTWEventInfo *info)
+{
+    uint64_t page_mask, aligned_addr;
+    SMMUTLBEntry *cached_entry = NULL;
+    SMMUTransTableInfo *tt;
+    int status;
+
+    /*
+     * Combined attributes used for TLB lookup, as only one stage is supported,
+     * it will hold attributes based on the enabled stage.
+     */
+    SMMUTransTableInfo tt_combined;
+
+    if (cfg->stage == SMMU_STAGE_1) {
+        /* Select stage1 translation table. */
+        tt = select_tt(cfg, addr);
+        if (!tt) {
+            info->type = SMMU_PTW_ERR_TRANSLATION;
+            info->stage = SMMU_STAGE_1;
+            return NULL;
+        }
+        tt_combined.granule_sz = tt->granule_sz;
+        tt_combined.tsz = tt->tsz;
+
+    } else {
+        /* Stage2. */
+        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
+        tt_combined.tsz = cfg->s2cfg.tsz;
+    }
+
+    /*
+     * TLB lookup looks for granule and input size for a translation stage,
+     * as only one stage is supported right now, choose the right values
+     * from the configuration.
+     */
+    page_mask = (1ULL << tt_combined.granule_sz) - 1;
+    aligned_addr = addr & ~page_mask;
+
+    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
+    if (cached_entry) {
+        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
+            info->type = SMMU_PTW_ERR_PERMISSION;
+            info->stage = cfg->stage;
+            return NULL;
+        }
+        return cached_entry;
+    }
+
+    cached_entry = g_new0(SMMUTLBEntry, 1);
+    status = smmu_ptw(cfg, aligned_addr, flag, cached_entry, info);
+    if (status) {
+            g_free(cached_entry);
+            return NULL;
+    }
+    smmu_iotlb_insert(bs, cfg, cached_entry);
+    return cached_entry;
+}
+
 /**
  * The bus number is used for lookup when SID based invalidation occurs.
  * In that case we lazily populate the SMMUPciBus array from the bus hash
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_flush_config(SMMUDevice *sdev)
     g_hash_table_remove(bc->configs, sdev);
 }
 
+/* Do translation with TLB lookup. */
+static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
+                                                 SMMUTransCfg *cfg,
+                                                 SMMUEventInfo *event,
+                                                 IOMMUAccessFlags flag,
+                                                 SMMUTLBEntry **out_entry)
+{
+    SMMUPTWEventInfo ptw_info = {};
+    SMMUState *bs = ARM_SMMU(s);
+    SMMUTLBEntry *cached_entry = NULL;
+
+    cached_entry = smmu_translate(bs, cfg, addr, flag, &ptw_info);
+    if (!cached_entry) {
+        /* All faults from PTW has S2 field. */
+        event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
+        switch (ptw_info.type) {
+        case SMMU_PTW_ERR_WALK_EABT:
+            event->type = SMMU_EVT_F_WALK_EABT;
+            event->u.f_walk_eabt.addr = addr;
+            event->u.f_walk_eabt.rnw = flag & 0x1;
+            event->u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
+                                          SMMU_CLASS_IN : SMMU_CLASS_TT;
+            event->u.f_walk_eabt.addr2 = ptw_info.addr;
+            break;
+        case SMMU_PTW_ERR_TRANSLATION:
+            if (PTW_RECORD_FAULT(cfg)) {
+                event->type = SMMU_EVT_F_TRANSLATION;
+                event->u.f_translation.addr = addr;
+                event->u.f_translation.addr2 = ptw_info.addr;
+                event->u.f_translation.class = SMMU_CLASS_IN;
+                event->u.f_translation.rnw = flag & 0x1;
+            }
+            break;
+        case SMMU_PTW_ERR_ADDR_SIZE:
+            if (PTW_RECORD_FAULT(cfg)) {
+                event->type = SMMU_EVT_F_ADDR_SIZE;
+                event->u.f_addr_size.addr = addr;
+                event->u.f_addr_size.addr2 = ptw_info.addr;
+                event->u.f_addr_size.class = SMMU_CLASS_IN;
+                event->u.f_addr_size.rnw = flag & 0x1;
+            }
+            break;
+        case SMMU_PTW_ERR_ACCESS:
+            if (PTW_RECORD_FAULT(cfg)) {
+                event->type = SMMU_EVT_F_ACCESS;
+                event->u.f_access.addr = addr;
+                event->u.f_access.addr2 = ptw_info.addr;
+                event->u.f_access.class = SMMU_CLASS_IN;
+                event->u.f_access.rnw = flag & 0x1;
+            }
+            break;
+        case SMMU_PTW_ERR_PERMISSION:
+            if (PTW_RECORD_FAULT(cfg)) {
+                event->type = SMMU_EVT_F_PERMISSION;
+                event->u.f_permission.addr = addr;
+                event->u.f_permission.addr2 = ptw_info.addr;
+                event->u.f_permission.class = SMMU_CLASS_IN;
+                event->u.f_permission.rnw = flag & 0x1;
+            }
+            break;
+        default:
+            g_assert_not_reached();
+        }
+        return SMMU_TRANS_ERROR;
+    }
+    *out_entry = cached_entry;
+    return SMMU_TRANS_SUCCESS;
+}
+
+/* Entry point to SMMU, does everything. */
 static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                                       IOMMUAccessFlags flag, int iommu_idx)
 {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
     SMMUEventInfo event = {.type = SMMU_EVT_NONE,
                            .sid = sid,
                            .inval_ste_allowed = false};
-    SMMUPTWEventInfo ptw_info = {};
     SMMUTranslationStatus status;
-    SMMUState *bs = ARM_SMMU(s);
-    uint64_t page_mask, aligned_addr;
-    SMMUTLBEntry *cached_entry = NULL;
-    SMMUTransTableInfo *tt;
     SMMUTransCfg *cfg = NULL;
     IOMMUTLBEntry entry = {
         .target_as = &address_space_memory,
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         .addr_mask = ~(hwaddr)0,
         .perm = IOMMU_NONE,
     };
-    /*
-     * Combined attributes used for TLB lookup, as only one stage is supported,
-     * it will hold attributes based on the enabled stage.
-     */
-    SMMUTransTableInfo tt_combined;
+    SMMUTLBEntry *cached_entry = NULL;
 
     qemu_mutex_lock(&s->mutex);
 
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         goto epilogue;
     }
 
-    if (cfg->stage == SMMU_STAGE_1) {
-        /* Select stage1 translation table. */
-        tt = select_tt(cfg, addr);
-        if (!tt) {
-            if (cfg->record_faults) {
-                event.type = SMMU_EVT_F_TRANSLATION;
-                event.u.f_translation.addr = addr;
-                event.u.f_translation.rnw = flag & 0x1;
-            }
-            status = SMMU_TRANS_ERROR;
-            goto epilogue;
-        }
-        tt_combined.granule_sz = tt->granule_sz;
-        tt_combined.tsz = tt->tsz;
-
-    } else {
-        /* Stage2. */
-        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
-        tt_combined.tsz = cfg->s2cfg.tsz;
-    }
-    /*
-     * TLB lookup looks for granule and input size for a translation stage,
-     * as only one stage is supported right now, choose the right values
-     * from the configuration.
-     */
-    page_mask = (1ULL << tt_combined.granule_sz) - 1;
-    aligned_addr = addr & ~page_mask;
-
-    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
-    if (cached_entry) {
-        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
-            status = SMMU_TRANS_ERROR;
-            /*
-             * We know that the TLB only contains either stage-1 or stage-2 as
-             * nesting is not supported. So it is sufficient to check the
-             * translation stage to know the TLB stage for now.
-             */
-            event.u.f_walk_eabt.s2 = (cfg->stage == SMMU_STAGE_2);
-            if (PTW_RECORD_FAULT(cfg)) {
-                event.type = SMMU_EVT_F_PERMISSION;
-                event.u.f_permission.addr = addr;
-                event.u.f_permission.rnw = flag & 0x1;
-            }
-        } else {
-            status = SMMU_TRANS_SUCCESS;
-        }
-        goto epilogue;
-    }
-
-    cached_entry = g_new0(SMMUTLBEntry, 1);
-
-    if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
-        /* All faults from PTW has S2 field. */
-        event.u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
-        g_free(cached_entry);
-        switch (ptw_info.type) {
-        case SMMU_PTW_ERR_WALK_EABT:
-            event.type = SMMU_EVT_F_WALK_EABT;
-            event.u.f_walk_eabt.addr = addr;
-            event.u.f_walk_eabt.rnw = flag & 0x1;
-            /* Stage-2 (only) is class IN while stage-1 is class TT */
-            event.u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
-                                         SMMU_CLASS_IN : SMMU_CLASS_TT;
-            event.u.f_walk_eabt.addr2 = ptw_info.addr;
-            break;
-        case SMMU_PTW_ERR_TRANSLATION:
-            if (PTW_RECORD_FAULT(cfg)) {
-                event.type = SMMU_EVT_F_TRANSLATION;
-                event.u.f_translation.addr = addr;
-                event.u.f_translation.addr2 = ptw_info.addr;
-                event.u.f_translation.class = SMMU_CLASS_IN;
-                event.u.f_translation.rnw = flag & 0x1;
-            }
-            break;
-        case SMMU_PTW_ERR_ADDR_SIZE:
-            if (PTW_RECORD_FAULT(cfg)) {
-                event.type = SMMU_EVT_F_ADDR_SIZE;
-                event.u.f_addr_size.addr = addr;
-                event.u.f_addr_size.addr2 = ptw_info.addr;
-                event.u.f_translation.class = SMMU_CLASS_IN;
-                event.u.f_addr_size.rnw = flag & 0x1;
-            }
-            break;
-        case SMMU_PTW_ERR_ACCESS:
-            if (PTW_RECORD_FAULT(cfg)) {
-                event.type = SMMU_EVT_F_ACCESS;
-                event.u.f_access.addr = addr;
-                event.u.f_access.addr2 = ptw_info.addr;
-                event.u.f_translation.class = SMMU_CLASS_IN;
-                event.u.f_access.rnw = flag & 0x1;
-            }
-            break;
-        case SMMU_PTW_ERR_PERMISSION:
-            if (PTW_RECORD_FAULT(cfg)) {
-                event.type = SMMU_EVT_F_PERMISSION;
-                event.u.f_permission.addr = addr;
-                event.u.f_permission.addr2 = ptw_info.addr;
-                event.u.f_translation.class = SMMU_CLASS_IN;
-                event.u.f_permission.rnw = flag & 0x1;
-            }
-            break;
-        default:
-            g_assert_not_reached();
-        }
-        status = SMMU_TRANS_ERROR;
-    } else {
-        smmu_iotlb_insert(bs, cfg, cached_entry);
-        status = SMMU_TRANS_SUCCESS;
-    }
+    status = smmuv3_do_translate(s, addr, cfg, &event, flag, &cached_entry);
 
 epilogue:
     qemu_mutex_unlock(&s->mutex);
@@ -XXX,XX +XXX,XX @@ epilogue:
                                     (addr & cached_entry->entry.addr_mask);
         entry.addr_mask = cached_entry->entry.addr_mask;
         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
-                                       entry.translated_addr, entry.perm);
+                                       entry.translated_addr, entry.perm,
+                                       cfg->stage);
         break;
     case SMMU_TRANS_DISABLE:
         entry.perm = flag;
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_get_ste(uint64_t addr) "STE addr: 0x%"PRIx64
 smmuv3_translate_disable(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x bypass (smmu disabled) iova:0x%"PRIx64" is_write=%d"
 smmuv3_translate_bypass(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x STE bypass iova:0x%"PRIx64" is_write=%d"
 smmuv3_translate_abort(const char *n, uint16_t sid, uint64_t addr, bool is_write) "%s sid=0x%x abort on iova:0x%"PRIx64" is_write=%d"
-smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm) "%s sid=0x%x iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x"
+smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm, int stage) "%s sid=0x%x iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x stage=%d"
 smmuv3_get_cd(uint64_t addr) "CD addr: 0x%"PRIx64
 smmuv3_decode_cd(uint32_t oas) "oas=%d"
 smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz, bool had) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d had:%d"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

ASID and VMID used to be uint16_t in the translation config, however,
in other contexts they can be int as -1 in case of TLB invalidation,
to represent all (don’t care).
When stage-2 was added asid was set to -1 in stage-2 and vmid to -1
in stage-1 configs. However, that meant they were set as (65536),
this was not an issue as nesting was not supported and no
commands/lookup uses both.

With nesting, it’s critical to get this right as translation must be
tagged correctly with ASID/VMID, and with ASID=-1 meaning stage-2.
Represent ASID/VMID everywhere as int.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-7-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h | 14 +++++++-------
 hw/arm/smmu-common.c         | 10 +++++-----
 hw/arm/smmuv3.c              |  4 ++--
 hw/arm/trace-events          | 18 +++++++++---------
 4 files changed, 23 insertions(+), 23 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUS2Cfg {
     bool record_faults;     /* Record fault events (S2R) */
     uint8_t granule_sz;     /* Granule page shift (based on S2TG) */
     uint8_t eff_ps;         /* Effective PA output range (based on S2PS) */
-    uint16_t vmid;          /* Virtual Machine ID (S2VMID) */
+    int vmid;               /* Virtual Machine ID (S2VMID) */
     uint64_t vttb;          /* Address of translation table base (S2TTB) */
 } SMMUS2Cfg;
 
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransCfg {
     uint64_t ttb;              /* TT base address */
     uint8_t oas;               /* output address width */
     uint8_t tbi;               /* Top Byte Ignore */
-    uint16_t asid;
+    int asid;
     SMMUTransTableInfo tt[2];
     /* Used by stage-2 only. */
     struct SMMUS2Cfg s2cfg;
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPciBus {
 
 typedef struct SMMUIOTLBKey {
     uint64_t iova;
-    uint16_t asid;
-    uint16_t vmid;
+    int asid;
+    int vmid;
     uint8_t tg;
     uint8_t level;
 } SMMUIOTLBKey;
@@ -XXX,XX +XXX,XX @@ SMMUDevice *smmu_find_sdev(SMMUState *s, uint32_t sid);
 SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
                                 SMMUTransTableInfo *tt, hwaddr iova);
 void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
-SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint16_t vmid, uint64_t iova,
+SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
                                 uint8_t tg, uint8_t level);
 void smmu_iotlb_inv_all(SMMUState *s);
-void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
-void smmu_iotlb_inv_vmid(SMMUState *s, uint16_t vmid);
+void smmu_iotlb_inv_asid(SMMUState *s, int asid);
+void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
 void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
 
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
            (k1->vmid == k2->vmid);
 }
 
-SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint16_t vmid, uint64_t iova,
+SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
                                 uint8_t tg, uint8_t level)
 {
     SMMUIOTLBKey key = {.asid = asid, .vmid = vmid, .iova = iova,
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_all(SMMUState *s)
 static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
                                          gpointer user_data)
 {
-    uint16_t asid = *(uint16_t *)user_data;
+    int asid = *(int *)user_data;
     SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
 
     return SMMU_IOTLB_ASID(*iotlb_key) == asid;
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
 static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
                                          gpointer user_data)
 {
-    uint16_t vmid = *(uint16_t *)user_data;
+    int vmid = *(int *)user_data;
     SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
 
     return SMMU_IOTLB_VMID(*iotlb_key) == vmid;
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                                 &info);
 }
 
-void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
+void smmu_iotlb_inv_asid(SMMUState *s, int asid)
 {
     trace_smmu_iotlb_inv_asid(asid);
     g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid, &asid);
 }
 
-void smmu_iotlb_inv_vmid(SMMUState *s, uint16_t vmid)
+void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
 {
     trace_smmu_iotlb_inv_vmid(vmid);
     g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid, &vmid);
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
         }
         case SMMU_CMD_TLBI_NH_ASID:
         {
-            uint16_t asid = CMD_ASID(&cmd);
+            int asid = CMD_ASID(&cmd);
 
             if (!STAGE1_SUPPORTED(s)) {
                 cmd_error = SMMU_CERROR_ILL;
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
             break;
         case SMMU_CMD_TLBI_S12_VMALL:
         {
-            uint16_t vmid = CMD_VMID(&cmd);
+            int vmid = CMD_VMID(&cmd);
 
             if (!STAGE2_SUPPORTED(s)) {
                 cmd_error = SMMU_CERROR_ILL;
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_ptw_page_pte(int stage, int level,  uint64_t iova, uint64_t baseaddr, uint6
 smmu_ptw_block_pte(int stage, int level, uint64_t baseaddr, uint64_t pteaddr, uint64_t pte, uint64_t iova, uint64_t gpa, int bsize_mb) "stage=%d level=%d base@=0x%"PRIx64" pte@=0x%"PRIx64" pte=0x%"PRIx64" iova=0x%"PRIx64" block address = 0x%"PRIx64" block size = %d MiB"
 smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "baseaddr=0x%"PRIx64" index=0x%x, pteaddr=0x%"PRIx64", pte=0x%"PRIx64
 smmu_iotlb_inv_all(void) "IOTLB invalidate all"
-smmu_iotlb_inv_asid(uint16_t asid) "IOTLB invalidate asid=%d"
-smmu_iotlb_inv_vmid(uint16_t vmid) "IOTLB invalidate vmid=%d"
-smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
+smmu_iotlb_inv_asid(int asid) "IOTLB invalidate asid=%d"
+smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
+smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
 smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
-smmu_iotlb_lookup_hit(uint16_t asid, uint16_t vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
-smmu_iotlb_lookup_miss(uint16_t asid, uint16_t vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
-smmu_iotlb_insert(uint16_t asid, uint16_t vmid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d vmid=%d addr=0x%"PRIx64" tg=%d level=%d"
+smmu_iotlb_lookup_hit(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
+smmu_iotlb_lookup_miss(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
+smmu_iotlb_insert(int asid, int vmid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d vmid=%d addr=0x%"PRIx64" tg=%d level=%d"
 
 # smmuv3.c
 smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
@@ -XXX,XX +XXX,XX @@ smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t p
 smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
 smmuv3_cmdq_tlbi_nh(void) ""
-smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
-smmuv3_cmdq_tlbi_s12_vmid(uint16_t vmid) "vmid=%d"
+smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
+smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
 smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
 smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
-smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint16_t vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
+smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
 
 # strongarm.c
 strongarm_uart_update_parameters(const char *label, int speed, char parity, int data_bits, int stop_bits) "%s speed=%d parity=%c data=%d stop=%d"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Soon, smmuv3_do_translate() will be used to translate the CD and the
TTBx, instead of re-writting the same logic to convert the returned
cached entry to an address, add a new macro CACHED_ENTRY_TO_ADDR.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-8-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h | 3 +++
 hw/arm/smmuv3.c              | 3 +--
 2 files changed, 4 insertions(+), 2 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@
 #define VMSA_IDXMSK(isz, strd, lvl)         ((1ULL << \
                                              VMSA_BIT_LVL(isz, strd, lvl)) - 1)
 
+#define CACHED_ENTRY_TO_ADDR(ent, addr)      ((ent)->entry.translated_addr + \
+                                             ((addr) & (ent)->entry.addr_mask))
+
 /*
  * Page table walk error types
  */
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ epilogue:
     switch (status) {
     case SMMU_TRANS_SUCCESS:
         entry.perm = cached_entry->entry.perm;
-        entry.translated_addr = cached_entry->entry.translated_addr +
-                                    (addr & cached_entry->entry.addr_mask);
+        entry.translated_addr = CACHED_ENTRY_TO_ADDR(cached_entry, addr);
         entry.addr_mask = cached_entry->entry.addr_mask;
         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
                                        entry.translated_addr, entry.perm,
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

According to ARM SMMU architecture specification (ARM IHI 0070 F.b),
In "5.2 Stream Table Entry":
 [51:6] S1ContextPtr
 If Config[1] == 1 (stage 2 enabled), this pointer is an IPA translated by
 stage 2 and the programmed value must be within the range of the IAS.

In "5.4.1 CD notes":
 The translation table walks performed from TTB0 or TTB1 are always performed
 in IPA space if stage 2 translations are enabled.

This patch implements translation of the S1 context descriptor pointer and
TTBx base addresses through the S2 stage (IPA -> PA)

smmuv3_do_translate() is updated to have one arg which is translation
class, this is useful to:
 - Decide wether a translation is stage-2 only or use the STE config.
 - Populate the class in case of faults, WALK_EABT is left unchanged
   for stage-1 as it is always IN, while stage-2 would match the
   used class (TT, IN, CD), this will change slightly when the ptw
   supports nested translation as it can also issue TT event with
   class IN.

In case for stage-2 only translation, used in the context of nested
translation, the stage and asid are saved and restored before and
after calling smmu_translate().

Translating CD or TTBx can fail for the following reasons:
1) Large address size: This is described in
   (3.4.3 Address sizes of SMMU-originated accesses)
   - For CD ptr larger than IAS, for SMMUv3.1, it can trigger either
     C_BAD_STE or Translation fault, we implement the latter as it
     requires no extra code.
   - For TTBx, if larger than the effective stage 1 output address size, it
     triggers C_BAD_CD.

2) Faults from PTWs (7.3 Event records)
   - F_ADDR_SIZE: large address size after first level causes stage 2 Address
     Size fault (Also in 3.4.3 Address sizes of SMMU-originated accesses)
   - F_PERMISSION: Same as an address translation. However, when
     CLASS == CD, the access is implicitly Data and a read.
   - F_ACCESS: Same as an address translation.
   - F_TRANSLATION: Same as an address translation.
   - F_WALK_EABT: Same as an address translation.
  These are already implemented in the PTW logic, so no extra handling
  required.

As in CD and TTBx translation context, the iova is not known, setting
the InputAddr was removed from "smmuv3_do_translate" and set after
from "smmuv3_translate" with the new function "smmuv3_fixup_event"

Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-9-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3.c | 120 +++++++++++++++++++++++++++++++++++++++++-------
 1 file changed, 103 insertions(+), 17 deletions(-)

diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static int smmu_get_ste(SMMUv3State *s, dma_addr_t addr, STE *buf,
 
 }
 
+static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
+                                                 SMMUTransCfg *cfg,
+                                                 SMMUEventInfo *event,
+                                                 IOMMUAccessFlags flag,
+                                                 SMMUTLBEntry **out_entry,
+                                                 SMMUTranslationClass class);
 /* @ssid > 0 not supported yet */
-static int smmu_get_cd(SMMUv3State *s, STE *ste, uint32_t ssid,
-                       CD *buf, SMMUEventInfo *event)
+static int smmu_get_cd(SMMUv3State *s, STE *ste, SMMUTransCfg *cfg,
+                       uint32_t ssid, CD *buf, SMMUEventInfo *event)
 {
     dma_addr_t addr = STE_CTXPTR(ste);
     int ret, i;
+    SMMUTranslationStatus status;
+    SMMUTLBEntry *entry;
 
     trace_smmuv3_get_cd(addr);
+
+    if (cfg->stage == SMMU_NESTED) {
+        status = smmuv3_do_translate(s, addr, cfg, event,
+                                     IOMMU_RO, &entry, SMMU_CLASS_CD);
+
+        /* Same PTW faults are reported but with CLASS = CD. */
+        if (status != SMMU_TRANS_SUCCESS) {
+            return -EINVAL;
+        }
+
+        addr = CACHED_ENTRY_TO_ADDR(entry, addr);
+    }
+
     /* TODO: guarantee 64-bit single-copy atomicity */
     ret = dma_memory_read(&address_space_memory, addr, buf, sizeof(*buf),
                           MEMTXATTRS_UNSPECIFIED);
@@ -XXX,XX +XXX,XX @@ static int smmu_find_ste(SMMUv3State *s, uint32_t sid, STE *ste,
     return 0;
 }
 
-static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
+static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
+                     CD *cd, SMMUEventInfo *event)
 {
     int ret = -EINVAL;
     int i;
+    SMMUTranslationStatus status;
+    SMMUTLBEntry *entry;
 
     if (!CD_VALID(cd) || !CD_AARCH64(cd)) {
         goto bad_cd;
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
 
         tt->tsz = tsz;
         tt->ttb = CD_TTB(cd, i);
+
         if (tt->ttb & ~(MAKE_64BIT_MASK(0, cfg->oas))) {
             goto bad_cd;
         }
+
+        /* Translate the TTBx, from IPA to PA if nesting is enabled. */
+        if (cfg->stage == SMMU_NESTED) {
+            status = smmuv3_do_translate(s, tt->ttb, cfg, event, IOMMU_RO,
+                                         &entry, SMMU_CLASS_TT);
+            /*
+             * Same PTW faults are reported but with CLASS = TT.
+             * If TTBx is larger than the effective stage 1 output addres
+             * size, it reports C_BAD_CD, which is handled by the above case.
+             */
+            if (status != SMMU_TRANS_SUCCESS) {
+                return -EINVAL;
+            }
+            tt->ttb = CACHED_ENTRY_TO_ADDR(entry, tt->ttb);
+        }
+
         tt->had = CD_HAD(cd, i);
         trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz, tt->had);
     }
@@ -XXX,XX +XXX,XX @@ static int smmuv3_decode_config(IOMMUMemoryRegion *mr, SMMUTransCfg *cfg,
         return 0;
     }
 
-    ret = smmu_get_cd(s, &ste, 0 /* ssid */, &cd, event);
+    ret = smmu_get_cd(s, &ste, cfg, 0 /* ssid */, &cd, event);
     if (ret) {
         return ret;
     }
 
-    return decode_cd(cfg, &cd, event);
+    return decode_cd(s, cfg, &cd, event);
 }
 
 /**
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
                                                  SMMUTransCfg *cfg,
                                                  SMMUEventInfo *event,
                                                  IOMMUAccessFlags flag,
-                                                 SMMUTLBEntry **out_entry)
+                                                 SMMUTLBEntry **out_entry,
+                                                 SMMUTranslationClass class)
 {
     SMMUPTWEventInfo ptw_info = {};
     SMMUState *bs = ARM_SMMU(s);
     SMMUTLBEntry *cached_entry = NULL;
+    int asid, stage;
+    bool desc_s2_translation = class != SMMU_CLASS_IN;
+
+    /*
+     * The function uses the argument class to identify which stage is used:
+     * - CLASS = IN: Means an input translation, determine the stage from STE.
+     * - CLASS = CD: Means the addr is an IPA of the CD, and it would be
+     *   translated using the stage-2.
+     * - CLASS = TT: Means the addr is an IPA of the stage-1 translation table
+     *   and it would be translated using the stage-2.
+     * For the last 2 cases instead of having intrusive changes in the common
+     * logic, we modify the cfg to be a stage-2 translation only in case of
+     * nested, and then restore it after.
+     */
+    if (desc_s2_translation) {
+        asid = cfg->asid;
+        stage = cfg->stage;
+        cfg->asid = -1;
+        cfg->stage = SMMU_STAGE_2;
+    }
 
     cached_entry = smmu_translate(bs, cfg, addr, flag, &ptw_info);
+
+    if (desc_s2_translation) {
+        cfg->asid = asid;
+        cfg->stage = stage;
+    }
+
     if (!cached_entry) {
         /* All faults from PTW has S2 field. */
         event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
         switch (ptw_info.type) {
         case SMMU_PTW_ERR_WALK_EABT:
             event->type = SMMU_EVT_F_WALK_EABT;
-            event->u.f_walk_eabt.addr = addr;
             event->u.f_walk_eabt.rnw = flag & 0x1;
             event->u.f_walk_eabt.class = (ptw_info.stage == SMMU_STAGE_2) ?
-                                          SMMU_CLASS_IN : SMMU_CLASS_TT;
+                                          class : SMMU_CLASS_TT;
             event->u.f_walk_eabt.addr2 = ptw_info.addr;
             break;
         case SMMU_PTW_ERR_TRANSLATION:
             if (PTW_RECORD_FAULT(cfg)) {
                 event->type = SMMU_EVT_F_TRANSLATION;
-                event->u.f_translation.addr = addr;
                 event->u.f_translation.addr2 = ptw_info.addr;
-                event->u.f_translation.class = SMMU_CLASS_IN;
+                event->u.f_translation.class = class;
                 event->u.f_translation.rnw = flag & 0x1;
             }
             break;
         case SMMU_PTW_ERR_ADDR_SIZE:
             if (PTW_RECORD_FAULT(cfg)) {
                 event->type = SMMU_EVT_F_ADDR_SIZE;
-                event->u.f_addr_size.addr = addr;
                 event->u.f_addr_size.addr2 = ptw_info.addr;
-                event->u.f_addr_size.class = SMMU_CLASS_IN;
+                event->u.f_addr_size.class = class;
                 event->u.f_addr_size.rnw = flag & 0x1;
             }
             break;
         case SMMU_PTW_ERR_ACCESS:
             if (PTW_RECORD_FAULT(cfg)) {
                 event->type = SMMU_EVT_F_ACCESS;
-                event->u.f_access.addr = addr;
                 event->u.f_access.addr2 = ptw_info.addr;
-                event->u.f_access.class = SMMU_CLASS_IN;
+                event->u.f_access.class = class;
                 event->u.f_access.rnw = flag & 0x1;
             }
             break;
         case SMMU_PTW_ERR_PERMISSION:
             if (PTW_RECORD_FAULT(cfg)) {
                 event->type = SMMU_EVT_F_PERMISSION;
-                event->u.f_permission.addr = addr;
                 event->u.f_permission.addr2 = ptw_info.addr;
-                event->u.f_permission.class = SMMU_CLASS_IN;
+                event->u.f_permission.class = class;
                 event->u.f_permission.rnw = flag & 0x1;
             }
             break;
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
     return SMMU_TRANS_SUCCESS;
 }
 
+/*
+ * Sets the InputAddr for an SMMU_TRANS_ERROR, as it can't be
+ * set from all contexts, as smmuv3_get_config() can return
+ * translation faults in case of nested translation (for CD
+ * and TTBx). But in that case the iova is not known.
+ */
+static void smmuv3_fixup_event(SMMUEventInfo *event, hwaddr iova)
+{
+    switch (event->type) {
+    case SMMU_EVT_F_WALK_EABT:
+    case SMMU_EVT_F_TRANSLATION:
+    case SMMU_EVT_F_ADDR_SIZE:
+    case SMMU_EVT_F_ACCESS:
+    case SMMU_EVT_F_PERMISSION:
+        event->u.f_walk_eabt.addr = iova;
+        break;
+    default:
+        break;
+    }
+}
+
 /* Entry point to SMMU, does everything. */
 static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
                                       IOMMUAccessFlags flag, int iommu_idx)
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         goto epilogue;
     }
 
-    status = smmuv3_do_translate(s, addr, cfg, &event, flag, &cached_entry);
+    status = smmuv3_do_translate(s, addr, cfg, &event, flag,
+                                 &cached_entry, SMMU_CLASS_IN);
 
 epilogue:
     qemu_mutex_unlock(&s->mutex);
@@ -XXX,XX +XXX,XX @@ epilogue:
                                      entry.perm);
         break;
     case SMMU_TRANS_ERROR:
+        smmuv3_fixup_event(&event, addr);
         qemu_log_mask(LOG_GUEST_ERROR,
                       "%s translation failed for iova=0x%"PRIx64" (%s)\n",
                       mr->parent_obj.name, addr, smmu_event_string(event.type));
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

In the next patch, combine_tlb() will be added which combines 2 TLB
entries into one for nested translations, which chooses the granule
and level from the smallest entry.

This means that with nested translation, an entry can be cached with
the granule of stage-2 and not stage-1.

However, currently, the lookup for an IOVA is done with input stage
granule, which is stage-1 for nested configuration, which will not
work with the above logic.
This patch reworks lookup in that case, so it falls back to stage-2
granule if no entry is found using stage-1 granule.

Also, drop aligning the iova to avoid over-aligning in case the iova
is cached with a smaller granule, the TLB lookup will align the iova
anyway for each granule and level, and the page table walker doesn't
consider the page offset bits.

Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-10-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-common.c | 64 +++++++++++++++++++++++++++++---------------
 1 file changed, 43 insertions(+), 21 deletions(-)

diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
     return key;
 }
 
-SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
-                                SMMUTransTableInfo *tt, hwaddr iova)
+static SMMUTLBEntry *smmu_iotlb_lookup_all_levels(SMMUState *bs,
+                                                  SMMUTransCfg *cfg,
+                                                  SMMUTransTableInfo *tt,
+                                                  hwaddr iova)
 {
     uint8_t tg = (tt->granule_sz - 10) / 2;
     uint8_t inputsize = 64 - tt->tsz;
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
         }
         level++;
     }
+    return entry;
+}
+
+/**
+ * smmu_iotlb_lookup - Look up for a TLB entry.
+ * @bs: SMMU state which includes the TLB instance
+ * @cfg: Configuration of the translation
+ * @tt: Translation table info (granule and tsz)
+ * @iova: IOVA address to lookup
+ *
+ * returns a valid entry on success, otherwise NULL.
+ * In case of nested translation, tt can be updated to include
+ * the granule of the found entry as it might different from
+ * the IOVA granule.
+ */
+SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+                                SMMUTransTableInfo *tt, hwaddr iova)
+{
+    SMMUTLBEntry *entry = NULL;
+
+    entry = smmu_iotlb_lookup_all_levels(bs, cfg, tt, iova);
+    /*
+     * For nested translation also try the s2 granule, as the TLB will insert
+     * it if the size of s2 tlb entry was smaller.
+     */
+    if (!entry && (cfg->stage == SMMU_NESTED) &&
+        (cfg->s2cfg.granule_sz != tt->granule_sz)) {
+        tt->granule_sz = cfg->s2cfg.granule_sz;
+        entry = smmu_iotlb_lookup_all_levels(bs, cfg, tt, iova);
+    }
 
     if (entry) {
         cfg->iotlb_hits++;
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
 SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
                              IOMMUAccessFlags flag, SMMUPTWEventInfo *info)
 {
-    uint64_t page_mask, aligned_addr;
     SMMUTLBEntry *cached_entry = NULL;
     SMMUTransTableInfo *tt;
     int status;
 
     /*
-     * Combined attributes used for TLB lookup, as only one stage is supported,
-     * it will hold attributes based on the enabled stage.
+     * Combined attributes used for TLB lookup, holds the attributes for
+     * the input stage.
      */
     SMMUTransTableInfo tt_combined;
 
-    if (cfg->stage == SMMU_STAGE_1) {
+    if (cfg->stage == SMMU_STAGE_2) {
+        /* Stage2. */
+        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
+        tt_combined.tsz = cfg->s2cfg.tsz;
+    } else {
         /* Select stage1 translation table. */
         tt = select_tt(cfg, addr);
         if (!tt) {
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
         }
         tt_combined.granule_sz = tt->granule_sz;
         tt_combined.tsz = tt->tsz;
-
-    } else {
-        /* Stage2. */
-        tt_combined.granule_sz = cfg->s2cfg.granule_sz;
-        tt_combined.tsz = cfg->s2cfg.tsz;
     }
 
-    /*
-     * TLB lookup looks for granule and input size for a translation stage,
-     * as only one stage is supported right now, choose the right values
-     * from the configuration.
-     */
-    page_mask = (1ULL << tt_combined.granule_sz) - 1;
-    aligned_addr = addr & ~page_mask;
-
-    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, aligned_addr);
+    cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, addr);
     if (cached_entry) {
         if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
             info->type = SMMU_PTW_ERR_PERMISSION;
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
     }
 
     cached_entry = g_new0(SMMUTLBEntry, 1);
-    status = smmu_ptw(cfg, aligned_addr, flag, cached_entry, info);
+    status = smmu_ptw(cfg, addr, flag, cached_entry, info);
     if (status) {
             g_free(cached_entry);
             return NULL;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

This patch adds support for nested (combined) TLB entries.
The main function combine_tlb() is not used here but in the next
patches, but to simplify the patches it is introduced first.

Main changes:
1) New field added in the SMMUTLBEntry struct: parent_perm, for
   nested TLB, holds the stage-2 permission, this can be used to know
   the origin of a permission fault from a cached entry as caching
   the “and” of the permissions loses this information.

SMMUPTWEventInfo is used to hold information about PTW faults so
   the event can be populated, the value of stage used to be set
   based on the current stage for TLB permission faults, however
   with the parent_perm, it is now set based on which perm has
   the missing permission

When nesting is not enabled it has the same value as perm which
   doesn't change the logic.

2) As combined TLB implementation is used, the combination logic
   chooses:
   - tg and level from the entry which has the smallest addr_mask.
   - Based on that the iova that would be cached is recalculated.
   - Translated_addr is chosen from stage-2.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-11-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  1 +
 hw/arm/smmu-common.c         | 37 ++++++++++++++++++++++++++++++++----
 2 files changed, 34 insertions(+), 4 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTLBEntry {
     IOMMUTLBEntry entry;
     uint8_t level;
     uint8_t granule;
+    IOMMUAccessFlags parent_perm;
 } SMMUTLBEntry;
 
 /* Stage-2 configuration. */
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
         tlbe->entry.translated_addr = gpa;
         tlbe->entry.iova = iova & ~mask;
         tlbe->entry.addr_mask = mask;
-        tlbe->entry.perm = PTE_AP_TO_PERM(ap);
+        tlbe->parent_perm = PTE_AP_TO_PERM(ap);
+        tlbe->entry.perm = tlbe->parent_perm;
         tlbe->level = level;
         tlbe->granule = granule_sz;
         return 0;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
         tlbe->entry.translated_addr = gpa;
         tlbe->entry.iova = ipa & ~mask;
         tlbe->entry.addr_mask = mask;
-        tlbe->entry.perm = s2ap;
+        tlbe->parent_perm = s2ap;
+        tlbe->entry.perm = tlbe->parent_perm;
         tlbe->level = level;
         tlbe->granule = granule_sz;
         return 0;
@@ -XXX,XX +XXX,XX @@ error:
     return -EINVAL;
 }
 
+/*
+ * combine S1 and S2 TLB entries into a single entry.
+ * As a result the S1 entry is overriden with combined data.
+ */
+static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
+                                                SMMUTLBEntry *tlbe_s2,
+                                                dma_addr_t iova,
+                                                SMMUTransCfg *cfg)
+{
+    if (tlbe_s2->entry.addr_mask < tlbe->entry.addr_mask) {
+        tlbe->entry.addr_mask = tlbe_s2->entry.addr_mask;
+        tlbe->granule = tlbe_s2->granule;
+        tlbe->level = tlbe_s2->level;
+    }
+
+    tlbe->entry.translated_addr = CACHED_ENTRY_TO_ADDR(tlbe_s2,
+                                    tlbe->entry.translated_addr);
+
+    tlbe->entry.iova = iova & ~tlbe->entry.addr_mask;
+    /* parent_perm has s2 perm while perm keeps s1 perm. */
+    tlbe->parent_perm = tlbe_s2->entry.perm;
+    return;
+}
+
 /**
  * smmu_ptw - Walk the page tables for an IOVA, according to @cfg
  *
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
 
     cached_entry = smmu_iotlb_lookup(bs, cfg, &tt_combined, addr);
     if (cached_entry) {
-        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
+        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm &
+            cached_entry->parent_perm & IOMMU_WO)) {
             info->type = SMMU_PTW_ERR_PERMISSION;
-            info->stage = cfg->stage;
+            info->stage = !(cached_entry->entry.perm & IOMMU_WO) ?
+                          SMMU_STAGE_1 :
+                          SMMU_STAGE_2;
             return NULL;
         }
         return cached_entry;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

When nested translation is requested, do the following:
- Translate stage-1 table address IPA into PA through stage-2.
- Translate stage-1 table walk output (IPA) through stage-2.
- Create a single TLB entry from stage-1 and stage-2 translations
  using logic introduced before.

smmu_ptw() has a new argument SMMUState which include the TLB as
stage-1 table address can be cached in there.

Also in smmu_ptw(), a separate path used for nesting to simplify the
code, although some logic can be combined.

With nested translation class of translation fault can be different,
from the class of the translation, as faults from translating stage-1
tables are considered as CLASS_TT and not CLASS_IN, a new member
"is_ipa_descriptor" added to "SMMUPTWEventInfo" to differ faults
from walking stage 1 translation table and faults from translating
an IPA for a transaction.

Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-12-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  7 ++--
 hw/arm/smmu-common.c         | 74 +++++++++++++++++++++++++++++++-----
 hw/arm/smmuv3.c              | 14 +++++++
 3 files changed, 82 insertions(+), 13 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPTWEventInfo {
     SMMUStage stage;
     SMMUPTWEventType type;
     dma_addr_t addr; /* fetched address that induced an abort, if any */
+    bool is_ipa_descriptor; /* src for fault in nested translation. */
 } SMMUPTWEventInfo;
 
 typedef struct SMMUTransTableInfo {
@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
  * smmu_ptw - Perform the page table walk for a given iova / access flags
  * pair, according to @cfg translation config
  */
-int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
-             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
-
+int smmu_ptw(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t iova,
+             IOMMUAccessFlags perm, SMMUTLBEntry *tlbe,
+             SMMUPTWEventInfo *info);
 
 /*
  * smmu_translate - Look for a translation in TLB, if not, do a PTW.
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
     return NULL;
 }
 
+/* Translate stage-1 table address using stage-2 page table. */
+static inline int translate_table_addr_ipa(SMMUState *bs,
+                                           dma_addr_t *table_addr,
+                                           SMMUTransCfg *cfg,
+                                           SMMUPTWEventInfo *info)
+{
+    dma_addr_t addr = *table_addr;
+    SMMUTLBEntry *cached_entry;
+    int asid;
+
+    /*
+     * The translation table walks performed from TTB0 or TTB1 are always
+     * performed in IPA space if stage 2 translations are enabled.
+     */
+    asid = cfg->asid;
+    cfg->stage = SMMU_STAGE_2;
+    cfg->asid = -1;
+    cached_entry = smmu_translate(bs, cfg, addr, IOMMU_RO, info);
+    cfg->asid = asid;
+    cfg->stage = SMMU_NESTED;
+
+    if (cached_entry) {
+        *table_addr = CACHED_ENTRY_TO_ADDR(cached_entry, addr);
+        return 0;
+    }
+
+    info->stage = SMMU_STAGE_2;
+    info->addr = addr;
+    info->is_ipa_descriptor = true;
+    return -EINVAL;
+}
+
 /**
  * smmu_ptw_64_s1 - VMSAv8-64 Walk of the page tables for a given IOVA
+ * @bs: smmu state which includes TLB instance
  * @cfg: translation config
  * @iova: iova to translate
  * @perm: access type
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
  * Upon success, @tlbe is filled with translated_addr and entry
  * permission rights.
  */
-static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
+static int smmu_ptw_64_s1(SMMUState *bs, SMMUTransCfg *cfg,
                           dma_addr_t iova, IOMMUAccessFlags perm,
                           SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUTransCfg *cfg,
                 goto error;
             }
             baseaddr = get_table_pte_address(pte, granule_sz);
+            if (cfg->stage == SMMU_NESTED) {
+                if (translate_table_addr_ipa(bs, &baseaddr, cfg, info)) {
+                    goto error;
+                }
+            }
             level++;
             continue;
         } else if (is_page_pte(pte, level)) {
@@ -XXX,XX +XXX,XX @@ error:
  * combine S1 and S2 TLB entries into a single entry.
  * As a result the S1 entry is overriden with combined data.
  */
-static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
-                                                SMMUTLBEntry *tlbe_s2,
-                                                dma_addr_t iova,
-                                                SMMUTransCfg *cfg)
+static void combine_tlb(SMMUTLBEntry *tlbe, SMMUTLBEntry *tlbe_s2,
+                        dma_addr_t iova, SMMUTransCfg *cfg)
 {
     if (tlbe_s2->entry.addr_mask < tlbe->entry.addr_mask) {
         tlbe->entry.addr_mask = tlbe_s2->entry.addr_mask;
@@ -XXX,XX +XXX,XX @@ static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
 /**
  * smmu_ptw - Walk the page tables for an IOVA, according to @cfg
  *
+ * @bs: smmu state which includes TLB instance
  * @cfg: translation configuration
  * @iova: iova to translate
  * @perm: tentative access type
@@ -XXX,XX +XXX,XX @@ static void __attribute__((unused)) combine_tlb(SMMUTLBEntry *tlbe,
  *
  * return 0 on success
  */
-int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
-             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
+int smmu_ptw(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t iova,
+             IOMMUAccessFlags perm, SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
+    int ret;
+    SMMUTLBEntry tlbe_s2;
+    dma_addr_t ipa;
+
     if (cfg->stage == SMMU_STAGE_1) {
-        return smmu_ptw_64_s1(cfg, iova, perm, tlbe, info);
+        return smmu_ptw_64_s1(bs, cfg, iova, perm, tlbe, info);
     } else if (cfg->stage == SMMU_STAGE_2) {
         /*
          * If bypassing stage 1(or unimplemented), the input address is passed
@@ -XXX,XX +XXX,XX @@ int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
         return smmu_ptw_64_s2(cfg, iova, perm, tlbe, info);
     }
 
-    g_assert_not_reached();
+    /* SMMU_NESTED. */
+    ret = smmu_ptw_64_s1(bs, cfg, iova, perm, tlbe, info);
+    if (ret) {
+        return ret;
+    }
+
+    ipa = CACHED_ENTRY_TO_ADDR(tlbe, iova);
+    ret = smmu_ptw_64_s2(cfg, ipa, perm, &tlbe_s2, info);
+    if (ret) {
+        return ret;
+    }
+
+    combine_tlb(tlbe, &tlbe_s2, iova, cfg);
+    return 0;
 }
 
 SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_translate(SMMUState *bs, SMMUTransCfg *cfg, dma_addr_t addr,
     }
 
     cached_entry = g_new0(SMMUTLBEntry, 1);
-    status = smmu_ptw(cfg, addr, flag, cached_entry, info);
+    status = smmu_ptw(bs, cfg, addr, flag, cached_entry, info);
     if (status) {
             g_free(cached_entry);
             return NULL;
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
     if (!cached_entry) {
         /* All faults from PTW has S2 field. */
         event->u.f_walk_eabt.s2 = (ptw_info.stage == SMMU_STAGE_2);
+        /*
+         * Fault class is set as follows based on "class" input to
+         * the function and to "ptw_info" from "smmu_translate()"
+         * For stage-1:
+         *   - EABT => CLASS_TT (hardcoded)
+         *   - other events => CLASS_IN (input to function)
+         * For stage-2 => CLASS_IN (input to function)
+         * For nested, for all events:
+         *  - CD fetch => CLASS_CD (input to function)
+         *  - walking stage 1 translation table  => CLASS_TT (from
+         *    is_ipa_descriptor or input in case of TTBx)
+         *  - s2 translation => CLASS_IN (input to function)
+         */
+        class = ptw_info.is_ipa_descriptor ? SMMU_CLASS_TT : class;
         switch (ptw_info.type) {
         case SMMU_PTW_ERR_WALK_EABT:
             event->type = SMMU_EVT_F_WALK_EABT;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

With nesting, we would need to invalidate IPAs without
over-invalidating stage-1 IOVAs. This can be done by
distinguishing IPAs in the TLBs by having ASID=-1.
To achieve that, rework the invalidation for IPAs to have a
separate function, while for IOVA invalidation ASID=-1 means
invalidate for all ASIDs.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-13-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  3 ++-
 hw/arm/smmu-common.c         | 47 ++++++++++++++++++++++++++++++++++++
 hw/arm/smmuv3.c              | 23 ++++++++++++------
 hw/arm/trace-events          |  2 +-
 4 files changed, 66 insertions(+), 9 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_asid(SMMUState *s, int asid);
 void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
 void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
-
+void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
+                        uint64_t num_pages, uint8_t ttl);
 /* Unmap the range of all the notifiers registered to any IOMMU mr */
 void smmu_inv_notifiers_all(SMMUState *s);
 
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid_vmid_iova(gpointer key, gpointer value,
            ((entry->iova & ~info->mask) == info->iova);
 }
 
+static gboolean smmu_hash_remove_by_vmid_ipa(gpointer key, gpointer value,
+                                             gpointer user_data)
+{
+    SMMUTLBEntry *iter = (SMMUTLBEntry *)value;
+    IOMMUTLBEntry *entry = &iter->entry;
+    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
+    SMMUIOTLBKey iotlb_key = *(SMMUIOTLBKey *)key;
+
+    if (SMMU_IOTLB_ASID(iotlb_key) >= 0) {
+        /* This is a stage-1 address. */
+        return false;
+    }
+    if (info->vmid != SMMU_IOTLB_VMID(iotlb_key)) {
+        return false;
+    }
+    return ((info->iova & ~entry->addr_mask) == entry->iova) ||
+           ((entry->iova & ~info->mask) == info->iova);
+}
+
 void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                          uint8_t tg, uint64_t num_pages, uint8_t ttl)
 {
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                                 &info);
 }
 
+/*
+ * Similar to smmu_iotlb_inv_iova(), but for Stage-2, ASID is always -1,
+ * in Stage-1 invalidation ASID = -1, means don't care.
+ */
+void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
+                        uint64_t num_pages, uint8_t ttl)
+{
+    uint8_t granule = tg ? tg * 2 + 10 : 12;
+    int asid = -1;
+
+   if (ttl && (num_pages == 1)) {
+        SMMUIOTLBKey key = smmu_get_iotlb_key(asid, vmid, ipa, tg, ttl);
+
+        if (g_hash_table_remove(s->iotlb, &key)) {
+            return;
+        }
+    }
+
+    SMMUIOTLBPageInvInfo info = {
+        .iova = ipa,
+        .vmid = vmid,
+        .mask = (num_pages << granule) - 1};
+
+    g_hash_table_foreach_remove(s->iotlb,
+                                smmu_hash_remove_by_vmid_ipa,
+                                &info);
+}
+
 void smmu_iotlb_inv_asid(SMMUState *s, int asid)
 {
     trace_smmu_iotlb_inv_asid(asid);
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
     }
 }
 
-static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
+static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
 {
     dma_addr_t end, addr = CMD_ADDR(cmd);
     uint8_t type = CMD_TYPE(cmd);
@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
     }
 
     if (!tg) {
-        trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf);
+        trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf, stage);
         smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1);
-        smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
+        if (stage == SMMU_STAGE_1) {
+            smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
+        } else {
+            smmu_iotlb_inv_ipa(s, vmid, addr, tg, 1, ttl);
+        }
         return;
     }
 
@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd)
         uint64_t mask = dma_aligned_pow2_mask(addr, end, 64);
 
         num_pages = (mask + 1) >> granule;
-        trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages, ttl, leaf);
+        trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages,
+                                 ttl, leaf, stage);
         smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages);
-        smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
+        if (stage == SMMU_STAGE_1) {
+            smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
+        } else {
+            smmu_iotlb_inv_ipa(s, vmid, addr, tg, num_pages, ttl);
+        }
         addr += mask + 1;
     }
 }
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
                 cmd_error = SMMU_CERROR_ILL;
                 break;
             }
-            smmuv3_range_inval(bs, &cmd);
+            smmuv3_range_inval(bs, &cmd, SMMU_STAGE_1);
             break;
         case SMMU_CMD_TLBI_S12_VMALL:
         {
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
              * As currently only either s1 or s2 are supported
              * we can reuse same function for s2.
              */
-            smmuv3_range_inval(bs, &cmd);
+            smmuv3_range_inval(bs, &cmd, SMMU_STAGE_2);
             break;
         case SMMU_CMD_TLBI_EL3_ALL:
         case SMMU_CMD_TLBI_EL3_VA:
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%x - end=0x%x"
 smmuv3_cmdq_cfgi_cd(uint32_t sid) "sid=0x%x"
 smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
-smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
+smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf, int stage) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d stage=%d"
 smmuv3_cmdq_tlbi_nh(void) ""
 smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
 smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Soon, Instead of doing TLB invalidation by ASID only, VMID will be
also required.
Add smmu_iotlb_inv_asid_vmid() which invalidates by both ASID and VMID.

However, at the moment this function is only used in SMMU_CMD_TLBI_NH_ASID
which is a stage-1 command, so passing VMID = -1 keeps the original
behaviour.

Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-14-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  2 +-
 hw/arm/smmu-common.c         | 20 +++++++++++++-------
 hw/arm/smmuv3.c              |  2 +-
 hw/arm/trace-events          |  2 +-
 4 files changed, 16 insertions(+), 10 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
 SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
                                 uint8_t tg, uint8_t level);
 void smmu_iotlb_inv_all(SMMUState *s);
-void smmu_iotlb_inv_asid(SMMUState *s, int asid);
+void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid);
 void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
 void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_all(SMMUState *s)
     g_hash_table_remove_all(s->iotlb);
 }
 
-static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
-                                         gpointer user_data)
+static gboolean smmu_hash_remove_by_asid_vmid(gpointer key, gpointer value,
+                                              gpointer user_data)
 {
-    int asid = *(int *)user_data;
+    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
     SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
 
-    return SMMU_IOTLB_ASID(*iotlb_key) == asid;
+    return (SMMU_IOTLB_ASID(*iotlb_key) == info->asid) &&
+           (SMMU_IOTLB_VMID(*iotlb_key) == info->vmid);
 }
 
 static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
                                 &info);
 }
 
-void smmu_iotlb_inv_asid(SMMUState *s, int asid)
+void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid)
 {
-    trace_smmu_iotlb_inv_asid(asid);
-    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid, &asid);
+    SMMUIOTLBPageInvInfo info = {
+        .asid = asid,
+        .vmid = vmid,
+    };
+
+    trace_smmu_iotlb_inv_asid_vmid(asid, vmid);
+    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_vmid, &info);
 }
 
 void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
 
             trace_smmuv3_cmdq_tlbi_nh_asid(asid);
             smmu_inv_notifiers_all(&s->smmu_state);
-            smmu_iotlb_inv_asid(bs, asid);
+            smmu_iotlb_inv_asid_vmid(bs, asid, -1);
             break;
         }
         case SMMU_CMD_TLBI_NH_ALL:
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_ptw_page_pte(int stage, int level,  uint64_t iova, uint64_t baseaddr, uint6
 smmu_ptw_block_pte(int stage, int level, uint64_t baseaddr, uint64_t pteaddr, uint64_t pte, uint64_t iova, uint64_t gpa, int bsize_mb) "stage=%d level=%d base@=0x%"PRIx64" pte@=0x%"PRIx64" pte=0x%"PRIx64" iova=0x%"PRIx64" block address = 0x%"PRIx64" block size = %d MiB"
 smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "baseaddr=0x%"PRIx64" index=0x%x, pteaddr=0x%"PRIx64", pte=0x%"PRIx64
 smmu_iotlb_inv_all(void) "IOTLB invalidate all"
-smmu_iotlb_inv_asid(int asid) "IOTLB invalidate asid=%d"
+smmu_iotlb_inv_asid_vmid(int asid, int vmid) "IOTLB invalidate asid=%d vmid=%d"
 smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
 smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
 smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Some commands need rework for nesting, as they used to assume S1
and S2 are mutually exclusive:

- CMD_TLBI_NH_ASID: Consider VMID if stage-2 is supported
- CMD_TLBI_NH_ALL: Consider VMID if stage-2 is supported, otherwise
  invalidate everything, this required a new vmid invalidation
  function for stage-1 only (ASID >= 0)

Also, rework trace events to reflect the new implementation.

Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-15-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  1 +
 hw/arm/smmu-common.c         | 16 ++++++++++++++++
 hw/arm/smmuv3.c              | 28 ++++++++++++++++++++++++++--
 hw/arm/trace-events          |  4 +++-
 4 files changed, 46 insertions(+), 3 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(int asid, int vmid, uint64_t iova,
 void smmu_iotlb_inv_all(SMMUState *s);
 void smmu_iotlb_inv_asid_vmid(SMMUState *s, int asid, int vmid);
 void smmu_iotlb_inv_vmid(SMMUState *s, int vmid);
+void smmu_iotlb_inv_vmid_s1(SMMUState *s, int vmid);
 void smmu_iotlb_inv_iova(SMMUState *s, int asid, int vmid, dma_addr_t iova,
                          uint8_t tg, uint64_t num_pages, uint8_t ttl);
 void smmu_iotlb_inv_ipa(SMMUState *s, int vmid, dma_addr_t ipa, uint8_t tg,
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_vmid(gpointer key, gpointer value,
     return SMMU_IOTLB_VMID(*iotlb_key) == vmid;
 }
 
+static gboolean smmu_hash_remove_by_vmid_s1(gpointer key, gpointer value,
+                                            gpointer user_data)
+{
+    int vmid = *(int *)user_data;
+    SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
+
+    return (SMMU_IOTLB_VMID(*iotlb_key) == vmid) &&
+           (SMMU_IOTLB_ASID(*iotlb_key) >= 0);
+}
+
 static gboolean smmu_hash_remove_by_asid_vmid_iova(gpointer key, gpointer value,
                                               gpointer user_data)
 {
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_inv_vmid(SMMUState *s, int vmid)
     g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid, &vmid);
 }
 
+inline void smmu_iotlb_inv_vmid_s1(SMMUState *s, int vmid)
+{
+    trace_smmu_iotlb_inv_vmid_s1(vmid);
+    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_vmid_s1, &vmid);
+}
+
 /* VMSAv8-64 Translation */
 
 /**
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
         case SMMU_CMD_TLBI_NH_ASID:
         {
             int asid = CMD_ASID(&cmd);
+            int vmid = -1;
 
             if (!STAGE1_SUPPORTED(s)) {
                 cmd_error = SMMU_CERROR_ILL;
                 break;
             }
 
+            /*
+             * VMID is only matched when stage 2 is supported, otherwise set it
+             * to -1 as the value used for stage-1 only VMIDs.
+             */
+            if (STAGE2_SUPPORTED(s)) {
+                vmid = CMD_VMID(&cmd);
+            }
+
             trace_smmuv3_cmdq_tlbi_nh_asid(asid);
             smmu_inv_notifiers_all(&s->smmu_state);
-            smmu_iotlb_inv_asid_vmid(bs, asid, -1);
+            smmu_iotlb_inv_asid_vmid(bs, asid, vmid);
             break;
         }
         case SMMU_CMD_TLBI_NH_ALL:
+        {
+            int vmid = -1;
+
             if (!STAGE1_SUPPORTED(s)) {
                 cmd_error = SMMU_CERROR_ILL;
                 break;
             }
+
+            /*
+             * If stage-2 is supported, invalidate for this VMID only, otherwise
+             * invalidate the whole thing.
+             */
+            if (STAGE2_SUPPORTED(s)) {
+                vmid = CMD_VMID(&cmd);
+                trace_smmuv3_cmdq_tlbi_nh(vmid);
+                smmu_iotlb_inv_vmid_s1(bs, vmid);
+                break;
+            }
             QEMU_FALLTHROUGH;
+        }
         case SMMU_CMD_TLBI_NSNH_ALL:
-            trace_smmuv3_cmdq_tlbi_nh();
+            trace_smmuv3_cmdq_tlbi_nsnh();
             smmu_inv_notifiers_all(&s->smmu_state);
             smmu_iotlb_inv_all(bs);
             break;
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_get_pte(uint64_t baseaddr, int index, uint64_t pteaddr, uint64_t pte) "base
 smmu_iotlb_inv_all(void) "IOTLB invalidate all"
 smmu_iotlb_inv_asid_vmid(int asid, int vmid) "IOTLB invalidate asid=%d vmid=%d"
 smmu_iotlb_inv_vmid(int vmid) "IOTLB invalidate vmid=%d"
+smmu_iotlb_inv_vmid_s1(int vmid) "IOTLB invalidate vmid=%d"
 smmu_iotlb_inv_iova(int asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
 smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
 smmu_iotlb_lookup_hit(int asid, int vmid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d vmid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_cd(uint32_t sid) "sid=0x%x"
 smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid=0x%x (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf, int stage) "vmid=%d asid=%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d stage=%d"
-smmuv3_cmdq_tlbi_nh(void) ""
+smmuv3_cmdq_tlbi_nh(int vmid) "vmid=%d"
+smmuv3_cmdq_tlbi_nsnh(void) ""
 smmuv3_cmdq_tlbi_nh_asid(int asid) "asid=%d"
 smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

IOMMUTLBEvent only understands IOVA, for stage-1 or stage-2
SMMU instances we consider the input address as the IOVA, but when
nesting is used, we can't mix stage-1 and stage-2 addresses, so for
nesting only stage-1 is considered the IOVA and would be notified.

Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-16-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3.c     | 39 +++++++++++++++++++++++++--------------
 hw/arm/trace-events |  2 +-
 2 files changed, 26 insertions(+), 15 deletions(-)

diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ epilogue:
  * @iova: iova
  * @tg: translation granule (if communicated through range invalidation)
  * @num_pages: number of @granule sized pages (if tg != 0), otherwise 1
+ * @stage: Which stage(1 or 2) is used
  */
 static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
                                IOMMUNotifier *n,
                                int asid, int vmid,
                                dma_addr_t iova, uint8_t tg,
-                               uint64_t num_pages)
+                               uint64_t num_pages, int stage)
 {
     SMMUDevice *sdev = container_of(mr, SMMUDevice, iommu);
+    SMMUEventInfo eventinfo = {.inval_ste_allowed = true};
+    SMMUTransCfg *cfg = smmuv3_get_config(sdev, &eventinfo);
     IOMMUTLBEvent event;
     uint8_t granule;
-    SMMUv3State *s = sdev->smmu;
+
+    if (!cfg) {
+        return;
+    }
+
+    /*
+     * stage is passed from TLB invalidation commands which can be either
+     * stage-1 or stage-2.
+     * However, IOMMUTLBEvent only understands IOVA, for stage-1 or stage-2
+     * SMMU instances we consider the input address as the IOVA, but when
+     * nesting is used, we can't mix stage-1 and stage-2 addresses, so for
+     * nesting only stage-1 is considered the IOVA and would be notified.
+     */
+    if ((stage == SMMU_STAGE_2) && (cfg->stage == SMMU_NESTED))
+        return;
 
     if (!tg) {
-        SMMUEventInfo eventinfo = {.inval_ste_allowed = true};
-        SMMUTransCfg *cfg = smmuv3_get_config(sdev, &eventinfo);
         SMMUTransTableInfo *tt;
 
-        if (!cfg) {
-            return;
-        }
-
         if (asid >= 0 && cfg->asid != asid) {
             return;
         }
@@ -XXX,XX +XXX,XX @@ static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
             return;
         }
 
-        if (STAGE1_SUPPORTED(s)) {
+        if (stage == SMMU_STAGE_1) {
             tt = select_tt(cfg, iova);
             if (!tt) {
                 return;
@@ -XXX,XX +XXX,XX @@ static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
 /* invalidate an asid/vmid/iova range tuple in all mr's */
 static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
                                       dma_addr_t iova, uint8_t tg,
-                                      uint64_t num_pages)
+                                      uint64_t num_pages, int stage)
 {
     SMMUDevice *sdev;
 
@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, int vmid,
         IOMMUNotifier *n;
 
         trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, vmid,
-                                        iova, tg, num_pages);
+                                        iova, tg, num_pages, stage);
 
         IOMMU_NOTIFIER_FOREACH(n, mr) {
-            smmuv3_notify_iova(mr, n, asid, vmid, iova, tg, num_pages);
+            smmuv3_notify_iova(mr, n, asid, vmid, iova, tg, num_pages, stage);
         }
     }
 }
@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
 
     if (!tg) {
         trace_smmuv3_range_inval(vmid, asid, addr, tg, 1, ttl, leaf, stage);
-        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1);
+        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, 1, stage);
         if (stage == SMMU_STAGE_1) {
             smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, 1, ttl);
         } else {
@@ -XXX,XX +XXX,XX @@ static void smmuv3_range_inval(SMMUState *s, Cmd *cmd, SMMUStage stage)
         num_pages = (mask + 1) >> granule;
         trace_smmuv3_range_inval(vmid, asid, addr, tg, num_pages,
                                  ttl, leaf, stage);
-        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages);
+        smmuv3_inv_notifiers_iova(s, asid, vmid, addr, tg, num_pages, stage);
         if (stage == SMMU_STAGE_1) {
             smmu_iotlb_inv_iova(s, asid, vmid, addr, tg, num_pages, ttl);
         } else {
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_tlbi_s12_vmid(int vmid) "vmid=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid=0x%x"
 smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
 smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
-smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
+smmuv3_inv_notifiers_iova(const char *name, int asid, int vmid, uint64_t iova, uint8_t tg, uint64_t num_pages, int stage) "iommu mr=%s asid=%d vmid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" stage=%d"
 
 # strongarm.c
 strongarm_uart_update_parameters(const char *label, int speed, char parity, int data_bits, int stop_bits) "%s speed=%d parity=%c data=%d stop=%d"
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Previously, to check if faults are enabled, it was sufficient to check
the current stage of translation and check the corresponding
record_faults flag.

However, with nesting, it is possible for stage-1 (nested) translation
to trigger a stage-2 fault, so we check SMMUPTWEventInfo as it would
have the correct stage set from the page table walk.

Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-17-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3.c | 15 ++++++++-------
 1 file changed, 8 insertions(+), 7 deletions(-)

diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@
 #include "smmuv3-internal.h"
 #include "smmu-internal.h"
 
-#define PTW_RECORD_FAULT(cfg)   (((cfg)->stage == SMMU_STAGE_1) ? \
-                                 (cfg)->record_faults : \
-                                 (cfg)->s2cfg.record_faults)
+#define PTW_RECORD_FAULT(ptw_info, cfg) (((ptw_info).stage == SMMU_STAGE_1 && \
+                                        (cfg)->record_faults) || \
+                                        ((ptw_info).stage == SMMU_STAGE_2 && \
+                                        (cfg)->s2cfg.record_faults))
 
 /**
  * smmuv3_trigger_irq - pulse @irq if enabled and update
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
             event->u.f_walk_eabt.addr2 = ptw_info.addr;
             break;
         case SMMU_PTW_ERR_TRANSLATION:
-            if (PTW_RECORD_FAULT(cfg)) {
+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
                 event->type = SMMU_EVT_F_TRANSLATION;
                 event->u.f_translation.addr2 = ptw_info.addr;
                 event->u.f_translation.class = class;
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
             }
             break;
         case SMMU_PTW_ERR_ADDR_SIZE:
-            if (PTW_RECORD_FAULT(cfg)) {
+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
                 event->type = SMMU_EVT_F_ADDR_SIZE;
                 event->u.f_addr_size.addr2 = ptw_info.addr;
                 event->u.f_addr_size.class = class;
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
             }
             break;
         case SMMU_PTW_ERR_ACCESS:
-            if (PTW_RECORD_FAULT(cfg)) {
+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
                 event->type = SMMU_EVT_F_ACCESS;
                 event->u.f_access.addr2 = ptw_info.addr;
                 event->u.f_access.class = class;
@@ -XXX,XX +XXX,XX @@ static SMMUTranslationStatus smmuv3_do_translate(SMMUv3State *s, hwaddr addr,
             }
             break;
         case SMMU_PTW_ERR_PERMISSION:
-            if (PTW_RECORD_FAULT(cfg)) {
+            if (PTW_RECORD_FAULT(ptw_info, cfg)) {
                 event->type = SMMU_EVT_F_PERMISSION;
                 event->u.f_permission.addr2 = ptw_info.addr;
                 event->u.f_permission.class = class;
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

Everything is in place, consolidate parsing of STE cfg and setting
translation stage.

Advertise nesting if stage requested is "nested".

Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-18-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3.c | 35 ++++++++++++++++++++++++++---------
 1 file changed, 26 insertions(+), 9 deletions(-)

diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
     /* Based on sys property, the stages supported in smmu will be advertised.*/
     if (s->stage && !strcmp("2", s->stage)) {
         s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S2P, 1);
+    } else if (s->stage && !strcmp("nested", s->stage)) {
+        s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S1P, 1);
+        s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S2P, 1);
     } else {
         s->idr[0] = FIELD_DP32(s->idr[0], IDR0, S1P, 1);
     }
@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
 
 static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
 {
-    cfg->stage = SMMU_STAGE_2;
-
     if (STE_S2AA64(ste) == 0x0) {
         qemu_log_mask(LOG_UNIMP,
                       "SMMUv3 AArch32 tables not supported\n");
@@ -XXX,XX +XXX,XX @@ bad_ste:
     return -EINVAL;
 }
 
+static void decode_ste_config(SMMUTransCfg *cfg, uint32_t config)
+{
+
+    if (STE_CFG_ABORT(config)) {
+        cfg->aborted = true;
+        return;
+    }
+    if (STE_CFG_BYPASS(config)) {
+        cfg->bypassed = true;
+        return;
+    }
+
+    if (STE_CFG_S1_ENABLED(config)) {
+        cfg->stage = SMMU_STAGE_1;
+    }
+
+    if (STE_CFG_S2_ENABLED(config)) {
+        cfg->stage |= SMMU_STAGE_2;
+    }
+}
+
 /* Returns < 0 in case of invalid STE, 0 otherwise */
 static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
                       STE *ste, SMMUEventInfo *event)
@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
 
     config = STE_CONFIG(ste);
 
-    if (STE_CFG_ABORT(config)) {
-        cfg->aborted = true;
-        return 0;
-    }
+    decode_ste_config(cfg, config);
 
-    if (STE_CFG_BYPASS(config)) {
-        cfg->bypassed = true;
+    if (cfg->aborted || cfg->bypassed) {
         return 0;
     }
 
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
 
     /* we support only those at the moment */
     cfg->aa64 = true;
-    cfg->stage = SMMU_STAGE_1;
 
     cfg->oas = oas2bits(CD_IPS(cd));
     cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
-- 
2.34.1

From: Mostafa Saleh <smostafa@google.com>

SMMUv3 OAS is currently hardcoded in the code to 44 bits, for nested
configurations that can be a problem, as stage-2 might be shared with
the CPU which might have different PARANGE, and according to SMMU manual
ARM IHI 0070F.b:
    6.3.6 SMMU_IDR5, OAS must match the system physical address size.

This patch doesn't change the SMMU OAS, but refactors the code to
make it easier to do that:
- Rely everywhere on IDR5 for reading OAS instead of using the
  SMMU_IDR5_OAS macro, so, it is easier just to change IDR5 and
  it propagages correctly.
- Add additional checks when OAS is greater than 48bits.
- Remove unused functions/macros: pa_range/MAX_PA.

Reviewed-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Mostafa Saleh <smostafa@google.com>
Reviewed-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240715084519.1189624-19-smostafa@google.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h | 13 -------------
 hw/arm/smmu-common.c     |  7 ++++---
 hw/arm/smmuv3.c          | 35 ++++++++++++++++++++++++++++-------
 3 files changed, 32 insertions(+), 23 deletions(-)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ static inline int oas2bits(int oas_field)
     return -1;
 }
 
-static inline int pa_range(STE *ste)
-{
-    int oas_field = MIN(STE_S2PS(ste), SMMU_IDR5_OAS);
-
-    if (!STE_S2AA64(ste)) {
-        return 40;
-    }
-
-    return oas2bits(oas_field);
-}
-
-#define MAX_PA(ste) ((1 << pa_range(ste)) - 1)
-
 /* CD fields */
 
 #define CD_VALID(x)   extract32((x)->word[0], 31, 1)
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s1(SMMUState *bs, SMMUTransCfg *cfg,
     inputsize = 64 - tt->tsz;
     level = 4 - (inputsize - 4) / stride;
     indexmask = VMSA_IDXMSK(inputsize, stride, level);
-    baseaddr = extract64(tt->ttb, 0, 48);
+
+    baseaddr = extract64(tt->ttb, 0, cfg->oas);
     baseaddr &= ~indexmask;
 
     while (level < VMSA_LEVELS) {
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64_s2(SMMUTransCfg *cfg,
      * Get the ttb from concatenated structure.
      * The offset is the idx * size of each ttb(number of ptes * (sizeof(pte))
      */
-    uint64_t baseaddr = extract64(cfg->s2cfg.vttb, 0, 48) + (1 << stride) *
-                                  idx * sizeof(uint64_t);
+    uint64_t baseaddr = extract64(cfg->s2cfg.vttb, 0, cfg->s2cfg.eff_ps) +
+                                  (1 << stride) * idx * sizeof(uint64_t);
     dma_addr_t indexmask = VMSA_IDXMSK(inputsize, stride, level);
 
     baseaddr &= ~indexmask;
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static bool s2t0sz_valid(SMMUTransCfg *cfg)
     }
 
     if (cfg->s2cfg.granule_sz == 16) {
-        return (cfg->s2cfg.tsz >= 64 - oas2bits(SMMU_IDR5_OAS));
+        return (cfg->s2cfg.tsz >= 64 - cfg->s2cfg.eff_ps);
     }
 
-    return (cfg->s2cfg.tsz >= MAX(64 - oas2bits(SMMU_IDR5_OAS), 16));
+    return (cfg->s2cfg.tsz >= MAX(64 - cfg->s2cfg.eff_ps, 16));
 }
 
 /*
@@ -XXX,XX +XXX,XX @@ static bool s2_pgtable_config_valid(uint8_t sl0, uint8_t t0sz, uint8_t gran)
     return nr_concat <= VMSA_MAX_S2_CONCAT;
 }
 
-static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
+static int decode_ste_s2_cfg(SMMUv3State *s, SMMUTransCfg *cfg,
+                             STE *ste)
 {
+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
+
     if (STE_S2AA64(ste) == 0x0) {
         qemu_log_mask(LOG_UNIMP,
                       "SMMUv3 AArch32 tables not supported\n");
@@ -XXX,XX +XXX,XX @@ static int decode_ste_s2_cfg(SMMUTransCfg *cfg, STE *ste)
     }
 
     /* For AA64, The effective S2PS size is capped to the OAS. */
-    cfg->s2cfg.eff_ps = oas2bits(MIN(STE_S2PS(ste), SMMU_IDR5_OAS));
+    cfg->s2cfg.eff_ps = oas2bits(MIN(STE_S2PS(ste), oas));
+    /*
+     * For SMMUv3.1 and later, when OAS == IAS == 52, the stage 2 input
+     * range is further limited to 48 bits unless STE.S2TG indicates a
+     * 64KB granule.
+     */
+    if (cfg->s2cfg.granule_sz != 16) {
+        cfg->s2cfg.eff_ps = MIN(cfg->s2cfg.eff_ps, 48);
+    }
     /*
      * It is ILLEGAL for the address in S2TTB to be outside the range
      * described by the effective S2PS value.
@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
                       STE *ste, SMMUEventInfo *event)
 {
     uint32_t config;
+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
     int ret;
 
     if (!STE_VALID(ste)) {
@@ -XXX,XX +XXX,XX @@ static int decode_ste(SMMUv3State *s, SMMUTransCfg *cfg,
          * Stage-1 OAS defaults to OAS even if not enabled as it would be used
          * in input address check for stage-2.
          */
-        cfg->oas = oas2bits(SMMU_IDR5_OAS);
-        ret = decode_ste_s2_cfg(cfg, ste);
+        cfg->oas = oas2bits(oas);
+        ret = decode_ste_s2_cfg(s, cfg, ste);
         if (ret) {
             goto bad_ste;
         }
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
     int i;
     SMMUTranslationStatus status;
     SMMUTLBEntry *entry;
+    uint8_t oas = FIELD_EX32(s->idr[5], IDR5, OAS);
 
     if (!CD_VALID(cd) || !CD_AARCH64(cd)) {
         goto bad_cd;
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
     cfg->aa64 = true;
 
     cfg->oas = oas2bits(CD_IPS(cd));
-    cfg->oas = MIN(oas2bits(SMMU_IDR5_OAS), cfg->oas);
+    cfg->oas = MIN(oas2bits(oas), cfg->oas);
     cfg->tbi = CD_TBI(cd);
     cfg->asid = CD_ASID(cd);
     cfg->affd = CD_AFFD(cd);
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUv3State *s, SMMUTransCfg *cfg,
             goto bad_cd;
         }
 
+        /*
+         * An address greater than 48 bits in size can only be output from a
+         * TTD when, in SMMUv3.1 and later, the effective IPS is 52 and a 64KB
+         * granule is in use for that translation table
+         */
+        if (tt->granule_sz != 16) {
+            cfg->oas = MIN(cfg->oas, 48);
+        }
         tt->tsz = tsz;
         tt->ttb = CD_TTB(cd, i);
 
-- 
2.34.1

From: Daniyal Khan <danikhan632@gmail.com>

We made a copy above because the fp exception flags
are not propagated back to the FPST register, but
then failed to use the copy.

Cc: qemu-stable@nongnu.org
Fixes: 558e956c719 ("target/arm: Implement FMOPA, FMOPS (non-widening)")
Signed-off-by: Daniyal Khan <danikhan632@gmail.com>
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240717060149.204788-2-richard.henderson@linaro.org
[rth: Split from a larger patch]
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/sme_helper.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/target/arm/tcg/sme_helper.c b/target/arm/tcg/sme_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/sme_helper.c
+++ b/target/arm/tcg/sme_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(sme_fmopa_s)(void *vza, void *vzn, void *vzm, void *vpn,
                         if (pb & 1) {
                             uint32_t *a = vza_row + H1_4(col);
                             uint32_t *m = vzm + H1_4(col);
-                            *a = float32_muladd(n, *m, *a, 0, vst);
+                            *a = float32_muladd(n, *m, *a, 0, &fpst);
                         }
                         col += 4;
                         pb >>= 4;
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

This operation has float16 inputs and thus must use
the FZ16 control not the FZ control.

Cc: qemu-stable@nongnu.org
Fixes: 3916841ac75 ("target/arm: Implement FMOPA, FMOPS (widening)")
Reported-by: Daniyal Khan <danikhan632@gmail.com>
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240717060149.204788-3-richard.henderson@linaro.org
Resolves: https://gitlab.com/qemu-project/qemu/-/issues/2374
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/translate-sme.c | 12 ++++++++----
 1 file changed, 8 insertions(+), 4 deletions(-)

diff --git a/target/arm/tcg/translate-sme.c b/target/arm/tcg/translate-sme.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-sme.c
+++ b/target/arm/tcg/translate-sme.c
@@ -XXX,XX +XXX,XX @@ static bool do_outprod(DisasContext *s, arg_op *a, MemOp esz,
 }
 
 static bool do_outprod_fpst(DisasContext *s, arg_op *a, MemOp esz,
+                            ARMFPStatusFlavour e_fpst,
                             gen_helper_gvec_5_ptr *fn)
 {
     int svl = streaming_vec_reg_size(s);
@@ -XXX,XX +XXX,XX @@ static bool do_outprod_fpst(DisasContext *s, arg_op *a, MemOp esz,
     zm = vec_full_reg_ptr(s, a->zm);
     pn = pred_full_reg_ptr(s, a->pn);
     pm = pred_full_reg_ptr(s, a->pm);
-    fpst = fpstatus_ptr(FPST_FPCR);
+    fpst = fpstatus_ptr(e_fpst);
 
     fn(za, zn, zm, pn, pm, fpst, tcg_constant_i32(desc));
     return true;
 }
 
-TRANS_FEAT(FMOPA_h, aa64_sme, do_outprod_fpst, a, MO_32, gen_helper_sme_fmopa_h)
-TRANS_FEAT(FMOPA_s, aa64_sme, do_outprod_fpst, a, MO_32, gen_helper_sme_fmopa_s)
-TRANS_FEAT(FMOPA_d, aa64_sme_f64f64, do_outprod_fpst, a, MO_64, gen_helper_sme_fmopa_d)
+TRANS_FEAT(FMOPA_h, aa64_sme, do_outprod_fpst, a,
+           MO_32, FPST_FPCR_F16, gen_helper_sme_fmopa_h)
+TRANS_FEAT(FMOPA_s, aa64_sme, do_outprod_fpst, a,
+           MO_32, FPST_FPCR, gen_helper_sme_fmopa_s)
+TRANS_FEAT(FMOPA_d, aa64_sme_f64f64, do_outprod_fpst, a,
+           MO_64, FPST_FPCR, gen_helper_sme_fmopa_d)
 
 /* TODO: FEAT_EBF16 */
 TRANS_FEAT(BFMOPA, aa64_sme, do_outprod, a, MO_32, gen_helper_sme_bfmopa)
-- 
2.34.1

From: Daniyal Khan <danikhan632@gmail.com>

Signed-off-by: Daniyal Khan <danikhan632@gmail.com>
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20240717060149.204788-4-richard.henderson@linaro.org
Message-Id: 172090222034.13953.16888708708822922098-1@git.sr.ht
[rth: Split test from a larger patch, tidy assembly]
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 tests/tcg/aarch64/sme-fmopa-1.c   | 63 +++++++++++++++++++++++++++++++
 tests/tcg/aarch64/sme-fmopa-2.c   | 56 +++++++++++++++++++++++++++
 tests/tcg/aarch64/sme-fmopa-3.c   | 63 +++++++++++++++++++++++++++++++
 tests/tcg/aarch64/Makefile.target |  5 ++-
 4 files changed, 185 insertions(+), 2 deletions(-)
 create mode 100644 tests/tcg/aarch64/sme-fmopa-1.c
 create mode 100644 tests/tcg/aarch64/sme-fmopa-2.c
 create mode 100644 tests/tcg/aarch64/sme-fmopa-3.c

diff --git a/tests/tcg/aarch64/sme-fmopa-1.c b/tests/tcg/aarch64/sme-fmopa-1.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/tests/tcg/aarch64/sme-fmopa-1.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SME outer product, 1 x 1.
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include <stdio.h>
+
+static void foo(float *dst)
+{
+    asm(".arch_extension sme\n\t"
+        "smstart\n\t"
+        "ptrue p0.s, vl4\n\t"
+        "fmov z0.s, #1.0\n\t"
+        /*
+         * An outer product of a vector of 1.0 by itself should be a matrix of 1.0.
+         * Note that we are using tile 1 here (za1.s) rather than tile 0.
+         */
+        "zero {za}\n\t"
+        "fmopa za1.s, p0/m, p0/m, z0.s, z0.s\n\t"
+        /*
+         * Read the first 4x4 sub-matrix of elements from tile 1:
+         * Note that za1h should be interchangeable here.
+         */
+        "mov w12, #0\n\t"
+        "mova z0.s, p0/m, za1v.s[w12, #0]\n\t"
+        "mova z1.s, p0/m, za1v.s[w12, #1]\n\t"
+        "mova z2.s, p0/m, za1v.s[w12, #2]\n\t"
+        "mova z3.s, p0/m, za1v.s[w12, #3]\n\t"
+        /*
+         * And store them to the input pointer (dst in the C code):
+         */
+        "st1w {z0.s}, p0, [%0]\n\t"
+        "add x0, x0, #16\n\t"
+        "st1w {z1.s}, p0, [x0]\n\t"
+        "add x0, x0, #16\n\t"
+        "st1w {z2.s}, p0, [x0]\n\t"
+        "add x0, x0, #16\n\t"
+        "st1w {z3.s}, p0, [x0]\n\t"
+        "smstop"
+        : : "r"(dst)
+        : "x12", "d0", "d1", "d2", "d3", "memory");
+}
+
+int main()
+{
+    float dst[16] = { };
+
+    foo(dst);
+
+    for (int i = 0; i < 16; i++) {
+        if (dst[i] != 1.0f) {
+            goto failure;
+        }
+    }
+    /* success */
+    return 0;
+
+ failure:
+    for (int i = 0; i < 16; i++) {
+        printf("%f%c", dst[i], i % 4 == 3 ? '\n' : ' ');
+    }
+    return 1;
+}
diff --git a/tests/tcg/aarch64/sme-fmopa-2.c b/tests/tcg/aarch64/sme-fmopa-2.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/tests/tcg/aarch64/sme-fmopa-2.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SME outer product, FZ vs FZ16
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include <stdint.h>
+#include <stdio.h>
+
+static void test_fmopa(uint32_t *result)
+{
+    asm(".arch_extension sme\n\t"
+        "smstart\n\t"               /* Z*, P* and ZArray cleared */
+        "ptrue p2.b, vl16\n\t"      /* Limit vector length to 16 */
+        "ptrue p5.b, vl16\n\t"
+        "movi d0, #0x00ff\n\t"      /* fp16 denormal */
+        "movi d16, #0x00ff\n\t"
+        "mov w15, #0x0001000000\n\t" /* FZ=1, FZ16=0 */
+        "msr fpcr, x15\n\t"
+        "fmopa za3.s, p2/m, p5/m, z16.h, z0.h\n\t"
+        "mov w15, #0\n\t"
+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
+        "mov w15, #2\n\t"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
+        "smstop"
+        : "+r"(result) :
+        : "x15", "x16", "p2", "p5", "d0", "d16", "memory");
+}
+
+int main(void)
+{
+    uint32_t result[4 * 4] = { };
+
+    test_fmopa(result);
+
+    if (result[0] != 0x2f7e0100) {
+        printf("Test failed: Incorrect output in first 4 bytes\n"
+               "Expected: %08x\n"
+               "Got:      %08x\n",
+               0x2f7e0100, result[0]);
+        return 1;
+    }
+
+    for (int i = 1; i < 16; ++i) {
+        if (result[i] != 0) {
+            printf("Test failed: Non-zero word at position %d\n", i);
+            return 1;
+        }
+    }
+
+    return 0;
+}
diff --git a/tests/tcg/aarch64/sme-fmopa-3.c b/tests/tcg/aarch64/sme-fmopa-3.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/tests/tcg/aarch64/sme-fmopa-3.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SME outer product, [ 1 2 3 4 ] squared
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include <stdio.h>
+#include <stdint.h>
+#include <string.h>
+#include <math.h>
+
+static const float i_1234[4] = {
+    1.0f, 2.0f, 3.0f, 4.0f
+};
+
+static const float expected[4] = {
+    4.515625f, 5.750000f, 6.984375f, 8.218750f
+};
+
+static void test_fmopa(float *result)
+{
+    asm(".arch_extension sme\n\t"
+        "smstart\n\t"               /* ZArray cleared */
+        "ptrue p2.b, vl16\n\t"      /* Limit vector length to 16 */
+        "ld1w {z0.s}, p2/z, [%1]\n\t"
+        "mov w15, #0\n\t"
+        "mov za3h.s[w15, 0], p2/m, z0.s\n\t"
+        "mov za3h.s[w15, 1], p2/m, z0.s\n\t"
+        "mov w15, #2\n\t"
+        "mov za3h.s[w15, 0], p2/m, z0.s\n\t"
+        "mov za3h.s[w15, 1], p2/m, z0.s\n\t"
+        "msr fpcr, xzr\n\t"
+        "fmopa za3.s, p2/m, p2/m, z0.h, z0.h\n\t"
+        "mov w15, #0\n\t"
+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
+        "mov w15, #2\n\t"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 0]}, p2, [%0]\n\t"
+        "add %0, %0, #16\n\t"
+        "st1w {za3h.s[w15, 1]}, p2, [%0]\n\t"
+        "smstop"
+        : "+r"(result) : "r"(i_1234)
+        : "x15", "x16", "p2", "d0", "memory");
+}
+
+int main(void)
+{
+    float result[4 * 4] = { };
+    int ret = 0;
+
+    test_fmopa(result);
+
+    for (int i = 0; i < 4; i++) {
+        float actual = result[i];
+        if (fabsf(actual - expected[i]) > 0.001f) {
+            printf("Test failed at element %d: Expected %f, got %f\n",
+                   i, expected[i], actual);
+            ret = 1;
+        }
+    }
+    return ret;
+}
diff --git a/tests/tcg/aarch64/Makefile.target b/tests/tcg/aarch64/Makefile.target
index XXXXXXX..XXXXXXX 100644
--- a/tests/tcg/aarch64/Makefile.target
+++ b/tests/tcg/aarch64/Makefile.target
@@ -XXX,XX +XXX,XX @@ endif
 
 # SME Tests
 ifneq ($(CROSS_AS_HAS_ARMV9_SME),)
-AARCH64_TESTS += sme-outprod1 sme-smopa-1 sme-smopa-2
-sme-outprod1 sme-smopa-1 sme-smopa-2: CFLAGS += $(CROSS_AS_HAS_ARMV9_SME)
+SME_TESTS = sme-outprod1 sme-smopa-1 sme-smopa-2 sme-fmopa-1 sme-fmopa-2 sme-fmopa-3
+AARCH64_TESTS += $(SME_TESTS)
+$(SME_TESTS): CFLAGS += $(CROSS_AS_HAS_ARMV9_SME)
 endif
 
 # System Registers Tests
-- 
2.34.1

From: Akihiko Odaki <akihiko.odaki@daynix.com>

Asahi Linux supports KVM but lacks PMU support.

Signed-off-by: Akihiko Odaki <akihiko.odaki@daynix.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20240716-pmu-v3-1-8c7c1858a227@daynix.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 tests/qtest/arm-cpu-features.c | 13 ++++++++-----
 1 file changed, 8 insertions(+), 5 deletions(-)

diff --git a/tests/qtest/arm-cpu-features.c b/tests/qtest/arm-cpu-features.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/qtest/arm-cpu-features.c
+++ b/tests/qtest/arm-cpu-features.c
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
     assert_set_feature(qts, "host", "kvm-no-adjvtime", false);
 
     if (g_str_equal(qtest_get_arch(), "aarch64")) {
+        bool kvm_supports_pmu;
         bool kvm_supports_steal_time;
         bool kvm_supports_sve;
         char max_name[8], name[8];
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
 
         assert_has_feature_enabled(qts, "host", "aarch64");
 
-        /* Enabling and disabling pmu should always work. */
-        assert_has_feature_enabled(qts, "host", "pmu");
-        assert_set_feature(qts, "host", "pmu", false);
-        assert_set_feature(qts, "host", "pmu", true);
-
         /*
          * Some features would be enabled by default, but they're disabled
          * because this instance of KVM doesn't support them. Test that the
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
         assert_has_feature(qts, "host", "sve");
 
         resp = do_query_no_props(qts, "host");
+        kvm_supports_pmu = resp_get_feature(resp, "pmu");
         kvm_supports_steal_time = resp_get_feature(resp, "kvm-steal-time");
         kvm_supports_sve = resp_get_feature(resp, "sve");
         vls = resp_get_sve_vls(resp);
         qobject_unref(resp);
 
+        if (kvm_supports_pmu) {
+            /* If we have pmu then we should be able to toggle it. */
+            assert_set_feature(qts, "host", "pmu", false);
+            assert_set_feature(qts, "host", "pmu", true);
+        }
+
         if (kvm_supports_steal_time) {
             /* If we have steal-time then we should be able to toggle it. */
             assert_set_feature(qts, "host", "kvm-steal-time", false);
-- 
2.34.1

From: Akihiko Odaki <akihiko.odaki@daynix.com>

hvf did not advance PC when raising an exception for most unhandled
system registers, but it mistakenly advanced PC when raising an
exception for GICv3 registers.

Cc: qemu-stable@nongnu.org
Fixes: a2260983c655 ("hvf: arm: Add support for GICv3")
Signed-off-by: Akihiko Odaki <akihiko.odaki@daynix.com>
Message-id: 20240716-pmu-v3-4-8c7c1858a227@daynix.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/hvf/hvf.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/target/arm/hvf/hvf.c b/target/arm/hvf/hvf.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/hvf/hvf.c
+++ b/target/arm/hvf/hvf.c
@@ -XXX,XX +XXX,XX @@ static int hvf_sysreg_read(CPUState *cpu, uint32_t reg, uint32_t rt)
         /* Call the TCG sysreg handler. This is only safe for GICv3 regs. */
         if (!hvf_sysreg_read_cp(cpu, reg, &val)) {
             hvf_raise_exception(cpu, EXCP_UDEF, syn_uncategorized());
+            return 1;
         }
         break;
     case SYSREG_DBGBVR0_EL1:
-- 
2.34.1

First arm pullreq of the cycle; this is mostly my softfloat NaN
handling series. (Lots more in my to-review queue, but I don't
like pullreqs growing too close to a hundred patches at a time :-))

thanks
-- PMM

The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:

Open 10.0 development tree (2024-12-10 17:41:17 +0000)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211

for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:

MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)

----------------------------------------------------------------
target-arm queue:
 * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
 * fpu: Make muladd NaN handling runtime-selected, not compile-time
 * fpu: Make default NaN pattern runtime-selected, not compile-time
 * fpu: Minor NaN-related cleanups
 * MAINTAINERS: email address updates

----------------------------------------------------------------
Bernhard Beschow (5):
      hw/net/lan9118: Extract lan9118_phy
      hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
      hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
      hw/net/lan9118_phy: Reuse MII constants
      hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement

Leif Lindholm (1):
      MAINTAINERS: update email address for Leif Lindholm

Peter Maydell (54):
      fpu: handle raising Invalid for infzero in pick_nan_muladd
      fpu: Check for default_nan_mode before calling pickNaNMulAdd
      softfloat: Allow runtime choice of inf * 0 + NaN result
      tests/fp: Explicitly set inf-zero-nan rule
      target/arm: Set FloatInfZeroNaNRule explicitly
      target/s390: Set FloatInfZeroNaNRule explicitly
      target/ppc: Set FloatInfZeroNaNRule explicitly
      target/mips: Set FloatInfZeroNaNRule explicitly
      target/sparc: Set FloatInfZeroNaNRule explicitly
      target/xtensa: Set FloatInfZeroNaNRule explicitly
      target/x86: Set FloatInfZeroNaNRule explicitly
      target/loongarch: Set FloatInfZeroNaNRule explicitly
      target/hppa: Set FloatInfZeroNaNRule explicitly
      softfloat: Pass have_snan to pickNaNMulAdd
      softfloat: Allow runtime choice of NaN propagation for muladd
      tests/fp: Explicitly set 3-NaN propagation rule
      target/arm: Set Float3NaNPropRule explicitly
      target/loongarch: Set Float3NaNPropRule explicitly
      target/ppc: Set Float3NaNPropRule explicitly
      target/s390x: Set Float3NaNPropRule explicitly
      target/sparc: Set Float3NaNPropRule explicitly
      target/mips: Set Float3NaNPropRule explicitly
      target/xtensa: Set Float3NaNPropRule explicitly
      target/i386: Set Float3NaNPropRule explicitly
      target/hppa: Set Float3NaNPropRule explicitly
      fpu: Remove use_first_nan field from float_status
      target/m68k: Don't pass NULL float_status to floatx80_default_nan()
      softfloat: Create floatx80 default NaN from parts64_default_nan
      target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
      target/m68k: In frem helper, initialize local float_status from env->fp_status
      target/m68k: Init local float_status from env fp_status in gdb get/set reg
      target/sparc: Initialize local scratch float_status from env->fp_status
      target/ppc: Use env->fp_status in helper_compute_fprf functions
      fpu: Allow runtime choice of default NaN value
      tests/fp: Set default NaN pattern explicitly
      target/microblaze: Set default NaN pattern explicitly
      target/i386: Set default NaN pattern explicitly
      target/hppa: Set default NaN pattern explicitly
      target/alpha: Set default NaN pattern explicitly
      target/arm: Set default NaN pattern explicitly
      target/loongarch: Set default NaN pattern explicitly
      target/m68k: Set default NaN pattern explicitly
      target/mips: Set default NaN pattern explicitly
      target/openrisc: Set default NaN pattern explicitly
      target/ppc: Set default NaN pattern explicitly
      target/sh4: Set default NaN pattern explicitly
      target/rx: Set default NaN pattern explicitly
      target/s390x: Set default NaN pattern explicitly
      target/sparc: Set default NaN pattern explicitly
      target/xtensa: Set default NaN pattern explicitly
      target/hexagon: Set default NaN pattern explicitly
      target/riscv: Set default NaN pattern explicitly
      target/tricore: Set default NaN pattern explicitly
      fpu: Remove default handling for dnan_pattern

Richard Henderson (11):
      target/arm: Copy entire float_status in is_ebf
      softfloat: Inline pickNaNMulAdd
      softfloat: Use goto for default nan case in pick_nan_muladd
      softfloat: Remove which from parts_pick_nan_muladd
      softfloat: Pad array size in pick_nan_muladd
      softfloat: Move propagateFloatx80NaN to softfloat.c
      softfloat: Use parts_pick_nan in propagateFloatx80NaN
      softfloat: Inline pickNaN
      softfloat: Share code between parts_pick_nan cases
      softfloat: Sink frac_cmp in parts_pick_nan until needed
      softfloat: Replace WHICH with RET in parts_pick_nan

Vikram Garhwal (1):
      MAINTAINERS: Add correct email address for Vikram Garhwal

From: Bernhard Beschow <shentey@gmail.com>

A very similar implementation of the same device exists in imx_fec. Prepare for
a common implementation by extracting a device model into its own files.

Some migration state has been moved into the new device model which breaks
migration compatibility for the following machines:
* smdkc210
* realview-*
* vexpress-*
* kzm
* mps2-*

While breaking migration ABI, fix the size of the MII registers to be 16 bit,
as defined by IEEE 802.3u.

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-2-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/lan9118_phy.h |  37 ++++++++
 hw/net/lan9118.c             | 137 +++++-----------------------
 hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
 hw/net/Kconfig               |   4 +
 hw/net/meson.build           |   1 +
 5 files changed, 233 insertions(+), 115 deletions(-)
 create mode 100644 include/hw/net/lan9118_phy.h
 create mode 100644 hw/net/lan9118_phy.c

diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/include/hw/net/lan9118_phy.h
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This work is licensed under the terms of the GNU GPL, version 2 or later.
+ * See the COPYING file in the top-level directory.
+ */
+
+#ifndef HW_NET_LAN9118_PHY_H
+#define HW_NET_LAN9118_PHY_H
+
+#include "qom/object.h"
+#include "hw/sysbus.h"
+
+#define TYPE_LAN9118_PHY "lan9118-phy"
+OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
+
+typedef struct Lan9118PhyState {
+    SysBusDevice parent_obj;
+
+    uint16_t status;
+    uint16_t control;
+    uint16_t advertise;
+    uint16_t ints;
+    uint16_t int_mask;
+    qemu_irq irq;
+    bool link_down;
+} Lan9118PhyState;
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
+void lan9118_phy_reset(Lan9118PhyState *s);
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
+
+#endif
diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118.c
+++ b/hw/net/lan9118.c
@@ -XXX,XX +XXX,XX @@
 #include "net/net.h"
 #include "net/eth.h"
 #include "hw/irq.h"
+#include "hw/net/lan9118_phy.h"
 #include "hw/net/lan9118.h"
 #include "hw/ptimer.h"
 #include "hw/qdev-properties.h"
@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
 #define MAC_CR_RXEN     0x00000004
 #define MAC_CR_RESERVED 0x7f404213
 
-#define PHY_INT_ENERGYON            0x80
-#define PHY_INT_AUTONEG_COMPLETE    0x40
-#define PHY_INT_FAULT               0x20
-#define PHY_INT_DOWN                0x10
-#define PHY_INT_AUTONEG_LP          0x08
-#define PHY_INT_PARFAULT            0x04
-#define PHY_INT_AUTONEG_PAGE        0x02
-
 #define GPT_TIMER_EN    0x20000000
 
 /*
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
     uint32_t mac_mii_data;
     uint32_t mac_flow;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
 
     int32_t eeprom_writable;
     uint8_t eeprom[128];
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
 
 static const VMStateDescription vmstate_lan9118 = {
     .name = "lan9118",
-    .version_id = 2,
-    .minimum_version_id = 1,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_PTIMER(timer, lan9118_state),
         VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
         VMSTATE_UINT32(mac_mii_acc, lan9118_state),
         VMSTATE_UINT32(mac_mii_data, lan9118_state),
         VMSTATE_UINT32(mac_flow, lan9118_state),
-        VMSTATE_UINT32(phy_status, lan9118_state),
-        VMSTATE_UINT32(phy_control, lan9118_state),
-        VMSTATE_UINT32(phy_advertise, lan9118_state),
-        VMSTATE_UINT32(phy_int, lan9118_state),
-        VMSTATE_UINT32(phy_int_mask, lan9118_state),
         VMSTATE_INT32(eeprom_writable, lan9118_state),
         VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
         VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
     lan9118_mac_changed(s);
 }
 
-static void phy_update_irq(lan9118_state *s)
+static void lan9118_update_irq(void *opaque, int n, int level)
 {
-    if (s->phy_int & s->phy_int_mask) {
+    lan9118_state *s = opaque;
+
+    if (level) {
         s->int_sts |= PHY_INT;
     } else {
         s->int_sts &= ~PHY_INT;
@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
     lan9118_update(s);
 }
 
-static void phy_update_link(lan9118_state *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    phy_update_irq(s);
-}
-
 static void lan9118_set_link(NetClientState *nc)
 {
-    phy_update_link(qemu_get_nic_opaque(nc));
-}
-
-static void phy_reset(lan9118_state *s)
-{
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    phy_update_link(s);
+    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static void lan9118_reset(DeviceState *d)
@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
     s->read_word_n = 0;
     s->write_word_n = 0;
 
-    phy_reset(s);
-
     s->eeprom_writable = 0;
     lan9118_reload_eeprom(s);
 }
@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
     uint32_t status;
 
     /* FIXME: Honor TX disable, and allow queueing of packets.  */
-    if (s->phy_control & 0x4000)  {
+    if (s->mii.control & 0x4000) {
         /* This assumes the receive routine doesn't touch the VLANClient.  */
         qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
     } else {
@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
     }
 }
 
-static uint32_t do_phy_read(lan9118_state *s, int reg)
-{
-    uint32_t val;
-
-    switch (reg) {
-    case 0: /* Basic Control */
-        return s->phy_control;
-    case 1: /* Basic Status */
-        return s->phy_status;
-    case 2: /* ID1 */
-        return 0x0007;
-    case 3: /* ID2 */
-        return 0xc0d1;
-    case 4: /* Auto-neg advertisement */
-        return s->phy_advertise;
-    case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
-    case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
-    case 29: /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        phy_update_irq(s);
-        return val;
-    case 30: /* Interrupt mask */
-        return s->phy_int_mask;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_read: PHY read reg %d\n", reg);
-        return 0;
-    }
-}
-
-static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
-{
-    switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
-            phy_reset(s);
-            break;
-        }
-        s->phy_control = val & 0x7980;
-        /* Complete autonegotiation immediately.  */
-        if (val & 0x1000) {
-            s->phy_status |= 0x0020;
-        }
-        break;
-    case 4: /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-        /* TODO 17, 18, 27, 31 */
-    case 30: /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        phy_update_irq(s);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
-    }
-}
-
 static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
 {
     switch (reg) {
@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
         if (val & 2) {
             DPRINTF("PHY write %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
-            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
+            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
         } else {
-            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
+            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
             DPRINTF("PHY read %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
         }
@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
         break;
     case CSR_PMT_CTRL:
         if (val & 0x400) {
-            phy_reset(s);
+            lan9118_phy_reset(&s->mii);
         }
         s->pmt_ctrl &= ~0x34e;
         s->pmt_ctrl |= (val & 0x34e);
@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
     const MemoryRegionOps *mem_ops =
             s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
 
+    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
                           "lan9118-mmio", 0x100);
     sysbus_init_mmio(sbd, &s->mmio);
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This code is licensed under the GNU GPL v2
+ *
+ * Contributions after 2012-01-13 are licensed under the terms of the
+ * GNU GPL, version 2 or (at your option) any later version.
+ */
+
+#include "qemu/osdep.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
+#include "hw/resettable.h"
+#include "migration/vmstate.h"
+#include "qemu/log.h"
+
+#define PHY_INT_ENERGYON            (1 << 7)
+#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+#define PHY_INT_FAULT               (1 << 5)
+#define PHY_INT_DOWN                (1 << 4)
+#define PHY_INT_AUTONEG_LP          (1 << 3)
+#define PHY_INT_PARFAULT            (1 << 2)
+#define PHY_INT_AUTONEG_PAGE        (1 << 1)
+
+static void lan9118_phy_update_irq(Lan9118PhyState *s)
+{
+    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
+}
+
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+{
+    uint16_t val;
+
+    switch (reg) {
+    case 0: /* Basic Control */
+        return s->control;
+    case 1: /* Basic Status */
+        return s->status;
+    case 2: /* ID1 */
+        return 0x0007;
+    case 3: /* ID2 */
+        return 0xc0d1;
+    case 4: /* Auto-neg advertisement */
+        return s->advertise;
+    case 5: /* Auto-neg Link Partner Ability */
+        return 0x0f71;
+    case 6: /* Auto-neg Expansion */
+        return 1;
+        /* TODO 17, 18, 27, 29, 30, 31 */
+    case 29: /* Interrupt source. */
+        val = s->ints;
+        s->ints = 0;
+        lan9118_phy_update_irq(s);
+        return val;
+    case 30: /* Interrupt mask */
+        return s->int_mask;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_read: PHY read reg %d\n", reg);
+        return 0;
+    }
+}
+
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+{
+    switch (reg) {
+    case 0: /* Basic Control */
+        if (val & 0x8000) {
+            lan9118_phy_reset(s);
+            break;
+        }
+        s->control = val & 0x7980;
+        /* Complete autonegotiation immediately. */
+        if (val & 0x1000) {
+            s->status |= 0x0020;
+        }
+        break;
+    case 4: /* Auto-neg advertisement */
+        s->advertise = (val & 0x2d7f) | 0x80;
+        break;
+        /* TODO 17, 18, 27, 31 */
+    case 30: /* Interrupt mask */
+        s->int_mask = val & 0xff;
+        lan9118_phy_update_irq(s);
+        break;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+    }
+}
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+{
+    s->link_down = link_down;
+
+    /* Autonegotiation status mirrors link status. */
+    if (link_down) {
+        s->status &= ~0x0024;
+        s->ints |= PHY_INT_DOWN;
+    } else {
+        s->status |= 0x0024;
+        s->ints |= PHY_INT_ENERGYON;
+        s->ints |= PHY_INT_AUTONEG_COMPLETE;
+    }
+    lan9118_phy_update_irq(s);
+}
+
+void lan9118_phy_reset(Lan9118PhyState *s)
+{
+    s->control = 0x3000;
+    s->status = 0x7809;
+    s->advertise = 0x01e1;
+    s->int_mask = 0;
+    s->ints = 0;
+    lan9118_phy_update_link(s, s->link_down);
+}
+
+static void lan9118_phy_reset_hold(Object *obj, ResetType type)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    lan9118_phy_reset(s);
+}
+
+static void lan9118_phy_init(Object *obj)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
+}
+
+static const VMStateDescription vmstate_lan9118_phy = {
+    .name = "lan9118-phy",
+    .version_id = 1,
+    .minimum_version_id = 1,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT16(control, Lan9118PhyState),
+        VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(advertise, Lan9118PhyState),
+        VMSTATE_UINT16(ints, Lan9118PhyState),
+        VMSTATE_UINT16(int_mask, Lan9118PhyState),
+        VMSTATE_BOOL(link_down, Lan9118PhyState),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static void lan9118_phy_class_init(ObjectClass *klass, void *data)
+{
+    ResettableClass *rc = RESETTABLE_CLASS(klass);
+    DeviceClass *dc = DEVICE_CLASS(klass);
+
+    rc->phases.hold = lan9118_phy_reset_hold;
+    dc->vmsd = &vmstate_lan9118_phy;
+}
+
+static const TypeInfo types[] = {
+    {
+        .name          = TYPE_LAN9118_PHY,
+        .parent        = TYPE_SYS_BUS_DEVICE,
+        .instance_size = sizeof(Lan9118PhyState),
+        .instance_init = lan9118_phy_init,
+        .class_init    = lan9118_phy_class_init,
+    }
+};
+
+DEFINE_TYPES(types)
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
 config SMC91C111
     bool
 
+config LAN9118_PHY
+    bool
+
 config LAN9118
     bool
+    select LAN9118_PHY
     select PTIMER
 
 config NE2000_ISA
diff --git a/hw/net/meson.build b/hw/net/meson.build
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/meson.build
+++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
 
 system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
 system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
+system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
 system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
 system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
 system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

imx_fec models the same PHY as lan9118_phy. The code is almost the same with
imx_fec having more logging and tracing. Merge these improvements into
lan9118_phy and reuse in imx_fec to fix the code duplication.

Some migration state how resides in the new device model which breaks migration
compatibility for the following machines:
* imx25-pdk
* sabrelite
* mcimx7d-sabre
* mcimx6ul-evk

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-3-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/imx_fec.h |   9 ++-
 hw/net/imx_fec.c         | 146 ++++-----------------------------------
 hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
 hw/net/Kconfig           |   1 +
 hw/net/trace-events      |  10 +--
 5 files changed, 85 insertions(+), 163 deletions(-)

diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/imx_fec.h
+++ b/include/hw/net/imx_fec.h
@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
 #define TYPE_IMX_ENET "imx.enet"
 
 #include "hw/sysbus.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
 #include "net/net.h"
 
 #define ENET_EIR               1
@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
     uint32_t tx_descriptor[ENET_TX_RING_NUM];
     uint32_t tx_ring_num;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
     uint32_t phy_num;
     bool phy_connected;
     struct IMXFECState *phy_consumer;
diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/imx_fec.c
+++ b/hw/net/imx_fec.c
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
 
 static const VMStateDescription vmstate_imx_eth = {
     .name = TYPE_IMX_FEC,
-    .version_id = 2,
-    .minimum_version_id = 2,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
         VMSTATE_UINT32(rx_descriptor, IMXFECState),
         VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
-        VMSTATE_UINT32(phy_status, IMXFECState),
-        VMSTATE_UINT32(phy_control, IMXFECState),
-        VMSTATE_UINT32(phy_advertise, IMXFECState),
-        VMSTATE_UINT32(phy_int, IMXFECState),
-        VMSTATE_UINT32(phy_int_mask, IMXFECState),
         VMSTATE_END_OF_LIST()
     },
     .subsections = (const VMStateDescription * const []) {
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
     },
 };
 
-#define PHY_INT_ENERGYON            (1 << 7)
-#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
-#define PHY_INT_FAULT               (1 << 5)
-#define PHY_INT_DOWN                (1 << 4)
-#define PHY_INT_AUTONEG_LP          (1 << 3)
-#define PHY_INT_PARFAULT            (1 << 2)
-#define PHY_INT_AUTONEG_PAGE        (1 << 1)
-
 static void imx_eth_update(IMXFECState *s);
 
 /*
@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
  * For now we don't handle any GPIO/interrupt line, so the OS will
  * have to poll for the PHY status.
  */
-static void imx_phy_update_irq(IMXFECState *s)
+static void imx_phy_update_irq(void *opaque, int n, int level)
 {
-    imx_eth_update(s);
-}
-
-static void imx_phy_update_link(IMXFECState *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        trace_imx_phy_update_link("down");
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        trace_imx_phy_update_link("up");
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    imx_phy_update_irq(s);
+    imx_eth_update(opaque);
 }
 
 static void imx_eth_set_link(NetClientState *nc)
 {
-    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
-}
-
-static void imx_phy_reset(IMXFECState *s)
-{
-    trace_imx_phy_reset();
-
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    imx_phy_update_link(s);
+    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static uint32_t imx_phy_read(IMXFECState *s, int reg)
 {
-    uint32_t val;
     uint32_t phy = reg / 32;
 
     if (!s->phy_connected) {
@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
 
     reg %= 32;
 
-    switch (reg) {
-    case 0:     /* Basic Control */
-        val = s->phy_control;
-        break;
-    case 1:     /* Basic Status */
-        val = s->phy_status;
-        break;
-    case 2:     /* ID1 */
-        val = 0x0007;
-        break;
-    case 3:     /* ID2 */
-        val = 0xc0d1;
-        break;
-    case 4:     /* Auto-neg advertisement */
-        val = s->phy_advertise;
-        break;
-    case 5:     /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
-        break;
-    case 6:     /* Auto-neg Expansion */
-        val = 1;
-        break;
-    case 29:    /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        imx_phy_update_irq(s);
-        break;
-    case 30:    /* Interrupt mask */
-        val = s->phy_int_mask;
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    }
-
-    trace_imx_phy_read(val, phy, reg);
-
-    return val;
+    return lan9118_phy_read(&s->mii, reg);
 }
 
 static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
 
     reg %= 32;
 
-    trace_imx_phy_write(val, phy, reg);
-
-    switch (reg) {
-    case 0:     /* Basic Control */
-        if (val & 0x8000) {
-            imx_phy_reset(s);
-        } else {
-            s->phy_control = val & 0x7980;
-            /* Complete autonegotiation immediately.  */
-            if (val & 0x1000) {
-                s->phy_status |= 0x0020;
-            }
-        }
-        break;
-    case 4:     /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-    case 30:    /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        imx_phy_update_irq(s);
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    }
+    lan9118_phy_write(&s->mii, reg, val);
 }
 
 static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
 
     s->rx_descriptor = 0;
     memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
-
-    /* We also reset the PHY */
-    imx_phy_reset(s);
 }
 
 static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
     sysbus_init_irq(sbd, &s->irq[0]);
     sysbus_init_irq(sbd, &s->irq[1]);
 
+    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     qemu_macaddr_default_if_unset(&s->conf.macaddr);
 
     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
  * Copyright (c) 2009 CodeSourcery, LLC.
  * Written by Paul Brook
  *
+ * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
+ *
  * This code is licensed under the GNU GPL v2
  *
  * Contributions after 2012-01-13 are licensed under the terms of the
@@ -XXX,XX +XXX,XX @@
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
 #include "qemu/log.h"
+#include "trace.h"
 
 #define PHY_INT_ENERGYON            (1 << 7)
 #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
 
     switch (reg) {
     case 0: /* Basic Control */
-        return s->control;
+        val = s->control;
+        break;
     case 1: /* Basic Status */
-        return s->status;
+        val = s->status;
+        break;
     case 2: /* ID1 */
-        return 0x0007;
+        val = 0x0007;
+        break;
     case 3: /* ID2 */
-        return 0xc0d1;
+        val = 0xc0d1;
+        break;
     case 4: /* Auto-neg advertisement */
-        return s->advertise;
+        val = s->advertise;
+        break;
     case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
+        val = 0x0f71;
+        break;
     case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
+        val = 1;
+        break;
     case 29: /* Interrupt source. */
         val = s->ints;
         s->ints = 0;
         lan9118_phy_update_irq(s);
-        return val;
+        break;
     case 30: /* Interrupt mask */
-        return s->int_mask;
+        val = s->int_mask;
+        break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        val = 0;
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_read: PHY read reg %d\n", reg);
-        return 0;
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        val = 0;
+        break;
     }
+
+    trace_lan9118_phy_read(val, reg);
+
+    return val;
 }
 
 void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
 {
+    trace_lan9118_phy_write(val, reg);
+
     switch (reg) {
     case 0: /* Basic Control */
         if (val & 0x8000) {
             lan9118_phy_reset(s);
-            break;
-        }
-        s->control = val & 0x7980;
-        /* Complete autonegotiation immediately. */
-        if (val & 0x1000) {
-            s->status |= 0x0020;
+        } else {
+            s->control = val & 0x7980;
+            /* Complete autonegotiation immediately. */
+            if (val & 0x1000) {
+                s->status |= 0x0020;
+            }
         }
         break;
     case 4: /* Auto-neg advertisement */
         s->advertise = (val & 0x2d7f) | 0x80;
         break;
-        /* TODO 17, 18, 27, 31 */
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
         lan9118_phy_update_irq(s);
         break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        break;
     }
 }
 
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
+        trace_lan9118_phy_update_link("down");
         s->status &= ~0x0024;
         s->ints |= PHY_INT_DOWN;
     } else {
+        trace_lan9118_phy_update_link("up");
         s->status |= 0x0024;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
 void lan9118_phy_reset(Lan9118PhyState *s)
 {
+    trace_lan9118_phy_reset();
+
     s->control = 0x3000;
     s->status = 0x7809;
     s->advertise = 0x01e1;
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
     .version_id = 1,
     .minimum_version_id = 1,
     .fields = (const VMStateField[]) {
-        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(advertise, Lan9118PhyState),
         VMSTATE_UINT16(ints, Lan9118PhyState),
         VMSTATE_UINT16(int_mask, Lan9118PhyState),
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
 
 config IMX_FEC
     bool
+    select LAN9118_PHY
 
 config CADENCE
     bool
diff --git a/hw/net/trace-events b/hw/net/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
 allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
 allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
 
+# lan9118_phy.c
+lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
+lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
+lan9118_phy_update_link(const char *s) "%s"
+lan9118_phy_reset(void) ""
+
 # lance.c
 lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
 lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
 i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
 
 # imx_fec.c
-imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
 imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
-imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
 imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
-imx_phy_update_link(const char *s) "%s"
-imx_phy_reset(void) ""
 imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
 imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
 imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
fixes the MSB of selector field to be zero, as specified in the datasheet.

Fixes: 2a424990170b "LAN9118 emulation"
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-4-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
         val = s->advertise;
         break;
     case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
+        val = 0x0fe1;
         break;
     case 6: /* Auto-neg Expansion */
         val = 1;
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Prefer named constants over magic values for better readability.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-5-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/mii.h |  6 +++++
 hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
 2 files changed, 46 insertions(+), 23 deletions(-)

diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/mii.h
+++ b/include/hw/net/mii.h
@@ -XXX,XX +XXX,XX @@
 #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
 #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
 
+#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
 #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
 #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
 #define MII_ANAR_TXFD       (1 << 8)
@@ -XXX,XX +XXX,XX @@
 #define MII_ANAR_10FD       (1 << 6)
 #define MII_ANAR_10         (1 << 5)
 #define MII_ANAR_CSMACD     (1 << 0)
+#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
 
 #define MII_ANLPAR_ACK      (1 << 14)
 #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
@@ -XXX,XX +XXX,XX @@
 #define RTL8201CP_PHYID1    0x0000
 #define RTL8201CP_PHYID2    0x8201
 
+/* SMSC LAN9118 */
+#define SMSCLAN9118_PHYID1  0x0007
+#define SMSCLAN9118_PHYID2  0xc0d1
+
 /* RealTek 8211E */
 #define RTL8211E_PHYID1     0x001c
 #define RTL8211E_PHYID2     0xc915
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
 
 #include "qemu/osdep.h"
 #include "hw/net/lan9118_phy.h"
+#include "hw/net/mii.h"
 #include "hw/irq.h"
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
     uint16_t val;
 
     switch (reg) {
-    case 0: /* Basic Control */
+    case MII_BMCR:
         val = s->control;
         break;
-    case 1: /* Basic Status */
+    case MII_BMSR:
         val = s->status;
         break;
-    case 2: /* ID1 */
-        val = 0x0007;
+    case MII_PHYID1:
+        val = SMSCLAN9118_PHYID1;
         break;
-    case 3: /* ID2 */
-        val = 0xc0d1;
+    case MII_PHYID2:
+        val = SMSCLAN9118_PHYID2;
         break;
-    case 4: /* Auto-neg advertisement */
+    case MII_ANAR:
         val = s->advertise;
         break;
-    case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0fe1;
+    case MII_ANLPAR:
+        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
+              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
+              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
         break;
-    case 6: /* Auto-neg Expansion */
-        val = 1;
+    case MII_ANER:
+        val = MII_ANER_NWAY;
         break;
     case 29: /* Interrupt source. */
         val = s->ints;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
     trace_lan9118_phy_write(val, reg);
 
     switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
+    case MII_BMCR:
+        if (val & MII_BMCR_RESET) {
             lan9118_phy_reset(s);
         } else {
-            s->control = val & 0x7980;
+            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
+                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
+                                MII_BMCR_CTST);
             /* Complete autonegotiation immediately. */
-            if (val & 0x1000) {
-                s->status |= 0x0020;
+            if (val & MII_BMCR_AUTOEN) {
+                s->status |= MII_BMSR_AN_COMP;
             }
         }
         break;
-    case 4: /* Auto-neg advertisement */
-        s->advertise = (val & 0x2d7f) | 0x80;
+    case MII_ANAR:
+        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
+                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
+                               MII_ANAR_SELECT))
+                     | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
         trace_lan9118_phy_update_link("down");
-        s->status &= ~0x0024;
+        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
         s->ints |= PHY_INT_DOWN;
     } else {
         trace_lan9118_phy_update_link("up");
-        s->status |= 0x0024;
+        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
     }
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
 {
     trace_lan9118_phy_reset();
 
-    s->control = 0x3000;
-    s->status = 0x7809;
-    s->advertise = 0x01e1;
+    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
+    s->status = MII_BMSR_100TX_FD
+                | MII_BMSR_100TX_HD
+                | MII_BMSR_10T_FD
+                | MII_BMSR_10T_HD
+                | MII_BMSR_AUTONEG
+                | MII_BMSR_EXTCAP;
+    s->advertise = MII_ANAR_TXFD
+                   | MII_ANAR_TX
+                   | MII_ANAR_10FD
+                   | MII_ANAR_10
+                   | MII_ANAR_CSMACD;
     s->int_mask = 0;
     s->ints = 0;
     lan9118_phy_update_link(s, s->link_down);
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

The real device advertises this mode and the device model already advertises
100 mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
make the model more realistic.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-6-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
         break;
     case MII_ANAR:
         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
-                               MII_ANAR_SELECT))
+                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
+                               MII_ANAR_10 | MII_ANAR_SELECT))
                      | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
-- 
2.34.1

For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
Invalid for the multiplication of 0 by infinity.  Currently we handle
this in the per-architecture ifdef ladder in pickNaNMulAdd().
However, since this isn't really architecture specific we can hoist
it up to the generic code.

For the cases where the infzero test in pickNaNMulAdd was
returning 2, we can delete the check entirely and allow the
code to fall into the normal pick-a-NaN handling, because this
will return 2 anyway (input 'c' being the only NaN in this case).
For the cases where infzero was returning 3 to indicate "return
the default NaN", we must retain that "return 3".

For Arm, this looks like it might be a behaviour change because we
used to set float_flag_invalid | float_flag_invalid_imz only if C is
a quiet NaN.  However, it is not, because Arm target code never looks
at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
already raised float_flag_invalid via the "abc_mask &
float_cmask_snan" check in pick_nan_muladd.

For any target architecture using the "default implementation" at the
bottom of the ifdef, this is a behaviour change but will be fixing a
bug (where we failed to raise the Invalid exception for (0 * inf +
QNaN).  The architectures using the default case are:
 * hppa
 * i386
 * sh4
 * tricore

The x86, Tricore and SH4 CPU architecture manuals are clear that this
should have raised Invalid; HPPA is a bit vaguer but still seems
clear enough.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 13 +++++++------
 fpu/softfloat-specialize.c.inc | 29 +----------------------------
 2 files changed, 8 insertions(+), 34 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             int ab_mask, int abc_mask)
 {
     int which;
+    bool infzero = (ab_mask == float_cmask_infzero);
 
     if (unlikely(abc_mask & float_cmask_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
-                          ab_mask == float_cmask_infzero, s);
+    if (infzero) {
+        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
+        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+    }
+
+    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
 
     if (s->default_nan_mode || which == 3) {
-        /*
-         * Note that this check is after pickNaNMulAdd so that function
-         * has an opportunity to set the Invalid flag for infzero.
-         */
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * the default NaN
      */
     if (infzero && is_qnan(c_cls)) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * case sets InvalidOp and returns the default NaN
          */
         if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
             return 3;
         }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-            return 2;
-        }
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
      * case sets InvalidOp and returns the input value 'c'
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
+
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * to return an input NaN if we have one (ie c) rather than generating
      * a default NaN
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
 
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_RISCV)
-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-    }
     return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 2;
     }
 #elif defined(TARGET_SPARC)
-    /* For (inf,0,nan) return c. */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     /* Prefer SNaN over QNaN, order C, B, A. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
      * an input NaN if we have one (ie c).
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     if (status->use_first_nan) {
         if (is_nan(a_cls)) {
             return 0;
-- 
2.34.1

If the target sets default_nan_mode then we're always going to return
the default NaN, and pickNaNMulAdd() no longer has any side effects.
For consistency with pickNaN(), check for default_nan_mode before
calling pickNaNMulAdd().

When we convert pickNaNMulAdd() to allow runtime selection of the NaN
propagation rule, this means we won't have to make the targets which
use default_nan_mode also set a propagation rule.

Since RiscV always uses default_nan_mode, this allows us to remove
its ifdef case from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 8 ++++++--
 fpu/softfloat-specialize.c.inc | 9 +++++++--
 2 files changed, 13 insertions(+), 4 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    if (s->default_nan_mode) {
+        which = 3;
+    } else {
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    }
 
-    if (s->default_nan_mode || which == 3) {
+    if (which == 3) {
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    /*
+     * We guarantee not to require the target to tell us how to
+     * pick a NaN if we're always returning the default NaN.
+     * But if we're not in default-NaN mode then the target must
+     * specify.
+     */
+    assert(!status->default_nan_mode);
 #if defined(TARGET_ARM)
     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
      * the default NaN
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     } else {
         return 1;
     }
-#elif defined(TARGET_RISCV)
-    return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
         return 3;
-- 
2.34.1

IEEE 758 does not define a fixed rule for what NaN to return in
the case of a fused multiply-add of inf * 0 + NaN. Different
architectures thus do different things:
 * some return the default NaN
 * some return the input NaN
 * Arm returns the default NaN if the input NaN is quiet,
   and the input NaN if it is signalling

We want to make this logic be runtime selected rather than
hardcoded into the binary, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

In this commit we add an enum for the propagation rule, the field in
float_status, and the corresponding getters and setters.  We change
pickNaNMulAdd to honour this, but because all targets still leave
this field at its default 0 value, the fallback logic will pick the
rule type with the old ifdef ladder.

Note that four architectures both use the muladd softfloat functions
and did not have a branch of the ifdef ladder to specify their
behaviour (and so were ending up with the "default" case, probably
wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
default_nan_mode, and so will never get into pickNaNMulAdd().  For
HPPA and i386 we retain the same behaviour as the old default-case,
which is to not ever return the default NaN.  This might not be
correct but it is not a behaviour change.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 ++++
 include/fpu/softfloat-types.h   | 23 +++++++++
 fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
 3 files changed, 95 insertions(+), 30 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
+                                             float_status *status)
+{
+    status->float_infzeronan_rule = rule;
+}
+
 static inline void set_flush_to_zero(bool val, float_status *status)
 {
     status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
+{
+    return status->float_infzeronan_rule;
+}
+
 static inline bool get_flush_to_zero(float_status *status)
 {
     return status->flush_to_zero;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * Rule for result of fused multiply-add 0 * Inf + NaN.
+ * This must be a NaN, but implementations differ on whether this
+ * is the input NaN or the default NaN.
+ *
+ * You don't need to set this if default_nan_mode is enabled.
+ * When not in default-NaN mode, it is an error for the target
+ * not to set the rule in float_status if it uses muladd, and we
+ * will assert if we need to handle an input NaN and no rule was
+ * selected.
+ */
+typedef enum __attribute__((__packed__)) {
+    /* No propagation rule specified */
+    float_infzeronan_none = 0,
+    /* Result is never the default NaN (so always the input NaN) */
+    float_infzeronan_dnan_never,
+    /* Result is always the default NaN */
+    float_infzeronan_dnan_always,
+    /* Result is the default NaN if the input NaN is quiet */
+    float_infzeronan_dnan_if_qnan,
+} FloatInfZeroNaNRule;
+
 /*
  * Floating Point Status. Individual architectures may maintain
  * several versions of float_status for different functions. The
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
     bool flush_to_zero;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * specify.
      */
     assert(!status->default_nan_mode);
+
+    if (rule == float_infzeronan_none) {
+        /*
+         * Temporarily fall back to ifdef ladder
+         */
 #if defined(TARGET_ARM)
-    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
-     * the default NaN
-     */
-    if (infzero && is_qnan(c_cls)) {
-        return 3;
+        /*
+         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+         * but (inf,zero,snan) returns the input NaN.
+         */
+        rule = float_infzeronan_dnan_if_qnan;
+#elif defined(TARGET_MIPS)
+        if (snan_bit_is_one(status)) {
+            /*
+             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+             * case sets InvalidOp and returns the default NaN
+             */
+            rule = float_infzeronan_dnan_always;
+        } else {
+            /*
+             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+             * case sets InvalidOp and returns the input value 'c'
+             */
+            rule = float_infzeronan_dnan_never;
+        }
+#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+        /*
+         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+         * case sets InvalidOp and returns the input value 'c'
+         */
+        /*
+         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+         * to return an input NaN if we have one (ie c) rather than generating
+         * a default NaN
+         */
+        rule = float_infzeronan_dnan_never;
+#elif defined(TARGET_S390X)
+        rule = float_infzeronan_dnan_always;
+#endif
     }
 
+    if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the default NaN here,
+         * and some return the input NaN.
+         */
+        switch (rule) {
+        case float_infzeronan_dnan_never:
+            return 2;
+        case float_infzeronan_dnan_always:
+            return 3;
+        case float_infzeronan_dnan_if_qnan:
+            return is_qnan(c_cls) ? 3 : 2;
+        default:
+            g_assert_not_reached();
+        }
+    }
+
+#if defined(TARGET_ARM)
+
     /* This looks different from the ARM ARM pseudocode, because the ARM ARM
      * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 #elif defined(TARGET_MIPS)
     if (snan_bit_is_one(status)) {
-        /*
-         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-         * case sets InvalidOp and returns the default NaN
-         */
-        if (infzero) {
-            return 3;
-        }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
         if (is_snan(a_cls)) {
             return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
             return 2;
         }
     } else {
-        /*
-         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 #elif defined(TARGET_LOONGARCH64)
-    /*
-     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-     * case sets InvalidOp and returns the input value 'c'
-     */
-
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_PPC)
-    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-     * to return an input NaN if we have one (ie c) rather than generating
-     * a default NaN
-     */
-
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_S390X)
-    if (infzero) {
-        return 3;
-    }
-
     if (is_snan(a_cls)) {
         return 0;
     } else if (is_snan(b_cls)) {
-- 
2.34.1

Explicitly set a rule in the softfloat tests for the inf-zero-nan
muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
and so we should select here the Arm rule of
float_infzeronan_dnan_if_qnan.

Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 5 +++++
 tests/fp/fp-test.c  | 5 +++++
 2 files changed, 10 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
 {
     bench_func_t f;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
 {
     unsigned int i;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
     verCases_maxErrorCount = n_max_errors;
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the Arm target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 4 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+ *    and the input NaN if it is signalling
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_ARM)
-        /*
-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
-         * but (inf,zero,snan) returns the input NaN.
-         */
-        rule = float_infzeronan_dnan_if_qnan;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             /*
              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for s390, so we
can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                  &env->fpu_status);
        /* fall through */
     case RESET_TYPE_S390_CPU_NORMAL:
         env->psw.mask &= ~PSW_MASK_RI;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * a default NaN
          */
         rule = float_infzeronan_dnan_never;
-#elif defined(TARGET_S390X)
-        rule = float_infzeronan_dnan_always;
 #endif
     }
 
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the PPC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 7 +++++++
 fpu/softfloat-specialize.c.inc | 7 +------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+     * to return an input NaN if we have one (ie c) rather than generating
+     * a default NaN
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
 
     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
         ppc_spr_t *spr = &env->spr_cb[i];
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              */
             rule = float_infzeronan_dnan_never;
         }
-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+#elif defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        /*
-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-         * to return an input NaN if we have one (ie c) rather than generating
-         * a default NaN
-         */
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the MIPS target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       |  9 +++++++++
 target/mips/msa.c              |  4 ++++
 fpu/softfloat-specialize.c.inc | 16 +---------------
 3 files changed, 14 insertions(+), 15 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
 static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
+    FloatInfZeroNaNRule izn_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
+    /*
+     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+     * case sets InvalidOp and returns the default NaN.
+     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'.
+     */
+    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
 
     /* set proper signanling bit meaning ("1" means "quiet") */
     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
+
+    /* Inf * 0 + NaN returns the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                              &env->active_tc.msa_fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            /*
-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-             * case sets InvalidOp and returns the default NaN
-             */
-            rule = float_infzeronan_dnan_always;
-        } else {
-            /*
-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-             * case sets InvalidOp and returns the input value 'c'
-             */
-            rule = float_infzeronan_dnan_never;
-        }
-#elif defined(TARGET_SPARC) || \
+#if defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the SPARC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_SPARC) || \
-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the xtensa target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
---
 target/xtensa/cpu.c            | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/cpu.c
+++ b/target/xtensa/cpu.c
@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
     reset_mmu(env);
     cs->halted = env->runstall;
 #endif
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
     set_no_signaling_nans(!dfpu, &env->fp_status);
     xtensa_use_first_nan(env, !dfpu);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the x86 target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 7 +++++++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
+    /*
+     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
+     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
+     * specified -- for 0 * inf + NaN the input NaN is selected, and if
+     * there are multiple input NaNs they are selected in the order a, b, c.
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * Temporarily fall back to ifdef ladder
          */
 #if defined(TARGET_HPPA) || \
-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+    defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the loongarch target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 5 +++++
 fpu/softfloat-specialize.c.inc    | 7 +------
 2 files changed, 6 insertions(+), 6 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
                             &env->fp_status);
     set_flush_to_zero(0, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_HPPA) || \
-    defined(TARGET_LOONGARCH)
-        /*
-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
+#if defined(TARGET_HPPA)
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the HPPA target,
so we can remove the ifdef from pickNaNMulAdd().

As this is the last target to be converted to explicitly setting
the rule, we can remove the fallback code in pickNaNMulAdd()
entirely.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       |  2 ++
 fpu/softfloat-specialize.c.inc | 13 +------------
 2 files changed, 3 insertions(+), 12 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
-
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      */
     assert(!status->default_nan_mode);
 
-    if (rule == float_infzeronan_none) {
-        /*
-         * Temporarily fall back to ifdef ladder
-         */
-#if defined(TARGET_HPPA)
-        rule = float_infzeronan_dnan_never;
-#endif
-    }
-
     if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the default NaN here,
          * and some return the input NaN.
          */
-        switch (rule) {
+        switch (status->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
             return 2;
         case float_infzeronan_dnan_always:
-- 
2.34.1

The new implementation of pickNaNMulAdd() will find it convenient
to know whether at least one of the three arguments to the muladd
was a signaling NaN. We already calculate that in the caller,
so pass it in as a new bool have_snan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 5 +++--
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
 {
     int which;
     bool infzero = (ab_mask == float_cmask_infzero);
+    bool have_snan = (abc_mask & float_cmask_snan);
 
-    if (unlikely(abc_mask & float_cmask_snan)) {
+    if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     if (s->default_nan_mode) {
         which = 3;
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
 *----------------------------------------------------------------------------*/
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, float_status *status)
+                         bool infzero, bool have_snan, float_status *status)
 {
     /*
      * We guarantee not to require the target to tell us how to
-- 
2.34.1

IEEE 758 does not define a fixed rule for which NaN to pick as the
result if both operands of a 3-operand fused multiply-add operation
are NaNs.  As a result different architectures have ended up with
different rules for propagating NaNs.

QEMU currently hardcodes the NaN propagation logic into the binary
because pickNaNMulAdd() has an ifdef ladder for different targets.
We want to make the propagation rule instead be selectable at
runtime, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

It's valid not to set a propagation rule if default_nan_mode is
enabled, because in that case there's no need to pick a NaN; all the
callers of pickNaNMulAdd() catch this case and skip calling it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h |  11 +++
 include/fpu/softfloat-types.h   |  55 +++++++++++
 fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
 3 files changed, 107 insertions(+), 126 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
+                                            float_status *status)
+{
+    status->float_3nan_prop_rule = rule;
+}
+
 static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
                                              float_status *status)
 {
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
+{
+    return status->float_3nan_prop_rule;
+}
+
 static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 {
     return status->float_infzeronan_rule;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ this code that are retained.
 #ifndef SOFTFLOAT_TYPES_H
 #define SOFTFLOAT_TYPES_H
 
+#include "hw/registerfields.h"
+
 /*
  * Software IEC/IEEE floating-point types.
  */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * 3-input NaN propagation rule, for fused multiply-add. Individual
+ * architectures have different rules for which input NaN is
+ * propagated to the output when there is more than one NaN on the
+ * input.
+ *
+ * If default_nan_mode is enabled then it is valid not to set a NaN
+ * propagation rule, because the softfloat code guarantees not to try
+ * to pick a NaN to propagate in default NaN mode.  When not in
+ * default-NaN mode, it is an error for the target not to set the rule
+ * in float_status if it uses a muladd, and we will assert if we need
+ * to handle an input NaN and no rule was selected.
+ *
+ * The naming scheme for Float3NaNPropRule values is:
+ *  float_3nan_prop_s_abc:
+ *    = "Prefer SNaN over QNaN, then operand A over B over C"
+ *  float_3nan_prop_abc:
+ *    = "Prefer A over B over C regardless of SNaN vs QNAN"
+ *
+ * For QEMU, the multiply-add operation is A * B + C.
+ */
+
+/*
+ * We set the Float3NaNPropRule enum values up so we can select the
+ * right value in pickNaNMulAdd in a data driven way.
+ */
+FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
+FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
+FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
+FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
+
+#define PROPRULE(X, Y, Z) \
+    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
+
+typedef enum __attribute__((__packed__)) {
+    float_3nan_prop_none = 0,     /* No propagation rule specified */
+    float_3nan_prop_abc = PROPRULE(0, 1, 2),
+    float_3nan_prop_acb = PROPRULE(0, 2, 1),
+    float_3nan_prop_bac = PROPRULE(1, 0, 2),
+    float_3nan_prop_bca = PROPRULE(1, 2, 0),
+    float_3nan_prop_cab = PROPRULE(2, 0, 1),
+    float_3nan_prop_cba = PROPRULE(2, 1, 0),
+    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
+} Float3NaNPropRule;
+
+#undef PROPRULE
+
 /*
  * Rule for result of fused multiply-add 0 * Inf + NaN.
  * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    Float3NaNPropRule float_3nan_prop_rule;
     FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, bool have_snan, float_status *status)
 {
+    FloatClass cls[3] = { a_cls, b_cls, c_cls };
+    Float3NaNPropRule rule = status->float_3nan_prop_rule;
+    int which;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
+    if (rule == float_3nan_prop_none) {
 #if defined(TARGET_ARM)
-
-    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
-     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
-     */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        /*
+         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+         */
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_MIPS)
-    if (snan_bit_is_one(status)) {
-        /* Prefer sNaN over qNaN, in the a, b, c order. */
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
+        if (snan_bit_is_one(status)) {
+            rule = float_3nan_prop_s_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_s_cab;
         }
-    } else {
-        /* Prefer sNaN over qNaN, in the c, a, b order. */
-        if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-    }
 #elif defined(TARGET_LOONGARCH64)
-    /* Prefer sNaN over qNaN, in the c, a, b order. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
-    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(c_cls)) {
-        return 2;
-    } else {
-        return 1;
-    }
+        /*
+         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+         */
+        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
-    if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-    /* Prefer SNaN over QNaN, order C, B, A. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 0;
-    }
+        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-    /*
-     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
-     * an input NaN if we have one (ie c).
-     */
-    if (status->use_first_nan) {
-        if (is_nan(a_cls)) {
-            return 0;
-        } else if (is_nan(b_cls)) {
-            return 1;
+        if (status->use_first_nan) {
+            rule = float_3nan_prop_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_cba;
         }
-    } else {
-        if (is_nan(c_cls)) {
-            return 2;
-        } else if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-    }
 #else
-    /* A default implementation: prefer a to b to c.
-     * This is unlikely to actually match any real implementation.
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_abc;
 #endif
+    }
+
+    assert(rule != float_3nan_prop_none);
+    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+        /* We have at least one SNaN input and should prefer it */
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_snan(cls[which]));
+    } else {
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_nan(cls[which]));
+    }
+    return which;
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

Explicitly set a rule in the softfloat tests for propagating NaNs in
the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
so we should select here the Arm rule of float_3nan_prop_s_cab.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 1 +
 tests/fp/fp-test.c  | 1 +
 2 files changed, 2 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 5 +++++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 6 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
+ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
+ *    but note that for QEMU muladd is a * b + c, whereas for
+ *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
  */
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_ARM)
-        /*
-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
-         */
-        rule = float_3nan_prop_s_cab;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             rule = float_3nan_prop_s_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for loongarch, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 1 +
 fpu/softfloat-specialize.c.inc    | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
      * case sets InvalidOp and returns the input value 'c'
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_LOONGARCH64)
-        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
         /*
          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for PPC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 6 ------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * NaN propagation for fused multiply-add:
+     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
+     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+     * whereas QEMU labels the operands as (a * b) + c.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
     /*
      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
      * to return an input NaN if we have one (ie c) rather than generating
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_PPC)
-        /*
-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-         */
-        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
         rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for s390x, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 1 +
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
         set_float_infzeronan_rule(float_infzeronan_dnan_always,
                                   &env->fpu_status);
        /* fall through */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_S390X)
-        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
         rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for SPARC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
+    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_SPARC)
-        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       | 4 ++++
 target/mips/msa.c              | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 3 files changed, 8 insertions(+), 7 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
     FloatInfZeroNaNRule izn_rule;
+    Float3NaNPropRule nan3_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab,
                              &env->active_tc.msa_fp_status);
 
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
+                             &env->active_tc.msa_fp_status);
+
     /* clear float_status exception flags */
     set_float_exception_flags(0, &env->active_tc.msa_fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            rule = float_3nan_prop_s_abc;
-        } else {
-            rule = float_3nan_prop_s_cab;
-        }
-#elif defined(TARGET_XTENSA)
+#if defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for xtensa, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
---
 target/xtensa/fpu_helper.c     | 2 ++
 fpu/softfloat-specialize.c.inc | 8 --------
 2 files changed, 2 insertions(+), 8 deletions(-)

diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/fpu_helper.c
+++ b/target/xtensa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
     set_use_first_nan(use_first, &env->fp_status);
     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                              &env->fp_status);
+    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
+                             &env->fp_status);
 }
 
 void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_XTENSA)
-        if (status->use_first_nan) {
-            rule = float_3nan_prop_abc;
-        } else {
-            rule = float_3nan_prop_cba;
-        }
-#else
         rule = float_3nan_prop_abc;
-#endif
     }
 
     assert(rule != float_3nan_prop_none);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for i386.  We had no
i386-specific behaviour in the old ifdef ladder, so we were using the
default "prefer a then b then c" fallback; this is actually the
correct per-the-spec handling for i386.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c | 1 +
 1 file changed, 1 insertion(+)

Set the Float3NaNPropRule explicitly for HPPA, and remove the
ifdef from pickNaNMulAdd().

HPPA is the only target that was using the default branch of the
ifdef ladder (other targets either do not use muladd or set
default_nan_mode), so we can remove the ifdef fallback entirely now
(allowing the "rule not set" case to fall into the default of the
switch statement and assert).

We add a TODO note that the HPPA rule is probably wrong; this is
not a behavioural change for this refactoring.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 4 ----
 2 files changed, 8 insertions(+), 4 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * TODO: The HPPA architecture reference only documents its NaN
+     * propagation rule for 2-operand operations. Testing on real hardware
+     * might be necessary to confirm whether this order for muladd is correct.
+     * Not preferring the SNaN is almost certainly incorrect as it diverges
+     * from the documented rules for 2-operand operations.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
-    if (rule == float_3nan_prop_none) {
-        rule = float_3nan_prop_abc;
-    }
-
     assert(rule != float_3nan_prop_none);
     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
         /* We have at least one SNaN input and should prefer it */
-- 
2.34.1

The use_first_nan field in float_status was an xtensa-specific way to
select at runtime from two different NaN propagation rules.  Now that
xtensa is using the target-agnostic NaN propagation rule selection
that we've just added, we can remove use_first_nan, because there is
no longer any code that reads it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 5 -----
 include/fpu/softfloat-types.h   | 1 -
 target/xtensa/fpu_helper.c      | 1 -
 3 files changed, 7 deletions(-)

Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
to get the NaN bit pattern to reset the FPU registers. This
works because it happens that our implementation of
floatx80_default_nan() doesn't actually look at the float_status
pointer except for TARGET_MIPS. However, this isn't guaranteed,
and to be able to remove the ifdef in floatx80_default_nan()
we're going to need a real float_status here.

Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
earlier, and thus can pass it to floatx80_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
---
 target/m68k/cpu.c | 12 +++++++-----
 1 file changed, 7 insertions(+), 5 deletions(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
     CPUState *cs = CPU(obj);
     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
     CPUM68KState *env = cpu_env(cs);
-    floatx80 nan = floatx80_default_nan(NULL);
+    floatx80 nan;
     int i;
 
     if (mcc->parent_phases.hold) {
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
 #else
     cpu_m68k_set_sr(env, SR_S | SR_I);
 #endif
-    for (i = 0; i < 8; i++) {
-        env->fregs[i].d = nan;
-    }
-    cpu_m68k_set_fpcr(env, 0);
     /*
      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+
+    nan = floatx80_default_nan(&env->fp_status);
+    for (i = 0; i < 8; i++) {
+        env->fregs[i].d = nan;
+    }
+    cpu_m68k_set_fpcr(env, 0);
     env->fpsr = 0;
 
     /* TODO: We should set PC from the interrupt vector.  */
-- 
2.34.1

We create our 128-bit default NaN by calling parts64_default_nan()
and then adjusting the result.  We can do the same trick for creating
the floatx80 default NaN, which lets us drop a target ifdef.

floatx80 is used only by:
 i386
 m68k
 arm nwfpe old floating-point emulation emulation support
    (which is essentially dead, especially the parts involving floatx80)
 PPC (only in the xsrqpxp instruction, which just rounds an input
    value by converting to floatx80 and back, so will never generate
    the default NaN)

The floatx80 default NaN as currently implemented is:
 m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
 i386: sign = 1, exp = 1...1, int = 1, frac = 10...0

These are the same as the parts64_default_nan for these architectures.

This is technically a possible behaviour change for arm linux-user
nwfpe emulation emulation, because the default NaN will now have the
sign bit clear.  But we were already generating a different floatx80
default NaN from the real kernel emulation we are supposedly
following, which appears to use an all-bits-1 value:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267

This won't affect the only "real" use of the nwfpe emulation, which
is ancient binaries that used it as part of the old floating point
calling convention; that only uses loads and stores of 32 and 64 bit
floats, not any of the floatx80 behaviour the original hardware had.
We also get the nwfpe float64 default NaN value wrong:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
so if we ever cared about this obscure corner the right fix would be
to correct that so nwfpe used its own default-NaN setting rather
than the Arm VFP one.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
 floatx80 floatx80_default_nan(float_status *status)
 {
     floatx80 r;
+    /*
+     * Extrapolate from the choices made by parts64_default_nan to fill
+     * in the floatx80 format. We assume that floatx80's explicit
+     * integer bit is always set (this is true for i386 and m68k,
+     * which are the only real users of this format).
+     */
+    FloatParts64 p64;
+    parts64_default_nan(&p64, status);
 
-    /* None of the targets that have snan_bit_is_one use floatx80.  */
-    assert(!snan_bit_is_one(status));
-#if defined(TARGET_M68K)
-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
-    r.high = 0x7FFF;
-#else
-    /* X86 */
-    r.low = UINT64_C(0xC000000000000000);
-    r.high = 0xFFFF;
-#endif
+    r.high = 0x7FFF | (p64.sign << 15);
+    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
     return r;
 }
 
-- 
2.34.1

In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
a zero-initialized float_status struct to float32_is_quiet_nan() and
float64_is_quiet_nan(), with the cryptic comment "for
snan_bit_is_one".

This pattern appears to have been copied from target/riscv, where it
is used because the functions there do not have ready access to the
CPU state struct. The comment presumably refers to the fact that the
main reason the is_quiet_nan() functions want the float_state is
because they want to know about the snan_bit_is_one config.

In the loongarch helpers, though, we have the CPU state struct
to hand. Use the usual env->fp_status here. This avoids our needing
to track that we need to update the initializer of the local
float_status structs when the core softfloat code adds new
options for targets to configure their behaviour.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
     } else if (float32_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float32_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
     } else if (float64_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float64_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
-- 
2.34.1

In the frem helper, we have a local float_status because we want to
execute the floatx80_div() with a custom rounding mode.  Instead of
zero-initializing the local float_status and then having to set it up
with the m68k standard behaviour (including the NaN propagation rule
and copying the rounding precision from env->fp_status), initialize
it as a complete copy of env->fp_status. This will avoid our having
to add new code in this function for every new config knob we add
to fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
---
 target/m68k/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/fpu_helper.c
+++ b/target/m68k/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
 
     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
     if (!floatx80_is_any_nan(fp_rem)) {
-        float_status fp_status = { };
+        /* Use local temporary fp_status to set different rounding mode */
+        float_status fp_status = env->fp_status;
         uint32_t quotient;
         int sign;
 
         /* Calculate quotient directly using round to nearest mode */
-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
         set_float_rounding_mode(float_round_nearest_even, &fp_status);
-        set_floatx80_rounding_precision(
-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
 
         sign = extractFloatx80Sign(fp_quot.d);
-- 
2.34.1

In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
from float64 to floatx80 using a scratch float_status, because we
don't want the conversion to affect the CPU's floating point exception
status. Currently we use a zero-initialized float_status. This will
get steadily more awkward as we add config knobs to float_status
that the target must initialize. Avoid having to add any of that
configuration here by instead initializing our local float_status
from the env->fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
---
 target/m68k/helper.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/target/m68k/helper.c b/target/m68k/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/helper.c
+++ b/target/m68k/helper.c
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
     }
     switch (n) {
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
         return 8;
     }
-- 
2.34.1

In the helper functions flcmps and flcmpd we use a scratch float_status
so that we don't change the CPU state if the comparison raises any
floating point exception flags. Instead of zero-initializing this
scratch float_status, initialize it as a copy of env->fp_status. This
avoids the need to explicitly initialize settings like the NaN
propagation rule or others we might add to softfloat in future.

To do this we need to pass the CPU env pointer in to the helper.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
---
 target/sparc/helper.h     | 4 ++--
 target/sparc/fop_helper.c | 8 ++++----
 target/sparc/translate.c  | 4 ++--
 3 files changed, 8 insertions(+), 8 deletions(-)

diff --git a/target/sparc/helper.h b/target/sparc/helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/helper.h
+++ b/target/sparc/helper.h
@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
 DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
+DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
+DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
 DEF_HELPER_2(raise_exception, noreturn, env, int)
 
 DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/fop_helper.c
+++ b/target/sparc/fop_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
     return finish_fcmp(env, r, GETPC());
 }
 
-uint32_t helper_flcmps(float32 src1, float32 src2)
+uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
 {
     /*
      * FLCMP never raises an exception nor modifies any FSR fields.
      * Perform the comparison with a dummy fp environment.
      */
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
     g_assert_not_reached();
 }
 
-uint32_t helper_flcmpd(float64 src1, float64 src2)
+uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
 {
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
diff --git a/target/sparc/translate.c b/target/sparc/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/translate.c
+++ b/target/sparc/translate.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
 
     src1 = gen_load_fpr_F(dc, a->rs1);
     src2 = gen_load_fpr_F(dc, a->rs2);
-    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
 
     src1 = gen_load_fpr_D(dc, a->rs1);
     src2 = gen_load_fpr_D(dc, a->rs2);
-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
-- 
2.34.1

In the helper_compute_fprf functions, we pass a dummy float_status
in to the is_signaling_nan() function. This is unnecessary, because
we have convenient access to the CPU env pointer here and that
is already set up with the correct values for the snan_bit_is_one
and no_signaling_nans config settings. is_signaling_nan() doesn't
ever update the fp_status with any exception flags, so there is
no reason not to use env->fp_status here.

Use env->fp_status instead of the dummy fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
---
 target/ppc/fpu_helper.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/fpu_helper.c
+++ b/target/ppc/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
     } else if (tp##_is_infinity(arg)) {                           \
         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
     } else {                                                      \
-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
+        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
             fprf = 0x00 << FPSCR_FPRF;                            \
         } else {                                                  \
             fprf = 0x11 << FPSCR_FPRF;                            \
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Now that float_status has a bunch of fp parameters,
it is easier to copy an existing structure than create
one from scratch.  Begin by copying the structure that
corresponds to the FPSR and make only the adjustments
required for BFloat16 semantics.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/vec_helper.c | 20 +++++++-------------
 1 file changed, 7 insertions(+), 13 deletions(-)

diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/vec_helper.c
+++ b/target/arm/tcg/vec_helper.c
@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
      * no effect on AArch32 instructions.
      */
     bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
-    *statusp = (float_status){
-        .tininess_before_rounding = float_tininess_before_rounding,
-        .float_rounding_mode = float_round_to_odd_inf,
-        .flush_to_zero = true,
-        .flush_inputs_to_zero = true,
-        .default_nan_mode = true,
-    };
+
+    *statusp = env->vfp.fp_status;
+    set_default_nan_mode(true, statusp);
 
     if (ebf) {
-        float_status *fpst = &env->vfp.fp_status;
-        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
-        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
-        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
-
         /* EBF=1 needs to do a step with round-to-odd semantics */
         *oddstatusp = *statusp;
         set_float_rounding_mode(float_round_to_odd, oddstatusp);
+    } else {
+        set_flush_to_zero(true, statusp);
+        set_flush_inputs_to_zero(true, statusp);
+        set_float_rounding_mode(float_round_to_odd_inf, statusp);
     }
-
     return ebf;
 }
 
-- 
2.34.1

Currently we hardcode the default NaN value in parts64_default_nan()
using a compile-time ifdef ladder. This is awkward for two cases:
 * for single-QEMU-binary we can't hard-code target-specifics like this
 * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
   (specifically the sign bit is different)

Add a field to float_status to specify the default NaN value; fall
back to the old ifdef behaviour if these are not set.

The default NaN value is specified by setting a uint8_t to a
pattern corresponding to the sign and upper fraction parts of
the NaN; the lower bits of the fraction are set from bit 0 of
the pattern.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 +++++++
 include/fpu/softfloat-types.h   | 10 ++++++
 fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
 3 files changed, 54 insertions(+), 22 deletions(-)

Set the default NaN pattern explicitly for the tests/fp code.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c     | 1 +
 tests/fp/fp-test-log2.c | 1 +
 tests/fp/fp-test.c      | 1 +
 3 files changed, 3 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+    set_float_default_nan_pattern(0b01000000, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test-log2.c
+++ b/tests/fp/fp-test-log2.c
@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
     int i;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_rounding_mode(float_round_nearest_even, &qsf);
 
     test.d = 0.0;
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
---
 target/microblaze/cpu.c        | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/microblaze/cpu.c
+++ b/target/microblaze/cpu.c
@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
      * this architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
 
 #if defined(CONFIG_USER_ONLY)
     /* start in user mode with interrupts enabled.  */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
-    || defined(TARGET_MICROBLAZE)
+#elif defined(TARGET_I386) || defined(TARGET_X86_64)
         /* Sign bit set, most significant frac bit set */
         dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 4 ++++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
+    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
+    set_float_default_nan_pattern(0b11000000, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
-        /* Sign bit set, most significant frac bit set */
-        dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
         /* Sign bit clear, msb-1 frac bit set */
         dnan_pattern = 0b00100000;
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 2 ++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 2 insertions(+), 3 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN: sign bit clear, msb-1 frac bit set */
+    set_float_default_nan_pattern(0b00100000, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_HPPA)
-        /* Sign bit clear, msb-1 frac bit set */
-        dnan_pattern = 0b00100000;
 #elif defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
-- 
2.34.1

Set the default NaN pattern explicitly for the arm target.
This includes setting it for the old linux-user nwfpe emulation.
For nwfpe, our default doesn't match the real kernel, but we
avoid making a behaviour change in this commit.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
---
 linux-user/arm/nwfpe/fpa11.c | 5 +++++
 target/arm/cpu.c             | 2 ++
 2 files changed, 7 insertions(+)

diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
index XXXXXXX..XXXXXXX 100644
--- a/linux-user/arm/nwfpe/fpa11.c
+++ b/linux-user/arm/nwfpe/fpa11.c
@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
    * this late date.
    */
   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
+  /*
+   * Use the same default NaN value as Arm VFP. This doesn't match
+   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
+   */
+  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
 }
 
 void SetRoundingMode(const unsigned int opcode)
diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
+ *  * Default NaN has sign bit clear, msb frac bit set
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+    set_float_default_nan_pattern(0b01000000, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
-- 
2.34.1

Set the default NaN pattern explicitly for m68k.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
---
 target/m68k/cpu.c              | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+    /* Default NaN: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     nan = floatx80_default_nan(&env->fp_status);
     for (i = 0; i < 8; i++) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
+#if defined(TARGET_SPARC)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
 #elif defined(TARGET_HEXAGON)
-- 
2.34.1

Set the default NaN pattern explicitly for MIPS. Note that this
is our only target which currently changes the default NaN
at runtime (which it was previously doing indirectly when it
changed the snan_bit_is_one setting).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h | 7 +++++++
 target/mips/msa.c        | 3 +++
 2 files changed, 10 insertions(+)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+    /*
+     * With nan2008, the default NaN value has the sign bit clear and the
+     * frac msb set; with the older mode, the sign bit is clear, and all
+     * frac bits except the msb are set.
+     */
+    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
+                                  &env->active_fpu.fp_status);
 
 }
 
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     /* Inf * 0 + NaN returns the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never,
                               &env->active_tc.msa_fp_status);
+    /* Default NaN: sign bit clear, frac msb set */
+    set_float_default_nan_pattern(0b01000000,
+                                  &env->active_tc.msa_fp_status);
 }
-- 
2.34.1

Set the default NaN pattern explicitly for SPARC, and remove
the ifdef from parts64_default_nan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 5 +----
 2 files changed, 3 insertions(+), 4 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN value: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC)
-        /* Sign bit clear, all frac bits set */
-        dnan_pattern = 0b01111111;
-#elif defined(TARGET_HEXAGON)
+#if defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
 #else
-- 
2.34.1

Set the default NaN pattern explicitly for hexagon.
Remove the ifdef from parts64_default_nan(); the only
remaining unconverted targets all use the default case.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
---
 target/hexagon/cpu.c           | 2 ++
 fpu/softfloat-specialize.c.inc | 5 -----
 2 files changed, 2 insertions(+), 5 deletions(-)

diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hexagon/cpu.c
+++ b/target/hexagon/cpu.c
@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
 
     set_default_nan_mode(1, &env->fp_status);
     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
+    /* Default NaN value: sign bit set, all frac bits set */
+    set_float_default_nan_pattern(0b11111111, &env->fp_status);
 }
 
 static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_HEXAGON)
-        /* Sign bit set, all frac bits set. */
-        dnan_pattern = 0b11111111;
-#else
         /*
          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
             /* sign bit clear, set frac msb */
             dnan_pattern = 0b01000000;
         }
-#endif
     }
     assert(dnan_pattern != 0);
 
-- 
2.34.1

Now that all our targets have bene converted to explicitly specify
their pattern for the default NaN value we can remove the remaining
fallback code in parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 14 --------------
 1 file changed, 14 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint64_t frac;
     uint8_t dnan_pattern = status->default_nan_pattern;
 
-    if (dnan_pattern == 0) {
-        /*
-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
-         * do not have floating-point.
-         */
-        if (snan_bit_is_one(status)) {
-            /* sign bit clear, set all frac bits other than msb */
-            dnan_pattern = 0b00111111;
-        } else {
-            /* sign bit clear, set frac msb */
-            dnan_pattern = 0b01000000;
-        }
-    }
     assert(dnan_pattern != 0);
 
     sign = dnan_pattern >> 7;
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaNMulAdd into its only caller.  This makes
one assert redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
[PMM: keep comment from old code in new location]
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
 fpu/softfloat-specialize.c.inc | 54 ----------------------------------
 2 files changed, 40 insertions(+), 55 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     }
 
     if (s->default_nan_mode) {
+        /*
+         * We guarantee not to require the target to tell us how to
+         * pick a NaN if we're always returning the default NaN.
+         * But if we're not in default-NaN mode then the target must
+         * specify.
+         */
         which = 3;
+    } else if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the
+         * default NaN here, and some return the input NaN.
+         */
+        switch (s->float_infzeronan_rule) {
+        case float_infzeronan_dnan_never:
+            which = 2;
+            break;
+        case float_infzeronan_dnan_always:
+            which = 3;
+            break;
+        case float_infzeronan_dnan_if_qnan:
+            which = is_qnan(c->cls) ? 3 : 2;
+            break;
+        default:
+            g_assert_not_reached();
+        }
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
+        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        Float3NaNPropRule rule = s->float_3nan_prop_rule;
+
+        assert(rule != float_3nan_prop_none);
+        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+            /* We have at least one SNaN input and should prefer it */
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_snan(cls[which]));
+        } else {
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_nan(cls[which]));
+        }
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a three-input operation.
-| For the moment we assume that no CPU needs the 'larger significand'
-| information.
-| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-*----------------------------------------------------------------------------*/
-static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, bool have_snan, float_status *status)
-{
-    FloatClass cls[3] = { a_cls, b_cls, c_cls };
-    Float3NaNPropRule rule = status->float_3nan_prop_rule;
-    int which;
-
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify.
-     */
-    assert(!status->default_nan_mode);
-
-    if (infzero) {
-        /*
-         * Inf * 0 + NaN -- some implementations return the default NaN here,
-         * and some return the input NaN.
-         */
-        switch (status->float_infzeronan_rule) {
-        case float_infzeronan_dnan_never:
-            return 2;
-        case float_infzeronan_dnan_always:
-            return 3;
-        case float_infzeronan_dnan_if_qnan:
-            return is_qnan(c_cls) ? 3 : 2;
-        default:
-            g_assert_not_reached();
-        }
-    }
-
-    assert(rule != float_3nan_prop_none);
-    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-        /* We have at least one SNaN input and should prefer it */
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_snan(cls[which]));
-    } else {
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_nan(cls[which]));
-    }
-    return which;
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remove "3" as a special case for which and simply
branch to return the desired value.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          * But if we're not in default-NaN mode then the target must
          * specify.
          */
-        which = 3;
+        goto default_nan;
     } else if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          */
         switch (s->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
-            which = 2;
             break;
         case float_infzeronan_dnan_always:
-            which = 3;
-            break;
+            goto default_nan;
         case float_infzeronan_dnan_if_qnan:
-            which = is_qnan(c->cls) ? 3 : 2;
+            if (is_qnan(c->cls)) {
+                goto default_nan;
+            }
             break;
         default:
             g_assert_not_reached();
         }
+        which = 2;
     } else {
         FloatClass cls[3] = { a->cls, b->cls, c->cls };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
     }
 
-    if (which == 3) {
-        parts_default_nan(a, s);
-        return a;
-    }
-
     switch (which) {
     case 0:
         break;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         parts_silence_nan(a, s);
     }
     return a;
+
+ default_nan:
+    parts_default_nan(a, s);
+    return a;
 }
 
 /*
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Assign the pointer return value to 'a' directly,
rather than going through an intermediary index.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
 1 file changed, 10 insertions(+), 22 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             FloatPartsN *c, float_status *s,
                                             int ab_mask, int abc_mask)
 {
-    int which;
     bool infzero = (ab_mask == float_cmask_infzero);
     bool have_snan = (abc_mask & float_cmask_snan);
+    FloatPartsN *ret;
 
     if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         default:
             g_assert_not_reached();
         }
-        which = 2;
+        ret = c;
     } else {
-        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        FloatPartsN *val[3] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
         if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
             /* We have at least one SNaN input and should prefer it */
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_snan(cls[which]));
+            } while (!is_snan(ret->cls));
         } else {
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_nan(cls[which]));
+            } while (!is_nan(ret->cls));
         }
     }
 
-    switch (which) {
-    case 0:
-        break;
-    case 1:
-        a = b;
-        break;
-    case 2:
-        a = c;
-        break;
-    default:
-        g_assert_not_reached();
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 
  default_nan:
     parts_default_nan(a, s);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

While all indices into val[] should be in [0-2], the mask
applied is two bits.  To help static analysis see there is
no possibility of read beyond the end of the array, pad the
array to 4 entries, with the final being (implicitly) NULL.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
         ret = c;
     } else {
-        FloatPartsN *val[3] = { a, b, c };
+        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

This function is part of the public interface and
is not "specialized" to any target in any way.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
 fpu/softfloat-specialize.c.inc | 52 ----------------------------------
 2 files changed, 52 insertions(+), 52 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
     *zExpPtr = 1 - shiftCount;
 }
 
+/*----------------------------------------------------------------------------
+| Takes two extended double-precision floating-point values `a' and `b', one
+| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
+| `b' is a signaling NaN, the invalid exception is raised.
+*----------------------------------------------------------------------------*/
+
+floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
+{
+    bool aIsLargerSignificand;
+    FloatClass a_cls, b_cls;
+
+    /* This is not complete, but is good enough for pickNaN.  */
+    a_cls = (!floatx80_is_any_nan(a)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(a, status)
+             ? float_class_snan
+             : float_class_qnan);
+    b_cls = (!floatx80_is_any_nan(b)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(b, status)
+             ? float_class_snan
+             : float_class_qnan);
+
+    if (is_snan(a_cls) || is_snan(b_cls)) {
+        float_raise(float_flag_invalid, status);
+    }
+
+    if (status->default_nan_mode) {
+        return floatx80_default_nan(status);
+    }
+
+    if (a.low < b.low) {
+        aIsLargerSignificand = 0;
+    } else if (b.low < a.low) {
+        aIsLargerSignificand = 1;
+    } else {
+        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
+    }
+
+    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
+        if (is_snan(b_cls)) {
+            return floatx80_silence_nan(b, status);
+        }
+        return b;
+    } else {
+        if (is_snan(a_cls)) {
+            return floatx80_silence_nan(a, status);
+        }
+        return a;
+    }
+}
+
 /*----------------------------------------------------------------------------
 | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
 | and extended significand formed by the concatenation of `zSig0' and `zSig1',
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
     return a;
 }
 
-/*----------------------------------------------------------------------------
-| Takes two extended double-precision floating-point values `a' and `b', one
-| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
-| `b' is a signaling NaN, the invalid exception is raised.
-*----------------------------------------------------------------------------*/
-
-floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-{
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
-
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
-        return floatx80_default_nan(status);
-    }
-
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Unpacking and repacking the parts may be slightly more work
than we did before, but we get to reuse more code.  For a
code path handling exceptional values, this is an improvement.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c | 43 +++++--------------------------------------
 1 file changed, 5 insertions(+), 38 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
 
 floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 {
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
+    FloatParts128 pa, pb, *pr;
 
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
+    if (!floatx80_unpack_canonical(&pa, a, status) ||
+        !floatx80_unpack_canonical(&pb, b, status)) {
         return floatx80_default_nan(status);
     }
 
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
+    pr = parts_pick_nan(&pa, &pb, status);
+    return floatx80_round_pack_canonical(pr, status);
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaN into its only caller.  This makes one assert
redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
 fpu/softfloat-specialize.c.inc | 96 ----------------------------------
 2 files changed, 73 insertions(+), 105 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
 static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
+    int cmp, which;
+
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
     if (s->default_nan_mode) {
         parts_default_nan(a, s);
-    } else {
-        int cmp = frac_cmp(a, b);
-        if (cmp == 0) {
-            cmp = a->sign < b->sign;
-        }
+        return a;
+    }
 
-        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
-            a = b;
-        }
+    cmp = frac_cmp(a, b);
+    if (cmp == 0) {
+        cmp = a->sign < b->sign;
+    }
+
+    switch (s->float_2nan_prop_rule) {
+    case float_2nan_prop_s_ab:
         if (is_snan(a->cls)) {
-            parts_silence_nan(a, s);
+            which = 0;
+        } else if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_qnan(a->cls)) {
+            which = 0;
+        } else {
+            which = 1;
         }
+        break;
+    case float_2nan_prop_s_ba:
+        if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_snan(a->cls)) {
+            which = 0;
+        } else if (is_qnan(b->cls)) {
+            which = 1;
+        } else {
+            which = 0;
+        }
+        break;
+    case float_2nan_prop_ab:
+        which = is_nan(a->cls) ? 0 : 1;
+        break;
+    case float_2nan_prop_ba:
+        which = is_nan(b->cls) ? 1 : 0;
+        break;
+    case float_2nan_prop_x87:
+        /*
+         * This implements x87 NaN propagation rules:
+         * SNaN + QNaN => return the QNaN
+         * two SNaNs => return the one with the larger significand, silenced
+         * two QNaNs => return the one with the larger significand
+         * SNaN and a non-NaN => return the SNaN, silenced
+         * QNaN and a non-NaN => return the QNaN
+         *
+         * If we get down to comparing significands and they are the same,
+         * return the NaN with the positive sign bit (if any).
+         */
+        if (is_snan(a->cls)) {
+            if (is_snan(b->cls)) {
+                which = cmp > 0 ? 0 : 1;
+            } else {
+                which = is_qnan(b->cls) ? 1 : 0;
+            }
+        } else if (is_qnan(a->cls)) {
+            if (is_snan(b->cls) || !is_qnan(b->cls)) {
+                which = 0;
+            } else {
+                which = cmp > 0 ? 0 : 1;
+            }
+        } else {
+            which = 1;
+        }
+        break;
+    default:
+        g_assert_not_reached();
+    }
+
+    if (which) {
+        a = b;
+    }
+    if (is_snan(a->cls)) {
+        parts_silence_nan(a, s);
     }
     return a;
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a two-input operation.
-| IEEE754 doesn't specify all the details of this, so the
-| algorithm is target-specific.
-| The routine is passed various bits of information about the
-| two NaNs and should return 0 to select NaN a and 1 for NaN b.
-| Note that signalling NaNs are always squashed to quiet NaNs
-| by the caller, by calling floatXX_silence_nan() before
-| returning them.
-|
-| aIsLargerSignificand is only valid if both a and b are NaNs
-| of some kind, and is true if a has the larger significand,
-| or if both a and b have the same significand but a is
-| positive but b is negative. It is only needed for the x87
-| tie-break rule.
-*----------------------------------------------------------------------------*/
-
-static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-                   bool aIsLargerSignificand, float_status *status)
-{
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify via set_float_2nan_prop_rule().
-     */
-    assert(!status->default_nan_mode);
-
-    switch (status->float_2nan_prop_rule) {
-    case float_2nan_prop_s_ab:
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_s_ba:
-        if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_ab:
-        if (is_nan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_ba:
-        if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_x87:
-        /*
-         * This implements x87 NaN propagation rules:
-         * SNaN + QNaN => return the QNaN
-         * two SNaNs => return the one with the larger significand, silenced
-         * two QNaNs => return the one with the larger significand
-         * SNaN and a non-NaN => return the SNaN, silenced
-         * QNaN and a non-NaN => return the QNaN
-         *
-         * If we get down to comparing significands and they are the same,
-         * return the NaN with the positive sign bit (if any).
-         */
-        if (is_snan(a_cls)) {
-            if (is_snan(b_cls)) {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-            return is_qnan(b_cls) ? 1 : 0;
-        } else if (is_qnan(a_cls)) {
-            if (is_snan(b_cls) || !is_qnan(b_cls)) {
-                return 0;
-            } else {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-        } else {
-            return 1;
-        }
-    default:
-        g_assert_not_reached();
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remember if there was an SNaN, and use that to simplify
float_2nan_prop_s_{ab,ba} to only the snan component.
Then, fall through to the corresponding
float_2nan_prop_{ab,ba} case to handle any remaining
nans, which must be quiet.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
 1 file changed, 12 insertions(+), 20 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Move the fractional comparison to the end of the
float_2nan_prop_x87 case.  This is not required for
any other 2nan propagation rule.  Reorganize the
x87 case itself to break out of the switch when the
fractional comparison is not required.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 19 +++++++++----------
 1 file changed, 9 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
         return a;
     }
 
-    cmp = frac_cmp(a, b);
-    if (cmp == 0) {
-        cmp = a->sign < b->sign;
-    }
-
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          * return the NaN with the positive sign bit (if any).
          */
         if (is_snan(a->cls)) {
-            if (is_snan(b->cls)) {
-                which = cmp > 0 ? 0 : 1;
-            } else {
+            if (!is_snan(b->cls)) {
                 which = is_qnan(b->cls) ? 1 : 0;
+                break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
                 which = 0;
-            } else {
-                which = cmp > 0 ? 0 : 1;
+                break;
             }
         } else {
             which = 1;
+            break;
         }
+        cmp = frac_cmp(a, b);
+        if (cmp == 0) {
+            cmp = a->sign < b->sign;
+        }
+        which = cmp > 0 ? 0 : 1;
         break;
     default:
         g_assert_not_reached();
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Replace the "index" selecting between A and B with a result variable
of the proper type.  This improves clarity within the function.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
 1 file changed, 13 insertions(+), 15 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
     bool have_snan = false;
-    int cmp, which;
+    FloatPartsN *ret;
+    int cmp;
 
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
-            which = is_snan(a->cls) ? 0 : 1;
+            ret = is_snan(a->cls) ? a : b;
             break;
         }
         /* fall through */
     case float_2nan_prop_ab:
-        which = is_nan(a->cls) ? 0 : 1;
+        ret = is_nan(a->cls) ? a : b;
         break;
     case float_2nan_prop_s_ba:
         if (have_snan) {
-            which = is_snan(b->cls) ? 1 : 0;
+            ret = is_snan(b->cls) ? b : a;
             break;
         }
         /* fall through */
     case float_2nan_prop_ba:
-        which = is_nan(b->cls) ? 1 : 0;
+        ret = is_nan(b->cls) ? b : a;
         break;
     case float_2nan_prop_x87:
         /*
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          */
         if (is_snan(a->cls)) {
             if (!is_snan(b->cls)) {
-                which = is_qnan(b->cls) ? 1 : 0;
+                ret = is_qnan(b->cls) ? b : a;
                 break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
-                which = 0;
+                ret = a;
                 break;
             }
         } else {
-            which = 1;
+            ret = b;
             break;
         }
         cmp = frac_cmp(a, b);
         if (cmp == 0) {
             cmp = a->sign < b->sign;
         }
-        which = cmp > 0 ? 0 : 1;
+        ret = cmp > 0 ? a : b;
         break;
     default:
         g_assert_not_reached();
     }
 
-    if (which) {
-        a = b;
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 }
 
 static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-- 
2.34.1

From: Leif Lindholm <quic_llindhol@quicinc.com>

I'm migrating to Qualcomm's new open source email infrastructure, so
update my email address, and update the mailmap to match.

Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 +-
 .mailmap    | 5 +++--
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
 SBSA-REF
 M: Radoslaw Biernacki <rad@semihalf.com>
 M: Peter Maydell <peter.maydell@linaro.org>
-R: Leif Lindholm <quic_llindhol@quicinc.com>
+R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
 R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
 L: qemu-arm@nongnu.org
 S: Maintained
diff --git a/.mailmap b/.mailmap
index XXXXXXX..XXXXXXX 100644
--- a/.mailmap
+++ b/.mailmap
@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
 Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
 James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
 Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
 Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
 Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
 Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
-- 
2.34.1

From: Vikram Garhwal <vikram.garhwal@bytedance.com>

Previously, maintainer role was paused due to inactive email id. Commit id:
c009d715721861984c4987bcc78b7ee183e86d75.

Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
 
 Xilinx CAN
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 F: hw/net/can/xlnx-*
 F: include/hw/net/xlnx-*
@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
 CAN bus subsystem and hardware
 M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 W: https://canbus.pages.fel.cvut.cz/
 F: net/can/*
-- 
2.34.1