Series comparison

-[PULL 00/27] target-arm queue
+[PULL 00/72] target-arm queue
-First arm pullreq for 5.2: Eric's SMMU stuff, and a bunch of
+First arm pullreq of the cycle; this is mostly my softfloat NaN
-cleanup/refactoring from me.
+handling series. (Lots more in my to-review queue, but I don't
 like pullreqs growing too close to a hundred patches at a time :-))
 thanks
 -- PMM
-The following changes since commit 8367a77c4d3f6e1e60890f5510304feb2c621611:
+The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:
-  Merge remote-tracking branch 'remotes/vivier2/tags/linux-user-for-5.2-pull-request' into staging (2020-08-23 16:34:43 +0100)
+  Open 10.0 development tree (2024-12-10 17:41:17 +0000)
 are available in the Git repository at:
-  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20200824
+  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211
-for you to fetch changes up to b34aa5129e9c3aff890b4f4bcc84962e94185629:
+for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:
-  target/arm: Use correct FPST for VCMLA, VCADD on fp16 (2020-08-24 10:15:12 +0100)
+  MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)
 ----------------------------------------------------------------
 target-arm queue:
- * hw/cpu/a9mpcore: Verify the machine use Cortex-A9 cores
+ * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
- * hw/arm/smmuv3: Implement SMMUv3.2 range-invalidation
+ * fpu: Make muladd NaN handling runtime-selected, not compile-time
- * docs/system/arm: Document the Xilinx Versal Virt board
+ * fpu: Make default NaN pattern runtime-selected, not compile-time
- * target/arm: Make M-profile NOCP take precedence over UNDEF
+ * fpu: Minor NaN-related cleanups
- * target/arm: Use correct FPST for VCMLA, VCADD on fp16
+ * MAINTAINERS: email address updates
  * target/arm: Various cleanups preparing for fp16 support
 ----------------------------------------------------------------
-Edgar E. Iglesias (1):
+Bernhard Beschow (5):
-      docs/system/arm: Document the Xilinx Versal Virt board
+      hw/net/lan9118: Extract lan9118_phy
       hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
       hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
       hw/net/lan9118_phy: Reuse MII constants
       hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
-Eric Auger (11):
+Leif Lindholm (1):
-      hw/arm/smmu-common: Factorize some code in smmu_ptw_64()
+      MAINTAINERS: update email address for Leif Lindholm
       hw/arm/smmu-common: Add IOTLB helpers
       hw/arm/smmu: Introduce smmu_get_iotlb_key()
       hw/arm/smmu: Introduce SMMUTLBEntry for PTW and IOTLB value
       hw/arm/smmu-common: Manage IOTLB block entries
       hw/arm/smmuv3: Introduce smmuv3_s1_range_inval() helper
       hw/arm/smmuv3: Get prepared for range invalidation
       hw/arm/smmuv3: Fix IIDR offset
       hw/arm/smmuv3: Let AIDR advertise SMMUv3.0 support
       hw/arm/smmuv3: Support HAD and advertise SMMUv3.1 support
       hw/arm/smmuv3: Advertise SMMUv3.2 range invalidation
-Peter Maydell (14):
+Peter Maydell (54):
-      target/arm: Pull handling of XScale insns out of disas_coproc_insn()
+      fpu: handle raising Invalid for infzero in pick_nan_muladd
-      target/arm: Separate decode from handling of coproc insns
+      fpu: Check for default_nan_mode before calling pickNaNMulAdd
-      target/arm: Convert A32 coprocessor insns to decodetree
+      softfloat: Allow runtime choice of inf * 0 + NaN result
-      target/arm: Tidy up disas_arm_insn()
+      tests/fp: Explicitly set inf-zero-nan rule
-      target/arm: Do M-profile NOCP checks early and via decodetree
+      target/arm: Set FloatInfZeroNaNRule explicitly
-      target/arm: Convert T32 coprocessor insns to decodetree
+      target/s390: Set FloatInfZeroNaNRule explicitly
-      target/arm: Remove ARCH macro
+      target/ppc: Set FloatInfZeroNaNRule explicitly
-      target/arm: Delete unused VFP_DREG macros
+      target/mips: Set FloatInfZeroNaNRule explicitly
-      target/arm/translate.c: Delete/amend incorrect comments
+      target/sparc: Set FloatInfZeroNaNRule explicitly
-      target/arm: Delete unused ARM_FEATURE_CRC
+      target/xtensa: Set FloatInfZeroNaNRule explicitly
-      target/arm: Replace A64 get_fpstatus_ptr() with generic fpstatus_ptr()
+      target/x86: Set FloatInfZeroNaNRule explicitly
-      target/arm: Make A32/T32 use new fpstatus_ptr() API
+      target/loongarch: Set FloatInfZeroNaNRule explicitly
-      target/arm: Implement FPST_STD_F16 fpstatus
+      target/hppa: Set FloatInfZeroNaNRule explicitly
-      target/arm: Use correct FPST for VCMLA, VCADD on fp16
+      softfloat: Pass have_snan to pickNaNMulAdd
       softfloat: Allow runtime choice of NaN propagation for muladd
       tests/fp: Explicitly set 3-NaN propagation rule
       target/arm: Set Float3NaNPropRule explicitly
       target/loongarch: Set Float3NaNPropRule explicitly
       target/ppc: Set Float3NaNPropRule explicitly
       target/s390x: Set Float3NaNPropRule explicitly
       target/sparc: Set Float3NaNPropRule explicitly
       target/mips: Set Float3NaNPropRule explicitly
       target/xtensa: Set Float3NaNPropRule explicitly
       target/i386: Set Float3NaNPropRule explicitly
       target/hppa: Set Float3NaNPropRule explicitly
       fpu: Remove use_first_nan field from float_status
       target/m68k: Don't pass NULL float_status to floatx80_default_nan()
       softfloat: Create floatx80 default NaN from parts64_default_nan
       target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
       target/m68k: In frem helper, initialize local float_status from env->fp_status
       target/m68k: Init local float_status from env fp_status in gdb get/set reg
       target/sparc: Initialize local scratch float_status from env->fp_status
       target/ppc: Use env->fp_status in helper_compute_fprf functions
       fpu: Allow runtime choice of default NaN value
       tests/fp: Set default NaN pattern explicitly
       target/microblaze: Set default NaN pattern explicitly
       target/i386: Set default NaN pattern explicitly
       target/hppa: Set default NaN pattern explicitly
       target/alpha: Set default NaN pattern explicitly
       target/arm: Set default NaN pattern explicitly
       target/loongarch: Set default NaN pattern explicitly
       target/m68k: Set default NaN pattern explicitly
       target/mips: Set default NaN pattern explicitly
       target/openrisc: Set default NaN pattern explicitly
       target/ppc: Set default NaN pattern explicitly
       target/sh4: Set default NaN pattern explicitly
       target/rx: Set default NaN pattern explicitly
       target/s390x: Set default NaN pattern explicitly
       target/sparc: Set default NaN pattern explicitly
       target/xtensa: Set default NaN pattern explicitly
       target/hexagon: Set default NaN pattern explicitly
       target/riscv: Set default NaN pattern explicitly
       target/tricore: Set default NaN pattern explicitly
       fpu: Remove default handling for dnan_pattern
-Philippe Mathieu-Daudé (1):
+Richard Henderson (11):
-      hw/cpu/a9mpcore: Verify the machine use Cortex-A9 cores
+      target/arm: Copy entire float_status in is_ebf
       softfloat: Inline pickNaNMulAdd
       softfloat: Use goto for default nan case in pick_nan_muladd
       softfloat: Remove which from parts_pick_nan_muladd
       softfloat: Pad array size in pick_nan_muladd
       softfloat: Move propagateFloatx80NaN to softfloat.c
       softfloat: Use parts_pick_nan in propagateFloatx80NaN
       softfloat: Inline pickNaN
       softfloat: Share code between parts_pick_nan cases
       softfloat: Sink frac_cmp in parts_pick_nan until needed
       softfloat: Replace WHICH with RET in parts_pick_nan
- docs/system/arm/xlnx-versal-virt.rst | 176 +++++++++++++++++++++++
+Vikram Garhwal (1):
- docs/system/target-arm.rst           |   1 +
+      MAINTAINERS: Add correct email address for Vikram Garhwal
  hw/arm/smmu-internal.h               |   8 ++
  hw/arm/smmuv3-internal.h             |  10 +-
  include/hw/arm/smmu-common.h         |  19 ++-
  include/hw/arm/smmuv3.h              |   1 +
  target/arm/cpu.h                     |  10 +-
  target/arm/translate-a64.h           |   1 -
  target/arm/translate.h               |  52 +++++++
  target/arm/a32.decode                |  19 +++
  target/arm/m-nocp.decode             |  42 ++++++
  target/arm/t32.decode                |  19 +++
  target/arm/vfp.decode                |   2 -
  hw/arm/smmu-common.c                 | 214 ++++++++++++++++++---------
  hw/arm/smmuv3.c                      | 142 +++++++++---------
  hw/cpu/a9mpcore.c                    |  12 +-
  target/arm/cpu.c                     |   3 +
  target/arm/helper.c                  |  29 ++++
  target/arm/translate-a64.c           |  89 +++++-------
  target/arm/translate-sve.c           |  34 ++---
  target/arm/translate.c               | 272 +++++++++++++++++------------------
  target/arm/vfp_helper.c              |   5 +
  MAINTAINERS                          |   3 +-
  hw/arm/trace-events                  |  12 +-
  target/arm/meson.build               |   1 +
  target/arm/translate-neon.c.inc      |  28 ++--
  target/arm/translate-vfp.c.inc       |  96 ++++++++-----
 files changed, 885 insertions(+), 415 deletions(-)
  create mode 100644 docs/system/arm/xlnx-versal-virt.rst
  create mode 100644 target/arm/m-nocp.decode
+ MAINTAINERS                       |   4 +-
+ include/fpu/softfloat-helpers.h   |  38 +++-
+ include/fpu/softfloat-types.h     |  89 +++++++-
+ include/hw/net/imx_fec.h          |   9 +-
+ include/hw/net/lan9118_phy.h      |  37 ++++
+ include/hw/net/mii.h              |   6 +
+ target/mips/fpu_helper.h          |  20 ++
+ target/sparc/helper.h             |   4 +-
+ fpu/softfloat.c                   |  19 ++
+ hw/net/imx_fec.c                  | 146 ++------------
+ hw/net/lan9118.c                  | 137 ++-----------
+ hw/net/lan9118_phy.c              | 222 ++++++++++++++++++++
+ linux-user/arm/nwfpe/fpa11.c      |   5 +
+ target/alpha/cpu.c                |   2 +
+ target/arm/cpu.c                  |  10 +
+ target/arm/tcg/vec_helper.c       |  20 +-
+ target/hexagon/cpu.c              |   2 +
+ target/hppa/fpu_helper.c          |  12 ++
+ target/i386/tcg/fpu_helper.c      |  12 ++
+ target/loongarch/tcg/fpu_helper.c |  14 +-
+ target/m68k/cpu.c                 |  14 +-
+ target/m68k/fpu_helper.c          |   6 +-
+ target/m68k/helper.c              |   6 +-
+ target/microblaze/cpu.c           |   2 +
+ target/mips/msa.c                 |  10 +
+ target/openrisc/cpu.c             |   2 +
+ target/ppc/cpu_init.c             |  19 ++
+ target/ppc/fpu_helper.c           |   3 +-
+ target/riscv/cpu.c                |   2 +
+ target/rx/cpu.c                   |   2 +
+ target/s390x/cpu.c                |   5 +
+ target/sh4/cpu.c                  |   2 +
+ target/sparc/cpu.c                |   6 +
+ target/sparc/fop_helper.c         |   8 +-
+ target/sparc/translate.c          |   4 +-
+ target/tricore/helper.c           |   2 +
+ target/xtensa/cpu.c               |   4 +
+ target/xtensa/fpu_helper.c        |   3 +-
+ tests/fp/fp-bench.c               |   7 +
+ tests/fp/fp-test-log2.c           |   1 +
+ tests/fp/fp-test.c                |   7 +
+ fpu/softfloat-parts.c.inc         | 152 +++++++++++---
+ fpu/softfloat-specialize.c.inc    | 412 ++------------------------------------
+ .mailmap                          |   5 +-
+ hw/net/Kconfig                    |   5 +
+ hw/net/meson.build                |   1 +
+ hw/net/trace-events               |  10 +-
+files changed, 778 insertions(+), 730 deletions(-)
+ create mode 100644 include/hw/net/lan9118_phy.h
+ create mode 100644 hw/net/lan9118_phy.c

-[PULL 08/27] hw/arm/smmuv3: Get prepared for range invalidation
+[PULL 01/72] hw/net/lan9118: Extract lan9118_phy
-From: Eric Auger <eric.auger@redhat.com>
+From: Bernhard Beschow <shentey@gmail.com>
-Enhance the smmu_iotlb_inv_iova() helper with range invalidation.
+A very similar implementation of the same device exists in imx_fec. Prepare for
-This uses the new fields passed in the NH_VA and NH_VAA commands:
+a common implementation by extracting a device model into its own files.
 the size of the range, the level and the granule.
-As NH_VA and NH_VAA both use those fields, their decoding and
+Some migration state has been moved into the new device model which breaks
-handling is factorized in a new smmuv3_s1_range_inval() helper.
+migration compatibility for the following machines:
 * smdkc210
 * realview-*
 * vexpress-*
 * kzm
 * mps2-*
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+While breaking migration ABI, fix the size of the MII registers to be 16 bit,
 as defined by IEEE 802.3u.
 Signed-off-by: Bernhard Beschow <shentey@gmail.com>
 Tested-by: Guenter Roeck <linux@roeck-us.net>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20200728150815.11446-8-eric.auger@redhat.com
+Message-id: 20241102125724.532843-2-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h     |  4 +++
+ include/hw/net/lan9118_phy.h |  37 ++++++++
- include/hw/arm/smmu-common.h |  3 +-
+ hw/net/lan9118.c             | 137 +++++-----------------------
- hw/arm/smmu-common.c         | 25 +++++++++++---
+ hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
- hw/arm/smmuv3.c              | 64 +++++++++++++++++++++++-------------
+ hw/net/Kconfig               |   4 +
- hw/arm/trace-events          |  4 +--
+ hw/net/meson.build           |   1 +
-files changed, 69 insertions(+), 31 deletions(-)
+files changed, 233 insertions(+), 115 deletions(-)
  create mode 100644 include/hw/net/lan9118_phy.h
  create mode 100644 hw/net/lan9118_phy.c
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
 new file mode 100644
 index XXXXXXX..XXXXXXX
 --- /dev/null
 +++ b/include/hw/net/lan9118_phy.h
@@ -XXX,XX +XXX,XX @@
 +/*
 + * SMSC LAN9118 PHY emulation
 + *
 + * Copyright (c) 2009 CodeSourcery, LLC.
 + * Written by Paul Brook
 + *
 + * This work is licensed under the terms of the GNU GPL, version 2 or later.
 + * See the COPYING file in the top-level directory.
 + */
 +
 +#ifndef HW_NET_LAN9118_PHY_H
 +#define HW_NET_LAN9118_PHY_H
 +
 +#include "qom/object.h"
 +#include "hw/sysbus.h"
 +
 +#define TYPE_LAN9118_PHY "lan9118-phy"
 +OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
 +
 +typedef struct Lan9118PhyState {
 +    SysBusDevice parent_obj;
 +
 +    uint16_t status;
 +    uint16_t control;
 +    uint16_t advertise;
 +    uint16_t ints;
 +    uint16_t int_mask;
 +    qemu_irq irq;
 +    bool link_down;
 +} Lan9118PhyState;
 +
 +void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
 +void lan9118_phy_reset(Lan9118PhyState *s);
 +uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
 +void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
 +
 +#endif
 diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/hw/net/lan9118.c
-+++ b/hw/arm/smmuv3-internal.h
++++ b/hw/net/lan9118.c
-@@ -XXX,XX +XXX,XX @@ enum { /* Command completion notification */
+@@ -XXX,XX +XXX,XX @@
- };
+ #include "net/net.h"
+ #include "net/eth.h"
- #define CMD_TYPE(x)         extract32((x)->word[0], 0 , 8)
+ #include "hw/irq.h"
-+#define CMD_NUM(x)          extract32((x)->word[0], 12 , 5)
++#include "hw/net/lan9118_phy.h"
-+#define CMD_SCALE(x)        extract32((x)->word[0], 20 , 5)
+ #include "hw/net/lan9118.h"
- #define CMD_SSEC(x)         extract32((x)->word[0], 10, 1)
+ #include "hw/ptimer.h"
- #define CMD_SSV(x)          extract32((x)->word[0], 11, 1)
+ #include "hw/qdev-properties.h"
- #define CMD_RESUME_AC(x)    extract32((x)->word[0], 12, 1)
+@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
-@@ -XXX,XX +XXX,XX @@ enum { /* Command completion notification */
+ #define MAC_CR_RXEN     0x00000004
- #define CMD_RESUME_STAG(x)  extract32((x)->word[2], 0 , 16)
+ #define MAC_CR_RESERVED 0x7f404213
- #define CMD_RESP(x)         extract32((x)->word[2], 11, 2)
- #define CMD_LEAF(x)         extract32((x)->word[2], 0 , 1)
+-#define PHY_INT_ENERGYON            0x80
-+#define CMD_TTL(x)          extract32((x)->word[2], 8 , 2)
+-#define PHY_INT_AUTONEG_COMPLETE    0x40
-+#define CMD_TG(x)           extract32((x)->word[2], 10, 2)
+-#define PHY_INT_FAULT               0x20
- #define CMD_STE_RANGE(x)    extract32((x)->word[2], 0 , 5)
+-#define PHY_INT_DOWN                0x10
- #define CMD_ADDR(x) ({                                        \
+-#define PHY_INT_AUTONEG_LP          0x08
-             uint64_t high = (uint64_t)(x)->word[3];           \
+-#define PHY_INT_PARFAULT            0x04
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+-#define PHY_INT_AUTONEG_PAGE        0x02
-index XXXXXXX..XXXXXXX 100644
+-
---- a/include/hw/arm/smmu-common.h
+ #define GPT_TIMER_EN    0x20000000
-+++ b/include/hw/arm/smmu-common.h
-@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
+ /*
-                                 uint8_t tg, uint8_t level);
+@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
- void smmu_iotlb_inv_all(SMMUState *s);
+     uint32_t mac_mii_data;
- void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
+     uint32_t mac_flow;
--void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova);
-+void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova,
+-    uint32_t phy_status;
-+                         uint8_t tg, uint64_t num_pages, uint8_t ttl);
+-    uint32_t phy_control;
+-    uint32_t phy_advertise;
- /* Unmap the range of all the notifiers registered to any IOMMU mr */
+-    uint32_t phy_int;
- void smmu_inv_notifiers_all(SMMUState *s);
+-    uint32_t phy_int_mask;
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
++    Lan9118PhyState mii;
-index XXXXXXX..XXXXXXX 100644
++    IRQState mii_irq;
---- a/hw/arm/smmu-common.c
-+++ b/hw/arm/smmu-common.c
+     int32_t eeprom_writable;
-@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid_iova(gpointer key, gpointer value,
+     uint8_t eeprom[128];
-     if (info->asid >= 0 && info->asid != SMMU_IOTLB_ASID(iotlb_key)) {
+@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
-         return false;
-     }
+ static const VMStateDescription vmstate_lan9118 = {
--    return (info->iova & ~entry->addr_mask) == entry->iova;
+     .name = "lan9118",
-+    return ((info->iova & ~entry->addr_mask) == entry->iova) ||
+-    .version_id = 2,
-+           ((entry->iova & ~info->mask) == info->iova);
+-    .minimum_version_id = 1,
 +    .version_id = 3,
 +    .minimum_version_id = 3,
      .fields = (const VMStateField[]) {
          VMSTATE_PTIMER(timer, lan9118_state),
          VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
          VMSTATE_UINT32(mac_mii_acc, lan9118_state),
          VMSTATE_UINT32(mac_mii_data, lan9118_state),
          VMSTATE_UINT32(mac_flow, lan9118_state),
 -        VMSTATE_UINT32(phy_status, lan9118_state),
 -        VMSTATE_UINT32(phy_control, lan9118_state),
 -        VMSTATE_UINT32(phy_advertise, lan9118_state),
 -        VMSTATE_UINT32(phy_int, lan9118_state),
 -        VMSTATE_UINT32(phy_int_mask, lan9118_state),
          VMSTATE_INT32(eeprom_writable, lan9118_state),
          VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
          VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
      lan9118_mac_changed(s);
  }
--inline void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova)
+-static void phy_update_irq(lan9118_state *s)
-+inline void
++static void lan9118_update_irq(void *opaque, int n, int level)
 +smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova,
 +                    uint8_t tg, uint64_t num_pages, uint8_t ttl)
  {
--    SMMUIOTLBPageInvInfo info = {.asid = asid, .iova = iova};
+-    if (s->phy_int & s->phy_int_mask) {
-+    if (ttl && (num_pages == 1)) {
++    lan9118_state *s = opaque;
-+        SMMUIOTLBKey key = smmu_get_iotlb_key(asid, iova, tg, ttl);
++
++    if (level) {
--    trace_smmu_iotlb_inv_iova(asid, iova);
+         s->int_sts |= PHY_INT;
--    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_iova, &info);
+     } else {
-+        g_hash_table_remove(s->iotlb, &key);
+         s->int_sts &= ~PHY_INT;
-+    } else {
+@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
-+        /* if tg is not set we use 4KB range invalidation */
+     lan9118_update(s);
 +        uint8_t granule = tg ? tg * 2 + 10 : 12;
 +
 +        SMMUIOTLBPageInvInfo info = {
 +            .asid = asid, .iova = iova,
 +            .mask = (num_pages * 1 << granule) - 1};
 +
 +        g_hash_table_foreach_remove(s->iotlb,
 +                                    smmu_hash_remove_by_asid_iova,
 +                                    &info);
 +    }
  }
- inline void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
+-static void phy_update_link(lan9118_state *s)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+-{
-index XXXXXXX..XXXXXXX 100644
+-    /* Autonegotiation status mirrors link status.  */
---- a/hw/arm/smmuv3.c
+-    if (qemu_get_queue(s->nic)->link_down) {
-+++ b/hw/arm/smmuv3.c
+-        s->phy_status &= ~0x0024;
-@@ -XXX,XX +XXX,XX @@ epilogue:
+-        s->phy_int |= PHY_INT_DOWN;
-  * @n: notifier to be called
+-    } else {
-  * @asid: address space ID or negative value if we don't care
+-        s->phy_status |= 0x0024;
-  * @iova: iova
+-        s->phy_int |= PHY_INT_ENERGYON;
-+ * @tg: translation granule (if communicated through range invalidation)
+-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-+ * @num_pages: number of @granule sized pages (if tg != 0), otherwise 1
+-    }
-  */
+-    phy_update_irq(s);
- static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
+-}
-                                IOMMUNotifier *n,
+-
--                               int asid,
+ static void lan9118_set_link(NetClientState *nc)
 -                               dma_addr_t iova)
 +                               int asid, dma_addr_t iova,
 +                               uint8_t tg, uint64_t num_pages)
  {
-     SMMUDevice *sdev = container_of(mr, SMMUDevice, iommu);
+-    phy_update_link(qemu_get_nic_opaque(nc));
--    SMMUEventInfo event = {.inval_ste_allowed = true};
+-}
--    SMMUTransTableInfo *tt;
+-
--    SMMUTransCfg *cfg;
+-static void phy_reset(lan9118_state *s)
-     IOMMUTLBEntry entry;
+-{
-+    uint8_t granule = tg;
+-    s->phy_status = 0x7809;
+-    s->phy_control = 0x3000;
--    cfg = smmuv3_get_config(sdev, &event);
+-    s->phy_advertise = 0x01e1;
--    if (!cfg) {
+-    s->phy_int_mask = 0;
--        return;
+-    s->phy_int = 0;
--    }
+-    phy_update_link(s);
-+    if (!tg) {
++    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
-+        SMMUEventInfo event = {.inval_ste_allowed = true};
++                            nc->link_down);
 +        SMMUTransCfg *cfg = smmuv3_get_config(sdev, &event);
 +        SMMUTransTableInfo *tt;
 -    if (asid >= 0 && cfg->asid != asid) {
 -        return;
 -    }
 +        if (!cfg) {
 +            return;
 +        }
 -    tt = select_tt(cfg, iova);
 -    if (!tt) {
 -        return;
 +        if (asid >= 0 && cfg->asid != asid) {
 +            return;
 +        }
 +
 +        tt = select_tt(cfg, iova);
 +        if (!tt) {
 +            return;
 +        }
 +        granule = tt->granule_sz;
      }
      entry.target_as = &address_space_memory;
      entry.iova = iova;
 -    entry.addr_mask = (1 << tt->granule_sz) - 1;
 +    entry.addr_mask = num_pages * (1 << granule) - 1;
      entry.perm = IOMMU_NONE;
      memory_region_notify_one(n, &entry);
  }
--/* invalidate an asid/iova tuple in all mr's */
+ static void lan9118_reset(DeviceState *d)
--static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
+@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
-+/* invalidate an asid/iova range tuple in all mr's */
+     s->read_word_n = 0;
-+static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova,
+     s->write_word_n = 0;
-+                                      uint8_t tg, uint64_t num_pages)
- {
+-    phy_reset(s);
-     SMMUDevice *sdev;
+-
+     s->eeprom_writable = 0;
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
+     lan9118_reload_eeprom(s);
-         IOMMUMemoryRegion *mr = &sdev->iommu;
+ }
-         IOMMUNotifier *n;
+@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
+     uint32_t status;
--        trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, iova);
-+        trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, iova,
+     /* FIXME: Honor TX disable, and allow queueing of packets.  */
-+                                        tg, num_pages);
+-    if (s->phy_control & 0x4000)  {
++    if (s->mii.control & 0x4000) {
-         IOMMU_NOTIFIER_FOREACH(n, mr) {
+         /* This assumes the receive routine doesn't touch the VLANClient.  */
--            smmuv3_notify_iova(mr, n, asid, iova);
+         qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
-+            smmuv3_notify_iova(mr, n, asid, iova, tg, num_pages);
+     } else {
-         }
+@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
      }
  }
- static void smmuv3_s1_range_inval(SMMUState *s, Cmd *cmd)
+-static uint32_t do_phy_read(lan9118_state *s, int reg)
 -{
 -    uint32_t val;
 -
 -    switch (reg) {
 -    case 0: /* Basic Control */
 -        return s->phy_control;
 -    case 1: /* Basic Status */
 -        return s->phy_status;
 -    case 2: /* ID1 */
 -        return 0x0007;
 -    case 3: /* ID2 */
 -        return 0xc0d1;
 -    case 4: /* Auto-neg advertisement */
 -        return s->phy_advertise;
 -    case 5: /* Auto-neg Link Partner Ability */
 -        return 0x0f71;
 -    case 6: /* Auto-neg Expansion */
 -        return 1;
 -        /* TODO 17, 18, 27, 29, 30, 31 */
 -    case 29: /* Interrupt source.  */
 -        val = s->phy_int;
 -        s->phy_int = 0;
 -        phy_update_irq(s);
 -        return val;
 -    case 30: /* Interrupt mask */
 -        return s->phy_int_mask;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_read: PHY read reg %d\n", reg);
 -        return 0;
 -    }
 -}
 -
 -static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
 -{
 -    switch (reg) {
 -    case 0: /* Basic Control */
 -        if (val & 0x8000) {
 -            phy_reset(s);
 -            break;
 -        }
 -        s->phy_control = val & 0x7980;
 -        /* Complete autonegotiation immediately.  */
 -        if (val & 0x1000) {
 -            s->phy_status |= 0x0020;
 -        }
 -        break;
 -    case 4: /* Auto-neg advertisement */
 -        s->phy_advertise = (val & 0x2d7f) | 0x80;
 -        break;
 -        /* TODO 17, 18, 27, 31 */
 -    case 30: /* Interrupt mask */
 -        s->phy_int_mask = val & 0xff;
 -        phy_update_irq(s);
 -        break;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 -    }
 -}
 -
  static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
  {
-+    uint8_t scale = 0, num = 0, ttl = 0;
+     switch (reg) {
-     dma_addr_t addr = CMD_ADDR(cmd);
+@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
-     uint8_t type = CMD_TYPE(cmd);
+         if (val & 2) {
-     uint16_t vmid = CMD_VMID(cmd);
+             DPRINTF("PHY write %d = 0x%04x\n",
-     bool leaf = CMD_LEAF(cmd);
+                     (val >> 6) & 0x1f, s->mac_mii_data);
-+    uint8_t tg = CMD_TG(cmd);
+-            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
-+    hwaddr num_pages = 1;
++            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
-     int asid = -1;
+         } else {
+-            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
-+    if (tg) {
++            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
-+        scale = CMD_SCALE(cmd);
+             DPRINTF("PHY read %d = 0x%04x\n",
-+        num = CMD_NUM(cmd);
+                     (val >> 6) & 0x1f, s->mac_mii_data);
-+        ttl = CMD_TTL(cmd);
+         }
-+        num_pages = (num + 1) * (1 << (scale));
+@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
          break;
      case CSR_PMT_CTRL:
          if (val & 0x400) {
 -            phy_reset(s);
 +            lan9118_phy_reset(&s->mii);
          }
          s->pmt_ctrl &= ~0x34e;
          s->pmt_ctrl |= (val & 0x34e);
@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
      const MemoryRegionOps *mem_ops =
              s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
 +    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
 +    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
 +    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
 +        return;
 +    }
-+
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
-     if (type == SMMU_CMD_TLBI_NH_VA) {
++
-         asid = CMD_ASID(cmd);
+     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
-     }
+                           "lan9118-mmio", 0x100);
--    trace_smmuv3_s1_range_inval(vmid, asid, addr, leaf);
+     sysbus_init_mmio(sbd, &s->mmio);
--    smmuv3_inv_notifiers_iova(s, asid, addr);
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
--    smmu_iotlb_inv_iova(s, asid, addr);
+new file mode 100644
-+    trace_smmuv3_s1_range_inval(vmid, asid, addr, tg, num_pages, ttl, leaf);
+index XXXXXXX..XXXXXXX
-+    smmuv3_inv_notifiers_iova(s, asid, addr, tg, num_pages);
+--- /dev/null
-+    smmu_iotlb_inv_iova(s, asid, addr, tg, num_pages, ttl);
++++ b/hw/net/lan9118_phy.c
- }
+@@ -XXX,XX +XXX,XX @@
++/*
- static int smmuv3_cmdq_consume(SMMUv3State *s)
++ * SMSC LAN9118 PHY emulation
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
++ *
 + * Copyright (c) 2009 CodeSourcery, LLC.
 + * Written by Paul Brook
 + *
 + * This code is licensed under the GNU GPL v2
 + *
 + * Contributions after 2012-01-13 are licensed under the terms of the
 + * GNU GPL, version 2 or (at your option) any later version.
 + */
 +
 +#include "qemu/osdep.h"
 +#include "hw/net/lan9118_phy.h"
 +#include "hw/irq.h"
 +#include "hw/resettable.h"
 +#include "migration/vmstate.h"
 +#include "qemu/log.h"
 +
 +#define PHY_INT_ENERGYON            (1 << 7)
 +#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
 +#define PHY_INT_FAULT               (1 << 5)
 +#define PHY_INT_DOWN                (1 << 4)
 +#define PHY_INT_AUTONEG_LP          (1 << 3)
 +#define PHY_INT_PARFAULT            (1 << 2)
 +#define PHY_INT_AUTONEG_PAGE        (1 << 1)
 +
 +static void lan9118_phy_update_irq(Lan9118PhyState *s)
 +{
 +    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
 +}
 +
 +uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
 +{
 +    uint16_t val;
 +
 +    switch (reg) {
 +    case 0: /* Basic Control */
 +        return s->control;
 +    case 1: /* Basic Status */
 +        return s->status;
 +    case 2: /* ID1 */
 +        return 0x0007;
 +    case 3: /* ID2 */
 +        return 0xc0d1;
 +    case 4: /* Auto-neg advertisement */
 +        return s->advertise;
 +    case 5: /* Auto-neg Link Partner Ability */
 +        return 0x0f71;
 +    case 6: /* Auto-neg Expansion */
 +        return 1;
 +        /* TODO 17, 18, 27, 29, 30, 31 */
 +    case 29: /* Interrupt source. */
 +        val = s->ints;
 +        s->ints = 0;
 +        lan9118_phy_update_irq(s);
 +        return val;
 +    case 30: /* Interrupt mask */
 +        return s->int_mask;
 +    default:
 +        qemu_log_mask(LOG_GUEST_ERROR,
 +                      "lan9118_phy_read: PHY read reg %d\n", reg);
 +        return 0;
 +    }
 +}
 +
 +void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
 +{
 +    switch (reg) {
 +    case 0: /* Basic Control */
 +        if (val & 0x8000) {
 +            lan9118_phy_reset(s);
 +            break;
 +        }
 +        s->control = val & 0x7980;
 +        /* Complete autonegotiation immediately. */
 +        if (val & 0x1000) {
 +            s->status |= 0x0020;
 +        }
 +        break;
 +    case 4: /* Auto-neg advertisement */
 +        s->advertise = (val & 0x2d7f) | 0x80;
 +        break;
 +        /* TODO 17, 18, 27, 31 */
 +    case 30: /* Interrupt mask */
 +        s->int_mask = val & 0xff;
 +        lan9118_phy_update_irq(s);
 +        break;
 +    default:
 +        qemu_log_mask(LOG_GUEST_ERROR,
 +                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 +    }
 +}
 +
 +void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 +{
 +    s->link_down = link_down;
 +
 +    /* Autonegotiation status mirrors link status. */
 +    if (link_down) {
 +        s->status &= ~0x0024;
 +        s->ints |= PHY_INT_DOWN;
 +    } else {
 +        s->status |= 0x0024;
 +        s->ints |= PHY_INT_ENERGYON;
 +        s->ints |= PHY_INT_AUTONEG_COMPLETE;
 +    }
 +    lan9118_phy_update_irq(s);
 +}
 +
 +void lan9118_phy_reset(Lan9118PhyState *s)
 +{
 +    s->control = 0x3000;
 +    s->status = 0x7809;
 +    s->advertise = 0x01e1;
 +    s->int_mask = 0;
 +    s->ints = 0;
 +    lan9118_phy_update_link(s, s->link_down);
 +}
 +
 +static void lan9118_phy_reset_hold(Object *obj, ResetType type)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    lan9118_phy_reset(s);
 +}
 +
 +static void lan9118_phy_init(Object *obj)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
 +}
 +
 +static const VMStateDescription vmstate_lan9118_phy = {
 +    .name = "lan9118-phy",
 +    .version_id = 1,
 +    .minimum_version_id = 1,
 +    .fields = (const VMStateField[]) {
 +        VMSTATE_UINT16(control, Lan9118PhyState),
 +        VMSTATE_UINT16(status, Lan9118PhyState),
 +        VMSTATE_UINT16(advertise, Lan9118PhyState),
 +        VMSTATE_UINT16(ints, Lan9118PhyState),
 +        VMSTATE_UINT16(int_mask, Lan9118PhyState),
 +        VMSTATE_BOOL(link_down, Lan9118PhyState),
 +        VMSTATE_END_OF_LIST()
 +    }
 +};
 +
 +static void lan9118_phy_class_init(ObjectClass *klass, void *data)
 +{
 +    ResettableClass *rc = RESETTABLE_CLASS(klass);
 +    DeviceClass *dc = DEVICE_CLASS(klass);
 +
 +    rc->phases.hold = lan9118_phy_reset_hold;
 +    dc->vmsd = &vmstate_lan9118_phy;
 +}
 +
 +static const TypeInfo types[] = {
 +    {
 +        .name          = TYPE_LAN9118_PHY,
 +        .parent        = TYPE_SYS_BUS_DEVICE,
 +        .instance_size = sizeof(Lan9118PhyState),
 +        .instance_init = lan9118_phy_init,
 +        .class_init    = lan9118_phy_class_init,
 +    }
 +};
 +
 +DEFINE_TYPES(types)
 diff --git a/hw/net/Kconfig b/hw/net/Kconfig
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/trace-events
+--- a/hw/net/Kconfig
-+++ b/hw/arm/trace-events
++++ b/hw/net/Kconfig
-@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
+@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
- smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
+ config SMC91C111
- smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid %d (hits=%d, misses=%d, hit rate=%d)"
+     bool
- smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid %d (hits=%d, misses=%d, hit rate=%d)"
--smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
++config LAN9118_PHY
-+smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
++    bool
- smmuv3_cmdq_tlbi_nh(void) ""
++
- smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
+ config LAN9118
- smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
+     bool
- smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
++    select LAN9118_PHY
- smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
+     select PTIMER
--smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint64_t iova) "iommu mr=%s asid=%d iova=0x%"PRIx64
-+smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
+ config NE2000_ISA
+diff --git a/hw/net/meson.build b/hw/net/meson.build
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/meson.build
 +++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
  system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
  system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
 +system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
  system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
  system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
  system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
 --
-.20.1
+.34.1

-[PULL 01/27] hw/cpu/a9mpcore: Verify the machine use Cortex-A9 cores
+[PULL 02/72] hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
-From: Philippe Mathieu-Daudé <f4bug@amsat.org>
+From: Bernhard Beschow <shentey@gmail.com>
-The 'Cortex-A9MPCore internal peripheral' block can only be
+imx_fec models the same PHY as lan9118_phy. The code is almost the same with
-used with Cortex A5 and A9 cores. As we don't model the A5
+imx_fec having more logging and tracing. Merge these improvements into
-yet, simply check the machine cpu core is a Cortex A9. If
+lan9118_phy and reuse in imx_fec to fix the code duplication.
 not return an error.
-Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Some migration state how resides in the new device model which breaks migration
-Reviewed-by: Alistair Francis <alistair.francis@wdc.com>
+compatibility for the following machines:
-Message-id: 20200709152337.15533-1-f4bug@amsat.org
+* imx25-pdk
 * sabrelite
 * mcimx7d-sabre
 * mcimx6ul-evk
 Signed-off-by: Bernhard Beschow <shentey@gmail.com>
 Tested-by: Guenter Roeck <linux@roeck-us.net>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20241102125724.532843-3-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/cpu/a9mpcore.c | 12 +++++++++++-
+ include/hw/net/imx_fec.h |   9 ++-
-file changed, 11 insertions(+), 1 deletion(-)
+ hw/net/imx_fec.c         | 146 ++++-----------------------------------
  hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
  hw/net/Kconfig           |   1 +
  hw/net/trace-events      |  10 +--
 files changed, 85 insertions(+), 163 deletions(-)
-diff --git a/hw/cpu/a9mpcore.c b/hw/cpu/a9mpcore.c
+diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
 index XXXXXXX..XXXXXXX 100644
---- a/hw/cpu/a9mpcore.c
+--- a/include/hw/net/imx_fec.h
-+++ b/hw/cpu/a9mpcore.c
++++ b/include/hw/net/imx_fec.h
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
- #include "hw/irq.h"
+ #define TYPE_IMX_ENET "imx.enet"
- #include "hw/qdev-properties.h"
- #include "hw/core/cpu.h"
+ #include "hw/sysbus.h"
-+#include "cpu.h"
++#include "hw/net/lan9118_phy.h"
++#include "hw/irq.h"
- #define A9_GIC_NUM_PRIORITY_BITS    5
+ #include "net/net.h"
-@@ -XXX,XX +XXX,XX @@ static void a9mp_priv_realize(DeviceState *dev, Error **errp)
+ #define ENET_EIR               1
-                  *wdtbusdev;
+@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
-     int i;
+     uint32_t tx_descriptor[ENET_TX_RING_NUM];
-     bool has_el3;
+     uint32_t tx_ring_num;
-+    CPUState *cpu0;
-     Object *cpuobj;
+-    uint32_t phy_status;
+-    uint32_t phy_control;
-+    cpu0 = qemu_get_cpu(0);
+-    uint32_t phy_advertise;
-+    cpuobj = OBJECT(cpu0);
+-    uint32_t phy_int;
-+    if (strcmp(object_get_typename(cpuobj), ARM_CPU_TYPE_NAME("cortex-a9"))) {
+-    uint32_t phy_int_mask;
-+        /* We might allow Cortex-A5 once we model it */
++    Lan9118PhyState mii;
-+        error_setg(errp,
++    IRQState mii_irq;
-+                   "Cortex-A9MPCore peripheral can only use Cortex-A9 CPU");
+     uint32_t phy_num;
      bool phy_connected;
      struct IMXFECState *phy_consumer;
 diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/imx_fec.c
 +++ b/hw/net/imx_fec.c
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
  static const VMStateDescription vmstate_imx_eth = {
      .name = TYPE_IMX_FEC,
 -    .version_id = 2,
 -    .minimum_version_id = 2,
 +    .version_id = 3,
 +    .minimum_version_id = 3,
      .fields = (const VMStateField[]) {
          VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
          VMSTATE_UINT32(rx_descriptor, IMXFECState),
          VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
 -        VMSTATE_UINT32(phy_status, IMXFECState),
 -        VMSTATE_UINT32(phy_control, IMXFECState),
 -        VMSTATE_UINT32(phy_advertise, IMXFECState),
 -        VMSTATE_UINT32(phy_int, IMXFECState),
 -        VMSTATE_UINT32(phy_int_mask, IMXFECState),
          VMSTATE_END_OF_LIST()
      },
      .subsections = (const VMStateDescription * const []) {
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
      },
  };
 -#define PHY_INT_ENERGYON            (1 << 7)
 -#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
 -#define PHY_INT_FAULT               (1 << 5)
 -#define PHY_INT_DOWN                (1 << 4)
 -#define PHY_INT_AUTONEG_LP          (1 << 3)
 -#define PHY_INT_PARFAULT            (1 << 2)
 -#define PHY_INT_AUTONEG_PAGE        (1 << 1)
 -
  static void imx_eth_update(IMXFECState *s);
  /*
@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
   * For now we don't handle any GPIO/interrupt line, so the OS will
   * have to poll for the PHY status.
   */
 -static void imx_phy_update_irq(IMXFECState *s)
 +static void imx_phy_update_irq(void *opaque, int n, int level)
  {
 -    imx_eth_update(s);
 -}
 -
 -static void imx_phy_update_link(IMXFECState *s)
 -{
 -    /* Autonegotiation status mirrors link status.  */
 -    if (qemu_get_queue(s->nic)->link_down) {
 -        trace_imx_phy_update_link("down");
 -        s->phy_status &= ~0x0024;
 -        s->phy_int |= PHY_INT_DOWN;
 -    } else {
 -        trace_imx_phy_update_link("up");
 -        s->phy_status |= 0x0024;
 -        s->phy_int |= PHY_INT_ENERGYON;
 -        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
 -    }
 -    imx_phy_update_irq(s);
 +    imx_eth_update(opaque);
  }
  static void imx_eth_set_link(NetClientState *nc)
  {
 -    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
 -}
 -
 -static void imx_phy_reset(IMXFECState *s)
 -{
 -    trace_imx_phy_reset();
 -
 -    s->phy_status = 0x7809;
 -    s->phy_control = 0x3000;
 -    s->phy_advertise = 0x01e1;
 -    s->phy_int_mask = 0;
 -    s->phy_int = 0;
 -    imx_phy_update_link(s);
 +    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
 +                            nc->link_down);
  }
  static uint32_t imx_phy_read(IMXFECState *s, int reg)
  {
 -    uint32_t val;
      uint32_t phy = reg / 32;
      if (!s->phy_connected) {
@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
      reg %= 32;
 -    switch (reg) {
 -    case 0:     /* Basic Control */
 -        val = s->phy_control;
 -        break;
 -    case 1:     /* Basic Status */
 -        val = s->phy_status;
 -        break;
 -    case 2:     /* ID1 */
 -        val = 0x0007;
 -        break;
 -    case 3:     /* ID2 */
 -        val = 0xc0d1;
 -        break;
 -    case 4:     /* Auto-neg advertisement */
 -        val = s->phy_advertise;
 -        break;
 -    case 5:     /* Auto-neg Link Partner Ability */
 -        val = 0x0f71;
 -        break;
 -    case 6:     /* Auto-neg Expansion */
 -        val = 1;
 -        break;
 -    case 29:    /* Interrupt source.  */
 -        val = s->phy_int;
 -        s->phy_int = 0;
 -        imx_phy_update_irq(s);
 -        break;
 -    case 30:    /* Interrupt mask */
 -        val = s->phy_int_mask;
 -        break;
 -    case 17:
 -    case 18:
 -    case 27:
 -    case 31:
 -        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
 -                      TYPE_IMX_FEC, __func__, reg);
 -        val = 0;
 -        break;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
 -                      TYPE_IMX_FEC, __func__, reg);
 -        val = 0;
 -        break;
 -    }
 -
 -    trace_imx_phy_read(val, phy, reg);
 -
 -    return val;
 +    return lan9118_phy_read(&s->mii, reg);
  }
  static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
      reg %= 32;
 -    trace_imx_phy_write(val, phy, reg);
 -
 -    switch (reg) {
 -    case 0:     /* Basic Control */
 -        if (val & 0x8000) {
 -            imx_phy_reset(s);
 -        } else {
 -            s->phy_control = val & 0x7980;
 -            /* Complete autonegotiation immediately.  */
 -            if (val & 0x1000) {
 -                s->phy_status |= 0x0020;
 -            }
 -        }
 -        break;
 -    case 4:     /* Auto-neg advertisement */
 -        s->phy_advertise = (val & 0x2d7f) | 0x80;
 -        break;
 -    case 30:    /* Interrupt mask */
 -        s->phy_int_mask = val & 0xff;
 -        imx_phy_update_irq(s);
 -        break;
 -    case 17:
 -    case 18:
 -    case 27:
 -    case 31:
 -        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
 -                      TYPE_IMX_FEC, __func__, reg);
 -        break;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
 -                      TYPE_IMX_FEC, __func__, reg);
 -        break;
 -    }
 +    lan9118_phy_write(&s->mii, reg, val);
  }
  static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
      s->rx_descriptor = 0;
      memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
 -
 -    /* We also reset the PHY */
 -    imx_phy_reset(s);
  }
  static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
      sysbus_init_irq(sbd, &s->irq[0]);
      sysbus_init_irq(sbd, &s->irq[1]);
 +    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
 +    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
 +    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
 +        return;
 +    }
-+
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
-     scudev = DEVICE(&s->scu);
++
-     qdev_prop_set_uint32(scudev, "num-cpu", s->num_cpu);
+     qemu_macaddr_default_if_unset(&s->conf.macaddr);
-     if (!sysbus_realize(SYS_BUS_DEVICE(&s->scu), errp)) {
-@@ -XXX,XX +XXX,XX @@ static void a9mp_priv_realize(DeviceState *dev, Error **errp)
+     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
-     /* Make the GIC's TZ support match the CPUs. We assume that
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
-      * either all the CPUs have TZ, or none do.
+index XXXXXXX..XXXXXXX 100644
-      */
+--- a/hw/net/lan9118_phy.c
--    cpuobj = OBJECT(qemu_get_cpu(0));
++++ b/hw/net/lan9118_phy.c
-     has_el3 = object_property_find(cpuobj, "has_el3", NULL) &&
+@@ -XXX,XX +XXX,XX @@
-         object_property_get_bool(cpuobj, "has_el3", &error_abort);
+  * Copyright (c) 2009 CodeSourcery, LLC.
-     qdev_prop_set_bit(gicdev, "has-security-extensions", has_el3);
+  * Written by Paul Brook
   *
 + * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
 + *
   * This code is licensed under the GNU GPL v2
   *
   * Contributions after 2012-01-13 are licensed under the terms of the
@@ -XXX,XX +XXX,XX @@
  #include "hw/resettable.h"
  #include "migration/vmstate.h"
  #include "qemu/log.h"
 +#include "trace.h"
  #define PHY_INT_ENERGYON            (1 << 7)
  #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
      switch (reg) {
      case 0: /* Basic Control */
 -        return s->control;
 +        val = s->control;
 +        break;
      case 1: /* Basic Status */
 -        return s->status;
 +        val = s->status;
 +        break;
      case 2: /* ID1 */
 -        return 0x0007;
 +        val = 0x0007;
 +        break;
      case 3: /* ID2 */
 -        return 0xc0d1;
 +        val = 0xc0d1;
 +        break;
      case 4: /* Auto-neg advertisement */
 -        return s->advertise;
 +        val = s->advertise;
 +        break;
      case 5: /* Auto-neg Link Partner Ability */
 -        return 0x0f71;
 +        val = 0x0f71;
 +        break;
      case 6: /* Auto-neg Expansion */
 -        return 1;
 -        /* TODO 17, 18, 27, 29, 30, 31 */
 +        val = 1;
 +        break;
      case 29: /* Interrupt source. */
          val = s->ints;
          s->ints = 0;
          lan9118_phy_update_irq(s);
 -        return val;
 +        break;
      case 30: /* Interrupt mask */
 -        return s->int_mask;
 +        val = s->int_mask;
 +        break;
 +    case 17:
 +    case 18:
 +    case 27:
 +    case 31:
 +        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
 +                      __func__, reg);
 +        val = 0;
 +        break;
      default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "lan9118_phy_read: PHY read reg %d\n", reg);
 -        return 0;
 +        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
 +                      __func__, reg);
 +        val = 0;
 +        break;
      }
 +
 +    trace_lan9118_phy_read(val, reg);
 +
 +    return val;
  }
  void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
  {
 +    trace_lan9118_phy_write(val, reg);
 +
      switch (reg) {
      case 0: /* Basic Control */
          if (val & 0x8000) {
              lan9118_phy_reset(s);
 -            break;
 -        }
 -        s->control = val & 0x7980;
 -        /* Complete autonegotiation immediately. */
 -        if (val & 0x1000) {
 -            s->status |= 0x0020;
 +        } else {
 +            s->control = val & 0x7980;
 +            /* Complete autonegotiation immediately. */
 +            if (val & 0x1000) {
 +                s->status |= 0x0020;
 +            }
          }
          break;
      case 4: /* Auto-neg advertisement */
          s->advertise = (val & 0x2d7f) | 0x80;
          break;
 -        /* TODO 17, 18, 27, 31 */
      case 30: /* Interrupt mask */
          s->int_mask = val & 0xff;
          lan9118_phy_update_irq(s);
          break;
 +    case 17:
 +    case 18:
 +    case 27:
 +    case 31:
 +        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
 +                      __func__, reg);
 +        break;
      default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 +        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
 +                      __func__, reg);
 +        break;
      }
  }
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
      /* Autonegotiation status mirrors link status. */
      if (link_down) {
 +        trace_lan9118_phy_update_link("down");
          s->status &= ~0x0024;
          s->ints |= PHY_INT_DOWN;
      } else {
 +        trace_lan9118_phy_update_link("up");
          s->status |= 0x0024;
          s->ints |= PHY_INT_ENERGYON;
          s->ints |= PHY_INT_AUTONEG_COMPLETE;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
  void lan9118_phy_reset(Lan9118PhyState *s)
  {
 +    trace_lan9118_phy_reset();
 +
      s->control = 0x3000;
      s->status = 0x7809;
      s->advertise = 0x01e1;
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
      .version_id = 1,
      .minimum_version_id = 1,
      .fields = (const VMStateField[]) {
 -        VMSTATE_UINT16(control, Lan9118PhyState),
          VMSTATE_UINT16(status, Lan9118PhyState),
 +        VMSTATE_UINT16(control, Lan9118PhyState),
          VMSTATE_UINT16(advertise, Lan9118PhyState),
          VMSTATE_UINT16(ints, Lan9118PhyState),
          VMSTATE_UINT16(int_mask, Lan9118PhyState),
 diff --git a/hw/net/Kconfig b/hw/net/Kconfig
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/Kconfig
 +++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
  config IMX_FEC
      bool
 +    select LAN9118_PHY
  config CADENCE
      bool
 diff --git a/hw/net/trace-events b/hw/net/trace-events
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/trace-events
 +++ b/hw/net/trace-events
@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
  allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
  allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
 +# lan9118_phy.c
 +lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
 +lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
 +lan9118_phy_update_link(const char *s) "%s"
 +lan9118_phy_reset(void) ""
 +
  # lance.c
  lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
  lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
  i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
  # imx_fec.c
 -imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
  imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
 -imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
  imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
 -imx_phy_update_link(const char *s) "%s"
 -imx_phy_reset(void) ""
  imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
  imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
  imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
 --
-.20.1
+.34.1

-New patch
+[PULL 03/72] hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
+From: Bernhard Beschow <shentey@gmail.com>
+Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
+fixes the MSB of selector field to be zero, as specified in the datasheet.
+Fixes: 2a424990170b "LAN9118 emulation"
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241102125724.532843-4-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/net/lan9118_phy.c | 2 +-
+file changed, 1 insertion(+), 1 deletion(-)
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+         val = s->advertise;
+         break;
+     case 5: /* Auto-neg Link Partner Ability */
+-        val = 0x0f71;
++        val = 0x0fe1;
+         break;
+     case 6: /* Auto-neg Expansion */
+         val = 1;
+--
+.34.1

-New patch
+[PULL 04/72] hw/net/lan9118_phy: Reuse MII constants
+From: Bernhard Beschow <shentey@gmail.com>
+Prefer named constants over magic values for better readability.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Message-id: 20241102125724.532843-5-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ include/hw/net/mii.h |  6 +++++
+ hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
+files changed, 46 insertions(+), 23 deletions(-)
+diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/hw/net/mii.h
++++ b/include/hw/net/mii.h
+@@ -XXX,XX +XXX,XX @@
+ #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
+ #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
++#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
+ #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
+ #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
+ #define MII_ANAR_TXFD       (1 << 8)
+@@ -XXX,XX +XXX,XX @@
+ #define MII_ANAR_10FD       (1 << 6)
+ #define MII_ANAR_10         (1 << 5)
+ #define MII_ANAR_CSMACD     (1 << 0)
++#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
+ #define MII_ANLPAR_ACK      (1 << 14)
+ #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
+@@ -XXX,XX +XXX,XX @@
+ #define RTL8201CP_PHYID1    0x0000
+ #define RTL8201CP_PHYID2    0x8201
++/* SMSC LAN9118 */
++#define SMSCLAN9118_PHYID1  0x0007
++#define SMSCLAN9118_PHYID2  0xc0d1
++
+ /* RealTek 8211E */
+ #define RTL8211E_PHYID1     0x001c
+ #define RTL8211E_PHYID2     0xc915
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@
+ #include "qemu/osdep.h"
+ #include "hw/net/lan9118_phy.h"
++#include "hw/net/mii.h"
+ #include "hw/irq.h"
+ #include "hw/resettable.h"
+ #include "migration/vmstate.h"
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+     uint16_t val;
+     switch (reg) {
+-    case 0: /* Basic Control */
++    case MII_BMCR:
+         val = s->control;
+         break;
+-    case 1: /* Basic Status */
++    case MII_BMSR:
+         val = s->status;
+         break;
+-    case 2: /* ID1 */
+-        val = 0x0007;
++    case MII_PHYID1:
++        val = SMSCLAN9118_PHYID1;
+         break;
+-    case 3: /* ID2 */
+-        val = 0xc0d1;
++    case MII_PHYID2:
++        val = SMSCLAN9118_PHYID2;
+         break;
+-    case 4: /* Auto-neg advertisement */
++    case MII_ANAR:
+         val = s->advertise;
+         break;
+-    case 5: /* Auto-neg Link Partner Ability */
+-        val = 0x0fe1;
++    case MII_ANLPAR:
++        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
++              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
++              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
+         break;
+-    case 6: /* Auto-neg Expansion */
+-        val = 1;
++    case MII_ANER:
++        val = MII_ANER_NWAY;
+         break;
+     case 29: /* Interrupt source. */
+         val = s->ints;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+     trace_lan9118_phy_write(val, reg);
+     switch (reg) {
+-    case 0: /* Basic Control */
+-        if (val & 0x8000) {
++    case MII_BMCR:
++        if (val & MII_BMCR_RESET) {
+             lan9118_phy_reset(s);
+         } else {
+-            s->control = val & 0x7980;
++            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
++                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
++                                MII_BMCR_CTST);
+             /* Complete autonegotiation immediately. */
+-            if (val & 0x1000) {
+-                s->status |= 0x0020;
++            if (val & MII_BMCR_AUTOEN) {
++                s->status |= MII_BMSR_AN_COMP;
+             }
+         }
+         break;
+-    case 4: /* Auto-neg advertisement */
+-        s->advertise = (val & 0x2d7f) | 0x80;
++    case MII_ANAR:
++        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
++                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
++                               MII_ANAR_SELECT))
++                     | MII_ANAR_TX;
+         break;
+     case 30: /* Interrupt mask */
+         s->int_mask = val & 0xff;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+     /* Autonegotiation status mirrors link status. */
+     if (link_down) {
+         trace_lan9118_phy_update_link("down");
+-        s->status &= ~0x0024;
++        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
+         s->ints |= PHY_INT_DOWN;
+     } else {
+         trace_lan9118_phy_update_link("up");
+-        s->status |= 0x0024;
++        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
+         s->ints |= PHY_INT_ENERGYON;
+         s->ints |= PHY_INT_AUTONEG_COMPLETE;
+     }
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
+ {
+     trace_lan9118_phy_reset();
+-    s->control = 0x3000;
+-    s->status = 0x7809;
+-    s->advertise = 0x01e1;
++    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
++    s->status = MII_BMSR_100TX_FD
++                | MII_BMSR_100TX_HD
++                | MII_BMSR_10T_FD
++                | MII_BMSR_10T_HD
++                | MII_BMSR_AUTONEG
++                | MII_BMSR_EXTCAP;
++    s->advertise = MII_ANAR_TXFD
++                   | MII_ANAR_TX
++                   | MII_ANAR_10FD
++                   | MII_ANAR_10
++                   | MII_ANAR_CSMACD;
+     s->int_mask = 0;
+     s->ints = 0;
+     lan9118_phy_update_link(s, s->link_down);
+--
+.34.1

-New patch
+[PULL 05/72] hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
+From: Bernhard Beschow <shentey@gmail.com>
+The real device advertises this mode and the device model already advertises
+mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
+make the model more realistic.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Message-id: 20241102125724.532843-6-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/net/lan9118_phy.c | 4 ++--
+file changed, 2 insertions(+), 2 deletions(-)
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+         break;
+     case MII_ANAR:
+         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
+-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
+-                               MII_ANAR_SELECT))
++                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
++                               MII_ANAR_10 | MII_ANAR_SELECT))
+                      | MII_ANAR_TX;
+         break;
+     case 30: /* Interrupt mask */
+--
+.34.1

-New patch
+[PULL 06/72] fpu: handle raising Invalid for infzero in pick_nan_muladd
+For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
+Invalid for the multiplication of 0 by infinity.  Currently we handle
+this in the per-architecture ifdef ladder in pickNaNMulAdd().
+However, since this isn't really architecture specific we can hoist
+it up to the generic code.
+For the cases where the infzero test in pickNaNMulAdd was
+returning 2, we can delete the check entirely and allow the
+code to fall into the normal pick-a-NaN handling, because this
+will return 2 anyway (input 'c' being the only NaN in this case).
+For the cases where infzero was returning 3 to indicate "return
+the default NaN", we must retain that "return 3".
+For Arm, this looks like it might be a behaviour change because we
+used to set float_flag_invalid | float_flag_invalid_imz only if C is
+a quiet NaN.  However, it is not, because Arm target code never looks
+at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
+already raised float_flag_invalid via the "abc_mask &
+float_cmask_snan" check in pick_nan_muladd.
+For any target architecture using the "default implementation" at the
+bottom of the ifdef, this is a behaviour change but will be fixing a
+bug (where we failed to raise the Invalid exception for (0 * inf +
+QNaN).  The architectures using the default case are:
+ * hppa
+ * i386
+ * sh4
+ * tricore
+The x86, Tricore and SH4 CPU architecture manuals are clear that this
+should have raised Invalid; HPPA is a bit vaguer but still seems
+clear enough.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 13 +++++++------
+ fpu/softfloat-specialize.c.inc | 29 +----------------------------
+files changed, 8 insertions(+), 34 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+                                             int ab_mask, int abc_mask)
+ {
+     int which;
++    bool infzero = (ab_mask == float_cmask_infzero);
+     if (unlikely(abc_mask & float_cmask_snan)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
+     }
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
+-                          ab_mask == float_cmask_infzero, s);
++    if (infzero) {
++        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
++        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
++    }
++
++    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+     if (s->default_nan_mode || which == 3) {
+-        /*
+-         * Note that this check is after pickNaNMulAdd so that function
+-         * has an opportunity to set the Invalid flag for infzero.
+-         */
+         parts_default_nan(a, s);
+         return a;
+     }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * the default NaN
+      */
+     if (infzero && is_qnan(c_cls)) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+         return 3;
+     }
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * case sets InvalidOp and returns the default NaN
+          */
+         if (infzero) {
+-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+             return 3;
+         }
+         /* Prefer sNaN over qNaN, in the a, b, c order. */
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+          */
+-        if (infzero) {
+-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-            return 2;
+-        }
+         /* Prefer sNaN over qNaN, in the c, a, b order. */
+         if (is_snan(c_cls)) {
+             return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+      * case sets InvalidOp and returns the input value 'c'
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
++
+     /* Prefer sNaN over qNaN, in the c, a, b order. */
+     if (is_snan(c_cls)) {
+         return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * to return an input NaN if we have one (ie c) rather than generating
+      * a default NaN
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         return 1;
+     }
+ #elif defined(TARGET_RISCV)
+-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-    }
+     return 3; /* default NaN */
+ #elif defined(TARGET_S390X)
+     if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+         return 3;
+     }
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         return 2;
+     }
+ #elif defined(TARGET_SPARC)
+-    /* For (inf,0,nan) return c. */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     /* Prefer SNaN over QNaN, order C, B, A. */
+     if (is_snan(c_cls)) {
+         return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
+      * an input NaN if we have one (ie c).
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     if (status->use_first_nan) {
+         if (is_nan(a_cls)) {
+             return 0;
+--
+.34.1

-New patch
+[PULL 07/72] fpu: Check for default_nan_mode before calling pickNaNMulAdd
+If the target sets default_nan_mode then we're always going to return
+the default NaN, and pickNaNMulAdd() no longer has any side effects.
+For consistency with pickNaN(), check for default_nan_mode before
+calling pickNaNMulAdd().
+When we convert pickNaNMulAdd() to allow runtime selection of the NaN
+propagation rule, this means we won't have to make the targets which
+use default_nan_mode also set a propagation rule.
+Since RiscV always uses default_nan_mode, this allows us to remove
+its ifdef case from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 8 ++++++--
+ fpu/softfloat-specialize.c.inc | 9 +++++++--
+files changed, 13 insertions(+), 4 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+     }
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    if (s->default_nan_mode) {
++        which = 3;
++    } else {
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    }
+-    if (s->default_nan_mode || which == 3) {
++    if (which == 3) {
+         parts_default_nan(a, s);
+         return a;
+     }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
++    /*
++     * We guarantee not to require the target to tell us how to
++     * pick a NaN if we're always returning the default NaN.
++     * But if we're not in default-NaN mode then the target must
++     * specify.
++     */
++    assert(!status->default_nan_mode);
+ #if defined(TARGET_ARM)
+     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
+      * the default NaN
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     } else {
+         return 1;
+     }
+-#elif defined(TARGET_RISCV)
+-    return 3; /* default NaN */
+ #elif defined(TARGET_S390X)
+     if (infzero) {
+         return 3;
+--
+.34.1

-[PULL 18/27] target/arm: Do M-profile NOCP checks early and via decodetree
+[PULL 08/72] softfloat: Allow runtime choice of inf * 0 + NaN result
-For M-profile CPUs, the architecture specifies that the NOCP
+IEEE 758 does not define a fixed rule for what NaN to return in
-exception when a coprocessor is not present or disabled should cover
+the case of a fused multiply-add of inf * 0 + NaN. Different
-the entire wide range of coprocessor-space encodings, and should take
+architectures thus do different things:
-precedence over UNDEF exceptions.  (This is the opposite of
+ * some return the default NaN
-A-profile, where checking for a disabled FPU has to happen last.)
+ * some return the input NaN
+ * Arm returns the default NaN if the input NaN is quiet,
-Implement this with decodetree patterns that cover the specified
+   and the input NaN if it is signalling
-ranges of the encoding space.  There are a few instructions (VLLDM,
-VLSTM, and in v8.1 also VSCCLRM) which are in copro-space but must
+We want to make this logic be runtime selected rather than
-not be NOCP'd: these must be handled also in the new m-nocp.decode so
+hardcoded into the binary, because:
-they take precedence.
+ * this will let us have multiple targets in one QEMU binary
+ * the Arm FEAT_AFP architectural feature includes letting
-This is a minor behaviour change: for unallocated insn patterns in
+   the guest select a NaN propagation rule at runtime
-the VFP area (cp=10,11) we will now NOCP rather than UNDEF when the
-FPU is disabled.
+In this commit we add an enum for the propagation rule, the field in
+float_status, and the corresponding getters and setters.  We change
-As well as giving us the correct architectural behaviour for v8.1M
+pickNaNMulAdd to honour this, but because all targets still leave
-and the recommended behaviour for v8.0M, this refactoring also
+this field at its default 0 value, the fallback logic will pick the
-removes the old NOCP handling from the remains of the 'legacy
+rule type with the old ifdef ladder.
-decoder' in disas_thumb2_insn(), paving the way for cleaning that up.
+Note that four architectures both use the muladd softfloat functions
-Since we don't currently have a v8.1M feature bit or any v8.1M CPUs,
+and did not have a branch of the ifdef ladder to specify their
-the minor changes to this logic that we'll need for v8.1M are marked
+behaviour (and so were ending up with the "default" case, probably
-up with TODO comments.
+wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
 default_nan_mode, and so will never get into pickNaNMulAdd().  For
 HPPA and i386 we retain the same behaviour as the old default-case,
 which is to not ever return the default NaN.  This might not be
 correct but it is not a behaviour change.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-6-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
 ---
- target/arm/m-nocp.decode       | 42 +++++++++++++++++++++++++++
+ include/fpu/softfloat-helpers.h | 11 ++++
- target/arm/vfp.decode          |  2 --
+ include/fpu/softfloat-types.h   | 23 +++++++++
- target/arm/translate.c         | 30 ++++++++++----------
+ fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
- target/arm/meson.build         |  1 +
+files changed, 95 insertions(+), 30 deletions(-)
- target/arm/translate-vfp.c.inc | 52 +++++++++++++++++++++++++++-------
-files changed, 100 insertions(+), 27 deletions(-)
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
- create mode 100644 target/arm/m-nocp.decode
+index XXXXXXX..XXXXXXX 100644
+--- a/include/fpu/softfloat-helpers.h
-diff --git a/target/arm/m-nocp.decode b/target/arm/m-nocp.decode
++++ b/include/fpu/softfloat-helpers.h
-new file mode 100644
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
-index XXXXXXX..XXXXXXX
+     status->float_2nan_prop_rule = rule;
---- /dev/null
+ }
-+++ b/target/arm/m-nocp.decode
-@@ -XXX,XX +XXX,XX @@
++static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
-+# M-profile UserFault.NOCP exception handling
++                                             float_status *status)
 +#
 +#  Copyright (c) 2020 Linaro, Ltd
 +#
 +# This library is free software; you can redistribute it and/or
 +# modify it under the terms of the GNU Lesser General Public
 +# License as published by the Free Software Foundation; either
 +# version 2.1 of the License, or (at your option) any later version.
 +#
 +# This library is distributed in the hope that it will be useful,
 +# but WITHOUT ANY WARRANTY; without even the implied warranty of
 +# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the GNU
 +# Lesser General Public License for more details.
 +#
 +# You should have received a copy of the GNU Lesser General Public
 +# License along with this library; if not, see <http://www.gnu.org/licenses/>.
 +
 +#
 +# This file is processed by scripts/decodetree.py
 +#
 +# For M-profile, the architecture specifies that NOCP UsageFaults
 +# should take precedence over UNDEF faults over the whole wide
 +# range of coprocessor-space encodings, with the exception of
 +# VLLDM and VLSTM. (Compare v8.1M IsCPInstruction() pseudocode and
 +# v8M Arm ARM rule R_QLGM.) This isn't mandatory for v8.0M but we choose
 +# to behave the same as v8.1M.
 +# This decode is handled before any others (and in particular before
 +# decoding FP instructions which are in the coprocessor space).
 +# If the coprocessor is not present or disabled then we will generate
 +# the NOCP exception; otherwise we let the insn through to the main decode.
 +
 +{
-+  # Special cases which do not take an early NOCP: VLLDM and VLSTM
++    status->float_infzeronan_rule = rule;
 +  VLLDM_VLSTM  1110 1100 001 l:1 rn:4 0000 1010 0000 0000
 +  # TODO: VSCCLRM (new in v8.1M) is similar:
 +  #VSCCLRM      1110 1100 1-01 1111 ---- 1011 ---- ---0
 +
 +  NOCP         111- 1110 ---- ---- ---- cp:4 ---- ----
 +  NOCP         111- 110- ---- ---- ---- cp:4 ---- ----
 +  # TODO: From v8.1M onwards we will also want this range to NOCP
 +  #NOCP_8_1     111- 1111 ---- ---- ---- ---- ---- ---- cp=10
 +}
-diff --git a/target/arm/vfp.decode b/target/arm/vfp.decode
++
  static inline void set_flush_to_zero(bool val, float_status *status)
  {
      status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
      return status->float_2nan_prop_rule;
  }
 +static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 +{
 +    return status->float_infzeronan_rule;
 +}
 +
  static inline bool get_flush_to_zero(float_status *status)
  {
      return status->flush_to_zero;
 diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/vfp.decode
+--- a/include/fpu/softfloat-types.h
-+++ b/target/arm/vfp.decode
++++ b/include/fpu/softfloat-types.h
-@@ -XXX,XX +XXX,XX @@ VCVT_sp_int  ---- 1110 1.11 110 s:1 .... 1010 rz:1 1.0 .... \
+@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
-              vd=%vd_sp vm=%vm_sp
+     float_2nan_prop_x87,
- VCVT_dp_int  ---- 1110 1.11 110 s:1 .... 1011 rz:1 1.0 .... \
+ } Float2NaNPropRule;
-              vd=%vd_sp vm=%vm_dp
--
++/*
--VLLDM_VLSTM  1110 1100 001 l:1 rn:4 0000 1010 0000 0000
++ * Rule for result of fused multiply-add 0 * Inf + NaN.
-diff --git a/target/arm/translate.c b/target/arm/translate.c
++ * This must be a NaN, but implementations differ on whether this
 + * is the input NaN or the default NaN.
 + *
 + * You don't need to set this if default_nan_mode is enabled.
 + * When not in default-NaN mode, it is an error for the target
 + * not to set the rule in float_status if it uses muladd, and we
 + * will assert if we need to handle an input NaN and no rule was
 + * selected.
 + */
 +typedef enum __attribute__((__packed__)) {
 +    /* No propagation rule specified */
 +    float_infzeronan_none = 0,
 +    /* Result is never the default NaN (so always the input NaN) */
 +    float_infzeronan_dnan_never,
 +    /* Result is always the default NaN */
 +    float_infzeronan_dnan_always,
 +    /* Result is the default NaN if the input NaN is quiet */
 +    float_infzeronan_dnan_if_qnan,
 +} FloatInfZeroNaNRule;
 +
  /*
   * Floating Point Status. Individual architectures may maintain
   * several versions of float_status for different functions. The
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
      bool flush_to_zero;
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/translate.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static TCGv_ptr vfp_reg_ptr(bool dp, int reg)
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
- #define ARM_CP_RW_BIT   (1 << 20)
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
- /* Include the VFP and Neon decoders */
+ {
-+#include "decode-m-nocp.c.inc"
++    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
- #include "translate-vfp.c.inc"
++
- #include "translate-neon.c.inc"
+     /*
+      * We guarantee not to require the target to tell us how to
-@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
+      * pick a NaN if we're always returning the default NaN.
-         ARCH(6T2);
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-     }
+      * specify.
+      */
-+    if (arm_dc_feature(s, ARM_FEATURE_M)) {
+     assert(!status->default_nan_mode);
-+        /*
++
-+         * NOCP takes precedence over any UNDEF for (almost) the
++    if (rule == float_infzeronan_none) {
-+         * entire wide range of coprocessor-space encodings, so check
++        /*
-+         * for it first before proceeding to actually decode eg VFP
++         * Temporarily fall back to ifdef ladder
-+         * insns. This decode also handles the few insns which are
++         */
-+         * in copro space but do not have NOCP checks (eg VLLDM, VLSTM).
+ #if defined(TARGET_ARM)
-+         */
+-    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
-+        if (disas_m_nocp(s, insn)) {
+-     * the default NaN
-+            return;
+-     */
 -    if (infzero && is_qnan(c_cls)) {
 -        return 3;
 +        /*
 +         * For ARM, the (inf,zero,qnan) case returns the default NaN,
 +         * but (inf,zero,snan) returns the input NaN.
 +         */
 +        rule = float_infzeronan_dnan_if_qnan;
 +#elif defined(TARGET_MIPS)
 +        if (snan_bit_is_one(status)) {
 +            /*
 +             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the default NaN
 +             */
 +            rule = float_infzeronan_dnan_always;
 +        } else {
 +            /*
 +             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the input value 'c'
 +             */
 +            rule = float_infzeronan_dnan_never;
 +        }
 +#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
 +    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
 +    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
 +        /*
 +         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 +         * case sets InvalidOp and returns the input value 'c'
 +         */
 +        /*
 +         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
 +         * to return an input NaN if we have one (ie c) rather than generating
 +         * a default NaN
 +         */
 +        rule = float_infzeronan_dnan_never;
 +#elif defined(TARGET_S390X)
 +        rule = float_infzeronan_dnan_always;
 +#endif
      }
 +    if (infzero) {
 +        /*
 +         * Inf * 0 + NaN -- some implementations return the default NaN here,
 +         * and some return the input NaN.
 +         */
 +        switch (rule) {
 +        case float_infzeronan_dnan_never:
 +            return 2;
 +        case float_infzeronan_dnan_always:
 +            return 3;
 +        case float_infzeronan_dnan_if_qnan:
 +            return is_qnan(c_cls) ? 3 : 2;
 +        default:
 +            g_assert_not_reached();
 +        }
 +    }
 +
-     if ((insn & 0xef000000) == 0xef000000) {
++#if defined(TARGET_ARM)
-         /*
++
-          * T32 encodings 0b111p_1111_qqqq_qqqq_qqqq_qqqq_qqqq_qqqq
+     /* This looks different from the ARM ARM pseudocode, because the ARM ARM
-@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
+      * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
-         /* Coprocessor.  */
+      */
-         if (arm_dc_feature(s, ARM_FEATURE_M)) {
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-             /* 0b111x_11xx_xxxx_xxxx_xxxx_xxxx_xxxx_xxxx */
+     }
--            if (extract32(insn, 24, 2) == 3) {
+ #elif defined(TARGET_MIPS)
--                goto illegal_op; /* op0 = 0b11 : unallocated */
+     if (snan_bit_is_one(status)) {
--            }
+-        /*
 -         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the default NaN
 -         */
 -        if (infzero) {
 -            return 3;
 -        }
          /* Prefer sNaN over qNaN, in the a, b, c order. */
          if (is_snan(a_cls)) {
              return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              return 2;
          }
      } else {
 -        /*
 -         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the input value 'c'
 -         */
          /* Prefer sNaN over qNaN, in the c, a, b order. */
          if (is_snan(c_cls)) {
              return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
  #elif defined(TARGET_LOONGARCH64)
 -    /*
 -     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 -     * case sets InvalidOp and returns the input value 'c'
 -     */
 -
--            if (((insn >> 8) & 0xe) == 10 &&
+     /* Prefer sNaN over qNaN, in the c, a, b order. */
--                dc_isar_feature(aa32_fpsp_v2, s)) {
+     if (is_snan(c_cls)) {
--                /* FP, and the CPU supports it */
+         return 2;
--                goto illegal_op;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
--            } else {
+         return 1;
--                /* All other insns: NOCP */
+     }
--                gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
+ #elif defined(TARGET_PPC)
--                                   syn_uncategorized(),
+-    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
--                                   default_exception_el(s));
+-     * to return an input NaN if we have one (ie c) rather than generating
--            }
+-     * a default NaN
--            break;
+-     */
-+            goto illegal_op;
+-
-         }
+     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-         if (((insn >> 24) & 3) == 3) {
+      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-             /* Neon DP, but failed disas_neon_dp() */
+      */
-diff --git a/target/arm/meson.build b/target/arm/meson.build
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-index XXXXXXX..XXXXXXX 100644
+         return 1;
---- a/target/arm/meson.build
+     }
-+++ b/target/arm/meson.build
+ #elif defined(TARGET_S390X)
-@@ -XXX,XX +XXX,XX @@ gen = [
+-    if (infzero) {
-   decodetree.process('neon-ls.decode', extra_args: '--static-decode=disas_neon_ls'),
+-        return 3;
-   decodetree.process('vfp.decode', extra_args: '--static-decode=disas_vfp'),
+-    }
-   decodetree.process('vfp-uncond.decode', extra_args: '--static-decode=disas_vfp_uncond'),
+-
-+  decodetree.process('m-nocp.decode', extra_args: '--static-decode=disas_m_nocp'),
+     if (is_snan(a_cls)) {
-   decodetree.process('a32.decode', extra_args: '--static-decode=disas_a32'),
+         return 0;
-   decodetree.process('a32-uncond.decode', extra_args: '--static-decode=disas_a32_uncond'),
+     } else if (is_snan(b_cls)) {
    decodetree.process('t32.decode', extra_args: '--static-decode=disas_t32'),
 diff --git a/target/arm/translate-vfp.c.inc b/target/arm/translate-vfp.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-vfp.c.inc
 +++ b/target/arm/translate-vfp.c.inc
@@ -XXX,XX +XXX,XX @@ static inline long vfp_f16_offset(unsigned reg, bool top)
  static bool full_vfp_access_check(DisasContext *s, bool ignore_vfp_enabled)
  {
      if (s->fp_excp_el) {
 -        if (arm_dc_feature(s, ARM_FEATURE_M)) {
 -            gen_exception_insn(s, s->pc_curr, EXCP_NOCP, syn_uncategorized(),
 -                               s->fp_excp_el);
 -        } else {
 -            gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
 -                               syn_fp_access_trap(1, 0xe, false),
 -                               s->fp_excp_el);
 -        }
 +        /* M-profile handled this earlier, in disas_m_nocp() */
 +        assert (!arm_dc_feature(s, ARM_FEATURE_M));
 +        gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
 +                           syn_fp_access_trap(1, 0xe, false),
 +                           s->fp_excp_el);
          return false;
      }
@@ -XXX,XX +XXX,XX @@ static bool trans_VLLDM_VLSTM(DisasContext *s, arg_VLLDM_VLSTM *a)
          !arm_dc_feature(s, ARM_FEATURE_V8)) {
          return false;
      }
 -    /* If not secure, UNDEF. */
 +    /*
 +     * If not secure, UNDEF. We must emit code for this
 +     * rather than returning false so that this takes
 +     * precedence over the m-nocp.decode NOCP fallback.
 +     */
      if (!s->v8m_secure) {
 -        return false;
 +        unallocated_encoding(s);
 +        return true;
      }
      /* If no fpu, NOP. */
      if (!dc_isar_feature(aa32_vfp, s)) {
@@ -XXX,XX +XXX,XX @@ static bool trans_VLLDM_VLSTM(DisasContext *s, arg_VLLDM_VLSTM *a)
      s->base.is_jmp = DISAS_UPDATE_EXIT;
      return true;
  }
 +
 +static bool trans_NOCP(DisasContext *s, arg_NOCP *a)
 +{
 +    /*
 +     * Handle M-profile early check for disabled coprocessor:
 +     * all we need to do here is emit the NOCP exception if
 +     * the coprocessor is disabled. Otherwise we return false
 +     * and the real VFP/etc decode will handle the insn.
 +     */
 +    assert(arm_dc_feature(s, ARM_FEATURE_M));
 +
 +    if (a->cp == 11) {
 +        a->cp = 10;
 +    }
 +    /* TODO: in v8.1M cp 8, 9, 14, 15 also are governed by the cp10 enable */
 +
 +    if (a->cp != 10) {
 +        gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
 +                           syn_uncategorized(), default_exception_el(s));
 +        return true;
 +    }
 +
 +    if (s->fp_excp_el != 0) {
 +        gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
 +                           syn_uncategorized(), s->fp_excp_el);
 +        return true;
 +    }
 +
 +    return false;
 +}
 --
-.20.1
+.34.1

-New patch
+[PULL 09/72] tests/fp: Explicitly set inf-zero-nan rule
+Explicitly set a rule in the softfloat tests for the inf-zero-nan
+muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
+and so we should select here the Arm rule of
+float_infzeronan_dnan_if_qnan.
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 5 +++++
+ tests/fp/fp-test.c  | 5 +++++
+files changed, 10 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+ {
+     bench_func_t f;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+ {
+     unsigned int i;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+     verCases_maxErrorCount = n_max_errors;
+--
+.34.1

-New patch
+[PULL 10/72] target/arm: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the Arm target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 3 +++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 4 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
++ *    and the input NaN if it is signalling
+  */
+ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_ARM)
+-        /*
+-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+-         * but (inf,zero,snan) returns the input NaN.
+-         */
+-        rule = float_infzeronan_dnan_if_qnan;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             /*
+              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 11/72] target/s390: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for s390, so we
+can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_infzeronan_rule(float_infzeronan_dnan_always,
++                                  &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * a default NaN
+          */
+         rule = float_infzeronan_dnan_never;
+-#elif defined(TARGET_S390X)
+-        rule = float_infzeronan_dnan_always;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 12/72] target/ppc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the PPC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 7 +------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
++     * to return an input NaN if we have one (ie c) rather than generating
++     * a default NaN
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+              */
+             rule = float_infzeronan_dnan_never;
+         }
+-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
++#elif defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+          */
+-        /*
+-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+-         * to return an input NaN if we have one (ie c) rather than generating
+-         * a default NaN
+-         */
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 13/72] target/mips: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the MIPS target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h       |  9 +++++++++
+ target/mips/msa.c              |  4 ++++
+ fpu/softfloat-specialize.c.inc | 16 +---------------
+files changed, 14 insertions(+), 15 deletions(-)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
+ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+ {
+     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
++    FloatInfZeroNaNRule izn_rule;
+     /*
+      * With nan2008, SNaNs are silenced in the usual way.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+      */
+     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
+     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
++    /*
++     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
++     * case sets InvalidOp and returns the default NaN.
++     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'.
++     */
++    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
++    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+ }
+ static inline void restore_fp_status(CPUMIPSState *env)
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* set proper signanling bit meaning ("1" means "quiet") */
+     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
++
++    /* Inf * 0 + NaN returns the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never,
++                              &env->active_tc.msa_fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_MIPS)
+-        if (snan_bit_is_one(status)) {
+-            /*
+-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the default NaN
+-             */
+-            rule = float_infzeronan_dnan_always;
+-        } else {
+-            /*
+-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the input value 'c'
+-             */
+-            rule = float_infzeronan_dnan_never;
+-        }
+-#elif defined(TARGET_SPARC) || \
++#if defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+--
+.34.1

-New patch
+[PULL 14/72] target/sparc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the SPARC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 +--
+files changed, 3 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_SPARC) || \
+-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 15/72] target/xtensa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the xtensa target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
+---
+ target/xtensa/cpu.c            | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/xtensa/cpu.c
++++ b/target/xtensa/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
+     reset_mmu(env);
+     cs->halted = env->runstall;
+ #endif
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_no_signaling_nans(!dfpu, &env->fp_status);
+     xtensa_use_first_nan(env, !dfpu);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 16/72] target/x86: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the x86 target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 8 insertions(+), 1 deletion(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
++    /*
++     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
++     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
++     * specified -- for 0 * inf + NaN the input NaN is selected, and if
++     * there are multiple input NaNs they are selected in the order a, b, c.
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * Temporarily fall back to ifdef ladder
+          */
+ #if defined(TARGET_HPPA) || \
+-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
++    defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+--
+.34.1

-New patch
+[PULL 17/72] target/loongarch: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the loongarch target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 5 +++++
+ fpu/softfloat-specialize.c.inc    | 7 +------
+files changed, 6 insertions(+), 6 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+                             &env->fp_status);
+     set_flush_to_zero(0, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_HPPA) || \
+-    defined(TARGET_LOONGARCH)
+-        /*
+-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+-         * case sets InvalidOp and returns the input value 'c'
+-         */
++#if defined(TARGET_HPPA)
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 18/72] target/hppa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the HPPA target,
+so we can remove the ifdef from pickNaNMulAdd().
+As this is the last target to be converted to explicitly setting
+the rule, we can remove the fallback code in pickNaNMulAdd()
+entirely.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       |  2 ++
+ fpu/softfloat-specialize.c.inc | 13 +------------
+files changed, 3 insertions(+), 12 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
+-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+-
+     /*
+      * We guarantee not to require the target to tell us how to
+      * pick a NaN if we're always returning the default NaN.
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      */
+     assert(!status->default_nan_mode);
+-    if (rule == float_infzeronan_none) {
+-        /*
+-         * Temporarily fall back to ifdef ladder
+-         */
+-#if defined(TARGET_HPPA)
+-        rule = float_infzeronan_dnan_never;
+-#endif
+-    }
+-
+     if (infzero) {
+         /*
+          * Inf * 0 + NaN -- some implementations return the default NaN here,
+          * and some return the input NaN.
+          */
+-        switch (rule) {
++        switch (status->float_infzeronan_rule) {
+         case float_infzeronan_dnan_never:
+             return 2;
+         case float_infzeronan_dnan_always:
+--
+.34.1

-[PULL 27/27] target/arm: Use correct FPST for VCMLA, VCADD on fp16
+[PULL 19/72] softfloat: Pass have_snan to pickNaNMulAdd
-When we implemented the VCMLA and VCADD insns we put in the
+The new implementation of pickNaNMulAdd() will find it convenient
-code to handle fp16, but left it using the standard fp status
+to know whether at least one of the three arguments to the muladd
-flags. Correct them to use FPST_STD_F16 for fp16 operations.
+was a signaling NaN. We already calculate that in the caller,
 so pass it in as a new bool have_snan.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
 Message-id: 20200806104453.30393-5-peter.maydell@linaro.org
 ---
- target/arm/translate-neon.c.inc | 6 +++---
+ fpu/softfloat-parts.c.inc      | 5 +++--
-file changed, 3 insertions(+), 3 deletions(-)
+ fpu/softfloat-specialize.c.inc | 2 +-
 files changed, 4 insertions(+), 3 deletions(-)
-diff --git a/target/arm/translate-neon.c.inc b/target/arm/translate-neon.c.inc
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate-neon.c.inc
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/translate-neon.c.inc
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
  {
      int which;
      bool infzero = (ab_mask == float_cmask_infzero);
 +    bool have_snan = (abc_mask & float_cmask_snan);
 -    if (unlikely(abc_mask & float_cmask_snan)) {
 +    if (unlikely(have_snan)) {
          float_raise(float_flag_invalid | float_flag_invalid_snan, s);
      }
-     opr_sz = (1 + a->q) * 8;
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
--    fpst = fpstatus_ptr(FPST_STD);
+     if (s->default_nan_mode) {
-+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
+         which = 3;
-     fn_gvec_ptr = a->size ? gen_helper_gvec_fcmlas : gen_helper_gvec_fcmlah;
+     } else {
-     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
+-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
-                        vfp_reg_offset(1, a->vn),
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCADD(DisasContext *s, arg_VCADD *a)
      }
-     opr_sz = (1 + a->q) * 8;
+     if (which == 3) {
--    fpst = fpstatus_ptr(FPST_STD);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
+index XXXXXXX..XXXXXXX 100644
-     fn_gvec_ptr = a->size ? gen_helper_gvec_fcadds : gen_helper_gvec_fcaddh;
+--- a/fpu/softfloat-specialize.c.inc
-     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
++++ b/fpu/softfloat-specialize.c.inc
-                        vfp_reg_offset(1, a->vn),
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA_scalar(DisasContext *s, arg_VCMLA_scalar *a)
+ | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-     fn_gvec_ptr = (a->size ? gen_helper_gvec_fcmlas_idx
+ *----------------------------------------------------------------------------*/
-                    : gen_helper_gvec_fcmlah_idx);
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-     opr_sz = (1 + a->q) * 8;
+-                         bool infzero, float_status *status)
--    fpst = fpstatus_ptr(FPST_STD);
++                         bool infzero, bool have_snan, float_status *status)
-+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
+ {
-     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
+     /*
-                        vfp_reg_offset(1, a->vn),
+      * We guarantee not to require the target to tell us how to
                         vfp_reg_offset(1, a->vm),
 --
-.20.1
+.34.1

-[PULL 15/27] target/arm: Separate decode from handling of coproc insns
+[PULL 20/72] softfloat: Allow runtime choice of NaN propagation for muladd
-As a prelude to making coproc insns use decodetree, split out the
+IEEE 758 does not define a fixed rule for which NaN to pick as the
-part of disas_coproc_insn() which does instruction decoding from the
+result if both operands of a 3-operand fused multiply-add operation
-part which does the actual work, and make do_coproc_insn() handle the
+are NaNs.  As a result different architectures have ended up with
-UNDEF-on-bad-permissions and similar cases itself rather than
+different rules for propagating NaNs.
-returning 1 to eventually percolate up to a callsite that calls
-unallocated_encoding() for it.
+QEMU currently hardcodes the NaN propagation logic into the binary
 because pickNaNMulAdd() has an ifdef ladder for different targets.
 We want to make the propagation rule instead be selectable at
 runtime, because:
  * this will let us have multiple targets in one QEMU binary
  * the Arm FEAT_AFP architectural feature includes letting
    the guest select a NaN propagation rule at runtime
 In this commit we add an enum for the propagation rule, the field in
 float_status, and the corresponding getters and setters.  We change
 pickNaNMulAdd to honour this, but because all targets still leave
 this field at its default 0 value, the fallback logic will pick the
 rule type with the old ifdef ladder.
 It's valid not to set a propagation rule if default_nan_mode is
 enabled, because in that case there's no need to pick a NaN; all the
 callers of pickNaNMulAdd() catch this case and skip calling it.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-3-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 76 ++++++++++++++++++++++++------------------
+ include/fpu/softfloat-helpers.h |  11 +++
-file changed, 44 insertions(+), 32 deletions(-)
+ include/fpu/softfloat-types.h   |  55 +++++++++++
+ fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+files changed, 107 insertions(+), 126 deletions(-)
 diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/include/fpu/softfloat-helpers.h
-+++ b/target/arm/translate.c
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@ void gen_gvec_uaba(unsigned vece, uint32_t rd_ofs, uint32_t rn_ofs,
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
-     tcg_gen_gvec_3(rd_ofs, rn_ofs, rm_ofs, opr_sz, max_sz, &ops[vece]);
+     status->float_2nan_prop_rule = rule;
  }
--static int disas_coproc_insn(DisasContext *s, uint32_t insn)
++static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
-+static void do_coproc_insn(DisasContext *s, int cpnum, int is64,
++                                            float_status *status)
-+                           int opc1, int crn, int crm, int opc2,
++{
-+                           bool isread, int rt, int rt2)
++    status->float_3nan_prop_rule = rule;
 +}
 +
  static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
                                               float_status *status)
  {
--    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
+@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
-     const ARMCPRegInfo *ri;
+     return status->float_2nan_prop_rule;
+ }
--    cpnum = (insn >> 8) & 0xf;
 +static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
 +{
 +    return status->float_3nan_prop_rule;
 +}
 +
  static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
  {
      return status->float_infzeronan_rule;
 diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-types.h
 +++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ this code that are retained.
  #ifndef SOFTFLOAT_TYPES_H
  #define SOFTFLOAT_TYPES_H
 +#include "hw/registerfields.h"
 +
  /*
   * Software IEC/IEEE floating-point types.
   */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
      float_2nan_prop_x87,
  } Float2NaNPropRule;
 +/*
 + * 3-input NaN propagation rule, for fused multiply-add. Individual
 + * architectures have different rules for which input NaN is
 + * propagated to the output when there is more than one NaN on the
 + * input.
 + *
 + * If default_nan_mode is enabled then it is valid not to set a NaN
 + * propagation rule, because the softfloat code guarantees not to try
 + * to pick a NaN to propagate in default NaN mode.  When not in
 + * default-NaN mode, it is an error for the target not to set the rule
 + * in float_status if it uses a muladd, and we will assert if we need
 + * to handle an input NaN and no rule was selected.
 + *
 + * The naming scheme for Float3NaNPropRule values is:
 + *  float_3nan_prop_s_abc:
 + *    = "Prefer SNaN over QNaN, then operand A over B over C"
 + *  float_3nan_prop_abc:
 + *    = "Prefer A over B over C regardless of SNaN vs QNAN"
 + *
 + * For QEMU, the multiply-add operation is A * B + C.
 + */
 +
 +/*
 + * We set the Float3NaNPropRule enum values up so we can select the
 + * right value in pickNaNMulAdd in a data driven way.
 + */
 +FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
 +FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
 +FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
 +FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
 +
 +#define PROPRULE(X, Y, Z) \
 +    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
 +
 +typedef enum __attribute__((__packed__)) {
 +    float_3nan_prop_none = 0,     /* No propagation rule specified */
 +    float_3nan_prop_abc = PROPRULE(0, 1, 2),
 +    float_3nan_prop_acb = PROPRULE(0, 2, 1),
 +    float_3nan_prop_bac = PROPRULE(1, 0, 2),
 +    float_3nan_prop_bca = PROPRULE(1, 2, 0),
 +    float_3nan_prop_cab = PROPRULE(2, 0, 1),
 +    float_3nan_prop_cba = PROPRULE(2, 1, 0),
 +    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
 +} Float3NaNPropRule;
 +
 +#undef PROPRULE
 +
  /*
   * Rule for result of fused multiply-add 0 * Inf + NaN.
   * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    Float3NaNPropRule float_3nan_prop_rule;
      FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
  static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                           bool infzero, bool have_snan, float_status *status)
  {
 +    FloatClass cls[3] = { a_cls, b_cls, c_cls };
 +    Float3NaNPropRule rule = status->float_3nan_prop_rule;
 +    int which;
 +
      /*
       * We guarantee not to require the target to tell us how to
       * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
 +    if (rule == float_3nan_prop_none) {
  #if defined(TARGET_ARM)
 -
--    is64 = (insn & (1 << 25)) == 0;
+-    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
--    if (!is64 && ((insn & (1 << 4)) == 0)) {
+-     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
--        /* cdp */
+-     */
--        return 1;
+-    if (is_snan(c_cls)) {
--    }
+-        return 2;
--
+-    } else if (is_snan(a_cls)) {
--    crm = insn & 0xf;
+-        return 0;
--    if (is64) {
+-    } else if (is_snan(b_cls)) {
--        crn = 0;
+-        return 1;
--        opc1 = (insn >> 4) & 0xf;
+-    } else if (is_qnan(c_cls)) {
--        opc2 = 0;
+-        return 2;
--        rt2 = (insn >> 16) & 0xf;
+-    } else if (is_qnan(a_cls)) {
--    } else {
+-        return 0;
--        crn = (insn >> 16) & 0xf;
+-    } else {
--        opc1 = (insn >> 21) & 7;
+-        return 1;
--        opc2 = (insn >> 5) & 7;
+-    }
--        rt2 = 0;
++        /*
--    }
++         * This looks different from the ARM ARM pseudocode, because the ARM ARM
--    isread = (insn >> 20) & 1;
++         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
--    rt = (insn >> 12) & 0xf;
++         */
--
++        rule = float_3nan_prop_s_cab;
-     ri = get_arm_cp_reginfo(s->cp_regs,
+ #elif defined(TARGET_MIPS)
-             ENCODE_CP_REG(cpnum, is64, s->ns, crn, crm, opc1, opc2));
+-    if (snan_bit_is_one(status)) {
-     if (ri) {
+-        /* Prefer sNaN over qNaN, in the a, b, c order. */
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+-        if (is_snan(a_cls)) {
+-            return 0;
-         /* Check access permissions */
+-        } else if (is_snan(b_cls)) {
-         if (!cp_access_ok(s->current_el, ri, isread)) {
+-            return 1;
--            return 1;
+-        } else if (is_snan(c_cls)) {
-+            unallocated_encoding(s);
+-            return 2;
-+            return;
+-        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 +        if (snan_bit_is_one(status)) {
 +            rule = float_3nan_prop_s_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_s_cab;
          }
+-    } else {
-         if (s->hstr_active || ri->accessfn ||
+-        /* Prefer sNaN over qNaN, in the c, a, b order. */
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+-        if (is_snan(c_cls)) {
-         /* Handle special cases first */
+-            return 2;
-         switch (ri->type & ~(ARM_CP_FLAG_MASK & ~ARM_CP_SPECIAL)) {
+-        } else if (is_snan(a_cls)) {
-         case ARM_CP_NOP:
+-            return 0;
--            return 0;
+-        } else if (is_snan(b_cls)) {
-+            return;
+-            return 1;
-         case ARM_CP_WFI:
+-        } else if (is_qnan(c_cls)) {
-             if (isread) {
+-            return 2;
--                return 1;
+-        } else if (is_qnan(a_cls)) {
-+                unallocated_encoding(s);
+-            return 0;
-+                return;
+-        } else {
-             }
+-            return 1;
-             gen_set_pc_im(s, s->base.pc_next);
+-        }
-             s->base.is_jmp = DISAS_WFI;
+-    }
--            return 0;
+ #elif defined(TARGET_LOONGARCH64)
-+            return;
+-    /* Prefer sNaN over qNaN, in the c, a, b order. */
-         default:
+-    if (is_snan(c_cls)) {
-             break;
+-        return 2;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else {
 -        return 1;
 -    }
 +        rule = float_3nan_prop_s_cab;
  #elif defined(TARGET_PPC)
 -    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 -     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 -     */
 -    if (is_nan(a_cls)) {
 -        return 0;
 -    } else if (is_nan(c_cls)) {
 -        return 2;
 -    } else {
 -        return 1;
 -    }
 +        /*
 +         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 +         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 +         */
 +        rule = float_3nan_prop_acb;
  #elif defined(TARGET_S390X)
 -    if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 2;
 -    }
 +        rule = float_3nan_prop_s_abc;
  #elif defined(TARGET_SPARC)
 -    /* Prefer SNaN over QNaN, order C, B, A. */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 0;
 -    }
 +        rule = float_3nan_prop_s_cba;
  #elif defined(TARGET_XTENSA)
 -    /*
 -     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
 -     * an input NaN if we have one (ie c).
 -     */
 -    if (status->use_first_nan) {
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else if (is_nan(b_cls)) {
 -            return 1;
 +        if (status->use_first_nan) {
 +            rule = float_3nan_prop_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_cba;
          }
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+-    } else {
-             /* Write */
+-        if (is_nan(c_cls)) {
-             if (ri->type & ARM_CP_CONST) {
+-            return 2;
-                 /* If not forbidden by access permissions, treat as WI */
+-        } else if (is_nan(b_cls)) {
--                return 0;
+-            return 1;
-+                return;
+-        } else {
-             }
+-            return 0;
+-        }
-             if (is64) {
+-    }
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+ #else
-             gen_lookup_tb(s);
+-    /* A default implementation: prefer a to b to c.
-         }
+-     * This is unlikely to actually match any real implementation.
+-     */
--        return 0;
+-    if (is_nan(a_cls)) {
-+        return;
+-        return 0;
-     }
+-    } else if (is_nan(b_cls)) {
+-        return 1;
-     /* Unknown register; this might be a guest error or a QEMU
+-    } else {
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+-        return 2;
-                       s->ns ? "non-secure" : "secure");
+-    }
-     }
++        rule = float_3nan_prop_abc;
+ #endif
 -    return 1;
 +    unallocated_encoding(s);
 +    return;
 +}
 +
 +static int disas_coproc_insn(DisasContext *s, uint32_t insn)
 +{
 +    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
 +
 +    cpnum = (insn >> 8) & 0xf;
 +
 +    is64 = (insn & (1 << 25)) == 0;
 +    if (!is64 && ((insn & (1 << 4)) == 0)) {
 +        /* cdp */
 +        return 1;
 +    }
 +
-+    crm = insn & 0xf;
++    assert(rule != float_3nan_prop_none);
-+    if (is64) {
++    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-+        crn = 0;
++        /* We have at least one SNaN input and should prefer it */
-+        opc1 = (insn >> 4) & 0xf;
++        do {
-+        opc2 = 0;
++            which = rule & R_3NAN_1ST_MASK;
-+        rt2 = (insn >> 16) & 0xf;
++            rule >>= R_3NAN_1ST_LENGTH;
 +        } while (!is_snan(cls[which]));
 +    } else {
-+        crn = (insn >> 16) & 0xf;
++        do {
-+        opc1 = (insn >> 21) & 7;
++            which = rule & R_3NAN_1ST_MASK;
-+        opc2 = (insn >> 5) & 7;
++            rule >>= R_3NAN_1ST_LENGTH;
-+        rt2 = 0;
++        } while (!is_nan(cls[which]));
 +    }
-+    isread = (insn >> 20) & 1;
++    return which;
 +    rt = (insn >> 12) & 0xf;
 +
 +    do_coproc_insn(s, cpnum, is64, opc1, crn, crm, opc2, isread, rt, rt2);
 +    return 0;
  }
- /* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
+ /*----------------------------------------------------------------------------
 --
-.20.1
+.34.1

-New patch
+[PULL 21/72] tests/fp: Explicitly set 3-NaN propagation rule
+Explicitly set a rule in the softfloat tests for propagating NaNs in
+the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
+so we should select here the Arm rule of float_3nan_prop_s_cab.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 1 +
+ tests/fp/fp-test.c  | 1 +
+files changed, 2 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-New patch
+[PULL 22/72] target/arm: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for Arm, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 5 +++++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 6 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
++ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
++ *    but note that for QEMU muladd is a * b + c, whereas for
++ *    the pseudocode function the arguments are in the order c, a, b.
+  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+  *    and the input NaN if it is signalling
+  */
+@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     }
+     if (rule == float_3nan_prop_none) {
+-#if defined(TARGET_ARM)
+-        /*
+-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+-         */
+-        rule = float_3nan_prop_s_cab;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             rule = float_3nan_prop_s_abc;
+         } else {
+--
+.34.1

-New patch
+[PULL 23/72] target/loongarch: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for loongarch, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 1 +
+ fpu/softfloat-specialize.c.inc    | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      * case sets InvalidOp and returns the input value 'c'
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_LOONGARCH64)
+-        rule = float_3nan_prop_s_cab;
+ #elif defined(TARGET_PPC)
+         /*
+          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+--
+.34.1

-New patch
+[PULL 24/72] target/ppc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for PPC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 6 ------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * NaN propagation for fused multiply-add:
++     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
++     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
++     * whereas QEMU labels the operands as (a * b) + c.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
+     /*
+      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+      * to return an input NaN if we have one (ie c) rather than generating
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_PPC)
+-        /*
+-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+-         */
+-        rule = float_3nan_prop_acb;
+ #elif defined(TARGET_S390X)
+         rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+--
+.34.1

-New patch
+[PULL 25/72] target/s390x: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for s390x, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 1 +
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
+        /* fall through */
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_S390X)
+-        rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+         rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+--
+.34.1

-New patch
+[PULL 26/72] target/sparc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for SPARC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
++    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_SPARC)
+-        rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+         if (status->use_first_nan) {
+             rule = float_3nan_prop_abc;
+--
+.34.1

-New patch
+[PULL 27/72] target/mips: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for Arm, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h       | 4 ++++
+ target/mips/msa.c              | 3 +++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 8 insertions(+), 7 deletions(-)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+ {
+     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
+     FloatInfZeroNaNRule izn_rule;
++    Float3NaNPropRule nan3_rule;
+     /*
+      * With nan2008, SNaNs are silenced in the usual way.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+      */
+     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
++    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
++    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
++
+ }
+ static inline void restore_fp_status(CPUMIPSState *env)
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab,
+                              &env->active_tc.msa_fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
++                             &env->active_tc.msa_fp_status);
++
+     /* clear float_status exception flags */
+     set_float_exception_flags(0, &env->active_tc.msa_fp_status);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     }
+     if (rule == float_3nan_prop_none) {
+-#if defined(TARGET_MIPS)
+-        if (snan_bit_is_one(status)) {
+-            rule = float_3nan_prop_s_abc;
+-        } else {
+-            rule = float_3nan_prop_s_cab;
+-        }
+-#elif defined(TARGET_XTENSA)
++#if defined(TARGET_XTENSA)
+         if (status->use_first_nan) {
+             rule = float_3nan_prop_abc;
+         } else {
+--
+.34.1

-[PULL 25/27] target/arm: Make A32/T32 use new fpstatus_ptr() API
+[PULL 28/72] target/xtensa: Set Float3NaNPropRule explicitly
-Make A32/T32 code use the new fpstatus_ptr() API:
+Set the Float3NaNPropRule explicitly for xtensa, and remove the
- get_fpstatus_ptr(0) -> fpstatus_ptr(FPST_FPCR)
+ifdef from pickNaNMulAdd().
  get_fpstatus_ptr(1) -> fpstatus_ptr(FPST_STD)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
 Message-id: 20200806104453.30393-3-peter.maydell@linaro.org
 ---
- target/arm/translate.c          | 13 ----------
+ target/xtensa/fpu_helper.c     | 2 ++
- target/arm/translate-neon.c.inc | 28 ++++++++++-----------
+ fpu/softfloat-specialize.c.inc | 8 --------
- target/arm/translate-vfp.c.inc  | 44 ++++++++++++++++-----------------
+files changed, 2 insertions(+), 8 deletions(-)
 files changed, 36 insertions(+), 49 deletions(-)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/xtensa/fpu_helper.c
-+++ b/target/arm/translate.c
++++ b/target/xtensa/fpu_helper.c
-@@ -XXX,XX +XXX,XX @@ static inline void gen_hlt(DisasContext *s, int imm)
+@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
-     unallocated_encoding(s);
+     set_use_first_nan(use_first, &env->fp_status);
      set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                               &env->fp_status);
 +    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
 +                             &env->fp_status);
  }
--static TCGv_ptr get_fpstatus_ptr(int neon)
+ void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
--{
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 -    TCGv_ptr statusptr = tcg_temp_new_ptr();
 -    int offset;
 -    if (neon) {
 -        offset = offsetof(CPUARMState, vfp.standard_fp_status);
 -    } else {
 -        offset = offsetof(CPUARMState, vfp.fp_status);
 -    }
 -    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
 -    return statusptr;
 -}
 -
  static inline long vfp_reg_offset(bool dp, unsigned reg)
  {
      if (dp) {
 diff --git a/target/arm/translate-neon.c.inc b/target/arm/translate-neon.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate-neon.c.inc
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/translate-neon.c.inc
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      }
-     opr_sz = (1 + a->q) * 8;
+     if (rule == float_3nan_prop_none) {
--    fpst = get_fpstatus_ptr(1);
+-#if defined(TARGET_XTENSA)
-+    fpst = fpstatus_ptr(FPST_STD);
+-        if (status->use_first_nan) {
-     fn_gvec_ptr = a->size ? gen_helper_gvec_fcmlas : gen_helper_gvec_fcmlah;
+-            rule = float_3nan_prop_abc;
-     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
+-        } else {
-                        vfp_reg_offset(1, a->vn),
+-            rule = float_3nan_prop_cba;
-@@ -XXX,XX +XXX,XX @@ static bool trans_VCADD(DisasContext *s, arg_VCADD *a)
+-        }
 -#else
          rule = float_3nan_prop_abc;
 -#endif
      }
-     opr_sz = (1 + a->q) * 8;
+     assert(rule != float_3nan_prop_none);
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      fn_gvec_ptr = a->size ? gen_helper_gvec_fcadds : gen_helper_gvec_fcaddh;
      tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                         vfp_reg_offset(1, a->vn),
@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA_scalar(DisasContext *s, arg_VCMLA_scalar *a)
      fn_gvec_ptr = (a->size ? gen_helper_gvec_fcmlas_idx
                     : gen_helper_gvec_fcmlah_idx);
      opr_sz = (1 + a->q) * 8;
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                         vfp_reg_offset(1, a->vn),
                         vfp_reg_offset(1, a->vm),
@@ -XXX,XX +XXX,XX @@ static bool trans_VDOT_scalar(DisasContext *s, arg_VDOT_scalar *a)
      fn_gvec = a->u ? gen_helper_gvec_udot_idx_b : gen_helper_gvec_sdot_idx_b;
      opr_sz = (1 + a->q) * 8;
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      tcg_gen_gvec_3_ool(vfp_reg_offset(1, a->vd),
                         vfp_reg_offset(1, a->vn),
                         vfp_reg_offset(1, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn,
          return true;
      }
 -    TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 +    TCGv_ptr fpstatus = fpstatus_ptr(FPST_STD);
      for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
          tmp = neon_load_reg(a->vn, pass);
          tmp2 = neon_load_reg(a->vm, pass);
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn,
                                  uint32_t rn_ofs, uint32_t rm_ofs,       \
                                  uint32_t oprsz, uint32_t maxsz)         \
      {                                                                   \
 -        TCGv_ptr fpst = get_fpstatus_ptr(1);                            \
 +        TCGv_ptr fpst = fpstatus_ptr(FPST_STD);                         \
          tcg_gen_gvec_3_ptr(rd_ofs, rn_ofs, rm_ofs, fpst,                \
                             oprsz, maxsz, 0, FUNC);                      \
          tcg_temp_free_ptr(fpst);                                        \
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp_pair(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn)
       * early. Since Q is 0 there are always just two passes, so instead
       * of a complicated loop over each pass we just unroll.
       */
 -    fpstatus = get_fpstatus_ptr(1);
 +    fpstatus = fpstatus_ptr(FPST_STD);
      tmp = neon_load_reg(a->vn, 0);
      tmp2 = neon_load_reg(a->vn, 1);
      fn(tmp, tmp, tmp2, fpstatus);
@@ -XXX,XX +XXX,XX @@ static bool do_fp_2sh(DisasContext *s, arg_2reg_shift *a,
          return true;
      }
 -    fpstatus = get_fpstatus_ptr(1);
 +    fpstatus = fpstatus_ptr(FPST_STD);
      shiftv = tcg_const_i32(a->shift);
      for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
          tmp = neon_load_reg(a->vm, pass);
@@ -XXX,XX +XXX,XX @@ static bool trans_VMLS_2sc(DisasContext *s, arg_2scalar *a)
  #define WRAP_FP_FN(WRAPNAME, FUNC)                              \
      static void WRAPNAME(TCGv_i32 rd, TCGv_i32 rn, TCGv_i32 rm) \
      {                                                           \
 -        TCGv_ptr fpstatus = get_fpstatus_ptr(1);                \
 +        TCGv_ptr fpstatus = fpstatus_ptr(FPST_STD);             \
          FUNC(rd, rn, rm, fpstatus);                             \
          tcg_temp_free_ptr(fpstatus);                            \
      }
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_F16_F32(DisasContext *s, arg_2misc *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(true);
 +    fpst = fpstatus_ptr(FPST_STD);
      ahp = get_ahp_flag();
      tmp = neon_load_reg(a->vm, 0);
      gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_F32_F16(DisasContext *s, arg_2misc *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(true);
 +    fpst = fpstatus_ptr(FPST_STD);
      ahp = get_ahp_flag();
      tmp3 = tcg_temp_new_i32();
      tmp = neon_load_reg(a->vm, 0);
@@ -XXX,XX +XXX,XX @@ static bool do_2misc_fp(DisasContext *s, arg_2misc *a,
          return true;
      }
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
          TCGv_i32 tmp = neon_load_reg(a->vm, pass);
          fn(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool do_vrint(DisasContext *s, arg_2misc *a, int rmode)
          return true;
      }
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
      gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
      for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
@@ -XXX,XX +XXX,XX @@ static bool do_vcvt(DisasContext *s, arg_2misc *a, int rmode, bool is_signed)
          return true;
      }
 -    fpst = get_fpstatus_ptr(1);
 +    fpst = fpstatus_ptr(FPST_STD);
      tcg_shift = tcg_const_i32(0);
      tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
      gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
 diff --git a/target/arm/translate-vfp.c.inc b/target/arm/translate-vfp.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-vfp.c.inc
 +++ b/target/arm/translate-vfp.c.inc
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINT(DisasContext *s, arg_VRINT *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rounding));
      gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT(DisasContext *s, arg_VCVT *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_shift = tcg_const_i32(0);
@@ -XXX,XX +XXX,XX @@ static bool do_vfp_3op_sp(DisasContext *s, VFPGen3OpSPFn *fn,
      f0 = tcg_temp_new_i32();
      f1 = tcg_temp_new_i32();
      fd = tcg_temp_new_i32();
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      neon_load_reg32(f0, vn);
      neon_load_reg32(f1, vm);
@@ -XXX,XX +XXX,XX @@ static bool do_vfp_3op_dp(DisasContext *s, VFPGen3OpDPFn *fn,
      f0 = tcg_temp_new_i64();
      f1 = tcg_temp_new_i64();
      fd = tcg_temp_new_i64();
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      neon_load_reg64(f0, vn);
      neon_load_reg64(f1, vm);
@@ -XXX,XX +XXX,XX @@ static bool do_vfm_sp(DisasContext *s, arg_VFMA_sp *a, bool neg_n, bool neg_d)
          /* VFNMA, VFNMS */
          gen_helper_vfp_negs(vd, vd);
      }
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_vfp_muladds(vd, vn, vm, vd, fpst);
      neon_store_reg32(vd, a->vd);
@@ -XXX,XX +XXX,XX @@ static bool do_vfm_dp(DisasContext *s, arg_VFMA_dp *a, bool neg_n, bool neg_d)
          /* VFNMA, VFNMS */
          gen_helper_vfp_negd(vd, vd);
      }
 -    fpst = get_fpstatus_ptr(0);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_vfp_muladdd(vd, vn, vm, vd, fpst);
      neon_store_reg64(vd, a->vd);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f32_f16(DisasContext *s, arg_VCVT_f32_f16 *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      ahp_mode = get_ahp_flag();
      tmp = tcg_temp_new_i32();
      /* The T bit tells us if we want the low or high 16 bits of Vm */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f64_f16(DisasContext *s, arg_VCVT_f64_f16 *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      ahp_mode = get_ahp_flag();
      tmp = tcg_temp_new_i32();
      /* The T bit tells us if we want the low or high 16 bits of Vm */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f16_f32(DisasContext *s, arg_VCVT_f16_f32 *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      ahp_mode = get_ahp_flag();
      tmp = tcg_temp_new_i32();
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f16_f64(DisasContext *s, arg_VCVT_f16_f64 *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      ahp_mode = get_ahp_flag();
      tmp = tcg_temp_new_i32();
      vm = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTR_sp(DisasContext *s, arg_VRINTR_sp *a)
      tmp = tcg_temp_new_i32();
      neon_load_reg32(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_rints(tmp, tmp, fpst);
      neon_store_reg32(tmp, a->vd);
      tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTR_dp(DisasContext *s, arg_VRINTR_dp *a)
      tmp = tcg_temp_new_i64();
      neon_load_reg64(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_rintd(tmp, tmp, fpst);
      neon_store_reg64(tmp, a->vd);
      tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTZ_sp(DisasContext *s, arg_VRINTZ_sp *a)
      tmp = tcg_temp_new_i32();
      neon_load_reg32(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_rmode = tcg_const_i32(float_round_to_zero);
      gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
      gen_helper_rints(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTZ_dp(DisasContext *s, arg_VRINTZ_dp *a)
      tmp = tcg_temp_new_i64();
      neon_load_reg64(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_rmode = tcg_const_i32(float_round_to_zero);
      gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
      gen_helper_rintd(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTX_sp(DisasContext *s, arg_VRINTX_sp *a)
      tmp = tcg_temp_new_i32();
      neon_load_reg32(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_rints_exact(tmp, tmp, fpst);
      neon_store_reg32(tmp, a->vd);
      tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTX_dp(DisasContext *s, arg_VRINTX_dp *a)
      tmp = tcg_temp_new_i64();
      neon_load_reg64(tmp, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      gen_helper_rintd_exact(tmp, tmp, fpst);
      neon_store_reg64(tmp, a->vd);
      tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_int_sp(DisasContext *s, arg_VCVT_int_sp *a)
      vm = tcg_temp_new_i32();
      neon_load_reg32(vm, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      if (a->s) {
          /* i32 -> f32 */
          gen_helper_vfp_sitos(vm, vm, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_int_dp(DisasContext *s, arg_VCVT_int_dp *a)
      vm = tcg_temp_new_i32();
      vd = tcg_temp_new_i64();
      neon_load_reg32(vm, a->vm);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      if (a->s) {
          /* i32 -> f64 */
          gen_helper_vfp_sitod(vd, vm, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_fix_sp(DisasContext *s, arg_VCVT_fix_sp *a)
      vd = tcg_temp_new_i32();
      neon_load_reg32(vd, a->vd);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      shift = tcg_const_i32(frac_bits);
      /* Switch on op:U:sx bits */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_fix_dp(DisasContext *s, arg_VCVT_fix_dp *a)
      vd = tcg_temp_new_i64();
      neon_load_reg64(vd, a->vd);
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      shift = tcg_const_i32(frac_bits);
      /* Switch on op:U:sx bits */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_sp_int(DisasContext *s, arg_VCVT_sp_int *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      vm = tcg_temp_new_i32();
      neon_load_reg32(vm, a->vm);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_dp_int(DisasContext *s, arg_VCVT_dp_int *a)
          return true;
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      vm = tcg_temp_new_i64();
      vd = tcg_temp_new_i32();
      neon_load_reg64(vm, a->vm);
 --
-.20.1
+.34.1

-New patch
+[PULL 29/72] target/i386: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for i386.  We had no
+i386-specific behaviour in the old ifdef ladder, so we were using the
+default "prefer a then b then c" fallback; this is actually the
+correct per-the-spec handling for i386.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c | 1 +
+file changed, 1 insertion(+)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      * there are multiple input NaNs they are selected in the order a, b, c.
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+--
+.34.1

-New patch
+[PULL 30/72] target/hppa: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for HPPA, and remove the
+ifdef from pickNaNMulAdd().
+HPPA is the only target that was using the default branch of the
+ifdef ladder (other targets either do not use muladd or set
+default_nan_mode), so we can remove the ifdef fallback entirely now
+(allowing the "rule not set" case to fall into the default of the
+switch statement and assert).
+We add a TODO note that the HPPA rule is probably wrong; this is
+not a behavioural change for this refactoring.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 4 ----
+files changed, 8 insertions(+), 4 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * TODO: The HPPA architecture reference only documents its NaN
++     * propagation rule for 2-operand operations. Testing on real hardware
++     * might be necessary to confirm whether this order for muladd is correct.
++     * Not preferring the SNaN is almost certainly incorrect as it diverges
++     * from the documented rules for 2-operand operations.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         }
+     }
+-    if (rule == float_3nan_prop_none) {
+-        rule = float_3nan_prop_abc;
+-    }
+-
+     assert(rule != float_3nan_prop_none);
+     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+         /* We have at least one SNaN input and should prefer it */
+--
+.34.1

-[PULL 19/27] target/arm: Convert T32 coprocessor insns to decodetree
+[PULL 31/72] fpu: Remove use_first_nan field from float_status
-Convert the T32 coprocessor instructions to decodetree.
+The use_first_nan field in float_status was an xtensa-specific way to
-As with the A32 conversion, this corrects an underdecoding
+select at runtime from two different NaN propagation rules.  Now that
-where we did not check that MRRC/MCRR [24:21] were 0b0010
+xtensa is using the target-agnostic NaN propagation rule selection
-and so treated some kinds of LDC/STC and MRRC/MCRR rather
+that we've just added, we can remove use_first_nan, because there is
-than UNDEFing them.
+no longer any code that reads it.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-7-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
 ---
- target/arm/t32.decode  | 19 +++++++++++++
+ include/fpu/softfloat-helpers.h | 5 -----
- target/arm/translate.c | 64 ++----------------------------------------
+ include/fpu/softfloat-types.h   | 1 -
-files changed, 21 insertions(+), 62 deletions(-)
+ target/xtensa/fpu_helper.c      | 1 -
 files changed, 7 deletions(-)
-diff --git a/target/arm/t32.decode b/target/arm/t32.decode
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/t32.decode
+--- a/include/fpu/softfloat-helpers.h
-+++ b/target/arm/t32.decode
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static inline void set_snan_bit_is_one(bool val, float_status *status)
- &sat             !extern rd rn satimm imm sh
+     status->snan_bit_is_one = val;
  &pkh             !extern rd rn rm imm tb
  &cps             !extern mode imod M A I F
 +&mcr             !extern cp opc1 crn crm opc2 rt
 +&mcrr            !extern cp opc1 crm rt rt2
  # Data-processing (register)
@@ -XXX,XX +XXX,XX @@ RFE              1110 1001 10.1 .... 1100000000000000         @rfe pu=1
  SRS              1110 1000 00.0 1101 1100 0000 000. ....      @srs pu=2
  SRS              1110 1001 10.0 1101 1100 0000 000. ....      @srs pu=1
 +# Coprocessor instructions
 +
 +# We decode MCR, MCR, MRRC and MCRR only, because for QEMU the
 +# other coprocessor instructions always UNDEF.
 +# The trans_ functions for these will ignore cp values 8..13 for v7 or
 +# earlier, and 0..13 for v8 and later, because those areas of the
 +# encoding space may be used for other things, such as VFP or Neon.
 +
 +@mcr             .... .... opc1:3 . crn:4 rt:4 cp:4 opc2:3 . crm:4
 +@mcrr            .... .... .... rt2:4 rt:4 cp:4 opc1:4 crm:4
 +
 +MCRR             1110 1100 0100 .... .... .... .... .... @mcrr
 +MRRC             1110 1100 0101 .... .... .... .... .... @mcrr
 +
 +MCR              1110 1110 ... 0 .... .... .... ... 1 .... @mcr
 +MRC              1110 1110 ... 1 .... .... .... ... 1 .... @mcr
 +
  # Branches
  %imm24           26:s1 13:1 11:1 16:10 0:11 !function=t32_branch24
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void do_coproc_insn(DisasContext *s, int cpnum, int is64,
      return;
  }
--static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+-static inline void set_use_first_nan(bool val, float_status *status)
 -{
--    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
+-    status->use_first_nan = val;
 -
 -    cpnum = (insn >> 8) & 0xf;
 -
 -    is64 = (insn & (1 << 25)) == 0;
 -    if (!is64 && ((insn & (1 << 4)) == 0)) {
 -        /* cdp */
 -        return 1;
 -    }
 -
 -    crm = insn & 0xf;
 -    if (is64) {
 -        crn = 0;
 -        opc1 = (insn >> 4) & 0xf;
 -        opc2 = 0;
 -        rt2 = (insn >> 16) & 0xf;
 -    } else {
 -        crn = (insn >> 16) & 0xf;
 -        opc1 = (insn >> 21) & 7;
 -        opc2 = (insn >> 5) & 7;
 -        rt2 = 0;
 -    }
 -    isread = (insn >> 20) & 1;
 -    rt = (insn >> 12) & 0xf;
 -
 -    do_coproc_insn(s, cpnum, is64, opc1, crn, crm, opc2, isread, rt, rt2);
 -    return 0;
 -}
 -
- /* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
+ static inline void set_no_signaling_nans(bool val, float_status *status)
  static void disas_xscale_insn(DisasContext *s, uint32_t insn)
  {
-@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
+     status->no_signaling_nans = val;
-         ((insn >> 28) == 0xe && disas_vfp(s, insn))) {
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
-         return;
+index XXXXXXX..XXXXXXX 100644
-     }
+--- a/include/fpu/softfloat-types.h
--    /* fall back to legacy decoder */
++++ b/include/fpu/softfloat-types.h
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
--    switch ((insn >> 25) & 0xf) {
+      * softfloat-specialize.inc.c)
--    case 0: case 1: case 2: case 3:
+      */
--        /* 16-bit instructions.  Should never happen.  */
+     bool snan_bit_is_one;
--        abort();
+-    bool use_first_nan;
--    case 6: case 7: case 14: case 15:
+     bool no_signaling_nans;
--        /* Coprocessor.  */
+     /* should overflowed results subtract re_bias to its exponent? */
--        if (arm_dc_feature(s, ARM_FEATURE_M)) {
+     bool rebias_overflow;
--            /* 0b111x_11xx_xxxx_xxxx_xxxx_xxxx_xxxx_xxxx */
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
--            goto illegal_op;
+index XXXXXXX..XXXXXXX 100644
--        }
+--- a/target/xtensa/fpu_helper.c
--        if (((insn >> 24) & 3) == 3) {
++++ b/target/xtensa/fpu_helper.c
--            /* Neon DP, but failed disas_neon_dp() */
+@@ -XXX,XX +XXX,XX @@ static const struct {
--            goto illegal_op;
--        } else if (((insn >> 8) & 0xe) == 10) {
+ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
--            /* VFP, but failed disas_vfp.  */
+ {
--            goto illegal_op;
+-    set_use_first_nan(use_first, &env->fp_status);
--        } else {
+     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
--            if (insn & (1 << 28))
+                              &env->fp_status);
--                goto illegal_op;
+     set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
 -            if (disas_coproc_insn(s, insn)) {
 -                goto illegal_op;
 -            }
 -        }
 -        break;
 -    case 12:
 -        goto illegal_op;
 -    default:
 -    illegal_op:
 -        unallocated_encoding(s);
 -    }
 +illegal_op:
 +    unallocated_encoding(s);
  }
  static void disas_thumb_insn(DisasContext *s, uint32_t insn)
 --
-.20.1
+.34.1

-New patch
+[PULL 32/72] target/m68k: Don't pass NULL float_status to floatx80_default_nan()
+Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
+to get the NaN bit pattern to reset the FPU registers. This
+works because it happens that our implementation of
+floatx80_default_nan() doesn't actually look at the float_status
+pointer except for TARGET_MIPS. However, this isn't guaranteed,
+and to be able to remove the ifdef in floatx80_default_nan()
+we're going to need a real float_status here.
+Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
+earlier, and thus can pass it to floatx80_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c | 12 +++++++-----
+file changed, 7 insertions(+), 5 deletions(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+     CPUState *cs = CPU(obj);
+     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
+     CPUM68KState *env = cpu_env(cs);
+-    floatx80 nan = floatx80_default_nan(NULL);
++    floatx80 nan;
+     int i;
+     if (mcc->parent_phases.hold) {
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+ #else
+     cpu_m68k_set_sr(env, SR_S | SR_I);
+ #endif
+-    for (i = 0; i < 8; i++) {
+-        env->fregs[i].d = nan;
+-    }
+-    cpu_m68k_set_fpcr(env, 0);
+     /*
+      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
+      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++
++    nan = floatx80_default_nan(&env->fp_status);
++    for (i = 0; i < 8; i++) {
++        env->fregs[i].d = nan;
++    }
++    cpu_m68k_set_fpcr(env, 0);
+     env->fpsr = 0;
+     /* TODO: We should set PC from the interrupt vector.  */
+--
+.34.1

-New patch
+[PULL 33/72] softfloat: Create floatx80 default NaN from parts64_default_nan
+We create our 128-bit default NaN by calling parts64_default_nan()
+and then adjusting the result.  We can do the same trick for creating
+the floatx80 default NaN, which lets us drop a target ifdef.
+floatx80 is used only by:
+ i386
+ m68k
+ arm nwfpe old floating-point emulation emulation support
+    (which is essentially dead, especially the parts involving floatx80)
+ PPC (only in the xsrqpxp instruction, which just rounds an input
+    value by converting to floatx80 and back, so will never generate
+    the default NaN)
+The floatx80 default NaN as currently implemented is:
+ m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
+ i386: sign = 1, exp = 1...1, int = 1, frac = 10...0
+These are the same as the parts64_default_nan for these architectures.
+This is technically a possible behaviour change for arm linux-user
+nwfpe emulation emulation, because the default NaN will now have the
+sign bit clear.  But we were already generating a different floatx80
+default NaN from the real kernel emulation we are supposedly
+following, which appears to use an all-bits-1 value:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267
+This won't affect the only "real" use of the nwfpe emulation, which
+is ancient binaries that used it as part of the old floating point
+calling convention; that only uses loads and stores of 32 and 64 bit
+floats, not any of the floatx80 behaviour the original hardware had.
+We also get the nwfpe float64 default NaN value wrong:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
+so if we ever cared about this obscure corner the right fix would be
+to correct that so nwfpe used its own default-NaN setting rather
+than the Arm VFP one.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
+---
+ fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
+file changed, 10 insertions(+), 10 deletions(-)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
+ floatx80 floatx80_default_nan(float_status *status)
+ {
+     floatx80 r;
++    /*
++     * Extrapolate from the choices made by parts64_default_nan to fill
++     * in the floatx80 format. We assume that floatx80's explicit
++     * integer bit is always set (this is true for i386 and m68k,
++     * which are the only real users of this format).
++     */
++    FloatParts64 p64;
++    parts64_default_nan(&p64, status);
+-    /* None of the targets that have snan_bit_is_one use floatx80.  */
+-    assert(!snan_bit_is_one(status));
+-#if defined(TARGET_M68K)
+-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
+-    r.high = 0x7FFF;
+-#else
+-    /* X86 */
+-    r.low = UINT64_C(0xC000000000000000);
+-    r.high = 0xFFFF;
+-#endif
++    r.high = 0x7FFF | (p64.sign << 15);
++    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
+     return r;
+ }
+--
+.34.1

-New patch
+[PULL 34/72] target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
+In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
+a zero-initialized float_status struct to float32_is_quiet_nan() and
+float64_is_quiet_nan(), with the cryptic comment "for
+snan_bit_is_one".
+This pattern appears to have been copied from target/riscv, where it
+is used because the functions there do not have ready access to the
+CPU state struct. The comment presumably refers to the fact that the
+main reason the is_quiet_nan() functions want the float_state is
+because they want to know about the snan_bit_is_one config.
+In the loongarch helpers, though, we have the CPU state struct
+to hand. Use the usual env->fp_status here. This avoids our needing
+to track that we need to update the initializer of the local
+float_status structs when the core softfloat code adds new
+options for targets to configure their behaviour.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
+     } else if (float32_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float32_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
+     } else if (float64_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float64_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+--
+.34.1

-New patch
+[PULL 35/72] target/m68k: In frem helper, initialize local float_status from env->fp_status
+In the frem helper, we have a local float_status because we want to
+execute the floatx80_div() with a custom rounding mode.  Instead of
+zero-initializing the local float_status and then having to set it up
+with the m68k standard behaviour (including the NaN propagation rule
+and copying the rounding precision from env->fp_status), initialize
+it as a complete copy of env->fp_status. This will avoid our having
+to add new code in this function for every new config knob we add
+to fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
+---
+ target/m68k/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/fpu_helper.c
++++ b/target/m68k/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
+     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
+     if (!floatx80_is_any_nan(fp_rem)) {
+-        float_status fp_status = { };
++        /* Use local temporary fp_status to set different rounding mode */
++        float_status fp_status = env->fp_status;
+         uint32_t quotient;
+         int sign;
+         /* Calculate quotient directly using round to nearest mode */
+-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
+         set_float_rounding_mode(float_round_nearest_even, &fp_status);
+-        set_floatx80_rounding_precision(
+-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
+         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
+         sign = extractFloatx80Sign(fp_quot.d);
+--
+.34.1

-New patch
+[PULL 36/72] target/m68k: Init local float_status from env fp_status in gdb get/set reg
+In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
+from float64 to floatx80 using a scratch float_status, because we
+don't want the conversion to affect the CPU's floating point exception
+status. Currently we use a zero-initialized float_status. This will
+get steadily more awkward as we add config knobs to float_status
+that the target must initialize. Avoid having to add any of that
+configuration here by instead initializing our local float_status
+from the env->fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
+---
+ target/m68k/helper.c | 6 ++++--
+file changed, 4 insertions(+), 2 deletions(-)
+diff --git a/target/m68k/helper.c b/target/m68k/helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/helper.c
++++ b/target/m68k/helper.c
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
+     }
+     switch (n) {
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
+         return 8;
+     }
+--
+.34.1

-[PULL 05/27] hw/arm/smmu: Introduce SMMUTLBEntry for PTW and IOTLB value
+[PULL 37/72] target/sparc: Initialize local scratch float_status from env->fp_status
-From: Eric Auger <eric.auger@redhat.com>
+In the helper functions flcmps and flcmpd we use a scratch float_status
 so that we don't change the CPU state if the comparison raises any
 floating point exception flags. Instead of zero-initializing this
 scratch float_status, initialize it as a copy of env->fp_status. This
 avoids the need to explicitly initialize settings like the NaN
 propagation rule or others we might add to softfloat in future.
-Introduce a specialized SMMUTLBEntry to store the result of
+To do this we need to pass the CPU env pointer in to the helper.
 the PTW and cache in the IOTLB. This structure extends the
 generic IOMMUTLBEntry struct with the level of the entry and
 the granule size.
-Those latter will be useful when implementing range invalidation.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
 ---
  target/sparc/helper.h     | 4 ++--
  target/sparc/fop_helper.c | 8 ++++----
  target/sparc/translate.c  | 4 ++--
 files changed, 8 insertions(+), 8 deletions(-)
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+diff --git a/target/sparc/helper.h b/target/sparc/helper.h
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20200728150815.11446-5-eric.auger@redhat.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  include/hw/arm/smmu-common.h | 12 +++++++++---
  hw/arm/smmu-common.c         | 32 +++++++++++++++++---------------
  hw/arm/smmuv3.c              | 10 +++++-----
 files changed, 31 insertions(+), 23 deletions(-)
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/target/sparc/helper.h
-+++ b/include/hw/arm/smmu-common.h
++++ b/target/sparc/helper.h
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransTableInfo {
+@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
-     uint8_t granule_sz;        /* granule page shift */
+ DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
- } SMMUTransTableInfo;
+ DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
+ DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
-+typedef struct SMMUTLBEntry {
+-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
-+    IOMMUTLBEntry entry;
+-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
-+    uint8_t level;
++DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
-+    uint8_t granule;
++DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
-+} SMMUTLBEntry;
+ DEF_HELPER_2(raise_exception, noreturn, env, int)
-+
- /*
+ DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
-  * Generic structure populated by derived SMMU devices
+diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
   * after decoding the configuration information and used as
@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
   * pair, according to @cfg translation config
   */
  int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
 -             IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
 +             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
  /**
   * select_tt - compute which translation table shall be used according to
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
  #define SMMU_IOTLB_MAX_SIZE 256
 -IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
 -void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry);
 +SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
 +void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
  SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova);
  void smmu_iotlb_inv_all(SMMUState *s);
  void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/target/sparc/fop_helper.c
-+++ b/hw/arm/smmu-common.c
++++ b/target/sparc/fop_helper.c
-@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova)
+@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
-     return key;
+     return finish_fcmp(env, r, GETPC());
  }
--IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+-uint32_t helper_flcmps(float32 src1, float32 src2)
--                                 hwaddr iova)
++uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
 +SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
 +                                hwaddr iova)
  {
-     SMMUIOTLBKey key = smmu_get_iotlb_key(cfg->asid, iova);
+     /*
--    IOMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
+      * FLCMP never raises an exception nor modifies any FSR fields.
-+    SMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
+      * Perform the comparison with a dummy fp environment.
+      */
-     if (entry) {
+-    float_status discard = { };
-         cfg->iotlb_hits++;
++    float_status discard = env->fp_status;
-@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+     FloatRelation r;
-     return entry;
      set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
      g_assert_not_reached();
  }
--void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
+-uint32_t helper_flcmpd(float64 src1, float64 src2)
-+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *new)
++uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
  {
-     SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
+-    float_status discard = { };
++    float_status discard = env->fp_status;
-@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
+     FloatRelation r;
-         smmu_iotlb_inv_all(bs);
-     }
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
+diff --git a/target/sparc/translate.c b/target/sparc/translate.c
--    *key = smmu_get_iotlb_key(cfg->asid, entry->iova);
+index XXXXXXX..XXXXXXX 100644
--    trace_smmu_iotlb_insert(cfg->asid, entry->iova);
+--- a/target/sparc/translate.c
--    g_hash_table_insert(bs->iotlb, key, entry);
++++ b/target/sparc/translate.c
-+    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova);
+@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
-+    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova);
-+    g_hash_table_insert(bs->iotlb, key, new);
+     src1 = gen_load_fpr_F(dc, a->rs1);
      src2 = gen_load_fpr_F(dc, a->rs2);
 -    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
 +    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
      return advance_pc(dc);
  }
- inline void smmu_iotlb_inv_all(SMMUState *s)
+@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
-@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
-  * @cfg: translation config
+     src1 = gen_load_fpr_D(dc, a->rs1);
-  * @iova: iova to translate
+     src2 = gen_load_fpr_D(dc, a->rs2);
-  * @perm: access type
+-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
-- * @tlbe: IOMMUTLBEntry (out)
++    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
-+ * @tlbe: SMMUTLBEntry (out)
+     return advance_pc(dc);
   * @info: handle to an error info
   *
   * Return 0 on success, < 0 on error. In case of error, @info is filled
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
   */
  static int smmu_ptw_64(SMMUTransCfg *cfg,
                         dma_addr_t iova, IOMMUAccessFlags perm,
 -                       IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 +                       SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
  {
      dma_addr_t baseaddr, indexmask;
      int stage = cfg->stage;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
      baseaddr = extract64(tt->ttb, 0, 48);
      baseaddr &= ~indexmask;
 -    tlbe->iova = iova;
 -    tlbe->addr_mask = (1 << granule_sz) - 1;
 +    tlbe->entry.iova = iova;
 +    tlbe->entry.addr_mask = (1 << granule_sz) - 1;
      while (level <= 3) {
          uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
              goto error;
          }
 -        tlbe->translated_addr = gpa + (iova & mask);
 -        tlbe->perm = PTE_AP_TO_PERM(ap);
 +        tlbe->entry.translated_addr = gpa + (iova & mask);
 +        tlbe->entry.perm = PTE_AP_TO_PERM(ap);
 +        tlbe->level = level;
 +        tlbe->granule = granule_sz;
          return 0;
      }
      info->type = SMMU_PTW_ERR_TRANSLATION;
  error:
 -    tlbe->perm = IOMMU_NONE;
 +    tlbe->entry.perm = IOMMU_NONE;
      return -EINVAL;
  }
-@@ -XXX,XX +XXX,XX @@ error:
-  * return 0 on success
-  */
- inline int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
--             IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
-+                    SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
- {
-     if (!cfg->aa64) {
-         /*
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
-index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
-+++ b/hw/arm/smmuv3.c
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
-     SMMUTranslationStatus status;
-     SMMUState *bs = ARM_SMMU(s);
-     uint64_t page_mask, aligned_addr;
--    IOMMUTLBEntry *cached_entry = NULL;
-+    SMMUTLBEntry *cached_entry = NULL;
-     SMMUTransTableInfo *tt;
-     SMMUTransCfg *cfg = NULL;
-     IOMMUTLBEntry entry = {
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
-     cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
-     if (cached_entry) {
--        if ((flag & IOMMU_WO) && !(cached_entry->perm & IOMMU_WO)) {
-+        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
-             status = SMMU_TRANS_ERROR;
-             if (event.record_trans_faults) {
-                 event.type = SMMU_EVT_F_PERMISSION;
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
-         goto epilogue;
-     }
--    cached_entry = g_new0(IOMMUTLBEntry, 1);
-+    cached_entry = g_new0(SMMUTLBEntry, 1);
-     if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
-         g_free(cached_entry);
-@@ -XXX,XX +XXX,XX @@ epilogue:
-     switch (status) {
-     case SMMU_TRANS_SUCCESS:
-         entry.perm = flag;
--        entry.translated_addr = cached_entry->translated_addr +
-+        entry.translated_addr = cached_entry->entry.translated_addr +
-                                     (addr & page_mask);
--        entry.addr_mask = cached_entry->addr_mask;
-+        entry.addr_mask = cached_entry->entry.addr_mask;
-         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
-                                        entry.translated_addr, entry.perm);
-         break;
 --
-.20.1
+.34.1

-New patch
+[PULL 38/72] target/ppc: Use env->fp_status in helper_compute_fprf functions
+In the helper_compute_fprf functions, we pass a dummy float_status
+in to the is_signaling_nan() function. This is unnecessary, because
+we have convenient access to the CPU env pointer here and that
+is already set up with the correct values for the snan_bit_is_one
+and no_signaling_nans config settings. is_signaling_nan() doesn't
+ever update the fp_status with any exception flags, so there is
+no reason not to use env->fp_status here.
+Use env->fp_status instead of the dummy fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
+---
+ target/ppc/fpu_helper.c | 3 +--
+file changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/fpu_helper.c
++++ b/target/ppc/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
+     } else if (tp##_is_infinity(arg)) {                           \
+         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
+     } else {                                                      \
+-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
+-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
++        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
+             fprf = 0x00 << FPSCR_FPRF;                            \
+         } else {                                                  \
+             fprf = 0x11 << FPSCR_FPRF;                            \
+--
+.34.1

-[PULL 12/27] hw/arm/smmuv3: Advertise SMMUv3.2 range invalidation
+[PULL 39/72] target/arm: Copy entire float_status in is_ebf
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Expose the RIL bit so that the guest driver uses range
+Now that float_status has a bunch of fp parameters,
-invalidation. Although RIL is a 3.2 features, We let
+it is easier to copy an existing structure than create
-the AIDR advertise SMMUv3.1 support as v3.x implementation
+one from scratch.  Begin by copying the structure that
-is allowed to implement features from v3.(x+1).
+corresponds to the FPSR and make only the adjustments
 required for BFloat16 semantics.
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20200728150815.11446-12-eric.auger@redhat.com
+Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h | 1 +
+ target/arm/tcg/vec_helper.c | 20 +++++++-------------
- hw/arm/smmuv3.c          | 1 +
+file changed, 7 insertions(+), 13 deletions(-)
 files changed, 2 insertions(+)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/target/arm/tcg/vec_helper.c
-+++ b/hw/arm/smmuv3-internal.h
++++ b/target/arm/tcg/vec_helper.c
-@@ -XXX,XX +XXX,XX @@ REG32(IDR1,                0x4)
+@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
- REG32(IDR2,                0x8)
+      * no effect on AArch32 instructions.
- REG32(IDR3,                0xc)
+      */
-      FIELD(IDR3, HAD,         2, 1);
+     bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
-+     FIELD(IDR3, RIL,        10, 1);
+-    *statusp = (float_status){
- REG32(IDR4,                0x10)
+-        .tininess_before_rounding = float_tininess_before_rounding,
- REG32(IDR5,                0x14)
+-        .float_rounding_mode = float_round_to_odd_inf,
-      FIELD(IDR5, OAS,         0, 3);
+-        .flush_to_zero = true,
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+-        .flush_inputs_to_zero = true,
-index XXXXXXX..XXXXXXX 100644
+-        .default_nan_mode = true,
---- a/hw/arm/smmuv3.c
+-    };
-+++ b/hw/arm/smmuv3.c
++
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
++    *statusp = env->vfp.fp_status;
-     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, EVENTQS, SMMU_EVENTQS);
++    set_default_nan_mode(true, statusp);
-     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, CMDQS,   SMMU_CMDQS);
+     if (ebf) {
-+    s->idr[3] = FIELD_DP32(s->idr[3], IDR3, RIL, 1);
+-        float_status *fpst = &env->vfp.fp_status;
-     s->idr[3] = FIELD_DP32(s->idr[3], IDR3, HAD, 1);
+-        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
+-        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
-    /* 4K and 64K granule support */
+-        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
 -
          /* EBF=1 needs to do a step with round-to-odd semantics */
          *oddstatusp = *statusp;
          set_float_rounding_mode(float_round_to_odd, oddstatusp);
 +    } else {
 +        set_flush_to_zero(true, statusp);
 +        set_flush_inputs_to_zero(true, statusp);
 +        set_float_rounding_mode(float_round_to_odd_inf, statusp);
      }
 -
      return ebf;
  }
 --
-.20.1
+.34.1

-[PULL 16/27] target/arm: Convert A32 coprocessor insns to decodetree
+[PULL 40/72] fpu: Allow runtime choice of default NaN value
-Convert the A32 coprocessor instructions to decodetree.
+Currently we hardcode the default NaN value in parts64_default_nan()
 using a compile-time ifdef ladder. This is awkward for two cases:
  * for single-QEMU-binary we can't hard-code target-specifics like this
  * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
    (specifically the sign bit is different)
-Note that this corrects an underdecoding: for the 64-bit access case
+Add a field to float_status to specify the default NaN value; fall
-(MRRC/MCRR) we did not check that bits [24:21] were 0b0010, so we
+back to the old ifdef behaviour if these are not set.
 would incorrectly treat LDC/STC as MRRC/MCRR rather than UNDEFing
 them.
-The decodetree versions of these insns assume the coprocessor
+The default NaN value is specified by setting a uint8_t to a
-is in the range 0..7 or 14..15. This is architecturally sensible
+pattern corresponding to the sign and upper fraction parts of
-(as per the comments) and OK in practice for QEMU because the only
+the NaN; the lower bits of the fraction are set from bit 0 of
-uses of the ARMCPRegInfo infrastructure we have that aren't
+the pattern.
 for coprocessors 14 or 15 are the pxa2xx use of coprocessor 6.
 We add an assertion to the define_one_arm_cp_reg_with_opaque()
 function to catch any accidental future attempts to use it to
 define coprocessor registers for invalid coprocessors.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-4-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
 ---
- target/arm/a32.decode  | 19 +++++++++++
+ include/fpu/softfloat-helpers.h | 11 +++++++
- target/arm/helper.c    | 29 +++++++++++++++++
+ include/fpu/softfloat-types.h   | 10 ++++++
- target/arm/translate.c | 74 +++++++++++++++++++++++++++++++++++-------
+ fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
-files changed, 111 insertions(+), 11 deletions(-)
+files changed, 54 insertions(+), 22 deletions(-)
-diff --git a/target/arm/a32.decode b/target/arm/a32.decode
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/a32.decode
+--- a/include/fpu/softfloat-helpers.h
-+++ b/target/arm/a32.decode
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
- &bfi             rd rn lsb msb
+     status->float_infzeronan_rule = rule;
- &sat             rd rn satimm imm sh
+ }
- &pkh             rd rn rm imm tb
-+&mcr             cp opc1 crn crm opc2 rt
++static inline void set_float_default_nan_pattern(uint8_t dnan_pattern,
-+&mcrr            cp opc1 crm rt rt2
++                                                 float_status *status)
  # Data-processing (register)
@@ -XXX,XX +XXX,XX @@ LDM_a32          ---- 100 b:1 i:1 u:1 w:1 1 rn:4 list:16   &ldst_block
  B                .... 1010 ........................           @branch
  BL               .... 1011 ........................           @branch
 +# Coprocessor instructions
 +
 +# We decode MCR, MCR, MRRC and MCRR only, because for QEMU the
 +# other coprocessor instructions always UNDEF.
 +# The trans_ functions for these will ignore cp values 8..13 for v7 or
 +# earlier, and 0..13 for v8 and later, because those areas of the
 +# encoding space may be used for other things, such as VFP or Neon.
 +
 +@mcr             ---- .... opc1:3 . crn:4 rt:4 cp:4 opc2:3 . crm:4 &mcr
 +@mcrr            ---- .... .... rt2:4 rt:4 cp:4 opc1:4 crm:4       &mcrr
 +
 +MCRR             .... 1100 0100 .... .... .... .... .... @mcrr
 +MRRC             .... 1100 0101 .... .... .... .... .... @mcrr
 +
 +MCR              .... 1110 ... 0 .... .... .... ... 1 .... @mcr
 +MRC              .... 1110 ... 1 .... .... .... ... 1 .... @mcr
 +
  # Supervisor call
  SVC              ---- 1111 imm:24                             &i
 diff --git a/target/arm/helper.c b/target/arm/helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/helper.c
 +++ b/target/arm/helper.c
@@ -XXX,XX +XXX,XX @@ void define_one_arm_cp_reg_with_opaque(ARMCPU *cpu,
      assert((r->state != ARM_CP_STATE_AA32) || (r->opc0 == 0));
      /* AArch64 regs are all 64 bit so ARM_CP_64BIT is meaningless */
      assert((r->state != ARM_CP_STATE_AA64) || !(r->type & ARM_CP_64BIT));
 +    /*
 +     * This API is only for Arm's system coprocessors (14 and 15) or
 +     * (M-profile or v7A-and-earlier only) for implementation defined
 +     * coprocessors in the range 0..7.  Our decode assumes this, since
 +     * 8..13 can be used for other insns including VFP and Neon. See
 +     * valid_cp() in translate.c.  Assert here that we haven't tried
 +     * to use an invalid coprocessor number.
 +     */
 +    switch (r->state) {
 +    case ARM_CP_STATE_BOTH:
 +        /* 0 has a special meaning, but otherwise the same rules as AA32. */
 +        if (r->cp == 0) {
 +            break;
 +        }
 +        /* fall through */
 +    case ARM_CP_STATE_AA32:
 +        if (arm_feature(&cpu->env, ARM_FEATURE_V8) &&
 +            !arm_feature(&cpu->env, ARM_FEATURE_M)) {
 +            assert(r->cp >= 14 && r->cp <= 15);
 +        } else {
 +            assert(r->cp < 8 || (r->cp >= 14 && r->cp <= 15));
 +        }
 +        break;
 +    case ARM_CP_STATE_AA64:
 +        assert(r->cp == 0 || r->cp == CP_REG_ARM64_SYSREG_CP);
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
      /* The AArch64 pseudocode CheckSystemAccess() specifies that op1
       * encodes a minimum access level for the register. We roll this
       * runtime check into our general permission check code, so check
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int t16_pop_list(DisasContext *s, int x)
  #include "decode-t32.c.inc"
  #include "decode-t16.c.inc"
 +static bool valid_cp(DisasContext *s, int cp)
 +{
-+    /*
++    status->default_nan_pattern = dnan_pattern;
 +     * Return true if this coprocessor field indicates something
 +     * that's really a possible coprocessor.
 +     * For v7 and earlier, coprocessors 8..15 were reserved for Arm use,
 +     * and of those only cp14 and cp15 were used for registers.
 +     * cp10 and cp11 were used for VFP and Neon, whose decode is
 +     * dealt with elsewhere. With the advent of fp16, cp9 is also
 +     * now part of VFP.
 +     * For v8A and later, the encoding has been tightened so that
 +     * only cp14 and cp15 are valid, and other values aren't considered
 +     * to be in the coprocessor-instruction space at all. v8M still
 +     * permits coprocessors 0..7.
 +     */
 +    if (arm_dc_feature(s, ARM_FEATURE_V8) &&
 +        !arm_dc_feature(s, ARM_FEATURE_M)) {
 +        return cp >= 14;
 +    }
 +    return cp < 8 || cp >= 14;
 +}
 +
-+static bool trans_MCR(DisasContext *s, arg_MCR *a)
+ static inline void set_flush_to_zero(bool val, float_status *status)
  {
      status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status
      return status->float_infzeronan_rule;
  }
 +static inline uint8_t get_float_default_nan_pattern(float_status *status)
 +{
-+    if (!valid_cp(s, a->cp)) {
++    return status->default_nan_pattern;
 +        return false;
 +    }
 +    do_coproc_insn(s, a->cp, false, a->opc1, a->crn, a->crm, a->opc2,
 +                   false, a->rt, 0);
 +    return true;
 +}
 +
-+static bool trans_MRC(DisasContext *s, arg_MRC *a)
+ static inline bool get_flush_to_zero(float_status *status)
-+{
+ {
-+    if (!valid_cp(s, a->cp)) {
+     return status->flush_to_zero;
-+        return false;
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-types.h
 +++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      /* should denormalised inputs go to zero and set the input_denormal flag? */
      bool flush_inputs_to_zero;
      bool default_nan_mode;
 +    /*
 +     * The pattern to use for the default NaN. Here the high bit specifies
 +     * the default NaN's sign bit, and bits 6..0 specify the high bits of the
 +     * fractional part. The low bits of the fractional part are copies of bit 0.
 +     * The exponent of the default NaN is (as for any NaN) always all 1s.
 +     * Note that a value of 0 here is not a valid NaN. The target must set
 +     * this to the correct non-zero value, or we will assert when trying to
 +     * create a default NaN.
 +     */
 +    uint8_t default_nan_pattern;
      /*
       * The flags below are not used on all specializations and may
       * constant fold away (see snan_bit_is_one()/no_signalling_nans() in
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
  {
      bool sign = 0;
      uint64_t frac;
 +    uint8_t dnan_pattern = status->default_nan_pattern;
 +    if (dnan_pattern == 0) {
  #if defined(TARGET_SPARC) || defined(TARGET_M68K)
 -    /* !snan_bit_is_one, set all bits */
 -    frac = (1ULL << DECOMPOSED_BINARY_POINT) - 1;
 -#elif defined(TARGET_I386) || defined(TARGET_X86_64) \
 +        /* Sign bit clear, all frac bits set */
 +        dnan_pattern = 0b01111111;
 +#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
      || defined(TARGET_MICROBLAZE)
 -    /* !snan_bit_is_one, set sign and msb */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    sign = 1;
 +        /* Sign bit set, most significant frac bit set */
 +        dnan_pattern = 0b11000000;
  #elif defined(TARGET_HPPA)
 -    /* snan_bit_is_one, set msb-1.  */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 2);
 +        /* Sign bit clear, msb-1 frac bit set */
 +        dnan_pattern = 0b00100000;
  #elif defined(TARGET_HEXAGON)
 -    sign = 1;
 -    frac = ~0ULL;
 +        /* Sign bit set, all frac bits set. */
 +        dnan_pattern = 0b11111111;
  #else
 -    /*
 -     * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 -     * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 -     * do not have floating-point.
 -     */
 -    if (snan_bit_is_one(status)) {
 -        /* set all bits other than msb */
 -        frac = (1ULL << (DECOMPOSED_BINARY_POINT - 1)) - 1;
 -    } else {
 -        /* set msb */
 -        frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    }
 +        /*
 +         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 +         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 +         * do not have floating-point.
 +         */
 +        if (snan_bit_is_one(status)) {
 +            /* sign bit clear, set all frac bits other than msb */
 +            dnan_pattern = 0b00111111;
 +        } else {
 +            /* sign bit clear, set frac msb */
 +            dnan_pattern = 0b01000000;
 +        }
  #endif
 +    }
-+    do_coproc_insn(s, a->cp, false, a->opc1, a->crn, a->crm, a->opc2,
++    assert(dnan_pattern != 0);
 +                   true, a->rt, 0);
 +    return true;
 +}
 +
-+static bool trans_MCRR(DisasContext *s, arg_MCRR *a)
++    sign = dnan_pattern >> 7;
-+{
++    /*
-+    if (!valid_cp(s, a->cp)) {
++     * Place default_nan_pattern [6:0] into bits [62:56],
-+        return false;
++     * and replecate bit [0] down into [55:0]
-+    }
++     */
-+    do_coproc_insn(s, a->cp, true, a->opc1, 0, a->crm, 0,
++    frac = deposit64(0, DECOMPOSED_BINARY_POINT - 7, 7, dnan_pattern);
-+                   false, a->rt, a->rt2);
++    frac = deposit64(frac, 0, DECOMPOSED_BINARY_POINT - 7, -(dnan_pattern & 1));
-+    return true;
-+}
+     *p = (FloatParts64) {
-+
+         .cls = float_class_qnan,
 +static bool trans_MRRC(DisasContext *s, arg_MRRC *a)
 +{
 +    if (!valid_cp(s, a->cp)) {
 +        return false;
 +    }
 +    do_coproc_insn(s, a->cp, true, a->opc1, 0, a->crm, 0,
 +                   true, a->rt, a->rt2);
 +    return true;
 +}
 +
  /* Helpers to swap operands for reverse-subtract.  */
  static void gen_rsb(TCGv_i32 dst, TCGv_i32 a, TCGv_i32 b)
  {
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
              disas_xscale_insn(s, insn);
              break;
          }
 -
 -        if ((cpnum & 0xe) == 10) {
 -            /* VFP, but failed disas_vfp.  */
 -            goto illegal_op;
 -        }
 -
 -        if (disas_coproc_insn(s, insn)) {
 -            /* Coprocessor.  */
 -            goto illegal_op;
 -        }
 -        break;
 +        /* fall through */
      }
      default:
      illegal_op:
 --
-.20.1
+.34.1

-New patch
+[PULL 41/72] tests/fp: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the tests/fp code.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c     | 1 +
+ tests/fp/fp-test-log2.c | 1 +
+ tests/fp/fp-test.c      | 1 +
+files changed, 3 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
++    set_float_default_nan_pattern(0b01000000, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test-log2.c
++++ b/tests/fp/fp-test-log2.c
+@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
+     int i;
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_rounding_mode(float_round_nearest_even, &qsf);
+     test.d = 0.0;
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-[PULL 20/27] target/arm: Remove ARCH macro
+[PULL 42/72] target/microblaze: Set default NaN pattern explicitly
-The ARCH() macro was used a lot in the legacy decoder, but
+Set the default NaN pattern explicitly, and remove the ifdef from
-there are now just two uses of it left. Since a macro which
+parts64_default_nan().
 expands out to a goto is liable to be confusing when reading
 code, replace the last two uses with a simple open-coded
 qeuivalent.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-8-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 14 +++++++++-----
+ target/microblaze/cpu.c        | 2 ++
-file changed, 9 insertions(+), 5 deletions(-)
+ fpu/softfloat-specialize.c.inc | 3 +--
 files changed, 3 insertions(+), 2 deletions(-)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/microblaze/cpu.c
-+++ b/target/arm/translate.c
++++ b/target/microblaze/cpu.c
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
- #define ENABLE_ARCH_7     arm_dc_feature(s, ARM_FEATURE_V7)
+      * this architecture.
- #define ENABLE_ARCH_8     arm_dc_feature(s, ARM_FEATURE_V8)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
--#define ARCH(x) do { if (!ENABLE_ARCH_##x) goto illegal_op; } while(0)
++    /* Default NaN: sign bit set, most significant frac bit set */
--
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
  #include "translate.h"
  #if defined(CONFIG_USER_ONLY)
-@@ -XXX,XX +XXX,XX @@ static bool trans_BLX_i(DisasContext *s, arg_BLX_i *a)
+     /* start in user mode with interrupts enabled.  */
- {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-     TCGv_i32 tmp;
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
--    /* For A32, ARCH(5) is checked near the start of the uncond block. */
++++ b/fpu/softfloat-specialize.c.inc
-+    /* For A32, ARM_FEATURE_V5 is checked near the start of the uncond block. */
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
-     if (s->thumb && (a->imm & 2)) {
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
-         return false;
+         /* Sign bit clear, all frac bits set */
-     }
+         dnan_pattern = 0b01111111;
-@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
-          * choose to UNDEF. In ARMv5 and above the space is used
+-    || defined(TARGET_MICROBLAZE)
-          * for miscellaneous unconditional instructions.
++#elif defined(TARGET_I386) || defined(TARGET_X86_64)
-          */
+         /* Sign bit set, most significant frac bit set */
--        ARCH(5);
+         dnan_pattern = 0b11000000;
-+        if (!arm_dc_feature(s, ARM_FEATURE_V5)) {
+ #elif defined(TARGET_HPPA)
 +            unallocated_encoding(s);
 +            return;
 +        }
          /* Unconditional instructions.  */
          /* TODO: Perhaps merge these into one decodetree output file.  */
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
              goto illegal_op;
          }
      } else if ((insn & 0xf800e800) != 0xf000e800)  {
 -        ARCH(6T2);
 +        if (!arm_dc_feature(s, ARM_FEATURE_THUMB2)) {
 +            unallocated_encoding(s);
 +            return;
 +        }
      }
      if (arm_dc_feature(s, ARM_FEATURE_M)) {
 --
-.20.1
+.34.1

-New patch
+[PULL 43/72] target/i386: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 4 ++++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 4 insertions(+), 3 deletions(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
++    /* Default NaN: sign bit set, most significant frac bit set */
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
++    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
++    set_float_default_nan_pattern(0b11000000, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
+-        /* Sign bit set, most significant frac bit set */
+-        dnan_pattern = 0b11000000;
+ #elif defined(TARGET_HPPA)
+         /* Sign bit clear, msb-1 frac bit set */
+         dnan_pattern = 0b00100000;
+--
+.34.1

-New patch
+[PULL 44/72] target/hppa: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 2 insertions(+), 3 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN: sign bit clear, msb-1 frac bit set */
++    set_float_default_nan_pattern(0b00100000, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HPPA)
+-        /* Sign bit clear, msb-1 frac bit set */
+-        dnan_pattern = 0b00100000;
+ #elif defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+--
+.34.1

-New patch
+[PULL 45/72] target/alpha: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the alpha target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-40-peter.maydell@linaro.org
+---
+ target/alpha/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/alpha/cpu.c b/target/alpha/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/alpha/cpu.c
++++ b/target/alpha/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void alpha_cpu_initfn(Object *obj)
+      * operand in Fa. That is float_2nan_prop_ba.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ #if defined(CONFIG_USER_ONLY)
+     env->flags = ENV_FLAG_PS_USER | ENV_FLAG_FEN;
+     cpu_alpha_store_fpcr(env, (uint64_t)(FPCR_INVD | FPCR_DZED | FPCR_OVFD
+--
+.34.1

-[PULL 26/27] target/arm: Implement FPST_STD_F16 fpstatus
+[PULL 46/72] target/arm: Set default NaN pattern explicitly
-Architecturally, Neon FP16 operations use the "standard FPSCR" like
+Set the default NaN pattern explicitly for the arm target.
-all other Neon operations.  However, this is defined in the Arm ARM
+This includes setting it for the old linux-user nwfpe emulation.
-pseudocode as "a fixed value, except that FZ16 (and AHP) follow the
+For nwfpe, our default doesn't match the real kernel, but we
-FPSCR bits". In QEMU, the softfloat float_status doesn't include
+avoid making a behaviour change in this commit.
 separate flush-to-zero for FP16 operations, so we must keep separate
 fp_status for "Neon non-FP16" and "Neon fp16" operations, in the
 same way we do already for the non-Neon "fp_status" vs "fp_status_f16".
 Add the extra float_status field to the CPU state structure,
 ensure it is correctly initialized and updated on FPSCR writes,
 and make fpstatus_ptr(FPST_STD_F16) return a pointer to it.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
 Message-id: 20200806104453.30393-4-peter.maydell@linaro.org
 ---
- target/arm/cpu.h        | 9 ++++++++-
+ linux-user/arm/nwfpe/fpa11.c | 5 +++++
- target/arm/translate.h  | 3 ++-
+ target/arm/cpu.c             | 2 ++
- target/arm/cpu.c        | 3 +++
+files changed, 7 insertions(+)
  target/arm/vfp_helper.c | 5 +++++
 files changed, 18 insertions(+), 2 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/linux-user/arm/nwfpe/fpa11.c
-+++ b/target/arm/cpu.h
++++ b/linux-user/arm/nwfpe/fpa11.c
-@@ -XXX,XX +XXX,XX @@ typedef struct CPUARMState {
+@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
-          *  fp_status: is the "normal" fp status.
+    * this late date.
-          *  fp_status_fp16: used for half-precision calculations
+    */
-          *  standard_fp_status : the ARM "Standard FPSCR Value"
+   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
-+         *  standard_fp_status_fp16 : used for half-precision
++  /*
-+         *       calculations with the ARM "Standard FPSCR Value"
++   * Use the same default NaN value as Arm VFP. This doesn't match
-          *
++   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
-          * Half-precision operations are governed by a separate
++   */
-          * flush-to-zero control bit in FPSCR:FZ16. We pass a separate
++  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
-@@ -XXX,XX +XXX,XX @@ typedef struct CPUARMState {
+ }
-          * Neon) which the architecture defines as controlled by the
-          * standard FPSCR value rather than the FPSCR.
+ void SetRoundingMode(const unsigned int opcode)
           *
 +         * The "standard FPSCR but for fp16 ops" is needed because
 +         * the "standard FPSCR" tracks the FPSCR.FZ16 bit rather than
 +         * using a fixed value for it.
 +         *
           * To avoid having to transfer exception bits around, we simply
           * say that the FPSCR cumulative exception flags are the logical
 -         * OR of the flags in the three fp statuses. This relies on the
 +         * OR of the flags in the four fp statuses. This relies on the
           * only thing which needs to read the exception flags being
           * an explicit FPSCR read.
           */
          float_status fp_status;
          float_status fp_status_f16;
          float_status standard_fp_status;
 +        float_status standard_fp_status_f16;
          /* ZCR_EL[1-3] */
          uint64_t zcr_el[4];
 diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.h
 +++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ static inline TCGv_ptr fpstatus_ptr(ARMFPStatusFlavour flavour)
          offset = offsetof(CPUARMState, vfp.standard_fp_status);
          break;
      case FPST_STD_F16:
 -        /* Not yet used or implemented: fall through to assert */
 +        offset = offsetof(CPUARMState, vfp.standard_fp_status_f16);
 +        break;
      default:
          g_assert_not_reached();
      }
 diff --git a/target/arm/cpu.c b/target/arm/cpu.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/cpu.c
 +++ b/target/arm/cpu.c
-@@ -XXX,XX +XXX,XX @@ static void arm_cpu_reset(DeviceState *dev)
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
-     set_flush_to_zero(1, &env->vfp.standard_fp_status);
+  *    the pseudocode function the arguments are in the order c, a, b.
-     set_flush_inputs_to_zero(1, &env->vfp.standard_fp_status);
+  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
-     set_default_nan_mode(1, &env->vfp.standard_fp_status);
+  *    and the input NaN if it is signalling
-+    set_default_nan_mode(1, &env->vfp.standard_fp_status_f16);
++ *  * Default NaN has sign bit clear, msb frac bit set
-     set_float_detect_tininess(float_tininess_before_rounding,
+  */
-                               &env->vfp.fp_status);
+ static void arm_set_default_fp_behaviours(float_status *s)
-     set_float_detect_tininess(float_tininess_before_rounding,
+ {
-                               &env->vfp.standard_fp_status);
+@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
-     set_float_detect_tininess(float_tininess_before_rounding,
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
-                               &env->vfp.fp_status_f16);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
-+    set_float_detect_tininess(float_tininess_before_rounding,
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
-+                              &env->vfp.standard_fp_status_f16);
++    set_float_default_nan_pattern(0b01000000, s);
  #ifndef CONFIG_USER_ONLY
      if (kvm_enabled()) {
          kvm_arm_reset_vcpu(cpu);
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/vfp_helper.c
 +++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
      /* FZ16 does not generate an input denormal exception.  */
      i |= (get_float_exception_flags(&env->vfp.fp_status_f16)
            & ~float_flag_input_denormal);
 +    i |= (get_float_exception_flags(&env->vfp.standard_fp_status_f16)
 +          & ~float_flag_input_denormal);
      return vfp_exceptbits_from_host(i);
  }
-@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
      if (changed & FPCR_FZ16) {
          bool ftz_enabled = val & FPCR_FZ16;
          set_flush_to_zero(ftz_enabled, &env->vfp.fp_status_f16);
 +        set_flush_to_zero(ftz_enabled, &env->vfp.standard_fp_status_f16);
          set_flush_inputs_to_zero(ftz_enabled, &env->vfp.fp_status_f16);
 +        set_flush_inputs_to_zero(ftz_enabled, &env->vfp.standard_fp_status_f16);
      }
      if (changed & FPCR_FZ) {
          bool ftz_enabled = val & FPCR_FZ;
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
      set_float_exception_flags(i, &env->vfp.fp_status);
      set_float_exception_flags(0, &env->vfp.fp_status_f16);
      set_float_exception_flags(0, &env->vfp.standard_fp_status);
 +    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
  }
  #else
 --
-.20.1
+.34.1

-New patch
+[PULL 47/72] target/loongarch: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for loongarch.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-42-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+--
+.34.1

-New patch
+[PULL 48/72] target/m68k: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for m68k.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c              | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++    /* Default NaN: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     nan = floatx80_default_nan(&env->fp_status);
+     for (i = 0; i < 8; i++) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
++#if defined(TARGET_SPARC)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+ #elif defined(TARGET_HEXAGON)
+--
+.34.1

-New patch
+[PULL 49/72] target/mips: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for MIPS. Note that this
+is our only target which currently changes the default NaN
+at runtime (which it was previously doing indirectly when it
+changed the snan_bit_is_one setting).
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h | 7 +++++++
+ target/mips/msa.c        | 3 +++
+files changed, 10 insertions(+)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
++    /*
++     * With nan2008, the default NaN value has the sign bit clear and the
++     * frac msb set; with the older mode, the sign bit is clear, and all
++     * frac bits except the msb are set.
++     */
++    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
++                                  &env->active_fpu.fp_status);
+ }
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* Inf * 0 + NaN returns the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                               &env->active_tc.msa_fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000,
++                                  &env->active_tc.msa_fp_status);
+ }
+--
+.34.1

-New patch
+[PULL 50/72] target/openrisc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for openrisc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-45-peter.maydell@linaro.org
+---
+ target/openrisc/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/openrisc/cpu.c b/target/openrisc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/openrisc/cpu.c
++++ b/target/openrisc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void openrisc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &cpu->env.fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000, &cpu->env.fp_status);
+ #ifndef CONFIG_USER_ONLY
+     cpu->env.picmr = 0x00000000;
+--
+.34.1

-New patch
+[PULL 51/72] target/ppc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for ppc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-46-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c | 4 ++++
+file changed, 4 insertions(+)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
++    /* Default NaN: sign bit clear, set frac msb */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
++    set_float_default_nan_pattern(0b01000000, &env->vec_status);
++
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+--
+.34.1

-New patch
+[PULL 52/72] target/sh4: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for sh4. Note that sh4
+is one of the only three targets (the others being HPPA and
+sometimes MIPS) that has snan_bit_is_one set.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-47-peter.maydell@linaro.org
+---
+ target/sh4/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/sh4/cpu.c b/target/sh4/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sh4/cpu.c
++++ b/target/sh4/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void superh_cpu_reset_hold(Object *obj, ResetType type)
+     set_flush_to_zero(1, &env->fp_status);
+ #endif
+     set_default_nan_mode(1, &env->fp_status);
++    /* sign bit clear, set all frac bits other than msb */
++    set_float_default_nan_pattern(0b00111111, &env->fp_status);
+ }
+ static void superh_cpu_disas_set_info(CPUState *cpu, disassemble_info *info)
+--
+.34.1

-New patch
+[PULL 53/72] target/rx: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for rx.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-48-peter.maydell@linaro.org
+---
+ target/rx/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/rx/cpu.c b/target/rx/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/rx/cpu.c
++++ b/target/rx/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void rx_cpu_reset_hold(Object *obj, ResetType type)
+      * then prefer dest over source", which is float_2nan_prop_s_ab.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN value: sign bit clear, set frac msb */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ }
+ static ObjectClass *rx_cpu_class_by_name(const char *cpu_model)
+--
+.34.1

-New patch
+[PULL 54/72] target/s390x: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for s390x.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-49-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
++        /* Default NaN value: sign bit clear, frac msb set */
++        set_float_default_nan_pattern(0b01000000, &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+--
+.34.1

-New patch
+[PULL 55/72] target/sparc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for SPARC, and remove
+the ifdef from parts64_default_nan.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 5 +----
+files changed, 3 insertions(+), 4 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN value: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC)
+-        /* Sign bit clear, all frac bits set */
+-        dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HEXAGON)
++#if defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+ #else
+--
+.34.1

-[PULL 22/27] target/arm/translate.c: Delete/amend incorrect comments
+[PULL 56/72] target/xtensa: Set default NaN pattern explicitly
-In arm_tr_init_disas_context() we have a FIXME comment that suggests
+Set the default NaN pattern explicitly for xtensa.
 "cpu_M0 can probably be the same as cpu_V0".  This isn't in fact
 possible: cpu_V0 is used as a temporary inside gen_iwmmxt_shift(),
 and that function is called in various places where cpu_M0 contains a
 live value (i.e.  between gen_op_iwmmxt_movq_M0_wRn() and
 gen_op_iwmmxt_movq_wRn_M0() calls).  Remove the comment.
 We also have a comment on the declarations of cpu_V0/V1/M0 which
 claims they're "for efficiency".  This isn't true with modern TCG, so
 replace this comment with one which notes that they're only used with
 the iwmmxt decode.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803132815.3861-1-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-51-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 4 ++--
+ target/xtensa/cpu.c | 2 ++
-file changed, 2 insertions(+), 2 deletions(-)
+file changed, 2 insertions(+)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/xtensa/cpu.c
-+++ b/target/arm/translate.c
++++ b/target/xtensa/cpu.c
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
- #define IS_USER(s) (s->user)
+     /* For inf * 0 + NaN, return the input NaN */
- #endif
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_no_signaling_nans(!dfpu, &env->fp_status);
--/* We reuse the same 64-bit temporaries for efficiency.  */
++    /* Default NaN value: sign bit clear, set frac msb */
-+/* These are TCG temporaries used only by the legacy iwMMXt decoder */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
- static TCGv_i64 cpu_V0, cpu_V1, cpu_M0;
+     xtensa_use_first_nan(env, !dfpu);
 +/* These are TCG globals which alias CPUARMState fields */
  static TCGv_i32 cpu_R[16];
  TCGv_i32 cpu_CF, cpu_NF, cpu_VF, cpu_ZF;
  TCGv_i64 cpu_exclusive_addr;
@@ -XXX,XX +XXX,XX @@ static void arm_tr_init_disas_context(DisasContextBase *dcbase, CPUState *cs)
      cpu_V0 = tcg_temp_new_i64();
      cpu_V1 = tcg_temp_new_i64();
 -    /* FIXME: cpu_M0 can probably be the same as cpu_V0.  */
      cpu_M0 = tcg_temp_new_i64();
  }
 --
-.20.1
+.34.1

-[PULL 14/27] target/arm: Pull handling of XScale insns out of disas_coproc_insn()
+[PULL 57/72] target/hexagon: Set default NaN pattern explicitly
-At the moment we check for XScale/iwMMXt insns inside
+Set the default NaN pattern explicitly for hexagon.
-disas_coproc_insn(): for CPUs with ARM_FEATURE_XSCALE all copro insns
+Remove the ifdef from parts64_default_nan(); the only
-with cp 0 or 1 are handled specially.  This works, but is an odd
+remaining unconverted targets all use the default case.
 place for this check, because disas_coproc_insn() is called from both
 the Arm and Thumb decoders but the XScale case never applies for
 Thumb (all the XScale CPUs were ARMv5, which has only Thumb1, not
 Thumb2 with the 32-bit coprocessor insn encodings).  It also makes it
 awkward to convert the real copro access insns to decodetree.
 Move the identification of XScale out to its own function
 which is only called from disas_arm_insn().
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-2-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 44 ++++++++++++++++++++++++++++--------------
+ target/hexagon/cpu.c           | 2 ++
-file changed, 29 insertions(+), 15 deletions(-)
+ fpu/softfloat-specialize.c.inc | 5 -----
 files changed, 2 insertions(+), 5 deletions(-)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/hexagon/cpu.c
-+++ b/target/arm/translate.c
++++ b/target/hexagon/cpu.c
-@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
-     cpnum = (insn >> 8) & 0xf;
+     set_default_nan_mode(1, &env->fp_status);
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
--    /* First check for coprocessor space used for XScale/iwMMXt insns */
++    /* Default NaN value: sign bit set, all frac bits set */
--    if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
++    set_float_default_nan_pattern(0b11111111, &env->fp_status);
 -        if (extract32(s->c15_cpar, cpnum, 1) == 0) {
 -            return 1;
 -        }
 -        if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
 -            return disas_iwmmxt_insn(s, insn);
 -        } else if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
 -            return disas_dsp_insn(s, insn);
 -        }
 -        return 1;
 -    }
 -
 -    /* Otherwise treat as a generic register access */
      is64 = (insn & (1 << 25)) == 0;
      if (!is64 && ((insn & (1 << 4)) == 0)) {
          /* cdp */
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
      return 1;
  }
-+/* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
+ static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
-+static void disas_xscale_insn(DisasContext *s, uint32_t insn)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-+{
+index XXXXXXX..XXXXXXX 100644
-+    int cpnum = (insn >> 8) & 0xf;
+--- a/fpu/softfloat-specialize.c.inc
-+
++++ b/fpu/softfloat-specialize.c.inc
-+    if (extract32(s->c15_cpar, cpnum, 1) == 0) {
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
-+        unallocated_encoding(s);
+     uint8_t dnan_pattern = status->default_nan_pattern;
-+    } else if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
-+        if (disas_iwmmxt_insn(s, insn)) {
+     if (dnan_pattern == 0) {
-+            unallocated_encoding(s);
+-#if defined(TARGET_HEXAGON)
-+        }
+-        /* Sign bit set, all frac bits set. */
-+    } else if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
+-        dnan_pattern = 0b11111111;
-+        if (disas_dsp_insn(s, insn)) {
+-#else
-+            unallocated_encoding(s);
+         /*
-+        }
+          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-+    }
+          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
-+}
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+             /* sign bit clear, set frac msb */
- /* Store a 64-bit value to a register pair.  Clobbers val.  */
+             dnan_pattern = 0b01000000;
  static void gen_storeq_reg(DisasContext *s, int rlow, int rhigh, TCGv_i64 val)
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
      case 0xc:
      case 0xd:
      case 0xe:
 -        if (((insn >> 8) & 0xe) == 10) {
 +    {
 +        /* First check for coprocessor space used for XScale/iwMMXt insns */
 +        int cpnum = (insn >> 8) & 0xf;
 +
 +        if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
 +            disas_xscale_insn(s, insn);
 +            break;
 +        }
 +
 +        if ((cpnum & 0xe) == 10) {
              /* VFP, but failed disas_vfp.  */
              goto illegal_op;
          }
-+
+-#endif
-         if (disas_coproc_insn(s, insn)) {
+     }
-             /* Coprocessor.  */
+     assert(dnan_pattern != 0);
-             goto illegal_op;
          }
          break;
 +    }
      default:
      illegal_op:
          unallocated_encoding(s);
 --
-.20.1
+.34.1

-[PULL 21/27] target/arm: Delete unused VFP_DREG macros
+[PULL 58/72] target/riscv: Set default NaN pattern explicitly
-As part of the Neon decodetree conversion we removed all
+Set the default NaN pattern explicitly for riscv.
 the uses of the VFP_DREG macros, but forgot to remove the
 macro definitions. Do so now.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20241202131347.498124-53-peter.maydell@linaro.org
 Message-id: 20200803124848.18295-1-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 15 ---------------
+ target/riscv/cpu.c | 2 ++
-file changed, 15 deletions(-)
+file changed, 2 insertions(+)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/riscv/cpu.c
-+++ b/target/arm/translate.c
++++ b/target/riscv/cpu.c
-@@ -XXX,XX +XXX,XX @@ static int disas_dsp_insn(DisasContext *s, uint32_t insn)
+@@ -XXX,XX +XXX,XX @@ static void riscv_cpu_reset_hold(Object *obj, ResetType type)
-     return 1;
+     cs->exception_index = RISCV_EXCP_NONE;
- }
+     env->load_res = -1;
+     set_default_nan_mode(1, &env->fp_status);
--#define VFP_REG_SHR(x, n) (((n) > 0) ? (x) >> (n) : (x) << -(n))
++    /* Default NaN value: sign bit clear, frac msb set */
--#define VFP_DREG(reg, insn, bigbit, smallbit) do { \
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
--    if (dc_isar_feature(aa32_simd_r32, s)) { \
+     env->vill = true;
--        reg = (((insn) >> (bigbit)) & 0x0f) \
 -              | (((insn) >> ((smallbit) - 4)) & 0x10); \
 -    } else { \
 -        if (insn & (1 << (smallbit))) \
 -            return 1; \
 -        reg = ((insn) >> (bigbit)) & 0x0f; \
 -    }} while (0)
 -
 -#define VFP_DREG_D(reg, insn) VFP_DREG(reg, insn, 12, 22)
 -#define VFP_DREG_N(reg, insn) VFP_DREG(reg, insn, 16,  7)
 -#define VFP_DREG_M(reg, insn) VFP_DREG(reg, insn,  0,  5)
 -
  static inline bool use_goto_tb(DisasContext *s, target_ulong dest)
  {
  #ifndef CONFIG_USER_ONLY
 --
-.20.1
+.34.1

-[PULL 17/27] target/arm: Tidy up disas_arm_insn()
+[PULL 59/72] target/tricore: Set default NaN pattern explicitly
-The only thing left in the "legacy decoder" is the handling
+Set the default NaN pattern explicitly for tricore.
 of disas_xscale_insn(), and we can simplify the code.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20200803111849.13368-5-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-54-peter.maydell@linaro.org
 ---
- target/arm/translate.c | 26 +++++++++-----------------
+ target/tricore/helper.c | 2 ++
-file changed, 9 insertions(+), 17 deletions(-)
+file changed, 2 insertions(+)
-diff --git a/target/arm/translate.c b/target/arm/translate.c
+diff --git a/target/tricore/helper.c b/target/tricore/helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate.c
+--- a/target/tricore/helper.c
-+++ b/target/arm/translate.c
++++ b/target/tricore/helper.c
-@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
+@@ -XXX,XX +XXX,XX @@ void fpu_set_state(CPUTriCoreState *env)
-         return;
+     set_flush_to_zero(1, &env->fp_status);
-     }
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
-     /* fall back to legacy decoder */
+     set_default_nan_mode(1, &env->fp_status);
--
++    /* Default NaN pattern: sign bit clear, frac msb set */
--    switch ((insn >> 24) & 0xf) {
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
 -    case 0xc:
 -    case 0xd:
 -    case 0xe:
 -    {
 -        /* First check for coprocessor space used for XScale/iwMMXt insns */
 -        int cpnum = (insn >> 8) & 0xf;
 -
 -        if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
 +    /* TODO: convert xscale/iwmmxt decoder to decodetree ?? */
 +    if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
 +        if (((insn & 0x0c000e00) == 0x0c000000)
 +            && ((insn & 0x03000000) != 0x03000000)) {
 +            /* Coprocessor insn, coprocessor 0 or 1 */
              disas_xscale_insn(s, insn);
 -            break;
 +            return;
          }
 -        /* fall through */
 -    }
 -    default:
 -    illegal_op:
 -        unallocated_encoding(s);
 -        break;
      }
 +
 +illegal_op:
 +    unallocated_encoding(s);
  }
- static bool thumb_insn_is_16bit(DisasContext *s, uint32_t pc, uint32_t insn)
+ uint32_t psw_read(CPUTriCoreState *env)
 --
-.20.1
+.34.1

-[PULL 23/27] target/arm: Delete unused ARM_FEATURE_CRC
+[PULL 60/72] fpu: Remove default handling for dnan_pattern
-In commit 962fcbf2efe57231a9f5df we converted the uses of the
+Now that all our targets have bene converted to explicitly specify
-ARM_FEATURE_CRC bit to use the aa32_crc32 isar_feature test
+their pattern for the default NaN value we can remove the remaining
-instead. However we forgot to remove the now-unused definition
+fallback code in parts64_default_nan().
 of the feature name in the enum. Delete it now.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
 Message-id: 20200805210848.6688-1-peter.maydell@linaro.org
 ---
- target/arm/cpu.h | 1 -
+ fpu/softfloat-specialize.c.inc | 14 --------------
-file changed, 1 deletion(-)
+file changed, 14 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/cpu.h
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ enum arm_features {
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
-     ARM_FEATURE_V8,
+     uint64_t frac;
-     ARM_FEATURE_AARCH64, /* supports 64 bit mode */
+     uint8_t dnan_pattern = status->default_nan_pattern;
-     ARM_FEATURE_CBAR, /* has cp15 CBAR */
--    ARM_FEATURE_CRC, /* ARMv8 CRC instructions */
+-    if (dnan_pattern == 0) {
-     ARM_FEATURE_CBAR_RO, /* has cp15 CBAR and it is read-only */
+-        /*
-     ARM_FEATURE_EL2, /* has EL2 Virtualization support */
+-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-     ARM_FEATURE_EL3, /* has EL3 Secure monitor support */
+-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 -         * do not have floating-point.
 -         */
 -        if (snan_bit_is_one(status)) {
 -            /* sign bit clear, set all frac bits other than msb */
 -            dnan_pattern = 0b00111111;
 -        } else {
 -            /* sign bit clear, set frac msb */
 -            dnan_pattern = 0b01000000;
 -        }
 -    }
      assert(dnan_pattern != 0);
      sign = dnan_pattern >> 7;
 --
-.20.1
+.34.1

-[PULL 07/27] hw/arm/smmuv3: Introduce smmuv3_s1_range_inval() helper
+[PULL 61/72] softfloat: Inline pickNaNMulAdd
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Let's introduce an helper for S1 IOVA range invalidation.
+Inline pickNaNMulAdd into its only caller.  This makes
-This will be used for NH_VA and NH_VAA commands. It decodes
+one assert redundant with the immediately preceding IF.
 the same fields, trace, calls the UNMAP notifiers and
 invalidate the corresponding IOTLB entries.
-At the moment, we do not support 3.2 range invalidation yet.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-So it reduces to a single IOVA invalidation.
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
-Note the leaf bit now is also decoded for the CMD_TLBI_NH_VAA
+[PMM: keep comment from old code in new location]
 command. At the moment it is only used for tracing.
 Signed-off-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20200728150815.11446-7-eric.auger@redhat.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3.c     | 36 +++++++++++++++++-------------------
+ fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
- hw/arm/trace-events |  3 +--
+ fpu/softfloat-specialize.c.inc | 54 ----------------------------------
-files changed, 18 insertions(+), 21 deletions(-)
+files changed, 40 insertions(+), 55 deletions(-)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
      }
      if (s->default_nan_mode) {
 +        /*
 +         * We guarantee not to require the target to tell us how to
 +         * pick a NaN if we're always returning the default NaN.
 +         * But if we're not in default-NaN mode then the target must
 +         * specify.
 +         */
          which = 3;
 +    } else if (infzero) {
 +        /*
 +         * Inf * 0 + NaN -- some implementations return the
 +         * default NaN here, and some return the input NaN.
 +         */
 +        switch (s->float_infzeronan_rule) {
 +        case float_infzeronan_dnan_never:
 +            which = 2;
 +            break;
 +        case float_infzeronan_dnan_always:
 +            which = 3;
 +            break;
 +        case float_infzeronan_dnan_if_qnan:
 +            which = is_qnan(c->cls) ? 3 : 2;
 +            break;
 +        default:
 +            g_assert_not_reached();
 +        }
      } else {
 -        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
 +        FloatClass cls[3] = { a->cls, b->cls, c->cls };
 +        Float3NaNPropRule rule = s->float_3nan_prop_rule;
 +
 +        assert(rule != float_3nan_prop_none);
 +        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
 +            /* We have at least one SNaN input and should prefer it */
 +            do {
 +                which = rule & R_3NAN_1ST_MASK;
 +                rule >>= R_3NAN_1ST_LENGTH;
 +            } while (!is_snan(cls[which]));
 +        } else {
 +            do {
 +                which = rule & R_3NAN_1ST_MASK;
 +                rule >>= R_3NAN_1ST_LENGTH;
 +            } while (!is_nan(cls[which]));
 +        }
      }
      if (which == 3) {
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
      }
  }
-+static void smmuv3_s1_range_inval(SMMUState *s, Cmd *cmd)
+-/*----------------------------------------------------------------------------
-+{
+-| Select which NaN to propagate for a three-input operation.
-+    dma_addr_t addr = CMD_ADDR(cmd);
+-| For the moment we assume that no CPU needs the 'larger significand'
-+    uint8_t type = CMD_TYPE(cmd);
+-| information.
-+    uint16_t vmid = CMD_VMID(cmd);
+-| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-+    bool leaf = CMD_LEAF(cmd);
+-*----------------------------------------------------------------------------*/
-+    int asid = -1;
+-static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-+
+-                         bool infzero, bool have_snan, float_status *status)
-+    if (type == SMMU_CMD_TLBI_NH_VA) {
+-{
-+        asid = CMD_ASID(cmd);
+-    FloatClass cls[3] = { a_cls, b_cls, c_cls };
-+    }
+-    Float3NaNPropRule rule = status->float_3nan_prop_rule;
-+    trace_smmuv3_s1_range_inval(vmid, asid, addr, leaf);
+-    int which;
 +    smmuv3_inv_notifiers_iova(s, asid, addr);
 +    smmu_iotlb_inv_iova(s, asid, addr);
 +}
 +
  static int smmuv3_cmdq_consume(SMMUv3State *s)
  {
      SMMUState *bs = ARM_SMMU(s);
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
              smmu_iotlb_inv_all(bs);
              break;
          case SMMU_CMD_TLBI_NH_VAA:
 -        {
 -            dma_addr_t addr = CMD_ADDR(&cmd);
 -            uint16_t vmid = CMD_VMID(&cmd);
 -
--            trace_smmuv3_cmdq_tlbi_nh_vaa(vmid, addr);
+-    /*
--            smmuv3_inv_notifiers_iova(bs, -1, addr);
+-     * We guarantee not to require the target to tell us how to
--            smmu_iotlb_inv_iova(bs, -1, addr);
+-     * pick a NaN if we're always returning the default NaN.
--            break;
+-     * But if we're not in default-NaN mode then the target must
 -     * specify.
 -     */
 -    assert(!status->default_nan_mode);
 -
 -    if (infzero) {
 -        /*
 -         * Inf * 0 + NaN -- some implementations return the default NaN here,
 -         * and some return the input NaN.
 -         */
 -        switch (status->float_infzeronan_rule) {
 -        case float_infzeronan_dnan_never:
 -            return 2;
 -        case float_infzeronan_dnan_always:
 -            return 3;
 -        case float_infzeronan_dnan_if_qnan:
 -            return is_qnan(c_cls) ? 3 : 2;
 -        default:
 -            g_assert_not_reached();
 -        }
-         case SMMU_CMD_TLBI_NH_VA:
+-    }
 -        {
 -            uint16_t asid = CMD_ASID(&cmd);
 -            uint16_t vmid = CMD_VMID(&cmd);
 -            dma_addr_t addr = CMD_ADDR(&cmd);
 -            bool leaf = CMD_LEAF(&cmd);
 -
--            trace_smmuv3_cmdq_tlbi_nh_va(vmid, asid, addr, leaf);
+-    assert(rule != float_3nan_prop_none);
--            smmuv3_inv_notifiers_iova(bs, asid, addr);
+-    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
--            smmu_iotlb_inv_iova(bs, asid, addr);
+-        /* We have at least one SNaN input and should prefer it */
-+            smmuv3_s1_range_inval(bs, &cmd);
+-        do {
-             break;
+-            which = rule & R_3NAN_1ST_MASK;
--        }
+-            rule >>= R_3NAN_1ST_LENGTH;
-         case SMMU_CMD_TLBI_EL3_ALL:
+-        } while (!is_snan(cls[which]));
-         case SMMU_CMD_TLBI_EL3_VA:
+-    } else {
-         case SMMU_CMD_TLBI_EL2_ALL:
+-        do {
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
+-            which = rule & R_3NAN_1ST_MASK;
-index XXXXXXX..XXXXXXX 100644
+-            rule >>= R_3NAN_1ST_LENGTH;
---- a/hw/arm/trace-events
+-        } while (!is_nan(cls[which]));
-+++ b/hw/arm/trace-events
+-    }
-@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
+-    return which;
- smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
+-}
- smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid %d (hits=%d, misses=%d, hit rate=%d)"
+-
- smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid %d (hits=%d, misses=%d, hit rate=%d)"
+ /*----------------------------------------------------------------------------
--smmuv3_cmdq_tlbi_nh_va(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
+ | Returns 1 if the double-precision floating-point value `a' is a quiet
--smmuv3_cmdq_tlbi_nh_vaa(int vmid, uint64_t addr) "vmid =%d addr=0x%"PRIx64
+ | NaN; otherwise returns 0.
 +smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
  smmuv3_cmdq_tlbi_nh(void) ""
  smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
  smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
 --
-.20.1
+.34.1

-[PULL 03/27] hw/arm/smmu-common: Add IOTLB helpers
+[PULL 62/72] softfloat: Use goto for default nan case in pick_nan_muladd
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Add two helpers: one to lookup for a given IOTLB entry and
+Remove "3" as a special case for which and simply
-one to insert a new entry. We also move the tracing there.
+branch to return the desired value.
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20200728150815.11446-3-eric.auger@redhat.com
+Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/hw/arm/smmu-common.h |  2 ++
+ fpu/softfloat-parts.c.inc | 20 ++++++++++----------
- hw/arm/smmu-common.c         | 36 ++++++++++++++++++++++++++++++++++++
+file changed, 10 insertions(+), 10 deletions(-)
  hw/arm/smmuv3.c              | 26 ++------------------------
  hw/arm/trace-events          |  5 +++--
 files changed, 43 insertions(+), 26 deletions(-)
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmu-common.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/arm/smmu-common.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+          * But if we're not in default-NaN mode then the target must
- #define SMMU_IOTLB_MAX_SIZE 256
+          * specify.
+          */
-+IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
+-        which = 3;
-+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry);
++        goto default_nan;
- void smmu_iotlb_inv_all(SMMUState *s);
+     } else if (infzero) {
- void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
+         /*
- void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova);
+          * Inf * 0 + NaN -- some implementations return the
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-index XXXXXXX..XXXXXXX 100644
+          */
---- a/hw/arm/smmu-common.c
+         switch (s->float_infzeronan_rule) {
-+++ b/hw/arm/smmu-common.c
+         case float_infzeronan_dnan_never:
-@@ -XXX,XX +XXX,XX @@
+-            which = 2;
+             break;
- /* IOTLB Management */
+         case float_infzeronan_dnan_always:
+-            which = 3;
-+IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+-            break;
-+                                 hwaddr iova)
++            goto default_nan;
-+{
+         case float_infzeronan_dnan_if_qnan:
-+    SMMUIOTLBKey key = {.asid = cfg->asid, .iova = iova};
+-            which = is_qnan(c->cls) ? 3 : 2;
-+    IOMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
++            if (is_qnan(c->cls)) {
-+
++                goto default_nan;
-+    if (entry) {
++            }
-+        cfg->iotlb_hits++;
+             break;
-+        trace_smmu_iotlb_lookup_hit(cfg->asid, iova,
+         default:
-+                                    cfg->iotlb_hits, cfg->iotlb_misses,
+             g_assert_not_reached();
-+                                    100 * cfg->iotlb_hits /
+         }
-+                                    (cfg->iotlb_hits + cfg->iotlb_misses));
++        which = 2;
-+    } else {
+     } else {
-+        cfg->iotlb_misses++;
+         FloatClass cls[3] = { a->cls, b->cls, c->cls };
-+        trace_smmu_iotlb_lookup_miss(cfg->asid, iova,
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
-+                                     cfg->iotlb_hits, cfg->iotlb_misses,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-+                                     100 * cfg->iotlb_hits /
+         }
 +                                     (cfg->iotlb_hits + cfg->iotlb_misses));
 +    }
 +    return entry;
 +}
 +
 +void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
 +{
 +    SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
 +
 +    if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
 +        smmu_iotlb_inv_all(bs);
 +    }
 +
 +    key->asid = cfg->asid;
 +    key->iova = entry->iova;
 +    trace_smmu_iotlb_insert(cfg->asid, entry->iova);
 +    g_hash_table_insert(bs->iotlb, key, entry);
 +}
 +
  inline void smmu_iotlb_inv_all(SMMUState *s)
  {
      trace_smmu_iotlb_inv_all();
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
          .addr_mask = ~(hwaddr)0,
          .perm = IOMMU_NONE,
      };
 -    SMMUIOTLBKey key, *new_key;
      qemu_mutex_lock(&s->mutex);
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
      page_mask = (1ULL << (tt->granule_sz)) - 1;
      aligned_addr = addr & ~page_mask;
 -    key.asid = cfg->asid;
 -    key.iova = aligned_addr;
 -
 -    cached_entry = g_hash_table_lookup(bs->iotlb, &key);
 +    cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
      if (cached_entry) {
 -        cfg->iotlb_hits++;
 -        trace_smmu_iotlb_cache_hit(cfg->asid, aligned_addr,
 -                                   cfg->iotlb_hits, cfg->iotlb_misses,
 -                                   100 * cfg->iotlb_hits /
 -                                   (cfg->iotlb_hits + cfg->iotlb_misses));
          if ((flag & IOMMU_WO) && !(cached_entry->perm & IOMMU_WO)) {
              status = SMMU_TRANS_ERROR;
              if (event.record_trans_faults) {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
          goto epilogue;
      }
--    cfg->iotlb_misses++;
+-    if (which == 3) {
--    trace_smmu_iotlb_cache_miss(cfg->asid, addr & ~page_mask,
+-        parts_default_nan(a, s);
--                                cfg->iotlb_hits, cfg->iotlb_misses,
+-        return a;
 -                                100 * cfg->iotlb_hits /
 -                                (cfg->iotlb_hits + cfg->iotlb_misses));
 -
 -    if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
 -        smmu_iotlb_inv_all(bs);
 -    }
 -
-     cached_entry = g_new0(IOMMUTLBEntry, 1);
+     switch (which) {
+     case 0:
-     if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
+         break;
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-         }
+         parts_silence_nan(a, s);
          status = SMMU_TRANS_ERROR;
      } else {
 -        new_key = g_new0(SMMUIOTLBKey, 1);
 -        new_key->asid = cfg->asid;
 -        new_key->iova = aligned_addr;
 -        g_hash_table_insert(bs->iotlb, new_key, cached_entry);
 +        smmu_iotlb_insert(bs, cfg, cached_entry);
          status = SMMU_TRANS_SUCCESS;
      }
+     return a;
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
++
-index XXXXXXX..XXXXXXX 100644
++ default_nan:
---- a/hw/arm/trace-events
++    parts_default_nan(a, s);
-+++ b/hw/arm/trace-events
++    return a;
-@@ -XXX,XX +XXX,XX @@ smmu_iotlb_inv_all(void) "IOTLB invalidate all"
+ }
- smmu_iotlb_inv_asid(uint16_t asid) "IOTLB invalidate asid=%d"
- smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
+ /*
  smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
 +smmu_iotlb_lookup_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 +smmu_iotlb_lookup_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 +smmu_iotlb_insert(uint16_t asid, uint64_t addr) "IOTLB ++ asid=%d addr=0x%"PRIx64
  # smmuv3.c
  smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_tlbi_nh_va(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d a
  smmuv3_cmdq_tlbi_nh_vaa(int vmid, uint64_t addr) "vmid =%d addr=0x%"PRIx64
  smmuv3_cmdq_tlbi_nh(void) ""
  smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
 -smmu_iotlb_cache_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 -smmu_iotlb_cache_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
  smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
  smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
  smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
 --
-.20.1
+.34.1

-[PULL 11/27] hw/arm/smmuv3: Support HAD and advertise SMMUv3.1 support
+[PULL 63/72] softfloat: Remove which from parts_pick_nan_muladd
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-HAD is a mandatory features with SMMUv3.1 if S1P is set, which is
+Assign the pointer return value to 'a' directly,
-our case. Other 3.1 mandatory features come with S2P which we don't
+rather than going through an intermediary index.
 have.
-So let's support HAD and advertise SMMUv3.1 support in AIDR.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-HAD support allows the CD to disable hierarchical attributes, ie.
+Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
 if the HAD0/1 bit is set, the APTable field of table descriptors
 walked through TTB0/1 is ignored.
 Signed-off-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Message-id: 20200728150815.11446-11-eric.auger@redhat.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h     | 2 ++
+ fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
- include/hw/arm/smmu-common.h | 1 +
+file changed, 10 insertions(+), 22 deletions(-)
  hw/arm/smmu-common.c         | 2 +-
  hw/arm/smmuv3.c              | 6 +++++-
  hw/arm/trace-events          | 2 +-
 files changed, 10 insertions(+), 3 deletions(-)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3-internal.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ REG32(IDR1,                0x4)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+                                             FloatPartsN *c, float_status *s,
- REG32(IDR2,                0x8)
+                                             int ab_mask, int abc_mask)
- REG32(IDR3,                0xc)
+ {
-+     FIELD(IDR3, HAD,         2, 1);
+-    int which;
- REG32(IDR4,                0x10)
+     bool infzero = (ab_mask == float_cmask_infzero);
- REG32(IDR5,                0x14)
+     bool have_snan = (abc_mask & float_cmask_snan);
-      FIELD(IDR5, OAS,         0, 3);
++    FloatPartsN *ret;
-@@ -XXX,XX +XXX,XX @@ static inline int pa_range(STE *ste)
-         lo = (x)->word[(sel) * 2 + 2] & ~0xfULL;            \
+     if (unlikely(have_snan)) {
-         hi | lo;                                            \
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-     })
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-+#define CD_HAD(x, sel)   extract32((x)->word[(sel) * 2 + 2], 1, 1)
+         default:
+             g_assert_not_reached();
  #define CD_TSZ(x, sel)   extract32((x)->word[0], (16 * (sel)) + 0, 6)
  #define CD_TG(x, sel)    extract32((x)->word[0], (16 * (sel)) + 6, 2)
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/hw/arm/smmu-common.h
 +++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransTableInfo {
      uint64_t ttb;              /* TT base address */
      uint8_t tsz;               /* input range, ie. 2^(64 -tsz)*/
      uint8_t granule_sz;        /* granule page shift */
 +    bool had;                  /* hierarchical attribute disable */
  } SMMUTransTableInfo;
  typedef struct SMMUTLBEntry {
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
          if (is_table_pte(pte, level)) {
              ap = PTE_APTABLE(pte);
 -            if (is_permission_fault(ap, perm)) {
 +            if (is_permission_fault(ap, perm) && !tt->had) {
                  info->type = SMMU_PTW_ERR_PERMISSION;
                  goto error;
              }
 diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmuv3.c
 +++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
      s->idr[1] = FIELD_DP32(s->idr[1], IDR1, EVENTQS, SMMU_EVENTQS);
      s->idr[1] = FIELD_DP32(s->idr[1], IDR1, CMDQS,   SMMU_CMDQS);
 +    s->idr[3] = FIELD_DP32(s->idr[3], IDR3, HAD, 1);
 +
     /* 4K and 64K granule support */
      s->idr[5] = FIELD_DP32(s->idr[5], IDR5, GRAN4K, 1);
      s->idr[5] = FIELD_DP32(s->idr[5], IDR5, GRAN64K, 1);
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
      s->features = 0;
      s->sid_split = 0;
 +    s->aidr = 0x1;
  }
  static int smmu_get_ste(SMMUv3State *s, dma_addr_t addr, STE *buf,
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
          if (tt->ttb & ~(MAKE_64BIT_MASK(0, cfg->oas))) {
              goto bad_cd;
          }
--        trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz);
+-        which = 2;
-+        tt->had = CD_HAD(cd, i);
++        ret = c;
-+        trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz, tt->had);
+     } else {
 -        FloatClass cls[3] = { a->cls, b->cls, c->cls };
 +        FloatPartsN *val[3] = { a, b, c };
          Float3NaNPropRule rule = s->float_3nan_prop_rule;
          assert(rule != float_3nan_prop_none);
          if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
              /* We have at least one SNaN input and should prefer it */
              do {
 -                which = rule & R_3NAN_1ST_MASK;
 +                ret = val[rule & R_3NAN_1ST_MASK];
                  rule >>= R_3NAN_1ST_LENGTH;
 -            } while (!is_snan(cls[which]));
 +            } while (!is_snan(ret->cls));
          } else {
              do {
 -                which = rule & R_3NAN_1ST_MASK;
 +                ret = val[rule & R_3NAN_1ST_MASK];
                  rule >>= R_3NAN_1ST_LENGTH;
 -            } while (!is_nan(cls[which]));
 +            } while (!is_nan(ret->cls));
          }
      }
-     event->record_trans_faults = CD_R(cd);
+-    switch (which) {
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
+-    case 0:
-index XXXXXXX..XXXXXXX 100644
+-        break;
---- a/hw/arm/trace-events
+-    case 1:
-+++ b/hw/arm/trace-events
+-        a = b;
-@@ -XXX,XX +XXX,XX @@ smmuv3_translate_abort(const char *n, uint16_t sid, uint64_t addr, bool is_write
+-        break;
- smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm) "%s sid=%d iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x"
+-    case 2:
- smmuv3_get_cd(uint64_t addr) "CD addr: 0x%"PRIx64
+-        a = c;
- smmuv3_decode_cd(uint32_t oas) "oas=%d"
+-        break;
--smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d"
+-    default:
-+smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz, bool had) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d had:%d"
+-        g_assert_not_reached();
- smmuv3_cmdq_cfgi_ste(int streamid) "streamid =%d"
++    if (is_snan(ret->cls)) {
- smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
++        parts_silence_nan(ret, s);
- smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
+     }
 -    if (is_snan(a->cls)) {
 -        parts_silence_nan(a, s);
 -    }
 -    return a;
 +    return ret;
   default_nan:
      parts_default_nan(a, s);
 --
-.20.1
+.34.1

-[PULL 09/27] hw/arm/smmuv3: Fix IIDR offset
+[PULL 64/72] softfloat: Pad array size in pick_nan_muladd
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-The SMMU IIDR register is at 0x018 offset.
+While all indices into val[] should be in [0-2], the mask
 applied is two bits.  To help static analysis see there is
 no possibility of read beyond the end of the array, pad the
 array to 4 entries, with the final being (implicitly) NULL.
-Fixes: 10a83cb9887 ("hw/arm/smmuv3: Skeleton")
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
 Message-id: 20200728150815.11446-9-eric.auger@redhat.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h | 2 +-
+ fpu/softfloat-parts.c.inc | 2 +-
 file changed, 1 insertion(+), 1 deletion(-)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmuv3-internal.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ REG32(IDR5,                0x14)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         }
- #define SMMU_IDR5_OAS 4
+         ret = c;
+     } else {
--REG32(IIDR,                0x1c)
+-        FloatPartsN *val[3] = { a, b, c };
-+REG32(IIDR,                0x18)
++        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
- REG32(CR0,                 0x20)
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
-     FIELD(CR0, SMMU_ENABLE,   0, 1)
-     FIELD(CR0, EVENTQEN,      2, 1)
+         assert(rule != float_3nan_prop_none);
 --
-.20.1
+.34.1

-[PULL 06/27] hw/arm/smmu-common: Manage IOTLB block entries
+[PULL 65/72] softfloat: Move propagateFloatx80NaN to softfloat.c
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-At the moment each entry in the IOTLB corresponds to a page sized
+This function is part of the public interface and
-mapping (4K, 16K or 64K), even if the page belongs to a mapped
+is not "specialized" to any target in any way.
 block. In case of block mapping this unefficiently consumes IOTLB
 entries.
-Change the value of the entry so that it reflects the actual
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 mapping it belongs to (block or page start address and size).
 Also the level/tg of the entry is encoded in the key. In subsequent
 patches we will enable range invalidation. This latter is able
 to provide the level/tg of the entry.
 Encoding the level/tg directly in the key will allow to invalidate
 using g_hash_table_remove() when num_pages equals to 1.
 Signed-off-by: Eric Auger <eric.auger@redhat.com>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20200728150815.11446-6-eric.auger@redhat.com
+Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmu-internal.h       |  7 ++++
+ fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
- include/hw/arm/smmu-common.h | 10 ++++--
+ fpu/softfloat-specialize.c.inc | 52 ----------------------------------
- hw/arm/smmu-common.c         | 67 ++++++++++++++++++++++++++----------
+files changed, 52 insertions(+), 52 deletions(-)
  hw/arm/smmuv3.c              |  6 ++--
  hw/arm/trace-events          |  2 +-
 files changed, 67 insertions(+), 25 deletions(-)
-diff --git a/hw/arm/smmu-internal.h b/hw/arm/smmu-internal.h
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-internal.h
+--- a/fpu/softfloat.c
-+++ b/hw/arm/smmu-internal.h
++++ b/fpu/softfloat.c
-@@ -XXX,XX +XXX,XX @@ uint64_t iova_level_offset(uint64_t iova, int inputsize,
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
      *zExpPtr = 1 - shiftCount;
  }
- #define SMMU_IOTLB_ASID(key) ((key).asid)
++/*----------------------------------------------------------------------------
 +| Takes two extended double-precision floating-point values `a' and `b', one
 +| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
 +| `b' is a signaling NaN, the invalid exception is raised.
 +*----------------------------------------------------------------------------*/
 +
-+typedef struct SMMUIOTLBPageInvInfo {
++floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-+    int asid;
++{
-+    uint64_t iova;
++    bool aIsLargerSignificand;
-+    uint64_t mask;
++    FloatClass a_cls, b_cls;
 +} SMMUIOTLBPageInvInfo;
 +
- #endif
++    /* This is not complete, but is good enough for pickNaN.  */
-diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
++    a_cls = (!floatx80_is_any_nan(a)
-index XXXXXXX..XXXXXXX 100644
++             ? float_class_normal
---- a/include/hw/arm/smmu-common.h
++             : floatx80_is_signaling_nan(a, status)
-+++ b/include/hw/arm/smmu-common.h
++             ? float_class_snan
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPciBus {
++             : float_class_qnan);
- typedef struct SMMUIOTLBKey {
++    b_cls = (!floatx80_is_any_nan(b)
-     uint64_t iova;
++             ? float_class_normal
-     uint16_t asid;
++             : floatx80_is_signaling_nan(b, status)
-+    uint8_t tg;
++             ? float_class_snan
-+    uint8_t level;
++             : float_class_qnan);
  } SMMUIOTLBKey;
  typedef struct SMMUState {
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
  #define SMMU_IOTLB_MAX_SIZE 256
 -SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
 +SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
 +                                SMMUTransTableInfo *tt, hwaddr iova);
  void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
 -SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova);
 +SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
 +                                uint8_t tg, uint8_t level);
  void smmu_iotlb_inv_all(SMMUState *s);
  void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
 -void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova);
 +void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova);
  /* Unmap the range of all the notifiers registered to any IOMMU mr */
  void smmu_inv_notifiers_all(SMMUState *s);
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static guint smmu_iotlb_key_hash(gconstpointer v)
      /* Jenkins hash */
      a = b = c = JHASH_INITVAL + sizeof(*key);
 -    a += key->asid;
 +    a += key->asid + key->level + key->tg;
      b += extract64(key->iova, 0, 32);
      c += extract64(key->iova, 32, 32);
@@ -XXX,XX +XXX,XX @@ static guint smmu_iotlb_key_hash(gconstpointer v)
  static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
  {
 -    const SMMUIOTLBKey *k1 = v1;
 -    const SMMUIOTLBKey *k2 = v2;
 +    SMMUIOTLBKey *k1 = (SMMUIOTLBKey *)v1, *k2 = (SMMUIOTLBKey *)v2;
 -    return (k1->asid == k2->asid) && (k1->iova == k2->iova);
 +    return (k1->asid == k2->asid) && (k1->iova == k2->iova) &&
 +           (k1->level == k2->level) && (k1->tg == k2->tg);
  }
 -SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova)
 +SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
 +                                uint8_t tg, uint8_t level)
  {
 -    SMMUIOTLBKey key = {.asid = asid, .iova = iova};
 +    SMMUIOTLBKey key = {.asid = asid, .iova = iova, .tg = tg, .level = level};
      return key;
  }
  SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
 -                                hwaddr iova)
 +                                SMMUTransTableInfo *tt, hwaddr iova)
  {
 -    SMMUIOTLBKey key = smmu_get_iotlb_key(cfg->asid, iova);
 -    SMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
 +    uint8_t tg = (tt->granule_sz - 10) / 2;
 +    uint8_t inputsize = 64 - tt->tsz;
 +    uint8_t stride = tt->granule_sz - 3;
 +    uint8_t level = 4 - (inputsize - 4) / stride;
 +    SMMUTLBEntry *entry = NULL;
 +
-+    while (level <= 3) {
++    if (is_snan(a_cls) || is_snan(b_cls)) {
-+        uint64_t subpage_size = 1ULL << level_shift(level, tt->granule_sz);
++        float_raise(float_flag_invalid, status);
-+        uint64_t mask = subpage_size - 1;
++    }
 +        SMMUIOTLBKey key;
 +
-+        key = smmu_get_iotlb_key(cfg->asid, iova & ~mask, tg, level);
++    if (status->default_nan_mode) {
-+        entry = g_hash_table_lookup(bs->iotlb, &key);
++        return floatx80_default_nan(status);
-+        if (entry) {
++    }
-+            break;
++
 +    if (a.low < b.low) {
 +        aIsLargerSignificand = 0;
 +    } else if (b.low < a.low) {
 +        aIsLargerSignificand = 1;
 +    } else {
 +        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
 +    }
 +
 +    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
 +        if (is_snan(b_cls)) {
 +            return floatx80_silence_nan(b, status);
 +        }
-+        level++;
++        return b;
 +    } else {
 +        if (is_snan(a_cls)) {
 +            return floatx80_silence_nan(a, status);
 +        }
 +        return a;
 +    }
-     if (entry) {
-         cfg->iotlb_hits++;
-@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
- void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *new)
- {
-     SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
-+    uint8_t tg = (new->granule - 10) / 2;
-     if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
-         smmu_iotlb_inv_all(bs);
-     }
--    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova);
--    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova);
-+    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova, tg, new->level);
-+    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova, tg, new->level);
-     g_hash_table_insert(bs->iotlb, key, new);
- }
-@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
-     return SMMU_IOTLB_ASID(*iotlb_key) == asid;
- }
--inline void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova)
-+static gboolean smmu_hash_remove_by_asid_iova(gpointer key, gpointer value,
-+                                              gpointer user_data)
- {
--    SMMUIOTLBKey key = smmu_get_iotlb_key(asid, iova);
-+    SMMUTLBEntry *iter = (SMMUTLBEntry *)value;
-+    IOMMUTLBEntry *entry = &iter->entry;
-+    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
-+    SMMUIOTLBKey iotlb_key = *(SMMUIOTLBKey *)key;
-+
-+    if (info->asid >= 0 && info->asid != SMMU_IOTLB_ASID(iotlb_key)) {
-+        return false;
-+    }
-+    return (info->iova & ~entry->addr_mask) == entry->iova;
 +}
 +
-+inline void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova)
+ /*----------------------------------------------------------------------------
-+{
+ | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
-+    SMMUIOTLBPageInvInfo info = {.asid = asid, .iova = iova};
+ | and extended significand formed by the concatenation of `zSig0' and `zSig1',
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-     trace_smmu_iotlb_inv_iova(asid, iova);
+index XXXXXXX..XXXXXXX 100644
--    g_hash_table_remove(s->iotlb, &key);
+--- a/fpu/softfloat-specialize.c.inc
-+    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_iova, &info);
++++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
      return a;
  }
- inline void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
+-/*----------------------------------------------------------------------------
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
+-| Takes two extended double-precision floating-point values `a' and `b', one
-     baseaddr = extract64(tt->ttb, 0, 48);
+-| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
-     baseaddr &= ~indexmask;
+-| `b' is a signaling NaN, the invalid exception is raised.
+-*----------------------------------------------------------------------------*/
 -    tlbe->entry.iova = iova;
 -    tlbe->entry.addr_mask = (1 << granule_sz) - 1;
 -
-     while (level <= 3) {
+-floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-         uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
+-{
-         uint64_t mask = subpage_size - 1;
+-    bool aIsLargerSignificand;
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
+-    FloatClass a_cls, b_cls;
-             goto error;
+-
-         }
+-    /* This is not complete, but is good enough for pickNaN.  */
+-    a_cls = (!floatx80_is_any_nan(a)
--        tlbe->entry.translated_addr = gpa + (iova & mask);
+-             ? float_class_normal
-+        tlbe->entry.translated_addr = gpa;
+-             : floatx80_is_signaling_nan(a, status)
-+        tlbe->entry.iova = iova & ~mask;
+-             ? float_class_snan
-+        tlbe->entry.addr_mask = mask;
+-             : float_class_qnan);
-         tlbe->entry.perm = PTE_AP_TO_PERM(ap);
+-    b_cls = (!floatx80_is_any_nan(b)
-         tlbe->level = level;
+-             ? float_class_normal
-         tlbe->granule = granule_sz;
+-             : floatx80_is_signaling_nan(b, status)
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
+-             ? float_class_snan
-index XXXXXXX..XXXXXXX 100644
+-             : float_class_qnan);
---- a/hw/arm/smmuv3.c
+-
-+++ b/hw/arm/smmuv3.c
+-    if (is_snan(a_cls) || is_snan(b_cls)) {
-@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
+-        float_raise(float_flag_invalid, status);
-     page_mask = (1ULL << (tt->granule_sz)) - 1;
+-    }
-     aligned_addr = addr & ~page_mask;
+-
+-    if (status->default_nan_mode) {
--    cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
+-        return floatx80_default_nan(status);
-+    cached_entry = smmu_iotlb_lookup(bs, cfg, tt, aligned_addr);
+-    }
-     if (cached_entry) {
+-
-         if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
+-    if (a.low < b.low) {
-             status = SMMU_TRANS_ERROR;
+-        aIsLargerSignificand = 0;
-@@ -XXX,XX +XXX,XX @@ epilogue:
+-    } else if (b.low < a.low) {
-     case SMMU_TRANS_SUCCESS:
+-        aIsLargerSignificand = 1;
-         entry.perm = flag;
+-    } else {
-         entry.translated_addr = cached_entry->entry.translated_addr +
+-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
--                                    (addr & page_mask);
+-    }
-+                                    (addr & cached_entry->entry.addr_mask);
+-
-         entry.addr_mask = cached_entry->entry.addr_mask;
+-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
+-        if (is_snan(b_cls)) {
-                                        entry.translated_addr, entry.perm);
+-            return floatx80_silence_nan(b, status);
-@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
+-        }
+-        return b;
-             trace_smmuv3_cmdq_tlbi_nh_vaa(vmid, addr);
+-    } else {
-             smmuv3_inv_notifiers_iova(bs, -1, addr);
+-        if (is_snan(a_cls)) {
--            smmu_iotlb_inv_all(bs);
+-            return floatx80_silence_nan(a, status);
-+            smmu_iotlb_inv_iova(bs, -1, addr);
+-        }
-             break;
+-        return a;
-         }
+-    }
-         case SMMU_CMD_TLBI_NH_VA:
+-}
-diff --git a/hw/arm/trace-events b/hw/arm/trace-events
+-
-index XXXXXXX..XXXXXXX 100644
+ /*----------------------------------------------------------------------------
---- a/hw/arm/trace-events
+ | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
-+++ b/hw/arm/trace-events
+ | NaN; otherwise returns 0.
@@ -XXX,XX +XXX,XX @@ smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr
  smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
  smmu_iotlb_lookup_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
  smmu_iotlb_lookup_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 -smmu_iotlb_insert(uint16_t asid, uint64_t addr) "IOTLB ++ asid=%d addr=0x%"PRIx64
 +smmu_iotlb_insert(uint16_t asid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d addr=0x%"PRIx64" tg=%d level=%d"
  # smmuv3.c
  smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
 --
-.20.1
+.34.1

-New patch
+[PULL 66/72] softfloat: Use parts_pick_nan in propagateFloatx80NaN
+From: Richard Henderson <richard.henderson@linaro.org>
+Unpacking and repacking the parts may be slightly more work
+than we did before, but we get to reuse more code.  For a
+code path handling exceptional values, this is an improvement.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ fpu/softfloat.c | 43 +++++--------------------------------------
+file changed, 5 insertions(+), 38 deletions(-)
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat.c
++++ b/fpu/softfloat.c
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
+ floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
+ {
+-    bool aIsLargerSignificand;
+-    FloatClass a_cls, b_cls;
++    FloatParts128 pa, pb, *pr;
+-    /* This is not complete, but is good enough for pickNaN.  */
+-    a_cls = (!floatx80_is_any_nan(a)
+-             ? float_class_normal
+-             : floatx80_is_signaling_nan(a, status)
+-             ? float_class_snan
+-             : float_class_qnan);
+-    b_cls = (!floatx80_is_any_nan(b)
+-             ? float_class_normal
+-             : floatx80_is_signaling_nan(b, status)
+-             ? float_class_snan
+-             : float_class_qnan);
+-
+-    if (is_snan(a_cls) || is_snan(b_cls)) {
+-        float_raise(float_flag_invalid, status);
+-    }
+-
+-    if (status->default_nan_mode) {
++    if (!floatx80_unpack_canonical(&pa, a, status) ||
++        !floatx80_unpack_canonical(&pb, b, status)) {
+         return floatx80_default_nan(status);
+     }
+-    if (a.low < b.low) {
+-        aIsLargerSignificand = 0;
+-    } else if (b.low < a.low) {
+-        aIsLargerSignificand = 1;
+-    } else {
+-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
+-    }
+-
+-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
+-        if (is_snan(b_cls)) {
+-            return floatx80_silence_nan(b, status);
+-        }
+-        return b;
+-    } else {
+-        if (is_snan(a_cls)) {
+-            return floatx80_silence_nan(a, status);
+-        }
+-        return a;
+-    }
++    pr = parts_pick_nan(&pa, &pb, status);
++    return floatx80_round_pack_canonical(pr, status);
+ }
+ /*----------------------------------------------------------------------------
+--
+.34.1

-[PULL 24/27] target/arm: Replace A64 get_fpstatus_ptr() with generic fpstatus_ptr()
+[PULL 67/72] softfloat: Inline pickNaN
-We currently have two versions of get_fpstatus_ptr(), which both take
+From: Richard Henderson <richard.henderson@linaro.org>
-an effectively boolean argument:
- * the one for A64 takes "bool is_f16" to distinguish fp16 from other ops
+Inline pickNaN into its only caller.  This makes one assert
- * the one for A32/T32 takes "int neon" to distinguish Neon from other ops
+redundant with the immediately preceding IF.
-This is confusing, and to implement ARMv8.2-FP16 the A32/T32 one will
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-need to make a four-way distinction between "non-Neon, FP16",
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-"non-Neon, single/double", "Neon, FP16" and "Neon, single/double".
+Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
 The A64 version will then be a strict subset of the A32/T32 version.
 To clean this all up, we want to go to a single implementation which
 takes an enum argument with values FPST_FPCR, FPST_STD,
 FPST_FPCR_F16, and FPST_STD_F16.  We rename the function to
 fpstatus_ptr() so that unconverted code gets a compilation error
 rather than silently passing the wrong thing to the new function.
 This commit implements that new API, and converts A64 to use it:
  get_fpstatus_ptr(false) -> fpstatus_ptr(FPST_FPCR)
  get_fpstatus_ptr(true) -> fpstatus_ptr(FPST_FPCR_F16)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
-Message-id: 20200806104453.30393-2-peter.maydell@linaro.org
 ---
- target/arm/translate-a64.h |  1 -
+ fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
- target/arm/translate.h     | 51 ++++++++++++++++++++++
+ fpu/softfloat-specialize.c.inc | 96 ----------------------------------
- target/arm/translate-a64.c | 89 +++++++++++++++-----------------------
+files changed, 73 insertions(+), 105 deletions(-)
- target/arm/translate-sve.c | 34 +++++++--------
-files changed, 103 insertions(+), 72 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 diff --git a/target/arm/translate-a64.h b/target/arm/translate-a64.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate-a64.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/translate-a64.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ TCGv_i64 cpu_reg_sp(DisasContext *s, int reg);
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
- TCGv_i64 read_cpu_reg(DisasContext *s, int reg, int sf);
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- TCGv_i64 read_cpu_reg_sp(DisasContext *s, int reg, int sf);
+                                      float_status *s)
- void write_fp_dreg(DisasContext *s, int reg, TCGv_i64 v);
+ {
--TCGv_ptr get_fpstatus_ptr(bool);
++    int cmp, which;
  bool logic_imm_decode_wmask(uint64_t *result, unsigned int immn,
                              unsigned int imms, unsigned int immr);
  bool sve_access_check(DisasContext *s);
 diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.h
 +++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ typedef void CryptoThreeOpIntFn(TCGv_ptr, TCGv_ptr, TCGv_i32);
  typedef void CryptoThreeOpFn(TCGv_ptr, TCGv_ptr, TCGv_ptr);
  typedef void AtomicThreeOpFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGArg, MemOp);
 +/*
 + * Enum for argument to fpstatus_ptr().
 + */
 +typedef enum ARMFPStatusFlavour {
 +    FPST_FPCR,
 +    FPST_FPCR_F16,
 +    FPST_STD,
 +    FPST_STD_F16,
 +} ARMFPStatusFlavour;
 +
-+/**
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+ * fpstatus_ptr: return TCGv_ptr to the specified fp_status field
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-+ *
+     }
-+ * We have multiple softfloat float_status fields in the Arm CPU state struct
-+ * (see the comment in cpu.h for details). Return a TCGv_ptr which has
+     if (s->default_nan_mode) {
-+ * been set up to point to the requested field in the CPU state struct.
+         parts_default_nan(a, s);
-+ * The options are:
+-    } else {
-+ *
+-        int cmp = frac_cmp(a, b);
-+ * FPST_FPCR
+-        if (cmp == 0) {
-+ *   for non-FP16 operations controlled by the FPCR
+-            cmp = a->sign < b->sign;
-+ * FPST_FPCR_F16
+-        }
-+ *   for operations controlled by the FPCR where FPCR.FZ16 is to be used
++        return a;
-+ * FPST_STD
++    }
-+ *   for A32/T32 Neon operations using the "standard FPSCR value"
-+ * FPST_STD_F16
+-        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
-+ *   as FPST_STD, but where FPCR.FZ16 is to be used
+-            a = b;
-+ */
+-        }
-+static inline TCGv_ptr fpstatus_ptr(ARMFPStatusFlavour flavour)
++    cmp = frac_cmp(a, b);
-+{
++    if (cmp == 0) {
-+    TCGv_ptr statusptr = tcg_temp_new_ptr();
++        cmp = a->sign < b->sign;
-+    int offset;
++    }
 +
-+    switch (flavour) {
++    switch (s->float_2nan_prop_rule) {
-+    case FPST_FPCR:
++    case float_2nan_prop_s_ab:
-+        offset = offsetof(CPUARMState, vfp.fp_status);
+         if (is_snan(a->cls)) {
-+        break;
+-            parts_silence_nan(a, s);
-+    case FPST_FPCR_F16:
++            which = 0;
-+        offset = offsetof(CPUARMState, vfp.fp_status_f16);
++        } else if (is_snan(b->cls)) {
-+        break;
++            which = 1;
-+    case FPST_STD:
++        } else if (is_qnan(a->cls)) {
-+        offset = offsetof(CPUARMState, vfp.standard_fp_status);
++            which = 0;
-+        break;
++        } else {
-+    case FPST_STD_F16:
++            which = 1;
-+        /* Not yet used or implemented: fall through to assert */
+         }
 +        break;
 +    case float_2nan_prop_s_ba:
 +        if (is_snan(b->cls)) {
 +            which = 1;
 +        } else if (is_snan(a->cls)) {
 +            which = 0;
 +        } else if (is_qnan(b->cls)) {
 +            which = 1;
 +        } else {
 +            which = 0;
 +        }
 +        break;
 +    case float_2nan_prop_ab:
 +        which = is_nan(a->cls) ? 0 : 1;
 +        break;
 +    case float_2nan_prop_ba:
 +        which = is_nan(b->cls) ? 1 : 0;
 +        break;
 +    case float_2nan_prop_x87:
 +        /*
 +         * This implements x87 NaN propagation rules:
 +         * SNaN + QNaN => return the QNaN
 +         * two SNaNs => return the one with the larger significand, silenced
 +         * two QNaNs => return the one with the larger significand
 +         * SNaN and a non-NaN => return the SNaN, silenced
 +         * QNaN and a non-NaN => return the QNaN
 +         *
 +         * If we get down to comparing significands and they are the same,
 +         * return the NaN with the positive sign bit (if any).
 +         */
 +        if (is_snan(a->cls)) {
 +            if (is_snan(b->cls)) {
 +                which = cmp > 0 ? 0 : 1;
 +            } else {
 +                which = is_qnan(b->cls) ? 1 : 0;
 +            }
 +        } else if (is_qnan(a->cls)) {
 +            if (is_snan(b->cls) || !is_qnan(b->cls)) {
 +                which = 0;
 +            } else {
 +                which = cmp > 0 ? 0 : 1;
 +            }
 +        } else {
 +            which = 1;
 +        }
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
-+    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
-+    return statusptr;
-+}
 +
- #endif /* TARGET_ARM_TRANSLATE_H */
++    if (which) {
-diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
++        a = b;
 +    }
 +    if (is_snan(a->cls)) {
 +        parts_silence_nan(a, s);
      }
      return a;
  }
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/translate-a64.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/translate-a64.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static void write_fp_sreg(DisasContext *s, int reg, TCGv_i32 v)
+@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
-     tcg_temp_free_i64(tmp);
+     }
  }
--TCGv_ptr get_fpstatus_ptr(bool is_f16)
+-/*----------------------------------------------------------------------------
 -| Select which NaN to propagate for a two-input operation.
 -| IEEE754 doesn't specify all the details of this, so the
 -| algorithm is target-specific.
 -| The routine is passed various bits of information about the
 -| two NaNs and should return 0 to select NaN a and 1 for NaN b.
 -| Note that signalling NaNs are always squashed to quiet NaNs
 -| by the caller, by calling floatXX_silence_nan() before
 -| returning them.
 -|
 -| aIsLargerSignificand is only valid if both a and b are NaNs
 -| of some kind, and is true if a has the larger significand,
 -| or if both a and b have the same significand but a is
 -| positive but b is negative. It is only needed for the x87
 -| tie-break rule.
 -*----------------------------------------------------------------------------*/
 -
 -static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 -                   bool aIsLargerSignificand, float_status *status)
 -{
--    TCGv_ptr statusptr = tcg_temp_new_ptr();
+-    /*
--    int offset;
+-     * We guarantee not to require the target to tell us how to
 -     * pick a NaN if we're always returning the default NaN.
 -     * But if we're not in default-NaN mode then the target must
 -     * specify via set_float_2nan_prop_rule().
 -     */
 -    assert(!status->default_nan_mode);
 -
--    /* In A64 all instructions (both FP and Neon) use the FPCR; there
+-    switch (status->float_2nan_prop_rule) {
--     * is no equivalent of the A32 Neon "standard FPSCR value".
+-    case float_2nan_prop_s_ab:
--     * However half-precision operations operate under a different
+-        if (is_snan(a_cls)) {
--     * FZ16 flag and use vfp.fp_status_f16 instead of vfp.fp_status.
+-            return 0;
--     */
+-        } else if (is_snan(b_cls)) {
--    if (is_f16) {
+-            return 1;
--        offset = offsetof(CPUARMState, vfp.fp_status_f16);
+-        } else if (is_qnan(a_cls)) {
--    } else {
+-            return 0;
--        offset = offsetof(CPUARMState, vfp.fp_status);
+-        } else {
 -            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_s_ba:
 -        if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_ab:
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_ba:
 -        if (is_nan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_x87:
 -        /*
 -         * This implements x87 NaN propagation rules:
 -         * SNaN + QNaN => return the QNaN
 -         * two SNaNs => return the one with the larger significand, silenced
 -         * two QNaNs => return the one with the larger significand
 -         * SNaN and a non-NaN => return the SNaN, silenced
 -         * QNaN and a non-NaN => return the QNaN
 -         *
 -         * If we get down to comparing significands and they are the same,
 -         * return the NaN with the positive sign bit (if any).
 -         */
 -        if (is_snan(a_cls)) {
 -            if (is_snan(b_cls)) {
 -                return aIsLargerSignificand ? 0 : 1;
 -            }
 -            return is_qnan(b_cls) ? 1 : 0;
 -        } else if (is_qnan(a_cls)) {
 -            if (is_snan(b_cls) || !is_qnan(b_cls)) {
 -                return 0;
 -            } else {
 -                return aIsLargerSignificand ? 0 : 1;
 -            }
 -        } else {
 -            return 1;
 -        }
 -    default:
 -        g_assert_not_reached();
 -    }
--    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
--    return statusptr;
 -}
 -
- /* Expand a 2-operand AdvSIMD vector operation using an expander function.  */
+ /*----------------------------------------------------------------------------
- static void gen_gvec_fn2(DisasContext *s, bool is_q, int rd, int rn,
+ | Returns 1 if the double-precision floating-point value `a' is a quiet
-                          GVecGen2Fn *gvec_fn, int vece)
+ | NaN; otherwise returns 0.
@@ -XXX,XX +XXX,XX @@ static void gen_gvec_op3_fpst(DisasContext *s, bool is_q, int rd, int rn,
                                int rm, bool is_fp16, int data,
                                gen_helper_gvec_3_ptr *fn)
  {
 -    TCGv_ptr fpst = get_fpstatus_ptr(is_fp16);
 +    TCGv_ptr fpst = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
      tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, rd),
                         vec_full_reg_offset(s, rn),
                         vec_full_reg_offset(s, rm), fpst,
@@ -XXX,XX +XXX,XX @@ static void handle_fp_compare(DisasContext *s, int size,
                                bool cmp_with_zero, bool signal_all_nans)
  {
      TCGv_i64 tcg_flags = tcg_temp_new_i64();
 -    TCGv_ptr fpst = get_fpstatus_ptr(size == MO_16);
 +    TCGv_ptr fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      if (size == MO_64) {
          TCGv_i64 tcg_vn, tcg_vm;
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
          tcg_gen_xori_i32(tcg_res, tcg_op, 0x8000);
          break;
      case 0x3: /* FSQRT */
 -        fpst = get_fpstatus_ptr(true);
 +        fpst = fpstatus_ptr(FPST_FPCR_F16);
          gen_helper_sqrt_f16(tcg_res, tcg_op, fpst);
          break;
      case 0x8: /* FRINTN */
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
      case 0xc: /* FRINTA */
      {
          TCGv_i32 tcg_rmode = tcg_const_i32(arm_rmode_to_sf(opcode & 7));
 -        fpst = get_fpstatus_ptr(true);
 +        fpst = fpstatus_ptr(FPST_FPCR_F16);
          gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
          gen_helper_advsimd_rinth(tcg_res, tcg_op, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
          break;
      }
      case 0xe: /* FRINTX */
 -        fpst = get_fpstatus_ptr(true);
 +        fpst = fpstatus_ptr(FPST_FPCR_F16);
          gen_helper_advsimd_rinth_exact(tcg_res, tcg_op, fpst);
          break;
      case 0xf: /* FRINTI */
 -        fpst = get_fpstatus_ptr(true);
 +        fpst = fpstatus_ptr(FPST_FPCR_F16);
          gen_helper_advsimd_rinth(tcg_res, tcg_op, fpst);
          break;
      default:
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_single(DisasContext *s, int opcode, int rd, int rn)
          g_assert_not_reached();
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      if (rmode >= 0) {
          TCGv_i32 tcg_rmode = tcg_const_i32(rmode);
          gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_double(DisasContext *s, int opcode, int rd, int rn)
          g_assert_not_reached();
      }
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      if (rmode >= 0) {
          TCGv_i32 tcg_rmode = tcg_const_i32(rmode);
          gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
              /* Single to half */
              TCGv_i32 tcg_rd = tcg_temp_new_i32();
              TCGv_i32 ahp = get_ahp_flag();
 -            TCGv_ptr fpst = get_fpstatus_ptr(false);
 +            TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
              gen_helper_vfp_fcvt_f32_to_f16(tcg_rd, tcg_rn, fpst, ahp);
              /* write_fp_sreg is OK here because top half of tcg_rd is zero */
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
              /* Double to single */
              gen_helper_vfp_fcvtsd(tcg_rd, tcg_rn, cpu_env);
          } else {
 -            TCGv_ptr fpst = get_fpstatus_ptr(false);
 +            TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
              TCGv_i32 ahp = get_ahp_flag();
              /* Double to half */
              gen_helper_vfp_fcvt_f64_to_f16(tcg_rd, tcg_rn, fpst, ahp);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
      case 0x3:
      {
          TCGv_i32 tcg_rn = read_fp_sreg(s, rn);
 -        TCGv_ptr tcg_fpst = get_fpstatus_ptr(false);
 +        TCGv_ptr tcg_fpst = fpstatus_ptr(FPST_FPCR);
          TCGv_i32 tcg_ahp = get_ahp_flag();
          tcg_gen_ext16u_i32(tcg_rn, tcg_rn);
          if (dtype == 0) {
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_single(DisasContext *s, int opcode,
      TCGv_ptr fpst;
      tcg_res = tcg_temp_new_i32();
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_op1 = read_fp_sreg(s, rn);
      tcg_op2 = read_fp_sreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_double(DisasContext *s, int opcode,
      TCGv_ptr fpst;
      tcg_res = tcg_temp_new_i64();
 -    fpst = get_fpstatus_ptr(false);
 +    fpst = fpstatus_ptr(FPST_FPCR);
      tcg_op1 = read_fp_dreg(s, rn);
      tcg_op2 = read_fp_dreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_half(DisasContext *s, int opcode,
      TCGv_ptr fpst;
      tcg_res = tcg_temp_new_i32();
 -    fpst = get_fpstatus_ptr(true);
 +    fpst = fpstatus_ptr(FPST_FPCR_F16);
      tcg_op1 = read_fp_hreg(s, rn);
      tcg_op2 = read_fp_hreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_single(DisasContext *s, bool o0, bool o1,
  {
      TCGv_i32 tcg_op1, tcg_op2, tcg_op3;
      TCGv_i32 tcg_res = tcg_temp_new_i32();
 -    TCGv_ptr fpst = get_fpstatus_ptr(false);
 +    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
      tcg_op1 = read_fp_sreg(s, rn);
      tcg_op2 = read_fp_sreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_double(DisasContext *s, bool o0, bool o1,
  {
      TCGv_i64 tcg_op1, tcg_op2, tcg_op3;
      TCGv_i64 tcg_res = tcg_temp_new_i64();
 -    TCGv_ptr fpst = get_fpstatus_ptr(false);
 +    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
      tcg_op1 = read_fp_dreg(s, rn);
      tcg_op2 = read_fp_dreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_half(DisasContext *s, bool o0, bool o1,
  {
      TCGv_i32 tcg_op1, tcg_op2, tcg_op3;
      TCGv_i32 tcg_res = tcg_temp_new_i32();
 -    TCGv_ptr fpst = get_fpstatus_ptr(true);
 +    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR_F16);
      tcg_op1 = read_fp_hreg(s, rn);
      tcg_op2 = read_fp_hreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fpfpcvt(DisasContext *s, int rd, int rn, int opcode,
      TCGv_i32 tcg_shift, tcg_single;
      TCGv_i64 tcg_double;
 -    tcg_fpstatus = get_fpstatus_ptr(type == 3);
 +    tcg_fpstatus = fpstatus_ptr(type == 3 ? FPST_FPCR_F16 : FPST_FPCR);
      tcg_shift = tcg_const_i32(64 - scale);
@@ -XXX,XX +XXX,XX @@ static void handle_fmov(DisasContext *s, int rd, int rn, int type, bool itof)
  static void handle_fjcvtzs(DisasContext *s, int rd, int rn)
  {
      TCGv_i64 t = read_fp_dreg(s, rn);
 -    TCGv_ptr fpstatus = get_fpstatus_ptr(false);
 +    TCGv_ptr fpstatus = fpstatus_ptr(FPST_FPCR);
      gen_helper_fjcvtzs(t, t, fpstatus);
@@ -XXX,XX +XXX,XX @@ static void disas_simd_across_lanes(DisasContext *s, uint32_t insn)
           * Note that correct NaN propagation requires that we do these
           * operations in exactly the order specified by the pseudocode.
           */
 -        TCGv_ptr fpst = get_fpstatus_ptr(size == MO_16);
 +        TCGv_ptr fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          int fpopcode = opcode | is_min << 4 | is_u << 5;
          int vmap = (1 << elements) - 1;
          TCGv_i32 tcg_res32 = do_reduction_op(s, fpopcode, rn, esize,
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_pairwise(DisasContext *s, uint32_t insn)
              return;
          }
 -        fpst = get_fpstatus_ptr(size == MO_16);
 +        fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          break;
      default:
          unallocated_encoding(s);
@@ -XXX,XX +XXX,XX @@ static void handle_simd_intfp_conv(DisasContext *s, int rd, int rn,
                                     int elements, int is_signed,
                                     int fracbits, int size)
  {
 -    TCGv_ptr tcg_fpst = get_fpstatus_ptr(size == MO_16);
 +    TCGv_ptr tcg_fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      TCGv_i32 tcg_shift = NULL;
      MemOp mop = size | (is_signed ? MO_SIGN : 0);
@@ -XXX,XX +XXX,XX @@ static void handle_simd_shift_fpint_conv(DisasContext *s, bool is_scalar,
      assert(!(is_scalar && is_q));
      tcg_rmode = tcg_const_i32(arm_rmode_to_sf(FPROUNDING_ZERO));
 -    tcg_fpstatus = get_fpstatus_ptr(size == MO_16);
 +    tcg_fpstatus = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      gen_helper_set_rmode(tcg_rmode, tcg_rmode, tcg_fpstatus);
      fracbits = (16 << size) - immhb;
      tcg_shift = tcg_const_i32(fracbits);
@@ -XXX,XX +XXX,XX @@ static void handle_3same_float(DisasContext *s, int size, int elements,
                                 int fpopcode, int rd, int rn, int rm)
  {
      int pass;
 -    TCGv_ptr fpst = get_fpstatus_ptr(false);
 +    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
      for (pass = 0; pass < elements; pass++) {
          if (size) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_three_reg_same_fp16(DisasContext *s,
          return;
      }
 -    fpst = get_fpstatus_ptr(true);
 +    fpst = fpstatus_ptr(FPST_FPCR_F16);
      tcg_op1 = read_fp_hreg(s, rn);
      tcg_op2 = read_fp_hreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_fcmp_zero(DisasContext *s, int opcode,
          return;
      }
 -    fpst = get_fpstatus_ptr(size == MO_16);
 +    fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      if (is_double) {
          TCGv_i64 tcg_op = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_reciprocal(DisasContext *s, int opcode,
                                      int size, int rn, int rd)
  {
      bool is_double = (size == 3);
 -    TCGv_ptr fpst = get_fpstatus_ptr(false);
 +    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
      if (is_double) {
          TCGv_i64 tcg_op = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_narrow(DisasContext *s, bool scalar,
              } else {
                  TCGv_i32 tcg_lo = tcg_temp_new_i32();
                  TCGv_i32 tcg_hi = tcg_temp_new_i32();
 -                TCGv_ptr fpst = get_fpstatus_ptr(false);
 +                TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
                  TCGv_i32 ahp = get_ahp_flag();
                  tcg_gen_extr_i64_i32(tcg_lo, tcg_hi, tcg_op);
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_two_reg_misc(DisasContext *s, uint32_t insn)
      if (is_fcvt) {
          tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
 -        tcg_fpstatus = get_fpstatus_ptr(false);
 +        tcg_fpstatus = fpstatus_ptr(FPST_FPCR);
          gen_helper_set_rmode(tcg_rmode, tcg_rmode, tcg_fpstatus);
      } else {
          tcg_rmode = NULL;
@@ -XXX,XX +XXX,XX @@ static void handle_simd_3same_pair(DisasContext *s, int is_q, int u, int opcode,
      /* Floating point operations need fpst */
      if (opcode >= 0x58) {
 -        fpst = get_fpstatus_ptr(false);
 +        fpst = fpstatus_ptr(FPST_FPCR);
      } else {
          fpst = NULL;
      }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_same_fp16(DisasContext *s, uint32_t insn)
          break;
      }
 -    fpst = get_fpstatus_ptr(true);
 +    fpst = fpstatus_ptr(FPST_FPCR_F16);
      if (pairwise) {
          int maxpass = is_q ? 8 : 4;
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_widening(DisasContext *s, int opcode, bool is_q,
          /* 16 -> 32 bit fp conversion */
          int srcelt = is_q ? 4 : 0;
          TCGv_i32 tcg_res[4];
 -        TCGv_ptr fpst = get_fpstatus_ptr(false);
 +        TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
          TCGv_i32 ahp = get_ahp_flag();
          for (pass = 0; pass < 4; pass++) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc(DisasContext *s, uint32_t insn)
      }
      if (need_fpstatus || need_rmode) {
 -        tcg_fpstatus = get_fpstatus_ptr(false);
 +        tcg_fpstatus = fpstatus_ptr(FPST_FPCR);
      } else {
          tcg_fpstatus = NULL;
      }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc_fp16(DisasContext *s, uint32_t insn)
      }
      if (need_rmode || need_fpst) {
 -        tcg_fpstatus = get_fpstatus_ptr(true);
 +        tcg_fpstatus = fpstatus_ptr(FPST_FPCR_F16);
      }
      if (need_rmode) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
      }
      if (is_fp) {
 -        fpst = get_fpstatus_ptr(is_fp16);
 +        fpst = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
      } else {
          fpst = NULL;
      }
 diff --git a/target/arm/translate-sve.c b/target/arm/translate-sve.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-sve.c
 +++ b/target/arm/translate-sve.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FMLA_zzxz(DisasContext *s, arg_FMLA_zzxz *a)
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FMUL_zzx(DisasContext *s, arg_FMUL_zzx *a)
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static void do_reduce(DisasContext *s, arg_rpr_esz *a,
      tcg_gen_addi_ptr(t_zn, cpu_env, vec_full_reg_offset(s, a->rn));
      tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, a->pg));
 -    status = get_fpstatus_ptr(a->esz == MO_16);
 +    status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      fn(temp, t_zn, t_pg, status, t_desc);
      tcg_temp_free_ptr(t_zn);
@@ -XXX,XX +XXX,XX @@ DO_VPZ(FMAXV, fmaxv)
  static void do_zz_fp(DisasContext *s, arg_rr_esz *a, gen_helper_gvec_2_ptr *fn)
  {
      unsigned vsz = vec_full_reg_size(s);
 -    TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +    TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      tcg_gen_gvec_2_ptr(vec_full_reg_offset(s, a->rd),
                         vec_full_reg_offset(s, a->rn),
@@ -XXX,XX +XXX,XX @@ static void do_ppz_fp(DisasContext *s, arg_rpr_esz *a,
                        gen_helper_gvec_3_ptr *fn)
  {
      unsigned vsz = vec_full_reg_size(s);
 -    TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +    TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      tcg_gen_gvec_3_ptr(pred_full_reg_offset(s, a->rd),
                         vec_full_reg_offset(s, a->rn),
@@ -XXX,XX +XXX,XX @@ static bool trans_FTMAD(DisasContext *s, arg_FTMAD *a)
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FADDA(DisasContext *s, arg_rprr_esz *a)
      t_pg = tcg_temp_new_ptr();
      tcg_gen_addi_ptr(t_rm, cpu_env, vec_full_reg_offset(s, a->rm));
      tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, a->pg));
 -    t_fpst = get_fpstatus_ptr(a->esz == MO_16);
 +    t_fpst = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
      t_desc = tcg_const_i32(simd_desc(vsz, vsz, 0));
      fns[a->esz - 1](t_val, t_val, t_rm, t_pg, t_fpst, t_desc);
@@ -XXX,XX +XXX,XX @@ static bool do_zzz_fp(DisasContext *s, arg_rrr_esz *a,
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_zpzz_fp(DisasContext *s, arg_rprr_esz *a,
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static void do_fp_scalar(DisasContext *s, int zd, int zn, int pg, bool is_fp16,
      tcg_gen_addi_ptr(t_zn, cpu_env, vec_full_reg_offset(s, zn));
      tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, pg));
 -    status = get_fpstatus_ptr(is_fp16);
 +    status = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
      desc = tcg_const_i32(simd_desc(vsz, vsz, 0));
      fn(t_zd, t_zn, t_pg, scalar, status, desc);
@@ -XXX,XX +XXX,XX @@ static bool do_fp_cmp(DisasContext *s, arg_rprr_esz *a,
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_4_ptr(pred_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCADD(DisasContext *s, arg_FCADD *a)
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_fmla(DisasContext *s, arg_rprrr_esz *a,
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_5_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCMLA_zpzzz(DisasContext *s, arg_FCMLA_zpzzz *a)
      }
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_5_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCMLA_zzxz(DisasContext *s, arg_FCMLA_zzxz *a)
      tcg_debug_assert(a->rd == a->ra);
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                             vec_full_reg_offset(s, a->rn),
                             vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_zpz_ptr(DisasContext *s, int rd, int rn, int pg,
  {
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
 -        TCGv_ptr status = get_fpstatus_ptr(is_fp16);
 +        TCGv_ptr status = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
          tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, rd),
                             vec_full_reg_offset(s, rn),
                             pred_full_reg_offset(s, pg),
@@ -XXX,XX +XXX,XX @@ static bool do_frint_mode(DisasContext *s, arg_rpr_esz *a, int mode)
      if (sve_access_check(s)) {
          unsigned vsz = vec_full_reg_size(s);
          TCGv_i32 tmode = tcg_const_i32(mode);
 -        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
 +        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
          gen_helper_set_rmode(tmode, tmode, status);
 --
-.20.1
+.34.1

-New patch
+[PULL 68/72] softfloat: Share code between parts_pick_nan cases
+From: Richard Henderson <richard.henderson@linaro.org>
+Remember if there was an SNaN, and use that to simplify
+float_2nan_prop_s_{ab,ba} to only the snan component.
+Then, fall through to the corresponding
+float_2nan_prop_{ab,ba} case to handle any remaining
+nans, which must be quiet.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
+file changed, 12 insertions(+), 20 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
+                                      float_status *s)
+ {
++    bool have_snan = false;
+     int cmp, which;
+     if (is_snan(a->cls) || is_snan(b->cls)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
++        have_snan = true;
+     }
+     if (s->default_nan_mode) {
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
+     switch (s->float_2nan_prop_rule) {
+     case float_2nan_prop_s_ab:
+-        if (is_snan(a->cls)) {
+-            which = 0;
+-        } else if (is_snan(b->cls)) {
+-            which = 1;
+-        } else if (is_qnan(a->cls)) {
+-            which = 0;
+-        } else {
+-            which = 1;
++        if (have_snan) {
++            which = is_snan(a->cls) ? 0 : 1;
++            break;
+         }
+-        break;
+-    case float_2nan_prop_s_ba:
+-        if (is_snan(b->cls)) {
+-            which = 1;
+-        } else if (is_snan(a->cls)) {
+-            which = 0;
+-        } else if (is_qnan(b->cls)) {
+-            which = 1;
+-        } else {
+-            which = 0;
+-        }
+-        break;
++        /* fall through */
+     case float_2nan_prop_ab:
+         which = is_nan(a->cls) ? 0 : 1;
+         break;
++    case float_2nan_prop_s_ba:
++        if (have_snan) {
++            which = is_snan(b->cls) ? 1 : 0;
++            break;
++        }
++        /* fall through */
+     case float_2nan_prop_ba:
+         which = is_nan(b->cls) ? 1 : 0;
+         break;
+--
+.34.1

-[PULL 02/27] hw/arm/smmu-common: Factorize some code in smmu_ptw_64()
+[PULL 69/72] softfloat: Sink frac_cmp in parts_pick_nan until needed
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Page and block PTE decoding can share some code. Let's
+Move the fractional comparison to the end of the
-first handle table PTE and factorize some code shared by
+float_2nan_prop_x87 case.  This is not required for
-page and block PTEs.
+any other 2nan propagation rule.  Reorganize the
 x87 case itself to break out of the switch when the
 fractional comparison is not required.
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20200728150815.11446-2-eric.auger@redhat.com
+Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmu-common.c | 48 ++++++++++++++++----------------------------
+ fpu/softfloat-parts.c.inc | 19 +++++++++----------
-file changed, 17 insertions(+), 31 deletions(-)
+file changed, 9 insertions(+), 10 deletions(-)
-diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-common.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmu-common.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-         uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
+         return a;
-         uint64_t mask = subpage_size - 1;
+     }
-         uint32_t offset = iova_level_offset(iova, inputsize, level, granule_sz);
--        uint64_t pte;
+-    cmp = frac_cmp(a, b);
-+        uint64_t pte, gpa;
+-    if (cmp == 0) {
-         dma_addr_t pte_addr = baseaddr + offset * sizeof(pte);
+-        cmp = a->sign < b->sign;
-         uint8_t ap;
+-    }
+-
-@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
+     switch (s->float_2nan_prop_rule) {
-         if (is_invalid_pte(pte) || is_reserved_pte(pte, level)) {
+     case float_2nan_prop_s_ab:
-             trace_smmu_ptw_invalid_pte(stage, level, baseaddr,
+         if (have_snan) {
-                                        pte_addr, offset, pte);
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
--            info->type = SMMU_PTW_ERR_TRANSLATION;
+          * return the NaN with the positive sign bit (if any).
--            goto error;
+          */
          if (is_snan(a->cls)) {
 -            if (is_snan(b->cls)) {
 -                which = cmp > 0 ? 0 : 1;
 -            } else {
 +            if (!is_snan(b->cls)) {
                  which = is_qnan(b->cls) ? 1 : 0;
 +                break;
              }
          } else if (is_qnan(a->cls)) {
              if (is_snan(b->cls) || !is_qnan(b->cls)) {
                  which = 0;
 -            } else {
 -                which = cmp > 0 ? 0 : 1;
 +                break;
              }
          } else {
              which = 1;
 +            break;
          }
++        cmp = frac_cmp(a, b);
--        if (is_page_pte(pte, level)) {
++        if (cmp == 0) {
--            uint64_t gpa = get_page_pte_address(pte, granule_sz);
++            cmp = a->sign < b->sign;
-+        if (is_table_pte(pte, level)) {
++        }
-+            ap = PTE_APTABLE(pte);
++        which = cmp > 0 ? 0 : 1;
+         break;
--            ap = PTE_AP(pte);
+     default:
-             if (is_permission_fault(ap, perm)) {
+         g_assert_not_reached();
                  info->type = SMMU_PTW_ERR_PERMISSION;
                  goto error;
              }
 -
 -            tlbe->translated_addr = gpa + (iova & mask);
 -            tlbe->perm = PTE_AP_TO_PERM(ap);
 +            baseaddr = get_table_pte_address(pte, granule_sz);
 +            level++;
 +            continue;
 +        } else if (is_page_pte(pte, level)) {
 +            gpa = get_page_pte_address(pte, granule_sz);
              trace_smmu_ptw_page_pte(stage, level, iova,
                                      baseaddr, pte_addr, pte, gpa);
 -            return 0;
 -        }
 -        if (is_block_pte(pte, level)) {
 +        } else {
              uint64_t block_size;
 -            hwaddr gpa = get_block_pte_address(pte, level, granule_sz,
 -                                               &block_size);
 -
 -            ap = PTE_AP(pte);
 -            if (is_permission_fault(ap, perm)) {
 -                info->type = SMMU_PTW_ERR_PERMISSION;
 -                goto error;
 -            }
 +            gpa = get_block_pte_address(pte, level, granule_sz,
 +                                        &block_size);
              trace_smmu_ptw_block_pte(stage, level, baseaddr,
                                       pte_addr, pte, iova, gpa,
                                       block_size >> 20);
 -
 -            tlbe->translated_addr = gpa + (iova & mask);
 -            tlbe->perm = PTE_AP_TO_PERM(ap);
 -            return 0;
          }
 -
 -        /* table pte */
 -        ap = PTE_APTABLE(pte);
 -
 +        ap = PTE_AP(pte);
          if (is_permission_fault(ap, perm)) {
              info->type = SMMU_PTW_ERR_PERMISSION;
              goto error;
          }
 -        baseaddr = get_table_pte_address(pte, granule_sz);
 -        level++;
 -    }
 +        tlbe->translated_addr = gpa + (iova & mask);
 +        tlbe->perm = PTE_AP_TO_PERM(ap);
 +        return 0;
 +    }
      info->type = SMMU_PTW_ERR_TRANSLATION;
  error:
 --
-.20.1
+.34.1

-[PULL 04/27] hw/arm/smmu: Introduce smmu_get_iotlb_key()
+[PULL 70/72] softfloat: Replace WHICH with RET in parts_pick_nan
-From: Eric Auger <eric.auger@redhat.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-Introduce the smmu_get_iotlb_key() helper and the
+Replace the "index" selecting between A and B with a result variable
-SMMU_IOTLB_ASID() macro. Also move smmu_get_iotlb_key and
+of the proper type.  This improves clarity within the function.
 smmu_iotlb_key_hash in the IOTLB related code section.
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20200728150815.11446-4-eric.auger@redhat.com
+Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmu-internal.h       |  1 +
+ fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
- include/hw/arm/smmu-common.h |  1 +
+file changed, 13 insertions(+), 15 deletions(-)
  hw/arm/smmu-common.c         | 66 ++++++++++++++++++++----------------
 files changed, 38 insertions(+), 30 deletions(-)
-diff --git a/hw/arm/smmu-internal.h b/hw/arm/smmu-internal.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmu-internal.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/smmu-internal.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ uint64_t iova_level_offset(uint64_t iova, int inputsize,
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-             MAKE_64BIT_MASK(0, gsz - 3);
+                                      float_status *s)
  {
      bool have_snan = false;
 -    int cmp, which;
 +    FloatPartsN *ret;
 +    int cmp;
      if (is_snan(a->cls) || is_snan(b->cls)) {
          float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
      switch (s->float_2nan_prop_rule) {
      case float_2nan_prop_s_ab:
          if (have_snan) {
 -            which = is_snan(a->cls) ? 0 : 1;
 +            ret = is_snan(a->cls) ? a : b;
              break;
          }
          /* fall through */
      case float_2nan_prop_ab:
 -        which = is_nan(a->cls) ? 0 : 1;
 +        ret = is_nan(a->cls) ? a : b;
          break;
      case float_2nan_prop_s_ba:
          if (have_snan) {
 -            which = is_snan(b->cls) ? 1 : 0;
 +            ret = is_snan(b->cls) ? b : a;
              break;
          }
          /* fall through */
      case float_2nan_prop_ba:
 -        which = is_nan(b->cls) ? 1 : 0;
 +        ret = is_nan(b->cls) ? b : a;
          break;
      case float_2nan_prop_x87:
          /*
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
           */
          if (is_snan(a->cls)) {
              if (!is_snan(b->cls)) {
 -                which = is_qnan(b->cls) ? 1 : 0;
 +                ret = is_qnan(b->cls) ? b : a;
                  break;
              }
          } else if (is_qnan(a->cls)) {
              if (is_snan(b->cls) || !is_qnan(b->cls)) {
 -                which = 0;
 +                ret = a;
                  break;
              }
          } else {
 -            which = 1;
 +            ret = b;
              break;
          }
          cmp = frac_cmp(a, b);
          if (cmp == 0) {
              cmp = a->sign < b->sign;
          }
 -        which = cmp > 0 ? 0 : 1;
 +        ret = cmp > 0 ? a : b;
          break;
      default:
          g_assert_not_reached();
      }
 -    if (which) {
 -        a = b;
 +    if (is_snan(ret->cls)) {
 +        parts_silence_nan(ret, s);
      }
 -    if (is_snan(a->cls)) {
 -        parts_silence_nan(a, s);
 -    }
 -    return a;
 +    return ret;
  }
-+#define SMMU_IOTLB_ASID(key) ((key).asid)
+ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
  #endif
 diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/hw/arm/smmu-common.h
 +++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
  IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
  void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry);
 +SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova);
  void smmu_iotlb_inv_all(SMMUState *s);
  void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
  void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova);
 diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/smmu-common.c
 +++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@
  /* IOTLB Management */
 +static guint smmu_iotlb_key_hash(gconstpointer v)
 +{
 +    SMMUIOTLBKey *key = (SMMUIOTLBKey *)v;
 +    uint32_t a, b, c;
 +
 +    /* Jenkins hash */
 +    a = b = c = JHASH_INITVAL + sizeof(*key);
 +    a += key->asid;
 +    b += extract64(key->iova, 0, 32);
 +    c += extract64(key->iova, 32, 32);
 +
 +    __jhash_mix(a, b, c);
 +    __jhash_final(a, b, c);
 +
 +    return c;
 +}
 +
 +static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
 +{
 +    const SMMUIOTLBKey *k1 = v1;
 +    const SMMUIOTLBKey *k2 = v2;
 +
 +    return (k1->asid == k2->asid) && (k1->iova == k2->iova);
 +}
 +
 +SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova)
 +{
 +    SMMUIOTLBKey key = {.asid = asid, .iova = iova};
 +
 +    return key;
 +}
 +
  IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
                                   hwaddr iova)
  {
 -    SMMUIOTLBKey key = {.asid = cfg->asid, .iova = iova};
 +    SMMUIOTLBKey key = smmu_get_iotlb_key(cfg->asid, iova);
      IOMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
      if (entry) {
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
          smmu_iotlb_inv_all(bs);
      }
 -    key->asid = cfg->asid;
 -    key->iova = entry->iova;
 +    *key = smmu_get_iotlb_key(cfg->asid, entry->iova);
      trace_smmu_iotlb_insert(cfg->asid, entry->iova);
      g_hash_table_insert(bs->iotlb, key, entry);
  }
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
      uint16_t asid = *(uint16_t *)user_data;
      SMMUIOTLBKey *iotlb_key = (SMMUIOTLBKey *)key;
 -    return iotlb_key->asid == asid;
 +    return SMMU_IOTLB_ASID(*iotlb_key) == asid;
  }
  inline void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova)
  {
 -    SMMUIOTLBKey key = {.asid = asid, .iova = iova};
 +    SMMUIOTLBKey key = smmu_get_iotlb_key(asid, iova);
      trace_smmu_iotlb_inv_iova(asid, iova);
      g_hash_table_remove(s->iotlb, &key);
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid)
      return NULL;
  }
 -static guint smmu_iotlb_key_hash(gconstpointer v)
 -{
 -    SMMUIOTLBKey *key = (SMMUIOTLBKey *)v;
 -    uint32_t a, b, c;
 -
 -    /* Jenkins hash */
 -    a = b = c = JHASH_INITVAL + sizeof(*key);
 -    a += key->asid;
 -    b += extract64(key->iova, 0, 32);
 -    c += extract64(key->iova, 32, 32);
 -
 -    __jhash_mix(a, b, c);
 -    __jhash_final(a, b, c);
 -
 -    return c;
 -}
 -
 -static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
 -{
 -    const SMMUIOTLBKey *k1 = v1;
 -    const SMMUIOTLBKey *k2 = v2;
 -
 -    return (k1->asid == k2->asid) && (k1->iova == k2->iova);
 -}
 -
  /* Unmap the whole notifier's range */
  static void smmu_unmap_notifier_range(IOMMUNotifier *n)
  {
 --
-.20.1
+.34.1

-[PULL 10/27] hw/arm/smmuv3: Let AIDR advertise SMMUv3.0 support
+[PULL 71/72] MAINTAINERS: update email address for Leif Lindholm
-From: Eric Auger <eric.auger@redhat.com>
+From: Leif Lindholm <quic_llindhol@quicinc.com>
-Add the support for AIDR register. It currently advertises
+I'm migrating to Qualcomm's new open source email infrastructure, so
-SMMU V3.0 spec.
+update my email address, and update the mailmap to match.
-Signed-off-by: Eric Auger <eric.auger@redhat.com>
+Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
-Message-id: 20200728150815.11446-10-eric.auger@redhat.com
+Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/smmuv3-internal.h | 1 +
+ MAINTAINERS | 2 +-
- include/hw/arm/smmuv3.h  | 1 +
+ .mailmap    | 5 +++--
- hw/arm/smmuv3.c          | 3 +++
+files changed, 4 insertions(+), 3 deletions(-)
 files changed, 5 insertions(+)
-diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
+diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/smmuv3-internal.h
+--- a/MAINTAINERS
-+++ b/hw/arm/smmuv3-internal.h
++++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ REG32(IDR5,                0x14)
+@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
- #define SMMU_IDR5_OAS 4
+ SBSA-REF
+ M: Radoslaw Biernacki <rad@semihalf.com>
- REG32(IIDR,                0x18)
+ M: Peter Maydell <peter.maydell@linaro.org>
-+REG32(AIDR,                0x1c)
+-R: Leif Lindholm <quic_llindhol@quicinc.com>
- REG32(CR0,                 0x20)
++R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
-     FIELD(CR0, SMMU_ENABLE,   0, 1)
+ R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
-     FIELD(CR0, EVENTQEN,      2, 1)
+ L: qemu-arm@nongnu.org
-diff --git a/include/hw/arm/smmuv3.h b/include/hw/arm/smmuv3.h
+ S: Maintained
 diff --git a/.mailmap b/.mailmap
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/arm/smmuv3.h
+--- a/.mailmap
-+++ b/include/hw/arm/smmuv3.h
++++ b/.mailmap
-@@ -XXX,XX +XXX,XX @@ typedef struct SMMUv3State {
+@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
+ Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
-     uint32_t idr[6];
+ James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
-     uint32_t iidr;
+ Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
-+    uint32_t aidr;
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-     uint32_t cr[3];
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
-     uint32_t cr0ack;
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
-     uint32_t statusr;
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
-diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
-index XXXXXXX..XXXXXXX 100644
+ Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
---- a/hw/arm/smmuv3.c
+ Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
-+++ b/hw/arm/smmuv3.c
+ Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
@@ -XXX,XX +XXX,XX @@ static MemTxResult smmu_readl(SMMUv3State *s, hwaddr offset,
      case A_IIDR:
          *data = s->iidr;
          return MEMTX_OK;
 +    case A_AIDR:
 +        *data = s->aidr;
 +        return MEMTX_OK;
      case A_CR0:
          *data = s->cr[0];
          return MEMTX_OK;
 --
-.20.1
+.34.1

-[PULL 13/27] docs/system/arm: Document the Xilinx Versal Virt board
+[PULL 72/72] MAINTAINERS: Add correct email address for Vikram Garhwal
-From: "Edgar E. Iglesias" <edgar.iglesias@xilinx.com>
+From: Vikram Garhwal <vikram.garhwal@bytedance.com>
-Document the Xilinx Versal Virt board.
+Previously, maintainer role was paused due to inactive email id. Commit id:
 c009d715721861984c4987bcc78b7ee183e86d75.
-Signed-off-by: Edgar E. Iglesias <edgar.iglesias@xilinx.com>
+Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
-Message-id: 20200803164749.301971-2-edgar.iglesias@gmail.com
+Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- docs/system/arm/xlnx-versal-virt.rst | 176 +++++++++++++++++++++++++++
+ MAINTAINERS | 2 ++
- docs/system/target-arm.rst           |   1 +
+file changed, 2 insertions(+)
  MAINTAINERS                          |   3 +-
 files changed, 179 insertions(+), 1 deletion(-)
  create mode 100644 docs/system/arm/xlnx-versal-virt.rst
-diff --git a/docs/system/arm/xlnx-versal-virt.rst b/docs/system/arm/xlnx-versal-virt.rst
-new file mode 100644
-index XXXXXXX..XXXXXXX
---- /dev/null
-+++ b/docs/system/arm/xlnx-versal-virt.rst
-@@ -XXX,XX +XXX,XX @@
-+Xilinx Versal Virt (``xlnx-versal-virt``)
-+=========================================
-+
-+Xilinx Versal is a family of heterogeneous multi-core SoCs
-+(System on Chip) that combine traditional hardened CPUs and I/O
-+peripherals in a Processing System (PS) with runtime programmable
-+FPGA logic (PL) and an Artificial Intelligence Engine (AIE).
-+
-+More details here:
-+https://www.xilinx.com/products/silicon-devices/acap/versal.html
-+
-+The family of Versal SoCs share a single architecture but come in
-+different parts with different speed grades, amounts of PL and
-+other differences.
-+
-+The Xilinx Versal Virt board in QEMU is a model of a virtual board
-+(does not exist in reality) with a virtual Versal SoC without I/O
-+limitations. Currently, we support the following cores and devices:
-+
-+Implemented CPU cores:
-+
-+- 2 ACPUs (ARM Cortex-A72)
-+
-+Implemented devices:
-+
-+- Interrupt controller (ARM GICv3)
-+- 2 UARTs (ARM PL011)
-+- An RTC (Versal built-in)
-+- 2 GEMs (Cadence MACB Ethernet MACs)
-+- 8 ADMA (Xilinx zDMA) channels
-+- 2 SD Controllers
-+- OCM (256KB of On Chip Memory)
-+- DDR memory
-+
-+QEMU does not yet model any other devices, including the PL and the AI Engine.
-+
-+Other differences between the hardware and the QEMU model:
-+
-+- QEMU allows the amount of DDR memory provided to be specified with the
-+  ``-m`` argument. If a DTB is provided on the command line then QEMU will
-+  edit it to include suitable entries describing the Versal DDR memory ranges.
-+
-+- QEMU provides 8 virtio-mmio virtio transports; these start at
-+  address ``0xa0000000`` and have IRQs from 111 and upwards.
-+
-+Running
-+"""""""
-+If the user provides an Operating System to be loaded, we expect users
-+to use the ``-kernel`` command line option.
-+
-+Users can load firmware or boot-loaders with the ``-device loader`` options.
-+
-+When loading an OS, QEMU generates a DTB and selects an appropriate address
-+where it gets loaded. This DTB will be passed to the kernel in register x0.
-+
-+If there's no ``-kernel`` option, we generate a DTB and place it at 0x1000
-+for boot-loaders or firmware to pick it up.
-+
-+If users want to provide their own DTB, they can use the ``-dtb`` option.
-+These DTBs will have their memory nodes modified to match QEMU's
-+selected ram_size option before they get passed to the kernel or FW.
-+
-+When loading an OS, we turn on QEMU's PSCI implementation with SMC
-+as the PSCI conduit. When there's no ``-kernel`` option, we assume the user
-+provides EL3 firmware to handle PSCI.
-+
-+A few examples:
-+
-+Direct Linux boot of a generic ARM64 upstream Linux kernel:
-+
-+.. code-block:: bash
-+
-+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 2G \
-+      -serial mon:stdio -display none \
-+      -kernel arch/arm64/boot/Image \
-+      -nic user -nic user \
-+      -device virtio-rng-device,bus=virtio-mmio-bus.0 \
-+      -drive if=none,index=0,file=hd0.qcow2,id=hd0,snapshot \
-+      -drive file=qemu_sd.qcow2,if=sd,index=0,snapshot \
-+      -device virtio-blk-device,drive=hd0 -append root=/dev/vda
-+
-+Direct Linux boot of PetaLinux 2019.2:
-+
-+.. code-block:: bash
-+
-+  $ qemu-system-aarch64  -M xlnx-versal-virt -m 2G \
-+      -serial mon:stdio -display none \
-+      -kernel petalinux-v2019.2/Image \
-+      -append "rdinit=/sbin/init console=ttyAMA0,115200n8 earlycon=pl011,mmio,0xFF000000,115200n8" \
-+      -net nic,model=cadence_gem,netdev=net0 -netdev user,id=net0 \
-+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
-+      -object rng-random,filename=/dev/urandom,id=rng0
-+
-+Boot PetaLinux 2019.2 via ARM Trusted Firmware (2018.3 because the 2019.2
-+version of ATF tries to configure the CCI which we don't model) and U-boot:
-+
-+.. code-block:: bash
-+
-+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 2G \
-+      -serial stdio -display none \
-+      -device loader,file=petalinux-v2018.3/bl31.elf,cpu-num=0 \
-+      -device loader,file=petalinux-v2019.2/u-boot.elf \
-+      -device loader,addr=0x20000000,file=petalinux-v2019.2/Image \
-+      -nic user -nic user \
-+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
-+      -object rng-random,filename=/dev/urandom,id=rng0
-+
-+Run the following at the U-Boot prompt:
-+
-+.. code-block:: bash
-+
-+  Versal>
-+  fdt addr $fdtcontroladdr
-+  fdt move $fdtcontroladdr 0x40000000
-+  fdt set /timer clock-frequency <0x3dfd240>
-+  setenv bootargs "rdinit=/sbin/init maxcpus=1 console=ttyAMA0,115200n8 earlycon=pl011,mmio,0xFF000000,115200n8"
-+  booti 20000000 - 40000000
-+  fdt addr $fdtcontroladdr
-+
-+Boot Linux as DOM0 on Xen via U-Boot:
-+
-+.. code-block:: bash
-+
-+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 4G \
-+      -serial stdio -display none \
-+      -device loader,file=petalinux-v2019.2/u-boot.elf,cpu-num=0 \
-+      -device loader,addr=0x30000000,file=linux/2018-04-24/xen \
-+      -device loader,addr=0x40000000,file=petalinux-v2019.2/Image \
-+      -nic user -nic user \
-+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
-+      -object rng-random,filename=/dev/urandom,id=rng0
-+
-+Run the following at the U-Boot prompt:
-+
-+.. code-block:: bash
-+
-+  Versal>
-+  fdt addr $fdtcontroladdr
-+  fdt move $fdtcontroladdr 0x20000000
-+  fdt set /timer clock-frequency <0x3dfd240>
-+  fdt set /chosen xen,xen-bootargs "console=dtuart dtuart=/uart@ff000000 dom0_mem=640M bootscrub=0 maxcpus=1 timer_slop=0"
-+  fdt set /chosen xen,dom0-bootargs "rdinit=/sbin/init clk_ignore_unused console=hvc0 maxcpus=1"
-+  fdt mknode /chosen dom0
-+  fdt set /chosen/dom0 compatible "xen,multiboot-module"
-+  fdt set /chosen/dom0 reg <0x00000000 0x40000000 0x0 0x03100000>
-+  booti 30000000 - 20000000
-+
-+Boot Linux as Dom0 on Xen via ARM Trusted Firmware and U-Boot:
-+
-+.. code-block:: bash
-+
-+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 4G \
-+      -serial stdio -display none \
-+      -device loader,file=petalinux-v2018.3/bl31.elf,cpu-num=0 \
-+      -device loader,file=petalinux-v2019.2/u-boot.elf \
-+      -device loader,addr=0x30000000,file=linux/2018-04-24/xen \
-+      -device loader,addr=0x40000000,file=petalinux-v2019.2/Image \
-+      -nic user -nic user \
-+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
-+      -object rng-random,filename=/dev/urandom,id=rng0
-+
-+Run the following at the U-Boot prompt:
-+
-+.. code-block:: bash
-+
-+  Versal>
-+  fdt addr $fdtcontroladdr
-+  fdt move $fdtcontroladdr 0x20000000
-+  fdt set /timer clock-frequency <0x3dfd240>
-+  fdt set /chosen xen,xen-bootargs "console=dtuart dtuart=/uart@ff000000 dom0_mem=640M bootscrub=0 maxcpus=1 timer_slop=0"
-+  fdt set /chosen xen,dom0-bootargs "rdinit=/sbin/init clk_ignore_unused console=hvc0 maxcpus=1"
-+  fdt mknode /chosen dom0
-+  fdt set /chosen/dom0 compatible "xen,multiboot-module"
-+  fdt set /chosen/dom0 reg <0x00000000 0x40000000 0x0 0x03100000>
-+  booti 30000000 - 20000000
-+
-diff --git a/docs/system/target-arm.rst b/docs/system/target-arm.rst
-index XXXXXXX..XXXXXXX 100644
---- a/docs/system/target-arm.rst
-+++ b/docs/system/target-arm.rst
-@@ -XXX,XX +XXX,XX @@ undocumented; you can get a complete list by running
-    arm/sx1
-    arm/stellaris
-    arm/virt
-+   arm/xlnx-versal-virt
- Arm CPU features
- ================
 diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
 --- a/MAINTAINERS
 +++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ F: hw/misc/zynq*
+@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
- F: include/hw/misc/zynq*
- X: hw/ssi/xilinx_*
+ Xilinx CAN
+ M: Francisco Iglesias <francisco.iglesias@amd.com>
--Xilinx ZynqMP
++M: Vikram Garhwal <vikram.garhwal@bytedance.com>
-+Xilinx ZynqMP and Versal
+ S: Maintained
- M: Alistair Francis <alistair@alistair23.me>
+ F: hw/net/can/xlnx-*
- M: Edgar E. Iglesias <edgar.iglesias@gmail.com>
+ F: include/hw/net/xlnx-*
- M: Peter Maydell <peter.maydell@linaro.org>
+@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
-@@ -XXX,XX +XXX,XX @@ F: include/hw/*/xlnx*.h
+ CAN bus subsystem and hardware
- F: include/hw/ssi/xilinx_spips.h
+ M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
- F: hw/display/dpcd.c
+ M: Francisco Iglesias <francisco.iglesias@amd.com>
- F: include/hw/display/dpcd.h
++M: Vikram Garhwal <vikram.garhwal@bytedance.com>
-+F: docs/system/arm/xlnx-versal-virt.rst
+ S: Maintained
+ W: https://canbus.pages.fel.cvut.cz/
- ARM ACPI Subsystem
+ F: net/can/*
  M: Shannon Zhao <shannon.zhaosl@gmail.com>
 --
-.20.1
+.34.1

First arm pullreq for 5.2: Eric's SMMU stuff, and a bunch of
cleanup/refactoring from me.

thanks
-- PMM

The following changes since commit 8367a77c4d3f6e1e60890f5510304feb2c621611:

Merge remote-tracking branch 'remotes/vivier2/tags/linux-user-for-5.2-pull-request' into staging (2020-08-23 16:34:43 +0100)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20200824

for you to fetch changes up to b34aa5129e9c3aff890b4f4bcc84962e94185629:

target/arm: Use correct FPST for VCMLA, VCADD on fp16 (2020-08-24 10:15:12 +0100)

----------------------------------------------------------------
target-arm queue:
 * hw/cpu/a9mpcore: Verify the machine use Cortex-A9 cores
 * hw/arm/smmuv3: Implement SMMUv3.2 range-invalidation
 * docs/system/arm: Document the Xilinx Versal Virt board
 * target/arm: Make M-profile NOCP take precedence over UNDEF
 * target/arm: Use correct FPST for VCMLA, VCADD on fp16
 * target/arm: Various cleanups preparing for fp16 support

----------------------------------------------------------------
Edgar E. Iglesias (1):
      docs/system/arm: Document the Xilinx Versal Virt board

Eric Auger (11):
      hw/arm/smmu-common: Factorize some code in smmu_ptw_64()
      hw/arm/smmu-common: Add IOTLB helpers
      hw/arm/smmu: Introduce smmu_get_iotlb_key()
      hw/arm/smmu: Introduce SMMUTLBEntry for PTW and IOTLB value
      hw/arm/smmu-common: Manage IOTLB block entries
      hw/arm/smmuv3: Introduce smmuv3_s1_range_inval() helper
      hw/arm/smmuv3: Get prepared for range invalidation
      hw/arm/smmuv3: Fix IIDR offset
      hw/arm/smmuv3: Let AIDR advertise SMMUv3.0 support
      hw/arm/smmuv3: Support HAD and advertise SMMUv3.1 support
      hw/arm/smmuv3: Advertise SMMUv3.2 range invalidation

Peter Maydell (14):
      target/arm: Pull handling of XScale insns out of disas_coproc_insn()
      target/arm: Separate decode from handling of coproc insns
      target/arm: Convert A32 coprocessor insns to decodetree
      target/arm: Tidy up disas_arm_insn()
      target/arm: Do M-profile NOCP checks early and via decodetree
      target/arm: Convert T32 coprocessor insns to decodetree
      target/arm: Remove ARCH macro
      target/arm: Delete unused VFP_DREG macros
      target/arm/translate.c: Delete/amend incorrect comments
      target/arm: Delete unused ARM_FEATURE_CRC
      target/arm: Replace A64 get_fpstatus_ptr() with generic fpstatus_ptr()
      target/arm: Make A32/T32 use new fpstatus_ptr() API
      target/arm: Implement FPST_STD_F16 fpstatus
      target/arm: Use correct FPST for VCMLA, VCADD on fp16

Philippe Mathieu-Daudé (1):
      hw/cpu/a9mpcore: Verify the machine use Cortex-A9 cores

From: Philippe Mathieu-Daudé <f4bug@amsat.org>

The 'Cortex-A9MPCore internal peripheral' block can only be
used with Cortex A5 and A9 cores. As we don't model the A5
yet, simply check the machine cpu core is a Cortex A9. If
not return an error.

Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Reviewed-by: Alistair Francis <alistair.francis@wdc.com>
Message-id: 20200709152337.15533-1-f4bug@amsat.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/cpu/a9mpcore.c | 12 +++++++++++-
 1 file changed, 11 insertions(+), 1 deletion(-)

diff --git a/hw/cpu/a9mpcore.c b/hw/cpu/a9mpcore.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/cpu/a9mpcore.c
+++ b/hw/cpu/a9mpcore.c
@@ -XXX,XX +XXX,XX @@
 #include "hw/irq.h"
 #include "hw/qdev-properties.h"
 #include "hw/core/cpu.h"
+#include "cpu.h"
 
 #define A9_GIC_NUM_PRIORITY_BITS    5
 
@@ -XXX,XX +XXX,XX @@ static void a9mp_priv_realize(DeviceState *dev, Error **errp)
                  *wdtbusdev;
     int i;
     bool has_el3;
+    CPUState *cpu0;
     Object *cpuobj;
 
+    cpu0 = qemu_get_cpu(0);
+    cpuobj = OBJECT(cpu0);
+    if (strcmp(object_get_typename(cpuobj), ARM_CPU_TYPE_NAME("cortex-a9"))) {
+        /* We might allow Cortex-A5 once we model it */
+        error_setg(errp,
+                   "Cortex-A9MPCore peripheral can only use Cortex-A9 CPU");
+        return;
+    }
+
     scudev = DEVICE(&s->scu);
     qdev_prop_set_uint32(scudev, "num-cpu", s->num_cpu);
     if (!sysbus_realize(SYS_BUS_DEVICE(&s->scu), errp)) {
@@ -XXX,XX +XXX,XX @@ static void a9mp_priv_realize(DeviceState *dev, Error **errp)
     /* Make the GIC's TZ support match the CPUs. We assume that
      * either all the CPUs have TZ, or none do.
      */
-    cpuobj = OBJECT(qemu_get_cpu(0));
     has_el3 = object_property_find(cpuobj, "has_el3", NULL) &&
         object_property_get_bool(cpuobj, "has_el3", &error_abort);
     qdev_prop_set_bit(gicdev, "has-security-extensions", has_el3);
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Page and block PTE decoding can share some code. Let's
first handle table PTE and factorize some code shared by
page and block PTEs.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-2-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-common.c | 48 ++++++++++++++++----------------------------
 1 file changed, 17 insertions(+), 31 deletions(-)

diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
         uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
         uint64_t mask = subpage_size - 1;
         uint32_t offset = iova_level_offset(iova, inputsize, level, granule_sz);
-        uint64_t pte;
+        uint64_t pte, gpa;
         dma_addr_t pte_addr = baseaddr + offset * sizeof(pte);
         uint8_t ap;
 
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
         if (is_invalid_pte(pte) || is_reserved_pte(pte, level)) {
             trace_smmu_ptw_invalid_pte(stage, level, baseaddr,
                                        pte_addr, offset, pte);
-            info->type = SMMU_PTW_ERR_TRANSLATION;
-            goto error;
+            break;
         }
 
-        if (is_page_pte(pte, level)) {
-            uint64_t gpa = get_page_pte_address(pte, granule_sz);
+        if (is_table_pte(pte, level)) {
+            ap = PTE_APTABLE(pte);
 
-            ap = PTE_AP(pte);
             if (is_permission_fault(ap, perm)) {
                 info->type = SMMU_PTW_ERR_PERMISSION;
                 goto error;
             }
-
-            tlbe->translated_addr = gpa + (iova & mask);
-            tlbe->perm = PTE_AP_TO_PERM(ap);
+            baseaddr = get_table_pte_address(pte, granule_sz);
+            level++;
+            continue;
+        } else if (is_page_pte(pte, level)) {
+            gpa = get_page_pte_address(pte, granule_sz);
             trace_smmu_ptw_page_pte(stage, level, iova,
                                     baseaddr, pte_addr, pte, gpa);
-            return 0;
-        }
-        if (is_block_pte(pte, level)) {
+        } else {
             uint64_t block_size;
-            hwaddr gpa = get_block_pte_address(pte, level, granule_sz,
-                                               &block_size);
-
-            ap = PTE_AP(pte);
-            if (is_permission_fault(ap, perm)) {
-                info->type = SMMU_PTW_ERR_PERMISSION;
-                goto error;
-            }
 
+            gpa = get_block_pte_address(pte, level, granule_sz,
+                                        &block_size);
             trace_smmu_ptw_block_pte(stage, level, baseaddr,
                                      pte_addr, pte, iova, gpa,
                                      block_size >> 20);
-
-            tlbe->translated_addr = gpa + (iova & mask);
-            tlbe->perm = PTE_AP_TO_PERM(ap);
-            return 0;
         }
-
-        /* table pte */
-        ap = PTE_APTABLE(pte);
-
+        ap = PTE_AP(pte);
         if (is_permission_fault(ap, perm)) {
             info->type = SMMU_PTW_ERR_PERMISSION;
             goto error;
         }
-        baseaddr = get_table_pte_address(pte, granule_sz);
-        level++;
-    }
 
+        tlbe->translated_addr = gpa + (iova & mask);
+        tlbe->perm = PTE_AP_TO_PERM(ap);
+        return 0;
+    }
     info->type = SMMU_PTW_ERR_TRANSLATION;
 
 error:
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Add two helpers: one to lookup for a given IOTLB entry and
one to insert a new entry. We also move the tracing there.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-3-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h |  2 ++
 hw/arm/smmu-common.c         | 36 ++++++++++++++++++++++++++++++++++++
 hw/arm/smmuv3.c              | 26 ++------------------------
 hw/arm/trace-events          |  5 +++--
 4 files changed, 43 insertions(+), 26 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
 
 #define SMMU_IOTLB_MAX_SIZE 256
 
+IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry);
 void smmu_iotlb_inv_all(SMMUState *s);
 void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
 void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova);
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@
 
 /* IOTLB Management */
 
+IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+                                 hwaddr iova)
+{
+    SMMUIOTLBKey key = {.asid = cfg->asid, .iova = iova};
+    IOMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
+
+    if (entry) {
+        cfg->iotlb_hits++;
+        trace_smmu_iotlb_lookup_hit(cfg->asid, iova,
+                                    cfg->iotlb_hits, cfg->iotlb_misses,
+                                    100 * cfg->iotlb_hits /
+                                    (cfg->iotlb_hits + cfg->iotlb_misses));
+    } else {
+        cfg->iotlb_misses++;
+        trace_smmu_iotlb_lookup_miss(cfg->asid, iova,
+                                     cfg->iotlb_hits, cfg->iotlb_misses,
+                                     100 * cfg->iotlb_hits /
+                                     (cfg->iotlb_hits + cfg->iotlb_misses));
+    }
+    return entry;
+}
+
+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
+{
+    SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
+
+    if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
+        smmu_iotlb_inv_all(bs);
+    }
+
+    key->asid = cfg->asid;
+    key->iova = entry->iova;
+    trace_smmu_iotlb_insert(cfg->asid, entry->iova);
+    g_hash_table_insert(bs->iotlb, key, entry);
+}
+
 inline void smmu_iotlb_inv_all(SMMUState *s)
 {
     trace_smmu_iotlb_inv_all();
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         .addr_mask = ~(hwaddr)0,
         .perm = IOMMU_NONE,
     };
-    SMMUIOTLBKey key, *new_key;
 
     qemu_mutex_lock(&s->mutex);
 
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
     page_mask = (1ULL << (tt->granule_sz)) - 1;
     aligned_addr = addr & ~page_mask;
 
-    key.asid = cfg->asid;
-    key.iova = aligned_addr;
-
-    cached_entry = g_hash_table_lookup(bs->iotlb, &key);
+    cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
     if (cached_entry) {
-        cfg->iotlb_hits++;
-        trace_smmu_iotlb_cache_hit(cfg->asid, aligned_addr,
-                                   cfg->iotlb_hits, cfg->iotlb_misses,
-                                   100 * cfg->iotlb_hits /
-                                   (cfg->iotlb_hits + cfg->iotlb_misses));
         if ((flag & IOMMU_WO) && !(cached_entry->perm & IOMMU_WO)) {
             status = SMMU_TRANS_ERROR;
             if (event.record_trans_faults) {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         goto epilogue;
     }
 
-    cfg->iotlb_misses++;
-    trace_smmu_iotlb_cache_miss(cfg->asid, addr & ~page_mask,
-                                cfg->iotlb_hits, cfg->iotlb_misses,
-                                100 * cfg->iotlb_hits /
-                                (cfg->iotlb_hits + cfg->iotlb_misses));
-
-    if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
-        smmu_iotlb_inv_all(bs);
-    }
-
     cached_entry = g_new0(IOMMUTLBEntry, 1);
 
     if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         }
         status = SMMU_TRANS_ERROR;
     } else {
-        new_key = g_new0(SMMUIOTLBKey, 1);
-        new_key->asid = cfg->asid;
-        new_key->iova = aligned_addr;
-        g_hash_table_insert(bs->iotlb, new_key, cached_entry);
+        smmu_iotlb_insert(bs, cfg, cached_entry);
         status = SMMU_TRANS_SUCCESS;
     }
 
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_iotlb_inv_all(void) "IOTLB invalidate all"
 smmu_iotlb_inv_asid(uint16_t asid) "IOTLB invalidate asid=%d"
 smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr=0x%"PRIx64
 smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
+smmu_iotlb_lookup_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
+smmu_iotlb_lookup_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
+smmu_iotlb_insert(uint16_t asid, uint64_t addr) "IOTLB ++ asid=%d addr=0x%"PRIx64
 
 # smmuv3.c
 smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_tlbi_nh_va(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d a
 smmuv3_cmdq_tlbi_nh_vaa(int vmid, uint64_t addr) "vmid =%d addr=0x%"PRIx64
 smmuv3_cmdq_tlbi_nh(void) ""
 smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
-smmu_iotlb_cache_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
-smmu_iotlb_cache_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
 smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
 smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Introduce the smmu_get_iotlb_key() helper and the
SMMU_IOTLB_ASID() macro. Also move smmu_get_iotlb_key and
smmu_iotlb_key_hash in the IOTLB related code section.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-4-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-internal.h       |  1 +
 include/hw/arm/smmu-common.h |  1 +
 hw/arm/smmu-common.c         | 66 ++++++++++++++++++++----------------
 3 files changed, 38 insertions(+), 30 deletions(-)

From: Eric Auger <eric.auger@redhat.com>

Introduce a specialized SMMUTLBEntry to store the result of
the PTW and cache in the IOTLB. This structure extends the
generic IOMMUTLBEntry struct with the level of the entry and
the granule size.

Those latter will be useful when implementing range invalidation.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-5-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/arm/smmu-common.h | 12 +++++++++---
 hw/arm/smmu-common.c         | 32 +++++++++++++++++---------------
 hw/arm/smmuv3.c              | 10 +++++-----
 3 files changed, 31 insertions(+), 23 deletions(-)

diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransTableInfo {
     uint8_t granule_sz;        /* granule page shift */
 } SMMUTransTableInfo;
 
+typedef struct SMMUTLBEntry {
+    IOMMUTLBEntry entry;
+    uint8_t level;
+    uint8_t granule;
+} SMMUTLBEntry;
+
 /*
  * Generic structure populated by derived SMMU devices
  * after decoding the configuration information and used as
@@ -XXX,XX +XXX,XX @@ static inline uint16_t smmu_get_sid(SMMUDevice *sdev)
  * pair, according to @cfg translation config
  */
 int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
-             IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
+             SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info);
 
 /**
  * select_tt - compute which translation table shall be used according to
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
 
 #define SMMU_IOTLB_MAX_SIZE 256
 
-IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
-void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry);
+SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
 SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova);
 void smmu_iotlb_inv_all(SMMUState *s);
 void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova)
     return key;
 }
 
-IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
-                                 hwaddr iova)
+SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+                                hwaddr iova)
 {
     SMMUIOTLBKey key = smmu_get_iotlb_key(cfg->asid, iova);
-    IOMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
+    SMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
 
     if (entry) {
         cfg->iotlb_hits++;
@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
     return entry;
 }
 
-void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
+void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *new)
 {
     SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
 
@@ -XXX,XX +XXX,XX @@ void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, IOMMUTLBEntry *entry)
         smmu_iotlb_inv_all(bs);
     }
 
-    *key = smmu_get_iotlb_key(cfg->asid, entry->iova);
-    trace_smmu_iotlb_insert(cfg->asid, entry->iova);
-    g_hash_table_insert(bs->iotlb, key, entry);
+    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova);
+    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova);
+    g_hash_table_insert(bs->iotlb, key, new);
 }
 
 inline void smmu_iotlb_inv_all(SMMUState *s)
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
  * @cfg: translation config
  * @iova: iova to translate
  * @perm: access type
- * @tlbe: IOMMUTLBEntry (out)
+ * @tlbe: SMMUTLBEntry (out)
  * @info: handle to an error info
  *
  * Return 0 on success, < 0 on error. In case of error, @info is filled
@@ -XXX,XX +XXX,XX @@ SMMUTransTableInfo *select_tt(SMMUTransCfg *cfg, dma_addr_t iova)
  */
 static int smmu_ptw_64(SMMUTransCfg *cfg,
                        dma_addr_t iova, IOMMUAccessFlags perm,
-                       IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
+                       SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
     dma_addr_t baseaddr, indexmask;
     int stage = cfg->stage;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
     baseaddr = extract64(tt->ttb, 0, 48);
     baseaddr &= ~indexmask;
 
-    tlbe->iova = iova;
-    tlbe->addr_mask = (1 << granule_sz) - 1;
+    tlbe->entry.iova = iova;
+    tlbe->entry.addr_mask = (1 << granule_sz) - 1;
 
     while (level <= 3) {
         uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
             goto error;
         }
 
-        tlbe->translated_addr = gpa + (iova & mask);
-        tlbe->perm = PTE_AP_TO_PERM(ap);
+        tlbe->entry.translated_addr = gpa + (iova & mask);
+        tlbe->entry.perm = PTE_AP_TO_PERM(ap);
+        tlbe->level = level;
+        tlbe->granule = granule_sz;
         return 0;
     }
     info->type = SMMU_PTW_ERR_TRANSLATION;
 
 error:
-    tlbe->perm = IOMMU_NONE;
+    tlbe->entry.perm = IOMMU_NONE;
     return -EINVAL;
 }
 
@@ -XXX,XX +XXX,XX @@ error:
  * return 0 on success
  */
 inline int smmu_ptw(SMMUTransCfg *cfg, dma_addr_t iova, IOMMUAccessFlags perm,
-             IOMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
+                    SMMUTLBEntry *tlbe, SMMUPTWEventInfo *info)
 {
     if (!cfg->aa64) {
         /*
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
     SMMUTranslationStatus status;
     SMMUState *bs = ARM_SMMU(s);
     uint64_t page_mask, aligned_addr;
-    IOMMUTLBEntry *cached_entry = NULL;
+    SMMUTLBEntry *cached_entry = NULL;
     SMMUTransTableInfo *tt;
     SMMUTransCfg *cfg = NULL;
     IOMMUTLBEntry entry = {
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
 
     cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
     if (cached_entry) {
-        if ((flag & IOMMU_WO) && !(cached_entry->perm & IOMMU_WO)) {
+        if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
             status = SMMU_TRANS_ERROR;
             if (event.record_trans_faults) {
                 event.type = SMMU_EVT_F_PERMISSION;
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
         goto epilogue;
     }
 
-    cached_entry = g_new0(IOMMUTLBEntry, 1);
+    cached_entry = g_new0(SMMUTLBEntry, 1);
 
     if (smmu_ptw(cfg, aligned_addr, flag, cached_entry, &ptw_info)) {
         g_free(cached_entry);
@@ -XXX,XX +XXX,XX @@ epilogue:
     switch (status) {
     case SMMU_TRANS_SUCCESS:
         entry.perm = flag;
-        entry.translated_addr = cached_entry->translated_addr +
+        entry.translated_addr = cached_entry->entry.translated_addr +
                                     (addr & page_mask);
-        entry.addr_mask = cached_entry->addr_mask;
+        entry.addr_mask = cached_entry->entry.addr_mask;
         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
                                        entry.translated_addr, entry.perm);
         break;
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

At the moment each entry in the IOTLB corresponds to a page sized
mapping (4K, 16K or 64K), even if the page belongs to a mapped
block. In case of block mapping this unefficiently consumes IOTLB
entries.

Change the value of the entry so that it reflects the actual
mapping it belongs to (block or page start address and size).

Also the level/tg of the entry is encoded in the key. In subsequent
patches we will enable range invalidation. This latter is able
to provide the level/tg of the entry.

Encoding the level/tg directly in the key will allow to invalidate
using g_hash_table_remove() when num_pages equals to 1.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-6-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmu-internal.h       |  7 ++++
 include/hw/arm/smmu-common.h | 10 ++++--
 hw/arm/smmu-common.c         | 67 ++++++++++++++++++++++++++----------
 hw/arm/smmuv3.c              |  6 ++--
 hw/arm/trace-events          |  2 +-
 5 files changed, 67 insertions(+), 25 deletions(-)

diff --git a/hw/arm/smmu-internal.h b/hw/arm/smmu-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-internal.h
+++ b/hw/arm/smmu-internal.h
@@ -XXX,XX +XXX,XX @@ uint64_t iova_level_offset(uint64_t iova, int inputsize,
 }
 
 #define SMMU_IOTLB_ASID(key) ((key).asid)
+
+typedef struct SMMUIOTLBPageInvInfo {
+    int asid;
+    uint64_t iova;
+    uint64_t mask;
+} SMMUIOTLBPageInvInfo;
+
 #endif
diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUPciBus {
 typedef struct SMMUIOTLBKey {
     uint64_t iova;
     uint16_t asid;
+    uint8_t tg;
+    uint8_t level;
 } SMMUIOTLBKey;
 
 typedef struct SMMUState {
@@ -XXX,XX +XXX,XX @@ IOMMUMemoryRegion *smmu_iommu_mr(SMMUState *s, uint32_t sid);
 
 #define SMMU_IOTLB_MAX_SIZE 256
 
-SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg, hwaddr iova);
+SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
+                                SMMUTransTableInfo *tt, hwaddr iova);
 void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *entry);
-SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova);
+SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
+                                uint8_t tg, uint8_t level);
 void smmu_iotlb_inv_all(SMMUState *s);
 void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
-void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova);
+void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova);
 
 /* Unmap the range of all the notifiers registered to any IOMMU mr */
 void smmu_inv_notifiers_all(SMMUState *s);
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static guint smmu_iotlb_key_hash(gconstpointer v)
 
     /* Jenkins hash */
     a = b = c = JHASH_INITVAL + sizeof(*key);
-    a += key->asid;
+    a += key->asid + key->level + key->tg;
     b += extract64(key->iova, 0, 32);
     c += extract64(key->iova, 32, 32);
 
@@ -XXX,XX +XXX,XX @@ static guint smmu_iotlb_key_hash(gconstpointer v)
 
 static gboolean smmu_iotlb_key_equal(gconstpointer v1, gconstpointer v2)
 {
-    const SMMUIOTLBKey *k1 = v1;
-    const SMMUIOTLBKey *k2 = v2;
+    SMMUIOTLBKey *k1 = (SMMUIOTLBKey *)v1, *k2 = (SMMUIOTLBKey *)v2;
 
-    return (k1->asid == k2->asid) && (k1->iova == k2->iova);
+    return (k1->asid == k2->asid) && (k1->iova == k2->iova) &&
+           (k1->level == k2->level) && (k1->tg == k2->tg);
 }
 
-SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova)
+SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
+                                uint8_t tg, uint8_t level)
 {
-    SMMUIOTLBKey key = {.asid = asid, .iova = iova};
+    SMMUIOTLBKey key = {.asid = asid, .iova = iova, .tg = tg, .level = level};
 
     return key;
 }
 
 SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
-                                hwaddr iova)
+                                SMMUTransTableInfo *tt, hwaddr iova)
 {
-    SMMUIOTLBKey key = smmu_get_iotlb_key(cfg->asid, iova);
-    SMMUTLBEntry *entry = g_hash_table_lookup(bs->iotlb, &key);
+    uint8_t tg = (tt->granule_sz - 10) / 2;
+    uint8_t inputsize = 64 - tt->tsz;
+    uint8_t stride = tt->granule_sz - 3;
+    uint8_t level = 4 - (inputsize - 4) / stride;
+    SMMUTLBEntry *entry = NULL;
+
+    while (level <= 3) {
+        uint64_t subpage_size = 1ULL << level_shift(level, tt->granule_sz);
+        uint64_t mask = subpage_size - 1;
+        SMMUIOTLBKey key;
+
+        key = smmu_get_iotlb_key(cfg->asid, iova & ~mask, tg, level);
+        entry = g_hash_table_lookup(bs->iotlb, &key);
+        if (entry) {
+            break;
+        }
+        level++;
+    }
 
     if (entry) {
         cfg->iotlb_hits++;
@@ -XXX,XX +XXX,XX @@ SMMUTLBEntry *smmu_iotlb_lookup(SMMUState *bs, SMMUTransCfg *cfg,
 void smmu_iotlb_insert(SMMUState *bs, SMMUTransCfg *cfg, SMMUTLBEntry *new)
 {
     SMMUIOTLBKey *key = g_new0(SMMUIOTLBKey, 1);
+    uint8_t tg = (new->granule - 10) / 2;
 
     if (g_hash_table_size(bs->iotlb) >= SMMU_IOTLB_MAX_SIZE) {
         smmu_iotlb_inv_all(bs);
     }
 
-    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova);
-    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova);
+    *key = smmu_get_iotlb_key(cfg->asid, new->entry.iova, tg, new->level);
+    trace_smmu_iotlb_insert(cfg->asid, new->entry.iova, tg, new->level);
     g_hash_table_insert(bs->iotlb, key, new);
 }
 
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid(gpointer key, gpointer value,
     return SMMU_IOTLB_ASID(*iotlb_key) == asid;
 }
 
-inline void smmu_iotlb_inv_iova(SMMUState *s, uint16_t asid, dma_addr_t iova)
+static gboolean smmu_hash_remove_by_asid_iova(gpointer key, gpointer value,
+                                              gpointer user_data)
 {
-    SMMUIOTLBKey key = smmu_get_iotlb_key(asid, iova);
+    SMMUTLBEntry *iter = (SMMUTLBEntry *)value;
+    IOMMUTLBEntry *entry = &iter->entry;
+    SMMUIOTLBPageInvInfo *info = (SMMUIOTLBPageInvInfo *)user_data;
+    SMMUIOTLBKey iotlb_key = *(SMMUIOTLBKey *)key;
+
+    if (info->asid >= 0 && info->asid != SMMU_IOTLB_ASID(iotlb_key)) {
+        return false;
+    }
+    return (info->iova & ~entry->addr_mask) == entry->iova;
+}
+
+inline void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova)
+{
+    SMMUIOTLBPageInvInfo info = {.asid = asid, .iova = iova};
 
     trace_smmu_iotlb_inv_iova(asid, iova);
-    g_hash_table_remove(s->iotlb, &key);
+    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_iova, &info);
 }
 
 inline void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
     baseaddr = extract64(tt->ttb, 0, 48);
     baseaddr &= ~indexmask;
 
-    tlbe->entry.iova = iova;
-    tlbe->entry.addr_mask = (1 << granule_sz) - 1;
-
     while (level <= 3) {
         uint64_t subpage_size = 1ULL << level_shift(level, granule_sz);
         uint64_t mask = subpage_size - 1;
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
             goto error;
         }
 
-        tlbe->entry.translated_addr = gpa + (iova & mask);
+        tlbe->entry.translated_addr = gpa;
+        tlbe->entry.iova = iova & ~mask;
+        tlbe->entry.addr_mask = mask;
         tlbe->entry.perm = PTE_AP_TO_PERM(ap);
         tlbe->level = level;
         tlbe->granule = granule_sz;
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static IOMMUTLBEntry smmuv3_translate(IOMMUMemoryRegion *mr, hwaddr addr,
     page_mask = (1ULL << (tt->granule_sz)) - 1;
     aligned_addr = addr & ~page_mask;
 
-    cached_entry = smmu_iotlb_lookup(bs, cfg, aligned_addr);
+    cached_entry = smmu_iotlb_lookup(bs, cfg, tt, aligned_addr);
     if (cached_entry) {
         if ((flag & IOMMU_WO) && !(cached_entry->entry.perm & IOMMU_WO)) {
             status = SMMU_TRANS_ERROR;
@@ -XXX,XX +XXX,XX @@ epilogue:
     case SMMU_TRANS_SUCCESS:
         entry.perm = flag;
         entry.translated_addr = cached_entry->entry.translated_addr +
-                                    (addr & page_mask);
+                                    (addr & cached_entry->entry.addr_mask);
         entry.addr_mask = cached_entry->entry.addr_mask;
         trace_smmuv3_translate_success(mr->parent_obj.name, sid, addr,
                                        entry.translated_addr, entry.perm);
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
 
             trace_smmuv3_cmdq_tlbi_nh_vaa(vmid, addr);
             smmuv3_inv_notifiers_iova(bs, -1, addr);
-            smmu_iotlb_inv_all(bs);
+            smmu_iotlb_inv_iova(bs, -1, addr);
             break;
         }
         case SMMU_CMD_TLBI_NH_VA:
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmu_iotlb_inv_iova(uint16_t asid, uint64_t addr) "IOTLB invalidate asid=%d addr
 smmu_inv_notifiers_mr(const char *name) "iommu mr=%s"
 smmu_iotlb_lookup_hit(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache HIT asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
 smmu_iotlb_lookup_miss(uint16_t asid, uint64_t addr, uint32_t hit, uint32_t miss, uint32_t p) "IOTLB cache MISS asid=%d addr=0x%"PRIx64" hit=%d miss=%d hit rate=%d"
-smmu_iotlb_insert(uint16_t asid, uint64_t addr) "IOTLB ++ asid=%d addr=0x%"PRIx64
+smmu_iotlb_insert(uint16_t asid, uint64_t addr, uint8_t tg, uint8_t level) "IOTLB ++ asid=%d addr=0x%"PRIx64" tg=%d level=%d"
 
 # smmuv3.c
 smmuv3_read_mmio(uint64_t addr, uint64_t val, unsigned size, uint32_t r) "addr: 0x%"PRIx64" val:0x%"PRIx64" size: 0x%x(%d)"
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Let's introduce an helper for S1 IOVA range invalidation.
This will be used for NH_VA and NH_VAA commands. It decodes
the same fields, trace, calls the UNMAP notifiers and
invalidate the corresponding IOTLB entries.

At the moment, we do not support 3.2 range invalidation yet.
So it reduces to a single IOVA invalidation.

Note the leaf bit now is also decoded for the CMD_TLBI_NH_VAA
command. At the moment it is only used for tracing.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-7-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3.c     | 36 +++++++++++++++++-------------------
 hw/arm/trace-events |  3 +--
 2 files changed, 18 insertions(+), 21 deletions(-)

diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
     }
 }
 
+static void smmuv3_s1_range_inval(SMMUState *s, Cmd *cmd)
+{
+    dma_addr_t addr = CMD_ADDR(cmd);
+    uint8_t type = CMD_TYPE(cmd);
+    uint16_t vmid = CMD_VMID(cmd);
+    bool leaf = CMD_LEAF(cmd);
+    int asid = -1;
+
+    if (type == SMMU_CMD_TLBI_NH_VA) {
+        asid = CMD_ASID(cmd);
+    }
+    trace_smmuv3_s1_range_inval(vmid, asid, addr, leaf);
+    smmuv3_inv_notifiers_iova(s, asid, addr);
+    smmu_iotlb_inv_iova(s, asid, addr);
+}
+
 static int smmuv3_cmdq_consume(SMMUv3State *s)
 {
     SMMUState *bs = ARM_SMMU(s);
@@ -XXX,XX +XXX,XX @@ static int smmuv3_cmdq_consume(SMMUv3State *s)
             smmu_iotlb_inv_all(bs);
             break;
         case SMMU_CMD_TLBI_NH_VAA:
-        {
-            dma_addr_t addr = CMD_ADDR(&cmd);
-            uint16_t vmid = CMD_VMID(&cmd);
-
-            trace_smmuv3_cmdq_tlbi_nh_vaa(vmid, addr);
-            smmuv3_inv_notifiers_iova(bs, -1, addr);
-            smmu_iotlb_inv_iova(bs, -1, addr);
-            break;
-        }
         case SMMU_CMD_TLBI_NH_VA:
-        {
-            uint16_t asid = CMD_ASID(&cmd);
-            uint16_t vmid = CMD_VMID(&cmd);
-            dma_addr_t addr = CMD_ADDR(&cmd);
-            bool leaf = CMD_LEAF(&cmd);
-
-            trace_smmuv3_cmdq_tlbi_nh_va(vmid, asid, addr, leaf);
-            smmuv3_inv_notifiers_iova(bs, asid, addr);
-            smmu_iotlb_inv_iova(bs, asid, addr);
+            smmuv3_s1_range_inval(bs, &cmd);
             break;
-        }
         case SMMU_CMD_TLBI_EL3_ALL:
         case SMMU_CMD_TLBI_EL3_VA:
         case SMMU_CMD_TLBI_EL2_ALL:
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
 smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
 smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid %d (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid %d (hits=%d, misses=%d, hit rate=%d)"
-smmuv3_cmdq_tlbi_nh_va(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
-smmuv3_cmdq_tlbi_nh_vaa(int vmid, uint64_t addr) "vmid =%d addr=0x%"PRIx64
+smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
 smmuv3_cmdq_tlbi_nh(void) ""
 smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Enhance the smmu_iotlb_inv_iova() helper with range invalidation.
This uses the new fields passed in the NH_VA and NH_VAA commands:
the size of the range, the level and the granule.

As NH_VA and NH_VAA both use those fields, their decoding and
handling is factorized in a new smmuv3_s1_range_inval() helper.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-8-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h     |  4 +++
 include/hw/arm/smmu-common.h |  3 +-
 hw/arm/smmu-common.c         | 25 +++++++++++---
 hw/arm/smmuv3.c              | 64 +++++++++++++++++++++++-------------
 hw/arm/trace-events          |  4 +--
 5 files changed, 69 insertions(+), 31 deletions(-)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ enum { /* Command completion notification */
 };
 
 #define CMD_TYPE(x)         extract32((x)->word[0], 0 , 8)
+#define CMD_NUM(x)          extract32((x)->word[0], 12 , 5)
+#define CMD_SCALE(x)        extract32((x)->word[0], 20 , 5)
 #define CMD_SSEC(x)         extract32((x)->word[0], 10, 1)
 #define CMD_SSV(x)          extract32((x)->word[0], 11, 1)
 #define CMD_RESUME_AC(x)    extract32((x)->word[0], 12, 1)
@@ -XXX,XX +XXX,XX @@ enum { /* Command completion notification */
 #define CMD_RESUME_STAG(x)  extract32((x)->word[2], 0 , 16)
 #define CMD_RESP(x)         extract32((x)->word[2], 11, 2)
 #define CMD_LEAF(x)         extract32((x)->word[2], 0 , 1)
+#define CMD_TTL(x)          extract32((x)->word[2], 8 , 2)
+#define CMD_TG(x)           extract32((x)->word[2], 10, 2)
 #define CMD_STE_RANGE(x)    extract32((x)->word[2], 0 , 5)
 #define CMD_ADDR(x) ({                                        \
             uint64_t high = (uint64_t)(x)->word[3];           \
diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ SMMUIOTLBKey smmu_get_iotlb_key(uint16_t asid, uint64_t iova,
                                 uint8_t tg, uint8_t level);
 void smmu_iotlb_inv_all(SMMUState *s);
 void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid);
-void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova);
+void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova,
+                         uint8_t tg, uint64_t num_pages, uint8_t ttl);
 
 /* Unmap the range of all the notifiers registered to any IOMMU mr */
 void smmu_inv_notifiers_all(SMMUState *s);
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static gboolean smmu_hash_remove_by_asid_iova(gpointer key, gpointer value,
     if (info->asid >= 0 && info->asid != SMMU_IOTLB_ASID(iotlb_key)) {
         return false;
     }
-    return (info->iova & ~entry->addr_mask) == entry->iova;
+    return ((info->iova & ~entry->addr_mask) == entry->iova) ||
+           ((entry->iova & ~info->mask) == info->iova);
 }
 
-inline void smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova)
+inline void
+smmu_iotlb_inv_iova(SMMUState *s, int asid, dma_addr_t iova,
+                    uint8_t tg, uint64_t num_pages, uint8_t ttl)
 {
-    SMMUIOTLBPageInvInfo info = {.asid = asid, .iova = iova};
+    if (ttl && (num_pages == 1)) {
+        SMMUIOTLBKey key = smmu_get_iotlb_key(asid, iova, tg, ttl);
 
-    trace_smmu_iotlb_inv_iova(asid, iova);
-    g_hash_table_foreach_remove(s->iotlb, smmu_hash_remove_by_asid_iova, &info);
+        g_hash_table_remove(s->iotlb, &key);
+    } else {
+        /* if tg is not set we use 4KB range invalidation */
+        uint8_t granule = tg ? tg * 2 + 10 : 12;
+
+        SMMUIOTLBPageInvInfo info = {
+            .asid = asid, .iova = iova,
+            .mask = (num_pages * 1 << granule) - 1};
+
+        g_hash_table_foreach_remove(s->iotlb,
+                                    smmu_hash_remove_by_asid_iova,
+                                    &info);
+    }
 }
 
 inline void smmu_iotlb_inv_asid(SMMUState *s, uint16_t asid)
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ epilogue:
  * @n: notifier to be called
  * @asid: address space ID or negative value if we don't care
  * @iova: iova
+ * @tg: translation granule (if communicated through range invalidation)
+ * @num_pages: number of @granule sized pages (if tg != 0), otherwise 1
  */
 static void smmuv3_notify_iova(IOMMUMemoryRegion *mr,
                                IOMMUNotifier *n,
-                               int asid,
-                               dma_addr_t iova)
+                               int asid, dma_addr_t iova,
+                               uint8_t tg, uint64_t num_pages)
 {
     SMMUDevice *sdev = container_of(mr, SMMUDevice, iommu);
-    SMMUEventInfo event = {.inval_ste_allowed = true};
-    SMMUTransTableInfo *tt;
-    SMMUTransCfg *cfg;
     IOMMUTLBEntry entry;
+    uint8_t granule = tg;
 
-    cfg = smmuv3_get_config(sdev, &event);
-    if (!cfg) {
-        return;
-    }
+    if (!tg) {
+        SMMUEventInfo event = {.inval_ste_allowed = true};
+        SMMUTransCfg *cfg = smmuv3_get_config(sdev, &event);
+        SMMUTransTableInfo *tt;
 
-    if (asid >= 0 && cfg->asid != asid) {
-        return;
-    }
+        if (!cfg) {
+            return;
+        }
 
-    tt = select_tt(cfg, iova);
-    if (!tt) {
-        return;
+        if (asid >= 0 && cfg->asid != asid) {
+            return;
+        }
+
+        tt = select_tt(cfg, iova);
+        if (!tt) {
+            return;
+        }
+        granule = tt->granule_sz;
     }
 
     entry.target_as = &address_space_memory;
     entry.iova = iova;
-    entry.addr_mask = (1 << tt->granule_sz) - 1;
+    entry.addr_mask = num_pages * (1 << granule) - 1;
     entry.perm = IOMMU_NONE;
 
     memory_region_notify_one(n, &entry);
 }
 
-/* invalidate an asid/iova tuple in all mr's */
-static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
+/* invalidate an asid/iova range tuple in all mr's */
+static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova,
+                                      uint8_t tg, uint64_t num_pages)
 {
     SMMUDevice *sdev;
 
@@ -XXX,XX +XXX,XX @@ static void smmuv3_inv_notifiers_iova(SMMUState *s, int asid, dma_addr_t iova)
         IOMMUMemoryRegion *mr = &sdev->iommu;
         IOMMUNotifier *n;
 
-        trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, iova);
+        trace_smmuv3_inv_notifiers_iova(mr->parent_obj.name, asid, iova,
+                                        tg, num_pages);
 
         IOMMU_NOTIFIER_FOREACH(n, mr) {
-            smmuv3_notify_iova(mr, n, asid, iova);
+            smmuv3_notify_iova(mr, n, asid, iova, tg, num_pages);
         }
     }
 }
 
 static void smmuv3_s1_range_inval(SMMUState *s, Cmd *cmd)
 {
+    uint8_t scale = 0, num = 0, ttl = 0;
     dma_addr_t addr = CMD_ADDR(cmd);
     uint8_t type = CMD_TYPE(cmd);
     uint16_t vmid = CMD_VMID(cmd);
     bool leaf = CMD_LEAF(cmd);
+    uint8_t tg = CMD_TG(cmd);
+    hwaddr num_pages = 1;
     int asid = -1;
 
+    if (tg) {
+        scale = CMD_SCALE(cmd);
+        num = CMD_NUM(cmd);
+        ttl = CMD_TTL(cmd);
+        num_pages = (num + 1) * (1 << (scale));
+    }
+
     if (type == SMMU_CMD_TLBI_NH_VA) {
         asid = CMD_ASID(cmd);
     }
-    trace_smmuv3_s1_range_inval(vmid, asid, addr, leaf);
-    smmuv3_inv_notifiers_iova(s, asid, addr);
-    smmu_iotlb_inv_iova(s, asid, addr);
+    trace_smmuv3_s1_range_inval(vmid, asid, addr, tg, num_pages, ttl, leaf);
+    smmuv3_inv_notifiers_iova(s, asid, addr, tg, num_pages);
+    smmu_iotlb_inv_iova(s, asid, addr, tg, num_pages, ttl);
 }
 
 static int smmuv3_cmdq_consume(SMMUv3State *s)
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
 smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
 smmuv3_config_cache_hit(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache HIT for sid %d (hits=%d, misses=%d, hit rate=%d)"
 smmuv3_config_cache_miss(uint32_t sid, uint32_t hits, uint32_t misses, uint32_t perc) "Config cache MISS for sid %d (hits=%d, misses=%d, hit rate=%d)"
-smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" leaf=%d"
+smmuv3_s1_range_inval(int vmid, int asid, uint64_t addr, uint8_t tg, uint64_t num_pages, uint8_t ttl, bool leaf) "vmid =%d asid =%d addr=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64" ttl=%d leaf=%d"
 smmuv3_cmdq_tlbi_nh(void) ""
 smmuv3_cmdq_tlbi_nh_asid(uint16_t asid) "asid=%d"
 smmuv3_config_cache_inv(uint32_t sid) "Config cache INV for sid %d"
 smmuv3_notify_flag_add(const char *iommu) "ADD SMMUNotifier node for iommu mr=%s"
 smmuv3_notify_flag_del(const char *iommu) "DEL SMMUNotifier node for iommu mr=%s"
-smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint64_t iova) "iommu mr=%s asid=%d iova=0x%"PRIx64
+smmuv3_inv_notifiers_iova(const char *name, uint16_t asid, uint64_t iova, uint8_t tg, uint64_t num_pages) "iommu mr=%s asid=%d iova=0x%"PRIx64" tg=%d num_pages=0x%"PRIx64
 
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Add the support for AIDR register. It currently advertises
SMMU V3.0 spec.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-10-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h | 1 +
 include/hw/arm/smmuv3.h  | 1 +
 hw/arm/smmuv3.c          | 3 +++
 3 files changed, 5 insertions(+)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ REG32(IDR5,                0x14)
 #define SMMU_IDR5_OAS 4
 
 REG32(IIDR,                0x18)
+REG32(AIDR,                0x1c)
 REG32(CR0,                 0x20)
     FIELD(CR0, SMMU_ENABLE,   0, 1)
     FIELD(CR0, EVENTQEN,      2, 1)
diff --git a/include/hw/arm/smmuv3.h b/include/hw/arm/smmuv3.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmuv3.h
+++ b/include/hw/arm/smmuv3.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUv3State {
 
     uint32_t idr[6];
     uint32_t iidr;
+    uint32_t aidr;
     uint32_t cr[3];
     uint32_t cr0ack;
     uint32_t statusr;
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static MemTxResult smmu_readl(SMMUv3State *s, hwaddr offset,
     case A_IIDR:
         *data = s->iidr;
         return MEMTX_OK;
+    case A_AIDR:
+        *data = s->aidr;
+        return MEMTX_OK;
     case A_CR0:
         *data = s->cr[0];
         return MEMTX_OK;
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

HAD is a mandatory features with SMMUv3.1 if S1P is set, which is
our case. Other 3.1 mandatory features come with S2P which we don't
have.

So let's support HAD and advertise SMMUv3.1 support in AIDR.

HAD support allows the CD to disable hierarchical attributes, ie.
if the HAD0/1 bit is set, the APTable field of table descriptors
walked through TTB0/1 is ignored.

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-11-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h     | 2 ++
 include/hw/arm/smmu-common.h | 1 +
 hw/arm/smmu-common.c         | 2 +-
 hw/arm/smmuv3.c              | 6 +++++-
 hw/arm/trace-events          | 2 +-
 5 files changed, 10 insertions(+), 3 deletions(-)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ REG32(IDR1,                0x4)
 
 REG32(IDR2,                0x8)
 REG32(IDR3,                0xc)
+     FIELD(IDR3, HAD,         2, 1);
 REG32(IDR4,                0x10)
 REG32(IDR5,                0x14)
      FIELD(IDR5, OAS,         0, 3);
@@ -XXX,XX +XXX,XX @@ static inline int pa_range(STE *ste)
         lo = (x)->word[(sel) * 2 + 2] & ~0xfULL;            \
         hi | lo;                                            \
     })
+#define CD_HAD(x, sel)   extract32((x)->word[(sel) * 2 + 2], 1, 1)
 
 #define CD_TSZ(x, sel)   extract32((x)->word[0], (16 * (sel)) + 0, 6)
 #define CD_TG(x, sel)    extract32((x)->word[0], (16 * (sel)) + 6, 2)
diff --git a/include/hw/arm/smmu-common.h b/include/hw/arm/smmu-common.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/arm/smmu-common.h
+++ b/include/hw/arm/smmu-common.h
@@ -XXX,XX +XXX,XX @@ typedef struct SMMUTransTableInfo {
     uint64_t ttb;              /* TT base address */
     uint8_t tsz;               /* input range, ie. 2^(64 -tsz)*/
     uint8_t granule_sz;        /* granule page shift */
+    bool had;                  /* hierarchical attribute disable */
 } SMMUTransTableInfo;
 
 typedef struct SMMUTLBEntry {
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -XXX,XX +XXX,XX @@ static int smmu_ptw_64(SMMUTransCfg *cfg,
         if (is_table_pte(pte, level)) {
             ap = PTE_APTABLE(pte);
 
-            if (is_permission_fault(ap, perm)) {
+            if (is_permission_fault(ap, perm) && !tt->had) {
                 info->type = SMMU_PTW_ERR_PERMISSION;
                 goto error;
             }
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, EVENTQS, SMMU_EVENTQS);
     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, CMDQS,   SMMU_CMDQS);
 
+    s->idr[3] = FIELD_DP32(s->idr[3], IDR3, HAD, 1);
+
    /* 4K and 64K granule support */
     s->idr[5] = FIELD_DP32(s->idr[5], IDR5, GRAN4K, 1);
     s->idr[5] = FIELD_DP32(s->idr[5], IDR5, GRAN64K, 1);
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
 
     s->features = 0;
     s->sid_split = 0;
+    s->aidr = 0x1;
 }
 
 static int smmu_get_ste(SMMUv3State *s, dma_addr_t addr, STE *buf,
@@ -XXX,XX +XXX,XX @@ static int decode_cd(SMMUTransCfg *cfg, CD *cd, SMMUEventInfo *event)
         if (tt->ttb & ~(MAKE_64BIT_MASK(0, cfg->oas))) {
             goto bad_cd;
         }
-        trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz);
+        tt->had = CD_HAD(cd, i);
+        trace_smmuv3_decode_cd_tt(i, tt->tsz, tt->ttb, tt->granule_sz, tt->had);
     }
 
     event->record_trans_faults = CD_R(cd);
diff --git a/hw/arm/trace-events b/hw/arm/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/trace-events
+++ b/hw/arm/trace-events
@@ -XXX,XX +XXX,XX @@ smmuv3_translate_abort(const char *n, uint16_t sid, uint64_t addr, bool is_write
 smmuv3_translate_success(const char *n, uint16_t sid, uint64_t iova, uint64_t translated, int perm) "%s sid=%d iova=0x%"PRIx64" translated=0x%"PRIx64" perm=0x%x"
 smmuv3_get_cd(uint64_t addr) "CD addr: 0x%"PRIx64
 smmuv3_decode_cd(uint32_t oas) "oas=%d"
-smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d"
+smmuv3_decode_cd_tt(int i, uint32_t tsz, uint64_t ttb, uint32_t granule_sz, bool had) "TT[%d]:tsz:%d ttb:0x%"PRIx64" granule_sz:%d had:%d"
 smmuv3_cmdq_cfgi_ste(int streamid) "streamid =%d"
 smmuv3_cmdq_cfgi_ste_range(int start, int end) "start=0x%d - end=0x%d"
 smmuv3_cmdq_cfgi_cd(uint32_t sid) "streamid = %d"
-- 
2.20.1

From: Eric Auger <eric.auger@redhat.com>

Expose the RIL bit so that the guest driver uses range
invalidation. Although RIL is a 3.2 features, We let
the AIDR advertise SMMUv3.1 support as v3.x implementation
is allowed to implement features from v3.(x+1).

Signed-off-by: Eric Auger <eric.auger@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20200728150815.11446-12-eric.auger@redhat.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/smmuv3-internal.h | 1 +
 hw/arm/smmuv3.c          | 1 +
 2 files changed, 2 insertions(+)

diff --git a/hw/arm/smmuv3-internal.h b/hw/arm/smmuv3-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3-internal.h
+++ b/hw/arm/smmuv3-internal.h
@@ -XXX,XX +XXX,XX @@ REG32(IDR1,                0x4)
 REG32(IDR2,                0x8)
 REG32(IDR3,                0xc)
      FIELD(IDR3, HAD,         2, 1);
+     FIELD(IDR3, RIL,        10, 1);
 REG32(IDR4,                0x10)
 REG32(IDR5,                0x14)
      FIELD(IDR5, OAS,         0, 3);
diff --git a/hw/arm/smmuv3.c b/hw/arm/smmuv3.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/smmuv3.c
+++ b/hw/arm/smmuv3.c
@@ -XXX,XX +XXX,XX @@ static void smmuv3_init_regs(SMMUv3State *s)
     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, EVENTQS, SMMU_EVENTQS);
     s->idr[1] = FIELD_DP32(s->idr[1], IDR1, CMDQS,   SMMU_CMDQS);
 
+    s->idr[3] = FIELD_DP32(s->idr[3], IDR3, RIL, 1);
     s->idr[3] = FIELD_DP32(s->idr[3], IDR3, HAD, 1);
 
    /* 4K and 64K granule support */
-- 
2.20.1

From: "Edgar E. Iglesias" <edgar.iglesias@xilinx.com>

Document the Xilinx Versal Virt board.

Signed-off-by: Edgar E. Iglesias <edgar.iglesias@xilinx.com>
Message-id: 20200803164749.301971-2-edgar.iglesias@gmail.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 docs/system/arm/xlnx-versal-virt.rst | 176 +++++++++++++++++++++++++++
 docs/system/target-arm.rst           |   1 +
 MAINTAINERS                          |   3 +-
 3 files changed, 179 insertions(+), 1 deletion(-)
 create mode 100644 docs/system/arm/xlnx-versal-virt.rst

diff --git a/docs/system/arm/xlnx-versal-virt.rst b/docs/system/arm/xlnx-versal-virt.rst
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/docs/system/arm/xlnx-versal-virt.rst
@@ -XXX,XX +XXX,XX @@
+Xilinx Versal Virt (``xlnx-versal-virt``)
+=========================================
+
+Xilinx Versal is a family of heterogeneous multi-core SoCs
+(System on Chip) that combine traditional hardened CPUs and I/O
+peripherals in a Processing System (PS) with runtime programmable
+FPGA logic (PL) and an Artificial Intelligence Engine (AIE).
+
+More details here:
+https://www.xilinx.com/products/silicon-devices/acap/versal.html
+
+The family of Versal SoCs share a single architecture but come in
+different parts with different speed grades, amounts of PL and
+other differences.
+
+The Xilinx Versal Virt board in QEMU is a model of a virtual board
+(does not exist in reality) with a virtual Versal SoC without I/O
+limitations. Currently, we support the following cores and devices:
+
+Implemented CPU cores:
+
+- 2 ACPUs (ARM Cortex-A72)
+
+Implemented devices:
+
+- Interrupt controller (ARM GICv3)
+- 2 UARTs (ARM PL011)
+- An RTC (Versal built-in)
+- 2 GEMs (Cadence MACB Ethernet MACs)
+- 8 ADMA (Xilinx zDMA) channels
+- 2 SD Controllers
+- OCM (256KB of On Chip Memory)
+- DDR memory
+
+QEMU does not yet model any other devices, including the PL and the AI Engine.
+
+Other differences between the hardware and the QEMU model:
+
+- QEMU allows the amount of DDR memory provided to be specified with the
+  ``-m`` argument. If a DTB is provided on the command line then QEMU will
+  edit it to include suitable entries describing the Versal DDR memory ranges.
+
+- QEMU provides 8 virtio-mmio virtio transports; these start at
+  address ``0xa0000000`` and have IRQs from 111 and upwards.
+
+Running
+"""""""
+If the user provides an Operating System to be loaded, we expect users
+to use the ``-kernel`` command line option.
+
+Users can load firmware or boot-loaders with the ``-device loader`` options.
+
+When loading an OS, QEMU generates a DTB and selects an appropriate address
+where it gets loaded. This DTB will be passed to the kernel in register x0.
+
+If there's no ``-kernel`` option, we generate a DTB and place it at 0x1000
+for boot-loaders or firmware to pick it up.
+
+If users want to provide their own DTB, they can use the ``-dtb`` option.
+These DTBs will have their memory nodes modified to match QEMU's
+selected ram_size option before they get passed to the kernel or FW.
+
+When loading an OS, we turn on QEMU's PSCI implementation with SMC
+as the PSCI conduit. When there's no ``-kernel`` option, we assume the user
+provides EL3 firmware to handle PSCI.
+
+A few examples:
+
+Direct Linux boot of a generic ARM64 upstream Linux kernel:
+
+.. code-block:: bash
+
+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 2G \
+      -serial mon:stdio -display none \
+      -kernel arch/arm64/boot/Image \
+      -nic user -nic user \
+      -device virtio-rng-device,bus=virtio-mmio-bus.0 \
+      -drive if=none,index=0,file=hd0.qcow2,id=hd0,snapshot \
+      -drive file=qemu_sd.qcow2,if=sd,index=0,snapshot \
+      -device virtio-blk-device,drive=hd0 -append root=/dev/vda
+
+Direct Linux boot of PetaLinux 2019.2:
+
+.. code-block:: bash
+
+  $ qemu-system-aarch64  -M xlnx-versal-virt -m 2G \
+      -serial mon:stdio -display none \
+      -kernel petalinux-v2019.2/Image \
+      -append "rdinit=/sbin/init console=ttyAMA0,115200n8 earlycon=pl011,mmio,0xFF000000,115200n8" \
+      -net nic,model=cadence_gem,netdev=net0 -netdev user,id=net0 \
+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
+      -object rng-random,filename=/dev/urandom,id=rng0
+
+Boot PetaLinux 2019.2 via ARM Trusted Firmware (2018.3 because the 2019.2
+version of ATF tries to configure the CCI which we don't model) and U-boot:
+
+.. code-block:: bash
+
+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 2G \
+      -serial stdio -display none \
+      -device loader,file=petalinux-v2018.3/bl31.elf,cpu-num=0 \
+      -device loader,file=petalinux-v2019.2/u-boot.elf \
+      -device loader,addr=0x20000000,file=petalinux-v2019.2/Image \
+      -nic user -nic user \
+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
+      -object rng-random,filename=/dev/urandom,id=rng0
+
+Run the following at the U-Boot prompt:
+
+.. code-block:: bash
+
+  Versal>
+  fdt addr $fdtcontroladdr
+  fdt move $fdtcontroladdr 0x40000000
+  fdt set /timer clock-frequency <0x3dfd240>
+  setenv bootargs "rdinit=/sbin/init maxcpus=1 console=ttyAMA0,115200n8 earlycon=pl011,mmio,0xFF000000,115200n8"
+  booti 20000000 - 40000000
+  fdt addr $fdtcontroladdr
+
+Boot Linux as DOM0 on Xen via U-Boot:
+
+.. code-block:: bash
+
+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 4G \
+      -serial stdio -display none \
+      -device loader,file=petalinux-v2019.2/u-boot.elf,cpu-num=0 \
+      -device loader,addr=0x30000000,file=linux/2018-04-24/xen \
+      -device loader,addr=0x40000000,file=petalinux-v2019.2/Image \
+      -nic user -nic user \
+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
+      -object rng-random,filename=/dev/urandom,id=rng0
+
+Run the following at the U-Boot prompt:
+
+.. code-block:: bash
+
+  Versal>
+  fdt addr $fdtcontroladdr
+  fdt move $fdtcontroladdr 0x20000000
+  fdt set /timer clock-frequency <0x3dfd240>
+  fdt set /chosen xen,xen-bootargs "console=dtuart dtuart=/uart@ff000000 dom0_mem=640M bootscrub=0 maxcpus=1 timer_slop=0"
+  fdt set /chosen xen,dom0-bootargs "rdinit=/sbin/init clk_ignore_unused console=hvc0 maxcpus=1"
+  fdt mknode /chosen dom0
+  fdt set /chosen/dom0 compatible "xen,multiboot-module"
+  fdt set /chosen/dom0 reg <0x00000000 0x40000000 0x0 0x03100000>
+  booti 30000000 - 20000000
+
+Boot Linux as Dom0 on Xen via ARM Trusted Firmware and U-Boot:
+
+.. code-block:: bash
+
+  $ qemu-system-aarch64 -M xlnx-versal-virt -m 4G \
+      -serial stdio -display none \
+      -device loader,file=petalinux-v2018.3/bl31.elf,cpu-num=0 \
+      -device loader,file=petalinux-v2019.2/u-boot.elf \
+      -device loader,addr=0x30000000,file=linux/2018-04-24/xen \
+      -device loader,addr=0x40000000,file=petalinux-v2019.2/Image \
+      -nic user -nic user \
+      -device virtio-rng-device,bus=virtio-mmio-bus.0,rng=rng0 \
+      -object rng-random,filename=/dev/urandom,id=rng0
+
+Run the following at the U-Boot prompt:
+
+.. code-block:: bash
+
+  Versal>
+  fdt addr $fdtcontroladdr
+  fdt move $fdtcontroladdr 0x20000000
+  fdt set /timer clock-frequency <0x3dfd240>
+  fdt set /chosen xen,xen-bootargs "console=dtuart dtuart=/uart@ff000000 dom0_mem=640M bootscrub=0 maxcpus=1 timer_slop=0"
+  fdt set /chosen xen,dom0-bootargs "rdinit=/sbin/init clk_ignore_unused console=hvc0 maxcpus=1"
+  fdt mknode /chosen dom0
+  fdt set /chosen/dom0 compatible "xen,multiboot-module"
+  fdt set /chosen/dom0 reg <0x00000000 0x40000000 0x0 0x03100000>
+  booti 30000000 - 20000000
+
diff --git a/docs/system/target-arm.rst b/docs/system/target-arm.rst
index XXXXXXX..XXXXXXX 100644
--- a/docs/system/target-arm.rst
+++ b/docs/system/target-arm.rst
@@ -XXX,XX +XXX,XX @@ undocumented; you can get a complete list by running
    arm/sx1
    arm/stellaris
    arm/virt
+   arm/xlnx-versal-virt
 
 Arm CPU features
 ================
diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: hw/misc/zynq*
 F: include/hw/misc/zynq*
 X: hw/ssi/xilinx_*
 
-Xilinx ZynqMP
+Xilinx ZynqMP and Versal
 M: Alistair Francis <alistair@alistair23.me>
 M: Edgar E. Iglesias <edgar.iglesias@gmail.com>
 M: Peter Maydell <peter.maydell@linaro.org>
@@ -XXX,XX +XXX,XX @@ F: include/hw/*/xlnx*.h
 F: include/hw/ssi/xilinx_spips.h
 F: hw/display/dpcd.c
 F: include/hw/display/dpcd.h
+F: docs/system/arm/xlnx-versal-virt.rst
 
 ARM ACPI Subsystem
 M: Shannon Zhao <shannon.zhaosl@gmail.com>
-- 
2.20.1

At the moment we check for XScale/iwMMXt insns inside
disas_coproc_insn(): for CPUs with ARM_FEATURE_XSCALE all copro insns
with cp 0 or 1 are handled specially.  This works, but is an odd
place for this check, because disas_coproc_insn() is called from both
the Arm and Thumb decoders but the XScale case never applies for
Thumb (all the XScale CPUs were ARMv5, which has only Thumb1, not
Thumb2 with the 32-bit coprocessor insn encodings).  It also makes it
awkward to convert the real copro access insns to decodetree.

Move the identification of XScale out to its own function
which is only called from disas_arm_insn().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-2-peter.maydell@linaro.org
---
 target/arm/translate.c | 44 ++++++++++++++++++++++++++++--------------
 1 file changed, 29 insertions(+), 15 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
 
     cpnum = (insn >> 8) & 0xf;
 
-    /* First check for coprocessor space used for XScale/iwMMXt insns */
-    if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
-        if (extract32(s->c15_cpar, cpnum, 1) == 0) {
-            return 1;
-        }
-        if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
-            return disas_iwmmxt_insn(s, insn);
-        } else if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
-            return disas_dsp_insn(s, insn);
-        }
-        return 1;
-    }
-
-    /* Otherwise treat as a generic register access */
     is64 = (insn & (1 << 25)) == 0;
     if (!is64 && ((insn & (1 << 4)) == 0)) {
         /* cdp */
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
     return 1;
 }
 
+/* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
+static void disas_xscale_insn(DisasContext *s, uint32_t insn)
+{
+    int cpnum = (insn >> 8) & 0xf;
+
+    if (extract32(s->c15_cpar, cpnum, 1) == 0) {
+        unallocated_encoding(s);
+    } else if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
+        if (disas_iwmmxt_insn(s, insn)) {
+            unallocated_encoding(s);
+        }
+    } else if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
+        if (disas_dsp_insn(s, insn)) {
+            unallocated_encoding(s);
+        }
+    }
+}
 
 /* Store a 64-bit value to a register pair.  Clobbers val.  */
 static void gen_storeq_reg(DisasContext *s, int rlow, int rhigh, TCGv_i64 val)
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
     case 0xc:
     case 0xd:
     case 0xe:
-        if (((insn >> 8) & 0xe) == 10) {
+    {
+        /* First check for coprocessor space used for XScale/iwMMXt insns */
+        int cpnum = (insn >> 8) & 0xf;
+
+        if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
+            disas_xscale_insn(s, insn);
+            break;
+        }
+
+        if ((cpnum & 0xe) == 10) {
             /* VFP, but failed disas_vfp.  */
             goto illegal_op;
         }
+
         if (disas_coproc_insn(s, insn)) {
             /* Coprocessor.  */
             goto illegal_op;
         }
         break;
+    }
     default:
     illegal_op:
         unallocated_encoding(s);
-- 
2.20.1

As a prelude to making coproc insns use decodetree, split out the
part of disas_coproc_insn() which does instruction decoding from the
part which does the actual work, and make do_coproc_insn() handle the
UNDEF-on-bad-permissions and similar cases itself rather than
returning 1 to eventually percolate up to a callsite that calls
unallocated_encoding() for it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-3-peter.maydell@linaro.org
---
 target/arm/translate.c | 76 ++++++++++++++++++++++++------------------
 1 file changed, 44 insertions(+), 32 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ void gen_gvec_uaba(unsigned vece, uint32_t rd_ofs, uint32_t rn_ofs,
     tcg_gen_gvec_3(rd_ofs, rn_ofs, rm_ofs, opr_sz, max_sz, &ops[vece]);
 }
 
-static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+static void do_coproc_insn(DisasContext *s, int cpnum, int is64,
+                           int opc1, int crn, int crm, int opc2,
+                           bool isread, int rt, int rt2)
 {
-    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
     const ARMCPRegInfo *ri;
 
-    cpnum = (insn >> 8) & 0xf;
-
-    is64 = (insn & (1 << 25)) == 0;
-    if (!is64 && ((insn & (1 << 4)) == 0)) {
-        /* cdp */
-        return 1;
-    }
-
-    crm = insn & 0xf;
-    if (is64) {
-        crn = 0;
-        opc1 = (insn >> 4) & 0xf;
-        opc2 = 0;
-        rt2 = (insn >> 16) & 0xf;
-    } else {
-        crn = (insn >> 16) & 0xf;
-        opc1 = (insn >> 21) & 7;
-        opc2 = (insn >> 5) & 7;
-        rt2 = 0;
-    }
-    isread = (insn >> 20) & 1;
-    rt = (insn >> 12) & 0xf;
-
     ri = get_arm_cp_reginfo(s->cp_regs,
             ENCODE_CP_REG(cpnum, is64, s->ns, crn, crm, opc1, opc2));
     if (ri) {
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
 
         /* Check access permissions */
         if (!cp_access_ok(s->current_el, ri, isread)) {
-            return 1;
+            unallocated_encoding(s);
+            return;
         }
 
         if (s->hstr_active || ri->accessfn ||
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
         /* Handle special cases first */
         switch (ri->type & ~(ARM_CP_FLAG_MASK & ~ARM_CP_SPECIAL)) {
         case ARM_CP_NOP:
-            return 0;
+            return;
         case ARM_CP_WFI:
             if (isread) {
-                return 1;
+                unallocated_encoding(s);
+                return;
             }
             gen_set_pc_im(s, s->base.pc_next);
             s->base.is_jmp = DISAS_WFI;
-            return 0;
+            return;
         default:
             break;
         }
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
             /* Write */
             if (ri->type & ARM_CP_CONST) {
                 /* If not forbidden by access permissions, treat as WI */
-                return 0;
+                return;
             }
 
             if (is64) {
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
             gen_lookup_tb(s);
         }
 
-        return 0;
+        return;
     }
 
     /* Unknown register; this might be a guest error or a QEMU
@@ -XXX,XX +XXX,XX @@ static int disas_coproc_insn(DisasContext *s, uint32_t insn)
                       s->ns ? "non-secure" : "secure");
     }
 
-    return 1;
+    unallocated_encoding(s);
+    return;
+}
+
+static int disas_coproc_insn(DisasContext *s, uint32_t insn)
+{
+    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
+
+    cpnum = (insn >> 8) & 0xf;
+
+    is64 = (insn & (1 << 25)) == 0;
+    if (!is64 && ((insn & (1 << 4)) == 0)) {
+        /* cdp */
+        return 1;
+    }
+
+    crm = insn & 0xf;
+    if (is64) {
+        crn = 0;
+        opc1 = (insn >> 4) & 0xf;
+        opc2 = 0;
+        rt2 = (insn >> 16) & 0xf;
+    } else {
+        crn = (insn >> 16) & 0xf;
+        opc1 = (insn >> 21) & 7;
+        opc2 = (insn >> 5) & 7;
+        rt2 = 0;
+    }
+    isread = (insn >> 20) & 1;
+    rt = (insn >> 12) & 0xf;
+
+    do_coproc_insn(s, cpnum, is64, opc1, crn, crm, opc2, isread, rt, rt2);
+    return 0;
 }
 
 /* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
-- 
2.20.1

Convert the A32 coprocessor instructions to decodetree.

Note that this corrects an underdecoding: for the 64-bit access case
(MRRC/MCRR) we did not check that bits [24:21] were 0b0010, so we
would incorrectly treat LDC/STC as MRRC/MCRR rather than UNDEFing
them.

The decodetree versions of these insns assume the coprocessor
is in the range 0..7 or 14..15. This is architecturally sensible
(as per the comments) and OK in practice for QEMU because the only
uses of the ARMCPRegInfo infrastructure we have that aren't
for coprocessors 14 or 15 are the pxa2xx use of coprocessor 6.
We add an assertion to the define_one_arm_cp_reg_with_opaque()
function to catch any accidental future attempts to use it to
define coprocessor registers for invalid coprocessors.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-4-peter.maydell@linaro.org
---
 target/arm/a32.decode  | 19 +++++++++++
 target/arm/helper.c    | 29 +++++++++++++++++
 target/arm/translate.c | 74 +++++++++++++++++++++++++++++++++++-------
 3 files changed, 111 insertions(+), 11 deletions(-)

diff --git a/target/arm/a32.decode b/target/arm/a32.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/a32.decode
+++ b/target/arm/a32.decode
@@ -XXX,XX +XXX,XX @@
 &bfi             rd rn lsb msb
 &sat             rd rn satimm imm sh
 &pkh             rd rn rm imm tb
+&mcr             cp opc1 crn crm opc2 rt
+&mcrr            cp opc1 crm rt rt2
 
 # Data-processing (register)
 
@@ -XXX,XX +XXX,XX @@ LDM_a32          ---- 100 b:1 i:1 u:1 w:1 1 rn:4 list:16   &ldst_block
 B                .... 1010 ........................           @branch
 BL               .... 1011 ........................           @branch
 
+# Coprocessor instructions
+
+# We decode MCR, MCR, MRRC and MCRR only, because for QEMU the
+# other coprocessor instructions always UNDEF.
+# The trans_ functions for these will ignore cp values 8..13 for v7 or
+# earlier, and 0..13 for v8 and later, because those areas of the
+# encoding space may be used for other things, such as VFP or Neon.
+
+@mcr             ---- .... opc1:3 . crn:4 rt:4 cp:4 opc2:3 . crm:4 &mcr
+@mcrr            ---- .... .... rt2:4 rt:4 cp:4 opc1:4 crm:4       &mcrr
+
+MCRR             .... 1100 0100 .... .... .... .... .... @mcrr
+MRRC             .... 1100 0101 .... .... .... .... .... @mcrr
+
+MCR              .... 1110 ... 0 .... .... .... ... 1 .... @mcr
+MRC              .... 1110 ... 1 .... .... .... ... 1 .... @mcr
+
 # Supervisor call
 
 SVC              ---- 1111 imm:24                             &i
diff --git a/target/arm/helper.c b/target/arm/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/helper.c
+++ b/target/arm/helper.c
@@ -XXX,XX +XXX,XX @@ void define_one_arm_cp_reg_with_opaque(ARMCPU *cpu,
     assert((r->state != ARM_CP_STATE_AA32) || (r->opc0 == 0));
     /* AArch64 regs are all 64 bit so ARM_CP_64BIT is meaningless */
     assert((r->state != ARM_CP_STATE_AA64) || !(r->type & ARM_CP_64BIT));
+    /*
+     * This API is only for Arm's system coprocessors (14 and 15) or
+     * (M-profile or v7A-and-earlier only) for implementation defined
+     * coprocessors in the range 0..7.  Our decode assumes this, since
+     * 8..13 can be used for other insns including VFP and Neon. See
+     * valid_cp() in translate.c.  Assert here that we haven't tried
+     * to use an invalid coprocessor number.
+     */
+    switch (r->state) {
+    case ARM_CP_STATE_BOTH:
+        /* 0 has a special meaning, but otherwise the same rules as AA32. */
+        if (r->cp == 0) {
+            break;
+        }
+        /* fall through */
+    case ARM_CP_STATE_AA32:
+        if (arm_feature(&cpu->env, ARM_FEATURE_V8) &&
+            !arm_feature(&cpu->env, ARM_FEATURE_M)) {
+            assert(r->cp >= 14 && r->cp <= 15);
+        } else {
+            assert(r->cp < 8 || (r->cp >= 14 && r->cp <= 15));
+        }
+        break;
+    case ARM_CP_STATE_AA64:
+        assert(r->cp == 0 || r->cp == CP_REG_ARM64_SYSREG_CP);
+        break;
+    default:
+        g_assert_not_reached();
+    }
     /* The AArch64 pseudocode CheckSystemAccess() specifies that op1
      * encodes a minimum access level for the register. We roll this
      * runtime check into our general permission check code, so check
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int t16_pop_list(DisasContext *s, int x)
 #include "decode-t32.c.inc"
 #include "decode-t16.c.inc"
 
+static bool valid_cp(DisasContext *s, int cp)
+{
+    /*
+     * Return true if this coprocessor field indicates something
+     * that's really a possible coprocessor.
+     * For v7 and earlier, coprocessors 8..15 were reserved for Arm use,
+     * and of those only cp14 and cp15 were used for registers.
+     * cp10 and cp11 were used for VFP and Neon, whose decode is
+     * dealt with elsewhere. With the advent of fp16, cp9 is also
+     * now part of VFP.
+     * For v8A and later, the encoding has been tightened so that
+     * only cp14 and cp15 are valid, and other values aren't considered
+     * to be in the coprocessor-instruction space at all. v8M still
+     * permits coprocessors 0..7.
+     */
+    if (arm_dc_feature(s, ARM_FEATURE_V8) &&
+        !arm_dc_feature(s, ARM_FEATURE_M)) {
+        return cp >= 14;
+    }
+    return cp < 8 || cp >= 14;
+}
+
+static bool trans_MCR(DisasContext *s, arg_MCR *a)
+{
+    if (!valid_cp(s, a->cp)) {
+        return false;
+    }
+    do_coproc_insn(s, a->cp, false, a->opc1, a->crn, a->crm, a->opc2,
+                   false, a->rt, 0);
+    return true;
+}
+
+static bool trans_MRC(DisasContext *s, arg_MRC *a)
+{
+    if (!valid_cp(s, a->cp)) {
+        return false;
+    }
+    do_coproc_insn(s, a->cp, false, a->opc1, a->crn, a->crm, a->opc2,
+                   true, a->rt, 0);
+    return true;
+}
+
+static bool trans_MCRR(DisasContext *s, arg_MCRR *a)
+{
+    if (!valid_cp(s, a->cp)) {
+        return false;
+    }
+    do_coproc_insn(s, a->cp, true, a->opc1, 0, a->crm, 0,
+                   false, a->rt, a->rt2);
+    return true;
+}
+
+static bool trans_MRRC(DisasContext *s, arg_MRRC *a)
+{
+    if (!valid_cp(s, a->cp)) {
+        return false;
+    }
+    do_coproc_insn(s, a->cp, true, a->opc1, 0, a->crm, 0,
+                   true, a->rt, a->rt2);
+    return true;
+}
+
 /* Helpers to swap operands for reverse-subtract.  */
 static void gen_rsb(TCGv_i32 dst, TCGv_i32 a, TCGv_i32 b)
 {
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
             disas_xscale_insn(s, insn);
             break;
         }
-
-        if ((cpnum & 0xe) == 10) {
-            /* VFP, but failed disas_vfp.  */
-            goto illegal_op;
-        }
-
-        if (disas_coproc_insn(s, insn)) {
-            /* Coprocessor.  */
-            goto illegal_op;
-        }
-        break;
+        /* fall through */
     }
     default:
     illegal_op:
-- 
2.20.1

The only thing left in the "legacy decoder" is the handling
of disas_xscale_insn(), and we can simplify the code.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-5-peter.maydell@linaro.org
---
 target/arm/translate.c | 26 +++++++++-----------------
 1 file changed, 9 insertions(+), 17 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
         return;
     }
     /* fall back to legacy decoder */
-
-    switch ((insn >> 24) & 0xf) {
-    case 0xc:
-    case 0xd:
-    case 0xe:
-    {
-        /* First check for coprocessor space used for XScale/iwMMXt insns */
-        int cpnum = (insn >> 8) & 0xf;
-
-        if (arm_dc_feature(s, ARM_FEATURE_XSCALE) && (cpnum < 2)) {
+    /* TODO: convert xscale/iwmmxt decoder to decodetree ?? */
+    if (arm_dc_feature(s, ARM_FEATURE_XSCALE)) {
+        if (((insn & 0x0c000e00) == 0x0c000000)
+            && ((insn & 0x03000000) != 0x03000000)) {
+            /* Coprocessor insn, coprocessor 0 or 1 */
             disas_xscale_insn(s, insn);
-            break;
+            return;
         }
-        /* fall through */
-    }
-    default:
-    illegal_op:
-        unallocated_encoding(s);
-        break;
     }
+
+illegal_op:
+    unallocated_encoding(s);
 }
 
 static bool thumb_insn_is_16bit(DisasContext *s, uint32_t pc, uint32_t insn)
-- 
2.20.1

For M-profile CPUs, the architecture specifies that the NOCP
exception when a coprocessor is not present or disabled should cover
the entire wide range of coprocessor-space encodings, and should take
precedence over UNDEF exceptions.  (This is the opposite of
A-profile, where checking for a disabled FPU has to happen last.)

Implement this with decodetree patterns that cover the specified
ranges of the encoding space.  There are a few instructions (VLLDM,
VLSTM, and in v8.1 also VSCCLRM) which are in copro-space but must
not be NOCP'd: these must be handled also in the new m-nocp.decode so
they take precedence.

This is a minor behaviour change: for unallocated insn patterns in
the VFP area (cp=10,11) we will now NOCP rather than UNDEF when the
FPU is disabled.

As well as giving us the correct architectural behaviour for v8.1M
and the recommended behaviour for v8.0M, this refactoring also
removes the old NOCP handling from the remains of the 'legacy
decoder' in disas_thumb2_insn(), paving the way for cleaning that up.

Since we don't currently have a v8.1M feature bit or any v8.1M CPUs,
the minor changes to this logic that we'll need for v8.1M are marked
up with TODO comments.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-6-peter.maydell@linaro.org
---
 target/arm/m-nocp.decode       | 42 +++++++++++++++++++++++++++
 target/arm/vfp.decode          |  2 --
 target/arm/translate.c         | 30 ++++++++++----------
 target/arm/meson.build         |  1 +
 target/arm/translate-vfp.c.inc | 52 +++++++++++++++++++++++++++-------
 5 files changed, 100 insertions(+), 27 deletions(-)
 create mode 100644 target/arm/m-nocp.decode

diff --git a/target/arm/m-nocp.decode b/target/arm/m-nocp.decode
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/target/arm/m-nocp.decode
@@ -XXX,XX +XXX,XX @@
+# M-profile UserFault.NOCP exception handling
+#
+#  Copyright (c) 2020 Linaro, Ltd
+#
+# This library is free software; you can redistribute it and/or
+# modify it under the terms of the GNU Lesser General Public
+# License as published by the Free Software Foundation; either
+# version 2.1 of the License, or (at your option) any later version.
+#
+# This library is distributed in the hope that it will be useful,
+# but WITHOUT ANY WARRANTY; without even the implied warranty of
+# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the GNU
+# Lesser General Public License for more details.
+#
+# You should have received a copy of the GNU Lesser General Public
+# License along with this library; if not, see <http://www.gnu.org/licenses/>.
+
+#
+# This file is processed by scripts/decodetree.py
+#
+# For M-profile, the architecture specifies that NOCP UsageFaults
+# should take precedence over UNDEF faults over the whole wide
+# range of coprocessor-space encodings, with the exception of
+# VLLDM and VLSTM. (Compare v8.1M IsCPInstruction() pseudocode and
+# v8M Arm ARM rule R_QLGM.) This isn't mandatory for v8.0M but we choose
+# to behave the same as v8.1M.
+# This decode is handled before any others (and in particular before
+# decoding FP instructions which are in the coprocessor space).
+# If the coprocessor is not present or disabled then we will generate
+# the NOCP exception; otherwise we let the insn through to the main decode.
+
+{
+  # Special cases which do not take an early NOCP: VLLDM and VLSTM
+  VLLDM_VLSTM  1110 1100 001 l:1 rn:4 0000 1010 0000 0000
+  # TODO: VSCCLRM (new in v8.1M) is similar:
+  #VSCCLRM      1110 1100 1-01 1111 ---- 1011 ---- ---0
+
+  NOCP         111- 1110 ---- ---- ---- cp:4 ---- ----
+  NOCP         111- 110- ---- ---- ---- cp:4 ---- ----
+  # TODO: From v8.1M onwards we will also want this range to NOCP
+  #NOCP_8_1     111- 1111 ---- ---- ---- ---- ---- ---- cp=10
+}
diff --git a/target/arm/vfp.decode b/target/arm/vfp.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp.decode
+++ b/target/arm/vfp.decode
@@ -XXX,XX +XXX,XX @@ VCVT_sp_int  ---- 1110 1.11 110 s:1 .... 1010 rz:1 1.0 .... \
              vd=%vd_sp vm=%vm_sp
 VCVT_dp_int  ---- 1110 1.11 110 s:1 .... 1011 rz:1 1.0 .... \
              vd=%vd_sp vm=%vm_dp
-
-VLLDM_VLSTM  1110 1100 001 l:1 rn:4 0000 1010 0000 0000
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static TCGv_ptr vfp_reg_ptr(bool dp, int reg)
 #define ARM_CP_RW_BIT   (1 << 20)
 
 /* Include the VFP and Neon decoders */
+#include "decode-m-nocp.c.inc"
 #include "translate-vfp.c.inc"
 #include "translate-neon.c.inc"
 
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
         ARCH(6T2);
     }
 
+    if (arm_dc_feature(s, ARM_FEATURE_M)) {
+        /*
+         * NOCP takes precedence over any UNDEF for (almost) the
+         * entire wide range of coprocessor-space encodings, so check
+         * for it first before proceeding to actually decode eg VFP
+         * insns. This decode also handles the few insns which are
+         * in copro space but do not have NOCP checks (eg VLLDM, VLSTM).
+         */
+        if (disas_m_nocp(s, insn)) {
+            return;
+        }
+    }
+
     if ((insn & 0xef000000) == 0xef000000) {
         /*
          * T32 encodings 0b111p_1111_qqqq_qqqq_qqqq_qqqq_qqqq_qqqq
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
         /* Coprocessor.  */
         if (arm_dc_feature(s, ARM_FEATURE_M)) {
             /* 0b111x_11xx_xxxx_xxxx_xxxx_xxxx_xxxx_xxxx */
-            if (extract32(insn, 24, 2) == 3) {
-                goto illegal_op; /* op0 = 0b11 : unallocated */
-            }
-
-            if (((insn >> 8) & 0xe) == 10 &&
-                dc_isar_feature(aa32_fpsp_v2, s)) {
-                /* FP, and the CPU supports it */
-                goto illegal_op;
-            } else {
-                /* All other insns: NOCP */
-                gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
-                                   syn_uncategorized(),
-                                   default_exception_el(s));
-            }
-            break;
+            goto illegal_op;
         }
         if (((insn >> 24) & 3) == 3) {
             /* Neon DP, but failed disas_neon_dp() */
diff --git a/target/arm/meson.build b/target/arm/meson.build
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/meson.build
+++ b/target/arm/meson.build
@@ -XXX,XX +XXX,XX @@ gen = [
   decodetree.process('neon-ls.decode', extra_args: '--static-decode=disas_neon_ls'),
   decodetree.process('vfp.decode', extra_args: '--static-decode=disas_vfp'),
   decodetree.process('vfp-uncond.decode', extra_args: '--static-decode=disas_vfp_uncond'),
+  decodetree.process('m-nocp.decode', extra_args: '--static-decode=disas_m_nocp'),
   decodetree.process('a32.decode', extra_args: '--static-decode=disas_a32'),
   decodetree.process('a32-uncond.decode', extra_args: '--static-decode=disas_a32_uncond'),
   decodetree.process('t32.decode', extra_args: '--static-decode=disas_t32'),
diff --git a/target/arm/translate-vfp.c.inc b/target/arm/translate-vfp.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-vfp.c.inc
+++ b/target/arm/translate-vfp.c.inc
@@ -XXX,XX +XXX,XX @@ static inline long vfp_f16_offset(unsigned reg, bool top)
 static bool full_vfp_access_check(DisasContext *s, bool ignore_vfp_enabled)
 {
     if (s->fp_excp_el) {
-        if (arm_dc_feature(s, ARM_FEATURE_M)) {
-            gen_exception_insn(s, s->pc_curr, EXCP_NOCP, syn_uncategorized(),
-                               s->fp_excp_el);
-        } else {
-            gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
-                               syn_fp_access_trap(1, 0xe, false),
-                               s->fp_excp_el);
-        }
+        /* M-profile handled this earlier, in disas_m_nocp() */
+        assert (!arm_dc_feature(s, ARM_FEATURE_M));
+        gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
+                           syn_fp_access_trap(1, 0xe, false),
+                           s->fp_excp_el);
         return false;
     }
 
@@ -XXX,XX +XXX,XX @@ static bool trans_VLLDM_VLSTM(DisasContext *s, arg_VLLDM_VLSTM *a)
         !arm_dc_feature(s, ARM_FEATURE_V8)) {
         return false;
     }
-    /* If not secure, UNDEF. */
+    /*
+     * If not secure, UNDEF. We must emit code for this
+     * rather than returning false so that this takes
+     * precedence over the m-nocp.decode NOCP fallback.
+     */
     if (!s->v8m_secure) {
-        return false;
+        unallocated_encoding(s);
+        return true;
     }
     /* If no fpu, NOP. */
     if (!dc_isar_feature(aa32_vfp, s)) {
@@ -XXX,XX +XXX,XX @@ static bool trans_VLLDM_VLSTM(DisasContext *s, arg_VLLDM_VLSTM *a)
     s->base.is_jmp = DISAS_UPDATE_EXIT;
     return true;
 }
+
+static bool trans_NOCP(DisasContext *s, arg_NOCP *a)
+{
+    /*
+     * Handle M-profile early check for disabled coprocessor:
+     * all we need to do here is emit the NOCP exception if
+     * the coprocessor is disabled. Otherwise we return false
+     * and the real VFP/etc decode will handle the insn.
+     */
+    assert(arm_dc_feature(s, ARM_FEATURE_M));
+
+    if (a->cp == 11) {
+        a->cp = 10;
+    }
+    /* TODO: in v8.1M cp 8, 9, 14, 15 also are governed by the cp10 enable */
+
+    if (a->cp != 10) {
+        gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
+                           syn_uncategorized(), default_exception_el(s));
+        return true;
+    }
+
+    if (s->fp_excp_el != 0) {
+        gen_exception_insn(s, s->pc_curr, EXCP_NOCP,
+                           syn_uncategorized(), s->fp_excp_el);
+        return true;
+    }
+
+    return false;
+}
-- 
2.20.1

Convert the T32 coprocessor instructions to decodetree.
As with the A32 conversion, this corrects an underdecoding
where we did not check that MRRC/MCRR [24:21] were 0b0010
and so treated some kinds of LDC/STC and MRRC/MCRR rather
than UNDEFing them.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-7-peter.maydell@linaro.org
---
 target/arm/t32.decode  | 19 +++++++++++++
 target/arm/translate.c | 64 ++----------------------------------------
 2 files changed, 21 insertions(+), 62 deletions(-)

diff --git a/target/arm/t32.decode b/target/arm/t32.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/t32.decode
+++ b/target/arm/t32.decode
@@ -XXX,XX +XXX,XX @@
 &sat             !extern rd rn satimm imm sh
 &pkh             !extern rd rn rm imm tb
 &cps             !extern mode imod M A I F
+&mcr             !extern cp opc1 crn crm opc2 rt
+&mcrr            !extern cp opc1 crm rt rt2
 
 # Data-processing (register)
 
@@ -XXX,XX +XXX,XX @@ RFE              1110 1001 10.1 .... 1100000000000000         @rfe pu=1
 SRS              1110 1000 00.0 1101 1100 0000 000. ....      @srs pu=2
 SRS              1110 1001 10.0 1101 1100 0000 000. ....      @srs pu=1
 
+# Coprocessor instructions
+
+# We decode MCR, MCR, MRRC and MCRR only, because for QEMU the
+# other coprocessor instructions always UNDEF.
+# The trans_ functions for these will ignore cp values 8..13 for v7 or
+# earlier, and 0..13 for v8 and later, because those areas of the
+# encoding space may be used for other things, such as VFP or Neon.
+
+@mcr             .... .... opc1:3 . crn:4 rt:4 cp:4 opc2:3 . crm:4
+@mcrr            .... .... .... rt2:4 rt:4 cp:4 opc1:4 crm:4
+
+MCRR             1110 1100 0100 .... .... .... .... .... @mcrr
+MRRC             1110 1100 0101 .... .... .... .... .... @mcrr
+
+MCR              1110 1110 ... 0 .... .... .... ... 1 .... @mcr
+MRC              1110 1110 ... 1 .... .... .... ... 1 .... @mcr
+
 # Branches
 
 %imm24           26:s1 13:1 11:1 16:10 0:11 !function=t32_branch24
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void do_coproc_insn(DisasContext *s, int cpnum, int is64,
     return;
 }
 
-static int disas_coproc_insn(DisasContext *s, uint32_t insn)
-{
-    int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
-
-    cpnum = (insn >> 8) & 0xf;
-
-    is64 = (insn & (1 << 25)) == 0;
-    if (!is64 && ((insn & (1 << 4)) == 0)) {
-        /* cdp */
-        return 1;
-    }
-
-    crm = insn & 0xf;
-    if (is64) {
-        crn = 0;
-        opc1 = (insn >> 4) & 0xf;
-        opc2 = 0;
-        rt2 = (insn >> 16) & 0xf;
-    } else {
-        crn = (insn >> 16) & 0xf;
-        opc1 = (insn >> 21) & 7;
-        opc2 = (insn >> 5) & 7;
-        rt2 = 0;
-    }
-    isread = (insn >> 20) & 1;
-    rt = (insn >> 12) & 0xf;
-
-    do_coproc_insn(s, cpnum, is64, opc1, crn, crm, opc2, isread, rt, rt2);
-    return 0;
-}
-
 /* Decode XScale DSP or iWMMXt insn (in the copro space, cp=0 or 1) */
 static void disas_xscale_insn(DisasContext *s, uint32_t insn)
 {
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
         ((insn >> 28) == 0xe && disas_vfp(s, insn))) {
         return;
     }
-    /* fall back to legacy decoder */
 
-    switch ((insn >> 25) & 0xf) {
-    case 0: case 1: case 2: case 3:
-        /* 16-bit instructions.  Should never happen.  */
-        abort();
-    case 6: case 7: case 14: case 15:
-        /* Coprocessor.  */
-        if (arm_dc_feature(s, ARM_FEATURE_M)) {
-            /* 0b111x_11xx_xxxx_xxxx_xxxx_xxxx_xxxx_xxxx */
-            goto illegal_op;
-        }
-        if (((insn >> 24) & 3) == 3) {
-            /* Neon DP, but failed disas_neon_dp() */
-            goto illegal_op;
-        } else if (((insn >> 8) & 0xe) == 10) {
-            /* VFP, but failed disas_vfp.  */
-            goto illegal_op;
-        } else {
-            if (insn & (1 << 28))
-                goto illegal_op;
-            if (disas_coproc_insn(s, insn)) {
-                goto illegal_op;
-            }
-        }
-        break;
-    case 12:
-        goto illegal_op;
-    default:
-    illegal_op:
-        unallocated_encoding(s);
-    }
+illegal_op:
+    unallocated_encoding(s);
 }
 
 static void disas_thumb_insn(DisasContext *s, uint32_t insn)
-- 
2.20.1

The ARCH() macro was used a lot in the legacy decoder, but
there are now just two uses of it left. Since a macro which
expands out to a goto is liable to be confusing when reading
code, replace the last two uses with a simple open-coded
qeuivalent.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803111849.13368-8-peter.maydell@linaro.org
---
 target/arm/translate.c | 14 +++++++++-----
 1 file changed, 9 insertions(+), 5 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@
 #define ENABLE_ARCH_7     arm_dc_feature(s, ARM_FEATURE_V7)
 #define ENABLE_ARCH_8     arm_dc_feature(s, ARM_FEATURE_V8)
 
-#define ARCH(x) do { if (!ENABLE_ARCH_##x) goto illegal_op; } while(0)
-
 #include "translate.h"
 
 #if defined(CONFIG_USER_ONLY)
@@ -XXX,XX +XXX,XX @@ static bool trans_BLX_i(DisasContext *s, arg_BLX_i *a)
 {
     TCGv_i32 tmp;
 
-    /* For A32, ARCH(5) is checked near the start of the uncond block. */
+    /* For A32, ARM_FEATURE_V5 is checked near the start of the uncond block. */
     if (s->thumb && (a->imm & 2)) {
         return false;
     }
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
          * choose to UNDEF. In ARMv5 and above the space is used
          * for miscellaneous unconditional instructions.
          */
-        ARCH(5);
+        if (!arm_dc_feature(s, ARM_FEATURE_V5)) {
+            unallocated_encoding(s);
+            return;
+        }
 
         /* Unconditional instructions.  */
         /* TODO: Perhaps merge these into one decodetree output file.  */
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
             goto illegal_op;
         }
     } else if ((insn & 0xf800e800) != 0xf000e800)  {
-        ARCH(6T2);
+        if (!arm_dc_feature(s, ARM_FEATURE_THUMB2)) {
+            unallocated_encoding(s);
+            return;
+        }
     }
 
     if (arm_dc_feature(s, ARM_FEATURE_M)) {
-- 
2.20.1

As part of the Neon decodetree conversion we removed all
the uses of the VFP_DREG macros, but forgot to remove the
macro definitions. Do so now.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20200803124848.18295-1-peter.maydell@linaro.org
---
 target/arm/translate.c | 15 ---------------
 1 file changed, 15 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_dsp_insn(DisasContext *s, uint32_t insn)
     return 1;
 }
 
-#define VFP_REG_SHR(x, n) (((n) > 0) ? (x) >> (n) : (x) << -(n))
-#define VFP_DREG(reg, insn, bigbit, smallbit) do { \
-    if (dc_isar_feature(aa32_simd_r32, s)) { \
-        reg = (((insn) >> (bigbit)) & 0x0f) \
-              | (((insn) >> ((smallbit) - 4)) & 0x10); \
-    } else { \
-        if (insn & (1 << (smallbit))) \
-            return 1; \
-        reg = ((insn) >> (bigbit)) & 0x0f; \
-    }} while (0)
-
-#define VFP_DREG_D(reg, insn) VFP_DREG(reg, insn, 12, 22)
-#define VFP_DREG_N(reg, insn) VFP_DREG(reg, insn, 16,  7)
-#define VFP_DREG_M(reg, insn) VFP_DREG(reg, insn,  0,  5)
-
 static inline bool use_goto_tb(DisasContext *s, target_ulong dest)
 {
 #ifndef CONFIG_USER_ONLY
-- 
2.20.1

In arm_tr_init_disas_context() we have a FIXME comment that suggests
"cpu_M0 can probably be the same as cpu_V0".  This isn't in fact
possible: cpu_V0 is used as a temporary inside gen_iwmmxt_shift(),
and that function is called in various places where cpu_M0 contains a
live value (i.e.  between gen_op_iwmmxt_movq_M0_wRn() and
gen_op_iwmmxt_movq_wRn_M0() calls).  Remove the comment.

We also have a comment on the declarations of cpu_V0/V1/M0 which
claims they're "for efficiency".  This isn't true with modern TCG, so
replace this comment with one which notes that they're only used with
the iwmmxt decode.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200803132815.3861-1-peter.maydell@linaro.org
---
 target/arm/translate.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@
 #define IS_USER(s) (s->user)
 #endif
 
-/* We reuse the same 64-bit temporaries for efficiency.  */
+/* These are TCG temporaries used only by the legacy iwMMXt decoder */
 static TCGv_i64 cpu_V0, cpu_V1, cpu_M0;
+/* These are TCG globals which alias CPUARMState fields */
 static TCGv_i32 cpu_R[16];
 TCGv_i32 cpu_CF, cpu_NF, cpu_VF, cpu_ZF;
 TCGv_i64 cpu_exclusive_addr;
@@ -XXX,XX +XXX,XX @@ static void arm_tr_init_disas_context(DisasContextBase *dcbase, CPUState *cs)
 
     cpu_V0 = tcg_temp_new_i64();
     cpu_V1 = tcg_temp_new_i64();
-    /* FIXME: cpu_M0 can probably be the same as cpu_V0.  */
     cpu_M0 = tcg_temp_new_i64();
 }
 
-- 
2.20.1

In commit 962fcbf2efe57231a9f5df we converted the uses of the
ARM_FEATURE_CRC bit to use the aa32_crc32 isar_feature test
instead. However we forgot to remove the now-unused definition
of the feature name in the enum. Delete it now.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Message-id: 20200805210848.6688-1-peter.maydell@linaro.org
---
 target/arm/cpu.h | 1 -
 1 file changed, 1 deletion(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ enum arm_features {
     ARM_FEATURE_V8,
     ARM_FEATURE_AARCH64, /* supports 64 bit mode */
     ARM_FEATURE_CBAR, /* has cp15 CBAR */
-    ARM_FEATURE_CRC, /* ARMv8 CRC instructions */
     ARM_FEATURE_CBAR_RO, /* has cp15 CBAR and it is read-only */
     ARM_FEATURE_EL2, /* has EL2 Virtualization support */
     ARM_FEATURE_EL3, /* has EL3 Secure monitor support */
-- 
2.20.1

We currently have two versions of get_fpstatus_ptr(), which both take
an effectively boolean argument:
 * the one for A64 takes "bool is_f16" to distinguish fp16 from other ops
 * the one for A32/T32 takes "int neon" to distinguish Neon from other ops

This is confusing, and to implement ARMv8.2-FP16 the A32/T32 one will
need to make a four-way distinction between "non-Neon, FP16",
"non-Neon, single/double", "Neon, FP16" and "Neon, single/double".
The A64 version will then be a strict subset of the A32/T32 version.

To clean this all up, we want to go to a single implementation which
takes an enum argument with values FPST_FPCR, FPST_STD,
FPST_FPCR_F16, and FPST_STD_F16.  We rename the function to
fpstatus_ptr() so that unconverted code gets a compilation error
rather than silently passing the wrong thing to the new function.

This commit implements that new API, and converts A64 to use it:
 get_fpstatus_ptr(false) -> fpstatus_ptr(FPST_FPCR)
 get_fpstatus_ptr(true) -> fpstatus_ptr(FPST_FPCR_F16)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20200806104453.30393-2-peter.maydell@linaro.org
---
 target/arm/translate-a64.h |  1 -
 target/arm/translate.h     | 51 ++++++++++++++++++++++
 target/arm/translate-a64.c | 89 +++++++++++++++-----------------------
 target/arm/translate-sve.c | 34 +++++++--------
 4 files changed, 103 insertions(+), 72 deletions(-)

diff --git a/target/arm/translate-a64.h b/target/arm/translate-a64.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-a64.h
+++ b/target/arm/translate-a64.h
@@ -XXX,XX +XXX,XX @@ TCGv_i64 cpu_reg_sp(DisasContext *s, int reg);
 TCGv_i64 read_cpu_reg(DisasContext *s, int reg, int sf);
 TCGv_i64 read_cpu_reg_sp(DisasContext *s, int reg, int sf);
 void write_fp_dreg(DisasContext *s, int reg, TCGv_i64 v);
-TCGv_ptr get_fpstatus_ptr(bool);
 bool logic_imm_decode_wmask(uint64_t *result, unsigned int immn,
                             unsigned int imms, unsigned int immr);
 bool sve_access_check(DisasContext *s);
diff --git a/target/arm/translate.h b/target/arm/translate.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.h
+++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ typedef void CryptoThreeOpIntFn(TCGv_ptr, TCGv_ptr, TCGv_i32);
 typedef void CryptoThreeOpFn(TCGv_ptr, TCGv_ptr, TCGv_ptr);
 typedef void AtomicThreeOpFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGArg, MemOp);
 
+/*
+ * Enum for argument to fpstatus_ptr().
+ */
+typedef enum ARMFPStatusFlavour {
+    FPST_FPCR,
+    FPST_FPCR_F16,
+    FPST_STD,
+    FPST_STD_F16,
+} ARMFPStatusFlavour;
+
+/**
+ * fpstatus_ptr: return TCGv_ptr to the specified fp_status field
+ *
+ * We have multiple softfloat float_status fields in the Arm CPU state struct
+ * (see the comment in cpu.h for details). Return a TCGv_ptr which has
+ * been set up to point to the requested field in the CPU state struct.
+ * The options are:
+ *
+ * FPST_FPCR
+ *   for non-FP16 operations controlled by the FPCR
+ * FPST_FPCR_F16
+ *   for operations controlled by the FPCR where FPCR.FZ16 is to be used
+ * FPST_STD
+ *   for A32/T32 Neon operations using the "standard FPSCR value"
+ * FPST_STD_F16
+ *   as FPST_STD, but where FPCR.FZ16 is to be used
+ */
+static inline TCGv_ptr fpstatus_ptr(ARMFPStatusFlavour flavour)
+{
+    TCGv_ptr statusptr = tcg_temp_new_ptr();
+    int offset;
+
+    switch (flavour) {
+    case FPST_FPCR:
+        offset = offsetof(CPUARMState, vfp.fp_status);
+        break;
+    case FPST_FPCR_F16:
+        offset = offsetof(CPUARMState, vfp.fp_status_f16);
+        break;
+    case FPST_STD:
+        offset = offsetof(CPUARMState, vfp.standard_fp_status);
+        break;
+    case FPST_STD_F16:
+        /* Not yet used or implemented: fall through to assert */
+    default:
+        g_assert_not_reached();
+    }
+    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
+    return statusptr;
+}
+
 #endif /* TARGET_ARM_TRANSLATE_H */
diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-a64.c
+++ b/target/arm/translate-a64.c
@@ -XXX,XX +XXX,XX @@ static void write_fp_sreg(DisasContext *s, int reg, TCGv_i32 v)
     tcg_temp_free_i64(tmp);
 }
 
-TCGv_ptr get_fpstatus_ptr(bool is_f16)
-{
-    TCGv_ptr statusptr = tcg_temp_new_ptr();
-    int offset;
-
-    /* In A64 all instructions (both FP and Neon) use the FPCR; there
-     * is no equivalent of the A32 Neon "standard FPSCR value".
-     * However half-precision operations operate under a different
-     * FZ16 flag and use vfp.fp_status_f16 instead of vfp.fp_status.
-     */
-    if (is_f16) {
-        offset = offsetof(CPUARMState, vfp.fp_status_f16);
-    } else {
-        offset = offsetof(CPUARMState, vfp.fp_status);
-    }
-    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
-    return statusptr;
-}
-
 /* Expand a 2-operand AdvSIMD vector operation using an expander function.  */
 static void gen_gvec_fn2(DisasContext *s, bool is_q, int rd, int rn,
                          GVecGen2Fn *gvec_fn, int vece)
@@ -XXX,XX +XXX,XX @@ static void gen_gvec_op3_fpst(DisasContext *s, bool is_q, int rd, int rn,
                               int rm, bool is_fp16, int data,
                               gen_helper_gvec_3_ptr *fn)
 {
-    TCGv_ptr fpst = get_fpstatus_ptr(is_fp16);
+    TCGv_ptr fpst = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
     tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, rd),
                        vec_full_reg_offset(s, rn),
                        vec_full_reg_offset(s, rm), fpst,
@@ -XXX,XX +XXX,XX @@ static void handle_fp_compare(DisasContext *s, int size,
                               bool cmp_with_zero, bool signal_all_nans)
 {
     TCGv_i64 tcg_flags = tcg_temp_new_i64();
-    TCGv_ptr fpst = get_fpstatus_ptr(size == MO_16);
+    TCGv_ptr fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
     if (size == MO_64) {
         TCGv_i64 tcg_vn, tcg_vm;
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
         tcg_gen_xori_i32(tcg_res, tcg_op, 0x8000);
         break;
     case 0x3: /* FSQRT */
-        fpst = get_fpstatus_ptr(true);
+        fpst = fpstatus_ptr(FPST_FPCR_F16);
         gen_helper_sqrt_f16(tcg_res, tcg_op, fpst);
         break;
     case 0x8: /* FRINTN */
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
     case 0xc: /* FRINTA */
     {
         TCGv_i32 tcg_rmode = tcg_const_i32(arm_rmode_to_sf(opcode & 7));
-        fpst = get_fpstatus_ptr(true);
+        fpst = fpstatus_ptr(FPST_FPCR_F16);
 
         gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
         gen_helper_advsimd_rinth(tcg_res, tcg_op, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_half(DisasContext *s, int opcode, int rd, int rn)
         break;
     }
     case 0xe: /* FRINTX */
-        fpst = get_fpstatus_ptr(true);
+        fpst = fpstatus_ptr(FPST_FPCR_F16);
         gen_helper_advsimd_rinth_exact(tcg_res, tcg_op, fpst);
         break;
     case 0xf: /* FRINTI */
-        fpst = get_fpstatus_ptr(true);
+        fpst = fpstatus_ptr(FPST_FPCR_F16);
         gen_helper_advsimd_rinth(tcg_res, tcg_op, fpst);
         break;
     default:
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_single(DisasContext *s, int opcode, int rd, int rn)
         g_assert_not_reached();
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     if (rmode >= 0) {
         TCGv_i32 tcg_rmode = tcg_const_i32(rmode);
         gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_1src_double(DisasContext *s, int opcode, int rd, int rn)
         g_assert_not_reached();
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     if (rmode >= 0) {
         TCGv_i32 tcg_rmode = tcg_const_i32(rmode);
         gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
             /* Single to half */
             TCGv_i32 tcg_rd = tcg_temp_new_i32();
             TCGv_i32 ahp = get_ahp_flag();
-            TCGv_ptr fpst = get_fpstatus_ptr(false);
+            TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
 
             gen_helper_vfp_fcvt_f32_to_f16(tcg_rd, tcg_rn, fpst, ahp);
             /* write_fp_sreg is OK here because top half of tcg_rd is zero */
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
             /* Double to single */
             gen_helper_vfp_fcvtsd(tcg_rd, tcg_rn, cpu_env);
         } else {
-            TCGv_ptr fpst = get_fpstatus_ptr(false);
+            TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
             TCGv_i32 ahp = get_ahp_flag();
             /* Double to half */
             gen_helper_vfp_fcvt_f64_to_f16(tcg_rd, tcg_rn, fpst, ahp);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_fcvt(DisasContext *s, int opcode,
     case 0x3:
     {
         TCGv_i32 tcg_rn = read_fp_sreg(s, rn);
-        TCGv_ptr tcg_fpst = get_fpstatus_ptr(false);
+        TCGv_ptr tcg_fpst = fpstatus_ptr(FPST_FPCR);
         TCGv_i32 tcg_ahp = get_ahp_flag();
         tcg_gen_ext16u_i32(tcg_rn, tcg_rn);
         if (dtype == 0) {
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_single(DisasContext *s, int opcode,
     TCGv_ptr fpst;
 
     tcg_res = tcg_temp_new_i32();
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     tcg_op1 = read_fp_sreg(s, rn);
     tcg_op2 = read_fp_sreg(s, rm);
 
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_double(DisasContext *s, int opcode,
     TCGv_ptr fpst;
 
     tcg_res = tcg_temp_new_i64();
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     tcg_op1 = read_fp_dreg(s, rn);
     tcg_op2 = read_fp_dreg(s, rm);
 
@@ -XXX,XX +XXX,XX @@ static void handle_fp_2src_half(DisasContext *s, int opcode,
     TCGv_ptr fpst;
 
     tcg_res = tcg_temp_new_i32();
-    fpst = get_fpstatus_ptr(true);
+    fpst = fpstatus_ptr(FPST_FPCR_F16);
     tcg_op1 = read_fp_hreg(s, rn);
     tcg_op2 = read_fp_hreg(s, rm);
 
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_single(DisasContext *s, bool o0, bool o1,
 {
     TCGv_i32 tcg_op1, tcg_op2, tcg_op3;
     TCGv_i32 tcg_res = tcg_temp_new_i32();
-    TCGv_ptr fpst = get_fpstatus_ptr(false);
+    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
 
     tcg_op1 = read_fp_sreg(s, rn);
     tcg_op2 = read_fp_sreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_double(DisasContext *s, bool o0, bool o1,
 {
     TCGv_i64 tcg_op1, tcg_op2, tcg_op3;
     TCGv_i64 tcg_res = tcg_temp_new_i64();
-    TCGv_ptr fpst = get_fpstatus_ptr(false);
+    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
 
     tcg_op1 = read_fp_dreg(s, rn);
     tcg_op2 = read_fp_dreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fp_3src_half(DisasContext *s, bool o0, bool o1,
 {
     TCGv_i32 tcg_op1, tcg_op2, tcg_op3;
     TCGv_i32 tcg_res = tcg_temp_new_i32();
-    TCGv_ptr fpst = get_fpstatus_ptr(true);
+    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR_F16);
 
     tcg_op1 = read_fp_hreg(s, rn);
     tcg_op2 = read_fp_hreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_fpfpcvt(DisasContext *s, int rd, int rn, int opcode,
     TCGv_i32 tcg_shift, tcg_single;
     TCGv_i64 tcg_double;
 
-    tcg_fpstatus = get_fpstatus_ptr(type == 3);
+    tcg_fpstatus = fpstatus_ptr(type == 3 ? FPST_FPCR_F16 : FPST_FPCR);
 
     tcg_shift = tcg_const_i32(64 - scale);
 
@@ -XXX,XX +XXX,XX @@ static void handle_fmov(DisasContext *s, int rd, int rn, int type, bool itof)
 static void handle_fjcvtzs(DisasContext *s, int rd, int rn)
 {
     TCGv_i64 t = read_fp_dreg(s, rn);
-    TCGv_ptr fpstatus = get_fpstatus_ptr(false);
+    TCGv_ptr fpstatus = fpstatus_ptr(FPST_FPCR);
 
     gen_helper_fjcvtzs(t, t, fpstatus);
 
@@ -XXX,XX +XXX,XX @@ static void disas_simd_across_lanes(DisasContext *s, uint32_t insn)
          * Note that correct NaN propagation requires that we do these
          * operations in exactly the order specified by the pseudocode.
          */
-        TCGv_ptr fpst = get_fpstatus_ptr(size == MO_16);
+        TCGv_ptr fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         int fpopcode = opcode | is_min << 4 | is_u << 5;
         int vmap = (1 << elements) - 1;
         TCGv_i32 tcg_res32 = do_reduction_op(s, fpopcode, rn, esize,
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_pairwise(DisasContext *s, uint32_t insn)
             return;
         }
 
-        fpst = get_fpstatus_ptr(size == MO_16);
+        fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         break;
     default:
         unallocated_encoding(s);
@@ -XXX,XX +XXX,XX @@ static void handle_simd_intfp_conv(DisasContext *s, int rd, int rn,
                                    int elements, int is_signed,
                                    int fracbits, int size)
 {
-    TCGv_ptr tcg_fpst = get_fpstatus_ptr(size == MO_16);
+    TCGv_ptr tcg_fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
     TCGv_i32 tcg_shift = NULL;
 
     MemOp mop = size | (is_signed ? MO_SIGN : 0);
@@ -XXX,XX +XXX,XX @@ static void handle_simd_shift_fpint_conv(DisasContext *s, bool is_scalar,
     assert(!(is_scalar && is_q));
 
     tcg_rmode = tcg_const_i32(arm_rmode_to_sf(FPROUNDING_ZERO));
-    tcg_fpstatus = get_fpstatus_ptr(size == MO_16);
+    tcg_fpstatus = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
     gen_helper_set_rmode(tcg_rmode, tcg_rmode, tcg_fpstatus);
     fracbits = (16 << size) - immhb;
     tcg_shift = tcg_const_i32(fracbits);
@@ -XXX,XX +XXX,XX @@ static void handle_3same_float(DisasContext *s, int size, int elements,
                                int fpopcode, int rd, int rn, int rm)
 {
     int pass;
-    TCGv_ptr fpst = get_fpstatus_ptr(false);
+    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
 
     for (pass = 0; pass < elements; pass++) {
         if (size) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_three_reg_same_fp16(DisasContext *s,
         return;
     }
 
-    fpst = get_fpstatus_ptr(true);
+    fpst = fpstatus_ptr(FPST_FPCR_F16);
 
     tcg_op1 = read_fp_hreg(s, rn);
     tcg_op2 = read_fp_hreg(s, rm);
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_fcmp_zero(DisasContext *s, int opcode,
         return;
     }
 
-    fpst = get_fpstatus_ptr(size == MO_16);
+    fpst = fpstatus_ptr(size == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
     if (is_double) {
         TCGv_i64 tcg_op = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_reciprocal(DisasContext *s, int opcode,
                                     int size, int rn, int rd)
 {
     bool is_double = (size == 3);
-    TCGv_ptr fpst = get_fpstatus_ptr(false);
+    TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
 
     if (is_double) {
         TCGv_i64 tcg_op = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_narrow(DisasContext *s, bool scalar,
             } else {
                 TCGv_i32 tcg_lo = tcg_temp_new_i32();
                 TCGv_i32 tcg_hi = tcg_temp_new_i32();
-                TCGv_ptr fpst = get_fpstatus_ptr(false);
+                TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
                 TCGv_i32 ahp = get_ahp_flag();
 
                 tcg_gen_extr_i64_i32(tcg_lo, tcg_hi, tcg_op);
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_two_reg_misc(DisasContext *s, uint32_t insn)
 
     if (is_fcvt) {
         tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
-        tcg_fpstatus = get_fpstatus_ptr(false);
+        tcg_fpstatus = fpstatus_ptr(FPST_FPCR);
         gen_helper_set_rmode(tcg_rmode, tcg_rmode, tcg_fpstatus);
     } else {
         tcg_rmode = NULL;
@@ -XXX,XX +XXX,XX @@ static void handle_simd_3same_pair(DisasContext *s, int is_q, int u, int opcode,
 
     /* Floating point operations need fpst */
     if (opcode >= 0x58) {
-        fpst = get_fpstatus_ptr(false);
+        fpst = fpstatus_ptr(FPST_FPCR);
     } else {
         fpst = NULL;
     }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_same_fp16(DisasContext *s, uint32_t insn)
         break;
     }
 
-    fpst = get_fpstatus_ptr(true);
+    fpst = fpstatus_ptr(FPST_FPCR_F16);
 
     if (pairwise) {
         int maxpass = is_q ? 8 : 4;
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_widening(DisasContext *s, int opcode, bool is_q,
         /* 16 -> 32 bit fp conversion */
         int srcelt = is_q ? 4 : 0;
         TCGv_i32 tcg_res[4];
-        TCGv_ptr fpst = get_fpstatus_ptr(false);
+        TCGv_ptr fpst = fpstatus_ptr(FPST_FPCR);
         TCGv_i32 ahp = get_ahp_flag();
 
         for (pass = 0; pass < 4; pass++) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc(DisasContext *s, uint32_t insn)
     }
 
     if (need_fpstatus || need_rmode) {
-        tcg_fpstatus = get_fpstatus_ptr(false);
+        tcg_fpstatus = fpstatus_ptr(FPST_FPCR);
     } else {
         tcg_fpstatus = NULL;
     }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc_fp16(DisasContext *s, uint32_t insn)
     }
 
     if (need_rmode || need_fpst) {
-        tcg_fpstatus = get_fpstatus_ptr(true);
+        tcg_fpstatus = fpstatus_ptr(FPST_FPCR_F16);
     }
 
     if (need_rmode) {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
     }
 
     if (is_fp) {
-        fpst = get_fpstatus_ptr(is_fp16);
+        fpst = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
     } else {
         fpst = NULL;
     }
diff --git a/target/arm/translate-sve.c b/target/arm/translate-sve.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-sve.c
+++ b/target/arm/translate-sve.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FMLA_zzxz(DisasContext *s, arg_FMLA_zzxz *a)
 
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FMUL_zzx(DisasContext *s, arg_FMUL_zzx *a)
 
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static void do_reduce(DisasContext *s, arg_rpr_esz *a,
 
     tcg_gen_addi_ptr(t_zn, cpu_env, vec_full_reg_offset(s, a->rn));
     tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, a->pg));
-    status = get_fpstatus_ptr(a->esz == MO_16);
+    status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
     fn(temp, t_zn, t_pg, status, t_desc);
     tcg_temp_free_ptr(t_zn);
@@ -XXX,XX +XXX,XX @@ DO_VPZ(FMAXV, fmaxv)
 static void do_zz_fp(DisasContext *s, arg_rr_esz *a, gen_helper_gvec_2_ptr *fn)
 {
     unsigned vsz = vec_full_reg_size(s);
-    TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+    TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
     tcg_gen_gvec_2_ptr(vec_full_reg_offset(s, a->rd),
                        vec_full_reg_offset(s, a->rn),
@@ -XXX,XX +XXX,XX @@ static void do_ppz_fp(DisasContext *s, arg_rpr_esz *a,
                       gen_helper_gvec_3_ptr *fn)
 {
     unsigned vsz = vec_full_reg_size(s);
-    TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+    TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
     tcg_gen_gvec_3_ptr(pred_full_reg_offset(s, a->rd),
                        vec_full_reg_offset(s, a->rn),
@@ -XXX,XX +XXX,XX @@ static bool trans_FTMAD(DisasContext *s, arg_FTMAD *a)
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FADDA(DisasContext *s, arg_rprr_esz *a)
     t_pg = tcg_temp_new_ptr();
     tcg_gen_addi_ptr(t_rm, cpu_env, vec_full_reg_offset(s, a->rm));
     tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, a->pg));
-    t_fpst = get_fpstatus_ptr(a->esz == MO_16);
+    t_fpst = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
     t_desc = tcg_const_i32(simd_desc(vsz, vsz, 0));
 
     fns[a->esz - 1](t_val, t_val, t_rm, t_pg, t_fpst, t_desc);
@@ -XXX,XX +XXX,XX @@ static bool do_zzz_fp(DisasContext *s, arg_rrr_esz *a,
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_zpzz_fp(DisasContext *s, arg_rprr_esz *a,
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static void do_fp_scalar(DisasContext *s, int zd, int zn, int pg, bool is_fp16,
     tcg_gen_addi_ptr(t_zn, cpu_env, vec_full_reg_offset(s, zn));
     tcg_gen_addi_ptr(t_pg, cpu_env, pred_full_reg_offset(s, pg));
 
-    status = get_fpstatus_ptr(is_fp16);
+    status = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
     desc = tcg_const_i32(simd_desc(vsz, vsz, 0));
     fn(t_zd, t_zn, t_pg, scalar, status, desc);
 
@@ -XXX,XX +XXX,XX @@ static bool do_fp_cmp(DisasContext *s, arg_rprr_esz *a,
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_4_ptr(pred_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCADD(DisasContext *s, arg_FCADD *a)
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_4_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_fmla(DisasContext *s, arg_rprrr_esz *a,
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_5_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCMLA_zpzzz(DisasContext *s, arg_FCMLA_zpzzz *a)
     }
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_5_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool trans_FCMLA_zzxz(DisasContext *s, arg_FCMLA_zzxz *a)
     tcg_debug_assert(a->rd == a->ra);
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, a->rd),
                            vec_full_reg_offset(s, a->rn),
                            vec_full_reg_offset(s, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_zpz_ptr(DisasContext *s, int rd, int rn, int pg,
 {
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
-        TCGv_ptr status = get_fpstatus_ptr(is_fp16);
+        TCGv_ptr status = fpstatus_ptr(is_fp16 ? FPST_FPCR_F16 : FPST_FPCR);
         tcg_gen_gvec_3_ptr(vec_full_reg_offset(s, rd),
                            vec_full_reg_offset(s, rn),
                            pred_full_reg_offset(s, pg),
@@ -XXX,XX +XXX,XX @@ static bool do_frint_mode(DisasContext *s, arg_rpr_esz *a, int mode)
     if (sve_access_check(s)) {
         unsigned vsz = vec_full_reg_size(s);
         TCGv_i32 tmode = tcg_const_i32(mode);
-        TCGv_ptr status = get_fpstatus_ptr(a->esz == MO_16);
+        TCGv_ptr status = fpstatus_ptr(a->esz == MO_16 ? FPST_FPCR_F16 : FPST_FPCR);
 
         gen_helper_set_rmode(tmode, tmode, status);
 
-- 
2.20.1

Make A32/T32 code use the new fpstatus_ptr() API:
 get_fpstatus_ptr(0) -> fpstatus_ptr(FPST_FPCR)
 get_fpstatus_ptr(1) -> fpstatus_ptr(FPST_STD)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20200806104453.30393-3-peter.maydell@linaro.org
---
 target/arm/translate.c          | 13 ----------
 target/arm/translate-neon.c.inc | 28 ++++++++++-----------
 target/arm/translate-vfp.c.inc  | 44 ++++++++++++++++-----------------
 3 files changed, 36 insertions(+), 49 deletions(-)

diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static inline void gen_hlt(DisasContext *s, int imm)
     unallocated_encoding(s);
 }
 
-static TCGv_ptr get_fpstatus_ptr(int neon)
-{
-    TCGv_ptr statusptr = tcg_temp_new_ptr();
-    int offset;
-    if (neon) {
-        offset = offsetof(CPUARMState, vfp.standard_fp_status);
-    } else {
-        offset = offsetof(CPUARMState, vfp.fp_status);
-    }
-    tcg_gen_addi_ptr(statusptr, cpu_env, offset);
-    return statusptr;
-}
-
 static inline long vfp_reg_offset(bool dp, unsigned reg)
 {
     if (dp) {
diff --git a/target/arm/translate-neon.c.inc b/target/arm/translate-neon.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.c.inc
+++ b/target/arm/translate-neon.c.inc
@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
     }
 
     opr_sz = (1 + a->q) * 8;
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     fn_gvec_ptr = a->size ? gen_helper_gvec_fcmlas : gen_helper_gvec_fcmlah;
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
@@ -XXX,XX +XXX,XX @@ static bool trans_VCADD(DisasContext *s, arg_VCADD *a)
     }
 
     opr_sz = (1 + a->q) * 8;
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     fn_gvec_ptr = a->size ? gen_helper_gvec_fcadds : gen_helper_gvec_fcaddh;
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA_scalar(DisasContext *s, arg_VCMLA_scalar *a)
     fn_gvec_ptr = (a->size ? gen_helper_gvec_fcmlas_idx
                    : gen_helper_gvec_fcmlah_idx);
     opr_sz = (1 + a->q) * 8;
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
                        vfp_reg_offset(1, a->vm),
@@ -XXX,XX +XXX,XX @@ static bool trans_VDOT_scalar(DisasContext *s, arg_VDOT_scalar *a)
 
     fn_gvec = a->u ? gen_helper_gvec_udot_idx_b : gen_helper_gvec_sdot_idx_b;
     opr_sz = (1 + a->q) * 8;
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     tcg_gen_gvec_3_ool(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
                        vfp_reg_offset(1, a->rm),
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn,
         return true;
     }
 
-    TCGv_ptr fpstatus = get_fpstatus_ptr(1);
+    TCGv_ptr fpstatus = fpstatus_ptr(FPST_STD);
     for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
         tmp = neon_load_reg(a->vn, pass);
         tmp2 = neon_load_reg(a->vm, pass);
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn,
                                 uint32_t rn_ofs, uint32_t rm_ofs,       \
                                 uint32_t oprsz, uint32_t maxsz)         \
     {                                                                   \
-        TCGv_ptr fpst = get_fpstatus_ptr(1);                            \
+        TCGv_ptr fpst = fpstatus_ptr(FPST_STD);                         \
         tcg_gen_gvec_3_ptr(rd_ofs, rn_ofs, rm_ofs, fpst,                \
                            oprsz, maxsz, 0, FUNC);                      \
         tcg_temp_free_ptr(fpst);                                        \
@@ -XXX,XX +XXX,XX @@ static bool do_3same_fp_pair(DisasContext *s, arg_3same *a, VFPGen3OpSPFn *fn)
      * early. Since Q is 0 there are always just two passes, so instead
      * of a complicated loop over each pass we just unroll.
      */
-    fpstatus = get_fpstatus_ptr(1);
+    fpstatus = fpstatus_ptr(FPST_STD);
     tmp = neon_load_reg(a->vn, 0);
     tmp2 = neon_load_reg(a->vn, 1);
     fn(tmp, tmp, tmp2, fpstatus);
@@ -XXX,XX +XXX,XX @@ static bool do_fp_2sh(DisasContext *s, arg_2reg_shift *a,
         return true;
     }
 
-    fpstatus = get_fpstatus_ptr(1);
+    fpstatus = fpstatus_ptr(FPST_STD);
     shiftv = tcg_const_i32(a->shift);
     for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
         tmp = neon_load_reg(a->vm, pass);
@@ -XXX,XX +XXX,XX @@ static bool trans_VMLS_2sc(DisasContext *s, arg_2scalar *a)
 #define WRAP_FP_FN(WRAPNAME, FUNC)                              \
     static void WRAPNAME(TCGv_i32 rd, TCGv_i32 rn, TCGv_i32 rm) \
     {                                                           \
-        TCGv_ptr fpstatus = get_fpstatus_ptr(1);                \
+        TCGv_ptr fpstatus = fpstatus_ptr(FPST_STD);             \
         FUNC(rd, rn, rm, fpstatus);                             \
         tcg_temp_free_ptr(fpstatus);                            \
     }
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_F16_F32(DisasContext *s, arg_2misc *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(true);
+    fpst = fpstatus_ptr(FPST_STD);
     ahp = get_ahp_flag();
     tmp = neon_load_reg(a->vm, 0);
     gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_F32_F16(DisasContext *s, arg_2misc *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(true);
+    fpst = fpstatus_ptr(FPST_STD);
     ahp = get_ahp_flag();
     tmp3 = tcg_temp_new_i32();
     tmp = neon_load_reg(a->vm, 0);
@@ -XXX,XX +XXX,XX @@ static bool do_2misc_fp(DisasContext *s, arg_2misc *a,
         return true;
     }
 
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
         TCGv_i32 tmp = neon_load_reg(a->vm, pass);
         fn(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool do_vrint(DisasContext *s, arg_2misc *a, int rmode)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
     gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
     for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
@@ -XXX,XX +XXX,XX @@ static bool do_vcvt(DisasContext *s, arg_2misc *a, int rmode, bool is_signed)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(1);
+    fpst = fpstatus_ptr(FPST_STD);
     tcg_shift = tcg_const_i32(0);
     tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
     gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
diff --git a/target/arm/translate-vfp.c.inc b/target/arm/translate-vfp.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-vfp.c.inc
+++ b/target/arm/translate-vfp.c.inc
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINT(DisasContext *s, arg_VRINT *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
 
     tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rounding));
     gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT(DisasContext *s, arg_VCVT *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
 
     tcg_shift = tcg_const_i32(0);
 
@@ -XXX,XX +XXX,XX @@ static bool do_vfp_3op_sp(DisasContext *s, VFPGen3OpSPFn *fn,
     f0 = tcg_temp_new_i32();
     f1 = tcg_temp_new_i32();
     fd = tcg_temp_new_i32();
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
 
     neon_load_reg32(f0, vn);
     neon_load_reg32(f1, vm);
@@ -XXX,XX +XXX,XX @@ static bool do_vfp_3op_dp(DisasContext *s, VFPGen3OpDPFn *fn,
     f0 = tcg_temp_new_i64();
     f1 = tcg_temp_new_i64();
     fd = tcg_temp_new_i64();
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
 
     neon_load_reg64(f0, vn);
     neon_load_reg64(f1, vm);
@@ -XXX,XX +XXX,XX @@ static bool do_vfm_sp(DisasContext *s, arg_VFMA_sp *a, bool neg_n, bool neg_d)
         /* VFNMA, VFNMS */
         gen_helper_vfp_negs(vd, vd);
     }
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_vfp_muladds(vd, vn, vm, vd, fpst);
     neon_store_reg32(vd, a->vd);
 
@@ -XXX,XX +XXX,XX @@ static bool do_vfm_dp(DisasContext *s, arg_VFMA_dp *a, bool neg_n, bool neg_d)
         /* VFNMA, VFNMS */
         gen_helper_vfp_negd(vd, vd);
     }
-    fpst = get_fpstatus_ptr(0);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_vfp_muladdd(vd, vn, vm, vd, fpst);
     neon_store_reg64(vd, a->vd);
 
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f32_f16(DisasContext *s, arg_VCVT_f32_f16 *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     ahp_mode = get_ahp_flag();
     tmp = tcg_temp_new_i32();
     /* The T bit tells us if we want the low or high 16 bits of Vm */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f64_f16(DisasContext *s, arg_VCVT_f64_f16 *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     ahp_mode = get_ahp_flag();
     tmp = tcg_temp_new_i32();
     /* The T bit tells us if we want the low or high 16 bits of Vm */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f16_f32(DisasContext *s, arg_VCVT_f16_f32 *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     ahp_mode = get_ahp_flag();
     tmp = tcg_temp_new_i32();
 
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_f16_f64(DisasContext *s, arg_VCVT_f16_f64 *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     ahp_mode = get_ahp_flag();
     tmp = tcg_temp_new_i32();
     vm = tcg_temp_new_i64();
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTR_sp(DisasContext *s, arg_VRINTR_sp *a)
 
     tmp = tcg_temp_new_i32();
     neon_load_reg32(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_rints(tmp, tmp, fpst);
     neon_store_reg32(tmp, a->vd);
     tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTR_dp(DisasContext *s, arg_VRINTR_dp *a)
 
     tmp = tcg_temp_new_i64();
     neon_load_reg64(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_rintd(tmp, tmp, fpst);
     neon_store_reg64(tmp, a->vd);
     tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTZ_sp(DisasContext *s, arg_VRINTZ_sp *a)
 
     tmp = tcg_temp_new_i32();
     neon_load_reg32(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     tcg_rmode = tcg_const_i32(float_round_to_zero);
     gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
     gen_helper_rints(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTZ_dp(DisasContext *s, arg_VRINTZ_dp *a)
 
     tmp = tcg_temp_new_i64();
     neon_load_reg64(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     tcg_rmode = tcg_const_i32(float_round_to_zero);
     gen_helper_set_rmode(tcg_rmode, tcg_rmode, fpst);
     gen_helper_rintd(tmp, tmp, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTX_sp(DisasContext *s, arg_VRINTX_sp *a)
 
     tmp = tcg_temp_new_i32();
     neon_load_reg32(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_rints_exact(tmp, tmp, fpst);
     neon_store_reg32(tmp, a->vd);
     tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTX_dp(DisasContext *s, arg_VRINTX_dp *a)
 
     tmp = tcg_temp_new_i64();
     neon_load_reg64(tmp, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     gen_helper_rintd_exact(tmp, tmp, fpst);
     neon_store_reg64(tmp, a->vd);
     tcg_temp_free_ptr(fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_int_sp(DisasContext *s, arg_VCVT_int_sp *a)
 
     vm = tcg_temp_new_i32();
     neon_load_reg32(vm, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     if (a->s) {
         /* i32 -> f32 */
         gen_helper_vfp_sitos(vm, vm, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_int_dp(DisasContext *s, arg_VCVT_int_dp *a)
     vm = tcg_temp_new_i32();
     vd = tcg_temp_new_i64();
     neon_load_reg32(vm, a->vm);
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     if (a->s) {
         /* i32 -> f64 */
         gen_helper_vfp_sitod(vd, vm, fpst);
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_fix_sp(DisasContext *s, arg_VCVT_fix_sp *a)
     vd = tcg_temp_new_i32();
     neon_load_reg32(vd, a->vd);
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     shift = tcg_const_i32(frac_bits);
 
     /* Switch on op:U:sx bits */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_fix_dp(DisasContext *s, arg_VCVT_fix_dp *a)
     vd = tcg_temp_new_i64();
     neon_load_reg64(vd, a->vd);
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     shift = tcg_const_i32(frac_bits);
 
     /* Switch on op:U:sx bits */
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_sp_int(DisasContext *s, arg_VCVT_sp_int *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     vm = tcg_temp_new_i32();
     neon_load_reg32(vm, a->vm);
 
@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_dp_int(DisasContext *s, arg_VCVT_dp_int *a)
         return true;
     }
 
-    fpst = get_fpstatus_ptr(false);
+    fpst = fpstatus_ptr(FPST_FPCR);
     vm = tcg_temp_new_i64();
     vd = tcg_temp_new_i32();
     neon_load_reg64(vm, a->vm);
-- 
2.20.1

Architecturally, Neon FP16 operations use the "standard FPSCR" like
all other Neon operations.  However, this is defined in the Arm ARM
pseudocode as "a fixed value, except that FZ16 (and AHP) follow the
FPSCR bits". In QEMU, the softfloat float_status doesn't include
separate flush-to-zero for FP16 operations, so we must keep separate
fp_status for "Neon non-FP16" and "Neon fp16" operations, in the
same way we do already for the non-Neon "fp_status" vs "fp_status_f16".

Add the extra float_status field to the CPU state structure,
ensure it is correctly initialized and updated on FPSCR writes,
and make fpstatus_ptr(FPST_STD_F16) return a pointer to it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20200806104453.30393-4-peter.maydell@linaro.org
---
 target/arm/cpu.h        | 9 ++++++++-
 target/arm/translate.h  | 3 ++-
 target/arm/cpu.c        | 3 +++
 target/arm/vfp_helper.c | 5 +++++
 4 files changed, 18 insertions(+), 2 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ typedef struct CPUARMState {
          *  fp_status: is the "normal" fp status.
          *  fp_status_fp16: used for half-precision calculations
          *  standard_fp_status : the ARM "Standard FPSCR Value"
+         *  standard_fp_status_fp16 : used for half-precision
+         *       calculations with the ARM "Standard FPSCR Value"
          *
          * Half-precision operations are governed by a separate
          * flush-to-zero control bit in FPSCR:FZ16. We pass a separate
@@ -XXX,XX +XXX,XX @@ typedef struct CPUARMState {
          * Neon) which the architecture defines as controlled by the
          * standard FPSCR value rather than the FPSCR.
          *
+         * The "standard FPSCR but for fp16 ops" is needed because
+         * the "standard FPSCR" tracks the FPSCR.FZ16 bit rather than
+         * using a fixed value for it.
+         *
          * To avoid having to transfer exception bits around, we simply
          * say that the FPSCR cumulative exception flags are the logical
-         * OR of the flags in the three fp statuses. This relies on the
+         * OR of the flags in the four fp statuses. This relies on the
          * only thing which needs to read the exception flags being
          * an explicit FPSCR read.
          */
         float_status fp_status;
         float_status fp_status_f16;
         float_status standard_fp_status;
+        float_status standard_fp_status_f16;
 
         /* ZCR_EL[1-3] */
         uint64_t zcr_el[4];
diff --git a/target/arm/translate.h b/target/arm/translate.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.h
+++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ static inline TCGv_ptr fpstatus_ptr(ARMFPStatusFlavour flavour)
         offset = offsetof(CPUARMState, vfp.standard_fp_status);
         break;
     case FPST_STD_F16:
-        /* Not yet used or implemented: fall through to assert */
+        offset = offsetof(CPUARMState, vfp.standard_fp_status_f16);
+        break;
     default:
         g_assert_not_reached();
     }
diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ static void arm_cpu_reset(DeviceState *dev)
     set_flush_to_zero(1, &env->vfp.standard_fp_status);
     set_flush_inputs_to_zero(1, &env->vfp.standard_fp_status);
     set_default_nan_mode(1, &env->vfp.standard_fp_status);
+    set_default_nan_mode(1, &env->vfp.standard_fp_status_f16);
     set_float_detect_tininess(float_tininess_before_rounding,
                               &env->vfp.fp_status);
     set_float_detect_tininess(float_tininess_before_rounding,
                               &env->vfp.standard_fp_status);
     set_float_detect_tininess(float_tininess_before_rounding,
                               &env->vfp.fp_status_f16);
+    set_float_detect_tininess(float_tininess_before_rounding,
+                              &env->vfp.standard_fp_status_f16);
 #ifndef CONFIG_USER_ONLY
     if (kvm_enabled()) {
         kvm_arm_reset_vcpu(cpu);
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
     /* FZ16 does not generate an input denormal exception.  */
     i |= (get_float_exception_flags(&env->vfp.fp_status_f16)
           & ~float_flag_input_denormal);
+    i |= (get_float_exception_flags(&env->vfp.standard_fp_status_f16)
+          & ~float_flag_input_denormal);
     return vfp_exceptbits_from_host(i);
 }
 
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
     if (changed & FPCR_FZ16) {
         bool ftz_enabled = val & FPCR_FZ16;
         set_flush_to_zero(ftz_enabled, &env->vfp.fp_status_f16);
+        set_flush_to_zero(ftz_enabled, &env->vfp.standard_fp_status_f16);
         set_flush_inputs_to_zero(ftz_enabled, &env->vfp.fp_status_f16);
+        set_flush_inputs_to_zero(ftz_enabled, &env->vfp.standard_fp_status_f16);
     }
     if (changed & FPCR_FZ) {
         bool ftz_enabled = val & FPCR_FZ;
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
     set_float_exception_flags(i, &env->vfp.fp_status);
     set_float_exception_flags(0, &env->vfp.fp_status_f16);
     set_float_exception_flags(0, &env->vfp.standard_fp_status);
+    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
 }
 
 #else
-- 
2.20.1

When we implemented the VCMLA and VCADD insns we put in the
code to handle fp16, but left it using the standard fp status
flags. Correct them to use FPST_STD_F16 for fp16 operations.

diff --git a/target/arm/translate-neon.c.inc b/target/arm/translate-neon.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.c.inc
+++ b/target/arm/translate-neon.c.inc
@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
     }
 
     opr_sz = (1 + a->q) * 8;
-    fpst = fpstatus_ptr(FPST_STD);
+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
     fn_gvec_ptr = a->size ? gen_helper_gvec_fcmlas : gen_helper_gvec_fcmlah;
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
@@ -XXX,XX +XXX,XX @@ static bool trans_VCADD(DisasContext *s, arg_VCADD *a)
     }
 
     opr_sz = (1 + a->q) * 8;
-    fpst = fpstatus_ptr(FPST_STD);
+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
     fn_gvec_ptr = a->size ? gen_helper_gvec_fcadds : gen_helper_gvec_fcaddh;
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
@@ -XXX,XX +XXX,XX @@ static bool trans_VCMLA_scalar(DisasContext *s, arg_VCMLA_scalar *a)
     fn_gvec_ptr = (a->size ? gen_helper_gvec_fcmlas_idx
                    : gen_helper_gvec_fcmlah_idx);
     opr_sz = (1 + a->q) * 8;
-    fpst = fpstatus_ptr(FPST_STD);
+    fpst = fpstatus_ptr(a->size == 0 ? FPST_STD_F16 : FPST_STD);
     tcg_gen_gvec_3_ptr(vfp_reg_offset(1, a->vd),
                        vfp_reg_offset(1, a->vn),
                        vfp_reg_offset(1, a->vm),
-- 
2.20.1

First arm pullreq of the cycle; this is mostly my softfloat NaN
handling series. (Lots more in my to-review queue, but I don't
like pullreqs growing too close to a hundred patches at a time :-))

thanks
-- PMM

The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:

Open 10.0 development tree (2024-12-10 17:41:17 +0000)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211

for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:

MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)

----------------------------------------------------------------
target-arm queue:
 * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
 * fpu: Make muladd NaN handling runtime-selected, not compile-time
 * fpu: Make default NaN pattern runtime-selected, not compile-time
 * fpu: Minor NaN-related cleanups
 * MAINTAINERS: email address updates

----------------------------------------------------------------
Bernhard Beschow (5):
      hw/net/lan9118: Extract lan9118_phy
      hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
      hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
      hw/net/lan9118_phy: Reuse MII constants
      hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement

Leif Lindholm (1):
      MAINTAINERS: update email address for Leif Lindholm

Peter Maydell (54):
      fpu: handle raising Invalid for infzero in pick_nan_muladd
      fpu: Check for default_nan_mode before calling pickNaNMulAdd
      softfloat: Allow runtime choice of inf * 0 + NaN result
      tests/fp: Explicitly set inf-zero-nan rule
      target/arm: Set FloatInfZeroNaNRule explicitly
      target/s390: Set FloatInfZeroNaNRule explicitly
      target/ppc: Set FloatInfZeroNaNRule explicitly
      target/mips: Set FloatInfZeroNaNRule explicitly
      target/sparc: Set FloatInfZeroNaNRule explicitly
      target/xtensa: Set FloatInfZeroNaNRule explicitly
      target/x86: Set FloatInfZeroNaNRule explicitly
      target/loongarch: Set FloatInfZeroNaNRule explicitly
      target/hppa: Set FloatInfZeroNaNRule explicitly
      softfloat: Pass have_snan to pickNaNMulAdd
      softfloat: Allow runtime choice of NaN propagation for muladd
      tests/fp: Explicitly set 3-NaN propagation rule
      target/arm: Set Float3NaNPropRule explicitly
      target/loongarch: Set Float3NaNPropRule explicitly
      target/ppc: Set Float3NaNPropRule explicitly
      target/s390x: Set Float3NaNPropRule explicitly
      target/sparc: Set Float3NaNPropRule explicitly
      target/mips: Set Float3NaNPropRule explicitly
      target/xtensa: Set Float3NaNPropRule explicitly
      target/i386: Set Float3NaNPropRule explicitly
      target/hppa: Set Float3NaNPropRule explicitly
      fpu: Remove use_first_nan field from float_status
      target/m68k: Don't pass NULL float_status to floatx80_default_nan()
      softfloat: Create floatx80 default NaN from parts64_default_nan
      target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
      target/m68k: In frem helper, initialize local float_status from env->fp_status
      target/m68k: Init local float_status from env fp_status in gdb get/set reg
      target/sparc: Initialize local scratch float_status from env->fp_status
      target/ppc: Use env->fp_status in helper_compute_fprf functions
      fpu: Allow runtime choice of default NaN value
      tests/fp: Set default NaN pattern explicitly
      target/microblaze: Set default NaN pattern explicitly
      target/i386: Set default NaN pattern explicitly
      target/hppa: Set default NaN pattern explicitly
      target/alpha: Set default NaN pattern explicitly
      target/arm: Set default NaN pattern explicitly
      target/loongarch: Set default NaN pattern explicitly
      target/m68k: Set default NaN pattern explicitly
      target/mips: Set default NaN pattern explicitly
      target/openrisc: Set default NaN pattern explicitly
      target/ppc: Set default NaN pattern explicitly
      target/sh4: Set default NaN pattern explicitly
      target/rx: Set default NaN pattern explicitly
      target/s390x: Set default NaN pattern explicitly
      target/sparc: Set default NaN pattern explicitly
      target/xtensa: Set default NaN pattern explicitly
      target/hexagon: Set default NaN pattern explicitly
      target/riscv: Set default NaN pattern explicitly
      target/tricore: Set default NaN pattern explicitly
      fpu: Remove default handling for dnan_pattern

Richard Henderson (11):
      target/arm: Copy entire float_status in is_ebf
      softfloat: Inline pickNaNMulAdd
      softfloat: Use goto for default nan case in pick_nan_muladd
      softfloat: Remove which from parts_pick_nan_muladd
      softfloat: Pad array size in pick_nan_muladd
      softfloat: Move propagateFloatx80NaN to softfloat.c
      softfloat: Use parts_pick_nan in propagateFloatx80NaN
      softfloat: Inline pickNaN
      softfloat: Share code between parts_pick_nan cases
      softfloat: Sink frac_cmp in parts_pick_nan until needed
      softfloat: Replace WHICH with RET in parts_pick_nan

Vikram Garhwal (1):
      MAINTAINERS: Add correct email address for Vikram Garhwal

From: Bernhard Beschow <shentey@gmail.com>

A very similar implementation of the same device exists in imx_fec. Prepare for
a common implementation by extracting a device model into its own files.

Some migration state has been moved into the new device model which breaks
migration compatibility for the following machines:
* smdkc210
* realview-*
* vexpress-*
* kzm
* mps2-*

While breaking migration ABI, fix the size of the MII registers to be 16 bit,
as defined by IEEE 802.3u.

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-2-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/lan9118_phy.h |  37 ++++++++
 hw/net/lan9118.c             | 137 +++++-----------------------
 hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
 hw/net/Kconfig               |   4 +
 hw/net/meson.build           |   1 +
 5 files changed, 233 insertions(+), 115 deletions(-)
 create mode 100644 include/hw/net/lan9118_phy.h
 create mode 100644 hw/net/lan9118_phy.c

diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/include/hw/net/lan9118_phy.h
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This work is licensed under the terms of the GNU GPL, version 2 or later.
+ * See the COPYING file in the top-level directory.
+ */
+
+#ifndef HW_NET_LAN9118_PHY_H
+#define HW_NET_LAN9118_PHY_H
+
+#include "qom/object.h"
+#include "hw/sysbus.h"
+
+#define TYPE_LAN9118_PHY "lan9118-phy"
+OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
+
+typedef struct Lan9118PhyState {
+    SysBusDevice parent_obj;
+
+    uint16_t status;
+    uint16_t control;
+    uint16_t advertise;
+    uint16_t ints;
+    uint16_t int_mask;
+    qemu_irq irq;
+    bool link_down;
+} Lan9118PhyState;
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
+void lan9118_phy_reset(Lan9118PhyState *s);
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
+
+#endif
diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118.c
+++ b/hw/net/lan9118.c
@@ -XXX,XX +XXX,XX @@
 #include "net/net.h"
 #include "net/eth.h"
 #include "hw/irq.h"
+#include "hw/net/lan9118_phy.h"
 #include "hw/net/lan9118.h"
 #include "hw/ptimer.h"
 #include "hw/qdev-properties.h"
@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
 #define MAC_CR_RXEN     0x00000004
 #define MAC_CR_RESERVED 0x7f404213
 
-#define PHY_INT_ENERGYON            0x80
-#define PHY_INT_AUTONEG_COMPLETE    0x40
-#define PHY_INT_FAULT               0x20
-#define PHY_INT_DOWN                0x10
-#define PHY_INT_AUTONEG_LP          0x08
-#define PHY_INT_PARFAULT            0x04
-#define PHY_INT_AUTONEG_PAGE        0x02
-
 #define GPT_TIMER_EN    0x20000000
 
 /*
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
     uint32_t mac_mii_data;
     uint32_t mac_flow;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
 
     int32_t eeprom_writable;
     uint8_t eeprom[128];
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
 
 static const VMStateDescription vmstate_lan9118 = {
     .name = "lan9118",
-    .version_id = 2,
-    .minimum_version_id = 1,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_PTIMER(timer, lan9118_state),
         VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
         VMSTATE_UINT32(mac_mii_acc, lan9118_state),
         VMSTATE_UINT32(mac_mii_data, lan9118_state),
         VMSTATE_UINT32(mac_flow, lan9118_state),
-        VMSTATE_UINT32(phy_status, lan9118_state),
-        VMSTATE_UINT32(phy_control, lan9118_state),
-        VMSTATE_UINT32(phy_advertise, lan9118_state),
-        VMSTATE_UINT32(phy_int, lan9118_state),
-        VMSTATE_UINT32(phy_int_mask, lan9118_state),
         VMSTATE_INT32(eeprom_writable, lan9118_state),
         VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
         VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
     lan9118_mac_changed(s);
 }
 
-static void phy_update_irq(lan9118_state *s)
+static void lan9118_update_irq(void *opaque, int n, int level)
 {
-    if (s->phy_int & s->phy_int_mask) {
+    lan9118_state *s = opaque;
+
+    if (level) {
         s->int_sts |= PHY_INT;
     } else {
         s->int_sts &= ~PHY_INT;
@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
     lan9118_update(s);
 }
 
-static void phy_update_link(lan9118_state *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    phy_update_irq(s);
-}
-
 static void lan9118_set_link(NetClientState *nc)
 {
-    phy_update_link(qemu_get_nic_opaque(nc));
-}
-
-static void phy_reset(lan9118_state *s)
-{
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    phy_update_link(s);
+    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static void lan9118_reset(DeviceState *d)
@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
     s->read_word_n = 0;
     s->write_word_n = 0;
 
-    phy_reset(s);
-
     s->eeprom_writable = 0;
     lan9118_reload_eeprom(s);
 }
@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
     uint32_t status;
 
     /* FIXME: Honor TX disable, and allow queueing of packets.  */
-    if (s->phy_control & 0x4000)  {
+    if (s->mii.control & 0x4000) {
         /* This assumes the receive routine doesn't touch the VLANClient.  */
         qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
     } else {
@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
     }
 }
 
-static uint32_t do_phy_read(lan9118_state *s, int reg)
-{
-    uint32_t val;
-
-    switch (reg) {
-    case 0: /* Basic Control */
-        return s->phy_control;
-    case 1: /* Basic Status */
-        return s->phy_status;
-    case 2: /* ID1 */
-        return 0x0007;
-    case 3: /* ID2 */
-        return 0xc0d1;
-    case 4: /* Auto-neg advertisement */
-        return s->phy_advertise;
-    case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
-    case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
-    case 29: /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        phy_update_irq(s);
-        return val;
-    case 30: /* Interrupt mask */
-        return s->phy_int_mask;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_read: PHY read reg %d\n", reg);
-        return 0;
-    }
-}
-
-static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
-{
-    switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
-            phy_reset(s);
-            break;
-        }
-        s->phy_control = val & 0x7980;
-        /* Complete autonegotiation immediately.  */
-        if (val & 0x1000) {
-            s->phy_status |= 0x0020;
-        }
-        break;
-    case 4: /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-        /* TODO 17, 18, 27, 31 */
-    case 30: /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        phy_update_irq(s);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
-    }
-}
-
 static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
 {
     switch (reg) {
@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
         if (val & 2) {
             DPRINTF("PHY write %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
-            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
+            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
         } else {
-            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
+            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
             DPRINTF("PHY read %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
         }
@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
         break;
     case CSR_PMT_CTRL:
         if (val & 0x400) {
-            phy_reset(s);
+            lan9118_phy_reset(&s->mii);
         }
         s->pmt_ctrl &= ~0x34e;
         s->pmt_ctrl |= (val & 0x34e);
@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
     const MemoryRegionOps *mem_ops =
             s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
 
+    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
                           "lan9118-mmio", 0x100);
     sysbus_init_mmio(sbd, &s->mmio);
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This code is licensed under the GNU GPL v2
+ *
+ * Contributions after 2012-01-13 are licensed under the terms of the
+ * GNU GPL, version 2 or (at your option) any later version.
+ */
+
+#include "qemu/osdep.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
+#include "hw/resettable.h"
+#include "migration/vmstate.h"
+#include "qemu/log.h"
+
+#define PHY_INT_ENERGYON            (1 << 7)
+#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+#define PHY_INT_FAULT               (1 << 5)
+#define PHY_INT_DOWN                (1 << 4)
+#define PHY_INT_AUTONEG_LP          (1 << 3)
+#define PHY_INT_PARFAULT            (1 << 2)
+#define PHY_INT_AUTONEG_PAGE        (1 << 1)
+
+static void lan9118_phy_update_irq(Lan9118PhyState *s)
+{
+    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
+}
+
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+{
+    uint16_t val;
+
+    switch (reg) {
+    case 0: /* Basic Control */
+        return s->control;
+    case 1: /* Basic Status */
+        return s->status;
+    case 2: /* ID1 */
+        return 0x0007;
+    case 3: /* ID2 */
+        return 0xc0d1;
+    case 4: /* Auto-neg advertisement */
+        return s->advertise;
+    case 5: /* Auto-neg Link Partner Ability */
+        return 0x0f71;
+    case 6: /* Auto-neg Expansion */
+        return 1;
+        /* TODO 17, 18, 27, 29, 30, 31 */
+    case 29: /* Interrupt source. */
+        val = s->ints;
+        s->ints = 0;
+        lan9118_phy_update_irq(s);
+        return val;
+    case 30: /* Interrupt mask */
+        return s->int_mask;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_read: PHY read reg %d\n", reg);
+        return 0;
+    }
+}
+
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+{
+    switch (reg) {
+    case 0: /* Basic Control */
+        if (val & 0x8000) {
+            lan9118_phy_reset(s);
+            break;
+        }
+        s->control = val & 0x7980;
+        /* Complete autonegotiation immediately. */
+        if (val & 0x1000) {
+            s->status |= 0x0020;
+        }
+        break;
+    case 4: /* Auto-neg advertisement */
+        s->advertise = (val & 0x2d7f) | 0x80;
+        break;
+        /* TODO 17, 18, 27, 31 */
+    case 30: /* Interrupt mask */
+        s->int_mask = val & 0xff;
+        lan9118_phy_update_irq(s);
+        break;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+    }
+}
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+{
+    s->link_down = link_down;
+
+    /* Autonegotiation status mirrors link status. */
+    if (link_down) {
+        s->status &= ~0x0024;
+        s->ints |= PHY_INT_DOWN;
+    } else {
+        s->status |= 0x0024;
+        s->ints |= PHY_INT_ENERGYON;
+        s->ints |= PHY_INT_AUTONEG_COMPLETE;
+    }
+    lan9118_phy_update_irq(s);
+}
+
+void lan9118_phy_reset(Lan9118PhyState *s)
+{
+    s->control = 0x3000;
+    s->status = 0x7809;
+    s->advertise = 0x01e1;
+    s->int_mask = 0;
+    s->ints = 0;
+    lan9118_phy_update_link(s, s->link_down);
+}
+
+static void lan9118_phy_reset_hold(Object *obj, ResetType type)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    lan9118_phy_reset(s);
+}
+
+static void lan9118_phy_init(Object *obj)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
+}
+
+static const VMStateDescription vmstate_lan9118_phy = {
+    .name = "lan9118-phy",
+    .version_id = 1,
+    .minimum_version_id = 1,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT16(control, Lan9118PhyState),
+        VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(advertise, Lan9118PhyState),
+        VMSTATE_UINT16(ints, Lan9118PhyState),
+        VMSTATE_UINT16(int_mask, Lan9118PhyState),
+        VMSTATE_BOOL(link_down, Lan9118PhyState),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static void lan9118_phy_class_init(ObjectClass *klass, void *data)
+{
+    ResettableClass *rc = RESETTABLE_CLASS(klass);
+    DeviceClass *dc = DEVICE_CLASS(klass);
+
+    rc->phases.hold = lan9118_phy_reset_hold;
+    dc->vmsd = &vmstate_lan9118_phy;
+}
+
+static const TypeInfo types[] = {
+    {
+        .name          = TYPE_LAN9118_PHY,
+        .parent        = TYPE_SYS_BUS_DEVICE,
+        .instance_size = sizeof(Lan9118PhyState),
+        .instance_init = lan9118_phy_init,
+        .class_init    = lan9118_phy_class_init,
+    }
+};
+
+DEFINE_TYPES(types)
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
 config SMC91C111
     bool
 
+config LAN9118_PHY
+    bool
+
 config LAN9118
     bool
+    select LAN9118_PHY
     select PTIMER
 
 config NE2000_ISA
diff --git a/hw/net/meson.build b/hw/net/meson.build
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/meson.build
+++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
 
 system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
 system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
+system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
 system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
 system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
 system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

imx_fec models the same PHY as lan9118_phy. The code is almost the same with
imx_fec having more logging and tracing. Merge these improvements into
lan9118_phy and reuse in imx_fec to fix the code duplication.

Some migration state how resides in the new device model which breaks migration
compatibility for the following machines:
* imx25-pdk
* sabrelite
* mcimx7d-sabre
* mcimx6ul-evk

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-3-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/imx_fec.h |   9 ++-
 hw/net/imx_fec.c         | 146 ++++-----------------------------------
 hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
 hw/net/Kconfig           |   1 +
 hw/net/trace-events      |  10 +--
 5 files changed, 85 insertions(+), 163 deletions(-)

diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/imx_fec.h
+++ b/include/hw/net/imx_fec.h
@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
 #define TYPE_IMX_ENET "imx.enet"
 
 #include "hw/sysbus.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
 #include "net/net.h"
 
 #define ENET_EIR               1
@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
     uint32_t tx_descriptor[ENET_TX_RING_NUM];
     uint32_t tx_ring_num;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
     uint32_t phy_num;
     bool phy_connected;
     struct IMXFECState *phy_consumer;
diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/imx_fec.c
+++ b/hw/net/imx_fec.c
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
 
 static const VMStateDescription vmstate_imx_eth = {
     .name = TYPE_IMX_FEC,
-    .version_id = 2,
-    .minimum_version_id = 2,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
         VMSTATE_UINT32(rx_descriptor, IMXFECState),
         VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
-        VMSTATE_UINT32(phy_status, IMXFECState),
-        VMSTATE_UINT32(phy_control, IMXFECState),
-        VMSTATE_UINT32(phy_advertise, IMXFECState),
-        VMSTATE_UINT32(phy_int, IMXFECState),
-        VMSTATE_UINT32(phy_int_mask, IMXFECState),
         VMSTATE_END_OF_LIST()
     },
     .subsections = (const VMStateDescription * const []) {
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
     },
 };
 
-#define PHY_INT_ENERGYON            (1 << 7)
-#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
-#define PHY_INT_FAULT               (1 << 5)
-#define PHY_INT_DOWN                (1 << 4)
-#define PHY_INT_AUTONEG_LP          (1 << 3)
-#define PHY_INT_PARFAULT            (1 << 2)
-#define PHY_INT_AUTONEG_PAGE        (1 << 1)
-
 static void imx_eth_update(IMXFECState *s);
 
 /*
@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
  * For now we don't handle any GPIO/interrupt line, so the OS will
  * have to poll for the PHY status.
  */
-static void imx_phy_update_irq(IMXFECState *s)
+static void imx_phy_update_irq(void *opaque, int n, int level)
 {
-    imx_eth_update(s);
-}
-
-static void imx_phy_update_link(IMXFECState *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        trace_imx_phy_update_link("down");
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        trace_imx_phy_update_link("up");
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    imx_phy_update_irq(s);
+    imx_eth_update(opaque);
 }
 
 static void imx_eth_set_link(NetClientState *nc)
 {
-    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
-}
-
-static void imx_phy_reset(IMXFECState *s)
-{
-    trace_imx_phy_reset();
-
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    imx_phy_update_link(s);
+    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static uint32_t imx_phy_read(IMXFECState *s, int reg)
 {
-    uint32_t val;
     uint32_t phy = reg / 32;
 
     if (!s->phy_connected) {
@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
 
     reg %= 32;
 
-    switch (reg) {
-    case 0:     /* Basic Control */
-        val = s->phy_control;
-        break;
-    case 1:     /* Basic Status */
-        val = s->phy_status;
-        break;
-    case 2:     /* ID1 */
-        val = 0x0007;
-        break;
-    case 3:     /* ID2 */
-        val = 0xc0d1;
-        break;
-    case 4:     /* Auto-neg advertisement */
-        val = s->phy_advertise;
-        break;
-    case 5:     /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
-        break;
-    case 6:     /* Auto-neg Expansion */
-        val = 1;
-        break;
-    case 29:    /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        imx_phy_update_irq(s);
-        break;
-    case 30:    /* Interrupt mask */
-        val = s->phy_int_mask;
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    }
-
-    trace_imx_phy_read(val, phy, reg);
-
-    return val;
+    return lan9118_phy_read(&s->mii, reg);
 }
 
 static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
 
     reg %= 32;
 
-    trace_imx_phy_write(val, phy, reg);
-
-    switch (reg) {
-    case 0:     /* Basic Control */
-        if (val & 0x8000) {
-            imx_phy_reset(s);
-        } else {
-            s->phy_control = val & 0x7980;
-            /* Complete autonegotiation immediately.  */
-            if (val & 0x1000) {
-                s->phy_status |= 0x0020;
-            }
-        }
-        break;
-    case 4:     /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-    case 30:    /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        imx_phy_update_irq(s);
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    }
+    lan9118_phy_write(&s->mii, reg, val);
 }
 
 static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
 
     s->rx_descriptor = 0;
     memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
-
-    /* We also reset the PHY */
-    imx_phy_reset(s);
 }
 
 static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
     sysbus_init_irq(sbd, &s->irq[0]);
     sysbus_init_irq(sbd, &s->irq[1]);
 
+    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     qemu_macaddr_default_if_unset(&s->conf.macaddr);
 
     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
  * Copyright (c) 2009 CodeSourcery, LLC.
  * Written by Paul Brook
  *
+ * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
+ *
  * This code is licensed under the GNU GPL v2
  *
  * Contributions after 2012-01-13 are licensed under the terms of the
@@ -XXX,XX +XXX,XX @@
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
 #include "qemu/log.h"
+#include "trace.h"
 
 #define PHY_INT_ENERGYON            (1 << 7)
 #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
 
     switch (reg) {
     case 0: /* Basic Control */
-        return s->control;
+        val = s->control;
+        break;
     case 1: /* Basic Status */
-        return s->status;
+        val = s->status;
+        break;
     case 2: /* ID1 */
-        return 0x0007;
+        val = 0x0007;
+        break;
     case 3: /* ID2 */
-        return 0xc0d1;
+        val = 0xc0d1;
+        break;
     case 4: /* Auto-neg advertisement */
-        return s->advertise;
+        val = s->advertise;
+        break;
     case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
+        val = 0x0f71;
+        break;
     case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
+        val = 1;
+        break;
     case 29: /* Interrupt source. */
         val = s->ints;
         s->ints = 0;
         lan9118_phy_update_irq(s);
-        return val;
+        break;
     case 30: /* Interrupt mask */
-        return s->int_mask;
+        val = s->int_mask;
+        break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        val = 0;
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_read: PHY read reg %d\n", reg);
-        return 0;
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        val = 0;
+        break;
     }
+
+    trace_lan9118_phy_read(val, reg);
+
+    return val;
 }
 
 void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
 {
+    trace_lan9118_phy_write(val, reg);
+
     switch (reg) {
     case 0: /* Basic Control */
         if (val & 0x8000) {
             lan9118_phy_reset(s);
-            break;
-        }
-        s->control = val & 0x7980;
-        /* Complete autonegotiation immediately. */
-        if (val & 0x1000) {
-            s->status |= 0x0020;
+        } else {
+            s->control = val & 0x7980;
+            /* Complete autonegotiation immediately. */
+            if (val & 0x1000) {
+                s->status |= 0x0020;
+            }
         }
         break;
     case 4: /* Auto-neg advertisement */
         s->advertise = (val & 0x2d7f) | 0x80;
         break;
-        /* TODO 17, 18, 27, 31 */
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
         lan9118_phy_update_irq(s);
         break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        break;
     }
 }
 
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
+        trace_lan9118_phy_update_link("down");
         s->status &= ~0x0024;
         s->ints |= PHY_INT_DOWN;
     } else {
+        trace_lan9118_phy_update_link("up");
         s->status |= 0x0024;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
 void lan9118_phy_reset(Lan9118PhyState *s)
 {
+    trace_lan9118_phy_reset();
+
     s->control = 0x3000;
     s->status = 0x7809;
     s->advertise = 0x01e1;
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
     .version_id = 1,
     .minimum_version_id = 1,
     .fields = (const VMStateField[]) {
-        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(advertise, Lan9118PhyState),
         VMSTATE_UINT16(ints, Lan9118PhyState),
         VMSTATE_UINT16(int_mask, Lan9118PhyState),
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
 
 config IMX_FEC
     bool
+    select LAN9118_PHY
 
 config CADENCE
     bool
diff --git a/hw/net/trace-events b/hw/net/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
 allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
 allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
 
+# lan9118_phy.c
+lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
+lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
+lan9118_phy_update_link(const char *s) "%s"
+lan9118_phy_reset(void) ""
+
 # lance.c
 lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
 lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
 i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
 
 # imx_fec.c
-imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
 imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
-imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
 imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
-imx_phy_update_link(const char *s) "%s"
-imx_phy_reset(void) ""
 imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
 imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
 imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
fixes the MSB of selector field to be zero, as specified in the datasheet.

Fixes: 2a424990170b "LAN9118 emulation"
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-4-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
         val = s->advertise;
         break;
     case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
+        val = 0x0fe1;
         break;
     case 6: /* Auto-neg Expansion */
         val = 1;
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Prefer named constants over magic values for better readability.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-5-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/mii.h |  6 +++++
 hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
 2 files changed, 46 insertions(+), 23 deletions(-)

diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/mii.h
+++ b/include/hw/net/mii.h
@@ -XXX,XX +XXX,XX @@
 #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
 #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
 
+#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
 #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
 #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
 #define MII_ANAR_TXFD       (1 << 8)
@@ -XXX,XX +XXX,XX @@
 #define MII_ANAR_10FD       (1 << 6)
 #define MII_ANAR_10         (1 << 5)
 #define MII_ANAR_CSMACD     (1 << 0)
+#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
 
 #define MII_ANLPAR_ACK      (1 << 14)
 #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
@@ -XXX,XX +XXX,XX @@
 #define RTL8201CP_PHYID1    0x0000
 #define RTL8201CP_PHYID2    0x8201
 
+/* SMSC LAN9118 */
+#define SMSCLAN9118_PHYID1  0x0007
+#define SMSCLAN9118_PHYID2  0xc0d1
+
 /* RealTek 8211E */
 #define RTL8211E_PHYID1     0x001c
 #define RTL8211E_PHYID2     0xc915
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
 
 #include "qemu/osdep.h"
 #include "hw/net/lan9118_phy.h"
+#include "hw/net/mii.h"
 #include "hw/irq.h"
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
     uint16_t val;
 
     switch (reg) {
-    case 0: /* Basic Control */
+    case MII_BMCR:
         val = s->control;
         break;
-    case 1: /* Basic Status */
+    case MII_BMSR:
         val = s->status;
         break;
-    case 2: /* ID1 */
-        val = 0x0007;
+    case MII_PHYID1:
+        val = SMSCLAN9118_PHYID1;
         break;
-    case 3: /* ID2 */
-        val = 0xc0d1;
+    case MII_PHYID2:
+        val = SMSCLAN9118_PHYID2;
         break;
-    case 4: /* Auto-neg advertisement */
+    case MII_ANAR:
         val = s->advertise;
         break;
-    case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0fe1;
+    case MII_ANLPAR:
+        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
+              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
+              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
         break;
-    case 6: /* Auto-neg Expansion */
-        val = 1;
+    case MII_ANER:
+        val = MII_ANER_NWAY;
         break;
     case 29: /* Interrupt source. */
         val = s->ints;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
     trace_lan9118_phy_write(val, reg);
 
     switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
+    case MII_BMCR:
+        if (val & MII_BMCR_RESET) {
             lan9118_phy_reset(s);
         } else {
-            s->control = val & 0x7980;
+            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
+                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
+                                MII_BMCR_CTST);
             /* Complete autonegotiation immediately. */
-            if (val & 0x1000) {
-                s->status |= 0x0020;
+            if (val & MII_BMCR_AUTOEN) {
+                s->status |= MII_BMSR_AN_COMP;
             }
         }
         break;
-    case 4: /* Auto-neg advertisement */
-        s->advertise = (val & 0x2d7f) | 0x80;
+    case MII_ANAR:
+        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
+                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
+                               MII_ANAR_SELECT))
+                     | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
         trace_lan9118_phy_update_link("down");
-        s->status &= ~0x0024;
+        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
         s->ints |= PHY_INT_DOWN;
     } else {
         trace_lan9118_phy_update_link("up");
-        s->status |= 0x0024;
+        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
     }
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
 {
     trace_lan9118_phy_reset();
 
-    s->control = 0x3000;
-    s->status = 0x7809;
-    s->advertise = 0x01e1;
+    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
+    s->status = MII_BMSR_100TX_FD
+                | MII_BMSR_100TX_HD
+                | MII_BMSR_10T_FD
+                | MII_BMSR_10T_HD
+                | MII_BMSR_AUTONEG
+                | MII_BMSR_EXTCAP;
+    s->advertise = MII_ANAR_TXFD
+                   | MII_ANAR_TX
+                   | MII_ANAR_10FD
+                   | MII_ANAR_10
+                   | MII_ANAR_CSMACD;
     s->int_mask = 0;
     s->ints = 0;
     lan9118_phy_update_link(s, s->link_down);
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

The real device advertises this mode and the device model already advertises
100 mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
make the model more realistic.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-6-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
         break;
     case MII_ANAR:
         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
-                               MII_ANAR_SELECT))
+                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
+                               MII_ANAR_10 | MII_ANAR_SELECT))
                      | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
-- 
2.34.1

For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
Invalid for the multiplication of 0 by infinity.  Currently we handle
this in the per-architecture ifdef ladder in pickNaNMulAdd().
However, since this isn't really architecture specific we can hoist
it up to the generic code.

For the cases where the infzero test in pickNaNMulAdd was
returning 2, we can delete the check entirely and allow the
code to fall into the normal pick-a-NaN handling, because this
will return 2 anyway (input 'c' being the only NaN in this case).
For the cases where infzero was returning 3 to indicate "return
the default NaN", we must retain that "return 3".

For Arm, this looks like it might be a behaviour change because we
used to set float_flag_invalid | float_flag_invalid_imz only if C is
a quiet NaN.  However, it is not, because Arm target code never looks
at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
already raised float_flag_invalid via the "abc_mask &
float_cmask_snan" check in pick_nan_muladd.

For any target architecture using the "default implementation" at the
bottom of the ifdef, this is a behaviour change but will be fixing a
bug (where we failed to raise the Invalid exception for (0 * inf +
QNaN).  The architectures using the default case are:
 * hppa
 * i386
 * sh4
 * tricore

The x86, Tricore and SH4 CPU architecture manuals are clear that this
should have raised Invalid; HPPA is a bit vaguer but still seems
clear enough.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 13 +++++++------
 fpu/softfloat-specialize.c.inc | 29 +----------------------------
 2 files changed, 8 insertions(+), 34 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             int ab_mask, int abc_mask)
 {
     int which;
+    bool infzero = (ab_mask == float_cmask_infzero);
 
     if (unlikely(abc_mask & float_cmask_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
-                          ab_mask == float_cmask_infzero, s);
+    if (infzero) {
+        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
+        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+    }
+
+    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
 
     if (s->default_nan_mode || which == 3) {
-        /*
-         * Note that this check is after pickNaNMulAdd so that function
-         * has an opportunity to set the Invalid flag for infzero.
-         */
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * the default NaN
      */
     if (infzero && is_qnan(c_cls)) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * case sets InvalidOp and returns the default NaN
          */
         if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
             return 3;
         }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-            return 2;
-        }
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
      * case sets InvalidOp and returns the input value 'c'
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
+
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * to return an input NaN if we have one (ie c) rather than generating
      * a default NaN
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
 
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_RISCV)
-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-    }
     return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 2;
     }
 #elif defined(TARGET_SPARC)
-    /* For (inf,0,nan) return c. */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     /* Prefer SNaN over QNaN, order C, B, A. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
      * an input NaN if we have one (ie c).
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     if (status->use_first_nan) {
         if (is_nan(a_cls)) {
             return 0;
-- 
2.34.1

If the target sets default_nan_mode then we're always going to return
the default NaN, and pickNaNMulAdd() no longer has any side effects.
For consistency with pickNaN(), check for default_nan_mode before
calling pickNaNMulAdd().

When we convert pickNaNMulAdd() to allow runtime selection of the NaN
propagation rule, this means we won't have to make the targets which
use default_nan_mode also set a propagation rule.

Since RiscV always uses default_nan_mode, this allows us to remove
its ifdef case from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 8 ++++++--
 fpu/softfloat-specialize.c.inc | 9 +++++++--
 2 files changed, 13 insertions(+), 4 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    if (s->default_nan_mode) {
+        which = 3;
+    } else {
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    }
 
-    if (s->default_nan_mode || which == 3) {
+    if (which == 3) {
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    /*
+     * We guarantee not to require the target to tell us how to
+     * pick a NaN if we're always returning the default NaN.
+     * But if we're not in default-NaN mode then the target must
+     * specify.
+     */
+    assert(!status->default_nan_mode);
 #if defined(TARGET_ARM)
     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
      * the default NaN
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     } else {
         return 1;
     }
-#elif defined(TARGET_RISCV)
-    return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
         return 3;
-- 
2.34.1

IEEE 758 does not define a fixed rule for what NaN to return in
the case of a fused multiply-add of inf * 0 + NaN. Different
architectures thus do different things:
 * some return the default NaN
 * some return the input NaN
 * Arm returns the default NaN if the input NaN is quiet,
   and the input NaN if it is signalling

We want to make this logic be runtime selected rather than
hardcoded into the binary, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

In this commit we add an enum for the propagation rule, the field in
float_status, and the corresponding getters and setters.  We change
pickNaNMulAdd to honour this, but because all targets still leave
this field at its default 0 value, the fallback logic will pick the
rule type with the old ifdef ladder.

Note that four architectures both use the muladd softfloat functions
and did not have a branch of the ifdef ladder to specify their
behaviour (and so were ending up with the "default" case, probably
wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
default_nan_mode, and so will never get into pickNaNMulAdd().  For
HPPA and i386 we retain the same behaviour as the old default-case,
which is to not ever return the default NaN.  This might not be
correct but it is not a behaviour change.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 ++++
 include/fpu/softfloat-types.h   | 23 +++++++++
 fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
 3 files changed, 95 insertions(+), 30 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
+                                             float_status *status)
+{
+    status->float_infzeronan_rule = rule;
+}
+
 static inline void set_flush_to_zero(bool val, float_status *status)
 {
     status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
+{
+    return status->float_infzeronan_rule;
+}
+
 static inline bool get_flush_to_zero(float_status *status)
 {
     return status->flush_to_zero;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * Rule for result of fused multiply-add 0 * Inf + NaN.
+ * This must be a NaN, but implementations differ on whether this
+ * is the input NaN or the default NaN.
+ *
+ * You don't need to set this if default_nan_mode is enabled.
+ * When not in default-NaN mode, it is an error for the target
+ * not to set the rule in float_status if it uses muladd, and we
+ * will assert if we need to handle an input NaN and no rule was
+ * selected.
+ */
+typedef enum __attribute__((__packed__)) {
+    /* No propagation rule specified */
+    float_infzeronan_none = 0,
+    /* Result is never the default NaN (so always the input NaN) */
+    float_infzeronan_dnan_never,
+    /* Result is always the default NaN */
+    float_infzeronan_dnan_always,
+    /* Result is the default NaN if the input NaN is quiet */
+    float_infzeronan_dnan_if_qnan,
+} FloatInfZeroNaNRule;
+
 /*
  * Floating Point Status. Individual architectures may maintain
  * several versions of float_status for different functions. The
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
     bool flush_to_zero;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * specify.
      */
     assert(!status->default_nan_mode);
+
+    if (rule == float_infzeronan_none) {
+        /*
+         * Temporarily fall back to ifdef ladder
+         */
 #if defined(TARGET_ARM)
-    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
-     * the default NaN
-     */
-    if (infzero && is_qnan(c_cls)) {
-        return 3;
+        /*
+         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+         * but (inf,zero,snan) returns the input NaN.
+         */
+        rule = float_infzeronan_dnan_if_qnan;
+#elif defined(TARGET_MIPS)
+        if (snan_bit_is_one(status)) {
+            /*
+             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+             * case sets InvalidOp and returns the default NaN
+             */
+            rule = float_infzeronan_dnan_always;
+        } else {
+            /*
+             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+             * case sets InvalidOp and returns the input value 'c'
+             */
+            rule = float_infzeronan_dnan_never;
+        }
+#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+        /*
+         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+         * case sets InvalidOp and returns the input value 'c'
+         */
+        /*
+         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+         * to return an input NaN if we have one (ie c) rather than generating
+         * a default NaN
+         */
+        rule = float_infzeronan_dnan_never;
+#elif defined(TARGET_S390X)
+        rule = float_infzeronan_dnan_always;
+#endif
     }
 
+    if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the default NaN here,
+         * and some return the input NaN.
+         */
+        switch (rule) {
+        case float_infzeronan_dnan_never:
+            return 2;
+        case float_infzeronan_dnan_always:
+            return 3;
+        case float_infzeronan_dnan_if_qnan:
+            return is_qnan(c_cls) ? 3 : 2;
+        default:
+            g_assert_not_reached();
+        }
+    }
+
+#if defined(TARGET_ARM)
+
     /* This looks different from the ARM ARM pseudocode, because the ARM ARM
      * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 #elif defined(TARGET_MIPS)
     if (snan_bit_is_one(status)) {
-        /*
-         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-         * case sets InvalidOp and returns the default NaN
-         */
-        if (infzero) {
-            return 3;
-        }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
         if (is_snan(a_cls)) {
             return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
             return 2;
         }
     } else {
-        /*
-         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 #elif defined(TARGET_LOONGARCH64)
-    /*
-     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-     * case sets InvalidOp and returns the input value 'c'
-     */
-
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_PPC)
-    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-     * to return an input NaN if we have one (ie c) rather than generating
-     * a default NaN
-     */
-
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_S390X)
-    if (infzero) {
-        return 3;
-    }
-
     if (is_snan(a_cls)) {
         return 0;
     } else if (is_snan(b_cls)) {
-- 
2.34.1

Explicitly set a rule in the softfloat tests for the inf-zero-nan
muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
and so we should select here the Arm rule of
float_infzeronan_dnan_if_qnan.

Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 5 +++++
 tests/fp/fp-test.c  | 5 +++++
 2 files changed, 10 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
 {
     bench_func_t f;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
 {
     unsigned int i;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
     verCases_maxErrorCount = n_max_errors;
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the Arm target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 4 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+ *    and the input NaN if it is signalling
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_ARM)
-        /*
-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
-         * but (inf,zero,snan) returns the input NaN.
-         */
-        rule = float_infzeronan_dnan_if_qnan;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             /*
              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for s390, so we
can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                  &env->fpu_status);
        /* fall through */
     case RESET_TYPE_S390_CPU_NORMAL:
         env->psw.mask &= ~PSW_MASK_RI;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * a default NaN
          */
         rule = float_infzeronan_dnan_never;
-#elif defined(TARGET_S390X)
-        rule = float_infzeronan_dnan_always;
 #endif
     }
 
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the PPC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 7 +++++++
 fpu/softfloat-specialize.c.inc | 7 +------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+     * to return an input NaN if we have one (ie c) rather than generating
+     * a default NaN
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
 
     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
         ppc_spr_t *spr = &env->spr_cb[i];
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              */
             rule = float_infzeronan_dnan_never;
         }
-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+#elif defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        /*
-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-         * to return an input NaN if we have one (ie c) rather than generating
-         * a default NaN
-         */
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the MIPS target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       |  9 +++++++++
 target/mips/msa.c              |  4 ++++
 fpu/softfloat-specialize.c.inc | 16 +---------------
 3 files changed, 14 insertions(+), 15 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
 static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
+    FloatInfZeroNaNRule izn_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
+    /*
+     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+     * case sets InvalidOp and returns the default NaN.
+     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'.
+     */
+    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
 
     /* set proper signanling bit meaning ("1" means "quiet") */
     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
+
+    /* Inf * 0 + NaN returns the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                              &env->active_tc.msa_fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            /*
-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-             * case sets InvalidOp and returns the default NaN
-             */
-            rule = float_infzeronan_dnan_always;
-        } else {
-            /*
-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-             * case sets InvalidOp and returns the input value 'c'
-             */
-            rule = float_infzeronan_dnan_never;
-        }
-#elif defined(TARGET_SPARC) || \
+#if defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the SPARC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_SPARC) || \
-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the xtensa target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
---
 target/xtensa/cpu.c            | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/cpu.c
+++ b/target/xtensa/cpu.c
@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
     reset_mmu(env);
     cs->halted = env->runstall;
 #endif
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
     set_no_signaling_nans(!dfpu, &env->fp_status);
     xtensa_use_first_nan(env, !dfpu);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the x86 target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 7 +++++++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
+    /*
+     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
+     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
+     * specified -- for 0 * inf + NaN the input NaN is selected, and if
+     * there are multiple input NaNs they are selected in the order a, b, c.
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * Temporarily fall back to ifdef ladder
          */
 #if defined(TARGET_HPPA) || \
-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+    defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the loongarch target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 5 +++++
 fpu/softfloat-specialize.c.inc    | 7 +------
 2 files changed, 6 insertions(+), 6 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
                             &env->fp_status);
     set_flush_to_zero(0, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_HPPA) || \
-    defined(TARGET_LOONGARCH)
-        /*
-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
+#if defined(TARGET_HPPA)
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the HPPA target,
so we can remove the ifdef from pickNaNMulAdd().

As this is the last target to be converted to explicitly setting
the rule, we can remove the fallback code in pickNaNMulAdd()
entirely.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       |  2 ++
 fpu/softfloat-specialize.c.inc | 13 +------------
 2 files changed, 3 insertions(+), 12 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
-
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      */
     assert(!status->default_nan_mode);
 
-    if (rule == float_infzeronan_none) {
-        /*
-         * Temporarily fall back to ifdef ladder
-         */
-#if defined(TARGET_HPPA)
-        rule = float_infzeronan_dnan_never;
-#endif
-    }
-
     if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the default NaN here,
          * and some return the input NaN.
          */
-        switch (rule) {
+        switch (status->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
             return 2;
         case float_infzeronan_dnan_always:
-- 
2.34.1

The new implementation of pickNaNMulAdd() will find it convenient
to know whether at least one of the three arguments to the muladd
was a signaling NaN. We already calculate that in the caller,
so pass it in as a new bool have_snan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 5 +++--
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
 {
     int which;
     bool infzero = (ab_mask == float_cmask_infzero);
+    bool have_snan = (abc_mask & float_cmask_snan);
 
-    if (unlikely(abc_mask & float_cmask_snan)) {
+    if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     if (s->default_nan_mode) {
         which = 3;
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
 *----------------------------------------------------------------------------*/
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, float_status *status)
+                         bool infzero, bool have_snan, float_status *status)
 {
     /*
      * We guarantee not to require the target to tell us how to
-- 
2.34.1

IEEE 758 does not define a fixed rule for which NaN to pick as the
result if both operands of a 3-operand fused multiply-add operation
are NaNs.  As a result different architectures have ended up with
different rules for propagating NaNs.

QEMU currently hardcodes the NaN propagation logic into the binary
because pickNaNMulAdd() has an ifdef ladder for different targets.
We want to make the propagation rule instead be selectable at
runtime, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

It's valid not to set a propagation rule if default_nan_mode is
enabled, because in that case there's no need to pick a NaN; all the
callers of pickNaNMulAdd() catch this case and skip calling it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h |  11 +++
 include/fpu/softfloat-types.h   |  55 +++++++++++
 fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
 3 files changed, 107 insertions(+), 126 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
+                                            float_status *status)
+{
+    status->float_3nan_prop_rule = rule;
+}
+
 static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
                                              float_status *status)
 {
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
+{
+    return status->float_3nan_prop_rule;
+}
+
 static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 {
     return status->float_infzeronan_rule;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ this code that are retained.
 #ifndef SOFTFLOAT_TYPES_H
 #define SOFTFLOAT_TYPES_H
 
+#include "hw/registerfields.h"
+
 /*
  * Software IEC/IEEE floating-point types.
  */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * 3-input NaN propagation rule, for fused multiply-add. Individual
+ * architectures have different rules for which input NaN is
+ * propagated to the output when there is more than one NaN on the
+ * input.
+ *
+ * If default_nan_mode is enabled then it is valid not to set a NaN
+ * propagation rule, because the softfloat code guarantees not to try
+ * to pick a NaN to propagate in default NaN mode.  When not in
+ * default-NaN mode, it is an error for the target not to set the rule
+ * in float_status if it uses a muladd, and we will assert if we need
+ * to handle an input NaN and no rule was selected.
+ *
+ * The naming scheme for Float3NaNPropRule values is:
+ *  float_3nan_prop_s_abc:
+ *    = "Prefer SNaN over QNaN, then operand A over B over C"
+ *  float_3nan_prop_abc:
+ *    = "Prefer A over B over C regardless of SNaN vs QNAN"
+ *
+ * For QEMU, the multiply-add operation is A * B + C.
+ */
+
+/*
+ * We set the Float3NaNPropRule enum values up so we can select the
+ * right value in pickNaNMulAdd in a data driven way.
+ */
+FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
+FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
+FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
+FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
+
+#define PROPRULE(X, Y, Z) \
+    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
+
+typedef enum __attribute__((__packed__)) {
+    float_3nan_prop_none = 0,     /* No propagation rule specified */
+    float_3nan_prop_abc = PROPRULE(0, 1, 2),
+    float_3nan_prop_acb = PROPRULE(0, 2, 1),
+    float_3nan_prop_bac = PROPRULE(1, 0, 2),
+    float_3nan_prop_bca = PROPRULE(1, 2, 0),
+    float_3nan_prop_cab = PROPRULE(2, 0, 1),
+    float_3nan_prop_cba = PROPRULE(2, 1, 0),
+    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
+} Float3NaNPropRule;
+
+#undef PROPRULE
+
 /*
  * Rule for result of fused multiply-add 0 * Inf + NaN.
  * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    Float3NaNPropRule float_3nan_prop_rule;
     FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, bool have_snan, float_status *status)
 {
+    FloatClass cls[3] = { a_cls, b_cls, c_cls };
+    Float3NaNPropRule rule = status->float_3nan_prop_rule;
+    int which;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
+    if (rule == float_3nan_prop_none) {
 #if defined(TARGET_ARM)
-
-    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
-     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
-     */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        /*
+         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+         */
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_MIPS)
-    if (snan_bit_is_one(status)) {
-        /* Prefer sNaN over qNaN, in the a, b, c order. */
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
+        if (snan_bit_is_one(status)) {
+            rule = float_3nan_prop_s_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_s_cab;
         }
-    } else {
-        /* Prefer sNaN over qNaN, in the c, a, b order. */
-        if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-    }
 #elif defined(TARGET_LOONGARCH64)
-    /* Prefer sNaN over qNaN, in the c, a, b order. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
-    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(c_cls)) {
-        return 2;
-    } else {
-        return 1;
-    }
+        /*
+         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+         */
+        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
-    if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-    /* Prefer SNaN over QNaN, order C, B, A. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 0;
-    }
+        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-    /*
-     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
-     * an input NaN if we have one (ie c).
-     */
-    if (status->use_first_nan) {
-        if (is_nan(a_cls)) {
-            return 0;
-        } else if (is_nan(b_cls)) {
-            return 1;
+        if (status->use_first_nan) {
+            rule = float_3nan_prop_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_cba;
         }
-    } else {
-        if (is_nan(c_cls)) {
-            return 2;
-        } else if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-    }
 #else
-    /* A default implementation: prefer a to b to c.
-     * This is unlikely to actually match any real implementation.
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_abc;
 #endif
+    }
+
+    assert(rule != float_3nan_prop_none);
+    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+        /* We have at least one SNaN input and should prefer it */
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_snan(cls[which]));
+    } else {
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_nan(cls[which]));
+    }
+    return which;
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

Explicitly set a rule in the softfloat tests for propagating NaNs in
the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
so we should select here the Arm rule of float_3nan_prop_s_cab.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 1 +
 tests/fp/fp-test.c  | 1 +
 2 files changed, 2 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 5 +++++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 6 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
+ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
+ *    but note that for QEMU muladd is a * b + c, whereas for
+ *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
  */
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_ARM)
-        /*
-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
-         */
-        rule = float_3nan_prop_s_cab;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             rule = float_3nan_prop_s_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for loongarch, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 1 +
 fpu/softfloat-specialize.c.inc    | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
      * case sets InvalidOp and returns the input value 'c'
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_LOONGARCH64)
-        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
         /*
          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for PPC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 6 ------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * NaN propagation for fused multiply-add:
+     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
+     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+     * whereas QEMU labels the operands as (a * b) + c.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
     /*
      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
      * to return an input NaN if we have one (ie c) rather than generating
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_PPC)
-        /*
-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-         */
-        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
         rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for s390x, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 1 +
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
         set_float_infzeronan_rule(float_infzeronan_dnan_always,
                                   &env->fpu_status);
        /* fall through */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_S390X)
-        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
         rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for SPARC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
+    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_SPARC)
-        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       | 4 ++++
 target/mips/msa.c              | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 3 files changed, 8 insertions(+), 7 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
     FloatInfZeroNaNRule izn_rule;
+    Float3NaNPropRule nan3_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab,
                              &env->active_tc.msa_fp_status);
 
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
+                             &env->active_tc.msa_fp_status);
+
     /* clear float_status exception flags */
     set_float_exception_flags(0, &env->active_tc.msa_fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            rule = float_3nan_prop_s_abc;
-        } else {
-            rule = float_3nan_prop_s_cab;
-        }
-#elif defined(TARGET_XTENSA)
+#if defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for xtensa, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
---
 target/xtensa/fpu_helper.c     | 2 ++
 fpu/softfloat-specialize.c.inc | 8 --------
 2 files changed, 2 insertions(+), 8 deletions(-)

diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/fpu_helper.c
+++ b/target/xtensa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
     set_use_first_nan(use_first, &env->fp_status);
     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                              &env->fp_status);
+    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
+                             &env->fp_status);
 }
 
 void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_XTENSA)
-        if (status->use_first_nan) {
-            rule = float_3nan_prop_abc;
-        } else {
-            rule = float_3nan_prop_cba;
-        }
-#else
         rule = float_3nan_prop_abc;
-#endif
     }
 
     assert(rule != float_3nan_prop_none);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for i386.  We had no
i386-specific behaviour in the old ifdef ladder, so we were using the
default "prefer a then b then c" fallback; this is actually the
correct per-the-spec handling for i386.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c | 1 +
 1 file changed, 1 insertion(+)

Set the Float3NaNPropRule explicitly for HPPA, and remove the
ifdef from pickNaNMulAdd().

HPPA is the only target that was using the default branch of the
ifdef ladder (other targets either do not use muladd or set
default_nan_mode), so we can remove the ifdef fallback entirely now
(allowing the "rule not set" case to fall into the default of the
switch statement and assert).

We add a TODO note that the HPPA rule is probably wrong; this is
not a behavioural change for this refactoring.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 4 ----
 2 files changed, 8 insertions(+), 4 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * TODO: The HPPA architecture reference only documents its NaN
+     * propagation rule for 2-operand operations. Testing on real hardware
+     * might be necessary to confirm whether this order for muladd is correct.
+     * Not preferring the SNaN is almost certainly incorrect as it diverges
+     * from the documented rules for 2-operand operations.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
-    if (rule == float_3nan_prop_none) {
-        rule = float_3nan_prop_abc;
-    }
-
     assert(rule != float_3nan_prop_none);
     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
         /* We have at least one SNaN input and should prefer it */
-- 
2.34.1

The use_first_nan field in float_status was an xtensa-specific way to
select at runtime from two different NaN propagation rules.  Now that
xtensa is using the target-agnostic NaN propagation rule selection
that we've just added, we can remove use_first_nan, because there is
no longer any code that reads it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 5 -----
 include/fpu/softfloat-types.h   | 1 -
 target/xtensa/fpu_helper.c      | 1 -
 3 files changed, 7 deletions(-)

Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
to get the NaN bit pattern to reset the FPU registers. This
works because it happens that our implementation of
floatx80_default_nan() doesn't actually look at the float_status
pointer except for TARGET_MIPS. However, this isn't guaranteed,
and to be able to remove the ifdef in floatx80_default_nan()
we're going to need a real float_status here.

Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
earlier, and thus can pass it to floatx80_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
---
 target/m68k/cpu.c | 12 +++++++-----
 1 file changed, 7 insertions(+), 5 deletions(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
     CPUState *cs = CPU(obj);
     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
     CPUM68KState *env = cpu_env(cs);
-    floatx80 nan = floatx80_default_nan(NULL);
+    floatx80 nan;
     int i;
 
     if (mcc->parent_phases.hold) {
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
 #else
     cpu_m68k_set_sr(env, SR_S | SR_I);
 #endif
-    for (i = 0; i < 8; i++) {
-        env->fregs[i].d = nan;
-    }
-    cpu_m68k_set_fpcr(env, 0);
     /*
      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+
+    nan = floatx80_default_nan(&env->fp_status);
+    for (i = 0; i < 8; i++) {
+        env->fregs[i].d = nan;
+    }
+    cpu_m68k_set_fpcr(env, 0);
     env->fpsr = 0;
 
     /* TODO: We should set PC from the interrupt vector.  */
-- 
2.34.1

We create our 128-bit default NaN by calling parts64_default_nan()
and then adjusting the result.  We can do the same trick for creating
the floatx80 default NaN, which lets us drop a target ifdef.

floatx80 is used only by:
 i386
 m68k
 arm nwfpe old floating-point emulation emulation support
    (which is essentially dead, especially the parts involving floatx80)
 PPC (only in the xsrqpxp instruction, which just rounds an input
    value by converting to floatx80 and back, so will never generate
    the default NaN)

The floatx80 default NaN as currently implemented is:
 m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
 i386: sign = 1, exp = 1...1, int = 1, frac = 10...0

These are the same as the parts64_default_nan for these architectures.

This is technically a possible behaviour change for arm linux-user
nwfpe emulation emulation, because the default NaN will now have the
sign bit clear.  But we were already generating a different floatx80
default NaN from the real kernel emulation we are supposedly
following, which appears to use an all-bits-1 value:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267

This won't affect the only "real" use of the nwfpe emulation, which
is ancient binaries that used it as part of the old floating point
calling convention; that only uses loads and stores of 32 and 64 bit
floats, not any of the floatx80 behaviour the original hardware had.
We also get the nwfpe float64 default NaN value wrong:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
so if we ever cared about this obscure corner the right fix would be
to correct that so nwfpe used its own default-NaN setting rather
than the Arm VFP one.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
 floatx80 floatx80_default_nan(float_status *status)
 {
     floatx80 r;
+    /*
+     * Extrapolate from the choices made by parts64_default_nan to fill
+     * in the floatx80 format. We assume that floatx80's explicit
+     * integer bit is always set (this is true for i386 and m68k,
+     * which are the only real users of this format).
+     */
+    FloatParts64 p64;
+    parts64_default_nan(&p64, status);
 
-    /* None of the targets that have snan_bit_is_one use floatx80.  */
-    assert(!snan_bit_is_one(status));
-#if defined(TARGET_M68K)
-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
-    r.high = 0x7FFF;
-#else
-    /* X86 */
-    r.low = UINT64_C(0xC000000000000000);
-    r.high = 0xFFFF;
-#endif
+    r.high = 0x7FFF | (p64.sign << 15);
+    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
     return r;
 }
 
-- 
2.34.1

In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
a zero-initialized float_status struct to float32_is_quiet_nan() and
float64_is_quiet_nan(), with the cryptic comment "for
snan_bit_is_one".

This pattern appears to have been copied from target/riscv, where it
is used because the functions there do not have ready access to the
CPU state struct. The comment presumably refers to the fact that the
main reason the is_quiet_nan() functions want the float_state is
because they want to know about the snan_bit_is_one config.

In the loongarch helpers, though, we have the CPU state struct
to hand. Use the usual env->fp_status here. This avoids our needing
to track that we need to update the initializer of the local
float_status structs when the core softfloat code adds new
options for targets to configure their behaviour.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
     } else if (float32_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float32_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
     } else if (float64_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float64_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
-- 
2.34.1

In the frem helper, we have a local float_status because we want to
execute the floatx80_div() with a custom rounding mode.  Instead of
zero-initializing the local float_status and then having to set it up
with the m68k standard behaviour (including the NaN propagation rule
and copying the rounding precision from env->fp_status), initialize
it as a complete copy of env->fp_status. This will avoid our having
to add new code in this function for every new config knob we add
to fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
---
 target/m68k/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/fpu_helper.c
+++ b/target/m68k/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
 
     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
     if (!floatx80_is_any_nan(fp_rem)) {
-        float_status fp_status = { };
+        /* Use local temporary fp_status to set different rounding mode */
+        float_status fp_status = env->fp_status;
         uint32_t quotient;
         int sign;
 
         /* Calculate quotient directly using round to nearest mode */
-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
         set_float_rounding_mode(float_round_nearest_even, &fp_status);
-        set_floatx80_rounding_precision(
-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
 
         sign = extractFloatx80Sign(fp_quot.d);
-- 
2.34.1

In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
from float64 to floatx80 using a scratch float_status, because we
don't want the conversion to affect the CPU's floating point exception
status. Currently we use a zero-initialized float_status. This will
get steadily more awkward as we add config knobs to float_status
that the target must initialize. Avoid having to add any of that
configuration here by instead initializing our local float_status
from the env->fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
---
 target/m68k/helper.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/target/m68k/helper.c b/target/m68k/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/helper.c
+++ b/target/m68k/helper.c
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
     }
     switch (n) {
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
         return 8;
     }
-- 
2.34.1

In the helper functions flcmps and flcmpd we use a scratch float_status
so that we don't change the CPU state if the comparison raises any
floating point exception flags. Instead of zero-initializing this
scratch float_status, initialize it as a copy of env->fp_status. This
avoids the need to explicitly initialize settings like the NaN
propagation rule or others we might add to softfloat in future.

To do this we need to pass the CPU env pointer in to the helper.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
---
 target/sparc/helper.h     | 4 ++--
 target/sparc/fop_helper.c | 8 ++++----
 target/sparc/translate.c  | 4 ++--
 3 files changed, 8 insertions(+), 8 deletions(-)

diff --git a/target/sparc/helper.h b/target/sparc/helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/helper.h
+++ b/target/sparc/helper.h
@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
 DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
+DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
+DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
 DEF_HELPER_2(raise_exception, noreturn, env, int)
 
 DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/fop_helper.c
+++ b/target/sparc/fop_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
     return finish_fcmp(env, r, GETPC());
 }
 
-uint32_t helper_flcmps(float32 src1, float32 src2)
+uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
 {
     /*
      * FLCMP never raises an exception nor modifies any FSR fields.
      * Perform the comparison with a dummy fp environment.
      */
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
     g_assert_not_reached();
 }
 
-uint32_t helper_flcmpd(float64 src1, float64 src2)
+uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
 {
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
diff --git a/target/sparc/translate.c b/target/sparc/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/translate.c
+++ b/target/sparc/translate.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
 
     src1 = gen_load_fpr_F(dc, a->rs1);
     src2 = gen_load_fpr_F(dc, a->rs2);
-    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
 
     src1 = gen_load_fpr_D(dc, a->rs1);
     src2 = gen_load_fpr_D(dc, a->rs2);
-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
-- 
2.34.1

In the helper_compute_fprf functions, we pass a dummy float_status
in to the is_signaling_nan() function. This is unnecessary, because
we have convenient access to the CPU env pointer here and that
is already set up with the correct values for the snan_bit_is_one
and no_signaling_nans config settings. is_signaling_nan() doesn't
ever update the fp_status with any exception flags, so there is
no reason not to use env->fp_status here.

Use env->fp_status instead of the dummy fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
---
 target/ppc/fpu_helper.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/fpu_helper.c
+++ b/target/ppc/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
     } else if (tp##_is_infinity(arg)) {                           \
         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
     } else {                                                      \
-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
+        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
             fprf = 0x00 << FPSCR_FPRF;                            \
         } else {                                                  \
             fprf = 0x11 << FPSCR_FPRF;                            \
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Now that float_status has a bunch of fp parameters,
it is easier to copy an existing structure than create
one from scratch.  Begin by copying the structure that
corresponds to the FPSR and make only the adjustments
required for BFloat16 semantics.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/vec_helper.c | 20 +++++++-------------
 1 file changed, 7 insertions(+), 13 deletions(-)

diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/vec_helper.c
+++ b/target/arm/tcg/vec_helper.c
@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
      * no effect on AArch32 instructions.
      */
     bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
-    *statusp = (float_status){
-        .tininess_before_rounding = float_tininess_before_rounding,
-        .float_rounding_mode = float_round_to_odd_inf,
-        .flush_to_zero = true,
-        .flush_inputs_to_zero = true,
-        .default_nan_mode = true,
-    };
+
+    *statusp = env->vfp.fp_status;
+    set_default_nan_mode(true, statusp);
 
     if (ebf) {
-        float_status *fpst = &env->vfp.fp_status;
-        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
-        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
-        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
-
         /* EBF=1 needs to do a step with round-to-odd semantics */
         *oddstatusp = *statusp;
         set_float_rounding_mode(float_round_to_odd, oddstatusp);
+    } else {
+        set_flush_to_zero(true, statusp);
+        set_flush_inputs_to_zero(true, statusp);
+        set_float_rounding_mode(float_round_to_odd_inf, statusp);
     }
-
     return ebf;
 }
 
-- 
2.34.1

Currently we hardcode the default NaN value in parts64_default_nan()
using a compile-time ifdef ladder. This is awkward for two cases:
 * for single-QEMU-binary we can't hard-code target-specifics like this
 * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
   (specifically the sign bit is different)

Add a field to float_status to specify the default NaN value; fall
back to the old ifdef behaviour if these are not set.

The default NaN value is specified by setting a uint8_t to a
pattern corresponding to the sign and upper fraction parts of
the NaN; the lower bits of the fraction are set from bit 0 of
the pattern.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 +++++++
 include/fpu/softfloat-types.h   | 10 ++++++
 fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
 3 files changed, 54 insertions(+), 22 deletions(-)

Set the default NaN pattern explicitly for the tests/fp code.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c     | 1 +
 tests/fp/fp-test-log2.c | 1 +
 tests/fp/fp-test.c      | 1 +
 3 files changed, 3 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+    set_float_default_nan_pattern(0b01000000, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test-log2.c
+++ b/tests/fp/fp-test-log2.c
@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
     int i;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_rounding_mode(float_round_nearest_even, &qsf);
 
     test.d = 0.0;
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
---
 target/microblaze/cpu.c        | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/microblaze/cpu.c
+++ b/target/microblaze/cpu.c
@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
      * this architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
 
 #if defined(CONFIG_USER_ONLY)
     /* start in user mode with interrupts enabled.  */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
-    || defined(TARGET_MICROBLAZE)
+#elif defined(TARGET_I386) || defined(TARGET_X86_64)
         /* Sign bit set, most significant frac bit set */
         dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 4 ++++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
+    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
+    set_float_default_nan_pattern(0b11000000, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
-        /* Sign bit set, most significant frac bit set */
-        dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
         /* Sign bit clear, msb-1 frac bit set */
         dnan_pattern = 0b00100000;
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 2 ++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 2 insertions(+), 3 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN: sign bit clear, msb-1 frac bit set */
+    set_float_default_nan_pattern(0b00100000, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_HPPA)
-        /* Sign bit clear, msb-1 frac bit set */
-        dnan_pattern = 0b00100000;
 #elif defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
-- 
2.34.1

Set the default NaN pattern explicitly for the arm target.
This includes setting it for the old linux-user nwfpe emulation.
For nwfpe, our default doesn't match the real kernel, but we
avoid making a behaviour change in this commit.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
---
 linux-user/arm/nwfpe/fpa11.c | 5 +++++
 target/arm/cpu.c             | 2 ++
 2 files changed, 7 insertions(+)

diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
index XXXXXXX..XXXXXXX 100644
--- a/linux-user/arm/nwfpe/fpa11.c
+++ b/linux-user/arm/nwfpe/fpa11.c
@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
    * this late date.
    */
   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
+  /*
+   * Use the same default NaN value as Arm VFP. This doesn't match
+   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
+   */
+  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
 }
 
 void SetRoundingMode(const unsigned int opcode)
diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
+ *  * Default NaN has sign bit clear, msb frac bit set
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+    set_float_default_nan_pattern(0b01000000, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
-- 
2.34.1

Set the default NaN pattern explicitly for m68k.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
---
 target/m68k/cpu.c              | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+    /* Default NaN: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     nan = floatx80_default_nan(&env->fp_status);
     for (i = 0; i < 8; i++) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
+#if defined(TARGET_SPARC)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
 #elif defined(TARGET_HEXAGON)
-- 
2.34.1

Set the default NaN pattern explicitly for MIPS. Note that this
is our only target which currently changes the default NaN
at runtime (which it was previously doing indirectly when it
changed the snan_bit_is_one setting).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h | 7 +++++++
 target/mips/msa.c        | 3 +++
 2 files changed, 10 insertions(+)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+    /*
+     * With nan2008, the default NaN value has the sign bit clear and the
+     * frac msb set; with the older mode, the sign bit is clear, and all
+     * frac bits except the msb are set.
+     */
+    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
+                                  &env->active_fpu.fp_status);
 
 }
 
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     /* Inf * 0 + NaN returns the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never,
                               &env->active_tc.msa_fp_status);
+    /* Default NaN: sign bit clear, frac msb set */
+    set_float_default_nan_pattern(0b01000000,
+                                  &env->active_tc.msa_fp_status);
 }
-- 
2.34.1

Set the default NaN pattern explicitly for SPARC, and remove
the ifdef from parts64_default_nan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 5 +----
 2 files changed, 3 insertions(+), 4 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN value: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC)
-        /* Sign bit clear, all frac bits set */
-        dnan_pattern = 0b01111111;
-#elif defined(TARGET_HEXAGON)
+#if defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
 #else
-- 
2.34.1

Set the default NaN pattern explicitly for hexagon.
Remove the ifdef from parts64_default_nan(); the only
remaining unconverted targets all use the default case.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
---
 target/hexagon/cpu.c           | 2 ++
 fpu/softfloat-specialize.c.inc | 5 -----
 2 files changed, 2 insertions(+), 5 deletions(-)

diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hexagon/cpu.c
+++ b/target/hexagon/cpu.c
@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
 
     set_default_nan_mode(1, &env->fp_status);
     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
+    /* Default NaN value: sign bit set, all frac bits set */
+    set_float_default_nan_pattern(0b11111111, &env->fp_status);
 }
 
 static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_HEXAGON)
-        /* Sign bit set, all frac bits set. */
-        dnan_pattern = 0b11111111;
-#else
         /*
          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
             /* sign bit clear, set frac msb */
             dnan_pattern = 0b01000000;
         }
-#endif
     }
     assert(dnan_pattern != 0);
 
-- 
2.34.1

Now that all our targets have bene converted to explicitly specify
their pattern for the default NaN value we can remove the remaining
fallback code in parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 14 --------------
 1 file changed, 14 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint64_t frac;
     uint8_t dnan_pattern = status->default_nan_pattern;
 
-    if (dnan_pattern == 0) {
-        /*
-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
-         * do not have floating-point.
-         */
-        if (snan_bit_is_one(status)) {
-            /* sign bit clear, set all frac bits other than msb */
-            dnan_pattern = 0b00111111;
-        } else {
-            /* sign bit clear, set frac msb */
-            dnan_pattern = 0b01000000;
-        }
-    }
     assert(dnan_pattern != 0);
 
     sign = dnan_pattern >> 7;
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaNMulAdd into its only caller.  This makes
one assert redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
[PMM: keep comment from old code in new location]
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
 fpu/softfloat-specialize.c.inc | 54 ----------------------------------
 2 files changed, 40 insertions(+), 55 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     }
 
     if (s->default_nan_mode) {
+        /*
+         * We guarantee not to require the target to tell us how to
+         * pick a NaN if we're always returning the default NaN.
+         * But if we're not in default-NaN mode then the target must
+         * specify.
+         */
         which = 3;
+    } else if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the
+         * default NaN here, and some return the input NaN.
+         */
+        switch (s->float_infzeronan_rule) {
+        case float_infzeronan_dnan_never:
+            which = 2;
+            break;
+        case float_infzeronan_dnan_always:
+            which = 3;
+            break;
+        case float_infzeronan_dnan_if_qnan:
+            which = is_qnan(c->cls) ? 3 : 2;
+            break;
+        default:
+            g_assert_not_reached();
+        }
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
+        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        Float3NaNPropRule rule = s->float_3nan_prop_rule;
+
+        assert(rule != float_3nan_prop_none);
+        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+            /* We have at least one SNaN input and should prefer it */
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_snan(cls[which]));
+        } else {
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_nan(cls[which]));
+        }
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a three-input operation.
-| For the moment we assume that no CPU needs the 'larger significand'
-| information.
-| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-*----------------------------------------------------------------------------*/
-static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, bool have_snan, float_status *status)
-{
-    FloatClass cls[3] = { a_cls, b_cls, c_cls };
-    Float3NaNPropRule rule = status->float_3nan_prop_rule;
-    int which;
-
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify.
-     */
-    assert(!status->default_nan_mode);
-
-    if (infzero) {
-        /*
-         * Inf * 0 + NaN -- some implementations return the default NaN here,
-         * and some return the input NaN.
-         */
-        switch (status->float_infzeronan_rule) {
-        case float_infzeronan_dnan_never:
-            return 2;
-        case float_infzeronan_dnan_always:
-            return 3;
-        case float_infzeronan_dnan_if_qnan:
-            return is_qnan(c_cls) ? 3 : 2;
-        default:
-            g_assert_not_reached();
-        }
-    }
-
-    assert(rule != float_3nan_prop_none);
-    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-        /* We have at least one SNaN input and should prefer it */
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_snan(cls[which]));
-    } else {
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_nan(cls[which]));
-    }
-    return which;
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remove "3" as a special case for which and simply
branch to return the desired value.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          * But if we're not in default-NaN mode then the target must
          * specify.
          */
-        which = 3;
+        goto default_nan;
     } else if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          */
         switch (s->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
-            which = 2;
             break;
         case float_infzeronan_dnan_always:
-            which = 3;
-            break;
+            goto default_nan;
         case float_infzeronan_dnan_if_qnan:
-            which = is_qnan(c->cls) ? 3 : 2;
+            if (is_qnan(c->cls)) {
+                goto default_nan;
+            }
             break;
         default:
             g_assert_not_reached();
         }
+        which = 2;
     } else {
         FloatClass cls[3] = { a->cls, b->cls, c->cls };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
     }
 
-    if (which == 3) {
-        parts_default_nan(a, s);
-        return a;
-    }
-
     switch (which) {
     case 0:
         break;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         parts_silence_nan(a, s);
     }
     return a;
+
+ default_nan:
+    parts_default_nan(a, s);
+    return a;
 }
 
 /*
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Assign the pointer return value to 'a' directly,
rather than going through an intermediary index.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
 1 file changed, 10 insertions(+), 22 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             FloatPartsN *c, float_status *s,
                                             int ab_mask, int abc_mask)
 {
-    int which;
     bool infzero = (ab_mask == float_cmask_infzero);
     bool have_snan = (abc_mask & float_cmask_snan);
+    FloatPartsN *ret;
 
     if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         default:
             g_assert_not_reached();
         }
-        which = 2;
+        ret = c;
     } else {
-        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        FloatPartsN *val[3] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
         if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
             /* We have at least one SNaN input and should prefer it */
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_snan(cls[which]));
+            } while (!is_snan(ret->cls));
         } else {
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_nan(cls[which]));
+            } while (!is_nan(ret->cls));
         }
     }
 
-    switch (which) {
-    case 0:
-        break;
-    case 1:
-        a = b;
-        break;
-    case 2:
-        a = c;
-        break;
-    default:
-        g_assert_not_reached();
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 
  default_nan:
     parts_default_nan(a, s);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

While all indices into val[] should be in [0-2], the mask
applied is two bits.  To help static analysis see there is
no possibility of read beyond the end of the array, pad the
array to 4 entries, with the final being (implicitly) NULL.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
         ret = c;
     } else {
-        FloatPartsN *val[3] = { a, b, c };
+        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

This function is part of the public interface and
is not "specialized" to any target in any way.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
 fpu/softfloat-specialize.c.inc | 52 ----------------------------------
 2 files changed, 52 insertions(+), 52 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
     *zExpPtr = 1 - shiftCount;
 }
 
+/*----------------------------------------------------------------------------
+| Takes two extended double-precision floating-point values `a' and `b', one
+| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
+| `b' is a signaling NaN, the invalid exception is raised.
+*----------------------------------------------------------------------------*/
+
+floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
+{
+    bool aIsLargerSignificand;
+    FloatClass a_cls, b_cls;
+
+    /* This is not complete, but is good enough for pickNaN.  */
+    a_cls = (!floatx80_is_any_nan(a)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(a, status)
+             ? float_class_snan
+             : float_class_qnan);
+    b_cls = (!floatx80_is_any_nan(b)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(b, status)
+             ? float_class_snan
+             : float_class_qnan);
+
+    if (is_snan(a_cls) || is_snan(b_cls)) {
+        float_raise(float_flag_invalid, status);
+    }
+
+    if (status->default_nan_mode) {
+        return floatx80_default_nan(status);
+    }
+
+    if (a.low < b.low) {
+        aIsLargerSignificand = 0;
+    } else if (b.low < a.low) {
+        aIsLargerSignificand = 1;
+    } else {
+        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
+    }
+
+    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
+        if (is_snan(b_cls)) {
+            return floatx80_silence_nan(b, status);
+        }
+        return b;
+    } else {
+        if (is_snan(a_cls)) {
+            return floatx80_silence_nan(a, status);
+        }
+        return a;
+    }
+}
+
 /*----------------------------------------------------------------------------
 | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
 | and extended significand formed by the concatenation of `zSig0' and `zSig1',
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
     return a;
 }
 
-/*----------------------------------------------------------------------------
-| Takes two extended double-precision floating-point values `a' and `b', one
-| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
-| `b' is a signaling NaN, the invalid exception is raised.
-*----------------------------------------------------------------------------*/
-
-floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-{
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
-
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
-        return floatx80_default_nan(status);
-    }
-
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Unpacking and repacking the parts may be slightly more work
than we did before, but we get to reuse more code.  For a
code path handling exceptional values, this is an improvement.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c | 43 +++++--------------------------------------
 1 file changed, 5 insertions(+), 38 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
 
 floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 {
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
+    FloatParts128 pa, pb, *pr;
 
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
+    if (!floatx80_unpack_canonical(&pa, a, status) ||
+        !floatx80_unpack_canonical(&pb, b, status)) {
         return floatx80_default_nan(status);
     }
 
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
+    pr = parts_pick_nan(&pa, &pb, status);
+    return floatx80_round_pack_canonical(pr, status);
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaN into its only caller.  This makes one assert
redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
 fpu/softfloat-specialize.c.inc | 96 ----------------------------------
 2 files changed, 73 insertions(+), 105 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
 static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
+    int cmp, which;
+
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
     if (s->default_nan_mode) {
         parts_default_nan(a, s);
-    } else {
-        int cmp = frac_cmp(a, b);
-        if (cmp == 0) {
-            cmp = a->sign < b->sign;
-        }
+        return a;
+    }
 
-        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
-            a = b;
-        }
+    cmp = frac_cmp(a, b);
+    if (cmp == 0) {
+        cmp = a->sign < b->sign;
+    }
+
+    switch (s->float_2nan_prop_rule) {
+    case float_2nan_prop_s_ab:
         if (is_snan(a->cls)) {
-            parts_silence_nan(a, s);
+            which = 0;
+        } else if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_qnan(a->cls)) {
+            which = 0;
+        } else {
+            which = 1;
         }
+        break;
+    case float_2nan_prop_s_ba:
+        if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_snan(a->cls)) {
+            which = 0;
+        } else if (is_qnan(b->cls)) {
+            which = 1;
+        } else {
+            which = 0;
+        }
+        break;
+    case float_2nan_prop_ab:
+        which = is_nan(a->cls) ? 0 : 1;
+        break;
+    case float_2nan_prop_ba:
+        which = is_nan(b->cls) ? 1 : 0;
+        break;
+    case float_2nan_prop_x87:
+        /*
+         * This implements x87 NaN propagation rules:
+         * SNaN + QNaN => return the QNaN
+         * two SNaNs => return the one with the larger significand, silenced
+         * two QNaNs => return the one with the larger significand
+         * SNaN and a non-NaN => return the SNaN, silenced
+         * QNaN and a non-NaN => return the QNaN
+         *
+         * If we get down to comparing significands and they are the same,
+         * return the NaN with the positive sign bit (if any).
+         */
+        if (is_snan(a->cls)) {
+            if (is_snan(b->cls)) {
+                which = cmp > 0 ? 0 : 1;
+            } else {
+                which = is_qnan(b->cls) ? 1 : 0;
+            }
+        } else if (is_qnan(a->cls)) {
+            if (is_snan(b->cls) || !is_qnan(b->cls)) {
+                which = 0;
+            } else {
+                which = cmp > 0 ? 0 : 1;
+            }
+        } else {
+            which = 1;
+        }
+        break;
+    default:
+        g_assert_not_reached();
+    }
+
+    if (which) {
+        a = b;
+    }
+    if (is_snan(a->cls)) {
+        parts_silence_nan(a, s);
     }
     return a;
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a two-input operation.
-| IEEE754 doesn't specify all the details of this, so the
-| algorithm is target-specific.
-| The routine is passed various bits of information about the
-| two NaNs and should return 0 to select NaN a and 1 for NaN b.
-| Note that signalling NaNs are always squashed to quiet NaNs
-| by the caller, by calling floatXX_silence_nan() before
-| returning them.
-|
-| aIsLargerSignificand is only valid if both a and b are NaNs
-| of some kind, and is true if a has the larger significand,
-| or if both a and b have the same significand but a is
-| positive but b is negative. It is only needed for the x87
-| tie-break rule.
-*----------------------------------------------------------------------------*/
-
-static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-                   bool aIsLargerSignificand, float_status *status)
-{
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify via set_float_2nan_prop_rule().
-     */
-    assert(!status->default_nan_mode);
-
-    switch (status->float_2nan_prop_rule) {
-    case float_2nan_prop_s_ab:
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_s_ba:
-        if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_ab:
-        if (is_nan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_ba:
-        if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_x87:
-        /*
-         * This implements x87 NaN propagation rules:
-         * SNaN + QNaN => return the QNaN
-         * two SNaNs => return the one with the larger significand, silenced
-         * two QNaNs => return the one with the larger significand
-         * SNaN and a non-NaN => return the SNaN, silenced
-         * QNaN and a non-NaN => return the QNaN
-         *
-         * If we get down to comparing significands and they are the same,
-         * return the NaN with the positive sign bit (if any).
-         */
-        if (is_snan(a_cls)) {
-            if (is_snan(b_cls)) {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-            return is_qnan(b_cls) ? 1 : 0;
-        } else if (is_qnan(a_cls)) {
-            if (is_snan(b_cls) || !is_qnan(b_cls)) {
-                return 0;
-            } else {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-        } else {
-            return 1;
-        }
-    default:
-        g_assert_not_reached();
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remember if there was an SNaN, and use that to simplify
float_2nan_prop_s_{ab,ba} to only the snan component.
Then, fall through to the corresponding
float_2nan_prop_{ab,ba} case to handle any remaining
nans, which must be quiet.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
 1 file changed, 12 insertions(+), 20 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Move the fractional comparison to the end of the
float_2nan_prop_x87 case.  This is not required for
any other 2nan propagation rule.  Reorganize the
x87 case itself to break out of the switch when the
fractional comparison is not required.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 19 +++++++++----------
 1 file changed, 9 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
         return a;
     }
 
-    cmp = frac_cmp(a, b);
-    if (cmp == 0) {
-        cmp = a->sign < b->sign;
-    }
-
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          * return the NaN with the positive sign bit (if any).
          */
         if (is_snan(a->cls)) {
-            if (is_snan(b->cls)) {
-                which = cmp > 0 ? 0 : 1;
-            } else {
+            if (!is_snan(b->cls)) {
                 which = is_qnan(b->cls) ? 1 : 0;
+                break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
                 which = 0;
-            } else {
-                which = cmp > 0 ? 0 : 1;
+                break;
             }
         } else {
             which = 1;
+            break;
         }
+        cmp = frac_cmp(a, b);
+        if (cmp == 0) {
+            cmp = a->sign < b->sign;
+        }
+        which = cmp > 0 ? 0 : 1;
         break;
     default:
         g_assert_not_reached();
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Replace the "index" selecting between A and B with a result variable
of the proper type.  This improves clarity within the function.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
 1 file changed, 13 insertions(+), 15 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
     bool have_snan = false;
-    int cmp, which;
+    FloatPartsN *ret;
+    int cmp;
 
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
-            which = is_snan(a->cls) ? 0 : 1;
+            ret = is_snan(a->cls) ? a : b;
             break;
         }
         /* fall through */
     case float_2nan_prop_ab:
-        which = is_nan(a->cls) ? 0 : 1;
+        ret = is_nan(a->cls) ? a : b;
         break;
     case float_2nan_prop_s_ba:
         if (have_snan) {
-            which = is_snan(b->cls) ? 1 : 0;
+            ret = is_snan(b->cls) ? b : a;
             break;
         }
         /* fall through */
     case float_2nan_prop_ba:
-        which = is_nan(b->cls) ? 1 : 0;
+        ret = is_nan(b->cls) ? b : a;
         break;
     case float_2nan_prop_x87:
         /*
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          */
         if (is_snan(a->cls)) {
             if (!is_snan(b->cls)) {
-                which = is_qnan(b->cls) ? 1 : 0;
+                ret = is_qnan(b->cls) ? b : a;
                 break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
-                which = 0;
+                ret = a;
                 break;
             }
         } else {
-            which = 1;
+            ret = b;
             break;
         }
         cmp = frac_cmp(a, b);
         if (cmp == 0) {
             cmp = a->sign < b->sign;
         }
-        which = cmp > 0 ? 0 : 1;
+        ret = cmp > 0 ? a : b;
         break;
     default:
         g_assert_not_reached();
     }
 
-    if (which) {
-        a = b;
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 }
 
 static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-- 
2.34.1

From: Leif Lindholm <quic_llindhol@quicinc.com>

I'm migrating to Qualcomm's new open source email infrastructure, so
update my email address, and update the mailmap to match.

Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 +-
 .mailmap    | 5 +++--
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
 SBSA-REF
 M: Radoslaw Biernacki <rad@semihalf.com>
 M: Peter Maydell <peter.maydell@linaro.org>
-R: Leif Lindholm <quic_llindhol@quicinc.com>
+R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
 R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
 L: qemu-arm@nongnu.org
 S: Maintained
diff --git a/.mailmap b/.mailmap
index XXXXXXX..XXXXXXX 100644
--- a/.mailmap
+++ b/.mailmap
@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
 Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
 James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
 Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
 Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
 Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
 Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
-- 
2.34.1

From: Vikram Garhwal <vikram.garhwal@bytedance.com>

Previously, maintainer role was paused due to inactive email id. Commit id:
c009d715721861984c4987bcc78b7ee183e86d75.

Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
 
 Xilinx CAN
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 F: hw/net/can/xlnx-*
 F: include/hw/net/xlnx-*
@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
 CAN bus subsystem and hardware
 M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 W: https://canbus.pages.fel.cvut.cz/
 F: net/can/*
-- 
2.34.1