Series comparison

-[PULL 00/24] target-arm queue
+[PULL 00/72] target-arm queue
-The following changes since commit 59084feb256c617063e0dbe7e64821ae8852d7cf:
+First arm pullreq of the cycle; this is mostly my softfloat NaN
 handling series. (Lots more in my to-review queue, but I don't
 like pullreqs growing too close to a hundred patches at a time :-))
-  Merge tag 'pull-aspeed-20240709' of https://github.com/legoater/qemu into staging (2024-07-09 07:13:55 -0700)
+thanks
 -- PMM
 The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:
   Open 10.0 development tree (2024-12-10 17:41:17 +0000)
 are available in the Git repository at:
-  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20240711
+  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211
-for you to fetch changes up to 7f49089158a4db644fcbadfa90cd3d30a4868735:
+for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:
-  target/arm: Convert PMULL to decodetree (2024-07-11 11:41:34 +0100)
+  MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)
 ----------------------------------------------------------------
 target-arm queue:
- * Refactor FPCR/FPSR handling in preparation for FEAT_AFP
+ * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
- * More decodetree conversions
+ * fpu: Make muladd NaN handling runtime-selected, not compile-time
- * target/arm: Use cpu_env in cpu_untagged_addr
+ * fpu: Make default NaN pattern runtime-selected, not compile-time
- * target/arm: Set arm_v7m_tcg_ops cpu_exec_halt to arm_cpu_exec_halt()
+ * fpu: Minor NaN-related cleanups
- * hw/char/pl011: Avoid division-by-zero in pl011_get_baudrate()
+ * MAINTAINERS: email address updates
  * hw/misc/bcm2835_thermal: Fix access size handling in bcm2835_thermal_ops
  * accel/tcg: Make TCGCPUOps::cpu_exec_halt mandatory
  * STM32L4x5: Handle USART interrupts correctly
 ----------------------------------------------------------------
-Inès Varhol (3):
+Bernhard Beschow (5):
-      hw/misc: In STM32L4x5 EXTI, consolidate 2 constants
+      hw/net/lan9118: Extract lan9118_phy
-      hw/misc: In STM32L4x5 EXTI, handle direct interrupts
+      hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
-      hw/arm: In STM32L4x5 SOC, connect USART devices to EXTI
+      hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
       hw/net/lan9118_phy: Reuse MII constants
       hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
-Peter Maydell (12):
+Leif Lindholm (1):
-      target/arm: Correct comments about M-profile FPSCR
+      MAINTAINERS: update email address for Leif Lindholm
       target/arm: Make vfp_get_fpscr() call vfp_get_{fpcr, fpsr}
       target/arm: Make vfp_set_fpscr() call vfp_set_{fpcr, fpsr}
       target/arm: Support migration when FPSR/FPCR won't fit in the FPSCR
       target/arm: Implement store_cpu_field_low32() macro
       target/arm: Store FPSR and FPCR in separate CPU state fields
       target/arm: Rename FPCR_ QC, NZCV macros to FPSR_
       target/arm: Rename FPSR_MASK and FPCR_MASK and define them symbolically
       target/arm: Allow FPCR bits that aren't in FPSCR
       target/arm: Set arm_v7m_tcg_ops cpu_exec_halt to arm_cpu_exec_halt()
       target: Set TCGCPUOps::cpu_exec_halt to target's has_work implementation
       accel/tcg: Make TCGCPUOps::cpu_exec_halt mandatory
-Richard Henderson (7):
+Peter Maydell (54):
-      target/arm: Use cpu_env in cpu_untagged_addr
+      fpu: handle raising Invalid for infzero in pick_nan_muladd
-      target/arm: Convert SMULL, UMULL, SMLAL, UMLAL, SMLSL, UMLSL to decodetree
+      fpu: Check for default_nan_mode before calling pickNaNMulAdd
-      target/arm: Convert SADDL, SSUBL, SABDL, SABAL, and unsigned to decodetree
+      softfloat: Allow runtime choice of inf * 0 + NaN result
-      target/arm: Convert SQDMULL, SQDMLAL, SQDMLSL to decodetree
+      tests/fp: Explicitly set inf-zero-nan rule
-      target/arm: Convert SADDW, SSUBW, UADDW, USUBW to decodetree
+      target/arm: Set FloatInfZeroNaNRule explicitly
-      target/arm: Convert ADDHN, SUBHN, RADDHN, RSUBHN to decodetree
+      target/s390: Set FloatInfZeroNaNRule explicitly
-      target/arm: Convert PMULL to decodetree
+      target/ppc: Set FloatInfZeroNaNRule explicitly
       target/mips: Set FloatInfZeroNaNRule explicitly
       target/sparc: Set FloatInfZeroNaNRule explicitly
       target/xtensa: Set FloatInfZeroNaNRule explicitly
       target/x86: Set FloatInfZeroNaNRule explicitly
       target/loongarch: Set FloatInfZeroNaNRule explicitly
       target/hppa: Set FloatInfZeroNaNRule explicitly
       softfloat: Pass have_snan to pickNaNMulAdd
       softfloat: Allow runtime choice of NaN propagation for muladd
       tests/fp: Explicitly set 3-NaN propagation rule
       target/arm: Set Float3NaNPropRule explicitly
       target/loongarch: Set Float3NaNPropRule explicitly
       target/ppc: Set Float3NaNPropRule explicitly
       target/s390x: Set Float3NaNPropRule explicitly
       target/sparc: Set Float3NaNPropRule explicitly
       target/mips: Set Float3NaNPropRule explicitly
       target/xtensa: Set Float3NaNPropRule explicitly
       target/i386: Set Float3NaNPropRule explicitly
       target/hppa: Set Float3NaNPropRule explicitly
       fpu: Remove use_first_nan field from float_status
       target/m68k: Don't pass NULL float_status to floatx80_default_nan()
       softfloat: Create floatx80 default NaN from parts64_default_nan
       target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
       target/m68k: In frem helper, initialize local float_status from env->fp_status
       target/m68k: Init local float_status from env fp_status in gdb get/set reg
       target/sparc: Initialize local scratch float_status from env->fp_status
       target/ppc: Use env->fp_status in helper_compute_fprf functions
       fpu: Allow runtime choice of default NaN value
       tests/fp: Set default NaN pattern explicitly
       target/microblaze: Set default NaN pattern explicitly
       target/i386: Set default NaN pattern explicitly
       target/hppa: Set default NaN pattern explicitly
       target/alpha: Set default NaN pattern explicitly
       target/arm: Set default NaN pattern explicitly
       target/loongarch: Set default NaN pattern explicitly
       target/m68k: Set default NaN pattern explicitly
       target/mips: Set default NaN pattern explicitly
       target/openrisc: Set default NaN pattern explicitly
       target/ppc: Set default NaN pattern explicitly
       target/sh4: Set default NaN pattern explicitly
       target/rx: Set default NaN pattern explicitly
       target/s390x: Set default NaN pattern explicitly
       target/sparc: Set default NaN pattern explicitly
       target/xtensa: Set default NaN pattern explicitly
       target/hexagon: Set default NaN pattern explicitly
       target/riscv: Set default NaN pattern explicitly
       target/tricore: Set default NaN pattern explicitly
       fpu: Remove default handling for dnan_pattern
-Zheyu Ma (2):
+Richard Henderson (11):
-      hw/char/pl011: Avoid division-by-zero in pl011_get_baudrate()
+      target/arm: Copy entire float_status in is_ebf
-      hw/misc/bcm2835_thermal: Fix access size handling in bcm2835_thermal_ops
+      softfloat: Inline pickNaNMulAdd
       softfloat: Use goto for default nan case in pick_nan_muladd
       softfloat: Remove which from parts_pick_nan_muladd
       softfloat: Pad array size in pick_nan_muladd
       softfloat: Move propagateFloatx80NaN to softfloat.c
       softfloat: Use parts_pick_nan in propagateFloatx80NaN
       softfloat: Inline pickNaN
       softfloat: Share code between parts_pick_nan cases
       softfloat: Sink frac_cmp in parts_pick_nan until needed
       softfloat: Replace WHICH with RET in parts_pick_nan
- include/hw/core/tcg-cpu-ops.h     |    9 +-
+Vikram Garhwal (1):
- include/hw/misc/stm32l4x5_exti.h  |    4 +-
+      MAINTAINERS: Add correct email address for Vikram Garhwal
  target/arm/cpu.h                  |  113 ++--
  target/arm/internals.h            |    3 +
  target/arm/tcg/translate-a32.h    |    7 +
  target/arm/tcg/translate.h        |    3 +-
  target/riscv/internals.h          |    3 +
  target/arm/tcg/a64.decode         |   77 +++
  accel/tcg/cpu-exec.c              |   11 +-
  hw/arm/stm32l4x5_soc.c            |   24 +-
  hw/char/pl011.c                   |   13 +-
  hw/misc/bcm2835_thermal.c         |    2 +
  hw/misc/stm32l4x5_exti.c          |   13 +-
  target/alpha/cpu.c                |    1 +
  target/arm/cpu.c                  |    2 +-
  target/arm/machine.c              |  135 ++++-
  target/arm/tcg/cpu-v7m.c          |    1 +
  target/arm/tcg/mve_helper.c       |   12 +-
  target/arm/tcg/translate-a64.c    | 1155 ++++++++++++-------------------------
  target/arm/tcg/translate-m-nocp.c |   22 +-
  target/arm/tcg/translate-vfp.c    |    4 +-
  target/arm/vfp_helper.c           |  187 +++---
  target/avr/cpu.c                  |    1 +
  target/cris/cpu.c                 |    2 +
  target/hppa/cpu.c                 |    1 +
  target/loongarch/cpu.c            |    1 +
  target/m68k/cpu.c                 |    1 +
  target/microblaze/cpu.c           |    1 +
  target/mips/cpu.c                 |    1 +
  target/openrisc/cpu.c             |    1 +
  target/ppc/cpu_init.c             |    2 +
  target/riscv/cpu.c                |    2 +-
  target/riscv/tcg/tcg-cpu.c        |    2 +
  target/rx/cpu.c                   |    1 +
  target/s390x/cpu.c                |    1 +
  target/sh4/cpu.c                  |    1 +
  target/sparc/cpu.c                |    1 +
  target/tricore/cpu.c              |    1 +
  target/xtensa/cpu.c               |    1 +
 files changed, 893 insertions(+), 929 deletions(-)
+ MAINTAINERS                       |   4 +-
+ include/fpu/softfloat-helpers.h   |  38 +++-
+ include/fpu/softfloat-types.h     |  89 +++++++-
+ include/hw/net/imx_fec.h          |   9 +-
+ include/hw/net/lan9118_phy.h      |  37 ++++
+ include/hw/net/mii.h              |   6 +
+ target/mips/fpu_helper.h          |  20 ++
+ target/sparc/helper.h             |   4 +-
+ fpu/softfloat.c                   |  19 ++
+ hw/net/imx_fec.c                  | 146 ++------------
+ hw/net/lan9118.c                  | 137 ++-----------
+ hw/net/lan9118_phy.c              | 222 ++++++++++++++++++++
+ linux-user/arm/nwfpe/fpa11.c      |   5 +
+ target/alpha/cpu.c                |   2 +
+ target/arm/cpu.c                  |  10 +
+ target/arm/tcg/vec_helper.c       |  20 +-
+ target/hexagon/cpu.c              |   2 +
+ target/hppa/fpu_helper.c          |  12 ++
+ target/i386/tcg/fpu_helper.c      |  12 ++
+ target/loongarch/tcg/fpu_helper.c |  14 +-
+ target/m68k/cpu.c                 |  14 +-
+ target/m68k/fpu_helper.c          |   6 +-
+ target/m68k/helper.c              |   6 +-
+ target/microblaze/cpu.c           |   2 +
+ target/mips/msa.c                 |  10 +
+ target/openrisc/cpu.c             |   2 +
+ target/ppc/cpu_init.c             |  19 ++
+ target/ppc/fpu_helper.c           |   3 +-
+ target/riscv/cpu.c                |   2 +
+ target/rx/cpu.c                   |   2 +
+ target/s390x/cpu.c                |   5 +
+ target/sh4/cpu.c                  |   2 +
+ target/sparc/cpu.c                |   6 +
+ target/sparc/fop_helper.c         |   8 +-
+ target/sparc/translate.c          |   4 +-
+ target/tricore/helper.c           |   2 +
+ target/xtensa/cpu.c               |   4 +
+ target/xtensa/fpu_helper.c        |   3 +-
+ tests/fp/fp-bench.c               |   7 +
+ tests/fp/fp-test-log2.c           |   1 +
+ tests/fp/fp-test.c                |   7 +
+ fpu/softfloat-parts.c.inc         | 152 +++++++++++---
+ fpu/softfloat-specialize.c.inc    | 412 ++------------------------------------
+ .mailmap                          |   5 +-
+ hw/net/Kconfig                    |   5 +
+ hw/net/meson.build                |   1 +
+ hw/net/trace-events               |  10 +-
+files changed, 778 insertions(+), 730 deletions(-)
+ create mode 100644 include/hw/net/lan9118_phy.h
+ create mode 100644 hw/net/lan9118_phy.c

-[PULL 23/24] target/arm: Convert ADDHN, SUBHN, RADDHN, RSUBHN to decodetree
+[PULL 01/72] hw/net/lan9118: Extract lan9118_phy
-From: Richard Henderson <richard.henderson@linaro.org>
+From: Bernhard Beschow <shentey@gmail.com>
-Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+A very similar implementation of the same device exists in imx_fec. Prepare for
 a common implementation by extracting a device model into its own files.
 Some migration state has been moved into the new device model which breaks
 migration compatibility for the following machines:
 * smdkc210
 * realview-*
 * vexpress-*
 * kzm
 * mps2-*
 While breaking migration ABI, fix the size of the MII registers to be 16 bit,
 as defined by IEEE 802.3u.
 Signed-off-by: Bernhard Beschow <shentey@gmail.com>
 Tested-by: Guenter Roeck <linux@roeck-us.net>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20240709000610.382391-6-richard.henderson@linaro.org
+Message-id: 20241102125724.532843-2-shentey@gmail.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |   5 ++
+ include/hw/net/lan9118_phy.h |  37 ++++++++
- target/arm/tcg/translate-a64.c | 127 +++++++++++++++------------------
+ hw/net/lan9118.c             | 137 +++++-----------------------
-files changed, 61 insertions(+), 71 deletions(-)
+ hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
  hw/net/Kconfig               |   4 +
  hw/net/meson.build           |   1 +
 files changed, 233 insertions(+), 115 deletions(-)
  create mode 100644 include/hw/net/lan9118_phy.h
  create mode 100644 hw/net/lan9118_phy.c
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
 new file mode 100644
 index XXXXXXX..XXXXXXX
 --- /dev/null
 +++ b/include/hw/net/lan9118_phy.h
@@ -XXX,XX +XXX,XX @@
 +/*
 + * SMSC LAN9118 PHY emulation
 + *
 + * Copyright (c) 2009 CodeSourcery, LLC.
 + * Written by Paul Brook
 + *
 + * This work is licensed under the terms of the GNU GPL, version 2 or later.
 + * See the COPYING file in the top-level directory.
 + */
 +
 +#ifndef HW_NET_LAN9118_PHY_H
 +#define HW_NET_LAN9118_PHY_H
 +
 +#include "qom/object.h"
 +#include "hw/sysbus.h"
 +
 +#define TYPE_LAN9118_PHY "lan9118-phy"
 +OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
 +
 +typedef struct Lan9118PhyState {
 +    SysBusDevice parent_obj;
 +
 +    uint16_t status;
 +    uint16_t control;
 +    uint16_t advertise;
 +    uint16_t ints;
 +    uint16_t int_mask;
 +    qemu_irq irq;
 +    bool link_down;
 +} Lan9118PhyState;
 +
 +void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
 +void lan9118_phy_reset(Lan9118PhyState *s);
 +uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
 +void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
 +
 +#endif
 diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/hw/net/lan9118.c
-+++ b/target/arm/tcg/a64.decode
++++ b/hw/net/lan9118.c
-@@ -XXX,XX +XXX,XX @@ UADDW           0.10 1110 ..1 ..... 00010 0 ..... ..... @qrrr_e
+@@ -XXX,XX +XXX,XX @@
- SSUBW           0.00 1110 ..1 ..... 00110 0 ..... ..... @qrrr_e
+ #include "net/net.h"
- USUBW           0.10 1110 ..1 ..... 00110 0 ..... ..... @qrrr_e
+ #include "net/eth.h"
+ #include "hw/irq.h"
-+ADDHN           0.00 1110 ..1 ..... 01000 0 ..... ..... @qrrr_e
++#include "hw/net/lan9118_phy.h"
-+RADDHN          0.10 1110 ..1 ..... 01000 0 ..... ..... @qrrr_e
+ #include "hw/net/lan9118.h"
-+SUBHN           0.00 1110 ..1 ..... 01100 0 ..... ..... @qrrr_e
+ #include "hw/ptimer.h"
-+RSUBHN          0.10 1110 ..1 ..... 01100 0 ..... ..... @qrrr_e
+ #include "hw/qdev-properties.h"
-+
+@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
- ### Advanced SIMD scalar x indexed element
+ #define MAC_CR_RXEN     0x00000004
+ #define MAC_CR_RESERVED 0x7f404213
- FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
-diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
+-#define PHY_INT_ENERGYON            0x80
-index XXXXXXX..XXXXXXX 100644
+-#define PHY_INT_AUTONEG_COMPLETE    0x40
---- a/target/arm/tcg/translate-a64.c
+-#define PHY_INT_FAULT               0x20
-+++ b/target/arm/tcg/translate-a64.c
+-#define PHY_INT_DOWN                0x10
-@@ -XXX,XX +XXX,XX @@ TRANS(UADDW, do_addsub_wide, a, 0, false)
+-#define PHY_INT_AUTONEG_LP          0x08
- TRANS(SSUBW, do_addsub_wide, a, MO_SIGN, true)
+-#define PHY_INT_PARFAULT            0x04
- TRANS(USUBW, do_addsub_wide, a, 0, true)
+-#define PHY_INT_AUTONEG_PAGE        0x02
+-
-+static bool do_addsub_highnarrow(DisasContext *s, arg_qrrr_e *a,
+ #define GPT_TIMER_EN    0x20000000
-+                                 bool sub, bool round)
 +{
 +    TCGv_i64 tcg_op0, tcg_op1;
 +    MemOp esz = a->esz;
 +    int half = 8 >> esz;
 +    bool top = a->q;
 +    int ebits = 8 << esz;
 +    uint64_t rbit = 1ull << (ebits - 1);
 +    int top_swap, top_half;
 +
 +    /* There are no 128x128->64 bit operations. */
 +    if (esz >= MO_64) {
 +        return false;
 +    }
 +    if (!fp_access_check(s)) {
 +        return true;
 +    }
 +    tcg_op0 = tcg_temp_new_i64();
 +    tcg_op1 = tcg_temp_new_i64();
 +
 +    /*
 +     * For top half inputs, iterate backward; forward for bottom half.
 +     * This means the store to the destination will not occur until
 +     * overlapping input inputs are consumed.
 +     */
 +    top_swap = top ? half - 1 : 0;
 +    top_half = top ? half : 0;
 +
 +    for (int elt_fwd = 0; elt_fwd < half; ++elt_fwd) {
 +        int elt = elt_fwd ^ top_swap;
 +
 +        read_vec_element(s, tcg_op1, a->rm, elt, esz + 1);
 +        read_vec_element(s, tcg_op0, a->rn, elt, esz + 1);
 +        if (sub) {
 +            tcg_gen_sub_i64(tcg_op0, tcg_op0, tcg_op1);
 +        } else {
 +            tcg_gen_add_i64(tcg_op0, tcg_op0, tcg_op1);
 +        }
 +        if (round) {
 +            tcg_gen_addi_i64(tcg_op0, tcg_op0, rbit);
 +        }
 +        tcg_gen_shri_i64(tcg_op0, tcg_op0, ebits);
 +        write_vec_element(s, tcg_op0, a->rd, elt + top_half, esz);
 +    }
 +    clear_vec_high(s, top, a->rd);
 +    return true;
 +}
 +
 +TRANS(ADDHN, do_addsub_highnarrow, a, false, false)
 +TRANS(SUBHN, do_addsub_highnarrow, a, true, false)
 +TRANS(RADDHN, do_addsub_highnarrow, a, false, true)
 +TRANS(RSUBHN, do_addsub_highnarrow, a, true, true)
 +
  /*
-  * Advanced SIMD scalar/vector x indexed element
+@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
-  */
+     uint32_t mac_mii_data;
-@@ -XXX,XX +XXX,XX @@ static void disas_simd_shift_imm(DisasContext *s, uint32_t insn)
+     uint32_t mac_flow;
 -    uint32_t phy_status;
 -    uint32_t phy_control;
 -    uint32_t phy_advertise;
 -    uint32_t phy_int;
 -    uint32_t phy_int_mask;
 +    Lan9118PhyState mii;
 +    IRQState mii_irq;
      int32_t eeprom_writable;
      uint8_t eeprom[128];
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
  static const VMStateDescription vmstate_lan9118 = {
      .name = "lan9118",
 -    .version_id = 2,
 -    .minimum_version_id = 1,
 +    .version_id = 3,
 +    .minimum_version_id = 3,
      .fields = (const VMStateField[]) {
          VMSTATE_PTIMER(timer, lan9118_state),
          VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
          VMSTATE_UINT32(mac_mii_acc, lan9118_state),
          VMSTATE_UINT32(mac_mii_data, lan9118_state),
          VMSTATE_UINT32(mac_flow, lan9118_state),
 -        VMSTATE_UINT32(phy_status, lan9118_state),
 -        VMSTATE_UINT32(phy_control, lan9118_state),
 -        VMSTATE_UINT32(phy_advertise, lan9118_state),
 -        VMSTATE_UINT32(phy_int, lan9118_state),
 -        VMSTATE_UINT32(phy_int_mask, lan9118_state),
          VMSTATE_INT32(eeprom_writable, lan9118_state),
          VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
          VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
      lan9118_mac_changed(s);
  }
 -static void phy_update_irq(lan9118_state *s)
 +static void lan9118_update_irq(void *opaque, int n, int level)
  {
 -    if (s->phy_int & s->phy_int_mask) {
 +    lan9118_state *s = opaque;
 +
 +    if (level) {
          s->int_sts |= PHY_INT;
      } else {
          s->int_sts &= ~PHY_INT;
@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
      lan9118_update(s);
  }
 -static void phy_update_link(lan9118_state *s)
 -{
 -    /* Autonegotiation status mirrors link status.  */
 -    if (qemu_get_queue(s->nic)->link_down) {
 -        s->phy_status &= ~0x0024;
 -        s->phy_int |= PHY_INT_DOWN;
 -    } else {
 -        s->phy_status |= 0x0024;
 -        s->phy_int |= PHY_INT_ENERGYON;
 -        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
 -    }
 -    phy_update_irq(s);
 -}
 -
  static void lan9118_set_link(NetClientState *nc)
  {
 -    phy_update_link(qemu_get_nic_opaque(nc));
 -}
 -
 -static void phy_reset(lan9118_state *s)
 -{
 -    s->phy_status = 0x7809;
 -    s->phy_control = 0x3000;
 -    s->phy_advertise = 0x01e1;
 -    s->phy_int_mask = 0;
 -    s->phy_int = 0;
 -    phy_update_link(s);
 +    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
 +                            nc->link_down);
  }
  static void lan9118_reset(DeviceState *d)
@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
      s->read_word_n = 0;
      s->write_word_n = 0;
 -    phy_reset(s);
 -
      s->eeprom_writable = 0;
      lan9118_reload_eeprom(s);
  }
@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
      uint32_t status;
      /* FIXME: Honor TX disable, and allow queueing of packets.  */
 -    if (s->phy_control & 0x4000)  {
 +    if (s->mii.control & 0x4000) {
          /* This assumes the receive routine doesn't touch the VLANClient.  */
          qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
      } else {
@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
      }
  }
--/* Generate code to do a "long" addition or subtraction, ie one done in
+-static uint32_t do_phy_read(lan9118_state *s, int reg)
 - * TCGv_i64 on vector lanes twice the width specified by size.
 - */
 -static void gen_neon_addl(int size, bool is_sub, TCGv_i64 tcg_res,
 -                          TCGv_i64 tcg_op1, TCGv_i64 tcg_op2)
 -{
--    static NeonGenTwo64OpFn * const fns[3][2] = {
+-    uint32_t val;
--        { gen_helper_neon_addl_u16, gen_helper_neon_subl_u16 },
+-
--        { gen_helper_neon_addl_u32, gen_helper_neon_subl_u32 },
+-    switch (reg) {
--        { tcg_gen_add_i64, tcg_gen_sub_i64 },
+-    case 0: /* Basic Control */
--    };
+-        return s->phy_control;
--    NeonGenTwo64OpFn *genfn;
+-    case 1: /* Basic Status */
--    assert(size < 3);
+-        return s->phy_status;
--
+-    case 2: /* ID1 */
--    genfn = fns[size][is_sub];
+-        return 0x0007;
--    genfn(tcg_res, tcg_op1, tcg_op2);
+-    case 3: /* ID2 */
 -        return 0xc0d1;
 -    case 4: /* Auto-neg advertisement */
 -        return s->phy_advertise;
 -    case 5: /* Auto-neg Link Partner Ability */
 -        return 0x0f71;
 -    case 6: /* Auto-neg Expansion */
 -        return 1;
 -        /* TODO 17, 18, 27, 29, 30, 31 */
 -    case 29: /* Interrupt source.  */
 -        val = s->phy_int;
 -        s->phy_int = 0;
 -        phy_update_irq(s);
 -        return val;
 -    case 30: /* Interrupt mask */
 -        return s->phy_int_mask;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_read: PHY read reg %d\n", reg);
 -        return 0;
 -    }
 -}
 -
--static void do_narrow_round_high_u32(TCGv_i32 res, TCGv_i64 in)
+-static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
 -{
--    tcg_gen_addi_i64(in, in, 1U << 31);
+-    switch (reg) {
--    tcg_gen_extrh_i64_i32(res, in);
+-    case 0: /* Basic Control */
 -        if (val & 0x8000) {
 -            phy_reset(s);
 -            break;
 -        }
 -        s->phy_control = val & 0x7980;
 -        /* Complete autonegotiation immediately.  */
 -        if (val & 0x1000) {
 -            s->phy_status |= 0x0020;
 -        }
 -        break;
 -    case 4: /* Auto-neg advertisement */
 -        s->phy_advertise = (val & 0x2d7f) | 0x80;
 -        break;
 -        /* TODO 17, 18, 27, 31 */
 -    case 30: /* Interrupt mask */
 -        s->phy_int_mask = val & 0xff;
 -        phy_update_irq(s);
 -        break;
 -    default:
 -        qemu_log_mask(LOG_GUEST_ERROR,
 -                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 -    }
 -}
 -
--static void handle_3rd_narrowing(DisasContext *s, int is_q, int is_u, int size,
+ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
--                                 int opcode, int rd, int rn, int rm)
+ {
--{
+     switch (reg) {
--    TCGv_i32 tcg_res[2];
+@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
--    int part = is_q ? 2 : 0;
+         if (val & 2) {
--    int pass;
+             DPRINTF("PHY write %d = 0x%04x\n",
--
+                     (val >> 6) & 0x1f, s->mac_mii_data);
--    for (pass = 0; pass < 2; pass++) {
+-            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
--        TCGv_i64 tcg_op1 = tcg_temp_new_i64();
++            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
--        TCGv_i64 tcg_op2 = tcg_temp_new_i64();
+         } else {
--        TCGv_i64 tcg_wideres = tcg_temp_new_i64();
+-            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
--        static NeonGenNarrowFn * const narrowfns[3][2] = {
++            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
--            { gen_helper_neon_narrow_high_u8,
+             DPRINTF("PHY read %d = 0x%04x\n",
--              gen_helper_neon_narrow_round_high_u8 },
+                     (val >> 6) & 0x1f, s->mac_mii_data);
--            { gen_helper_neon_narrow_high_u16,
+         }
--              gen_helper_neon_narrow_round_high_u16 },
+@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
--            { tcg_gen_extrh_i64_i32, do_narrow_round_high_u32 },
+         break;
--        };
+     case CSR_PMT_CTRL:
--        NeonGenNarrowFn *gennarrow = narrowfns[size][is_u];
+         if (val & 0x400) {
--
+-            phy_reset(s);
--        read_vec_element(s, tcg_op1, rn, pass, MO_64);
++            lan9118_phy_reset(&s->mii);
--        read_vec_element(s, tcg_op2, rm, pass, MO_64);
+         }
--
+         s->pmt_ctrl &= ~0x34e;
--        gen_neon_addl(size, (opcode == 6), tcg_wideres, tcg_op1, tcg_op2);
+         s->pmt_ctrl |= (val & 0x34e);
--
+@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
--        tcg_res[pass] = tcg_temp_new_i32();
+     const MemoryRegionOps *mem_ops =
--        gennarrow(tcg_res[pass], tcg_wideres);
+             s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
--    }
--
++    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
--    for (pass = 0; pass < 2; pass++) {
++    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
--        write_vec_element_i32(s, tcg_res[pass], rd, pass + part, MO_32);
++    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
--    }
++        return;
--    clear_vec_high(s, is_q, rd);
++    }
--}
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
--
++
- /* AdvSIMD three different
+     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
-  *   31  30  29 28       24 23  22  21 20  16 15    12 11 10 9    5 4    0
+                           "lan9118-mmio", 0x100);
-  * +---+---+---+-----------+------+---+------+--------+-----+------+------+
+     sysbus_init_mmio(sbd, &s->mmio);
-@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
-     int rd = extract32(insn, 0, 5);
+new file mode 100644
+index XXXXXXX..XXXXXXX
-     switch (opcode) {
+--- /dev/null
--    case 4: /* ADDHN, ADDHN2, RADDHN, RADDHN2 */
++++ b/hw/net/lan9118_phy.c
--    case 6: /* SUBHN, SUBHN2, RSUBHN, RSUBHN2 */
+@@ -XXX,XX +XXX,XX @@
--        /* 128 x 128 -> 64 */
++/*
--        if (size == 3) {
++ * SMSC LAN9118 PHY emulation
--            unallocated_encoding(s);
++ *
--            return;
++ * Copyright (c) 2009 CodeSourcery, LLC.
--        }
++ * Written by Paul Brook
--        if (!fp_access_check(s)) {
++ *
--            return;
++ * This code is licensed under the GNU GPL v2
--        }
++ *
--        handle_3rd_narrowing(s, is_q, is_u, size, opcode, rd, rn, rm);
++ * Contributions after 2012-01-13 are licensed under the terms of the
--        break;
++ * GNU GPL, version 2 or (at your option) any later version.
-     case 14: /* PMULL, PMULL2 */
++ */
-         if (is_u) {
++
-             unallocated_encoding(s);
++#include "qemu/osdep.h"
-@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
++#include "hw/net/lan9118_phy.h"
-     case 1: /* SADDW, SADDW2, UADDW, UADDW2 */
++#include "hw/irq.h"
-     case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
++#include "hw/resettable.h"
-     case 3: /* SSUBW, SSUBW2, USUBW, USUBW2 */
++#include "migration/vmstate.h"
-+    case 4: /* ADDHN, ADDHN2, RADDHN, RADDHN2 */
++#include "qemu/log.h"
-     case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
++
-+    case 6: /* SUBHN, SUBHN2, RSUBHN, RSUBHN2 */
++#define PHY_INT_ENERGYON            (1 << 7)
-     case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
++#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
-     case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
++#define PHY_INT_FAULT               (1 << 5)
-     case 9: /* SQDMLAL, SQDMLAL2 */
++#define PHY_INT_DOWN                (1 << 4)
 +#define PHY_INT_AUTONEG_LP          (1 << 3)
 +#define PHY_INT_PARFAULT            (1 << 2)
 +#define PHY_INT_AUTONEG_PAGE        (1 << 1)
 +
 +static void lan9118_phy_update_irq(Lan9118PhyState *s)
 +{
 +    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
 +}
 +
 +uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
 +{
 +    uint16_t val;
 +
 +    switch (reg) {
 +    case 0: /* Basic Control */
 +        return s->control;
 +    case 1: /* Basic Status */
 +        return s->status;
 +    case 2: /* ID1 */
 +        return 0x0007;
 +    case 3: /* ID2 */
 +        return 0xc0d1;
 +    case 4: /* Auto-neg advertisement */
 +        return s->advertise;
 +    case 5: /* Auto-neg Link Partner Ability */
 +        return 0x0f71;
 +    case 6: /* Auto-neg Expansion */
 +        return 1;
 +        /* TODO 17, 18, 27, 29, 30, 31 */
 +    case 29: /* Interrupt source. */
 +        val = s->ints;
 +        s->ints = 0;
 +        lan9118_phy_update_irq(s);
 +        return val;
 +    case 30: /* Interrupt mask */
 +        return s->int_mask;
 +    default:
 +        qemu_log_mask(LOG_GUEST_ERROR,
 +                      "lan9118_phy_read: PHY read reg %d\n", reg);
 +        return 0;
 +    }
 +}
 +
 +void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
 +{
 +    switch (reg) {
 +    case 0: /* Basic Control */
 +        if (val & 0x8000) {
 +            lan9118_phy_reset(s);
 +            break;
 +        }
 +        s->control = val & 0x7980;
 +        /* Complete autonegotiation immediately. */
 +        if (val & 0x1000) {
 +            s->status |= 0x0020;
 +        }
 +        break;
 +    case 4: /* Auto-neg advertisement */
 +        s->advertise = (val & 0x2d7f) | 0x80;
 +        break;
 +        /* TODO 17, 18, 27, 31 */
 +    case 30: /* Interrupt mask */
 +        s->int_mask = val & 0xff;
 +        lan9118_phy_update_irq(s);
 +        break;
 +    default:
 +        qemu_log_mask(LOG_GUEST_ERROR,
 +                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
 +    }
 +}
 +
 +void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 +{
 +    s->link_down = link_down;
 +
 +    /* Autonegotiation status mirrors link status. */
 +    if (link_down) {
 +        s->status &= ~0x0024;
 +        s->ints |= PHY_INT_DOWN;
 +    } else {
 +        s->status |= 0x0024;
 +        s->ints |= PHY_INT_ENERGYON;
 +        s->ints |= PHY_INT_AUTONEG_COMPLETE;
 +    }
 +    lan9118_phy_update_irq(s);
 +}
 +
 +void lan9118_phy_reset(Lan9118PhyState *s)
 +{
 +    s->control = 0x3000;
 +    s->status = 0x7809;
 +    s->advertise = 0x01e1;
 +    s->int_mask = 0;
 +    s->ints = 0;
 +    lan9118_phy_update_link(s, s->link_down);
 +}
 +
 +static void lan9118_phy_reset_hold(Object *obj, ResetType type)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    lan9118_phy_reset(s);
 +}
 +
 +static void lan9118_phy_init(Object *obj)
 +{
 +    Lan9118PhyState *s = LAN9118_PHY(obj);
 +
 +    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
 +}
 +
 +static const VMStateDescription vmstate_lan9118_phy = {
 +    .name = "lan9118-phy",
 +    .version_id = 1,
 +    .minimum_version_id = 1,
 +    .fields = (const VMStateField[]) {
 +        VMSTATE_UINT16(control, Lan9118PhyState),
 +        VMSTATE_UINT16(status, Lan9118PhyState),
 +        VMSTATE_UINT16(advertise, Lan9118PhyState),
 +        VMSTATE_UINT16(ints, Lan9118PhyState),
 +        VMSTATE_UINT16(int_mask, Lan9118PhyState),
 +        VMSTATE_BOOL(link_down, Lan9118PhyState),
 +        VMSTATE_END_OF_LIST()
 +    }
 +};
 +
 +static void lan9118_phy_class_init(ObjectClass *klass, void *data)
 +{
 +    ResettableClass *rc = RESETTABLE_CLASS(klass);
 +    DeviceClass *dc = DEVICE_CLASS(klass);
 +
 +    rc->phases.hold = lan9118_phy_reset_hold;
 +    dc->vmsd = &vmstate_lan9118_phy;
 +}
 +
 +static const TypeInfo types[] = {
 +    {
 +        .name          = TYPE_LAN9118_PHY,
 +        .parent        = TYPE_SYS_BUS_DEVICE,
 +        .instance_size = sizeof(Lan9118PhyState),
 +        .instance_init = lan9118_phy_init,
 +        .class_init    = lan9118_phy_class_init,
 +    }
 +};
 +
 +DEFINE_TYPES(types)
 diff --git a/hw/net/Kconfig b/hw/net/Kconfig
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/Kconfig
 +++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
  config SMC91C111
      bool
 +config LAN9118_PHY
 +    bool
 +
  config LAN9118
      bool
 +    select LAN9118_PHY
      select PTIMER
  config NE2000_ISA
 diff --git a/hw/net/meson.build b/hw/net/meson.build
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/net/meson.build
 +++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
  system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
  system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
 +system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
  system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
  system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
  system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
 --
 .34.1

-New patch
+[PULL 02/72] hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
+From: Bernhard Beschow <shentey@gmail.com>
+imx_fec models the same PHY as lan9118_phy. The code is almost the same with
+imx_fec having more logging and tracing. Merge these improvements into
+lan9118_phy and reuse in imx_fec to fix the code duplication.
+Some migration state how resides in the new device model which breaks migration
+compatibility for the following machines:
+* imx25-pdk
+* sabrelite
+* mcimx7d-sabre
+* mcimx6ul-evk
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241102125724.532843-3-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ include/hw/net/imx_fec.h |   9 ++-
+ hw/net/imx_fec.c         | 146 ++++-----------------------------------
+ hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
+ hw/net/Kconfig           |   1 +
+ hw/net/trace-events      |  10 +--
+files changed, 85 insertions(+), 163 deletions(-)
+diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/hw/net/imx_fec.h
++++ b/include/hw/net/imx_fec.h
+@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
+ #define TYPE_IMX_ENET "imx.enet"
+ #include "hw/sysbus.h"
++#include "hw/net/lan9118_phy.h"
++#include "hw/irq.h"
+ #include "net/net.h"
+ #define ENET_EIR               1
+@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
+     uint32_t tx_descriptor[ENET_TX_RING_NUM];
+     uint32_t tx_ring_num;
+-    uint32_t phy_status;
+-    uint32_t phy_control;
+-    uint32_t phy_advertise;
+-    uint32_t phy_int;
+-    uint32_t phy_int_mask;
++    Lan9118PhyState mii;
++    IRQState mii_irq;
+     uint32_t phy_num;
+     bool phy_connected;
+     struct IMXFECState *phy_consumer;
+diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/imx_fec.c
++++ b/hw/net/imx_fec.c
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
+ static const VMStateDescription vmstate_imx_eth = {
+     .name = TYPE_IMX_FEC,
+-    .version_id = 2,
+-    .minimum_version_id = 2,
++    .version_id = 3,
++    .minimum_version_id = 3,
+     .fields = (const VMStateField[]) {
+         VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
+         VMSTATE_UINT32(rx_descriptor, IMXFECState),
+         VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
+-        VMSTATE_UINT32(phy_status, IMXFECState),
+-        VMSTATE_UINT32(phy_control, IMXFECState),
+-        VMSTATE_UINT32(phy_advertise, IMXFECState),
+-        VMSTATE_UINT32(phy_int, IMXFECState),
+-        VMSTATE_UINT32(phy_int_mask, IMXFECState),
+         VMSTATE_END_OF_LIST()
+     },
+     .subsections = (const VMStateDescription * const []) {
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
+     },
+ };
+-#define PHY_INT_ENERGYON            (1 << 7)
+-#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+-#define PHY_INT_FAULT               (1 << 5)
+-#define PHY_INT_DOWN                (1 << 4)
+-#define PHY_INT_AUTONEG_LP          (1 << 3)
+-#define PHY_INT_PARFAULT            (1 << 2)
+-#define PHY_INT_AUTONEG_PAGE        (1 << 1)
+-
+ static void imx_eth_update(IMXFECState *s);
+ /*
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
+  * For now we don't handle any GPIO/interrupt line, so the OS will
+  * have to poll for the PHY status.
+  */
+-static void imx_phy_update_irq(IMXFECState *s)
++static void imx_phy_update_irq(void *opaque, int n, int level)
+ {
+-    imx_eth_update(s);
+-}
+-
+-static void imx_phy_update_link(IMXFECState *s)
+-{
+-    /* Autonegotiation status mirrors link status.  */
+-    if (qemu_get_queue(s->nic)->link_down) {
+-        trace_imx_phy_update_link("down");
+-        s->phy_status &= ~0x0024;
+-        s->phy_int |= PHY_INT_DOWN;
+-    } else {
+-        trace_imx_phy_update_link("up");
+-        s->phy_status |= 0x0024;
+-        s->phy_int |= PHY_INT_ENERGYON;
+-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
+-    }
+-    imx_phy_update_irq(s);
++    imx_eth_update(opaque);
+ }
+ static void imx_eth_set_link(NetClientState *nc)
+ {
+-    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
+-}
+-
+-static void imx_phy_reset(IMXFECState *s)
+-{
+-    trace_imx_phy_reset();
+-
+-    s->phy_status = 0x7809;
+-    s->phy_control = 0x3000;
+-    s->phy_advertise = 0x01e1;
+-    s->phy_int_mask = 0;
+-    s->phy_int = 0;
+-    imx_phy_update_link(s);
++    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
++                            nc->link_down);
+ }
+ static uint32_t imx_phy_read(IMXFECState *s, int reg)
+ {
+-    uint32_t val;
+     uint32_t phy = reg / 32;
+     if (!s->phy_connected) {
+@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
+     reg %= 32;
+-    switch (reg) {
+-    case 0:     /* Basic Control */
+-        val = s->phy_control;
+-        break;
+-    case 1:     /* Basic Status */
+-        val = s->phy_status;
+-        break;
+-    case 2:     /* ID1 */
+-        val = 0x0007;
+-        break;
+-    case 3:     /* ID2 */
+-        val = 0xc0d1;
+-        break;
+-    case 4:     /* Auto-neg advertisement */
+-        val = s->phy_advertise;
+-        break;
+-    case 5:     /* Auto-neg Link Partner Ability */
+-        val = 0x0f71;
+-        break;
+-    case 6:     /* Auto-neg Expansion */
+-        val = 1;
+-        break;
+-    case 29:    /* Interrupt source.  */
+-        val = s->phy_int;
+-        s->phy_int = 0;
+-        imx_phy_update_irq(s);
+-        break;
+-    case 30:    /* Interrupt mask */
+-        val = s->phy_int_mask;
+-        break;
+-    case 17:
+-    case 18:
+-    case 27:
+-    case 31:
+-        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        val = 0;
+-        break;
+-    default:
+-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        val = 0;
+-        break;
+-    }
+-
+-    trace_imx_phy_read(val, phy, reg);
+-
+-    return val;
++    return lan9118_phy_read(&s->mii, reg);
+ }
+ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
+@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
+     reg %= 32;
+-    trace_imx_phy_write(val, phy, reg);
+-
+-    switch (reg) {
+-    case 0:     /* Basic Control */
+-        if (val & 0x8000) {
+-            imx_phy_reset(s);
+-        } else {
+-            s->phy_control = val & 0x7980;
+-            /* Complete autonegotiation immediately.  */
+-            if (val & 0x1000) {
+-                s->phy_status |= 0x0020;
+-            }
+-        }
+-        break;
+-    case 4:     /* Auto-neg advertisement */
+-        s->phy_advertise = (val & 0x2d7f) | 0x80;
+-        break;
+-    case 30:    /* Interrupt mask */
+-        s->phy_int_mask = val & 0xff;
+-        imx_phy_update_irq(s);
+-        break;
+-    case 17:
+-    case 18:
+-    case 27:
+-    case 31:
+-        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        break;
+-    default:
+-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
+-                      TYPE_IMX_FEC, __func__, reg);
+-        break;
+-    }
++    lan9118_phy_write(&s->mii, reg, val);
+ }
+ static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
+     s->rx_descriptor = 0;
+     memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
+-
+-    /* We also reset the PHY */
+-    imx_phy_reset(s);
+ }
+ static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
+@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
+     sysbus_init_irq(sbd, &s->irq[0]);
+     sysbus_init_irq(sbd, &s->irq[1]);
++    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
++    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
++    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
++        return;
++    }
++    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
++
+     qemu_macaddr_default_if_unset(&s->conf.macaddr);
+     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@
+  * Copyright (c) 2009 CodeSourcery, LLC.
+  * Written by Paul Brook
+  *
++ * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
++ *
+  * This code is licensed under the GNU GPL v2
+  *
+  * Contributions after 2012-01-13 are licensed under the terms of the
+@@ -XXX,XX +XXX,XX @@
+ #include "hw/resettable.h"
+ #include "migration/vmstate.h"
+ #include "qemu/log.h"
++#include "trace.h"
+ #define PHY_INT_ENERGYON            (1 << 7)
+ #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+     switch (reg) {
+     case 0: /* Basic Control */
+-        return s->control;
++        val = s->control;
++        break;
+     case 1: /* Basic Status */
+-        return s->status;
++        val = s->status;
++        break;
+     case 2: /* ID1 */
+-        return 0x0007;
++        val = 0x0007;
++        break;
+     case 3: /* ID2 */
+-        return 0xc0d1;
++        val = 0xc0d1;
++        break;
+     case 4: /* Auto-neg advertisement */
+-        return s->advertise;
++        val = s->advertise;
++        break;
+     case 5: /* Auto-neg Link Partner Ability */
+-        return 0x0f71;
++        val = 0x0f71;
++        break;
+     case 6: /* Auto-neg Expansion */
+-        return 1;
+-        /* TODO 17, 18, 27, 29, 30, 31 */
++        val = 1;
++        break;
+     case 29: /* Interrupt source. */
+         val = s->ints;
+         s->ints = 0;
+         lan9118_phy_update_irq(s);
+-        return val;
++        break;
+     case 30: /* Interrupt mask */
+-        return s->int_mask;
++        val = s->int_mask;
++        break;
++    case 17:
++    case 18:
++    case 27:
++    case 31:
++        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
++                      __func__, reg);
++        val = 0;
++        break;
+     default:
+-        qemu_log_mask(LOG_GUEST_ERROR,
+-                      "lan9118_phy_read: PHY read reg %d\n", reg);
+-        return 0;
++        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
++                      __func__, reg);
++        val = 0;
++        break;
+     }
++
++    trace_lan9118_phy_read(val, reg);
++
++    return val;
+ }
+ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+ {
++    trace_lan9118_phy_write(val, reg);
++
+     switch (reg) {
+     case 0: /* Basic Control */
+         if (val & 0x8000) {
+             lan9118_phy_reset(s);
+-            break;
+-        }
+-        s->control = val & 0x7980;
+-        /* Complete autonegotiation immediately. */
+-        if (val & 0x1000) {
+-            s->status |= 0x0020;
++        } else {
++            s->control = val & 0x7980;
++            /* Complete autonegotiation immediately. */
++            if (val & 0x1000) {
++                s->status |= 0x0020;
++            }
+         }
+         break;
+     case 4: /* Auto-neg advertisement */
+         s->advertise = (val & 0x2d7f) | 0x80;
+         break;
+-        /* TODO 17, 18, 27, 31 */
+     case 30: /* Interrupt mask */
+         s->int_mask = val & 0xff;
+         lan9118_phy_update_irq(s);
+         break;
++    case 17:
++    case 18:
++    case 27:
++    case 31:
++        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
++                      __func__, reg);
++        break;
+     default:
+-        qemu_log_mask(LOG_GUEST_ERROR,
+-                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
++        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
++                      __func__, reg);
++        break;
+     }
+ }
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+     /* Autonegotiation status mirrors link status. */
+     if (link_down) {
++        trace_lan9118_phy_update_link("down");
+         s->status &= ~0x0024;
+         s->ints |= PHY_INT_DOWN;
+     } else {
++        trace_lan9118_phy_update_link("up");
+         s->status |= 0x0024;
+         s->ints |= PHY_INT_ENERGYON;
+         s->ints |= PHY_INT_AUTONEG_COMPLETE;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+ void lan9118_phy_reset(Lan9118PhyState *s)
+ {
++    trace_lan9118_phy_reset();
++
+     s->control = 0x3000;
+     s->status = 0x7809;
+     s->advertise = 0x01e1;
+@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
+     .version_id = 1,
+     .minimum_version_id = 1,
+     .fields = (const VMStateField[]) {
+-        VMSTATE_UINT16(control, Lan9118PhyState),
+         VMSTATE_UINT16(status, Lan9118PhyState),
++        VMSTATE_UINT16(control, Lan9118PhyState),
+         VMSTATE_UINT16(advertise, Lan9118PhyState),
+         VMSTATE_UINT16(ints, Lan9118PhyState),
+         VMSTATE_UINT16(int_mask, Lan9118PhyState),
+diff --git a/hw/net/Kconfig b/hw/net/Kconfig
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/Kconfig
++++ b/hw/net/Kconfig
+@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
+ config IMX_FEC
+     bool
++    select LAN9118_PHY
+ config CADENCE
+     bool
+diff --git a/hw/net/trace-events b/hw/net/trace-events
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/trace-events
++++ b/hw/net/trace-events
+@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
+ allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
+ allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
++# lan9118_phy.c
++lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
++lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
++lan9118_phy_update_link(const char *s) "%s"
++lan9118_phy_reset(void) ""
++
+ # lance.c
+ lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
+ lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
+@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
+ i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
+ # imx_fec.c
+-imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
+ imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
+-imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
+ imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
+-imx_phy_update_link(const char *s) "%s"
+-imx_phy_reset(void) ""
+ imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
+ imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
+ imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
+--
+.34.1

-New patch
+[PULL 03/72] hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
+From: Bernhard Beschow <shentey@gmail.com>
+Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
+fixes the MSB of selector field to be zero, as specified in the datasheet.
+Fixes: 2a424990170b "LAN9118 emulation"
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241102125724.532843-4-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/net/lan9118_phy.c | 2 +-
+file changed, 1 insertion(+), 1 deletion(-)
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+         val = s->advertise;
+         break;
+     case 5: /* Auto-neg Link Partner Ability */
+-        val = 0x0f71;
++        val = 0x0fe1;
+         break;
+     case 6: /* Auto-neg Expansion */
+         val = 1;
+--
+.34.1

-New patch
+[PULL 04/72] hw/net/lan9118_phy: Reuse MII constants
+From: Bernhard Beschow <shentey@gmail.com>
+Prefer named constants over magic values for better readability.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Message-id: 20241102125724.532843-5-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ include/hw/net/mii.h |  6 +++++
+ hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
+files changed, 46 insertions(+), 23 deletions(-)
+diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/hw/net/mii.h
++++ b/include/hw/net/mii.h
+@@ -XXX,XX +XXX,XX @@
+ #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
+ #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
++#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
+ #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
+ #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
+ #define MII_ANAR_TXFD       (1 << 8)
+@@ -XXX,XX +XXX,XX @@
+ #define MII_ANAR_10FD       (1 << 6)
+ #define MII_ANAR_10         (1 << 5)
+ #define MII_ANAR_CSMACD     (1 << 0)
++#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
+ #define MII_ANLPAR_ACK      (1 << 14)
+ #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
+@@ -XXX,XX +XXX,XX @@
+ #define RTL8201CP_PHYID1    0x0000
+ #define RTL8201CP_PHYID2    0x8201
++/* SMSC LAN9118 */
++#define SMSCLAN9118_PHYID1  0x0007
++#define SMSCLAN9118_PHYID2  0xc0d1
++
+ /* RealTek 8211E */
+ #define RTL8211E_PHYID1     0x001c
+ #define RTL8211E_PHYID2     0xc915
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@
+ #include "qemu/osdep.h"
+ #include "hw/net/lan9118_phy.h"
++#include "hw/net/mii.h"
+ #include "hw/irq.h"
+ #include "hw/resettable.h"
+ #include "migration/vmstate.h"
+@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+     uint16_t val;
+     switch (reg) {
+-    case 0: /* Basic Control */
++    case MII_BMCR:
+         val = s->control;
+         break;
+-    case 1: /* Basic Status */
++    case MII_BMSR:
+         val = s->status;
+         break;
+-    case 2: /* ID1 */
+-        val = 0x0007;
++    case MII_PHYID1:
++        val = SMSCLAN9118_PHYID1;
+         break;
+-    case 3: /* ID2 */
+-        val = 0xc0d1;
++    case MII_PHYID2:
++        val = SMSCLAN9118_PHYID2;
+         break;
+-    case 4: /* Auto-neg advertisement */
++    case MII_ANAR:
+         val = s->advertise;
+         break;
+-    case 5: /* Auto-neg Link Partner Ability */
+-        val = 0x0fe1;
++    case MII_ANLPAR:
++        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
++              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
++              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
+         break;
+-    case 6: /* Auto-neg Expansion */
+-        val = 1;
++    case MII_ANER:
++        val = MII_ANER_NWAY;
+         break;
+     case 29: /* Interrupt source. */
+         val = s->ints;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+     trace_lan9118_phy_write(val, reg);
+     switch (reg) {
+-    case 0: /* Basic Control */
+-        if (val & 0x8000) {
++    case MII_BMCR:
++        if (val & MII_BMCR_RESET) {
+             lan9118_phy_reset(s);
+         } else {
+-            s->control = val & 0x7980;
++            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
++                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
++                                MII_BMCR_CTST);
+             /* Complete autonegotiation immediately. */
+-            if (val & 0x1000) {
+-                s->status |= 0x0020;
++            if (val & MII_BMCR_AUTOEN) {
++                s->status |= MII_BMSR_AN_COMP;
+             }
+         }
+         break;
+-    case 4: /* Auto-neg advertisement */
+-        s->advertise = (val & 0x2d7f) | 0x80;
++    case MII_ANAR:
++        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
++                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
++                               MII_ANAR_SELECT))
++                     | MII_ANAR_TX;
+         break;
+     case 30: /* Interrupt mask */
+         s->int_mask = val & 0xff;
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+     /* Autonegotiation status mirrors link status. */
+     if (link_down) {
+         trace_lan9118_phy_update_link("down");
+-        s->status &= ~0x0024;
++        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
+         s->ints |= PHY_INT_DOWN;
+     } else {
+         trace_lan9118_phy_update_link("up");
+-        s->status |= 0x0024;
++        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
+         s->ints |= PHY_INT_ENERGYON;
+         s->ints |= PHY_INT_AUTONEG_COMPLETE;
+     }
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
+ {
+     trace_lan9118_phy_reset();
+-    s->control = 0x3000;
+-    s->status = 0x7809;
+-    s->advertise = 0x01e1;
++    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
++    s->status = MII_BMSR_100TX_FD
++                | MII_BMSR_100TX_HD
++                | MII_BMSR_10T_FD
++                | MII_BMSR_10T_HD
++                | MII_BMSR_AUTONEG
++                | MII_BMSR_EXTCAP;
++    s->advertise = MII_ANAR_TXFD
++                   | MII_ANAR_TX
++                   | MII_ANAR_10FD
++                   | MII_ANAR_10
++                   | MII_ANAR_CSMACD;
+     s->int_mask = 0;
+     s->ints = 0;
+     lan9118_phy_update_link(s, s->link_down);
+--
+.34.1

-New patch
+[PULL 05/72] hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement
+From: Bernhard Beschow <shentey@gmail.com>
+The real device advertises this mode and the device model already advertises
+mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
+make the model more realistic.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Bernhard Beschow <shentey@gmail.com>
+Tested-by: Guenter Roeck <linux@roeck-us.net>
+Message-id: 20241102125724.532843-6-shentey@gmail.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/net/lan9118_phy.c | 4 ++--
+file changed, 2 insertions(+), 2 deletions(-)
+diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/net/lan9118_phy.c
++++ b/hw/net/lan9118_phy.c
+@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+         break;
+     case MII_ANAR:
+         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
+-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
+-                               MII_ANAR_SELECT))
++                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
++                               MII_ANAR_10 | MII_ANAR_SELECT))
+                      | MII_ANAR_TX;
+         break;
+     case 30: /* Interrupt mask */
+--
+.34.1

-New patch
+[PULL 06/72] fpu: handle raising Invalid for infzero in pick_nan_muladd
+For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
+Invalid for the multiplication of 0 by infinity.  Currently we handle
+this in the per-architecture ifdef ladder in pickNaNMulAdd().
+However, since this isn't really architecture specific we can hoist
+it up to the generic code.
+For the cases where the infzero test in pickNaNMulAdd was
+returning 2, we can delete the check entirely and allow the
+code to fall into the normal pick-a-NaN handling, because this
+will return 2 anyway (input 'c' being the only NaN in this case).
+For the cases where infzero was returning 3 to indicate "return
+the default NaN", we must retain that "return 3".
+For Arm, this looks like it might be a behaviour change because we
+used to set float_flag_invalid | float_flag_invalid_imz only if C is
+a quiet NaN.  However, it is not, because Arm target code never looks
+at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
+already raised float_flag_invalid via the "abc_mask &
+float_cmask_snan" check in pick_nan_muladd.
+For any target architecture using the "default implementation" at the
+bottom of the ifdef, this is a behaviour change but will be fixing a
+bug (where we failed to raise the Invalid exception for (0 * inf +
+QNaN).  The architectures using the default case are:
+ * hppa
+ * i386
+ * sh4
+ * tricore
+The x86, Tricore and SH4 CPU architecture manuals are clear that this
+should have raised Invalid; HPPA is a bit vaguer but still seems
+clear enough.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 13 +++++++------
+ fpu/softfloat-specialize.c.inc | 29 +----------------------------
+files changed, 8 insertions(+), 34 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+                                             int ab_mask, int abc_mask)
+ {
+     int which;
++    bool infzero = (ab_mask == float_cmask_infzero);
+     if (unlikely(abc_mask & float_cmask_snan)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
+     }
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
+-                          ab_mask == float_cmask_infzero, s);
++    if (infzero) {
++        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
++        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
++    }
++
++    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+     if (s->default_nan_mode || which == 3) {
+-        /*
+-         * Note that this check is after pickNaNMulAdd so that function
+-         * has an opportunity to set the Invalid flag for infzero.
+-         */
+         parts_default_nan(a, s);
+         return a;
+     }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * the default NaN
+      */
+     if (infzero && is_qnan(c_cls)) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+         return 3;
+     }
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * case sets InvalidOp and returns the default NaN
+          */
+         if (infzero) {
+-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+             return 3;
+         }
+         /* Prefer sNaN over qNaN, in the a, b, c order. */
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+          */
+-        if (infzero) {
+-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-            return 2;
+-        }
+         /* Prefer sNaN over qNaN, in the c, a, b order. */
+         if (is_snan(c_cls)) {
+             return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+      * case sets InvalidOp and returns the input value 'c'
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
++
+     /* Prefer sNaN over qNaN, in the c, a, b order. */
+     if (is_snan(c_cls)) {
+         return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * to return an input NaN if we have one (ie c) rather than generating
+      * a default NaN
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         return 1;
+     }
+ #elif defined(TARGET_RISCV)
+-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-    }
+     return 3; /* default NaN */
+ #elif defined(TARGET_S390X)
+     if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+         return 3;
+     }
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         return 2;
+     }
+ #elif defined(TARGET_SPARC)
+-    /* For (inf,0,nan) return c. */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     /* Prefer SNaN over QNaN, order C, B, A. */
+     if (is_snan(c_cls)) {
+         return 2;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
+      * an input NaN if we have one (ie c).
+      */
+-    if (infzero) {
+-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
+-        return 2;
+-    }
+     if (status->use_first_nan) {
+         if (is_nan(a_cls)) {
+             return 0;
+--
+.34.1

-New patch
+[PULL 07/72] fpu: Check for default_nan_mode before calling pickNaNMulAdd
+If the target sets default_nan_mode then we're always going to return
+the default NaN, and pickNaNMulAdd() no longer has any side effects.
+For consistency with pickNaN(), check for default_nan_mode before
+calling pickNaNMulAdd().
+When we convert pickNaNMulAdd() to allow runtime selection of the NaN
+propagation rule, this means we won't have to make the targets which
+use default_nan_mode also set a propagation rule.
+Since RiscV always uses default_nan_mode, this allows us to remove
+its ifdef case from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 8 ++++++--
+ fpu/softfloat-specialize.c.inc | 9 +++++++--
+files changed, 13 insertions(+), 4 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+     }
+-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    if (s->default_nan_mode) {
++        which = 3;
++    } else {
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++    }
+-    if (s->default_nan_mode || which == 3) {
++    if (which == 3) {
+         parts_default_nan(a, s);
+         return a;
+     }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
++    /*
++     * We guarantee not to require the target to tell us how to
++     * pick a NaN if we're always returning the default NaN.
++     * But if we're not in default-NaN mode then the target must
++     * specify.
++     */
++    assert(!status->default_nan_mode);
+ #if defined(TARGET_ARM)
+     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
+      * the default NaN
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     } else {
+         return 1;
+     }
+-#elif defined(TARGET_RISCV)
+-    return 3; /* default NaN */
+ #elif defined(TARGET_S390X)
+     if (infzero) {
+         return 3;
+--
+.34.1

-New patch
+[PULL 08/72] softfloat: Allow runtime choice of inf * 0 + NaN result
+IEEE 758 does not define a fixed rule for what NaN to return in
 the case of a fused multiply-add of inf * 0 + NaN. Different
 architectures thus do different things:
  * some return the default NaN
  * some return the input NaN
  * Arm returns the default NaN if the input NaN is quiet,
    and the input NaN if it is signalling
 We want to make this logic be runtime selected rather than
 hardcoded into the binary, because:
  * this will let us have multiple targets in one QEMU binary
  * the Arm FEAT_AFP architectural feature includes letting
    the guest select a NaN propagation rule at runtime
 In this commit we add an enum for the propagation rule, the field in
 float_status, and the corresponding getters and setters.  We change
 pickNaNMulAdd to honour this, but because all targets still leave
 this field at its default 0 value, the fallback logic will pick the
 rule type with the old ifdef ladder.
 Note that four architectures both use the muladd softfloat functions
 and did not have a branch of the ifdef ladder to specify their
 behaviour (and so were ending up with the "default" case, probably
 wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
 default_nan_mode, and so will never get into pickNaNMulAdd().  For
 HPPA and i386 we retain the same behaviour as the old default-case,
 which is to not ever return the default NaN.  This might not be
 correct but it is not a behaviour change.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
 ---
  include/fpu/softfloat-helpers.h | 11 ++++
  include/fpu/softfloat-types.h   | 23 +++++++++
  fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
 files changed, 95 insertions(+), 30 deletions(-)
 diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-helpers.h
 +++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
      status->float_2nan_prop_rule = rule;
  }
 +static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
 +                                             float_status *status)
 +{
 +    status->float_infzeronan_rule = rule;
 +}
 +
  static inline void set_flush_to_zero(bool val, float_status *status)
  {
      status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
      return status->float_2nan_prop_rule;
  }
 +static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 +{
 +    return status->float_infzeronan_rule;
 +}
 +
  static inline bool get_flush_to_zero(float_status *status)
  {
      return status->flush_to_zero;
 diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-types.h
 +++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
      float_2nan_prop_x87,
  } Float2NaNPropRule;
 +/*
 + * Rule for result of fused multiply-add 0 * Inf + NaN.
 + * This must be a NaN, but implementations differ on whether this
 + * is the input NaN or the default NaN.
 + *
 + * You don't need to set this if default_nan_mode is enabled.
 + * When not in default-NaN mode, it is an error for the target
 + * not to set the rule in float_status if it uses muladd, and we
 + * will assert if we need to handle an input NaN and no rule was
 + * selected.
 + */
 +typedef enum __attribute__((__packed__)) {
 +    /* No propagation rule specified */
 +    float_infzeronan_none = 0,
 +    /* Result is never the default NaN (so always the input NaN) */
 +    float_infzeronan_dnan_never,
 +    /* Result is always the default NaN */
 +    float_infzeronan_dnan_always,
 +    /* Result is the default NaN if the input NaN is quiet */
 +    float_infzeronan_dnan_if_qnan,
 +} FloatInfZeroNaNRule;
 +
  /*
   * Floating Point Status. Individual architectures may maintain
   * several versions of float_status for different functions. The
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
      bool flush_to_zero;
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
  static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                           bool infzero, float_status *status)
  {
 +    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
 +
      /*
       * We guarantee not to require the target to tell us how to
       * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
       * specify.
       */
      assert(!status->default_nan_mode);
 +
 +    if (rule == float_infzeronan_none) {
 +        /*
 +         * Temporarily fall back to ifdef ladder
 +         */
  #if defined(TARGET_ARM)
 -    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
 -     * the default NaN
 -     */
 -    if (infzero && is_qnan(c_cls)) {
 -        return 3;
 +        /*
 +         * For ARM, the (inf,zero,qnan) case returns the default NaN,
 +         * but (inf,zero,snan) returns the input NaN.
 +         */
 +        rule = float_infzeronan_dnan_if_qnan;
 +#elif defined(TARGET_MIPS)
 +        if (snan_bit_is_one(status)) {
 +            /*
 +             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the default NaN
 +             */
 +            rule = float_infzeronan_dnan_always;
 +        } else {
 +            /*
 +             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 +             * case sets InvalidOp and returns the input value 'c'
 +             */
 +            rule = float_infzeronan_dnan_never;
 +        }
 +#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
 +    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
 +    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
 +        /*
 +         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 +         * case sets InvalidOp and returns the input value 'c'
 +         */
 +        /*
 +         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
 +         * to return an input NaN if we have one (ie c) rather than generating
 +         * a default NaN
 +         */
 +        rule = float_infzeronan_dnan_never;
 +#elif defined(TARGET_S390X)
 +        rule = float_infzeronan_dnan_always;
 +#endif
      }
 +    if (infzero) {
 +        /*
 +         * Inf * 0 + NaN -- some implementations return the default NaN here,
 +         * and some return the input NaN.
 +         */
 +        switch (rule) {
 +        case float_infzeronan_dnan_never:
 +            return 2;
 +        case float_infzeronan_dnan_always:
 +            return 3;
 +        case float_infzeronan_dnan_if_qnan:
 +            return is_qnan(c_cls) ? 3 : 2;
 +        default:
 +            g_assert_not_reached();
 +        }
 +    }
 +
 +#if defined(TARGET_ARM)
 +
      /* This looks different from the ARM ARM pseudocode, because the ARM ARM
       * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
       */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      }
  #elif defined(TARGET_MIPS)
      if (snan_bit_is_one(status)) {
 -        /*
 -         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the default NaN
 -         */
 -        if (infzero) {
 -            return 3;
 -        }
          /* Prefer sNaN over qNaN, in the a, b, c order. */
          if (is_snan(a_cls)) {
              return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              return 2;
          }
      } else {
 -        /*
 -         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
 -         * case sets InvalidOp and returns the input value 'c'
 -         */
          /* Prefer sNaN over qNaN, in the c, a, b order. */
          if (is_snan(c_cls)) {
              return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
  #elif defined(TARGET_LOONGARCH64)
 -    /*
 -     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
 -     * case sets InvalidOp and returns the input value 'c'
 -     */
 -
      /* Prefer sNaN over qNaN, in the c, a, b order. */
      if (is_snan(c_cls)) {
          return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          return 1;
      }
  #elif defined(TARGET_PPC)
 -    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
 -     * to return an input NaN if we have one (ie c) rather than generating
 -     * a default NaN
 -     */
 -
      /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
       * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
       */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          return 1;
      }
  #elif defined(TARGET_S390X)
 -    if (infzero) {
 -        return 3;
 -    }
 -
      if (is_snan(a_cls)) {
          return 0;
      } else if (is_snan(b_cls)) {
 --
 .34.1

-New patch
+[PULL 09/72] tests/fp: Explicitly set inf-zero-nan rule
+Explicitly set a rule in the softfloat tests for the inf-zero-nan
+muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
+and so we should select here the Arm rule of
+float_infzeronan_dnan_if_qnan.
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 5 +++++
+ tests/fp/fp-test.c  | 5 +++++
+files changed, 10 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+ {
+     bench_func_t f;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+ {
+     unsigned int i;
++    /*
++     * These implementation-defined choices for various things IEEE
++     * doesn't specify match those used by the Arm architecture.
++     */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+     verCases_maxErrorCount = n_max_errors;
+--
+.34.1

-New patch
+[PULL 10/72] target/arm: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the Arm target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 3 +++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 4 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
++ *    and the input NaN if it is signalling
+  */
+ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_ARM)
+-        /*
+-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+-         * but (inf,zero,snan) returns the input NaN.
+-         */
+-        rule = float_infzeronan_dnan_if_qnan;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             /*
+              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 11/72] target/s390: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for s390, so we
+can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_infzeronan_rule(float_infzeronan_dnan_always,
++                                  &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * a default NaN
+          */
+         rule = float_infzeronan_dnan_never;
+-#elif defined(TARGET_S390X)
+-        rule = float_infzeronan_dnan_always;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 12/72] target/ppc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the PPC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 7 +------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
++     * to return an input NaN if we have one (ie c) rather than generating
++     * a default NaN
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+              */
+             rule = float_infzeronan_dnan_never;
+         }
+-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
++#elif defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+          */
+-        /*
+-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+-         * to return an input NaN if we have one (ie c) rather than generating
+-         * a default NaN
+-         */
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 13/72] target/mips: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the MIPS target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h       |  9 +++++++++
+ target/mips/msa.c              |  4 ++++
+ fpu/softfloat-specialize.c.inc | 16 +---------------
+files changed, 14 insertions(+), 15 deletions(-)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
+ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+ {
+     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
++    FloatInfZeroNaNRule izn_rule;
+     /*
+      * With nan2008, SNaNs are silenced in the usual way.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+      */
+     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
+     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
++    /*
++     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
++     * case sets InvalidOp and returns the default NaN.
++     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'.
++     */
++    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
++    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+ }
+ static inline void restore_fp_status(CPUMIPSState *env)
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* set proper signanling bit meaning ("1" means "quiet") */
+     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
++
++    /* Inf * 0 + NaN returns the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never,
++                              &env->active_tc.msa_fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_MIPS)
+-        if (snan_bit_is_one(status)) {
+-            /*
+-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the default NaN
+-             */
+-            rule = float_infzeronan_dnan_always;
+-        } else {
+-            /*
+-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+-             * case sets InvalidOp and returns the input value 'c'
+-             */
+-            rule = float_infzeronan_dnan_never;
+-        }
+-#elif defined(TARGET_SPARC) || \
++#if defined(TARGET_SPARC) || \
+     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+--
+.34.1

-New patch
+[PULL 14/72] target/sparc: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the SPARC target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 +--
+files changed, 3 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_SPARC) || \
+-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 15/72] target/xtensa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the xtensa target,
+so we can remove the ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
+---
+ target/xtensa/cpu.c            | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/xtensa/cpu.c
++++ b/target/xtensa/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
+     reset_mmu(env);
+     cs->halted = env->runstall;
+ #endif
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_no_signaling_nans(!dfpu, &env->fp_status);
+     xtensa_use_first_nan(env, !dfpu);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
++#if defined(TARGET_HPPA) || \
+     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+--
+.34.1

-New patch
+[PULL 16/72] target/x86: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the x86 target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 7 +++++++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 8 insertions(+), 1 deletion(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
++    /*
++     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
++     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
++     * specified -- for 0 * inf + NaN the input NaN is selected, and if
++     * there are multiple input NaNs they are selected in the order a, b, c.
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+          * Temporarily fall back to ifdef ladder
+          */
+ #if defined(TARGET_HPPA) || \
+-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
++    defined(TARGET_LOONGARCH)
+         /*
+          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+          * case sets InvalidOp and returns the input value 'c'
+--
+.34.1

-New patch
+[PULL 17/72] target/loongarch: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the loongarch target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 5 +++++
+ fpu/softfloat-specialize.c.inc    | 7 +------
+files changed, 6 insertions(+), 6 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+                             &env->fp_status);
+     set_flush_to_zero(0, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
++     * case sets InvalidOp and returns the input value 'c'
++     */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         /*
+          * Temporarily fall back to ifdef ladder
+          */
+-#if defined(TARGET_HPPA) || \
+-    defined(TARGET_LOONGARCH)
+-        /*
+-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+-         * case sets InvalidOp and returns the input value 'c'
+-         */
++#if defined(TARGET_HPPA)
+         rule = float_infzeronan_dnan_never;
+ #endif
+     }
+--
+.34.1

-New patch
+[PULL 18/72] target/hppa: Set FloatInfZeroNaNRule explicitly
+Set the FloatInfZeroNaNRule explicitly for the HPPA target,
+so we can remove the ifdef from pickNaNMulAdd().
+As this is the last target to be converted to explicitly setting
+the rule, we can remove the fallback code in pickNaNMulAdd()
+entirely.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       |  2 ++
+ fpu/softfloat-specialize.c.inc | 13 +------------
+files changed, 3 insertions(+), 12 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /* For inf * 0 + NaN, return the input NaN */
++    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+                          bool infzero, float_status *status)
+ {
+-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+-
+     /*
+      * We guarantee not to require the target to tell us how to
+      * pick a NaN if we're always returning the default NaN.
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+      */
+     assert(!status->default_nan_mode);
+-    if (rule == float_infzeronan_none) {
+-        /*
+-         * Temporarily fall back to ifdef ladder
+-         */
+-#if defined(TARGET_HPPA)
+-        rule = float_infzeronan_dnan_never;
+-#endif
+-    }
+-
+     if (infzero) {
+         /*
+          * Inf * 0 + NaN -- some implementations return the default NaN here,
+          * and some return the input NaN.
+          */
+-        switch (rule) {
++        switch (status->float_infzeronan_rule) {
+         case float_infzeronan_dnan_never:
+             return 2;
+         case float_infzeronan_dnan_always:
+--
+.34.1

-New patch
+[PULL 19/72] softfloat: Pass have_snan to pickNaNMulAdd
+The new implementation of pickNaNMulAdd() will find it convenient
+to know whether at least one of the three arguments to the muladd
+was a signaling NaN. We already calculate that in the caller,
+so pass it in as a new bool have_snan.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
+---
+ fpu/softfloat-parts.c.inc      | 5 +++--
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 4 insertions(+), 3 deletions(-)
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-parts.c.inc
++++ b/fpu/softfloat-parts.c.inc
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+ {
+     int which;
+     bool infzero = (ab_mask == float_cmask_infzero);
++    bool have_snan = (abc_mask & float_cmask_snan);
+-    if (unlikely(abc_mask & float_cmask_snan)) {
++    if (unlikely(have_snan)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
+     }
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+     if (s->default_nan_mode) {
+         which = 3;
+     } else {
+-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
++        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
+     }
+     if (which == 3) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
+ | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
+ *----------------------------------------------------------------------------*/
+ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+-                         bool infzero, float_status *status)
++                         bool infzero, bool have_snan, float_status *status)
+ {
+     /*
+      * We guarantee not to require the target to tell us how to
+--
+.34.1

-[PULL 04/24] target/arm: Support migration when FPSR/FPCR won't fit in the FPSCR
+[PULL 20/72] softfloat: Allow runtime choice of NaN propagation for muladd
-To support FPSR and FPCR bits that don't exist in the AArch32 FPSCR
+IEEE 758 does not define a fixed rule for which NaN to pick as the
-view of floating point control and status (such as the FEAT_AFP ones),
+result if both operands of a 3-operand fused multiply-add operation
-we need to make sure those bits can be migrated. This commit allows
+are NaNs.  As a result different architectures have ended up with
-that, whilst maintaining backwards and forwards migration compatibility
+different rules for propagating NaNs.
-for CPUs where there are no such bits:
+QEMU currently hardcodes the NaN propagation logic into the binary
-On sending:
+because pickNaNMulAdd() has an ifdef ladder for different targets.
- * If either the FPCR or the FPSR include set bits that are not
+We want to make the propagation rule instead be selectable at
-   visible in the AArch32 FPSCR view of floating point control/status
+runtime, because:
-   then we send the FPCR and FPSR as two separate fields in a new
+ * this will let us have multiple targets in one QEMU binary
-   cpu/vfp/fpcr_fpsr subsection, and we send a 0 for the old
+ * the Arm FEAT_AFP architectural feature includes letting
-   FPSCR field in cpu/vfp
+   the guest select a NaN propagation rule at runtime
- * Otherwise, we don't send the fpcr_fpsr subsection, and we send
-   an FPSCR-format value in cpu/vfp as we did previously
+In this commit we add an enum for the propagation rule, the field in
+float_status, and the corresponding getters and setters.  We change
-On receiving:
+pickNaNMulAdd to honour this, but because all targets still leave
- * if we see a non-zero FPSCR field, that is the right information
+this field at its default 0 value, the fallback logic will pick the
- * if we see a fpcr_fpsr subsection then that has the information
+rule type with the old ifdef ladder.
- * if we see neither, then FPSCR/FPCR/FPSR are all zero on the source;
-   cpu_pre_load() ensures the CPU state defaults to that
+It's valid not to set a propagation rule if default_nan_mode is
- * if we see both, then the migration source is buggy or malicious;
+enabled, because in that case there's no need to pick a NaN; all the
-   either the fpcr_fpsr or the FPSCR will "win" depending which
+callers of pickNaNMulAdd() catch this case and skip calling it.
    is first in the migration stream; we don't care which that is
 We make the new FPCR and FPSR on-the-wire data be 64 bits, because
 architecturally these registers are that wide, and this avoids the
 need to engage in further migration-compatibility contortions in
 future if some new architecture revision defines bits in the high
 half of either register.
 (We won't ever send the new migration subsection until we add support
 for a CPU feature which enables setting overlapping FPCR bits, like
 FEAT_AFP.)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-5-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
 ---
- target/arm/machine.c | 134 ++++++++++++++++++++++++++++++++++++++++++-
+ include/fpu/softfloat-helpers.h |  11 +++
-file changed, 132 insertions(+), 2 deletions(-)
+ include/fpu/softfloat-types.h   |  55 +++++++++++
+ fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
-diff --git a/target/arm/machine.c b/target/arm/machine.c
+files changed, 107 insertions(+), 126 deletions(-)
 diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/machine.c
+--- a/include/fpu/softfloat-helpers.h
-+++ b/target/arm/machine.c
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@ static bool vfp_needed(void *opaque)
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
-             : cpu_isar_feature(aa32_vfp_simd, cpu));
+     status->float_2nan_prop_rule = rule;
  }
-+static bool vfp_fpcr_fpsr_needed(void *opaque)
++static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
 +                                            float_status *status)
 +{
-+    /*
++    status->float_3nan_prop_rule = rule;
 +     * If either the FPCR or the FPSR include set bits that are not
 +     * visible in the AArch32 FPSCR view of floating point control/status
 +     * then we must send the FPCR and FPSR as two separate fields in the
 +     * cpu/vfp/fpcr_fpsr subsection, and we will send a 0 for the old
 +     * FPSCR field in cpu/vfp.
 +     *
 +     * If all the set bits are representable in an AArch32 FPSCR then we
 +     * send that value as the cpu/vfp FPSCR field, and don't send the
 +     * cpu/vfp/fpcr_fpsr subsection.
 +     *
 +     * On incoming migration, if the cpu/vfp FPSCR field is non-zero we
 +     * use it, and if the fpcr_fpsr subsection is present we use that.
 +     * (The subsection will never be present with a non-zero FPSCR field,
 +     * and if FPSCR is zero and the subsection is not present that means
 +     * that FPSCR/FPSR/FPCR are zero.)
 +     *
 +     * This preserves migration compatibility with older QEMU versions,
 +     * in both directions.
 +     */
 +    ARMCPU *cpu = opaque;
 +    CPUARMState *env = &cpu->env;
 +
 +    return (vfp_get_fpcr(env) & ~FPCR_MASK) || (vfp_get_fpsr(env) & ~FPSR_MASK);
 +}
 +
- static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
+ static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
-                      const VMStateField *field)
+                                              float_status *status)
  {
-@@ -XXX,XX +XXX,XX @@ static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
+@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
-     CPUARMState *env = &cpu->env;
+     return status->float_2nan_prop_rule;
-     uint32_t val = qemu_get_be32(f);
+ }
--    vfp_set_fpscr(env, val);
++static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
-+    if (val) {
++{
-+        /* 0 means we might have the data in the fpcr_fpsr subsection */
++    return status->float_3nan_prop_rule;
-+        vfp_set_fpscr(env, val);
++}
 +
  static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
  {
      return status->float_infzeronan_rule;
 diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
 index XXXXXXX..XXXXXXX 100644
 --- a/include/fpu/softfloat-types.h
 +++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ this code that are retained.
  #ifndef SOFTFLOAT_TYPES_H
  #define SOFTFLOAT_TYPES_H
 +#include "hw/registerfields.h"
 +
  /*
   * Software IEC/IEEE floating-point types.
   */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
      float_2nan_prop_x87,
  } Float2NaNPropRule;
 +/*
 + * 3-input NaN propagation rule, for fused multiply-add. Individual
 + * architectures have different rules for which input NaN is
 + * propagated to the output when there is more than one NaN on the
 + * input.
 + *
 + * If default_nan_mode is enabled then it is valid not to set a NaN
 + * propagation rule, because the softfloat code guarantees not to try
 + * to pick a NaN to propagate in default NaN mode.  When not in
 + * default-NaN mode, it is an error for the target not to set the rule
 + * in float_status if it uses a muladd, and we will assert if we need
 + * to handle an input NaN and no rule was selected.
 + *
 + * The naming scheme for Float3NaNPropRule values is:
 + *  float_3nan_prop_s_abc:
 + *    = "Prefer SNaN over QNaN, then operand A over B over C"
 + *  float_3nan_prop_abc:
 + *    = "Prefer A over B over C regardless of SNaN vs QNAN"
 + *
 + * For QEMU, the multiply-add operation is A * B + C.
 + */
 +
 +/*
 + * We set the Float3NaNPropRule enum values up so we can select the
 + * right value in pickNaNMulAdd in a data driven way.
 + */
 +FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
 +FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
 +FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
 +FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
 +
 +#define PROPRULE(X, Y, Z) \
 +    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
 +
 +typedef enum __attribute__((__packed__)) {
 +    float_3nan_prop_none = 0,     /* No propagation rule specified */
 +    float_3nan_prop_abc = PROPRULE(0, 1, 2),
 +    float_3nan_prop_acb = PROPRULE(0, 2, 1),
 +    float_3nan_prop_bac = PROPRULE(1, 0, 2),
 +    float_3nan_prop_bca = PROPRULE(1, 2, 0),
 +    float_3nan_prop_cab = PROPRULE(2, 0, 1),
 +    float_3nan_prop_cba = PROPRULE(2, 1, 0),
 +    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
 +    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
 +} Float3NaNPropRule;
 +
 +#undef PROPRULE
 +
  /*
   * Rule for result of fused multiply-add 0 * Inf + NaN.
   * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
      FloatRoundMode float_rounding_mode;
      FloatX80RoundPrec floatx80_rounding_precision;
      Float2NaNPropRule float_2nan_prop_rule;
 +    Float3NaNPropRule float_3nan_prop_rule;
      FloatInfZeroNaNRule float_infzeronan_rule;
      bool tininess_before_rounding;
      /* should denormalised results go to zero and set the inexact flag? */
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
  static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                           bool infzero, bool have_snan, float_status *status)
  {
 +    FloatClass cls[3] = { a_cls, b_cls, c_cls };
 +    Float3NaNPropRule rule = status->float_3nan_prop_rule;
 +    int which;
 +
      /*
       * We guarantee not to require the target to tell us how to
       * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          }
      }
 +    if (rule == float_3nan_prop_none) {
  #if defined(TARGET_ARM)
 -
 -    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
 -     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
 -     */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else {
 -        return 1;
 -    }
 +        /*
 +         * This looks different from the ARM ARM pseudocode, because the ARM ARM
 +         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
 +         */
 +        rule = float_3nan_prop_s_cab;
  #elif defined(TARGET_MIPS)
 -    if (snan_bit_is_one(status)) {
 -        /* Prefer sNaN over qNaN, in the a, b, c order. */
 -        if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_snan(c_cls)) {
 -            return 2;
 -        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 +        if (snan_bit_is_one(status)) {
 +            rule = float_3nan_prop_s_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_s_cab;
          }
 -    } else {
 -        /* Prefer sNaN over qNaN, in the c, a, b order. */
 -        if (is_snan(c_cls)) {
 -            return 2;
 -        } else if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_qnan(c_cls)) {
 -            return 2;
 -        } else if (is_qnan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -    }
  #elif defined(TARGET_LOONGARCH64)
 -    /* Prefer sNaN over qNaN, in the c, a, b order. */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else {
 -        return 1;
 -    }
 +        rule = float_3nan_prop_s_cab;
  #elif defined(TARGET_PPC)
 -    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 -     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 -     */
 -    if (is_nan(a_cls)) {
 -        return 0;
 -    } else if (is_nan(c_cls)) {
 -        return 2;
 -    } else {
 -        return 1;
 -    }
 +        /*
 +         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
 +         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
 +         */
 +        rule = float_3nan_prop_acb;
  #elif defined(TARGET_S390X)
 -    if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 2;
 -    }
 +        rule = float_3nan_prop_s_abc;
  #elif defined(TARGET_SPARC)
 -    /* Prefer SNaN over QNaN, order C, B, A. */
 -    if (is_snan(c_cls)) {
 -        return 2;
 -    } else if (is_snan(b_cls)) {
 -        return 1;
 -    } else if (is_snan(a_cls)) {
 -        return 0;
 -    } else if (is_qnan(c_cls)) {
 -        return 2;
 -    } else if (is_qnan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 0;
 -    }
 +        rule = float_3nan_prop_s_cba;
  #elif defined(TARGET_XTENSA)
 -    /*
 -     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
 -     * an input NaN if we have one (ie c).
 -     */
 -    if (status->use_first_nan) {
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else if (is_nan(b_cls)) {
 -            return 1;
 +        if (status->use_first_nan) {
 +            rule = float_3nan_prop_abc;
          } else {
 -            return 2;
 +            rule = float_3nan_prop_cba;
          }
 -    } else {
 -        if (is_nan(c_cls)) {
 -            return 2;
 -        } else if (is_nan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -    }
  #else
 -    /* A default implementation: prefer a to b to c.
 -     * This is unlikely to actually match any real implementation.
 -     */
 -    if (is_nan(a_cls)) {
 -        return 0;
 -    } else if (is_nan(b_cls)) {
 -        return 1;
 -    } else {
 -        return 2;
 -    }
 +        rule = float_3nan_prop_abc;
  #endif
 +    }
-     return 0;
++
 +    assert(rule != float_3nan_prop_none);
 +    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
 +        /* We have at least one SNaN input and should prefer it */
 +        do {
 +            which = rule & R_3NAN_1ST_MASK;
 +            rule >>= R_3NAN_1ST_LENGTH;
 +        } while (!is_snan(cls[which]));
 +    } else {
 +        do {
 +            which = rule & R_3NAN_1ST_MASK;
 +            rule >>= R_3NAN_1ST_LENGTH;
 +        } while (!is_nan(cls[which]));
 +    }
 +    return which;
  }
-@@ -XXX,XX +XXX,XX @@ static int put_fpscr(QEMUFile *f, void *opaque, size_t size,
+ /*----------------------------------------------------------------------------
  {
      ARMCPU *cpu = opaque;
      CPUARMState *env = &cpu->env;
 +    uint32_t fpscr = vfp_fpcr_fpsr_needed(opaque) ? 0 : vfp_get_fpscr(env);
 -    qemu_put_be32(f, vfp_get_fpscr(env));
 +    qemu_put_be32(f, fpscr);
      return 0;
  }
@@ -XXX,XX +XXX,XX @@ static const VMStateInfo vmstate_fpscr = {
      .put = put_fpscr,
  };
 +static int get_fpcr(QEMUFile *f, void *opaque, size_t size,
 +                     const VMStateField *field)
 +{
 +    ARMCPU *cpu = opaque;
 +    CPUARMState *env = &cpu->env;
 +    uint64_t val = qemu_get_be64(f);
 +
 +    vfp_set_fpcr(env, val);
 +    return 0;
 +}
 +
 +static int put_fpcr(QEMUFile *f, void *opaque, size_t size,
 +                     const VMStateField *field, JSONWriter *vmdesc)
 +{
 +    ARMCPU *cpu = opaque;
 +    CPUARMState *env = &cpu->env;
 +
 +    qemu_put_be64(f, vfp_get_fpcr(env));
 +    return 0;
 +}
 +
 +static const VMStateInfo vmstate_fpcr = {
 +    .name = "fpcr",
 +    .get = get_fpcr,
 +    .put = put_fpcr,
 +};
 +
 +static int get_fpsr(QEMUFile *f, void *opaque, size_t size,
 +                     const VMStateField *field)
 +{
 +    ARMCPU *cpu = opaque;
 +    CPUARMState *env = &cpu->env;
 +    uint64_t val = qemu_get_be64(f);
 +
 +    vfp_set_fpsr(env, val);
 +    return 0;
 +}
 +
 +static int put_fpsr(QEMUFile *f, void *opaque, size_t size,
 +                     const VMStateField *field, JSONWriter *vmdesc)
 +{
 +    ARMCPU *cpu = opaque;
 +    CPUARMState *env = &cpu->env;
 +
 +    qemu_put_be64(f, vfp_get_fpsr(env));
 +    return 0;
 +}
 +
 +static const VMStateInfo vmstate_fpsr = {
 +    .name = "fpsr",
 +    .get = get_fpsr,
 +    .put = put_fpsr,
 +};
 +
 +static const VMStateDescription vmstate_vfp_fpcr_fpsr = {
 +    .name = "cpu/vfp/fpcr_fpsr",
 +    .version_id = 1,
 +    .minimum_version_id = 1,
 +    .needed = vfp_fpcr_fpsr_needed,
 +    .fields = (const VMStateField[]) {
 +        {
 +            .name = "fpcr",
 +            .version_id = 0,
 +            .size = sizeof(uint64_t),
 +            .info = &vmstate_fpcr,
 +            .flags = VMS_SINGLE,
 +            .offset = 0,
 +        },
 +        {
 +            .name = "fpsr",
 +            .version_id = 0,
 +            .size = sizeof(uint64_t),
 +            .info = &vmstate_fpsr,
 +            .flags = VMS_SINGLE,
 +            .offset = 0,
 +        },
 +        VMSTATE_END_OF_LIST()
 +    },
 +};
 +
  static const VMStateDescription vmstate_vfp = {
      .name = "cpu/vfp",
      .version_id = 3,
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_vfp = {
              .offset = 0,
          },
          VMSTATE_END_OF_LIST()
 +    },
 +    .subsections = (const VMStateDescription * const []) {
 +        &vmstate_vfp_fpcr_fpsr,
 +        NULL
      }
  };
@@ -XXX,XX +XXX,XX @@ static int cpu_pre_load(void *opaque)
      ARMCPU *cpu = opaque;
      CPUARMState *env = &cpu->env;
 +    /*
 +     * In an inbound migration where on the source FPSCR/FPSR/FPCR are 0,
 +     * there will be no fpcr_fpsr subsection so we won't call vfp_set_fpcr()
 +     * and vfp_set_fpsr() from get_fpcr() and get_fpsr(); also the get_fpscr()
 +     * function will not call vfp_set_fpscr() because it will see a 0 in the
 +     * inbound data. Ensure that in this case we have a correctly set up
 +     * zero FPSCR/FPCR/FPSR.
 +     *
 +     * This is not strictly needed because FPSCR is zero out of reset, but
 +     * it avoids the possibility of future confusing migration bugs if some
 +     * future architecture change makes the reset value non-zero.
 +     */
 +    vfp_set_fpscr(env, 0);
 +
      /*
       * Pre-initialize irq_line_state to a value that's never valid as
       * real data, so cpu_post_load() can tell whether we've seen the
 --
 .34.1

-New patch
+[PULL 21/72] tests/fp: Explicitly set 3-NaN propagation rule
+Explicitly set a rule in the softfloat tests for propagating NaNs in
+the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
+so we should select here the Arm rule of float_3nan_prop_s_cab.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c | 1 +
+ tests/fp/fp-test.c  | 1 +
+files changed, 2 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+     f = bench_funcs[operation][precision];
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      * doesn't specify match those used by the Arm architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-New patch
+[PULL 22/72] target/arm: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for Arm, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
+---
+ target/arm/cpu.c               | 5 +++++
+ fpu/softfloat-specialize.c.inc | 8 +-------
+files changed, 6 insertions(+), 7 deletions(-)
+diff --git a/target/arm/cpu.c b/target/arm/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/cpu.c
++++ b/target/arm/cpu.c
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
+  *  * tininess-before-rounding
+  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
+  *    operand A over operand B (see FPProcessNaNs() pseudocode)
++ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
++ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
++ *    but note that for QEMU muladd is a * b + c, whereas for
++ *    the pseudocode function the arguments are in the order c, a, b.
+  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+  *    and the input NaN if it is signalling
+  */
+@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
+ {
+     set_float_detect_tininess(float_tininess_before_rounding, s);
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+     }
+     if (rule == float_3nan_prop_none) {
+-#if defined(TARGET_ARM)
+-        /*
+-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+-         */
+-        rule = float_3nan_prop_s_cab;
+-#elif defined(TARGET_MIPS)
++#if defined(TARGET_MIPS)
+         if (snan_bit_is_one(status)) {
+             rule = float_3nan_prop_s_abc;
+         } else {
+--
+.34.1

-New patch
+[PULL 23/72] target/loongarch: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for loongarch, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 1 +
+ fpu/softfloat-specialize.c.inc    | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      * case sets InvalidOp and returns the input value 'c'
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_LOONGARCH64)
+-        rule = float_3nan_prop_s_cab;
+ #elif defined(TARGET_PPC)
+         /*
+          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+--
+.34.1

-New patch
+[PULL 24/72] target/ppc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for PPC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c          | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 6 ------
+files changed, 8 insertions(+), 6 deletions(-)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
++    /*
++     * NaN propagation for fused multiply-add:
++     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
++     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
++     * whereas QEMU labels the operands as (a * b) + c.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
++    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
+     /*
+      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+      * to return an input NaN if we have one (ie c) rather than generating
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_PPC)
+-        /*
+-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+-         */
+-        rule = float_3nan_prop_acb;
+ #elif defined(TARGET_S390X)
+         rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+--
+.34.1

-New patch
+[PULL 25/72] target/s390x: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for s390x, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c             | 1 +
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_detect_tininess(float_tininess_before_rounding,
+                                   &env->fpu_status);
+         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
++        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
+        /* fall through */
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_S390X)
+-        rule = float_3nan_prop_s_abc;
+ #elif defined(TARGET_SPARC)
+         rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+--
+.34.1

-New patch
+[PULL 26/72] target/sparc: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for SPARC, and remove the
+ifdef from pickNaNMulAdd().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 --
+files changed, 2 insertions(+), 2 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+      * the CPU state struct so it won't get zeroed on reset.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
++    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
++    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         } else {
+             rule = float_3nan_prop_s_cab;
+         }
+-#elif defined(TARGET_SPARC)
+-        rule = float_3nan_prop_s_cba;
+ #elif defined(TARGET_XTENSA)
+         if (status->use_first_nan) {
+             rule = float_3nan_prop_abc;
+--
+.34.1

-[PULL 01/24] target/arm: Correct comments about M-profile FPSCR
+[PULL 27/72] target/mips: Set Float3NaNPropRule explicitly
-The M-profile FPSCR LTPSIZE is bits [18:16]; this is the same
+Set the Float3NaNPropRule explicitly for Arm, and remove the
-field as A-profile FPSCR Len, not Stride. Correct the comment
+ifdef from pickNaNMulAdd().
 in vfp_get_fpscr().
 We also implemented M-profile FPSCR.QC, but forgot to delete
 a TODO comment from vfp_set_fpscr(); remove it now.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-2-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
 ---
- target/arm/vfp_helper.c | 5 ++---
+ target/mips/fpu_helper.h       | 4 ++++
-file changed, 2 insertions(+), 3 deletions(-)
+ target/mips/msa.c              | 3 +++
  fpu/softfloat-specialize.c.inc | 8 +-------
 files changed, 8 insertions(+), 7 deletions(-)
-diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/vfp_helper.c
+--- a/target/mips/fpu_helper.h
-+++ b/target/arm/vfp_helper.c
++++ b/target/mips/fpu_helper.h
-@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
-             | (env->vfp.vec_stride << 20);
+ {
      bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
      FloatInfZeroNaNRule izn_rule;
 +    Float3NaNPropRule nan3_rule;
      /*
--     * M-profile LTPSIZE overlaps A-profile Stride; whichever of the
+      * With nan2008, SNaNs are silenced in the usual way.
--     * two is not applicable to this CPU will always be zero.
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
 +     * M-profile LTPSIZE is the same bits [18:16] as A-profile Len; whichever
 +     * of the two is not applicable to this CPU will always be zero.
       */
-     fpscr |= env->v7m.ltpsize << 16;
+     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
-@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
++    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
-         /*
++    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
-          * The bit we set within fpscr_q is arbitrary; the register as a
++
-          * whole being zero/non-zero is what counts.
+ }
--         * TODO: M-profile MVE also has a QC bit.
-          */
+ static inline void restore_fp_status(CPUMIPSState *env)
-         env->vfp.qc[0] = val & FPCR_QC;
+diff --git a/target/mips/msa.c b/target/mips/msa.c
-         env->vfp.qc[1] = 0;
+index XXXXXXX..XXXXXXX 100644
 --- a/target/mips/msa.c
 +++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
      set_float_2nan_prop_rule(float_2nan_prop_s_ab,
                               &env->active_tc.msa_fp_status);
 +    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
 +                             &env->active_tc.msa_fp_status);
 +
      /* clear float_status exception flags */
      set_float_exception_flags(0, &env->active_tc.msa_fp_status);
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      }
      if (rule == float_3nan_prop_none) {
 -#if defined(TARGET_MIPS)
 -        if (snan_bit_is_one(status)) {
 -            rule = float_3nan_prop_s_abc;
 -        } else {
 -            rule = float_3nan_prop_s_cab;
 -        }
 -#elif defined(TARGET_XTENSA)
 +#if defined(TARGET_XTENSA)
          if (status->use_first_nan) {
              rule = float_3nan_prop_abc;
          } else {
 --
 .34.1

-[PULL 09/24] target/arm: Allow FPCR bits that aren't in FPSCR
+[PULL 28/72] target/xtensa: Set Float3NaNPropRule explicitly
-In order to allow FPCR bits that aren't in the FPSCR (like the new
+Set the Float3NaNPropRule explicitly for xtensa, and remove the
-bits that are defined for FEAT_AFP), we need to make sure that writes
+ifdef from pickNaNMulAdd().
 to the FPSCR only write to the bits of FPCR that are architecturally
 mapped, and not the others.
 Implement this with a new function vfp_set_fpcr_masked() which
 takes a mask of which bits to update.
 (We could do the same for FPSR, but we leave that until we actually
 are likely to need it.)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-10-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
 ---
- target/arm/vfp_helper.c | 54 ++++++++++++++++++++++++++---------------
+ target/xtensa/fpu_helper.c     | 2 ++
-file changed, 34 insertions(+), 20 deletions(-)
+ fpu/softfloat-specialize.c.inc | 8 --------
 files changed, 2 insertions(+), 8 deletions(-)
-diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/vfp_helper.c
+--- a/target/xtensa/fpu_helper.c
-+++ b/target/arm/vfp_helper.c
++++ b/target/xtensa/fpu_helper.c
-@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
+@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
-     set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
+     set_use_first_nan(use_first, &env->fp_status);
      set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                               &env->fp_status);
 +    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
 +                             &env->fp_status);
  }
--static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
+ void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
-+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val, uint32_t mask)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
- {
+index XXXXXXX..XXXXXXX 100644
-     uint64_t changed = env->vfp.fpcr;
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
-     changed ^= val;
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
 +    changed &= mask;
      if (changed & (3 << 22)) {
          int i = (val >> 22) & 3;
          switch (i) {
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
  {
  }
 -static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
 +static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val, uint32_t mask)
  {
  }
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
      env->vfp.fpsr = val;
  }
 -void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 +static void vfp_set_fpcr_masked(CPUARMState *env, uint32_t val, uint32_t mask)
  {
 +    /*
 +     * We only set FPCR bits defined by mask, and leave the others alone.
 +     * We assume the mask is sensible (e.g. doesn't try to set only
 +     * part of a field)
 +     */
      ARMCPU *cpu = env_archcpu(env);
      /* When ARMv8.2-FP16 is not supported, FZ16 is RES0.  */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
          val &= ~FPCR_FZ16;
      }
--    vfp_set_fpcr_to_host(env, val);
+     if (rule == float_3nan_prop_none) {
-+    vfp_set_fpcr_to_host(env, val, mask);
+-#if defined(TARGET_XTENSA)
+-        if (status->use_first_nan) {
--    if (!arm_feature(env, ARM_FEATURE_M)) {
+-            rule = float_3nan_prop_abc;
--        /*
+-        } else {
--         * Short-vector length and stride; on M-profile these bits
+-            rule = float_3nan_prop_cba;
--         * are used for different purposes.
+-        }
--         * We can't make this conditional be "if MVFR0.FPShVec != 0",
+-#else
--         * because in v7A no-short-vector-support cores still had to
+         rule = float_3nan_prop_abc;
--         * allow Stride/Len to be written with the only effect that
+-#endif
 -         * some insns are required to UNDEF if the guest sets them.
 -         */
 -        env->vfp.vec_len = extract32(val, 16, 3);
 -        env->vfp.vec_stride = extract32(val, 20, 2);
 -    } else if (cpu_isar_feature(aa32_mve, cpu)) {
 -        env->v7m.ltpsize = extract32(val, FPCR_LTPSIZE_SHIFT,
 -                                     FPCR_LTPSIZE_LENGTH);
 +    if (mask & (FPCR_LEN_MASK | FPCR_STRIDE_MASK)) {
 +        if (!arm_feature(env, ARM_FEATURE_M)) {
 +            /*
 +             * Short-vector length and stride; on M-profile these bits
 +             * are used for different purposes.
 +             * We can't make this conditional be "if MVFR0.FPShVec != 0",
 +             * because in v7A no-short-vector-support cores still had to
 +             * allow Stride/Len to be written with the only effect that
 +             * some insns are required to UNDEF if the guest sets them.
 +             */
 +            env->vfp.vec_len = extract32(val, 16, 3);
 +            env->vfp.vec_stride = extract32(val, 20, 2);
 +        } else if (cpu_isar_feature(aa32_mve, cpu)) {
 +            env->v7m.ltpsize = extract32(val, FPCR_LTPSIZE_SHIFT,
 +                                         FPCR_LTPSIZE_LENGTH);
 +        }
      }
-     /*
+     assert(rule != float_3nan_prop_none);
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
       * bits.
       */
      val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
 -    env->vfp.fpcr = val;
 +    env->vfp.fpcr &= ~mask;
 +    env->vfp.fpcr |= val;
 +}
 +
 +void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 +{
 +    vfp_set_fpcr_masked(env, val, MAKE_64BIT_MASK(0, 32));
  }
  void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
  {
 -    vfp_set_fpcr(env, val & FPSCR_FPCR_MASK);
 +    vfp_set_fpcr_masked(env, val, FPSCR_FPCR_MASK);
      vfp_set_fpsr(env, val & FPSCR_FPSR_MASK);
  }
 --
 .34.1

-New patch
+[PULL 29/72] target/i386: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for i386.  We had no
+i386-specific behaviour in the old ifdef ladder, so we were using the
+default "prefer a then b then c" fallback; this is actually the
+correct per-the-spec handling for i386.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c | 1 +
+file changed, 1 insertion(+)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      * there are multiple input NaNs they are selected in the order a, b, c.
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+--
+.34.1

-New patch
+[PULL 30/72] target/hppa: Set Float3NaNPropRule explicitly
+Set the Float3NaNPropRule explicitly for HPPA, and remove the
+ifdef from pickNaNMulAdd().
+HPPA is the only target that was using the default branch of the
+ifdef ladder (other targets either do not use muladd or set
+default_nan_mode), so we can remove the ifdef fallback entirely now
+(allowing the "rule not set" case to fall into the default of the
+switch statement and assert).
+We add a TODO note that the HPPA rule is probably wrong; this is
+not a behavioural change for this refactoring.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 8 ++++++++
+ fpu/softfloat-specialize.c.inc | 4 ----
+files changed, 8 insertions(+), 4 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+      * HPPA does note implement a CPU reset method at all...
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
++    /*
++     * TODO: The HPPA architecture reference only documents its NaN
++     * propagation rule for 2-operand operations. Testing on real hardware
++     * might be necessary to confirm whether this order for muladd is correct.
++     * Not preferring the SNaN is almost certainly incorrect as it diverges
++     * from the documented rules for 2-operand operations.
++     */
++    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+ }
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
+         }
+     }
+-    if (rule == float_3nan_prop_none) {
+-        rule = float_3nan_prop_abc;
+-    }
+-
+     assert(rule != float_3nan_prop_none);
+     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+         /* We have at least one SNaN input and should prefer it */
+--
+.34.1

-New patch
+[PULL 31/72] fpu: Remove use_first_nan field from float_status
+The use_first_nan field in float_status was an xtensa-specific way to
+select at runtime from two different NaN propagation rules.  Now that
+xtensa is using the target-agnostic NaN propagation rule selection
+that we've just added, we can remove use_first_nan, because there is
+no longer any code that reads it.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
+---
+ include/fpu/softfloat-helpers.h | 5 -----
+ include/fpu/softfloat-types.h   | 1 -
+ target/xtensa/fpu_helper.c      | 1 -
+files changed, 7 deletions(-)
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/fpu/softfloat-helpers.h
++++ b/include/fpu/softfloat-helpers.h
+@@ -XXX,XX +XXX,XX @@ static inline void set_snan_bit_is_one(bool val, float_status *status)
+     status->snan_bit_is_one = val;
+ }
+-static inline void set_use_first_nan(bool val, float_status *status)
+-{
+-    status->use_first_nan = val;
+-}
+-
+ static inline void set_no_signaling_nans(bool val, float_status *status)
+ {
+     status->no_signaling_nans = val;
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
+index XXXXXXX..XXXXXXX 100644
+--- a/include/fpu/softfloat-types.h
++++ b/include/fpu/softfloat-types.h
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
+      * softfloat-specialize.inc.c)
+      */
+     bool snan_bit_is_one;
+-    bool use_first_nan;
+     bool no_signaling_nans;
+     /* should overflowed results subtract re_bias to its exponent? */
+     bool rebias_overflow;
+diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/xtensa/fpu_helper.c
++++ b/target/xtensa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ static const struct {
+ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
+ {
+-    set_use_first_nan(use_first, &env->fp_status);
+     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
+                              &env->fp_status);
+     set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
+--
+.34.1

-New patch
+[PULL 32/72] target/m68k: Don't pass NULL float_status to floatx80_default_nan()
+Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
+to get the NaN bit pattern to reset the FPU registers. This
+works because it happens that our implementation of
+floatx80_default_nan() doesn't actually look at the float_status
+pointer except for TARGET_MIPS. However, this isn't guaranteed,
+and to be able to remove the ifdef in floatx80_default_nan()
+we're going to need a real float_status here.
+Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
+earlier, and thus can pass it to floatx80_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c | 12 +++++++-----
+file changed, 7 insertions(+), 5 deletions(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+     CPUState *cs = CPU(obj);
+     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
+     CPUM68KState *env = cpu_env(cs);
+-    floatx80 nan = floatx80_default_nan(NULL);
++    floatx80 nan;
+     int i;
+     if (mcc->parent_phases.hold) {
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+ #else
+     cpu_m68k_set_sr(env, SR_S | SR_I);
+ #endif
+-    for (i = 0; i < 8; i++) {
+-        env->fregs[i].d = nan;
+-    }
+-    cpu_m68k_set_fpcr(env, 0);
+     /*
+      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
+      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++
++    nan = floatx80_default_nan(&env->fp_status);
++    for (i = 0; i < 8; i++) {
++        env->fregs[i].d = nan;
++    }
++    cpu_m68k_set_fpcr(env, 0);
+     env->fpsr = 0;
+     /* TODO: We should set PC from the interrupt vector.  */
+--
+.34.1

-New patch
+[PULL 33/72] softfloat: Create floatx80 default NaN from parts64_default_nan
+We create our 128-bit default NaN by calling parts64_default_nan()
+and then adjusting the result.  We can do the same trick for creating
+the floatx80 default NaN, which lets us drop a target ifdef.
+floatx80 is used only by:
+ i386
+ m68k
+ arm nwfpe old floating-point emulation emulation support
+    (which is essentially dead, especially the parts involving floatx80)
+ PPC (only in the xsrqpxp instruction, which just rounds an input
+    value by converting to floatx80 and back, so will never generate
+    the default NaN)
+The floatx80 default NaN as currently implemented is:
+ m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
+ i386: sign = 1, exp = 1...1, int = 1, frac = 10...0
+These are the same as the parts64_default_nan for these architectures.
+This is technically a possible behaviour change for arm linux-user
+nwfpe emulation emulation, because the default NaN will now have the
+sign bit clear.  But we were already generating a different floatx80
+default NaN from the real kernel emulation we are supposedly
+following, which appears to use an all-bits-1 value:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267
+This won't affect the only "real" use of the nwfpe emulation, which
+is ancient binaries that used it as part of the old floating point
+calling convention; that only uses loads and stores of 32 and 64 bit
+floats, not any of the floatx80 behaviour the original hardware had.
+We also get the nwfpe float64 default NaN value wrong:
+ https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
+so if we ever cared about this obscure corner the right fix would be
+to correct that so nwfpe used its own default-NaN setting rather
+than the Arm VFP one.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
+---
+ fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
+file changed, 10 insertions(+), 10 deletions(-)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
+ floatx80 floatx80_default_nan(float_status *status)
+ {
+     floatx80 r;
++    /*
++     * Extrapolate from the choices made by parts64_default_nan to fill
++     * in the floatx80 format. We assume that floatx80's explicit
++     * integer bit is always set (this is true for i386 and m68k,
++     * which are the only real users of this format).
++     */
++    FloatParts64 p64;
++    parts64_default_nan(&p64, status);
+-    /* None of the targets that have snan_bit_is_one use floatx80.  */
+-    assert(!snan_bit_is_one(status));
+-#if defined(TARGET_M68K)
+-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
+-    r.high = 0x7FFF;
+-#else
+-    /* X86 */
+-    r.low = UINT64_C(0xC000000000000000);
+-    r.high = 0xFFFF;
+-#endif
++    r.high = 0x7FFF | (p64.sign << 15);
++    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
+     return r;
+ }
+--
+.34.1

-New patch
+[PULL 34/72] target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
+In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
+a zero-initialized float_status struct to float32_is_quiet_nan() and
+float64_is_quiet_nan(), with the cryptic comment "for
+snan_bit_is_one".
+This pattern appears to have been copied from target/riscv, where it
+is used because the functions there do not have ready access to the
+CPU state struct. The comment presumably refers to the fact that the
+main reason the is_quiet_nan() functions want the float_state is
+because they want to know about the snan_bit_is_one config.
+In the loongarch helpers, though, we have the CPU state struct
+to hand. Use the usual env->fp_status here. This avoids our needing
+to track that we need to update the initializer of the local
+float_status structs when the core softfloat code adds new
+options for targets to configure their behaviour.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
+     } else if (float32_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float32_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
+     } else if (float64_is_zero_or_denormal(f)) {
+         return sign ? 1 << 4 : 1 << 8;
+     } else if (float64_is_any_nan(f)) {
+-        float_status s = { }; /* for snan_bit_is_one */
+-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
++        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
+     } else {
+         return sign ? 1 << 3 : 1 << 7;
+     }
+--
+.34.1

-New patch
+[PULL 35/72] target/m68k: In frem helper, initialize local float_status from env->fp_status
+In the frem helper, we have a local float_status because we want to
+execute the floatx80_div() with a custom rounding mode.  Instead of
+zero-initializing the local float_status and then having to set it up
+with the m68k standard behaviour (including the NaN propagation rule
+and copying the rounding precision from env->fp_status), initialize
+it as a complete copy of env->fp_status. This will avoid our having
+to add new code in this function for every new config knob we add
+to fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
+---
+ target/m68k/fpu_helper.c | 6 ++----
+file changed, 2 insertions(+), 4 deletions(-)
+diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/fpu_helper.c
++++ b/target/m68k/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
+     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
+     if (!floatx80_is_any_nan(fp_rem)) {
+-        float_status fp_status = { };
++        /* Use local temporary fp_status to set different rounding mode */
++        float_status fp_status = env->fp_status;
+         uint32_t quotient;
+         int sign;
+         /* Calculate quotient directly using round to nearest mode */
+-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
+         set_float_rounding_mode(float_round_nearest_even, &fp_status);
+-        set_floatx80_rounding_precision(
+-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
+         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
+         sign = extractFloatx80Sign(fp_quot.d);
+--
+.34.1

-New patch
+[PULL 36/72] target/m68k: Init local float_status from env fp_status in gdb get/set reg
+In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
+from float64 to floatx80 using a scratch float_status, because we
+don't want the conversion to affect the CPU's floating point exception
+status. Currently we use a zero-initialized float_status. This will
+get steadily more awkward as we add config knobs to float_status
+that the target must initialize. Avoid having to add any of that
+configuration here by instead initializing our local float_status
+from the env->fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
+---
+ target/m68k/helper.c | 6 ++++--
+file changed, 4 insertions(+), 2 deletions(-)
+diff --git a/target/m68k/helper.c b/target/m68k/helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/helper.c
++++ b/target/m68k/helper.c
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
+     }
+     switch (n) {
+@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
+     CPUM68KState *env = &cpu->env;
+     if (n < 8) {
+-        float_status s = {};
++        /* Use scratch float_status so any exceptions don't change CPU state */
++        float_status s = env->fp_status;
+         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
+         return 8;
+     }
+--
+.34.1

-[PULL 07/24] target/arm: Rename FPCR_ QC, NZCV macros to FPSR_
+[PULL 37/72] target/sparc: Initialize local scratch float_status from env->fp_status
-The QC, N, Z, C, V bits live in the FPSR, not the FPCR. Rename the
+In the helper functions flcmps and flcmpd we use a scratch float_status
-macros that define these bits accordingly.
+so that we don't change the CPU state if the comparison raises any
 floating point exception flags. Instead of zero-initializing this
 scratch float_status, initialize it as a copy of env->fp_status. This
 avoids the need to explicitly initialize settings like the NaN
 propagation rule or others we might add to softfloat in future.
 To do this we need to pass the CPU env pointer in to the helper.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-8-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
 ---
- target/arm/cpu.h                  | 17 ++++++++++-------
+ target/sparc/helper.h     | 4 ++--
- target/arm/tcg/mve_helper.c       |  8 ++++----
+ target/sparc/fop_helper.c | 8 ++++----
- target/arm/tcg/translate-m-nocp.c | 16 ++++++++--------
+ target/sparc/translate.c  | 4 ++--
- target/arm/tcg/translate-vfp.c    |  2 +-
+files changed, 8 insertions(+), 8 deletions(-)
  target/arm/vfp_helper.c           |  8 ++++----
 files changed, 27 insertions(+), 24 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/target/sparc/helper.h b/target/sparc/helper.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/target/sparc/helper.h
-+++ b/target/arm/cpu.h
++++ b/target/sparc/helper.h
-@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
+@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
- #define FPSR_MASK 0xf800009f
+ DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
- #define FPCR_MASK 0x07ff9f00
+ DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
+ DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
-+/* FPCR bits */
+-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
- #define FPCR_IOE    (1 << 8)    /* Invalid Operation exception trap enable */
+-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
- #define FPCR_DZE    (1 << 9)    /* Divide by Zero exception trap enable */
++DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
- #define FPCR_OFE    (1 << 10)   /* Overflow exception trap enable */
++DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
-@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
+ DEF_HELPER_2(raise_exception, noreturn, env, int)
- #define FPCR_FZ     (1 << 24)   /* Flush-to-zero enable bit */
- #define FPCR_DN     (1 << 25)   /* Default NaN enable bit */
+ DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
- #define FPCR_AHP    (1 << 26)   /* Alternative half-precision */
+diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
 -#define FPCR_QC     (1 << 27)   /* Cumulative saturation bit */
 -#define FPCR_V      (1 << 28)   /* FP overflow flag */
 -#define FPCR_C      (1 << 29)   /* FP carry flag */
 -#define FPCR_Z      (1 << 30)   /* FP zero flag */
 -#define FPCR_N      (1 << 31)   /* FP negative flag */
  #define FPCR_LTPSIZE_SHIFT 16   /* LTPSIZE, M-profile only */
  #define FPCR_LTPSIZE_MASK (7 << FPCR_LTPSIZE_SHIFT)
  #define FPCR_LTPSIZE_LENGTH 3
 -#define FPCR_NZCV_MASK (FPCR_N | FPCR_Z | FPCR_C | FPCR_V)
 -#define FPCR_NZCVQC_MASK (FPCR_NZCV_MASK | FPCR_QC)
 +/* FPSR bits */
 +#define FPSR_QC     (1 << 27)   /* Cumulative saturation bit */
 +#define FPSR_V      (1 << 28)   /* FP overflow flag */
 +#define FPSR_C      (1 << 29)   /* FP carry flag */
 +#define FPSR_Z      (1 << 30)   /* FP zero flag */
 +#define FPSR_N      (1 << 31)   /* FP negative flag */
 +
 +#define FPSR_NZCV_MASK (FPSR_N | FPSR_Z | FPSR_C | FPSR_V)
 +#define FPSR_NZCVQC_MASK (FPSR_NZCV_MASK | FPSR_QC)
  /**
   * vfp_get_fpsr: read the AArch64 FPSR
 diff --git a/target/arm/tcg/mve_helper.c b/target/arm/tcg/mve_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/mve_helper.c
+--- a/target/sparc/fop_helper.c
-+++ b/target/arm/tcg/mve_helper.c
++++ b/target/sparc/fop_helper.c
-@@ -XXX,XX +XXX,XX @@ static void do_vadc(CPUARMState *env, uint32_t *d, uint32_t *n, uint32_t *m,
+@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
+     return finish_fcmp(env, r, GETPC());
      if (update_flags) {
          /* Store C, clear NZV. */
 -        env->vfp.fpsr &= ~FPCR_NZCV_MASK;
 -        env->vfp.fpsr |= carry_in * FPCR_C;
 +        env->vfp.fpsr &= ~FPSR_NZCV_MASK;
 +        env->vfp.fpsr |= carry_in * FPSR_C;
      }
      mve_advance_vpt(env);
  }
- void HELPER(mve_vadc)(CPUARMState *env, void *vd, void *vn, void *vm)
+-uint32_t helper_flcmps(float32 src1, float32 src2)
 +uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
  {
--    bool carry_in = env->vfp.fpsr & FPCR_C;
+     /*
-+    bool carry_in = env->vfp.fpsr & FPSR_C;
+      * FLCMP never raises an exception nor modifies any FSR fields.
-     do_vadc(env, vd, vn, vm, 0, carry_in, false);
+      * Perform the comparison with a dummy fp environment.
       */
 -    float_status discard = { };
 +    float_status discard = env->fp_status;
      FloatRelation r;
      set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
      g_assert_not_reached();
  }
- void HELPER(mve_vsbc)(CPUARMState *env, void *vd, void *vn, void *vm)
+-uint32_t helper_flcmpd(float64 src1, float64 src2)
 +uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
  {
--    bool carry_in = env->vfp.fpsr & FPCR_C;
+-    float_status discard = { };
-+    bool carry_in = env->vfp.fpsr & FPSR_C;
++    float_status discard = env->fp_status;
-     do_vadc(env, vd, vn, vm, -1, carry_in, false);
+     FloatRelation r;
      set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
 diff --git a/target/sparc/translate.c b/target/sparc/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/sparc/translate.c
 +++ b/target/sparc/translate.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
      src1 = gen_load_fpr_F(dc, a->rs1);
      src2 = gen_load_fpr_F(dc, a->rs2);
 -    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
 +    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
      return advance_pc(dc);
  }
-diff --git a/target/arm/tcg/translate-m-nocp.c b/target/arm/tcg/translate-m-nocp.c
+@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
-index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-m-nocp.c
+     src1 = gen_load_fpr_D(dc, a->rs1);
-+++ b/target/arm/tcg/translate-m-nocp.c
+     src2 = gen_load_fpr_D(dc, a->rs2);
-@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
+-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
-         if (dc_isar_feature(aa32_mve, s)) {
++    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
-             /* QC is only present for MVE; otherwise RES0 */
+     return advance_pc(dc);
              TCGv_i32 qc = tcg_temp_new_i32();
 -            tcg_gen_andi_i32(qc, tmp, FPCR_QC);
 +            tcg_gen_andi_i32(qc, tmp, FPSR_QC);
              /*
               * The 4 vfp.qc[] fields need only be "zero" vs "non-zero";
               * here writing the same value into all elements is simplest.
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
              tcg_gen_gvec_dup_i32(MO_32, offsetof(CPUARMState, vfp.qc),
 , 16, qc);
          }
 -        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
          fpscr = load_cpu_field_low32(vfp.fpsr);
 -        tcg_gen_andi_i32(fpscr, fpscr, ~FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(fpscr, fpscr, ~FPSR_NZCV_MASK);
          tcg_gen_or_i32(fpscr, fpscr, tmp);
          store_cpu_field_low32(fpscr, vfp.fpsr);
          break;
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
          tcg_gen_deposit_i32(control, control, sfpa,
                              R_V7M_CONTROL_SFPA_SHIFT, 1);
          store_cpu_field(control, v7m.control[M_REG_S]);
 -        tcg_gen_andi_i32(tmp, tmp, ~FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(tmp, tmp, ~FPSR_NZCV_MASK);
          gen_helper_vfp_set_fpscr(tcg_env, tmp);
          s->base.is_jmp = DISAS_UPDATE_NOCHAIN;
          break;
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
      case ARM_VFP_FPSCR_NZCVQC:
          tmp = tcg_temp_new_i32();
          gen_helper_vfp_get_fpscr(tmp, tcg_env);
 -        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCVQC_MASK);
 +        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCVQC_MASK);
          storefn(s, opaque, tmp, true);
          break;
      case QEMU_VFP_FPSCR_NZCV:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
           * helper call for the "VMRS to CPSR.NZCV" insn.
           */
          tmp = load_cpu_field_low32(vfp.fpsr);
 -        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
          storefn(s, opaque, tmp, true);
          break;
      case ARM_VFP_FPCXT_S:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
          tmp = tcg_temp_new_i32();
          sfpa = tcg_temp_new_i32();
          gen_helper_vfp_get_fpscr(tmp, tcg_env);
 -        tcg_gen_andi_i32(tmp, tmp, ~FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(tmp, tmp, ~FPSR_NZCV_MASK);
          control = load_cpu_field(v7m.control[M_REG_S]);
          tcg_gen_andi_i32(sfpa, control, R_V7M_CONTROL_SFPA_MASK);
          tcg_gen_shli_i32(sfpa, sfpa, 31 - R_V7M_CONTROL_SFPA_SHIFT);
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
          sfpa = tcg_temp_new_i32();
          fpscr = tcg_temp_new_i32();
          gen_helper_vfp_get_fpscr(fpscr, tcg_env);
 -        tcg_gen_andi_i32(tmp, fpscr, ~FPCR_NZCV_MASK);
 +        tcg_gen_andi_i32(tmp, fpscr, ~FPSR_NZCV_MASK);
          control = load_cpu_field(v7m.control[M_REG_S]);
          tcg_gen_andi_i32(sfpa, control, R_V7M_CONTROL_SFPA_MASK);
          tcg_gen_shli_i32(sfpa, sfpa, 31 - R_V7M_CONTROL_SFPA_SHIFT);
 diff --git a/target/arm/tcg/translate-vfp.c b/target/arm/tcg/translate-vfp.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/translate-vfp.c
 +++ b/target/arm/tcg/translate-vfp.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VMSR_VMRS(DisasContext *s, arg_VMSR_VMRS *a)
          case ARM_VFP_FPSCR:
              if (a->rt == 15) {
                  tmp = load_cpu_field_low32(vfp.fpsr);
 -                tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
 +                tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
              } else {
                  tmp = tcg_temp_new_i32();
                  gen_helper_vfp_get_fpscr(tmp, tcg_env);
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/vfp_helper.c
 +++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env)
      fpsr |= vfp_get_fpsr_from_host(env);
      i = env->vfp.qc[0] | env->vfp.qc[1] | env->vfp.qc[2] | env->vfp.qc[3];
 -    fpsr |= i ? FPCR_QC : 0;
 +    fpsr |= i ? FPSR_QC : 0;
      return fpsr;
  }
-@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
-          * The bit we set within vfp.qc[] is arbitrary; the array as a
-          * whole being zero/non-zero is what counts.
-          */
--        env->vfp.qc[0] = val & FPCR_QC;
-+        env->vfp.qc[0] = val & FPSR_QC;
-         env->vfp.qc[1] = 0;
-         env->vfp.qc[2] = 0;
-         env->vfp.qc[3] = 0;
-@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
-      * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
-      * and zero any of the other FPSR bits.
-      */
--    val &= FPCR_NZCV_MASK;
-+    val &= FPSR_NZCV_MASK;
-     env->vfp.fpsr = val;
- }
-@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vjcvt)(float64 value, CPUARMState *env)
-     uint32_t z = (pair >> 32) == 0;
-     /* Store Z, clear NCV, in FPSCR.NZCV.  */
--    env->vfp.fpsr = (env->vfp.fpsr & ~FPCR_NZCV_MASK) | (z * FPCR_Z);
-+    env->vfp.fpsr = (env->vfp.fpsr & ~FPSR_NZCV_MASK) | (z * FPSR_Z);
-     return result;
- }
 --
 .34.1

-New patch
+[PULL 38/72] target/ppc: Use env->fp_status in helper_compute_fprf functions
+In the helper_compute_fprf functions, we pass a dummy float_status
+in to the is_signaling_nan() function. This is unnecessary, because
+we have convenient access to the CPU env pointer here and that
+is already set up with the correct values for the snan_bit_is_one
+and no_signaling_nans config settings. is_signaling_nan() doesn't
+ever update the fp_status with any exception flags, so there is
+no reason not to use env->fp_status here.
+Use env->fp_status instead of the dummy fp_status.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
+---
+ target/ppc/fpu_helper.c | 3 +--
+file changed, 1 insertion(+), 2 deletions(-)
+diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/fpu_helper.c
++++ b/target/ppc/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
+     } else if (tp##_is_infinity(arg)) {                           \
+         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
+     } else {                                                      \
+-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
+-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
++        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
+             fprf = 0x00 << FPSCR_FPRF;                            \
+         } else {                                                  \
+             fprf = 0x11 << FPSCR_FPRF;                            \
+--
+.34.1

-[PULL 17/24] hw/misc: In STM32L4x5 EXTI, handle direct interrupts
+[PULL 39/72] target/arm: Copy entire float_status in is_ebf
-From: Inès Varhol <ines.varhol@telecom-paris.fr>
+From: Richard Henderson <richard.henderson@linaro.org>
-The previous implementation for EXTI interrupts only handled
+Now that float_status has a bunch of fp parameters,
-"configurable" interrupts, like those originating from STM32L4x5 SYSCFG
+it is easier to copy an existing structure than create
-(the only device currently connected to the EXTI up until now).
+one from scratch.  Begin by copying the structure that
 corresponds to the FPSR and make only the adjustments
 required for BFloat16 semantics.
-In order to connect STM32L4x5 USART to the EXTI, this commit adds
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-handling for direct interrupts (interrupts without configurable edge).
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
 Message-id: 20240707085927.122867-3-ines.varhol@telecom-paris.fr
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/misc/stm32l4x5_exti.c | 7 +++++++
+ target/arm/tcg/vec_helper.c | 20 +++++++-------------
-file changed, 7 insertions(+)
+file changed, 7 insertions(+), 13 deletions(-)
-diff --git a/hw/misc/stm32l4x5_exti.c b/hw/misc/stm32l4x5_exti.c
+diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/misc/stm32l4x5_exti.c
+--- a/target/arm/tcg/vec_helper.c
-+++ b/hw/misc/stm32l4x5_exti.c
++++ b/target/arm/tcg/vec_helper.c
-@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_set_irq(void *opaque, int irq, int level)
+@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
-         return;
+      * no effect on AArch32 instructions.
       */
      bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
 -    *statusp = (float_status){
 -        .tininess_before_rounding = float_tininess_before_rounding,
 -        .float_rounding_mode = float_round_to_odd_inf,
 -        .flush_to_zero = true,
 -        .flush_inputs_to_zero = true,
 -        .default_nan_mode = true,
 -    };
 +
 +    *statusp = env->vfp.fp_status;
 +    set_default_nan_mode(true, statusp);
      if (ebf) {
 -        float_status *fpst = &env->vfp.fp_status;
 -        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
 -        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
 -        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
 -
          /* EBF=1 needs to do a step with round-to-odd semantics */
          *oddstatusp = *statusp;
          set_float_rounding_mode(float_round_to_odd, oddstatusp);
 +    } else {
 +        set_flush_to_zero(true, statusp);
 +        set_flush_inputs_to_zero(true, statusp);
 +        set_float_rounding_mode(float_round_to_odd_inf, statusp);
      }
+-
-+    /* In case of a direct line interrupt */
+     return ebf;
-+    if (extract32(exti_romask[bank], irq, 1)) {
+ }
 +        qemu_set_irq(s->irq[oirq], level);
 +        return;
 +    }
 +
 +    /* In case of a configurable interrupt */
      if ((level && extract32(s->rtsr[bank], irq, 1)) ||
          (!level && extract32(s->ftsr[bank], irq, 1))) {
 --
 .34.1

-[PULL 03/24] target/arm: Make vfp_set_fpscr() call vfp_set_{fpcr, fpsr}
+[PULL 40/72] fpu: Allow runtime choice of default NaN value
-Make vfp_set_fpscr() call vfp_set_fpsr() and vfp_set_fpcr()
+Currently we hardcode the default NaN value in parts64_default_nan()
-instead of the other way around.
+using a compile-time ifdef ladder. This is awkward for two cases:
  * for single-QEMU-binary we can't hard-code target-specifics like this
  * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
    (specifically the sign bit is different)
-The masking we do when getting and setting vfp.xregs[ARM_VFP_FPSCR]
+Add a field to float_status to specify the default NaN value; fall
-is a little awkward, but we are going to change where we store the
+back to the old ifdef behaviour if these are not set.
-underlying FPSR and FPCR information in a later commit, so it will
-go away then.
+The default NaN value is specified by setting a uint8_t to a
 pattern corresponding to the sign and upper fraction parts of
 the NaN; the lower bits of the fraction are set from bit 0 of
 the pattern.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-4-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
 ---
- target/arm/cpu.h        |  22 +++++----
+ include/fpu/softfloat-helpers.h | 11 +++++++
- target/arm/vfp_helper.c | 100 ++++++++++++++++++++++++++--------------
+ include/fpu/softfloat-types.h   | 10 ++++++
-files changed, 78 insertions(+), 44 deletions(-)
+ fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
 files changed, 54 insertions(+), 22 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/include/fpu/softfloat-helpers.h
-+++ b/target/arm/cpu.h
++++ b/include/fpu/softfloat-helpers.h
-@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env);
+@@ -XXX,XX +XXX,XX @@ static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
-  */
+     status->float_infzeronan_rule = rule;
  uint32_t vfp_get_fpcr(CPUARMState *env);
 -static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
 -{
 -    uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPSR_MASK) | (val & FPSR_MASK);
 -    vfp_set_fpscr(env, new_fpscr);
 -}
 +/**
 + * vfp_set_fpsr: write the AArch64 FPSR
 + * @env: CPU context
 + * @value: new value
 + */
 +void vfp_set_fpsr(CPUARMState *env, uint32_t value);
 -static inline void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 -{
 -    uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPCR_MASK) | (val & FPCR_MASK);
 -    vfp_set_fpscr(env, new_fpscr);
 -}
 +/**
 + * vfp_set_fpcr: write the AArch64 FPCR
 + * @env: CPU context
 + * @value: new value
 + */
 +void vfp_set_fpcr(CPUARMState *env, uint32_t value);
  enum arm_cpu_mode {
    ARM_CPU_MODE_USR = 0x10,
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/vfp_helper.c
 +++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
      return vfp_exceptbits_from_host(i);
  }
--static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
++static inline void set_float_default_nan_pattern(uint8_t dnan_pattern,
-+static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
++                                                 float_status *status)
 +{
-+    /*
++    status->default_nan_pattern = dnan_pattern;
 +     * The exception flags are ORed together when we read fpscr so we
 +     * only need to preserve the current state in one of our
 +     * float_status values.
 +     */
 +    int i = vfp_exceptbits_to_host(val);
 +    set_float_exception_flags(i, &env->vfp.fp_status);
 +    set_float_exception_flags(0, &env->vfp.fp_status_f16);
 +    set_float_exception_flags(0, &env->vfp.standard_fp_status);
 +    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
 +}
 +
-+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
+ static inline void set_flush_to_zero(bool val, float_status *status)
  {
--    int i;
+     status->flush_to_zero = val;
-     uint32_t changed = env->vfp.xregs[ARM_VFP_FPSCR];
+@@ -XXX,XX +XXX,XX @@ static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status
+     return status->float_infzeronan_rule;
      changed ^= val;
      if (changed & (3 << 22)) {
 -        i = (val >> 22) & 3;
 +        int i = (val >> 22) & 3;
          switch (i) {
          case FPROUNDING_TIEEVEN:
              i = float_round_nearest_even;
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
          set_default_nan_mode(dnan_enabled, &env->vfp.fp_status);
          set_default_nan_mode(dnan_enabled, &env->vfp.fp_status_f16);
      }
 -
 -    /*
 -     * The exception flags are ORed together when we read fpscr so we
 -     * only need to preserve the current state in one of our
 -     * float_status values.
 -     */
 -    i = vfp_exceptbits_to_host(val);
 -    set_float_exception_flags(i, &env->vfp.fp_status);
 -    set_float_exception_flags(0, &env->vfp.fp_status_f16);
 -    set_float_exception_flags(0, &env->vfp.standard_fp_status);
 -    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
  }
- #else
++static inline uint8_t get_float_default_nan_pattern(float_status *status)
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
      return 0;
  }
 -static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
 +static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
 +{
++    return status->default_nan_pattern;
 +}
 +
-+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
+ static inline bool get_flush_to_zero(float_status *status)
  {
- }
+     return status->flush_to_zero;
+diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
-@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpscr(CPUARMState *env)
+index XXXXXXX..XXXXXXX 100644
-     return HELPER(vfp_get_fpscr)(env);
+--- a/include/fpu/softfloat-types.h
- }
++++ b/include/fpu/softfloat-types.h
+@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
--void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
+     /* should denormalised inputs go to zero and set the input_denormal flag? */
-+void vfp_set_fpsr(CPUARMState *env, uint32_t val)
+     bool flush_inputs_to_zero;
-+{
+     bool default_nan_mode;
-+    ARMCPU *cpu = env_archcpu(env);
++    /*
 +     * The pattern to use for the default NaN. Here the high bit specifies
 +     * the default NaN's sign bit, and bits 6..0 specify the high bits of the
 +     * fractional part. The low bits of the fractional part are copies of bit 0.
 +     * The exponent of the default NaN is (as for any NaN) always all 1s.
 +     * Note that a value of 0 here is not a valid NaN. The target must set
 +     * this to the correct non-zero value, or we will assert when trying to
 +     * create a default NaN.
 +     */
 +    uint8_t default_nan_pattern;
      /*
       * The flags below are not used on all specializations and may
       * constant fold away (see snan_bit_is_one()/no_signalling_nans() in
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
 --- a/fpu/softfloat-specialize.c.inc
 +++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
  {
      bool sign = 0;
      uint64_t frac;
 +    uint8_t dnan_pattern = status->default_nan_pattern;
 +    if (dnan_pattern == 0) {
  #if defined(TARGET_SPARC) || defined(TARGET_M68K)
 -    /* !snan_bit_is_one, set all bits */
 -    frac = (1ULL << DECOMPOSED_BINARY_POINT) - 1;
 -#elif defined(TARGET_I386) || defined(TARGET_X86_64) \
 +        /* Sign bit clear, all frac bits set */
 +        dnan_pattern = 0b01111111;
 +#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
      || defined(TARGET_MICROBLAZE)
 -    /* !snan_bit_is_one, set sign and msb */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    sign = 1;
 +        /* Sign bit set, most significant frac bit set */
 +        dnan_pattern = 0b11000000;
  #elif defined(TARGET_HPPA)
 -    /* snan_bit_is_one, set msb-1.  */
 -    frac = 1ULL << (DECOMPOSED_BINARY_POINT - 2);
 +        /* Sign bit clear, msb-1 frac bit set */
 +        dnan_pattern = 0b00100000;
  #elif defined(TARGET_HEXAGON)
 -    sign = 1;
 -    frac = ~0ULL;
 +        /* Sign bit set, all frac bits set. */
 +        dnan_pattern = 0b11111111;
  #else
 -    /*
 -     * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 -     * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 -     * do not have floating-point.
 -     */
 -    if (snan_bit_is_one(status)) {
 -        /* set all bits other than msb */
 -        frac = (1ULL << (DECOMPOSED_BINARY_POINT - 1)) - 1;
 -    } else {
 -        /* set msb */
 -        frac = 1ULL << (DECOMPOSED_BINARY_POINT - 1);
 -    }
 +        /*
 +         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
 +         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
 +         * do not have floating-point.
 +         */
 +        if (snan_bit_is_one(status)) {
 +            /* sign bit clear, set all frac bits other than msb */
 +            dnan_pattern = 0b00111111;
 +        } else {
 +            /* sign bit clear, set frac msb */
 +            dnan_pattern = 0b01000000;
 +        }
  #endif
 +    }
 +    assert(dnan_pattern != 0);
 +
-+    vfp_set_fpsr_to_host(env, val);
++    sign = dnan_pattern >> 7;
 +
 +    if (arm_feature(env, ARM_FEATURE_NEON) ||
 +        cpu_isar_feature(aa32_mve, cpu)) {
 +        /*
 +         * The bit we set within vfp.qc[] is arbitrary; the array as a
 +         * whole being zero/non-zero is what counts.
 +         */
 +        env->vfp.qc[0] = val & FPCR_QC;
 +        env->vfp.qc[1] = 0;
 +        env->vfp.qc[2] = 0;
 +        env->vfp.qc[3] = 0;
 +    }
 +
 +    /*
-+     * The only FPSR bits we keep in vfp.xregs[FPSCR] are NZCV:
++     * Place default_nan_pattern [6:0] into bits [62:56],
-+     * the exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
++     * and replecate bit [0] down into [55:0]
 +     * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
 +     * and zero any of the other FPSR bits (but preserve the FPCR
 +     * bits).
 +     */
-+    val &= FPCR_NZCV_MASK;
++    frac = deposit64(0, DECOMPOSED_BINARY_POINT - 7, 7, dnan_pattern);
-+    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPSR_MASK;
++    frac = deposit64(frac, 0, DECOMPOSED_BINARY_POINT - 7, -(dnan_pattern & 1));
-+    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
-+}
+     *p = (FloatParts64) {
-+
+         .cls = float_class_qnan,
 +void vfp_set_fpcr(CPUARMState *env, uint32_t val)
  {
      ARMCPU *cpu = env_archcpu(env);
@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
          val &= ~FPCR_FZ16;
      }
 -    vfp_set_fpscr_to_host(env, val);
 +    vfp_set_fpcr_to_host(env, val);
      if (!arm_feature(env, ARM_FEATURE_M)) {
          /*
@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
                                       FPCR_LTPSIZE_LENGTH);
      }
 -    if (arm_feature(env, ARM_FEATURE_NEON) ||
 -        cpu_isar_feature(aa32_mve, cpu)) {
 -        /*
 -         * The bit we set within fpscr_q is arbitrary; the register as a
 -         * whole being zero/non-zero is what counts.
 -         */
 -        env->vfp.qc[0] = val & FPCR_QC;
 -        env->vfp.qc[1] = 0;
 -        env->vfp.qc[2] = 0;
 -        env->vfp.qc[3] = 0;
 -    }
 -
      /*
       * We don't implement trapped exception handling, so the
       * trap enable bits, IDE|IXE|UFE|OFE|DZE|IOE are all RAZ/WI (not RES0!)
       *
 -     * The exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
 -     * fp_status; QC, Len and Stride are stored separately earlier.
 -     * Clear out all of those and the RES0 bits: only NZCV, AHP, DN,
 -     * FZ, RMode and FZ16 are kept in vfp.xregs[FPSCR].
 +     * The FPCR bits we keep in vfp.xregs[FPSCR] are AHP, DN, FZ, RMode
 +     * and FZ16. Len, Stride and LTPSIZE we just handled. Store those bits
 +     * there, and zero any of the other FPCR bits and the RES0 and RAZ/WI
 +     * bits.
       */
 -    env->vfp.xregs[ARM_VFP_FPSCR] = val & 0xf7c80000;
 +    val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
 +    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_MASK;
 +    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
 +}
 +
 +void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
 +{
 +    vfp_set_fpcr(env, val & FPCR_MASK);
 +    vfp_set_fpsr(env, val & FPSR_MASK);
  }
  void vfp_set_fpscr(CPUARMState *env, uint32_t val)
 --
 .34.1

-New patch
+[PULL 41/72] tests/fp: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the tests/fp code.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
+---
+ tests/fp/fp-bench.c     | 1 +
+ tests/fp/fp-test-log2.c | 1 +
+ tests/fp/fp-test.c      | 1 +
+files changed, 3 insertions(+)
+diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-bench.c
++++ b/tests/fp/fp-bench.c
+@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
++    set_float_default_nan_pattern(0b01000000, &soft_status);
+     f = bench_funcs[operation][precision];
+     g_assert(f);
+diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test-log2.c
++++ b/tests/fp/fp-test-log2.c
+@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
+     int i;
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_rounding_mode(float_round_nearest_even, &qsf);
+     test.d = 0.0;
+diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
+index XXXXXXX..XXXXXXX 100644
+--- a/tests/fp/fp-test.c
++++ b/tests/fp/fp-test.c
+@@ -XXX,XX +XXX,XX @@ void run_test(void)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
++    set_float_default_nan_pattern(0b01000000, &qsf);
+     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
+     genCases_setLevel(test_level);
+--
+.34.1

-New patch
+[PULL 42/72] target/microblaze: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
+---
+ target/microblaze/cpu.c        | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 +--
+files changed, 3 insertions(+), 2 deletions(-)
+diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/microblaze/cpu.c
++++ b/target/microblaze/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
+      * this architecture.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN: sign bit set, most significant frac bit set */
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
+ #if defined(CONFIG_USER_ONLY)
+     /* start in user mode with interrupts enabled.  */
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
+-    || defined(TARGET_MICROBLAZE)
++#elif defined(TARGET_I386) || defined(TARGET_X86_64)
+         /* Sign bit set, most significant frac bit set */
+         dnan_pattern = 0b11000000;
+ #elif defined(TARGET_HPPA)
+--
+.34.1

-New patch
+[PULL 43/72] target/i386: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
+---
+ target/i386/tcg/fpu_helper.c   | 4 ++++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 4 insertions(+), 3 deletions(-)
+diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/i386/tcg/fpu_helper.c
++++ b/target/i386/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
++    /* Default NaN: sign bit set, most significant frac bit set */
++    set_float_default_nan_pattern(0b11000000, &env->fp_status);
++    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
++    set_float_default_nan_pattern(0b11000000, &env->sse_status);
+ }
+ static inline uint8_t save_exception_flags(CPUX86State *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
+-        /* Sign bit set, most significant frac bit set */
+-        dnan_pattern = 0b11000000;
+ #elif defined(TARGET_HPPA)
+         /* Sign bit clear, msb-1 frac bit set */
+         dnan_pattern = 0b00100000;
+--
+.34.1

-New patch
+[PULL 44/72] target/hppa: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly, and remove the ifdef from
+parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
+---
+ target/hppa/fpu_helper.c       | 2 ++
+ fpu/softfloat-specialize.c.inc | 3 ---
+files changed, 2 insertions(+), 3 deletions(-)
+diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/hppa/fpu_helper.c
++++ b/target/hppa/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
+     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN: sign bit clear, msb-1 frac bit set */
++    set_float_default_nan_pattern(0b00100000, &env->fp_status);
+ }
+ void cpu_hppa_loaded_fr0(CPUHPPAState *env)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+ #if defined(TARGET_SPARC) || defined(TARGET_M68K)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HPPA)
+-        /* Sign bit clear, msb-1 frac bit set */
+-        dnan_pattern = 0b00100000;
+ #elif defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+--
+.34.1

-New patch
+[PULL 45/72] target/alpha: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for the alpha target.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-40-peter.maydell@linaro.org
+---
+ target/alpha/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/alpha/cpu.c b/target/alpha/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/alpha/cpu.c
++++ b/target/alpha/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void alpha_cpu_initfn(Object *obj)
+      * operand in Fa. That is float_2nan_prop_ba.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ #if defined(CONFIG_USER_ONLY)
+     env->flags = ENV_FLAG_PS_USER | ENV_FLAG_FEN;
+     cpu_alpha_store_fpcr(env, (uint64_t)(FPCR_INVD | FPCR_DZED | FPCR_OVFD
+--
+.34.1

-[PULL 13/24] target/arm: Set arm_v7m_tcg_ops cpu_exec_halt to arm_cpu_exec_halt()
+[PULL 46/72] target/arm: Set default NaN pattern explicitly
-In commit a96edb687e76 we set the cpu_exec_halt field of the
+Set the default NaN pattern explicitly for the arm target.
-TCGCPUOps arm_tcg_ops to arm_cpu_exec_halt(), but we left the
+This includes setting it for the old linux-user nwfpe emulation.
-arm_v7m_tcg_ops struct unchanged.  That isn't wrong, because for
+For nwfpe, our default doesn't match the real kernel, but we
-M-profile FEAT_WFxT doesn't exist and the default handling for "no
+avoid making a behaviour change in this commit.
 cpu_exec_halt method" is correct, but it's perhaps a little
 confusing.  We would also like to make setting the cpu_exec_halt
 method mandatory.
 Initialize arm_v7m_tcg_ops cpu_exec_halt to the same function we use
 for A-profile.  (On M-profile we never set up the wfxt timer so there
 is no change in behaviour here.)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
 ---
- target/arm/internals.h   | 3 +++
+ linux-user/arm/nwfpe/fpa11.c | 5 +++++
- target/arm/cpu.c         | 2 +-
+ target/arm/cpu.c             | 2 ++
- target/arm/tcg/cpu-v7m.c | 1 +
+files changed, 7 insertions(+)
 files changed, 5 insertions(+), 1 deletion(-)
-diff --git a/target/arm/internals.h b/target/arm/internals.h
+diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/internals.h
+--- a/linux-user/arm/nwfpe/fpa11.c
-+++ b/target/arm/internals.h
++++ b/linux-user/arm/nwfpe/fpa11.c
-@@ -XXX,XX +XXX,XX @@ void arm_restore_state_to_opc(CPUState *cs,
+@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
+    * this late date.
- #ifdef CONFIG_TCG
+    */
- void arm_cpu_synchronize_from_tb(CPUState *cs, const TranslationBlock *tb);
+   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
-+
++  /*
-+/* Our implementation of TCGCPUOps::cpu_exec_halt */
++   * Use the same default NaN value as Arm VFP. This doesn't match
-+bool arm_cpu_exec_halt(CPUState *cs);
++   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
- #endif /* CONFIG_TCG */
++   */
++  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
- typedef enum ARMFPRounding {
+ }
  void SetRoundingMode(const unsigned int opcode)
 diff --git a/target/arm/cpu.c b/target/arm/cpu.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/cpu.c
 +++ b/target/arm/cpu.c
-@@ -XXX,XX +XXX,XX @@ static bool arm_cpu_virtio_is_big_endian(CPUState *cs)
+@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
   *    the pseudocode function the arguments are in the order c, a, b.
   *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
   *    and the input NaN if it is signalling
 + *  * Default NaN has sign bit clear, msb frac bit set
   */
  static void arm_set_default_fp_behaviours(float_status *s)
  {
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
      set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
      set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
      set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 +    set_float_default_nan_pattern(0b01000000, s);
  }
- #ifdef CONFIG_TCG
+ static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
 -static bool arm_cpu_exec_halt(CPUState *cs)
 +bool arm_cpu_exec_halt(CPUState *cs)
  {
      bool leave_halt = cpu_has_work(cs);
 diff --git a/target/arm/tcg/cpu-v7m.c b/target/arm/tcg/cpu-v7m.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/cpu-v7m.c
 +++ b/target/arm/tcg/cpu-v7m.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps arm_v7m_tcg_ops = {
  #else
      .tlb_fill = arm_cpu_tlb_fill,
      .cpu_exec_interrupt = arm_v7m_cpu_exec_interrupt,
 +    .cpu_exec_halt = arm_cpu_exec_halt,
      .do_interrupt = arm_v7m_cpu_do_interrupt,
      .do_transaction_failed = arm_cpu_do_transaction_failed,
      .do_unaligned_access = arm_cpu_do_unaligned_access,
 --
 .34.1

-New patch
+[PULL 47/72] target/loongarch: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for loongarch.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-42-peter.maydell@linaro.org
+---
+ target/loongarch/tcg/fpu_helper.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/loongarch/tcg/fpu_helper.c
++++ b/target/loongarch/tcg/fpu_helper.c
+@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
+      */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
++    /* Default NaN: sign bit clear, msb frac bit set */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
+ }
+ int ieee_ex_to_loongarch(int xcpt)
+--
+.34.1

-New patch
+[PULL 48/72] target/m68k: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for m68k.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
+---
+ target/m68k/cpu.c              | 2 ++
+ fpu/softfloat-specialize.c.inc | 2 +-
+files changed, 3 insertions(+), 1 deletion(-)
+diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/m68k/cpu.c
++++ b/target/m68k/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
+      * preceding paragraph for nonsignaling NaNs.
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
++    /* Default NaN: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     nan = floatx80_default_nan(&env->fp_status);
+     for (i = 0; i < 8; i++) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
++#if defined(TARGET_SPARC)
+         /* Sign bit clear, all frac bits set */
+         dnan_pattern = 0b01111111;
+ #elif defined(TARGET_HEXAGON)
+--
+.34.1

-New patch
+[PULL 49/72] target/mips: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for MIPS. Note that this
+is our only target which currently changes the default NaN
+at runtime (which it was previously doing indirectly when it
+changed the snan_bit_is_one setting).
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
+---
+ target/mips/fpu_helper.h | 7 +++++++
+ target/mips/msa.c        | 3 +++
+files changed, 10 insertions(+)
+diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/fpu_helper.h
++++ b/target/mips/fpu_helper.h
+@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
+     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
++    /*
++     * With nan2008, the default NaN value has the sign bit clear and the
++     * frac msb set; with the older mode, the sign bit is clear, and all
++     * frac bits except the msb are set.
++     */
++    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
++                                  &env->active_fpu.fp_status);
+ }
+diff --git a/target/mips/msa.c b/target/mips/msa.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/mips/msa.c
++++ b/target/mips/msa.c
+@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
+     /* Inf * 0 + NaN returns the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                               &env->active_tc.msa_fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000,
++                                  &env->active_tc.msa_fp_status);
+ }
+--
+.34.1

-New patch
+[PULL 50/72] target/openrisc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for openrisc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-45-peter.maydell@linaro.org
+---
+ target/openrisc/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/openrisc/cpu.c b/target/openrisc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/openrisc/cpu.c
++++ b/target/openrisc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void openrisc_cpu_reset_hold(Object *obj, ResetType type)
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &cpu->env.fp_status);
++    /* Default NaN: sign bit clear, frac msb set */
++    set_float_default_nan_pattern(0b01000000, &cpu->env.fp_status);
+ #ifndef CONFIG_USER_ONLY
+     cpu->env.picmr = 0x00000000;
+--
+.34.1

-New patch
+[PULL 51/72] target/ppc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for ppc.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-46-peter.maydell@linaro.org
+---
+ target/ppc/cpu_init.c | 4 ++++
+file changed, 4 insertions(+)
+diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/ppc/cpu_init.c
++++ b/target/ppc/cpu_init.c
+@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
++    /* Default NaN: sign bit clear, set frac msb */
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
++    set_float_default_nan_pattern(0b01000000, &env->vec_status);
++
+     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
+         ppc_spr_t *spr = &env->spr_cb[i];
+--
+.34.1

-New patch
+[PULL 52/72] target/sh4: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for sh4. Note that sh4
+is one of the only three targets (the others being HPPA and
+sometimes MIPS) that has snan_bit_is_one set.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-47-peter.maydell@linaro.org
+---
+ target/sh4/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/sh4/cpu.c b/target/sh4/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sh4/cpu.c
++++ b/target/sh4/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void superh_cpu_reset_hold(Object *obj, ResetType type)
+     set_flush_to_zero(1, &env->fp_status);
+ #endif
+     set_default_nan_mode(1, &env->fp_status);
++    /* sign bit clear, set all frac bits other than msb */
++    set_float_default_nan_pattern(0b00111111, &env->fp_status);
+ }
+ static void superh_cpu_disas_set_info(CPUState *cpu, disassemble_info *info)
+--
+.34.1

-[PULL 08/24] target/arm: Rename FPSR_MASK and FPCR_MASK and define them symbolically
+[PULL 53/72] target/rx: Set default NaN pattern explicitly
-Now that we store FPSR and FPCR separately, the FPSR_MASK and
+Set the default NaN pattern explicitly for rx.
 FPCR_MASK macros are slightly confusingly named and the comment
 describing them is out of date.  Rename them to FPSCR_FPSR_MASK and
 FPSCR_FPCR_MASK, document that they are the mask of which FPSCR bits
 are architecturally mapped to which AArch64 register, and define them
 symbolically rather than as hex values.  (This latter requires
 defining some extra macros for bits which we haven't previously
 defined.)
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-9-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-48-peter.maydell@linaro.org
 ---
- target/arm/cpu.h        | 41 ++++++++++++++++++++++++++++++++++-------
+ target/rx/cpu.c | 2 ++
- target/arm/machine.c    |  3 ++-
+file changed, 2 insertions(+)
  target/arm/vfp_helper.c |  7 ++++---
 files changed, 40 insertions(+), 11 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/target/rx/cpu.c b/target/rx/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/target/rx/cpu.c
-+++ b/target/arm/cpu.h
++++ b/target/rx/cpu.c
-@@ -XXX,XX +XXX,XX @@ static inline void xpsr_write(CPUARMState *env, uint32_t val, uint32_t mask)
+@@ -XXX,XX +XXX,XX @@ static void rx_cpu_reset_hold(Object *obj, ResetType type)
- uint32_t vfp_get_fpscr(CPUARMState *env);
+      * then prefer dest over source", which is float_2nan_prop_s_ab.
- void vfp_set_fpscr(CPUARMState *env, uint32_t val);
+      */
+     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
--/* FPCR, Floating Point Control Register
++    /* Default NaN value: sign bit clear, set frac msb */
-- * FPSR, Floating Poiht Status Register
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
 +/*
 + * FPCR, Floating Point Control Register
 + * FPSR, Floating Point Status Register
   *
 - * For A64 the FPSCR is split into two logically distinct registers,
 - * FPCR and FPSR. However since they still use non-overlapping bits
 - * we store the underlying state in fpscr and just mask on read/write.
 + * For A64 floating point control and status bits are stored in
 + * two logically distinct registers, FPCR and FPSR. We store these
 + * in QEMU in vfp.fpcr and vfp.fpsr.
 + * For A32 there was only one register, FPSCR. The bits are arranged
 + * such that FPSCR bits map to FPCR or FPSR bits in the same bit positions,
 + * so we can use appropriate masking to handle FPSCR reads and writes.
 + * Note that the FPCR has some bits which are not visible in the
 + * AArch32 view (for FEAT_AFP). Writing the FPSCR leaves these unchanged.
   */
 -#define FPSR_MASK 0xf800009f
 -#define FPCR_MASK 0x07ff9f00
  /* FPCR bits */
  #define FPCR_IOE    (1 << 8)    /* Invalid Operation exception trap enable */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
  #define FPCR_UFE    (1 << 11)   /* Underflow exception trap enable */
  #define FPCR_IXE    (1 << 12)   /* Inexact exception trap enable */
  #define FPCR_IDE    (1 << 15)   /* Input Denormal exception trap enable */
 +#define FPCR_LEN_MASK (7 << 16) /* LEN, A-profile only */
  #define FPCR_FZ16   (1 << 19)   /* ARMv8.2+, FP16 flush-to-zero */
 +#define FPCR_STRIDE_MASK (3 << 20) /* Stride */
  #define FPCR_RMODE_MASK (3 << 22) /* Rounding mode */
  #define FPCR_FZ     (1 << 24)   /* Flush-to-zero enable bit */
  #define FPCR_DN     (1 << 25)   /* Default NaN enable bit */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
  #define FPCR_LTPSIZE_MASK (7 << FPCR_LTPSIZE_SHIFT)
  #define FPCR_LTPSIZE_LENGTH 3
 +/* Cumulative exception trap enable bits */
 +#define FPCR_EEXC_MASK (FPCR_IOE | FPCR_DZE | FPCR_OFE | FPCR_UFE | FPCR_IXE | FPCR_IDE)
 +
  /* FPSR bits */
 +#define FPSR_IOC    (1 << 0)    /* Invalid Operation cumulative exception */
 +#define FPSR_DZC    (1 << 1)    /* Divide by Zero cumulative exception */
 +#define FPSR_OFC    (1 << 2)    /* Overflow cumulative exception */
 +#define FPSR_UFC    (1 << 3)    /* Underflow cumulative exception */
 +#define FPSR_IXC    (1 << 4)    /* Inexact cumulative exception */
 +#define FPSR_IDC    (1 << 7)    /* Input Denormal cumulative exception */
  #define FPSR_QC     (1 << 27)   /* Cumulative saturation bit */
  #define FPSR_V      (1 << 28)   /* FP overflow flag */
  #define FPSR_C      (1 << 29)   /* FP carry flag */
  #define FPSR_Z      (1 << 30)   /* FP zero flag */
  #define FPSR_N      (1 << 31)   /* FP negative flag */
 +/* Cumulative exception status bits */
 +#define FPSR_CEXC_MASK (FPSR_IOC | FPSR_DZC | FPSR_OFC | FPSR_UFC | FPSR_IXC | FPSR_IDC)
 +
  #define FPSR_NZCV_MASK (FPSR_N | FPSR_Z | FPSR_C | FPSR_V)
  #define FPSR_NZCVQC_MASK (FPSR_NZCV_MASK | FPSR_QC)
 +/* A32 FPSCR bits which architecturally map to FPSR bits */
 +#define FPSCR_FPSR_MASK (FPSR_NZCVQC_MASK | FPSR_CEXC_MASK)
 +/* A32 FPSCR bits which architecturally map to FPCR bits */
 +#define FPSCR_FPCR_MASK (FPCR_EEXC_MASK | FPCR_LEN_MASK | FPCR_FZ16 | \
 +                         FPCR_STRIDE_MASK | FPCR_RMODE_MASK | \
 +                         FPCR_FZ | FPCR_DN | FPCR_AHP)
 +/* These masks don't overlap: each bit lives in only one place */
 +QEMU_BUILD_BUG_ON(FPSCR_FPSR_MASK & FPSCR_FPCR_MASK);
 +
  /**
   * vfp_get_fpsr: read the AArch64 FPSR
   * @env: CPU context
 diff --git a/target/arm/machine.c b/target/arm/machine.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/machine.c
 +++ b/target/arm/machine.c
@@ -XXX,XX +XXX,XX @@ static bool vfp_fpcr_fpsr_needed(void *opaque)
      ARMCPU *cpu = opaque;
      CPUARMState *env = &cpu->env;
 -    return (vfp_get_fpcr(env) & ~FPCR_MASK) || (vfp_get_fpsr(env) & ~FPSR_MASK);
 +    return (vfp_get_fpcr(env) & ~FPSCR_FPCR_MASK) ||
 +        (vfp_get_fpsr(env) & ~FPSCR_FPSR_MASK);
  }
- static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
+ static ObjectClass *rx_cpu_class_by_name(const char *cpu_model)
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/vfp_helper.c
 +++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env)
  uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
  {
 -    return (vfp_get_fpcr(env) & FPCR_MASK) | (vfp_get_fpsr(env) & FPSR_MASK);
 +    return (vfp_get_fpcr(env) & FPSCR_FPCR_MASK) |
 +        (vfp_get_fpsr(env) & FPSCR_FPSR_MASK);
  }
  uint32_t vfp_get_fpscr(CPUARMState *env)
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
  void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
  {
 -    vfp_set_fpcr(env, val & FPCR_MASK);
 -    vfp_set_fpsr(env, val & FPSR_MASK);
 +    vfp_set_fpcr(env, val & FPSCR_FPCR_MASK);
 +    vfp_set_fpsr(env, val & FPSCR_FPSR_MASK);
  }
  void vfp_set_fpscr(CPUARMState *env, uint32_t val)
 --
 .34.1

-New patch
+[PULL 54/72] target/s390x: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for s390x.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-49-peter.maydell@linaro.org
+---
+ target/s390x/cpu.c | 2 ++
+file changed, 2 insertions(+)
+diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/s390x/cpu.c
++++ b/target/s390x/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
+         set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
+         set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                   &env->fpu_status);
++        /* Default NaN value: sign bit clear, frac msb set */
++        set_float_default_nan_pattern(0b01000000, &env->fpu_status);
+        /* fall through */
+     case RESET_TYPE_S390_CPU_NORMAL:
+         env->psw.mask &= ~PSW_MASK_RI;
+--
+.34.1

-New patch
+[PULL 55/72] target/sparc: Set default NaN pattern explicitly
+Set the default NaN pattern explicitly for SPARC, and remove
+the ifdef from parts64_default_nan.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
+---
+ target/sparc/cpu.c             | 2 ++
+ fpu/softfloat-specialize.c.inc | 5 +----
+files changed, 3 insertions(+), 4 deletions(-)
+diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/sparc/cpu.c
++++ b/target/sparc/cpu.c
+@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
+     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
+     /* For inf * 0 + NaN, return the input NaN */
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
++    /* Default NaN value: sign bit clear, all frac bits set */
++    set_float_default_nan_pattern(0b01111111, &env->fp_status);
+     cpu_exec_realizefn(cs, &local_err);
+     if (local_err != NULL) {
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint8_t dnan_pattern = status->default_nan_pattern;
+     if (dnan_pattern == 0) {
+-#if defined(TARGET_SPARC)
+-        /* Sign bit clear, all frac bits set */
+-        dnan_pattern = 0b01111111;
+-#elif defined(TARGET_HEXAGON)
++#if defined(TARGET_HEXAGON)
+         /* Sign bit set, all frac bits set. */
+         dnan_pattern = 0b11111111;
+ #else
+--
+.34.1

-[PULL 05/24] target/arm: Implement store_cpu_field_low32() macro
+[PULL 56/72] target/xtensa: Set default NaN pattern explicitly
-We already have a load_cpu_field_low32() to load the low half of a
+Set the default NaN pattern explicitly for xtensa.
 -bit CPU struct field to a TCGv_i32; however we haven't yet needed
 the store equivalent.  We'll want that in the next patch, so
 implement it.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-6-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-51-peter.maydell@linaro.org
 ---
- target/arm/tcg/translate-a32.h | 7 +++++++
+ target/xtensa/cpu.c | 2 ++
-file changed, 7 insertions(+)
+file changed, 2 insertions(+)
-diff --git a/target/arm/tcg/translate-a32.h b/target/arm/tcg/translate-a32.h
+diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-a32.h
+--- a/target/xtensa/cpu.c
-+++ b/target/arm/tcg/translate-a32.h
++++ b/target/xtensa/cpu.c
-@@ -XXX,XX +XXX,XX @@ void store_cpu_offset(TCGv_i32 var, int offset, int size);
+@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
-                          sizeof_field(CPUARMState, name));              \
+     /* For inf * 0 + NaN, return the input NaN */
-     })
+     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+     set_no_signaling_nans(!dfpu, &env->fp_status);
-+/* Store to the low half of a 64-bit field from a TCGv_i32 */
++    /* Default NaN value: sign bit clear, set frac msb */
-+#define store_cpu_field_low32(val, name)                                \
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
-+    ({                                                                  \
+     xtensa_use_first_nan(env, !dfpu);
-+        QEMU_BUILD_BUG_ON(sizeof_field(CPUARMState, name) != 8);        \
+ }
 +        store_cpu_offset(val, offsetoflow32(CPUARMState, name), 4);     \
 +    })
 +
  #define store_cpu_field_constant(val, name) \
      store_cpu_field(tcg_constant_i32(val), name)
 --
 .34.1

-[PULL 02/24] target/arm: Make vfp_get_fpscr() call vfp_get_{fpcr, fpsr}
+[PULL 57/72] target/hexagon: Set default NaN pattern explicitly
-In AArch32, the floating point control and status bits are all in a
+Set the default NaN pattern explicitly for hexagon.
-single register, FPSCR.  In AArch64, these were split into separate
+Remove the ifdef from parts64_default_nan(); the only
-FPCR and FPSR registers, but the bit layouts remained the same, with
+remaining unconverted targets all use the default case.
 no overlaps, so that you could construct an FPSCR value by ORing FPCR
 and FPSR, or equivalently could produce FPSR and FPCR by masking an
 FPSCR value.  For QEMU's implementation, we opted to use masking to
 produce FPSR and FPCR, because we started with an AArch32
 implementation of FPSCR.
 The addition of the (AArch64-only) FEAT_AFP adds new bits to the FPCR
 which overlap with some bits in the FPSR.  This means we'll no longer
 be able to consider the FPSCR-encoded value as the primary one, but
 instead need to treat FPSR/FPCR as the primary encoding and construct
 the FPSCR from those.  (This remains possible because the FEAT_AFP
 bits in FPCR don't appear in the FPSCR.)
 As the first step in this refactoring, make vfp_get_fpscr() call
 vfp_get_fpcr() and vfp_get_fpsr(), instead of the other way around.
 Note that vfp_get_fpcsr_from_host() returns only bits in the FPSR
 (for the cumulative fp exception bits), so we can simply rename
 it without needing to add a new function for getting FPCR bits.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-3-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
 ---
- target/arm/cpu.h        | 24 +++++++++++++++---------
+ target/hexagon/cpu.c           | 2 ++
- target/arm/vfp_helper.c | 34 ++++++++++++++++++++++------------
+ fpu/softfloat-specialize.c.inc | 5 -----
-files changed, 37 insertions(+), 21 deletions(-)
+files changed, 2 insertions(+), 5 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/target/hexagon/cpu.c
-+++ b/target/arm/cpu.h
++++ b/target/hexagon/cpu.c
-@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
+@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
- #define FPCR_NZCV_MASK (FPCR_N | FPCR_Z | FPCR_C | FPCR_V)
- #define FPCR_NZCVQC_MASK (FPCR_NZCV_MASK | FPCR_QC)
+     set_default_nan_mode(1, &env->fp_status);
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
--static inline uint32_t vfp_get_fpsr(CPUARMState *env)
++    /* Default NaN value: sign bit set, all frac bits set */
--{
++    set_float_default_nan_pattern(0b11111111, &env->fp_status);
 -    return vfp_get_fpscr(env) & FPSR_MASK;
 -}
 +/**
 + * vfp_get_fpsr: read the AArch64 FPSR
 + * @env: CPU context
 + *
 + * Return the current AArch64 FPSR value
 + */
 +uint32_t vfp_get_fpsr(CPUARMState *env);
 +
 +/**
 + * vfp_get_fpcr: read the AArch64 FPCR
 + * @env: CPU context
 + *
 + * Return the current AArch64 FPCR value
 + */
 +uint32_t vfp_get_fpcr(CPUARMState *env);
  static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
  {
@@ -XXX,XX +XXX,XX @@ static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
      vfp_set_fpscr(env, new_fpscr);
  }
--static inline uint32_t vfp_get_fpcr(CPUARMState *env)
+ static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
--{
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 -    return vfp_get_fpscr(env) & FPCR_MASK;
 -}
 -
  static inline void vfp_set_fpcr(CPUARMState *env, uint32_t val)
  {
      uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPCR_MASK) | (val & FPCR_MASK);
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/vfp_helper.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/vfp_helper.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ static inline int vfp_exceptbits_to_host(int target_bits)
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
-     return host_bits;
+     uint8_t dnan_pattern = status->default_nan_pattern;
- }
+     if (dnan_pattern == 0) {
--static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
+-#if defined(TARGET_HEXAGON)
-+static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
+-        /* Sign bit set, all frac bits set. */
- {
+-        dnan_pattern = 0b11111111;
-     uint32_t i;
+-#else
+         /*
-@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
+          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
+          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
- #else
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+             /* sign bit clear, set frac msb */
--static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
+             dnan_pattern = 0b01000000;
-+static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
+         }
- {
+-#endif
-     return 0;
+     }
- }
+     assert(dnan_pattern != 0);
-@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
  #endif
 -uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
 +uint32_t vfp_get_fpcr(CPUARMState *env)
  {
 -    uint32_t i, fpscr;
 -
 -    fpscr = env->vfp.xregs[ARM_VFP_FPSCR]
 -            | (env->vfp.vec_len << 16)
 -            | (env->vfp.vec_stride << 20);
 +    uint32_t fpcr = (env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_MASK)
 +        | (env->vfp.vec_len << 16)
 +        | (env->vfp.vec_stride << 20);
      /*
       * M-profile LTPSIZE is the same bits [18:16] as A-profile Len; whichever
       * of the two is not applicable to this CPU will always be zero.
       */
 -    fpscr |= env->v7m.ltpsize << 16;
 +    fpcr |= env->v7m.ltpsize << 16;
 -    fpscr |= vfp_get_fpscr_from_host(env);
 +    return fpcr;
 +}
 +
 +uint32_t vfp_get_fpsr(CPUARMState *env)
 +{
 +    uint32_t fpsr = env->vfp.xregs[ARM_VFP_FPSCR] & FPSR_MASK;
 +    uint32_t i;
 +
 +    fpsr |= vfp_get_fpsr_from_host(env);
      i = env->vfp.qc[0] | env->vfp.qc[1] | env->vfp.qc[2] | env->vfp.qc[3];
 -    fpscr |= i ? FPCR_QC : 0;
 +    fpsr |= i ? FPCR_QC : 0;
 +    return fpsr;
 +}
 -    return fpscr;
 +uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
 +{
 +    return (vfp_get_fpcr(env) & FPCR_MASK) | (vfp_get_fpsr(env) & FPSR_MASK);
  }
  uint32_t vfp_get_fpscr(CPUARMState *env)
 --
 .34.1

-[PULL 14/24] target: Set TCGCPUOps::cpu_exec_halt to target's has_work implementation
+[PULL 58/72] target/riscv: Set default NaN pattern explicitly
-Currently the TCGCPUOps::cpu_exec_halt method is optional, and if it
+Set the default NaN pattern explicitly for riscv.
 is not set then the default is to call the CPUClass::has_work
 method (which has an identical function signature).
 We would like to make the cpu_exec_halt method mandatory so we can
 remove the runtime check and fallback handling.  In preparation for
 that, make all the targets which don't need special handling in their
 cpu_exec_halt set it to their cpu_has_work implementation instead of
 leaving it unset.  (This is every target except for arm and i386.)
 In the riscv case this requires us to make the function not
 be local to the source file it's defined in.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20241202131347.498124-53-peter.maydell@linaro.org
 ---
- target/riscv/internals.h   | 3 +++
+ target/riscv/cpu.c | 2 ++
- target/alpha/cpu.c         | 1 +
+file changed, 2 insertions(+)
  target/avr/cpu.c           | 1 +
  target/cris/cpu.c          | 2 ++
  target/hppa/cpu.c          | 1 +
  target/loongarch/cpu.c     | 1 +
  target/m68k/cpu.c          | 1 +
  target/microblaze/cpu.c    | 1 +
  target/mips/cpu.c          | 1 +
  target/openrisc/cpu.c      | 1 +
  target/ppc/cpu_init.c      | 2 ++
  target/riscv/cpu.c         | 2 +-
  target/riscv/tcg/tcg-cpu.c | 2 ++
  target/rx/cpu.c            | 1 +
  target/s390x/cpu.c         | 1 +
  target/sh4/cpu.c           | 1 +
  target/sparc/cpu.c         | 1 +
  target/tricore/cpu.c       | 1 +
  target/xtensa/cpu.c        | 1 +
 files changed, 24 insertions(+), 1 deletion(-)
-diff --git a/target/riscv/internals.h b/target/riscv/internals.h
-index XXXXXXX..XXXXXXX 100644
---- a/target/riscv/internals.h
-+++ b/target/riscv/internals.h
-@@ -XXX,XX +XXX,XX @@ static inline float16 check_nanbox_h(CPURISCVState *env, uint64_t f)
-     }
- }
-+/* Our implementation of CPUClass::has_work */
-+bool riscv_cpu_has_work(CPUState *cs);
-+
- #endif
-diff --git a/target/alpha/cpu.c b/target/alpha/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/alpha/cpu.c
-+++ b/target/alpha/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps alpha_tcg_ops = {
- #else
-     .tlb_fill = alpha_cpu_tlb_fill,
-     .cpu_exec_interrupt = alpha_cpu_exec_interrupt,
-+    .cpu_exec_halt = alpha_cpu_has_work,
-     .do_interrupt = alpha_cpu_do_interrupt,
-     .do_transaction_failed = alpha_cpu_do_transaction_failed,
-     .do_unaligned_access = alpha_cpu_do_unaligned_access,
-diff --git a/target/avr/cpu.c b/target/avr/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/avr/cpu.c
-+++ b/target/avr/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps avr_tcg_ops = {
-     .synchronize_from_tb = avr_cpu_synchronize_from_tb,
-     .restore_state_to_opc = avr_restore_state_to_opc,
-     .cpu_exec_interrupt = avr_cpu_exec_interrupt,
-+    .cpu_exec_halt = avr_cpu_has_work,
-     .tlb_fill = avr_cpu_tlb_fill,
-     .do_interrupt = avr_cpu_do_interrupt,
- };
-diff --git a/target/cris/cpu.c b/target/cris/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/cris/cpu.c
-+++ b/target/cris/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps crisv10_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = cris_cpu_tlb_fill,
-     .cpu_exec_interrupt = cris_cpu_exec_interrupt,
-+    .cpu_exec_halt = cris_cpu_has_work,
-     .do_interrupt = crisv10_cpu_do_interrupt,
- #endif /* !CONFIG_USER_ONLY */
- };
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps crisv32_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = cris_cpu_tlb_fill,
-     .cpu_exec_interrupt = cris_cpu_exec_interrupt,
-+    .cpu_exec_halt = cris_cpu_has_work,
-     .do_interrupt = cris_cpu_do_interrupt,
- #endif /* !CONFIG_USER_ONLY */
- };
-diff --git a/target/hppa/cpu.c b/target/hppa/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/hppa/cpu.c
-+++ b/target/hppa/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps hppa_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = hppa_cpu_tlb_fill,
-     .cpu_exec_interrupt = hppa_cpu_exec_interrupt,
-+    .cpu_exec_halt = hppa_cpu_has_work,
-     .do_interrupt = hppa_cpu_do_interrupt,
-     .do_unaligned_access = hppa_cpu_do_unaligned_access,
-     .do_transaction_failed = hppa_cpu_do_transaction_failed,
-diff --git a/target/loongarch/cpu.c b/target/loongarch/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/loongarch/cpu.c
-+++ b/target/loongarch/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps loongarch_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = loongarch_cpu_tlb_fill,
-     .cpu_exec_interrupt = loongarch_cpu_exec_interrupt,
-+    .cpu_exec_halt = loongarch_cpu_has_work,
-     .do_interrupt = loongarch_cpu_do_interrupt,
-     .do_transaction_failed = loongarch_cpu_do_transaction_failed,
- #endif
-diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/m68k/cpu.c
-+++ b/target/m68k/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps m68k_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = m68k_cpu_tlb_fill,
-     .cpu_exec_interrupt = m68k_cpu_exec_interrupt,
-+    .cpu_exec_halt = m68k_cpu_has_work,
-     .do_interrupt = m68k_cpu_do_interrupt,
-     .do_transaction_failed = m68k_cpu_transaction_failed,
- #endif /* !CONFIG_USER_ONLY */
-diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/microblaze/cpu.c
-+++ b/target/microblaze/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps mb_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = mb_cpu_tlb_fill,
-     .cpu_exec_interrupt = mb_cpu_exec_interrupt,
-+    .cpu_exec_halt = mb_cpu_has_work,
-     .do_interrupt = mb_cpu_do_interrupt,
-     .do_transaction_failed = mb_cpu_transaction_failed,
-     .do_unaligned_access = mb_cpu_do_unaligned_access,
-diff --git a/target/mips/cpu.c b/target/mips/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/mips/cpu.c
-+++ b/target/mips/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps mips_tcg_ops = {
- #if !defined(CONFIG_USER_ONLY)
-     .tlb_fill = mips_cpu_tlb_fill,
-     .cpu_exec_interrupt = mips_cpu_exec_interrupt,
-+    .cpu_exec_halt = mips_cpu_has_work,
-     .do_interrupt = mips_cpu_do_interrupt,
-     .do_transaction_failed = mips_cpu_do_transaction_failed,
-     .do_unaligned_access = mips_cpu_do_unaligned_access,
-diff --git a/target/openrisc/cpu.c b/target/openrisc/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/openrisc/cpu.c
-+++ b/target/openrisc/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps openrisc_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = openrisc_cpu_tlb_fill,
-     .cpu_exec_interrupt = openrisc_cpu_exec_interrupt,
-+    .cpu_exec_halt = openrisc_cpu_has_work,
-     .do_interrupt = openrisc_cpu_do_interrupt,
- #endif /* !CONFIG_USER_ONLY */
- };
-diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/ppc/cpu_init.c
-+++ b/target/ppc/cpu_init.c
-@@ -XXX,XX +XXX,XX @@
-+
- /*
-  *  PowerPC CPU initialization for qemu.
-  *
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps ppc_tcg_ops = {
- #else
-   .tlb_fill = ppc_cpu_tlb_fill,
-   .cpu_exec_interrupt = ppc_cpu_exec_interrupt,
-+  .cpu_exec_halt = ppc_cpu_has_work,
-   .do_interrupt = ppc_cpu_do_interrupt,
-   .cpu_exec_enter = ppc_cpu_exec_enter,
-   .cpu_exec_exit = ppc_cpu_exec_exit,
 diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/riscv/cpu.c
 +++ b/target/riscv/cpu.c
-@@ -XXX,XX +XXX,XX @@ static vaddr riscv_cpu_get_pc(CPUState *cs)
+@@ -XXX,XX +XXX,XX @@ static void riscv_cpu_reset_hold(Object *obj, ResetType type)
-     return env->pc;
+     cs->exception_index = RISCV_EXCP_NONE;
- }
+     env->load_res = -1;
+     set_default_nan_mode(1, &env->fp_status);
--static bool riscv_cpu_has_work(CPUState *cs)
++    /* Default NaN value: sign bit clear, frac msb set */
-+bool riscv_cpu_has_work(CPUState *cs)
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
- {
+     env->vill = true;
  #ifndef CONFIG_USER_ONLY
-     RISCVCPU *cpu = RISCV_CPU(cs);
-diff --git a/target/riscv/tcg/tcg-cpu.c b/target/riscv/tcg/tcg-cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/riscv/tcg/tcg-cpu.c
-+++ b/target/riscv/tcg/tcg-cpu.c
-@@ -XXX,XX +XXX,XX @@
- #include "exec/exec-all.h"
- #include "tcg-cpu.h"
- #include "cpu.h"
-+#include "internals.h"
- #include "pmu.h"
- #include "time_helper.h"
- #include "qapi/error.h"
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps riscv_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = riscv_cpu_tlb_fill,
-     .cpu_exec_interrupt = riscv_cpu_exec_interrupt,
-+    .cpu_exec_halt = riscv_cpu_has_work,
-     .do_interrupt = riscv_cpu_do_interrupt,
-     .do_transaction_failed = riscv_cpu_do_transaction_failed,
-     .do_unaligned_access = riscv_cpu_do_unaligned_access,
-diff --git a/target/rx/cpu.c b/target/rx/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/rx/cpu.c
-+++ b/target/rx/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps rx_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .cpu_exec_interrupt = rx_cpu_exec_interrupt,
-+    .cpu_exec_halt = rx_cpu_has_work,
-     .do_interrupt = rx_cpu_do_interrupt,
- #endif /* !CONFIG_USER_ONLY */
- };
-diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/s390x/cpu.c
-+++ b/target/s390x/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps s390_tcg_ops = {
- #else
-     .tlb_fill = s390_cpu_tlb_fill,
-     .cpu_exec_interrupt = s390_cpu_exec_interrupt,
-+    .cpu_exec_halt = s390_cpu_has_work,
-     .do_interrupt = s390_cpu_do_interrupt,
-     .debug_excp_handler = s390x_cpu_debug_excp_handler,
-     .do_unaligned_access = s390x_cpu_do_unaligned_access,
-diff --git a/target/sh4/cpu.c b/target/sh4/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/sh4/cpu.c
-+++ b/target/sh4/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps superh_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = superh_cpu_tlb_fill,
-     .cpu_exec_interrupt = superh_cpu_exec_interrupt,
-+    .cpu_exec_halt = superh_cpu_has_work,
-     .do_interrupt = superh_cpu_do_interrupt,
-     .do_unaligned_access = superh_cpu_do_unaligned_access,
-     .io_recompile_replay_branch = superh_io_recompile_replay_branch,
-diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/sparc/cpu.c
-+++ b/target/sparc/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps sparc_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = sparc_cpu_tlb_fill,
-     .cpu_exec_interrupt = sparc_cpu_exec_interrupt,
-+    .cpu_exec_halt = sparc_cpu_has_work,
-     .do_interrupt = sparc_cpu_do_interrupt,
-     .do_transaction_failed = sparc_cpu_do_transaction_failed,
-     .do_unaligned_access = sparc_cpu_do_unaligned_access,
-diff --git a/target/tricore/cpu.c b/target/tricore/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/tricore/cpu.c
-+++ b/target/tricore/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps tricore_tcg_ops = {
-     .synchronize_from_tb = tricore_cpu_synchronize_from_tb,
-     .restore_state_to_opc = tricore_restore_state_to_opc,
-     .tlb_fill = tricore_cpu_tlb_fill,
-+    .cpu_exec_halt = tricore_cpu_has_work,
- };
- static void tricore_cpu_class_init(ObjectClass *c, void *data)
-diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
-index XXXXXXX..XXXXXXX 100644
---- a/target/xtensa/cpu.c
-+++ b/target/xtensa/cpu.c
-@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps xtensa_tcg_ops = {
- #ifndef CONFIG_USER_ONLY
-     .tlb_fill = xtensa_cpu_tlb_fill,
-     .cpu_exec_interrupt = xtensa_cpu_exec_interrupt,
-+    .cpu_exec_halt = xtensa_cpu_has_work,
-     .do_interrupt = xtensa_cpu_do_interrupt,
-     .do_transaction_failed = xtensa_cpu_do_transaction_failed,
-     .do_unaligned_access = xtensa_cpu_do_unaligned_access,
 --
 .34.1

-[PULL 06/24] target/arm: Store FPSR and FPCR in separate CPU state fields
+[PULL 59/72] target/tricore: Set default NaN pattern explicitly
-Now that we have refactored the set/get functions so that the FPSCR
+Set the default NaN pattern explicitly for tricore.
 format is no longer the authoritative one, we can keep FPSR and FPCR
 in separate CPU state fields.
 As well as the get and set functions, we also have a scattering of
 places in the code which directly access vfp.xregs[ARM_VFP_FPSCR] to
 extract single fields which are stored there.  These all change to
 directly access either vfp.fpsr or vfp.fpcr, depending on the
 location of the field.  (Most commonly, this is the NZCV flags.)
 We make the field in the CPU state struct 64 bits, because
 architecturally FPSR and FPCR are 64 bits.  However we leave the
 types of the arguments and return values of the get/set functions as
 bits, since we don't need to make that change with the current
 architecture and various callsites would be unable to handle
 set bits in the high half (for instance the gdbstub protocol
 assumes they're only 32 bit registers).
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20240628142347.1283015-7-peter.maydell@linaro.org
+Message-id: 20241202131347.498124-54-peter.maydell@linaro.org
 ---
- target/arm/cpu.h                  |  7 +++++++
+ target/tricore/helper.c | 2 ++
- target/arm/tcg/translate.h        |  3 +--
+file changed, 2 insertions(+)
  target/arm/tcg/mve_helper.c       | 12 ++++++------
  target/arm/tcg/translate-m-nocp.c |  6 +++---
  target/arm/tcg/translate-vfp.c    |  2 +-
  target/arm/vfp_helper.c           | 25 ++++++++++---------------
 files changed, 28 insertions(+), 27 deletions(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/target/tricore/helper.c b/target/tricore/helper.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/target/tricore/helper.c
-+++ b/target/arm/cpu.h
++++ b/target/tricore/helper.c
-@@ -XXX,XX +XXX,XX @@ typedef struct CPUArchState {
+@@ -XXX,XX +XXX,XX @@ void fpu_set_state(CPUTriCoreState *env)
-         int vec_len;
+     set_flush_to_zero(1, &env->fp_status);
-         int vec_stride;
+     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
+     set_default_nan_mode(1, &env->fp_status);
-+        /*
++    /* Default NaN pattern: sign bit clear, frac msb set */
-+         * Floating point status and control registers. Some bits are
++    set_float_default_nan_pattern(0b01000000, &env->fp_status);
 +         * stored separately in other fields or in the float_status below.
 +         */
 +        uint64_t fpsr;
 +        uint64_t fpcr;
 +
          uint32_t xregs[16];
          /* Scratch space for aa32 neon expansion.  */
 diff --git a/target/arm/tcg/translate.h b/target/arm/tcg/translate.h
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/translate.h
 +++ b/target/arm/tcg/translate.h
@@ -XXX,XX +XXX,XX @@ static inline TCGv_i32 get_ahp_flag(void)
  {
      TCGv_i32 ret = tcg_temp_new_i32();
 -    tcg_gen_ld_i32(ret, tcg_env,
 -                   offsetof(CPUARMState, vfp.xregs[ARM_VFP_FPSCR]));
 +    tcg_gen_ld_i32(ret, tcg_env, offsetoflow32(CPUARMState, vfp.fpcr));
      tcg_gen_extract_i32(ret, ret, 26, 1);
      return ret;
 diff --git a/target/arm/tcg/mve_helper.c b/target/arm/tcg/mve_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/mve_helper.c
 +++ b/target/arm/tcg/mve_helper.c
@@ -XXX,XX +XXX,XX @@ static void do_vadc(CPUARMState *env, uint32_t *d, uint32_t *n, uint32_t *m,
      if (update_flags) {
          /* Store C, clear NZV. */
 -        env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_NZCV_MASK;
 -        env->vfp.xregs[ARM_VFP_FPSCR] |= carry_in * FPCR_C;
 +        env->vfp.fpsr &= ~FPCR_NZCV_MASK;
 +        env->vfp.fpsr |= carry_in * FPCR_C;
      }
      mve_advance_vpt(env);
  }
- void HELPER(mve_vadc)(CPUARMState *env, void *vd, void *vn, void *vm)
+ uint32_t psw_read(CPUTriCoreState *env)
  {
 -    bool carry_in = env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_C;
 +    bool carry_in = env->vfp.fpsr & FPCR_C;
      do_vadc(env, vd, vn, vm, 0, carry_in, false);
  }
  void HELPER(mve_vsbc)(CPUARMState *env, void *vd, void *vn, void *vm)
  {
 -    bool carry_in = env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_C;
 +    bool carry_in = env->vfp.fpsr & FPCR_C;
      do_vadc(env, vd, vn, vm, -1, carry_in, false);
  }
@@ -XXX,XX +XXX,XX @@ static void do_vcvt_sh(CPUARMState *env, void *vd, void *vm, int top)
      uint32_t *m = vm;
      uint16_t r;
      uint16_t mask = mve_element_mask(env);
 -    bool ieee = !(env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_AHP);
 +    bool ieee = !(env->vfp.fpcr & FPCR_AHP);
      unsigned e;
      float_status *fpst;
      float_status scratch_fpst;
@@ -XXX,XX +XXX,XX @@ static void do_vcvt_hs(CPUARMState *env, void *vd, void *vm, int top)
      uint16_t *m = vm;
      uint32_t r;
      uint16_t mask = mve_element_mask(env);
 -    bool ieee = !(env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_AHP);
 +    bool ieee = !(env->vfp.fpcr & FPCR_AHP);
      unsigned e;
      float_status *fpst;
      float_status scratch_fpst;
 diff --git a/target/arm/tcg/translate-m-nocp.c b/target/arm/tcg/translate-m-nocp.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/translate-m-nocp.c
 +++ b/target/arm/tcg/translate-m-nocp.c
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
 , 16, qc);
          }
          tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
 -        fpscr = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
 +        fpscr = load_cpu_field_low32(vfp.fpsr);
          tcg_gen_andi_i32(fpscr, fpscr, ~FPCR_NZCV_MASK);
          tcg_gen_or_i32(fpscr, fpscr, tmp);
 -        store_cpu_field(fpscr, vfp.xregs[ARM_VFP_FPSCR]);
 +        store_cpu_field_low32(fpscr, vfp.fpsr);
          break;
      }
      case ARM_VFP_FPCXT_NS:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
           * Read just NZCV; this is a special case to avoid the
           * helper call for the "VMRS to CPSR.NZCV" insn.
           */
 -        tmp = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
 +        tmp = load_cpu_field_low32(vfp.fpsr);
          tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
          storefn(s, opaque, tmp, true);
          break;
 diff --git a/target/arm/tcg/translate-vfp.c b/target/arm/tcg/translate-vfp.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/translate-vfp.c
 +++ b/target/arm/tcg/translate-vfp.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VMSR_VMRS(DisasContext *s, arg_VMSR_VMRS *a)
              break;
          case ARM_VFP_FPSCR:
              if (a->rt == 15) {
 -                tmp = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
 +                tmp = load_cpu_field_low32(vfp.fpsr);
                  tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
              } else {
                  tmp = tcg_temp_new_i32();
 diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/vfp_helper.c
 +++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
  static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
  {
 -    uint32_t changed = env->vfp.xregs[ARM_VFP_FPSCR];
 +    uint64_t changed = env->vfp.fpcr;
      changed ^= val;
      if (changed & (3 << 22)) {
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
  uint32_t vfp_get_fpcr(CPUARMState *env)
  {
 -    uint32_t fpcr = (env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_MASK)
 +    uint32_t fpcr = env->vfp.fpcr
          | (env->vfp.vec_len << 16)
          | (env->vfp.vec_stride << 20);
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpcr(CPUARMState *env)
  uint32_t vfp_get_fpsr(CPUARMState *env)
  {
 -    uint32_t fpsr = env->vfp.xregs[ARM_VFP_FPSCR] & FPSR_MASK;
 +    uint32_t fpsr = env->vfp.fpsr;
      uint32_t i;
      fpsr |= vfp_get_fpsr_from_host(env);
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
      }
      /*
 -     * The only FPSR bits we keep in vfp.xregs[FPSCR] are NZCV:
 +     * The only FPSR bits we keep in vfp.fpsr are NZCV:
       * the exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
       * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
 -     * and zero any of the other FPSR bits (but preserve the FPCR
 -     * bits).
 +     * and zero any of the other FPSR bits.
       */
      val &= FPCR_NZCV_MASK;
 -    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPSR_MASK;
 -    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
 +    env->vfp.fpsr = val;
  }
  void vfp_set_fpcr(CPUARMState *env, uint32_t val)
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
       * We don't implement trapped exception handling, so the
       * trap enable bits, IDE|IXE|UFE|OFE|DZE|IOE are all RAZ/WI (not RES0!)
       *
 -     * The FPCR bits we keep in vfp.xregs[FPSCR] are AHP, DN, FZ, RMode
 +     * The FPCR bits we keep in vfp.fpcr are AHP, DN, FZ, RMode
       * and FZ16. Len, Stride and LTPSIZE we just handled. Store those bits
       * there, and zero any of the other FPCR bits and the RES0 and RAZ/WI
       * bits.
       */
      val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
 -    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_MASK;
 -    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
 +    env->vfp.fpcr = val;
  }
  void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void softfloat_to_vfp_compare(CPUARMState *env, FloatRelation cmp)
      default:
          g_assert_not_reached();
      }
 -    env->vfp.xregs[ARM_VFP_FPSCR] =
 -        deposit32(env->vfp.xregs[ARM_VFP_FPSCR], 28, 4, flags);
 +    env->vfp.fpsr = deposit64(env->vfp.fpsr, 28, 4, flags); /* NZCV */
  }
  /* XXX: check quiet/signaling case */
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vjcvt)(float64 value, CPUARMState *env)
      uint32_t z = (pair >> 32) == 0;
      /* Store Z, clear NCV, in FPSCR.NZCV.  */
 -    env->vfp.xregs[ARM_VFP_FPSCR]
 -        = (env->vfp.xregs[ARM_VFP_FPSCR] & ~CPSR_NZCV) | (z * CPSR_Z);
 +    env->vfp.fpsr = (env->vfp.fpsr & ~FPCR_NZCV_MASK) | (z * FPCR_Z);
      return result;
  }
 --
 .34.1

-New patch
+[PULL 60/72] fpu: Remove default handling for dnan_pattern
+Now that all our targets have bene converted to explicitly specify
+their pattern for the default NaN value we can remove the remaining
+fallback code in parts64_default_nan().
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
+---
+ fpu/softfloat-specialize.c.inc | 14 --------------
+file changed, 14 deletions(-)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat-specialize.c.inc
++++ b/fpu/softfloat-specialize.c.inc
+@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
+     uint64_t frac;
+     uint8_t dnan_pattern = status->default_nan_pattern;
+-    if (dnan_pattern == 0) {
+-        /*
+-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
+-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
+-         * do not have floating-point.
+-         */
+-        if (snan_bit_is_one(status)) {
+-            /* sign bit clear, set all frac bits other than msb */
+-            dnan_pattern = 0b00111111;
+-        } else {
+-            /* sign bit clear, set frac msb */
+-            dnan_pattern = 0b01000000;
+-        }
+-    }
+     assert(dnan_pattern != 0);
+     sign = dnan_pattern >> 7;
+--
+.34.1

-[PULL 21/24] target/arm: Convert SQDMULL, SQDMLAL, SQDMLSL to decodetree
+[PULL 61/72] softfloat: Inline pickNaNMulAdd
 From: Richard Henderson <richard.henderson@linaro.org>
+Inline pickNaNMulAdd into its only caller.  This makes
+one assert redundant with the immediately preceding IF.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240709000610.382391-4-richard.henderson@linaro.org
+Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
 [PMM: keep comment from old code in new location]
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |  33 ++
+ fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
- target/arm/tcg/translate-a64.c | 604 ++++++---------------------------
+ fpu/softfloat-specialize.c.inc | 54 ----------------------------------
-files changed, 138 insertions(+), 499 deletions(-)
+files changed, 40 insertions(+), 55 deletions(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/tcg/a64.decode
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ SQRDMULH_s      0111 1110 ..1 ..... 10110 1 ..... ..... @rrr_e
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
- SQRDMLAH_s      0111 1110 ..0 ..... 10000 1 ..... ..... @rrr_e
+     }
- SQRDMLSH_s      0111 1110 ..0 ..... 10001 1 ..... ..... @rrr_e
+     if (s->default_nan_mode) {
-+# Decode scalar x scalar as scalar x indexed, with index 0.
++        /*
-+SQDMULL_si      0101 1110 011 rm:5  11010 0 rn:5  rd:5  &rrx_e idx=0 esz=1
++         * We guarantee not to require the target to tell us how to
-+SQDMULL_si      0101 1110 101 rm:5  11010 0 rn:5  rd:5  &rrx_e idx=0 esz=2
++         * pick a NaN if we're always returning the default NaN.
-+SQDMLAL_si      0101 1110 011 rm:5  10010 0 rn:5  rd:5  &rrx_e idx=0 esz=1
++         * But if we're not in default-NaN mode then the target must
-+SQDMLAL_si      0101 1110 101 rm:5  10010 0 rn:5  rd:5  &rrx_e idx=0 esz=2
++         * specify.
-+SQDMLSL_si      0101 1110 011 rm:5  10110 0 rn:5  rd:5  &rrx_e idx=0 esz=1
++         */
-+SQDMLSL_si      0101 1110 101 rm:5  10110 0 rn:5  rd:5  &rrx_e idx=0 esz=2
+         which = 3;
 +    } else if (infzero) {
 +        /*
 +         * Inf * 0 + NaN -- some implementations return the
 +         * default NaN here, and some return the input NaN.
 +         */
 +        switch (s->float_infzeronan_rule) {
 +        case float_infzeronan_dnan_never:
 +            which = 2;
 +            break;
 +        case float_infzeronan_dnan_always:
 +            which = 3;
 +            break;
 +        case float_infzeronan_dnan_if_qnan:
 +            which = is_qnan(c->cls) ? 3 : 2;
 +            break;
 +        default:
 +            g_assert_not_reached();
 +        }
      } else {
 -        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
 +        FloatClass cls[3] = { a->cls, b->cls, c->cls };
 +        Float3NaNPropRule rule = s->float_3nan_prop_rule;
 +
- ### Advanced SIMD scalar pairwise
++        assert(rule != float_3nan_prop_none);
++        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
- FADDP_s         0101 1110 0011 0000 1101 10 ..... ..... @rr_h
++            /* We have at least one SNaN input and should prefer it */
-@@ -XXX,XX +XXX,XX @@ UABAL_v         0.10 1110 ..1 ..... 01010 0 ..... ..... @qrrr_e
++            do {
- SABDL_v         0.00 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
++                which = rule & R_3NAN_1ST_MASK;
- UABDL_v         0.10 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
++                rule >>= R_3NAN_1ST_LENGTH;
++            } while (!is_snan(cls[which]));
-+SQDMULL_v       0.00 1110 011 ..... 11010 0 ..... ..... @qrrr_h
++        } else {
-+SQDMULL_v       0.00 1110 101 ..... 11010 0 ..... ..... @qrrr_s
++            do {
-+SQDMLAL_v       0.00 1110 011 ..... 10010 0 ..... ..... @qrrr_h
++                which = rule & R_3NAN_1ST_MASK;
-+SQDMLAL_v       0.00 1110 101 ..... 10010 0 ..... ..... @qrrr_s
++                rule >>= R_3NAN_1ST_LENGTH;
-+SQDMLSL_v       0.00 1110 011 ..... 10110 0 ..... ..... @qrrr_h
++            } while (!is_nan(cls[which]));
-+SQDMLSL_v       0.00 1110 101 ..... 10110 0 ..... ..... @qrrr_s
++        }
-+
+     }
- ### Advanced SIMD scalar x indexed element
+     if (which == 3) {
- FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ SQRDMLAH_si     0111 1111 10 .. .... 1101 . 0 ..... .....   @rrx_s
  SQRDMLSH_si     0111 1111 01 .. .... 1111 . 0 ..... .....   @rrx_h
  SQRDMLSH_si     0111 1111 10 .. .... 1111 . 0 ..... .....   @rrx_s
 +SQDMULL_si      0101 1111 01 .. .... 1011 . 0 ..... .....   @rrx_h
 +SQDMULL_si      0101 1111 10 . ..... 1011 . 0 ..... .....   @rrx_s
 +
 +SQDMLAL_si      0101 1111 01 .. .... 0011 . 0 ..... .....   @rrx_h
 +SQDMLAL_si      0101 1111 10 . ..... 0011 . 0 ..... .....   @rrx_s
 +
 +SQDMLSL_si      0101 1111 01 .. .... 0111 . 0 ..... .....   @rrx_h
 +SQDMLSL_si      0101 1111 10 . ..... 0111 . 0 ..... .....   @rrx_s
 +
  ### Advanced SIMD vector x indexed element
  FMUL_vi         0.00 1111 00 .. .... 1001 . 0 ..... .....   @qrrx_h
@@ -XXX,XX +XXX,XX @@ SMLSL_vi        0.00 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
  UMLSL_vi        0.10 1111 01 .. .... 0110 . 0 ..... .....   @qrrx_h
  UMLSL_vi        0.10 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
 +SQDMULL_vi      0.00 1111 01 .. .... 1011 . 0 ..... .....   @qrrx_h
 +SQDMULL_vi      0.00 1111 10 . ..... 1011 . 0 ..... .....   @qrrx_s
 +
 +SQDMLAL_vi      0.00 1111 01 .. .... 0011 . 0 ..... .....   @qrrx_h
 +SQDMLAL_vi      0.00 1111 10 . ..... 0011 . 0 ..... .....   @qrrx_s
 +
 +SQDMLSL_vi      0.00 1111 01 .. .... 0111 . 0 ..... .....   @qrrx_h
 +SQDMLSL_vi      0.00 1111 10 . ..... 0111 . 0 ..... .....   @qrrx_s
 +
  # Floating-point conditional select
  FCSEL           0001 1110 .. 1 rm:5 cond:4 11 rn:5 rd:5     esz=%esz_hsd
 diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-a64.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/tcg/translate-a64.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ TRANS(UABAL_v, do_3op_widening,
+@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
        a->esz, a->q, a->rd, a->rn, a->rm, -1,
        gen_uaba_i64, true)
 +static void gen_sqdmull_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    tcg_gen_mul_i64(d, n, m);
 +    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, d);
 +}
 +
 +static void gen_sqdmull_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    tcg_gen_mul_i64(d, n, m);
 +    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, d);
 +}
 +
 +static void gen_sqdmlal_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +
 +    tcg_gen_mul_i64(t, n, m);
 +    gen_helper_neon_addl_saturate_s32(t, tcg_env, t, t);
 +    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, t);
 +}
 +
 +static void gen_sqdmlal_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +
 +    tcg_gen_mul_i64(t, n, m);
 +    gen_helper_neon_addl_saturate_s64(t, tcg_env, t, t);
 +    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, t);
 +}
 +
 +static void gen_sqdmlsl_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +
 +    tcg_gen_mul_i64(t, n, m);
 +    gen_helper_neon_addl_saturate_s32(t, tcg_env, t, t);
 +    tcg_gen_neg_i64(t, t);
 +    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, t);
 +}
 +
 +static void gen_sqdmlsl_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +
 +    tcg_gen_mul_i64(t, n, m);
 +    gen_helper_neon_addl_saturate_s64(t, tcg_env, t, t);
 +    tcg_gen_neg_i64(t, t);
 +    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, t);
 +}
 +
 +TRANS(SQDMULL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
 +TRANS(SQDMLAL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
 +TRANS(SQDMLSL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
 +
 +TRANS(SQDMULL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
 +TRANS(SQDMLAL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
 +TRANS(SQDMLSL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
 +
  /*
   * Advanced SIMD scalar/vector x indexed element
   */
@@ -XXX,XX +XXX,XX @@ static bool do_env_scalar3_idx_hs(DisasContext *s, arg_rrx_e *a,
  TRANS_FEAT(SQRDMLAH_si, aa64_rdm, do_env_scalar3_idx_hs, a, &f_scalar_sqrdmlah)
  TRANS_FEAT(SQRDMLSH_si, aa64_rdm, do_env_scalar3_idx_hs, a, &f_scalar_sqrdmlsh)
 +static bool do_scalar_muladd_widening_idx(DisasContext *s, arg_rrx_e *a,
 +                                          NeonGenTwo64OpFn *fn, bool acc)
 +{
 +    if (fp_access_check(s)) {
 +        TCGv_i64 t0 = tcg_temp_new_i64();
 +        TCGv_i64 t1 = tcg_temp_new_i64();
 +        TCGv_i64 t2 = tcg_temp_new_i64();
 +        unsigned vsz, dofs;
 +
 +        if (acc) {
 +            read_vec_element(s, t0, a->rd, 0, a->esz + 1);
 +        }
 +        read_vec_element(s, t1, a->rn, 0, a->esz | MO_SIGN);
 +        read_vec_element(s, t2, a->rm, a->idx, a->esz | MO_SIGN);
 +        fn(t0, t1, t2);
 +
 +        /* Clear the whole register first, then store scalar. */
 +        vsz = vec_full_reg_size(s);
 +        dofs = vec_full_reg_offset(s, a->rd);
 +        tcg_gen_gvec_dup_imm(MO_64, dofs, vsz, vsz, 0);
 +        write_vec_element(s, t0, a->rd, 0, a->esz + 1);
 +    }
 +    return true;
 +}
 +
 +TRANS(SQDMULL_si, do_scalar_muladd_widening_idx, a,
 +      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
 +TRANS(SQDMLAL_si, do_scalar_muladd_widening_idx, a,
 +      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
 +TRANS(SQDMLSL_si, do_scalar_muladd_widening_idx, a,
 +      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
 +
  static bool do_fp3_vector_idx(DisasContext *s, arg_qrrx_e *a,
                                gen_helper_gvec_3_ptr * const fns[3])
  {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_shift_imm(DisasContext *s, uint32_t insn)
      }
  }
--/* AdvSIMD scalar three different
+-/*----------------------------------------------------------------------------
-- *  31 30  29 28       24 23  22  21 20  16 15    12 11 10 9    5 4    0
+-| Select which NaN to propagate for a three-input operation.
-- * +-----+---+-----------+------+---+------+--------+-----+------+------+
+-| For the moment we assume that no CPU needs the 'larger significand'
-- * | 0 1 | U | 1 1 1 1 0 | size | 1 |  Rm  | opcode | 0 0 |  Rn  |  Rd  |
+-| information.
-- * +-----+---+-----------+------+---+------+--------+-----+------+------+
+-| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-- */
+-*----------------------------------------------------------------------------*/
--static void disas_simd_scalar_three_reg_diff(DisasContext *s, uint32_t insn)
+-static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
 -                         bool infzero, bool have_snan, float_status *status)
 -{
--    bool is_u = extract32(insn, 29, 1);
+-    FloatClass cls[3] = { a_cls, b_cls, c_cls };
--    int size = extract32(insn, 22, 2);
+-    Float3NaNPropRule rule = status->float_3nan_prop_rule;
--    int opcode = extract32(insn, 12, 4);
+-    int which;
 -    int rm = extract32(insn, 16, 5);
 -    int rn = extract32(insn, 5, 5);
 -    int rd = extract32(insn, 0, 5);
 -
--    if (is_u) {
+-    /*
--        unallocated_encoding(s);
+-     * We guarantee not to require the target to tell us how to
--        return;
+-     * pick a NaN if we're always returning the default NaN.
--    }
+-     * But if we're not in default-NaN mode then the target must
 -     * specify.
 -     */
 -    assert(!status->default_nan_mode);
 -
--    switch (opcode) {
+-    if (infzero) {
--    case 0x9: /* SQDMLAL, SQDMLAL2 */
+-        /*
--    case 0xb: /* SQDMLSL, SQDMLSL2 */
+-         * Inf * 0 + NaN -- some implementations return the default NaN here,
--    case 0xd: /* SQDMULL, SQDMULL2 */
+-         * and some return the input NaN.
--        if (size == 0 || size == 3) {
+-         */
--            unallocated_encoding(s);
+-        switch (status->float_infzeronan_rule) {
--            return;
+-        case float_infzeronan_dnan_never:
--        }
+-            return 2;
--        break;
+-        case float_infzeronan_dnan_always:
--    default:
+-            return 3;
--        unallocated_encoding(s);
+-        case float_infzeronan_dnan_if_qnan:
--        return;
+-            return is_qnan(c_cls) ? 3 : 2;
 -    }
 -
 -    if (!fp_access_check(s)) {
 -        return;
 -    }
 -
 -    if (size == 2) {
 -        TCGv_i64 tcg_op1 = tcg_temp_new_i64();
 -        TCGv_i64 tcg_op2 = tcg_temp_new_i64();
 -        TCGv_i64 tcg_res = tcg_temp_new_i64();
 -
 -        read_vec_element(s, tcg_op1, rn, 0, MO_32 | MO_SIGN);
 -        read_vec_element(s, tcg_op2, rm, 0, MO_32 | MO_SIGN);
 -
 -        tcg_gen_mul_i64(tcg_res, tcg_op1, tcg_op2);
 -        gen_helper_neon_addl_saturate_s64(tcg_res, tcg_env, tcg_res, tcg_res);
 -
 -        switch (opcode) {
 -        case 0xd: /* SQDMULL, SQDMULL2 */
 -            break;
 -        case 0xb: /* SQDMLSL, SQDMLSL2 */
 -            tcg_gen_neg_i64(tcg_res, tcg_res);
 -            /* fall through */
 -        case 0x9: /* SQDMLAL, SQDMLAL2 */
 -            read_vec_element(s, tcg_op1, rd, 0, MO_64);
 -            gen_helper_neon_addl_saturate_s64(tcg_res, tcg_env,
 -                                              tcg_res, tcg_op1);
 -            break;
 -        default:
 -            g_assert_not_reached();
 -        }
+-    }
 -
--        write_fp_dreg(s, rd, tcg_res);
+-    assert(rule != float_3nan_prop_none);
 -    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
 -        /* We have at least one SNaN input and should prefer it */
 -        do {
 -            which = rule & R_3NAN_1ST_MASK;
 -            rule >>= R_3NAN_1ST_LENGTH;
 -        } while (!is_snan(cls[which]));
 -    } else {
--        TCGv_i32 tcg_op1 = read_fp_hreg(s, rn);
+-        do {
--        TCGv_i32 tcg_op2 = read_fp_hreg(s, rm);
+-            which = rule & R_3NAN_1ST_MASK;
--        TCGv_i64 tcg_res = tcg_temp_new_i64();
+-            rule >>= R_3NAN_1ST_LENGTH;
--
+-        } while (!is_nan(cls[which]));
 -        gen_helper_neon_mull_s16(tcg_res, tcg_op1, tcg_op2);
 -        gen_helper_neon_addl_saturate_s32(tcg_res, tcg_env, tcg_res, tcg_res);
 -
 -        switch (opcode) {
 -        case 0xd: /* SQDMULL, SQDMULL2 */
 -            break;
 -        case 0xb: /* SQDMLSL, SQDMLSL2 */
 -            gen_helper_neon_negl_u32(tcg_res, tcg_res);
 -            /* fall through */
 -        case 0x9: /* SQDMLAL, SQDMLAL2 */
 -        {
 -            TCGv_i64 tcg_op3 = tcg_temp_new_i64();
 -            read_vec_element(s, tcg_op3, rd, 0, MO_32);
 -            gen_helper_neon_addl_saturate_s32(tcg_res, tcg_env,
 -                                              tcg_res, tcg_op3);
 -            break;
 -        }
 -        default:
 -            g_assert_not_reached();
 -        }
 -
 -        tcg_gen_ext32u_i64(tcg_res, tcg_res);
 -        write_fp_dreg(s, rd, tcg_res);
 -    }
+-    return which;
 -}
 -
- static void handle_2misc_64(DisasContext *s, int opcode, bool u,
+ /*----------------------------------------------------------------------------
-                             TCGv_i64 tcg_rd, TCGv_i64 tcg_rn,
+ | Returns 1 if the double-precision floating-point value `a' is a quiet
-                             TCGv_i32 tcg_rmode, TCGv_ptr tcg_fpstatus)
+ | NaN; otherwise returns 0.
@@ -XXX,XX +XXX,XX @@ static void gen_neon_addl(int size, bool is_sub, TCGv_i64 tcg_res,
      genfn(tcg_res, tcg_op1, tcg_op2);
  }
 -static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
 -                                int opcode, int rd, int rn, int rm)
 -{
 -    /* 3-reg-different widening insns: 64 x 64 -> 128 */
 -    TCGv_i64 tcg_res[2];
 -    int pass, accop;
 -
 -    tcg_res[0] = tcg_temp_new_i64();
 -    tcg_res[1] = tcg_temp_new_i64();
 -
 -    /* Does this op do an adding accumulate, a subtracting accumulate,
 -     * or no accumulate at all?
 -     */
 -    switch (opcode) {
 -    case 5:
 -    case 8:
 -    case 9:
 -        accop = 1;
 -        break;
 -    case 10:
 -    case 11:
 -        accop = -1;
 -        break;
 -    default:
 -        accop = 0;
 -        break;
 -    }
 -
 -    if (accop != 0) {
 -        read_vec_element(s, tcg_res[0], rd, 0, MO_64);
 -        read_vec_element(s, tcg_res[1], rd, 1, MO_64);
 -    }
 -
 -    /* size == 2 means two 32x32->64 operations; this is worth special
 -     * casing because we can generally handle it inline.
 -     */
 -    if (size == 2) {
 -        for (pass = 0; pass < 2; pass++) {
 -            TCGv_i64 tcg_op1 = tcg_temp_new_i64();
 -            TCGv_i64 tcg_op2 = tcg_temp_new_i64();
 -            TCGv_i64 tcg_passres;
 -            MemOp memop = MO_32 | (is_u ? 0 : MO_SIGN);
 -
 -            int elt = pass + is_q * 2;
 -
 -            read_vec_element(s, tcg_op1, rn, elt, memop);
 -            read_vec_element(s, tcg_op2, rm, elt, memop);
 -
 -            if (accop == 0) {
 -                tcg_passres = tcg_res[pass];
 -            } else {
 -                tcg_passres = tcg_temp_new_i64();
 -            }
 -
 -            switch (opcode) {
 -            case 9: /* SQDMLAL, SQDMLAL2 */
 -            case 11: /* SQDMLSL, SQDMLSL2 */
 -            case 13: /* SQDMULL, SQDMULL2 */
 -                tcg_gen_mul_i64(tcg_passres, tcg_op1, tcg_op2);
 -                gen_helper_neon_addl_saturate_s64(tcg_passres, tcg_env,
 -                                                  tcg_passres, tcg_passres);
 -                break;
 -            default:
 -            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 -            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 -            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 -            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -                g_assert_not_reached();
 -            }
 -
 -            if (accop != 0) {
 -                /* saturating accumulate ops */
 -                if (accop < 0) {
 -                    tcg_gen_neg_i64(tcg_passres, tcg_passres);
 -                }
 -                gen_helper_neon_addl_saturate_s64(tcg_res[pass], tcg_env,
 -                                                  tcg_res[pass], tcg_passres);
 -            }
 -        }
 -    } else {
 -        /* size 0 or 1, generally helper functions */
 -        for (pass = 0; pass < 2; pass++) {
 -            TCGv_i32 tcg_op1 = tcg_temp_new_i32();
 -            TCGv_i32 tcg_op2 = tcg_temp_new_i32();
 -            TCGv_i64 tcg_passres;
 -            int elt = pass + is_q * 2;
 -
 -            read_vec_element_i32(s, tcg_op1, rn, elt, MO_32);
 -            read_vec_element_i32(s, tcg_op2, rm, elt, MO_32);
 -
 -            if (accop == 0) {
 -                tcg_passres = tcg_res[pass];
 -            } else {
 -                tcg_passres = tcg_temp_new_i64();
 -            }
 -
 -            switch (opcode) {
 -            case 9: /* SQDMLAL, SQDMLAL2 */
 -            case 11: /* SQDMLSL, SQDMLSL2 */
 -            case 13: /* SQDMULL, SQDMULL2 */
 -                assert(size == 1);
 -                gen_helper_neon_mull_s16(tcg_passres, tcg_op1, tcg_op2);
 -                gen_helper_neon_addl_saturate_s32(tcg_passres, tcg_env,
 -                                                  tcg_passres, tcg_passres);
 -                break;
 -            default:
 -            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 -            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 -            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 -            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -                g_assert_not_reached();
 -            }
 -
 -            if (accop != 0) {
 -                /* saturating accumulate ops */
 -                if (accop < 0) {
 -                    gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
 -                }
 -                gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
 -                                                  tcg_res[pass],
 -                                                  tcg_passres);
 -            }
 -        }
 -    }
 -
 -    write_vec_element(s, tcg_res[0], rd, 0, MO_64);
 -    write_vec_element(s, tcg_res[1], rd, 1, MO_64);
 -}
 -
  static void handle_3rd_wide(DisasContext *s, int is_q, int is_u, int size,
                              int opcode, int rd, int rn, int rm)
  {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
              break;
          }
          return;
 -    case 9: /* SQDMLAL, SQDMLAL2 */
 -    case 11: /* SQDMLSL, SQDMLSL2 */
 -    case 13: /* SQDMULL, SQDMULL2 */
 -        if (is_u || size == 0) {
 -            unallocated_encoding(s);
 -            return;
 -        }
 -        /* 64 x 64 -> 128 */
 -        if (size == 3) {
 -            unallocated_encoding(s);
 -            return;
 -        }
 -        if (!fp_access_check(s)) {
 -            return;
 -        }
 -
 -        handle_3rd_widening(s, is_q, is_u, size, opcode, rd, rn, rm);
 -        break;
      default:
      case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
      case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
      case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
      case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
      case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +    case 9: /* SQDMLAL, SQDMLAL2 */
      case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 +    case 11: /* SQDMLSL, SQDMLSL2 */
      case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
 +    case 13: /* SQDMULL, SQDMULL2 */
          /* opcode 15 not allocated */
          unallocated_encoding(s);
          break;
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc_fp16(DisasContext *s, uint32_t insn)
      }
  }
 -/* AdvSIMD scalar x indexed element
 - *  31 30  29 28       24 23  22 21  20  19  16 15 12  11  10 9    5 4    0
 - * +-----+---+-----------+------+---+---+------+-----+---+---+------+------+
 - * | 0 1 | U | 1 1 1 1 1 | size | L | M |  Rm  | opc | H | 0 |  Rn  |  Rd  |
 - * +-----+---+-----------+------+---+---+------+-----+---+---+------+------+
 - * AdvSIMD vector x indexed element
 - *   31  30  29 28       24 23  22 21  20  19  16 15 12  11  10 9    5 4    0
 - * +---+---+---+-----------+------+---+---+------+-----+---+---+------+------+
 - * | 0 | Q | U | 0 1 1 1 1 | size | L | M |  Rm  | opc | H | 0 |  Rn  |  Rd  |
 - * +---+---+---+-----------+------+---+---+------+-----+---+---+------+------+
 - */
 -static void disas_simd_indexed(DisasContext *s, uint32_t insn)
 -{
 -    /* This encoding has two kinds of instruction:
 -     *  normal, where we perform elt x idxelt => elt for each
 -     *     element in the vector
 -     *  long, where we perform elt x idxelt and generate a result of
 -     *     double the width of the input element
 -     * The long ops have a 'part' specifier (ie come in INSN, INSN2 pairs).
 -     */
 -    bool is_scalar = extract32(insn, 28, 1);
 -    bool is_q = extract32(insn, 30, 1);
 -    bool u = extract32(insn, 29, 1);
 -    int size = extract32(insn, 22, 2);
 -    int l = extract32(insn, 21, 1);
 -    int m = extract32(insn, 20, 1);
 -    /* Note that the Rm field here is only 4 bits, not 5 as it usually is */
 -    int rm = extract32(insn, 16, 4);
 -    int opcode = extract32(insn, 12, 4);
 -    int h = extract32(insn, 11, 1);
 -    int rn = extract32(insn, 5, 5);
 -    int rd = extract32(insn, 0, 5);
 -    int index;
 -
 -    switch (16 * u + opcode) {
 -    case 0x03: /* SQDMLAL, SQDMLAL2 */
 -    case 0x07: /* SQDMLSL, SQDMLSL2 */
 -    case 0x0b: /* SQDMULL, SQDMULL2 */
 -        break;
 -    default:
 -    case 0x00: /* FMLAL */
 -    case 0x01: /* FMLA */
 -    case 0x02: /* SMLAL, SMLAL2 */
 -    case 0x04: /* FMLSL */
 -    case 0x05: /* FMLS */
 -    case 0x06: /* SMLSL, SMLSL2 */
 -    case 0x08: /* MUL */
 -    case 0x09: /* FMUL */
 -    case 0x0a: /* SMULL, SMULL2 */
 -    case 0x0c: /* SQDMULH */
 -    case 0x0d: /* SQRDMULH */
 -    case 0x0e: /* SDOT */
 -    case 0x0f: /* SUDOT / BFDOT / USDOT / BFMLAL */
 -    case 0x10: /* MLA */
 -    case 0x11: /* FCMLA #0 */
 -    case 0x12: /* UMLAL, UMLAL2 */
 -    case 0x13: /* FCMLA #90 */
 -    case 0x14: /* MLS */
 -    case 0x15: /* FCMLA #180 */
 -    case 0x16: /* UMLSL, UMLSL2 */
 -    case 0x17: /* FCMLA #270 */
 -    case 0x18: /* FMLAL2 */
 -    case 0x19: /* FMULX */
 -    case 0x1a: /* UMULL, UMULL2 */
 -    case 0x1c: /* FMLSL2 */
 -    case 0x1d: /* SQRDMLAH */
 -    case 0x1e: /* UDOT */
 -    case 0x1f: /* SQRDMLSH */
 -        unallocated_encoding(s);
 -        return;
 -    }
 -
 -    /* Given MemOp size, adjust register and indexing.  */
 -    switch (size) {
 -    case MO_8:
 -    case MO_64:
 -        unallocated_encoding(s);
 -        return;
 -    case MO_16:
 -        index = h << 2 | l << 1 | m;
 -        break;
 -    case MO_32:
 -        index = h << 1 | l;
 -        rm |= m << 4;
 -        break;
 -    default:
 -        g_assert_not_reached();
 -    }
 -
 -    if (!fp_access_check(s)) {
 -        return;
 -    }
 -
 -    if (size == 3) {
 -        g_assert_not_reached();
 -    } else {
 -        /* long ops: 16x16->32 or 32x32->64 */
 -        TCGv_i64 tcg_res[2];
 -        int pass;
 -        bool satop = extract32(opcode, 0, 1);
 -        MemOp memop = MO_32;
 -
 -        if (satop || !u) {
 -            memop |= MO_SIGN;
 -        }
 -
 -        if (size == 2) {
 -            TCGv_i64 tcg_idx = tcg_temp_new_i64();
 -
 -            read_vec_element(s, tcg_idx, rm, index, memop);
 -
 -            for (pass = 0; pass < (is_scalar ? 1 : 2); pass++) {
 -                TCGv_i64 tcg_op = tcg_temp_new_i64();
 -                TCGv_i64 tcg_passres;
 -                int passelt;
 -
 -                if (is_scalar) {
 -                    passelt = 0;
 -                } else {
 -                    passelt = pass + (is_q * 2);
 -                }
 -
 -                read_vec_element(s, tcg_op, rn, passelt, memop);
 -
 -                tcg_res[pass] = tcg_temp_new_i64();
 -
 -                if (opcode == 0xa || opcode == 0xb) {
 -                    /* Non-accumulating ops */
 -                    tcg_passres = tcg_res[pass];
 -                } else {
 -                    tcg_passres = tcg_temp_new_i64();
 -                }
 -
 -                tcg_gen_mul_i64(tcg_passres, tcg_op, tcg_idx);
 -
 -                if (satop) {
 -                    /* saturating, doubling */
 -                    gen_helper_neon_addl_saturate_s64(tcg_passres, tcg_env,
 -                                                      tcg_passres, tcg_passres);
 -                }
 -
 -                if (opcode == 0xa || opcode == 0xb) {
 -                    continue;
 -                }
 -
 -                /* Accumulating op: handle accumulate step */
 -                read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
 -
 -                switch (opcode) {
 -                case 0x7: /* SQDMLSL, SQDMLSL2 */
 -                    tcg_gen_neg_i64(tcg_passres, tcg_passres);
 -                    /* fall through */
 -                case 0x3: /* SQDMLAL, SQDMLAL2 */
 -                    gen_helper_neon_addl_saturate_s64(tcg_res[pass], tcg_env,
 -                                                      tcg_res[pass],
 -                                                      tcg_passres);
 -                    break;
 -                default:
 -                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -                    g_assert_not_reached();
 -                }
 -            }
 -
 -            clear_vec_high(s, !is_scalar, rd);
 -        } else {
 -            TCGv_i32 tcg_idx = tcg_temp_new_i32();
 -
 -            assert(size == 1);
 -            read_vec_element_i32(s, tcg_idx, rm, index, size);
 -
 -            if (!is_scalar) {
 -                /* The simplest way to handle the 16x16 indexed ops is to
 -                 * duplicate the index into both halves of the 32 bit tcg_idx
 -                 * and then use the usual Neon helpers.
 -                 */
 -                tcg_gen_deposit_i32(tcg_idx, tcg_idx, tcg_idx, 16, 16);
 -            }
 -
 -            for (pass = 0; pass < (is_scalar ? 1 : 2); pass++) {
 -                TCGv_i32 tcg_op = tcg_temp_new_i32();
 -                TCGv_i64 tcg_passres;
 -
 -                if (is_scalar) {
 -                    read_vec_element_i32(s, tcg_op, rn, pass, size);
 -                } else {
 -                    read_vec_element_i32(s, tcg_op, rn,
 -                                         pass + (is_q * 2), MO_32);
 -                }
 -
 -                tcg_res[pass] = tcg_temp_new_i64();
 -
 -                if (opcode == 0xa || opcode == 0xb) {
 -                    /* Non-accumulating ops */
 -                    tcg_passres = tcg_res[pass];
 -                } else {
 -                    tcg_passres = tcg_temp_new_i64();
 -                }
 -
 -                if (memop & MO_SIGN) {
 -                    gen_helper_neon_mull_s16(tcg_passres, tcg_op, tcg_idx);
 -                } else {
 -                    gen_helper_neon_mull_u16(tcg_passres, tcg_op, tcg_idx);
 -                }
 -                if (satop) {
 -                    gen_helper_neon_addl_saturate_s32(tcg_passres, tcg_env,
 -                                                      tcg_passres, tcg_passres);
 -                }
 -
 -                if (opcode == 0xa || opcode == 0xb) {
 -                    continue;
 -                }
 -
 -                /* Accumulating op: handle accumulate step */
 -                read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
 -
 -                switch (opcode) {
 -                case 0x7: /* SQDMLSL, SQDMLSL2 */
 -                    gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
 -                    /* fall through */
 -                case 0x3: /* SQDMLAL, SQDMLAL2 */
 -                    gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
 -                                                      tcg_res[pass],
 -                                                      tcg_passres);
 -                    break;
 -                default:
 -                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -                    g_assert_not_reached();
 -                }
 -            }
 -
 -            if (is_scalar) {
 -                tcg_gen_ext32u_i64(tcg_res[0], tcg_res[0]);
 -            }
 -        }
 -
 -        if (is_scalar) {
 -            tcg_res[1] = tcg_constant_i64(0);
 -        }
 -
 -        for (pass = 0; pass < 2; pass++) {
 -            write_vec_element(s, tcg_res[pass], rd, pass, MO_64);
 -        }
 -    }
 -}
 -
  /* C3.6 Data processing - SIMD, inc Crypto
   *
   * As the decode gets a little complex we are using a table based
@@ -XXX,XX +XXX,XX @@ static const AArch64DecodeTable data_proc_simd[] = {
      { 0x0e200000, 0x9f200c00, disas_simd_three_reg_diff },
      { 0x0e200800, 0x9f3e0c00, disas_simd_two_reg_misc },
      { 0x0e300800, 0x9f3e0c00, disas_simd_across_lanes },
 -    { 0x0f000000, 0x9f000400, disas_simd_indexed }, /* vector indexed */
      /* simd_mod_imm decode is a subset of simd_shift_imm, so must precede it */
      { 0x0f000400, 0x9ff80400, disas_simd_mod_imm },
      { 0x0f000400, 0x9f800400, disas_simd_shift_imm },
      { 0x0e000000, 0xbf208c00, disas_simd_tb },
      { 0x0e000800, 0xbf208c00, disas_simd_zip_trn },
      { 0x2e000000, 0xbf208400, disas_simd_ext },
 -    { 0x5e200000, 0xdf200c00, disas_simd_scalar_three_reg_diff },
      { 0x5e200800, 0xdf3e0c00, disas_simd_scalar_two_reg_misc },
 -    { 0x5f000000, 0xdf000400, disas_simd_indexed }, /* scalar indexed */
      { 0x5f000400, 0xdf800400, disas_simd_scalar_shift_imm },
      { 0x0e780800, 0x8f7e0c00, disas_simd_two_reg_misc_fp16 },
      { 0x00000000, 0x00000000, NULL }
 --
 .34.1

-[PULL 16/24] hw/misc: In STM32L4x5 EXTI, consolidate 2 constants
+[PULL 62/72] softfloat: Use goto for default nan case in pick_nan_muladd
-From: Inès Varhol <ines.varhol@telecom-paris.fr>
+From: Richard Henderson <richard.henderson@linaro.org>
-Up until now, the EXTI implementation had 16 inbound GPIOs connected to
+Remove "3" as a special case for which and simply
-the 16 outbound GPIOs of STM32L4x5 SYSCFG.
+branch to return the desired value.
 The EXTI actually handles 40 lines (namely 5 from STM32L4x5 USART
 devices which are already implemented in QEMU).
 In order to connect USART devices to EXTI, this commit consolidates
 constants `EXTI_NUM_INTERRUPT_OUT_LINES` (40) and
 `EXTI_NUM_GPIO_EVENT_IN_LINES` (16) into `EXTI_NUM_LINES` (40).
-Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240707085927.122867-2-ines.varhol@telecom-paris.fr
+Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/hw/misc/stm32l4x5_exti.h | 4 ++--
+ fpu/softfloat-parts.c.inc | 20 ++++++++++----------
- hw/misc/stm32l4x5_exti.c         | 6 ++----
+file changed, 10 insertions(+), 10 deletions(-)
 files changed, 4 insertions(+), 6 deletions(-)
-diff --git a/include/hw/misc/stm32l4x5_exti.h b/include/hw/misc/stm32l4x5_exti.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/misc/stm32l4x5_exti.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/include/hw/misc/stm32l4x5_exti.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
- #define TYPE_STM32L4X5_EXTI "stm32l4x5-exti"
+          * But if we're not in default-NaN mode then the target must
- OBJECT_DECLARE_SIMPLE_TYPE(Stm32l4x5ExtiState, STM32L4X5_EXTI)
+          * specify.
+          */
--#define EXTI_NUM_INTERRUPT_OUT_LINES 40
+-        which = 3;
-+#define EXTI_NUM_LINES 40
++        goto default_nan;
- #define EXTI_NUM_REGISTER 2
+     } else if (infzero) {
+         /*
- struct Stm32l4x5ExtiState {
+          * Inf * 0 + NaN -- some implementations return the
-@@ -XXX,XX +XXX,XX @@ struct Stm32l4x5ExtiState {
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+          */
-     /* used for edge detection */
+         switch (s->float_infzeronan_rule) {
-     uint32_t irq_levels[EXTI_NUM_REGISTER];
+         case float_infzeronan_dnan_never:
--    qemu_irq irq[EXTI_NUM_INTERRUPT_OUT_LINES];
+-            which = 2;
-+    qemu_irq irq[EXTI_NUM_LINES];
+             break;
- };
+         case float_infzeronan_dnan_always:
+-            which = 3;
- #endif
+-            break;
-diff --git a/hw/misc/stm32l4x5_exti.c b/hw/misc/stm32l4x5_exti.c
++            goto default_nan;
-index XXXXXXX..XXXXXXX 100644
+         case float_infzeronan_dnan_if_qnan:
---- a/hw/misc/stm32l4x5_exti.c
+-            which = is_qnan(c->cls) ? 3 : 2;
-+++ b/hw/misc/stm32l4x5_exti.c
++            if (is_qnan(c->cls)) {
-@@ -XXX,XX +XXX,XX @@
++                goto default_nan;
- #define EXTI_SWIER2 0x30
++            }
- #define EXTI_PR2    0x34
+             break;
+         default:
--#define EXTI_NUM_GPIO_EVENT_IN_LINES 16
+             g_assert_not_reached();
- #define EXTI_MAX_IRQ_PER_BANK 32
+         }
- #define EXTI_IRQS_BANK0  32
++        which = 2;
- #define EXTI_IRQS_BANK1  8
+     } else {
-@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_init(Object *obj)
+         FloatClass cls[3] = { a->cls, b->cls, c->cls };
- {
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
-     Stm32l4x5ExtiState *s = STM32L4X5_EXTI(obj);
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
+         }
 -    for (size_t i = 0; i < EXTI_NUM_INTERRUPT_OUT_LINES; i++) {
 +    for (size_t i = 0; i < EXTI_NUM_LINES; i++) {
          sysbus_init_irq(SYS_BUS_DEVICE(obj), &s->irq[i]);
      }
-@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_init(Object *obj)
+-    if (which == 3) {
-                           TYPE_STM32L4X5_EXTI, 0x400);
+-        parts_default_nan(a, s);
-     sysbus_init_mmio(SYS_BUS_DEVICE(obj), &s->mmio);
+-        return a;
+-    }
--    qdev_init_gpio_in(DEVICE(obj), stm32l4x5_exti_set_irq,
+-
--                      EXTI_NUM_GPIO_EVENT_IN_LINES);
+     switch (which) {
-+    qdev_init_gpio_in(DEVICE(obj), stm32l4x5_exti_set_irq, EXTI_NUM_LINES);
+     case 0:
          break;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          parts_silence_nan(a, s);
      }
      return a;
 +
 + default_nan:
 +    parts_default_nan(a, s);
 +    return a;
  }
- static const VMStateDescription vmstate_stm32l4x5_exti = {
+ /*
 --
 .34.1

-[PULL 18/24] hw/arm: In STM32L4x5 SOC, connect USART devices to EXTI
+[PULL 63/72] softfloat: Remove which from parts_pick_nan_muladd
-From: Inès Varhol <ines.varhol@telecom-paris.fr>
+From: Richard Henderson <richard.henderson@linaro.org>
-The USART devices were previously connecting their outbound IRQs
+Assign the pointer return value to 'a' directly,
-directly to the CPU because the EXTI wasn't handling direct lines
+rather than going through an intermediary index.
 interrupts.
 Now the USART connects to the EXTI inbound GPIOs, and the EXTI connects
 its IRQs to the CPU.
 The existing QTest for the USART (tests/qtest/stm32l4x5_usart-test.c)
 checks that USART1_IRQ in the CPU is pending when expected so it
 confirms that the connection through the EXTI still works.
-Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240707085927.122867-4-ines.varhol@telecom-paris.fr
+Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/stm32l4x5_soc.c | 24 +++++++++++-------------
+ fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
-file changed, 11 insertions(+), 13 deletions(-)
+file changed, 10 insertions(+), 22 deletions(-)
-diff --git a/hw/arm/stm32l4x5_soc.c b/hw/arm/stm32l4x5_soc.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/stm32l4x5_soc.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/arm/stm32l4x5_soc.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ static const int exti_irq[NUM_EXTI_IRQ] = {
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
- #define RCC_BASE_ADDRESS 0x40021000
+                                             FloatPartsN *c, float_status *s,
- #define RCC_IRQ 5
+                                             int ab_mask, int abc_mask)
 +#define EXTI_USART1_IRQ 26
 +#define EXTI_UART4_IRQ 29
 +#define EXTI_LPUART1_IRQ 31
 +
  static const int exti_or_gates_out[NUM_EXTI_OR_GATES] = {
 , 40, 63, 1,
  };
@@ -XXX,XX +XXX,XX @@ static const hwaddr uart_addr[] = {
  #define LPUART_BASE_ADDRESS 0x40008000
 -static const int usart_irq[] = { 37, 38, 39 };
 -static const int uart_irq[] = { 52, 53 };
 -#define LPUART_IRQ 70
 -
  static void stm32l4x5_soc_initfn(Object *obj)
  {
-     Stm32l4x5SocState *s = STM32L4X5_SOC(obj);
+-    int which;
-@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
+     bool infzero = (ab_mask == float_cmask_infzero);
      bool have_snan = (abc_mask & float_cmask_snan);
 +    FloatPartsN *ret;
      if (unlikely(have_snan)) {
          float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          default:
              g_assert_not_reached();
          }
 -        which = 2;
 +        ret = c;
      } else {
 -        FloatClass cls[3] = { a->cls, b->cls, c->cls };
 +        FloatPartsN *val[3] = { a, b, c };
          Float3NaNPropRule rule = s->float_3nan_prop_rule;
          assert(rule != float_3nan_prop_none);
          if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
              /* We have at least one SNaN input and should prefer it */
              do {
 -                which = rule & R_3NAN_1ST_MASK;
 +                ret = val[rule & R_3NAN_1ST_MASK];
                  rule >>= R_3NAN_1ST_LENGTH;
 -            } while (!is_snan(cls[which]));
 +            } while (!is_snan(ret->cls));
          } else {
              do {
 -                which = rule & R_3NAN_1ST_MASK;
 +                ret = val[rule & R_3NAN_1ST_MASK];
                  rule >>= R_3NAN_1ST_LENGTH;
 -            } while (!is_nan(cls[which]));
 +            } while (!is_nan(ret->cls));
          }
      }
-+    /* Connect SYSCFG to EXTI */
+-    switch (which) {
-     for (unsigned i = 0; i < GPIO_NUM_PINS; i++) {
+-    case 0:
-         qdev_connect_gpio_out(DEVICE(&s->syscfg), i,
+-        break;
-                               qdev_get_gpio_in(DEVICE(&s->exti), i));
+-    case 1:
-@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
+-        a = b;
-             return;
+-        break;
-         }
+-    case 2:
-         sysbus_mmio_map(busdev, 0, usart_addr[i]);
+-        a = c;
--        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, usart_irq[i]));
+-        break;
-+        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
+-    default:
-+                                                       EXTI_USART1_IRQ + i));
+-        g_assert_not_reached();
 +    if (is_snan(ret->cls)) {
 +        parts_silence_nan(ret, s);
      }
+-    if (is_snan(a->cls)) {
--    /*
+-        parts_silence_nan(a, s);
--     * TODO: Connect the USARTs, UARTs and LPUART to the EXTI once the EXTI
+-    }
--     * can handle other gpio-in than the gpios. (e.g. Direct Lines for the
+-    return a;
--     * usarts)
++    return ret;
--     */
--
+  default_nan:
-     /* UART devices */
+     parts_default_nan(a, s);
      for (int i = 0; i < STM_NUM_UARTS; i++) {
          g_autofree char *name = g_strdup_printf("uart%d-out", STM_NUM_USARTS + i + 1);
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
              return;
          }
          sysbus_mmio_map(busdev, 0, uart_addr[i]);
 -        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, uart_irq[i]));
 +        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
 +                                                       EXTI_UART4_IRQ + i));
      }
      /* LPUART device*/
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
          return;
      }
      sysbus_mmio_map(busdev, 0, LPUART_BASE_ADDRESS);
 -    sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, LPUART_IRQ));
 +    sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
 +                                                   EXTI_LPUART1_IRQ));
      /* APB1 BUS */
      create_unimplemented_device("TIM2",      0x40000000, 0x400);
 --
 .34.1

-[PULL 12/24] target/arm: Use cpu_env in cpu_untagged_addr
+[PULL 64/72] softfloat: Pad array size in pick_nan_muladd
 From: Richard Henderson <richard.henderson@linaro.org>
-In a completely artifical memset benchmark object_dynamic_cast_assert
+While all indices into val[] should be in [0-2], the mask
-dominates the profile, even above guest address resolution and
+applied is two bits.  To help static analysis see there is
-the underlying host memset.
+no possibility of read beyond the end of the array, pad the
 array to 4 entries, with the final being (implicitly) NULL.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240702154911.1667418-1-richard.henderson@linaro.org
+Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/cpu.h | 4 ++--
+ fpu/softfloat-parts.c.inc | 2 +-
-file changed, 2 insertions(+), 2 deletions(-)
+file changed, 1 insertion(+), 1 deletion(-)
-diff --git a/target/arm/cpu.h b/target/arm/cpu.h
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/cpu.h
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/cpu.h
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ extern const uint64_t pred_esz_masks[5];
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-  */
+         }
- static inline target_ulong cpu_untagged_addr(CPUState *cs, target_ulong x)
+         ret = c;
- {
+     } else {
--    ARMCPU *cpu = ARM_CPU(cs);
+-        FloatPartsN *val[3] = { a, b, c };
--    if (cpu->env.tagged_addr_enable) {
++        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
-+    CPUARMState *env = cpu_env(cs);
+         Float3NaNPropRule rule = s->float_3nan_prop_rule;
-+    if (env->tagged_addr_enable) {
-         /*
+         assert(rule != float_3nan_prop_none);
           * TBI is enabled for userspace but not kernelspace addresses.
           * Only clear the tag if bit 55 is clear.
 --
 .34.1

-[PULL 22/24] target/arm: Convert SADDW, SSUBW, UADDW, USUBW to decodetree
+[PULL 65/72] softfloat: Move propagateFloatx80NaN to softfloat.c
 From: Richard Henderson <richard.henderson@linaro.org>
+This function is part of the public interface and
+is not "specialized" to any target in any way.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20240709000610.382391-5-richard.henderson@linaro.org
+Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |  5 ++
+ fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
- target/arm/tcg/translate-a64.c | 86 +++++++++++++++++-----------------
+ fpu/softfloat-specialize.c.inc | 52 ----------------------------------
-files changed, 48 insertions(+), 43 deletions(-)
+files changed, 52 insertions(+), 52 deletions(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/fpu/softfloat.c
-+++ b/target/arm/tcg/a64.decode
++++ b/fpu/softfloat.c
-@@ -XXX,XX +XXX,XX @@ SQDMLAL_v       0.00 1110 101 ..... 10010 0 ..... ..... @qrrr_s
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
- SQDMLSL_v       0.00 1110 011 ..... 10110 0 ..... ..... @qrrr_h
+     *zExpPtr = 1 - shiftCount;
- SQDMLSL_v       0.00 1110 101 ..... 10110 0 ..... ..... @qrrr_s
+ }
-+SADDW           0.00 1110 ..1 ..... 00010 0 ..... ..... @qrrr_e
++/*----------------------------------------------------------------------------
-+UADDW           0.10 1110 ..1 ..... 00010 0 ..... ..... @qrrr_e
++| Takes two extended double-precision floating-point values `a' and `b', one
-+SSUBW           0.00 1110 ..1 ..... 00110 0 ..... ..... @qrrr_e
++| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
-+USUBW           0.10 1110 ..1 ..... 00110 0 ..... ..... @qrrr_e
++| `b' is a signaling NaN, the invalid exception is raised.
 +*----------------------------------------------------------------------------*/
 +
- ### Advanced SIMD scalar x indexed element
++floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
  FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
 diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/tcg/translate-a64.c
 +++ b/target/arm/tcg/translate-a64.c
@@ -XXX,XX +XXX,XX @@ TRANS(SQDMLSL_vi, do_3op_widening,
        a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
        a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
 +static bool do_addsub_wide(DisasContext *s, arg_qrrr_e *a,
 +                           MemOp sign, bool sub)
 +{
-+    TCGv_i64 tcg_op0, tcg_op1;
++    bool aIsLargerSignificand;
-+    MemOp esz = a->esz;
++    FloatClass a_cls, b_cls;
 +    int half = 8 >> esz;
 +    bool top = a->q;
 +    int top_swap = top ? 0 : half - 1;
 +    int top_half = top ? half : 0;
 +
-+    /* There are no 64x64->128 bit operations. */
++    /* This is not complete, but is good enough for pickNaN.  */
-+    if (esz >= MO_64) {
++    a_cls = (!floatx80_is_any_nan(a)
-+        return false;
++             ? float_class_normal
 +             : floatx80_is_signaling_nan(a, status)
 +             ? float_class_snan
 +             : float_class_qnan);
 +    b_cls = (!floatx80_is_any_nan(b)
 +             ? float_class_normal
 +             : floatx80_is_signaling_nan(b, status)
 +             ? float_class_snan
 +             : float_class_qnan);
 +
 +    if (is_snan(a_cls) || is_snan(b_cls)) {
 +        float_raise(float_flag_invalid, status);
 +    }
-+    if (!fp_access_check(s)) {
++
-+        return true;
++    if (status->default_nan_mode) {
 +        return floatx80_default_nan(status);
 +    }
-+    tcg_op0 = tcg_temp_new_i64();
-+    tcg_op1 = tcg_temp_new_i64();
 +
-+    for (int elt_fwd = 0; elt_fwd < half; ++elt_fwd) {
++    if (a.low < b.low) {
-+        int elt = elt_fwd ^ top_swap;
++        aIsLargerSignificand = 0;
 +    } else if (b.low < a.low) {
 +        aIsLargerSignificand = 1;
 +    } else {
 +        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
 +    }
 +
-+        read_vec_element(s, tcg_op1, a->rm, elt + top_half, esz | sign);
++    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-+        read_vec_element(s, tcg_op0, a->rn, elt, esz + 1);
++        if (is_snan(b_cls)) {
-+        if (sub) {
++            return floatx80_silence_nan(b, status);
 +            tcg_gen_sub_i64(tcg_op0, tcg_op0, tcg_op1);
 +        } else {
 +            tcg_gen_add_i64(tcg_op0, tcg_op0, tcg_op1);
 +        }
-+        write_vec_element(s, tcg_op0, a->rd, elt, esz + 1);
++        return b;
 +    } else {
 +        if (is_snan(a_cls)) {
 +            return floatx80_silence_nan(a, status);
 +        }
 +        return a;
 +    }
-+    clear_vec_high(s, 1, a->rd);
-+    return true;
 +}
 +
-+TRANS(SADDW, do_addsub_wide, a, MO_SIGN, false)
+ /*----------------------------------------------------------------------------
-+TRANS(UADDW, do_addsub_wide, a, 0, false)
+ | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
-+TRANS(SSUBW, do_addsub_wide, a, MO_SIGN, true)
+ | and extended significand formed by the concatenation of `zSig0' and `zSig1',
-+TRANS(USUBW, do_addsub_wide, a, 0, true)
+diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
-+
+index XXXXXXX..XXXXXXX 100644
- /*
+--- a/fpu/softfloat-specialize.c.inc
-  * Advanced SIMD scalar/vector x indexed element
++++ b/fpu/softfloat-specialize.c.inc
-  */
+@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
-@@ -XXX,XX +XXX,XX @@ static void gen_neon_addl(int size, bool is_sub, TCGv_i64 tcg_res,
+     return a;
      genfn(tcg_res, tcg_op1, tcg_op2);
  }
--static void handle_3rd_wide(DisasContext *s, int is_q, int is_u, int size,
+-/*----------------------------------------------------------------------------
--                            int opcode, int rd, int rn, int rm)
+-| Takes two extended double-precision floating-point values `a' and `b', one
 -| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
 -| `b' is a signaling NaN, the invalid exception is raised.
 -*----------------------------------------------------------------------------*/
 -
 -floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 -{
--    TCGv_i64 tcg_res[2];
+-    bool aIsLargerSignificand;
--    int part = is_q ? 2 : 0;
+-    FloatClass a_cls, b_cls;
 -    int pass;
 -
--    for (pass = 0; pass < 2; pass++) {
+-    /* This is not complete, but is good enough for pickNaN.  */
--        TCGv_i64 tcg_op1 = tcg_temp_new_i64();
+-    a_cls = (!floatx80_is_any_nan(a)
--        TCGv_i32 tcg_op2 = tcg_temp_new_i32();
+-             ? float_class_normal
--        TCGv_i64 tcg_op2_wide = tcg_temp_new_i64();
+-             : floatx80_is_signaling_nan(a, status)
--        static NeonGenWidenFn * const widenfns[3][2] = {
+-             ? float_class_snan
--            { gen_helper_neon_widen_s8, gen_helper_neon_widen_u8 },
+-             : float_class_qnan);
--            { gen_helper_neon_widen_s16, gen_helper_neon_widen_u16 },
+-    b_cls = (!floatx80_is_any_nan(b)
--            { tcg_gen_ext_i32_i64, tcg_gen_extu_i32_i64 },
+-             ? float_class_normal
--        };
+-             : floatx80_is_signaling_nan(b, status)
--        NeonGenWidenFn *widenfn = widenfns[size][is_u];
+-             ? float_class_snan
 -             : float_class_qnan);
 -
--        read_vec_element(s, tcg_op1, rn, pass, MO_64);
+-    if (is_snan(a_cls) || is_snan(b_cls)) {
--        read_vec_element_i32(s, tcg_op2, rm, part + pass, MO_32);
+-        float_raise(float_flag_invalid, status);
 -        widenfn(tcg_op2_wide, tcg_op2);
 -        tcg_res[pass] = tcg_temp_new_i64();
 -        gen_neon_addl(size, (opcode == 3),
 -                      tcg_res[pass], tcg_op1, tcg_op2_wide);
 -    }
 -
--    for (pass = 0; pass < 2; pass++) {
+-    if (status->default_nan_mode) {
--        write_vec_element(s, tcg_res[pass], rd, pass, MO_64);
+-        return floatx80_default_nan(status);
 -    }
 -
 -    if (a.low < b.low) {
 -        aIsLargerSignificand = 0;
 -    } else if (b.low < a.low) {
 -        aIsLargerSignificand = 1;
 -    } else {
 -        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
 -    }
 -
 -    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
 -        if (is_snan(b_cls)) {
 -            return floatx80_silence_nan(b, status);
 -        }
 -        return b;
 -    } else {
 -        if (is_snan(a_cls)) {
 -            return floatx80_silence_nan(a, status);
 -        }
 -        return a;
 -    }
 -}
 -
- static void do_narrow_round_high_u32(TCGv_i32 res, TCGv_i64 in)
+ /*----------------------------------------------------------------------------
- {
+ | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
-     tcg_gen_addi_i64(in, in, 1U << 31);
+ | NaN; otherwise returns 0.
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
      int rd = extract32(insn, 0, 5);
      switch (opcode) {
 -    case 1: /* SADDW, SADDW2, UADDW, UADDW2 */
 -    case 3: /* SSUBW, SSUBW2, USUBW, USUBW2 */
 -        /* 64 x 128 -> 128 */
 -        if (size == 3) {
 -            unallocated_encoding(s);
 -            return;
 -        }
 -        if (!fp_access_check(s)) {
 -            return;
 -        }
 -        handle_3rd_wide(s, is_q, is_u, size, opcode, rd, rn, rm);
 -        break;
      case 4: /* ADDHN, ADDHN2, RADDHN, RADDHN2 */
      case 6: /* SUBHN, SUBHN2, RSUBHN, RSUBHN2 */
          /* 128 x 128 -> 64 */
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
          return;
      default:
      case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 +    case 1: /* SADDW, SADDW2, UADDW, UADDW2 */
      case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 +    case 3: /* SSUBW, SSUBW2, USUBW, USUBW2 */
      case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
      case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
      case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 --
 .34.1

-New patch
+[PULL 66/72] softfloat: Use parts_pick_nan in propagateFloatx80NaN
+From: Richard Henderson <richard.henderson@linaro.org>
+Unpacking and repacking the parts may be slightly more work
+than we did before, but we get to reuse more code.  For a
+code path handling exceptional values, this is an improvement.
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ fpu/softfloat.c | 43 +++++--------------------------------------
+file changed, 5 insertions(+), 38 deletions(-)
+diff --git a/fpu/softfloat.c b/fpu/softfloat.c
+index XXXXXXX..XXXXXXX 100644
+--- a/fpu/softfloat.c
++++ b/fpu/softfloat.c
+@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
+ floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
+ {
+-    bool aIsLargerSignificand;
+-    FloatClass a_cls, b_cls;
++    FloatParts128 pa, pb, *pr;
+-    /* This is not complete, but is good enough for pickNaN.  */
+-    a_cls = (!floatx80_is_any_nan(a)
+-             ? float_class_normal
+-             : floatx80_is_signaling_nan(a, status)
+-             ? float_class_snan
+-             : float_class_qnan);
+-    b_cls = (!floatx80_is_any_nan(b)
+-             ? float_class_normal
+-             : floatx80_is_signaling_nan(b, status)
+-             ? float_class_snan
+-             : float_class_qnan);
+-
+-    if (is_snan(a_cls) || is_snan(b_cls)) {
+-        float_raise(float_flag_invalid, status);
+-    }
+-
+-    if (status->default_nan_mode) {
++    if (!floatx80_unpack_canonical(&pa, a, status) ||
++        !floatx80_unpack_canonical(&pb, b, status)) {
+         return floatx80_default_nan(status);
+     }
+-    if (a.low < b.low) {
+-        aIsLargerSignificand = 0;
+-    } else if (b.low < a.low) {
+-        aIsLargerSignificand = 1;
+-    } else {
+-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
+-    }
+-
+-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
+-        if (is_snan(b_cls)) {
+-            return floatx80_silence_nan(b, status);
+-        }
+-        return b;
+-    } else {
+-        if (is_snan(a_cls)) {
+-            return floatx80_silence_nan(a, status);
+-        }
+-        return a;
+-    }
++    pr = parts_pick_nan(&pa, &pb, status);
++    return floatx80_round_pack_canonical(pr, status);
+ }
+ /*----------------------------------------------------------------------------
+--
+.34.1

-[PULL 24/24] target/arm: Convert PMULL to decodetree
+[PULL 67/72] softfloat: Inline pickNaN
 From: Richard Henderson <richard.henderson@linaro.org>
+Inline pickNaN into its only caller.  This makes one assert
+redundant with the immediately preceding IF.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
 Message-id: 20240709000610.382391-7-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |  3 ++
+ fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
- target/arm/tcg/translate-a64.c | 94 +++++-----------------------------
+ fpu/softfloat-specialize.c.inc | 96 ----------------------------------
-files changed, 15 insertions(+), 82 deletions(-)
+files changed, 73 insertions(+), 105 deletions(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/tcg/a64.decode
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ RADDHN          0.10 1110 ..1 ..... 01000 0 ..... ..... @qrrr_e
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
- SUBHN           0.00 1110 ..1 ..... 01100 0 ..... ..... @qrrr_e
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- RSUBHN          0.10 1110 ..1 ..... 01100 0 ..... ..... @qrrr_e
+                                      float_status *s)
+ {
-+PMULL_p8        0.00 1110 001 ..... 11100 0 ..... ..... @qrrr_b
++    int cmp, which;
 +PMULL_p64       0.00 1110 111 ..... 11100 0 ..... ..... @qrrr_b
 +
- ### Advanced SIMD scalar x indexed element
+     if (is_snan(a->cls) || is_snan(b->cls)) {
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
- FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
+     }
-diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
      if (s->default_nan_mode) {
          parts_default_nan(a, s);
 -    } else {
 -        int cmp = frac_cmp(a, b);
 -        if (cmp == 0) {
 -            cmp = a->sign < b->sign;
 -        }
 +        return a;
 +    }
 -        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
 -            a = b;
 -        }
 +    cmp = frac_cmp(a, b);
 +    if (cmp == 0) {
 +        cmp = a->sign < b->sign;
 +    }
 +
 +    switch (s->float_2nan_prop_rule) {
 +    case float_2nan_prop_s_ab:
          if (is_snan(a->cls)) {
 -            parts_silence_nan(a, s);
 +            which = 0;
 +        } else if (is_snan(b->cls)) {
 +            which = 1;
 +        } else if (is_qnan(a->cls)) {
 +            which = 0;
 +        } else {
 +            which = 1;
          }
 +        break;
 +    case float_2nan_prop_s_ba:
 +        if (is_snan(b->cls)) {
 +            which = 1;
 +        } else if (is_snan(a->cls)) {
 +            which = 0;
 +        } else if (is_qnan(b->cls)) {
 +            which = 1;
 +        } else {
 +            which = 0;
 +        }
 +        break;
 +    case float_2nan_prop_ab:
 +        which = is_nan(a->cls) ? 0 : 1;
 +        break;
 +    case float_2nan_prop_ba:
 +        which = is_nan(b->cls) ? 1 : 0;
 +        break;
 +    case float_2nan_prop_x87:
 +        /*
 +         * This implements x87 NaN propagation rules:
 +         * SNaN + QNaN => return the QNaN
 +         * two SNaNs => return the one with the larger significand, silenced
 +         * two QNaNs => return the one with the larger significand
 +         * SNaN and a non-NaN => return the SNaN, silenced
 +         * QNaN and a non-NaN => return the QNaN
 +         *
 +         * If we get down to comparing significands and they are the same,
 +         * return the NaN with the positive sign bit (if any).
 +         */
 +        if (is_snan(a->cls)) {
 +            if (is_snan(b->cls)) {
 +                which = cmp > 0 ? 0 : 1;
 +            } else {
 +                which = is_qnan(b->cls) ? 1 : 0;
 +            }
 +        } else if (is_qnan(a->cls)) {
 +            if (is_snan(b->cls) || !is_qnan(b->cls)) {
 +                which = 0;
 +            } else {
 +                which = cmp > 0 ? 0 : 1;
 +            }
 +        } else {
 +            which = 1;
 +        }
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +
 +    if (which) {
 +        a = b;
 +    }
 +    if (is_snan(a->cls)) {
 +        parts_silence_nan(a, s);
      }
      return a;
  }
 diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/translate-a64.c
+--- a/fpu/softfloat-specialize.c.inc
-+++ b/target/arm/tcg/translate-a64.c
++++ b/fpu/softfloat-specialize.c.inc
-@@ -XXX,XX +XXX,XX @@ TRANS(SUBHN, do_addsub_highnarrow, a, true, false)
+@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
  TRANS(RADDHN, do_addsub_highnarrow, a, false, true)
  TRANS(RSUBHN, do_addsub_highnarrow, a, true, true)
 +static bool do_pmull(DisasContext *s, arg_qrrr_e *a, gen_helper_gvec_3 *fn)
 +{
 +    if (fp_access_check(s)) {
 +        /* The Q field specifies lo/hi half input for these insns.  */
 +        gen_gvec_op3_ool(s, true, a->rd, a->rn, a->rm, a->q, fn);
 +    }
 +    return true;
 +}
 +
 +TRANS(PMULL_p8, do_pmull, a, gen_helper_neon_pmull_h)
 +TRANS_FEAT(PMULL_p64, aa64_pmull, do_pmull, a, gen_helper_gvec_pmull_q)
 +
  /*
   * Advanced SIMD scalar/vector x indexed element
   */
@@ -XXX,XX +XXX,XX @@ static void disas_simd_shift_imm(DisasContext *s, uint32_t insn)
      }
  }
--/* AdvSIMD three different
+-/*----------------------------------------------------------------------------
-- *   31  30  29 28       24 23  22  21 20  16 15    12 11 10 9    5 4    0
+-| Select which NaN to propagate for a two-input operation.
-- * +---+---+---+-----------+------+---+------+--------+-----+------+------+
+-| IEEE754 doesn't specify all the details of this, so the
-- * | 0 | Q | U | 0 1 1 1 0 | size | 1 |  Rm  | opcode | 0 0 |  Rn  |  Rd  |
+-| algorithm is target-specific.
-- * +---+---+---+-----------+------+---+------+--------+-----+------+------+
+-| The routine is passed various bits of information about the
-- */
+-| two NaNs and should return 0 to select NaN a and 1 for NaN b.
--static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
+-| Note that signalling NaNs are always squashed to quiet NaNs
 -| by the caller, by calling floatXX_silence_nan() before
 -| returning them.
 -|
 -| aIsLargerSignificand is only valid if both a and b are NaNs
 -| of some kind, and is true if a has the larger significand,
 -| or if both a and b have the same significand but a is
 -| positive but b is negative. It is only needed for the x87
 -| tie-break rule.
 -*----------------------------------------------------------------------------*/
 -
 -static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 -                   bool aIsLargerSignificand, float_status *status)
 -{
--    /* Instructions in this group fall into three basic classes
+-    /*
--     * (in each case with the operation working on each element in
+-     * We guarantee not to require the target to tell us how to
--     * the input vectors):
+-     * pick a NaN if we're always returning the default NaN.
--     * (1) widening 64 x 64 -> 128 (with possibly Vd as an extra
+-     * But if we're not in default-NaN mode then the target must
--     *     128 bit input)
+-     * specify via set_float_2nan_prop_rule().
 -     * (2) wide 64 x 128 -> 128
 -     * (3) narrowing 128 x 128 -> 64
 -     * Here we do initial decode, catch unallocated cases and
 -     * dispatch to separate functions for each class.
 -     */
--    int is_q = extract32(insn, 30, 1);
+-    assert(!status->default_nan_mode);
 -    int is_u = extract32(insn, 29, 1);
 -    int size = extract32(insn, 22, 2);
 -    int opcode = extract32(insn, 12, 4);
 -    int rm = extract32(insn, 16, 5);
 -    int rn = extract32(insn, 5, 5);
 -    int rd = extract32(insn, 0, 5);
 -
--    switch (opcode) {
+-    switch (status->float_2nan_prop_rule) {
--    case 14: /* PMULL, PMULL2 */
+-    case float_2nan_prop_s_ab:
--        if (is_u) {
+-        if (is_snan(a_cls)) {
--            unallocated_encoding(s);
+-            return 0;
--            return;
+-        } else if (is_snan(b_cls)) {
--        }
+-            return 1;
--        switch (size) {
+-        } else if (is_qnan(a_cls)) {
--        case 0: /* PMULL.P8 */
+-            return 0;
--            if (!fp_access_check(s)) {
+-        } else {
--                return;
+-            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_s_ba:
 -        if (is_snan(b_cls)) {
 -            return 1;
 -        } else if (is_snan(a_cls)) {
 -            return 0;
 -        } else if (is_qnan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_ab:
 -        if (is_nan(a_cls)) {
 -            return 0;
 -        } else {
 -            return 1;
 -        }
 -        break;
 -    case float_2nan_prop_ba:
 -        if (is_nan(b_cls)) {
 -            return 1;
 -        } else {
 -            return 0;
 -        }
 -        break;
 -    case float_2nan_prop_x87:
 -        /*
 -         * This implements x87 NaN propagation rules:
 -         * SNaN + QNaN => return the QNaN
 -         * two SNaNs => return the one with the larger significand, silenced
 -         * two QNaNs => return the one with the larger significand
 -         * SNaN and a non-NaN => return the SNaN, silenced
 -         * QNaN and a non-NaN => return the QNaN
 -         *
 -         * If we get down to comparing significands and they are the same,
 -         * return the NaN with the positive sign bit (if any).
 -         */
 -        if (is_snan(a_cls)) {
 -            if (is_snan(b_cls)) {
 -                return aIsLargerSignificand ? 0 : 1;
 -            }
--            /* The Q field specifies lo/hi half input for this insn.  */
+-            return is_qnan(b_cls) ? 1 : 0;
--            gen_gvec_op3_ool(s, true, rd, rn, rm, is_q,
+-        } else if (is_qnan(a_cls)) {
--                             gen_helper_neon_pmull_h);
+-            if (is_snan(b_cls) || !is_qnan(b_cls)) {
--            break;
+-                return 0;
--
+-            } else {
--        case 3: /* PMULL.P64 */
+-                return aIsLargerSignificand ? 0 : 1;
 -            if (!dc_isar_feature(aa64_pmull, s)) {
 -                unallocated_encoding(s);
 -                return;
 -            }
--            if (!fp_access_check(s)) {
+-        } else {
--                return;
+-            return 1;
--            }
+-        }
 -            /* The Q field specifies lo/hi half input for this insn.  */
 -            gen_gvec_op3_ool(s, true, rd, rn, rm, is_q,
 -                             gen_helper_gvec_pmull_q);
 -            break;
 -
 -        default:
 -            unallocated_encoding(s);
 -            break;
 -        }
 -        return;
 -    default:
--    case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
+-        g_assert_not_reached();
 -    case 1: /* SADDW, SADDW2, UADDW, UADDW2 */
 -    case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 -    case 3: /* SSUBW, SSUBW2, USUBW, USUBW2 */
 -    case 4: /* ADDHN, ADDHN2, RADDHN, RADDHN2 */
 -    case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -    case 6: /* SUBHN, SUBHN2, RSUBHN, RSUBHN2 */
 -    case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -    case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -    case 9: /* SQDMLAL, SQDMLAL2 */
 -    case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -    case 11: /* SQDMLSL, SQDMLSL2 */
 -    case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
 -    case 13: /* SQDMULL, SQDMULL2 */
 -        /* opcode 15 not allocated */
 -        unallocated_encoding(s);
 -        break;
 -    }
 -}
 -
- static void handle_2misc_widening(DisasContext *s, int opcode, bool is_q,
+ /*----------------------------------------------------------------------------
-                                   int size, int rn, int rd)
+ | Returns 1 if the double-precision floating-point value `a' is a quiet
- {
+ | NaN; otherwise returns 0.
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc_fp16(DisasContext *s, uint32_t insn)
   */
  static const AArch64DecodeTable data_proc_simd[] = {
      /* pattern  ,  mask     ,  fn                        */
 -    { 0x0e200000, 0x9f200c00, disas_simd_three_reg_diff },
      { 0x0e200800, 0x9f3e0c00, disas_simd_two_reg_misc },
      { 0x0e300800, 0x9f3e0c00, disas_simd_across_lanes },
      /* simd_mod_imm decode is a subset of simd_shift_imm, so must precede it */
 --
 .34.1

-[PULL 19/24] target/arm: Convert SMULL, UMULL, SMLAL, UMLAL, SMLSL, UMLSL to decodetree
+[PULL 68/72] softfloat: Share code between parts_pick_nan cases
 From: Richard Henderson <richard.henderson@linaro.org>
+Remember if there was an SNaN, and use that to simplify
+float_2nan_prop_s_{ab,ba} to only the snan component.
+Then, fall through to the corresponding
+float_2nan_prop_{ab,ba} case to handle any remaining
+nans, which must be quiet.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20240709000610.382391-2-richard.henderson@linaro.org
+Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |  22 ++++
+ fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
- target/arm/tcg/translate-a64.c | 184 ++++++++++++++++++++++++---------
+file changed, 12 insertions(+), 20 deletions(-)
 files changed, 156 insertions(+), 50 deletions(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/tcg/a64.decode
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ FCADD_270       0.10 1110 ..0 ..... 11110 1 ..... ..... @qrrr_e
+@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
+ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- FCMLA_v         0 q:1 10 1110 esz:2 0 rm:5 110 rot:2 1 rn:5 rd:5
+                                      float_status *s)
+ {
-+SMULL_v         0.00 1110 ..1 ..... 11000 0 ..... ..... @qrrr_e
++    bool have_snan = false;
-+UMULL_v         0.10 1110 ..1 ..... 11000 0 ..... ..... @qrrr_e
+     int cmp, which;
-+SMLAL_v         0.00 1110 ..1 ..... 10000 0 ..... ..... @qrrr_e
-+UMLAL_v         0.10 1110 ..1 ..... 10000 0 ..... ..... @qrrr_e
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+SMLSL_v         0.00 1110 ..1 ..... 10100 0 ..... ..... @qrrr_e
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
-+UMLSL_v         0.10 1110 ..1 ..... 10100 0 ..... ..... @qrrr_e
++        have_snan = true;
-+
+     }
- ### Advanced SIMD scalar x indexed element
+     if (s->default_nan_mode) {
- FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-@@ -XXX,XX +XXX,XX @@ FCMLA_vi        0 0 10 1111 01 idx:1 rm:5 0 rot:2 1 0 0 rn:5 rd:5 esz=1 q=0
- FCMLA_vi        0 1 10 1111 01 . rm:5 0 rot:2 1 . 0 rn:5 rd:5 esz=1 idx=%hl q=1
+     switch (s->float_2nan_prop_rule) {
- FCMLA_vi        0 1 10 1111 10 0 rm:5 0 rot:2 1 idx:1 0 rn:5 rd:5 esz=2 q=1
+     case float_2nan_prop_s_ab:
+-        if (is_snan(a->cls)) {
-+SMULL_vi        0.00 1111 01 .. .... 1010 . 0 ..... .....   @qrrx_h
+-            which = 0;
-+SMULL_vi        0.00 1111 10 . ..... 1010 . 0 ..... .....   @qrrx_s
+-        } else if (is_snan(b->cls)) {
-+UMULL_vi        0.10 1111 01 .. .... 1010 . 0 ..... .....   @qrrx_h
+-            which = 1;
-+UMULL_vi        0.10 1111 10 . ..... 1010 . 0 ..... .....   @qrrx_s
+-        } else if (is_qnan(a->cls)) {
-+
+-            which = 0;
-+SMLAL_vi        0.00 1111 01 .. .... 0010 . 0 ..... .....   @qrrx_h
+-        } else {
-+SMLAL_vi        0.00 1111 10 . ..... 0010 . 0 ..... .....   @qrrx_s
+-            which = 1;
-+UMLAL_vi        0.10 1111 01 .. .... 0010 . 0 ..... .....   @qrrx_h
++        if (have_snan) {
-+UMLAL_vi        0.10 1111 10 . ..... 0010 . 0 ..... .....   @qrrx_s
++            which = is_snan(a->cls) ? 0 : 1;
-+
++            break;
-+SMLSL_vi        0.00 1111 01 .. .... 0110 . 0 ..... .....   @qrrx_h
+         }
-+SMLSL_vi        0.00 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
+-        break;
-+UMLSL_vi        0.10 1111 01 .. .... 0110 . 0 ..... .....   @qrrx_h
+-    case float_2nan_prop_s_ba:
-+UMLSL_vi        0.10 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
+-        if (is_snan(b->cls)) {
-+
+-            which = 1;
- # Floating-point conditional select
+-        } else if (is_snan(a->cls)) {
+-            which = 0;
- FCSEL           0001 1110 .. 1 rm:5 cond:4 11 rn:5 rd:5     esz=%esz_hsd
+-        } else if (is_qnan(b->cls)) {
-diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
+-            which = 1;
-index XXXXXXX..XXXXXXX 100644
+-        } else {
---- a/target/arm/tcg/translate-a64.c
+-            which = 0;
 +++ b/target/arm/tcg/translate-a64.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FCMLA_v(DisasContext *s, arg_FCMLA_v *a)
      return true;
  }
 +/*
 + * Widening vector x vector/indexed.
 + *
 + * These read from the top or bottom half of a 128-bit vector.
 + * After widening, optionally accumulate with a 128-bit vector.
 + * Implement these inline, as the number of elements are limited
 + * and the related SVE and SME operations on larger vectors use
 + * even/odd elements instead of top/bottom half.
 + *
 + * If idx >= 0, operand 2 is indexed, otherwise vector.
 + * If acc, operand 0 is loaded with rd.
 + */
 +
 +/* For low half, iterating up. */
 +static bool do_3op_widening(DisasContext *s, MemOp memop, int top,
 +                            int rd, int rn, int rm, int idx,
 +                            NeonGenTwo64OpFn *fn, bool acc)
 +{
 +    TCGv_i64 tcg_op0 = tcg_temp_new_i64();
 +    TCGv_i64 tcg_op1 = tcg_temp_new_i64();
 +    TCGv_i64 tcg_op2 = tcg_temp_new_i64();
 +    MemOp esz = memop & MO_SIZE;
 +    int half = 8 >> esz;
 +    int top_swap, top_half;
 +
 +    /* There are no 64x64->128 bit operations. */
 +    if (esz >= MO_64) {
 +        return false;
 +    }
 +    if (!fp_access_check(s)) {
 +        return true;
 +    }
 +
 +    if (idx >= 0) {
 +        read_vec_element(s, tcg_op2, rm, idx, memop);
 +    }
 +
 +    /*
 +     * For top half inputs, iterate forward; backward for bottom half.
 +     * This means the store to the destination will not occur until
 +     * overlapping input inputs are consumed.
 +     * Use top_swap to conditionally invert the forward iteration index.
 +     */
 +    top_swap = top ? 0 : half - 1;
 +    top_half = top ? half : 0;
 +
 +    for (int elt_fwd = 0; elt_fwd < half; ++elt_fwd) {
 +        int elt = elt_fwd ^ top_swap;
 +
 +        read_vec_element(s, tcg_op1, rn, elt + top_half, memop);
 +        if (idx < 0) {
 +            read_vec_element(s, tcg_op2, rm, elt + top_half, memop);
 +        }
 +        if (acc) {
 +            read_vec_element(s, tcg_op0, rd, elt, memop + 1);
 +        }
 +        fn(tcg_op0, tcg_op1, tcg_op2);
 +        write_vec_element(s, tcg_op0, rd, elt, esz + 1);
 +    }
 +    clear_vec_high(s, 1, rd);
 +    return true;
 +}
 +
 +static void gen_muladd_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +    tcg_gen_mul_i64(t, n, m);
 +    tcg_gen_add_i64(d, d, t);
 +}
 +
 +static void gen_mulsub_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +    tcg_gen_mul_i64(t, n, m);
 +    tcg_gen_sub_i64(d, d, t);
 +}
 +
 +TRANS(SMULL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_mul_i64, false)
 +TRANS(UMULL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_mul_i64, false)
 +TRANS(SMLAL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_muladd_i64, true)
 +TRANS(UMLAL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_muladd_i64, true)
 +TRANS(SMLSL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_mulsub_i64, true)
 +TRANS(UMLSL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_mulsub_i64, true)
 +
 +TRANS(SMULL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      tcg_gen_mul_i64, false)
 +TRANS(UMULL_vi, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, a->idx,
 +      tcg_gen_mul_i64, false)
 +TRANS(SMLAL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      gen_muladd_i64, true)
 +TRANS(UMLAL_vi, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, a->idx,
 +      gen_muladd_i64, true)
 +TRANS(SMLSL_vi, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
 +      gen_mulsub_i64, true)
 +TRANS(UMLSL_vi, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, a->idx,
 +      gen_mulsub_i64, true)
 +
  /*
   * Advanced SIMD scalar/vector x indexed element
   */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
                                      tcg_op1, tcg_op2, tcg_tmp1, tcg_tmp2);
                  break;
              }
 -            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 -                tcg_gen_mul_i64(tcg_passres, tcg_op1, tcg_op2);
 -                break;
              case 9: /* SQDMLAL, SQDMLAL2 */
              case 11: /* SQDMLSL, SQDMLSL2 */
              case 13: /* SQDMULL, SQDMULL2 */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
                                                    tcg_passres, tcg_passres);
                  break;
              default:
 +            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 +            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
                  g_assert_not_reached();
              }
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
                      }
                  }
                  break;
 -            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 -                if (size == 0) {
 -                    if (is_u) {
 -                        gen_helper_neon_mull_u8(tcg_passres, tcg_op1, tcg_op2);
 -                    } else {
 -                        gen_helper_neon_mull_s8(tcg_passres, tcg_op1, tcg_op2);
 -                    }
 -                } else {
 -                    if (is_u) {
 -                        gen_helper_neon_mull_u16(tcg_passres, tcg_op1, tcg_op2);
 -                    } else {
 -                        gen_helper_neon_mull_s16(tcg_passres, tcg_op1, tcg_op2);
 -                    }
 -                }
 -                break;
              case 9: /* SQDMLAL, SQDMLAL2 */
              case 11: /* SQDMLSL, SQDMLSL2 */
              case 13: /* SQDMULL, SQDMULL2 */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
                                                    tcg_passres, tcg_passres);
                  break;
              default:
 +            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 +            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
                  g_assert_not_reached();
              }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
      case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
      case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
      case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -    case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -    case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -    case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
          /* 64 x 64 -> 128 */
          if (size == 3) {
              unallocated_encoding(s);
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
          handle_3rd_widening(s, is_q, is_u, size, opcode, rd, rn, rm);
          break;
      default:
 +    case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +    case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 +    case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
          /* opcode 15 not allocated */
          unallocated_encoding(s);
          break;
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
      int index;
      switch (16 * u + opcode) {
 -    case 0x02: /* SMLAL, SMLAL2 */
 -    case 0x12: /* UMLAL, UMLAL2 */
 -    case 0x06: /* SMLSL, SMLSL2 */
 -    case 0x16: /* UMLSL, UMLSL2 */
 -    case 0x0a: /* SMULL, SMULL2 */
 -    case 0x1a: /* UMULL, UMULL2 */
 -        if (is_scalar) {
 -            unallocated_encoding(s);
 -            return;
 -        }
 -        break;
-     case 0x03: /* SQDMLAL, SQDMLAL2 */
++        /* fall through */
-     case 0x07: /* SQDMLSL, SQDMLSL2 */
+     case float_2nan_prop_ab:
-     case 0x0b: /* SQDMULL, SQDMULL2 */
+         which = is_nan(a->cls) ? 0 : 1;
-@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
+         break;
-     default:
++    case float_2nan_prop_s_ba:
-     case 0x00: /* FMLAL */
++        if (have_snan) {
-     case 0x01: /* FMLA */
++            which = is_snan(b->cls) ? 1 : 0;
-+    case 0x02: /* SMLAL, SMLAL2 */
++            break;
-     case 0x04: /* FMLSL */
++        }
-     case 0x05: /* FMLS */
++        /* fall through */
-+    case 0x06: /* SMLSL, SMLSL2 */
+     case float_2nan_prop_ba:
-     case 0x08: /* MUL */
+         which = is_nan(b->cls) ? 1 : 0;
-     case 0x09: /* FMUL */
+         break;
 +    case 0x0a: /* SMULL, SMULL2 */
      case 0x0c: /* SQDMULH */
      case 0x0d: /* SQRDMULH */
      case 0x0e: /* SDOT */
      case 0x0f: /* SUDOT / BFDOT / USDOT / BFMLAL */
      case 0x10: /* MLA */
      case 0x11: /* FCMLA #0 */
 +    case 0x12: /* UMLAL, UMLAL2 */
      case 0x13: /* FCMLA #90 */
      case 0x14: /* MLS */
      case 0x15: /* FCMLA #180 */
 +    case 0x16: /* UMLSL, UMLSL2 */
      case 0x17: /* FCMLA #270 */
      case 0x18: /* FMLAL2 */
      case 0x19: /* FMULX */
 +    case 0x1a: /* UMULL, UMULL2 */
      case 0x1c: /* FMLSL2 */
      case 0x1d: /* SQRDMLAH */
      case 0x1e: /* UDOT */
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
                  read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
                  switch (opcode) {
 -                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -                    tcg_gen_add_i64(tcg_res[pass], tcg_res[pass], tcg_passres);
 -                    break;
 -                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -                    tcg_gen_sub_i64(tcg_res[pass], tcg_res[pass], tcg_passres);
 -                    break;
                  case 0x7: /* SQDMLSL, SQDMLSL2 */
                      tcg_gen_neg_i64(tcg_passres, tcg_passres);
                      /* fall through */
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
                                                        tcg_passres);
                      break;
                  default:
 +                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
                      g_assert_not_reached();
                  }
              }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
                  read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
                  switch (opcode) {
 -                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 -                    gen_helper_neon_addl_u32(tcg_res[pass], tcg_res[pass],
 -                                             tcg_passres);
 -                    break;
 -                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
 -                    gen_helper_neon_subl_u32(tcg_res[pass], tcg_res[pass],
 -                                             tcg_passres);
 -                    break;
                  case 0x7: /* SQDMLSL, SQDMLSL2 */
                      gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
                      /* fall through */
@@ -XXX,XX +XXX,XX @@ static void disas_simd_indexed(DisasContext *s, uint32_t insn)
                                                        tcg_passres);
                      break;
                  default:
 +                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
 +                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
                      g_assert_not_reached();
                  }
              }
 --
 .34.1

-[PULL 20/24] target/arm: Convert SADDL, SSUBL, SABDL, SABAL, and unsigned to decodetree
+[PULL 69/72] softfloat: Sink frac_cmp in parts_pick_nan until needed
 From: Richard Henderson <richard.henderson@linaro.org>
+Move the fractional comparison to the end of the
+float_2nan_prop_x87 case.  This is not required for
+any other 2nan propagation rule.  Reorganize the
+x87 case itself to break out of the switch when the
+fractional comparison is not required.
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20240709000610.382391-3-richard.henderson@linaro.org
+Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/tcg/a64.decode      |   9 ++
+ fpu/softfloat-parts.c.inc | 19 +++++++++----------
- target/arm/tcg/translate-a64.c | 150 +++++++++++++++++----------------
+file changed, 9 insertions(+), 10 deletions(-)
 files changed, 87 insertions(+), 72 deletions(-)
-diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/tcg/a64.decode
+--- a/fpu/softfloat-parts.c.inc
-+++ b/target/arm/tcg/a64.decode
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ UMLAL_v         0.10 1110 ..1 ..... 10000 0 ..... ..... @qrrr_e
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- SMLSL_v         0.00 1110 ..1 ..... 10100 0 ..... ..... @qrrr_e
+         return a;
- UMLSL_v         0.10 1110 ..1 ..... 10100 0 ..... ..... @qrrr_e
+     }
-+SADDL_v         0.00 1110 ..1 ..... 00000 0 ..... ..... @qrrr_e
+-    cmp = frac_cmp(a, b);
-+UADDL_v         0.10 1110 ..1 ..... 00000 0 ..... ..... @qrrr_e
+-    if (cmp == 0) {
-+SSUBL_v         0.00 1110 ..1 ..... 00100 0 ..... ..... @qrrr_e
+-        cmp = a->sign < b->sign;
-+USUBL_v         0.10 1110 ..1 ..... 00100 0 ..... ..... @qrrr_e
+-    }
-+SABAL_v         0.00 1110 ..1 ..... 01010 0 ..... ..... @qrrr_e
+-
-+UABAL_v         0.10 1110 ..1 ..... 01010 0 ..... ..... @qrrr_e
+     switch (s->float_2nan_prop_rule) {
-+SABDL_v         0.00 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
+     case float_2nan_prop_s_ab:
-+UABDL_v         0.10 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
+         if (have_snan) {
-+
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- ### Advanced SIMD scalar x indexed element
+          * return the NaN with the positive sign bit (if any).
+          */
- FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
+         if (is_snan(a->cls)) {
-diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
+-            if (is_snan(b->cls)) {
-index XXXXXXX..XXXXXXX 100644
+-                which = cmp > 0 ? 0 : 1;
---- a/target/arm/tcg/translate-a64.c
+-            } else {
-+++ b/target/arm/tcg/translate-a64.c
++            if (!is_snan(b->cls)) {
-@@ -XXX,XX +XXX,XX @@ TRANS(UMLSL_vi, do_3op_widening,
+                 which = is_qnan(b->cls) ? 1 : 0;
-       a->esz, a->q, a->rd, a->rn, a->rm, a->idx,
++                break;
        gen_mulsub_i64, true)
 +static void gen_sabd_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t1 = tcg_temp_new_i64();
 +    TCGv_i64 t2 = tcg_temp_new_i64();
 +
 +    tcg_gen_sub_i64(t1, n, m);
 +    tcg_gen_sub_i64(t2, m, n);
 +    tcg_gen_movcond_i64(TCG_COND_GE, d, n, m, t1, t2);
 +}
 +
 +static void gen_uabd_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t1 = tcg_temp_new_i64();
 +    TCGv_i64 t2 = tcg_temp_new_i64();
 +
 +    tcg_gen_sub_i64(t1, n, m);
 +    tcg_gen_sub_i64(t2, m, n);
 +    tcg_gen_movcond_i64(TCG_COND_GEU, d, n, m, t1, t2);
 +}
 +
 +static void gen_saba_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +    gen_sabd_i64(t, n, m);
 +    tcg_gen_add_i64(d, d, t);
 +}
 +
 +static void gen_uaba_i64(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
 +{
 +    TCGv_i64 t = tcg_temp_new_i64();
 +    gen_uabd_i64(t, n, m);
 +    tcg_gen_add_i64(d, d, t);
 +}
 +
 +TRANS(SADDL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_add_i64, false)
 +TRANS(UADDL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_add_i64, false)
 +TRANS(SSUBL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_sub_i64, false)
 +TRANS(USUBL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      tcg_gen_sub_i64, false)
 +TRANS(SABDL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_sabd_i64, false)
 +TRANS(UABDL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_uabd_i64, false)
 +TRANS(SABAL_v, do_3op_widening,
 +      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_saba_i64, true)
 +TRANS(UABAL_v, do_3op_widening,
 +      a->esz, a->q, a->rd, a->rn, a->rm, -1,
 +      gen_uaba_i64, true)
 +
  /*
   * Advanced SIMD scalar/vector x indexed element
   */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
              }
+         } else if (is_qnan(a->cls)) {
-             switch (opcode) {
+             if (is_snan(b->cls) || !is_qnan(b->cls)) {
--            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
+                 which = 0;
--                tcg_gen_add_i64(tcg_passres, tcg_op1, tcg_op2);
+-            } else {
--                break;
+-                which = cmp > 0 ? 0 : 1;
--            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
++                break;
 -                tcg_gen_sub_i64(tcg_passres, tcg_op1, tcg_op2);
 -                break;
 -            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -            {
 -                TCGv_i64 tcg_tmp1 = tcg_temp_new_i64();
 -                TCGv_i64 tcg_tmp2 = tcg_temp_new_i64();
 -
 -                tcg_gen_sub_i64(tcg_tmp1, tcg_op1, tcg_op2);
 -                tcg_gen_sub_i64(tcg_tmp2, tcg_op2, tcg_op1);
 -                tcg_gen_movcond_i64(is_u ? TCG_COND_GEU : TCG_COND_GE,
 -                                    tcg_passres,
 -                                    tcg_op1, tcg_op2, tcg_tmp1, tcg_tmp2);
 -                break;
 -            }
              case 9: /* SQDMLAL, SQDMLAL2 */
              case 11: /* SQDMLSL, SQDMLSL2 */
              case 13: /* SQDMULL, SQDMULL2 */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
              case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
              case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
              case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 +            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 +            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 +            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 +            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
                  g_assert_not_reached();
              }
+         } else {
--            if (opcode == 9 || opcode == 11) {
+             which = 1;
-+            if (accop != 0) {
++            break;
                  /* saturating accumulate ops */
                  if (accop < 0) {
                      tcg_gen_neg_i64(tcg_passres, tcg_passres);
                  }
                  gen_helper_neon_addl_saturate_s64(tcg_res[pass], tcg_env,
                                                    tcg_res[pass], tcg_passres);
 -            } else if (accop > 0) {
 -                tcg_gen_add_i64(tcg_res[pass], tcg_res[pass], tcg_passres);
 -            } else if (accop < 0) {
 -                tcg_gen_sub_i64(tcg_res[pass], tcg_res[pass], tcg_passres);
              }
          }
-     } else {
++        cmp = frac_cmp(a, b);
-@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
++        if (cmp == 0) {
-             }
++            cmp = a->sign < b->sign;
++        }
-             switch (opcode) {
++        which = cmp > 0 ? 0 : 1;
 -            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 -            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 -            {
 -                TCGv_i64 tcg_op2_64 = tcg_temp_new_i64();
 -                static NeonGenWidenFn * const widenfns[2][2] = {
 -                    { gen_helper_neon_widen_s8, gen_helper_neon_widen_u8 },
 -                    { gen_helper_neon_widen_s16, gen_helper_neon_widen_u16 },
 -                };
 -                NeonGenWidenFn *widenfn = widenfns[size][is_u];
 -
 -                widenfn(tcg_op2_64, tcg_op2);
 -                widenfn(tcg_passres, tcg_op1);
 -                gen_neon_addl(size, (opcode == 2), tcg_passres,
 -                              tcg_passres, tcg_op2_64);
 -                break;
 -            }
 -            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
 -                if (size == 0) {
 -                    if (is_u) {
 -                        gen_helper_neon_abdl_u16(tcg_passres, tcg_op1, tcg_op2);
 -                    } else {
 -                        gen_helper_neon_abdl_s16(tcg_passres, tcg_op1, tcg_op2);
 -                    }
 -                } else {
 -                    if (is_u) {
 -                        gen_helper_neon_abdl_u32(tcg_passres, tcg_op1, tcg_op2);
 -                    } else {
 -                        gen_helper_neon_abdl_s32(tcg_passres, tcg_op1, tcg_op2);
 -                    }
 -                }
 -                break;
              case 9: /* SQDMLAL, SQDMLAL2 */
              case 11: /* SQDMLSL, SQDMLSL2 */
              case 13: /* SQDMULL, SQDMULL2 */
@@ -XXX,XX +XXX,XX @@ static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
              case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
              case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
              case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
 +            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 +            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 +            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 +            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
                  g_assert_not_reached();
              }
              if (accop != 0) {
 -                if (opcode == 9 || opcode == 11) {
 -                    /* saturating accumulate ops */
 -                    if (accop < 0) {
 -                        gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
 -                    }
 -                    gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
 -                                                      tcg_res[pass],
 -                                                      tcg_passres);
 -                } else {
 -                    gen_neon_addl(size, (accop < 0), tcg_res[pass],
 -                                  tcg_res[pass], tcg_passres);
 +                /* saturating accumulate ops */
 +                if (accop < 0) {
 +                    gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
                  }
 +                gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
 +                                                  tcg_res[pass],
 +                                                  tcg_passres);
              }
          }
      }
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
              unallocated_encoding(s);
              return;
          }
 -        /* fall through */
 -    case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
 -    case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 -    case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 -    case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
          /* 64 x 64 -> 128 */
          if (size == 3) {
              unallocated_encoding(s);
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
          handle_3rd_widening(s, is_q, is_u, size, opcode, rd, rn, rm);
          break;
      default:
-+    case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
+         g_assert_not_reached();
 +    case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
 +    case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
 +    case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
      case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
      case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
      case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
 --
 .34.1

-[PULL 10/24] hw/char/pl011: Avoid division-by-zero in pl011_get_baudrate()
+[PULL 70/72] softfloat: Replace WHICH with RET in parts_pick_nan
-From: Zheyu Ma <zheyuma97@gmail.com>
+From: Richard Henderson <richard.henderson@linaro.org>
-In pl011_get_baudrate(), when we calculate the baudrate we can
+Replace the "index" selecting between A and B with a result variable
-accidentally divide by zero. This happens because although (as the
+of the proper type.  This improves clarity within the function.
 specification requires) we treat UARTIBRD = 0 as invalid, we aren't
 correctly limiting UARTIBRD and UARTFBRD values to the 16-bit and 6-bit
 ranges the hardware allows, and so some non-zero values of UARTIBRD can
 result in a zero divisor.
-Enforce the correct register field widths on guest writes and on inbound
+Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 migration to avoid the division by zero.
 ASAN log:
 ==2973125==ERROR: AddressSanitizer: FPE on unknown address 0x55f72629b348
 (pc 0x55f72629b348 bp 0x7fffa24d0e00 sp 0x7fffa24d0d60 T0)
      #0 0x55f72629b348 in pl011_get_baudrate hw/char/pl011.c:255:17
      #1 0x55f726298d94 in pl011_trace_baudrate_change hw/char/pl011.c:260:33
      #2 0x55f726296fc8 in pl011_write hw/char/pl011.c:378:9
 Reproducer:
 cat << EOF | qemu-system-aarch64 -display \
 none -machine accel=qtest, -m 512M -machine realview-pb-a8 -qtest stdio
 writeq 0x1000b024 0xf8000000
 EOF
 Suggested-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Zheyu Ma <zheyuma97@gmail.com>
 Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
-Message-id: 20240702155752.3022007-1-zheyuma97@gmail.com
+Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/char/pl011.c | 13 +++++++++++--
+ fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
-file changed, 11 insertions(+), 2 deletions(-)
+file changed, 13 insertions(+), 15 deletions(-)
-diff --git a/hw/char/pl011.c b/hw/char/pl011.c
+diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
 index XXXXXXX..XXXXXXX 100644
---- a/hw/char/pl011.c
+--- a/fpu/softfloat-parts.c.inc
-+++ b/hw/char/pl011.c
++++ b/fpu/softfloat-parts.c.inc
-@@ -XXX,XX +XXX,XX @@ DeviceState *pl011_create(hwaddr addr, qemu_irq irq, Chardev *chr)
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
- #define CR_DTR      (1 << 10)
+                                      float_status *s)
- #define CR_LBE      (1 << 7)
+ {
+     bool have_snan = false;
-+/* Integer Baud Rate Divider, UARTIBRD */
+-    int cmp, which;
-+#define IBRD_MASK 0x3f
++    FloatPartsN *ret;
-+
++    int cmp;
-+/* Fractional Baud Rate Divider, UARTFBRD */
-+#define FBRD_MASK 0xffff
+     if (is_snan(a->cls) || is_snan(b->cls)) {
-+
+         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
- static const unsigned char pl011_id_arm[8] =
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-   { 0x11, 0x10, 0x14, 0x00, 0x0d, 0xf0, 0x05, 0xb1 };
+     switch (s->float_2nan_prop_rule) {
- static const unsigned char pl011_id_luminary[8] =
+     case float_2nan_prop_s_ab:
-@@ -XXX,XX +XXX,XX @@ static void pl011_write(void *opaque, hwaddr offset,
+         if (have_snan) {
-         s->ilpr = value;
+-            which = is_snan(a->cls) ? 0 : 1;
 +            ret = is_snan(a->cls) ? a : b;
              break;
          }
          /* fall through */
      case float_2nan_prop_ab:
 -        which = is_nan(a->cls) ? 0 : 1;
 +        ret = is_nan(a->cls) ? a : b;
          break;
-     case 9: /* UARTIBRD */
+     case float_2nan_prop_s_ba:
--        s->ibrd = value;
+         if (have_snan) {
-+        s->ibrd = value & IBRD_MASK;
+-            which = is_snan(b->cls) ? 1 : 0;
-         pl011_trace_baudrate_change(s);
++            ret = is_snan(b->cls) ? b : a;
              break;
          }
          /* fall through */
      case float_2nan_prop_ba:
 -        which = is_nan(b->cls) ? 1 : 0;
 +        ret = is_nan(b->cls) ? b : a;
          break;
-     case 10: /* UARTFBRD */
+     case float_2nan_prop_x87:
--        s->fbrd = value;
+         /*
-+        s->fbrd = value & FBRD_MASK;
+@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
-         pl011_trace_baudrate_change(s);
+          */
          if (is_snan(a->cls)) {
              if (!is_snan(b->cls)) {
 -                which = is_qnan(b->cls) ? 1 : 0;
 +                ret = is_qnan(b->cls) ? b : a;
                  break;
              }
          } else if (is_qnan(a->cls)) {
              if (is_snan(b->cls) || !is_qnan(b->cls)) {
 -                which = 0;
 +                ret = a;
                  break;
              }
          } else {
 -            which = 1;
 +            ret = b;
              break;
          }
          cmp = frac_cmp(a, b);
          if (cmp == 0) {
              cmp = a->sign < b->sign;
          }
 -        which = cmp > 0 ? 0 : 1;
 +        ret = cmp > 0 ? a : b;
          break;
-     case 11: /* UARTLCR_H */
+     default:
-@@ -XXX,XX +XXX,XX @@ static int pl011_post_load(void *opaque, int version_id)
+         g_assert_not_reached();
          s->read_pos = 0;
      }
-+    s->ibrd &= IBRD_MASK;
+-    if (which) {
-+    s->fbrd &= FBRD_MASK;
+-        a = b;
-+
++    if (is_snan(ret->cls)) {
-     return 0;
++        parts_silence_nan(ret, s);
      }
 -    if (is_snan(a->cls)) {
 -        parts_silence_nan(a, s);
 -    }
 -    return a;
 +    return ret;
  }
+ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
 --
 .34.1

-[PULL 15/24] accel/tcg: Make TCGCPUOps::cpu_exec_halt mandatory
+[PULL 71/72] MAINTAINERS: update email address for Leif Lindholm
-Now that all targets set TCGCPUOps::cpu_exec_halt, we can make it
+From: Leif Lindholm <quic_llindhol@quicinc.com>
 mandatory and remove the fallback handling that calls cpu_has_work.
+I'm migrating to Qualcomm's new open source email infrastructure, so
+update my email address, and update the mailmap to match.
+Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
+Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
+Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
+Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
+Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
 ---
- include/hw/core/tcg-cpu-ops.h |  9 ++++++---
+ MAINTAINERS | 2 +-
- accel/tcg/cpu-exec.c          | 11 +++++------
+ .mailmap    | 5 +++--
-files changed, 11 insertions(+), 9 deletions(-)
+files changed, 4 insertions(+), 3 deletions(-)
-diff --git a/include/hw/core/tcg-cpu-ops.h b/include/hw/core/tcg-cpu-ops.h
+diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
---- a/include/hw/core/tcg-cpu-ops.h
+--- a/MAINTAINERS
-+++ b/include/hw/core/tcg-cpu-ops.h
++++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ struct TCGCPUOps {
+@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
-      * to do when the CPU is in the halted state.
+ SBSA-REF
-      *
+ M: Radoslaw Biernacki <rad@semihalf.com>
-      * Return true to indicate that the CPU should now leave halt, false
+ M: Peter Maydell <peter.maydell@linaro.org>
--     * if it should remain in the halted state.
+-R: Leif Lindholm <quic_llindhol@quicinc.com>
-+     * if it should remain in the halted state. (This should generally
++R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
-+     * be the same value that cpu_has_work() would return.)
+ R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
-      *
+ L: qemu-arm@nongnu.org
--     * If this method is not provided, the default is to do nothing, and
+ S: Maintained
--     * to leave halt if cpu_has_work() returns true.
+diff --git a/.mailmap b/.mailmap
 +     * This method must be provided. If the target does not need to
 +     * do anything special for halt, the same function used for its
 +     * CPUClass::has_work method can be used here, as they have the
 +     * same function signature.
       */
      bool (*cpu_exec_halt)(CPUState *cpu);
      /**
 diff --git a/accel/tcg/cpu-exec.c b/accel/tcg/cpu-exec.c
 index XXXXXXX..XXXXXXX 100644
---- a/accel/tcg/cpu-exec.c
+--- a/.mailmap
-+++ b/accel/tcg/cpu-exec.c
++++ b/.mailmap
-@@ -XXX,XX +XXX,XX @@ static inline bool cpu_handle_halt(CPUState *cpu)
+@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
- #ifndef CONFIG_USER_ONLY
+ Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
-     if (cpu->halted) {
+ James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
-         const TCGCPUOps *tcg_ops = cpu->cc->tcg_ops;
+ Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
--        bool leave_halt;
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-+        bool leave_halt = tcg_ops->cpu_exec_halt(cpu);
+-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
--        if (tcg_ops->cpu_exec_halt) {
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
--            leave_halt = tcg_ops->cpu_exec_halt(cpu);
++Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
--        } else {
+ Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
--            leave_halt = cpu_has_work(cpu);
+ Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
--        }
+ Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
          if (!leave_halt) {
              return true;
          }
@@ -XXX,XX +XXX,XX @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp)
      static bool tcg_target_initialized;
      if (!tcg_target_initialized) {
 +        /* Check mandatory TCGCPUOps handlers */
 +#ifndef CONFIG_USER_ONLY
 +        assert(cpu->cc->tcg_ops->cpu_exec_halt);
 +#endif /* !CONFIG_USER_ONLY */
          cpu->cc->tcg_ops->initialize();
          tcg_target_initialized = true;
      }
 --
 .34.1

-[PULL 11/24] hw/misc/bcm2835_thermal: Fix access size handling in bcm2835_thermal_ops
+[PULL 72/72] MAINTAINERS: Add correct email address for Vikram Garhwal
-From: Zheyu Ma <zheyuma97@gmail.com>
+From: Vikram Garhwal <vikram.garhwal@bytedance.com>
-The current implementation of bcm2835_thermal_ops sets
+Previously, maintainer role was paused due to inactive email id. Commit id:
-impl.max_access_size and valid.min_access_size to 4, but leaves
+c009d715721861984c4987bcc78b7ee183e86d75.
 impl.min_access_size and valid.max_access_size unset, defaulting to 1.
 This causes issues when the memory system is presented with an access
 of size 2 at an offset of 3, leading to an attempt to synthesize it as
 a pair of byte accesses at offsets 3 and 4, which trips an assert.
-Additionally, the lack of valid.max_access_size setting causes another
+Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
-issue: the memory system tries to synthesize a read using a 4-byte
+Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
-access at offset 3 even though the device doesn't allow unaligned
+Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
 accesses.
 This patch addresses these issues by explicitly setting both
 impl.min_access_size and valid.max_access_size to 4, ensuring proper
 handling of access sizes.
 Error log:
 ERROR:hw/misc/bcm2835_thermal.c:55:bcm2835_thermal_read: code should not be reached
 Bail out! ERROR:hw/misc/bcm2835_thermal.c:55:bcm2835_thermal_read: code should not be reached
 Aborted
 Reproducer:
 cat << EOF | qemu-system-aarch64 -display \
 none -machine accel=qtest, -m 512M -machine raspi3b -m 1G -qtest stdio
 readw 0x3f212003
 EOF
 Signed-off-by: Zheyu Ma <zheyuma97@gmail.com>
 Message-id: 20240702154042.3018932-1-zheyuma97@gmail.com
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/misc/bcm2835_thermal.c | 2 ++
+ MAINTAINERS | 2 ++
 file changed, 2 insertions(+)
-diff --git a/hw/misc/bcm2835_thermal.c b/hw/misc/bcm2835_thermal.c
+diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
---- a/hw/misc/bcm2835_thermal.c
+--- a/MAINTAINERS
-+++ b/hw/misc/bcm2835_thermal.c
++++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ static void bcm2835_thermal_write(void *opaque, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
- static const MemoryRegionOps bcm2835_thermal_ops = {
-     .read = bcm2835_thermal_read,
+ Xilinx CAN
-     .write = bcm2835_thermal_write,
+ M: Francisco Iglesias <francisco.iglesias@amd.com>
-+    .impl.min_access_size = 4,
++M: Vikram Garhwal <vikram.garhwal@bytedance.com>
-     .impl.max_access_size = 4,
+ S: Maintained
-     .valid.min_access_size = 4,
+ F: hw/net/can/xlnx-*
-+    .valid.max_access_size = 4,
+ F: include/hw/net/xlnx-*
-     .endianness = DEVICE_NATIVE_ENDIAN,
+@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
- };
+ CAN bus subsystem and hardware
+ M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
  M: Francisco Iglesias <francisco.iglesias@amd.com>
 +M: Vikram Garhwal <vikram.garhwal@bytedance.com>
  S: Maintained
  W: https://canbus.pages.fel.cvut.cz/
  F: net/can/*
 --
 .34.1

The following changes since commit 59084feb256c617063e0dbe7e64821ae8852d7cf:

Merge tag 'pull-aspeed-20240709' of https://github.com/legoater/qemu into staging (2024-07-09 07:13:55 -0700)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20240711

for you to fetch changes up to 7f49089158a4db644fcbadfa90cd3d30a4868735:

target/arm: Convert PMULL to decodetree (2024-07-11 11:41:34 +0100)

----------------------------------------------------------------
target-arm queue:
 * Refactor FPCR/FPSR handling in preparation for FEAT_AFP
 * More decodetree conversions
 * target/arm: Use cpu_env in cpu_untagged_addr
 * target/arm: Set arm_v7m_tcg_ops cpu_exec_halt to arm_cpu_exec_halt()
 * hw/char/pl011: Avoid division-by-zero in pl011_get_baudrate()
 * hw/misc/bcm2835_thermal: Fix access size handling in bcm2835_thermal_ops
 * accel/tcg: Make TCGCPUOps::cpu_exec_halt mandatory
 * STM32L4x5: Handle USART interrupts correctly

----------------------------------------------------------------
Inès Varhol (3):
      hw/misc: In STM32L4x5 EXTI, consolidate 2 constants
      hw/misc: In STM32L4x5 EXTI, handle direct interrupts
      hw/arm: In STM32L4x5 SOC, connect USART devices to EXTI

Peter Maydell (12):
      target/arm: Correct comments about M-profile FPSCR
      target/arm: Make vfp_get_fpscr() call vfp_get_{fpcr, fpsr}
      target/arm: Make vfp_set_fpscr() call vfp_set_{fpcr, fpsr}
      target/arm: Support migration when FPSR/FPCR won't fit in the FPSCR
      target/arm: Implement store_cpu_field_low32() macro
      target/arm: Store FPSR and FPCR in separate CPU state fields
      target/arm: Rename FPCR_ QC, NZCV macros to FPSR_
      target/arm: Rename FPSR_MASK and FPCR_MASK and define them symbolically
      target/arm: Allow FPCR bits that aren't in FPSCR
      target/arm: Set arm_v7m_tcg_ops cpu_exec_halt to arm_cpu_exec_halt()
      target: Set TCGCPUOps::cpu_exec_halt to target's has_work implementation
      accel/tcg: Make TCGCPUOps::cpu_exec_halt mandatory

Richard Henderson (7):
      target/arm: Use cpu_env in cpu_untagged_addr
      target/arm: Convert SMULL, UMULL, SMLAL, UMLAL, SMLSL, UMLSL to decodetree
      target/arm: Convert SADDL, SSUBL, SABDL, SABAL, and unsigned to decodetree
      target/arm: Convert SQDMULL, SQDMLAL, SQDMLSL to decodetree
      target/arm: Convert SADDW, SSUBW, UADDW, USUBW to decodetree
      target/arm: Convert ADDHN, SUBHN, RADDHN, RSUBHN to decodetree
      target/arm: Convert PMULL to decodetree

Zheyu Ma (2):
      hw/char/pl011: Avoid division-by-zero in pl011_get_baudrate()
      hw/misc/bcm2835_thermal: Fix access size handling in bcm2835_thermal_ops

The M-profile FPSCR LTPSIZE is bits [18:16]; this is the same
field as A-profile FPSCR Len, not Stride. Correct the comment
in vfp_get_fpscr().

We also implemented M-profile FPSCR.QC, but forgot to delete
a TODO comment from vfp_set_fpscr(); remove it now.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-2-peter.maydell@linaro.org
---
 target/arm/vfp_helper.c | 5 ++---
 1 file changed, 2 insertions(+), 3 deletions(-)

diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
             | (env->vfp.vec_stride << 20);
 
     /*
-     * M-profile LTPSIZE overlaps A-profile Stride; whichever of the
-     * two is not applicable to this CPU will always be zero.
+     * M-profile LTPSIZE is the same bits [18:16] as A-profile Len; whichever
+     * of the two is not applicable to this CPU will always be zero.
      */
     fpscr |= env->v7m.ltpsize << 16;
 
@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
         /*
          * The bit we set within fpscr_q is arbitrary; the register as a
          * whole being zero/non-zero is what counts.
-         * TODO: M-profile MVE also has a QC bit.
          */
         env->vfp.qc[0] = val & FPCR_QC;
         env->vfp.qc[1] = 0;
-- 
2.34.1

In AArch32, the floating point control and status bits are all in a
single register, FPSCR.  In AArch64, these were split into separate
FPCR and FPSR registers, but the bit layouts remained the same, with
no overlaps, so that you could construct an FPSCR value by ORing FPCR
and FPSR, or equivalently could produce FPSR and FPCR by masking an
FPSCR value.  For QEMU's implementation, we opted to use masking to
produce FPSR and FPCR, because we started with an AArch32
implementation of FPSCR.

The addition of the (AArch64-only) FEAT_AFP adds new bits to the FPCR
which overlap with some bits in the FPSR.  This means we'll no longer
be able to consider the FPSCR-encoded value as the primary one, but
instead need to treat FPSR/FPCR as the primary encoding and construct
the FPSCR from those.  (This remains possible because the FEAT_AFP
bits in FPCR don't appear in the FPSCR.)

As the first step in this refactoring, make vfp_get_fpscr() call
vfp_get_fpcr() and vfp_get_fpsr(), instead of the other way around.

Note that vfp_get_fpcsr_from_host() returns only bits in the FPSR
(for the cumulative fp exception bits), so we can simply rename
it without needing to add a new function for getting FPCR bits.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-3-peter.maydell@linaro.org
---
 target/arm/cpu.h        | 24 +++++++++++++++---------
 target/arm/vfp_helper.c | 34 ++++++++++++++++++++++------------
 2 files changed, 37 insertions(+), 21 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 #define FPCR_NZCV_MASK (FPCR_N | FPCR_Z | FPCR_C | FPCR_V)
 #define FPCR_NZCVQC_MASK (FPCR_NZCV_MASK | FPCR_QC)
 
-static inline uint32_t vfp_get_fpsr(CPUARMState *env)
-{
-    return vfp_get_fpscr(env) & FPSR_MASK;
-}
+/**
+ * vfp_get_fpsr: read the AArch64 FPSR
+ * @env: CPU context
+ *
+ * Return the current AArch64 FPSR value
+ */
+uint32_t vfp_get_fpsr(CPUARMState *env);
+
+/**
+ * vfp_get_fpcr: read the AArch64 FPCR
+ * @env: CPU context
+ *
+ * Return the current AArch64 FPCR value
+ */
+uint32_t vfp_get_fpcr(CPUARMState *env);
 
 static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
 {
@@ -XXX,XX +XXX,XX @@ static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
     vfp_set_fpscr(env, new_fpscr);
 }
 
-static inline uint32_t vfp_get_fpcr(CPUARMState *env)
-{
-    return vfp_get_fpscr(env) & FPCR_MASK;
-}
-
 static inline void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 {
     uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPCR_MASK) | (val & FPCR_MASK);
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static inline int vfp_exceptbits_to_host(int target_bits)
     return host_bits;
 }
 
-static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
+static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
 {
     uint32_t i;
 
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
 
 #else
 
-static uint32_t vfp_get_fpscr_from_host(CPUARMState *env)
+static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
 {
     return 0;
 }
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
 
 #endif
 
-uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
+uint32_t vfp_get_fpcr(CPUARMState *env)
 {
-    uint32_t i, fpscr;
-
-    fpscr = env->vfp.xregs[ARM_VFP_FPSCR]
-            | (env->vfp.vec_len << 16)
-            | (env->vfp.vec_stride << 20);
+    uint32_t fpcr = (env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_MASK)
+        | (env->vfp.vec_len << 16)
+        | (env->vfp.vec_stride << 20);
 
     /*
      * M-profile LTPSIZE is the same bits [18:16] as A-profile Len; whichever
      * of the two is not applicable to this CPU will always be zero.
      */
-    fpscr |= env->v7m.ltpsize << 16;
+    fpcr |= env->v7m.ltpsize << 16;
 
-    fpscr |= vfp_get_fpscr_from_host(env);
+    return fpcr;
+}
+
+uint32_t vfp_get_fpsr(CPUARMState *env)
+{
+    uint32_t fpsr = env->vfp.xregs[ARM_VFP_FPSCR] & FPSR_MASK;
+    uint32_t i;
+
+    fpsr |= vfp_get_fpsr_from_host(env);
 
     i = env->vfp.qc[0] | env->vfp.qc[1] | env->vfp.qc[2] | env->vfp.qc[3];
-    fpscr |= i ? FPCR_QC : 0;
+    fpsr |= i ? FPCR_QC : 0;
+    return fpsr;
+}
 
-    return fpscr;
+uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
+{
+    return (vfp_get_fpcr(env) & FPCR_MASK) | (vfp_get_fpsr(env) & FPSR_MASK);
 }
 
 uint32_t vfp_get_fpscr(CPUARMState *env)
-- 
2.34.1

Make vfp_set_fpscr() call vfp_set_fpsr() and vfp_set_fpcr()
instead of the other way around.

The masking we do when getting and setting vfp.xregs[ARM_VFP_FPSCR]
is a little awkward, but we are going to change where we store the
underlying FPSR and FPCR information in a later commit, so it will
go away then.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-4-peter.maydell@linaro.org
---
 target/arm/cpu.h        |  22 +++++----
 target/arm/vfp_helper.c | 100 ++++++++++++++++++++++++++--------------
 2 files changed, 78 insertions(+), 44 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env);
  */
 uint32_t vfp_get_fpcr(CPUARMState *env);
 
-static inline void vfp_set_fpsr(CPUARMState *env, uint32_t val)
-{
-    uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPSR_MASK) | (val & FPSR_MASK);
-    vfp_set_fpscr(env, new_fpscr);
-}
+/**
+ * vfp_set_fpsr: write the AArch64 FPSR
+ * @env: CPU context
+ * @value: new value
+ */
+void vfp_set_fpsr(CPUARMState *env, uint32_t value);
 
-static inline void vfp_set_fpcr(CPUARMState *env, uint32_t val)
-{
-    uint32_t new_fpscr = (vfp_get_fpscr(env) & ~FPCR_MASK) | (val & FPCR_MASK);
-    vfp_set_fpscr(env, new_fpscr);
-}
+/**
+ * vfp_set_fpcr: write the AArch64 FPCR
+ * @env: CPU context
+ * @value: new value
+ */
+void vfp_set_fpcr(CPUARMState *env, uint32_t value);
 
 enum arm_cpu_mode {
   ARM_CPU_MODE_USR = 0x10,
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
     return vfp_exceptbits_from_host(i);
 }
 
-static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
+static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
+{
+    /*
+     * The exception flags are ORed together when we read fpscr so we
+     * only need to preserve the current state in one of our
+     * float_status values.
+     */
+    int i = vfp_exceptbits_to_host(val);
+    set_float_exception_flags(i, &env->vfp.fp_status);
+    set_float_exception_flags(0, &env->vfp.fp_status_f16);
+    set_float_exception_flags(0, &env->vfp.standard_fp_status);
+    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
+}
+
+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
 {
-    int i;
     uint32_t changed = env->vfp.xregs[ARM_VFP_FPSCR];
 
     changed ^= val;
     if (changed & (3 << 22)) {
-        i = (val >> 22) & 3;
+        int i = (val >> 22) & 3;
         switch (i) {
         case FPROUNDING_TIEEVEN:
             i = float_round_nearest_even;
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
         set_default_nan_mode(dnan_enabled, &env->vfp.fp_status);
         set_default_nan_mode(dnan_enabled, &env->vfp.fp_status_f16);
     }
-
-    /*
-     * The exception flags are ORed together when we read fpscr so we
-     * only need to preserve the current state in one of our
-     * float_status values.
-     */
-    i = vfp_exceptbits_to_host(val);
-    set_float_exception_flags(i, &env->vfp.fp_status);
-    set_float_exception_flags(0, &env->vfp.fp_status_f16);
-    set_float_exception_flags(0, &env->vfp.standard_fp_status);
-    set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
 }
 
 #else
@@ -XXX,XX +XXX,XX @@ static uint32_t vfp_get_fpsr_from_host(CPUARMState *env)
     return 0;
 }
 
-static void vfp_set_fpscr_to_host(CPUARMState *env, uint32_t val)
+static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
+{
+}
+
+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
 {
 }
 
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpscr(CPUARMState *env)
     return HELPER(vfp_get_fpscr)(env);
 }
 
-void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
+void vfp_set_fpsr(CPUARMState *env, uint32_t val)
+{
+    ARMCPU *cpu = env_archcpu(env);
+
+    vfp_set_fpsr_to_host(env, val);
+
+    if (arm_feature(env, ARM_FEATURE_NEON) ||
+        cpu_isar_feature(aa32_mve, cpu)) {
+        /*
+         * The bit we set within vfp.qc[] is arbitrary; the array as a
+         * whole being zero/non-zero is what counts.
+         */
+        env->vfp.qc[0] = val & FPCR_QC;
+        env->vfp.qc[1] = 0;
+        env->vfp.qc[2] = 0;
+        env->vfp.qc[3] = 0;
+    }
+
+    /*
+     * The only FPSR bits we keep in vfp.xregs[FPSCR] are NZCV:
+     * the exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
+     * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
+     * and zero any of the other FPSR bits (but preserve the FPCR
+     * bits).
+     */
+    val &= FPCR_NZCV_MASK;
+    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPSR_MASK;
+    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
+}
+
+void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 {
     ARMCPU *cpu = env_archcpu(env);
 
@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
         val &= ~FPCR_FZ16;
     }
 
-    vfp_set_fpscr_to_host(env, val);
+    vfp_set_fpcr_to_host(env, val);
 
     if (!arm_feature(env, ARM_FEATURE_M)) {
         /*
@@ -XXX,XX +XXX,XX @@ void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
                                      FPCR_LTPSIZE_LENGTH);
     }
 
-    if (arm_feature(env, ARM_FEATURE_NEON) ||
-        cpu_isar_feature(aa32_mve, cpu)) {
-        /*
-         * The bit we set within fpscr_q is arbitrary; the register as a
-         * whole being zero/non-zero is what counts.
-         */
-        env->vfp.qc[0] = val & FPCR_QC;
-        env->vfp.qc[1] = 0;
-        env->vfp.qc[2] = 0;
-        env->vfp.qc[3] = 0;
-    }
-
     /*
      * We don't implement trapped exception handling, so the
      * trap enable bits, IDE|IXE|UFE|OFE|DZE|IOE are all RAZ/WI (not RES0!)
      *
-     * The exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
-     * fp_status; QC, Len and Stride are stored separately earlier.
-     * Clear out all of those and the RES0 bits: only NZCV, AHP, DN,
-     * FZ, RMode and FZ16 are kept in vfp.xregs[FPSCR].
+     * The FPCR bits we keep in vfp.xregs[FPSCR] are AHP, DN, FZ, RMode
+     * and FZ16. Len, Stride and LTPSIZE we just handled. Store those bits
+     * there, and zero any of the other FPCR bits and the RES0 and RAZ/WI
+     * bits.
      */
-    env->vfp.xregs[ARM_VFP_FPSCR] = val & 0xf7c80000;
+    val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
+    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_MASK;
+    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
+}
+
+void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
+{
+    vfp_set_fpcr(env, val & FPCR_MASK);
+    vfp_set_fpsr(env, val & FPSR_MASK);
 }
 
 void vfp_set_fpscr(CPUARMState *env, uint32_t val)
-- 
2.34.1

To support FPSR and FPCR bits that don't exist in the AArch32 FPSCR
view of floating point control and status (such as the FEAT_AFP ones),
we need to make sure those bits can be migrated. This commit allows
that, whilst maintaining backwards and forwards migration compatibility
for CPUs where there are no such bits:

On sending:
 * If either the FPCR or the FPSR include set bits that are not
   visible in the AArch32 FPSCR view of floating point control/status
   then we send the FPCR and FPSR as two separate fields in a new
   cpu/vfp/fpcr_fpsr subsection, and we send a 0 for the old
   FPSCR field in cpu/vfp
 * Otherwise, we don't send the fpcr_fpsr subsection, and we send
   an FPSCR-format value in cpu/vfp as we did previously

On receiving:
 * if we see a non-zero FPSCR field, that is the right information
 * if we see a fpcr_fpsr subsection then that has the information
 * if we see neither, then FPSCR/FPCR/FPSR are all zero on the source;
   cpu_pre_load() ensures the CPU state defaults to that
 * if we see both, then the migration source is buggy or malicious;
   either the fpcr_fpsr or the FPSCR will "win" depending which
   is first in the migration stream; we don't care which that is

We make the new FPCR and FPSR on-the-wire data be 64 bits, because
architecturally these registers are that wide, and this avoids the
need to engage in further migration-compatibility contortions in
future if some new architecture revision defines bits in the high
half of either register.

(We won't ever send the new migration subsection until we add support
for a CPU feature which enables setting overlapping FPCR bits, like
FEAT_AFP.)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-5-peter.maydell@linaro.org
---
 target/arm/machine.c | 134 ++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 132 insertions(+), 2 deletions(-)

diff --git a/target/arm/machine.c b/target/arm/machine.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/machine.c
+++ b/target/arm/machine.c
@@ -XXX,XX +XXX,XX @@ static bool vfp_needed(void *opaque)
             : cpu_isar_feature(aa32_vfp_simd, cpu));
 }
 
+static bool vfp_fpcr_fpsr_needed(void *opaque)
+{
+    /*
+     * If either the FPCR or the FPSR include set bits that are not
+     * visible in the AArch32 FPSCR view of floating point control/status
+     * then we must send the FPCR and FPSR as two separate fields in the
+     * cpu/vfp/fpcr_fpsr subsection, and we will send a 0 for the old
+     * FPSCR field in cpu/vfp.
+     *
+     * If all the set bits are representable in an AArch32 FPSCR then we
+     * send that value as the cpu/vfp FPSCR field, and don't send the
+     * cpu/vfp/fpcr_fpsr subsection.
+     *
+     * On incoming migration, if the cpu/vfp FPSCR field is non-zero we
+     * use it, and if the fpcr_fpsr subsection is present we use that.
+     * (The subsection will never be present with a non-zero FPSCR field,
+     * and if FPSCR is zero and the subsection is not present that means
+     * that FPSCR/FPSR/FPCR are zero.)
+     *
+     * This preserves migration compatibility with older QEMU versions,
+     * in both directions.
+     */
+    ARMCPU *cpu = opaque;
+    CPUARMState *env = &cpu->env;
+
+    return (vfp_get_fpcr(env) & ~FPCR_MASK) || (vfp_get_fpsr(env) & ~FPSR_MASK);
+}
+
 static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
                      const VMStateField *field)
 {
@@ -XXX,XX +XXX,XX @@ static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
     CPUARMState *env = &cpu->env;
     uint32_t val = qemu_get_be32(f);
 
-    vfp_set_fpscr(env, val);
+    if (val) {
+        /* 0 means we might have the data in the fpcr_fpsr subsection */
+        vfp_set_fpscr(env, val);
+    }
     return 0;
 }
 
@@ -XXX,XX +XXX,XX @@ static int put_fpscr(QEMUFile *f, void *opaque, size_t size,
 {
     ARMCPU *cpu = opaque;
     CPUARMState *env = &cpu->env;
+    uint32_t fpscr = vfp_fpcr_fpsr_needed(opaque) ? 0 : vfp_get_fpscr(env);
 
-    qemu_put_be32(f, vfp_get_fpscr(env));
+    qemu_put_be32(f, fpscr);
     return 0;
 }
 
@@ -XXX,XX +XXX,XX @@ static const VMStateInfo vmstate_fpscr = {
     .put = put_fpscr,
 };
 
+static int get_fpcr(QEMUFile *f, void *opaque, size_t size,
+                     const VMStateField *field)
+{
+    ARMCPU *cpu = opaque;
+    CPUARMState *env = &cpu->env;
+    uint64_t val = qemu_get_be64(f);
+
+    vfp_set_fpcr(env, val);
+    return 0;
+}
+
+static int put_fpcr(QEMUFile *f, void *opaque, size_t size,
+                     const VMStateField *field, JSONWriter *vmdesc)
+{
+    ARMCPU *cpu = opaque;
+    CPUARMState *env = &cpu->env;
+
+    qemu_put_be64(f, vfp_get_fpcr(env));
+    return 0;
+}
+
+static const VMStateInfo vmstate_fpcr = {
+    .name = "fpcr",
+    .get = get_fpcr,
+    .put = put_fpcr,
+};
+
+static int get_fpsr(QEMUFile *f, void *opaque, size_t size,
+                     const VMStateField *field)
+{
+    ARMCPU *cpu = opaque;
+    CPUARMState *env = &cpu->env;
+    uint64_t val = qemu_get_be64(f);
+
+    vfp_set_fpsr(env, val);
+    return 0;
+}
+
+static int put_fpsr(QEMUFile *f, void *opaque, size_t size,
+                     const VMStateField *field, JSONWriter *vmdesc)
+{
+    ARMCPU *cpu = opaque;
+    CPUARMState *env = &cpu->env;
+
+    qemu_put_be64(f, vfp_get_fpsr(env));
+    return 0;
+}
+
+static const VMStateInfo vmstate_fpsr = {
+    .name = "fpsr",
+    .get = get_fpsr,
+    .put = put_fpsr,
+};
+
+static const VMStateDescription vmstate_vfp_fpcr_fpsr = {
+    .name = "cpu/vfp/fpcr_fpsr",
+    .version_id = 1,
+    .minimum_version_id = 1,
+    .needed = vfp_fpcr_fpsr_needed,
+    .fields = (const VMStateField[]) {
+        {
+            .name = "fpcr",
+            .version_id = 0,
+            .size = sizeof(uint64_t),
+            .info = &vmstate_fpcr,
+            .flags = VMS_SINGLE,
+            .offset = 0,
+        },
+        {
+            .name = "fpsr",
+            .version_id = 0,
+            .size = sizeof(uint64_t),
+            .info = &vmstate_fpsr,
+            .flags = VMS_SINGLE,
+            .offset = 0,
+        },
+        VMSTATE_END_OF_LIST()
+    },
+};
+
 static const VMStateDescription vmstate_vfp = {
     .name = "cpu/vfp",
     .version_id = 3,
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_vfp = {
             .offset = 0,
         },
         VMSTATE_END_OF_LIST()
+    },
+    .subsections = (const VMStateDescription * const []) {
+        &vmstate_vfp_fpcr_fpsr,
+        NULL
     }
 };
 
@@ -XXX,XX +XXX,XX @@ static int cpu_pre_load(void *opaque)
     ARMCPU *cpu = opaque;
     CPUARMState *env = &cpu->env;
 
+    /*
+     * In an inbound migration where on the source FPSCR/FPSR/FPCR are 0,
+     * there will be no fpcr_fpsr subsection so we won't call vfp_set_fpcr()
+     * and vfp_set_fpsr() from get_fpcr() and get_fpsr(); also the get_fpscr()
+     * function will not call vfp_set_fpscr() because it will see a 0 in the
+     * inbound data. Ensure that in this case we have a correctly set up
+     * zero FPSCR/FPCR/FPSR.
+     *
+     * This is not strictly needed because FPSCR is zero out of reset, but
+     * it avoids the possibility of future confusing migration bugs if some
+     * future architecture change makes the reset value non-zero.
+     */
+    vfp_set_fpscr(env, 0);
+
     /*
      * Pre-initialize irq_line_state to a value that's never valid as
      * real data, so cpu_post_load() can tell whether we've seen the
-- 
2.34.1

We already have a load_cpu_field_low32() to load the low half of a
64-bit CPU struct field to a TCGv_i32; however we haven't yet needed
the store equivalent.  We'll want that in the next patch, so
implement it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-6-peter.maydell@linaro.org
---
 target/arm/tcg/translate-a32.h | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/target/arm/tcg/translate-a32.h b/target/arm/tcg/translate-a32.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-a32.h
+++ b/target/arm/tcg/translate-a32.h
@@ -XXX,XX +XXX,XX @@ void store_cpu_offset(TCGv_i32 var, int offset, int size);
                          sizeof_field(CPUARMState, name));              \
     })
 
+/* Store to the low half of a 64-bit field from a TCGv_i32 */
+#define store_cpu_field_low32(val, name)                                \
+    ({                                                                  \
+        QEMU_BUILD_BUG_ON(sizeof_field(CPUARMState, name) != 8);        \
+        store_cpu_offset(val, offsetoflow32(CPUARMState, name), 4);     \
+    })
+
 #define store_cpu_field_constant(val, name) \
     store_cpu_field(tcg_constant_i32(val), name)
 
-- 
2.34.1

Now that we have refactored the set/get functions so that the FPSCR
format is no longer the authoritative one, we can keep FPSR and FPCR
in separate CPU state fields.

As well as the get and set functions, we also have a scattering of
places in the code which directly access vfp.xregs[ARM_VFP_FPSCR] to
extract single fields which are stored there.  These all change to
directly access either vfp.fpsr or vfp.fpcr, depending on the
location of the field.  (Most commonly, this is the NZCV flags.)

We make the field in the CPU state struct 64 bits, because
architecturally FPSR and FPCR are 64 bits.  However we leave the
types of the arguments and return values of the get/set functions as
32 bits, since we don't need to make that change with the current
architecture and various callsites would be unable to handle
set bits in the high half (for instance the gdbstub protocol
assumes they're only 32 bit registers).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-7-peter.maydell@linaro.org
---
 target/arm/cpu.h                  |  7 +++++++
 target/arm/tcg/translate.h        |  3 +--
 target/arm/tcg/mve_helper.c       | 12 ++++++------
 target/arm/tcg/translate-m-nocp.c |  6 +++---
 target/arm/tcg/translate-vfp.c    |  2 +-
 target/arm/vfp_helper.c           | 25 ++++++++++---------------
 6 files changed, 28 insertions(+), 27 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ typedef struct CPUArchState {
         int vec_len;
         int vec_stride;
 
+        /*
+         * Floating point status and control registers. Some bits are
+         * stored separately in other fields or in the float_status below.
+         */
+        uint64_t fpsr;
+        uint64_t fpcr;
+
         uint32_t xregs[16];
 
         /* Scratch space for aa32 neon expansion.  */
diff --git a/target/arm/tcg/translate.h b/target/arm/tcg/translate.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate.h
+++ b/target/arm/tcg/translate.h
@@ -XXX,XX +XXX,XX @@ static inline TCGv_i32 get_ahp_flag(void)
 {
     TCGv_i32 ret = tcg_temp_new_i32();
 
-    tcg_gen_ld_i32(ret, tcg_env,
-                   offsetof(CPUARMState, vfp.xregs[ARM_VFP_FPSCR]));
+    tcg_gen_ld_i32(ret, tcg_env, offsetoflow32(CPUARMState, vfp.fpcr));
     tcg_gen_extract_i32(ret, ret, 26, 1);
 
     return ret;
diff --git a/target/arm/tcg/mve_helper.c b/target/arm/tcg/mve_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/mve_helper.c
+++ b/target/arm/tcg/mve_helper.c
@@ -XXX,XX +XXX,XX @@ static void do_vadc(CPUARMState *env, uint32_t *d, uint32_t *n, uint32_t *m,
 
     if (update_flags) {
         /* Store C, clear NZV. */
-        env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_NZCV_MASK;
-        env->vfp.xregs[ARM_VFP_FPSCR] |= carry_in * FPCR_C;
+        env->vfp.fpsr &= ~FPCR_NZCV_MASK;
+        env->vfp.fpsr |= carry_in * FPCR_C;
     }
     mve_advance_vpt(env);
 }
 
 void HELPER(mve_vadc)(CPUARMState *env, void *vd, void *vn, void *vm)
 {
-    bool carry_in = env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_C;
+    bool carry_in = env->vfp.fpsr & FPCR_C;
     do_vadc(env, vd, vn, vm, 0, carry_in, false);
 }
 
 void HELPER(mve_vsbc)(CPUARMState *env, void *vd, void *vn, void *vm)
 {
-    bool carry_in = env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_C;
+    bool carry_in = env->vfp.fpsr & FPCR_C;
     do_vadc(env, vd, vn, vm, -1, carry_in, false);
 }
 
@@ -XXX,XX +XXX,XX @@ static void do_vcvt_sh(CPUARMState *env, void *vd, void *vm, int top)
     uint32_t *m = vm;
     uint16_t r;
     uint16_t mask = mve_element_mask(env);
-    bool ieee = !(env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_AHP);
+    bool ieee = !(env->vfp.fpcr & FPCR_AHP);
     unsigned e;
     float_status *fpst;
     float_status scratch_fpst;
@@ -XXX,XX +XXX,XX @@ static void do_vcvt_hs(CPUARMState *env, void *vd, void *vm, int top)
     uint16_t *m = vm;
     uint32_t r;
     uint16_t mask = mve_element_mask(env);
-    bool ieee = !(env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_AHP);
+    bool ieee = !(env->vfp.fpcr & FPCR_AHP);
     unsigned e;
     float_status *fpst;
     float_status scratch_fpst;
diff --git a/target/arm/tcg/translate-m-nocp.c b/target/arm/tcg/translate-m-nocp.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-m-nocp.c
+++ b/target/arm/tcg/translate-m-nocp.c
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
                                  16, 16, qc);
         }
         tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
-        fpscr = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
+        fpscr = load_cpu_field_low32(vfp.fpsr);
         tcg_gen_andi_i32(fpscr, fpscr, ~FPCR_NZCV_MASK);
         tcg_gen_or_i32(fpscr, fpscr, tmp);
-        store_cpu_field(fpscr, vfp.xregs[ARM_VFP_FPSCR]);
+        store_cpu_field_low32(fpscr, vfp.fpsr);
         break;
     }
     case ARM_VFP_FPCXT_NS:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
          * Read just NZCV; this is a special case to avoid the
          * helper call for the "VMRS to CPSR.NZCV" insn.
          */
-        tmp = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
+        tmp = load_cpu_field_low32(vfp.fpsr);
         tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
         storefn(s, opaque, tmp, true);
         break;
diff --git a/target/arm/tcg/translate-vfp.c b/target/arm/tcg/translate-vfp.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-vfp.c
+++ b/target/arm/tcg/translate-vfp.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VMSR_VMRS(DisasContext *s, arg_VMSR_VMRS *a)
             break;
         case ARM_VFP_FPSCR:
             if (a->rt == 15) {
-                tmp = load_cpu_field(vfp.xregs[ARM_VFP_FPSCR]);
+                tmp = load_cpu_field_low32(vfp.fpsr);
                 tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
             } else {
                 tmp = tcg_temp_new_i32();
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
 
 static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
 {
-    uint32_t changed = env->vfp.xregs[ARM_VFP_FPSCR];
+    uint64_t changed = env->vfp.fpcr;
 
     changed ^= val;
     if (changed & (3 << 22)) {
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
 
 uint32_t vfp_get_fpcr(CPUARMState *env)
 {
-    uint32_t fpcr = (env->vfp.xregs[ARM_VFP_FPSCR] & FPCR_MASK)
+    uint32_t fpcr = env->vfp.fpcr
         | (env->vfp.vec_len << 16)
         | (env->vfp.vec_stride << 20);
 
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpcr(CPUARMState *env)
 
 uint32_t vfp_get_fpsr(CPUARMState *env)
 {
-    uint32_t fpsr = env->vfp.xregs[ARM_VFP_FPSCR] & FPSR_MASK;
+    uint32_t fpsr = env->vfp.fpsr;
     uint32_t i;
 
     fpsr |= vfp_get_fpsr_from_host(env);
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
     }
 
     /*
-     * The only FPSR bits we keep in vfp.xregs[FPSCR] are NZCV:
+     * The only FPSR bits we keep in vfp.fpsr are NZCV:
      * the exception flags IOC|DZC|OFC|UFC|IXC|IDC are stored in
      * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
-     * and zero any of the other FPSR bits (but preserve the FPCR
-     * bits).
+     * and zero any of the other FPSR bits.
      */
     val &= FPCR_NZCV_MASK;
-    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPSR_MASK;
-    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
+    env->vfp.fpsr = val;
 }
 
 void vfp_set_fpcr(CPUARMState *env, uint32_t val)
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
      * We don't implement trapped exception handling, so the
      * trap enable bits, IDE|IXE|UFE|OFE|DZE|IOE are all RAZ/WI (not RES0!)
      *
-     * The FPCR bits we keep in vfp.xregs[FPSCR] are AHP, DN, FZ, RMode
+     * The FPCR bits we keep in vfp.fpcr are AHP, DN, FZ, RMode
      * and FZ16. Len, Stride and LTPSIZE we just handled. Store those bits
      * there, and zero any of the other FPCR bits and the RES0 and RAZ/WI
      * bits.
      */
     val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
-    env->vfp.xregs[ARM_VFP_FPSCR] &= ~FPCR_MASK;
-    env->vfp.xregs[ARM_VFP_FPSCR] |= val;
+    env->vfp.fpcr = val;
 }
 
 void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void softfloat_to_vfp_compare(CPUARMState *env, FloatRelation cmp)
     default:
         g_assert_not_reached();
     }
-    env->vfp.xregs[ARM_VFP_FPSCR] =
-        deposit32(env->vfp.xregs[ARM_VFP_FPSCR], 28, 4, flags);
+    env->vfp.fpsr = deposit64(env->vfp.fpsr, 28, 4, flags); /* NZCV */
 }
 
 /* XXX: check quiet/signaling case */
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vjcvt)(float64 value, CPUARMState *env)
     uint32_t z = (pair >> 32) == 0;
 
     /* Store Z, clear NCV, in FPSCR.NZCV.  */
-    env->vfp.xregs[ARM_VFP_FPSCR]
-        = (env->vfp.xregs[ARM_VFP_FPSCR] & ~CPSR_NZCV) | (z * CPSR_Z);
+    env->vfp.fpsr = (env->vfp.fpsr & ~FPCR_NZCV_MASK) | (z * FPCR_Z);
 
     return result;
 }
-- 
2.34.1

The QC, N, Z, C, V bits live in the FPSR, not the FPCR. Rename the
macros that define these bits accordingly.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-8-peter.maydell@linaro.org
---
 target/arm/cpu.h                  | 17 ++++++++++-------
 target/arm/tcg/mve_helper.c       |  8 ++++----
 target/arm/tcg/translate-m-nocp.c | 16 ++++++++--------
 target/arm/tcg/translate-vfp.c    |  2 +-
 target/arm/vfp_helper.c           |  8 ++++----
 5 files changed, 27 insertions(+), 24 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 #define FPSR_MASK 0xf800009f
 #define FPCR_MASK 0x07ff9f00
 
+/* FPCR bits */
 #define FPCR_IOE    (1 << 8)    /* Invalid Operation exception trap enable */
 #define FPCR_DZE    (1 << 9)    /* Divide by Zero exception trap enable */
 #define FPCR_OFE    (1 << 10)   /* Overflow exception trap enable */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 #define FPCR_FZ     (1 << 24)   /* Flush-to-zero enable bit */
 #define FPCR_DN     (1 << 25)   /* Default NaN enable bit */
 #define FPCR_AHP    (1 << 26)   /* Alternative half-precision */
-#define FPCR_QC     (1 << 27)   /* Cumulative saturation bit */
-#define FPCR_V      (1 << 28)   /* FP overflow flag */
-#define FPCR_C      (1 << 29)   /* FP carry flag */
-#define FPCR_Z      (1 << 30)   /* FP zero flag */
-#define FPCR_N      (1 << 31)   /* FP negative flag */
 
 #define FPCR_LTPSIZE_SHIFT 16   /* LTPSIZE, M-profile only */
 #define FPCR_LTPSIZE_MASK (7 << FPCR_LTPSIZE_SHIFT)
 #define FPCR_LTPSIZE_LENGTH 3
 
-#define FPCR_NZCV_MASK (FPCR_N | FPCR_Z | FPCR_C | FPCR_V)
-#define FPCR_NZCVQC_MASK (FPCR_NZCV_MASK | FPCR_QC)
+/* FPSR bits */
+#define FPSR_QC     (1 << 27)   /* Cumulative saturation bit */
+#define FPSR_V      (1 << 28)   /* FP overflow flag */
+#define FPSR_C      (1 << 29)   /* FP carry flag */
+#define FPSR_Z      (1 << 30)   /* FP zero flag */
+#define FPSR_N      (1 << 31)   /* FP negative flag */
+
+#define FPSR_NZCV_MASK (FPSR_N | FPSR_Z | FPSR_C | FPSR_V)
+#define FPSR_NZCVQC_MASK (FPSR_NZCV_MASK | FPSR_QC)
 
 /**
  * vfp_get_fpsr: read the AArch64 FPSR
diff --git a/target/arm/tcg/mve_helper.c b/target/arm/tcg/mve_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/mve_helper.c
+++ b/target/arm/tcg/mve_helper.c
@@ -XXX,XX +XXX,XX @@ static void do_vadc(CPUARMState *env, uint32_t *d, uint32_t *n, uint32_t *m,
 
     if (update_flags) {
         /* Store C, clear NZV. */
-        env->vfp.fpsr &= ~FPCR_NZCV_MASK;
-        env->vfp.fpsr |= carry_in * FPCR_C;
+        env->vfp.fpsr &= ~FPSR_NZCV_MASK;
+        env->vfp.fpsr |= carry_in * FPSR_C;
     }
     mve_advance_vpt(env);
 }
 
 void HELPER(mve_vadc)(CPUARMState *env, void *vd, void *vn, void *vm)
 {
-    bool carry_in = env->vfp.fpsr & FPCR_C;
+    bool carry_in = env->vfp.fpsr & FPSR_C;
     do_vadc(env, vd, vn, vm, 0, carry_in, false);
 }
 
 void HELPER(mve_vsbc)(CPUARMState *env, void *vd, void *vn, void *vm)
 {
-    bool carry_in = env->vfp.fpsr & FPCR_C;
+    bool carry_in = env->vfp.fpsr & FPSR_C;
     do_vadc(env, vd, vn, vm, -1, carry_in, false);
 }
 
diff --git a/target/arm/tcg/translate-m-nocp.c b/target/arm/tcg/translate-m-nocp.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-m-nocp.c
+++ b/target/arm/tcg/translate-m-nocp.c
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
         if (dc_isar_feature(aa32_mve, s)) {
             /* QC is only present for MVE; otherwise RES0 */
             TCGv_i32 qc = tcg_temp_new_i32();
-            tcg_gen_andi_i32(qc, tmp, FPCR_QC);
+            tcg_gen_andi_i32(qc, tmp, FPSR_QC);
             /*
              * The 4 vfp.qc[] fields need only be "zero" vs "non-zero";
              * here writing the same value into all elements is simplest.
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
             tcg_gen_gvec_dup_i32(MO_32, offsetof(CPUARMState, vfp.qc),
                                  16, 16, qc);
         }
-        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
         fpscr = load_cpu_field_low32(vfp.fpsr);
-        tcg_gen_andi_i32(fpscr, fpscr, ~FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(fpscr, fpscr, ~FPSR_NZCV_MASK);
         tcg_gen_or_i32(fpscr, fpscr, tmp);
         store_cpu_field_low32(fpscr, vfp.fpsr);
         break;
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_write(DisasContext *s, int regno,
         tcg_gen_deposit_i32(control, control, sfpa,
                             R_V7M_CONTROL_SFPA_SHIFT, 1);
         store_cpu_field(control, v7m.control[M_REG_S]);
-        tcg_gen_andi_i32(tmp, tmp, ~FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(tmp, tmp, ~FPSR_NZCV_MASK);
         gen_helper_vfp_set_fpscr(tcg_env, tmp);
         s->base.is_jmp = DISAS_UPDATE_NOCHAIN;
         break;
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
     case ARM_VFP_FPSCR_NZCVQC:
         tmp = tcg_temp_new_i32();
         gen_helper_vfp_get_fpscr(tmp, tcg_env);
-        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCVQC_MASK);
+        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCVQC_MASK);
         storefn(s, opaque, tmp, true);
         break;
     case QEMU_VFP_FPSCR_NZCV:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
          * helper call for the "VMRS to CPSR.NZCV" insn.
          */
         tmp = load_cpu_field_low32(vfp.fpsr);
-        tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
         storefn(s, opaque, tmp, true);
         break;
     case ARM_VFP_FPCXT_S:
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
         tmp = tcg_temp_new_i32();
         sfpa = tcg_temp_new_i32();
         gen_helper_vfp_get_fpscr(tmp, tcg_env);
-        tcg_gen_andi_i32(tmp, tmp, ~FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(tmp, tmp, ~FPSR_NZCV_MASK);
         control = load_cpu_field(v7m.control[M_REG_S]);
         tcg_gen_andi_i32(sfpa, control, R_V7M_CONTROL_SFPA_MASK);
         tcg_gen_shli_i32(sfpa, sfpa, 31 - R_V7M_CONTROL_SFPA_SHIFT);
@@ -XXX,XX +XXX,XX @@ static bool gen_M_fp_sysreg_read(DisasContext *s, int regno,
         sfpa = tcg_temp_new_i32();
         fpscr = tcg_temp_new_i32();
         gen_helper_vfp_get_fpscr(fpscr, tcg_env);
-        tcg_gen_andi_i32(tmp, fpscr, ~FPCR_NZCV_MASK);
+        tcg_gen_andi_i32(tmp, fpscr, ~FPSR_NZCV_MASK);
         control = load_cpu_field(v7m.control[M_REG_S]);
         tcg_gen_andi_i32(sfpa, control, R_V7M_CONTROL_SFPA_MASK);
         tcg_gen_shli_i32(sfpa, sfpa, 31 - R_V7M_CONTROL_SFPA_SHIFT);
diff --git a/target/arm/tcg/translate-vfp.c b/target/arm/tcg/translate-vfp.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-vfp.c
+++ b/target/arm/tcg/translate-vfp.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VMSR_VMRS(DisasContext *s, arg_VMSR_VMRS *a)
         case ARM_VFP_FPSCR:
             if (a->rt == 15) {
                 tmp = load_cpu_field_low32(vfp.fpsr);
-                tcg_gen_andi_i32(tmp, tmp, FPCR_NZCV_MASK);
+                tcg_gen_andi_i32(tmp, tmp, FPSR_NZCV_MASK);
             } else {
                 tmp = tcg_temp_new_i32();
                 gen_helper_vfp_get_fpscr(tmp, tcg_env);
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env)
     fpsr |= vfp_get_fpsr_from_host(env);
 
     i = env->vfp.qc[0] | env->vfp.qc[1] | env->vfp.qc[2] | env->vfp.qc[3];
-    fpsr |= i ? FPCR_QC : 0;
+    fpsr |= i ? FPSR_QC : 0;
     return fpsr;
 }
 
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
          * The bit we set within vfp.qc[] is arbitrary; the array as a
          * whole being zero/non-zero is what counts.
          */
-        env->vfp.qc[0] = val & FPCR_QC;
+        env->vfp.qc[0] = val & FPSR_QC;
         env->vfp.qc[1] = 0;
         env->vfp.qc[2] = 0;
         env->vfp.qc[3] = 0;
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
      * fp_status, and QC is in vfp.qc[]. Store the NZCV bits there,
      * and zero any of the other FPSR bits.
      */
-    val &= FPCR_NZCV_MASK;
+    val &= FPSR_NZCV_MASK;
     env->vfp.fpsr = val;
 }
 
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(vjcvt)(float64 value, CPUARMState *env)
     uint32_t z = (pair >> 32) == 0;
 
     /* Store Z, clear NCV, in FPSCR.NZCV.  */
-    env->vfp.fpsr = (env->vfp.fpsr & ~FPCR_NZCV_MASK) | (z * FPCR_Z);
+    env->vfp.fpsr = (env->vfp.fpsr & ~FPSR_NZCV_MASK) | (z * FPSR_Z);
 
     return result;
 }
-- 
2.34.1

Now that we store FPSR and FPCR separately, the FPSR_MASK and
FPCR_MASK macros are slightly confusingly named and the comment
describing them is out of date.  Rename them to FPSCR_FPSR_MASK and
FPSCR_FPCR_MASK, document that they are the mask of which FPSCR bits
are architecturally mapped to which AArch64 register, and define them
symbolically rather than as hex values.  (This latter requires
defining some extra macros for bits which we haven't previously
defined.)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-9-peter.maydell@linaro.org
---
 target/arm/cpu.h        | 41 ++++++++++++++++++++++++++++++++++-------
 target/arm/machine.c    |  3 ++-
 target/arm/vfp_helper.c |  7 ++++---
 3 files changed, 40 insertions(+), 11 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ static inline void xpsr_write(CPUARMState *env, uint32_t val, uint32_t mask)
 uint32_t vfp_get_fpscr(CPUARMState *env);
 void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 
-/* FPCR, Floating Point Control Register
- * FPSR, Floating Poiht Status Register
+/*
+ * FPCR, Floating Point Control Register
+ * FPSR, Floating Point Status Register
  *
- * For A64 the FPSCR is split into two logically distinct registers,
- * FPCR and FPSR. However since they still use non-overlapping bits
- * we store the underlying state in fpscr and just mask on read/write.
+ * For A64 floating point control and status bits are stored in
+ * two logically distinct registers, FPCR and FPSR. We store these
+ * in QEMU in vfp.fpcr and vfp.fpsr.
+ * For A32 there was only one register, FPSCR. The bits are arranged
+ * such that FPSCR bits map to FPCR or FPSR bits in the same bit positions,
+ * so we can use appropriate masking to handle FPSCR reads and writes.
+ * Note that the FPCR has some bits which are not visible in the
+ * AArch32 view (for FEAT_AFP). Writing the FPSCR leaves these unchanged.
  */
-#define FPSR_MASK 0xf800009f
-#define FPCR_MASK 0x07ff9f00
 
 /* FPCR bits */
 #define FPCR_IOE    (1 << 8)    /* Invalid Operation exception trap enable */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 #define FPCR_UFE    (1 << 11)   /* Underflow exception trap enable */
 #define FPCR_IXE    (1 << 12)   /* Inexact exception trap enable */
 #define FPCR_IDE    (1 << 15)   /* Input Denormal exception trap enable */
+#define FPCR_LEN_MASK (7 << 16) /* LEN, A-profile only */
 #define FPCR_FZ16   (1 << 19)   /* ARMv8.2+, FP16 flush-to-zero */
+#define FPCR_STRIDE_MASK (3 << 20) /* Stride */
 #define FPCR_RMODE_MASK (3 << 22) /* Rounding mode */
 #define FPCR_FZ     (1 << 24)   /* Flush-to-zero enable bit */
 #define FPCR_DN     (1 << 25)   /* Default NaN enable bit */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpscr(CPUARMState *env, uint32_t val);
 #define FPCR_LTPSIZE_MASK (7 << FPCR_LTPSIZE_SHIFT)
 #define FPCR_LTPSIZE_LENGTH 3
 
+/* Cumulative exception trap enable bits */
+#define FPCR_EEXC_MASK (FPCR_IOE | FPCR_DZE | FPCR_OFE | FPCR_UFE | FPCR_IXE | FPCR_IDE)
+
 /* FPSR bits */
+#define FPSR_IOC    (1 << 0)    /* Invalid Operation cumulative exception */
+#define FPSR_DZC    (1 << 1)    /* Divide by Zero cumulative exception */
+#define FPSR_OFC    (1 << 2)    /* Overflow cumulative exception */
+#define FPSR_UFC    (1 << 3)    /* Underflow cumulative exception */
+#define FPSR_IXC    (1 << 4)    /* Inexact cumulative exception */
+#define FPSR_IDC    (1 << 7)    /* Input Denormal cumulative exception */
 #define FPSR_QC     (1 << 27)   /* Cumulative saturation bit */
 #define FPSR_V      (1 << 28)   /* FP overflow flag */
 #define FPSR_C      (1 << 29)   /* FP carry flag */
 #define FPSR_Z      (1 << 30)   /* FP zero flag */
 #define FPSR_N      (1 << 31)   /* FP negative flag */
 
+/* Cumulative exception status bits */
+#define FPSR_CEXC_MASK (FPSR_IOC | FPSR_DZC | FPSR_OFC | FPSR_UFC | FPSR_IXC | FPSR_IDC)
+
 #define FPSR_NZCV_MASK (FPSR_N | FPSR_Z | FPSR_C | FPSR_V)
 #define FPSR_NZCVQC_MASK (FPSR_NZCV_MASK | FPSR_QC)
 
+/* A32 FPSCR bits which architecturally map to FPSR bits */
+#define FPSCR_FPSR_MASK (FPSR_NZCVQC_MASK | FPSR_CEXC_MASK)
+/* A32 FPSCR bits which architecturally map to FPCR bits */
+#define FPSCR_FPCR_MASK (FPCR_EEXC_MASK | FPCR_LEN_MASK | FPCR_FZ16 | \
+                         FPCR_STRIDE_MASK | FPCR_RMODE_MASK | \
+                         FPCR_FZ | FPCR_DN | FPCR_AHP)
+/* These masks don't overlap: each bit lives in only one place */
+QEMU_BUILD_BUG_ON(FPSCR_FPSR_MASK & FPSCR_FPCR_MASK);
+
 /**
  * vfp_get_fpsr: read the AArch64 FPSR
  * @env: CPU context
diff --git a/target/arm/machine.c b/target/arm/machine.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/machine.c
+++ b/target/arm/machine.c
@@ -XXX,XX +XXX,XX @@ static bool vfp_fpcr_fpsr_needed(void *opaque)
     ARMCPU *cpu = opaque;
     CPUARMState *env = &cpu->env;
 
-    return (vfp_get_fpcr(env) & ~FPCR_MASK) || (vfp_get_fpsr(env) & ~FPSR_MASK);
+    return (vfp_get_fpcr(env) & ~FPSCR_FPCR_MASK) ||
+        (vfp_get_fpsr(env) & ~FPSCR_FPSR_MASK);
 }
 
 static int get_fpscr(QEMUFile *f, void *opaque, size_t size,
diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t vfp_get_fpsr(CPUARMState *env)
 
 uint32_t HELPER(vfp_get_fpscr)(CPUARMState *env)
 {
-    return (vfp_get_fpcr(env) & FPCR_MASK) | (vfp_get_fpsr(env) & FPSR_MASK);
+    return (vfp_get_fpcr(env) & FPSCR_FPCR_MASK) |
+        (vfp_get_fpsr(env) & FPSCR_FPSR_MASK);
 }
 
 uint32_t vfp_get_fpscr(CPUARMState *env)
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
 
 void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
 {
-    vfp_set_fpcr(env, val & FPCR_MASK);
-    vfp_set_fpsr(env, val & FPSR_MASK);
+    vfp_set_fpcr(env, val & FPSCR_FPCR_MASK);
+    vfp_set_fpsr(env, val & FPSCR_FPSR_MASK);
 }
 
 void vfp_set_fpscr(CPUARMState *env, uint32_t val)
-- 
2.34.1

In order to allow FPCR bits that aren't in the FPSCR (like the new
bits that are defined for FEAT_AFP), we need to make sure that writes
to the FPSCR only write to the bits of FPCR that are architecturally
mapped, and not the others.

Implement this with a new function vfp_set_fpcr_masked() which
takes a mask of which bits to update.

(We could do the same for FPSR, but we leave that until we actually
are likely to need it.)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20240628142347.1283015-10-peter.maydell@linaro.org
---
 target/arm/vfp_helper.c | 54 ++++++++++++++++++++++++++---------------
 1 file changed, 34 insertions(+), 20 deletions(-)

diff --git a/target/arm/vfp_helper.c b/target/arm/vfp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/vfp_helper.c
+++ b/target/arm/vfp_helper.c
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
     set_float_exception_flags(0, &env->vfp.standard_fp_status_f16);
 }
 
-static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val, uint32_t mask)
 {
     uint64_t changed = env->vfp.fpcr;
 
     changed ^= val;
+    changed &= mask;
     if (changed & (3 << 22)) {
         int i = (val >> 22) & 3;
         switch (i) {
@@ -XXX,XX +XXX,XX @@ static void vfp_set_fpsr_to_host(CPUARMState *env, uint32_t val)
 {
 }
 
-static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val)
+static void vfp_set_fpcr_to_host(CPUARMState *env, uint32_t val, uint32_t mask)
 {
 }
 
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpsr(CPUARMState *env, uint32_t val)
     env->vfp.fpsr = val;
 }
 
-void vfp_set_fpcr(CPUARMState *env, uint32_t val)
+static void vfp_set_fpcr_masked(CPUARMState *env, uint32_t val, uint32_t mask)
 {
+    /*
+     * We only set FPCR bits defined by mask, and leave the others alone.
+     * We assume the mask is sensible (e.g. doesn't try to set only
+     * part of a field)
+     */
     ARMCPU *cpu = env_archcpu(env);
 
     /* When ARMv8.2-FP16 is not supported, FZ16 is RES0.  */
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
         val &= ~FPCR_FZ16;
     }
 
-    vfp_set_fpcr_to_host(env, val);
+    vfp_set_fpcr_to_host(env, val, mask);
 
-    if (!arm_feature(env, ARM_FEATURE_M)) {
-        /*
-         * Short-vector length and stride; on M-profile these bits
-         * are used for different purposes.
-         * We can't make this conditional be "if MVFR0.FPShVec != 0",
-         * because in v7A no-short-vector-support cores still had to
-         * allow Stride/Len to be written with the only effect that
-         * some insns are required to UNDEF if the guest sets them.
-         */
-        env->vfp.vec_len = extract32(val, 16, 3);
-        env->vfp.vec_stride = extract32(val, 20, 2);
-    } else if (cpu_isar_feature(aa32_mve, cpu)) {
-        env->v7m.ltpsize = extract32(val, FPCR_LTPSIZE_SHIFT,
-                                     FPCR_LTPSIZE_LENGTH);
+    if (mask & (FPCR_LEN_MASK | FPCR_STRIDE_MASK)) {
+        if (!arm_feature(env, ARM_FEATURE_M)) {
+            /*
+             * Short-vector length and stride; on M-profile these bits
+             * are used for different purposes.
+             * We can't make this conditional be "if MVFR0.FPShVec != 0",
+             * because in v7A no-short-vector-support cores still had to
+             * allow Stride/Len to be written with the only effect that
+             * some insns are required to UNDEF if the guest sets them.
+             */
+            env->vfp.vec_len = extract32(val, 16, 3);
+            env->vfp.vec_stride = extract32(val, 20, 2);
+        } else if (cpu_isar_feature(aa32_mve, cpu)) {
+            env->v7m.ltpsize = extract32(val, FPCR_LTPSIZE_SHIFT,
+                                         FPCR_LTPSIZE_LENGTH);
+        }
     }
 
     /*
@@ -XXX,XX +XXX,XX @@ void vfp_set_fpcr(CPUARMState *env, uint32_t val)
      * bits.
      */
     val &= FPCR_AHP | FPCR_DN | FPCR_FZ | FPCR_RMODE_MASK | FPCR_FZ16;
-    env->vfp.fpcr = val;
+    env->vfp.fpcr &= ~mask;
+    env->vfp.fpcr |= val;
+}
+
+void vfp_set_fpcr(CPUARMState *env, uint32_t val)
+{
+    vfp_set_fpcr_masked(env, val, MAKE_64BIT_MASK(0, 32));
 }
 
 void HELPER(vfp_set_fpscr)(CPUARMState *env, uint32_t val)
 {
-    vfp_set_fpcr(env, val & FPSCR_FPCR_MASK);
+    vfp_set_fpcr_masked(env, val, FPSCR_FPCR_MASK);
     vfp_set_fpsr(env, val & FPSCR_FPSR_MASK);
 }
 
-- 
2.34.1

From: Zheyu Ma <zheyuma97@gmail.com>

In pl011_get_baudrate(), when we calculate the baudrate we can
accidentally divide by zero. This happens because although (as the
specification requires) we treat UARTIBRD = 0 as invalid, we aren't
correctly limiting UARTIBRD and UARTFBRD values to the 16-bit and 6-bit
ranges the hardware allows, and so some non-zero values of UARTIBRD can
result in a zero divisor.

Enforce the correct register field widths on guest writes and on inbound
migration to avoid the division by zero.

ASAN log:
==2973125==ERROR: AddressSanitizer: FPE on unknown address 0x55f72629b348
(pc 0x55f72629b348 bp 0x7fffa24d0e00 sp 0x7fffa24d0d60 T0)
     #0 0x55f72629b348 in pl011_get_baudrate hw/char/pl011.c:255:17
     #1 0x55f726298d94 in pl011_trace_baudrate_change hw/char/pl011.c:260:33
     #2 0x55f726296fc8 in pl011_write hw/char/pl011.c:378:9

Reproducer:
cat << EOF | qemu-system-aarch64 -display \
none -machine accel=qtest, -m 512M -machine realview-pb-a8 -qtest stdio
writeq 0x1000b024 0xf8000000
EOF

Suggested-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Zheyu Ma <zheyuma97@gmail.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20240702155752.3022007-1-zheyuma97@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/char/pl011.c | 13 +++++++++++--
 1 file changed, 11 insertions(+), 2 deletions(-)

diff --git a/hw/char/pl011.c b/hw/char/pl011.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/char/pl011.c
+++ b/hw/char/pl011.c
@@ -XXX,XX +XXX,XX @@ DeviceState *pl011_create(hwaddr addr, qemu_irq irq, Chardev *chr)
 #define CR_DTR      (1 << 10)
 #define CR_LBE      (1 << 7)
 
+/* Integer Baud Rate Divider, UARTIBRD */
+#define IBRD_MASK 0x3f
+
+/* Fractional Baud Rate Divider, UARTFBRD */
+#define FBRD_MASK 0xffff
+
 static const unsigned char pl011_id_arm[8] =
   { 0x11, 0x10, 0x14, 0x00, 0x0d, 0xf0, 0x05, 0xb1 };
 static const unsigned char pl011_id_luminary[8] =
@@ -XXX,XX +XXX,XX @@ static void pl011_write(void *opaque, hwaddr offset,
         s->ilpr = value;
         break;
     case 9: /* UARTIBRD */
-        s->ibrd = value;
+        s->ibrd = value & IBRD_MASK;
         pl011_trace_baudrate_change(s);
         break;
     case 10: /* UARTFBRD */
-        s->fbrd = value;
+        s->fbrd = value & FBRD_MASK;
         pl011_trace_baudrate_change(s);
         break;
     case 11: /* UARTLCR_H */
@@ -XXX,XX +XXX,XX @@ static int pl011_post_load(void *opaque, int version_id)
         s->read_pos = 0;
     }
 
+    s->ibrd &= IBRD_MASK;
+    s->fbrd &= FBRD_MASK;
+
     return 0;
 }
 
-- 
2.34.1

From: Zheyu Ma <zheyuma97@gmail.com>

The current implementation of bcm2835_thermal_ops sets
impl.max_access_size and valid.min_access_size to 4, but leaves
impl.min_access_size and valid.max_access_size unset, defaulting to 1.
This causes issues when the memory system is presented with an access
of size 2 at an offset of 3, leading to an attempt to synthesize it as
a pair of byte accesses at offsets 3 and 4, which trips an assert.

Additionally, the lack of valid.max_access_size setting causes another
issue: the memory system tries to synthesize a read using a 4-byte
access at offset 3 even though the device doesn't allow unaligned
accesses.

This patch addresses these issues by explicitly setting both
impl.min_access_size and valid.max_access_size to 4, ensuring proper
handling of access sizes.

Error log:
ERROR:hw/misc/bcm2835_thermal.c:55:bcm2835_thermal_read: code should not be reached
Bail out! ERROR:hw/misc/bcm2835_thermal.c:55:bcm2835_thermal_read: code should not be reached
Aborted

Reproducer:
cat << EOF | qemu-system-aarch64 -display \
none -machine accel=qtest, -m 512M -machine raspi3b -m 1G -qtest stdio
readw 0x3f212003
EOF

Signed-off-by: Zheyu Ma <zheyuma97@gmail.com>
Message-id: 20240702154042.3018932-1-zheyuma97@gmail.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/misc/bcm2835_thermal.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/hw/misc/bcm2835_thermal.c b/hw/misc/bcm2835_thermal.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/misc/bcm2835_thermal.c
+++ b/hw/misc/bcm2835_thermal.c
@@ -XXX,XX +XXX,XX @@ static void bcm2835_thermal_write(void *opaque, hwaddr addr,
 static const MemoryRegionOps bcm2835_thermal_ops = {
     .read = bcm2835_thermal_read,
     .write = bcm2835_thermal_write,
+    .impl.min_access_size = 4,
     .impl.max_access_size = 4,
     .valid.min_access_size = 4,
+    .valid.max_access_size = 4,
     .endianness = DEVICE_NATIVE_ENDIAN,
 };
 
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

In a completely artifical memset benchmark object_dynamic_cast_assert
dominates the profile, even above guest address resolution and
the underlying host memset.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20240702154911.1667418-1-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/cpu.h | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ extern const uint64_t pred_esz_masks[5];
  */
 static inline target_ulong cpu_untagged_addr(CPUState *cs, target_ulong x)
 {
-    ARMCPU *cpu = ARM_CPU(cs);
-    if (cpu->env.tagged_addr_enable) {
+    CPUARMState *env = cpu_env(cs);
+    if (env->tagged_addr_enable) {
         /*
          * TBI is enabled for userspace but not kernelspace addresses.
          * Only clear the tag if bit 55 is clear.
-- 
2.34.1

In commit a96edb687e76 we set the cpu_exec_halt field of the
TCGCPUOps arm_tcg_ops to arm_cpu_exec_halt(), but we left the
arm_v7m_tcg_ops struct unchanged.  That isn't wrong, because for
M-profile FEAT_WFxT doesn't exist and the default handling for "no
cpu_exec_halt method" is correct, but it's perhaps a little
confusing.  We would also like to make setting the cpu_exec_halt
method mandatory.

Initialize arm_v7m_tcg_ops cpu_exec_halt to the same function we use
for A-profile.  (On M-profile we never set up the wfxt timer so there
is no change in behaviour here.)

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
---
 target/arm/internals.h   | 3 +++
 target/arm/cpu.c         | 2 +-
 target/arm/tcg/cpu-v7m.c | 1 +
 3 files changed, 5 insertions(+), 1 deletion(-)

diff --git a/target/arm/internals.h b/target/arm/internals.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/internals.h
+++ b/target/arm/internals.h
@@ -XXX,XX +XXX,XX @@ void arm_restore_state_to_opc(CPUState *cs,
 
 #ifdef CONFIG_TCG
 void arm_cpu_synchronize_from_tb(CPUState *cs, const TranslationBlock *tb);
+
+/* Our implementation of TCGCPUOps::cpu_exec_halt */
+bool arm_cpu_exec_halt(CPUState *cs);
 #endif /* CONFIG_TCG */
 
 typedef enum ARMFPRounding {
diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ static bool arm_cpu_virtio_is_big_endian(CPUState *cs)
 }
 
 #ifdef CONFIG_TCG
-static bool arm_cpu_exec_halt(CPUState *cs)
+bool arm_cpu_exec_halt(CPUState *cs)
 {
     bool leave_halt = cpu_has_work(cs);
 
diff --git a/target/arm/tcg/cpu-v7m.c b/target/arm/tcg/cpu-v7m.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/cpu-v7m.c
+++ b/target/arm/tcg/cpu-v7m.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps arm_v7m_tcg_ops = {
 #else
     .tlb_fill = arm_cpu_tlb_fill,
     .cpu_exec_interrupt = arm_v7m_cpu_exec_interrupt,
+    .cpu_exec_halt = arm_cpu_exec_halt,
     .do_interrupt = arm_v7m_cpu_do_interrupt,
     .do_transaction_failed = arm_cpu_do_transaction_failed,
     .do_unaligned_access = arm_cpu_do_unaligned_access,
-- 
2.34.1

Currently the TCGCPUOps::cpu_exec_halt method is optional, and if it
is not set then the default is to call the CPUClass::has_work
method (which has an identical function signature).

We would like to make the cpu_exec_halt method mandatory so we can
remove the runtime check and fallback handling.  In preparation for
that, make all the targets which don't need special handling in their
cpu_exec_halt set it to their cpu_has_work implementation instead of
leaving it unset.  (This is every target except for arm and i386.)

In the riscv case this requires us to make the function not
be local to the source file it's defined in.

diff --git a/target/riscv/internals.h b/target/riscv/internals.h
index XXXXXXX..XXXXXXX 100644
--- a/target/riscv/internals.h
+++ b/target/riscv/internals.h
@@ -XXX,XX +XXX,XX @@ static inline float16 check_nanbox_h(CPURISCVState *env, uint64_t f)
     }
 }
 
+/* Our implementation of CPUClass::has_work */
+bool riscv_cpu_has_work(CPUState *cs);
+
 #endif
diff --git a/target/alpha/cpu.c b/target/alpha/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/alpha/cpu.c
+++ b/target/alpha/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps alpha_tcg_ops = {
 #else
     .tlb_fill = alpha_cpu_tlb_fill,
     .cpu_exec_interrupt = alpha_cpu_exec_interrupt,
+    .cpu_exec_halt = alpha_cpu_has_work,
     .do_interrupt = alpha_cpu_do_interrupt,
     .do_transaction_failed = alpha_cpu_do_transaction_failed,
     .do_unaligned_access = alpha_cpu_do_unaligned_access,
diff --git a/target/avr/cpu.c b/target/avr/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/avr/cpu.c
+++ b/target/avr/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps avr_tcg_ops = {
     .synchronize_from_tb = avr_cpu_synchronize_from_tb,
     .restore_state_to_opc = avr_restore_state_to_opc,
     .cpu_exec_interrupt = avr_cpu_exec_interrupt,
+    .cpu_exec_halt = avr_cpu_has_work,
     .tlb_fill = avr_cpu_tlb_fill,
     .do_interrupt = avr_cpu_do_interrupt,
 };
diff --git a/target/cris/cpu.c b/target/cris/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/cris/cpu.c
+++ b/target/cris/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps crisv10_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = cris_cpu_tlb_fill,
     .cpu_exec_interrupt = cris_cpu_exec_interrupt,
+    .cpu_exec_halt = cris_cpu_has_work,
     .do_interrupt = crisv10_cpu_do_interrupt,
 #endif /* !CONFIG_USER_ONLY */
 };
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps crisv32_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = cris_cpu_tlb_fill,
     .cpu_exec_interrupt = cris_cpu_exec_interrupt,
+    .cpu_exec_halt = cris_cpu_has_work,
     .do_interrupt = cris_cpu_do_interrupt,
 #endif /* !CONFIG_USER_ONLY */
 };
diff --git a/target/hppa/cpu.c b/target/hppa/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/cpu.c
+++ b/target/hppa/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps hppa_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = hppa_cpu_tlb_fill,
     .cpu_exec_interrupt = hppa_cpu_exec_interrupt,
+    .cpu_exec_halt = hppa_cpu_has_work,
     .do_interrupt = hppa_cpu_do_interrupt,
     .do_unaligned_access = hppa_cpu_do_unaligned_access,
     .do_transaction_failed = hppa_cpu_do_transaction_failed,
diff --git a/target/loongarch/cpu.c b/target/loongarch/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/cpu.c
+++ b/target/loongarch/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps loongarch_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = loongarch_cpu_tlb_fill,
     .cpu_exec_interrupt = loongarch_cpu_exec_interrupt,
+    .cpu_exec_halt = loongarch_cpu_has_work,
     .do_interrupt = loongarch_cpu_do_interrupt,
     .do_transaction_failed = loongarch_cpu_do_transaction_failed,
 #endif
diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps m68k_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = m68k_cpu_tlb_fill,
     .cpu_exec_interrupt = m68k_cpu_exec_interrupt,
+    .cpu_exec_halt = m68k_cpu_has_work,
     .do_interrupt = m68k_cpu_do_interrupt,
     .do_transaction_failed = m68k_cpu_transaction_failed,
 #endif /* !CONFIG_USER_ONLY */
diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/microblaze/cpu.c
+++ b/target/microblaze/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps mb_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = mb_cpu_tlb_fill,
     .cpu_exec_interrupt = mb_cpu_exec_interrupt,
+    .cpu_exec_halt = mb_cpu_has_work,
     .do_interrupt = mb_cpu_do_interrupt,
     .do_transaction_failed = mb_cpu_transaction_failed,
     .do_unaligned_access = mb_cpu_do_unaligned_access,
diff --git a/target/mips/cpu.c b/target/mips/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/cpu.c
+++ b/target/mips/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps mips_tcg_ops = {
 #if !defined(CONFIG_USER_ONLY)
     .tlb_fill = mips_cpu_tlb_fill,
     .cpu_exec_interrupt = mips_cpu_exec_interrupt,
+    .cpu_exec_halt = mips_cpu_has_work,
     .do_interrupt = mips_cpu_do_interrupt,
     .do_transaction_failed = mips_cpu_do_transaction_failed,
     .do_unaligned_access = mips_cpu_do_unaligned_access,
diff --git a/target/openrisc/cpu.c b/target/openrisc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/openrisc/cpu.c
+++ b/target/openrisc/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps openrisc_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = openrisc_cpu_tlb_fill,
     .cpu_exec_interrupt = openrisc_cpu_exec_interrupt,
+    .cpu_exec_halt = openrisc_cpu_has_work,
     .do_interrupt = openrisc_cpu_do_interrupt,
 #endif /* !CONFIG_USER_ONLY */
 };
diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@
+
 /*
  *  PowerPC CPU initialization for qemu.
  *
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps ppc_tcg_ops = {
 #else
   .tlb_fill = ppc_cpu_tlb_fill,
   .cpu_exec_interrupt = ppc_cpu_exec_interrupt,
+  .cpu_exec_halt = ppc_cpu_has_work,
   .do_interrupt = ppc_cpu_do_interrupt,
   .cpu_exec_enter = ppc_cpu_exec_enter,
   .cpu_exec_exit = ppc_cpu_exec_exit,
diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/riscv/cpu.c
+++ b/target/riscv/cpu.c
@@ -XXX,XX +XXX,XX @@ static vaddr riscv_cpu_get_pc(CPUState *cs)
     return env->pc;
 }
 
-static bool riscv_cpu_has_work(CPUState *cs)
+bool riscv_cpu_has_work(CPUState *cs)
 {
 #ifndef CONFIG_USER_ONLY
     RISCVCPU *cpu = RISCV_CPU(cs);
diff --git a/target/riscv/tcg/tcg-cpu.c b/target/riscv/tcg/tcg-cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/riscv/tcg/tcg-cpu.c
+++ b/target/riscv/tcg/tcg-cpu.c
@@ -XXX,XX +XXX,XX @@
 #include "exec/exec-all.h"
 #include "tcg-cpu.h"
 #include "cpu.h"
+#include "internals.h"
 #include "pmu.h"
 #include "time_helper.h"
 #include "qapi/error.h"
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps riscv_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = riscv_cpu_tlb_fill,
     .cpu_exec_interrupt = riscv_cpu_exec_interrupt,
+    .cpu_exec_halt = riscv_cpu_has_work,
     .do_interrupt = riscv_cpu_do_interrupt,
     .do_transaction_failed = riscv_cpu_do_transaction_failed,
     .do_unaligned_access = riscv_cpu_do_unaligned_access,
diff --git a/target/rx/cpu.c b/target/rx/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/rx/cpu.c
+++ b/target/rx/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps rx_tcg_ops = {
 
 #ifndef CONFIG_USER_ONLY
     .cpu_exec_interrupt = rx_cpu_exec_interrupt,
+    .cpu_exec_halt = rx_cpu_has_work,
     .do_interrupt = rx_cpu_do_interrupt,
 #endif /* !CONFIG_USER_ONLY */
 };
diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps s390_tcg_ops = {
 #else
     .tlb_fill = s390_cpu_tlb_fill,
     .cpu_exec_interrupt = s390_cpu_exec_interrupt,
+    .cpu_exec_halt = s390_cpu_has_work,
     .do_interrupt = s390_cpu_do_interrupt,
     .debug_excp_handler = s390x_cpu_debug_excp_handler,
     .do_unaligned_access = s390x_cpu_do_unaligned_access,
diff --git a/target/sh4/cpu.c b/target/sh4/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sh4/cpu.c
+++ b/target/sh4/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps superh_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = superh_cpu_tlb_fill,
     .cpu_exec_interrupt = superh_cpu_exec_interrupt,
+    .cpu_exec_halt = superh_cpu_has_work,
     .do_interrupt = superh_cpu_do_interrupt,
     .do_unaligned_access = superh_cpu_do_unaligned_access,
     .io_recompile_replay_branch = superh_io_recompile_replay_branch,
diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps sparc_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = sparc_cpu_tlb_fill,
     .cpu_exec_interrupt = sparc_cpu_exec_interrupt,
+    .cpu_exec_halt = sparc_cpu_has_work,
     .do_interrupt = sparc_cpu_do_interrupt,
     .do_transaction_failed = sparc_cpu_do_transaction_failed,
     .do_unaligned_access = sparc_cpu_do_unaligned_access,
diff --git a/target/tricore/cpu.c b/target/tricore/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/tricore/cpu.c
+++ b/target/tricore/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps tricore_tcg_ops = {
     .synchronize_from_tb = tricore_cpu_synchronize_from_tb,
     .restore_state_to_opc = tricore_restore_state_to_opc,
     .tlb_fill = tricore_cpu_tlb_fill,
+    .cpu_exec_halt = tricore_cpu_has_work,
 };
 
 static void tricore_cpu_class_init(ObjectClass *c, void *data)
diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/cpu.c
+++ b/target/xtensa/cpu.c
@@ -XXX,XX +XXX,XX @@ static const TCGCPUOps xtensa_tcg_ops = {
 #ifndef CONFIG_USER_ONLY
     .tlb_fill = xtensa_cpu_tlb_fill,
     .cpu_exec_interrupt = xtensa_cpu_exec_interrupt,
+    .cpu_exec_halt = xtensa_cpu_has_work,
     .do_interrupt = xtensa_cpu_do_interrupt,
     .do_transaction_failed = xtensa_cpu_do_transaction_failed,
     .do_unaligned_access = xtensa_cpu_do_unaligned_access,
-- 
2.34.1

Now that all targets set TCGCPUOps::cpu_exec_halt, we can make it
mandatory and remove the fallback handling that calls cpu_has_work.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
---
 include/hw/core/tcg-cpu-ops.h |  9 ++++++---
 accel/tcg/cpu-exec.c          | 11 +++++------
 2 files changed, 11 insertions(+), 9 deletions(-)

diff --git a/include/hw/core/tcg-cpu-ops.h b/include/hw/core/tcg-cpu-ops.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/core/tcg-cpu-ops.h
+++ b/include/hw/core/tcg-cpu-ops.h
@@ -XXX,XX +XXX,XX @@ struct TCGCPUOps {
      * to do when the CPU is in the halted state.
      *
      * Return true to indicate that the CPU should now leave halt, false
-     * if it should remain in the halted state.
+     * if it should remain in the halted state. (This should generally
+     * be the same value that cpu_has_work() would return.)
      *
-     * If this method is not provided, the default is to do nothing, and
-     * to leave halt if cpu_has_work() returns true.
+     * This method must be provided. If the target does not need to
+     * do anything special for halt, the same function used for its
+     * CPUClass::has_work method can be used here, as they have the
+     * same function signature.
      */
     bool (*cpu_exec_halt)(CPUState *cpu);
     /**
diff --git a/accel/tcg/cpu-exec.c b/accel/tcg/cpu-exec.c
index XXXXXXX..XXXXXXX 100644
--- a/accel/tcg/cpu-exec.c
+++ b/accel/tcg/cpu-exec.c
@@ -XXX,XX +XXX,XX @@ static inline bool cpu_handle_halt(CPUState *cpu)
 #ifndef CONFIG_USER_ONLY
     if (cpu->halted) {
         const TCGCPUOps *tcg_ops = cpu->cc->tcg_ops;
-        bool leave_halt;
+        bool leave_halt = tcg_ops->cpu_exec_halt(cpu);
 
-        if (tcg_ops->cpu_exec_halt) {
-            leave_halt = tcg_ops->cpu_exec_halt(cpu);
-        } else {
-            leave_halt = cpu_has_work(cpu);
-        }
         if (!leave_halt) {
             return true;
         }
@@ -XXX,XX +XXX,XX @@ bool tcg_exec_realizefn(CPUState *cpu, Error **errp)
     static bool tcg_target_initialized;
 
     if (!tcg_target_initialized) {
+        /* Check mandatory TCGCPUOps handlers */
+#ifndef CONFIG_USER_ONLY
+        assert(cpu->cc->tcg_ops->cpu_exec_halt);
+#endif /* !CONFIG_USER_ONLY */
         cpu->cc->tcg_ops->initialize();
         tcg_target_initialized = true;
     }
-- 
2.34.1

From: Inès Varhol <ines.varhol@telecom-paris.fr>

Up until now, the EXTI implementation had 16 inbound GPIOs connected to
the 16 outbound GPIOs of STM32L4x5 SYSCFG.
The EXTI actually handles 40 lines (namely 5 from STM32L4x5 USART
devices which are already implemented in QEMU).
In order to connect USART devices to EXTI, this commit consolidates
constants `EXTI_NUM_INTERRUPT_OUT_LINES` (40) and
`EXTI_NUM_GPIO_EVENT_IN_LINES` (16) into `EXTI_NUM_LINES` (40).

Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240707085927.122867-2-ines.varhol@telecom-paris.fr
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/misc/stm32l4x5_exti.h | 4 ++--
 hw/misc/stm32l4x5_exti.c         | 6 ++----
 2 files changed, 4 insertions(+), 6 deletions(-)

diff --git a/include/hw/misc/stm32l4x5_exti.h b/include/hw/misc/stm32l4x5_exti.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/misc/stm32l4x5_exti.h
+++ b/include/hw/misc/stm32l4x5_exti.h
@@ -XXX,XX +XXX,XX @@
 #define TYPE_STM32L4X5_EXTI "stm32l4x5-exti"
 OBJECT_DECLARE_SIMPLE_TYPE(Stm32l4x5ExtiState, STM32L4X5_EXTI)
 
-#define EXTI_NUM_INTERRUPT_OUT_LINES 40
+#define EXTI_NUM_LINES 40
 #define EXTI_NUM_REGISTER 2
 
 struct Stm32l4x5ExtiState {
@@ -XXX,XX +XXX,XX @@ struct Stm32l4x5ExtiState {
 
     /* used for edge detection */
     uint32_t irq_levels[EXTI_NUM_REGISTER];
-    qemu_irq irq[EXTI_NUM_INTERRUPT_OUT_LINES];
+    qemu_irq irq[EXTI_NUM_LINES];
 };
 
 #endif
diff --git a/hw/misc/stm32l4x5_exti.c b/hw/misc/stm32l4x5_exti.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/misc/stm32l4x5_exti.c
+++ b/hw/misc/stm32l4x5_exti.c
@@ -XXX,XX +XXX,XX @@
 #define EXTI_SWIER2 0x30
 #define EXTI_PR2    0x34
 
-#define EXTI_NUM_GPIO_EVENT_IN_LINES 16
 #define EXTI_MAX_IRQ_PER_BANK 32
 #define EXTI_IRQS_BANK0  32
 #define EXTI_IRQS_BANK1  8
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_init(Object *obj)
 {
     Stm32l4x5ExtiState *s = STM32L4X5_EXTI(obj);
 
-    for (size_t i = 0; i < EXTI_NUM_INTERRUPT_OUT_LINES; i++) {
+    for (size_t i = 0; i < EXTI_NUM_LINES; i++) {
         sysbus_init_irq(SYS_BUS_DEVICE(obj), &s->irq[i]);
     }
 
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_init(Object *obj)
                           TYPE_STM32L4X5_EXTI, 0x400);
     sysbus_init_mmio(SYS_BUS_DEVICE(obj), &s->mmio);
 
-    qdev_init_gpio_in(DEVICE(obj), stm32l4x5_exti_set_irq,
-                      EXTI_NUM_GPIO_EVENT_IN_LINES);
+    qdev_init_gpio_in(DEVICE(obj), stm32l4x5_exti_set_irq, EXTI_NUM_LINES);
 }
 
 static const VMStateDescription vmstate_stm32l4x5_exti = {
-- 
2.34.1

From: Inès Varhol <ines.varhol@telecom-paris.fr>

The previous implementation for EXTI interrupts only handled
"configurable" interrupts, like those originating from STM32L4x5 SYSCFG
(the only device currently connected to the EXTI up until now).

In order to connect STM32L4x5 USART to the EXTI, this commit adds
handling for direct interrupts (interrupts without configurable edge).

Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
Message-id: 20240707085927.122867-3-ines.varhol@telecom-paris.fr
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/misc/stm32l4x5_exti.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/hw/misc/stm32l4x5_exti.c b/hw/misc/stm32l4x5_exti.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/misc/stm32l4x5_exti.c
+++ b/hw/misc/stm32l4x5_exti.c
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_exti_set_irq(void *opaque, int irq, int level)
         return;
     }
 
+    /* In case of a direct line interrupt */
+    if (extract32(exti_romask[bank], irq, 1)) {
+        qemu_set_irq(s->irq[oirq], level);
+        return;
+    }
+
+    /* In case of a configurable interrupt */
     if ((level && extract32(s->rtsr[bank], irq, 1)) ||
         (!level && extract32(s->ftsr[bank], irq, 1))) {
 
-- 
2.34.1

From: Inès Varhol <ines.varhol@telecom-paris.fr>

The USART devices were previously connecting their outbound IRQs
directly to the CPU because the EXTI wasn't handling direct lines
interrupts.
Now the USART connects to the EXTI inbound GPIOs, and the EXTI connects
its IRQs to the CPU.
The existing QTest for the USART (tests/qtest/stm32l4x5_usart-test.c)
checks that USART1_IRQ in the CPU is pending when expected so it
confirms that the connection through the EXTI still works.

Signed-off-by: Inès Varhol <ines.varhol@telecom-paris.fr>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240707085927.122867-4-ines.varhol@telecom-paris.fr
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/stm32l4x5_soc.c | 24 +++++++++++-------------
 1 file changed, 11 insertions(+), 13 deletions(-)

diff --git a/hw/arm/stm32l4x5_soc.c b/hw/arm/stm32l4x5_soc.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/stm32l4x5_soc.c
+++ b/hw/arm/stm32l4x5_soc.c
@@ -XXX,XX +XXX,XX @@ static const int exti_irq[NUM_EXTI_IRQ] = {
 #define RCC_BASE_ADDRESS 0x40021000
 #define RCC_IRQ 5
 
+#define EXTI_USART1_IRQ 26
+#define EXTI_UART4_IRQ 29
+#define EXTI_LPUART1_IRQ 31
+
 static const int exti_or_gates_out[NUM_EXTI_OR_GATES] = {
     23, 40, 63, 1,
 };
@@ -XXX,XX +XXX,XX @@ static const hwaddr uart_addr[] = {
 
 #define LPUART_BASE_ADDRESS 0x40008000
 
-static const int usart_irq[] = { 37, 38, 39 };
-static const int uart_irq[] = { 52, 53 };
-#define LPUART_IRQ 70
-
 static void stm32l4x5_soc_initfn(Object *obj)
 {
     Stm32l4x5SocState *s = STM32L4X5_SOC(obj);
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
         }
     }
 
+    /* Connect SYSCFG to EXTI */
     for (unsigned i = 0; i < GPIO_NUM_PINS; i++) {
         qdev_connect_gpio_out(DEVICE(&s->syscfg), i,
                               qdev_get_gpio_in(DEVICE(&s->exti), i));
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
             return;
         }
         sysbus_mmio_map(busdev, 0, usart_addr[i]);
-        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, usart_irq[i]));
+        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
+                                                       EXTI_USART1_IRQ + i));
     }
 
-    /*
-     * TODO: Connect the USARTs, UARTs and LPUART to the EXTI once the EXTI
-     * can handle other gpio-in than the gpios. (e.g. Direct Lines for the
-     * usarts)
-     */
-
     /* UART devices */
     for (int i = 0; i < STM_NUM_UARTS; i++) {
         g_autofree char *name = g_strdup_printf("uart%d-out", STM_NUM_USARTS + i + 1);
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
             return;
         }
         sysbus_mmio_map(busdev, 0, uart_addr[i]);
-        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, uart_irq[i]));
+        sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
+                                                       EXTI_UART4_IRQ + i));
     }
 
     /* LPUART device*/
@@ -XXX,XX +XXX,XX @@ static void stm32l4x5_soc_realize(DeviceState *dev_soc, Error **errp)
         return;
     }
     sysbus_mmio_map(busdev, 0, LPUART_BASE_ADDRESS);
-    sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(armv7m, LPUART_IRQ));
+    sysbus_connect_irq(busdev, 0, qdev_get_gpio_in(DEVICE(&s->exti),
+                                                   EXTI_LPUART1_IRQ));
 
     /* APB1 BUS */
     create_unimplemented_device("TIM2",      0x40000000, 0x400);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-2-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |  22 ++++
 target/arm/tcg/translate-a64.c | 184 ++++++++++++++++++++++++---------
 2 files changed, 156 insertions(+), 50 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-3-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |   9 ++
 target/arm/tcg/translate-a64.c | 150 +++++++++++++++++----------------
 2 files changed, 87 insertions(+), 72 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-4-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |  33 ++
 target/arm/tcg/translate-a64.c | 604 ++++++---------------------------
 2 files changed, 138 insertions(+), 499 deletions(-)

diff --git a/target/arm/tcg/a64.decode b/target/arm/tcg/a64.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/a64.decode
+++ b/target/arm/tcg/a64.decode
@@ -XXX,XX +XXX,XX @@ SQRDMULH_s      0111 1110 ..1 ..... 10110 1 ..... ..... @rrr_e
 SQRDMLAH_s      0111 1110 ..0 ..... 10000 1 ..... ..... @rrr_e
 SQRDMLSH_s      0111 1110 ..0 ..... 10001 1 ..... ..... @rrr_e
 
+# Decode scalar x scalar as scalar x indexed, with index 0.
+SQDMULL_si      0101 1110 011 rm:5  11010 0 rn:5  rd:5  &rrx_e idx=0 esz=1
+SQDMULL_si      0101 1110 101 rm:5  11010 0 rn:5  rd:5  &rrx_e idx=0 esz=2
+SQDMLAL_si      0101 1110 011 rm:5  10010 0 rn:5  rd:5  &rrx_e idx=0 esz=1
+SQDMLAL_si      0101 1110 101 rm:5  10010 0 rn:5  rd:5  &rrx_e idx=0 esz=2
+SQDMLSL_si      0101 1110 011 rm:5  10110 0 rn:5  rd:5  &rrx_e idx=0 esz=1
+SQDMLSL_si      0101 1110 101 rm:5  10110 0 rn:5  rd:5  &rrx_e idx=0 esz=2
+
 ### Advanced SIMD scalar pairwise
 
 FADDP_s         0101 1110 0011 0000 1101 10 ..... ..... @rr_h
@@ -XXX,XX +XXX,XX @@ UABAL_v         0.10 1110 ..1 ..... 01010 0 ..... ..... @qrrr_e
 SABDL_v         0.00 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
 UABDL_v         0.10 1110 ..1 ..... 01110 0 ..... ..... @qrrr_e
 
+SQDMULL_v       0.00 1110 011 ..... 11010 0 ..... ..... @qrrr_h
+SQDMULL_v       0.00 1110 101 ..... 11010 0 ..... ..... @qrrr_s
+SQDMLAL_v       0.00 1110 011 ..... 10010 0 ..... ..... @qrrr_h
+SQDMLAL_v       0.00 1110 101 ..... 10010 0 ..... ..... @qrrr_s
+SQDMLSL_v       0.00 1110 011 ..... 10110 0 ..... ..... @qrrr_h
+SQDMLSL_v       0.00 1110 101 ..... 10110 0 ..... ..... @qrrr_s
+
 ### Advanced SIMD scalar x indexed element
 
 FMUL_si         0101 1111 00 .. .... 1001 . 0 ..... .....   @rrx_h
@@ -XXX,XX +XXX,XX @@ SQRDMLAH_si     0111 1111 10 .. .... 1101 . 0 ..... .....   @rrx_s
 SQRDMLSH_si     0111 1111 01 .. .... 1111 . 0 ..... .....   @rrx_h
 SQRDMLSH_si     0111 1111 10 .. .... 1111 . 0 ..... .....   @rrx_s
 
+SQDMULL_si      0101 1111 01 .. .... 1011 . 0 ..... .....   @rrx_h
+SQDMULL_si      0101 1111 10 . ..... 1011 . 0 ..... .....   @rrx_s
+
+SQDMLAL_si      0101 1111 01 .. .... 0011 . 0 ..... .....   @rrx_h
+SQDMLAL_si      0101 1111 10 . ..... 0011 . 0 ..... .....   @rrx_s
+
+SQDMLSL_si      0101 1111 01 .. .... 0111 . 0 ..... .....   @rrx_h
+SQDMLSL_si      0101 1111 10 . ..... 0111 . 0 ..... .....   @rrx_s
+
 ### Advanced SIMD vector x indexed element
 
 FMUL_vi         0.00 1111 00 .. .... 1001 . 0 ..... .....   @qrrx_h
@@ -XXX,XX +XXX,XX @@ SMLSL_vi        0.00 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
 UMLSL_vi        0.10 1111 01 .. .... 0110 . 0 ..... .....   @qrrx_h
 UMLSL_vi        0.10 1111 10 . ..... 0110 . 0 ..... .....   @qrrx_s
 
+SQDMULL_vi      0.00 1111 01 .. .... 1011 . 0 ..... .....   @qrrx_h
+SQDMULL_vi      0.00 1111 10 . ..... 1011 . 0 ..... .....   @qrrx_s
+
+SQDMLAL_vi      0.00 1111 01 .. .... 0011 . 0 ..... .....   @qrrx_h
+SQDMLAL_vi      0.00 1111 10 . ..... 0011 . 0 ..... .....   @qrrx_s
+
+SQDMLSL_vi      0.00 1111 01 .. .... 0111 . 0 ..... .....   @qrrx_h
+SQDMLSL_vi      0.00 1111 10 . ..... 0111 . 0 ..... .....   @qrrx_s
+
 # Floating-point conditional select
 
 FCSEL           0001 1110 .. 1 rm:5 cond:4 11 rn:5 rd:5     esz=%esz_hsd
diff --git a/target/arm/tcg/translate-a64.c b/target/arm/tcg/translate-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/translate-a64.c
+++ b/target/arm/tcg/translate-a64.c
@@ -XXX,XX +XXX,XX @@ TRANS(UABAL_v, do_3op_widening,
       a->esz, a->q, a->rd, a->rn, a->rm, -1,
       gen_uaba_i64, true)
 
+static void gen_sqdmull_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    tcg_gen_mul_i64(d, n, m);
+    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, d);
+}
+
+static void gen_sqdmull_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    tcg_gen_mul_i64(d, n, m);
+    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, d);
+}
+
+static void gen_sqdmlal_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    TCGv_i64 t = tcg_temp_new_i64();
+
+    tcg_gen_mul_i64(t, n, m);
+    gen_helper_neon_addl_saturate_s32(t, tcg_env, t, t);
+    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, t);
+}
+
+static void gen_sqdmlal_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    TCGv_i64 t = tcg_temp_new_i64();
+
+    tcg_gen_mul_i64(t, n, m);
+    gen_helper_neon_addl_saturate_s64(t, tcg_env, t, t);
+    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, t);
+}
+
+static void gen_sqdmlsl_h(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    TCGv_i64 t = tcg_temp_new_i64();
+
+    tcg_gen_mul_i64(t, n, m);
+    gen_helper_neon_addl_saturate_s32(t, tcg_env, t, t);
+    tcg_gen_neg_i64(t, t);
+    gen_helper_neon_addl_saturate_s32(d, tcg_env, d, t);
+}
+
+static void gen_sqdmlsl_s(TCGv_i64 d, TCGv_i64 n, TCGv_i64 m)
+{
+    TCGv_i64 t = tcg_temp_new_i64();
+
+    tcg_gen_mul_i64(t, n, m);
+    gen_helper_neon_addl_saturate_s64(t, tcg_env, t, t);
+    tcg_gen_neg_i64(t, t);
+    gen_helper_neon_addl_saturate_s64(d, tcg_env, d, t);
+}
+
+TRANS(SQDMULL_v, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
+      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
+TRANS(SQDMLAL_v, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
+      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
+TRANS(SQDMLSL_v, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, -1,
+      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
+
+TRANS(SQDMULL_vi, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
+      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
+TRANS(SQDMLAL_vi, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
+      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
+TRANS(SQDMLSL_vi, do_3op_widening,
+      a->esz | MO_SIGN, a->q, a->rd, a->rn, a->rm, a->idx,
+      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
+
 /*
  * Advanced SIMD scalar/vector x indexed element
  */
@@ -XXX,XX +XXX,XX @@ static bool do_env_scalar3_idx_hs(DisasContext *s, arg_rrx_e *a,
 TRANS_FEAT(SQRDMLAH_si, aa64_rdm, do_env_scalar3_idx_hs, a, &f_scalar_sqrdmlah)
 TRANS_FEAT(SQRDMLSH_si, aa64_rdm, do_env_scalar3_idx_hs, a, &f_scalar_sqrdmlsh)
 
+static bool do_scalar_muladd_widening_idx(DisasContext *s, arg_rrx_e *a,
+                                          NeonGenTwo64OpFn *fn, bool acc)
+{
+    if (fp_access_check(s)) {
+        TCGv_i64 t0 = tcg_temp_new_i64();
+        TCGv_i64 t1 = tcg_temp_new_i64();
+        TCGv_i64 t2 = tcg_temp_new_i64();
+        unsigned vsz, dofs;
+
+        if (acc) {
+            read_vec_element(s, t0, a->rd, 0, a->esz + 1);
+        }
+        read_vec_element(s, t1, a->rn, 0, a->esz | MO_SIGN);
+        read_vec_element(s, t2, a->rm, a->idx, a->esz | MO_SIGN);
+        fn(t0, t1, t2);
+
+        /* Clear the whole register first, then store scalar. */
+        vsz = vec_full_reg_size(s);
+        dofs = vec_full_reg_offset(s, a->rd);
+        tcg_gen_gvec_dup_imm(MO_64, dofs, vsz, vsz, 0);
+        write_vec_element(s, t0, a->rd, 0, a->esz + 1);
+    }
+    return true;
+}
+
+TRANS(SQDMULL_si, do_scalar_muladd_widening_idx, a,
+      a->esz == MO_16 ? gen_sqdmull_h : gen_sqdmull_s, false)
+TRANS(SQDMLAL_si, do_scalar_muladd_widening_idx, a,
+      a->esz == MO_16 ? gen_sqdmlal_h : gen_sqdmlal_s, true)
+TRANS(SQDMLSL_si, do_scalar_muladd_widening_idx, a,
+      a->esz == MO_16 ? gen_sqdmlsl_h : gen_sqdmlsl_s, true)
+
 static bool do_fp3_vector_idx(DisasContext *s, arg_qrrx_e *a,
                               gen_helper_gvec_3_ptr * const fns[3])
 {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_scalar_shift_imm(DisasContext *s, uint32_t insn)
     }
 }
 
-/* AdvSIMD scalar three different
- *  31 30  29 28       24 23  22  21 20  16 15    12 11 10 9    5 4    0
- * +-----+---+-----------+------+---+------+--------+-----+------+------+
- * | 0 1 | U | 1 1 1 1 0 | size | 1 |  Rm  | opcode | 0 0 |  Rn  |  Rd  |
- * +-----+---+-----------+------+---+------+--------+-----+------+------+
- */
-static void disas_simd_scalar_three_reg_diff(DisasContext *s, uint32_t insn)
-{
-    bool is_u = extract32(insn, 29, 1);
-    int size = extract32(insn, 22, 2);
-    int opcode = extract32(insn, 12, 4);
-    int rm = extract32(insn, 16, 5);
-    int rn = extract32(insn, 5, 5);
-    int rd = extract32(insn, 0, 5);
-
-    if (is_u) {
-        unallocated_encoding(s);
-        return;
-    }
-
-    switch (opcode) {
-    case 0x9: /* SQDMLAL, SQDMLAL2 */
-    case 0xb: /* SQDMLSL, SQDMLSL2 */
-    case 0xd: /* SQDMULL, SQDMULL2 */
-        if (size == 0 || size == 3) {
-            unallocated_encoding(s);
-            return;
-        }
-        break;
-    default:
-        unallocated_encoding(s);
-        return;
-    }
-
-    if (!fp_access_check(s)) {
-        return;
-    }
-
-    if (size == 2) {
-        TCGv_i64 tcg_op1 = tcg_temp_new_i64();
-        TCGv_i64 tcg_op2 = tcg_temp_new_i64();
-        TCGv_i64 tcg_res = tcg_temp_new_i64();
-
-        read_vec_element(s, tcg_op1, rn, 0, MO_32 | MO_SIGN);
-        read_vec_element(s, tcg_op2, rm, 0, MO_32 | MO_SIGN);
-
-        tcg_gen_mul_i64(tcg_res, tcg_op1, tcg_op2);
-        gen_helper_neon_addl_saturate_s64(tcg_res, tcg_env, tcg_res, tcg_res);
-
-        switch (opcode) {
-        case 0xd: /* SQDMULL, SQDMULL2 */
-            break;
-        case 0xb: /* SQDMLSL, SQDMLSL2 */
-            tcg_gen_neg_i64(tcg_res, tcg_res);
-            /* fall through */
-        case 0x9: /* SQDMLAL, SQDMLAL2 */
-            read_vec_element(s, tcg_op1, rd, 0, MO_64);
-            gen_helper_neon_addl_saturate_s64(tcg_res, tcg_env,
-                                              tcg_res, tcg_op1);
-            break;
-        default:
-            g_assert_not_reached();
-        }
-
-        write_fp_dreg(s, rd, tcg_res);
-    } else {
-        TCGv_i32 tcg_op1 = read_fp_hreg(s, rn);
-        TCGv_i32 tcg_op2 = read_fp_hreg(s, rm);
-        TCGv_i64 tcg_res = tcg_temp_new_i64();
-
-        gen_helper_neon_mull_s16(tcg_res, tcg_op1, tcg_op2);
-        gen_helper_neon_addl_saturate_s32(tcg_res, tcg_env, tcg_res, tcg_res);
-
-        switch (opcode) {
-        case 0xd: /* SQDMULL, SQDMULL2 */
-            break;
-        case 0xb: /* SQDMLSL, SQDMLSL2 */
-            gen_helper_neon_negl_u32(tcg_res, tcg_res);
-            /* fall through */
-        case 0x9: /* SQDMLAL, SQDMLAL2 */
-        {
-            TCGv_i64 tcg_op3 = tcg_temp_new_i64();
-            read_vec_element(s, tcg_op3, rd, 0, MO_32);
-            gen_helper_neon_addl_saturate_s32(tcg_res, tcg_env,
-                                              tcg_res, tcg_op3);
-            break;
-        }
-        default:
-            g_assert_not_reached();
-        }
-
-        tcg_gen_ext32u_i64(tcg_res, tcg_res);
-        write_fp_dreg(s, rd, tcg_res);
-    }
-}
-
 static void handle_2misc_64(DisasContext *s, int opcode, bool u,
                             TCGv_i64 tcg_rd, TCGv_i64 tcg_rn,
                             TCGv_i32 tcg_rmode, TCGv_ptr tcg_fpstatus)
@@ -XXX,XX +XXX,XX @@ static void gen_neon_addl(int size, bool is_sub, TCGv_i64 tcg_res,
     genfn(tcg_res, tcg_op1, tcg_op2);
 }
 
-static void handle_3rd_widening(DisasContext *s, int is_q, int is_u, int size,
-                                int opcode, int rd, int rn, int rm)
-{
-    /* 3-reg-different widening insns: 64 x 64 -> 128 */
-    TCGv_i64 tcg_res[2];
-    int pass, accop;
-
-    tcg_res[0] = tcg_temp_new_i64();
-    tcg_res[1] = tcg_temp_new_i64();
-
-    /* Does this op do an adding accumulate, a subtracting accumulate,
-     * or no accumulate at all?
-     */
-    switch (opcode) {
-    case 5:
-    case 8:
-    case 9:
-        accop = 1;
-        break;
-    case 10:
-    case 11:
-        accop = -1;
-        break;
-    default:
-        accop = 0;
-        break;
-    }
-
-    if (accop != 0) {
-        read_vec_element(s, tcg_res[0], rd, 0, MO_64);
-        read_vec_element(s, tcg_res[1], rd, 1, MO_64);
-    }
-
-    /* size == 2 means two 32x32->64 operations; this is worth special
-     * casing because we can generally handle it inline.
-     */
-    if (size == 2) {
-        for (pass = 0; pass < 2; pass++) {
-            TCGv_i64 tcg_op1 = tcg_temp_new_i64();
-            TCGv_i64 tcg_op2 = tcg_temp_new_i64();
-            TCGv_i64 tcg_passres;
-            MemOp memop = MO_32 | (is_u ? 0 : MO_SIGN);
-
-            int elt = pass + is_q * 2;
-
-            read_vec_element(s, tcg_op1, rn, elt, memop);
-            read_vec_element(s, tcg_op2, rm, elt, memop);
-
-            if (accop == 0) {
-                tcg_passres = tcg_res[pass];
-            } else {
-                tcg_passres = tcg_temp_new_i64();
-            }
-
-            switch (opcode) {
-            case 9: /* SQDMLAL, SQDMLAL2 */
-            case 11: /* SQDMLSL, SQDMLSL2 */
-            case 13: /* SQDMULL, SQDMULL2 */
-                tcg_gen_mul_i64(tcg_passres, tcg_op1, tcg_op2);
-                gen_helper_neon_addl_saturate_s64(tcg_passres, tcg_env,
-                                                  tcg_passres, tcg_passres);
-                break;
-            default:
-            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
-            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
-            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
-            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
-            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
-            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
-            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
-                g_assert_not_reached();
-            }
-
-            if (accop != 0) {
-                /* saturating accumulate ops */
-                if (accop < 0) {
-                    tcg_gen_neg_i64(tcg_passres, tcg_passres);
-                }
-                gen_helper_neon_addl_saturate_s64(tcg_res[pass], tcg_env,
-                                                  tcg_res[pass], tcg_passres);
-            }
-        }
-    } else {
-        /* size 0 or 1, generally helper functions */
-        for (pass = 0; pass < 2; pass++) {
-            TCGv_i32 tcg_op1 = tcg_temp_new_i32();
-            TCGv_i32 tcg_op2 = tcg_temp_new_i32();
-            TCGv_i64 tcg_passres;
-            int elt = pass + is_q * 2;
-
-            read_vec_element_i32(s, tcg_op1, rn, elt, MO_32);
-            read_vec_element_i32(s, tcg_op2, rm, elt, MO_32);
-
-            if (accop == 0) {
-                tcg_passres = tcg_res[pass];
-            } else {
-                tcg_passres = tcg_temp_new_i64();
-            }
-
-            switch (opcode) {
-            case 9: /* SQDMLAL, SQDMLAL2 */
-            case 11: /* SQDMLSL, SQDMLSL2 */
-            case 13: /* SQDMULL, SQDMULL2 */
-                assert(size == 1);
-                gen_helper_neon_mull_s16(tcg_passres, tcg_op1, tcg_op2);
-                gen_helper_neon_addl_saturate_s32(tcg_passres, tcg_env,
-                                                  tcg_passres, tcg_passres);
-                break;
-            default:
-            case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
-            case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
-            case 12: /* UMULL, UMULL2, SMULL, SMULL2 */
-            case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
-            case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
-            case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
-            case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
-                g_assert_not_reached();
-            }
-
-            if (accop != 0) {
-                /* saturating accumulate ops */
-                if (accop < 0) {
-                    gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
-                }
-                gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
-                                                  tcg_res[pass],
-                                                  tcg_passres);
-            }
-        }
-    }
-
-    write_vec_element(s, tcg_res[0], rd, 0, MO_64);
-    write_vec_element(s, tcg_res[1], rd, 1, MO_64);
-}
-
 static void handle_3rd_wide(DisasContext *s, int is_q, int is_u, int size,
                             int opcode, int rd, int rn, int rm)
 {
@@ -XXX,XX +XXX,XX @@ static void disas_simd_three_reg_diff(DisasContext *s, uint32_t insn)
             break;
         }
         return;
-    case 9: /* SQDMLAL, SQDMLAL2 */
-    case 11: /* SQDMLSL, SQDMLSL2 */
-    case 13: /* SQDMULL, SQDMULL2 */
-        if (is_u || size == 0) {
-            unallocated_encoding(s);
-            return;
-        }
-        /* 64 x 64 -> 128 */
-        if (size == 3) {
-            unallocated_encoding(s);
-            return;
-        }
-        if (!fp_access_check(s)) {
-            return;
-        }
-
-        handle_3rd_widening(s, is_q, is_u, size, opcode, rd, rn, rm);
-        break;
     default:
     case 0: /* SADDL, SADDL2, UADDL, UADDL2 */
     case 2: /* SSUBL, SSUBL2, USUBL, USUBL2 */
     case 5: /* SABAL, SABAL2, UABAL, UABAL2 */
     case 7: /* SABDL, SABDL2, UABDL, UABDL2 */
     case 8: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
+    case 9: /* SQDMLAL, SQDMLAL2 */
     case 10: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
+    case 11: /* SQDMLSL, SQDMLSL2 */
     case 12: /* SMULL, SMULL2, UMULL, UMULL2 */
+    case 13: /* SQDMULL, SQDMULL2 */
         /* opcode 15 not allocated */
         unallocated_encoding(s);
         break;
@@ -XXX,XX +XXX,XX @@ static void disas_simd_two_reg_misc_fp16(DisasContext *s, uint32_t insn)
     }
 }
 
-/* AdvSIMD scalar x indexed element
- *  31 30  29 28       24 23  22 21  20  19  16 15 12  11  10 9    5 4    0
- * +-----+---+-----------+------+---+---+------+-----+---+---+------+------+
- * | 0 1 | U | 1 1 1 1 1 | size | L | M |  Rm  | opc | H | 0 |  Rn  |  Rd  |
- * +-----+---+-----------+------+---+---+------+-----+---+---+------+------+
- * AdvSIMD vector x indexed element
- *   31  30  29 28       24 23  22 21  20  19  16 15 12  11  10 9    5 4    0
- * +---+---+---+-----------+------+---+---+------+-----+---+---+------+------+
- * | 0 | Q | U | 0 1 1 1 1 | size | L | M |  Rm  | opc | H | 0 |  Rn  |  Rd  |
- * +---+---+---+-----------+------+---+---+------+-----+---+---+------+------+
- */
-static void disas_simd_indexed(DisasContext *s, uint32_t insn)
-{
-    /* This encoding has two kinds of instruction:
-     *  normal, where we perform elt x idxelt => elt for each
-     *     element in the vector
-     *  long, where we perform elt x idxelt and generate a result of
-     *     double the width of the input element
-     * The long ops have a 'part' specifier (ie come in INSN, INSN2 pairs).
-     */
-    bool is_scalar = extract32(insn, 28, 1);
-    bool is_q = extract32(insn, 30, 1);
-    bool u = extract32(insn, 29, 1);
-    int size = extract32(insn, 22, 2);
-    int l = extract32(insn, 21, 1);
-    int m = extract32(insn, 20, 1);
-    /* Note that the Rm field here is only 4 bits, not 5 as it usually is */
-    int rm = extract32(insn, 16, 4);
-    int opcode = extract32(insn, 12, 4);
-    int h = extract32(insn, 11, 1);
-    int rn = extract32(insn, 5, 5);
-    int rd = extract32(insn, 0, 5);
-    int index;
-
-    switch (16 * u + opcode) {
-    case 0x03: /* SQDMLAL, SQDMLAL2 */
-    case 0x07: /* SQDMLSL, SQDMLSL2 */
-    case 0x0b: /* SQDMULL, SQDMULL2 */
-        break;
-    default:
-    case 0x00: /* FMLAL */
-    case 0x01: /* FMLA */
-    case 0x02: /* SMLAL, SMLAL2 */
-    case 0x04: /* FMLSL */
-    case 0x05: /* FMLS */
-    case 0x06: /* SMLSL, SMLSL2 */
-    case 0x08: /* MUL */
-    case 0x09: /* FMUL */
-    case 0x0a: /* SMULL, SMULL2 */
-    case 0x0c: /* SQDMULH */
-    case 0x0d: /* SQRDMULH */
-    case 0x0e: /* SDOT */
-    case 0x0f: /* SUDOT / BFDOT / USDOT / BFMLAL */
-    case 0x10: /* MLA */
-    case 0x11: /* FCMLA #0 */
-    case 0x12: /* UMLAL, UMLAL2 */
-    case 0x13: /* FCMLA #90 */
-    case 0x14: /* MLS */
-    case 0x15: /* FCMLA #180 */
-    case 0x16: /* UMLSL, UMLSL2 */
-    case 0x17: /* FCMLA #270 */
-    case 0x18: /* FMLAL2 */
-    case 0x19: /* FMULX */
-    case 0x1a: /* UMULL, UMULL2 */
-    case 0x1c: /* FMLSL2 */
-    case 0x1d: /* SQRDMLAH */
-    case 0x1e: /* UDOT */
-    case 0x1f: /* SQRDMLSH */
-        unallocated_encoding(s);
-        return;
-    }
-
-    /* Given MemOp size, adjust register and indexing.  */
-    switch (size) {
-    case MO_8:
-    case MO_64:
-        unallocated_encoding(s);
-        return;
-    case MO_16:
-        index = h << 2 | l << 1 | m;
-        break;
-    case MO_32:
-        index = h << 1 | l;
-        rm |= m << 4;
-        break;
-    default:
-        g_assert_not_reached();
-    }
-
-    if (!fp_access_check(s)) {
-        return;
-    }
-
-    if (size == 3) {
-        g_assert_not_reached();
-    } else {
-        /* long ops: 16x16->32 or 32x32->64 */
-        TCGv_i64 tcg_res[2];
-        int pass;
-        bool satop = extract32(opcode, 0, 1);
-        MemOp memop = MO_32;
-
-        if (satop || !u) {
-            memop |= MO_SIGN;
-        }
-
-        if (size == 2) {
-            TCGv_i64 tcg_idx = tcg_temp_new_i64();
-
-            read_vec_element(s, tcg_idx, rm, index, memop);
-
-            for (pass = 0; pass < (is_scalar ? 1 : 2); pass++) {
-                TCGv_i64 tcg_op = tcg_temp_new_i64();
-                TCGv_i64 tcg_passres;
-                int passelt;
-
-                if (is_scalar) {
-                    passelt = 0;
-                } else {
-                    passelt = pass + (is_q * 2);
-                }
-
-                read_vec_element(s, tcg_op, rn, passelt, memop);
-
-                tcg_res[pass] = tcg_temp_new_i64();
-
-                if (opcode == 0xa || opcode == 0xb) {
-                    /* Non-accumulating ops */
-                    tcg_passres = tcg_res[pass];
-                } else {
-                    tcg_passres = tcg_temp_new_i64();
-                }
-
-                tcg_gen_mul_i64(tcg_passres, tcg_op, tcg_idx);
-
-                if (satop) {
-                    /* saturating, doubling */
-                    gen_helper_neon_addl_saturate_s64(tcg_passres, tcg_env,
-                                                      tcg_passres, tcg_passres);
-                }
-
-                if (opcode == 0xa || opcode == 0xb) {
-                    continue;
-                }
-
-                /* Accumulating op: handle accumulate step */
-                read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
-
-                switch (opcode) {
-                case 0x7: /* SQDMLSL, SQDMLSL2 */
-                    tcg_gen_neg_i64(tcg_passres, tcg_passres);
-                    /* fall through */
-                case 0x3: /* SQDMLAL, SQDMLAL2 */
-                    gen_helper_neon_addl_saturate_s64(tcg_res[pass], tcg_env,
-                                                      tcg_res[pass],
-                                                      tcg_passres);
-                    break;
-                default:
-                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
-                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
-                    g_assert_not_reached();
-                }
-            }
-
-            clear_vec_high(s, !is_scalar, rd);
-        } else {
-            TCGv_i32 tcg_idx = tcg_temp_new_i32();
-
-            assert(size == 1);
-            read_vec_element_i32(s, tcg_idx, rm, index, size);
-
-            if (!is_scalar) {
-                /* The simplest way to handle the 16x16 indexed ops is to
-                 * duplicate the index into both halves of the 32 bit tcg_idx
-                 * and then use the usual Neon helpers.
-                 */
-                tcg_gen_deposit_i32(tcg_idx, tcg_idx, tcg_idx, 16, 16);
-            }
-
-            for (pass = 0; pass < (is_scalar ? 1 : 2); pass++) {
-                TCGv_i32 tcg_op = tcg_temp_new_i32();
-                TCGv_i64 tcg_passres;
-
-                if (is_scalar) {
-                    read_vec_element_i32(s, tcg_op, rn, pass, size);
-                } else {
-                    read_vec_element_i32(s, tcg_op, rn,
-                                         pass + (is_q * 2), MO_32);
-                }
-
-                tcg_res[pass] = tcg_temp_new_i64();
-
-                if (opcode == 0xa || opcode == 0xb) {
-                    /* Non-accumulating ops */
-                    tcg_passres = tcg_res[pass];
-                } else {
-                    tcg_passres = tcg_temp_new_i64();
-                }
-
-                if (memop & MO_SIGN) {
-                    gen_helper_neon_mull_s16(tcg_passres, tcg_op, tcg_idx);
-                } else {
-                    gen_helper_neon_mull_u16(tcg_passres, tcg_op, tcg_idx);
-                }
-                if (satop) {
-                    gen_helper_neon_addl_saturate_s32(tcg_passres, tcg_env,
-                                                      tcg_passres, tcg_passres);
-                }
-
-                if (opcode == 0xa || opcode == 0xb) {
-                    continue;
-                }
-
-                /* Accumulating op: handle accumulate step */
-                read_vec_element(s, tcg_res[pass], rd, pass, MO_64);
-
-                switch (opcode) {
-                case 0x7: /* SQDMLSL, SQDMLSL2 */
-                    gen_helper_neon_negl_u32(tcg_passres, tcg_passres);
-                    /* fall through */
-                case 0x3: /* SQDMLAL, SQDMLAL2 */
-                    gen_helper_neon_addl_saturate_s32(tcg_res[pass], tcg_env,
-                                                      tcg_res[pass],
-                                                      tcg_passres);
-                    break;
-                default:
-                case 0x2: /* SMLAL, SMLAL2, UMLAL, UMLAL2 */
-                case 0x6: /* SMLSL, SMLSL2, UMLSL, UMLSL2 */
-                    g_assert_not_reached();
-                }
-            }
-
-            if (is_scalar) {
-                tcg_gen_ext32u_i64(tcg_res[0], tcg_res[0]);
-            }
-        }
-
-        if (is_scalar) {
-            tcg_res[1] = tcg_constant_i64(0);
-        }
-
-        for (pass = 0; pass < 2; pass++) {
-            write_vec_element(s, tcg_res[pass], rd, pass, MO_64);
-        }
-    }
-}
-
 /* C3.6 Data processing - SIMD, inc Crypto
  *
  * As the decode gets a little complex we are using a table based
@@ -XXX,XX +XXX,XX @@ static const AArch64DecodeTable data_proc_simd[] = {
     { 0x0e200000, 0x9f200c00, disas_simd_three_reg_diff },
     { 0x0e200800, 0x9f3e0c00, disas_simd_two_reg_misc },
     { 0x0e300800, 0x9f3e0c00, disas_simd_across_lanes },
-    { 0x0f000000, 0x9f000400, disas_simd_indexed }, /* vector indexed */
     /* simd_mod_imm decode is a subset of simd_shift_imm, so must precede it */
     { 0x0f000400, 0x9ff80400, disas_simd_mod_imm },
     { 0x0f000400, 0x9f800400, disas_simd_shift_imm },
     { 0x0e000000, 0xbf208c00, disas_simd_tb },
     { 0x0e000800, 0xbf208c00, disas_simd_zip_trn },
     { 0x2e000000, 0xbf208400, disas_simd_ext },
-    { 0x5e200000, 0xdf200c00, disas_simd_scalar_three_reg_diff },
     { 0x5e200800, 0xdf3e0c00, disas_simd_scalar_two_reg_misc },
-    { 0x5f000000, 0xdf000400, disas_simd_indexed }, /* scalar indexed */
     { 0x5f000400, 0xdf800400, disas_simd_scalar_shift_imm },
     { 0x0e780800, 0x8f7e0c00, disas_simd_two_reg_misc_fp16 },
     { 0x00000000, 0x00000000, NULL }
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-5-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |  5 ++
 target/arm/tcg/translate-a64.c | 86 +++++++++++++++++-----------------
 2 files changed, 48 insertions(+), 43 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-6-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |   5 ++
 target/arm/tcg/translate-a64.c | 127 +++++++++++++++------------------
 2 files changed, 61 insertions(+), 71 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20240709000610.382391-7-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/a64.decode      |  3 ++
 target/arm/tcg/translate-a64.c | 94 +++++-----------------------------
 2 files changed, 15 insertions(+), 82 deletions(-)

First arm pullreq of the cycle; this is mostly my softfloat NaN
handling series. (Lots more in my to-review queue, but I don't
like pullreqs growing too close to a hundred patches at a time :-))

thanks
-- PMM

The following changes since commit 97f2796a3736ed37a1b85dc1c76a6c45b829dd17:

Open 10.0 development tree (2024-12-10 17:41:17 +0000)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20241211

for you to fetch changes up to 1abe28d519239eea5cf9620bb13149423e5665f8:

MAINTAINERS: Add correct email address for Vikram Garhwal (2024-12-11 15:31:09 +0000)

----------------------------------------------------------------
target-arm queue:
 * hw/net/lan9118: Extract PHY model, reuse with imx_fec, fix bugs
 * fpu: Make muladd NaN handling runtime-selected, not compile-time
 * fpu: Make default NaN pattern runtime-selected, not compile-time
 * fpu: Minor NaN-related cleanups
 * MAINTAINERS: email address updates

----------------------------------------------------------------
Bernhard Beschow (5):
      hw/net/lan9118: Extract lan9118_phy
      hw/net/lan9118_phy: Reuse in imx_fec and consolidate implementations
      hw/net/lan9118_phy: Fix off-by-one error in MII_ANLPAR register
      hw/net/lan9118_phy: Reuse MII constants
      hw/net/lan9118_phy: Add missing 100 mbps full duplex advertisement

Leif Lindholm (1):
      MAINTAINERS: update email address for Leif Lindholm

Peter Maydell (54):
      fpu: handle raising Invalid for infzero in pick_nan_muladd
      fpu: Check for default_nan_mode before calling pickNaNMulAdd
      softfloat: Allow runtime choice of inf * 0 + NaN result
      tests/fp: Explicitly set inf-zero-nan rule
      target/arm: Set FloatInfZeroNaNRule explicitly
      target/s390: Set FloatInfZeroNaNRule explicitly
      target/ppc: Set FloatInfZeroNaNRule explicitly
      target/mips: Set FloatInfZeroNaNRule explicitly
      target/sparc: Set FloatInfZeroNaNRule explicitly
      target/xtensa: Set FloatInfZeroNaNRule explicitly
      target/x86: Set FloatInfZeroNaNRule explicitly
      target/loongarch: Set FloatInfZeroNaNRule explicitly
      target/hppa: Set FloatInfZeroNaNRule explicitly
      softfloat: Pass have_snan to pickNaNMulAdd
      softfloat: Allow runtime choice of NaN propagation for muladd
      tests/fp: Explicitly set 3-NaN propagation rule
      target/arm: Set Float3NaNPropRule explicitly
      target/loongarch: Set Float3NaNPropRule explicitly
      target/ppc: Set Float3NaNPropRule explicitly
      target/s390x: Set Float3NaNPropRule explicitly
      target/sparc: Set Float3NaNPropRule explicitly
      target/mips: Set Float3NaNPropRule explicitly
      target/xtensa: Set Float3NaNPropRule explicitly
      target/i386: Set Float3NaNPropRule explicitly
      target/hppa: Set Float3NaNPropRule explicitly
      fpu: Remove use_first_nan field from float_status
      target/m68k: Don't pass NULL float_status to floatx80_default_nan()
      softfloat: Create floatx80 default NaN from parts64_default_nan
      target/loongarch: Use normal float_status in fclass_s and fclass_d helpers
      target/m68k: In frem helper, initialize local float_status from env->fp_status
      target/m68k: Init local float_status from env fp_status in gdb get/set reg
      target/sparc: Initialize local scratch float_status from env->fp_status
      target/ppc: Use env->fp_status in helper_compute_fprf functions
      fpu: Allow runtime choice of default NaN value
      tests/fp: Set default NaN pattern explicitly
      target/microblaze: Set default NaN pattern explicitly
      target/i386: Set default NaN pattern explicitly
      target/hppa: Set default NaN pattern explicitly
      target/alpha: Set default NaN pattern explicitly
      target/arm: Set default NaN pattern explicitly
      target/loongarch: Set default NaN pattern explicitly
      target/m68k: Set default NaN pattern explicitly
      target/mips: Set default NaN pattern explicitly
      target/openrisc: Set default NaN pattern explicitly
      target/ppc: Set default NaN pattern explicitly
      target/sh4: Set default NaN pattern explicitly
      target/rx: Set default NaN pattern explicitly
      target/s390x: Set default NaN pattern explicitly
      target/sparc: Set default NaN pattern explicitly
      target/xtensa: Set default NaN pattern explicitly
      target/hexagon: Set default NaN pattern explicitly
      target/riscv: Set default NaN pattern explicitly
      target/tricore: Set default NaN pattern explicitly
      fpu: Remove default handling for dnan_pattern

Richard Henderson (11):
      target/arm: Copy entire float_status in is_ebf
      softfloat: Inline pickNaNMulAdd
      softfloat: Use goto for default nan case in pick_nan_muladd
      softfloat: Remove which from parts_pick_nan_muladd
      softfloat: Pad array size in pick_nan_muladd
      softfloat: Move propagateFloatx80NaN to softfloat.c
      softfloat: Use parts_pick_nan in propagateFloatx80NaN
      softfloat: Inline pickNaN
      softfloat: Share code between parts_pick_nan cases
      softfloat: Sink frac_cmp in parts_pick_nan until needed
      softfloat: Replace WHICH with RET in parts_pick_nan

Vikram Garhwal (1):
      MAINTAINERS: Add correct email address for Vikram Garhwal

From: Bernhard Beschow <shentey@gmail.com>

A very similar implementation of the same device exists in imx_fec. Prepare for
a common implementation by extracting a device model into its own files.

Some migration state has been moved into the new device model which breaks
migration compatibility for the following machines:
* smdkc210
* realview-*
* vexpress-*
* kzm
* mps2-*

While breaking migration ABI, fix the size of the MII registers to be 16 bit,
as defined by IEEE 802.3u.

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-2-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/lan9118_phy.h |  37 ++++++++
 hw/net/lan9118.c             | 137 +++++-----------------------
 hw/net/lan9118_phy.c         | 169 +++++++++++++++++++++++++++++++++++
 hw/net/Kconfig               |   4 +
 hw/net/meson.build           |   1 +
 5 files changed, 233 insertions(+), 115 deletions(-)
 create mode 100644 include/hw/net/lan9118_phy.h
 create mode 100644 hw/net/lan9118_phy.c

diff --git a/include/hw/net/lan9118_phy.h b/include/hw/net/lan9118_phy.h
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/include/hw/net/lan9118_phy.h
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This work is licensed under the terms of the GNU GPL, version 2 or later.
+ * See the COPYING file in the top-level directory.
+ */
+
+#ifndef HW_NET_LAN9118_PHY_H
+#define HW_NET_LAN9118_PHY_H
+
+#include "qom/object.h"
+#include "hw/sysbus.h"
+
+#define TYPE_LAN9118_PHY "lan9118-phy"
+OBJECT_DECLARE_SIMPLE_TYPE(Lan9118PhyState, LAN9118_PHY)
+
+typedef struct Lan9118PhyState {
+    SysBusDevice parent_obj;
+
+    uint16_t status;
+    uint16_t control;
+    uint16_t advertise;
+    uint16_t ints;
+    uint16_t int_mask;
+    qemu_irq irq;
+    bool link_down;
+} Lan9118PhyState;
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down);
+void lan9118_phy_reset(Lan9118PhyState *s);
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg);
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val);
+
+#endif
diff --git a/hw/net/lan9118.c b/hw/net/lan9118.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118.c
+++ b/hw/net/lan9118.c
@@ -XXX,XX +XXX,XX @@
 #include "net/net.h"
 #include "net/eth.h"
 #include "hw/irq.h"
+#include "hw/net/lan9118_phy.h"
 #include "hw/net/lan9118.h"
 #include "hw/ptimer.h"
 #include "hw/qdev-properties.h"
@@ -XXX,XX +XXX,XX @@ do { printf("lan9118: " fmt , ## __VA_ARGS__); } while (0)
 #define MAC_CR_RXEN     0x00000004
 #define MAC_CR_RESERVED 0x7f404213
 
-#define PHY_INT_ENERGYON            0x80
-#define PHY_INT_AUTONEG_COMPLETE    0x40
-#define PHY_INT_FAULT               0x20
-#define PHY_INT_DOWN                0x10
-#define PHY_INT_AUTONEG_LP          0x08
-#define PHY_INT_PARFAULT            0x04
-#define PHY_INT_AUTONEG_PAGE        0x02
-
 #define GPT_TIMER_EN    0x20000000
 
 /*
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
     uint32_t mac_mii_data;
     uint32_t mac_flow;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
 
     int32_t eeprom_writable;
     uint8_t eeprom[128];
@@ -XXX,XX +XXX,XX @@ struct lan9118_state {
 
 static const VMStateDescription vmstate_lan9118 = {
     .name = "lan9118",
-    .version_id = 2,
-    .minimum_version_id = 1,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_PTIMER(timer, lan9118_state),
         VMSTATE_UINT32(irq_cfg, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118 = {
         VMSTATE_UINT32(mac_mii_acc, lan9118_state),
         VMSTATE_UINT32(mac_mii_data, lan9118_state),
         VMSTATE_UINT32(mac_flow, lan9118_state),
-        VMSTATE_UINT32(phy_status, lan9118_state),
-        VMSTATE_UINT32(phy_control, lan9118_state),
-        VMSTATE_UINT32(phy_advertise, lan9118_state),
-        VMSTATE_UINT32(phy_int, lan9118_state),
-        VMSTATE_UINT32(phy_int_mask, lan9118_state),
         VMSTATE_INT32(eeprom_writable, lan9118_state),
         VMSTATE_UINT8_ARRAY(eeprom, lan9118_state, 128),
         VMSTATE_INT32(tx_fifo_size, lan9118_state),
@@ -XXX,XX +XXX,XX @@ static void lan9118_reload_eeprom(lan9118_state *s)
     lan9118_mac_changed(s);
 }
 
-static void phy_update_irq(lan9118_state *s)
+static void lan9118_update_irq(void *opaque, int n, int level)
 {
-    if (s->phy_int & s->phy_int_mask) {
+    lan9118_state *s = opaque;
+
+    if (level) {
         s->int_sts |= PHY_INT;
     } else {
         s->int_sts &= ~PHY_INT;
@@ -XXX,XX +XXX,XX @@ static void phy_update_irq(lan9118_state *s)
     lan9118_update(s);
 }
 
-static void phy_update_link(lan9118_state *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    phy_update_irq(s);
-}
-
 static void lan9118_set_link(NetClientState *nc)
 {
-    phy_update_link(qemu_get_nic_opaque(nc));
-}
-
-static void phy_reset(lan9118_state *s)
-{
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    phy_update_link(s);
+    lan9118_phy_update_link(&LAN9118(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static void lan9118_reset(DeviceState *d)
@@ -XXX,XX +XXX,XX @@ static void lan9118_reset(DeviceState *d)
     s->read_word_n = 0;
     s->write_word_n = 0;
 
-    phy_reset(s);
-
     s->eeprom_writable = 0;
     lan9118_reload_eeprom(s);
 }
@@ -XXX,XX +XXX,XX @@ static void do_tx_packet(lan9118_state *s)
     uint32_t status;
 
     /* FIXME: Honor TX disable, and allow queueing of packets.  */
-    if (s->phy_control & 0x4000)  {
+    if (s->mii.control & 0x4000) {
         /* This assumes the receive routine doesn't touch the VLANClient.  */
         qemu_receive_packet(qemu_get_queue(s->nic), s->txp->data, s->txp->len);
     } else {
@@ -XXX,XX +XXX,XX @@ static void tx_fifo_push(lan9118_state *s, uint32_t val)
     }
 }
 
-static uint32_t do_phy_read(lan9118_state *s, int reg)
-{
-    uint32_t val;
-
-    switch (reg) {
-    case 0: /* Basic Control */
-        return s->phy_control;
-    case 1: /* Basic Status */
-        return s->phy_status;
-    case 2: /* ID1 */
-        return 0x0007;
-    case 3: /* ID2 */
-        return 0xc0d1;
-    case 4: /* Auto-neg advertisement */
-        return s->phy_advertise;
-    case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
-    case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
-    case 29: /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        phy_update_irq(s);
-        return val;
-    case 30: /* Interrupt mask */
-        return s->phy_int_mask;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_read: PHY read reg %d\n", reg);
-        return 0;
-    }
-}
-
-static void do_phy_write(lan9118_state *s, int reg, uint32_t val)
-{
-    switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
-            phy_reset(s);
-            break;
-        }
-        s->phy_control = val & 0x7980;
-        /* Complete autonegotiation immediately.  */
-        if (val & 0x1000) {
-            s->phy_status |= 0x0020;
-        }
-        break;
-    case 4: /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-        /* TODO 17, 18, 27, 31 */
-    case 30: /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        phy_update_irq(s);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "do_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
-    }
-}
-
 static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
 {
     switch (reg) {
@@ -XXX,XX +XXX,XX @@ static void do_mac_write(lan9118_state *s, int reg, uint32_t val)
         if (val & 2) {
             DPRINTF("PHY write %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
-            do_phy_write(s, (val >> 6) & 0x1f, s->mac_mii_data);
+            lan9118_phy_write(&s->mii, (val >> 6) & 0x1f, s->mac_mii_data);
         } else {
-            s->mac_mii_data = do_phy_read(s, (val >> 6) & 0x1f);
+            s->mac_mii_data = lan9118_phy_read(&s->mii, (val >> 6) & 0x1f);
             DPRINTF("PHY read %d = 0x%04x\n",
                     (val >> 6) & 0x1f, s->mac_mii_data);
         }
@@ -XXX,XX +XXX,XX @@ static void lan9118_writel(void *opaque, hwaddr offset,
         break;
     case CSR_PMT_CTRL:
         if (val & 0x400) {
-            phy_reset(s);
+            lan9118_phy_reset(&s->mii);
         }
         s->pmt_ctrl &= ~0x34e;
         s->pmt_ctrl |= (val & 0x34e);
@@ -XXX,XX +XXX,XX @@ static void lan9118_realize(DeviceState *dev, Error **errp)
     const MemoryRegionOps *mem_ops =
             s->mode_16bit ? &lan9118_16bit_mem_ops : &lan9118_mem_ops;
 
+    qemu_init_irq(&s->mii_irq, lan9118_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     memory_region_init_io(&s->mmio, OBJECT(dev), mem_ops, s,
                           "lan9118-mmio", 0x100);
     sysbus_init_mmio(sbd, &s->mmio);
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
new file mode 100644
index XXXXXXX..XXXXXXX
--- /dev/null
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
+/*
+ * SMSC LAN9118 PHY emulation
+ *
+ * Copyright (c) 2009 CodeSourcery, LLC.
+ * Written by Paul Brook
+ *
+ * This code is licensed under the GNU GPL v2
+ *
+ * Contributions after 2012-01-13 are licensed under the terms of the
+ * GNU GPL, version 2 or (at your option) any later version.
+ */
+
+#include "qemu/osdep.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
+#include "hw/resettable.h"
+#include "migration/vmstate.h"
+#include "qemu/log.h"
+
+#define PHY_INT_ENERGYON            (1 << 7)
+#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
+#define PHY_INT_FAULT               (1 << 5)
+#define PHY_INT_DOWN                (1 << 4)
+#define PHY_INT_AUTONEG_LP          (1 << 3)
+#define PHY_INT_PARFAULT            (1 << 2)
+#define PHY_INT_AUTONEG_PAGE        (1 << 1)
+
+static void lan9118_phy_update_irq(Lan9118PhyState *s)
+{
+    qemu_set_irq(s->irq, !!(s->ints & s->int_mask));
+}
+
+uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
+{
+    uint16_t val;
+
+    switch (reg) {
+    case 0: /* Basic Control */
+        return s->control;
+    case 1: /* Basic Status */
+        return s->status;
+    case 2: /* ID1 */
+        return 0x0007;
+    case 3: /* ID2 */
+        return 0xc0d1;
+    case 4: /* Auto-neg advertisement */
+        return s->advertise;
+    case 5: /* Auto-neg Link Partner Ability */
+        return 0x0f71;
+    case 6: /* Auto-neg Expansion */
+        return 1;
+        /* TODO 17, 18, 27, 29, 30, 31 */
+    case 29: /* Interrupt source. */
+        val = s->ints;
+        s->ints = 0;
+        lan9118_phy_update_irq(s);
+        return val;
+    case 30: /* Interrupt mask */
+        return s->int_mask;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_read: PHY read reg %d\n", reg);
+        return 0;
+    }
+}
+
+void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
+{
+    switch (reg) {
+    case 0: /* Basic Control */
+        if (val & 0x8000) {
+            lan9118_phy_reset(s);
+            break;
+        }
+        s->control = val & 0x7980;
+        /* Complete autonegotiation immediately. */
+        if (val & 0x1000) {
+            s->status |= 0x0020;
+        }
+        break;
+    case 4: /* Auto-neg advertisement */
+        s->advertise = (val & 0x2d7f) | 0x80;
+        break;
+        /* TODO 17, 18, 27, 31 */
+    case 30: /* Interrupt mask */
+        s->int_mask = val & 0xff;
+        lan9118_phy_update_irq(s);
+        break;
+    default:
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+    }
+}
+
+void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
+{
+    s->link_down = link_down;
+
+    /* Autonegotiation status mirrors link status. */
+    if (link_down) {
+        s->status &= ~0x0024;
+        s->ints |= PHY_INT_DOWN;
+    } else {
+        s->status |= 0x0024;
+        s->ints |= PHY_INT_ENERGYON;
+        s->ints |= PHY_INT_AUTONEG_COMPLETE;
+    }
+    lan9118_phy_update_irq(s);
+}
+
+void lan9118_phy_reset(Lan9118PhyState *s)
+{
+    s->control = 0x3000;
+    s->status = 0x7809;
+    s->advertise = 0x01e1;
+    s->int_mask = 0;
+    s->ints = 0;
+    lan9118_phy_update_link(s, s->link_down);
+}
+
+static void lan9118_phy_reset_hold(Object *obj, ResetType type)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    lan9118_phy_reset(s);
+}
+
+static void lan9118_phy_init(Object *obj)
+{
+    Lan9118PhyState *s = LAN9118_PHY(obj);
+
+    qdev_init_gpio_out(DEVICE(s), &s->irq, 1);
+}
+
+static const VMStateDescription vmstate_lan9118_phy = {
+    .name = "lan9118-phy",
+    .version_id = 1,
+    .minimum_version_id = 1,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT16(control, Lan9118PhyState),
+        VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(advertise, Lan9118PhyState),
+        VMSTATE_UINT16(ints, Lan9118PhyState),
+        VMSTATE_UINT16(int_mask, Lan9118PhyState),
+        VMSTATE_BOOL(link_down, Lan9118PhyState),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static void lan9118_phy_class_init(ObjectClass *klass, void *data)
+{
+    ResettableClass *rc = RESETTABLE_CLASS(klass);
+    DeviceClass *dc = DEVICE_CLASS(klass);
+
+    rc->phases.hold = lan9118_phy_reset_hold;
+    dc->vmsd = &vmstate_lan9118_phy;
+}
+
+static const TypeInfo types[] = {
+    {
+        .name          = TYPE_LAN9118_PHY,
+        .parent        = TYPE_SYS_BUS_DEVICE,
+        .instance_size = sizeof(Lan9118PhyState),
+        .instance_init = lan9118_phy_init,
+        .class_init    = lan9118_phy_class_init,
+    }
+};
+
+DEFINE_TYPES(types)
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config VMXNET3_PCI
 config SMC91C111
     bool
 
+config LAN9118_PHY
+    bool
+
 config LAN9118
     bool
+    select LAN9118_PHY
     select PTIMER
 
 config NE2000_ISA
diff --git a/hw/net/meson.build b/hw/net/meson.build
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/meson.build
+++ b/hw/net/meson.build
@@ -XXX,XX +XXX,XX @@ system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('vmxnet3.c'))
 
 system_ss.add(when: 'CONFIG_SMC91C111', if_true: files('smc91c111.c'))
 system_ss.add(when: 'CONFIG_LAN9118', if_true: files('lan9118.c'))
+system_ss.add(when: 'CONFIG_LAN9118_PHY', if_true: files('lan9118_phy.c'))
 system_ss.add(when: 'CONFIG_NE2000_ISA', if_true: files('ne2000-isa.c'))
 system_ss.add(when: 'CONFIG_OPENCORES_ETH', if_true: files('opencores_eth.c'))
 system_ss.add(when: 'CONFIG_XGMAC', if_true: files('xgmac.c'))
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

imx_fec models the same PHY as lan9118_phy. The code is almost the same with
imx_fec having more logging and tracing. Merge these improvements into
lan9118_phy and reuse in imx_fec to fix the code duplication.

Some migration state how resides in the new device model which breaks migration
compatibility for the following machines:
* imx25-pdk
* sabrelite
* mcimx7d-sabre
* mcimx6ul-evk

Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-3-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/imx_fec.h |   9 ++-
 hw/net/imx_fec.c         | 146 ++++-----------------------------------
 hw/net/lan9118_phy.c     |  82 ++++++++++++++++------
 hw/net/Kconfig           |   1 +
 hw/net/trace-events      |  10 +--
 5 files changed, 85 insertions(+), 163 deletions(-)

diff --git a/include/hw/net/imx_fec.h b/include/hw/net/imx_fec.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/imx_fec.h
+++ b/include/hw/net/imx_fec.h
@@ -XXX,XX +XXX,XX @@ OBJECT_DECLARE_SIMPLE_TYPE(IMXFECState, IMX_FEC)
 #define TYPE_IMX_ENET "imx.enet"
 
 #include "hw/sysbus.h"
+#include "hw/net/lan9118_phy.h"
+#include "hw/irq.h"
 #include "net/net.h"
 
 #define ENET_EIR               1
@@ -XXX,XX +XXX,XX @@ struct IMXFECState {
     uint32_t tx_descriptor[ENET_TX_RING_NUM];
     uint32_t tx_ring_num;
 
-    uint32_t phy_status;
-    uint32_t phy_control;
-    uint32_t phy_advertise;
-    uint32_t phy_int;
-    uint32_t phy_int_mask;
+    Lan9118PhyState mii;
+    IRQState mii_irq;
     uint32_t phy_num;
     bool phy_connected;
     struct IMXFECState *phy_consumer;
diff --git a/hw/net/imx_fec.c b/hw/net/imx_fec.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/imx_fec.c
+++ b/hw/net/imx_fec.c
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth_txdescs = {
 
 static const VMStateDescription vmstate_imx_eth = {
     .name = TYPE_IMX_FEC,
-    .version_id = 2,
-    .minimum_version_id = 2,
+    .version_id = 3,
+    .minimum_version_id = 3,
     .fields = (const VMStateField[]) {
         VMSTATE_UINT32_ARRAY(regs, IMXFECState, ENET_MAX),
         VMSTATE_UINT32(rx_descriptor, IMXFECState),
         VMSTATE_UINT32(tx_descriptor[0], IMXFECState),
-        VMSTATE_UINT32(phy_status, IMXFECState),
-        VMSTATE_UINT32(phy_control, IMXFECState),
-        VMSTATE_UINT32(phy_advertise, IMXFECState),
-        VMSTATE_UINT32(phy_int, IMXFECState),
-        VMSTATE_UINT32(phy_int_mask, IMXFECState),
         VMSTATE_END_OF_LIST()
     },
     .subsections = (const VMStateDescription * const []) {
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_imx_eth = {
     },
 };
 
-#define PHY_INT_ENERGYON            (1 << 7)
-#define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
-#define PHY_INT_FAULT               (1 << 5)
-#define PHY_INT_DOWN                (1 << 4)
-#define PHY_INT_AUTONEG_LP          (1 << 3)
-#define PHY_INT_PARFAULT            (1 << 2)
-#define PHY_INT_AUTONEG_PAGE        (1 << 1)
-
 static void imx_eth_update(IMXFECState *s);
 
 /*
@@ -XXX,XX +XXX,XX @@ static void imx_eth_update(IMXFECState *s);
  * For now we don't handle any GPIO/interrupt line, so the OS will
  * have to poll for the PHY status.
  */
-static void imx_phy_update_irq(IMXFECState *s)
+static void imx_phy_update_irq(void *opaque, int n, int level)
 {
-    imx_eth_update(s);
-}
-
-static void imx_phy_update_link(IMXFECState *s)
-{
-    /* Autonegotiation status mirrors link status.  */
-    if (qemu_get_queue(s->nic)->link_down) {
-        trace_imx_phy_update_link("down");
-        s->phy_status &= ~0x0024;
-        s->phy_int |= PHY_INT_DOWN;
-    } else {
-        trace_imx_phy_update_link("up");
-        s->phy_status |= 0x0024;
-        s->phy_int |= PHY_INT_ENERGYON;
-        s->phy_int |= PHY_INT_AUTONEG_COMPLETE;
-    }
-    imx_phy_update_irq(s);
+    imx_eth_update(opaque);
 }
 
 static void imx_eth_set_link(NetClientState *nc)
 {
-    imx_phy_update_link(IMX_FEC(qemu_get_nic_opaque(nc)));
-}
-
-static void imx_phy_reset(IMXFECState *s)
-{
-    trace_imx_phy_reset();
-
-    s->phy_status = 0x7809;
-    s->phy_control = 0x3000;
-    s->phy_advertise = 0x01e1;
-    s->phy_int_mask = 0;
-    s->phy_int = 0;
-    imx_phy_update_link(s);
+    lan9118_phy_update_link(&IMX_FEC(qemu_get_nic_opaque(nc))->mii,
+                            nc->link_down);
 }
 
 static uint32_t imx_phy_read(IMXFECState *s, int reg)
 {
-    uint32_t val;
     uint32_t phy = reg / 32;
 
     if (!s->phy_connected) {
@@ -XXX,XX +XXX,XX @@ static uint32_t imx_phy_read(IMXFECState *s, int reg)
 
     reg %= 32;
 
-    switch (reg) {
-    case 0:     /* Basic Control */
-        val = s->phy_control;
-        break;
-    case 1:     /* Basic Status */
-        val = s->phy_status;
-        break;
-    case 2:     /* ID1 */
-        val = 0x0007;
-        break;
-    case 3:     /* ID2 */
-        val = 0xc0d1;
-        break;
-    case 4:     /* Auto-neg advertisement */
-        val = s->phy_advertise;
-        break;
-    case 5:     /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
-        break;
-    case 6:     /* Auto-neg Expansion */
-        val = 1;
-        break;
-    case 29:    /* Interrupt source.  */
-        val = s->phy_int;
-        s->phy_int = 0;
-        imx_phy_update_irq(s);
-        break;
-    case 30:    /* Interrupt mask */
-        val = s->phy_int_mask;
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy]%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        val = 0;
-        break;
-    }
-
-    trace_imx_phy_read(val, phy, reg);
-
-    return val;
+    return lan9118_phy_read(&s->mii, reg);
 }
 
 static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
@@ -XXX,XX +XXX,XX @@ static void imx_phy_write(IMXFECState *s, int reg, uint32_t val)
 
     reg %= 32;
 
-    trace_imx_phy_write(val, phy, reg);
-
-    switch (reg) {
-    case 0:     /* Basic Control */
-        if (val & 0x8000) {
-            imx_phy_reset(s);
-        } else {
-            s->phy_control = val & 0x7980;
-            /* Complete autonegotiation immediately.  */
-            if (val & 0x1000) {
-                s->phy_status |= 0x0020;
-            }
-        }
-        break;
-    case 4:     /* Auto-neg advertisement */
-        s->phy_advertise = (val & 0x2d7f) | 0x80;
-        break;
-    case 30:    /* Interrupt mask */
-        s->phy_int_mask = val & 0xff;
-        imx_phy_update_irq(s);
-        break;
-    case 17:
-    case 18:
-    case 27:
-    case 31:
-        qemu_log_mask(LOG_UNIMP, "[%s.phy)%s: reg %d not implemented\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    default:
-        qemu_log_mask(LOG_GUEST_ERROR, "[%s.phy]%s: Bad address at offset %d\n",
-                      TYPE_IMX_FEC, __func__, reg);
-        break;
-    }
+    lan9118_phy_write(&s->mii, reg, val);
 }
 
 static void imx_fec_read_bd(IMXFECBufDesc *bd, dma_addr_t addr)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_reset(DeviceState *d)
 
     s->rx_descriptor = 0;
     memset(s->tx_descriptor, 0, sizeof(s->tx_descriptor));
-
-    /* We also reset the PHY */
-    imx_phy_reset(s);
 }
 
 static uint32_t imx_default_read(IMXFECState *s, uint32_t index)
@@ -XXX,XX +XXX,XX @@ static void imx_eth_realize(DeviceState *dev, Error **errp)
     sysbus_init_irq(sbd, &s->irq[0]);
     sysbus_init_irq(sbd, &s->irq[1]);
 
+    qemu_init_irq(&s->mii_irq, imx_phy_update_irq, s, 0);
+    object_initialize_child(OBJECT(s), "mii", &s->mii, TYPE_LAN9118_PHY);
+    if (!sysbus_realize_and_unref(SYS_BUS_DEVICE(&s->mii), errp)) {
+        return;
+    }
+    qdev_connect_gpio_out(DEVICE(&s->mii), 0, &s->mii_irq);
+
     qemu_macaddr_default_if_unset(&s->conf.macaddr);
 
     s->nic = qemu_new_nic(&imx_eth_net_info, &s->conf,
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
  * Copyright (c) 2009 CodeSourcery, LLC.
  * Written by Paul Brook
  *
+ * Copyright (c) 2013 Jean-Christophe Dubois. <jcd@tribudubois.net>
+ *
  * This code is licensed under the GNU GPL v2
  *
  * Contributions after 2012-01-13 are licensed under the terms of the
@@ -XXX,XX +XXX,XX @@
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
 #include "qemu/log.h"
+#include "trace.h"
 
 #define PHY_INT_ENERGYON            (1 << 7)
 #define PHY_INT_AUTONEG_COMPLETE    (1 << 6)
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
 
     switch (reg) {
     case 0: /* Basic Control */
-        return s->control;
+        val = s->control;
+        break;
     case 1: /* Basic Status */
-        return s->status;
+        val = s->status;
+        break;
     case 2: /* ID1 */
-        return 0x0007;
+        val = 0x0007;
+        break;
     case 3: /* ID2 */
-        return 0xc0d1;
+        val = 0xc0d1;
+        break;
     case 4: /* Auto-neg advertisement */
-        return s->advertise;
+        val = s->advertise;
+        break;
     case 5: /* Auto-neg Link Partner Ability */
-        return 0x0f71;
+        val = 0x0f71;
+        break;
     case 6: /* Auto-neg Expansion */
-        return 1;
-        /* TODO 17, 18, 27, 29, 30, 31 */
+        val = 1;
+        break;
     case 29: /* Interrupt source. */
         val = s->ints;
         s->ints = 0;
         lan9118_phy_update_irq(s);
-        return val;
+        break;
     case 30: /* Interrupt mask */
-        return s->int_mask;
+        val = s->int_mask;
+        break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        val = 0;
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_read: PHY read reg %d\n", reg);
-        return 0;
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        val = 0;
+        break;
     }
+
+    trace_lan9118_phy_read(val, reg);
+
+    return val;
 }
 
 void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
 {
+    trace_lan9118_phy_write(val, reg);
+
     switch (reg) {
     case 0: /* Basic Control */
         if (val & 0x8000) {
             lan9118_phy_reset(s);
-            break;
-        }
-        s->control = val & 0x7980;
-        /* Complete autonegotiation immediately. */
-        if (val & 0x1000) {
-            s->status |= 0x0020;
+        } else {
+            s->control = val & 0x7980;
+            /* Complete autonegotiation immediately. */
+            if (val & 0x1000) {
+                s->status |= 0x0020;
+            }
         }
         break;
     case 4: /* Auto-neg advertisement */
         s->advertise = (val & 0x2d7f) | 0x80;
         break;
-        /* TODO 17, 18, 27, 31 */
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
         lan9118_phy_update_irq(s);
         break;
+    case 17:
+    case 18:
+    case 27:
+    case 31:
+        qemu_log_mask(LOG_UNIMP, "%s: reg %d not implemented\n",
+                      __func__, reg);
+        break;
     default:
-        qemu_log_mask(LOG_GUEST_ERROR,
-                      "lan9118_phy_write: PHY write reg %d = 0x%04x\n", reg, val);
+        qemu_log_mask(LOG_GUEST_ERROR, "%s: Bad address at offset %d\n",
+                      __func__, reg);
+        break;
     }
 }
 
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
+        trace_lan9118_phy_update_link("down");
         s->status &= ~0x0024;
         s->ints |= PHY_INT_DOWN;
     } else {
+        trace_lan9118_phy_update_link("up");
         s->status |= 0x0024;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
 
 void lan9118_phy_reset(Lan9118PhyState *s)
 {
+    trace_lan9118_phy_reset();
+
     s->control = 0x3000;
     s->status = 0x7809;
     s->advertise = 0x01e1;
@@ -XXX,XX +XXX,XX @@ static const VMStateDescription vmstate_lan9118_phy = {
     .version_id = 1,
     .minimum_version_id = 1,
     .fields = (const VMStateField[]) {
-        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(status, Lan9118PhyState),
+        VMSTATE_UINT16(control, Lan9118PhyState),
         VMSTATE_UINT16(advertise, Lan9118PhyState),
         VMSTATE_UINT16(ints, Lan9118PhyState),
         VMSTATE_UINT16(int_mask, Lan9118PhyState),
diff --git a/hw/net/Kconfig b/hw/net/Kconfig
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/Kconfig
+++ b/hw/net/Kconfig
@@ -XXX,XX +XXX,XX @@ config ALLWINNER_SUN8I_EMAC
 
 config IMX_FEC
     bool
+    select LAN9118_PHY
 
 config CADENCE
     bool
diff --git a/hw/net/trace-events b/hw/net/trace-events
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -XXX,XX +XXX,XX @@ allwinner_sun8i_emac_set_link(bool active) "Set link: active=%u"
 allwinner_sun8i_emac_read(uint64_t offset, uint64_t val) "MMIO read: offset=0x%" PRIx64 " value=0x%" PRIx64
 allwinner_sun8i_emac_write(uint64_t offset, uint64_t val) "MMIO write: offset=0x%" PRIx64 " value=0x%" PRIx64
 
+# lan9118_phy.c
+lan9118_phy_read(uint16_t val, int reg) "[0x%02x] -> 0x%04" PRIx16
+lan9118_phy_write(uint16_t val, int reg) "[0x%02x] <- 0x%04" PRIx16
+lan9118_phy_update_link(const char *s) "%s"
+lan9118_phy_reset(void) ""
+
 # lance.c
 lance_mem_readw(uint64_t addr, uint32_t ret) "addr=0x%"PRIx64"val=0x%04x"
 lance_mem_writew(uint64_t addr, uint32_t val) "addr=0x%"PRIx64"val=0x%04x"
@@ -XXX,XX +XXX,XX @@ i82596_set_multicast(uint16_t count) "Added %d multicast entries"
 i82596_channel_attention(void *s) "%p: Received CHANNEL ATTENTION"
 
 # imx_fec.c
-imx_phy_read(uint32_t val, int phy, int reg) "0x%04"PRIx32" <= phy[%d].reg[%d]"
 imx_phy_read_num(int phy, int configured) "read request from unconfigured phy %d (configured %d)"
-imx_phy_write(uint32_t val, int phy, int reg) "0x%04"PRIx32" => phy[%d].reg[%d]"
 imx_phy_write_num(int phy, int configured) "write request to unconfigured phy %d (configured %d)"
-imx_phy_update_link(const char *s) "%s"
-imx_phy_reset(void) ""
 imx_fec_read_bd(uint64_t addr, int flags, int len, int data) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x"
 imx_enet_read_bd(uint64_t addr, int flags, int len, int data, int options, int status) "tx_bd 0x%"PRIx64" flags 0x%04x len %d data 0x%08x option 0x%04x status 0x%04x"
 imx_eth_tx_bd_busy(void) "tx_bd ran out of descriptors to transmit"
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Turns 0x70 into 0xe0 (== 0x70 << 1) which adds the missing MII_ANLPAR_TX and
fixes the MSB of selector field to be zero, as specified in the datasheet.

Fixes: 2a424990170b "LAN9118 emulation"
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241102125724.532843-4-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
         val = s->advertise;
         break;
     case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0f71;
+        val = 0x0fe1;
         break;
     case 6: /* Auto-neg Expansion */
         val = 1;
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

Prefer named constants over magic values for better readability.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-5-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/hw/net/mii.h |  6 +++++
 hw/net/lan9118_phy.c | 63 ++++++++++++++++++++++++++++----------------
 2 files changed, 46 insertions(+), 23 deletions(-)

diff --git a/include/hw/net/mii.h b/include/hw/net/mii.h
index XXXXXXX..XXXXXXX 100644
--- a/include/hw/net/mii.h
+++ b/include/hw/net/mii.h
@@ -XXX,XX +XXX,XX @@
 #define MII_BMSR_JABBER     (1 << 1)  /* Jabber detected */
 #define MII_BMSR_EXTCAP     (1 << 0)  /* Ext-reg capability */
 
+#define MII_ANAR_RFAULT     (1 << 13) /* Say we can detect faults */
 #define MII_ANAR_PAUSE_ASYM (1 << 11) /* Try for asymmetric pause */
 #define MII_ANAR_PAUSE      (1 << 10) /* Try for pause */
 #define MII_ANAR_TXFD       (1 << 8)
@@ -XXX,XX +XXX,XX @@
 #define MII_ANAR_10FD       (1 << 6)
 #define MII_ANAR_10         (1 << 5)
 #define MII_ANAR_CSMACD     (1 << 0)
+#define MII_ANAR_SELECT     (0x001f)  /* Selector bits */
 
 #define MII_ANLPAR_ACK      (1 << 14)
 #define MII_ANLPAR_PAUSEASY (1 << 11) /* can pause asymmetrically */
@@ -XXX,XX +XXX,XX @@
 #define RTL8201CP_PHYID1    0x0000
 #define RTL8201CP_PHYID2    0x8201
 
+/* SMSC LAN9118 */
+#define SMSCLAN9118_PHYID1  0x0007
+#define SMSCLAN9118_PHYID2  0xc0d1
+
 /* RealTek 8211E */
 #define RTL8211E_PHYID1     0x001c
 #define RTL8211E_PHYID2     0xc915
diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@
 
 #include "qemu/osdep.h"
 #include "hw/net/lan9118_phy.h"
+#include "hw/net/mii.h"
 #include "hw/irq.h"
 #include "hw/resettable.h"
 #include "migration/vmstate.h"
@@ -XXX,XX +XXX,XX @@ uint16_t lan9118_phy_read(Lan9118PhyState *s, int reg)
     uint16_t val;
 
     switch (reg) {
-    case 0: /* Basic Control */
+    case MII_BMCR:
         val = s->control;
         break;
-    case 1: /* Basic Status */
+    case MII_BMSR:
         val = s->status;
         break;
-    case 2: /* ID1 */
-        val = 0x0007;
+    case MII_PHYID1:
+        val = SMSCLAN9118_PHYID1;
         break;
-    case 3: /* ID2 */
-        val = 0xc0d1;
+    case MII_PHYID2:
+        val = SMSCLAN9118_PHYID2;
         break;
-    case 4: /* Auto-neg advertisement */
+    case MII_ANAR:
         val = s->advertise;
         break;
-    case 5: /* Auto-neg Link Partner Ability */
-        val = 0x0fe1;
+    case MII_ANLPAR:
+        val = MII_ANLPAR_PAUSEASY | MII_ANLPAR_PAUSE | MII_ANLPAR_T4 |
+              MII_ANLPAR_TXFD | MII_ANLPAR_TX | MII_ANLPAR_10FD |
+              MII_ANLPAR_10 | MII_ANLPAR_CSMACD;
         break;
-    case 6: /* Auto-neg Expansion */
-        val = 1;
+    case MII_ANER:
+        val = MII_ANER_NWAY;
         break;
     case 29: /* Interrupt source. */
         val = s->ints;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
     trace_lan9118_phy_write(val, reg);
 
     switch (reg) {
-    case 0: /* Basic Control */
-        if (val & 0x8000) {
+    case MII_BMCR:
+        if (val & MII_BMCR_RESET) {
             lan9118_phy_reset(s);
         } else {
-            s->control = val & 0x7980;
+            s->control = val & (MII_BMCR_LOOPBACK | MII_BMCR_SPEED100 |
+                                MII_BMCR_AUTOEN | MII_BMCR_PDOWN | MII_BMCR_FD |
+                                MII_BMCR_CTST);
             /* Complete autonegotiation immediately. */
-            if (val & 0x1000) {
-                s->status |= 0x0020;
+            if (val & MII_BMCR_AUTOEN) {
+                s->status |= MII_BMSR_AN_COMP;
             }
         }
         break;
-    case 4: /* Auto-neg advertisement */
-        s->advertise = (val & 0x2d7f) | 0x80;
+    case MII_ANAR:
+        s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
+                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
+                               MII_ANAR_SELECT))
+                     | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
         s->int_mask = val & 0xff;
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_update_link(Lan9118PhyState *s, bool link_down)
     /* Autonegotiation status mirrors link status. */
     if (link_down) {
         trace_lan9118_phy_update_link("down");
-        s->status &= ~0x0024;
+        s->status &= ~(MII_BMSR_AN_COMP | MII_BMSR_LINK_ST);
         s->ints |= PHY_INT_DOWN;
     } else {
         trace_lan9118_phy_update_link("up");
-        s->status |= 0x0024;
+        s->status |= MII_BMSR_AN_COMP | MII_BMSR_LINK_ST;
         s->ints |= PHY_INT_ENERGYON;
         s->ints |= PHY_INT_AUTONEG_COMPLETE;
     }
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_reset(Lan9118PhyState *s)
 {
     trace_lan9118_phy_reset();
 
-    s->control = 0x3000;
-    s->status = 0x7809;
-    s->advertise = 0x01e1;
+    s->control = MII_BMCR_AUTOEN | MII_BMCR_SPEED100;
+    s->status = MII_BMSR_100TX_FD
+                | MII_BMSR_100TX_HD
+                | MII_BMSR_10T_FD
+                | MII_BMSR_10T_HD
+                | MII_BMSR_AUTONEG
+                | MII_BMSR_EXTCAP;
+    s->advertise = MII_ANAR_TXFD
+                   | MII_ANAR_TX
+                   | MII_ANAR_10FD
+                   | MII_ANAR_10
+                   | MII_ANAR_CSMACD;
     s->int_mask = 0;
     s->ints = 0;
     lan9118_phy_update_link(s, s->link_down);
-- 
2.34.1

From: Bernhard Beschow <shentey@gmail.com>

The real device advertises this mode and the device model already advertises
100 mbps half duplex and 10 mbps full+half duplex. So advertise this mode to
make the model more realistic.

Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Bernhard Beschow <shentey@gmail.com>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Message-id: 20241102125724.532843-6-shentey@gmail.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/net/lan9118_phy.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/hw/net/lan9118_phy.c b/hw/net/lan9118_phy.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/net/lan9118_phy.c
+++ b/hw/net/lan9118_phy.c
@@ -XXX,XX +XXX,XX @@ void lan9118_phy_write(Lan9118PhyState *s, int reg, uint16_t val)
         break;
     case MII_ANAR:
         s->advertise = (val & (MII_ANAR_RFAULT | MII_ANAR_PAUSE_ASYM |
-                               MII_ANAR_PAUSE | MII_ANAR_10FD | MII_ANAR_10 |
-                               MII_ANAR_SELECT))
+                               MII_ANAR_PAUSE | MII_ANAR_TXFD | MII_ANAR_10FD |
+                               MII_ANAR_10 | MII_ANAR_SELECT))
                      | MII_ANAR_TX;
         break;
     case 30: /* Interrupt mask */
-- 
2.34.1

For IEEE fused multiply-add, the (0 * inf) + NaN case should raise
Invalid for the multiplication of 0 by infinity.  Currently we handle
this in the per-architecture ifdef ladder in pickNaNMulAdd().
However, since this isn't really architecture specific we can hoist
it up to the generic code.

For the cases where the infzero test in pickNaNMulAdd was
returning 2, we can delete the check entirely and allow the
code to fall into the normal pick-a-NaN handling, because this
will return 2 anyway (input 'c' being the only NaN in this case).
For the cases where infzero was returning 3 to indicate "return
the default NaN", we must retain that "return 3".

For Arm, this looks like it might be a behaviour change because we
used to set float_flag_invalid | float_flag_invalid_imz only if C is
a quiet NaN.  However, it is not, because Arm target code never looks
at float_flag_invalid_imz, and for the (0 * inf) + SNaN case we
already raised float_flag_invalid via the "abc_mask &
float_cmask_snan" check in pick_nan_muladd.

For any target architecture using the "default implementation" at the
bottom of the ifdef, this is a behaviour change but will be fixing a
bug (where we failed to raise the Invalid exception for (0 * inf +
QNaN).  The architectures using the default case are:
 * hppa
 * i386
 * sh4
 * tricore

The x86, Tricore and SH4 CPU architecture manuals are clear that this
should have raised Invalid; HPPA is a bit vaguer but still seems
clear enough.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-2-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 13 +++++++------
 fpu/softfloat-specialize.c.inc | 29 +----------------------------
 2 files changed, 8 insertions(+), 34 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             int ab_mask, int abc_mask)
 {
     int which;
+    bool infzero = (ab_mask == float_cmask_infzero);
 
     if (unlikely(abc_mask & float_cmask_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls,
-                          ab_mask == float_cmask_infzero, s);
+    if (infzero) {
+        /* This is (0 * inf) + NaN or (inf * 0) + NaN */
+        float_raise(float_flag_invalid | float_flag_invalid_imz, s);
+    }
+
+    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
 
     if (s->default_nan_mode || which == 3) {
-        /*
-         * Note that this check is after pickNaNMulAdd so that function
-         * has an opportunity to set the Invalid flag for infzero.
-         */
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * the default NaN
      */
     if (infzero && is_qnan(c_cls)) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * case sets InvalidOp and returns the default NaN
          */
         if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
             return 3;
         }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        if (infzero) {
-            float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-            return 2;
-        }
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
      * case sets InvalidOp and returns the input value 'c'
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
+
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * to return an input NaN if we have one (ie c) rather than generating
      * a default NaN
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
 
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_RISCV)
-    /* For RISC-V, InvalidOp is set when multiplicands are Inf and zero */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-    }
     return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
         return 3;
     }
 
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 2;
     }
 #elif defined(TARGET_SPARC)
-    /* For (inf,0,nan) return c. */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     /* Prefer SNaN over QNaN, order C, B, A. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
      * an input NaN if we have one (ie c).
      */
-    if (infzero) {
-        float_raise(float_flag_invalid | float_flag_invalid_imz, status);
-        return 2;
-    }
     if (status->use_first_nan) {
         if (is_nan(a_cls)) {
             return 0;
-- 
2.34.1

If the target sets default_nan_mode then we're always going to return
the default NaN, and pickNaNMulAdd() no longer has any side effects.
For consistency with pickNaN(), check for default_nan_mode before
calling pickNaNMulAdd().

When we convert pickNaNMulAdd() to allow runtime selection of the NaN
propagation rule, this means we won't have to make the targets which
use default_nan_mode also set a propagation rule.

Since RiscV always uses default_nan_mode, this allows us to remove
its ifdef case from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-3-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 8 ++++++--
 fpu/softfloat-specialize.c.inc | 9 +++++++--
 2 files changed, 13 insertions(+), 4 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         float_raise(float_flag_invalid | float_flag_invalid_imz, s);
     }
 
-    which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    if (s->default_nan_mode) {
+        which = 3;
+    } else {
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+    }
 
-    if (s->default_nan_mode || which == 3) {
+    if (which == 3) {
         parts_default_nan(a, s);
         return a;
     }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    /*
+     * We guarantee not to require the target to tell us how to
+     * pick a NaN if we're always returning the default NaN.
+     * But if we're not in default-NaN mode then the target must
+     * specify.
+     */
+    assert(!status->default_nan_mode);
 #if defined(TARGET_ARM)
     /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
      * the default NaN
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     } else {
         return 1;
     }
-#elif defined(TARGET_RISCV)
-    return 3; /* default NaN */
 #elif defined(TARGET_S390X)
     if (infzero) {
         return 3;
-- 
2.34.1

IEEE 758 does not define a fixed rule for what NaN to return in
the case of a fused multiply-add of inf * 0 + NaN. Different
architectures thus do different things:
 * some return the default NaN
 * some return the input NaN
 * Arm returns the default NaN if the input NaN is quiet,
   and the input NaN if it is signalling

We want to make this logic be runtime selected rather than
hardcoded into the binary, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

In this commit we add an enum for the propagation rule, the field in
float_status, and the corresponding getters and setters.  We change
pickNaNMulAdd to honour this, but because all targets still leave
this field at its default 0 value, the fallback logic will pick the
rule type with the old ifdef ladder.

Note that four architectures both use the muladd softfloat functions
and did not have a branch of the ifdef ladder to specify their
behaviour (and so were ending up with the "default" case, probably
wrongly): i386, HPPA, SH4 and Tricore.  SH4 and Tricore both set
default_nan_mode, and so will never get into pickNaNMulAdd().  For
HPPA and i386 we retain the same behaviour as the old default-case,
which is to not ever return the default NaN.  This might not be
correct but it is not a behaviour change.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-4-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 ++++
 include/fpu/softfloat-types.h   | 23 +++++++++
 fpu/softfloat-specialize.c.inc  | 91 ++++++++++++++++++++++-----------
 3 files changed, 95 insertions(+), 30 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
+                                             float_status *status)
+{
+    status->float_infzeronan_rule = rule;
+}
+
 static inline void set_flush_to_zero(bool val, float_status *status)
 {
     status->flush_to_zero = val;
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
+{
+    return status->float_infzeronan_rule;
+}
+
 static inline bool get_flush_to_zero(float_status *status)
 {
     return status->flush_to_zero;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * Rule for result of fused multiply-add 0 * Inf + NaN.
+ * This must be a NaN, but implementations differ on whether this
+ * is the input NaN or the default NaN.
+ *
+ * You don't need to set this if default_nan_mode is enabled.
+ * When not in default-NaN mode, it is an error for the target
+ * not to set the rule in float_status if it uses muladd, and we
+ * will assert if we need to handle an input NaN and no rule was
+ * selected.
+ */
+typedef enum __attribute__((__packed__)) {
+    /* No propagation rule specified */
+    float_infzeronan_none = 0,
+    /* Result is never the default NaN (so always the input NaN) */
+    float_infzeronan_dnan_never,
+    /* Result is always the default NaN */
+    float_infzeronan_dnan_always,
+    /* Result is the default NaN if the input NaN is quiet */
+    float_infzeronan_dnan_if_qnan,
+} FloatInfZeroNaNRule;
+
 /*
  * Floating Point Status. Individual architectures may maintain
  * several versions of float_status for different functions. The
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
     bool flush_to_zero;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
+    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      * specify.
      */
     assert(!status->default_nan_mode);
+
+    if (rule == float_infzeronan_none) {
+        /*
+         * Temporarily fall back to ifdef ladder
+         */
 #if defined(TARGET_ARM)
-    /* For ARM, the (inf,zero,qnan) case sets InvalidOp and returns
-     * the default NaN
-     */
-    if (infzero && is_qnan(c_cls)) {
-        return 3;
+        /*
+         * For ARM, the (inf,zero,qnan) case returns the default NaN,
+         * but (inf,zero,snan) returns the input NaN.
+         */
+        rule = float_infzeronan_dnan_if_qnan;
+#elif defined(TARGET_MIPS)
+        if (snan_bit_is_one(status)) {
+            /*
+             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+             * case sets InvalidOp and returns the default NaN
+             */
+            rule = float_infzeronan_dnan_always;
+        } else {
+            /*
+             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+             * case sets InvalidOp and returns the input value 'c'
+             */
+            rule = float_infzeronan_dnan_never;
+        }
+#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+        /*
+         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+         * case sets InvalidOp and returns the input value 'c'
+         */
+        /*
+         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+         * to return an input NaN if we have one (ie c) rather than generating
+         * a default NaN
+         */
+        rule = float_infzeronan_dnan_never;
+#elif defined(TARGET_S390X)
+        rule = float_infzeronan_dnan_always;
+#endif
     }
 
+    if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the default NaN here,
+         * and some return the input NaN.
+         */
+        switch (rule) {
+        case float_infzeronan_dnan_never:
+            return 2;
+        case float_infzeronan_dnan_always:
+            return 3;
+        case float_infzeronan_dnan_if_qnan:
+            return is_qnan(c_cls) ? 3 : 2;
+        default:
+            g_assert_not_reached();
+        }
+    }
+
+#if defined(TARGET_ARM)
+
     /* This looks different from the ARM ARM pseudocode, because the ARM ARM
      * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 #elif defined(TARGET_MIPS)
     if (snan_bit_is_one(status)) {
-        /*
-         * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-         * case sets InvalidOp and returns the default NaN
-         */
-        if (infzero) {
-            return 3;
-        }
         /* Prefer sNaN over qNaN, in the a, b, c order. */
         if (is_snan(a_cls)) {
             return 0;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
             return 2;
         }
     } else {
-        /*
-         * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
         /* Prefer sNaN over qNaN, in the c, a, b order. */
         if (is_snan(c_cls)) {
             return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 #elif defined(TARGET_LOONGARCH64)
-    /*
-     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-     * case sets InvalidOp and returns the input value 'c'
-     */
-
     /* Prefer sNaN over qNaN, in the c, a, b order. */
     if (is_snan(c_cls)) {
         return 2;
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_PPC)
-    /* For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-     * to return an input NaN if we have one (ie c) rather than generating
-     * a default NaN
-     */
-
     /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
      * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
      */
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         return 1;
     }
 #elif defined(TARGET_S390X)
-    if (infzero) {
-        return 3;
-    }
-
     if (is_snan(a_cls)) {
         return 0;
     } else if (is_snan(b_cls)) {
-- 
2.34.1

Explicitly set a rule in the softfloat tests for the inf-zero-nan
muladd special case.  In meson.build we put -DTARGET_ARM in fpcflags,
and so we should select here the Arm rule of
float_infzeronan_dnan_if_qnan.

Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241202131347.498124-5-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 5 +++++
 tests/fp/fp-test.c  | 5 +++++
 2 files changed, 10 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
 {
     bench_func_t f;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
 {
     unsigned int i;
 
+    /*
+     * These implementation-defined choices for various things IEEE
+     * doesn't specify match those used by the Arm architecture.
+     */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
     verCases_maxErrorCount = n_max_errors;
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the Arm target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-6-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 4 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
+ *    and the input NaN if it is signalling
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_ARM)
-        /*
-         * For ARM, the (inf,zero,qnan) case returns the default NaN,
-         * but (inf,zero,snan) returns the input NaN.
-         */
-        rule = float_infzeronan_dnan_if_qnan;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             /*
              * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for s390, so we
can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-7-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_infzeronan_rule(float_infzeronan_dnan_always,
+                                  &env->fpu_status);
        /* fall through */
     case RESET_TYPE_S390_CPU_NORMAL:
         env->psw.mask &= ~PSW_MASK_RI;
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * a default NaN
          */
         rule = float_infzeronan_dnan_never;
-#elif defined(TARGET_S390X)
-        rule = float_infzeronan_dnan_always;
 #endif
     }
 
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the PPC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-8-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 7 +++++++
 fpu/softfloat-specialize.c.inc | 7 +------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
+     * to return an input NaN if we have one (ie c) rather than generating
+     * a default NaN
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->vec_status);
 
     for (i = 0; i < ARRAY_SIZE(env->spr_cb); i++) {
         ppc_spr_t *spr = &env->spr_cb[i];
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
              */
             rule = float_infzeronan_dnan_never;
         }
-#elif defined(TARGET_PPC) || defined(TARGET_SPARC) || \
+#elif defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
          */
-        /*
-         * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
-         * to return an input NaN if we have one (ie c) rather than generating
-         * a default NaN
-         */
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the MIPS target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-9-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       |  9 +++++++++
 target/mips/msa.c              |  4 ++++
 fpu/softfloat-specialize.c.inc | 16 +---------------
 3 files changed, 14 insertions(+), 15 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_flush_mode(CPUMIPSState *env)
 static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
+    FloatInfZeroNaNRule izn_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     set_snan_bit_is_one(!nan2008, &env->active_fpu.fp_status);
     set_default_nan_mode(!nan2008, &env->active_fpu.fp_status);
+    /*
+     * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
+     * case sets InvalidOp and returns the default NaN.
+     * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'.
+     */
+    izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
+    set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
 
     /* set proper signanling bit meaning ("1" means "quiet") */
     set_snan_bit_is_one(0, &env->active_tc.msa_fp_status);
+
+    /* Inf * 0 + NaN returns the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never,
+                              &env->active_tc.msa_fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            /*
-             * For MIPS systems that conform to IEEE754-1985, the (inf,zero,nan)
-             * case sets InvalidOp and returns the default NaN
-             */
-            rule = float_infzeronan_dnan_always;
-        } else {
-            /*
-             * For MIPS systems that conform to IEEE754-2008, the (inf,zero,nan)
-             * case sets InvalidOp and returns the input value 'c'
-             */
-            rule = float_infzeronan_dnan_never;
-        }
-#elif defined(TARGET_SPARC) || \
+#if defined(TARGET_SPARC) || \
     defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the SPARC target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-10-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_SPARC) || \
-    defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the xtensa target,
so we can remove the ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-11-peter.maydell@linaro.org
---
 target/xtensa/cpu.c            | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/xtensa/cpu.c b/target/xtensa/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/cpu.c
+++ b/target/xtensa/cpu.c
@@ -XXX,XX +XXX,XX @@ static void xtensa_cpu_reset_hold(Object *obj, ResetType type)
     reset_mmu(env);
     cs->halted = env->runstall;
 #endif
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
     set_no_signaling_nans(!dfpu, &env->fp_status);
     xtensa_use_first_nan(env, !dfpu);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_XTENSA) || defined(TARGET_HPPA) || \
+#if defined(TARGET_HPPA) || \
     defined(TARGET_I386) || defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the x86 target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-12-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 7 +++++++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->mmx_status);
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->sse_status);
+    /*
+     * Only SSE has multiply-add instructions. In the SDM Section 14.5.2
+     * "Fused-Multiply-ADD (FMA) Numeric Behavior" the NaN handling is
+     * specified -- for 0 * inf + NaN the input NaN is selected, and if
+     * there are multiple input NaNs they are selected in the order a, b, c.
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
          * Temporarily fall back to ifdef ladder
          */
 #if defined(TARGET_HPPA) || \
-    defined(TARGET_I386) || defined(TARGET_LOONGARCH)
+    defined(TARGET_LOONGARCH)
         /*
          * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
          * case sets InvalidOp and returns the input value 'c'
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the loongarch target.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-13-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 5 +++++
 fpu/softfloat-specialize.c.inc    | 7 +------
 2 files changed, 6 insertions(+), 6 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
                             &env->fp_status);
     set_flush_to_zero(0, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
+     * case sets InvalidOp and returns the input value 'c'
+     */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         /*
          * Temporarily fall back to ifdef ladder
          */
-#if defined(TARGET_HPPA) || \
-    defined(TARGET_LOONGARCH)
-        /*
-         * For LoongArch systems that conform to IEEE754-2008, the (inf,zero,nan)
-         * case sets InvalidOp and returns the input value 'c'
-         */
+#if defined(TARGET_HPPA)
         rule = float_infzeronan_dnan_never;
 #endif
     }
-- 
2.34.1

Set the FloatInfZeroNaNRule explicitly for the HPPA target,
so we can remove the ifdef from pickNaNMulAdd().

As this is the last target to be converted to explicitly setting
the rule, we can remove the fallback code in pickNaNMulAdd()
entirely.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-14-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       |  2 ++
 fpu/softfloat-specialize.c.inc | 13 +------------
 2 files changed, 3 insertions(+), 12 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /* For inf * 0 + NaN, return the input NaN */
+    set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, float_status *status)
 {
-    FloatInfZeroNaNRule rule = status->float_infzeronan_rule;
-
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
      */
     assert(!status->default_nan_mode);
 
-    if (rule == float_infzeronan_none) {
-        /*
-         * Temporarily fall back to ifdef ladder
-         */
-#if defined(TARGET_HPPA)
-        rule = float_infzeronan_dnan_never;
-#endif
-    }
-
     if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the default NaN here,
          * and some return the input NaN.
          */
-        switch (rule) {
+        switch (status->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
             return 2;
         case float_infzeronan_dnan_always:
-- 
2.34.1

The new implementation of pickNaNMulAdd() will find it convenient
to know whether at least one of the three arguments to the muladd
was a signaling NaN. We already calculate that in the caller,
so pass it in as a new bool have_snan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-15-peter.maydell@linaro.org
---
 fpu/softfloat-parts.c.inc      | 5 +++--
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
 {
     int which;
     bool infzero = (ab_mask == float_cmask_infzero);
+    bool have_snan = (abc_mask & float_cmask_snan);
 
-    if (unlikely(abc_mask & float_cmask_snan)) {
+    if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     if (s->default_nan_mode) {
         which = 3;
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, s);
+        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 | Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
 *----------------------------------------------------------------------------*/
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, float_status *status)
+                         bool infzero, bool have_snan, float_status *status)
 {
     /*
      * We guarantee not to require the target to tell us how to
-- 
2.34.1

IEEE 758 does not define a fixed rule for which NaN to pick as the
result if both operands of a 3-operand fused multiply-add operation
are NaNs.  As a result different architectures have ended up with
different rules for propagating NaNs.

QEMU currently hardcodes the NaN propagation logic into the binary
because pickNaNMulAdd() has an ifdef ladder for different targets.
We want to make the propagation rule instead be selectable at
runtime, because:
 * this will let us have multiple targets in one QEMU binary
 * the Arm FEAT_AFP architectural feature includes letting
   the guest select a NaN propagation rule at runtime

It's valid not to set a propagation rule if default_nan_mode is
enabled, because in that case there's no need to pick a NaN; all the
callers of pickNaNMulAdd() catch this case and skip calling it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-16-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h |  11 +++
 include/fpu/softfloat-types.h   |  55 +++++++++++
 fpu/softfloat-specialize.c.inc  | 167 ++++++++------------------------
 3 files changed, 107 insertions(+), 126 deletions(-)

diff --git a/include/fpu/softfloat-helpers.h b/include/fpu/softfloat-helpers.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-helpers.h
+++ b/include/fpu/softfloat-helpers.h
@@ -XXX,XX +XXX,XX @@ static inline void set_float_2nan_prop_rule(Float2NaNPropRule rule,
     status->float_2nan_prop_rule = rule;
 }
 
+static inline void set_float_3nan_prop_rule(Float3NaNPropRule rule,
+                                            float_status *status)
+{
+    status->float_3nan_prop_rule = rule;
+}
+
 static inline void set_float_infzeronan_rule(FloatInfZeroNaNRule rule,
                                              float_status *status)
 {
@@ -XXX,XX +XXX,XX @@ static inline Float2NaNPropRule get_float_2nan_prop_rule(float_status *status)
     return status->float_2nan_prop_rule;
 }
 
+static inline Float3NaNPropRule get_float_3nan_prop_rule(float_status *status)
+{
+    return status->float_3nan_prop_rule;
+}
+
 static inline FloatInfZeroNaNRule get_float_infzeronan_rule(float_status *status)
 {
     return status->float_infzeronan_rule;
diff --git a/include/fpu/softfloat-types.h b/include/fpu/softfloat-types.h
index XXXXXXX..XXXXXXX 100644
--- a/include/fpu/softfloat-types.h
+++ b/include/fpu/softfloat-types.h
@@ -XXX,XX +XXX,XX @@ this code that are retained.
 #ifndef SOFTFLOAT_TYPES_H
 #define SOFTFLOAT_TYPES_H
 
+#include "hw/registerfields.h"
+
 /*
  * Software IEC/IEEE floating-point types.
  */
@@ -XXX,XX +XXX,XX @@ typedef enum __attribute__((__packed__)) {
     float_2nan_prop_x87,
 } Float2NaNPropRule;
 
+/*
+ * 3-input NaN propagation rule, for fused multiply-add. Individual
+ * architectures have different rules for which input NaN is
+ * propagated to the output when there is more than one NaN on the
+ * input.
+ *
+ * If default_nan_mode is enabled then it is valid not to set a NaN
+ * propagation rule, because the softfloat code guarantees not to try
+ * to pick a NaN to propagate in default NaN mode.  When not in
+ * default-NaN mode, it is an error for the target not to set the rule
+ * in float_status if it uses a muladd, and we will assert if we need
+ * to handle an input NaN and no rule was selected.
+ *
+ * The naming scheme for Float3NaNPropRule values is:
+ *  float_3nan_prop_s_abc:
+ *    = "Prefer SNaN over QNaN, then operand A over B over C"
+ *  float_3nan_prop_abc:
+ *    = "Prefer A over B over C regardless of SNaN vs QNAN"
+ *
+ * For QEMU, the multiply-add operation is A * B + C.
+ */
+
+/*
+ * We set the Float3NaNPropRule enum values up so we can select the
+ * right value in pickNaNMulAdd in a data driven way.
+ */
+FIELD(3NAN, 1ST, 0, 2)   /* which operand is most preferred ? */
+FIELD(3NAN, 2ND, 2, 2)   /* which operand is next most preferred ? */
+FIELD(3NAN, 3RD, 4, 2)   /* which operand is least preferred ? */
+FIELD(3NAN, SNAN, 6, 1)  /* do we prefer SNaN over QNaN ? */
+
+#define PROPRULE(X, Y, Z) \
+    ((X << R_3NAN_1ST_SHIFT) | (Y << R_3NAN_2ND_SHIFT) | (Z << R_3NAN_3RD_SHIFT))
+
+typedef enum __attribute__((__packed__)) {
+    float_3nan_prop_none = 0,     /* No propagation rule specified */
+    float_3nan_prop_abc = PROPRULE(0, 1, 2),
+    float_3nan_prop_acb = PROPRULE(0, 2, 1),
+    float_3nan_prop_bac = PROPRULE(1, 0, 2),
+    float_3nan_prop_bca = PROPRULE(1, 2, 0),
+    float_3nan_prop_cab = PROPRULE(2, 0, 1),
+    float_3nan_prop_cba = PROPRULE(2, 1, 0),
+    float_3nan_prop_s_abc = float_3nan_prop_abc | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_acb = float_3nan_prop_acb | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bac = float_3nan_prop_bac | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_bca = float_3nan_prop_bca | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cab = float_3nan_prop_cab | R_3NAN_SNAN_MASK,
+    float_3nan_prop_s_cba = float_3nan_prop_cba | R_3NAN_SNAN_MASK,
+} Float3NaNPropRule;
+
+#undef PROPRULE
+
 /*
  * Rule for result of fused multiply-add 0 * Inf + NaN.
  * This must be a NaN, but implementations differ on whether this
@@ -XXX,XX +XXX,XX @@ typedef struct float_status {
     FloatRoundMode float_rounding_mode;
     FloatX80RoundPrec floatx80_rounding_precision;
     Float2NaNPropRule float_2nan_prop_rule;
+    Float3NaNPropRule float_3nan_prop_rule;
     FloatInfZeroNaNRule float_infzeronan_rule;
     bool tininess_before_rounding;
     /* should denormalised results go to zero and set the inexact flag? */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
 static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
                          bool infzero, bool have_snan, float_status *status)
 {
+    FloatClass cls[3] = { a_cls, b_cls, c_cls };
+    Float3NaNPropRule rule = status->float_3nan_prop_rule;
+    int which;
+
     /*
      * We guarantee not to require the target to tell us how to
      * pick a NaN if we're always returning the default NaN.
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
+    if (rule == float_3nan_prop_none) {
 #if defined(TARGET_ARM)
-
-    /* This looks different from the ARM ARM pseudocode, because the ARM ARM
-     * puts the operands to a fused mac operation (a*b)+c in the order c,a,b.
-     */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        /*
+         * This looks different from the ARM ARM pseudocode, because the ARM ARM
+         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
+         */
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_MIPS)
-    if (snan_bit_is_one(status)) {
-        /* Prefer sNaN over qNaN, in the a, b, c order. */
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
+        if (snan_bit_is_one(status)) {
+            rule = float_3nan_prop_s_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_s_cab;
         }
-    } else {
-        /* Prefer sNaN over qNaN, in the c, a, b order. */
-        if (is_snan(c_cls)) {
-            return 2;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(c_cls)) {
-            return 2;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-    }
 #elif defined(TARGET_LOONGARCH64)
-    /* Prefer sNaN over qNaN, in the c, a, b order. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else {
-        return 1;
-    }
+        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
-    /* If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(c_cls)) {
-        return 2;
-    } else {
-        return 1;
-    }
+        /*
+         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
+         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+         */
+        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
-    if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_qnan(a_cls)) {
-        return 0;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-    /* Prefer SNaN over QNaN, order C, B, A. */
-    if (is_snan(c_cls)) {
-        return 2;
-    } else if (is_snan(b_cls)) {
-        return 1;
-    } else if (is_snan(a_cls)) {
-        return 0;
-    } else if (is_qnan(c_cls)) {
-        return 2;
-    } else if (is_qnan(b_cls)) {
-        return 1;
-    } else {
-        return 0;
-    }
+        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-    /*
-     * For Xtensa, the (inf,zero,nan) case sets InvalidOp and returns
-     * an input NaN if we have one (ie c).
-     */
-    if (status->use_first_nan) {
-        if (is_nan(a_cls)) {
-            return 0;
-        } else if (is_nan(b_cls)) {
-            return 1;
+        if (status->use_first_nan) {
+            rule = float_3nan_prop_abc;
         } else {
-            return 2;
+            rule = float_3nan_prop_cba;
         }
-    } else {
-        if (is_nan(c_cls)) {
-            return 2;
-        } else if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-    }
 #else
-    /* A default implementation: prefer a to b to c.
-     * This is unlikely to actually match any real implementation.
-     */
-    if (is_nan(a_cls)) {
-        return 0;
-    } else if (is_nan(b_cls)) {
-        return 1;
-    } else {
-        return 2;
-    }
+        rule = float_3nan_prop_abc;
 #endif
+    }
+
+    assert(rule != float_3nan_prop_none);
+    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+        /* We have at least one SNaN input and should prefer it */
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_snan(cls[which]));
+    } else {
+        do {
+            which = rule & R_3NAN_1ST_MASK;
+            rule >>= R_3NAN_1ST_LENGTH;
+        } while (!is_nan(cls[which]));
+    }
+    return which;
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

Explicitly set a rule in the softfloat tests for propagating NaNs in
the muladd case.  In meson.build we put -DTARGET_ARM in fpcflags, and
so we should select here the Arm rule of float_3nan_prop_s_cab.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-17-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c | 1 +
 tests/fp/fp-test.c  | 1 +
 2 files changed, 2 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
 
     f = bench_funcs[operation][precision];
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      * doesn't specify match those used by the Arm architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-18-peter.maydell@linaro.org
---
 target/arm/cpu.c               | 5 +++++
 fpu/softfloat-specialize.c.inc | 8 +-------
 2 files changed, 6 insertions(+), 7 deletions(-)

diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *  * tininess-before-rounding
  *  * 2-input NaN propagation prefers SNaN over QNaN, and then
  *    operand A over operand B (see FPProcessNaNs() pseudocode)
+ *  * 3-input NaN propagation prefers SNaN over QNaN, and then
+ *    operand C over A over B (see FPProcessNaNs3() pseudocode,
+ *    but note that for QEMU muladd is a * b + c, whereas for
+ *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
  */
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
 {
     set_float_detect_tininess(float_tininess_before_rounding, s);
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
 }
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_ARM)
-        /*
-         * This looks different from the ARM ARM pseudocode, because the ARM ARM
-         * puts the operands to a fused mac operation (a*b)+c in the order c,a,b
-         */
-        rule = float_3nan_prop_s_cab;
-#elif defined(TARGET_MIPS)
+#if defined(TARGET_MIPS)
         if (snan_bit_is_one(status)) {
             rule = float_3nan_prop_s_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for loongarch, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-19-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 1 +
 fpu/softfloat-specialize.c.inc    | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void restore_fp_status(CPULoongArchState *env)
      * case sets InvalidOp and returns the input value 'c'
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab, &env->fp_status);
 }
 
 int ieee_ex_to_loongarch(int xcpt)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_LOONGARCH64)
-        rule = float_3nan_prop_s_cab;
 #elif defined(TARGET_PPC)
         /*
          * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for PPC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-20-peter.maydell@linaro.org
---
 target/ppc/cpu_init.c          | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 6 ------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/target/ppc/cpu_init.c b/target/ppc/cpu_init.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/cpu_init.c
+++ b/target/ppc/cpu_init.c
@@ -XXX,XX +XXX,XX @@ static void ppc_cpu_reset_hold(Object *obj, ResetType type)
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->vec_status);
+    /*
+     * NaN propagation for fused multiply-add:
+     * if fRA is a NaN return it; otherwise if fRB is a NaN return it;
+     * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
+     * whereas QEMU labels the operands as (a * b) + c.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->fp_status);
+    set_float_3nan_prop_rule(float_3nan_prop_acb, &env->vec_status);
     /*
      * For PPC, the (inf,zero,qnan) case sets InvalidOp, but we prefer
      * to return an input NaN if we have one (ie c) rather than generating
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_PPC)
-        /*
-         * If fRA is a NaN return it; otherwise if fRB is a NaN return it;
-         * otherwise return fRC. Note that muladd on PPC is (fRA * fRC) + frB
-         */
-        rule = float_3nan_prop_acb;
 #elif defined(TARGET_S390X)
         rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for s390x, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-21-peter.maydell@linaro.org
---
 target/s390x/cpu.c             | 1 +
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/s390x/cpu.c b/target/s390x/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/cpu.c
+++ b/target/s390x/cpu.c
@@ -XXX,XX +XXX,XX @@ static void s390_cpu_reset_hold(Object *obj, ResetType type)
         set_float_detect_tininess(float_tininess_before_rounding,
                                   &env->fpu_status);
         set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fpu_status);
+        set_float_3nan_prop_rule(float_3nan_prop_s_abc, &env->fpu_status);
         set_float_infzeronan_rule(float_infzeronan_dnan_always,
                                   &env->fpu_status);
        /* fall through */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_S390X)
-        rule = float_3nan_prop_s_abc;
 #elif defined(TARGET_SPARC)
         rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
-- 
2.34.1

Set the Float3NaNPropRule explicitly for SPARC, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-22-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 2 --
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
      * the CPU state struct so it won't get zeroed on reset.
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &env->fp_status);
+    /* For fused-multiply add, prefer SNaN over QNaN, then C->B->A */
+    set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         } else {
             rule = float_3nan_prop_s_cab;
         }
-#elif defined(TARGET_SPARC)
-        rule = float_3nan_prop_s_cba;
 #elif defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
-- 
2.34.1

Set the Float3NaNPropRule explicitly for Arm, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-23-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h       | 4 ++++
 target/mips/msa.c              | 3 +++
 fpu/softfloat-specialize.c.inc | 8 +-------
 3 files changed, 8 insertions(+), 7 deletions(-)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
 {
     bool nan2008 = env->active_fpu.fcr31 & (1 << FCR31_NAN2008);
     FloatInfZeroNaNRule izn_rule;
+    Float3NaNPropRule nan3_rule;
 
     /*
      * With nan2008, SNaNs are silenced in the usual way.
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
      */
     izn_rule = nan2008 ? float_infzeronan_dnan_never : float_infzeronan_dnan_always;
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
+    nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
+    set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+
 }
 
 static inline void restore_fp_status(CPUMIPSState *env)
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab,
                              &env->active_tc.msa_fp_status);
 
+    set_float_3nan_prop_rule(float_3nan_prop_s_cab,
+                             &env->active_tc.msa_fp_status);
+
     /* clear float_status exception flags */
     set_float_exception_flags(0, &env->active_tc.msa_fp_status);
 
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_MIPS)
-        if (snan_bit_is_one(status)) {
-            rule = float_3nan_prop_s_abc;
-        } else {
-            rule = float_3nan_prop_s_cab;
-        }
-#elif defined(TARGET_XTENSA)
+#if defined(TARGET_XTENSA)
         if (status->use_first_nan) {
             rule = float_3nan_prop_abc;
         } else {
-- 
2.34.1

Set the Float3NaNPropRule explicitly for xtensa, and remove the
ifdef from pickNaNMulAdd().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-24-peter.maydell@linaro.org
---
 target/xtensa/fpu_helper.c     | 2 ++
 fpu/softfloat-specialize.c.inc | 8 --------
 2 files changed, 2 insertions(+), 8 deletions(-)

diff --git a/target/xtensa/fpu_helper.c b/target/xtensa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/fpu_helper.c
+++ b/target/xtensa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void xtensa_use_first_nan(CPUXtensaState *env, bool use_first)
     set_use_first_nan(use_first, &env->fp_status);
     set_float_2nan_prop_rule(use_first ? float_2nan_prop_ab : float_2nan_prop_ba,
                              &env->fp_status);
+    set_float_3nan_prop_rule(use_first ? float_3nan_prop_abc : float_3nan_prop_cba,
+                             &env->fp_status);
 }
 
 void HELPER(wur_fpu2k_fcr)(CPUXtensaState *env, uint32_t v)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
     }
 
     if (rule == float_3nan_prop_none) {
-#if defined(TARGET_XTENSA)
-        if (status->use_first_nan) {
-            rule = float_3nan_prop_abc;
-        } else {
-            rule = float_3nan_prop_cba;
-        }
-#else
         rule = float_3nan_prop_abc;
-#endif
     }
 
     assert(rule != float_3nan_prop_none);
-- 
2.34.1

Set the Float3NaNPropRule explicitly for i386.  We had no
i386-specific behaviour in the old ifdef ladder, so we were using the
default "prefer a then b then c" fallback; this is actually the
correct per-the-spec handling for i386.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-25-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c | 1 +
 1 file changed, 1 insertion(+)

Set the Float3NaNPropRule explicitly for HPPA, and remove the
ifdef from pickNaNMulAdd().

HPPA is the only target that was using the default branch of the
ifdef ladder (other targets either do not use muladd or set
default_nan_mode), so we can remove the ifdef fallback entirely now
(allowing the "rule not set" case to fall into the default of the
switch statement and assert).

We add a TODO note that the HPPA rule is probably wrong; this is
not a behavioural change for this refactoring.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-26-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 8 ++++++++
 fpu/softfloat-specialize.c.inc | 4 ----
 2 files changed, 8 insertions(+), 4 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
      * HPPA does note implement a CPU reset method at all...
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &env->fp_status);
+    /*
+     * TODO: The HPPA architecture reference only documents its NaN
+     * propagation rule for 2-operand operations. Testing on real hardware
+     * might be necessary to confirm whether this order for muladd is correct.
+     * Not preferring the SNaN is almost certainly incorrect as it diverges
+     * from the documented rules for 2-operand operations.
+     */
+    set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
         }
     }
 
-    if (rule == float_3nan_prop_none) {
-        rule = float_3nan_prop_abc;
-    }
-
     assert(rule != float_3nan_prop_none);
     if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
         /* We have at least one SNaN input and should prefer it */
-- 
2.34.1

The use_first_nan field in float_status was an xtensa-specific way to
select at runtime from two different NaN propagation rules.  Now that
xtensa is using the target-agnostic NaN propagation rule selection
that we've just added, we can remove use_first_nan, because there is
no longer any code that reads it.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-27-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 5 -----
 include/fpu/softfloat-types.h   | 1 -
 target/xtensa/fpu_helper.c      | 1 -
 3 files changed, 7 deletions(-)

Currently m68k_cpu_reset_hold() calls floatx80_default_nan(NULL)
to get the NaN bit pattern to reset the FPU registers. This
works because it happens that our implementation of
floatx80_default_nan() doesn't actually look at the float_status
pointer except for TARGET_MIPS. However, this isn't guaranteed,
and to be able to remove the ifdef in floatx80_default_nan()
we're going to need a real float_status here.

Rearrange m68k_cpu_reset_hold() so that we initialize env->fp_status
earlier, and thus can pass it to floatx80_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-28-peter.maydell@linaro.org
---
 target/m68k/cpu.c | 12 +++++++-----
 1 file changed, 7 insertions(+), 5 deletions(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
     CPUState *cs = CPU(obj);
     M68kCPUClass *mcc = M68K_CPU_GET_CLASS(obj);
     CPUM68KState *env = cpu_env(cs);
-    floatx80 nan = floatx80_default_nan(NULL);
+    floatx80 nan;
     int i;
 
     if (mcc->parent_phases.hold) {
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
 #else
     cpu_m68k_set_sr(env, SR_S | SR_I);
 #endif
-    for (i = 0; i < 8; i++) {
-        env->fregs[i].d = nan;
-    }
-    cpu_m68k_set_fpcr(env, 0);
     /*
      * M68000 FAMILY PROGRAMMER'S REFERENCE MANUAL
      * 3.4 FLOATING-POINT INSTRUCTION DETAILS
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+
+    nan = floatx80_default_nan(&env->fp_status);
+    for (i = 0; i < 8; i++) {
+        env->fregs[i].d = nan;
+    }
+    cpu_m68k_set_fpcr(env, 0);
     env->fpsr = 0;
 
     /* TODO: We should set PC from the interrupt vector.  */
-- 
2.34.1

We create our 128-bit default NaN by calling parts64_default_nan()
and then adjusting the result.  We can do the same trick for creating
the floatx80 default NaN, which lets us drop a target ifdef.

floatx80 is used only by:
 i386
 m68k
 arm nwfpe old floating-point emulation emulation support
    (which is essentially dead, especially the parts involving floatx80)
 PPC (only in the xsrqpxp instruction, which just rounds an input
    value by converting to floatx80 and back, so will never generate
    the default NaN)

The floatx80 default NaN as currently implemented is:
 m68k: sign = 0, exp = 1...1, int = 1, frac = 1....1
 i386: sign = 1, exp = 1...1, int = 1, frac = 10...0

These are the same as the parts64_default_nan for these architectures.

This is technically a possible behaviour change for arm linux-user
nwfpe emulation emulation, because the default NaN will now have the
sign bit clear.  But we were already generating a different floatx80
default NaN from the real kernel emulation we are supposedly
following, which appears to use an all-bits-1 value:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L267

This won't affect the only "real" use of the nwfpe emulation, which
is ancient binaries that used it as part of the old floating point
calling convention; that only uses loads and stores of 32 and 64 bit
floats, not any of the floatx80 behaviour the original hardware had.
We also get the nwfpe float64 default NaN value wrong:
 https://elixir.bootlin.com/linux/v6.12/source/arch/arm/nwfpe/softfloat-specialize#L166
so if we ever cared about this obscure corner the right fix would be
to correct that so nwfpe used its own default-NaN setting rather
than the Arm VFP one.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-29-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts128_silence_nan(FloatParts128 *p, float_status *status)
 floatx80 floatx80_default_nan(float_status *status)
 {
     floatx80 r;
+    /*
+     * Extrapolate from the choices made by parts64_default_nan to fill
+     * in the floatx80 format. We assume that floatx80's explicit
+     * integer bit is always set (this is true for i386 and m68k,
+     * which are the only real users of this format).
+     */
+    FloatParts64 p64;
+    parts64_default_nan(&p64, status);
 
-    /* None of the targets that have snan_bit_is_one use floatx80.  */
-    assert(!snan_bit_is_one(status));
-#if defined(TARGET_M68K)
-    r.low = UINT64_C(0xFFFFFFFFFFFFFFFF);
-    r.high = 0x7FFF;
-#else
-    /* X86 */
-    r.low = UINT64_C(0xC000000000000000);
-    r.high = 0xFFFF;
-#endif
+    r.high = 0x7FFF | (p64.sign << 15);
+    r.low = (1ULL << DECOMPOSED_BINARY_POINT) | p64.frac;
     return r;
 }
 
-- 
2.34.1

In target/loongarch's helper_fclass_s() and helper_fclass_d() we pass
a zero-initialized float_status struct to float32_is_quiet_nan() and
float64_is_quiet_nan(), with the cryptic comment "for
snan_bit_is_one".

This pattern appears to have been copied from target/riscv, where it
is used because the functions there do not have ready access to the
CPU state struct. The comment presumably refers to the fact that the
main reason the is_quiet_nan() functions want the float_state is
because they want to know about the snan_bit_is_one config.

In the loongarch helpers, though, we have the CPU state struct
to hand. Use the usual env->fp_status here. This avoids our needing
to track that we need to update the initializer of the local
float_status structs when the core softfloat code adds new
options for targets to configure their behaviour.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-30-peter.maydell@linaro.org
---
 target/loongarch/tcg/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/loongarch/tcg/fpu_helper.c b/target/loongarch/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/loongarch/tcg/fpu_helper.c
+++ b/target/loongarch/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_s(CPULoongArchState *env, uint64_t fj)
     } else if (float32_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float32_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float32_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float32_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
@@ -XXX,XX +XXX,XX @@ uint64_t helper_fclass_d(CPULoongArchState *env, uint64_t fj)
     } else if (float64_is_zero_or_denormal(f)) {
         return sign ? 1 << 4 : 1 << 8;
     } else if (float64_is_any_nan(f)) {
-        float_status s = { }; /* for snan_bit_is_one */
-        return float64_is_quiet_nan(f, &s) ? 1 << 1 : 1 << 0;
+        return float64_is_quiet_nan(f, &env->fp_status) ? 1 << 1 : 1 << 0;
     } else {
         return sign ? 1 << 3 : 1 << 7;
     }
-- 
2.34.1

In the frem helper, we have a local float_status because we want to
execute the floatx80_div() with a custom rounding mode.  Instead of
zero-initializing the local float_status and then having to set it up
with the m68k standard behaviour (including the NaN propagation rule
and copying the rounding precision from env->fp_status), initialize
it as a complete copy of env->fp_status. This will avoid our having
to add new code in this function for every new config knob we add
to fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-31-peter.maydell@linaro.org
---
 target/m68k/fpu_helper.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/target/m68k/fpu_helper.c b/target/m68k/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/fpu_helper.c
+++ b/target/m68k/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(frem)(CPUM68KState *env, FPReg *res, FPReg *val0, FPReg *val1)
 
     fp_rem = floatx80_rem(val1->d, val0->d, &env->fp_status);
     if (!floatx80_is_any_nan(fp_rem)) {
-        float_status fp_status = { };
+        /* Use local temporary fp_status to set different rounding mode */
+        float_status fp_status = env->fp_status;
         uint32_t quotient;
         int sign;
 
         /* Calculate quotient directly using round to nearest mode */
-        set_float_2nan_prop_rule(float_2nan_prop_ab, &fp_status);
         set_float_rounding_mode(float_round_nearest_even, &fp_status);
-        set_floatx80_rounding_precision(
-            get_floatx80_rounding_precision(&env->fp_status), &fp_status);
         fp_quot.d = floatx80_div(val1->d, val0->d, &fp_status);
 
         sign = extractFloatx80Sign(fp_quot.d);
-- 
2.34.1

In cf_fpu_gdb_get_reg() and cf_fpu_gdb_set_reg() we do the conversion
from float64 to floatx80 using a scratch float_status, because we
don't want the conversion to affect the CPU's floating point exception
status. Currently we use a zero-initialized float_status. This will
get steadily more awkward as we add config knobs to float_status
that the target must initialize. Avoid having to add any of that
configuration here by instead initializing our local float_status
from the env->fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-32-peter.maydell@linaro.org
---
 target/m68k/helper.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/target/m68k/helper.c b/target/m68k/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/helper.c
+++ b/target/m68k/helper.c
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_get_reg(CPUState *cs, GByteArray *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         return gdb_get_reg64(mem_buf, floatx80_to_float64(env->fregs[n].d, &s));
     }
     switch (n) {
@@ -XXX,XX +XXX,XX @@ static int cf_fpu_gdb_set_reg(CPUState *cs, uint8_t *mem_buf, int n)
     CPUM68KState *env = &cpu->env;
 
     if (n < 8) {
-        float_status s = {};
+        /* Use scratch float_status so any exceptions don't change CPU state */
+        float_status s = env->fp_status;
         env->fregs[n].d = float64_to_floatx80(ldq_be_p(mem_buf), &s);
         return 8;
     }
-- 
2.34.1

In the helper functions flcmps and flcmpd we use a scratch float_status
so that we don't change the CPU state if the comparison raises any
floating point exception flags. Instead of zero-initializing this
scratch float_status, initialize it as a copy of env->fp_status. This
avoids the need to explicitly initialize settings like the NaN
propagation rule or others we might add to softfloat in future.

To do this we need to pass the CPU env pointer in to the helper.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-33-peter.maydell@linaro.org
---
 target/sparc/helper.h     | 4 ++--
 target/sparc/fop_helper.c | 8 ++++----
 target/sparc/translate.c  | 4 ++--
 3 files changed, 8 insertions(+), 8 deletions(-)

diff --git a/target/sparc/helper.h b/target/sparc/helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/helper.h
+++ b/target/sparc/helper.h
@@ -XXX,XX +XXX,XX @@ DEF_HELPER_FLAGS_3(fcmpd, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmped, TCG_CALL_NO_WG, i32, env, f64, f64)
 DEF_HELPER_FLAGS_3(fcmpq, TCG_CALL_NO_WG, i32, env, i128, i128)
 DEF_HELPER_FLAGS_3(fcmpeq, TCG_CALL_NO_WG, i32, env, i128, i128)
-DEF_HELPER_FLAGS_2(flcmps, TCG_CALL_NO_RWG_SE, i32, f32, f32)
-DEF_HELPER_FLAGS_2(flcmpd, TCG_CALL_NO_RWG_SE, i32, f64, f64)
+DEF_HELPER_FLAGS_3(flcmps, TCG_CALL_NO_RWG_SE, i32, env, f32, f32)
+DEF_HELPER_FLAGS_3(flcmpd, TCG_CALL_NO_RWG_SE, i32, env, f64, f64)
 DEF_HELPER_2(raise_exception, noreturn, env, int)
 
 DEF_HELPER_FLAGS_3(faddd, TCG_CALL_NO_WG, f64, env, f64, f64)
diff --git a/target/sparc/fop_helper.c b/target/sparc/fop_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/fop_helper.c
+++ b/target/sparc/fop_helper.c
@@ -XXX,XX +XXX,XX @@ uint32_t helper_fcmpeq(CPUSPARCState *env, Int128 src1, Int128 src2)
     return finish_fcmp(env, r, GETPC());
 }
 
-uint32_t helper_flcmps(float32 src1, float32 src2)
+uint32_t helper_flcmps(CPUSPARCState *env, float32 src1, float32 src2)
 {
     /*
      * FLCMP never raises an exception nor modifies any FSR fields.
      * Perform the comparison with a dummy fp environment.
      */
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
@@ -XXX,XX +XXX,XX @@ uint32_t helper_flcmps(float32 src1, float32 src2)
     g_assert_not_reached();
 }
 
-uint32_t helper_flcmpd(float64 src1, float64 src2)
+uint32_t helper_flcmpd(CPUSPARCState *env, float64 src1, float64 src2)
 {
-    float_status discard = { };
+    float_status discard = env->fp_status;
     FloatRelation r;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ba, &discard);
diff --git a/target/sparc/translate.c b/target/sparc/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/translate.c
+++ b/target/sparc/translate.c
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPs(DisasContext *dc, arg_FLCMPs *a)
 
     src1 = gen_load_fpr_F(dc, a->rs1);
     src2 = gen_load_fpr_F(dc, a->rs2);
-    gen_helper_flcmps(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmps(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
@@ -XXX,XX +XXX,XX @@ static bool trans_FLCMPd(DisasContext *dc, arg_FLCMPd *a)
 
     src1 = gen_load_fpr_D(dc, a->rs1);
     src2 = gen_load_fpr_D(dc, a->rs2);
-    gen_helper_flcmpd(cpu_fcc[a->cc], src1, src2);
+    gen_helper_flcmpd(cpu_fcc[a->cc], tcg_env, src1, src2);
     return advance_pc(dc);
 }
 
-- 
2.34.1

In the helper_compute_fprf functions, we pass a dummy float_status
in to the is_signaling_nan() function. This is unnecessary, because
we have convenient access to the CPU env pointer here and that
is already set up with the correct values for the snan_bit_is_one
and no_signaling_nans config settings. is_signaling_nan() doesn't
ever update the fp_status with any exception flags, so there is
no reason not to use env->fp_status here.

Use env->fp_status instead of the dummy fp_status.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-34-peter.maydell@linaro.org
---
 target/ppc/fpu_helper.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/ppc/fpu_helper.c b/target/ppc/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/fpu_helper.c
+++ b/target/ppc/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void helper_compute_fprf_##tp(CPUPPCState *env, tp arg)           \
     } else if (tp##_is_infinity(arg)) {                           \
         fprf = neg ? 0x09 << FPSCR_FPRF : 0x05 << FPSCR_FPRF;     \
     } else {                                                      \
-        float_status dummy = { };  /* snan_bit_is_one = 0 */      \
-        if (tp##_is_signaling_nan(arg, &dummy)) {                 \
+        if (tp##_is_signaling_nan(arg, &env->fp_status)) {        \
             fprf = 0x00 << FPSCR_FPRF;                            \
         } else {                                                  \
             fprf = 0x11 << FPSCR_FPRF;                            \
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Now that float_status has a bunch of fp parameters,
it is easier to copy an existing structure than create
one from scratch.  Begin by copying the structure that
corresponds to the FPSR and make only the adjustments
required for BFloat16 semantics.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-2-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/tcg/vec_helper.c | 20 +++++++-------------
 1 file changed, 7 insertions(+), 13 deletions(-)

diff --git a/target/arm/tcg/vec_helper.c b/target/arm/tcg/vec_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/tcg/vec_helper.c
+++ b/target/arm/tcg/vec_helper.c
@@ -XXX,XX +XXX,XX @@ bool is_ebf(CPUARMState *env, float_status *statusp, float_status *oddstatusp)
      * no effect on AArch32 instructions.
      */
     bool ebf = is_a64(env) && env->vfp.fpcr & FPCR_EBF;
-    *statusp = (float_status){
-        .tininess_before_rounding = float_tininess_before_rounding,
-        .float_rounding_mode = float_round_to_odd_inf,
-        .flush_to_zero = true,
-        .flush_inputs_to_zero = true,
-        .default_nan_mode = true,
-    };
+
+    *statusp = env->vfp.fp_status;
+    set_default_nan_mode(true, statusp);
 
     if (ebf) {
-        float_status *fpst = &env->vfp.fp_status;
-        set_flush_to_zero(get_flush_to_zero(fpst), statusp);
-        set_flush_inputs_to_zero(get_flush_inputs_to_zero(fpst), statusp);
-        set_float_rounding_mode(get_float_rounding_mode(fpst), statusp);
-
         /* EBF=1 needs to do a step with round-to-odd semantics */
         *oddstatusp = *statusp;
         set_float_rounding_mode(float_round_to_odd, oddstatusp);
+    } else {
+        set_flush_to_zero(true, statusp);
+        set_flush_inputs_to_zero(true, statusp);
+        set_float_rounding_mode(float_round_to_odd_inf, statusp);
     }
-
     return ebf;
 }
 
-- 
2.34.1

Currently we hardcode the default NaN value in parts64_default_nan()
using a compile-time ifdef ladder. This is awkward for two cases:
 * for single-QEMU-binary we can't hard-code target-specifics like this
 * for Arm FEAT_AFP the default NaN value depends on FPCR.AH
   (specifically the sign bit is different)

Add a field to float_status to specify the default NaN value; fall
back to the old ifdef behaviour if these are not set.

The default NaN value is specified by setting a uint8_t to a
pattern corresponding to the sign and upper fraction parts of
the NaN; the lower bits of the fraction are set from bit 0 of
the pattern.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-35-peter.maydell@linaro.org
---
 include/fpu/softfloat-helpers.h | 11 +++++++
 include/fpu/softfloat-types.h   | 10 ++++++
 fpu/softfloat-specialize.c.inc  | 55 ++++++++++++++++++++-------------
 3 files changed, 54 insertions(+), 22 deletions(-)

Set the default NaN pattern explicitly for the tests/fp code.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-36-peter.maydell@linaro.org
---
 tests/fp/fp-bench.c     | 1 +
 tests/fp/fp-test-log2.c | 1 +
 tests/fp/fp-test.c      | 1 +
 3 files changed, 3 insertions(+)

diff --git a/tests/fp/fp-bench.c b/tests/fp/fp-bench.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-bench.c
+++ b/tests/fp/fp-bench.c
@@ -XXX,XX +XXX,XX @@ static void run_bench(void)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &soft_status);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &soft_status);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &soft_status);
+    set_float_default_nan_pattern(0b01000000, &soft_status);
 
     f = bench_funcs[operation][precision];
     g_assert(f);
diff --git a/tests/fp/fp-test-log2.c b/tests/fp/fp-test-log2.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test-log2.c
+++ b/tests/fp/fp-test-log2.c
@@ -XXX,XX +XXX,XX @@ int main(int ac, char **av)
     int i;
 
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_rounding_mode(float_round_nearest_even, &qsf);
 
     test.d = 0.0;
diff --git a/tests/fp/fp-test.c b/tests/fp/fp-test.c
index XXXXXXX..XXXXXXX 100644
--- a/tests/fp/fp-test.c
+++ b/tests/fp/fp-test.c
@@ -XXX,XX +XXX,XX @@ void run_test(void)
      */
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, &qsf);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, &qsf);
+    set_float_default_nan_pattern(0b01000000, &qsf);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, &qsf);
 
     genCases_setLevel(test_level);
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-37-peter.maydell@linaro.org
---
 target/microblaze/cpu.c        | 2 ++
 fpu/softfloat-specialize.c.inc | 3 +--
 2 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/target/microblaze/cpu.c b/target/microblaze/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/microblaze/cpu.c
+++ b/target/microblaze/cpu.c
@@ -XXX,XX +XXX,XX @@ static void mb_cpu_reset_hold(Object *obj, ResetType type)
      * this architecture.
      */
     set_float_2nan_prop_rule(float_2nan_prop_x87, &env->fp_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
 
 #if defined(CONFIG_USER_ONLY)
     /* start in user mode with interrupts enabled.  */
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)    \
-    || defined(TARGET_MICROBLAZE)
+#elif defined(TARGET_I386) || defined(TARGET_X86_64)
         /* Sign bit set, most significant frac bit set */
         dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-38-peter.maydell@linaro.org
---
 target/i386/tcg/fpu_helper.c   | 4 ++++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/target/i386/tcg/fpu_helper.c b/target/i386/tcg/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/i386/tcg/fpu_helper.c
+++ b/target/i386/tcg/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void cpu_init_fp_statuses(CPUX86State *env)
      */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->sse_status);
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->sse_status);
+    /* Default NaN: sign bit set, most significant frac bit set */
+    set_float_default_nan_pattern(0b11000000, &env->fp_status);
+    set_float_default_nan_pattern(0b11000000, &env->mmx_status);
+    set_float_default_nan_pattern(0b11000000, &env->sse_status);
 }
 
 static inline uint8_t save_exception_flags(CPUX86State *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_I386) || defined(TARGET_X86_64)
-        /* Sign bit set, most significant frac bit set */
-        dnan_pattern = 0b11000000;
 #elif defined(TARGET_HPPA)
         /* Sign bit clear, msb-1 frac bit set */
         dnan_pattern = 0b00100000;
-- 
2.34.1

Set the default NaN pattern explicitly, and remove the ifdef from
parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-39-peter.maydell@linaro.org
---
 target/hppa/fpu_helper.c       | 2 ++
 fpu/softfloat-specialize.c.inc | 3 ---
 2 files changed, 2 insertions(+), 3 deletions(-)

diff --git a/target/hppa/fpu_helper.c b/target/hppa/fpu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hppa/fpu_helper.c
+++ b/target/hppa/fpu_helper.c
@@ -XXX,XX +XXX,XX @@ void HELPER(loaded_fr0)(CPUHPPAState *env)
     set_float_3nan_prop_rule(float_3nan_prop_abc, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN: sign bit clear, msb-1 frac bit set */
+    set_float_default_nan_pattern(0b00100000, &env->fp_status);
 }
 
 void cpu_hppa_loaded_fr0(CPUHPPAState *env)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
 #if defined(TARGET_SPARC) || defined(TARGET_M68K)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
-#elif defined(TARGET_HPPA)
-        /* Sign bit clear, msb-1 frac bit set */
-        dnan_pattern = 0b00100000;
 #elif defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
-- 
2.34.1

Set the default NaN pattern explicitly for the arm target.
This includes setting it for the old linux-user nwfpe emulation.
For nwfpe, our default doesn't match the real kernel, but we
avoid making a behaviour change in this commit.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-41-peter.maydell@linaro.org
---
 linux-user/arm/nwfpe/fpa11.c | 5 +++++
 target/arm/cpu.c             | 2 ++
 2 files changed, 7 insertions(+)

diff --git a/linux-user/arm/nwfpe/fpa11.c b/linux-user/arm/nwfpe/fpa11.c
index XXXXXXX..XXXXXXX 100644
--- a/linux-user/arm/nwfpe/fpa11.c
+++ b/linux-user/arm/nwfpe/fpa11.c
@@ -XXX,XX +XXX,XX @@ void resetFPA11(void)
    * this late date.
    */
   set_float_2nan_prop_rule(float_2nan_prop_s_ab, &fpa11->fp_status);
+  /*
+   * Use the same default NaN value as Arm VFP. This doesn't match
+   * the Linux kernel's nwfpe emulation, which uses an all-1s value.
+   */
+  set_float_default_nan_pattern(0b01000000, &fpa11->fp_status);
 }
 
 void SetRoundingMode(const unsigned int opcode)
diff --git a/target/arm/cpu.c b/target/arm/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.c
+++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ void arm_register_el_change_hook(ARMCPU *cpu, ARMELChangeHookFn *hook,
  *    the pseudocode function the arguments are in the order c, a, b.
  *  * 0 * Inf + NaN returns the default NaN if the input NaN is quiet,
  *    and the input NaN if it is signalling
+ *  * Default NaN has sign bit clear, msb frac bit set
  */
 static void arm_set_default_fp_behaviours(float_status *s)
 {
@@ -XXX,XX +XXX,XX @@ static void arm_set_default_fp_behaviours(float_status *s)
     set_float_2nan_prop_rule(float_2nan_prop_s_ab, s);
     set_float_3nan_prop_rule(float_3nan_prop_s_cab, s);
     set_float_infzeronan_rule(float_infzeronan_dnan_if_qnan, s);
+    set_float_default_nan_pattern(0b01000000, s);
 }
 
 static void cp_reg_reset(gpointer key, gpointer value, gpointer opaque)
-- 
2.34.1

Set the default NaN pattern explicitly for m68k.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-43-peter.maydell@linaro.org
---
 target/m68k/cpu.c              | 2 ++
 fpu/softfloat-specialize.c.inc | 2 +-
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/target/m68k/cpu.c b/target/m68k/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/m68k/cpu.c
+++ b/target/m68k/cpu.c
@@ -XXX,XX +XXX,XX @@ static void m68k_cpu_reset_hold(Object *obj, ResetType type)
      * preceding paragraph for nonsignaling NaNs.
      */
     set_float_2nan_prop_rule(float_2nan_prop_ab, &env->fp_status);
+    /* Default NaN: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     nan = floatx80_default_nan(&env->fp_status);
     for (i = 0; i < 8; i++) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC) || defined(TARGET_M68K)
+#if defined(TARGET_SPARC)
         /* Sign bit clear, all frac bits set */
         dnan_pattern = 0b01111111;
 #elif defined(TARGET_HEXAGON)
-- 
2.34.1

Set the default NaN pattern explicitly for MIPS. Note that this
is our only target which currently changes the default NaN
at runtime (which it was previously doing indirectly when it
changed the snan_bit_is_one setting).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-44-peter.maydell@linaro.org
---
 target/mips/fpu_helper.h | 7 +++++++
 target/mips/msa.c        | 3 +++
 2 files changed, 10 insertions(+)

diff --git a/target/mips/fpu_helper.h b/target/mips/fpu_helper.h
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/fpu_helper.h
+++ b/target/mips/fpu_helper.h
@@ -XXX,XX +XXX,XX @@ static inline void restore_snan_bit_mode(CPUMIPSState *env)
     set_float_infzeronan_rule(izn_rule, &env->active_fpu.fp_status);
     nan3_rule = nan2008 ? float_3nan_prop_s_cab : float_3nan_prop_s_abc;
     set_float_3nan_prop_rule(nan3_rule, &env->active_fpu.fp_status);
+    /*
+     * With nan2008, the default NaN value has the sign bit clear and the
+     * frac msb set; with the older mode, the sign bit is clear, and all
+     * frac bits except the msb are set.
+     */
+    set_float_default_nan_pattern(nan2008 ? 0b01000000 : 0b00111111,
+                                  &env->active_fpu.fp_status);
 
 }
 
diff --git a/target/mips/msa.c b/target/mips/msa.c
index XXXXXXX..XXXXXXX 100644
--- a/target/mips/msa.c
+++ b/target/mips/msa.c
@@ -XXX,XX +XXX,XX @@ void msa_reset(CPUMIPSState *env)
     /* Inf * 0 + NaN returns the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never,
                               &env->active_tc.msa_fp_status);
+    /* Default NaN: sign bit clear, frac msb set */
+    set_float_default_nan_pattern(0b01000000,
+                                  &env->active_tc.msa_fp_status);
 }
-- 
2.34.1

Set the default NaN pattern explicitly for SPARC, and remove
the ifdef from parts64_default_nan.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-50-peter.maydell@linaro.org
---
 target/sparc/cpu.c             | 2 ++
 fpu/softfloat-specialize.c.inc | 5 +----
 2 files changed, 3 insertions(+), 4 deletions(-)

diff --git a/target/sparc/cpu.c b/target/sparc/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/sparc/cpu.c
+++ b/target/sparc/cpu.c
@@ -XXX,XX +XXX,XX @@ static void sparc_cpu_realizefn(DeviceState *dev, Error **errp)
     set_float_3nan_prop_rule(float_3nan_prop_s_cba, &env->fp_status);
     /* For inf * 0 + NaN, return the input NaN */
     set_float_infzeronan_rule(float_infzeronan_dnan_never, &env->fp_status);
+    /* Default NaN value: sign bit clear, all frac bits set */
+    set_float_default_nan_pattern(0b01111111, &env->fp_status);
 
     cpu_exec_realizefn(cs, &local_err);
     if (local_err != NULL) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_SPARC)
-        /* Sign bit clear, all frac bits set */
-        dnan_pattern = 0b01111111;
-#elif defined(TARGET_HEXAGON)
+#if defined(TARGET_HEXAGON)
         /* Sign bit set, all frac bits set. */
         dnan_pattern = 0b11111111;
 #else
-- 
2.34.1

Set the default NaN pattern explicitly for hexagon.
Remove the ifdef from parts64_default_nan(); the only
remaining unconverted targets all use the default case.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-52-peter.maydell@linaro.org
---
 target/hexagon/cpu.c           | 2 ++
 fpu/softfloat-specialize.c.inc | 5 -----
 2 files changed, 2 insertions(+), 5 deletions(-)

diff --git a/target/hexagon/cpu.c b/target/hexagon/cpu.c
index XXXXXXX..XXXXXXX 100644
--- a/target/hexagon/cpu.c
+++ b/target/hexagon/cpu.c
@@ -XXX,XX +XXX,XX @@ static void hexagon_cpu_reset_hold(Object *obj, ResetType type)
 
     set_default_nan_mode(1, &env->fp_status);
     set_float_detect_tininess(float_tininess_before_rounding, &env->fp_status);
+    /* Default NaN value: sign bit set, all frac bits set */
+    set_float_default_nan_pattern(0b11111111, &env->fp_status);
 }
 
 static void hexagon_cpu_disas_set_info(CPUState *s, disassemble_info *info)
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint8_t dnan_pattern = status->default_nan_pattern;
 
     if (dnan_pattern == 0) {
-#if defined(TARGET_HEXAGON)
-        /* Sign bit set, all frac bits set. */
-        dnan_pattern = 0b11111111;
-#else
         /*
          * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
          * S390, SH4, TriCore, and Xtensa.  Our other supported targets
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
             /* sign bit clear, set frac msb */
             dnan_pattern = 0b01000000;
         }
-#endif
     }
     assert(dnan_pattern != 0);
 
-- 
2.34.1

Now that all our targets have bene converted to explicitly specify
their pattern for the default NaN value we can remove the remaining
fallback code in parts64_default_nan().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241202131347.498124-55-peter.maydell@linaro.org
---
 fpu/softfloat-specialize.c.inc | 14 --------------
 1 file changed, 14 deletions(-)

diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static void parts64_default_nan(FloatParts64 *p, float_status *status)
     uint64_t frac;
     uint8_t dnan_pattern = status->default_nan_pattern;
 
-    if (dnan_pattern == 0) {
-        /*
-         * This case is true for Alpha, ARM, MIPS, OpenRISC, PPC, RISC-V,
-         * S390, SH4, TriCore, and Xtensa.  Our other supported targets
-         * do not have floating-point.
-         */
-        if (snan_bit_is_one(status)) {
-            /* sign bit clear, set all frac bits other than msb */
-            dnan_pattern = 0b00111111;
-        } else {
-            /* sign bit clear, set frac msb */
-            dnan_pattern = 0b01000000;
-        }
-    }
     assert(dnan_pattern != 0);
 
     sign = dnan_pattern >> 7;
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaNMulAdd into its only caller.  This makes
one assert redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-3-richard.henderson@linaro.org
[PMM: keep comment from old code in new location]
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 41 +++++++++++++++++++++++++-
 fpu/softfloat-specialize.c.inc | 54 ----------------------------------
 2 files changed, 40 insertions(+), 55 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
     }
 
     if (s->default_nan_mode) {
+        /*
+         * We guarantee not to require the target to tell us how to
+         * pick a NaN if we're always returning the default NaN.
+         * But if we're not in default-NaN mode then the target must
+         * specify.
+         */
         which = 3;
+    } else if (infzero) {
+        /*
+         * Inf * 0 + NaN -- some implementations return the
+         * default NaN here, and some return the input NaN.
+         */
+        switch (s->float_infzeronan_rule) {
+        case float_infzeronan_dnan_never:
+            which = 2;
+            break;
+        case float_infzeronan_dnan_always:
+            which = 3;
+            break;
+        case float_infzeronan_dnan_if_qnan:
+            which = is_qnan(c->cls) ? 3 : 2;
+            break;
+        default:
+            g_assert_not_reached();
+        }
     } else {
-        which = pickNaNMulAdd(a->cls, b->cls, c->cls, infzero, have_snan, s);
+        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        Float3NaNPropRule rule = s->float_3nan_prop_rule;
+
+        assert(rule != float_3nan_prop_none);
+        if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
+            /* We have at least one SNaN input and should prefer it */
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_snan(cls[which]));
+        } else {
+            do {
+                which = rule & R_3NAN_1ST_MASK;
+                rule >>= R_3NAN_1ST_LENGTH;
+            } while (!is_nan(cls[which]));
+        }
     }
 
     if (which == 3) {
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ static int pickNaN(FloatClass a_cls, FloatClass b_cls,
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a three-input operation.
-| For the moment we assume that no CPU needs the 'larger significand'
-| information.
-| Return values : 0 : a; 1 : b; 2 : c; 3 : default-NaN
-*----------------------------------------------------------------------------*/
-static int pickNaNMulAdd(FloatClass a_cls, FloatClass b_cls, FloatClass c_cls,
-                         bool infzero, bool have_snan, float_status *status)
-{
-    FloatClass cls[3] = { a_cls, b_cls, c_cls };
-    Float3NaNPropRule rule = status->float_3nan_prop_rule;
-    int which;
-
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify.
-     */
-    assert(!status->default_nan_mode);
-
-    if (infzero) {
-        /*
-         * Inf * 0 + NaN -- some implementations return the default NaN here,
-         * and some return the input NaN.
-         */
-        switch (status->float_infzeronan_rule) {
-        case float_infzeronan_dnan_never:
-            return 2;
-        case float_infzeronan_dnan_always:
-            return 3;
-        case float_infzeronan_dnan_if_qnan:
-            return is_qnan(c_cls) ? 3 : 2;
-        default:
-            g_assert_not_reached();
-        }
-    }
-
-    assert(rule != float_3nan_prop_none);
-    if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
-        /* We have at least one SNaN input and should prefer it */
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_snan(cls[which]));
-    } else {
-        do {
-            which = rule & R_3NAN_1ST_MASK;
-            rule >>= R_3NAN_1ST_LENGTH;
-        } while (!is_nan(cls[which]));
-    }
-    return which;
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remove "3" as a special case for which and simply
branch to return the desired value.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-4-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          * But if we're not in default-NaN mode then the target must
          * specify.
          */
-        which = 3;
+        goto default_nan;
     } else if (infzero) {
         /*
          * Inf * 0 + NaN -- some implementations return the
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
          */
         switch (s->float_infzeronan_rule) {
         case float_infzeronan_dnan_never:
-            which = 2;
             break;
         case float_infzeronan_dnan_always:
-            which = 3;
-            break;
+            goto default_nan;
         case float_infzeronan_dnan_if_qnan:
-            which = is_qnan(c->cls) ? 3 : 2;
+            if (is_qnan(c->cls)) {
+                goto default_nan;
+            }
             break;
         default:
             g_assert_not_reached();
         }
+        which = 2;
     } else {
         FloatClass cls[3] = { a->cls, b->cls, c->cls };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
     }
 
-    if (which == 3) {
-        parts_default_nan(a, s);
-        return a;
-    }
-
     switch (which) {
     case 0:
         break;
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         parts_silence_nan(a, s);
     }
     return a;
+
+ default_nan:
+    parts_default_nan(a, s);
+    return a;
 }
 
 /*
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Assign the pointer return value to 'a' directly,
rather than going through an intermediary index.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-5-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++----------------------
 1 file changed, 10 insertions(+), 22 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
                                             FloatPartsN *c, float_status *s,
                                             int ab_mask, int abc_mask)
 {
-    int which;
     bool infzero = (ab_mask == float_cmask_infzero);
     bool have_snan = (abc_mask & float_cmask_snan);
+    FloatPartsN *ret;
 
     if (unlikely(have_snan)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         default:
             g_assert_not_reached();
         }
-        which = 2;
+        ret = c;
     } else {
-        FloatClass cls[3] = { a->cls, b->cls, c->cls };
+        FloatPartsN *val[3] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
         if (have_snan && (rule & R_3NAN_SNAN_MASK)) {
             /* We have at least one SNaN input and should prefer it */
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_snan(cls[which]));
+            } while (!is_snan(ret->cls));
         } else {
             do {
-                which = rule & R_3NAN_1ST_MASK;
+                ret = val[rule & R_3NAN_1ST_MASK];
                 rule >>= R_3NAN_1ST_LENGTH;
-            } while (!is_nan(cls[which]));
+            } while (!is_nan(ret->cls));
         }
     }
 
-    switch (which) {
-    case 0:
-        break;
-    case 1:
-        a = b;
-        break;
-    case 2:
-        a = c;
-        break;
-    default:
-        g_assert_not_reached();
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 
  default_nan:
     parts_default_nan(a, s);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

While all indices into val[] should be in [0-2], the mask
applied is two bits.  To help static analysis see there is
no possibility of read beyond the end of the array, pad the
array to 4 entries, with the final being (implicitly) NULL.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-6-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
         }
         ret = c;
     } else {
-        FloatPartsN *val[3] = { a, b, c };
+        FloatPartsN *val[R_3NAN_1ST_MASK + 1] = { a, b, c };
         Float3NaNPropRule rule = s->float_3nan_prop_rule;
 
         assert(rule != float_3nan_prop_none);
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

This function is part of the public interface and
is not "specialized" to any target in any way.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-7-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c                | 52 ++++++++++++++++++++++++++++++++++
 fpu/softfloat-specialize.c.inc | 52 ----------------------------------
 2 files changed, 52 insertions(+), 52 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
     *zExpPtr = 1 - shiftCount;
 }
 
+/*----------------------------------------------------------------------------
+| Takes two extended double-precision floating-point values `a' and `b', one
+| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
+| `b' is a signaling NaN, the invalid exception is raised.
+*----------------------------------------------------------------------------*/
+
+floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
+{
+    bool aIsLargerSignificand;
+    FloatClass a_cls, b_cls;
+
+    /* This is not complete, but is good enough for pickNaN.  */
+    a_cls = (!floatx80_is_any_nan(a)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(a, status)
+             ? float_class_snan
+             : float_class_qnan);
+    b_cls = (!floatx80_is_any_nan(b)
+             ? float_class_normal
+             : floatx80_is_signaling_nan(b, status)
+             ? float_class_snan
+             : float_class_qnan);
+
+    if (is_snan(a_cls) || is_snan(b_cls)) {
+        float_raise(float_flag_invalid, status);
+    }
+
+    if (status->default_nan_mode) {
+        return floatx80_default_nan(status);
+    }
+
+    if (a.low < b.low) {
+        aIsLargerSignificand = 0;
+    } else if (b.low < a.low) {
+        aIsLargerSignificand = 1;
+    } else {
+        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
+    }
+
+    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
+        if (is_snan(b_cls)) {
+            return floatx80_silence_nan(b, status);
+        }
+        return b;
+    } else {
+        if (is_snan(a_cls)) {
+            return floatx80_silence_nan(a, status);
+        }
+        return a;
+    }
+}
+
 /*----------------------------------------------------------------------------
 | Takes an abstract floating-point value having sign `zSign', exponent `zExp',
 | and extended significand formed by the concatenation of `zSig0' and `zSig1',
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ floatx80 floatx80_silence_nan(floatx80 a, float_status *status)
     return a;
 }
 
-/*----------------------------------------------------------------------------
-| Takes two extended double-precision floating-point values `a' and `b', one
-| of which is a NaN, and returns the appropriate NaN result.  If either `a' or
-| `b' is a signaling NaN, the invalid exception is raised.
-*----------------------------------------------------------------------------*/
-
-floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
-{
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
-
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
-        return floatx80_default_nan(status);
-    }
-
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the quadruple-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Unpacking and repacking the parts may be slightly more work
than we did before, but we get to reuse more code.  For a
code path handling exceptional values, this is an improvement.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20241203203949.483774-8-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat.c | 43 +++++--------------------------------------
 1 file changed, 5 insertions(+), 38 deletions(-)

diff --git a/fpu/softfloat.c b/fpu/softfloat.c
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat.c
+++ b/fpu/softfloat.c
@@ -XXX,XX +XXX,XX @@ void normalizeFloatx80Subnormal(uint64_t aSig, int32_t *zExpPtr,
 
 floatx80 propagateFloatx80NaN(floatx80 a, floatx80 b, float_status *status)
 {
-    bool aIsLargerSignificand;
-    FloatClass a_cls, b_cls;
+    FloatParts128 pa, pb, *pr;
 
-    /* This is not complete, but is good enough for pickNaN.  */
-    a_cls = (!floatx80_is_any_nan(a)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(a, status)
-             ? float_class_snan
-             : float_class_qnan);
-    b_cls = (!floatx80_is_any_nan(b)
-             ? float_class_normal
-             : floatx80_is_signaling_nan(b, status)
-             ? float_class_snan
-             : float_class_qnan);
-
-    if (is_snan(a_cls) || is_snan(b_cls)) {
-        float_raise(float_flag_invalid, status);
-    }
-
-    if (status->default_nan_mode) {
+    if (!floatx80_unpack_canonical(&pa, a, status) ||
+        !floatx80_unpack_canonical(&pb, b, status)) {
         return floatx80_default_nan(status);
     }
 
-    if (a.low < b.low) {
-        aIsLargerSignificand = 0;
-    } else if (b.low < a.low) {
-        aIsLargerSignificand = 1;
-    } else {
-        aIsLargerSignificand = (a.high < b.high) ? 1 : 0;
-    }
-
-    if (pickNaN(a_cls, b_cls, aIsLargerSignificand, status)) {
-        if (is_snan(b_cls)) {
-            return floatx80_silence_nan(b, status);
-        }
-        return b;
-    } else {
-        if (is_snan(a_cls)) {
-            return floatx80_silence_nan(a, status);
-        }
-        return a;
-    }
+    pr = parts_pick_nan(&pa, &pb, status);
+    return floatx80_round_pack_canonical(pr, status);
 }
 
 /*----------------------------------------------------------------------------
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Inline pickNaN into its only caller.  This makes one assert
redundant with the immediately preceding IF.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-9-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc      | 82 +++++++++++++++++++++++++----
 fpu/softfloat-specialize.c.inc | 96 ----------------------------------
 2 files changed, 73 insertions(+), 105 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static void partsN(return_nan)(FloatPartsN *a, float_status *s)
 static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
+    int cmp, which;
+
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
     }
 
     if (s->default_nan_mode) {
         parts_default_nan(a, s);
-    } else {
-        int cmp = frac_cmp(a, b);
-        if (cmp == 0) {
-            cmp = a->sign < b->sign;
-        }
+        return a;
+    }
 
-        if (pickNaN(a->cls, b->cls, cmp > 0, s)) {
-            a = b;
-        }
+    cmp = frac_cmp(a, b);
+    if (cmp == 0) {
+        cmp = a->sign < b->sign;
+    }
+
+    switch (s->float_2nan_prop_rule) {
+    case float_2nan_prop_s_ab:
         if (is_snan(a->cls)) {
-            parts_silence_nan(a, s);
+            which = 0;
+        } else if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_qnan(a->cls)) {
+            which = 0;
+        } else {
+            which = 1;
         }
+        break;
+    case float_2nan_prop_s_ba:
+        if (is_snan(b->cls)) {
+            which = 1;
+        } else if (is_snan(a->cls)) {
+            which = 0;
+        } else if (is_qnan(b->cls)) {
+            which = 1;
+        } else {
+            which = 0;
+        }
+        break;
+    case float_2nan_prop_ab:
+        which = is_nan(a->cls) ? 0 : 1;
+        break;
+    case float_2nan_prop_ba:
+        which = is_nan(b->cls) ? 1 : 0;
+        break;
+    case float_2nan_prop_x87:
+        /*
+         * This implements x87 NaN propagation rules:
+         * SNaN + QNaN => return the QNaN
+         * two SNaNs => return the one with the larger significand, silenced
+         * two QNaNs => return the one with the larger significand
+         * SNaN and a non-NaN => return the SNaN, silenced
+         * QNaN and a non-NaN => return the QNaN
+         *
+         * If we get down to comparing significands and they are the same,
+         * return the NaN with the positive sign bit (if any).
+         */
+        if (is_snan(a->cls)) {
+            if (is_snan(b->cls)) {
+                which = cmp > 0 ? 0 : 1;
+            } else {
+                which = is_qnan(b->cls) ? 1 : 0;
+            }
+        } else if (is_qnan(a->cls)) {
+            if (is_snan(b->cls) || !is_qnan(b->cls)) {
+                which = 0;
+            } else {
+                which = cmp > 0 ? 0 : 1;
+            }
+        } else {
+            which = 1;
+        }
+        break;
+    default:
+        g_assert_not_reached();
+    }
+
+    if (which) {
+        a = b;
+    }
+    if (is_snan(a->cls)) {
+        parts_silence_nan(a, s);
     }
     return a;
 }
diff --git a/fpu/softfloat-specialize.c.inc b/fpu/softfloat-specialize.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-specialize.c.inc
+++ b/fpu/softfloat-specialize.c.inc
@@ -XXX,XX +XXX,XX @@ bool float32_is_signaling_nan(float32 a_, float_status *status)
     }
 }
 
-/*----------------------------------------------------------------------------
-| Select which NaN to propagate for a two-input operation.
-| IEEE754 doesn't specify all the details of this, so the
-| algorithm is target-specific.
-| The routine is passed various bits of information about the
-| two NaNs and should return 0 to select NaN a and 1 for NaN b.
-| Note that signalling NaNs are always squashed to quiet NaNs
-| by the caller, by calling floatXX_silence_nan() before
-| returning them.
-|
-| aIsLargerSignificand is only valid if both a and b are NaNs
-| of some kind, and is true if a has the larger significand,
-| or if both a and b have the same significand but a is
-| positive but b is negative. It is only needed for the x87
-| tie-break rule.
-*----------------------------------------------------------------------------*/
-
-static int pickNaN(FloatClass a_cls, FloatClass b_cls,
-                   bool aIsLargerSignificand, float_status *status)
-{
-    /*
-     * We guarantee not to require the target to tell us how to
-     * pick a NaN if we're always returning the default NaN.
-     * But if we're not in default-NaN mode then the target must
-     * specify via set_float_2nan_prop_rule().
-     */
-    assert(!status->default_nan_mode);
-
-    switch (status->float_2nan_prop_rule) {
-    case float_2nan_prop_s_ab:
-        if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_qnan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_s_ba:
-        if (is_snan(b_cls)) {
-            return 1;
-        } else if (is_snan(a_cls)) {
-            return 0;
-        } else if (is_qnan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_ab:
-        if (is_nan(a_cls)) {
-            return 0;
-        } else {
-            return 1;
-        }
-        break;
-    case float_2nan_prop_ba:
-        if (is_nan(b_cls)) {
-            return 1;
-        } else {
-            return 0;
-        }
-        break;
-    case float_2nan_prop_x87:
-        /*
-         * This implements x87 NaN propagation rules:
-         * SNaN + QNaN => return the QNaN
-         * two SNaNs => return the one with the larger significand, silenced
-         * two QNaNs => return the one with the larger significand
-         * SNaN and a non-NaN => return the SNaN, silenced
-         * QNaN and a non-NaN => return the QNaN
-         *
-         * If we get down to comparing significands and they are the same,
-         * return the NaN with the positive sign bit (if any).
-         */
-        if (is_snan(a_cls)) {
-            if (is_snan(b_cls)) {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-            return is_qnan(b_cls) ? 1 : 0;
-        } else if (is_qnan(a_cls)) {
-            if (is_snan(b_cls) || !is_qnan(b_cls)) {
-                return 0;
-            } else {
-                return aIsLargerSignificand ? 0 : 1;
-            }
-        } else {
-            return 1;
-        }
-    default:
-        g_assert_not_reached();
-    }
-}
-
 /*----------------------------------------------------------------------------
 | Returns 1 if the double-precision floating-point value `a' is a quiet
 | NaN; otherwise returns 0.
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Remember if there was an SNaN, and use that to simplify
float_2nan_prop_s_{ab,ba} to only the snan component.
Then, fall through to the corresponding
float_2nan_prop_{ab,ba} case to handle any remaining
nans, which must be quiet.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-10-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 32 ++++++++++++--------------------
 1 file changed, 12 insertions(+), 20 deletions(-)

From: Richard Henderson <richard.henderson@linaro.org>

Move the fractional comparison to the end of the
float_2nan_prop_x87 case.  This is not required for
any other 2nan propagation rule.  Reorganize the
x87 case itself to break out of the switch when the
fractional comparison is not required.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20241203203949.483774-11-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 19 +++++++++----------
 1 file changed, 9 insertions(+), 10 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
         return a;
     }
 
-    cmp = frac_cmp(a, b);
-    if (cmp == 0) {
-        cmp = a->sign < b->sign;
-    }
-
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          * return the NaN with the positive sign bit (if any).
          */
         if (is_snan(a->cls)) {
-            if (is_snan(b->cls)) {
-                which = cmp > 0 ? 0 : 1;
-            } else {
+            if (!is_snan(b->cls)) {
                 which = is_qnan(b->cls) ? 1 : 0;
+                break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
                 which = 0;
-            } else {
-                which = cmp > 0 ? 0 : 1;
+                break;
             }
         } else {
             which = 1;
+            break;
         }
+        cmp = frac_cmp(a, b);
+        if (cmp == 0) {
+            cmp = a->sign < b->sign;
+        }
+        which = cmp > 0 ? 0 : 1;
         break;
     default:
         g_assert_not_reached();
-- 
2.34.1

From: Richard Henderson <richard.henderson@linaro.org>

Replace the "index" selecting between A and B with a result variable
of the proper type.  This improves clarity within the function.

Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241203203949.483774-12-richard.henderson@linaro.org
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 fpu/softfloat-parts.c.inc | 28 +++++++++++++---------------
 1 file changed, 13 insertions(+), 15 deletions(-)

diff --git a/fpu/softfloat-parts.c.inc b/fpu/softfloat-parts.c.inc
index XXXXXXX..XXXXXXX 100644
--- a/fpu/softfloat-parts.c.inc
+++ b/fpu/softfloat-parts.c.inc
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
                                      float_status *s)
 {
     bool have_snan = false;
-    int cmp, which;
+    FloatPartsN *ret;
+    int cmp;
 
     if (is_snan(a->cls) || is_snan(b->cls)) {
         float_raise(float_flag_invalid | float_flag_invalid_snan, s);
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
     switch (s->float_2nan_prop_rule) {
     case float_2nan_prop_s_ab:
         if (have_snan) {
-            which = is_snan(a->cls) ? 0 : 1;
+            ret = is_snan(a->cls) ? a : b;
             break;
         }
         /* fall through */
     case float_2nan_prop_ab:
-        which = is_nan(a->cls) ? 0 : 1;
+        ret = is_nan(a->cls) ? a : b;
         break;
     case float_2nan_prop_s_ba:
         if (have_snan) {
-            which = is_snan(b->cls) ? 1 : 0;
+            ret = is_snan(b->cls) ? b : a;
             break;
         }
         /* fall through */
     case float_2nan_prop_ba:
-        which = is_nan(b->cls) ? 1 : 0;
+        ret = is_nan(b->cls) ? b : a;
         break;
     case float_2nan_prop_x87:
         /*
@@ -XXX,XX +XXX,XX @@ static FloatPartsN *partsN(pick_nan)(FloatPartsN *a, FloatPartsN *b,
          */
         if (is_snan(a->cls)) {
             if (!is_snan(b->cls)) {
-                which = is_qnan(b->cls) ? 1 : 0;
+                ret = is_qnan(b->cls) ? b : a;
                 break;
             }
         } else if (is_qnan(a->cls)) {
             if (is_snan(b->cls) || !is_qnan(b->cls)) {
-                which = 0;
+                ret = a;
                 break;
             }
         } else {
-            which = 1;
+            ret = b;
             break;
         }
         cmp = frac_cmp(a, b);
         if (cmp == 0) {
             cmp = a->sign < b->sign;
         }
-        which = cmp > 0 ? 0 : 1;
+        ret = cmp > 0 ? a : b;
         break;
     default:
         g_assert_not_reached();
     }
 
-    if (which) {
-        a = b;
+    if (is_snan(ret->cls)) {
+        parts_silence_nan(ret, s);
     }
-    if (is_snan(a->cls)) {
-        parts_silence_nan(a, s);
-    }
-    return a;
+    return ret;
 }
 
 static FloatPartsN *partsN(pick_nan_muladd)(FloatPartsN *a, FloatPartsN *b,
-- 
2.34.1

From: Leif Lindholm <quic_llindhol@quicinc.com>

I'm migrating to Qualcomm's new open source email infrastructure, so
update my email address, and update the mailmap to match.

Signed-off-by: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
Reviewed-by: Leif Lindholm <quic_llindhol@quicinc.com>
Reviewed-by: Brian Cain <brian.cain@oss.qualcomm.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Tested-by: Philippe Mathieu-Daudé <philmd@linaro.org>
Message-id: 20241205114047.1125842-1-leif.lindholm@oss.qualcomm.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 +-
 .mailmap    | 5 +++--
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: include/hw/ssi/imx_spi.h
 SBSA-REF
 M: Radoslaw Biernacki <rad@semihalf.com>
 M: Peter Maydell <peter.maydell@linaro.org>
-R: Leif Lindholm <quic_llindhol@quicinc.com>
+R: Leif Lindholm <leif.lindholm@oss.qualcomm.com>
 R: Marcin Juszkiewicz <marcin.juszkiewicz@linaro.org>
 L: qemu-arm@nongnu.org
 S: Maintained
diff --git a/.mailmap b/.mailmap
index XXXXXXX..XXXXXXX 100644
--- a/.mailmap
+++ b/.mailmap
@@ -XXX,XX +XXX,XX @@ Huacai Chen <chenhuacai@kernel.org> <chenhc@lemote.com>
 Huacai Chen <chenhuacai@kernel.org> <chenhuacai@loongson.cn>
 James Hogan <jhogan@kernel.org> <james.hogan@imgtec.com>
 Juan Quintela <quintela@trasno.org> <quintela@redhat.com>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif.lindholm@linaro.org>
-Leif Lindholm <quic_llindhol@quicinc.com> <leif@nuviainc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <quic_llindhol@quicinc.com>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif.lindholm@linaro.org>
+Leif Lindholm <leif.lindholm@oss.qualcomm.com> <leif@nuviainc.com>
 Luc Michel <luc@lmichel.fr> <luc.michel@git.antfield.fr>
 Luc Michel <luc@lmichel.fr> <luc.michel@greensocs.com>
 Luc Michel <luc@lmichel.fr> <lmichel@kalray.eu>
-- 
2.34.1

From: Vikram Garhwal <vikram.garhwal@bytedance.com>

Previously, maintainer role was paused due to inactive email id. Commit id:
c009d715721861984c4987bcc78b7ee183e86d75.

Signed-off-by: Vikram Garhwal <vikram.garhwal@bytedance.com>
Reviewed-by: Francisco Iglesias <francisco.iglesias@amd.com>
Message-id: 20241204184205.12952-1-vikram.garhwal@bytedance.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 MAINTAINERS | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: tests/qtest/fuzz-sb16-test.c
 
 Xilinx CAN
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 F: hw/net/can/xlnx-*
 F: include/hw/net/xlnx-*
@@ -XXX,XX +XXX,XX @@ F: include/hw/rx/
 CAN bus subsystem and hardware
 M: Pavel Pisa <pisa@cmp.felk.cvut.cz>
 M: Francisco Iglesias <francisco.iglesias@amd.com>
+M: Vikram Garhwal <vikram.garhwal@bytedance.com>
 S: Maintained
 W: https://canbus.pages.fel.cvut.cz/
 F: net/can/*
-- 
2.34.1