Series comparison

-[Qemu-devel] [PULL 00/25] target-arm queue
+[PULL 00/42] target-arm queue
-target-arm queue. This has the "plumb txattrs through various
+The following changes since commit 61fee7f45955cd0bf9b79be9fa9c7ebabb5e6a85:
 bits of exec.c" patches, and a collection of bug fixes from
 various people.
-thanks
+  Merge remote-tracking branch 'remotes/philmd-gitlab/tags/acceptance-testing-20200622' into staging (2020-06-22 20:50:10 +0100)
 -- PMM
 The following changes since commit a3ac12fba028df90f7b3dbec924995c126c41022:
   Merge remote-tracking branch 'remotes/ehabkost/tags/numa-next-pull-request' into staging (2018-05-31 11:12:36 +0100)
 are available in the Git repository at:
-  git://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20180531
+  https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20200623
-for you to fetch changes up to 49d1dca0520ea71bc21867fab6647f474fcf857b:
+for you to fetch changes up to 539533b85fbd269f777bed931de8ccae1dd837e9:
-  KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice (2018-05-31 14:52:53 +0100)
+  arm/virt: Add memory hot remove support (2020-06-23 11:39:48 +0100)
 ----------------------------------------------------------------
 target-arm queue:
- * target/arm: Honour FPCR.FZ in FRECPX
+ * util/oslib-posix : qemu_init_exec_dir implementation for Mac
- * MAINTAINERS: Add entries for newer MPS2 boards and devices
+ * target/arm: Last parts of neon decodetree conversion
- * hw/intc/arm_gicv3: Fix APxR<n> register dispatching
+ * hw/arm/virt: Add 5.0 HW compat props
- * arm_gicv3_kvm: fix bug in writing zero bits back to the in-kernel
+ * hw/watchdog/cmsdk-apb-watchdog: Add trace event for lock status
-   GIC state
+ * mps2: Add CMSDK APB watchdog, FPGAIO block, S2I devices and I2C devices
- * tcg: Fix helper function vs host abi for float16
+ * mps2: Add some unimplemented-device stubs for audio and GPIO
- * arm: fix qemu crash on startup with -bios option
+ * mps2-tz: Use the ARM SBCon two-wire serial bus interface
- * arm: fix malloc type mismatch
+ * target/arm: Check supported KVM features globally (not per vCPU)
- * xlnx-zdma: Correct mem leaks and memset to zero on desc unaligned errors
+ * tests/qtest/arm-cpu-features: Add feature setting tests
- * Correct CPACR reset value for v7 cores
+ * arm/virt: Add memory hot remove support
  * memory.h: Improve IOMMU related documentation
  * exec: Plumb transaction attributes through various functions in
    preparation for allowing IOMMUs to see them
  * vmstate.h: Provide VMSTATE_BOOL_SUB_ARRAY
  * ARM: ACPI: Fix use-after-free due to memory realloc
  * KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice
 ----------------------------------------------------------------
-Francisco Iglesias (1):
+Andrew Jones (2):
-      xlnx-zdma: Correct mem leaks and memset to zero on desc unaligned errors
+      hw/arm/virt: Add 5.0 HW compat props
       tests/qtest/arm-cpu-features: Add feature setting tests
-Igor Mammedov (1):
+David CARLIER (1):
-      arm: fix qemu crash on startup with -bios option
+      util/oslib-posix : qemu_init_exec_dir implementation for Mac
-Jan Kiszka (1):
+Peter Maydell (23):
-      hw/intc/arm_gicv3: Fix APxR<n> register dispatching
+      target/arm: Convert Neon 2-reg-misc VREV64 to decodetree
       target/arm: Convert Neon 2-reg-misc pairwise ops to decodetree
       target/arm: Convert VZIP, VUZP to decodetree
       target/arm: Convert Neon narrowing moves to decodetree
       target/arm: Convert Neon 2-reg-misc VSHLL to decodetree
       target/arm: Convert Neon VCVT f16/f32 insns to decodetree
       target/arm: Convert vectorised 2-reg-misc Neon ops to decodetree
       target/arm: Convert Neon 2-reg-misc crypto operations to decodetree
       target/arm: Rename NeonGenOneOpFn to NeonGenOne64OpFn
       target/arm: Fix capitalization in NeonGenTwo{Single, Double}OPFn typedefs
       target/arm: Make gen_swap_half() take separate src and dest
       target/arm: Convert Neon 2-reg-misc VREV32 and VREV16 to decodetree
       target/arm: Convert remaining simple 2-reg-misc Neon ops
       target/arm: Convert Neon VQABS, VQNEG to decodetree
       target/arm: Convert simple fp Neon 2-reg-misc insns
       target/arm: Convert Neon 2-reg-misc fp-compare-with-zero insns to decodetree
       target/arm: Convert Neon 2-reg-misc VRINT insns to decodetree
       target/arm: Convert Neon 2-reg-misc VCVT insns to decodetree
       target/arm: Convert Neon VSWP to decodetree
       target/arm: Convert Neon VTRN to decodetree
       target/arm: Move some functions used only in translate-neon.inc.c to that file
       target/arm: Remove unnecessary gen_io_end() calls
       target/arm: Remove dead code relating to SABA and UABA
-Paolo Bonzini (1):
+Philippe Mathieu-Daudé (15):
-      arm: fix malloc type mismatch
+      hw/watchdog/cmsdk-apb-watchdog: Add trace event for lock status
       hw/i2c/versatile_i2c: Add definitions for register addresses
       hw/i2c/versatile_i2c: Add SCL/SDA definitions
       hw/i2c: Add header for ARM SBCon two-wire serial bus interface
       hw/arm: Use TYPE_VERSATILE_I2C instead of hardcoded string
       hw/arm/mps2: Document CMSDK/FPGA APB subsystem sections
       hw/arm/mps2: Rename CMSDK AHB peripheral region
       hw/arm/mps2: Add CMSDK APB watchdog device
       hw/arm/mps2: Add CMSDK AHB GPIO peripherals as unimplemented devices
       hw/arm/mps2: Map the FPGA I/O block
       hw/arm/mps2: Add SPI devices
       hw/arm/mps2: Add I2C devices
       hw/arm/mps2: Add audio I2S interface as unimplemented device
       hw/arm/mps2-tz: Use the ARM SBCon two-wire serial bus interface
       target/arm: Check supported KVM features globally (not per vCPU)
-Peter Maydell (17):
+Shameer Kolothum (1):
-      target/arm: Honour FPCR.FZ in FRECPX
+      arm/virt: Add memory hot remove support
       MAINTAINERS: Add entries for newer MPS2 boards and devices
       Correct CPACR reset value for v7 cores
       memory.h: Improve IOMMU related documentation
       Make tb_invalidate_phys_addr() take a MemTxAttrs argument
       Make address_space_translate{, _cached}() take a MemTxAttrs argument
       Make address_space_map() take a MemTxAttrs argument
       Make address_space_access_valid() take a MemTxAttrs argument
       Make flatview_extend_translation() take a MemTxAttrs argument
       Make memory_region_access_valid() take a MemTxAttrs argument
       Make MemoryRegion valid.accepts callback take a MemTxAttrs argument
       Make flatview_access_valid() take a MemTxAttrs argument
       Make flatview_translate() take a MemTxAttrs argument
       Make address_space_get_iotlb_entry() take a MemTxAttrs argument
       Make flatview_do_translate() take a MemTxAttrs argument
       Make address_space_translate_iommu take a MemTxAttrs argument
       vmstate.h: Provide VMSTATE_BOOL_SUB_ARRAY
-Richard Henderson (1):
+ include/hw/i2c/arm_sbcon_i2c.h   |   35 ++
-      tcg: Fix helper function vs host abi for float16
+ target/arm/cpu.h                 |    2 +-
  target/arm/kvm_arm.h             |   21 +-
  target/arm/translate.h           |    8 +-
  target/arm/neon-dp.decode        |  106 ++++
  hw/acpi/generic_event_device.c   |   29 +
  hw/arm/mps2-tz.c                 |   23 +-
  hw/arm/mps2.c                    |   65 ++-
  hw/arm/realview.c                |    3 +-
  hw/arm/versatilepb.c             |    3 +-
  hw/arm/vexpress.c                |    3 +-
  hw/arm/virt.c                    |   63 +-
  hw/i2c/versatile_i2c.c           |   38 +-
  hw/watchdog/cmsdk-apb-watchdog.c |    1 +
  target/arm/cpu.c                 |    2 +-
  target/arm/cpu64.c               |   10 +-
  target/arm/kvm.c                 |    4 +-
  target/arm/kvm64.c               |   14 +-
  target/arm/translate-a64.c       |   20 +-
  target/arm/translate-neon.inc.c  | 1191 +++++++++++++++++++++++++++++++++++++-
  target/arm/translate-vfp.inc.c   |    7 +-
  target/arm/translate.c           | 1064 +---------------------------------
  tests/qtest/arm-cpu-features.c   |   38 +-
  util/oslib-posix.c               |   15 +
  MAINTAINERS                      |    1 +
  hw/arm/Kconfig                   |    8 +-
  hw/watchdog/trace-events         |    1 +
 files changed, 1624 insertions(+), 1151 deletions(-)
  create mode 100644 include/hw/i2c/arm_sbcon_i2c.h
-Shannon Zhao (3):
-      arm_gicv3_kvm: increase clroffset accordingly
-      ARM: ACPI: Fix use-after-free due to memory realloc
-      KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice
- include/exec/exec-all.h        |   5 +-
- include/exec/helper-head.h     |   2 +-
- include/exec/memory-internal.h |   3 +-
- include/exec/memory.h          | 128 +++++++++++++++++++++++++++++++++++------
- include/migration/vmstate.h    |   3 +
- include/sysemu/dma.h           |   6 +-
- accel/tcg/translate-all.c      |   4 +-
- exec.c                         |  95 ++++++++++++++++++------------
- hw/arm/boot.c                  |  18 +++---
- hw/arm/virt-acpi-build.c       |  20 +++++--
- hw/dma/xlnx-zdma.c             |  10 +++-
- hw/hppa/dino.c                 |   3 +-
- hw/intc/arm_gic_kvm.c          |   1 -
- hw/intc/arm_gicv3_cpuif.c      |  12 ++--
- hw/intc/arm_gicv3_kvm.c        |   2 +-
- hw/nvram/fw_cfg.c              |  12 ++--
- hw/s390x/s390-pci-inst.c       |   3 +-
- hw/scsi/esp.c                  |   3 +-
- hw/vfio/common.c               |   3 +-
- hw/virtio/vhost.c              |   3 +-
- hw/xen/xen_pt_msi.c            |   3 +-
- memory.c                       |  12 ++--
- memory_ldst.inc.c              |  18 +++---
- target/arm/gdbstub.c           |   3 +-
- target/arm/helper-a64.c        |  41 +++++++------
- target/arm/helper.c            |  90 ++++++++++++++++-------------
- target/ppc/mmu-hash64.c        |   3 +-
- target/riscv/helper.c          |   2 +-
- target/s390x/diag.c            |   6 +-
- target/s390x/excp_helper.c     |   3 +-
- target/s390x/mmu_helper.c      |   3 +-
- target/s390x/sigp.c            |   3 +-
- target/xtensa/op_helper.c      |   3 +-
- MAINTAINERS                    |   9 ++-
-files changed, 353 insertions(+), 182 deletions(-)

-New patch
+[PULL 01/42] hw/arm/virt: Add 5.0 HW compat props
+From: Andrew Jones <drjones@redhat.com>
+Cc: Cornelia Huck <cohuck@redhat.com>
+Signed-off-by: Andrew Jones <drjones@redhat.com>
+Reviewed-by: Cornelia Huck <cohuck@redhat.com>
+Message-id: 20200616140803.25515-1-drjones@redhat.com
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/arm/virt.c | 1 +
+file changed, 1 insertion(+)
+diff --git a/hw/arm/virt.c b/hw/arm/virt.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/arm/virt.c
++++ b/hw/arm/virt.c
+@@ -XXX,XX +XXX,XX @@ DEFINE_VIRT_MACHINE_AS_LATEST(5, 1)
+ static void virt_machine_5_0_options(MachineClass *mc)
+ {
+     virt_machine_5_1_options(mc);
++    compat_props_add(mc->compat_props, hw_compat_5_0, hw_compat_5_0_len);
+ }
+ DEFINE_VIRT_MACHINE(5, 0)
+--
+.20.1

-New patch
+[PULL 02/42] util/oslib-posix : qemu_init_exec_dir implementation for Mac
+From: David CARLIER <devnexen@gmail.com>
+From 3025a0ce3fdf7d3559fc35a52c659f635f5c750c Mon Sep 17 00:00:00 2001
+From: David Carlier <devnexen@gmail.com>
+Date: Tue, 26 May 2020 21:35:27 +0100
+Subject: [PATCH] util/oslib-posix : qemu_init_exec_dir implementation for Mac
+Using dyld API to get the full path of the current process.
+Signed-off-by: David Carlier <devnexen@gmail.com>
+Message-id: CA+XhMqxwC10XHVs4Z-JfE0-WLAU3ztDuU9QKVi31mjr59HWCxg@mail.gmail.com
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ util/oslib-posix.c | 15 +++++++++++++++
+file changed, 15 insertions(+)
+diff --git a/util/oslib-posix.c b/util/oslib-posix.c
+index XXXXXXX..XXXXXXX 100644
+--- a/util/oslib-posix.c
++++ b/util/oslib-posix.c
+@@ -XXX,XX +XXX,XX @@
+ #include <lwp.h>
+ #endif
++#ifdef __APPLE__
++#include <mach-o/dyld.h>
++#endif
++
+ #include "qemu/mmap-alloc.h"
+ #ifdef CONFIG_DEBUG_STACK_USAGE
+@@ -XXX,XX +XXX,XX @@ void qemu_init_exec_dir(const char *argv0)
+             p = buf;
+         }
+     }
++#elif defined(__APPLE__)
++    {
++        char fpath[PATH_MAX];
++        uint32_t len = sizeof(fpath);
++        if (_NSGetExecutablePath(fpath, &len) == 0) {
++            p = realpath(fpath, buf);
++            if (!p) {
++                return;
++            }
++        }
++    }
+ #endif
+     /* If we don't have any way of figuring out the actual executable
+        location then try argv[0].  */
+--
+.20.1

-New patch
+[PULL 03/42] target/arm: Convert Neon 2-reg-misc VREV64 to decodetree
+Convert the Neon VREV64 insn from the 2-reg-misc grouping to decodetree.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20200616170844.13318-2-peter.maydell@linaro.org
+---
+ target/arm/neon-dp.decode       | 12 ++++++++
+ target/arm/translate-neon.inc.c | 50 +++++++++++++++++++++++++++++++++
+ target/arm/translate.c          | 24 ++--------------
+files changed, 64 insertions(+), 22 deletions(-)
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/neon-dp.decode
++++ b/target/arm/neon-dp.decode
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+                  vm=%vm_dp vd=%vd_dp size=1
+     VDUP_scalar  1111 001 1 1 . 11 index:1 100 .... 11 000 q:1 . 0 .... \
+                  vm=%vm_dp vd=%vd_dp size=2
++
++    ##################################################################
++    # 2-reg-misc grouping:
++    # 1111 001 11 D 11 size:2 opc1:2 Vd:4 0 opc2:4 q:1 M 0 Vm:4
++    ##################################################################
++
++    &2misc vd vm q size
++
++    @2misc       .... ... .. . .. size:2 .. .... . .... q:1 . . .... \
++                 &2misc vm=%vm_dp vd=%vd_dp
++
++    VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
+   ]
+   # Subgroup for size != 0b11
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate-neon.inc.c
++++ b/target/arm/translate-neon.inc.c
+@@ -XXX,XX +XXX,XX @@ static bool trans_VDUP_scalar(DisasContext *s, arg_VDUP_scalar *a)
+                          a->q ? 16 : 8, a->q ? 16 : 8);
+     return true;
+ }
++
++static bool trans_VREV64(DisasContext *s, arg_VREV64 *a)
++{
++    int pass, half;
++
++    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
++        return false;
++    }
++
++    /* UNDEF accesses to D16-D31 if they don't exist. */
++    if (!dc_isar_feature(aa32_simd_r32, s) &&
++        ((a->vd | a->vm) & 0x10)) {
++        return false;
++    }
++
++    if ((a->vd | a->vm) & a->q) {
++        return false;
++    }
++
++    if (a->size == 3) {
++        return false;
++    }
++
++    if (!vfp_access_check(s)) {
++        return true;
++    }
++
++    for (pass = 0; pass < (a->q ? 2 : 1); pass++) {
++        TCGv_i32 tmp[2];
++
++        for (half = 0; half < 2; half++) {
++            tmp[half] = neon_load_reg(a->vm, pass * 2 + half);
++            switch (a->size) {
++            case 0:
++                tcg_gen_bswap32_i32(tmp[half], tmp[half]);
++                break;
++            case 1:
++                gen_swap_half(tmp[half]);
++                break;
++            case 2:
++                break;
++            default:
++                g_assert_not_reached();
++            }
++        }
++        neon_store_reg(a->vd, pass * 2, tmp[1]);
++        neon_store_reg(a->vd, pass * 2 + 1, tmp[0]);
++    }
++    return true;
++}
+diff --git a/target/arm/translate.c b/target/arm/translate.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate.c
++++ b/target/arm/translate.c
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+                 }
+                 switch (op) {
+                 case NEON_2RM_VREV64:
+-                    for (pass = 0; pass < (q ? 2 : 1); pass++) {
+-                        tmp = neon_load_reg(rm, pass * 2);
+-                        tmp2 = neon_load_reg(rm, pass * 2 + 1);
+-                        switch (size) {
+-                        case 0: tcg_gen_bswap32_i32(tmp, tmp); break;
+-                        case 1: gen_swap_half(tmp); break;
+-                        case 2: /* no-op */ break;
+-                        default: abort();
+-                        }
+-                        neon_store_reg(rd, pass * 2 + 1, tmp);
+-                        if (size == 2) {
+-                            neon_store_reg(rd, pass * 2, tmp2);
+-                        } else {
+-                            switch (size) {
+-                            case 0: tcg_gen_bswap32_i32(tmp2, tmp2); break;
+-                            case 1: gen_swap_half(tmp2); break;
+-                            default: abort();
+-                            }
+-                            neon_store_reg(rd, pass * 2, tmp2);
+-                        }
+-                    }
+-                    break;
++                    /* handled by decodetree */
++                    return 1;
+                 case NEON_2RM_VPADDL: case NEON_2RM_VPADDL_U:
+                 case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
+                     for (pass = 0; pass < q + 1; pass++) {
+--
+.20.1

-New patch
+[PULL 04/42] target/arm: Convert Neon 2-reg-misc pairwise ops to decodetree
+Convert the pairwise ops VPADDL and VPADAL in the 2-reg-misc grouping
 to decodetree.
 At this point we can get rid of the weird CPU_V001 #define that was
 used to avoid having to explicitly list all the arguments being
 passed to some TCG gen/helper functions.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-3-peter.maydell@linaro.org
 ---
  target/arm/neon-dp.decode       |   6 ++
  target/arm/translate-neon.inc.c | 149 ++++++++++++++++++++++++++++++++
  target/arm/translate.c          |  35 +-------
 files changed, 157 insertions(+), 33 deletions(-)
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
                   &2misc vm=%vm_dp vd=%vd_dp
      VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
 +
 +    VPADDL_S     1111 001 11 . 11 .. 00 .... 0 0100 . . 0 .... @2misc
 +    VPADDL_U     1111 001 11 . 11 .. 00 .... 0 0101 . . 0 .... @2misc
 +
 +    VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
 +    VPADAL_U     1111 001 11 . 11 .. 00 .... 0 1101 . . 0 .... @2misc
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VREV64(DisasContext *s, arg_VREV64 *a)
      }
      return true;
  }
 +
 +static bool do_2misc_pairwise(DisasContext *s, arg_2misc *a,
 +                              NeonGenWidenFn *widenfn,
 +                              NeonGenTwo64OpFn *opfn,
 +                              NeonGenTwo64OpFn *accfn)
 +{
 +    /*
 +     * Pairwise long operations: widen both halves of the pair,
 +     * combine the pairs with the opfn, and then possibly accumulate
 +     * into the destination with the accfn.
 +     */
 +    int pass;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!widenfn) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    for (pass = 0; pass < a->q + 1; pass++) {
 +        TCGv_i32 tmp;
 +        TCGv_i64 rm0_64, rm1_64, rd_64;
 +
 +        rm0_64 = tcg_temp_new_i64();
 +        rm1_64 = tcg_temp_new_i64();
 +        rd_64 = tcg_temp_new_i64();
 +        tmp = neon_load_reg(a->vm, pass * 2);
 +        widenfn(rm0_64, tmp);
 +        tcg_temp_free_i32(tmp);
 +        tmp = neon_load_reg(a->vm, pass * 2 + 1);
 +        widenfn(rm1_64, tmp);
 +        tcg_temp_free_i32(tmp);
 +        opfn(rd_64, rm0_64, rm1_64);
 +        tcg_temp_free_i64(rm0_64);
 +        tcg_temp_free_i64(rm1_64);
 +
 +        if (accfn) {
 +            TCGv_i64 tmp64 = tcg_temp_new_i64();
 +            neon_load_reg64(tmp64, a->vd + pass);
 +            accfn(rd_64, tmp64, rd_64);
 +            tcg_temp_free_i64(tmp64);
 +        }
 +        neon_store_reg64(rd_64, a->vd + pass);
 +        tcg_temp_free_i64(rd_64);
 +    }
 +    return true;
 +}
 +
 +static bool trans_VPADDL_S(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenWidenFn * const widenfn[] = {
 +        gen_helper_neon_widen_s8,
 +        gen_helper_neon_widen_s16,
 +        tcg_gen_ext_i32_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const opfn[] = {
 +        gen_helper_neon_paddl_u16,
 +        gen_helper_neon_paddl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +
 +    return do_2misc_pairwise(s, a, widenfn[a->size], opfn[a->size], NULL);
 +}
 +
 +static bool trans_VPADDL_U(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenWidenFn * const widenfn[] = {
 +        gen_helper_neon_widen_u8,
 +        gen_helper_neon_widen_u16,
 +        tcg_gen_extu_i32_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const opfn[] = {
 +        gen_helper_neon_paddl_u16,
 +        gen_helper_neon_paddl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +
 +    return do_2misc_pairwise(s, a, widenfn[a->size], opfn[a->size], NULL);
 +}
 +
 +static bool trans_VPADAL_S(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenWidenFn * const widenfn[] = {
 +        gen_helper_neon_widen_s8,
 +        gen_helper_neon_widen_s16,
 +        tcg_gen_ext_i32_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const opfn[] = {
 +        gen_helper_neon_paddl_u16,
 +        gen_helper_neon_paddl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const accfn[] = {
 +        gen_helper_neon_addl_u16,
 +        gen_helper_neon_addl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +
 +    return do_2misc_pairwise(s, a, widenfn[a->size], opfn[a->size],
 +                             accfn[a->size]);
 +}
 +
 +static bool trans_VPADAL_U(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenWidenFn * const widenfn[] = {
 +        gen_helper_neon_widen_u8,
 +        gen_helper_neon_widen_u16,
 +        tcg_gen_extu_i32_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const opfn[] = {
 +        gen_helper_neon_paddl_u16,
 +        gen_helper_neon_paddl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +    static NeonGenTwo64OpFn * const accfn[] = {
 +        gen_helper_neon_addl_u16,
 +        gen_helper_neon_addl_u32,
 +        tcg_gen_add_i64,
 +        NULL,
 +    };
 +
 +    return do_2misc_pairwise(s, a, widenfn[a->size], opfn[a->size],
 +                             accfn[a->size]);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_exception_return(DisasContext *s, TCGv_i32 pc)
      gen_rfe(s, pc, load_cpu_field(spsr));
  }
 -#define CPU_V001 cpu_V0, cpu_V0, cpu_V1
 -
  static int gen_neon_unzip(int rd, int rm, int size, int q)
  {
      TCGv_ptr pd, pm;
@@ -XXX,XX +XXX,XX @@ static inline void gen_neon_widen(TCGv_i64 dest, TCGv_i32 src, int size, int u)
      tcg_temp_free_i32(src);
  }
 -static inline void gen_neon_addl(int size)
 -{
 -    switch (size) {
 -    case 0: gen_helper_neon_addl_u16(CPU_V001); break;
 -    case 1: gen_helper_neon_addl_u32(CPU_V001); break;
 -    case 2: tcg_gen_add_i64(CPU_V001); break;
 -    default: abort();
 -    }
 -}
 -
  static void gen_neon_narrow_op(int op, int u, int size,
                                 TCGv_i32 dest, TCGv_i64 src)
  {
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  }
                  switch (op) {
                  case NEON_2RM_VREV64:
 -                    /* handled by decodetree */
 -                    return 1;
                  case NEON_2RM_VPADDL: case NEON_2RM_VPADDL_U:
                  case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
 -                    for (pass = 0; pass < q + 1; pass++) {
 -                        tmp = neon_load_reg(rm, pass * 2);
 -                        gen_neon_widen(cpu_V0, tmp, size, op & 1);
 -                        tmp = neon_load_reg(rm, pass * 2 + 1);
 -                        gen_neon_widen(cpu_V1, tmp, size, op & 1);
 -                        switch (size) {
 -                        case 0: gen_helper_neon_paddl_u16(CPU_V001); break;
 -                        case 1: gen_helper_neon_paddl_u32(CPU_V001); break;
 -                        case 2: tcg_gen_add_i64(CPU_V001); break;
 -                        default: abort();
 -                        }
 -                        if (op >= NEON_2RM_VPADAL) {
 -                            /* Accumulate.  */
 -                            neon_load_reg64(cpu_V1, rd + pass);
 -                            gen_neon_addl(size);
 -                        }
 -                        neon_store_reg64(cpu_V0, rd + pass);
 -                    }
 -                    break;
 +                    /* handled by decodetree */
 +                    return 1;
                  case NEON_2RM_VTRN:
                      if (size == 2) {
                          int n;
 --
 .20.1

-New patch
+[PULL 05/42] target/arm: Convert VZIP, VUZP to decodetree
+Convert the Neon VZIP and VUZP insns in the 2-reg-misc group to
 decodetree.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-4-peter.maydell@linaro.org
 ---
  target/arm/neon-dp.decode       |  3 ++
  target/arm/translate-neon.inc.c | 74 ++++++++++++++++++++++++++
  target/arm/translate.c          | 92 +--------------------------------
 files changed, 79 insertions(+), 90 deletions(-)
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
      VPADAL_U     1111 001 11 . 11 .. 00 .... 0 1101 . . 0 .... @2misc
 +
 +    VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
 +    VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VPADAL_U(DisasContext *s, arg_2misc *a)
      return do_2misc_pairwise(s, a, widenfn[a->size], opfn[a->size],
                               accfn[a->size]);
  }
 +
 +typedef void ZipFn(TCGv_ptr, TCGv_ptr);
 +
 +static bool do_zip_uzp(DisasContext *s, arg_2misc *a,
 +                       ZipFn *fn)
 +{
 +    TCGv_ptr pd, pm;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!fn) {
 +        /* Bad size or size/q combination */
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    pd = vfp_reg_ptr(true, a->vd);
 +    pm = vfp_reg_ptr(true, a->vm);
 +    fn(pd, pm);
 +    tcg_temp_free_ptr(pd);
 +    tcg_temp_free_ptr(pm);
 +    return true;
 +}
 +
 +static bool trans_VUZP(DisasContext *s, arg_2misc *a)
 +{
 +    static ZipFn * const fn[2][4] = {
 +        {
 +            gen_helper_neon_unzip8,
 +            gen_helper_neon_unzip16,
 +            NULL,
 +            NULL,
 +        }, {
 +            gen_helper_neon_qunzip8,
 +            gen_helper_neon_qunzip16,
 +            gen_helper_neon_qunzip32,
 +            NULL,
 +        }
 +    };
 +    return do_zip_uzp(s, a, fn[a->q][a->size]);
 +}
 +
 +static bool trans_VZIP(DisasContext *s, arg_2misc *a)
 +{
 +    static ZipFn * const fn[2][4] = {
 +        {
 +            gen_helper_neon_zip8,
 +            gen_helper_neon_zip16,
 +            NULL,
 +            NULL,
 +        }, {
 +            gen_helper_neon_qzip8,
 +            gen_helper_neon_qzip16,
 +            gen_helper_neon_qzip32,
 +            NULL,
 +        }
 +    };
 +    return do_zip_uzp(s, a, fn[a->q][a->size]);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_exception_return(DisasContext *s, TCGv_i32 pc)
      gen_rfe(s, pc, load_cpu_field(spsr));
  }
 -static int gen_neon_unzip(int rd, int rm, int size, int q)
 -{
 -    TCGv_ptr pd, pm;
 -
 -    if (!q && size == 2) {
 -        return 1;
 -    }
 -    pd = vfp_reg_ptr(true, rd);
 -    pm = vfp_reg_ptr(true, rm);
 -    if (q) {
 -        switch (size) {
 -        case 0:
 -            gen_helper_neon_qunzip8(pd, pm);
 -            break;
 -        case 1:
 -            gen_helper_neon_qunzip16(pd, pm);
 -            break;
 -        case 2:
 -            gen_helper_neon_qunzip32(pd, pm);
 -            break;
 -        default:
 -            abort();
 -        }
 -    } else {
 -        switch (size) {
 -        case 0:
 -            gen_helper_neon_unzip8(pd, pm);
 -            break;
 -        case 1:
 -            gen_helper_neon_unzip16(pd, pm);
 -            break;
 -        default:
 -            abort();
 -        }
 -    }
 -    tcg_temp_free_ptr(pd);
 -    tcg_temp_free_ptr(pm);
 -    return 0;
 -}
 -
 -static int gen_neon_zip(int rd, int rm, int size, int q)
 -{
 -    TCGv_ptr pd, pm;
 -
 -    if (!q && size == 2) {
 -        return 1;
 -    }
 -    pd = vfp_reg_ptr(true, rd);
 -    pm = vfp_reg_ptr(true, rm);
 -    if (q) {
 -        switch (size) {
 -        case 0:
 -            gen_helper_neon_qzip8(pd, pm);
 -            break;
 -        case 1:
 -            gen_helper_neon_qzip16(pd, pm);
 -            break;
 -        case 2:
 -            gen_helper_neon_qzip32(pd, pm);
 -            break;
 -        default:
 -            abort();
 -        }
 -    } else {
 -        switch (size) {
 -        case 0:
 -            gen_helper_neon_zip8(pd, pm);
 -            break;
 -        case 1:
 -            gen_helper_neon_zip16(pd, pm);
 -            break;
 -        default:
 -            abort();
 -        }
 -    }
 -    tcg_temp_free_ptr(pd);
 -    tcg_temp_free_ptr(pm);
 -    return 0;
 -}
 -
  static void gen_neon_trn_u8(TCGv_i32 t0, TCGv_i32 t1)
  {
      TCGv_i32 rd, tmp;
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VREV64:
                  case NEON_2RM_VPADDL: case NEON_2RM_VPADDL_U:
                  case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
 +                case NEON_2RM_VUZP:
 +                case NEON_2RM_VZIP:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                          goto elementwise;
                      }
                      break;
 -                case NEON_2RM_VUZP:
 -                    if (gen_neon_unzip(rd, rm, size, q)) {
 -                        return 1;
 -                    }
 -                    break;
 -                case NEON_2RM_VZIP:
 -                    if (gen_neon_zip(rd, rm, size, q)) {
 -                        return 1;
 -                    }
 -                    break;
                  case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
                      /* also VQMOVUN; op field and mnemonics don't line up */
                      if (rm & 1) {
 --
 .20.1

-New patch
+[PULL 06/42] target/arm: Convert Neon narrowing moves to decodetree
+Convert the Neon narrowing moves VMQNV, VQMOVN, VQMOVUN in the 2-reg-misc
 group to decodetree.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-5-peter.maydell@linaro.org
 ---
  target/arm/neon-dp.decode       |  9 ++++
  target/arm/translate-neon.inc.c | 59 ++++++++++++++++++++++++
  target/arm/translate.c          | 81 +--------------------------------
 files changed, 70 insertions(+), 79 deletions(-)
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      @2misc       .... ... .. . .. size:2 .. .... . .... q:1 . . .... \
                   &2misc vm=%vm_dp vd=%vd_dp
 +    @2misc_q0    .... ... .. . .. size:2 .. .... . .... . . . .... \
 +                 &2misc vm=%vm_dp vd=%vd_dp q=0
      VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
      VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
 +
 +    VMOVN        1111 001 11 . 11 .. 10 .... 0 0100 0 . 0 .... @2misc_q0
 +    # VQMOVUN: unsigned result (source is always signed)
 +    VQMOVUN      1111 001 11 . 11 .. 10 .... 0 0100 1 . 0 .... @2misc_q0
 +    # VQMOVN: signed result, source may be signed (_S) or unsigned (_U)
 +    VQMOVN_S     1111 001 11 . 11 .. 10 .... 0 0101 0 . 0 .... @2misc_q0
 +    VQMOVN_U     1111 001 11 . 11 .. 10 .... 0 0101 1 . 0 .... @2misc_q0
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VZIP(DisasContext *s, arg_2misc *a)
      };
      return do_zip_uzp(s, a, fn[a->q][a->size]);
  }
 +
 +static bool do_vmovn(DisasContext *s, arg_2misc *a,
 +                     NeonGenNarrowEnvFn *narrowfn)
 +{
 +    TCGv_i64 rm;
 +    TCGv_i32 rd0, rd1;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (a->vm & 1) {
 +        return false;
 +    }
 +
 +    if (!narrowfn) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    rm = tcg_temp_new_i64();
 +    rd0 = tcg_temp_new_i32();
 +    rd1 = tcg_temp_new_i32();
 +
 +    neon_load_reg64(rm, a->vm);
 +    narrowfn(rd0, cpu_env, rm);
 +    neon_load_reg64(rm, a->vm + 1);
 +    narrowfn(rd1, cpu_env, rm);
 +    neon_store_reg(a->vd, 0, rd0);
 +    neon_store_reg(a->vd, 1, rd1);
 +    tcg_temp_free_i64(rm);
 +    return true;
 +}
 +
 +#define DO_VMOVN(INSN, FUNC)                                    \
 +    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
 +    {                                                           \
 +        static NeonGenNarrowEnvFn * const narrowfn[] = {        \
 +            FUNC##8,                                            \
 +            FUNC##16,                                           \
 +            FUNC##32,                                           \
 +            NULL,                                               \
 +        };                                                      \
 +        return do_vmovn(s, a, narrowfn[a->size]);               \
 +    }
 +
 +DO_VMOVN(VMOVN, gen_neon_narrow_u)
 +DO_VMOVN(VQMOVUN, gen_helper_neon_unarrow_sat)
 +DO_VMOVN(VQMOVN_S, gen_helper_neon_narrow_sat_s)
 +DO_VMOVN(VQMOVN_U, gen_helper_neon_narrow_sat_u)
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
      tcg_temp_free_i32(rd);
  }
 -static inline void gen_neon_narrow(int size, TCGv_i32 dest, TCGv_i64 src)
 -{
 -    switch (size) {
 -    case 0: gen_helper_neon_narrow_u8(dest, src); break;
 -    case 1: gen_helper_neon_narrow_u16(dest, src); break;
 -    case 2: tcg_gen_extrl_i64_i32(dest, src); break;
 -    default: abort();
 -    }
 -}
 -
 -static inline void gen_neon_narrow_sats(int size, TCGv_i32 dest, TCGv_i64 src)
 -{
 -    switch (size) {
 -    case 0: gen_helper_neon_narrow_sat_s8(dest, cpu_env, src); break;
 -    case 1: gen_helper_neon_narrow_sat_s16(dest, cpu_env, src); break;
 -    case 2: gen_helper_neon_narrow_sat_s32(dest, cpu_env, src); break;
 -    default: abort();
 -    }
 -}
 -
 -static inline void gen_neon_narrow_satu(int size, TCGv_i32 dest, TCGv_i64 src)
 -{
 -    switch (size) {
 -    case 0: gen_helper_neon_narrow_sat_u8(dest, cpu_env, src); break;
 -    case 1: gen_helper_neon_narrow_sat_u16(dest, cpu_env, src); break;
 -    case 2: gen_helper_neon_narrow_sat_u32(dest, cpu_env, src); break;
 -    default: abort();
 -    }
 -}
 -
 -static inline void gen_neon_unarrow_sats(int size, TCGv_i32 dest, TCGv_i64 src)
 -{
 -    switch (size) {
 -    case 0: gen_helper_neon_unarrow_sat8(dest, cpu_env, src); break;
 -    case 1: gen_helper_neon_unarrow_sat16(dest, cpu_env, src); break;
 -    case 2: gen_helper_neon_unarrow_sat32(dest, cpu_env, src); break;
 -    default: abort();
 -    }
 -}
 -
  static inline void gen_neon_widen(TCGv_i64 dest, TCGv_i32 src, int size, int u)
  {
      if (u) {
@@ -XXX,XX +XXX,XX @@ static inline void gen_neon_widen(TCGv_i64 dest, TCGv_i32 src, int size, int u)
      tcg_temp_free_i32(src);
  }
 -static void gen_neon_narrow_op(int op, int u, int size,
 -                               TCGv_i32 dest, TCGv_i64 src)
 -{
 -    if (op) {
 -        if (u) {
 -            gen_neon_unarrow_sats(size, dest, src);
 -        } else {
 -            gen_neon_narrow(size, dest, src);
 -        }
 -    } else {
 -        if (u) {
 -            gen_neon_narrow_satu(size, dest, src);
 -        } else {
 -            gen_neon_narrow_sats(size, dest, src);
 -        }
 -    }
 -}
 -
  /* Symbolic constants for op fields for Neon 2-register miscellaneous.
   * The values correspond to bits [17:16,10:7]; see the ARM ARM DDI0406B
   * table A7-13.
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                      !arm_dc_feature(s, ARM_FEATURE_V8)) {
                      return 1;
                  }
 -                if ((op != NEON_2RM_VMOVN && op != NEON_2RM_VQMOVN) &&
 -                    q && ((rm | rd) & 1)) {
 +                if (q && ((rm | rd) & 1)) {
                      return 1;
                  }
                  switch (op) {
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
                  case NEON_2RM_VUZP:
                  case NEON_2RM_VZIP:
 +                case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                          goto elementwise;
                      }
                      break;
 -                case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
 -                    /* also VQMOVUN; op field and mnemonics don't line up */
 -                    if (rm & 1) {
 -                        return 1;
 -                    }
 -                    tmp2 = NULL;
 -                    for (pass = 0; pass < 2; pass++) {
 -                        neon_load_reg64(cpu_V0, rm + pass);
 -                        tmp = tcg_temp_new_i32();
 -                        gen_neon_narrow_op(op == NEON_2RM_VMOVN, q, size,
 -                                           tmp, cpu_V0);
 -                        if (pass == 0) {
 -                            tmp2 = tmp;
 -                        } else {
 -                            neon_store_reg(rd, 0, tmp2);
 -                            neon_store_reg(rd, 1, tmp);
 -                        }
 -                    }
 -                    break;
                  case NEON_2RM_VSHLL:
                      if (q || (rd & 1)) {
                          return 1;
 --
 .20.1

-[Qemu-devel] [PULL 19/25] Make flatview_translate() take a MemTxAttrs argument
+[PULL 07/42] target/arm: Convert Neon 2-reg-misc VSHLL to decodetree
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the VSHLL insn in the 2-reg-misc Neon group to decodetree.
 add MemTxAttrs as an argument to flatview_translate(); all its
 callers now have attrs available.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-11-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-6-peter.maydell@linaro.org
 ---
- include/exec/memory.h |  7 ++++---
+ target/arm/neon-dp.decode       |  2 ++
- exec.c                | 17 +++++++++--------
+ target/arm/translate-neon.inc.c | 52 +++++++++++++++++++++++++++++++++
-files changed, 13 insertions(+), 11 deletions(-)
+ target/arm/translate.c          | 35 +---------------------
 files changed, 55 insertions(+), 34 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/neon-dp.decode
-+++ b/include/exec/memory.h
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
-  */
+     # VQMOVN: signed result, source may be signed (_S) or unsigned (_U)
- MemoryRegion *flatview_translate(FlatView *fv,
+     VQMOVN_S     1111 001 11 . 11 .. 10 .... 0 0101 0 . 0 .... @2misc_q0
-                                  hwaddr addr, hwaddr *xlat,
+     VQMOVN_U     1111 001 11 . 11 .. 10 .... 0 0101 1 . 0 .... @2misc_q0
--                                 hwaddr *len, bool is_write);
++
-+                                 hwaddr *len, bool is_write,
++    VSHLL        1111 001 11 . 11 .. 10 .... 0 0110 0 . 0 .... @2misc_q0
-+                                 MemTxAttrs attrs);
+   ]
- static inline MemoryRegion *address_space_translate(AddressSpace *as,
+   # Subgroup for size != 0b11
-                                                     hwaddr addr, hwaddr *xlat,
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ static inline MemoryRegion *address_space_translate(AddressSpace *as,
+index XXXXXXX..XXXXXXX 100644
-                                                     MemTxAttrs attrs)
+--- a/target/arm/translate-neon.inc.c
- {
++++ b/target/arm/translate-neon.inc.c
-     return flatview_translate(address_space_to_flatview(as),
+@@ -XXX,XX +XXX,XX @@ DO_VMOVN(VMOVN, gen_neon_narrow_u)
--                              addr, xlat, len, is_write);
+ DO_VMOVN(VQMOVUN, gen_helper_neon_unarrow_sat)
-+                              addr, xlat, len, is_write, attrs);
+ DO_VMOVN(VQMOVN_S, gen_helper_neon_narrow_sat_s)
  DO_VMOVN(VQMOVN_U, gen_helper_neon_narrow_sat_u)
 +
 +static bool trans_VSHLL(DisasContext *s, arg_2misc *a)
 +{
 +    TCGv_i32 rm0, rm1;
 +    TCGv_i64 rd;
 +    static NeonGenWidenFn * const widenfns[] = {
 +        gen_helper_neon_widen_u8,
 +        gen_helper_neon_widen_u16,
 +        tcg_gen_extu_i32_i64,
 +        NULL,
 +    };
 +    NeonGenWidenFn *widenfn = widenfns[a->size];
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (a->vd & 1) {
 +        return false;
 +    }
 +
 +    if (!widenfn) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    rd = tcg_temp_new_i64();
 +
 +    rm0 = neon_load_reg(a->vm, 0);
 +    rm1 = neon_load_reg(a->vm, 1);
 +
 +    widenfn(rd, rm0);
 +    tcg_gen_shli_i64(rd, rd, 8 << a->size);
 +    neon_store_reg64(rd, a->vd);
 +    widenfn(rd, rm1);
 +    tcg_gen_shli_i64(rd, rd, 8 << a->size);
 +    neon_store_reg64(rd, a->vd + 1);
 +
 +    tcg_temp_free_i64(rd);
 +    tcg_temp_free_i32(rm0);
 +    tcg_temp_free_i32(rm1);
 +    return true;
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
      tcg_temp_free_i32(rd);
  }
- /* address_space_access_valid: check for validity of accessing an address
+-static inline void gen_neon_widen(TCGv_i64 dest, TCGv_i32 src, int size, int u)
-@@ -XXX,XX +XXX,XX @@ MemTxResult address_space_read(AddressSpace *as, hwaddr addr,
+-{
-             rcu_read_lock();
+-    if (u) {
-             fv = address_space_to_flatview(as);
+-        switch (size) {
-             l = len;
+-        case 0: gen_helper_neon_widen_u8(dest, src); break;
--            mr = flatview_translate(fv, addr, &addr1, &l, false);
+-        case 1: gen_helper_neon_widen_u16(dest, src); break;
-+            mr = flatview_translate(fv, addr, &addr1, &l, false, attrs);
+-        case 2: tcg_gen_extu_i32_i64(dest, src); break;
-             if (len == l && memory_access_is_direct(mr, false)) {
+-        default: abort();
-                 ptr = qemu_map_ram_ptr(mr->ram_block, addr1);
+-        }
-                 memcpy(buf, ptr, len);
+-    } else {
-diff --git a/exec.c b/exec.c
+-        switch (size) {
-index XXXXXXX..XXXXXXX 100644
+-        case 0: gen_helper_neon_widen_s8(dest, src); break;
---- a/exec.c
+-        case 1: gen_helper_neon_widen_s16(dest, src); break;
-+++ b/exec.c
+-        case 2: tcg_gen_ext_i32_i64(dest, src); break;
-@@ -XXX,XX +XXX,XX @@ iotlb_fail:
+-        default: abort();
+-        }
- /* Called from RCU critical section */
+-    }
- MemoryRegion *flatview_translate(FlatView *fv, hwaddr addr, hwaddr *xlat,
+-    tcg_temp_free_i32(src);
--                                 hwaddr *plen, bool is_write)
+-}
-+                                 hwaddr *plen, bool is_write,
+-
-+                                 MemTxAttrs attrs)
+ /* Symbolic constants for op fields for Neon 2-register miscellaneous.
- {
+  * The values correspond to bits [17:16,10:7]; see the ARM ARM DDI0406B
-     MemoryRegion *mr;
+  * table A7-13.
-     MemoryRegionSection section;
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-@@ -XXX,XX +XXX,XX @@ static MemTxResult flatview_write_continue(FlatView *fv, hwaddr addr,
+                 case NEON_2RM_VUZP:
-         }
+                 case NEON_2RM_VZIP:
+                 case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
-         l = len;
++                case NEON_2RM_VSHLL:
--        mr = flatview_translate(fv, addr, &addr1, &l, true);
+                     /* handled by decodetree */
-+        mr = flatview_translate(fv, addr, &addr1, &l, true, attrs);
+                     return 1;
-     }
+                 case NEON_2RM_VTRN:
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-     return result;
+                         goto elementwise;
-@@ -XXX,XX +XXX,XX @@ static MemTxResult flatview_write(FlatView *fv, hwaddr addr, MemTxAttrs attrs,
+                     }
-     MemTxResult result = MEMTX_OK;
+                     break;
+-                case NEON_2RM_VSHLL:
-     l = len;
+-                    if (q || (rd & 1)) {
--    mr = flatview_translate(fv, addr, &addr1, &l, true);
+-                        return 1;
-+    mr = flatview_translate(fv, addr, &addr1, &l, true, attrs);
+-                    }
-     result = flatview_write_continue(fv, addr, attrs, buf, len,
+-                    tmp = neon_load_reg(rm, 0);
-                                      addr1, l, mr);
+-                    tmp2 = neon_load_reg(rm, 1);
+-                    for (pass = 0; pass < 2; pass++) {
-@@ -XXX,XX +XXX,XX @@ MemTxResult flatview_read_continue(FlatView *fv, hwaddr addr,
+-                        if (pass == 1)
-         }
+-                            tmp = tmp2;
+-                        gen_neon_widen(cpu_V0, tmp, size, 1);
-         l = len;
+-                        tcg_gen_shli_i64(cpu_V0, cpu_V0, 8 << size);
--        mr = flatview_translate(fv, addr, &addr1, &l, false);
+-                        neon_store_reg64(cpu_V0, rd + pass);
-+        mr = flatview_translate(fv, addr, &addr1, &l, false, attrs);
+-                    }
-     }
+-                    break;
+                 case NEON_2RM_VCVT_F16_F32:
-     return result;
+                 {
-@@ -XXX,XX +XXX,XX @@ static MemTxResult flatview_read(FlatView *fv, hwaddr addr,
+                     TCGv_ptr fpst;
      MemoryRegion *mr;
      l = len;
 -    mr = flatview_translate(fv, addr, &addr1, &l, false);
 +    mr = flatview_translate(fv, addr, &addr1, &l, false, attrs);
      return flatview_read_continue(fv, addr, attrs, buf, len,
                                    addr1, l, mr);
  }
@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
      while (len > 0) {
          l = len;
 -        mr = flatview_translate(fv, addr, &xlat, &l, is_write);
 +        mr = flatview_translate(fv, addr, &xlat, &l, is_write, attrs);
          if (!memory_access_is_direct(mr, is_write)) {
              l = memory_access_size(mr, l, addr);
              if (!memory_region_access_valid(mr, xlat, l, is_write, attrs)) {
@@ -XXX,XX +XXX,XX @@ flatview_extend_translation(FlatView *fv, hwaddr addr,
          len = target_len;
          this_mr = flatview_translate(fv, addr, &xlat,
 -                                                   &len, is_write);
 +                                     &len, is_write, attrs);
          if (this_mr != mr || xlat != base + done) {
              return done;
          }
@@ -XXX,XX +XXX,XX @@ void *address_space_map(AddressSpace *as,
      l = len;
      rcu_read_lock();
      fv = address_space_to_flatview(as);
 -    mr = flatview_translate(fv, addr, &xlat, &l, is_write);
 +    mr = flatview_translate(fv, addr, &xlat, &l, is_write, attrs);
      if (!memory_access_is_direct(mr, is_write)) {
          if (atomic_xchg(&bounce.in_use, true)) {
 --
-.17.1
+.20.1

-New patch
+[PULL 08/42] target/arm: Convert Neon VCVT f16/f32 insns to decodetree
+Convert the Neon insns in the 2-reg-misc group which are
 VCVT between f32 and f16 to decodetree.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-7-peter.maydell@linaro.org
 ---
  target/arm/neon-dp.decode       |  3 ++
  target/arm/translate-neon.inc.c | 96 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 65 ++--------------------
 files changed, 102 insertions(+), 62 deletions(-)
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VQMOVN_U     1111 001 11 . 11 .. 10 .... 0 0101 1 . 0 .... @2misc_q0
      VSHLL        1111 001 11 . 11 .. 10 .... 0 0110 0 . 0 .... @2misc_q0
 +
 +    VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
 +    VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VSHLL(DisasContext *s, arg_2misc *a)
      tcg_temp_free_i32(rm1);
      return true;
  }
 +
 +static bool trans_VCVT_F16_F32(DisasContext *s, arg_2misc *a)
 +{
 +    TCGv_ptr fpst;
 +    TCGv_i32 ahp, tmp, tmp2, tmp3;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON) ||
 +        !dc_isar_feature(aa32_fp16_spconv, s)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if ((a->vm & 1) || (a->size != 1)) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    fpst = get_fpstatus_ptr(true);
 +    ahp = get_ahp_flag();
 +    tmp = neon_load_reg(a->vm, 0);
 +    gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
 +    tmp2 = neon_load_reg(a->vm, 1);
 +    gen_helper_vfp_fcvt_f32_to_f16(tmp2, tmp2, fpst, ahp);
 +    tcg_gen_shli_i32(tmp2, tmp2, 16);
 +    tcg_gen_or_i32(tmp2, tmp2, tmp);
 +    tcg_temp_free_i32(tmp);
 +    tmp = neon_load_reg(a->vm, 2);
 +    gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
 +    tmp3 = neon_load_reg(a->vm, 3);
 +    neon_store_reg(a->vd, 0, tmp2);
 +    gen_helper_vfp_fcvt_f32_to_f16(tmp3, tmp3, fpst, ahp);
 +    tcg_gen_shli_i32(tmp3, tmp3, 16);
 +    tcg_gen_or_i32(tmp3, tmp3, tmp);
 +    neon_store_reg(a->vd, 1, tmp3);
 +    tcg_temp_free_i32(tmp);
 +    tcg_temp_free_i32(ahp);
 +    tcg_temp_free_ptr(fpst);
 +
 +    return true;
 +}
 +
 +static bool trans_VCVT_F32_F16(DisasContext *s, arg_2misc *a)
 +{
 +    TCGv_ptr fpst;
 +    TCGv_i32 ahp, tmp, tmp2, tmp3;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON) ||
 +        !dc_isar_feature(aa32_fp16_spconv, s)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if ((a->vd & 1) || (a->size != 1)) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    fpst = get_fpstatus_ptr(true);
 +    ahp = get_ahp_flag();
 +    tmp3 = tcg_temp_new_i32();
 +    tmp = neon_load_reg(a->vm, 0);
 +    tmp2 = neon_load_reg(a->vm, 1);
 +    tcg_gen_ext16u_i32(tmp3, tmp);
 +    gen_helper_vfp_fcvt_f16_to_f32(tmp3, tmp3, fpst, ahp);
 +    neon_store_reg(a->vd, 0, tmp3);
 +    tcg_gen_shri_i32(tmp, tmp, 16);
 +    gen_helper_vfp_fcvt_f16_to_f32(tmp, tmp, fpst, ahp);
 +    neon_store_reg(a->vd, 1, tmp);
 +    tmp3 = tcg_temp_new_i32();
 +    tcg_gen_ext16u_i32(tmp3, tmp2);
 +    gen_helper_vfp_fcvt_f16_to_f32(tmp3, tmp3, fpst, ahp);
 +    neon_store_reg(a->vd, 2, tmp3);
 +    tcg_gen_shri_i32(tmp2, tmp2, 16);
 +    gen_helper_vfp_fcvt_f16_to_f32(tmp2, tmp2, fpst, ahp);
 +    neon_store_reg(a->vd, 3, tmp2);
 +    tcg_temp_free_i32(ahp);
 +    tcg_temp_free_ptr(fpst);
 +
 +    return true;
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
      int pass;
      int u;
      int vec_size;
 -    TCGv_i32 tmp, tmp2, tmp3;
 +    TCGv_i32 tmp, tmp2;
      if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
          return 1;
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VZIP:
                  case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
                  case NEON_2RM_VSHLL:
 +                case NEON_2RM_VCVT_F16_F32:
 +                case NEON_2RM_VCVT_F32_F16:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                          goto elementwise;
                      }
                      break;
 -                case NEON_2RM_VCVT_F16_F32:
 -                {
 -                    TCGv_ptr fpst;
 -                    TCGv_i32 ahp;
 -
 -                    if (!dc_isar_feature(aa32_fp16_spconv, s) ||
 -                        q || (rm & 1)) {
 -                        return 1;
 -                    }
 -                    fpst = get_fpstatus_ptr(true);
 -                    ahp = get_ahp_flag();
 -                    tmp = neon_load_reg(rm, 0);
 -                    gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
 -                    tmp2 = neon_load_reg(rm, 1);
 -                    gen_helper_vfp_fcvt_f32_to_f16(tmp2, tmp2, fpst, ahp);
 -                    tcg_gen_shli_i32(tmp2, tmp2, 16);
 -                    tcg_gen_or_i32(tmp2, tmp2, tmp);
 -                    tcg_temp_free_i32(tmp);
 -                    tmp = neon_load_reg(rm, 2);
 -                    gen_helper_vfp_fcvt_f32_to_f16(tmp, tmp, fpst, ahp);
 -                    tmp3 = neon_load_reg(rm, 3);
 -                    neon_store_reg(rd, 0, tmp2);
 -                    gen_helper_vfp_fcvt_f32_to_f16(tmp3, tmp3, fpst, ahp);
 -                    tcg_gen_shli_i32(tmp3, tmp3, 16);
 -                    tcg_gen_or_i32(tmp3, tmp3, tmp);
 -                    neon_store_reg(rd, 1, tmp3);
 -                    tcg_temp_free_i32(tmp);
 -                    tcg_temp_free_i32(ahp);
 -                    tcg_temp_free_ptr(fpst);
 -                    break;
 -                }
 -                case NEON_2RM_VCVT_F32_F16:
 -                {
 -                    TCGv_ptr fpst;
 -                    TCGv_i32 ahp;
 -                    if (!dc_isar_feature(aa32_fp16_spconv, s) ||
 -                        q || (rd & 1)) {
 -                        return 1;
 -                    }
 -                    fpst = get_fpstatus_ptr(true);
 -                    ahp = get_ahp_flag();
 -                    tmp3 = tcg_temp_new_i32();
 -                    tmp = neon_load_reg(rm, 0);
 -                    tmp2 = neon_load_reg(rm, 1);
 -                    tcg_gen_ext16u_i32(tmp3, tmp);
 -                    gen_helper_vfp_fcvt_f16_to_f32(tmp3, tmp3, fpst, ahp);
 -                    neon_store_reg(rd, 0, tmp3);
 -                    tcg_gen_shri_i32(tmp, tmp, 16);
 -                    gen_helper_vfp_fcvt_f16_to_f32(tmp, tmp, fpst, ahp);
 -                    neon_store_reg(rd, 1, tmp);
 -                    tmp3 = tcg_temp_new_i32();
 -                    tcg_gen_ext16u_i32(tmp3, tmp2);
 -                    gen_helper_vfp_fcvt_f16_to_f32(tmp3, tmp3, fpst, ahp);
 -                    neon_store_reg(rd, 2, tmp3);
 -                    tcg_gen_shri_i32(tmp2, tmp2, 16);
 -                    gen_helper_vfp_fcvt_f16_to_f32(tmp2, tmp2, fpst, ahp);
 -                    neon_store_reg(rd, 3, tmp2);
 -                    tcg_temp_free_i32(ahp);
 -                    tcg_temp_free_ptr(fpst);
 -                    break;
 -                }
                  case NEON_2RM_AESE: case NEON_2RM_AESMC:
                      if (!dc_isar_feature(aa32_aes, s) || ((rm | rd) & 1)) {
                          return 1;
 --
 .20.1

-New patch
+[PULL 09/42] target/arm: Convert vectorised 2-reg-misc Neon ops to decodetree
+Convert to decodetree the insns in the Neon 2-reg-misc grouping which
+we implement using gvec.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20200616170844.13318-8-peter.maydell@linaro.org
+---
+ target/arm/neon-dp.decode       | 11 +++++++
+ target/arm/translate-neon.inc.c | 55 +++++++++++++++++++++++++++++++++
+ target/arm/translate.c          | 35 +++++----------------
+files changed, 74 insertions(+), 27 deletions(-)
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/neon-dp.decode
++++ b/target/arm/neon-dp.decode
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+     VPADDL_S     1111 001 11 . 11 .. 00 .... 0 0100 . . 0 .... @2misc
+     VPADDL_U     1111 001 11 . 11 .. 00 .... 0 0101 . . 0 .... @2misc
++    VMVN         1111 001 11 . 11 .. 00 .... 0 1011 . . 0 .... @2misc
++
+     VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
+     VPADAL_U     1111 001 11 . 11 .. 00 .... 0 1101 . . 0 .... @2misc
++    VCGT0        1111 001 11 . 11 .. 01 .... 0 0000 . . 0 .... @2misc
++    VCGE0        1111 001 11 . 11 .. 01 .... 0 0001 . . 0 .... @2misc
++    VCEQ0        1111 001 11 . 11 .. 01 .... 0 0010 . . 0 .... @2misc
++    VCLE0        1111 001 11 . 11 .. 01 .... 0 0011 . . 0 .... @2misc
++    VCLT0        1111 001 11 . 11 .. 01 .... 0 0100 . . 0 .... @2misc
++
++    VABS         1111 001 11 . 11 .. 01 .... 0 0110 . . 0 .... @2misc
++    VNEG         1111 001 11 . 11 .. 01 .... 0 0111 . . 0 .... @2misc
++
+     VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
+     VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate-neon.inc.c
++++ b/target/arm/translate-neon.inc.c
+@@ -XXX,XX +XXX,XX @@ static bool trans_VCVT_F32_F16(DisasContext *s, arg_2misc *a)
+     return true;
+ }
++
++static bool do_2misc_vec(DisasContext *s, arg_2misc *a, GVecGen2Fn *fn)
++{
++    int vec_size = a->q ? 16 : 8;
++    int rd_ofs = neon_reg_offset(a->vd, 0);
++    int rm_ofs = neon_reg_offset(a->vm, 0);
++
++    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
++        return false;
++    }
++
++    /* UNDEF accesses to D16-D31 if they don't exist. */
++    if (!dc_isar_feature(aa32_simd_r32, s) &&
++        ((a->vd | a->vm) & 0x10)) {
++        return false;
++    }
++
++    if (a->size == 3) {
++        return false;
++    }
++
++    if ((a->vd | a->vm) & a->q) {
++        return false;
++    }
++
++    if (!vfp_access_check(s)) {
++        return true;
++    }
++
++    fn(a->size, rd_ofs, rm_ofs, vec_size, vec_size);
++
++    return true;
++}
++
++#define DO_2MISC_VEC(INSN, FN)                                  \
++    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
++    {                                                           \
++        return do_2misc_vec(s, a, FN);                          \
++    }
++
++DO_2MISC_VEC(VNEG, tcg_gen_gvec_neg)
++DO_2MISC_VEC(VABS, tcg_gen_gvec_abs)
++DO_2MISC_VEC(VCEQ0, gen_gvec_ceq0)
++DO_2MISC_VEC(VCGT0, gen_gvec_cgt0)
++DO_2MISC_VEC(VCLE0, gen_gvec_cle0)
++DO_2MISC_VEC(VCGE0, gen_gvec_cge0)
++DO_2MISC_VEC(VCLT0, gen_gvec_clt0)
++
++static bool trans_VMVN(DisasContext *s, arg_2misc *a)
++{
++    if (a->size != 0) {
++        return false;
++    }
++    return do_2misc_vec(s, a, tcg_gen_gvec_not);
++}
+diff --git a/target/arm/translate.c b/target/arm/translate.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate.c
++++ b/target/arm/translate.c
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+     int size;
+     int pass;
+     int u;
+-    int vec_size;
+     TCGv_i32 tmp, tmp2;
+     if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+     VFP_DREG_D(rd, insn);
+     VFP_DREG_M(rm, insn);
+     size = (insn >> 20) & 3;
+-    vec_size = q ? 16 : 8;
+     rd_ofs = neon_reg_offset(rd, 0);
+     rm_ofs = neon_reg_offset(rm, 0);
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+                 case NEON_2RM_VSHLL:
+                 case NEON_2RM_VCVT_F16_F32:
+                 case NEON_2RM_VCVT_F32_F16:
++                case NEON_2RM_VMVN:
++                case NEON_2RM_VNEG:
++                case NEON_2RM_VABS:
++                case NEON_2RM_VCEQ0:
++                case NEON_2RM_VCGT0:
++                case NEON_2RM_VCLE0:
++                case NEON_2RM_VCGE0:
++                case NEON_2RM_VCLT0:
+                     /* handled by decodetree */
+                     return 1;
+                 case NEON_2RM_VTRN:
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+                                        q ? gen_helper_crypto_sha256su0
+                                        : gen_helper_crypto_sha1su1);
+                     break;
+-                case NEON_2RM_VMVN:
+-                    tcg_gen_gvec_not(0, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VNEG:
+-                    tcg_gen_gvec_neg(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VABS:
+-                    tcg_gen_gvec_abs(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-
+-                case NEON_2RM_VCEQ0:
+-                    gen_gvec_ceq0(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VCGT0:
+-                    gen_gvec_cgt0(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VCLE0:
+-                    gen_gvec_cle0(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VCGE0:
+-                    gen_gvec_cge0(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+-                case NEON_2RM_VCLT0:
+-                    gen_gvec_clt0(size, rd_ofs, rm_ofs, vec_size, vec_size);
+-                    break;
+                 default:
+                 elementwise:
+--
+.20.1

-New patch
+[PULL 10/42] target/arm: Convert Neon 2-reg-misc crypto operations to decodetree
+Convert the Neon-2-reg misc crypto ops (AESE, AESMC, SHA1H, SHA1SU1)
+to decodetree.
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20200616170844.13318-9-peter.maydell@linaro.org
+---
+ target/arm/neon-dp.decode       | 12 ++++++++
+ target/arm/translate-neon.inc.c | 42 ++++++++++++++++++++++++++
+ target/arm/translate.c          | 52 +++------------------------------
+files changed, 58 insertions(+), 48 deletions(-)
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/neon-dp.decode
++++ b/target/arm/neon-dp.decode
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+                  &2misc vm=%vm_dp vd=%vd_dp
+     @2misc_q0    .... ... .. . .. size:2 .. .... . .... . . . .... \
+                  &2misc vm=%vm_dp vd=%vd_dp q=0
++    @2misc_q1    .... ... .. . .. size:2 .. .... . .... . . . .... \
++                 &2misc vm=%vm_dp vd=%vd_dp q=1
+     VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
+     VPADDL_S     1111 001 11 . 11 .. 00 .... 0 0100 . . 0 .... @2misc
+     VPADDL_U     1111 001 11 . 11 .. 00 .... 0 0101 . . 0 .... @2misc
++    AESE         1111 001 11 . 11 .. 00 .... 0 0110 0 . 0 .... @2misc_q1
++    AESD         1111 001 11 . 11 .. 00 .... 0 0110 1 . 0 .... @2misc_q1
++    AESMC        1111 001 11 . 11 .. 00 .... 0 0111 0 . 0 .... @2misc_q1
++    AESIMC       1111 001 11 . 11 .. 00 .... 0 0111 1 . 0 .... @2misc_q1
++
+     VMVN         1111 001 11 . 11 .. 00 .... 0 1011 . . 0 .... @2misc
+     VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+     VCLE0        1111 001 11 . 11 .. 01 .... 0 0011 . . 0 .... @2misc
+     VCLT0        1111 001 11 . 11 .. 01 .... 0 0100 . . 0 .... @2misc
++    SHA1H        1111 001 11 . 11 .. 01 .... 0 0101 1 . 0 .... @2misc_q1
++
+     VABS         1111 001 11 . 11 .. 01 .... 0 0110 . . 0 .... @2misc
+     VNEG         1111 001 11 . 11 .. 01 .... 0 0111 . . 0 .... @2misc
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+     VSHLL        1111 001 11 . 11 .. 10 .... 0 0110 0 . 0 .... @2misc_q0
++    SHA1SU1      1111 001 11 . 11 .. 10 .... 0 0111 0 . 0 .... @2misc_q1
++    SHA256SU0    1111 001 11 . 11 .. 10 .... 0 0111 1 . 0 .... @2misc_q1
++
+     VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
+     VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
+   ]
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate-neon.inc.c
++++ b/target/arm/translate-neon.inc.c
+@@ -XXX,XX +XXX,XX @@ static bool trans_VMVN(DisasContext *s, arg_2misc *a)
+     }
+     return do_2misc_vec(s, a, tcg_gen_gvec_not);
+ }
++
++#define WRAP_2M_3_OOL_FN(WRAPNAME, FUNC, DATA)                          \
++    static void WRAPNAME(unsigned vece, uint32_t rd_ofs,                \
++                         uint32_t rm_ofs, uint32_t oprsz,               \
++                         uint32_t maxsz)                                \
++    {                                                                   \
++        tcg_gen_gvec_3_ool(rd_ofs, rd_ofs, rm_ofs, oprsz, maxsz,        \
++                           DATA, FUNC);                                 \
++    }
++
++#define WRAP_2M_2_OOL_FN(WRAPNAME, FUNC, DATA)                          \
++    static void WRAPNAME(unsigned vece, uint32_t rd_ofs,                \
++                         uint32_t rm_ofs, uint32_t oprsz,               \
++                         uint32_t maxsz)                                \
++    {                                                                   \
++        tcg_gen_gvec_2_ool(rd_ofs, rm_ofs, oprsz, maxsz, DATA, FUNC);   \
++    }
++
++WRAP_2M_3_OOL_FN(gen_AESE, gen_helper_crypto_aese, 0)
++WRAP_2M_3_OOL_FN(gen_AESD, gen_helper_crypto_aese, 1)
++WRAP_2M_2_OOL_FN(gen_AESMC, gen_helper_crypto_aesmc, 0)
++WRAP_2M_2_OOL_FN(gen_AESIMC, gen_helper_crypto_aesmc, 1)
++WRAP_2M_2_OOL_FN(gen_SHA1H, gen_helper_crypto_sha1h, 0)
++WRAP_2M_2_OOL_FN(gen_SHA1SU1, gen_helper_crypto_sha1su1, 0)
++WRAP_2M_2_OOL_FN(gen_SHA256SU0, gen_helper_crypto_sha256su0, 0)
++
++#define DO_2M_CRYPTO(INSN, FEATURE, SIZE)                       \
++    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
++    {                                                           \
++        if (!dc_isar_feature(FEATURE, s) || a->size != SIZE) {  \
++            return false;                                       \
++        }                                                       \
++        return do_2misc_vec(s, a, gen_##INSN);                  \
++    }
++
++DO_2M_CRYPTO(AESE, aa32_aes, 0)
++DO_2M_CRYPTO(AESD, aa32_aes, 0)
++DO_2M_CRYPTO(AESMC, aa32_aes, 0)
++DO_2M_CRYPTO(AESIMC, aa32_aes, 0)
++DO_2M_CRYPTO(SHA1H, aa32_sha1, 2)
++DO_2M_CRYPTO(SHA1SU1, aa32_sha1, 2)
++DO_2M_CRYPTO(SHA256SU0, aa32_sha2, 2)
+diff --git a/target/arm/translate.c b/target/arm/translate.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate.c
++++ b/target/arm/translate.c
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+ {
+     int op;
+     int q;
+-    int rd, rm, rd_ofs, rm_ofs;
++    int rd, rm;
+     int size;
+     int pass;
+     int u;
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+     VFP_DREG_D(rd, insn);
+     VFP_DREG_M(rm, insn);
+     size = (insn >> 20) & 3;
+-    rd_ofs = neon_reg_offset(rd, 0);
+-    rm_ofs = neon_reg_offset(rm, 0);
+     if ((insn & (1 << 23)) == 0) {
+         /* Three register same length: handled by decodetree */
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+                 case NEON_2RM_VCLE0:
+                 case NEON_2RM_VCGE0:
+                 case NEON_2RM_VCLT0:
++                case NEON_2RM_AESE: case NEON_2RM_AESMC:
++                case NEON_2RM_SHA1H:
++                case NEON_2RM_SHA1SU1:
+                     /* handled by decodetree */
+                     return 1;
+                 case NEON_2RM_VTRN:
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
+                         goto elementwise;
+                     }
+                     break;
+-                case NEON_2RM_AESE: case NEON_2RM_AESMC:
+-                    if (!dc_isar_feature(aa32_aes, s) || ((rm | rd) & 1)) {
+-                        return 1;
+-                    }
+-                    /*
+-                     * Bit 6 is the lowest opcode bit; it distinguishes
+-                     * between encryption (AESE/AESMC) and decryption
+-                     * (AESD/AESIMC).
+-                     */
+-                    if (op == NEON_2RM_AESE) {
+-                        tcg_gen_gvec_3_ool(vfp_reg_offset(true, rd),
+-                                           vfp_reg_offset(true, rd),
+-                                           vfp_reg_offset(true, rm),
+-                                           16, 16, extract32(insn, 6, 1),
+-                                           gen_helper_crypto_aese);
+-                    } else {
+-                        tcg_gen_gvec_2_ool(vfp_reg_offset(true, rd),
+-                                           vfp_reg_offset(true, rm),
+-                                           16, 16, extract32(insn, 6, 1),
+-                                           gen_helper_crypto_aesmc);
+-                    }
+-                    break;
+-                case NEON_2RM_SHA1H:
+-                    if (!dc_isar_feature(aa32_sha1, s) || ((rm | rd) & 1)) {
+-                        return 1;
+-                    }
+-                    tcg_gen_gvec_2_ool(rd_ofs, rm_ofs, 16, 16, 0,
+-                                       gen_helper_crypto_sha1h);
+-                    break;
+-                case NEON_2RM_SHA1SU1:
+-                    if ((rm | rd) & 1) {
+-                            return 1;
+-                    }
+-                    /* bit 6 (q): set -> SHA256SU0, cleared -> SHA1SU1 */
+-                    if (q) {
+-                        if (!dc_isar_feature(aa32_sha2, s)) {
+-                            return 1;
+-                        }
+-                    } else if (!dc_isar_feature(aa32_sha1, s)) {
+-                        return 1;
+-                    }
+-                    tcg_gen_gvec_2_ool(rd_ofs, rm_ofs, 16, 16, 0,
+-                                       q ? gen_helper_crypto_sha256su0
+-                                       : gen_helper_crypto_sha1su1);
+-                    break;
+                 default:
+                 elementwise:
+--
+.20.1

-[Qemu-devel] [PULL 23/25] vmstate.h: Provide VMSTATE_BOOL_SUB_ARRAY
+[PULL 11/42] target/arm: Rename NeonGenOneOpFn to NeonGenOne64OpFn
-Provide a VMSTATE_BOOL_SUB_ARRAY to go with VMSTATE_UINT8_SUB_ARRAY
+The NeonGenOneOpFn typedef breaks with the pattern of the other
-and friends.
+NeonGen*Fn typedefs, because it is a TCGv_i64 -> TCGv_i64 operation
 but it does not have '64' in its name. Rename it to NeonGenOne64OpFn,
 so that the old name is available for a TCGv_i32 -> TCGv_i32 operation
 (which we will need in a subsequent commit).
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-23-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-10-peter.maydell@linaro.org
 ---
- include/migration/vmstate.h | 3 +++
+ target/arm/translate.h     | 2 +-
-file changed, 3 insertions(+)
+ target/arm/translate-a64.c | 4 ++--
 files changed, 3 insertions(+), 3 deletions(-)
-diff --git a/include/migration/vmstate.h b/include/migration/vmstate.h
+diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/migration/vmstate.h
+--- a/target/arm/translate.h
-+++ b/include/migration/vmstate.h
++++ b/target/arm/translate.h
-@@ -XXX,XX +XXX,XX @@ extern const VMStateInfo vmstate_info_qtailq;
+@@ -XXX,XX +XXX,XX @@ typedef void NeonGenWidenFn(TCGv_i64, TCGv_i32);
- #define VMSTATE_BOOL_ARRAY(_f, _s, _n)                               \
+ typedef void NeonGenTwoOpWidenFn(TCGv_i64, TCGv_i32, TCGv_i32);
-     VMSTATE_BOOL_ARRAY_V(_f, _s, _n, 0)
+ typedef void NeonGenTwoSingleOPFn(TCGv_i32, TCGv_i32, TCGv_i32, TCGv_ptr);
+ typedef void NeonGenTwoDoubleOPFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGv_ptr);
-+#define VMSTATE_BOOL_SUB_ARRAY(_f, _s, _start, _num)                \
+-typedef void NeonGenOneOpFn(TCGv_i64, TCGv_i64);
-+    VMSTATE_SUB_ARRAY(_f, _s, _start, _num, 0, vmstate_info_bool, bool)
++typedef void NeonGenOne64OpFn(TCGv_i64, TCGv_i64);
-+
+ typedef void CryptoTwoOpFn(TCGv_ptr, TCGv_ptr);
- #define VMSTATE_UINT16_ARRAY_V(_f, _s, _n, _v)                         \
+ typedef void CryptoThreeOpIntFn(TCGv_ptr, TCGv_ptr, TCGv_i32);
-     VMSTATE_ARRAY(_f, _s, _n, _v, vmstate_info_uint16, uint16_t)
+ typedef void CryptoThreeOpFn(TCGv_ptr, TCGv_ptr, TCGv_ptr);
+diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-a64.c
 +++ b/target/arm/translate-a64.c
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_pairwise(DisasContext *s, int opcode, bool u,
      } else {
          for (pass = 0; pass < maxpass; pass++) {
              TCGv_i64 tcg_op = tcg_temp_new_i64();
 -            NeonGenOneOpFn *genfn;
 -            static NeonGenOneOpFn * const fns[2][2] = {
 +            NeonGenOne64OpFn *genfn;
 +            static NeonGenOne64OpFn * const fns[2][2] = {
                  { gen_helper_neon_addlp_s8,  gen_helper_neon_addlp_u8 },
                  { gen_helper_neon_addlp_s16,  gen_helper_neon_addlp_u16 },
              };
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 17/25] Make MemoryRegion valid.accepts callback take a MemTxAttrs argument
+[PULL 12/42] target/arm: Fix capitalization in NeonGenTwo{Single, Double}OPFn typedefs
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+All the other typedefs like these spell "Op" with a lowercase 'p';
-add MemTxAttrs as an argument to the MemoryRegion valid.accepts
+remane the NeonGenTwoSingleOPFn and NeonGenTwoDoubleOPFn typedefs to
-callback. We'll need this for subpage_accepts().
+match.
 We could take the approach we used with the read and write
 callbacks and add new a new _with_attrs version, but since there
 are so few implementations of the accepts hook we just change
 them all.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-9-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-11-peter.maydell@linaro.org
 ---
- include/exec/memory.h |  3 ++-
+ target/arm/translate.h          | 4 ++--
- exec.c                |  9 ++++++---
+ target/arm/translate-a64.c      | 4 ++--
- hw/hppa/dino.c        |  3 ++-
+ target/arm/translate-neon.inc.c | 2 +-
- hw/nvram/fw_cfg.c     | 12 ++++++++----
+files changed, 5 insertions(+), 5 deletions(-)
  hw/scsi/esp.c         |  3 ++-
  hw/xen/xen_pt_msi.c   |  3 ++-
  memory.c              |  5 +++--
 files changed, 25 insertions(+), 13 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/translate.h
-+++ b/include/exec/memory.h
++++ b/target/arm/translate.h
-@@ -XXX,XX +XXX,XX @@ struct MemoryRegionOps {
+@@ -XXX,XX +XXX,XX @@ typedef void NeonGenNarrowFn(TCGv_i32, TCGv_i64);
-          * as a machine check exception).
+ typedef void NeonGenNarrowEnvFn(TCGv_i32, TCGv_ptr, TCGv_i64);
-          */
+ typedef void NeonGenWidenFn(TCGv_i64, TCGv_i32);
-         bool (*accepts)(void *opaque, hwaddr addr,
+ typedef void NeonGenTwoOpWidenFn(TCGv_i64, TCGv_i32, TCGv_i32);
--                        unsigned size, bool is_write);
+-typedef void NeonGenTwoSingleOPFn(TCGv_i32, TCGv_i32, TCGv_i32, TCGv_ptr);
-+                        unsigned size, bool is_write,
+-typedef void NeonGenTwoDoubleOPFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGv_ptr);
-+                        MemTxAttrs attrs);
++typedef void NeonGenTwoSingleOpFn(TCGv_i32, TCGv_i32, TCGv_i32, TCGv_ptr);
-     } valid;
++typedef void NeonGenTwoDoubleOpFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGv_ptr);
-     /* Internal implementation constraints: */
+ typedef void NeonGenOne64OpFn(TCGv_i64, TCGv_i64);
-     struct {
+ typedef void CryptoTwoOpFn(TCGv_ptr, TCGv_ptr);
-diff --git a/exec.c b/exec.c
+ typedef void CryptoThreeOpIntFn(TCGv_ptr, TCGv_ptr, TCGv_i32);
 diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/translate-a64.c
-+++ b/exec.c
++++ b/target/arm/translate-a64.c
-@@ -XXX,XX +XXX,XX @@ static void notdirty_mem_write(void *opaque, hwaddr ram_addr,
+@@ -XXX,XX +XXX,XX @@ static void handle_2misc_fcmp_zero(DisasContext *s, int opcode,
          TCGv_i64 tcg_op = tcg_temp_new_i64();
          TCGv_i64 tcg_zero = tcg_const_i64(0);
          TCGv_i64 tcg_res = tcg_temp_new_i64();
 -        NeonGenTwoDoubleOPFn *genfn;
 +        NeonGenTwoDoubleOpFn *genfn;
          bool swap = false;
          int pass;
@@ -XXX,XX +XXX,XX @@ static void handle_2misc_fcmp_zero(DisasContext *s, int opcode,
          TCGv_i32 tcg_op = tcg_temp_new_i32();
          TCGv_i32 tcg_zero = tcg_const_i32(0);
          TCGv_i32 tcg_res = tcg_temp_new_i32();
 -        NeonGenTwoSingleOPFn *genfn;
 +        NeonGenTwoSingleOpFn *genfn;
          bool swap = false;
          int pass, maxpasses;
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VSHLL_U_2sh(DisasContext *s, arg_2reg_shift *a)
  }
- static bool notdirty_mem_accepts(void *opaque, hwaddr addr,
+ static bool do_fp_2sh(DisasContext *s, arg_2reg_shift *a,
--                                 unsigned size, bool is_write)
+-                      NeonGenTwoSingleOPFn *fn)
-+                                 unsigned size, bool is_write,
++                      NeonGenTwoSingleOpFn *fn)
 +                                 MemTxAttrs attrs)
  {
-     return is_write;
+     /* FP operations in 2-reg-and-shift group */
- }
+     TCGv_i32 tmp, shiftv;
@@ -XXX,XX +XXX,XX @@ static MemTxResult subpage_write(void *opaque, hwaddr addr,
  }
  static bool subpage_accepts(void *opaque, hwaddr addr,
 -                            unsigned len, bool is_write)
 +                            unsigned len, bool is_write,
 +                            MemTxAttrs attrs)
  {
      subpage_t *subpage = opaque;
  #if defined(DEBUG_SUBPAGE)
@@ -XXX,XX +XXX,XX @@ static void readonly_mem_write(void *opaque, hwaddr addr,
  }
  static bool readonly_mem_accepts(void *opaque, hwaddr addr,
 -                                 unsigned size, bool is_write)
 +                                 unsigned size, bool is_write,
 +                                 MemTxAttrs attrs)
  {
      return is_write;
  }
 diff --git a/hw/hppa/dino.c b/hw/hppa/dino.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/hppa/dino.c
 +++ b/hw/hppa/dino.c
@@ -XXX,XX +XXX,XX @@ static void gsc_to_pci_forwarding(DinoState *s)
  }
  static bool dino_chip_mem_valid(void *opaque, hwaddr addr,
 -                                unsigned size, bool is_write)
 +                                unsigned size, bool is_write,
 +                                MemTxAttrs attrs)
  {
      switch (addr) {
      case DINO_IAR0:
 diff --git a/hw/nvram/fw_cfg.c b/hw/nvram/fw_cfg.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/nvram/fw_cfg.c
 +++ b/hw/nvram/fw_cfg.c
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_dma_mem_write(void *opaque, hwaddr addr,
  }
  static bool fw_cfg_dma_mem_valid(void *opaque, hwaddr addr,
 -                                  unsigned size, bool is_write)
 +                                 unsigned size, bool is_write,
 +                                 MemTxAttrs attrs)
  {
      return !is_write || ((size == 4 && (addr == 0 || addr == 4)) ||
                           (size == 8 && addr == 0));
  }
  static bool fw_cfg_data_mem_valid(void *opaque, hwaddr addr,
 -                                  unsigned size, bool is_write)
 +                                  unsigned size, bool is_write,
 +                                  MemTxAttrs attrs)
  {
      return addr == 0;
  }
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_ctl_mem_write(void *opaque, hwaddr addr,
  }
  static bool fw_cfg_ctl_mem_valid(void *opaque, hwaddr addr,
 -                                 unsigned size, bool is_write)
 +                                 unsigned size, bool is_write,
 +                                 MemTxAttrs attrs)
  {
      return is_write && size == 2;
  }
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_comb_write(void *opaque, hwaddr addr,
  }
  static bool fw_cfg_comb_valid(void *opaque, hwaddr addr,
 -                                  unsigned size, bool is_write)
 +                              unsigned size, bool is_write,
 +                              MemTxAttrs attrs)
  {
      return (size == 1) || (is_write && size == 2);
  }
 diff --git a/hw/scsi/esp.c b/hw/scsi/esp.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/scsi/esp.c
 +++ b/hw/scsi/esp.c
@@ -XXX,XX +XXX,XX @@ void esp_reg_write(ESPState *s, uint32_t saddr, uint64_t val)
  }
  static bool esp_mem_accepts(void *opaque, hwaddr addr,
 -                            unsigned size, bool is_write)
 +                            unsigned size, bool is_write,
 +                            MemTxAttrs attrs)
  {
      return (size == 1) || (is_write && size == 4);
  }
 diff --git a/hw/xen/xen_pt_msi.c b/hw/xen/xen_pt_msi.c
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/xen/xen_pt_msi.c
 +++ b/hw/xen/xen_pt_msi.c
@@ -XXX,XX +XXX,XX @@ static uint64_t pci_msix_read(void *opaque, hwaddr addr,
  }
  static bool pci_msix_accepts(void *opaque, hwaddr addr,
 -                             unsigned size, bool is_write)
 +                             unsigned size, bool is_write,
 +                             MemTxAttrs attrs)
  {
      return !(addr & (size - 1));
  }
 diff --git a/memory.c b/memory.c
 index XXXXXXX..XXXXXXX 100644
 --- a/memory.c
 +++ b/memory.c
@@ -XXX,XX +XXX,XX @@ static void unassigned_mem_write(void *opaque, hwaddr addr,
  }
  static bool unassigned_mem_accepts(void *opaque, hwaddr addr,
 -                                   unsigned size, bool is_write)
 +                                   unsigned size, bool is_write,
 +                                   MemTxAttrs attrs)
  {
      return false;
  }
@@ -XXX,XX +XXX,XX @@ bool memory_region_access_valid(MemoryRegion *mr,
      access_size = MAX(MIN(size, access_size_max), access_size_min);
      for (i = 0; i < size; i += access_size) {
          if (!mr->ops->valid.accepts(mr->opaque, addr + i, access_size,
 -                                    is_write)) {
 +                                    is_write, attrs)) {
              return false;
          }
      }
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 13/25] Make address_space_map() take a MemTxAttrs argument
+[PULL 13/42] target/arm: Make gen_swap_half() take separate src and dest
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Make gen_swap_half() take a source and destination TCGv_i32 rather
-add MemTxAttrs as an argument to address_space_map().
+than modifying the input TCGv_i32; we're going to want to be able to
-Its callers either have an attrs value to hand, or don't care
+use it with the more flexible function signature, and this also
-and can use MEMTXATTRS_UNSPECIFIED.
+brings it into line with other functions like gen_rev16() and
 gen_revsh().
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-5-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-12-peter.maydell@linaro.org
 ---
- include/exec/memory.h   | 3 ++-
+ target/arm/translate-neon.inc.c |  2 +-
- include/sysemu/dma.h    | 3 ++-
+ target/arm/translate.c          | 10 +++++-----
- exec.c                  | 6 ++++--
+files changed, 6 insertions(+), 6 deletions(-)
  target/ppc/mmu-hash64.c | 3 ++-
 files changed, 10 insertions(+), 5 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/translate-neon.inc.c
-+++ b/include/exec/memory.h
++++ b/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len, bool is_
+@@ -XXX,XX +XXX,XX @@ static bool trans_VREV64(DisasContext *s, arg_VREV64 *a)
-  * @addr: address within that address space
+                 tcg_gen_bswap32_i32(tmp[half], tmp[half]);
-  * @plen: pointer to length of buffer; updated on return
+                 break;
-  * @is_write: indicates the transfer direction
+             case 1:
-+ * @attrs: memory attributes
+-                gen_swap_half(tmp[half]);
-  */
++                gen_swap_half(tmp[half], tmp[half]);
- void *address_space_map(AddressSpace *as, hwaddr addr,
+                 break;
--                        hwaddr *plen, bool is_write);
+             case 2:
-+                        hwaddr *plen, bool is_write, MemTxAttrs attrs);
+                 break;
+diff --git a/target/arm/translate.c b/target/arm/translate.c
  /* address_space_unmap: Unmaps a memory region previously mapped by address_space_map()
   *
 diff --git a/include/sysemu/dma.h b/include/sysemu/dma.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/sysemu/dma.h
+--- a/target/arm/translate.c
-+++ b/include/sysemu/dma.h
++++ b/target/arm/translate.c
-@@ -XXX,XX +XXX,XX @@ static inline void *dma_memory_map(AddressSpace *as,
+@@ -XXX,XX +XXX,XX @@ static void gen_revsh(TCGv_i32 dest, TCGv_i32 var)
      hwaddr xlen = *len;
      void *p;
 -    p = address_space_map(as, addr, &xlen, dir == DMA_DIRECTION_FROM_DEVICE);
 +    p = address_space_map(as, addr, &xlen, dir == DMA_DIRECTION_FROM_DEVICE,
 +                          MEMTXATTRS_UNSPECIFIED);
      *len = xlen;
      return p;
  }
-diff --git a/exec.c b/exec.c
-index XXXXXXX..XXXXXXX 100644
+ /* Swap low and high halfwords.  */
---- a/exec.c
+-static void gen_swap_half(TCGv_i32 var)
-+++ b/exec.c
++static void gen_swap_half(TCGv_i32 dest, TCGv_i32 var)
@@ -XXX,XX +XXX,XX @@ flatview_extend_translation(FlatView *fv, hwaddr addr,
  void *address_space_map(AddressSpace *as,
                          hwaddr addr,
                          hwaddr *plen,
 -                        bool is_write)
 +                        bool is_write,
 +                        MemTxAttrs attrs)
  {
-     hwaddr len = *plen;
+-    tcg_gen_rotri_i32(var, var, 16);
-     hwaddr l, xlat;
++    tcg_gen_rotri_i32(dest, var, 16);
@@ -XXX,XX +XXX,XX @@ void *cpu_physical_memory_map(hwaddr addr,
                                hwaddr *plen,
                                int is_write)
  {
 -    return address_space_map(&address_space_memory, addr, plen, is_write);
 +    return address_space_map(&address_space_memory, addr, plen, is_write,
 +                             MEMTXATTRS_UNSPECIFIED);
  }
- void cpu_physical_memory_unmap(void *buffer, hwaddr len,
+ /* Dual 16-bit add.  Result placed in t0 and t1 is marked as dead.
-diff --git a/target/ppc/mmu-hash64.c b/target/ppc/mmu-hash64.c
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-index XXXXXXX..XXXXXXX 100644
+                         case NEON_2RM_VREV32:
---- a/target/ppc/mmu-hash64.c
+                             switch (size) {
-+++ b/target/ppc/mmu-hash64.c
+                             case 0: tcg_gen_bswap32_i32(tmp, tmp); break;
-@@ -XXX,XX +XXX,XX @@ const ppc_hash_pte64_t *ppc_hash64_map_hptes(PowerPCCPU *cpu,
+-                            case 1: gen_swap_half(tmp); break;
-         return NULL;
++                            case 1: gen_swap_half(tmp, tmp); break;
                              default: abort();
                              }
                              break;
@@ -XXX,XX +XXX,XX @@ static bool op_smlad(DisasContext *s, arg_rrrr *a, bool m_swap, bool sub)
      t1 = load_reg(s, a->rn);
      t2 = load_reg(s, a->rm);
      if (m_swap) {
 -        gen_swap_half(t2);
 +        gen_swap_half(t2, t2);
      }
+     gen_smul_dual(t1, t2);
--    hptes = address_space_map(CPU(cpu)->as, base + pte_offset, &plen, false);
-+    hptes = address_space_map(CPU(cpu)->as, base + pte_offset, &plen, false,
+@@ -XXX,XX +XXX,XX @@ static bool op_smlald(DisasContext *s, arg_rrrr *a, bool m_swap, bool sub)
-+                              MEMTXATTRS_UNSPECIFIED);
+     t1 = load_reg(s, a->rn);
-     if (plen < (n * HASH_PTE_SIZE_64)) {
+     t2 = load_reg(s, a->rm);
-         hw_error("%s: Unable to map all requested HPTEs\n", __func__);
+     if (m_swap) {
 -        gen_swap_half(t2);
 +        gen_swap_half(t2, t2);
      }
+     gen_smul_dual(t1, t2);
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 21/25] Make flatview_do_translate() take a MemTxAttrs argument
+[PULL 14/42] target/arm: Convert Neon 2-reg-misc VREV32 and VREV16 to decodetree
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the VREV32 and VREV16 insns in the Neon 2-reg-misc group
-add MemTxAttrs as an argument to flatview_do_translate().
+to decodetree.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-13-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-13-peter.maydell@linaro.org
 ---
- exec.c | 9 ++++++---
+ target/arm/translate.h          |  1 +
-file changed, 6 insertions(+), 3 deletions(-)
+ target/arm/neon-dp.decode       |  2 ++
  target/arm/translate-neon.inc.c | 55 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 12 ++-----
 files changed, 60 insertions(+), 10 deletions(-)
-diff --git a/exec.c b/exec.c
+diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/translate.h
-+++ b/exec.c
++++ b/target/arm/translate.h
-@@ -XXX,XX +XXX,XX @@ unassigned:
+@@ -XXX,XX +XXX,XX @@ typedef void GVecGen4Fn(unsigned, uint32_t, uint32_t, uint32_t,
-  * @is_write: whether the translation operation is for write
+                         uint32_t, uint32_t, uint32_t);
-  * @is_mmio: whether this can be MMIO, set true if it can
-  * @target_as: the address space targeted by the IOMMU
+ /* Function prototype for gen_ functions for calling Neon helpers */
-+ * @attrs: memory transaction attributes
++typedef void NeonGenOneOpFn(TCGv_i32, TCGv_i32);
-  *
+ typedef void NeonGenOneOpEnvFn(TCGv_i32, TCGv_ptr, TCGv_i32);
-  * This function is called from RCU critical section
+ typedef void NeonGenTwoOpFn(TCGv_i32, TCGv_i32, TCGv_i32);
-  */
+ typedef void NeonGenTwoOpEnvFn(TCGv_i32, TCGv_ptr, TCGv_i32, TCGv_i32);
-@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
-                                                  hwaddr *page_mask_out,
+index XXXXXXX..XXXXXXX 100644
-                                                  bool is_write,
+--- a/target/arm/neon-dp.decode
-                                                  bool is_mmio,
++++ b/target/arm/neon-dp.decode
--                                                 AddressSpace **target_as)
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
-+                                                 AddressSpace **target_as,
+                  &2misc vm=%vm_dp vd=%vd_dp q=1
-+                                                 MemTxAttrs attrs)
- {
+     VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
-     MemoryRegionSection *section;
++    VREV32       1111 001 11 . 11 .. 00 .... 0 0001 . . 0 .... @2misc
-     IOMMUMemoryRegion *iommu_mr;
++    VREV16       1111 001 11 . 11 .. 00 .... 0 0010 . . 0 .... @2misc
-@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
-      * but page mask.
+     VPADDL_S     1111 001 11 . 11 .. 00 .... 0 0100 . . 0 .... @2misc
-      */
+     VPADDL_U     1111 001 11 . 11 .. 00 .... 0 0101 . . 0 .... @2misc
-     section = flatview_do_translate(address_space_to_flatview(as), addr, &xlat,
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
--                                    NULL, &page_mask, is_write, false, &as);
+index XXXXXXX..XXXXXXX 100644
-+                                    NULL, &page_mask, is_write, false, &as,
+--- a/target/arm/translate-neon.inc.c
-+                                    attrs);
++++ b/target/arm/translate-neon.inc.c
+@@ -XXX,XX +XXX,XX @@ DO_2M_CRYPTO(AESIMC, aa32_aes, 0)
-     /* Illegal translation */
+ DO_2M_CRYPTO(SHA1H, aa32_sha1, 2)
-     if (section.mr == &io_mem_unassigned) {
+ DO_2M_CRYPTO(SHA1SU1, aa32_sha1, 2)
-@@ -XXX,XX +XXX,XX @@ MemoryRegion *flatview_translate(FlatView *fv, hwaddr addr, hwaddr *xlat,
+ DO_2M_CRYPTO(SHA256SU0, aa32_sha2, 2)
++
-     /* This can be MMIO, so setup MMIO bit. */
++static bool do_2misc(DisasContext *s, arg_2misc *a, NeonGenOneOpFn *fn)
-     section = flatview_do_translate(fv, addr, xlat, plen, NULL,
++{
--                                    is_write, true, &as);
++    int pass;
-+                                    is_write, true, &as, attrs);
++
-     mr = section.mr;
++    /* Handle a 2-reg-misc operation by iterating 32 bits at a time */
++    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
-     if (xen_enabled() && memory_access_is_direct(mr, is_write)) {
++        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (!fn) {
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
 +        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
 +        fn(tmp, tmp);
 +        neon_store_reg(a->vd, pass, tmp);
 +    }
 +
 +    return true;
 +}
 +
 +static bool trans_VREV32(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenOneOpFn * const fn[] = {
 +        tcg_gen_bswap32_i32,
 +        gen_swap_half,
 +        NULL,
 +        NULL,
 +    };
 +    return do_2misc(s, a, fn[a->size]);
 +}
 +
 +static bool trans_VREV16(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 0) {
 +        return false;
 +    }
 +    return do_2misc(s, a, gen_rev16);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_AESE: case NEON_2RM_AESMC:
                  case NEON_2RM_SHA1H:
                  case NEON_2RM_SHA1SU1:
 +                case NEON_2RM_VREV32:
 +                case NEON_2RM_VREV16:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                      for (pass = 0; pass < (q ? 4 : 2); pass++) {
                          tmp = neon_load_reg(rm, pass);
                          switch (op) {
 -                        case NEON_2RM_VREV32:
 -                            switch (size) {
 -                            case 0: tcg_gen_bswap32_i32(tmp, tmp); break;
 -                            case 1: gen_swap_half(tmp, tmp); break;
 -                            default: abort();
 -                            }
 -                            break;
 -                        case NEON_2RM_VREV16:
 -                            gen_rev16(tmp, tmp);
 -                            break;
                          case NEON_2RM_VCLS:
                              switch (size) {
                              case 0: gen_helper_neon_cls_s8(tmp, tmp); break;
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 16/25] Make memory_region_access_valid() take a MemTxAttrs argument
+[PULL 15/42] target/arm: Convert remaining simple 2-reg-misc Neon ops
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the remaining ops in the Neon 2-reg-misc group which
-add MemTxAttrs as an argument to memory_region_access_valid().
+can be implemented simply with our do_2misc() helper.
 Its callers either have an attrs value to hand, or don't care
 and can use MEMTXATTRS_UNSPECIFIED.
 The callsite in flatview_access_valid() is part of a recursive
 loop flatview_access_valid() -> memory_region_access_valid() ->
  subpage_accepts() -> flatview_access_valid(); we make it pass
 MEMTXATTRS_UNSPECIFIED for now, until the next several commits
 have plumbed an attrs parameter through the rest of the loop
 and we can add an attrs parameter to flatview_access_valid().
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-8-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-14-peter.maydell@linaro.org
 ---
- include/exec/memory-internal.h | 3 ++-
+ target/arm/neon-dp.decode       | 10 +++++
- exec.c                         | 4 +++-
+ target/arm/translate-neon.inc.c | 69 +++++++++++++++++++++++++++++++++
- hw/s390x/s390-pci-inst.c       | 3 ++-
+ target/arm/translate.c          | 38 ++++--------------
- memory.c                       | 7 ++++---
+files changed, 86 insertions(+), 31 deletions(-)
 files changed, 11 insertions(+), 6 deletions(-)
-diff --git a/include/exec/memory-internal.h b/include/exec/memory-internal.h
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory-internal.h
+--- a/target/arm/neon-dp.decode
-+++ b/include/exec/memory-internal.h
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ void flatview_unref(FlatView *view);
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
- extern const MemoryRegionOps unassigned_mem_ops;
+     AESMC        1111 001 11 . 11 .. 00 .... 0 0111 0 . 0 .... @2misc_q1
+     AESIMC       1111 001 11 . 11 .. 00 .... 0 0111 1 . 0 .... @2misc_q1
- bool memory_region_access_valid(MemoryRegion *mr, hwaddr addr,
--                                unsigned size, bool is_write);
++    VCLS         1111 001 11 . 11 .. 00 .... 0 1000 . . 0 .... @2misc
-+                                unsigned size, bool is_write,
++    VCLZ         1111 001 11 . 11 .. 00 .... 0 1001 . . 0 .... @2misc
-+                                MemTxAttrs attrs);
++    VCNT         1111 001 11 . 11 .. 00 .... 0 1010 . . 0 .... @2misc
++
- void flatview_add_to_dispatch(FlatView *fv, MemoryRegionSection *section);
+     VMVN         1111 001 11 . 11 .. 00 .... 0 1011 . . 0 .... @2misc
- AddressSpaceDispatch *address_space_dispatch_new(FlatView *fv);
-diff --git a/exec.c b/exec.c
+     VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VABS         1111 001 11 . 11 .. 01 .... 0 0110 . . 0 .... @2misc
      VNEG         1111 001 11 . 11 .. 01 .... 0 0111 . . 0 .... @2misc
 +    VABS_F       1111 001 11 . 11 .. 01 .... 0 1110 . . 0 .... @2misc
 +    VNEG_F       1111 001 11 . 11 .. 01 .... 0 1111 . . 0 .... @2misc
 +
      VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
      VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
      VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
 +
 +    VRECPE       1111 001 11 . 11 .. 11 .... 0 1000 . . 0 .... @2misc
 +    VRSQRTE      1111 001 11 . 11 .. 11 .... 0 1001 . . 0 .... @2misc
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/translate-neon.inc.c
-+++ b/exec.c
++++ b/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
+@@ -XXX,XX +XXX,XX @@ static bool trans_VREV16(DisasContext *s, arg_2misc *a)
-         mr = flatview_translate(fv, addr, &xlat, &l, is_write);
+     }
-         if (!memory_access_is_direct(mr, is_write)) {
+     return do_2misc(s, a, gen_rev16);
-             l = memory_access_size(mr, l, addr);
+ }
--            if (!memory_region_access_valid(mr, xlat, l, is_write)) {
++
-+            /* When our callers all have attrs we'll pass them through here */
++static bool trans_VCLS(DisasContext *s, arg_2misc *a)
-+            if (!memory_region_access_valid(mr, xlat, l, is_write,
++{
-+                                            MEMTXATTRS_UNSPECIFIED)) {
++    static NeonGenOneOpFn * const fn[] = {
-                 return false;
++        gen_helper_neon_cls_s8,
-             }
++        gen_helper_neon_cls_s16,
-         }
++        gen_helper_neon_cls_s32,
-diff --git a/hw/s390x/s390-pci-inst.c b/hw/s390x/s390-pci-inst.c
++        NULL,
 +    };
 +    return do_2misc(s, a, fn[a->size]);
 +}
 +
 +static void do_VCLZ_32(TCGv_i32 rd, TCGv_i32 rm)
 +{
 +    tcg_gen_clzi_i32(rd, rm, 32);
 +}
 +
 +static bool trans_VCLZ(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenOneOpFn * const fn[] = {
 +        gen_helper_neon_clz_u8,
 +        gen_helper_neon_clz_u16,
 +        do_VCLZ_32,
 +        NULL,
 +    };
 +    return do_2misc(s, a, fn[a->size]);
 +}
 +
 +static bool trans_VCNT(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 0) {
 +        return false;
 +    }
 +    return do_2misc(s, a, gen_helper_neon_cnt_u8);
 +}
 +
 +static bool trans_VABS_F(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 2) {
 +        return false;
 +    }
 +    /* TODO: FP16 : size == 1 */
 +    return do_2misc(s, a, gen_helper_vfp_abss);
 +}
 +
 +static bool trans_VNEG_F(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 2) {
 +        return false;
 +    }
 +    /* TODO: FP16 : size == 1 */
 +    return do_2misc(s, a, gen_helper_vfp_negs);
 +}
 +
 +static bool trans_VRECPE(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 2) {
 +        return false;
 +    }
 +    return do_2misc(s, a, gen_helper_recpe_u32);
 +}
 +
 +static bool trans_VRSQRTE(DisasContext *s, arg_2misc *a)
 +{
 +    if (a->size != 2) {
 +        return false;
 +    }
 +    return do_2misc(s, a, gen_helper_rsqrte_u32);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/s390x/s390-pci-inst.c
+--- a/target/arm/translate.c
-+++ b/hw/s390x/s390-pci-inst.c
++++ b/target/arm/translate.c
-@@ -XXX,XX +XXX,XX @@ int pcistb_service_call(S390CPU *cpu, uint8_t r1, uint8_t r3, uint64_t gaddr,
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-     mr = s390_get_subregion(mr, offset, len);
+                 case NEON_2RM_SHA1SU1:
-     offset -= mr->addr;
+                 case NEON_2RM_VREV32:
+                 case NEON_2RM_VREV16:
--    if (!memory_region_access_valid(mr, offset, len, true)) {
++                case NEON_2RM_VCLS:
-+    if (!memory_region_access_valid(mr, offset, len, true,
++                case NEON_2RM_VCLZ:
-+                                    MEMTXATTRS_UNSPECIFIED)) {
++                case NEON_2RM_VCNT:
-         s390_program_interrupt(env, PGM_OPERAND, 6, ra);
++                case NEON_2RM_VABS_F:
-         return 0;
++                case NEON_2RM_VNEG_F:
-     }
++                case NEON_2RM_VRECPE:
-diff --git a/memory.c b/memory.c
++                case NEON_2RM_VRSQRTE:
-index XXXXXXX..XXXXXXX 100644
+                     /* handled by decodetree */
---- a/memory.c
+                     return 1;
-+++ b/memory.c
+                 case NEON_2RM_VTRN:
-@@ -XXX,XX +XXX,XX @@ static const MemoryRegionOps ram_device_mem_ops = {
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
- bool memory_region_access_valid(MemoryRegion *mr,
+                     for (pass = 0; pass < (q ? 4 : 2); pass++) {
-                                 hwaddr addr,
+                         tmp = neon_load_reg(rm, pass);
-                                 unsigned size,
+                         switch (op) {
--                                bool is_write)
+-                        case NEON_2RM_VCLS:
-+                                bool is_write,
+-                            switch (size) {
-+                                MemTxAttrs attrs)
+-                            case 0: gen_helper_neon_cls_s8(tmp, tmp); break;
- {
+-                            case 1: gen_helper_neon_cls_s16(tmp, tmp); break;
-     int access_size_min, access_size_max;
+-                            case 2: gen_helper_neon_cls_s32(tmp, tmp); break;
-     int access_size, i;
+-                            default: abort();
-@@ -XXX,XX +XXX,XX @@ MemTxResult memory_region_dispatch_read(MemoryRegion *mr,
+-                            }
- {
+-                            break;
-     MemTxResult r;
+-                        case NEON_2RM_VCLZ:
+-                            switch (size) {
--    if (!memory_region_access_valid(mr, addr, size, false)) {
+-                            case 0: gen_helper_neon_clz_u8(tmp, tmp); break;
-+    if (!memory_region_access_valid(mr, addr, size, false, attrs)) {
+-                            case 1: gen_helper_neon_clz_u16(tmp, tmp); break;
-         *pval = unassigned_mem_read(mr, addr, size);
+-                            case 2: tcg_gen_clzi_i32(tmp, tmp, 32); break;
-         return MEMTX_DECODE_ERROR;
+-                            default: abort();
-     }
+-                            }
-@@ -XXX,XX +XXX,XX @@ MemTxResult memory_region_dispatch_write(MemoryRegion *mr,
+-                            break;
-                                          unsigned size,
+-                        case NEON_2RM_VCNT:
-                                          MemTxAttrs attrs)
+-                            gen_helper_neon_cnt_u8(tmp, tmp);
- {
+-                            break;
--    if (!memory_region_access_valid(mr, addr, size, true)) {
+                         case NEON_2RM_VQABS:
-+    if (!memory_region_access_valid(mr, addr, size, true, attrs)) {
+                             switch (size) {
-         unassigned_mem_write(mr, addr, data, size);
+                             case 0:
-         return MEMTX_DECODE_ERROR;
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-     }
+                             tcg_temp_free_ptr(fpstatus);
                              break;
                          }
 -                        case NEON_2RM_VABS_F:
 -                            gen_helper_vfp_abss(tmp, tmp);
 -                            break;
 -                        case NEON_2RM_VNEG_F:
 -                            gen_helper_vfp_negs(tmp, tmp);
 -                            break;
                          case NEON_2RM_VSWP:
                              tmp2 = neon_load_reg(rd, pass);
                              neon_store_reg(rm, pass, tmp2);
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                              tcg_temp_free_ptr(fpst);
                              break;
                          }
 -                        case NEON_2RM_VRECPE:
 -                            gen_helper_recpe_u32(tmp, tmp);
 -                            break;
 -                        case NEON_2RM_VRSQRTE:
 -                            gen_helper_rsqrte_u32(tmp, tmp);
 -                            break;
                          case NEON_2RM_VRECPE_F:
                          {
                              TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 15/25] Make flatview_extend_translation() take a MemTxAttrs argument
+[PULL 16/42] target/arm: Convert Neon VQABS, VQNEG to decodetree
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the Neon VQABS and VQNEG insns to decodetree.
-add MemTxAttrs as an argument to flatview_extend_translation().
+Since these are the only ones which need cpu_env passing to
-Its callers either have an attrs value to hand, or don't care
+the helper, we wrap the helper rather than creating a whole
-and can use MEMTXATTRS_UNSPECIFIED.
+new do_2misc_env() function.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-7-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-15-peter.maydell@linaro.org
 ---
- exec.c | 15 ++++++++++-----
+ target/arm/neon-dp.decode       |  3 +++
-file changed, 10 insertions(+), 5 deletions(-)
+ target/arm/translate-neon.inc.c | 35 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 30 ++--------------------------
 files changed, 40 insertions(+), 28 deletions(-)
-diff --git a/exec.c b/exec.c
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/neon-dp.decode
-+++ b/exec.c
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
+     VPADAL_S     1111 001 11 . 11 .. 00 .... 0 1100 . . 0 .... @2misc
- static hwaddr
+     VPADAL_U     1111 001 11 . 11 .. 00 .... 0 1101 . . 0 .... @2misc
- flatview_extend_translation(FlatView *fv, hwaddr addr,
--                                 hwaddr target_len,
++    VQABS        1111 001 11 . 11 .. 00 .... 0 1110 . . 0 .... @2misc
--                                 MemoryRegion *mr, hwaddr base, hwaddr len,
++    VQNEG        1111 001 11 . 11 .. 00 .... 0 1111 . . 0 .... @2misc
--                                 bool is_write)
++
-+                            hwaddr target_len,
+     VCGT0        1111 001 11 . 11 .. 01 .... 0 0000 . . 0 .... @2misc
-+                            MemoryRegion *mr, hwaddr base, hwaddr len,
+     VCGE0        1111 001 11 . 11 .. 01 .... 0 0001 . . 0 .... @2misc
-+                            bool is_write, MemTxAttrs attrs)
+     VCEQ0        1111 001 11 . 11 .. 01 .... 0 0010 . . 0 .... @2misc
- {
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
-     hwaddr done = 0;
+index XXXXXXX..XXXXXXX 100644
-     hwaddr xlat;
+--- a/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ void *address_space_map(AddressSpace *as,
++++ b/target/arm/translate-neon.inc.c
+@@ -XXX,XX +XXX,XX @@ static bool trans_VRSQRTE(DisasContext *s, arg_2misc *a)
-     memory_region_ref(mr);
+     }
-     *plen = flatview_extend_translation(fv, addr, len, mr, xlat,
+     return do_2misc(s, a, gen_helper_rsqrte_u32);
--                                             l, is_write);
+ }
-+                                        l, is_write, attrs);
++
-     ptr = qemu_ram_ptr_length(mr->ram_block, xlat, plen, true);
++#define WRAP_1OP_ENV_FN(WRAPNAME, FUNC) \
-     rcu_read_unlock();
++    static void WRAPNAME(TCGv_i32 d, TCGv_i32 m)        \
++    {                                                   \
-@@ -XXX,XX +XXX,XX @@ int64_t address_space_cache_init(MemoryRegionCache *cache,
++        FUNC(d, cpu_env, m);                            \
-     mr = cache->mrs.mr;
++    }
-     memory_region_ref(mr);
++
-     if (memory_access_is_direct(mr, is_write)) {
++WRAP_1OP_ENV_FN(gen_VQABS_s8, gen_helper_neon_qabs_s8)
-+        /* We don't care about the memory attributes here as we're only
++WRAP_1OP_ENV_FN(gen_VQABS_s16, gen_helper_neon_qabs_s16)
-+         * doing this if we found actual RAM, which behaves the same
++WRAP_1OP_ENV_FN(gen_VQABS_s32, gen_helper_neon_qabs_s32)
-+         * regardless of attributes; so UNSPECIFIED is fine.
++WRAP_1OP_ENV_FN(gen_VQNEG_s8, gen_helper_neon_qneg_s8)
-+         */
++WRAP_1OP_ENV_FN(gen_VQNEG_s16, gen_helper_neon_qneg_s16)
-         l = flatview_extend_translation(cache->fv, addr, len, mr,
++WRAP_1OP_ENV_FN(gen_VQNEG_s32, gen_helper_neon_qneg_s32)
--                                        cache->xlat, l, is_write);
++
-+                                        cache->xlat, l, is_write,
++static bool trans_VQABS(DisasContext *s, arg_2misc *a)
-+                                        MEMTXATTRS_UNSPECIFIED);
++{
-         cache->ptr = qemu_ram_ptr_length(mr->ram_block, cache->xlat, &l, true);
++    static NeonGenOneOpFn * const fn[] = {
-     } else {
++        gen_VQABS_s8,
-         cache->ptr = NULL;
++        gen_VQABS_s16,
 +        gen_VQABS_s32,
 +        NULL,
 +    };
 +    return do_2misc(s, a, fn[a->size]);
 +}
 +
 +static bool trans_VQNEG(DisasContext *s, arg_2misc *a)
 +{
 +    static NeonGenOneOpFn * const fn[] = {
 +        gen_VQNEG_s8,
 +        gen_VQNEG_s16,
 +        gen_VQNEG_s32,
 +        NULL,
 +    };
 +    return do_2misc(s, a, fn[a->size]);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VNEG_F:
                  case NEON_2RM_VRECPE:
                  case NEON_2RM_VRSQRTE:
 +                case NEON_2RM_VQABS:
 +                case NEON_2RM_VQNEG:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                      for (pass = 0; pass < (q ? 4 : 2); pass++) {
                          tmp = neon_load_reg(rm, pass);
                          switch (op) {
 -                        case NEON_2RM_VQABS:
 -                            switch (size) {
 -                            case 0:
 -                                gen_helper_neon_qabs_s8(tmp, cpu_env, tmp);
 -                                break;
 -                            case 1:
 -                                gen_helper_neon_qabs_s16(tmp, cpu_env, tmp);
 -                                break;
 -                            case 2:
 -                                gen_helper_neon_qabs_s32(tmp, cpu_env, tmp);
 -                                break;
 -                            default: abort();
 -                            }
 -                            break;
 -                        case NEON_2RM_VQNEG:
 -                            switch (size) {
 -                            case 0:
 -                                gen_helper_neon_qneg_s8(tmp, cpu_env, tmp);
 -                                break;
 -                            case 1:
 -                                gen_helper_neon_qneg_s16(tmp, cpu_env, tmp);
 -                                break;
 -                            case 2:
 -                                gen_helper_neon_qneg_s32(tmp, cpu_env, tmp);
 -                                break;
 -                            default: abort();
 -                            }
 -                            break;
                          case NEON_2RM_VCGT0_F:
                          {
                              TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 --
-.17.1
+.20.1

-New patch
+[PULL 17/42] target/arm: Convert simple fp Neon 2-reg-misc insns
+Convert the Neon 2-reg-misc insns which are implemented with
 simple calls to functions that take the input, output and
 fpstatus pointer.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-16-peter.maydell@linaro.org
 ---
  target/arm/translate.h          |  1 +
  target/arm/neon-dp.decode       |  8 +++++
  target/arm/translate-neon.inc.c | 62 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 56 ++++-------------------------
 files changed, 78 insertions(+), 49 deletions(-)
 diff --git a/target/arm/translate.h b/target/arm/translate.h
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.h
 +++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ typedef void NeonGenNarrowFn(TCGv_i32, TCGv_i64);
  typedef void NeonGenNarrowEnvFn(TCGv_i32, TCGv_ptr, TCGv_i64);
  typedef void NeonGenWidenFn(TCGv_i64, TCGv_i32);
  typedef void NeonGenTwoOpWidenFn(TCGv_i64, TCGv_i32, TCGv_i32);
 +typedef void NeonGenOneSingleOpFn(TCGv_i32, TCGv_i32, TCGv_ptr);
  typedef void NeonGenTwoSingleOpFn(TCGv_i32, TCGv_i32, TCGv_i32, TCGv_ptr);
  typedef void NeonGenTwoDoubleOpFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGv_ptr);
  typedef void NeonGenOne64OpFn(TCGv_i64, TCGv_i64);
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      SHA1SU1      1111 001 11 . 11 .. 10 .... 0 0111 0 . 0 .... @2misc_q1
      SHA256SU0    1111 001 11 . 11 .. 10 .... 0 0111 1 . 0 .... @2misc_q1
 +    VRINTX       1111 001 11 . 11 .. 10 .... 0 1001 . . 0 .... @2misc
 +
      VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
      VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
      VRECPE       1111 001 11 . 11 .. 11 .... 0 1000 . . 0 .... @2misc
      VRSQRTE      1111 001 11 . 11 .. 11 .... 0 1001 . . 0 .... @2misc
 +    VRECPE_F     1111 001 11 . 11 .. 11 .... 0 1010 . . 0 .... @2misc
 +    VRSQRTE_F    1111 001 11 . 11 .. 11 .... 0 1011 . . 0 .... @2misc
 +    VCVT_FS      1111 001 11 . 11 .. 11 .... 0 1100 . . 0 .... @2misc
 +    VCVT_FU      1111 001 11 . 11 .. 11 .... 0 1101 . . 0 .... @2misc
 +    VCVT_SF      1111 001 11 . 11 .. 11 .... 0 1110 . . 0 .... @2misc
 +    VCVT_UF      1111 001 11 . 11 .. 11 .... 0 1111 . . 0 .... @2misc
    ]
    # Subgroup for size != 0b11
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VQNEG(DisasContext *s, arg_2misc *a)
      };
      return do_2misc(s, a, fn[a->size]);
  }
 +
 +static bool do_2misc_fp(DisasContext *s, arg_2misc *a,
 +                        NeonGenOneSingleOpFn *fn)
 +{
 +    int pass;
 +    TCGv_ptr fpst;
 +
 +    /* Handle a 2-reg-misc operation by iterating 32 bits at a time */
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (a->size != 2) {
 +        /* TODO: FP16 will be the size == 1 case */
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    fpst = get_fpstatus_ptr(1);
 +    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
 +        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
 +        fn(tmp, tmp, fpst);
 +        neon_store_reg(a->vd, pass, tmp);
 +    }
 +    tcg_temp_free_ptr(fpst);
 +
 +    return true;
 +}
 +
 +#define DO_2MISC_FP(INSN, FUNC)                                 \
 +    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
 +    {                                                           \
 +        return do_2misc_fp(s, a, FUNC);                         \
 +    }
 +
 +DO_2MISC_FP(VRECPE_F, gen_helper_recpe_f32)
 +DO_2MISC_FP(VRSQRTE_F, gen_helper_rsqrte_f32)
 +DO_2MISC_FP(VCVT_FS, gen_helper_vfp_sitos)
 +DO_2MISC_FP(VCVT_FU, gen_helper_vfp_uitos)
 +DO_2MISC_FP(VCVT_SF, gen_helper_vfp_tosizs)
 +DO_2MISC_FP(VCVT_UF, gen_helper_vfp_touizs)
 +
 +static bool trans_VRINTX(DisasContext *s, arg_2misc *a)
 +{
 +    if (!arm_dc_feature(s, ARM_FEATURE_V8)) {
 +        return false;
 +    }
 +    return do_2misc_fp(s, a, gen_helper_rints_exact);
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VRSQRTE:
                  case NEON_2RM_VQABS:
                  case NEON_2RM_VQNEG:
 +                case NEON_2RM_VRECPE_F:
 +                case NEON_2RM_VRSQRTE_F:
 +                case NEON_2RM_VCVT_FS:
 +                case NEON_2RM_VCVT_FU:
 +                case NEON_2RM_VCVT_SF:
 +                case NEON_2RM_VCVT_UF:
 +                case NEON_2RM_VRINTX:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                              tcg_temp_free_i32(tcg_rmode);
                              break;
                          }
 -                        case NEON_2RM_VRINTX:
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_rints_exact(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
                          case NEON_2RM_VCVTAU:
                          case NEON_2RM_VCVTAS:
                          case NEON_2RM_VCVTNU:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                              tcg_temp_free_ptr(fpst);
                              break;
                          }
 -                        case NEON_2RM_VRECPE_F:
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_recpe_f32(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
 -                        case NEON_2RM_VRSQRTE_F:
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_rsqrte_f32(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
 -                        case NEON_2RM_VCVT_FS: /* VCVT.F32.S32 */
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_vfp_sitos(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
 -                        case NEON_2RM_VCVT_FU: /* VCVT.F32.U32 */
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_vfp_uitos(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
 -                        case NEON_2RM_VCVT_SF: /* VCVT.S32.F32 */
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_vfp_tosizs(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
 -                        case NEON_2RM_VCVT_UF: /* VCVT.U32.F32 */
 -                        {
 -                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
 -                            gen_helper_vfp_touizs(tmp, tmp, fpstatus);
 -                            tcg_temp_free_ptr(fpstatus);
 -                            break;
 -                        }
                          default:
                              /* Reserved op values were caught by the
                               * neon_2rm_sizes[] check earlier.
 --
 .20.1

-[Qemu-devel] [PULL 12/25] Make address_space_translate{, _cached}() take a MemTxAttrs argument
+[PULL 18/42] target/arm: Convert Neon 2-reg-misc fp-compare-with-zero insns to decodetree
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the fp-compare-with-zero insns in the Neon 2-reg-misc group to
-add MemTxAttrs as an argument to address_space_translate()
+decodetree.
 and address_space_translate_cached(). Callers either have an
 attrs value to hand, or don't care and can use MEMTXATTRS_UNSPECIFIED.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-4-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-17-peter.maydell@linaro.org
 ---
- include/exec/memory.h     |  4 +++-
+ target/arm/neon-dp.decode       |  6 ++++
- accel/tcg/translate-all.c |  2 +-
+ target/arm/translate-neon.inc.c | 28 ++++++++++++++++++
- exec.c                    | 14 +++++++++-----
+ target/arm/translate.c          | 50 ++++-----------------------------
- hw/vfio/common.c          |  3 ++-
+files changed, 39 insertions(+), 45 deletions(-)
  memory_ldst.inc.c         | 18 +++++++++---------
  target/riscv/helper.c     |  2 +-
 files changed, 25 insertions(+), 18 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/neon-dp.decode
-+++ b/include/exec/memory.h
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
-  * #MemoryRegion.
+     VABS         1111 001 11 . 11 .. 01 .... 0 0110 . . 0 .... @2misc
-  * @len: pointer to length
+     VNEG         1111 001 11 . 11 .. 01 .... 0 0111 . . 0 .... @2misc
-  * @is_write: indicates the transfer direction
-+ * @attrs: memory attributes
++    VCGT0_F      1111 001 11 . 11 .. 01 .... 0 1000 . . 0 .... @2misc
-  */
++    VCGE0_F      1111 001 11 . 11 .. 01 .... 0 1001 . . 0 .... @2misc
- MemoryRegion *flatview_translate(FlatView *fv,
++    VCEQ0_F      1111 001 11 . 11 .. 01 .... 0 1010 . . 0 .... @2misc
-                                  hwaddr addr, hwaddr *xlat,
++    VCLE0_F      1111 001 11 . 11 .. 01 .... 0 1011 . . 0 .... @2misc
-@@ -XXX,XX +XXX,XX @@ MemoryRegion *flatview_translate(FlatView *fv,
++    VCLT0_F      1111 001 11 . 11 .. 01 .... 0 1100 . . 0 .... @2misc
++
- static inline MemoryRegion *address_space_translate(AddressSpace *as,
+     VABS_F       1111 001 11 . 11 .. 01 .... 0 1110 . . 0 .... @2misc
-                                                     hwaddr addr, hwaddr *xlat,
+     VNEG_F       1111 001 11 . 11 .. 01 .... 0 1111 . . 0 .... @2misc
--                                                    hwaddr *len, bool is_write)
-+                                                    hwaddr *len, bool is_write,
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 +                                                    MemTxAttrs attrs)
  {
      return flatview_translate(address_space_to_flatview(as),
                                addr, xlat, len, is_write);
 diff --git a/accel/tcg/translate-all.c b/accel/tcg/translate-all.c
 index XXXXXXX..XXXXXXX 100644
---- a/accel/tcg/translate-all.c
+--- a/target/arm/translate-neon.inc.c
-+++ b/accel/tcg/translate-all.c
++++ b/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs)
+@@ -XXX,XX +XXX,XX @@ static bool trans_VRINTX(DisasContext *s, arg_2misc *a)
-     hwaddr l = 1;
+     }
+     return do_2misc_fp(s, a, gen_helper_rints_exact);
-     rcu_read_lock();
+ }
--    mr = address_space_translate(as, addr, &addr, &l, false);
++
-+    mr = address_space_translate(as, addr, &addr, &l, false, attrs);
++#define WRAP_FP_CMP0_FWD(WRAPNAME, FUNC)                        \
-     if (!(memory_region_is_ram(mr)
++    static void WRAPNAME(TCGv_i32 d, TCGv_i32 m, TCGv_ptr fpst) \
-           || memory_region_is_romd(mr))) {
++    {                                                           \
-         rcu_read_unlock();
++        TCGv_i32 zero = tcg_const_i32(0);                       \
-diff --git a/exec.c b/exec.c
++        FUNC(d, m, zero, fpst);                                 \
 +        tcg_temp_free_i32(zero);                                \
 +    }
 +#define WRAP_FP_CMP0_REV(WRAPNAME, FUNC)                        \
 +    static void WRAPNAME(TCGv_i32 d, TCGv_i32 m, TCGv_ptr fpst) \
 +    {                                                           \
 +        TCGv_i32 zero = tcg_const_i32(0);                       \
 +        FUNC(d, zero, m, fpst);                                 \
 +        tcg_temp_free_i32(zero);                                \
 +    }
 +
 +#define DO_FP_CMP0(INSN, FUNC, REV)                             \
 +    WRAP_FP_CMP0_##REV(gen_##INSN, FUNC)                        \
 +    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
 +    {                                                           \
 +        return do_2misc_fp(s, a, gen_##INSN);                   \
 +    }
 +
 +DO_FP_CMP0(VCGT0_F, gen_helper_neon_cgt_f32, FWD)
 +DO_FP_CMP0(VCGE0_F, gen_helper_neon_cge_f32, FWD)
 +DO_FP_CMP0(VCEQ0_F, gen_helper_neon_ceq_f32, FWD)
 +DO_FP_CMP0(VCLE0_F, gen_helper_neon_cge_f32, REV)
 +DO_FP_CMP0(VCLT0_F, gen_helper_neon_cgt_f32, REV)
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/translate.c
-+++ b/exec.c
++++ b/target/arm/translate.c
-@@ -XXX,XX +XXX,XX @@ static inline void cpu_physical_memory_write_rom_internal(AddressSpace *as,
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-     rcu_read_lock();
+                 case NEON_2RM_VCVT_SF:
-     while (len > 0) {
+                 case NEON_2RM_VCVT_UF:
-         l = len;
+                 case NEON_2RM_VRINTX:
--        mr = address_space_translate(as, addr, &addr1, &l, true);
++                case NEON_2RM_VCGT0_F:
-+        mr = address_space_translate(as, addr, &addr1, &l, true,
++                case NEON_2RM_VCGE0_F:
-+                                     MEMTXATTRS_UNSPECIFIED);
++                case NEON_2RM_VCEQ0_F:
++                case NEON_2RM_VCLE0_F:
-         if (!(memory_region_is_ram(mr) ||
++                case NEON_2RM_VCLT0_F:
-               memory_region_is_romd(mr))) {
+                     /* handled by decodetree */
-@@ -XXX,XX +XXX,XX @@ void address_space_cache_destroy(MemoryRegionCache *cache)
+                     return 1;
-  */
+                 case NEON_2RM_VTRN:
- static inline MemoryRegion *address_space_translate_cached(
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-     MemoryRegionCache *cache, hwaddr addr, hwaddr *xlat,
+                     for (pass = 0; pass < (q ? 4 : 2); pass++) {
--    hwaddr *plen, bool is_write)
+                         tmp = neon_load_reg(rm, pass);
-+    hwaddr *plen, bool is_write, MemTxAttrs attrs)
+                         switch (op) {
- {
+-                        case NEON_2RM_VCGT0_F:
-     MemoryRegionSection section;
+-                        {
-     MemoryRegion *mr;
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-@@ -XXX,XX +XXX,XX @@ address_space_read_cached_slow(MemoryRegionCache *cache, hwaddr addr,
+-                            tmp2 = tcg_const_i32(0);
-     MemoryRegion *mr;
+-                            gen_helper_neon_cgt_f32(tmp, tmp, tmp2, fpstatus);
+-                            tcg_temp_free_i32(tmp2);
-     l = len;
+-                            tcg_temp_free_ptr(fpstatus);
--    mr = address_space_translate_cached(cache, addr, &addr1, &l, false);
+-                            break;
-+    mr = address_space_translate_cached(cache, addr, &addr1, &l, false,
+-                        }
-+                                        MEMTXATTRS_UNSPECIFIED);
+-                        case NEON_2RM_VCGE0_F:
-     flatview_read_continue(cache->fv,
+-                        {
-                            addr, MEMTXATTRS_UNSPECIFIED, buf, len,
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            addr1, l, mr);
+-                            tmp2 = tcg_const_i32(0);
-@@ -XXX,XX +XXX,XX @@ address_space_write_cached_slow(MemoryRegionCache *cache, hwaddr addr,
+-                            gen_helper_neon_cge_f32(tmp, tmp, tmp2, fpstatus);
-     MemoryRegion *mr;
+-                            tcg_temp_free_i32(tmp2);
+-                            tcg_temp_free_ptr(fpstatus);
-     l = len;
+-                            break;
--    mr = address_space_translate_cached(cache, addr, &addr1, &l, true);
+-                        }
-+    mr = address_space_translate_cached(cache, addr, &addr1, &l, true,
+-                        case NEON_2RM_VCEQ0_F:
-+                                        MEMTXATTRS_UNSPECIFIED);
+-                        {
-     flatview_write_continue(cache->fv,
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                             addr, MEMTXATTRS_UNSPECIFIED, buf, len,
+-                            tmp2 = tcg_const_i32(0);
-                             addr1, l, mr);
+-                            gen_helper_neon_ceq_f32(tmp, tmp, tmp2, fpstatus);
-@@ -XXX,XX +XXX,XX @@ bool cpu_physical_memory_is_io(hwaddr phys_addr)
+-                            tcg_temp_free_i32(tmp2);
+-                            tcg_temp_free_ptr(fpstatus);
-     rcu_read_lock();
+-                            break;
-     mr = address_space_translate(&address_space_memory,
+-                        }
--                                 phys_addr, &phys_addr, &l, false);
+-                        case NEON_2RM_VCLE0_F:
-+                                 phys_addr, &phys_addr, &l, false,
+-                        {
-+                                 MEMTXATTRS_UNSPECIFIED);
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
+-                            tmp2 = tcg_const_i32(0);
-     res = !(memory_region_is_ram(mr) || memory_region_is_romd(mr));
+-                            gen_helper_neon_cge_f32(tmp, tmp2, tmp, fpstatus);
-     rcu_read_unlock();
+-                            tcg_temp_free_i32(tmp2);
-diff --git a/hw/vfio/common.c b/hw/vfio/common.c
+-                            tcg_temp_free_ptr(fpstatus);
-index XXXXXXX..XXXXXXX 100644
+-                            break;
---- a/hw/vfio/common.c
+-                        }
-+++ b/hw/vfio/common.c
+-                        case NEON_2RM_VCLT0_F:
-@@ -XXX,XX +XXX,XX @@ static bool vfio_get_vaddr(IOMMUTLBEntry *iotlb, void **vaddr,
+-                        {
-      */
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-     mr = address_space_translate(&address_space_memory,
+-                            tmp2 = tcg_const_i32(0);
-                                  iotlb->translated_addr,
+-                            gen_helper_neon_cgt_f32(tmp, tmp2, tmp, fpstatus);
--                                 &xlat, &len, writable);
+-                            tcg_temp_free_i32(tmp2);
-+                                 &xlat, &len, writable,
+-                            tcg_temp_free_ptr(fpstatus);
-+                                 MEMTXATTRS_UNSPECIFIED);
+-                            break;
-     if (!memory_region_is_ram(mr)) {
+-                        }
-         error_report("iommu map to non memory area %"HWADDR_PRIx"",
+                         case NEON_2RM_VSWP:
-                      xlat);
+                             tmp2 = neon_load_reg(rd, pass);
-diff --git a/memory_ldst.inc.c b/memory_ldst.inc.c
+                             neon_store_reg(rm, pass, tmp2);
 index XXXXXXX..XXXXXXX 100644
 --- a/memory_ldst.inc.c
 +++ b/memory_ldst.inc.c
@@ -XXX,XX +XXX,XX @@ static inline uint32_t glue(address_space_ldl_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, false);
 +    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
      if (l < 4 || !IS_DIRECT(mr, false)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ static inline uint64_t glue(address_space_ldq_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, false);
 +    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
      if (l < 8 || !IS_DIRECT(mr, false)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ uint32_t glue(address_space_ldub, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, false);
 +    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
      if (!IS_DIRECT(mr, false)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ static inline uint32_t glue(address_space_lduw_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, false);
 +    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
      if (l < 2 || !IS_DIRECT(mr, false)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ void glue(address_space_stl_notdirty, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, true);
 +    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
      if (l < 4 || !IS_DIRECT(mr, true)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ static inline void glue(address_space_stl_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, true);
 +    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
      if (l < 4 || !IS_DIRECT(mr, true)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ void glue(address_space_stb, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, true);
 +    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
      if (!IS_DIRECT(mr, true)) {
          release_lock |= prepare_mmio_access(mr);
          r = memory_region_dispatch_write(mr, addr1, val, 1, attrs);
@@ -XXX,XX +XXX,XX @@ static inline void glue(address_space_stw_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, true);
 +    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
      if (l < 2 || !IS_DIRECT(mr, true)) {
          release_lock |= prepare_mmio_access(mr);
@@ -XXX,XX +XXX,XX @@ static void glue(address_space_stq_internal, SUFFIX)(ARG1_DECL,
      bool release_lock = false;
      RCU_READ_LOCK();
 -    mr = TRANSLATE(addr, &addr1, &l, true);
 +    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
      if (l < 8 || !IS_DIRECT(mr, true)) {
          release_lock |= prepare_mmio_access(mr);
 diff --git a/target/riscv/helper.c b/target/riscv/helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/riscv/helper.c
 +++ b/target/riscv/helper.c
@@ -XXX,XX +XXX,XX @@ restart:
                  MemoryRegion *mr;
                  hwaddr l = sizeof(target_ulong), addr1;
                  mr = address_space_translate(cs->as, pte_addr,
 -                    &addr1, &l, false);
 +                    &addr1, &l, false, MEMTXATTRS_UNSPECIFIED);
                  if (memory_access_is_direct(mr, true)) {
                      target_ulong *pte_pa =
                          qemu_map_ram_ptr(mr->ram_block, addr1);
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 10/25] memory.h: Improve IOMMU related documentation
+[PULL 19/42] target/arm: Convert Neon 2-reg-misc VRINT insns to decodetree
-Add more detail to the documentation for memory_region_init_iommu()
+Convert the Neon 2-reg-misc VRINT insns to decodetree.
-and other IOMMU-related functions and data structures.
+Giving these insns their own do_vrint() function allows us
 to change the rounding mode just once at the start and end
 rather than doing it for every element in the vector.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
+Message-id: 20200616170844.13318-18-peter.maydell@linaro.org
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 20180521140402.23318-2-peter.maydell@linaro.org
 ---
- include/exec/memory.h | 105 ++++++++++++++++++++++++++++++++++++++----
+ target/arm/neon-dp.decode       |  8 +++++
-file changed, 95 insertions(+), 10 deletions(-)
+ target/arm/translate-neon.inc.c | 61 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 31 +++--------------
 files changed, 74 insertions(+), 26 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/neon-dp.decode
-+++ b/include/exec/memory.h
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ enum IOMMUMemoryRegionAttr {
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
-     IOMMU_ATTR_SPAPR_TCE_FD
+     SHA1SU1      1111 001 11 . 11 .. 10 .... 0 0111 0 . 0 .... @2misc_q1
- };
+     SHA256SU0    1111 001 11 . 11 .. 10 .... 0 0111 1 . 0 .... @2misc_q1
-+/**
++    VRINTN       1111 001 11 . 11 .. 10 .... 0 1000 . . 0 .... @2misc
-+ * IOMMUMemoryRegionClass:
+     VRINTX       1111 001 11 . 11 .. 10 .... 0 1001 . . 0 .... @2misc
-+ *
++    VRINTA       1111 001 11 . 11 .. 10 .... 0 1010 . . 0 .... @2misc
-+ * All IOMMU implementations need to subclass TYPE_IOMMU_MEMORY_REGION
++    VRINTZ       1111 001 11 . 11 .. 10 .... 0 1011 . . 0 .... @2misc
-+ * and provide an implementation of at least the @translate method here
-+ * to handle requests to the memory region. Other methods are optional.
+     VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
-+ *
++
-+ * The IOMMU implementation must use the IOMMU notifier infrastructure
++    VRINTM       1111 001 11 . 11 .. 10 .... 0 1101 . . 0 .... @2misc
-+ * to report whenever mappings are changed, by calling
++
-+ * memory_region_notify_iommu() (or, if necessary, by calling
+     VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
-+ * memory_region_notify_one() for each registered notifier).
-+ */
++    VRINTP       1111 001 11 . 11 .. 10 .... 0 1111 . . 0 .... @2misc
- typedef struct IOMMUMemoryRegionClass {
++
-     /* private */
+     VRECPE       1111 001 11 . 11 .. 11 .... 0 1000 . . 0 .... @2misc
-     struct DeviceClass parent_class;
+     VRSQRTE      1111 001 11 . 11 .. 11 .... 0 1001 . . 0 .... @2misc
+     VRECPE_F     1111 001 11 . 11 .. 11 .... 0 1010 . . 0 .... @2misc
-     /*
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
--     * Return a TLB entry that contains a given address. Flag should
+index XXXXXXX..XXXXXXX 100644
--     * be the access permission of this translation operation. We can
+--- a/target/arm/translate-neon.inc.c
--     * set flag to IOMMU_NONE to mean that we don't need any
++++ b/target/arm/translate-neon.inc.c
--     * read/write permission checks, like, when for region replay.
+@@ -XXX,XX +XXX,XX @@ DO_FP_CMP0(VCGE0_F, gen_helper_neon_cge_f32, FWD)
-+     * Return a TLB entry that contains a given address.
+ DO_FP_CMP0(VCEQ0_F, gen_helper_neon_ceq_f32, FWD)
-+     *
+ DO_FP_CMP0(VCLE0_F, gen_helper_neon_cge_f32, REV)
-+     * The IOMMUAccessFlags indicated via @flag are optional and may
+ DO_FP_CMP0(VCLT0_F, gen_helper_neon_cgt_f32, REV)
-+     * be specified as IOMMU_NONE to indicate that the caller needs
++
-+     * the full translation information for both reads and writes. If
++static bool do_vrint(DisasContext *s, arg_2misc *a, int rmode)
-+     * the access flags are specified then the IOMMU implementation
++{
-+     * may use this as an optimization, to stop doing a page table
++    /*
-+     * walk as soon as it knows that the requested permissions are not
++     * Handle a VRINT* operation by iterating 32 bits at a time,
-+     * allowed. If IOMMU_NONE is passed then the IOMMU must do the
++     * with a specified rounding mode in operation.
 +     * full page table walk and report the permissions in the returned
 +     * IOMMUTLBEntry. (Note that this implies that an IOMMU may not
 +     * return different mappings for reads and writes.)
 +     *
 +     * The returned information remains valid while the caller is
 +     * holding the big QEMU lock or is inside an RCU critical section;
 +     * if the caller wishes to cache the mapping beyond that it must
 +     * register an IOMMU notifier so it can invalidate its cached
 +     * information when the IOMMU mapping changes.
 +     *
 +     * @iommu: the IOMMUMemoryRegion
 +     * @hwaddr: address to be translated within the memory region
 +     * @flag: requested access permissions
       */
      IOMMUTLBEntry (*translate)(IOMMUMemoryRegion *iommu, hwaddr addr,
                                 IOMMUAccessFlags flag);
 -    /* Returns minimum supported page size */
 +    /* Returns minimum supported page size in bytes.
 +     * If this method is not provided then the minimum is assumed to
 +     * be TARGET_PAGE_SIZE.
 +     *
 +     * @iommu: the IOMMUMemoryRegion
 +     */
-     uint64_t (*get_min_page_size)(IOMMUMemoryRegion *iommu);
++    int pass;
--    /* Called when IOMMU Notifier flag changed */
++    TCGv_ptr fpst;
-+    /* Called when IOMMU Notifier flag changes (ie when the set of
++    TCGv_i32 tcg_rmode;
-+     * events which IOMMU users are requesting notification for changes).
++
-+     * Optional method -- need not be provided if the IOMMU does not
++    if (!arm_dc_feature(s, ARM_FEATURE_NEON) ||
-+     * need to know exactly which events must be notified.
++        !arm_dc_feature(s, ARM_FEATURE_V8)) {
-+     *
++        return false;
-+     * @iommu: the IOMMUMemoryRegion
++    }
-+     * @old_flags: events which previously needed to be notified
++
-+     * @new_flags: events which now need to be notified
++    /* UNDEF accesses to D16-D31 if they don't exist. */
-+     */
++    if (!dc_isar_feature(aa32_simd_r32, s) &&
-     void (*notify_flag_changed)(IOMMUMemoryRegion *iommu,
++        ((a->vd | a->vm) & 0x10)) {
-                                 IOMMUNotifierFlag old_flags,
++        return false;
-                                 IOMMUNotifierFlag new_flags);
++    }
--    /* Set this up to provide customized IOMMU replay function */
++
-+    /* Called to handle memory_region_iommu_replay().
++    if (a->size != 2) {
-+     *
++        /* TODO: FP16 will be the size == 1 case */
-+     * The default implementation of memory_region_iommu_replay() is to
++        return false;
-+     * call the IOMMU translate method for every page in the address space
++    }
-+     * with flag == IOMMU_NONE and then call the notifier if translate
++
-+     * returns a valid mapping. If this method is implemented then it
++    if ((a->vd | a->vm) & a->q) {
-+     * overrides the default behaviour, and must provide the full semantics
++        return false;
-+     * of memory_region_iommu_replay(), by calling @notifier for every
++    }
-+     * translation present in the IOMMU.
++
-+     *
++    if (!vfp_access_check(s)) {
-+     * Optional method -- an IOMMU only needs to provide this method
++        return true;
-+     * if the default is inefficient or produces undesirable side effects.
++    }
-+     *
++
-+     * Note: this is not related to record-and-replay functionality.
++    fpst = get_fpstatus_ptr(1);
-+     */
++    tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
-     void (*replay)(IOMMUMemoryRegion *iommu, IOMMUNotifier *notifier);
++    gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
++    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
--    /* Get IOMMU misc attributes */
++        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
--    int (*get_attr)(IOMMUMemoryRegion *iommu, enum IOMMUMemoryRegionAttr,
++        gen_helper_rints(tmp, tmp, fpst);
-+    /* Get IOMMU misc attributes. This is an optional method that
++        neon_store_reg(a->vd, pass, tmp);
-+     * can be used to allow users of the IOMMU to get implementation-specific
++    }
-+     * information. The IOMMU implements this method to handle calls
++    gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
-+     * by IOMMU users to memory_region_iommu_get_attr() by filling in
++    tcg_temp_free_i32(tcg_rmode);
-+     * the arbitrary data pointer for any IOMMUMemoryRegionAttr values that
++    tcg_temp_free_ptr(fpst);
-+     * the IOMMU supports. If the method is unimplemented then
++
-+     * memory_region_iommu_get_attr() will always return -EINVAL.
++    return true;
-+     *
++}
-+     * @iommu: the IOMMUMemoryRegion
++
-+     * @attr: attribute being queried
++#define DO_VRINT(INSN, RMODE)                                   \
-+     * @data: memory to fill in with the attribute data
++    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
-+     *
++    {                                                           \
-+     * Returns 0 on success, or a negative errno; in particular
++        return do_vrint(s, a, RMODE);                           \
-+     * returns -EINVAL for unrecognized or unimplemented attribute types.
++    }
-+     */
++
-+    int (*get_attr)(IOMMUMemoryRegion *iommu, enum IOMMUMemoryRegionAttr attr,
++DO_VRINT(VRINTN, FPROUNDING_TIEEVEN)
-                     void *data);
++DO_VRINT(VRINTA, FPROUNDING_TIEAWAY)
- } IOMMUMemoryRegionClass;
++DO_VRINT(VRINTZ, FPROUNDING_ZERO)
++DO_VRINT(VRINTM, FPROUNDING_NEGINF)
-@@ -XXX,XX +XXX,XX @@ static inline void memory_region_init_reservation(MemoryRegion *mr,
++DO_VRINT(VRINTP, FPROUNDING_POSINF)
-  * An IOMMU region translates addresses and forwards accesses to a target
+diff --git a/target/arm/translate.c b/target/arm/translate.c
-  * memory region.
+index XXXXXXX..XXXXXXX 100644
-  *
+--- a/target/arm/translate.c
-+ * The IOMMU implementation must define a subclass of TYPE_IOMMU_MEMORY_REGION.
++++ b/target/arm/translate.c
-+ * @_iommu_mr should be a pointer to enough memory for an instance of
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-+ * that subclass, @instance_size is the size of that subclass, and
+                 case NEON_2RM_VCEQ0_F:
-+ * @mrtypename is its name. This function will initialize @_iommu_mr as an
+                 case NEON_2RM_VCLE0_F:
-+ * instance of the subclass, and its methods will then be called to handle
+                 case NEON_2RM_VCLT0_F:
-+ * accesses to the memory region. See the documentation of
++                case NEON_2RM_VRINTN:
-+ * #IOMMUMemoryRegionClass for further details.
++                case NEON_2RM_VRINTA:
-+ *
++                case NEON_2RM_VRINTM:
-  * @_iommu_mr: the #IOMMUMemoryRegion to be initialized
++                case NEON_2RM_VRINTP:
-  * @instance_size: the IOMMUMemoryRegion subclass instance size
++                case NEON_2RM_VRINTZ:
-  * @mrtypename: the type name of the #IOMMUMemoryRegion
+                     /* handled by decodetree */
-@@ -XXX,XX +XXX,XX @@ void memory_region_register_iommu_notifier(MemoryRegion *mr,
+                     return 1;
-  * a notifier with the minimum page granularity returned by
+                 case NEON_2RM_VTRN:
-  * mr->iommu_ops->get_page_size().
+@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-  *
+                             }
-+ * Note: this is not related to record-and-replay functionality.
+                             neon_store_reg(rm, pass, tmp2);
-+ *
+                             break;
-  * @iommu_mr: the memory region to observe
+-                        case NEON_2RM_VRINTN:
-  * @n: the notifier to which to replay iommu mappings
+-                        case NEON_2RM_VRINTA:
-  */
+-                        case NEON_2RM_VRINTM:
-@@ -XXX,XX +XXX,XX @@ void memory_region_iommu_replay(IOMMUMemoryRegion *iommu_mr, IOMMUNotifier *n);
+-                        case NEON_2RM_VRINTP:
-  * memory_region_iommu_replay_all: replay existing IOMMU translations
+-                        case NEON_2RM_VRINTZ:
-  * to all the notifiers registered.
+-                        {
-  *
+-                            TCGv_i32 tcg_rmode;
-+ * Note: this is not related to record-and-replay functionality.
+-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-+ *
+-                            int rmode;
-  * @iommu_mr: the memory region to observe
+-
-  */
+-                            if (op == NEON_2RM_VRINTZ) {
- void memory_region_iommu_replay_all(IOMMUMemoryRegion *iommu_mr);
+-                                rmode = FPROUNDING_ZERO;
-@@ -XXX,XX +XXX,XX @@ void memory_region_unregister_iommu_notifier(MemoryRegion *mr,
+-                            } else {
-  * memory_region_iommu_get_attr: return an IOMMU attr if get_attr() is
+-                                rmode = fp_decode_rm[((op & 0x6) >> 1) ^ 1];
-  * defined on the IOMMU.
+-                            }
-  *
+-
-- * Returns 0 if succeded, error code otherwise.
+-                            tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
-+ * Returns 0 on success, or a negative errno otherwise. In particular,
+-                            gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode,
-+ * -EINVAL indicates that the IOMMU does not support the requested
+-                                                      cpu_env);
-+ * attribute.
+-                            gen_helper_rints(tmp, tmp, fpstatus);
-  *
+-                            gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode,
-  * @iommu_mr: the memory region
+-                                                      cpu_env);
-  * @attr: the requested attribute
+-                            tcg_temp_free_ptr(fpstatus);
 -                            tcg_temp_free_i32(tcg_rmode);
 -                            break;
 -                        }
                          case NEON_2RM_VCVTAU:
                          case NEON_2RM_VCVTAS:
                          case NEON_2RM_VCVTNU:
 --
-.17.1
+.20.1

-New patch
+[PULL 20/42] target/arm: Convert Neon 2-reg-misc VCVT insns to decodetree
+Convert the VCVT instructions in the 2-reg-misc grouping to
 decodetree.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
 Message-id: 20200616170844.13318-19-peter.maydell@linaro.org
 ---
  target/arm/neon-dp.decode       |  9 +++++
  target/arm/translate-neon.inc.c | 70 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          | 70 ++++-----------------------------
 files changed, 87 insertions(+), 62 deletions(-)
 diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/neon-dp.decode
 +++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
      VRINTP       1111 001 11 . 11 .. 10 .... 0 1111 . . 0 .... @2misc
 +    VCVTAS       1111 001 11 . 11 .. 11 .... 0 0000 . . 0 .... @2misc
 +    VCVTAU       1111 001 11 . 11 .. 11 .... 0 0001 . . 0 .... @2misc
 +    VCVTNS       1111 001 11 . 11 .. 11 .... 0 0010 . . 0 .... @2misc
 +    VCVTNU       1111 001 11 . 11 .. 11 .... 0 0011 . . 0 .... @2misc
 +    VCVTPS       1111 001 11 . 11 .. 11 .... 0 0100 . . 0 .... @2misc
 +    VCVTPU       1111 001 11 . 11 .. 11 .... 0 0101 . . 0 .... @2misc
 +    VCVTMS       1111 001 11 . 11 .. 11 .... 0 0110 . . 0 .... @2misc
 +    VCVTMU       1111 001 11 . 11 .. 11 .... 0 0111 . . 0 .... @2misc
 +
      VRECPE       1111 001 11 . 11 .. 11 .... 0 1000 . . 0 .... @2misc
      VRSQRTE      1111 001 11 . 11 .. 11 .... 0 1001 . . 0 .... @2misc
      VRECPE_F     1111 001 11 . 11 .. 11 .... 0 1010 . . 0 .... @2misc
 diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate-neon.inc.c
 +++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ DO_VRINT(VRINTA, FPROUNDING_TIEAWAY)
  DO_VRINT(VRINTZ, FPROUNDING_ZERO)
  DO_VRINT(VRINTM, FPROUNDING_NEGINF)
  DO_VRINT(VRINTP, FPROUNDING_POSINF)
 +
 +static bool do_vcvt(DisasContext *s, arg_2misc *a, int rmode, bool is_signed)
 +{
 +    /*
 +     * Handle a VCVT* operation by iterating 32 bits at a time,
 +     * with a specified rounding mode in operation.
 +     */
 +    int pass;
 +    TCGv_ptr fpst;
 +    TCGv_i32 tcg_rmode, tcg_shift;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON) ||
 +        !arm_dc_feature(s, ARM_FEATURE_V8)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (a->size != 2) {
 +        /* TODO: FP16 will be the size == 1 case */
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    fpst = get_fpstatus_ptr(1);
 +    tcg_shift = tcg_const_i32(0);
 +    tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
 +    gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
 +    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
 +        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
 +        if (is_signed) {
 +            gen_helper_vfp_tosls(tmp, tmp, tcg_shift, fpst);
 +        } else {
 +            gen_helper_vfp_touls(tmp, tmp, tcg_shift, fpst);
 +        }
 +        neon_store_reg(a->vd, pass, tmp);
 +    }
 +    gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode, cpu_env);
 +    tcg_temp_free_i32(tcg_rmode);
 +    tcg_temp_free_i32(tcg_shift);
 +    tcg_temp_free_ptr(fpst);
 +
 +    return true;
 +}
 +
 +#define DO_VCVT(INSN, RMODE, SIGNED)                            \
 +    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
 +    {                                                           \
 +        return do_vcvt(s, a, RMODE, SIGNED);                    \
 +    }
 +
 +DO_VCVT(VCVTAU, FPROUNDING_TIEAWAY, false)
 +DO_VCVT(VCVTAS, FPROUNDING_TIEAWAY, true)
 +DO_VCVT(VCVTNU, FPROUNDING_TIEEVEN, false)
 +DO_VCVT(VCVTNS, FPROUNDING_TIEEVEN, true)
 +DO_VCVT(VCVTPU, FPROUNDING_POSINF, false)
 +DO_VCVT(VCVTPS, FPROUNDING_POSINF, true)
 +DO_VCVT(VCVTMU, FPROUNDING_NEGINF, false)
 +DO_VCVT(VCVTMS, FPROUNDING_NEGINF, true)
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
  #define NEON_2RM_VCVT_SF 62
  #define NEON_2RM_VCVT_UF 63
 -static bool neon_2rm_is_v8_op(int op)
 -{
 -    /* Return true if this neon 2reg-misc op is ARMv8 and up */
 -    switch (op) {
 -    case NEON_2RM_VRINTN:
 -    case NEON_2RM_VRINTA:
 -    case NEON_2RM_VRINTM:
 -    case NEON_2RM_VRINTP:
 -    case NEON_2RM_VRINTZ:
 -    case NEON_2RM_VRINTX:
 -    case NEON_2RM_VCVTAU:
 -    case NEON_2RM_VCVTAS:
 -    case NEON_2RM_VCVTNU:
 -    case NEON_2RM_VCVTNS:
 -    case NEON_2RM_VCVTPU:
 -    case NEON_2RM_VCVTPS:
 -    case NEON_2RM_VCVTMU:
 -    case NEON_2RM_VCVTMS:
 -        return true;
 -    default:
 -        return false;
 -    }
 -}
 -
  /* Each entry in this array has bit n set if the insn allows
   * size value n (otherwise it will UNDEF). Since unallocated
   * op values will have no bits set they always UNDEF.
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  if ((neon_2rm_sizes[op] & (1 << size)) == 0) {
                      return 1;
                  }
 -                if (neon_2rm_is_v8_op(op) &&
 -                    !arm_dc_feature(s, ARM_FEATURE_V8)) {
 -                    return 1;
 -                }
                  if (q && ((rm | rd) & 1)) {
                      return 1;
                  }
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VRINTM:
                  case NEON_2RM_VRINTP:
                  case NEON_2RM_VRINTZ:
 +                case NEON_2RM_VCVTAU:
 +                case NEON_2RM_VCVTAS:
 +                case NEON_2RM_VCVTNU:
 +                case NEON_2RM_VCVTNS:
 +                case NEON_2RM_VCVTPU:
 +                case NEON_2RM_VCVTPS:
 +                case NEON_2RM_VCVTMU:
 +                case NEON_2RM_VCVTMS:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                              }
                              neon_store_reg(rm, pass, tmp2);
                              break;
 -                        case NEON_2RM_VCVTAU:
 -                        case NEON_2RM_VCVTAS:
 -                        case NEON_2RM_VCVTNU:
 -                        case NEON_2RM_VCVTNS:
 -                        case NEON_2RM_VCVTPU:
 -                        case NEON_2RM_VCVTPS:
 -                        case NEON_2RM_VCVTMU:
 -                        case NEON_2RM_VCVTMS:
 -                        {
 -                            bool is_signed = !extract32(insn, 7, 1);
 -                            TCGv_ptr fpst = get_fpstatus_ptr(1);
 -                            TCGv_i32 tcg_rmode, tcg_shift;
 -                            int rmode = fp_decode_rm[extract32(insn, 8, 2)];
 -
 -                            tcg_shift = tcg_const_i32(0);
 -                            tcg_rmode = tcg_const_i32(arm_rmode_to_sf(rmode));
 -                            gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode,
 -                                                      cpu_env);
 -
 -                            if (is_signed) {
 -                                gen_helper_vfp_tosls(tmp, tmp,
 -                                                     tcg_shift, fpst);
 -                            } else {
 -                                gen_helper_vfp_touls(tmp, tmp,
 -                                                     tcg_shift, fpst);
 -                            }
 -
 -                            gen_helper_set_neon_rmode(tcg_rmode, tcg_rmode,
 -                                                      cpu_env);
 -                            tcg_temp_free_i32(tcg_rmode);
 -                            tcg_temp_free_i32(tcg_shift);
 -                            tcg_temp_free_ptr(fpst);
 -                            break;
 -                        }
                          default:
                              /* Reserved op values were caught by the
                               * neon_2rm_sizes[] check earlier.
 --
 .20.1

-[Qemu-devel] [PULL 01/25] target/arm: Honour FPCR.FZ in FRECPX
+[PULL 21/42] target/arm: Convert Neon VSWP to decodetree
-The FRECPX instructions should (like most other floating point operations)
+Convert the Neon VSWP insn to decodetree. Since the new implementation
-honour the FPCR.FZ bit which specifies whether input denormals should
+doesn't have to share a pass-loop with the other 2-reg-misc operations
-be flushed to zero (or FZ16 for the half-precision version).
+we can implement the swap with 64-bit accesses rather than 32-bits
-We forgot to implement this, which doesn't affect the results (since
+(which brings us into line with the pseudocode and is more efficient).
 the calculation doesn't actually care about the mantissa bits) but did
 mean we were failing to set the FPSR.IDC bit.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521172712.19930-1-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-20-peter.maydell@linaro.org
 ---
- target/arm/helper-a64.c | 6 ++++++
+ target/arm/neon-dp.decode       |  2 ++
-file changed, 6 insertions(+)
+ target/arm/translate-neon.inc.c | 41 +++++++++++++++++++++++++++++++++
  target/arm/translate.c          |  5 +---
 files changed, 44 insertions(+), 4 deletions(-)
-diff --git a/target/arm/helper-a64.c b/target/arm/helper-a64.c
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/helper-a64.c
+--- a/target/arm/neon-dp.decode
-+++ b/target/arm/helper-a64.c
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ float16 HELPER(frecpx_f16)(float16 a, void *fpstp)
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
-         return nan;
+     VABS_F       1111 001 11 . 11 .. 01 .... 0 1110 . . 0 .... @2misc
-     }
+     VNEG_F       1111 001 11 . 11 .. 01 .... 0 1111 . . 0 .... @2misc
-+    a = float16_squash_input_denormal(a, fpst);
++    VSWP         1111 001 11 . 11 .. 10 .... 0 0000 . . 0 .... @2misc
 +
-     val16 = float16_val(a);
+     VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
-     sbit = 0x8000 & val16;
+     VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
-     exp = extract32(val16, 10, 5);
-@@ -XXX,XX +XXX,XX @@ float32 HELPER(frecpx_f32)(float32 a, void *fpstp)
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
-         return nan;
+index XXXXXXX..XXXXXXX 100644
-     }
+--- a/target/arm/translate-neon.inc.c
++++ b/target/arm/translate-neon.inc.c
-+    a = float32_squash_input_denormal(a, fpst);
+@@ -XXX,XX +XXX,XX @@ DO_VCVT(VCVTPU, FPROUNDING_POSINF, false)
  DO_VCVT(VCVTPS, FPROUNDING_POSINF, true)
  DO_VCVT(VCVTMU, FPROUNDING_NEGINF, false)
  DO_VCVT(VCVTMS, FPROUNDING_NEGINF, true)
 +
-     val32 = float32_val(a);
++static bool trans_VSWP(DisasContext *s, arg_2misc *a)
-     sbit = 0x80000000ULL & val32;
++{
-     exp = extract32(val32, 23, 8);
++    TCGv_i64 rm, rd;
-@@ -XXX,XX +XXX,XX @@ float64 HELPER(frecpx_f64)(float64 a, void *fpstp)
++    int pass;
          return nan;
      }
 +    a = float64_squash_input_denormal(a, fpst);
 +
-     val64 = float64_val(a);
++    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
-     sbit = 0x8000000000000000ULL & val64;
++        return false;
-     exp = extract64(float64_val(a), 52, 11);
++    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if (a->size != 0) {
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    rm = tcg_temp_new_i64();
 +    rd = tcg_temp_new_i64();
 +    for (pass = 0; pass < (a->q ? 2 : 1); pass++) {
 +        neon_load_reg64(rm, a->vm + pass);
 +        neon_load_reg64(rd, a->vd + pass);
 +        neon_store_reg64(rm, a->vd + pass);
 +        neon_store_reg64(rd, a->vm + pass);
 +    }
 +    tcg_temp_free_i64(rm);
 +    tcg_temp_free_i64(rd);
 +
 +    return true;
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                  case NEON_2RM_VCVTPS:
                  case NEON_2RM_VCVTMU:
                  case NEON_2RM_VCVTMS:
 +                case NEON_2RM_VSWP:
                      /* handled by decodetree */
                      return 1;
                  case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                      for (pass = 0; pass < (q ? 4 : 2); pass++) {
                          tmp = neon_load_reg(rm, pass);
                          switch (op) {
 -                        case NEON_2RM_VSWP:
 -                            tmp2 = neon_load_reg(rd, pass);
 -                            neon_store_reg(rm, pass, tmp2);
 -                            break;
                          case NEON_2RM_VTRN:
                              tmp2 = neon_load_reg(rd, pass);
                              switch (size) {
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 18/25] Make flatview_access_valid() take a MemTxAttrs argument
+[PULL 22/42] target/arm: Convert Neon VTRN to decodetree
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Convert the Neon VTRN insn to decodetree. This is the last insn in the
-add MemTxAttrs as an argument to flatview_access_valid().
+Neon data-processing group, so we can remove all the now-unused old
-Its callers now all have an attrs value to hand, so we can
+decoder framework.
-correct our earlier temporary use of MEMTXATTRS_UNSPECIFIED.
 It's possible that there's a more efficient implementation of
 VTRN, but for this conversion we just copy the existing approach.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-10-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-21-peter.maydell@linaro.org
 ---
- exec.c | 12 +++++-------
+ target/arm/neon-dp.decode       |   2 +-
-file changed, 5 insertions(+), 7 deletions(-)
+ target/arm/translate-neon.inc.c |  90 ++++++++
  target/arm/translate.c          | 363 +-------------------------------
 files changed, 93 insertions(+), 362 deletions(-)
-diff --git a/exec.c b/exec.c
+diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/neon-dp.decode
-+++ b/exec.c
++++ b/target/arm/neon-dp.decode
-@@ -XXX,XX +XXX,XX @@ static MemTxResult flatview_read(FlatView *fv, hwaddr addr,
+@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
- static MemTxResult flatview_write(FlatView *fv, hwaddr addr, MemTxAttrs attrs,
+     VNEG_F       1111 001 11 . 11 .. 01 .... 0 1111 . . 0 .... @2misc
-                                   const uint8_t *buf, int len);
- static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
+     VSWP         1111 001 11 . 11 .. 10 .... 0 0000 . . 0 .... @2misc
--                                  bool is_write);
+-
-+                                  bool is_write, MemTxAttrs attrs);
++    VTRN         1111 001 11 . 11 .. 10 .... 0 0001 . . 0 .... @2misc
+     VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
- static MemTxResult subpage_read(void *opaque, hwaddr addr, uint64_t *data,
+     VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
-                                 unsigned len, MemTxAttrs attrs)
-@@ -XXX,XX +XXX,XX @@ static bool subpage_accepts(void *opaque, hwaddr addr,
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
- #endif
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate-neon.inc.c
-     return flatview_access_valid(subpage->fv, addr + subpage->base,
++++ b/target/arm/translate-neon.inc.c
--                                 len, is_write);
+@@ -XXX,XX +XXX,XX @@ static bool trans_VSWP(DisasContext *s, arg_2misc *a)
-+                                 len, is_write, attrs);
      return true;
  }
++static void gen_neon_trn_u8(TCGv_i32 t0, TCGv_i32 t1)
- static const MemoryRegionOps subpage_ops = {
++{
-@@ -XXX,XX +XXX,XX @@ static void cpu_notify_map_clients(void)
++    TCGv_i32 rd, tmp;
 +
 +    rd = tcg_temp_new_i32();
 +    tmp = tcg_temp_new_i32();
 +
 +    tcg_gen_shli_i32(rd, t0, 8);
 +    tcg_gen_andi_i32(rd, rd, 0xff00ff00);
 +    tcg_gen_andi_i32(tmp, t1, 0x00ff00ff);
 +    tcg_gen_or_i32(rd, rd, tmp);
 +
 +    tcg_gen_shri_i32(t1, t1, 8);
 +    tcg_gen_andi_i32(t1, t1, 0x00ff00ff);
 +    tcg_gen_andi_i32(tmp, t0, 0xff00ff00);
 +    tcg_gen_or_i32(t1, t1, tmp);
 +    tcg_gen_mov_i32(t0, rd);
 +
 +    tcg_temp_free_i32(tmp);
 +    tcg_temp_free_i32(rd);
 +}
 +
 +static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
 +{
 +    TCGv_i32 rd, tmp;
 +
 +    rd = tcg_temp_new_i32();
 +    tmp = tcg_temp_new_i32();
 +
 +    tcg_gen_shli_i32(rd, t0, 16);
 +    tcg_gen_andi_i32(tmp, t1, 0xffff);
 +    tcg_gen_or_i32(rd, rd, tmp);
 +    tcg_gen_shri_i32(t1, t1, 16);
 +    tcg_gen_andi_i32(tmp, t0, 0xffff0000);
 +    tcg_gen_or_i32(t1, t1, tmp);
 +    tcg_gen_mov_i32(t0, rd);
 +
 +    tcg_temp_free_i32(tmp);
 +    tcg_temp_free_i32(rd);
 +}
 +
 +static bool trans_VTRN(DisasContext *s, arg_2misc *a)
 +{
 +    TCGv_i32 tmp, tmp2;
 +    int pass;
 +
 +    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 +        return false;
 +    }
 +
 +    /* UNDEF accesses to D16-D31 if they don't exist. */
 +    if (!dc_isar_feature(aa32_simd_r32, s) &&
 +        ((a->vd | a->vm) & 0x10)) {
 +        return false;
 +    }
 +
 +    if ((a->vd | a->vm) & a->q) {
 +        return false;
 +    }
 +
 +    if (a->size == 3) {
 +        return false;
 +    }
 +
 +    if (!vfp_access_check(s)) {
 +        return true;
 +    }
 +
 +    if (a->size == 2) {
 +        for (pass = 0; pass < (a->q ? 4 : 2); pass += 2) {
 +            tmp = neon_load_reg(a->vm, pass);
 +            tmp2 = neon_load_reg(a->vd, pass + 1);
 +            neon_store_reg(a->vm, pass, tmp2);
 +            neon_store_reg(a->vd, pass + 1, tmp);
 +        }
 +    } else {
 +        for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
 +            tmp = neon_load_reg(a->vm, pass);
 +            tmp2 = neon_load_reg(a->vd, pass);
 +            if (a->size == 0) {
 +                gen_neon_trn_u8(tmp, tmp2);
 +            } else {
 +                gen_neon_trn_u16(tmp, tmp2);
 +            }
 +            neon_store_reg(a->vm, pass, tmp2);
 +            neon_store_reg(a->vd, pass, tmp);
 +        }
 +    }
 +    return true;
 +}
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/translate.c
 +++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_exception_return(DisasContext *s, TCGv_i32 pc)
      gen_rfe(s, pc, load_cpu_field(spsr));
  }
- static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
+-static void gen_neon_trn_u8(TCGv_i32 t0, TCGv_i32 t1)
--                                  bool is_write)
+-{
-+                                  bool is_write, MemTxAttrs attrs)
+-    TCGv_i32 rd, tmp;
 -
 -    rd = tcg_temp_new_i32();
 -    tmp = tcg_temp_new_i32();
 -
 -    tcg_gen_shli_i32(rd, t0, 8);
 -    tcg_gen_andi_i32(rd, rd, 0xff00ff00);
 -    tcg_gen_andi_i32(tmp, t1, 0x00ff00ff);
 -    tcg_gen_or_i32(rd, rd, tmp);
 -
 -    tcg_gen_shri_i32(t1, t1, 8);
 -    tcg_gen_andi_i32(t1, t1, 0x00ff00ff);
 -    tcg_gen_andi_i32(tmp, t0, 0xff00ff00);
 -    tcg_gen_or_i32(t1, t1, tmp);
 -    tcg_gen_mov_i32(t0, rd);
 -
 -    tcg_temp_free_i32(tmp);
 -    tcg_temp_free_i32(rd);
 -}
 -
 -static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
 -{
 -    TCGv_i32 rd, tmp;
 -
 -    rd = tcg_temp_new_i32();
 -    tmp = tcg_temp_new_i32();
 -
 -    tcg_gen_shli_i32(rd, t0, 16);
 -    tcg_gen_andi_i32(tmp, t1, 0xffff);
 -    tcg_gen_or_i32(rd, rd, tmp);
 -    tcg_gen_shri_i32(t1, t1, 16);
 -    tcg_gen_andi_i32(tmp, t0, 0xffff0000);
 -    tcg_gen_or_i32(t1, t1, tmp);
 -    tcg_gen_mov_i32(t0, rd);
 -
 -    tcg_temp_free_i32(tmp);
 -    tcg_temp_free_i32(rd);
 -}
 -
 -/* Symbolic constants for op fields for Neon 2-register miscellaneous.
 - * The values correspond to bits [17:16,10:7]; see the ARM ARM DDI0406B
 - * table A7-13.
 - */
 -#define NEON_2RM_VREV64 0
 -#define NEON_2RM_VREV32 1
 -#define NEON_2RM_VREV16 2
 -#define NEON_2RM_VPADDL 4
 -#define NEON_2RM_VPADDL_U 5
 -#define NEON_2RM_AESE 6 /* Includes AESD */
 -#define NEON_2RM_AESMC 7 /* Includes AESIMC */
 -#define NEON_2RM_VCLS 8
 -#define NEON_2RM_VCLZ 9
 -#define NEON_2RM_VCNT 10
 -#define NEON_2RM_VMVN 11
 -#define NEON_2RM_VPADAL 12
 -#define NEON_2RM_VPADAL_U 13
 -#define NEON_2RM_VQABS 14
 -#define NEON_2RM_VQNEG 15
 -#define NEON_2RM_VCGT0 16
 -#define NEON_2RM_VCGE0 17
 -#define NEON_2RM_VCEQ0 18
 -#define NEON_2RM_VCLE0 19
 -#define NEON_2RM_VCLT0 20
 -#define NEON_2RM_SHA1H 21
 -#define NEON_2RM_VABS 22
 -#define NEON_2RM_VNEG 23
 -#define NEON_2RM_VCGT0_F 24
 -#define NEON_2RM_VCGE0_F 25
 -#define NEON_2RM_VCEQ0_F 26
 -#define NEON_2RM_VCLE0_F 27
 -#define NEON_2RM_VCLT0_F 28
 -#define NEON_2RM_VABS_F 30
 -#define NEON_2RM_VNEG_F 31
 -#define NEON_2RM_VSWP 32
 -#define NEON_2RM_VTRN 33
 -#define NEON_2RM_VUZP 34
 -#define NEON_2RM_VZIP 35
 -#define NEON_2RM_VMOVN 36 /* Includes VQMOVN, VQMOVUN */
 -#define NEON_2RM_VQMOVN 37 /* Includes VQMOVUN */
 -#define NEON_2RM_VSHLL 38
 -#define NEON_2RM_SHA1SU1 39 /* Includes SHA256SU0 */
 -#define NEON_2RM_VRINTN 40
 -#define NEON_2RM_VRINTX 41
 -#define NEON_2RM_VRINTA 42
 -#define NEON_2RM_VRINTZ 43
 -#define NEON_2RM_VCVT_F16_F32 44
 -#define NEON_2RM_VRINTM 45
 -#define NEON_2RM_VCVT_F32_F16 46
 -#define NEON_2RM_VRINTP 47
 -#define NEON_2RM_VCVTAU 48
 -#define NEON_2RM_VCVTAS 49
 -#define NEON_2RM_VCVTNU 50
 -#define NEON_2RM_VCVTNS 51
 -#define NEON_2RM_VCVTPU 52
 -#define NEON_2RM_VCVTPS 53
 -#define NEON_2RM_VCVTMU 54
 -#define NEON_2RM_VCVTMS 55
 -#define NEON_2RM_VRECPE 56
 -#define NEON_2RM_VRSQRTE 57
 -#define NEON_2RM_VRECPE_F 58
 -#define NEON_2RM_VRSQRTE_F 59
 -#define NEON_2RM_VCVT_FS 60
 -#define NEON_2RM_VCVT_FU 61
 -#define NEON_2RM_VCVT_SF 62
 -#define NEON_2RM_VCVT_UF 63
 -
 -/* Each entry in this array has bit n set if the insn allows
 - * size value n (otherwise it will UNDEF). Since unallocated
 - * op values will have no bits set they always UNDEF.
 - */
 -static const uint8_t neon_2rm_sizes[] = {
 -    [NEON_2RM_VREV64] = 0x7,
 -    [NEON_2RM_VREV32] = 0x3,
 -    [NEON_2RM_VREV16] = 0x1,
 -    [NEON_2RM_VPADDL] = 0x7,
 -    [NEON_2RM_VPADDL_U] = 0x7,
 -    [NEON_2RM_AESE] = 0x1,
 -    [NEON_2RM_AESMC] = 0x1,
 -    [NEON_2RM_VCLS] = 0x7,
 -    [NEON_2RM_VCLZ] = 0x7,
 -    [NEON_2RM_VCNT] = 0x1,
 -    [NEON_2RM_VMVN] = 0x1,
 -    [NEON_2RM_VPADAL] = 0x7,
 -    [NEON_2RM_VPADAL_U] = 0x7,
 -    [NEON_2RM_VQABS] = 0x7,
 -    [NEON_2RM_VQNEG] = 0x7,
 -    [NEON_2RM_VCGT0] = 0x7,
 -    [NEON_2RM_VCGE0] = 0x7,
 -    [NEON_2RM_VCEQ0] = 0x7,
 -    [NEON_2RM_VCLE0] = 0x7,
 -    [NEON_2RM_VCLT0] = 0x7,
 -    [NEON_2RM_SHA1H] = 0x4,
 -    [NEON_2RM_VABS] = 0x7,
 -    [NEON_2RM_VNEG] = 0x7,
 -    [NEON_2RM_VCGT0_F] = 0x4,
 -    [NEON_2RM_VCGE0_F] = 0x4,
 -    [NEON_2RM_VCEQ0_F] = 0x4,
 -    [NEON_2RM_VCLE0_F] = 0x4,
 -    [NEON_2RM_VCLT0_F] = 0x4,
 -    [NEON_2RM_VABS_F] = 0x4,
 -    [NEON_2RM_VNEG_F] = 0x4,
 -    [NEON_2RM_VSWP] = 0x1,
 -    [NEON_2RM_VTRN] = 0x7,
 -    [NEON_2RM_VUZP] = 0x7,
 -    [NEON_2RM_VZIP] = 0x7,
 -    [NEON_2RM_VMOVN] = 0x7,
 -    [NEON_2RM_VQMOVN] = 0x7,
 -    [NEON_2RM_VSHLL] = 0x7,
 -    [NEON_2RM_SHA1SU1] = 0x4,
 -    [NEON_2RM_VRINTN] = 0x4,
 -    [NEON_2RM_VRINTX] = 0x4,
 -    [NEON_2RM_VRINTA] = 0x4,
 -    [NEON_2RM_VRINTZ] = 0x4,
 -    [NEON_2RM_VCVT_F16_F32] = 0x2,
 -    [NEON_2RM_VRINTM] = 0x4,
 -    [NEON_2RM_VCVT_F32_F16] = 0x2,
 -    [NEON_2RM_VRINTP] = 0x4,
 -    [NEON_2RM_VCVTAU] = 0x4,
 -    [NEON_2RM_VCVTAS] = 0x4,
 -    [NEON_2RM_VCVTNU] = 0x4,
 -    [NEON_2RM_VCVTNS] = 0x4,
 -    [NEON_2RM_VCVTPU] = 0x4,
 -    [NEON_2RM_VCVTPS] = 0x4,
 -    [NEON_2RM_VCVTMU] = 0x4,
 -    [NEON_2RM_VCVTMS] = 0x4,
 -    [NEON_2RM_VRECPE] = 0x4,
 -    [NEON_2RM_VRSQRTE] = 0x4,
 -    [NEON_2RM_VRECPE_F] = 0x4,
 -    [NEON_2RM_VRSQRTE_F] = 0x4,
 -    [NEON_2RM_VCVT_FS] = 0x4,
 -    [NEON_2RM_VCVT_FU] = 0x4,
 -    [NEON_2RM_VCVT_SF] = 0x4,
 -    [NEON_2RM_VCVT_UF] = 0x4,
 -};
 -
  static void gen_gvec_fn3_qc(uint32_t rd_ofs, uint32_t rn_ofs, uint32_t rm_ofs,
                              uint32_t opr_sz, uint32_t max_sz,
                              gen_helper_gvec_3_ptr *fn)
@@ -XXX,XX +XXX,XX @@ void gen_gvec_uaba(unsigned vece, uint32_t rd_ofs, uint32_t rn_ofs,
      tcg_gen_gvec_3(rd_ofs, rn_ofs, rm_ofs, opr_sz, max_sz, &ops[vece]);
  }
 -/* Translate a NEON data processing instruction.  Return nonzero if the
 -   instruction is invalid.
 -   We process data in a mixture of 32-bit and 64-bit chunks.
 -   Mostly we use 32-bit chunks so we can use normal scalar instructions.  */
 -
 -static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
 -{
 -    int op;
 -    int q;
 -    int rd, rm;
 -    int size;
 -    int pass;
 -    int u;
 -    TCGv_i32 tmp, tmp2;
 -
 -    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
 -        return 1;
 -    }
 -
 -    /* FIXME: this access check should not take precedence over UNDEF
 -     * for invalid encodings; we will generate incorrect syndrome information
 -     * for attempts to execute invalid vfp/neon encodings with FP disabled.
 -     */
 -    if (s->fp_excp_el) {
 -        gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
 -                           syn_simd_access_trap(1, 0xe, false), s->fp_excp_el);
 -        return 0;
 -    }
 -
 -    if (!s->vfp_enabled)
 -      return 1;
 -    q = (insn & (1 << 6)) != 0;
 -    u = (insn >> 24) & 1;
 -    VFP_DREG_D(rd, insn);
 -    VFP_DREG_M(rm, insn);
 -    size = (insn >> 20) & 3;
 -
 -    if ((insn & (1 << 23)) == 0) {
 -        /* Three register same length: handled by decodetree */
 -        return 1;
 -    } else if (insn & (1 << 4)) {
 -        /* Two registers and shift or reg and imm: handled by decodetree */
 -        return 1;
 -    } else { /* (insn & 0x00800010 == 0x00800000) */
 -        if (size != 3) {
 -            /*
 -             * Three registers of different lengths, or two registers and
 -             * a scalar: handled by decodetree
 -             */
 -            return 1;
 -        } else { /* size == 3 */
 -            if (!u) {
 -                /* Extract: handled by decodetree */
 -                return 1;
 -            } else if ((insn & (1 << 11)) == 0) {
 -                /* Two register misc.  */
 -                op = ((insn >> 12) & 0x30) | ((insn >> 7) & 0xf);
 -                size = (insn >> 18) & 3;
 -                /* UNDEF for unknown op values and bad op-size combinations */
 -                if ((neon_2rm_sizes[op] & (1 << size)) == 0) {
 -                    return 1;
 -                }
 -                if (q && ((rm | rd) & 1)) {
 -                    return 1;
 -                }
 -                switch (op) {
 -                case NEON_2RM_VREV64:
 -                case NEON_2RM_VPADDL: case NEON_2RM_VPADDL_U:
 -                case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
 -                case NEON_2RM_VUZP:
 -                case NEON_2RM_VZIP:
 -                case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
 -                case NEON_2RM_VSHLL:
 -                case NEON_2RM_VCVT_F16_F32:
 -                case NEON_2RM_VCVT_F32_F16:
 -                case NEON_2RM_VMVN:
 -                case NEON_2RM_VNEG:
 -                case NEON_2RM_VABS:
 -                case NEON_2RM_VCEQ0:
 -                case NEON_2RM_VCGT0:
 -                case NEON_2RM_VCLE0:
 -                case NEON_2RM_VCGE0:
 -                case NEON_2RM_VCLT0:
 -                case NEON_2RM_AESE: case NEON_2RM_AESMC:
 -                case NEON_2RM_SHA1H:
 -                case NEON_2RM_SHA1SU1:
 -                case NEON_2RM_VREV32:
 -                case NEON_2RM_VREV16:
 -                case NEON_2RM_VCLS:
 -                case NEON_2RM_VCLZ:
 -                case NEON_2RM_VCNT:
 -                case NEON_2RM_VABS_F:
 -                case NEON_2RM_VNEG_F:
 -                case NEON_2RM_VRECPE:
 -                case NEON_2RM_VRSQRTE:
 -                case NEON_2RM_VQABS:
 -                case NEON_2RM_VQNEG:
 -                case NEON_2RM_VRECPE_F:
 -                case NEON_2RM_VRSQRTE_F:
 -                case NEON_2RM_VCVT_FS:
 -                case NEON_2RM_VCVT_FU:
 -                case NEON_2RM_VCVT_SF:
 -                case NEON_2RM_VCVT_UF:
 -                case NEON_2RM_VRINTX:
 -                case NEON_2RM_VCGT0_F:
 -                case NEON_2RM_VCGE0_F:
 -                case NEON_2RM_VCEQ0_F:
 -                case NEON_2RM_VCLE0_F:
 -                case NEON_2RM_VCLT0_F:
 -                case NEON_2RM_VRINTN:
 -                case NEON_2RM_VRINTA:
 -                case NEON_2RM_VRINTM:
 -                case NEON_2RM_VRINTP:
 -                case NEON_2RM_VRINTZ:
 -                case NEON_2RM_VCVTAU:
 -                case NEON_2RM_VCVTAS:
 -                case NEON_2RM_VCVTNU:
 -                case NEON_2RM_VCVTNS:
 -                case NEON_2RM_VCVTPU:
 -                case NEON_2RM_VCVTPS:
 -                case NEON_2RM_VCVTMU:
 -                case NEON_2RM_VCVTMS:
 -                case NEON_2RM_VSWP:
 -                    /* handled by decodetree */
 -                    return 1;
 -                case NEON_2RM_VTRN:
 -                    if (size == 2) {
 -                        int n;
 -                        for (n = 0; n < (q ? 4 : 2); n += 2) {
 -                            tmp = neon_load_reg(rm, n);
 -                            tmp2 = neon_load_reg(rd, n + 1);
 -                            neon_store_reg(rm, n, tmp2);
 -                            neon_store_reg(rd, n + 1, tmp);
 -                        }
 -                    } else {
 -                        goto elementwise;
 -                    }
 -                    break;
 -
 -                default:
 -                elementwise:
 -                    for (pass = 0; pass < (q ? 4 : 2); pass++) {
 -                        tmp = neon_load_reg(rm, pass);
 -                        switch (op) {
 -                        case NEON_2RM_VTRN:
 -                            tmp2 = neon_load_reg(rd, pass);
 -                            switch (size) {
 -                            case 0: gen_neon_trn_u8(tmp, tmp2); break;
 -                            case 1: gen_neon_trn_u16(tmp, tmp2); break;
 -                            default: abort();
 -                            }
 -                            neon_store_reg(rm, pass, tmp2);
 -                            break;
 -                        default:
 -                            /* Reserved op values were caught by the
 -                             * neon_2rm_sizes[] check earlier.
 -                             */
 -                            abort();
 -                        }
 -                        neon_store_reg(rd, pass, tmp);
 -                    }
 -                    break;
 -                }
 -            } else {
 -                /* VTBL, VTBX, VDUP: handled by decodetree */
 -                return 1;
 -            }
 -        }
 -    }
 -    return 0;
 -}
 -
  static int disas_coproc_insn(DisasContext *s, uint32_t insn)
  {
-     MemoryRegion *mr;
+     int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
-     hwaddr l, xlat;
+@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
          mr = flatview_translate(fv, addr, &xlat, &l, is_write);
          if (!memory_access_is_direct(mr, is_write)) {
              l = memory_access_size(mr, l, addr);
 -            /* When our callers all have attrs we'll pass them through here */
 -            if (!memory_region_access_valid(mr, xlat, l, is_write,
 -                                            MEMTXATTRS_UNSPECIFIED)) {
 +            if (!memory_region_access_valid(mr, xlat, l, is_write, attrs)) {
                  return false;
              }
          }
-@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr,
+         /* fall back to legacy decoder */
-     rcu_read_lock();
+-        if (((insn >> 25) & 7) == 1) {
-     fv = address_space_to_flatview(as);
+-            /* NEON Data processing.  */
--    result = flatview_access_valid(fv, addr, len, is_write);
+-            if (disas_neon_data_insn(s, insn)) {
-+    result = flatview_access_valid(fv, addr, len, is_write, attrs);
+-                goto illegal_op;
-     rcu_read_unlock();
+-            }
-     return result;
+-            return;
- }
+-        }
          if ((insn & 0x0e000f00) == 0x0c000100) {
              if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
                  /* iWMMXt register transfer.  */
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
              break;
          }
          if (((insn >> 24) & 3) == 3) {
 -            /* Translate into the equivalent ARM encoding.  */
 -            insn = (insn & 0xe2ffffff) | ((insn & (1 << 28)) >> 4) | (1 << 28);
 -            if (disas_neon_data_insn(s, insn)) {
 -                goto illegal_op;
 -            }
 +            /* Neon DP, but failed disas_neon_dp() */
 +            goto illegal_op;
          } else if (((insn >> 8) & 0xe) == 10) {
              /* VFP, but failed disas_vfp.  */
              goto illegal_op;
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 14/25] Make address_space_access_valid() take a MemTxAttrs argument
+[PULL 23/42] target/arm: Move some functions used only in translate-neon.inc.c to that file
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+The functions neon_element_offset(), neon_load_element(),
-add MemTxAttrs as an argument to address_space_access_valid().
+neon_load_element64(), neon_store_element() and
-Its callers either have an attrs value to hand, or don't care
+neon_store_element64() are used only in the translate-neon.inc.c
-and can use MEMTXATTRS_UNSPECIFIED.
+file, so move their definitions there.
 Since the .inc.c file is #included in translate.c this doesn't make
 much difference currently, but it's a more logical place to put the
 functions and it might be helpful if we ever decide to try to make
 the .inc.c files genuinely separate compilation units.
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-6-peter.maydell@linaro.org
+Message-id: 20200616170844.13318-22-peter.maydell@linaro.org
 ---
- include/exec/memory.h      | 4 +++-
+ target/arm/translate-neon.inc.c | 101 ++++++++++++++++++++++++++++++++
- include/sysemu/dma.h       | 3 ++-
+ target/arm/translate.c          | 101 --------------------------------
- exec.c                     | 3 ++-
+files changed, 101 insertions(+), 101 deletions(-)
- target/s390x/diag.c        | 6 ++++--
- target/s390x/excp_helper.c | 3 ++-
+diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
  target/s390x/mmu_helper.c  | 3 ++-
  target/s390x/sigp.c        | 3 ++-
 files changed, 17 insertions(+), 8 deletions(-)
 diff --git a/include/exec/memory.h b/include/exec/memory.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/translate-neon.inc.c
-+++ b/include/exec/memory.h
++++ b/target/arm/translate-neon.inc.c
-@@ -XXX,XX +XXX,XX @@ static inline MemoryRegion *address_space_translate(AddressSpace *as,
+@@ -XXX,XX +XXX,XX @@ static inline int rsub_8(DisasContext *s, int x)
-  * @addr: address within that address space
+ #include "decode-neon-ls.inc.c"
-  * @len: length of the area to be checked
+ #include "decode-neon-shared.inc.c"
-  * @is_write: indicates the transfer direction
-+ * @attrs: memory attributes
++/* Return the offset of a 2**SIZE piece of a NEON register, at index ELE,
-  */
++ * where 0 is the least significant end of the register.
--bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len, bool is_write);
++ */
-+bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len,
++static inline long
-+                                bool is_write, MemTxAttrs attrs);
++neon_element_offset(int reg, int element, MemOp size)
++{
- /* address_space_map: map a physical memory region into a host virtual address
++    int element_size = 1 << size;
-  *
++    int ofs = element * element_size;
-diff --git a/include/sysemu/dma.h b/include/sysemu/dma.h
++#ifdef HOST_WORDS_BIGENDIAN
 +    /* Calculate the offset assuming fully little-endian,
 +     * then XOR to account for the order of the 8-byte units.
 +     */
 +    if (element_size < 8) {
 +        ofs ^= 8 - element_size;
 +    }
 +#endif
 +    return neon_reg_offset(reg, 0) + ofs;
 +}
 +
 +static void neon_load_element(TCGv_i32 var, int reg, int ele, MemOp mop)
 +{
 +    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
 +
 +    switch (mop) {
 +    case MO_UB:
 +        tcg_gen_ld8u_i32(var, cpu_env, offset);
 +        break;
 +    case MO_UW:
 +        tcg_gen_ld16u_i32(var, cpu_env, offset);
 +        break;
 +    case MO_UL:
 +        tcg_gen_ld_i32(var, cpu_env, offset);
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +}
 +
 +static void neon_load_element64(TCGv_i64 var, int reg, int ele, MemOp mop)
 +{
 +    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
 +
 +    switch (mop) {
 +    case MO_UB:
 +        tcg_gen_ld8u_i64(var, cpu_env, offset);
 +        break;
 +    case MO_UW:
 +        tcg_gen_ld16u_i64(var, cpu_env, offset);
 +        break;
 +    case MO_UL:
 +        tcg_gen_ld32u_i64(var, cpu_env, offset);
 +        break;
 +    case MO_Q:
 +        tcg_gen_ld_i64(var, cpu_env, offset);
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +}
 +
 +static void neon_store_element(int reg, int ele, MemOp size, TCGv_i32 var)
 +{
 +    long offset = neon_element_offset(reg, ele, size);
 +
 +    switch (size) {
 +    case MO_8:
 +        tcg_gen_st8_i32(var, cpu_env, offset);
 +        break;
 +    case MO_16:
 +        tcg_gen_st16_i32(var, cpu_env, offset);
 +        break;
 +    case MO_32:
 +        tcg_gen_st_i32(var, cpu_env, offset);
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +}
 +
 +static void neon_store_element64(int reg, int ele, MemOp size, TCGv_i64 var)
 +{
 +    long offset = neon_element_offset(reg, ele, size);
 +
 +    switch (size) {
 +    case MO_8:
 +        tcg_gen_st8_i64(var, cpu_env, offset);
 +        break;
 +    case MO_16:
 +        tcg_gen_st16_i64(var, cpu_env, offset);
 +        break;
 +    case MO_32:
 +        tcg_gen_st32_i64(var, cpu_env, offset);
 +        break;
 +    case MO_64:
 +        tcg_gen_st_i64(var, cpu_env, offset);
 +        break;
 +    default:
 +        g_assert_not_reached();
 +    }
 +}
 +
  static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
  {
      int opr_sz;
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
---- a/include/sysemu/dma.h
+--- a/target/arm/translate.c
-+++ b/include/sysemu/dma.h
++++ b/target/arm/translate.c
-@@ -XXX,XX +XXX,XX @@ static inline bool dma_memory_valid(AddressSpace *as,
+@@ -XXX,XX +XXX,XX @@ neon_reg_offset (int reg, int n)
-                                     DMADirection dir)
+     return vfp_reg_offset(0, sreg);
  {
      return address_space_access_valid(as, addr, len,
 -                                      dir == DMA_DIRECTION_FROM_DEVICE);
 +                                      dir == DMA_DIRECTION_FROM_DEVICE,
 +                                      MEMTXATTRS_UNSPECIFIED);
  }
- static inline int dma_memory_rw_relaxed(AddressSpace *as, dma_addr_t addr,
+-/* Return the offset of a 2**SIZE piece of a NEON register, at index ELE,
-diff --git a/exec.c b/exec.c
+- * where 0 is the least significant end of the register.
-index XXXXXXX..XXXXXXX 100644
+- */
---- a/exec.c
+-static inline long
-+++ b/exec.c
+-neon_element_offset(int reg, int element, MemOp size)
-@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
+-{
 -    int element_size = 1 << size;
 -    int ofs = element * element_size;
 -#ifdef HOST_WORDS_BIGENDIAN
 -    /* Calculate the offset assuming fully little-endian,
 -     * then XOR to account for the order of the 8-byte units.
 -     */
 -    if (element_size < 8) {
 -        ofs ^= 8 - element_size;
 -    }
 -#endif
 -    return neon_reg_offset(reg, 0) + ofs;
 -}
 -
  static TCGv_i32 neon_load_reg(int reg, int pass)
  {
      TCGv_i32 tmp = tcg_temp_new_i32();
@@ -XXX,XX +XXX,XX @@ static TCGv_i32 neon_load_reg(int reg, int pass)
      return tmp;
  }
- bool address_space_access_valid(AddressSpace *as, hwaddr addr,
+-static void neon_load_element(TCGv_i32 var, int reg, int ele, MemOp mop)
--                                int len, bool is_write)
+-{
-+                                int len, bool is_write,
+-    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
-+                                MemTxAttrs attrs)
+-
- {
+-    switch (mop) {
-     FlatView *fv;
+-    case MO_UB:
-     bool result;
+-        tcg_gen_ld8u_i32(var, cpu_env, offset);
-diff --git a/target/s390x/diag.c b/target/s390x/diag.c
+-        break;
-index XXXXXXX..XXXXXXX 100644
+-    case MO_UW:
---- a/target/s390x/diag.c
+-        tcg_gen_ld16u_i32(var, cpu_env, offset);
-+++ b/target/s390x/diag.c
+-        break;
-@@ -XXX,XX +XXX,XX @@ void handle_diag_308(CPUS390XState *env, uint64_t r1, uint64_t r3, uintptr_t ra)
+-    case MO_UL:
-             return;
+-        tcg_gen_ld_i32(var, cpu_env, offset);
-         }
+-        break;
-         if (!address_space_access_valid(&address_space_memory, addr,
+-    default:
--                                        sizeof(IplParameterBlock), false)) {
+-        g_assert_not_reached();
-+                                        sizeof(IplParameterBlock), false,
+-    }
-+                                        MEMTXATTRS_UNSPECIFIED)) {
+-}
-             s390_program_interrupt(env, PGM_ADDRESSING, ILEN_AUTO, ra);
+-
-             return;
+-static void neon_load_element64(TCGv_i64 var, int reg, int ele, MemOp mop)
-         }
+-{
-@@ -XXX,XX +XXX,XX @@ out:
+-    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
-             return;
+-
-         }
+-    switch (mop) {
-         if (!address_space_access_valid(&address_space_memory, addr,
+-    case MO_UB:
--                                        sizeof(IplParameterBlock), true)) {
+-        tcg_gen_ld8u_i64(var, cpu_env, offset);
-+                                        sizeof(IplParameterBlock), true,
+-        break;
-+                                        MEMTXATTRS_UNSPECIFIED)) {
+-    case MO_UW:
-             s390_program_interrupt(env, PGM_ADDRESSING, ILEN_AUTO, ra);
+-        tcg_gen_ld16u_i64(var, cpu_env, offset);
-             return;
+-        break;
-         }
+-    case MO_UL:
-diff --git a/target/s390x/excp_helper.c b/target/s390x/excp_helper.c
+-        tcg_gen_ld32u_i64(var, cpu_env, offset);
-index XXXXXXX..XXXXXXX 100644
+-        break;
---- a/target/s390x/excp_helper.c
+-    case MO_Q:
-+++ b/target/s390x/excp_helper.c
+-        tcg_gen_ld_i64(var, cpu_env, offset);
-@@ -XXX,XX +XXX,XX @@ int s390_cpu_handle_mmu_fault(CPUState *cs, vaddr orig_vaddr, int size,
+-        break;
+-    default:
-     /* check out of RAM access */
+-        g_assert_not_reached();
-     if (!address_space_access_valid(&address_space_memory, raddr,
+-    }
--                                    TARGET_PAGE_SIZE, rw)) {
+-}
-+                                    TARGET_PAGE_SIZE, rw,
+-
-+                                    MEMTXATTRS_UNSPECIFIED)) {
+ static void neon_store_reg(int reg, int pass, TCGv_i32 var)
-         DPRINTF("%s: raddr %" PRIx64 " > ram_size %" PRIx64 "\n", __func__,
+ {
-                 (uint64_t)raddr, (uint64_t)ram_size);
+     tcg_gen_st_i32(var, cpu_env, neon_reg_offset(reg, pass));
-         trigger_pgm_exception(env, PGM_ADDRESSING, ILEN_AUTO);
+     tcg_temp_free_i32(var);
-diff --git a/target/s390x/mmu_helper.c b/target/s390x/mmu_helper.c
+ }
-index XXXXXXX..XXXXXXX 100644
---- a/target/s390x/mmu_helper.c
+-static void neon_store_element(int reg, int ele, MemOp size, TCGv_i32 var)
-+++ b/target/s390x/mmu_helper.c
+-{
-@@ -XXX,XX +XXX,XX @@ static int translate_pages(S390CPU *cpu, vaddr addr, int nr_pages,
+-    long offset = neon_element_offset(reg, ele, size);
-             return ret;
+-
-         }
+-    switch (size) {
-         if (!address_space_access_valid(&address_space_memory, pages[i],
+-    case MO_8:
--                                        TARGET_PAGE_SIZE, is_write)) {
+-        tcg_gen_st8_i32(var, cpu_env, offset);
-+                                        TARGET_PAGE_SIZE, is_write,
+-        break;
-+                                        MEMTXATTRS_UNSPECIFIED)) {
+-    case MO_16:
-             trigger_access_exception(env, PGM_ADDRESSING, ILEN_AUTO, 0);
+-        tcg_gen_st16_i32(var, cpu_env, offset);
-             return -EFAULT;
+-        break;
-         }
+-    case MO_32:
-diff --git a/target/s390x/sigp.c b/target/s390x/sigp.c
+-        tcg_gen_st_i32(var, cpu_env, offset);
-index XXXXXXX..XXXXXXX 100644
+-        break;
---- a/target/s390x/sigp.c
+-    default:
-+++ b/target/s390x/sigp.c
+-        g_assert_not_reached();
-@@ -XXX,XX +XXX,XX @@ static void sigp_set_prefix(CPUState *cs, run_on_cpu_data arg)
+-    }
-     cpu_synchronize_state(cs);
+-}
+-
-     if (!address_space_access_valid(&address_space_memory, addr,
+-static void neon_store_element64(int reg, int ele, MemOp size, TCGv_i64 var)
--                                    sizeof(struct LowCore), false)) {
+-{
-+                                    sizeof(struct LowCore), false,
+-    long offset = neon_element_offset(reg, ele, size);
-+                                    MEMTXATTRS_UNSPECIFIED)) {
+-
-         set_sigp_status(si, SIGP_STAT_INVALID_PARAMETER);
+-    switch (size) {
-         return;
+-    case MO_8:
-     }
+-        tcg_gen_st8_i64(var, cpu_env, offset);
 -        break;
 -    case MO_16:
 -        tcg_gen_st16_i64(var, cpu_env, offset);
 -        break;
 -    case MO_32:
 -        tcg_gen_st32_i64(var, cpu_env, offset);
 -        break;
 -    case MO_64:
 -        tcg_gen_st_i64(var, cpu_env, offset);
 -        break;
 -    default:
 -        g_assert_not_reached();
 -    }
 -}
 -
  static inline void neon_load_reg64(TCGv_i64 var, int reg)
  {
      tcg_gen_ld_i64(var, cpu_env, vfp_reg_offset(1, reg));
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 20/25] Make address_space_get_iotlb_entry() take a MemTxAttrs argument
+[PULL 24/42] target/arm: Remove unnecessary gen_io_end() calls
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+Since commit ba3e7926691ed3 it has been unnecessary for target code
-add MemTxAttrs as an argument to address_space_get_iotlb_entry().
+to call gen_io_end() after an IO instruction in icount mode; it is
 sufficient to call gen_io_start() before it and to force the end of
 the TB.
 Many now-unnecessary calls to gen_io_end() were removed in commit
 e9b10c6491153b, but some were missed or accidentally added later.
 Remove unneeded calls from the arm target:
  * the call in the handling of exception-return-via-LDM is
    unnecessary, and the code is already forcing end-of-TB
  * the call in the VFP access check code is more complicated:
    we weren't ending the TB, so we need to add the code to
    force that by setting DISAS_UPDATE
  * the doc comment for ARM_CP_IO doesn't need to mention
    gen_io_end() any more
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
 Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-12-peter.maydell@linaro.org
+Reviewed-by: Pavel Dovgalyuk <Pavel.Dovgaluk@ispras.ru>
 Message-id: 20200619170324.12093-1-peter.maydell@linaro.org
 ---
- include/exec/memory.h | 2 +-
+ target/arm/cpu.h               | 2 +-
- exec.c                | 2 +-
+ target/arm/translate-vfp.inc.c | 7 +++----
- hw/virtio/vhost.c     | 3 ++-
+ target/arm/translate.c         | 3 ---
-files changed, 4 insertions(+), 3 deletions(-)
+files changed, 4 insertions(+), 8 deletions(-)
-diff --git a/include/exec/memory.h b/include/exec/memory.h
+diff --git a/target/arm/cpu.h b/target/arm/cpu.h
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/memory.h
+--- a/target/arm/cpu.h
-+++ b/include/exec/memory.h
++++ b/target/arm/cpu.h
-@@ -XXX,XX +XXX,XX @@ void address_space_cache_destroy(MemoryRegionCache *cache);
+@@ -XXX,XX +XXX,XX @@ static inline uint64_t cpreg_to_kvm_id(uint32_t cpregid)
-  * entry. Should be called from an RCU critical section.
+  * migration or KVM state synchronization. (Typically this is for "registers"
-  */
+  * which are actually used as instructions for cache maintenance and so on.)
- IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
+  * IO indicates that this register does I/O and therefore its accesses
--                                            bool is_write);
+- * need to be surrounded by gen_io_start()/gen_io_end(). In particular,
-+                                            bool is_write, MemTxAttrs attrs);
++ * need to be marked with gen_io_start() and also end the TB. In particular,
+  * registers which implement clocks or timers require this.
- /* address_space_translate: translate an address range into an address space
+  * RAISES_EXC is for when the read or write hook might raise an exception;
-  * into a MemoryRegion and an address range into that section.  Should be
+  * the generated code will synchronize the CPU state before calling the hook
-diff --git a/exec.c b/exec.c
+diff --git a/target/arm/translate-vfp.inc.c b/target/arm/translate-vfp.inc.c
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/target/arm/translate-vfp.inc.c
-+++ b/exec.c
++++ b/target/arm/translate-vfp.inc.c
-@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
+@@ -XXX,XX +XXX,XX @@ static bool full_vfp_access_check(DisasContext *s, bool ignore_vfp_enabled)
+         if (s->v7m_lspact) {
- /* Called from RCU critical section */
+             /*
- IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
+              * Lazy state saving affects external memory and also the NVIC,
--                                            bool is_write)
+-             * so we must mark it as an IO operation for icount.
-+                                            bool is_write, MemTxAttrs attrs)
++             * so we must mark it as an IO operation for icount (and cause
- {
++             * this to be the last insn in the TB).
-     MemoryRegionSection section;
+              */
-     hwaddr xlat, page_mask;
+             if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
-diff --git a/hw/virtio/vhost.c b/hw/virtio/vhost.c
++                s->base.is_jmp = DISAS_UPDATE;
                  gen_io_start();
              }
              gen_helper_v7m_preserve_fp_state(cpu_env);
 -            if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
 -                gen_io_end();
 -            }
              /*
               * If the preserve_fp_state helper doesn't throw an exception
               * then it will clear LSPACT; we don't need to repeat this for
 diff --git a/target/arm/translate.c b/target/arm/translate.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/virtio/vhost.c
+--- a/target/arm/translate.c
-+++ b/hw/virtio/vhost.c
++++ b/target/arm/translate.c
-@@ -XXX,XX +XXX,XX @@ int vhost_device_iotlb_miss(struct vhost_dev *dev, uint64_t iova, int write)
+@@ -XXX,XX +XXX,XX @@ static bool do_ldm(DisasContext *s, arg_ldst_block *a, int min_n)
-     trace_vhost_iotlb_miss(dev, 1);
+             gen_io_start();
+         }
-     iotlb = address_space_get_iotlb_entry(dev->vdev->dma_as,
+         gen_helper_cpsr_write_eret(cpu_env, tmp);
--                                          iova, write);
+-        if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
-+                                          iova, write,
+-            gen_io_end();
-+                                          MEMTXATTRS_UNSPECIFIED);
+-        }
-     if (iotlb.target_as != NULL) {
+         tcg_temp_free_i32(tmp);
-         ret = vhost_memory_region_lookup(dev, iotlb.translated_addr,
+         /* Must exit loop to check un-masked IRQs */
-                                          &uaddr, &len);
+         s->base.is_jmp = DISAS_EXIT;
 --
-.17.1
+.20.1

-New patch
+[PULL 25/42] target/arm: Remove dead code relating to SABA and UABA
+In commit cfdb2c0c95ae9205b0 ("target/arm: Vectorize SABA/UABA") we
+replaced the old handling of SABA/UABA with a vectorized implementation
+which returns early rather than falling into the loop-ever-elements
+code. We forgot to delete the part of the old looping code that
+did the accumulate step, and Coverity correctly warns (CID 1428955)
+that this code is now dead. Delete it.
+Fixes: cfdb2c0c95ae9205b0
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
+Message-id: 20200619171547.29780-1-peter.maydell@linaro.org
+---
+ target/arm/translate-a64.c | 12 ------------
+file changed, 12 deletions(-)
+diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
+index XXXXXXX..XXXXXXX 100644
+--- a/target/arm/translate-a64.c
++++ b/target/arm/translate-a64.c
+@@ -XXX,XX +XXX,XX @@ static void disas_simd_3same_int(DisasContext *s, uint32_t insn)
+                 genfn(tcg_res, tcg_op1, tcg_op2);
+             }
+-            if (opcode == 0xf) {
+-                /* SABA, UABA: accumulating ops */
+-                static NeonGenTwoOpFn * const fns[3] = {
+-                    gen_helper_neon_add_u8,
+-                    gen_helper_neon_add_u16,
+-                    tcg_gen_add_i32,
+-                };
+-
+-                read_vec_element_i32(s, tcg_op1, rd, pass, MO_32);
+-                fns[size](tcg_res, tcg_op1, tcg_res);
+-            }
+-
+             write_vec_element_i32(s, tcg_res, rd, pass, MO_32);
+             tcg_temp_free_i32(tcg_res);
+--
+.20.1

-New patch
+[PULL 26/42] hw/watchdog/cmsdk-apb-watchdog: Add trace event for lock status
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Add a trace event to see when a guest disable/enable the watchdog.
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-2-f4bug@amsat.org
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/watchdog/cmsdk-apb-watchdog.c | 1 +
+ hw/watchdog/trace-events         | 1 +
+files changed, 2 insertions(+)
+diff --git a/hw/watchdog/cmsdk-apb-watchdog.c b/hw/watchdog/cmsdk-apb-watchdog.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/watchdog/cmsdk-apb-watchdog.c
++++ b/hw/watchdog/cmsdk-apb-watchdog.c
+@@ -XXX,XX +XXX,XX @@ static void cmsdk_apb_watchdog_write(void *opaque, hwaddr offset,
+         break;
+     case A_WDOGLOCK:
+         s->lock = (value != WDOG_UNLOCK_VALUE);
++        trace_cmsdk_apb_watchdog_lock(s->lock);
+         break;
+     case A_WDOGITCR:
+         if (s->is_luminary) {
+diff --git a/hw/watchdog/trace-events b/hw/watchdog/trace-events
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/watchdog/trace-events
++++ b/hw/watchdog/trace-events
+@@ -XXX,XX +XXX,XX @@
+ cmsdk_apb_watchdog_read(uint64_t offset, uint64_t data, unsigned size) "CMSDK APB watchdog read: offset 0x%" PRIx64 " data 0x%" PRIx64 " size %u"
+ cmsdk_apb_watchdog_write(uint64_t offset, uint64_t data, unsigned size) "CMSDK APB watchdog write: offset 0x%" PRIx64 " data 0x%" PRIx64 " size %u"
+ cmsdk_apb_watchdog_reset(void) "CMSDK APB watchdog: reset"
++cmsdk_apb_watchdog_lock(uint32_t lock) "CMSDK APB watchdog: lock %" PRIu32
+--
+.20.1

-New patch
+[PULL 27/42] hw/i2c/versatile_i2c: Add definitions for register addresses
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Use self-explicit definitions instead of magic values.
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-3-f4bug@amsat.org
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/i2c/versatile_i2c.c | 14 ++++++++++----
+file changed, 10 insertions(+), 4 deletions(-)
+diff --git a/hw/i2c/versatile_i2c.c b/hw/i2c/versatile_i2c.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/i2c/versatile_i2c.c
++++ b/hw/i2c/versatile_i2c.c
+@@ -XXX,XX +XXX,XX @@
+ #include "qemu/osdep.h"
+ #include "hw/sysbus.h"
+ #include "hw/i2c/bitbang_i2c.h"
++#include "hw/registerfields.h"
+ #include "qemu/log.h"
+ #include "qemu/module.h"
+@@ -XXX,XX +XXX,XX @@ typedef struct VersatileI2CState {
+     int in;
+ } VersatileI2CState;
++REG32(CONTROL_GET, 0)
++REG32(CONTROL_SET, 0)
++REG32(CONTROL_CLR, 4)
++
+ static uint64_t versatile_i2c_read(void *opaque, hwaddr offset,
+                                    unsigned size)
+ {
+     VersatileI2CState *s = (VersatileI2CState *)opaque;
+-    if (offset == 0) {
++    switch (offset) {
++    case A_CONTROL_SET:
+         return (s->out & 1) | (s->in << 1);
+-    } else {
++    default:
+         qemu_log_mask(LOG_GUEST_ERROR,
+                       "%s: Bad offset 0x%x\n", __func__, (int)offset);
+         return -1;
+@@ -XXX,XX +XXX,XX @@ static void versatile_i2c_write(void *opaque, hwaddr offset,
+     VersatileI2CState *s = (VersatileI2CState *)opaque;
+     switch (offset) {
+-    case 0:
++    case A_CONTROL_SET:
+         s->out |= value & 3;
+         break;
+-    case 4:
++    case A_CONTROL_CLR:
+         s->out &= ~value;
+         break;
+     default:
+--
+.20.1

-[Qemu-devel] [PULL 05/25] tcg: Fix helper function vs host abi for float16
+[PULL 28/42] hw/i2c/versatile_i2c: Add SCL/SDA definitions
-From: Richard Henderson <richard.henderson@linaro.org>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-Depending on the host abi, float16, aka uint16_t, values are
+Use self-explicit definitions instead of magic values.
 passed and returned either zero-extended in the host register
 or with garbage at the top of the host register.
-The tcg code generator has so far been assuming garbage, as that
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
-matches the x86 abi, but this is incorrect for other host abis.
+Message-id: 20200617072539.32686-4-f4bug@amsat.org
 Further, target/arm has so far been assuming zero-extended results,
 so that it may store the 16-bit value into a 32-bit slot with the
 high 16-bits already clear.
 Rectify both problems by mapping "f16" in the helper definition
 to uint32_t instead of (a typedef for) uint16_t.  This forces
 the host compiler to assume garbage in the upper 16 bits on input
 and to zero-extend the result on output.
 Cc: qemu-stable@nongnu.org
 Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
 Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
 Tested-by: Laurent Desnogues <laurent.desnogues@gmail.com>
 Message-id: 20180522175629.24932-1-richard.henderson@linaro.org
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- include/exec/helper-head.h |  2 +-
+ hw/i2c/versatile_i2c.c | 7 +++++--
- target/arm/helper-a64.c    | 35 +++++++++--------
+file changed, 5 insertions(+), 2 deletions(-)
  target/arm/helper.c        | 80 +++++++++++++++++++-------------------
 files changed, 59 insertions(+), 58 deletions(-)
-diff --git a/include/exec/helper-head.h b/include/exec/helper-head.h
+diff --git a/hw/i2c/versatile_i2c.c b/hw/i2c/versatile_i2c.c
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/helper-head.h
+--- a/hw/i2c/versatile_i2c.c
-+++ b/include/exec/helper-head.h
++++ b/hw/i2c/versatile_i2c.c
-@@ -XXX,XX +XXX,XX @@
+@@ -XXX,XX +XXX,XX @@ REG32(CONTROL_GET, 0)
- #define dh_ctype_int int
+ REG32(CONTROL_SET, 0)
- #define dh_ctype_i64 uint64_t
+ REG32(CONTROL_CLR, 4)
- #define dh_ctype_s64 int64_t
--#define dh_ctype_f16 float16
++#define SCL BIT(0)
-+#define dh_ctype_f16 uint32_t
++#define SDA BIT(1)
- #define dh_ctype_f32 float32
++
- #define dh_ctype_f64 float64
+ static uint64_t versatile_i2c_read(void *opaque, hwaddr offset,
- #define dh_ctype_ptr void *
+                                    unsigned size)
-diff --git a/target/arm/helper-a64.c b/target/arm/helper-a64.c
+ {
-index XXXXXXX..XXXXXXX 100644
+@@ -XXX,XX +XXX,XX @@ static void versatile_i2c_write(void *opaque, hwaddr offset,
---- a/target/arm/helper-a64.c
+         qemu_log_mask(LOG_GUEST_ERROR,
-+++ b/target/arm/helper-a64.c
+                       "%s: Bad offset 0x%x\n", __func__, (int)offset);
-@@ -XXX,XX +XXX,XX @@ static inline uint32_t float_rel_to_flags(int res)
+     }
-     return flags;
+-    bitbang_i2c_set(&s->bitbang, BITBANG_I2C_SCL, (s->out & 1) != 0);
 -    s->in = bitbang_i2c_set(&s->bitbang, BITBANG_I2C_SDA, (s->out & 2) != 0);
 +    bitbang_i2c_set(&s->bitbang, BITBANG_I2C_SCL, (s->out & SCL) != 0);
 +    s->in = bitbang_i2c_set(&s->bitbang, BITBANG_I2C_SDA, (s->out & SDA) != 0);
  }
--uint64_t HELPER(vfp_cmph_a64)(float16 x, float16 y, void *fp_status)
+ static const MemoryRegionOps versatile_i2c_ops = {
 +uint64_t HELPER(vfp_cmph_a64)(uint32_t x, uint32_t y, void *fp_status)
  {
      return float_rel_to_flags(float16_compare_quiet(x, y, fp_status));
  }
 -uint64_t HELPER(vfp_cmpeh_a64)(float16 x, float16 y, void *fp_status)
 +uint64_t HELPER(vfp_cmpeh_a64)(uint32_t x, uint32_t y, void *fp_status)
  {
      return float_rel_to_flags(float16_compare(x, y, fp_status));
  }
@@ -XXX,XX +XXX,XX @@ uint64_t HELPER(neon_cgt_f64)(float64 a, float64 b, void *fpstp)
  #define float64_three make_float64(0x4008000000000000ULL)
  #define float64_one_point_five make_float64(0x3FF8000000000000ULL)
 -float16 HELPER(recpsf_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(recpsf_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
@@ -XXX,XX +XXX,XX @@ float64 HELPER(recpsf_f64)(float64 a, float64 b, void *fpstp)
      return float64_muladd(a, b, float64_two, 0, fpst);
  }
 -float16 HELPER(rsqrtsf_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(rsqrtsf_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
@@ -XXX,XX +XXX,XX @@ uint64_t HELPER(neon_addlp_u16)(uint64_t a)
  }
  /* Floating-point reciprocal exponent - see FPRecpX in ARM ARM */
 -float16 HELPER(frecpx_f16)(float16 a, void *fpstp)
 +uint32_t HELPER(frecpx_f16)(uint32_t a, void *fpstp)
  {
      float_status *fpst = fpstp;
      uint16_t val16, sbit;
@@ -XXX,XX +XXX,XX @@ void HELPER(casp_be_parallel)(CPUARMState *env, uint32_t rs, uint64_t addr,
  #define ADVSIMD_HELPER(name, suffix) HELPER(glue(glue(advsimd_, name), suffix))
  #define ADVSIMD_HALFOP(name) \
 -float16 ADVSIMD_HELPER(name, h)(float16 a, float16 b, void *fpstp) \
 +uint32_t ADVSIMD_HELPER(name, h)(uint32_t a, uint32_t b, void *fpstp) \
  { \
      float_status *fpst = fpstp; \
      return float16_ ## name(a, b, fpst);    \
@@ -XXX,XX +XXX,XX @@ ADVSIMD_HALFOP(mulx)
  ADVSIMD_TWOHALFOP(mulx)
  /* fused multiply-accumulate */
 -float16 HELPER(advsimd_muladdh)(float16 a, float16 b, float16 c, void *fpstp)
 +uint32_t HELPER(advsimd_muladdh)(uint32_t a, uint32_t b, uint32_t c,
 +                                 void *fpstp)
  {
      float_status *fpst = fpstp;
      return float16_muladd(a, b, c, 0, fpst);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_muladd2h)(uint32_t two_a, uint32_t two_b,
  #define ADVSIMD_CMPRES(test) (test) ? 0xffff : 0
 -uint32_t HELPER(advsimd_ceq_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(advsimd_ceq_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
      int compare = float16_compare_quiet(a, b, fpst);
      return ADVSIMD_CMPRES(compare == float_relation_equal);
  }
 -uint32_t HELPER(advsimd_cge_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(advsimd_cge_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
      int compare = float16_compare(a, b, fpst);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_cge_f16)(float16 a, float16 b, void *fpstp)
                            compare == float_relation_equal);
  }
 -uint32_t HELPER(advsimd_cgt_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(advsimd_cgt_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
      int compare = float16_compare(a, b, fpst);
      return ADVSIMD_CMPRES(compare == float_relation_greater);
  }
 -uint32_t HELPER(advsimd_acge_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(advsimd_acge_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
      float16 f0 = float16_abs(a);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_acge_f16)(float16 a, float16 b, void *fpstp)
                            compare == float_relation_equal);
  }
 -uint32_t HELPER(advsimd_acgt_f16)(float16 a, float16 b, void *fpstp)
 +uint32_t HELPER(advsimd_acgt_f16)(uint32_t a, uint32_t b, void *fpstp)
  {
      float_status *fpst = fpstp;
      float16 f0 = float16_abs(a);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_acgt_f16)(float16 a, float16 b, void *fpstp)
  }
  /* round to integral */
 -float16 HELPER(advsimd_rinth_exact)(float16 x, void *fp_status)
 +uint32_t HELPER(advsimd_rinth_exact)(uint32_t x, void *fp_status)
  {
      return float16_round_to_int(x, fp_status);
  }
 -float16 HELPER(advsimd_rinth)(float16 x, void *fp_status)
 +uint32_t HELPER(advsimd_rinth)(uint32_t x, void *fp_status)
  {
      int old_flags = get_float_exception_flags(fp_status), new_flags;
      float16 ret;
@@ -XXX,XX +XXX,XX @@ float16 HELPER(advsimd_rinth)(float16 x, void *fp_status)
   * setting the mode appropriately before calling the helper.
   */
 -uint32_t HELPER(advsimd_f16tosinth)(float16 a, void *fpstp)
 +uint32_t HELPER(advsimd_f16tosinth)(uint32_t a, void *fpstp)
  {
      float_status *fpst = fpstp;
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_f16tosinth)(float16 a, void *fpstp)
      return float16_to_int16(a, fpst);
  }
 -uint32_t HELPER(advsimd_f16touinth)(float16 a, void *fpstp)
 +uint32_t HELPER(advsimd_f16touinth)(uint32_t a, void *fpstp)
  {
      float_status *fpst = fpstp;
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_f16touinth)(float16 a, void *fpstp)
   * Square Root and Reciprocal square root
   */
 -float16 HELPER(sqrt_f16)(float16 a, void *fpstp)
 +uint32_t HELPER(sqrt_f16)(uint32_t a, void *fpstp)
  {
      float_status *s = fpstp;
 diff --git a/target/arm/helper.c b/target/arm/helper.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/helper.c
 +++ b/target/arm/helper.c
@@ -XXX,XX +XXX,XX @@ DO_VFP_cmp(d, float64)
  /* Integer to float and float to integer conversions */
 -#define CONV_ITOF(name, fsz, sign) \
 -    float##fsz HELPER(name)(uint32_t x, void *fpstp) \
 -{ \
 -    float_status *fpst = fpstp; \
 -    return sign##int32_to_##float##fsz((sign##int32_t)x, fpst); \
 +#define CONV_ITOF(name, ftype, fsz, sign)                           \
 +ftype HELPER(name)(uint32_t x, void *fpstp)                         \
 +{                                                                   \
 +    float_status *fpst = fpstp;                                     \
 +    return sign##int32_to_##float##fsz((sign##int32_t)x, fpst);     \
  }
 -#define CONV_FTOI(name, fsz, sign, round) \
 -uint32_t HELPER(name)(float##fsz x, void *fpstp) \
 -{ \
 -    float_status *fpst = fpstp; \
 -    if (float##fsz##_is_any_nan(x)) { \
 -        float_raise(float_flag_invalid, fpst); \
 -        return 0; \
 -    } \
 -    return float##fsz##_to_##sign##int32##round(x, fpst); \
 +#define CONV_FTOI(name, ftype, fsz, sign, round)                \
 +uint32_t HELPER(name)(ftype x, void *fpstp)                     \
 +{                                                               \
 +    float_status *fpst = fpstp;                                 \
 +    if (float##fsz##_is_any_nan(x)) {                           \
 +        float_raise(float_flag_invalid, fpst);                  \
 +        return 0;                                               \
 +    }                                                           \
 +    return float##fsz##_to_##sign##int32##round(x, fpst);       \
  }
 -#define FLOAT_CONVS(name, p, fsz, sign) \
 -CONV_ITOF(vfp_##name##to##p, fsz, sign) \
 -CONV_FTOI(vfp_to##name##p, fsz, sign, ) \
 -CONV_FTOI(vfp_to##name##z##p, fsz, sign, _round_to_zero)
 +#define FLOAT_CONVS(name, p, ftype, fsz, sign)            \
 +    CONV_ITOF(vfp_##name##to##p, ftype, fsz, sign)        \
 +    CONV_FTOI(vfp_to##name##p, ftype, fsz, sign, )        \
 +    CONV_FTOI(vfp_to##name##z##p, ftype, fsz, sign, _round_to_zero)
 -FLOAT_CONVS(si, h, 16, )
 -FLOAT_CONVS(si, s, 32, )
 -FLOAT_CONVS(si, d, 64, )
 -FLOAT_CONVS(ui, h, 16, u)
 -FLOAT_CONVS(ui, s, 32, u)
 -FLOAT_CONVS(ui, d, 64, u)
 +FLOAT_CONVS(si, h, uint32_t, 16, )
 +FLOAT_CONVS(si, s, float32, 32, )
 +FLOAT_CONVS(si, d, float64, 64, )
 +FLOAT_CONVS(ui, h, uint32_t, 16, u)
 +FLOAT_CONVS(ui, s, float32, 32, u)
 +FLOAT_CONVS(ui, d, float64, 64, u)
  #undef CONV_ITOF
  #undef CONV_FTOI
@@ -XXX,XX +XXX,XX @@ static float16 do_postscale_fp16(float64 f, int shift, float_status *fpst)
      return float64_to_float16(float64_scalbn(f, -shift, fpst), true, fpst);
  }
 -float16 HELPER(vfp_sltoh)(uint32_t x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_sltoh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return do_postscale_fp16(int32_to_float64(x, fpst), shift, fpst);
  }
 -float16 HELPER(vfp_ultoh)(uint32_t x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_ultoh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return do_postscale_fp16(uint32_to_float64(x, fpst), shift, fpst);
  }
 -float16 HELPER(vfp_sqtoh)(uint64_t x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_sqtoh)(uint64_t x, uint32_t shift, void *fpst)
  {
      return do_postscale_fp16(int64_to_float64(x, fpst), shift, fpst);
  }
 -float16 HELPER(vfp_uqtoh)(uint64_t x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_uqtoh)(uint64_t x, uint32_t shift, void *fpst)
  {
      return do_postscale_fp16(uint64_to_float64(x, fpst), shift, fpst);
  }
@@ -XXX,XX +XXX,XX @@ static float64 do_prescale_fp16(float16 f, int shift, float_status *fpst)
      }
  }
 -uint32_t HELPER(vfp_toshh)(float16 x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_toshh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_int16(do_prescale_fp16(x, shift, fpst), fpst);
  }
 -uint32_t HELPER(vfp_touhh)(float16 x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_touhh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_uint16(do_prescale_fp16(x, shift, fpst), fpst);
  }
 -uint32_t HELPER(vfp_toslh)(float16 x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_toslh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_int32(do_prescale_fp16(x, shift, fpst), fpst);
  }
 -uint32_t HELPER(vfp_toulh)(float16 x, uint32_t shift, void *fpst)
 +uint32_t HELPER(vfp_toulh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_uint32(do_prescale_fp16(x, shift, fpst), fpst);
  }
 -uint64_t HELPER(vfp_tosqh)(float16 x, uint32_t shift, void *fpst)
 +uint64_t HELPER(vfp_tosqh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_int64(do_prescale_fp16(x, shift, fpst), fpst);
  }
 -uint64_t HELPER(vfp_touqh)(float16 x, uint32_t shift, void *fpst)
 +uint64_t HELPER(vfp_touqh)(uint32_t x, uint32_t shift, void *fpst)
  {
      return float64_to_uint64(do_prescale_fp16(x, shift, fpst), fpst);
  }
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(set_neon_rmode)(uint32_t rmode, CPUARMState *env)
  }
  /* Half precision conversions.  */
 -float32 HELPER(vfp_fcvt_f16_to_f32)(float16 a, void *fpstp, uint32_t ahp_mode)
 +float32 HELPER(vfp_fcvt_f16_to_f32)(uint32_t a, void *fpstp, uint32_t ahp_mode)
  {
      /* Squash FZ16 to 0 for the duration of conversion.  In this case,
       * it would affect flushing input denormals.
@@ -XXX,XX +XXX,XX @@ float32 HELPER(vfp_fcvt_f16_to_f32)(float16 a, void *fpstp, uint32_t ahp_mode)
      return r;
  }
 -float16 HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
 +uint32_t HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
  {
      /* Squash FZ16 to 0 for the duration of conversion.  In this case,
       * it would affect flushing output denormals.
@@ -XXX,XX +XXX,XX @@ float16 HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
      return r;
  }
 -float64 HELPER(vfp_fcvt_f16_to_f64)(float16 a, void *fpstp, uint32_t ahp_mode)
 +float64 HELPER(vfp_fcvt_f16_to_f64)(uint32_t a, void *fpstp, uint32_t ahp_mode)
  {
      /* Squash FZ16 to 0 for the duration of conversion.  In this case,
       * it would affect flushing input denormals.
@@ -XXX,XX +XXX,XX @@ float64 HELPER(vfp_fcvt_f16_to_f64)(float16 a, void *fpstp, uint32_t ahp_mode)
      return r;
  }
 -float16 HELPER(vfp_fcvt_f64_to_f16)(float64 a, void *fpstp, uint32_t ahp_mode)
 +uint32_t HELPER(vfp_fcvt_f64_to_f16)(float64 a, void *fpstp, uint32_t ahp_mode)
  {
      /* Squash FZ16 to 0 for the duration of conversion.  In this case,
       * it would affect flushing output denormals.
@@ -XXX,XX +XXX,XX @@ static bool round_to_inf(float_status *fpst, bool sign_bit)
      g_assert_not_reached();
  }
 -float16 HELPER(recpe_f16)(float16 input, void *fpstp)
 +uint32_t HELPER(recpe_f16)(uint32_t input, void *fpstp)
  {
      float_status *fpst = fpstp;
      float16 f16 = float16_squash_input_denormal(input, fpst);
@@ -XXX,XX +XXX,XX @@ static uint64_t recip_sqrt_estimate(int *exp , int exp_off, uint64_t frac)
      return extract64(estimate, 0, 8) << 44;
  }
 -float16 HELPER(rsqrte_f16)(float16 input, void *fpstp)
 +uint32_t HELPER(rsqrte_f16)(uint32_t input, void *fpstp)
  {
      float_status *s = fpstp;
      float16 f16 = float16_squash_input_denormal(input, s);
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 02/25] MAINTAINERS: Add entries for newer MPS2 boards and devices
+[PULL 29/42] hw/i2c: Add header for ARM SBCon two-wire serial bus interface
-Add entries to MAINTAINERS to cover the newer MPS2 boards and
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
 the new devices they use.
+'ARM SBCon two-wire serial bus interface' is the official
+name describing the pair of registers used to bitbanging
+I2C in the Versatile boards.
+Make the private VersatileI2CState structure as public
+ArmSbconI2CState.
+Add the TYPE_ARM_SBCON_I2C, alias to our current
+TYPE_VERSATILE_I2C model.
+Rename the memory region description as 'arm_sbcon_i2c'.
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-5-f4bug@amsat.org
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Message-id: 20180518153157.14899-1-peter.maydell@linaro.org
 ---
- MAINTAINERS | 9 +++++++--
+ include/hw/i2c/arm_sbcon_i2c.h | 35 ++++++++++++++++++++++++++++++++++
-file changed, 7 insertions(+), 2 deletions(-)
+ hw/i2c/versatile_i2c.c         | 17 +++++------------
  MAINTAINERS                    |  1 +
 files changed, 41 insertions(+), 12 deletions(-)
  create mode 100644 include/hw/i2c/arm_sbcon_i2c.h
+diff --git a/include/hw/i2c/arm_sbcon_i2c.h b/include/hw/i2c/arm_sbcon_i2c.h
+new file mode 100644
+index XXXXXXX..XXXXXXX
+--- /dev/null
++++ b/include/hw/i2c/arm_sbcon_i2c.h
+@@ -XXX,XX +XXX,XX @@
++/*
++ * ARM SBCon two-wire serial bus interface (I2C bitbang)
++ *   a.k.a.
++ * ARM Versatile I2C controller
++ *
++ * Copyright (c) 2006-2007 CodeSourcery.
++ * Copyright (c) 2012 Oskar Andero <oskar.andero@gmail.com>
++ * Copyright (C) 2020 Philippe Mathieu-Daudé <f4bug@amsat.org>
++ *
++ * SPDX-License-Identifier: GPL-2.0-or-later
++ */
++#ifndef HW_I2C_ARM_SBCON_H
++#define HW_I2C_ARM_SBCON_H
++
++#include "hw/sysbus.h"
++#include "hw/i2c/bitbang_i2c.h"
++
++#define TYPE_VERSATILE_I2C "versatile_i2c"
++#define TYPE_ARM_SBCON_I2C TYPE_VERSATILE_I2C
++
++#define ARM_SBCON_I2C(obj) \
++    OBJECT_CHECK(ArmSbconI2CState, (obj), TYPE_ARM_SBCON_I2C)
++
++typedef struct ArmSbconI2CState {
++    /*< private >*/
++    SysBusDevice parent_obj;
++    /*< public >*/
++
++    MemoryRegion iomem;
++    bitbang_i2c_interface bitbang;
++    int out;
++    int in;
++} ArmSbconI2CState;
++
++#endif /* HW_I2C_ARM_SBCON_H */
+diff --git a/hw/i2c/versatile_i2c.c b/hw/i2c/versatile_i2c.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/i2c/versatile_i2c.c
++++ b/hw/i2c/versatile_i2c.c
+@@ -XXX,XX +XXX,XX @@
+ /*
+- * ARM Versatile I2C controller
++ * ARM SBCon two-wire serial bus interface (I2C bitbang)
++ * a.k.a. ARM Versatile I2C controller
+  *
+  * Copyright (c) 2006-2007 CodeSourcery.
+  * Copyright (c) 2012 Oskar Andero <oskar.andero@gmail.com>
+@@ -XXX,XX +XXX,XX @@
+  */
+ #include "qemu/osdep.h"
+-#include "hw/sysbus.h"
+-#include "hw/i2c/bitbang_i2c.h"
++#include "hw/i2c/arm_sbcon_i2c.h"
+ #include "hw/registerfields.h"
+ #include "qemu/log.h"
+ #include "qemu/module.h"
+-#define TYPE_VERSATILE_I2C "versatile_i2c"
+ #define VERSATILE_I2C(obj) \
+     OBJECT_CHECK(VersatileI2CState, (obj), TYPE_VERSATILE_I2C)
+-typedef struct VersatileI2CState {
+-    SysBusDevice parent_obj;
++typedef ArmSbconI2CState VersatileI2CState;
+-    MemoryRegion iomem;
+-    bitbang_i2c_interface bitbang;
+-    int out;
+-    int in;
+-} VersatileI2CState;
+ REG32(CONTROL_GET, 0)
+ REG32(CONTROL_SET, 0)
+@@ -XXX,XX +XXX,XX @@ static void versatile_i2c_init(Object *obj)
+     bus = i2c_init_bus(dev, "i2c");
+     bitbang_i2c_init(&s->bitbang, bus);
+     memory_region_init_io(&s->iomem, obj, &versatile_i2c_ops, s,
+-                          "versatile_i2c", 0x1000);
++                          "arm_sbcon_i2c", 0x1000);
+     sysbus_init_mmio(sbd, &s->iomem);
+ }
 diff --git a/MAINTAINERS b/MAINTAINERS
 index XXXXXXX..XXXXXXX 100644
 --- a/MAINTAINERS
 +++ b/MAINTAINERS
-@@ -XXX,XX +XXX,XX @@ F: hw/timer/cmsdk-apb-timer.c
- F: include/hw/timer/cmsdk-apb-timer.h
- F: hw/char/cmsdk-apb-uart.c
- F: include/hw/char/cmsdk-apb-uart.h
-+F: hw/misc/tz-ppc.c
-+F: include/hw/misc/tz-ppc.h
- ARM cores
- M: Peter Maydell <peter.maydell@linaro.org>
 @@ -XXX,XX +XXX,XX @@ M: Peter Maydell <peter.maydell@linaro.org>
  L: qemu-arm@nongnu.org
  S: Maintained
- F: hw/arm/mps2.c
+ F: hw/*/versatile*
--F: hw/misc/mps2-scc.c
++F: include/hw/i2c/arm_sbcon_i2c.h
--F: include/hw/misc/mps2-scc.h
+ F: hw/misc/arm_sysctl.c
-+F: hw/arm/mps2-tz.c
+ F: docs/system/arm/versatile.rst
-+F: hw/misc/mps2-*.c
 +F: include/hw/misc/mps2-*.h
 +F: hw/arm/iotkit.c
 +F: include/hw/arm/iotkit.h
  Musicpal
  M: Jan Kiszka <jan.kiszka@web.de>
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 03/25] hw/intc/arm_gicv3: Fix APxR<n> register dispatching
+[PULL 30/42] hw/arm: Use TYPE_VERSATILE_I2C instead of hardcoded string
-From: Jan Kiszka <jan.kiszka@siemens.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-There was a nasty flip in identifying which register group an access is
+By using the TYPE_* definitions for devices, we can:
-targeting. The issue caused spuriously raised priorities of the guest
+ - quickly find where devices are used with 'git-grep'
-when handing CPUs over in the Jailhouse hypervisor.
+ - easily rename a device (one-line change).
-Cc: qemu-stable@nongnu.org
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
-Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com>
+Message-id: 20200617072539.32686-6-f4bug@amsat.org
 Message-id: 28b927d3-da58-bce4-cc13-bfec7f9b1cb9@siemens.com
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/intc/arm_gicv3_cpuif.c | 12 ++++++------
+ hw/arm/realview.c    | 3 ++-
-file changed, 6 insertions(+), 6 deletions(-)
+ hw/arm/versatilepb.c | 3 ++-
  hw/arm/vexpress.c    | 3 ++-
 files changed, 6 insertions(+), 3 deletions(-)
-diff --git a/hw/intc/arm_gicv3_cpuif.c b/hw/intc/arm_gicv3_cpuif.c
+diff --git a/hw/arm/realview.c b/hw/arm/realview.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/intc/arm_gicv3_cpuif.c
+--- a/hw/arm/realview.c
-+++ b/hw/intc/arm_gicv3_cpuif.c
++++ b/hw/arm/realview.c
-@@ -XXX,XX +XXX,XX @@ static uint64_t icv_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
+@@ -XXX,XX +XXX,XX @@
- {
+ #include "hw/cpu/a9mpcore.h"
-     GICv3CPUState *cs = icc_cs_from_env(env);
+ #include "hw/intc/realview_gic.h"
-     int regno = ri->opc2 & 3;
+ #include "hw/irq.h"
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
++#include "hw/i2c/arm_sbcon_i2c.h"
-+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
-     uint64_t value = cs->ich_apr[grp][regno];
+ #define SMP_BOOT_ADDR 0xe0000000
+ #define SMP_BOOTREG_ADDR 0x10000030
-     trace_gicv3_icv_ap_read(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
+@@ -XXX,XX +XXX,XX @@ static void realview_init(MachineState *machine,
-@@ -XXX,XX +XXX,XX @@ static void icv_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
+         }
- {
+     }
-     GICv3CPUState *cs = icc_cs_from_env(env);
-     int regno = ri->opc2 & 3;
+-    dev = sysbus_create_simple("versatile_i2c", 0x10002000, NULL);
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
++    dev = sysbus_create_simple(TYPE_VERSATILE_I2C, 0x10002000, NULL);
-+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
+     i2c = (I2CBus *)qdev_get_child_bus(dev, "i2c");
+     i2c_create_slave(i2c, "ds1338", 0x68);
-     trace_gicv3_icv_ap_write(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
+diff --git a/hw/arm/versatilepb.c b/hw/arm/versatilepb.c
-@@ -XXX,XX +XXX,XX @@ static uint64_t icc_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
+index XXXXXXX..XXXXXXX 100644
-     uint64_t value;
+--- a/hw/arm/versatilepb.c
++++ b/hw/arm/versatilepb.c
-     int regno = ri->opc2 & 3;
+@@ -XXX,XX +XXX,XX @@
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1;
+ #include "sysemu/sysemu.h"
-+    int grp = (ri->crm & 1) ? GICV3_G1 : GICV3_G0;
+ #include "hw/pci/pci.h"
+ #include "hw/i2c/i2c.h"
-     if (icv_access(env, grp == GICV3_G0 ? HCR_FMO : HCR_IMO)) {
++#include "hw/i2c/arm_sbcon_i2c.h"
-         return icv_ap_read(env, ri);
+ #include "hw/irq.h"
-@@ -XXX,XX +XXX,XX @@ static void icc_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
+ #include "hw/boards.h"
-     GICv3CPUState *cs = icc_cs_from_env(env);
+ #include "exec/address-spaces.h"
+@@ -XXX,XX +XXX,XX @@ static void versatile_init(MachineState *machine, int board_id)
-     int regno = ri->opc2 & 3;
+     /* Add PL031 Real Time Clock. */
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1;
+     sysbus_create_simple("pl031", 0x101e8000, pic[10]);
-+    int grp = (ri->crm & 1) ? GICV3_G1 : GICV3_G0;
+-    dev = sysbus_create_simple("versatile_i2c", 0x10002000, NULL);
-     if (icv_access(env, grp == GICV3_G0 ? HCR_FMO : HCR_IMO)) {
++    dev = sysbus_create_simple(TYPE_VERSATILE_I2C, 0x10002000, NULL);
-         icv_ap_write(env, ri, value);
+     i2c = (I2CBus *)qdev_get_child_bus(dev, "i2c");
-@@ -XXX,XX +XXX,XX @@ static uint64_t ich_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
+     i2c_create_slave(i2c, "ds1338", 0x68);
- {
-     GICv3CPUState *cs = icc_cs_from_env(env);
+diff --git a/hw/arm/vexpress.c b/hw/arm/vexpress.c
-     int regno = ri->opc2 & 3;
+index XXXXXXX..XXXXXXX 100644
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+--- a/hw/arm/vexpress.c
-+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
++++ b/hw/arm/vexpress.c
-     uint64_t value;
+@@ -XXX,XX +XXX,XX @@
+ #include "hw/char/pl011.h"
-     value = cs->ich_apr[grp][regno];
+ #include "hw/cpu/a9mpcore.h"
-@@ -XXX,XX +XXX,XX @@ static void ich_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
+ #include "hw/cpu/a15mpcore.h"
- {
++#include "hw/i2c/arm_sbcon_i2c.h"
-     GICv3CPUState *cs = icc_cs_from_env(env);
-     int regno = ri->opc2 & 3;
+ #define VEXPRESS_BOARD_ID 0x8e0
--    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+ #define VEXPRESS_FLASH_SIZE (64 * 1024 * 1024)
-+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
+@@ -XXX,XX +XXX,XX @@ static void vexpress_common_init(MachineState *machine)
+     sysbus_create_simple("sp804", map[VE_TIMER01], pic[2]);
-     trace_gicv3_ich_ap_write(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
+     sysbus_create_simple("sp804", map[VE_TIMER23], pic[3]);
 -    dev = sysbus_create_simple("versatile_i2c", map[VE_SERIALDVI], NULL);
 +    dev = sysbus_create_simple(TYPE_VERSATILE_I2C, map[VE_SERIALDVI], NULL);
      i2c = (I2CBus *)qdev_get_child_bus(dev, "i2c");
      i2c_create_slave(i2c, "sii9022", 0x39);
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 24/25] ARM: ACPI: Fix use-after-free due to memory realloc
+[PULL 31/42] hw/arm/mps2: Document CMSDK/FPGA APB subsystem sections
-From: Shannon Zhao <zhaoshenglong@huawei.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-acpi_data_push uses g_array_set_size to resize the memory size. If there
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
-is no enough contiguous memory, the address will be changed. So previous
+Message-id: 20200617072539.32686-7-f4bug@amsat.org
-pointer could not be used any more. It must update the pointer and use
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 the new one.
 Also, previous codes wrongly use le32 conversion of iort->node_offset
 for subsequent computations that will result incorrect value if host is
 not litlle endian. So use the non-converted one instead.
 Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 1527663951-14552-1-git-send-email-zhaoshenglong@huawei.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/virt-acpi-build.c | 20 +++++++++++++++-----
+ hw/arm/mps2.c | 5 ++++-
-file changed, 15 insertions(+), 5 deletions(-)
+file changed, 4 insertions(+), 1 deletion(-)
-diff --git a/hw/arm/virt-acpi-build.c b/hw/arm/virt-acpi-build.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/virt-acpi-build.c
+--- a/hw/arm/mps2.c
-+++ b/hw/arm/virt-acpi-build.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
+@@ -XXX,XX +XXX,XX @@ typedef struct {
-     AcpiIortItsGroup *its;
+     MemoryRegion blockram_m2;
-     AcpiIortTable *iort;
+     MemoryRegion blockram_m3;
-     AcpiIortSmmu3 *smmu;
+     MemoryRegion sram;
--    size_t node_size, iort_length, smmu_offset = 0;
++    /* FPGA APB subsystem */
-+    size_t node_size, iort_node_offset, iort_length, smmu_offset = 0;
+     MPS2SCC scc;
-     AcpiIortRC *rc;
++    /* CMSDK APB subsystem */
+     CMSDKAPBDualTimer dualtimer;
-     iort = acpi_data_push(table_data, sizeof(*iort));
+ } MPS2MachineState;
-@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
-     iort_length = sizeof(*iort);
+         g_assert_not_reached();
      iort->node_count = cpu_to_le32(nb_nodes);
 -    iort->node_offset = cpu_to_le32(sizeof(*iort));
 +    /*
 +     * Use a copy in case table_data->data moves during acpi_data_push
 +     * operations.
 +     */
 +    iort_node_offset = sizeof(*iort);
 +    iort->node_offset = cpu_to_le32(iort_node_offset);
      /* ITS group node */
      node_size =  sizeof(*its) + sizeof(uint32_t);
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
          int irq =  vms->irqmap[VIRT_SMMU];
          /* SMMUv3 node */
 -        smmu_offset = iort->node_offset + node_size;
 +        smmu_offset = iort_node_offset + node_size;
          node_size = sizeof(*smmu) + sizeof(*idmap);
          iort_length += node_size;
          smmu = acpi_data_push(table_data, node_size);
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
          idmap->id_count = cpu_to_le32(0xFFFF);
          idmap->output_base = 0;
          /* output IORT node is the ITS group node (the first node) */
 -        idmap->output_reference = cpu_to_le32(iort->node_offset);
 +        idmap->output_reference = cpu_to_le32(iort_node_offset);
      }
-     /* Root Complex Node */
++    /* CMSDK APB subsystem */
-@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
+     cmsdk_apb_timer_create(0x40000000, qdev_get_gpio_in(armv7m, 8), SYSCLK_FRQ);
-         idmap->output_reference = cpu_to_le32(smmu_offset);
+     cmsdk_apb_timer_create(0x40001000, qdev_get_gpio_in(armv7m, 9), SYSCLK_FRQ);
-     } else {
+-
-         /* output IORT node is the ITS group node (the first node) */
+     object_initialize_child(OBJECT(mms), "dualtimer", &mms->dualtimer,
--        idmap->output_reference = cpu_to_le32(iort->node_offset);
+                             TYPE_CMSDK_APB_DUALTIMER);
-+        idmap->output_reference = cpu_to_le32(iort_node_offset);
+     qdev_prop_set_uint32(DEVICE(&mms->dualtimer), "pclk-frq", SYSCLK_FRQ);
-     }
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+                        qdev_get_gpio_in(armv7m, 10));
-+    /*
+     sysbus_mmio_map(SYS_BUS_DEVICE(&mms->dualtimer), 0, 0x40002000);
-+     * Update the pointer address in case table_data->data moves during above
-+     * acpi_data_push operations.
++    /* FPGA APB subsystem */
-+     */
+     object_initialize_child(OBJECT(mms), "scc", &mms->scc, TYPE_MPS2_SCC);
-+    iort = (AcpiIortTable *)(table_data->data + iort_start);
+     sccdev = DEVICE(&mms->scc);
-     iort->length = cpu_to_le32(iort_length);
+     qdev_prop_set_uint32(sccdev, "scc-cfg4", 0x2);
      build_header(linker, table_data, (void *)(table_data->data + iort_start),
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 22/25] Make address_space_translate_iommu take a MemTxAttrs argument
+[PULL 32/42] hw/arm/mps2: Rename CMSDK AHB peripheral region
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
 add MemTxAttrs as an argument to address_space_translate_iommu().
+To differenciate with the CMSDK APB peripheral region,
+rename this region 'CMSDK AHB peripheral region'.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-8-f4bug@amsat.org
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
-Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Message-id: 20180521140402.23318-14-peter.maydell@linaro.org
 ---
- exec.c | 8 +++++---
+ hw/arm/mps2.c | 3 ++-
-file changed, 5 insertions(+), 3 deletions(-)
+file changed, 2 insertions(+), 1 deletion(-)
-diff --git a/exec.c b/exec.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/exec.c
+--- a/hw/arm/mps2.c
-+++ b/exec.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ address_space_translate_internal(AddressSpaceDispatch *d, hwaddr addr, hwaddr *x
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
-  * @is_write: whether the translation operation is for write
+      */
-  * @is_mmio: whether this can be MMIO, set true if it can
+     create_unimplemented_device("CMSDK APB peripheral region @0x40000000",
-  * @target_as: the address space targeted by the IOMMU
+x40000000, 0x00010000);
-+ * @attrs: transaction attributes
+-    create_unimplemented_device("CMSDK peripheral region @0x40010000",
-  *
++    create_unimplemented_device("CMSDK AHB peripheral region @0x40010000",
-  * This function is called from RCU critical section.  It is the common
+x40010000, 0x00010000);
-  * part of flatview_do_translate and address_space_translate_cached.
+     create_unimplemented_device("Extra peripheral region @0x40020000",
-@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection address_space_translate_iommu(IOMMUMemoryRegion *iomm
+x40020000, 0x00010000);
-                                                          hwaddr *page_mask_out,
++
-                                                          bool is_write,
+     create_unimplemented_device("RESERVED 4", 0x40030000, 0x001D0000);
-                                                          bool is_mmio,
+     create_unimplemented_device("VGA", 0x41000000, 0x0200000);
 -                                                         AddressSpace **target_as)
 +                                                         AddressSpace **target_as,
 +                                                         MemTxAttrs attrs)
  {
      MemoryRegionSection *section;
      hwaddr page_mask = (hwaddr)-1;
@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
          return address_space_translate_iommu(iommu_mr, xlat,
                                               plen_out, page_mask_out,
                                               is_write, is_mmio,
 -                                             target_as);
 +                                             target_as, attrs);
      }
      if (page_mask_out) {
          /* Not behind an IOMMU, use default page size. */
@@ -XXX,XX +XXX,XX @@ static inline MemoryRegion *address_space_translate_cached(
      section = address_space_translate_iommu(iommu_mr, xlat, plen,
                                              NULL, is_write, true,
 -                                            &target_as);
 +                                            &target_as, attrs);
      return section.mr;
  }
 --
-.17.1
+.20.1

-New patch
+[PULL 33/42] hw/arm/mps2: Add CMSDK APB watchdog device
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
+We already model the CMSDK APB watchdog device, let's use it!
+Suggested-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-9-f4bug@amsat.org
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/arm/mps2.c  | 7 +++++++
+ hw/arm/Kconfig | 1 +
+files changed, 8 insertions(+)
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/arm/mps2.c
++++ b/hw/arm/mps2.c
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+     sysbus_connect_irq(SYS_BUS_DEVICE(&mms->dualtimer), 0,
+                        qdev_get_gpio_in(armv7m, 10));
+     sysbus_mmio_map(SYS_BUS_DEVICE(&mms->dualtimer), 0, 0x40002000);
++    object_initialize_child(OBJECT(mms), "watchdog", &mms->watchdog,
++                            TYPE_CMSDK_APB_WATCHDOG);
++    qdev_prop_set_uint32(DEVICE(&mms->watchdog), "wdogclk-frq", SYSCLK_FRQ);
++    sysbus_realize(SYS_BUS_DEVICE(&mms->watchdog), &error_fatal);
++    sysbus_connect_irq(SYS_BUS_DEVICE(&mms->watchdog), 0,
++                       qdev_get_gpio_in_named(armv7m, "NMI", 0));
++    sysbus_mmio_map(SYS_BUS_DEVICE(&mms->watchdog), 0, 0x40008000);
+     /* FPGA APB subsystem */
+     object_initialize_child(OBJECT(mms), "scc", &mms->scc, TYPE_MPS2_SCC);
+diff --git a/hw/arm/Kconfig b/hw/arm/Kconfig
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/arm/Kconfig
++++ b/hw/arm/Kconfig
+@@ -XXX,XX +XXX,XX @@ config MPS2
+     select PL080    # DMA controller
+     select SPLIT_IRQ
+     select UNIMP
++    select CMSDK_APB_WATCHDOG
+ config FSL_IMX7
+     bool
+--
+.20.1

-New patch
+[PULL 34/42] hw/arm/mps2: Add CMSDK AHB GPIO peripherals as unimplemented devices
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Register the GPIO peripherals as unimplemented to better
+follow their accesses, for example booting Zephyr:
+  ----------------
+  IN: arm_mps2_pinmux_init
+x00001160:  f64f 0231  movw     r2, #0xf831
+x00001164:  4b06       ldr      r3, [pc, #0x18]
+x00001166:  2000       movs     r0, #0
+x00001168:  619a       str      r2, [r3, #0x18]
+x0000116a:  f24c 426f  movw     r2, #0xc46f
+x0000116e:  f503 5380  add.w    r3, r3, #0x1000
+x00001172:  619a       str      r2, [r3, #0x18]
+x00001174:  f44f 529e  mov.w    r2, #0x13c0
+x00001178:  f503 5380  add.w    r3, r3, #0x1000
+x0000117c:  619a       str      r2, [r3, #0x18]
+x0000117e:  4770       bx       lr
+  cmsdk-ahb-gpio: unimplemented device write (size 4, value 0xf831, offset 0x18)
+  cmsdk-ahb-gpio: unimplemented device write (size 4, value 0xc46f, offset 0x18)
+  cmsdk-ahb-gpio: unimplemented device write (size 4, value 0x13c0, offset 0x18)
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Message-id: 20200617072539.32686-10-f4bug@amsat.org
+Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
+---
+ hw/arm/mps2.c | 8 ++++++--
+file changed, 6 insertions(+), 2 deletions(-)
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
+index XXXXXXX..XXXXXXX 100644
+--- a/hw/arm/mps2.c
++++ b/hw/arm/mps2.c
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+     MemoryRegion *system_memory = get_system_memory();
+     MachineClass *mc = MACHINE_GET_CLASS(machine);
+     DeviceState *armv7m, *sccdev;
++    int i;
+     if (strcmp(machine->cpu_type, mc->default_cpu_type) != 0) {
+         error_report("This board can only be used with CPU %s",
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+          */
+         Object *orgate;
+         DeviceState *orgate_dev;
+-        int i;
+         orgate = object_new(TYPE_OR_IRQ);
+         object_property_set_int(orgate, 6, "num-lines", &error_fatal);
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+          */
+         Object *orgate;
+         DeviceState *orgate_dev;
+-        int i;
+         orgate = object_new(TYPE_OR_IRQ);
+         object_property_set_int(orgate, 10, "num-lines", &error_fatal);
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
+     default:
+         g_assert_not_reached();
+     }
++    for (i = 0; i < 4; i++) {
++        static const hwaddr gpiobase[] = {0x40010000, 0x40011000,
++                                          0x40012000, 0x40013000};
++        create_unimplemented_device("cmsdk-ahb-gpio", gpiobase[i], 0x1000);
++    }
+     /* CMSDK APB subsystem */
+     cmsdk_apb_timer_create(0x40000000, qdev_get_gpio_in(armv7m, 8), SYSCLK_FRQ);
+--
+.20.1

-[Qemu-devel] [PULL 25/25] KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice
+[PULL 35/42] hw/arm/mps2: Map the FPGA I/O block
-From: Shannon Zhao <zhaoshenglong@huawei.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-kvm_irqchip_create called by kvm_init will call kvm_init_irq_routing to
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
-initialize global capability variables. If we call kvm_init_irq_routing in
+Message-id: 20200617072539.32686-11-f4bug@amsat.org
-GIC realize function, previous allocated memory will leak.
+Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Fix this by deleting the unnecessary call.
 Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 1527750994-14360-1-git-send-email-zhaoshenglong@huawei.com
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/intc/arm_gic_kvm.c   | 1 -
+ hw/arm/mps2.c | 9 +++++++++
- hw/intc/arm_gicv3_kvm.c | 1 -
+file changed, 9 insertions(+)
 files changed, 2 deletions(-)
-diff --git a/hw/intc/arm_gic_kvm.c b/hw/intc/arm_gic_kvm.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/intc/arm_gic_kvm.c
+--- a/hw/arm/mps2.c
-+++ b/hw/intc/arm_gic_kvm.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ static void kvm_arm_gic_realize(DeviceState *dev, Error **errp)
+@@ -XXX,XX +XXX,XX @@
+ #include "hw/timer/cmsdk-apb-timer.h"
-     if (kvm_has_gsi_routing()) {
+ #include "hw/timer/cmsdk-apb-dualtimer.h"
-         /* set up irq routing */
+ #include "hw/misc/mps2-scc.h"
--        kvm_init_irq_routing(kvm_state);
++#include "hw/misc/mps2-fpgaio.h"
-         for (i = 0; i < s->num_irq - GIC_INTERNAL; ++i) {
+ #include "hw/net/lan9118.h"
-             kvm_irqchip_add_irq_route(kvm_state, i, 0, i);
+ #include "net/net.h"
-         }
++#include "hw/watchdog/cmsdk-apb-watchdog.h"
-diff --git a/hw/intc/arm_gicv3_kvm.c b/hw/intc/arm_gicv3_kvm.c
-index XXXXXXX..XXXXXXX 100644
+ typedef enum MPS2FPGAType {
---- a/hw/intc/arm_gicv3_kvm.c
+     FPGA_AN385,
-+++ b/hw/intc/arm_gicv3_kvm.c
+@@ -XXX,XX +XXX,XX @@ typedef struct {
-@@ -XXX,XX +XXX,XX @@ static void kvm_arm_gicv3_realize(DeviceState *dev, Error **errp)
+     MemoryRegion sram;
+     /* FPGA APB subsystem */
-     if (kvm_has_gsi_routing()) {
+     MPS2SCC scc;
-         /* set up irq routing */
++    MPS2FPGAIO fpgaio;
--        kvm_init_irq_routing(kvm_state);
+     /* CMSDK APB subsystem */
-         for (i = 0; i < s->num_irq - GIC_INTERNAL; ++i) {
+     CMSDKAPBDualTimer dualtimer;
-             kvm_irqchip_add_irq_route(kvm_state, i, 0, i);
++    CMSDKAPBWatchdog watchdog;
-         }
+ } MPS2MachineState;
  #define TYPE_MPS2_MACHINE "mps2"
@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
      qdev_prop_set_uint32(sccdev, "scc-id", mmc->scc_id);
      sysbus_realize(SYS_BUS_DEVICE(&mms->scc), &error_fatal);
      sysbus_mmio_map(SYS_BUS_DEVICE(sccdev), 0, 0x4002f000);
 +    object_initialize_child(OBJECT(mms), "fpgaio",
 +                            &mms->fpgaio, TYPE_MPS2_FPGAIO);
 +    qdev_prop_set_uint32(DEVICE(&mms->fpgaio), "prescale-clk", 25000000);
 +    sysbus_realize(SYS_BUS_DEVICE(&mms->fpgaio), &error_fatal);
 +    sysbus_mmio_map(SYS_BUS_DEVICE(&mms->fpgaio), 0, 0x40028000);
      /* In hardware this is a LAN9220; the LAN9118 is software compatible
       * except that it doesn't support the checksum-offload feature.
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 08/25] xlnx-zdma: Correct mem leaks and memset to zero on desc unaligned errors
+[PULL 36/42] hw/arm/mps2: Add SPI devices
-From: Francisco Iglesias <frasse.iglesias@gmail.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-Coverity found that the string return by 'object_get_canonical_path' was not
+From 'Application Note AN385', chapter 3.9, SPI:
 being freed at two locations in the model (CID 1391294 and CID 1391293) and
 also that a memset was being called with a value greater than the max of a byte
 on the second argument (CID 1391286). This patch corrects this by adding the
 freeing of the strings and also changing to memset to zero instead on
 descriptor unaligned errors.
-Signed-off-by: Francisco Iglesias <frasse.iglesias@gmail.com>
+  The SMM implements five PL022 SPI modules.
-Reviewed-by: Edgar E. Iglesias <edgar.iglesias@xilinx.com>
-Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
+Two pairs of modules share the same OR-gated IRQ.
-Message-id: 20180528184859.3530-1-frasse.iglesias@gmail.com
 Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
 Message-id: 20200617072539.32686-12-f4bug@amsat.org
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/dma/xlnx-zdma.c | 10 +++++++---
+ hw/arm/mps2.c  | 24 ++++++++++++++++++++++++
-file changed, 7 insertions(+), 3 deletions(-)
+ hw/arm/Kconfig |  6 +++---
 files changed, 27 insertions(+), 3 deletions(-)
-diff --git a/hw/dma/xlnx-zdma.c b/hw/dma/xlnx-zdma.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/dma/xlnx-zdma.c
+--- a/hw/arm/mps2.c
-+++ b/hw/dma/xlnx-zdma.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ static bool zdma_load_descriptor(XlnxZDMA *s, uint64_t addr, void *buf)
+@@ -XXX,XX +XXX,XX @@
-         qemu_log_mask(LOG_GUEST_ERROR,
+ #include "hw/timer/cmsdk-apb-dualtimer.h"
-                       "zdma: unaligned descriptor at %" PRIx64,
+ #include "hw/misc/mps2-scc.h"
-                       addr);
+ #include "hw/misc/mps2-fpgaio.h"
--        memset(buf, 0xdeadbeef, sizeof(XlnxZDMADescr));
++#include "hw/ssi/pl022.h"
-+        memset(buf, 0x0, sizeof(XlnxZDMADescr));
+ #include "hw/net/lan9118.h"
-         s->error = true;
+ #include "net/net.h"
-         return false;
+ #include "hw/watchdog/cmsdk-apb-watchdog.h"
-     }
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
-@@ -XXX,XX +XXX,XX @@ static uint64_t zdma_read(void *opaque, hwaddr addr, unsigned size)
+     qdev_prop_set_uint32(DEVICE(&mms->fpgaio), "prescale-clk", 25000000);
-     RegisterInfo *r = &s->regs_info[addr / 4];
+     sysbus_realize(SYS_BUS_DEVICE(&mms->fpgaio), &error_fatal);
+     sysbus_mmio_map(SYS_BUS_DEVICE(&mms->fpgaio), 0, 0x40028000);
-     if (!r->data) {
++    sysbus_create_simple(TYPE_PL022, 0x40025000,        /* External ADC */
-+        gchar *path = object_get_canonical_path(OBJECT(s));
++                         qdev_get_gpio_in(armv7m, 22));
-         qemu_log("%s: Decode error: read from %" HWADDR_PRIx "\n",
++    for (i = 0; i < 2; i++) {
--                 object_get_canonical_path(OBJECT(s)),
++        static const int spi_irqno[] = {11, 24};
-+                 path,
++        static const hwaddr spibase[] = {0x40020000,    /* APB */
-                  addr);
++                                         0x40021000,    /* LCD */
-+        g_free(path);
++                                         0x40026000,    /* Shield0 */
-         ARRAY_FIELD_DP32(s->regs, ZDMA_CH_ISR, INV_APB, true);
++                                         0x40027000};   /* Shield1 */
-         zdma_ch_imr_update_irq(s);
++        DeviceState *orgate_dev;
-         return 0;
++        Object *orgate;
-@@ -XXX,XX +XXX,XX @@ static void zdma_write(void *opaque, hwaddr addr, uint64_t value,
++        int j;
-     RegisterInfo *r = &s->regs_info[addr / 4];
++
++        orgate = object_new(TYPE_OR_IRQ);
-     if (!r->data) {
++        object_property_set_int(orgate, 2, "num-lines", &error_fatal);
-+        gchar *path = object_get_canonical_path(OBJECT(s));
++        orgate_dev = DEVICE(orgate);
-         qemu_log("%s: Decode error: write to %" HWADDR_PRIx "=%" PRIx64 "\n",
++        qdev_realize(orgate_dev, NULL, &error_fatal);
--                 object_get_canonical_path(OBJECT(s)),
++        qdev_connect_gpio_out(orgate_dev, 0,
-+                 path,
++                              qdev_get_gpio_in(armv7m, spi_irqno[i]));
-                  addr, value);
++        for (j = 0; j < 2; j++) {
-+        g_free(path);
++            sysbus_create_simple(TYPE_PL022, spibase[2 * i + j],
-         ARRAY_FIELD_DP32(s->regs, ZDMA_CH_ISR, INV_APB, true);
++                                 qdev_get_gpio_in(orgate_dev, j));
-         zdma_ch_imr_update_irq(s);
++        }
-         return;
++    }
      /* In hardware this is a LAN9220; the LAN9118 is software compatible
       * except that it doesn't support the checksum-offload feature.
 diff --git a/hw/arm/Kconfig b/hw/arm/Kconfig
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/Kconfig
 +++ b/hw/arm/Kconfig
@@ -XXX,XX +XXX,XX @@ config HIGHBANK
      select ARM_TIMER # sp804
      select ARM_V7M
      select PL011 # UART
 -    select PL022 # Serial port
 +    select PL022 # SPI
      select PL031 # RTC
      select PL061 # GPIO
      select PL310 # cache controller
@@ -XXX,XX +XXX,XX @@ config STELLARIS
      select CMSDK_APB_WATCHDOG
      select I2C
      select PL011 # UART
 -    select PL022 # Serial port
 +    select PL022 # SPI
      select PL061 # GPIO
      select SSD0303 # OLED display
      select SSD0323 # OLED display
@@ -XXX,XX +XXX,XX @@ config MPS2
      select MPS2_FPGAIO
      select MPS2_SCC
      select OR_IRQ
 -    select PL022    # Serial port
 +    select PL022    # SPI
      select PL080    # DMA controller
      select SPLIT_IRQ
      select UNIMP
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 07/25] arm: fix malloc type mismatch
+[PULL 37/42] hw/arm/mps2: Add I2C devices
-From: Paolo Bonzini <pbonzini@redhat.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-cpregs_keys is an uint32_t* so the allocation should use uint32_t.
+From 'Application Note AN385', chapter 3.14:
 g_new is even better because it is type-safe.
-Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
+  The SMM implements a simple SBCon interface based on I2C.
 There are 4 SBCon interfaces on the FPGA APB subsystem.
 Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
 Message-id: 20200617072539.32686-13-f4bug@amsat.org
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- target/arm/gdbstub.c | 3 +--
+ hw/arm/mps2.c  | 8 ++++++++
-file changed, 1 insertion(+), 2 deletions(-)
+ hw/arm/Kconfig | 1 +
 files changed, 9 insertions(+)
-diff --git a/target/arm/gdbstub.c b/target/arm/gdbstub.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/gdbstub.c
+--- a/hw/arm/mps2.c
-+++ b/target/arm/gdbstub.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ int arm_gen_dynamic_xml(CPUState *cs)
+@@ -XXX,XX +XXX,XX @@
-     RegisterSysregXmlParam param = {cs, s};
+ #include "hw/misc/mps2-scc.h"
+ #include "hw/misc/mps2-fpgaio.h"
-     cpu->dyn_xml.num_cpregs = 0;
+ #include "hw/ssi/pl022.h"
--    cpu->dyn_xml.cpregs_keys = g_malloc(sizeof(uint32_t *) *
++#include "hw/i2c/arm_sbcon_i2c.h"
--                                        g_hash_table_size(cpu->cp_regs));
+ #include "hw/net/lan9118.h"
-+    cpu->dyn_xml.cpregs_keys = g_new(uint32_t, g_hash_table_size(cpu->cp_regs));
+ #include "net/net.h"
-     g_string_printf(s, "<?xml version=\"1.0\"?>");
+ #include "hw/watchdog/cmsdk-apb-watchdog.h"
-     g_string_append_printf(s, "<!DOCTYPE target SYSTEM \"gdb-target.dtd\">");
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
-     g_string_append_printf(s, "<feature name=\"org.qemu.gdb.arm.sys.regs\">");
+                                  qdev_get_gpio_in(orgate_dev, j));
          }
      }
 +    for (i = 0; i < 4; i++) {
 +        static const hwaddr i2cbase[] = {0x40022000,    /* Touch */
 +                                         0x40023000,    /* Audio */
 +                                         0x40029000,    /* Shield0 */
 +                                         0x4002a000};   /* Shield1 */
 +        sysbus_create_simple(TYPE_ARM_SBCON_I2C, i2cbase[i], NULL);
 +    }
      /* In hardware this is a LAN9220; the LAN9118 is software compatible
       * except that it doesn't support the checksum-offload feature.
 diff --git a/hw/arm/Kconfig b/hw/arm/Kconfig
 index XXXXXXX..XXXXXXX 100644
 --- a/hw/arm/Kconfig
 +++ b/hw/arm/Kconfig
@@ -XXX,XX +XXX,XX @@ config MPS2
      select SPLIT_IRQ
      select UNIMP
      select CMSDK_APB_WATCHDOG
 +    select VERSATILE_I2C
  config FSL_IMX7
      bool
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 04/25] arm_gicv3_kvm: increase clroffset accordingly
+[PULL 38/42] hw/arm/mps2: Add audio I2S interface as unimplemented device
-From: Shannon Zhao <zhaoshenglong@huawei.com>
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
-It forgot to increase clroffset during the loop. So it only clear the
+Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
-first 4 bytes.
+Message-id: 20200617072539.32686-14-f4bug@amsat.org
 Fixes: 367b9f527becdd20ddf116e17a3c0c2bbc486920
 Cc: qemu-stable@nongnu.org
 Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 1527047633-12368-1-git-send-email-zhaoshenglong@huawei.com
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/intc/arm_gicv3_kvm.c | 1 +
+ hw/arm/mps2.c | 1 +
 file changed, 1 insertion(+)
-diff --git a/hw/intc/arm_gicv3_kvm.c b/hw/intc/arm_gicv3_kvm.c
+diff --git a/hw/arm/mps2.c b/hw/arm/mps2.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/intc/arm_gicv3_kvm.c
+--- a/hw/arm/mps2.c
-+++ b/hw/intc/arm_gicv3_kvm.c
++++ b/hw/arm/mps2.c
-@@ -XXX,XX +XXX,XX @@ static void kvm_dist_putbmp(GICv3State *s, uint32_t offset,
+@@ -XXX,XX +XXX,XX @@ static void mps2_common_init(MachineState *machine)
-         if (clroffset != 0) {
+x4002a000};   /* Shield1 */
-             reg = 0;
+         sysbus_create_simple(TYPE_ARM_SBCON_I2C, i2cbase[i], NULL);
-             kvm_gicd_access(s, clroffset, &reg, true);
+     }
-+            clroffset += 4;
++    create_unimplemented_device("i2s", 0x40024000, 0x400);
-         }
-         reg = *gic_bmp_ptr32(bmp, irq);
+     /* In hardware this is a LAN9220; the LAN9118 is software compatible
-         kvm_gicd_access(s, offset, &reg, true);
+      * except that it doesn't support the checksum-offload feature.
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 09/25] Correct CPACR reset value for v7 cores
+[PULL 39/42] hw/arm/mps2-tz: Use the ARM SBCon two-wire serial bus interface
-In commit f0aff255700 we made cpacr_write() enforce that some CPACR
+From: Philippe Mathieu-Daudé <f4bug@amsat.org>
 bits are RAZ/WI and some are RAO/WI for ARMv7 cores. Unfortunately
 we forgot to also update the register's reset value. The effect
 was that (a) a guest that read CPACR on reset would not see ones in
 the RAO bits, and (b) if you did a migration before the guest did
 a write to the CPACR then the migration would fail because the
 destination would enforce the RAO bits and then complain that they
 didn't match the zero value from the source.
-Implement reset for the CPACR using a custom reset function
+From 'Application Note AN521', chapter 4.7:
 that just calls cpacr_write(), to avoid having to duplicate
 the logic for which bits are RAO.
-This bug would affect migration for TCG CPUs which are ARMv7
+  The SMM implements four SBCon serial modules:
 with VFP but without one of Neon or VFPv3.
-Reported-by: Cédric Le Goater <clg@kaod.org>
+  One SBCon module for use by the Color LCD touch interface.
   One SBCon module to configure the audio controller.
   Two general purpose SBCon modules, that connect to the
   Expansion headers J7 and J8, are intended for use with the
   V2C-Shield1 which provide an I2C interface on the headers.
 Signed-off-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
 Message-id: 20200617072539.32686-15-f4bug@amsat.org
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Tested-by: Cédric Le Goater <clg@kaod.org>
-Message-id: 20180522173713.26282-1-peter.maydell@linaro.org
 ---
- target/arm/helper.c | 10 +++++++++-
+ hw/arm/mps2-tz.c | 23 ++++++++++++++++++-----
-file changed, 9 insertions(+), 1 deletion(-)
+file changed, 18 insertions(+), 5 deletions(-)
-diff --git a/target/arm/helper.c b/target/arm/helper.c
+diff --git a/hw/arm/mps2-tz.c b/hw/arm/mps2-tz.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/arm/helper.c
+--- a/hw/arm/mps2-tz.c
-+++ b/target/arm/helper.c
++++ b/hw/arm/mps2-tz.c
-@@ -XXX,XX +XXX,XX @@ static void cpacr_write(CPUARMState *env, const ARMCPRegInfo *ri,
+@@ -XXX,XX +XXX,XX @@
-     env->cp15.cpacr_el1 = value;
+ #include "hw/arm/armsse.h"
  #include "hw/dma/pl080.h"
  #include "hw/ssi/pl022.h"
 +#include "hw/i2c/arm_sbcon_i2c.h"
  #include "hw/net/lan9118.h"
  #include "net/net.h"
  #include "hw/core/split-irq.h"
@@ -XXX,XX +XXX,XX @@ typedef struct {
      TZPPC ppc[5];
      TZMPC ssram_mpc[3];
      PL022State spi[5];
 -    UnimplementedDeviceState i2c[4];
 +    ArmSbconI2CState i2c[4];
      UnimplementedDeviceState i2s_audio;
      UnimplementedDeviceState gpio[4];
      UnimplementedDeviceState gfx;
@@ -XXX,XX +XXX,XX @@ static MemoryRegion *make_spi(MPS2TZMachineState *mms, void *opaque,
      return sysbus_mmio_get_region(s, 0);
  }
-+static void cpacr_reset(CPUARMState *env, const ARMCPRegInfo *ri)
++static MemoryRegion *make_i2c(MPS2TZMachineState *mms, void *opaque,
 +                              const char *name, hwaddr size)
 +{
-+    /* Call cpacr_write() so that we reset with the correct RAO bits set
++    ArmSbconI2CState *i2c = opaque;
-+     * for our CPU features.
++    SysBusDevice *s;
-+     */
++
-+    cpacr_write(env, ri, 0);
++    object_initialize_child(OBJECT(mms), name, i2c, TYPE_ARM_SBCON_I2C);
 +    s = SYS_BUS_DEVICE(i2c);
 +    sysbus_realize(s, &error_fatal);
 +    return sysbus_mmio_get_region(s, 0);
 +}
 +
- static CPAccessResult cpacr_access(CPUARMState *env, const ARMCPRegInfo *ri,
+ static void mps2tz_common_init(MachineState *machine)
                                     bool isread)
  {
-@@ -XXX,XX +XXX,XX @@ static const ARMCPRegInfo v6_cp_reginfo[] = {
+     MPS2TZMachineState *mms = MPS2TZ_MACHINE(machine);
-     { .name = "CPACR", .state = ARM_CP_STATE_BOTH, .opc0 = 3,
+@@ -XXX,XX +XXX,XX @@ static void mps2tz_common_init(MachineState *machine)
-       .crn = 1, .crm = 0, .opc1 = 0, .opc2 = 2, .accessfn = cpacr_access,
+                 { "uart2", make_uart, &mms->uart[2], 0x40202000, 0x1000 },
-       .access = PL1_RW, .fieldoffset = offsetof(CPUARMState, cp15.cpacr_el1),
+                 { "uart3", make_uart, &mms->uart[3], 0x40203000, 0x1000 },
--      .resetvalue = 0, .writefn = cpacr_write },
+                 { "uart4", make_uart, &mms->uart[4], 0x40204000, 0x1000 },
-+      .resetfn = cpacr_reset, .writefn = cpacr_write },
+-                { "i2c0", make_unimp_dev, &mms->i2c[0], 0x40207000, 0x1000 },
-     REGINFO_SENTINEL
+-                { "i2c1", make_unimp_dev, &mms->i2c[1], 0x40208000, 0x1000 },
- };
+-                { "i2c2", make_unimp_dev, &mms->i2c[2], 0x4020c000, 0x1000 },
+-                { "i2c3", make_unimp_dev, &mms->i2c[3], 0x4020d000, 0x1000 },
 +                { "i2c0", make_i2c, &mms->i2c[0], 0x40207000, 0x1000 },
 +                { "i2c1", make_i2c, &mms->i2c[1], 0x40208000, 0x1000 },
 +                { "i2c2", make_i2c, &mms->i2c[2], 0x4020c000, 0x1000 },
 +                { "i2c3", make_i2c, &mms->i2c[3], 0x4020d000, 0x1000 },
              },
          }, {
              .name = "apb_ppcexp2",
 --
-.17.1
+.20.1

-New patch
+[PULL 40/42] target/arm: Check supported KVM features globally (not per vCPU)
+From: Philippe Mathieu-Daudé <philmd@redhat.com>
 Since commit d70c996df23f, when enabling the PMU we get:
   $ qemu-system-aarch64 -cpu host,pmu=on -M virt,accel=kvm,gic-version=3
   Segmentation fault (core dumped)
   Thread 1 "qemu-system-aar" received signal SIGSEGV, Segmentation fault.
 x0000aaaaaae356d0 in kvm_ioctl (s=0x0, type=44547) at accel/kvm/kvm-all.c:2588
 ret = ioctl(s->fd, type, arg);
   (gdb) bt
   #0  0x0000aaaaaae356d0 in kvm_ioctl (s=0x0, type=44547) at accel/kvm/kvm-all.c:2588
   #1  0x0000aaaaaae31568 in kvm_check_extension (s=0x0, extension=126) at accel/kvm/kvm-all.c:916
   #2  0x0000aaaaaafce254 in kvm_arm_pmu_supported (cpu=0xaaaaac214ab0) at target/arm/kvm.c:213
   #3  0x0000aaaaaafc0f94 in arm_set_pmu (obj=0xaaaaac214ab0, value=true, errp=0xffffffffe438) at target/arm/cpu.c:1111
   #4  0x0000aaaaab5533ac in property_set_bool (obj=0xaaaaac214ab0, v=0xaaaaac223a80, name=0xaaaaac11a970 "pmu", opaque=0xaaaaac222730, errp=0xffffffffe438) at qom/object.c:2170
   #5  0x0000aaaaab5512f0 in object_property_set (obj=0xaaaaac214ab0, v=0xaaaaac223a80, name=0xaaaaac11a970 "pmu", errp=0xffffffffe438) at qom/object.c:1328
   #6  0x0000aaaaab551e10 in object_property_parse (obj=0xaaaaac214ab0, string=0xaaaaac11b4c0 "on", name=0xaaaaac11a970 "pmu", errp=0xffffffffe438) at qom/object.c:1561
   #7  0x0000aaaaab54ee8c in object_apply_global_props (obj=0xaaaaac214ab0, props=0xaaaaac018e20, errp=0xaaaaabd6fd88 <error_fatal>) at qom/object.c:407
   #8  0x0000aaaaab1dd5a4 in qdev_prop_set_globals (dev=0xaaaaac214ab0) at hw/core/qdev-properties.c:1218
   #9  0x0000aaaaab1d9fac in device_post_init (obj=0xaaaaac214ab0) at hw/core/qdev.c:1050
   ...
   #15 0x0000aaaaab54f310 in object_initialize_with_type (obj=0xaaaaac214ab0, size=52208, type=0xaaaaabe237f0) at qom/object.c:512
   #16 0x0000aaaaab54fa24 in object_new_with_type (type=0xaaaaabe237f0) at qom/object.c:687
   #17 0x0000aaaaab54fa80 in object_new (typename=0xaaaaabe23970 "host-arm-cpu") at qom/object.c:702
   #18 0x0000aaaaaaf04a74 in machvirt_init (machine=0xaaaaac0a8550) at hw/arm/virt.c:1770
   #19 0x0000aaaaab1e8720 in machine_run_board_init (machine=0xaaaaac0a8550) at hw/core/machine.c:1138
   #20 0x0000aaaaaaf95394 in qemu_init (argc=5, argv=0xffffffffea58, envp=0xffffffffea88) at softmmu/vl.c:4348
   #21 0x0000aaaaaada3f74 in main (argc=<optimized out>, argv=<optimized out>, envp=<optimized out>) at softmmu/main.c:48
 This is because in frame #2, cpu->kvm_state is still NULL
 (the vCPU is not yet realized).
 KVM has a hard requirement of all cores supporting the same
 feature set. We only need to check if the accelerator supports
 a feature, not each vCPU individually.
 Fix by removing the 'CPUState *cpu' argument from the
 kvm_arm_<FEATURE>_supported() functions.
 Fixes: d70c996df23f ('Use CPUState::kvm_state in kvm_arm_pmu_supported')
 Reported-by: Haibo Xu <haibo.xu@linaro.org>
 Reviewed-by: Andrew Jones <drjones@redhat.com>
 Acked-by: Paolo Bonzini <pbonzini@redhat.com>
 Signed-off-by: Philippe Mathieu-Daudé <philmd@redhat.com>
 Suggested-by: Paolo Bonzini <pbonzini@redhat.com>
 Signed-off-by: Philippe Mathieu-Daudé <philmd@redhat.com>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
  target/arm/kvm_arm.h | 21 +++++++++------------
  target/arm/cpu.c     |  2 +-
  target/arm/cpu64.c   | 10 +++++-----
  target/arm/kvm.c     |  4 ++--
  target/arm/kvm64.c   | 14 +++++---------
 files changed, 22 insertions(+), 29 deletions(-)
 diff --git a/target/arm/kvm_arm.h b/target/arm/kvm_arm.h
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/kvm_arm.h
 +++ b/target/arm/kvm_arm.h
@@ -XXX,XX +XXX,XX @@ void kvm_arm_add_vcpu_properties(Object *obj);
  /**
   * kvm_arm_aarch32_supported:
 - * @cs: CPUState
   *
 - * Returns: true if the KVM VCPU can enable AArch32 mode
 + * Returns: true if KVM can enable AArch32 mode
   * and false otherwise.
   */
 -bool kvm_arm_aarch32_supported(CPUState *cs);
 +bool kvm_arm_aarch32_supported(void);
  /**
   * kvm_arm_pmu_supported:
 - * @cs: CPUState
   *
 - * Returns: true if the KVM VCPU can enable its PMU
 + * Returns: true if KVM can enable the PMU
   * and false otherwise.
   */
 -bool kvm_arm_pmu_supported(CPUState *cs);
 +bool kvm_arm_pmu_supported(void);
  /**
   * kvm_arm_sve_supported:
 - * @cs: CPUState
   *
 - * Returns true if the KVM VCPU can enable SVE and false otherwise.
 + * Returns true if KVM can enable SVE and false otherwise.
   */
 -bool kvm_arm_sve_supported(CPUState *cs);
 +bool kvm_arm_sve_supported(void);
  /**
   * kvm_arm_get_max_vm_ipa_size:
@@ -XXX,XX +XXX,XX @@ static inline void kvm_arm_set_cpu_features_from_host(ARMCPU *cpu)
  static inline void kvm_arm_add_vcpu_properties(Object *obj) {}
 -static inline bool kvm_arm_aarch32_supported(CPUState *cs)
 +static inline bool kvm_arm_aarch32_supported(void)
  {
      return false;
  }
 -static inline bool kvm_arm_pmu_supported(CPUState *cs)
 +static inline bool kvm_arm_pmu_supported(void)
  {
      return false;
  }
 -static inline bool kvm_arm_sve_supported(CPUState *cs)
 +static inline bool kvm_arm_sve_supported(void)
  {
      return false;
  }
 diff --git a/target/arm/cpu.c b/target/arm/cpu.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/cpu.c
 +++ b/target/arm/cpu.c
@@ -XXX,XX +XXX,XX @@ static void arm_set_pmu(Object *obj, bool value, Error **errp)
      ARMCPU *cpu = ARM_CPU(obj);
      if (value) {
 -        if (kvm_enabled() && !kvm_arm_pmu_supported(CPU(cpu))) {
 +        if (kvm_enabled() && !kvm_arm_pmu_supported()) {
              error_setg(errp, "'pmu' feature not supported by KVM on this host");
              return;
          }
 diff --git a/target/arm/cpu64.c b/target/arm/cpu64.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/cpu64.c
 +++ b/target/arm/cpu64.c
@@ -XXX,XX +XXX,XX @@ void arm_cpu_sve_finalize(ARMCPU *cpu, Error **errp)
      /* Collect the set of vector lengths supported by KVM. */
      bitmap_zero(kvm_supported, ARM_MAX_VQ);
 -    if (kvm_enabled() && kvm_arm_sve_supported(CPU(cpu))) {
 +    if (kvm_enabled() && kvm_arm_sve_supported()) {
          kvm_arm_sve_get_vls(CPU(cpu), kvm_supported);
      } else if (kvm_enabled()) {
          assert(!cpu_isar_feature(aa64_sve, cpu));
@@ -XXX,XX +XXX,XX @@ static void cpu_max_set_sve_max_vq(Object *obj, Visitor *v, const char *name,
          return;
      }
 -    if (kvm_enabled() && !kvm_arm_sve_supported(CPU(cpu))) {
 +    if (kvm_enabled() && !kvm_arm_sve_supported()) {
          error_setg(errp, "cannot set sve-max-vq");
          error_append_hint(errp, "SVE not supported by KVM on this host\n");
          return;
@@ -XXX,XX +XXX,XX @@ static void cpu_arm_set_sve_vq(Object *obj, Visitor *v, const char *name,
          return;
      }
 -    if (value && kvm_enabled() && !kvm_arm_sve_supported(CPU(cpu))) {
 +    if (value && kvm_enabled() && !kvm_arm_sve_supported()) {
          error_setg(errp, "cannot enable %s", name);
          error_append_hint(errp, "SVE not supported by KVM on this host\n");
          return;
@@ -XXX,XX +XXX,XX @@ static void cpu_arm_set_sve(Object *obj, Visitor *v, const char *name,
          return;
      }
 -    if (value && kvm_enabled() && !kvm_arm_sve_supported(CPU(cpu))) {
 +    if (value && kvm_enabled() && !kvm_arm_sve_supported()) {
          error_setg(errp, "'sve' feature not supported by KVM on this host");
          return;
      }
@@ -XXX,XX +XXX,XX @@ static void aarch64_cpu_set_aarch64(Object *obj, bool value, Error **errp)
       * uniform execution state like do_interrupt.
       */
      if (value == false) {
 -        if (!kvm_enabled() || !kvm_arm_aarch32_supported(CPU(cpu))) {
 +        if (!kvm_enabled() || !kvm_arm_aarch32_supported()) {
              error_setg(errp, "'aarch64' feature cannot be disabled "
                               "unless KVM is enabled and 32-bit EL1 "
                               "is supported");
 diff --git a/target/arm/kvm.c b/target/arm/kvm.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/kvm.c
 +++ b/target/arm/kvm.c
@@ -XXX,XX +XXX,XX @@ void kvm_arm_add_vcpu_properties(Object *obj)
      }
  }
 -bool kvm_arm_pmu_supported(CPUState *cpu)
 +bool kvm_arm_pmu_supported(void)
  {
 -    return kvm_check_extension(cpu->kvm_state, KVM_CAP_ARM_PMU_V3);
 +    return kvm_check_extension(kvm_state, KVM_CAP_ARM_PMU_V3);
  }
  int kvm_arm_get_max_vm_ipa_size(MachineState *ms)
 diff --git a/target/arm/kvm64.c b/target/arm/kvm64.c
 index XXXXXXX..XXXXXXX 100644
 --- a/target/arm/kvm64.c
 +++ b/target/arm/kvm64.c
@@ -XXX,XX +XXX,XX @@ bool kvm_arm_get_host_cpu_features(ARMHostCPUFeatures *ahcf)
      return true;
  }
 -bool kvm_arm_aarch32_supported(CPUState *cpu)
 +bool kvm_arm_aarch32_supported(void)
  {
 -    KVMState *s = KVM_STATE(current_accel());
 -
 -    return kvm_check_extension(s, KVM_CAP_ARM_EL1_32BIT);
 +    return kvm_check_extension(kvm_state, KVM_CAP_ARM_EL1_32BIT);
  }
 -bool kvm_arm_sve_supported(CPUState *cpu)
 +bool kvm_arm_sve_supported(void)
  {
 -    KVMState *s = KVM_STATE(current_accel());
 -
 -    return kvm_check_extension(s, KVM_CAP_ARM_SVE);
 +    return kvm_check_extension(kvm_state, KVM_CAP_ARM_SVE);
  }
  QEMU_BUILD_BUG_ON(KVM_ARM64_SVE_VQ_MIN != 1);
@@ -XXX,XX +XXX,XX @@ int kvm_arch_init_vcpu(CPUState *cs)
          env->features &= ~(1ULL << ARM_FEATURE_PMU);
      }
      if (cpu_isar_feature(aa64_sve, cpu)) {
 -        assert(kvm_arm_sve_supported(cs));
 +        assert(kvm_arm_sve_supported());
          cpu->kvm_init_features[0] |= 1 << KVM_ARM_VCPU_SVE;
      }
 --
 .20.1

-[Qemu-devel] [PULL 06/25] arm: fix qemu crash on startup with -bios option
+[PULL 41/42] tests/qtest/arm-cpu-features: Add feature setting tests
-From: Igor Mammedov <imammedo@redhat.com>
+From: Andrew Jones <drjones@redhat.com>
-When QEMU is started with following CLI
+Some cpu features may be enabled and disabled for all configurations
- -machine virt,gic-version=3,accel=kvm -cpu host -bios AAVMF_CODE.fd
+that support the feature. Let's test that.
 it crashes with abort at
  accel/kvm/kvm-all.c:2164:
  KVM_SET_DEVICE_ATTR failed: Group 6 attr 0x000000000000c665: Invalid argument
-Which is caused by implicit dependency of kvm_arm_gicv3_reset() on
+A recent regression[*] inspired adding these tests.
 arm_gicv3_icc_reset() where the later is called by CPU reset
 reset callback.
-However commit:
+[*] '-cpu host,pmu=on' caused a segfault
 b77f6c arm/boot: split load_dtb() from arm_load_kernel()
 broke CPU reset callback registration in case
-  arm_load_kernel()
+Signed-off-by: Andrew Jones <drjones@redhat.com>
-      ...
+Signed-off-by: Philippe Mathieu-Daudé <philmd@redhat.com>
-      if (!info->kernel_filename || info->firmware_loaded)
+Message-id: 20200623090622.30365-2-philmd@redhat.com
+Message-Id: <20200623082310.17577-1-drjones@redhat.com>
 branch is taken, i.e. it's sufficient to provide a firmware
 or do not provide kernel on CLI to skip cpu reset callback
 registration, where before offending commit the callback
 has been registered unconditionally.
 Fix it by registering the callback right at the beginning of
 arm_load_kernel() unconditionally instead of doing it at the end.
 NOTE:
  we probably should eliminate that dependency anyways as well as
  separate arch CPU reset parts from arm_load_kernel() into CPU
  itself, but that refactoring that I probably would have to do
  anyways later for CPU hotplug to work.
 Reported-by: Auger Eric <eric.auger@redhat.com>
 Signed-off-by: Igor Mammedov <imammedo@redhat.com>
 Reviewed-by: Eric Auger <eric.auger@redhat.com>
 Tested-by: Eric Auger <eric.auger@redhat.com>
 Message-id: 1527070950-208350-1-git-send-email-imammedo@redhat.com
 Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
 ---
- hw/arm/boot.c | 18 +++++++++---------
+ tests/qtest/arm-cpu-features.c | 38 ++++++++++++++++++++++++++++++----
-file changed, 9 insertions(+), 9 deletions(-)
+file changed, 34 insertions(+), 4 deletions(-)
-diff --git a/hw/arm/boot.c b/hw/arm/boot.c
+diff --git a/tests/qtest/arm-cpu-features.c b/tests/qtest/arm-cpu-features.c
 index XXXXXXX..XXXXXXX 100644
---- a/hw/arm/boot.c
+--- a/tests/qtest/arm-cpu-features.c
-+++ b/hw/arm/boot.c
++++ b/tests/qtest/arm-cpu-features.c
-@@ -XXX,XX +XXX,XX @@ void arm_load_kernel(ARMCPU *cpu, struct arm_boot_info *info)
+@@ -XXX,XX +XXX,XX @@ static bool resp_get_feature(QDict *resp, const char *feature)
-     static const ARMInsnFixup *primary_loader;
+     qobject_unref(_resp);                                              \
-     AddressSpace *as = arm_boot_address_space(cpu, info);
+ })
-+    /* CPU objects (unlike devices) are not automatically reset on system
+-#define assert_feature(qts, cpu_type, feature, expected_value)         \
-+     * reset, so we must always register a handler to do so. If we're
++#define resp_assert_feature(resp, feature, expected_value)             \
-+     * actually loading a kernel, the handler is also responsible for
+ ({                                                                     \
-+     * arranging that we start it correctly.
+-    QDict *_resp, *_props;                                             \
-+     */
++    QDict *_props;                                                     \
-+    for (cs = first_cpu; cs; cs = CPU_NEXT(cs)) {
+                                                                        \
-+        qemu_register_reset(do_cpu_reset, ARM_CPU(cs));
+-    _resp = do_query_no_props(qts, cpu_type);                          \
-+    }
+     g_assert(_resp);                                                   \
      g_assert(resp_has_props(_resp));                                   \
      _props = resp_get_props(_resp);                                    \
      g_assert(qdict_get(_props, feature));                              \
      g_assert(qdict_get_bool(_props, feature) == (expected_value));     \
 +})
 +
-     /* The board code is not supposed to set secure_board_setup unless
++#define assert_feature(qts, cpu_type, feature, expected_value)         \
-      * running its code in secure mode is actually possible, and KVM
++({                                                                     \
-      * doesn't support secure.
++    QDict *_resp;                                                      \
-@@ -XXX,XX +XXX,XX @@ void arm_load_kernel(ARMCPU *cpu, struct arm_boot_info *info)
++                                                                       \
-         ARM_CPU(cs)->env.boot_info = info;
++    _resp = do_query_no_props(qts, cpu_type);                          \
 +    g_assert(_resp);                                                   \
 +    resp_assert_feature(_resp, feature, expected_value);               \
 +    qobject_unref(_resp);                                              \
 +})
 +
 +#define assert_set_feature(qts, cpu_type, feature, value)              \
 +({                                                                     \
 +    const char *_fmt = (value) ? "{ %s: true }" : "{ %s: false }";     \
 +    QDict *_resp;                                                      \
 +                                                                       \
 +    _resp = do_query(qts, cpu_type, _fmt, feature);                    \
 +    g_assert(_resp);                                                   \
 +    resp_assert_feature(_resp, feature, value);                        \
      qobject_unref(_resp);                                              \
  })
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion(const void *data)
      assert_error(qts, "host", "The CPU type 'host' requires KVM", NULL);
      /* Test expected feature presence/absence for some cpu types */
 -    assert_has_feature_enabled(qts, "max", "pmu");
      assert_has_feature_enabled(qts, "cortex-a15", "pmu");
      assert_has_not_feature(qts, "cortex-a15", "aarch64");
 +    /* Enabling and disabling pmu should always work. */
 +    assert_has_feature_enabled(qts, "max", "pmu");
 +    assert_set_feature(qts, "max", "pmu", false);
 +    assert_set_feature(qts, "max", "pmu", true);
 +
      assert_has_not_feature(qts, "max", "kvm-no-adjvtime");
      if (g_str_equal(qtest_get_arch(), "aarch64")) {
@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
          return;
      }
--    /* CPU objects (unlike devices) are not automatically reset on system
++    /* Enabling and disabling kvm-no-adjvtime should always work. */
--     * reset, so we must always register a handler to do so. If we're
+     assert_has_feature_disabled(qts, "host", "kvm-no-adjvtime");
--     * actually loading a kernel, the handler is also responsible for
++    assert_set_feature(qts, "host", "kvm-no-adjvtime", true);
--     * arranging that we start it correctly.
++    assert_set_feature(qts, "host", "kvm-no-adjvtime", false);
--     */
--    for (cs = first_cpu; cs; cs = CPU_NEXT(cs)) {
+     if (g_str_equal(qtest_get_arch(), "aarch64")) {
--        qemu_register_reset(do_cpu_reset, ARM_CPU(cs));
+         bool kvm_supports_sve;
--    }
+@@ -XXX,XX +XXX,XX @@ static void test_query_cpu_model_expansion_kvm(const void *data)
--
+         char *error;
-     if (!info->skip_dtb_autoload && have_dtb(info)) {
-         if (arm_load_dtb(info->dtb_start, info, info->dtb_limit, as) < 0) {
+         assert_has_feature_enabled(qts, "host", "aarch64");
-             exit(1);
++
 +        /* Enabling and disabling pmu should always work. */
          assert_has_feature_enabled(qts, "host", "pmu");
 +        assert_set_feature(qts, "host", "pmu", false);
 +        assert_set_feature(qts, "host", "pmu", true);
          assert_error(qts, "cortex-a15",
              "We cannot guarantee the CPU type 'cortex-a15' works "
 --
-.17.1
+.20.1

-[Qemu-devel] [PULL 11/25] Make tb_invalidate_phys_addr() take a MemTxAttrs argument
+[PULL 42/42] arm/virt: Add memory hot remove support
-As part of plumbing MemTxAttrs down to the IOMMU translate method,
+From: Shameer Kolothum <shameerali.kolothum.thodi@huawei.com>
 add MemTxAttrs as an argument to tb_invalidate_phys_addr().
 Its callers either have an attrs value to hand, or don't care
 and can use MEMTXATTRS_UNSPECIFIED.
+This adds support for memory(pc-dimm) hot remove on arm/virt that
+uses acpi ged device.
+NVDIMM hot removal is not yet supported.
+Signed-off-by: Shameer Kolothum <shameerali.kolothum.thodi@huawei.com>
+Message-id: 20200622124157.20360-1-shameerali.kolothum.thodi@huawei.com
+Reviewed-by: Eric Auger <eric.auger@redhat.com>
+Tested-by: Eric Auger <eric.auger@redhat.com>
 Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
-Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
-Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
-Message-id: 20180521140402.23318-3-peter.maydell@linaro.org
 ---
- include/exec/exec-all.h   | 5 +++--
+ hw/acpi/generic_event_device.c | 29 ++++++++++++++++
- accel/tcg/translate-all.c | 2 +-
+ hw/arm/virt.c                  | 62 ++++++++++++++++++++++++++++++++--
- exec.c                    | 2 +-
+files changed, 89 insertions(+), 2 deletions(-)
  target/xtensa/op_helper.c | 3 ++-
 files changed, 7 insertions(+), 5 deletions(-)
-diff --git a/include/exec/exec-all.h b/include/exec/exec-all.h
+diff --git a/hw/acpi/generic_event_device.c b/hw/acpi/generic_event_device.c
 index XXXXXXX..XXXXXXX 100644
---- a/include/exec/exec-all.h
+--- a/hw/acpi/generic_event_device.c
-+++ b/include/exec/exec-all.h
++++ b/hw/acpi/generic_event_device.c
-@@ -XXX,XX +XXX,XX @@ void tlb_set_page_with_attrs(CPUState *cpu, target_ulong vaddr,
+@@ -XXX,XX +XXX,XX @@ static void acpi_ged_device_plug_cb(HotplugHandler *hotplug_dev,
  void tlb_set_page(CPUState *cpu, target_ulong vaddr,
                    hwaddr paddr, int prot,
                    int mmu_idx, target_ulong size);
 -void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr);
 +void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs);
  void probe_write(CPUArchState *env, target_ulong addr, int size, int mmu_idx,
                   uintptr_t retaddr);
  #else
@@ -XXX,XX +XXX,XX @@ static inline void tlb_flush_by_mmuidx_all_cpus_synced(CPUState *cpu,
                                                         uint16_t idxmap)
  {
  }
 -static inline void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr)
 +static inline void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr,
 +                                           MemTxAttrs attrs)
  {
  }
  #endif
 diff --git a/accel/tcg/translate-all.c b/accel/tcg/translate-all.c
 index XXXXXXX..XXXXXXX 100644
 --- a/accel/tcg/translate-all.c
 +++ b/accel/tcg/translate-all.c
@@ -XXX,XX +XXX,XX @@ static TranslationBlock *tb_find_pc(uintptr_t tc_ptr)
  }
  #if !defined(CONFIG_USER_ONLY)
 -void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr)
 +void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs)
  {
      ram_addr_t ram_addr;
      MemoryRegion *mr;
 diff --git a/exec.c b/exec.c
 index XXXXXXX..XXXXXXX 100644
 --- a/exec.c
 +++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static void breakpoint_invalidate(CPUState *cpu, target_ulong pc)
      if (phys != -1) {
          /* Locks grabbed by tb_invalidate_phys_addr */
          tb_invalidate_phys_addr(cpu->cpu_ases[asidx].as,
 -                                phys | (pc & ~TARGET_PAGE_MASK));
 +                                phys | (pc & ~TARGET_PAGE_MASK), attrs);
      }
  }
- #endif
-diff --git a/target/xtensa/op_helper.c b/target/xtensa/op_helper.c
++static void acpi_ged_unplug_request_cb(HotplugHandler *hotplug_dev,
 +                                       DeviceState *dev, Error **errp)
 +{
 +    AcpiGedState *s = ACPI_GED(hotplug_dev);
 +
 +    if ((object_dynamic_cast(OBJECT(dev), TYPE_PC_DIMM) &&
 +                       !(object_dynamic_cast(OBJECT(dev), TYPE_NVDIMM)))) {
 +        acpi_memory_unplug_request_cb(hotplug_dev, &s->memhp_state, dev, errp);
 +    } else {
 +        error_setg(errp, "acpi: device unplug request for unsupported device"
 +                   " type: %s", object_get_typename(OBJECT(dev)));
 +    }
 +}
 +
 +static void acpi_ged_unplug_cb(HotplugHandler *hotplug_dev,
 +                               DeviceState *dev, Error **errp)
 +{
 +    AcpiGedState *s = ACPI_GED(hotplug_dev);
 +
 +    if (object_dynamic_cast(OBJECT(dev), TYPE_PC_DIMM)) {
 +        acpi_memory_unplug_cb(&s->memhp_state, dev, errp);
 +    } else {
 +        error_setg(errp, "acpi: device unplug for unsupported device"
 +                   " type: %s", object_get_typename(OBJECT(dev)));
 +    }
 +}
 +
  static void acpi_ged_send_event(AcpiDeviceIf *adev, AcpiEventStatusBits ev)
  {
      AcpiGedState *s = ACPI_GED(adev);
@@ -XXX,XX +XXX,XX @@ static void acpi_ged_class_init(ObjectClass *class, void *data)
      dc->vmsd = &vmstate_acpi_ged;
      hc->plug = acpi_ged_device_plug_cb;
 +    hc->unplug_request = acpi_ged_unplug_request_cb;
 +    hc->unplug = acpi_ged_unplug_cb;
      adevc->send_event = acpi_ged_send_event;
  }
 diff --git a/hw/arm/virt.c b/hw/arm/virt.c
 index XXXXXXX..XXXXXXX 100644
---- a/target/xtensa/op_helper.c
+--- a/hw/arm/virt.c
-+++ b/target/xtensa/op_helper.c
++++ b/hw/arm/virt.c
-@@ -XXX,XX +XXX,XX @@ static void tb_invalidate_virtual_addr(CPUXtensaState *env, uint32_t vaddr)
+@@ -XXX,XX +XXX,XX @@ static void virt_machine_device_plug_cb(HotplugHandler *hotplug_dev,
      int ret = xtensa_get_physical_addr(env, false, vaddr, 2, 0,
              &paddr, &page_size, &access);
      if (ret == 0) {
 -        tb_invalidate_phys_addr(&address_space_memory, paddr);
 +        tb_invalidate_phys_addr(&address_space_memory, paddr,
 +                                MEMTXATTRS_UNSPECIFIED);
      }
  }
++static void virt_dimm_unplug_request(HotplugHandler *hotplug_dev,
++                                     DeviceState *dev, Error **errp)
++{
++    VirtMachineState *vms = VIRT_MACHINE(hotplug_dev);
++    Error *local_err = NULL;
++
++    if (!vms->acpi_dev) {
++        error_setg(&local_err,
++                   "memory hotplug is not enabled: missing acpi-ged device");
++        goto out;
++    }
++
++    if (object_dynamic_cast(OBJECT(dev), TYPE_NVDIMM)) {
++        error_setg(&local_err,
++                   "nvdimm device hot unplug is not supported yet.");
++        goto out;
++    }
++
++    hotplug_handler_unplug_request(HOTPLUG_HANDLER(vms->acpi_dev), dev,
++                                   &local_err);
++out:
++    error_propagate(errp, local_err);
++}
++
++static void virt_dimm_unplug(HotplugHandler *hotplug_dev,
++                             DeviceState *dev, Error **errp)
++{
++    VirtMachineState *vms = VIRT_MACHINE(hotplug_dev);
++    Error *local_err = NULL;
++
++    hotplug_handler_unplug(HOTPLUG_HANDLER(vms->acpi_dev), dev, &local_err);
++    if (local_err) {
++        goto out;
++    }
++
++    pc_dimm_unplug(PC_DIMM(dev), MACHINE(vms));
++    qdev_unrealize(dev);
++
++out:
++    error_propagate(errp, local_err);
++}
++
+ static void virt_machine_device_unplug_request_cb(HotplugHandler *hotplug_dev,
+                                           DeviceState *dev, Error **errp)
+ {
+-    error_setg(errp, "device unplug request for unsupported device"
+-               " type: %s", object_get_typename(OBJECT(dev)));
++    if (object_dynamic_cast(OBJECT(dev), TYPE_PC_DIMM)) {
++        virt_dimm_unplug_request(hotplug_dev, dev, errp);
++    } else {
++        error_setg(errp, "device unplug request for unsupported device"
++                   " type: %s", object_get_typename(OBJECT(dev)));
++    }
++}
++
++static void virt_machine_device_unplug_cb(HotplugHandler *hotplug_dev,
++                                          DeviceState *dev, Error **errp)
++{
++    if (object_dynamic_cast(OBJECT(dev), TYPE_PC_DIMM)) {
++        virt_dimm_unplug(hotplug_dev, dev, errp);
++    } else {
++        error_setg(errp, "virt: device unplug for unsupported device"
++                   " type: %s", object_get_typename(OBJECT(dev)));
++    }
+ }
+ static HotplugHandler *virt_machine_get_hotplug_handler(MachineState *machine,
+@@ -XXX,XX +XXX,XX @@ static void virt_machine_class_init(ObjectClass *oc, void *data)
+     hc->pre_plug = virt_machine_device_pre_plug_cb;
+     hc->plug = virt_machine_device_plug_cb;
+     hc->unplug_request = virt_machine_device_unplug_request_cb;
++    hc->unplug = virt_machine_device_unplug_cb;
+     mc->numa_mem_supported = true;
+     mc->nvdimm_supported = true;
+     mc->auto_enable_numa_with_memhp = true;
 --
-.17.1
+.20.1

target-arm queue. This has the "plumb txattrs through various
bits of exec.c" patches, and a collection of bug fixes from
various people.

thanks
-- PMM

The following changes since commit a3ac12fba028df90f7b3dbec924995c126c41022:

Merge remote-tracking branch 'remotes/ehabkost/tags/numa-next-pull-request' into staging (2018-05-31 11:12:36 +0100)

are available in the Git repository at:

git://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20180531

for you to fetch changes up to 49d1dca0520ea71bc21867fab6647f474fcf857b:

KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice (2018-05-31 14:52:53 +0100)

----------------------------------------------------------------
target-arm queue:
 * target/arm: Honour FPCR.FZ in FRECPX
 * MAINTAINERS: Add entries for newer MPS2 boards and devices
 * hw/intc/arm_gicv3: Fix APxR<n> register dispatching
 * arm_gicv3_kvm: fix bug in writing zero bits back to the in-kernel
   GIC state
 * tcg: Fix helper function vs host abi for float16
 * arm: fix qemu crash on startup with -bios option
 * arm: fix malloc type mismatch
 * xlnx-zdma: Correct mem leaks and memset to zero on desc unaligned errors
 * Correct CPACR reset value for v7 cores
 * memory.h: Improve IOMMU related documentation
 * exec: Plumb transaction attributes through various functions in
   preparation for allowing IOMMUs to see them
 * vmstate.h: Provide VMSTATE_BOOL_SUB_ARRAY
 * ARM: ACPI: Fix use-after-free due to memory realloc
 * KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice

----------------------------------------------------------------
Francisco Iglesias (1):
      xlnx-zdma: Correct mem leaks and memset to zero on desc unaligned errors

Igor Mammedov (1):
      arm: fix qemu crash on startup with -bios option

Jan Kiszka (1):
      hw/intc/arm_gicv3: Fix APxR<n> register dispatching

Paolo Bonzini (1):
      arm: fix malloc type mismatch

Peter Maydell (17):
      target/arm: Honour FPCR.FZ in FRECPX
      MAINTAINERS: Add entries for newer MPS2 boards and devices
      Correct CPACR reset value for v7 cores
      memory.h: Improve IOMMU related documentation
      Make tb_invalidate_phys_addr() take a MemTxAttrs argument
      Make address_space_translate{, _cached}() take a MemTxAttrs argument
      Make address_space_map() take a MemTxAttrs argument
      Make address_space_access_valid() take a MemTxAttrs argument
      Make flatview_extend_translation() take a MemTxAttrs argument
      Make memory_region_access_valid() take a MemTxAttrs argument
      Make MemoryRegion valid.accepts callback take a MemTxAttrs argument
      Make flatview_access_valid() take a MemTxAttrs argument
      Make flatview_translate() take a MemTxAttrs argument
      Make address_space_get_iotlb_entry() take a MemTxAttrs argument
      Make flatview_do_translate() take a MemTxAttrs argument
      Make address_space_translate_iommu take a MemTxAttrs argument
      vmstate.h: Provide VMSTATE_BOOL_SUB_ARRAY

Richard Henderson (1):
      tcg: Fix helper function vs host abi for float16

Shannon Zhao (3):
      arm_gicv3_kvm: increase clroffset accordingly
      ARM: ACPI: Fix use-after-free due to memory realloc
      KVM: GIC: Fix memory leak due to calling kvm_init_irq_routing twice

The FRECPX instructions should (like most other floating point operations)
honour the FPCR.FZ bit which specifies whether input denormals should
be flushed to zero (or FZ16 for the half-precision version).
We forgot to implement this, which doesn't affect the results (since
the calculation doesn't actually care about the mantissa bits) but did
mean we were failing to set the FPSR.IDC bit.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521172712.19930-1-peter.maydell@linaro.org
---
 target/arm/helper-a64.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/target/arm/helper-a64.c b/target/arm/helper-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/helper-a64.c
+++ b/target/arm/helper-a64.c
@@ -XXX,XX +XXX,XX @@ float16 HELPER(frecpx_f16)(float16 a, void *fpstp)
         return nan;
     }
 
+    a = float16_squash_input_denormal(a, fpst);
+
     val16 = float16_val(a);
     sbit = 0x8000 & val16;
     exp = extract32(val16, 10, 5);
@@ -XXX,XX +XXX,XX @@ float32 HELPER(frecpx_f32)(float32 a, void *fpstp)
         return nan;
     }
 
+    a = float32_squash_input_denormal(a, fpst);
+
     val32 = float32_val(a);
     sbit = 0x80000000ULL & val32;
     exp = extract32(val32, 23, 8);
@@ -XXX,XX +XXX,XX @@ float64 HELPER(frecpx_f64)(float64 a, void *fpstp)
         return nan;
     }
 
+    a = float64_squash_input_denormal(a, fpst);
+
     val64 = float64_val(a);
     sbit = 0x8000000000000000ULL & val64;
     exp = extract64(float64_val(a), 52, 11);
-- 
2.17.1

Add entries to MAINTAINERS to cover the newer MPS2 boards and
the new devices they use.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Message-id: 20180518153157.14899-1-peter.maydell@linaro.org
---
 MAINTAINERS | 9 +++++++--
 1 file changed, 7 insertions(+), 2 deletions(-)

diff --git a/MAINTAINERS b/MAINTAINERS
index XXXXXXX..XXXXXXX 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -XXX,XX +XXX,XX @@ F: hw/timer/cmsdk-apb-timer.c
 F: include/hw/timer/cmsdk-apb-timer.h
 F: hw/char/cmsdk-apb-uart.c
 F: include/hw/char/cmsdk-apb-uart.h
+F: hw/misc/tz-ppc.c
+F: include/hw/misc/tz-ppc.h
 
 ARM cores
 M: Peter Maydell <peter.maydell@linaro.org>
@@ -XXX,XX +XXX,XX @@ M: Peter Maydell <peter.maydell@linaro.org>
 L: qemu-arm@nongnu.org
 S: Maintained
 F: hw/arm/mps2.c
-F: hw/misc/mps2-scc.c
-F: include/hw/misc/mps2-scc.h
+F: hw/arm/mps2-tz.c
+F: hw/misc/mps2-*.c
+F: include/hw/misc/mps2-*.h
+F: hw/arm/iotkit.c
+F: include/hw/arm/iotkit.h
 
 Musicpal
 M: Jan Kiszka <jan.kiszka@web.de>
-- 
2.17.1

From: Jan Kiszka <jan.kiszka@siemens.com>

There was a nasty flip in identifying which register group an access is
targeting. The issue caused spuriously raised priorities of the guest
when handing CPUs over in the Jailhouse hypervisor.

Cc: qemu-stable@nongnu.org
Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com>
Message-id: 28b927d3-da58-bce4-cc13-bfec7f9b1cb9@siemens.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/intc/arm_gicv3_cpuif.c | 12 ++++++------
 1 file changed, 6 insertions(+), 6 deletions(-)

diff --git a/hw/intc/arm_gicv3_cpuif.c b/hw/intc/arm_gicv3_cpuif.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/intc/arm_gicv3_cpuif.c
+++ b/hw/intc/arm_gicv3_cpuif.c
@@ -XXX,XX +XXX,XX @@ static uint64_t icv_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
 {
     GICv3CPUState *cs = icc_cs_from_env(env);
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
     uint64_t value = cs->ich_apr[grp][regno];
 
     trace_gicv3_icv_ap_read(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
@@ -XXX,XX +XXX,XX @@ static void icv_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
 {
     GICv3CPUState *cs = icc_cs_from_env(env);
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
 
     trace_gicv3_icv_ap_write(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
 
@@ -XXX,XX +XXX,XX @@ static uint64_t icc_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
     uint64_t value;
 
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1;
+    int grp = (ri->crm & 1) ? GICV3_G1 : GICV3_G0;
 
     if (icv_access(env, grp == GICV3_G0 ? HCR_FMO : HCR_IMO)) {
         return icv_ap_read(env, ri);
@@ -XXX,XX +XXX,XX @@ static void icc_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
     GICv3CPUState *cs = icc_cs_from_env(env);
 
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1;
+    int grp = (ri->crm & 1) ? GICV3_G1 : GICV3_G0;
 
     if (icv_access(env, grp == GICV3_G0 ? HCR_FMO : HCR_IMO)) {
         icv_ap_write(env, ri, value);
@@ -XXX,XX +XXX,XX @@ static uint64_t ich_ap_read(CPUARMState *env, const ARMCPRegInfo *ri)
 {
     GICv3CPUState *cs = icc_cs_from_env(env);
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
     uint64_t value;
 
     value = cs->ich_apr[grp][regno];
@@ -XXX,XX +XXX,XX @@ static void ich_ap_write(CPUARMState *env, const ARMCPRegInfo *ri,
 {
     GICv3CPUState *cs = icc_cs_from_env(env);
     int regno = ri->opc2 & 3;
-    int grp = ri->crm & 1 ? GICV3_G0 : GICV3_G1NS;
+    int grp = (ri->crm & 1) ? GICV3_G1NS : GICV3_G0;
 
     trace_gicv3_ich_ap_write(ri->crm & 1, regno, gicv3_redist_affid(cs), value);
 
-- 
2.17.1

From: Shannon Zhao <zhaoshenglong@huawei.com>

It forgot to increase clroffset during the loop. So it only clear the
first 4 bytes.

Fixes: 367b9f527becdd20ddf116e17a3c0c2bbc486920
Cc: qemu-stable@nongnu.org
Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Message-id: 1527047633-12368-1-git-send-email-zhaoshenglong@huawei.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/intc/arm_gicv3_kvm.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/hw/intc/arm_gicv3_kvm.c b/hw/intc/arm_gicv3_kvm.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/intc/arm_gicv3_kvm.c
+++ b/hw/intc/arm_gicv3_kvm.c
@@ -XXX,XX +XXX,XX @@ static void kvm_dist_putbmp(GICv3State *s, uint32_t offset,
         if (clroffset != 0) {
             reg = 0;
             kvm_gicd_access(s, clroffset, &reg, true);
+            clroffset += 4;
         }
         reg = *gic_bmp_ptr32(bmp, irq);
         kvm_gicd_access(s, offset, &reg, true);
-- 
2.17.1

From: Richard Henderson <richard.henderson@linaro.org>

Depending on the host abi, float16, aka uint16_t, values are
passed and returned either zero-extended in the host register
or with garbage at the top of the host register.

The tcg code generator has so far been assuming garbage, as that
matches the x86 abi, but this is incorrect for other host abis.
Further, target/arm has so far been assuming zero-extended results,
so that it may store the 16-bit value into a 32-bit slot with the
high 16-bits already clear.

Rectify both problems by mapping "f16" in the helper definition
to uint32_t instead of (a typedef for) uint16_t.  This forces
the host compiler to assume garbage in the upper 16 bits on input
and to zero-extend the result on output.

Cc: qemu-stable@nongnu.org
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Tested-by: Laurent Desnogues <laurent.desnogues@gmail.com>
Message-id: 20180522175629.24932-1-richard.henderson@linaro.org
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 include/exec/helper-head.h |  2 +-
 target/arm/helper-a64.c    | 35 +++++++++--------
 target/arm/helper.c        | 80 +++++++++++++++++++-------------------
 3 files changed, 59 insertions(+), 58 deletions(-)

diff --git a/include/exec/helper-head.h b/include/exec/helper-head.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/helper-head.h
+++ b/include/exec/helper-head.h
@@ -XXX,XX +XXX,XX @@
 #define dh_ctype_int int
 #define dh_ctype_i64 uint64_t
 #define dh_ctype_s64 int64_t
-#define dh_ctype_f16 float16
+#define dh_ctype_f16 uint32_t
 #define dh_ctype_f32 float32
 #define dh_ctype_f64 float64
 #define dh_ctype_ptr void *
diff --git a/target/arm/helper-a64.c b/target/arm/helper-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/helper-a64.c
+++ b/target/arm/helper-a64.c
@@ -XXX,XX +XXX,XX @@ static inline uint32_t float_rel_to_flags(int res)
     return flags;
 }
 
-uint64_t HELPER(vfp_cmph_a64)(float16 x, float16 y, void *fp_status)
+uint64_t HELPER(vfp_cmph_a64)(uint32_t x, uint32_t y, void *fp_status)
 {
     return float_rel_to_flags(float16_compare_quiet(x, y, fp_status));
 }
 
-uint64_t HELPER(vfp_cmpeh_a64)(float16 x, float16 y, void *fp_status)
+uint64_t HELPER(vfp_cmpeh_a64)(uint32_t x, uint32_t y, void *fp_status)
 {
     return float_rel_to_flags(float16_compare(x, y, fp_status));
 }
@@ -XXX,XX +XXX,XX @@ uint64_t HELPER(neon_cgt_f64)(float64 a, float64 b, void *fpstp)
 #define float64_three make_float64(0x4008000000000000ULL)
 #define float64_one_point_five make_float64(0x3FF8000000000000ULL)
 
-float16 HELPER(recpsf_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(recpsf_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
 
@@ -XXX,XX +XXX,XX @@ float64 HELPER(recpsf_f64)(float64 a, float64 b, void *fpstp)
     return float64_muladd(a, b, float64_two, 0, fpst);
 }
 
-float16 HELPER(rsqrtsf_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(rsqrtsf_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
 
@@ -XXX,XX +XXX,XX @@ uint64_t HELPER(neon_addlp_u16)(uint64_t a)
 }
 
 /* Floating-point reciprocal exponent - see FPRecpX in ARM ARM */
-float16 HELPER(frecpx_f16)(float16 a, void *fpstp)
+uint32_t HELPER(frecpx_f16)(uint32_t a, void *fpstp)
 {
     float_status *fpst = fpstp;
     uint16_t val16, sbit;
@@ -XXX,XX +XXX,XX @@ void HELPER(casp_be_parallel)(CPUARMState *env, uint32_t rs, uint64_t addr,
 #define ADVSIMD_HELPER(name, suffix) HELPER(glue(glue(advsimd_, name), suffix))
 
 #define ADVSIMD_HALFOP(name) \
-float16 ADVSIMD_HELPER(name, h)(float16 a, float16 b, void *fpstp) \
+uint32_t ADVSIMD_HELPER(name, h)(uint32_t a, uint32_t b, void *fpstp) \
 { \
     float_status *fpst = fpstp; \
     return float16_ ## name(a, b, fpst);    \
@@ -XXX,XX +XXX,XX @@ ADVSIMD_HALFOP(mulx)
 ADVSIMD_TWOHALFOP(mulx)
 
 /* fused multiply-accumulate */
-float16 HELPER(advsimd_muladdh)(float16 a, float16 b, float16 c, void *fpstp)
+uint32_t HELPER(advsimd_muladdh)(uint32_t a, uint32_t b, uint32_t c,
+                                 void *fpstp)
 {
     float_status *fpst = fpstp;
     return float16_muladd(a, b, c, 0, fpst);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_muladd2h)(uint32_t two_a, uint32_t two_b,
 
 #define ADVSIMD_CMPRES(test) (test) ? 0xffff : 0
 
-uint32_t HELPER(advsimd_ceq_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(advsimd_ceq_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
     int compare = float16_compare_quiet(a, b, fpst);
     return ADVSIMD_CMPRES(compare == float_relation_equal);
 }
 
-uint32_t HELPER(advsimd_cge_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(advsimd_cge_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
     int compare = float16_compare(a, b, fpst);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_cge_f16)(float16 a, float16 b, void *fpstp)
                           compare == float_relation_equal);
 }
 
-uint32_t HELPER(advsimd_cgt_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(advsimd_cgt_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
     int compare = float16_compare(a, b, fpst);
     return ADVSIMD_CMPRES(compare == float_relation_greater);
 }
 
-uint32_t HELPER(advsimd_acge_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(advsimd_acge_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
     float16 f0 = float16_abs(a);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_acge_f16)(float16 a, float16 b, void *fpstp)
                           compare == float_relation_equal);
 }
 
-uint32_t HELPER(advsimd_acgt_f16)(float16 a, float16 b, void *fpstp)
+uint32_t HELPER(advsimd_acgt_f16)(uint32_t a, uint32_t b, void *fpstp)
 {
     float_status *fpst = fpstp;
     float16 f0 = float16_abs(a);
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_acgt_f16)(float16 a, float16 b, void *fpstp)
 }
 
 /* round to integral */
-float16 HELPER(advsimd_rinth_exact)(float16 x, void *fp_status)
+uint32_t HELPER(advsimd_rinth_exact)(uint32_t x, void *fp_status)
 {
     return float16_round_to_int(x, fp_status);
 }
 
-float16 HELPER(advsimd_rinth)(float16 x, void *fp_status)
+uint32_t HELPER(advsimd_rinth)(uint32_t x, void *fp_status)
 {
     int old_flags = get_float_exception_flags(fp_status), new_flags;
     float16 ret;
@@ -XXX,XX +XXX,XX @@ float16 HELPER(advsimd_rinth)(float16 x, void *fp_status)
  * setting the mode appropriately before calling the helper.
  */
 
-uint32_t HELPER(advsimd_f16tosinth)(float16 a, void *fpstp)
+uint32_t HELPER(advsimd_f16tosinth)(uint32_t a, void *fpstp)
 {
     float_status *fpst = fpstp;
 
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_f16tosinth)(float16 a, void *fpstp)
     return float16_to_int16(a, fpst);
 }
 
-uint32_t HELPER(advsimd_f16touinth)(float16 a, void *fpstp)
+uint32_t HELPER(advsimd_f16touinth)(uint32_t a, void *fpstp)
 {
     float_status *fpst = fpstp;
 
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(advsimd_f16touinth)(float16 a, void *fpstp)
  * Square Root and Reciprocal square root
  */
 
-float16 HELPER(sqrt_f16)(float16 a, void *fpstp)
+uint32_t HELPER(sqrt_f16)(uint32_t a, void *fpstp)
 {
     float_status *s = fpstp;
 
diff --git a/target/arm/helper.c b/target/arm/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/helper.c
+++ b/target/arm/helper.c
@@ -XXX,XX +XXX,XX @@ DO_VFP_cmp(d, float64)
 
 /* Integer to float and float to integer conversions */
 
-#define CONV_ITOF(name, fsz, sign) \
-    float##fsz HELPER(name)(uint32_t x, void *fpstp) \
-{ \
-    float_status *fpst = fpstp; \
-    return sign##int32_to_##float##fsz((sign##int32_t)x, fpst); \
+#define CONV_ITOF(name, ftype, fsz, sign)                           \
+ftype HELPER(name)(uint32_t x, void *fpstp)                         \
+{                                                                   \
+    float_status *fpst = fpstp;                                     \
+    return sign##int32_to_##float##fsz((sign##int32_t)x, fpst);     \
 }
 
-#define CONV_FTOI(name, fsz, sign, round) \
-uint32_t HELPER(name)(float##fsz x, void *fpstp) \
-{ \
-    float_status *fpst = fpstp; \
-    if (float##fsz##_is_any_nan(x)) { \
-        float_raise(float_flag_invalid, fpst); \
-        return 0; \
-    } \
-    return float##fsz##_to_##sign##int32##round(x, fpst); \
+#define CONV_FTOI(name, ftype, fsz, sign, round)                \
+uint32_t HELPER(name)(ftype x, void *fpstp)                     \
+{                                                               \
+    float_status *fpst = fpstp;                                 \
+    if (float##fsz##_is_any_nan(x)) {                           \
+        float_raise(float_flag_invalid, fpst);                  \
+        return 0;                                               \
+    }                                                           \
+    return float##fsz##_to_##sign##int32##round(x, fpst);       \
 }
 
-#define FLOAT_CONVS(name, p, fsz, sign) \
-CONV_ITOF(vfp_##name##to##p, fsz, sign) \
-CONV_FTOI(vfp_to##name##p, fsz, sign, ) \
-CONV_FTOI(vfp_to##name##z##p, fsz, sign, _round_to_zero)
+#define FLOAT_CONVS(name, p, ftype, fsz, sign)            \
+    CONV_ITOF(vfp_##name##to##p, ftype, fsz, sign)        \
+    CONV_FTOI(vfp_to##name##p, ftype, fsz, sign, )        \
+    CONV_FTOI(vfp_to##name##z##p, ftype, fsz, sign, _round_to_zero)
 
-FLOAT_CONVS(si, h, 16, )
-FLOAT_CONVS(si, s, 32, )
-FLOAT_CONVS(si, d, 64, )
-FLOAT_CONVS(ui, h, 16, u)
-FLOAT_CONVS(ui, s, 32, u)
-FLOAT_CONVS(ui, d, 64, u)
+FLOAT_CONVS(si, h, uint32_t, 16, )
+FLOAT_CONVS(si, s, float32, 32, )
+FLOAT_CONVS(si, d, float64, 64, )
+FLOAT_CONVS(ui, h, uint32_t, 16, u)
+FLOAT_CONVS(ui, s, float32, 32, u)
+FLOAT_CONVS(ui, d, float64, 64, u)
 
 #undef CONV_ITOF
 #undef CONV_FTOI
@@ -XXX,XX +XXX,XX @@ static float16 do_postscale_fp16(float64 f, int shift, float_status *fpst)
     return float64_to_float16(float64_scalbn(f, -shift, fpst), true, fpst);
 }
 
-float16 HELPER(vfp_sltoh)(uint32_t x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_sltoh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return do_postscale_fp16(int32_to_float64(x, fpst), shift, fpst);
 }
 
-float16 HELPER(vfp_ultoh)(uint32_t x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_ultoh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return do_postscale_fp16(uint32_to_float64(x, fpst), shift, fpst);
 }
 
-float16 HELPER(vfp_sqtoh)(uint64_t x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_sqtoh)(uint64_t x, uint32_t shift, void *fpst)
 {
     return do_postscale_fp16(int64_to_float64(x, fpst), shift, fpst);
 }
 
-float16 HELPER(vfp_uqtoh)(uint64_t x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_uqtoh)(uint64_t x, uint32_t shift, void *fpst)
 {
     return do_postscale_fp16(uint64_to_float64(x, fpst), shift, fpst);
 }
@@ -XXX,XX +XXX,XX @@ static float64 do_prescale_fp16(float16 f, int shift, float_status *fpst)
     }
 }
 
-uint32_t HELPER(vfp_toshh)(float16 x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_toshh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_int16(do_prescale_fp16(x, shift, fpst), fpst);
 }
 
-uint32_t HELPER(vfp_touhh)(float16 x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_touhh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_uint16(do_prescale_fp16(x, shift, fpst), fpst);
 }
 
-uint32_t HELPER(vfp_toslh)(float16 x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_toslh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_int32(do_prescale_fp16(x, shift, fpst), fpst);
 }
 
-uint32_t HELPER(vfp_toulh)(float16 x, uint32_t shift, void *fpst)
+uint32_t HELPER(vfp_toulh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_uint32(do_prescale_fp16(x, shift, fpst), fpst);
 }
 
-uint64_t HELPER(vfp_tosqh)(float16 x, uint32_t shift, void *fpst)
+uint64_t HELPER(vfp_tosqh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_int64(do_prescale_fp16(x, shift, fpst), fpst);
 }
 
-uint64_t HELPER(vfp_touqh)(float16 x, uint32_t shift, void *fpst)
+uint64_t HELPER(vfp_touqh)(uint32_t x, uint32_t shift, void *fpst)
 {
     return float64_to_uint64(do_prescale_fp16(x, shift, fpst), fpst);
 }
@@ -XXX,XX +XXX,XX @@ uint32_t HELPER(set_neon_rmode)(uint32_t rmode, CPUARMState *env)
 }
 
 /* Half precision conversions.  */
-float32 HELPER(vfp_fcvt_f16_to_f32)(float16 a, void *fpstp, uint32_t ahp_mode)
+float32 HELPER(vfp_fcvt_f16_to_f32)(uint32_t a, void *fpstp, uint32_t ahp_mode)
 {
     /* Squash FZ16 to 0 for the duration of conversion.  In this case,
      * it would affect flushing input denormals.
@@ -XXX,XX +XXX,XX @@ float32 HELPER(vfp_fcvt_f16_to_f32)(float16 a, void *fpstp, uint32_t ahp_mode)
     return r;
 }
 
-float16 HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
+uint32_t HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
 {
     /* Squash FZ16 to 0 for the duration of conversion.  In this case,
      * it would affect flushing output denormals.
@@ -XXX,XX +XXX,XX @@ float16 HELPER(vfp_fcvt_f32_to_f16)(float32 a, void *fpstp, uint32_t ahp_mode)
     return r;
 }
 
-float64 HELPER(vfp_fcvt_f16_to_f64)(float16 a, void *fpstp, uint32_t ahp_mode)
+float64 HELPER(vfp_fcvt_f16_to_f64)(uint32_t a, void *fpstp, uint32_t ahp_mode)
 {
     /* Squash FZ16 to 0 for the duration of conversion.  In this case,
      * it would affect flushing input denormals.
@@ -XXX,XX +XXX,XX @@ float64 HELPER(vfp_fcvt_f16_to_f64)(float16 a, void *fpstp, uint32_t ahp_mode)
     return r;
 }
 
-float16 HELPER(vfp_fcvt_f64_to_f16)(float64 a, void *fpstp, uint32_t ahp_mode)
+uint32_t HELPER(vfp_fcvt_f64_to_f16)(float64 a, void *fpstp, uint32_t ahp_mode)
 {
     /* Squash FZ16 to 0 for the duration of conversion.  In this case,
      * it would affect flushing output denormals.
@@ -XXX,XX +XXX,XX @@ static bool round_to_inf(float_status *fpst, bool sign_bit)
     g_assert_not_reached();
 }
 
-float16 HELPER(recpe_f16)(float16 input, void *fpstp)
+uint32_t HELPER(recpe_f16)(uint32_t input, void *fpstp)
 {
     float_status *fpst = fpstp;
     float16 f16 = float16_squash_input_denormal(input, fpst);
@@ -XXX,XX +XXX,XX @@ static uint64_t recip_sqrt_estimate(int *exp , int exp_off, uint64_t frac)
     return extract64(estimate, 0, 8) << 44;
 }
 
-float16 HELPER(rsqrte_f16)(float16 input, void *fpstp)
+uint32_t HELPER(rsqrte_f16)(uint32_t input, void *fpstp)
 {
     float_status *s = fpstp;
     float16 f16 = float16_squash_input_denormal(input, s);
-- 
2.17.1

From: Igor Mammedov <imammedo@redhat.com>

When QEMU is started with following CLI
 -machine virt,gic-version=3,accel=kvm -cpu host -bios AAVMF_CODE.fd
it crashes with abort at
 accel/kvm/kvm-all.c:2164:
 KVM_SET_DEVICE_ATTR failed: Group 6 attr 0x000000000000c665: Invalid argument

Which is caused by implicit dependency of kvm_arm_gicv3_reset() on
arm_gicv3_icc_reset() where the later is called by CPU reset
reset callback.

However commit:
 3b77f6c arm/boot: split load_dtb() from arm_load_kernel()
broke CPU reset callback registration in case

arm_load_kernel()
      ...
      if (!info->kernel_filename || info->firmware_loaded)

branch is taken, i.e. it's sufficient to provide a firmware
or do not provide kernel on CLI to skip cpu reset callback
registration, where before offending commit the callback
has been registered unconditionally.

Fix it by registering the callback right at the beginning of
arm_load_kernel() unconditionally instead of doing it at the end.

NOTE:
 we probably should eliminate that dependency anyways as well as
 separate arch CPU reset parts from arm_load_kernel() into CPU
 itself, but that refactoring that I probably would have to do
 anyways later for CPU hotplug to work.

Reported-by: Auger Eric <eric.auger@redhat.com>
Signed-off-by: Igor Mammedov <imammedo@redhat.com>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Tested-by: Eric Auger <eric.auger@redhat.com>
Message-id: 1527070950-208350-1-git-send-email-imammedo@redhat.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/boot.c | 18 +++++++++---------
 1 file changed, 9 insertions(+), 9 deletions(-)

diff --git a/hw/arm/boot.c b/hw/arm/boot.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/boot.c
+++ b/hw/arm/boot.c
@@ -XXX,XX +XXX,XX @@ void arm_load_kernel(ARMCPU *cpu, struct arm_boot_info *info)
     static const ARMInsnFixup *primary_loader;
     AddressSpace *as = arm_boot_address_space(cpu, info);
 
+    /* CPU objects (unlike devices) are not automatically reset on system
+     * reset, so we must always register a handler to do so. If we're
+     * actually loading a kernel, the handler is also responsible for
+     * arranging that we start it correctly.
+     */
+    for (cs = first_cpu; cs; cs = CPU_NEXT(cs)) {
+        qemu_register_reset(do_cpu_reset, ARM_CPU(cs));
+    }
+
     /* The board code is not supposed to set secure_board_setup unless
      * running its code in secure mode is actually possible, and KVM
      * doesn't support secure.
@@ -XXX,XX +XXX,XX @@ void arm_load_kernel(ARMCPU *cpu, struct arm_boot_info *info)
         ARM_CPU(cs)->env.boot_info = info;
     }
 
-    /* CPU objects (unlike devices) are not automatically reset on system
-     * reset, so we must always register a handler to do so. If we're
-     * actually loading a kernel, the handler is also responsible for
-     * arranging that we start it correctly.
-     */
-    for (cs = first_cpu; cs; cs = CPU_NEXT(cs)) {
-        qemu_register_reset(do_cpu_reset, ARM_CPU(cs));
-    }
-
     if (!info->skip_dtb_autoload && have_dtb(info)) {
         if (arm_load_dtb(info->dtb_start, info, info->dtb_limit, as) < 0) {
             exit(1);
-- 
2.17.1

From: Paolo Bonzini <pbonzini@redhat.com>

cpregs_keys is an uint32_t* so the allocation should use uint32_t.
g_new is even better because it is type-safe.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 target/arm/gdbstub.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/target/arm/gdbstub.c b/target/arm/gdbstub.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/gdbstub.c
+++ b/target/arm/gdbstub.c
@@ -XXX,XX +XXX,XX @@ int arm_gen_dynamic_xml(CPUState *cs)
     RegisterSysregXmlParam param = {cs, s};
 
     cpu->dyn_xml.num_cpregs = 0;
-    cpu->dyn_xml.cpregs_keys = g_malloc(sizeof(uint32_t *) *
-                                        g_hash_table_size(cpu->cp_regs));
+    cpu->dyn_xml.cpregs_keys = g_new(uint32_t, g_hash_table_size(cpu->cp_regs));
     g_string_printf(s, "<?xml version=\"1.0\"?>");
     g_string_append_printf(s, "<!DOCTYPE target SYSTEM \"gdb-target.dtd\">");
     g_string_append_printf(s, "<feature name=\"org.qemu.gdb.arm.sys.regs\">");
-- 
2.17.1

From: Francisco Iglesias <frasse.iglesias@gmail.com>

Coverity found that the string return by 'object_get_canonical_path' was not
being freed at two locations in the model (CID 1391294 and CID 1391293) and
also that a memset was being called with a value greater than the max of a byte
on the second argument (CID 1391286). This patch corrects this by adding the
freeing of the strings and also changing to memset to zero instead on
descriptor unaligned errors.

Signed-off-by: Francisco Iglesias <frasse.iglesias@gmail.com>
Reviewed-by: Edgar E. Iglesias <edgar.iglesias@xilinx.com>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Message-id: 20180528184859.3530-1-frasse.iglesias@gmail.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/dma/xlnx-zdma.c | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/hw/dma/xlnx-zdma.c b/hw/dma/xlnx-zdma.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/dma/xlnx-zdma.c
+++ b/hw/dma/xlnx-zdma.c
@@ -XXX,XX +XXX,XX @@ static bool zdma_load_descriptor(XlnxZDMA *s, uint64_t addr, void *buf)
         qemu_log_mask(LOG_GUEST_ERROR,
                       "zdma: unaligned descriptor at %" PRIx64,
                       addr);
-        memset(buf, 0xdeadbeef, sizeof(XlnxZDMADescr));
+        memset(buf, 0x0, sizeof(XlnxZDMADescr));
         s->error = true;
         return false;
     }
@@ -XXX,XX +XXX,XX @@ static uint64_t zdma_read(void *opaque, hwaddr addr, unsigned size)
     RegisterInfo *r = &s->regs_info[addr / 4];
 
     if (!r->data) {
+        gchar *path = object_get_canonical_path(OBJECT(s));
         qemu_log("%s: Decode error: read from %" HWADDR_PRIx "\n",
-                 object_get_canonical_path(OBJECT(s)),
+                 path,
                  addr);
+        g_free(path);
         ARRAY_FIELD_DP32(s->regs, ZDMA_CH_ISR, INV_APB, true);
         zdma_ch_imr_update_irq(s);
         return 0;
@@ -XXX,XX +XXX,XX @@ static void zdma_write(void *opaque, hwaddr addr, uint64_t value,
     RegisterInfo *r = &s->regs_info[addr / 4];
 
     if (!r->data) {
+        gchar *path = object_get_canonical_path(OBJECT(s));
         qemu_log("%s: Decode error: write to %" HWADDR_PRIx "=%" PRIx64 "\n",
-                 object_get_canonical_path(OBJECT(s)),
+                 path,
                  addr, value);
+        g_free(path);
         ARRAY_FIELD_DP32(s->regs, ZDMA_CH_ISR, INV_APB, true);
         zdma_ch_imr_update_irq(s);
         return;
-- 
2.17.1

In commit f0aff255700 we made cpacr_write() enforce that some CPACR
bits are RAZ/WI and some are RAO/WI for ARMv7 cores. Unfortunately
we forgot to also update the register's reset value. The effect
was that (a) a guest that read CPACR on reset would not see ones in
the RAO bits, and (b) if you did a migration before the guest did
a write to the CPACR then the migration would fail because the
destination would enforce the RAO bits and then complain that they
didn't match the zero value from the source.

Implement reset for the CPACR using a custom reset function
that just calls cpacr_write(), to avoid having to duplicate
the logic for which bits are RAO.

This bug would affect migration for TCG CPUs which are ARMv7
with VFP but without one of Neon or VFPv3.

Reported-by: Cédric Le Goater <clg@kaod.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Tested-by: Cédric Le Goater <clg@kaod.org>
Message-id: 20180522173713.26282-1-peter.maydell@linaro.org
---
 target/arm/helper.c | 10 +++++++++-
 1 file changed, 9 insertions(+), 1 deletion(-)

diff --git a/target/arm/helper.c b/target/arm/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/helper.c
+++ b/target/arm/helper.c
@@ -XXX,XX +XXX,XX @@ static void cpacr_write(CPUARMState *env, const ARMCPRegInfo *ri,
     env->cp15.cpacr_el1 = value;
 }
 
+static void cpacr_reset(CPUARMState *env, const ARMCPRegInfo *ri)
+{
+    /* Call cpacr_write() so that we reset with the correct RAO bits set
+     * for our CPU features.
+     */
+    cpacr_write(env, ri, 0);
+}
+
 static CPAccessResult cpacr_access(CPUARMState *env, const ARMCPRegInfo *ri,
                                    bool isread)
 {
@@ -XXX,XX +XXX,XX @@ static const ARMCPRegInfo v6_cp_reginfo[] = {
     { .name = "CPACR", .state = ARM_CP_STATE_BOTH, .opc0 = 3,
       .crn = 1, .crm = 0, .opc1 = 0, .opc2 = 2, .accessfn = cpacr_access,
       .access = PL1_RW, .fieldoffset = offsetof(CPUARMState, cp15.cpacr_el1),
-      .resetvalue = 0, .writefn = cpacr_write },
+      .resetfn = cpacr_reset, .writefn = cpacr_write },
     REGINFO_SENTINEL
 };
 
-- 
2.17.1

Add more detail to the documentation for memory_region_init_iommu()
and other IOMMU-related functions and data structures.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Message-id: 20180521140402.23318-2-peter.maydell@linaro.org
---
 include/exec/memory.h | 105 ++++++++++++++++++++++++++++++++++++++----
 1 file changed, 95 insertions(+), 10 deletions(-)

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ enum IOMMUMemoryRegionAttr {
     IOMMU_ATTR_SPAPR_TCE_FD
 };
 
+/**
+ * IOMMUMemoryRegionClass:
+ *
+ * All IOMMU implementations need to subclass TYPE_IOMMU_MEMORY_REGION
+ * and provide an implementation of at least the @translate method here
+ * to handle requests to the memory region. Other methods are optional.
+ *
+ * The IOMMU implementation must use the IOMMU notifier infrastructure
+ * to report whenever mappings are changed, by calling
+ * memory_region_notify_iommu() (or, if necessary, by calling
+ * memory_region_notify_one() for each registered notifier).
+ */
 typedef struct IOMMUMemoryRegionClass {
     /* private */
     struct DeviceClass parent_class;
 
     /*
-     * Return a TLB entry that contains a given address. Flag should
-     * be the access permission of this translation operation. We can
-     * set flag to IOMMU_NONE to mean that we don't need any
-     * read/write permission checks, like, when for region replay.
+     * Return a TLB entry that contains a given address.
+     *
+     * The IOMMUAccessFlags indicated via @flag are optional and may
+     * be specified as IOMMU_NONE to indicate that the caller needs
+     * the full translation information for both reads and writes. If
+     * the access flags are specified then the IOMMU implementation
+     * may use this as an optimization, to stop doing a page table
+     * walk as soon as it knows that the requested permissions are not
+     * allowed. If IOMMU_NONE is passed then the IOMMU must do the
+     * full page table walk and report the permissions in the returned
+     * IOMMUTLBEntry. (Note that this implies that an IOMMU may not
+     * return different mappings for reads and writes.)
+     *
+     * The returned information remains valid while the caller is
+     * holding the big QEMU lock or is inside an RCU critical section;
+     * if the caller wishes to cache the mapping beyond that it must
+     * register an IOMMU notifier so it can invalidate its cached
+     * information when the IOMMU mapping changes.
+     *
+     * @iommu: the IOMMUMemoryRegion
+     * @hwaddr: address to be translated within the memory region
+     * @flag: requested access permissions
      */
     IOMMUTLBEntry (*translate)(IOMMUMemoryRegion *iommu, hwaddr addr,
                                IOMMUAccessFlags flag);
-    /* Returns minimum supported page size */
+    /* Returns minimum supported page size in bytes.
+     * If this method is not provided then the minimum is assumed to
+     * be TARGET_PAGE_SIZE.
+     *
+     * @iommu: the IOMMUMemoryRegion
+     */
     uint64_t (*get_min_page_size)(IOMMUMemoryRegion *iommu);
-    /* Called when IOMMU Notifier flag changed */
+    /* Called when IOMMU Notifier flag changes (ie when the set of
+     * events which IOMMU users are requesting notification for changes).
+     * Optional method -- need not be provided if the IOMMU does not
+     * need to know exactly which events must be notified.
+     *
+     * @iommu: the IOMMUMemoryRegion
+     * @old_flags: events which previously needed to be notified
+     * @new_flags: events which now need to be notified
+     */
     void (*notify_flag_changed)(IOMMUMemoryRegion *iommu,
                                 IOMMUNotifierFlag old_flags,
                                 IOMMUNotifierFlag new_flags);
-    /* Set this up to provide customized IOMMU replay function */
+    /* Called to handle memory_region_iommu_replay().
+     *
+     * The default implementation of memory_region_iommu_replay() is to
+     * call the IOMMU translate method for every page in the address space
+     * with flag == IOMMU_NONE and then call the notifier if translate
+     * returns a valid mapping. If this method is implemented then it
+     * overrides the default behaviour, and must provide the full semantics
+     * of memory_region_iommu_replay(), by calling @notifier for every
+     * translation present in the IOMMU.
+     *
+     * Optional method -- an IOMMU only needs to provide this method
+     * if the default is inefficient or produces undesirable side effects.
+     *
+     * Note: this is not related to record-and-replay functionality.
+     */
     void (*replay)(IOMMUMemoryRegion *iommu, IOMMUNotifier *notifier);
 
-    /* Get IOMMU misc attributes */
-    int (*get_attr)(IOMMUMemoryRegion *iommu, enum IOMMUMemoryRegionAttr,
+    /* Get IOMMU misc attributes. This is an optional method that
+     * can be used to allow users of the IOMMU to get implementation-specific
+     * information. The IOMMU implements this method to handle calls
+     * by IOMMU users to memory_region_iommu_get_attr() by filling in
+     * the arbitrary data pointer for any IOMMUMemoryRegionAttr values that
+     * the IOMMU supports. If the method is unimplemented then
+     * memory_region_iommu_get_attr() will always return -EINVAL.
+     *
+     * @iommu: the IOMMUMemoryRegion
+     * @attr: attribute being queried
+     * @data: memory to fill in with the attribute data
+     *
+     * Returns 0 on success, or a negative errno; in particular
+     * returns -EINVAL for unrecognized or unimplemented attribute types.
+     */
+    int (*get_attr)(IOMMUMemoryRegion *iommu, enum IOMMUMemoryRegionAttr attr,
                     void *data);
 } IOMMUMemoryRegionClass;
 
@@ -XXX,XX +XXX,XX @@ static inline void memory_region_init_reservation(MemoryRegion *mr,
  * An IOMMU region translates addresses and forwards accesses to a target
  * memory region.
  *
+ * The IOMMU implementation must define a subclass of TYPE_IOMMU_MEMORY_REGION.
+ * @_iommu_mr should be a pointer to enough memory for an instance of
+ * that subclass, @instance_size is the size of that subclass, and
+ * @mrtypename is its name. This function will initialize @_iommu_mr as an
+ * instance of the subclass, and its methods will then be called to handle
+ * accesses to the memory region. See the documentation of
+ * #IOMMUMemoryRegionClass for further details.
+ *
  * @_iommu_mr: the #IOMMUMemoryRegion to be initialized
  * @instance_size: the IOMMUMemoryRegion subclass instance size
  * @mrtypename: the type name of the #IOMMUMemoryRegion
@@ -XXX,XX +XXX,XX @@ void memory_region_register_iommu_notifier(MemoryRegion *mr,
  * a notifier with the minimum page granularity returned by
  * mr->iommu_ops->get_page_size().
  *
+ * Note: this is not related to record-and-replay functionality.
+ *
  * @iommu_mr: the memory region to observe
  * @n: the notifier to which to replay iommu mappings
  */
@@ -XXX,XX +XXX,XX @@ void memory_region_iommu_replay(IOMMUMemoryRegion *iommu_mr, IOMMUNotifier *n);
  * memory_region_iommu_replay_all: replay existing IOMMU translations
  * to all the notifiers registered.
  *
+ * Note: this is not related to record-and-replay functionality.
+ *
  * @iommu_mr: the memory region to observe
  */
 void memory_region_iommu_replay_all(IOMMUMemoryRegion *iommu_mr);
@@ -XXX,XX +XXX,XX @@ void memory_region_unregister_iommu_notifier(MemoryRegion *mr,
  * memory_region_iommu_get_attr: return an IOMMU attr if get_attr() is
  * defined on the IOMMU.
  *
- * Returns 0 if succeded, error code otherwise.
+ * Returns 0 on success, or a negative errno otherwise. In particular,
+ * -EINVAL indicates that the IOMMU does not support the requested
+ * attribute.
  *
  * @iommu_mr: the memory region
  * @attr: the requested attribute
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to tb_invalidate_phys_addr().
Its callers either have an attrs value to hand, or don't care
and can use MEMTXATTRS_UNSPECIFIED.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Message-id: 20180521140402.23318-3-peter.maydell@linaro.org
---
 include/exec/exec-all.h   | 5 +++--
 accel/tcg/translate-all.c | 2 +-
 exec.c                    | 2 +-
 target/xtensa/op_helper.c | 3 ++-
 4 files changed, 7 insertions(+), 5 deletions(-)

diff --git a/include/exec/exec-all.h b/include/exec/exec-all.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/exec-all.h
+++ b/include/exec/exec-all.h
@@ -XXX,XX +XXX,XX @@ void tlb_set_page_with_attrs(CPUState *cpu, target_ulong vaddr,
 void tlb_set_page(CPUState *cpu, target_ulong vaddr,
                   hwaddr paddr, int prot,
                   int mmu_idx, target_ulong size);
-void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr);
+void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs);
 void probe_write(CPUArchState *env, target_ulong addr, int size, int mmu_idx,
                  uintptr_t retaddr);
 #else
@@ -XXX,XX +XXX,XX @@ static inline void tlb_flush_by_mmuidx_all_cpus_synced(CPUState *cpu,
                                                        uint16_t idxmap)
 {
 }
-static inline void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr)
+static inline void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr,
+                                           MemTxAttrs attrs)
 {
 }
 #endif
diff --git a/accel/tcg/translate-all.c b/accel/tcg/translate-all.c
index XXXXXXX..XXXXXXX 100644
--- a/accel/tcg/translate-all.c
+++ b/accel/tcg/translate-all.c
@@ -XXX,XX +XXX,XX @@ static TranslationBlock *tb_find_pc(uintptr_t tc_ptr)
 }
 
 #if !defined(CONFIG_USER_ONLY)
-void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr)
+void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs)
 {
     ram_addr_t ram_addr;
     MemoryRegion *mr;
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static void breakpoint_invalidate(CPUState *cpu, target_ulong pc)
     if (phys != -1) {
         /* Locks grabbed by tb_invalidate_phys_addr */
         tb_invalidate_phys_addr(cpu->cpu_ases[asidx].as,
-                                phys | (pc & ~TARGET_PAGE_MASK));
+                                phys | (pc & ~TARGET_PAGE_MASK), attrs);
     }
 }
 #endif
diff --git a/target/xtensa/op_helper.c b/target/xtensa/op_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/xtensa/op_helper.c
+++ b/target/xtensa/op_helper.c
@@ -XXX,XX +XXX,XX @@ static void tb_invalidate_virtual_addr(CPUXtensaState *env, uint32_t vaddr)
     int ret = xtensa_get_physical_addr(env, false, vaddr, 2, 0,
             &paddr, &page_size, &access);
     if (ret == 0) {
-        tb_invalidate_phys_addr(&address_space_memory, paddr);
+        tb_invalidate_phys_addr(&address_space_memory, paddr,
+                                MEMTXATTRS_UNSPECIFIED);
     }
 }
 
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to address_space_translate()
and address_space_translate_cached(). Callers either have an
attrs value to hand, or don't care and can use MEMTXATTRS_UNSPECIFIED.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521140402.23318-4-peter.maydell@linaro.org
---
 include/exec/memory.h     |  4 +++-
 accel/tcg/translate-all.c |  2 +-
 exec.c                    | 14 +++++++++-----
 hw/vfio/common.c          |  3 ++-
 memory_ldst.inc.c         | 18 +++++++++---------
 target/riscv/helper.c     |  2 +-
 6 files changed, 25 insertions(+), 18 deletions(-)

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
  * #MemoryRegion.
  * @len: pointer to length
  * @is_write: indicates the transfer direction
+ * @attrs: memory attributes
  */
 MemoryRegion *flatview_translate(FlatView *fv,
                                  hwaddr addr, hwaddr *xlat,
@@ -XXX,XX +XXX,XX @@ MemoryRegion *flatview_translate(FlatView *fv,
 
 static inline MemoryRegion *address_space_translate(AddressSpace *as,
                                                     hwaddr addr, hwaddr *xlat,
-                                                    hwaddr *len, bool is_write)
+                                                    hwaddr *len, bool is_write,
+                                                    MemTxAttrs attrs)
 {
     return flatview_translate(address_space_to_flatview(as),
                               addr, xlat, len, is_write);
diff --git a/accel/tcg/translate-all.c b/accel/tcg/translate-all.c
index XXXXXXX..XXXXXXX 100644
--- a/accel/tcg/translate-all.c
+++ b/accel/tcg/translate-all.c
@@ -XXX,XX +XXX,XX @@ void tb_invalidate_phys_addr(AddressSpace *as, hwaddr addr, MemTxAttrs attrs)
     hwaddr l = 1;
 
     rcu_read_lock();
-    mr = address_space_translate(as, addr, &addr, &l, false);
+    mr = address_space_translate(as, addr, &addr, &l, false, attrs);
     if (!(memory_region_is_ram(mr)
           || memory_region_is_romd(mr))) {
         rcu_read_unlock();
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static inline void cpu_physical_memory_write_rom_internal(AddressSpace *as,
     rcu_read_lock();
     while (len > 0) {
         l = len;
-        mr = address_space_translate(as, addr, &addr1, &l, true);
+        mr = address_space_translate(as, addr, &addr1, &l, true,
+                                     MEMTXATTRS_UNSPECIFIED);
 
         if (!(memory_region_is_ram(mr) ||
               memory_region_is_romd(mr))) {
@@ -XXX,XX +XXX,XX @@ void address_space_cache_destroy(MemoryRegionCache *cache)
  */
 static inline MemoryRegion *address_space_translate_cached(
     MemoryRegionCache *cache, hwaddr addr, hwaddr *xlat,
-    hwaddr *plen, bool is_write)
+    hwaddr *plen, bool is_write, MemTxAttrs attrs)
 {
     MemoryRegionSection section;
     MemoryRegion *mr;
@@ -XXX,XX +XXX,XX @@ address_space_read_cached_slow(MemoryRegionCache *cache, hwaddr addr,
     MemoryRegion *mr;
 
     l = len;
-    mr = address_space_translate_cached(cache, addr, &addr1, &l, false);
+    mr = address_space_translate_cached(cache, addr, &addr1, &l, false,
+                                        MEMTXATTRS_UNSPECIFIED);
     flatview_read_continue(cache->fv,
                            addr, MEMTXATTRS_UNSPECIFIED, buf, len,
                            addr1, l, mr);
@@ -XXX,XX +XXX,XX @@ address_space_write_cached_slow(MemoryRegionCache *cache, hwaddr addr,
     MemoryRegion *mr;
 
     l = len;
-    mr = address_space_translate_cached(cache, addr, &addr1, &l, true);
+    mr = address_space_translate_cached(cache, addr, &addr1, &l, true,
+                                        MEMTXATTRS_UNSPECIFIED);
     flatview_write_continue(cache->fv,
                             addr, MEMTXATTRS_UNSPECIFIED, buf, len,
                             addr1, l, mr);
@@ -XXX,XX +XXX,XX @@ bool cpu_physical_memory_is_io(hwaddr phys_addr)
 
     rcu_read_lock();
     mr = address_space_translate(&address_space_memory,
-                                 phys_addr, &phys_addr, &l, false);
+                                 phys_addr, &phys_addr, &l, false,
+                                 MEMTXATTRS_UNSPECIFIED);
 
     res = !(memory_region_is_ram(mr) || memory_region_is_romd(mr));
     rcu_read_unlock();
diff --git a/hw/vfio/common.c b/hw/vfio/common.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/vfio/common.c
+++ b/hw/vfio/common.c
@@ -XXX,XX +XXX,XX @@ static bool vfio_get_vaddr(IOMMUTLBEntry *iotlb, void **vaddr,
      */
     mr = address_space_translate(&address_space_memory,
                                  iotlb->translated_addr,
-                                 &xlat, &len, writable);
+                                 &xlat, &len, writable,
+                                 MEMTXATTRS_UNSPECIFIED);
     if (!memory_region_is_ram(mr)) {
         error_report("iommu map to non memory area %"HWADDR_PRIx"",
                      xlat);
diff --git a/memory_ldst.inc.c b/memory_ldst.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/memory_ldst.inc.c
+++ b/memory_ldst.inc.c
@@ -XXX,XX +XXX,XX @@ static inline uint32_t glue(address_space_ldl_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, false);
+    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
     if (l < 4 || !IS_DIRECT(mr, false)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ static inline uint64_t glue(address_space_ldq_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, false);
+    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
     if (l < 8 || !IS_DIRECT(mr, false)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ uint32_t glue(address_space_ldub, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, false);
+    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
     if (!IS_DIRECT(mr, false)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ static inline uint32_t glue(address_space_lduw_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, false);
+    mr = TRANSLATE(addr, &addr1, &l, false, attrs);
     if (l < 2 || !IS_DIRECT(mr, false)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ void glue(address_space_stl_notdirty, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, true);
+    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
     if (l < 4 || !IS_DIRECT(mr, true)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ static inline void glue(address_space_stl_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, true);
+    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
     if (l < 4 || !IS_DIRECT(mr, true)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ void glue(address_space_stb, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, true);
+    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
     if (!IS_DIRECT(mr, true)) {
         release_lock |= prepare_mmio_access(mr);
         r = memory_region_dispatch_write(mr, addr1, val, 1, attrs);
@@ -XXX,XX +XXX,XX @@ static inline void glue(address_space_stw_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, true);
+    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
     if (l < 2 || !IS_DIRECT(mr, true)) {
         release_lock |= prepare_mmio_access(mr);
 
@@ -XXX,XX +XXX,XX @@ static void glue(address_space_stq_internal, SUFFIX)(ARG1_DECL,
     bool release_lock = false;
 
     RCU_READ_LOCK();
-    mr = TRANSLATE(addr, &addr1, &l, true);
+    mr = TRANSLATE(addr, &addr1, &l, true, attrs);
     if (l < 8 || !IS_DIRECT(mr, true)) {
         release_lock |= prepare_mmio_access(mr);
 
diff --git a/target/riscv/helper.c b/target/riscv/helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/riscv/helper.c
+++ b/target/riscv/helper.c
@@ -XXX,XX +XXX,XX @@ restart:
                 MemoryRegion *mr;
                 hwaddr l = sizeof(target_ulong), addr1;
                 mr = address_space_translate(cs->as, pte_addr,
-                    &addr1, &l, false);
+                    &addr1, &l, false, MEMTXATTRS_UNSPECIFIED);
                 if (memory_access_is_direct(mr, true)) {
                     target_ulong *pte_pa =
                         qemu_map_ram_ptr(mr->ram_block, addr1);
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to address_space_map().
Its callers either have an attrs value to hand, or don't care
and can use MEMTXATTRS_UNSPECIFIED.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521140402.23318-5-peter.maydell@linaro.org
---
 include/exec/memory.h   | 3 ++-
 include/sysemu/dma.h    | 3 ++-
 exec.c                  | 6 ++++--
 target/ppc/mmu-hash64.c | 3 ++-
 4 files changed, 10 insertions(+), 5 deletions(-)

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len, bool is_
  * @addr: address within that address space
  * @plen: pointer to length of buffer; updated on return
  * @is_write: indicates the transfer direction
+ * @attrs: memory attributes
  */
 void *address_space_map(AddressSpace *as, hwaddr addr,
-                        hwaddr *plen, bool is_write);
+                        hwaddr *plen, bool is_write, MemTxAttrs attrs);
 
 /* address_space_unmap: Unmaps a memory region previously mapped by address_space_map()
  *
diff --git a/include/sysemu/dma.h b/include/sysemu/dma.h
index XXXXXXX..XXXXXXX 100644
--- a/include/sysemu/dma.h
+++ b/include/sysemu/dma.h
@@ -XXX,XX +XXX,XX @@ static inline void *dma_memory_map(AddressSpace *as,
     hwaddr xlen = *len;
     void *p;
 
-    p = address_space_map(as, addr, &xlen, dir == DMA_DIRECTION_FROM_DEVICE);
+    p = address_space_map(as, addr, &xlen, dir == DMA_DIRECTION_FROM_DEVICE,
+                          MEMTXATTRS_UNSPECIFIED);
     *len = xlen;
     return p;
 }
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ flatview_extend_translation(FlatView *fv, hwaddr addr,
 void *address_space_map(AddressSpace *as,
                         hwaddr addr,
                         hwaddr *plen,
-                        bool is_write)
+                        bool is_write,
+                        MemTxAttrs attrs)
 {
     hwaddr len = *plen;
     hwaddr l, xlat;
@@ -XXX,XX +XXX,XX @@ void *cpu_physical_memory_map(hwaddr addr,
                               hwaddr *plen,
                               int is_write)
 {
-    return address_space_map(&address_space_memory, addr, plen, is_write);
+    return address_space_map(&address_space_memory, addr, plen, is_write,
+                             MEMTXATTRS_UNSPECIFIED);
 }
 
 void cpu_physical_memory_unmap(void *buffer, hwaddr len,
diff --git a/target/ppc/mmu-hash64.c b/target/ppc/mmu-hash64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/ppc/mmu-hash64.c
+++ b/target/ppc/mmu-hash64.c
@@ -XXX,XX +XXX,XX @@ const ppc_hash_pte64_t *ppc_hash64_map_hptes(PowerPCCPU *cpu,
         return NULL;
     }
 
-    hptes = address_space_map(CPU(cpu)->as, base + pte_offset, &plen, false);
+    hptes = address_space_map(CPU(cpu)->as, base + pte_offset, &plen, false,
+                              MEMTXATTRS_UNSPECIFIED);
     if (plen < (n * HASH_PTE_SIZE_64)) {
         hw_error("%s: Unable to map all requested HPTEs\n", __func__);
     }
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to address_space_access_valid().
Its callers either have an attrs value to hand, or don't care
and can use MEMTXATTRS_UNSPECIFIED.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521140402.23318-6-peter.maydell@linaro.org
---
 include/exec/memory.h      | 4 +++-
 include/sysemu/dma.h       | 3 ++-
 exec.c                     | 3 ++-
 target/s390x/diag.c        | 6 ++++--
 target/s390x/excp_helper.c | 3 ++-
 target/s390x/mmu_helper.c  | 3 ++-
 target/s390x/sigp.c        | 3 ++-
 7 files changed, 17 insertions(+), 8 deletions(-)

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ static inline MemoryRegion *address_space_translate(AddressSpace *as,
  * @addr: address within that address space
  * @len: length of the area to be checked
  * @is_write: indicates the transfer direction
+ * @attrs: memory attributes
  */
-bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len, bool is_write);
+bool address_space_access_valid(AddressSpace *as, hwaddr addr, int len,
+                                bool is_write, MemTxAttrs attrs);
 
 /* address_space_map: map a physical memory region into a host virtual address
  *
diff --git a/include/sysemu/dma.h b/include/sysemu/dma.h
index XXXXXXX..XXXXXXX 100644
--- a/include/sysemu/dma.h
+++ b/include/sysemu/dma.h
@@ -XXX,XX +XXX,XX @@ static inline bool dma_memory_valid(AddressSpace *as,
                                     DMADirection dir)
 {
     return address_space_access_valid(as, addr, len,
-                                      dir == DMA_DIRECTION_FROM_DEVICE);
+                                      dir == DMA_DIRECTION_FROM_DEVICE,
+                                      MEMTXATTRS_UNSPECIFIED);
 }
 
 static inline int dma_memory_rw_relaxed(AddressSpace *as, dma_addr_t addr,
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
 }
 
 bool address_space_access_valid(AddressSpace *as, hwaddr addr,
-                                int len, bool is_write)
+                                int len, bool is_write,
+                                MemTxAttrs attrs)
 {
     FlatView *fv;
     bool result;
diff --git a/target/s390x/diag.c b/target/s390x/diag.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/diag.c
+++ b/target/s390x/diag.c
@@ -XXX,XX +XXX,XX @@ void handle_diag_308(CPUS390XState *env, uint64_t r1, uint64_t r3, uintptr_t ra)
             return;
         }
         if (!address_space_access_valid(&address_space_memory, addr,
-                                        sizeof(IplParameterBlock), false)) {
+                                        sizeof(IplParameterBlock), false,
+                                        MEMTXATTRS_UNSPECIFIED)) {
             s390_program_interrupt(env, PGM_ADDRESSING, ILEN_AUTO, ra);
             return;
         }
@@ -XXX,XX +XXX,XX @@ out:
             return;
         }
         if (!address_space_access_valid(&address_space_memory, addr,
-                                        sizeof(IplParameterBlock), true)) {
+                                        sizeof(IplParameterBlock), true,
+                                        MEMTXATTRS_UNSPECIFIED)) {
             s390_program_interrupt(env, PGM_ADDRESSING, ILEN_AUTO, ra);
             return;
         }
diff --git a/target/s390x/excp_helper.c b/target/s390x/excp_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/excp_helper.c
+++ b/target/s390x/excp_helper.c
@@ -XXX,XX +XXX,XX @@ int s390_cpu_handle_mmu_fault(CPUState *cs, vaddr orig_vaddr, int size,
 
     /* check out of RAM access */
     if (!address_space_access_valid(&address_space_memory, raddr,
-                                    TARGET_PAGE_SIZE, rw)) {
+                                    TARGET_PAGE_SIZE, rw,
+                                    MEMTXATTRS_UNSPECIFIED)) {
         DPRINTF("%s: raddr %" PRIx64 " > ram_size %" PRIx64 "\n", __func__,
                 (uint64_t)raddr, (uint64_t)ram_size);
         trigger_pgm_exception(env, PGM_ADDRESSING, ILEN_AUTO);
diff --git a/target/s390x/mmu_helper.c b/target/s390x/mmu_helper.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/mmu_helper.c
+++ b/target/s390x/mmu_helper.c
@@ -XXX,XX +XXX,XX @@ static int translate_pages(S390CPU *cpu, vaddr addr, int nr_pages,
             return ret;
         }
         if (!address_space_access_valid(&address_space_memory, pages[i],
-                                        TARGET_PAGE_SIZE, is_write)) {
+                                        TARGET_PAGE_SIZE, is_write,
+                                        MEMTXATTRS_UNSPECIFIED)) {
             trigger_access_exception(env, PGM_ADDRESSING, ILEN_AUTO, 0);
             return -EFAULT;
         }
diff --git a/target/s390x/sigp.c b/target/s390x/sigp.c
index XXXXXXX..XXXXXXX 100644
--- a/target/s390x/sigp.c
+++ b/target/s390x/sigp.c
@@ -XXX,XX +XXX,XX @@ static void sigp_set_prefix(CPUState *cs, run_on_cpu_data arg)
     cpu_synchronize_state(cs);
 
     if (!address_space_access_valid(&address_space_memory, addr,
-                                    sizeof(struct LowCore), false)) {
+                                    sizeof(struct LowCore), false,
+                                    MEMTXATTRS_UNSPECIFIED)) {
         set_sigp_status(si, SIGP_STAT_INVALID_PARAMETER);
         return;
     }
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to flatview_extend_translation().
Its callers either have an attrs value to hand, or don't care
and can use MEMTXATTRS_UNSPECIFIED.

diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr,
 
 static hwaddr
 flatview_extend_translation(FlatView *fv, hwaddr addr,
-                                 hwaddr target_len,
-                                 MemoryRegion *mr, hwaddr base, hwaddr len,
-                                 bool is_write)
+                            hwaddr target_len,
+                            MemoryRegion *mr, hwaddr base, hwaddr len,
+                            bool is_write, MemTxAttrs attrs)
 {
     hwaddr done = 0;
     hwaddr xlat;
@@ -XXX,XX +XXX,XX @@ void *address_space_map(AddressSpace *as,
 
     memory_region_ref(mr);
     *plen = flatview_extend_translation(fv, addr, len, mr, xlat,
-                                             l, is_write);
+                                        l, is_write, attrs);
     ptr = qemu_ram_ptr_length(mr->ram_block, xlat, plen, true);
     rcu_read_unlock();
 
@@ -XXX,XX +XXX,XX @@ int64_t address_space_cache_init(MemoryRegionCache *cache,
     mr = cache->mrs.mr;
     memory_region_ref(mr);
     if (memory_access_is_direct(mr, is_write)) {
+        /* We don't care about the memory attributes here as we're only
+         * doing this if we found actual RAM, which behaves the same
+         * regardless of attributes; so UNSPECIFIED is fine.
+         */
         l = flatview_extend_translation(cache->fv, addr, len, mr,
-                                        cache->xlat, l, is_write);
+                                        cache->xlat, l, is_write,
+                                        MEMTXATTRS_UNSPECIFIED);
         cache->ptr = qemu_ram_ptr_length(mr->ram_block, cache->xlat, &l, true);
     } else {
         cache->ptr = NULL;
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to memory_region_access_valid().
Its callers either have an attrs value to hand, or don't care
and can use MEMTXATTRS_UNSPECIFIED.

The callsite in flatview_access_valid() is part of a recursive
loop flatview_access_valid() -> memory_region_access_valid() ->
 subpage_accepts() -> flatview_access_valid(); we make it pass
MEMTXATTRS_UNSPECIFIED for now, until the next several commits
have plumbed an attrs parameter through the rest of the loop
and we can add an attrs parameter to flatview_access_valid().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521140402.23318-8-peter.maydell@linaro.org
---
 include/exec/memory-internal.h | 3 ++-
 exec.c                         | 4 +++-
 hw/s390x/s390-pci-inst.c       | 3 ++-
 memory.c                       | 7 ++++---
 4 files changed, 11 insertions(+), 6 deletions(-)

diff --git a/include/exec/memory-internal.h b/include/exec/memory-internal.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory-internal.h
+++ b/include/exec/memory-internal.h
@@ -XXX,XX +XXX,XX @@ void flatview_unref(FlatView *view);
 extern const MemoryRegionOps unassigned_mem_ops;
 
 bool memory_region_access_valid(MemoryRegion *mr, hwaddr addr,
-                                unsigned size, bool is_write);
+                                unsigned size, bool is_write,
+                                MemTxAttrs attrs);
 
 void flatview_add_to_dispatch(FlatView *fv, MemoryRegionSection *section);
 AddressSpaceDispatch *address_space_dispatch_new(FlatView *fv);
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
         mr = flatview_translate(fv, addr, &xlat, &l, is_write);
         if (!memory_access_is_direct(mr, is_write)) {
             l = memory_access_size(mr, l, addr);
-            if (!memory_region_access_valid(mr, xlat, l, is_write)) {
+            /* When our callers all have attrs we'll pass them through here */
+            if (!memory_region_access_valid(mr, xlat, l, is_write,
+                                            MEMTXATTRS_UNSPECIFIED)) {
                 return false;
             }
         }
diff --git a/hw/s390x/s390-pci-inst.c b/hw/s390x/s390-pci-inst.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/s390x/s390-pci-inst.c
+++ b/hw/s390x/s390-pci-inst.c
@@ -XXX,XX +XXX,XX @@ int pcistb_service_call(S390CPU *cpu, uint8_t r1, uint8_t r3, uint64_t gaddr,
     mr = s390_get_subregion(mr, offset, len);
     offset -= mr->addr;
 
-    if (!memory_region_access_valid(mr, offset, len, true)) {
+    if (!memory_region_access_valid(mr, offset, len, true,
+                                    MEMTXATTRS_UNSPECIFIED)) {
         s390_program_interrupt(env, PGM_OPERAND, 6, ra);
         return 0;
     }
diff --git a/memory.c b/memory.c
index XXXXXXX..XXXXXXX 100644
--- a/memory.c
+++ b/memory.c
@@ -XXX,XX +XXX,XX @@ static const MemoryRegionOps ram_device_mem_ops = {
 bool memory_region_access_valid(MemoryRegion *mr,
                                 hwaddr addr,
                                 unsigned size,
-                                bool is_write)
+                                bool is_write,
+                                MemTxAttrs attrs)
 {
     int access_size_min, access_size_max;
     int access_size, i;
@@ -XXX,XX +XXX,XX @@ MemTxResult memory_region_dispatch_read(MemoryRegion *mr,
 {
     MemTxResult r;
 
-    if (!memory_region_access_valid(mr, addr, size, false)) {
+    if (!memory_region_access_valid(mr, addr, size, false, attrs)) {
         *pval = unassigned_mem_read(mr, addr, size);
         return MEMTX_DECODE_ERROR;
     }
@@ -XXX,XX +XXX,XX @@ MemTxResult memory_region_dispatch_write(MemoryRegion *mr,
                                          unsigned size,
                                          MemTxAttrs attrs)
 {
-    if (!memory_region_access_valid(mr, addr, size, true)) {
+    if (!memory_region_access_valid(mr, addr, size, true, attrs)) {
         unassigned_mem_write(mr, addr, data, size);
         return MEMTX_DECODE_ERROR;
     }
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to the MemoryRegion valid.accepts
callback. We'll need this for subpage_accepts().

We could take the approach we used with the read and write
callbacks and add new a new _with_attrs version, but since there
are so few implementations of the accepts hook we just change
them all.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20180521140402.23318-9-peter.maydell@linaro.org
---
 include/exec/memory.h |  3 ++-
 exec.c                |  9 ++++++---
 hw/hppa/dino.c        |  3 ++-
 hw/nvram/fw_cfg.c     | 12 ++++++++----
 hw/scsi/esp.c         |  3 ++-
 hw/xen/xen_pt_msi.c   |  3 ++-
 memory.c              |  5 +++--
 7 files changed, 25 insertions(+), 13 deletions(-)

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ struct MemoryRegionOps {
          * as a machine check exception).
          */
         bool (*accepts)(void *opaque, hwaddr addr,
-                        unsigned size, bool is_write);
+                        unsigned size, bool is_write,
+                        MemTxAttrs attrs);
     } valid;
     /* Internal implementation constraints: */
     struct {
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static void notdirty_mem_write(void *opaque, hwaddr ram_addr,
 }
 
 static bool notdirty_mem_accepts(void *opaque, hwaddr addr,
-                                 unsigned size, bool is_write)
+                                 unsigned size, bool is_write,
+                                 MemTxAttrs attrs)
 {
     return is_write;
 }
@@ -XXX,XX +XXX,XX @@ static MemTxResult subpage_write(void *opaque, hwaddr addr,
 }
 
 static bool subpage_accepts(void *opaque, hwaddr addr,
-                            unsigned len, bool is_write)
+                            unsigned len, bool is_write,
+                            MemTxAttrs attrs)
 {
     subpage_t *subpage = opaque;
 #if defined(DEBUG_SUBPAGE)
@@ -XXX,XX +XXX,XX @@ static void readonly_mem_write(void *opaque, hwaddr addr,
 }
 
 static bool readonly_mem_accepts(void *opaque, hwaddr addr,
-                                 unsigned size, bool is_write)
+                                 unsigned size, bool is_write,
+                                 MemTxAttrs attrs)
 {
     return is_write;
 }
diff --git a/hw/hppa/dino.c b/hw/hppa/dino.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/hppa/dino.c
+++ b/hw/hppa/dino.c
@@ -XXX,XX +XXX,XX @@ static void gsc_to_pci_forwarding(DinoState *s)
 }
 
 static bool dino_chip_mem_valid(void *opaque, hwaddr addr,
-                                unsigned size, bool is_write)
+                                unsigned size, bool is_write,
+                                MemTxAttrs attrs)
 {
     switch (addr) {
     case DINO_IAR0:
diff --git a/hw/nvram/fw_cfg.c b/hw/nvram/fw_cfg.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/nvram/fw_cfg.c
+++ b/hw/nvram/fw_cfg.c
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_dma_mem_write(void *opaque, hwaddr addr,
 }
 
 static bool fw_cfg_dma_mem_valid(void *opaque, hwaddr addr,
-                                  unsigned size, bool is_write)
+                                 unsigned size, bool is_write,
+                                 MemTxAttrs attrs)
 {
     return !is_write || ((size == 4 && (addr == 0 || addr == 4)) ||
                          (size == 8 && addr == 0));
 }
 
 static bool fw_cfg_data_mem_valid(void *opaque, hwaddr addr,
-                                  unsigned size, bool is_write)
+                                  unsigned size, bool is_write,
+                                  MemTxAttrs attrs)
 {
     return addr == 0;
 }
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_ctl_mem_write(void *opaque, hwaddr addr,
 }
 
 static bool fw_cfg_ctl_mem_valid(void *opaque, hwaddr addr,
-                                 unsigned size, bool is_write)
+                                 unsigned size, bool is_write,
+                                 MemTxAttrs attrs)
 {
     return is_write && size == 2;
 }
@@ -XXX,XX +XXX,XX @@ static void fw_cfg_comb_write(void *opaque, hwaddr addr,
 }
 
 static bool fw_cfg_comb_valid(void *opaque, hwaddr addr,
-                                  unsigned size, bool is_write)
+                              unsigned size, bool is_write,
+                              MemTxAttrs attrs)
 {
     return (size == 1) || (is_write && size == 2);
 }
diff --git a/hw/scsi/esp.c b/hw/scsi/esp.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/scsi/esp.c
+++ b/hw/scsi/esp.c
@@ -XXX,XX +XXX,XX @@ void esp_reg_write(ESPState *s, uint32_t saddr, uint64_t val)
 }
 
 static bool esp_mem_accepts(void *opaque, hwaddr addr,
-                            unsigned size, bool is_write)
+                            unsigned size, bool is_write,
+                            MemTxAttrs attrs)
 {
     return (size == 1) || (is_write && size == 4);
 }
diff --git a/hw/xen/xen_pt_msi.c b/hw/xen/xen_pt_msi.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/xen/xen_pt_msi.c
+++ b/hw/xen/xen_pt_msi.c
@@ -XXX,XX +XXX,XX @@ static uint64_t pci_msix_read(void *opaque, hwaddr addr,
 }
 
 static bool pci_msix_accepts(void *opaque, hwaddr addr,
-                             unsigned size, bool is_write)
+                             unsigned size, bool is_write,
+                             MemTxAttrs attrs)
 {
     return !(addr & (size - 1));
 }
diff --git a/memory.c b/memory.c
index XXXXXXX..XXXXXXX 100644
--- a/memory.c
+++ b/memory.c
@@ -XXX,XX +XXX,XX @@ static void unassigned_mem_write(void *opaque, hwaddr addr,
 }
 
 static bool unassigned_mem_accepts(void *opaque, hwaddr addr,
-                                   unsigned size, bool is_write)
+                                   unsigned size, bool is_write,
+                                   MemTxAttrs attrs)
 {
     return false;
 }
@@ -XXX,XX +XXX,XX @@ bool memory_region_access_valid(MemoryRegion *mr,
     access_size = MAX(MIN(size, access_size_max), access_size_min);
     for (i = 0; i < size; i += access_size) {
         if (!mr->ops->valid.accepts(mr->opaque, addr + i, access_size,
-                                    is_write)) {
+                                    is_write, attrs)) {
             return false;
         }
     }
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to flatview_access_valid().
Its callers now all have an attrs value to hand, so we can
correct our earlier temporary use of MEMTXATTRS_UNSPECIFIED.

diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static MemTxResult flatview_read(FlatView *fv, hwaddr addr,
 static MemTxResult flatview_write(FlatView *fv, hwaddr addr, MemTxAttrs attrs,
                                   const uint8_t *buf, int len);
 static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
-                                  bool is_write);
+                                  bool is_write, MemTxAttrs attrs);
 
 static MemTxResult subpage_read(void *opaque, hwaddr addr, uint64_t *data,
                                 unsigned len, MemTxAttrs attrs)
@@ -XXX,XX +XXX,XX @@ static bool subpage_accepts(void *opaque, hwaddr addr,
 #endif
 
     return flatview_access_valid(subpage->fv, addr + subpage->base,
-                                 len, is_write);
+                                 len, is_write, attrs);
 }
 
 static const MemoryRegionOps subpage_ops = {
@@ -XXX,XX +XXX,XX @@ static void cpu_notify_map_clients(void)
 }
 
 static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
-                                  bool is_write)
+                                  bool is_write, MemTxAttrs attrs)
 {
     MemoryRegion *mr;
     hwaddr l, xlat;
@@ -XXX,XX +XXX,XX @@ static bool flatview_access_valid(FlatView *fv, hwaddr addr, int len,
         mr = flatview_translate(fv, addr, &xlat, &l, is_write);
         if (!memory_access_is_direct(mr, is_write)) {
             l = memory_access_size(mr, l, addr);
-            /* When our callers all have attrs we'll pass them through here */
-            if (!memory_region_access_valid(mr, xlat, l, is_write,
-                                            MEMTXATTRS_UNSPECIFIED)) {
+            if (!memory_region_access_valid(mr, xlat, l, is_write, attrs)) {
                 return false;
             }
         }
@@ -XXX,XX +XXX,XX @@ bool address_space_access_valid(AddressSpace *as, hwaddr addr,
 
     rcu_read_lock();
     fv = address_space_to_flatview(as);
-    result = flatview_access_valid(fv, addr, len, is_write);
+    result = flatview_access_valid(fv, addr, len, is_write, attrs);
     rcu_read_unlock();
     return result;
 }
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to flatview_translate(); all its
callers now have attrs available.

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to address_space_get_iotlb_entry().

diff --git a/include/exec/memory.h b/include/exec/memory.h
index XXXXXXX..XXXXXXX 100644
--- a/include/exec/memory.h
+++ b/include/exec/memory.h
@@ -XXX,XX +XXX,XX @@ void address_space_cache_destroy(MemoryRegionCache *cache);
  * entry. Should be called from an RCU critical section.
  */
 IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
-                                            bool is_write);
+                                            bool is_write, MemTxAttrs attrs);
 
 /* address_space_translate: translate an address range into an address space
  * into a MemoryRegion and an address range into that section.  Should be
diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
 
 /* Called from RCU critical section */
 IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
-                                            bool is_write)
+                                            bool is_write, MemTxAttrs attrs)
 {
     MemoryRegionSection section;
     hwaddr xlat, page_mask;
diff --git a/hw/virtio/vhost.c b/hw/virtio/vhost.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/virtio/vhost.c
+++ b/hw/virtio/vhost.c
@@ -XXX,XX +XXX,XX @@ int vhost_device_iotlb_miss(struct vhost_dev *dev, uint64_t iova, int write)
     trace_vhost_iotlb_miss(dev, 1);
 
     iotlb = address_space_get_iotlb_entry(dev->vdev->dma_as,
-                                          iova, write);
+                                          iova, write,
+                                          MEMTXATTRS_UNSPECIFIED);
     if (iotlb.target_as != NULL) {
         ret = vhost_memory_region_lookup(dev, iotlb.translated_addr,
                                          &uaddr, &len);
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to flatview_do_translate().

diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ unassigned:
  * @is_write: whether the translation operation is for write
  * @is_mmio: whether this can be MMIO, set true if it can
  * @target_as: the address space targeted by the IOMMU
+ * @attrs: memory transaction attributes
  *
  * This function is called from RCU critical section
  */
@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
                                                  hwaddr *page_mask_out,
                                                  bool is_write,
                                                  bool is_mmio,
-                                                 AddressSpace **target_as)
+                                                 AddressSpace **target_as,
+                                                 MemTxAttrs attrs)
 {
     MemoryRegionSection *section;
     IOMMUMemoryRegion *iommu_mr;
@@ -XXX,XX +XXX,XX @@ IOMMUTLBEntry address_space_get_iotlb_entry(AddressSpace *as, hwaddr addr,
      * but page mask.
      */
     section = flatview_do_translate(address_space_to_flatview(as), addr, &xlat,
-                                    NULL, &page_mask, is_write, false, &as);
+                                    NULL, &page_mask, is_write, false, &as,
+                                    attrs);
 
     /* Illegal translation */
     if (section.mr == &io_mem_unassigned) {
@@ -XXX,XX +XXX,XX @@ MemoryRegion *flatview_translate(FlatView *fv, hwaddr addr, hwaddr *xlat,
 
     /* This can be MMIO, so setup MMIO bit. */
     section = flatview_do_translate(fv, addr, xlat, plen, NULL,
-                                    is_write, true, &as);
+                                    is_write, true, &as, attrs);
     mr = section.mr;
 
     if (xen_enabled() && memory_access_is_direct(mr, is_write)) {
-- 
2.17.1

As part of plumbing MemTxAttrs down to the IOMMU translate method,
add MemTxAttrs as an argument to address_space_translate_iommu().

diff --git a/exec.c b/exec.c
index XXXXXXX..XXXXXXX 100644
--- a/exec.c
+++ b/exec.c
@@ -XXX,XX +XXX,XX @@ address_space_translate_internal(AddressSpaceDispatch *d, hwaddr addr, hwaddr *x
  * @is_write: whether the translation operation is for write
  * @is_mmio: whether this can be MMIO, set true if it can
  * @target_as: the address space targeted by the IOMMU
+ * @attrs: transaction attributes
  *
  * This function is called from RCU critical section.  It is the common
  * part of flatview_do_translate and address_space_translate_cached.
@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection address_space_translate_iommu(IOMMUMemoryRegion *iomm
                                                          hwaddr *page_mask_out,
                                                          bool is_write,
                                                          bool is_mmio,
-                                                         AddressSpace **target_as)
+                                                         AddressSpace **target_as,
+                                                         MemTxAttrs attrs)
 {
     MemoryRegionSection *section;
     hwaddr page_mask = (hwaddr)-1;
@@ -XXX,XX +XXX,XX @@ static MemoryRegionSection flatview_do_translate(FlatView *fv,
         return address_space_translate_iommu(iommu_mr, xlat,
                                              plen_out, page_mask_out,
                                              is_write, is_mmio,
-                                             target_as);
+                                             target_as, attrs);
     }
     if (page_mask_out) {
         /* Not behind an IOMMU, use default page size. */
@@ -XXX,XX +XXX,XX @@ static inline MemoryRegion *address_space_translate_cached(
 
     section = address_space_translate_iommu(iommu_mr, xlat, plen,
                                             NULL, is_write, true,
-                                            &target_as);
+                                            &target_as, attrs);
     return section.mr;
 }
 
-- 
2.17.1

From: Shannon Zhao <zhaoshenglong@huawei.com>

acpi_data_push uses g_array_set_size to resize the memory size. If there
is no enough contiguous memory, the address will be changed. So previous
pointer could not be used any more. It must update the pointer and use
the new one.

Also, previous codes wrongly use le32 conversion of iort->node_offset
for subsequent computations that will result incorrect value if host is
not litlle endian. So use the non-converted one instead.

Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Message-id: 1527663951-14552-1-git-send-email-zhaoshenglong@huawei.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/arm/virt-acpi-build.c | 20 +++++++++++++++-----
 1 file changed, 15 insertions(+), 5 deletions(-)

diff --git a/hw/arm/virt-acpi-build.c b/hw/arm/virt-acpi-build.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/arm/virt-acpi-build.c
+++ b/hw/arm/virt-acpi-build.c
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
     AcpiIortItsGroup *its;
     AcpiIortTable *iort;
     AcpiIortSmmu3 *smmu;
-    size_t node_size, iort_length, smmu_offset = 0;
+    size_t node_size, iort_node_offset, iort_length, smmu_offset = 0;
     AcpiIortRC *rc;
 
     iort = acpi_data_push(table_data, sizeof(*iort));
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
 
     iort_length = sizeof(*iort);
     iort->node_count = cpu_to_le32(nb_nodes);
-    iort->node_offset = cpu_to_le32(sizeof(*iort));
+    /*
+     * Use a copy in case table_data->data moves during acpi_data_push
+     * operations.
+     */
+    iort_node_offset = sizeof(*iort);
+    iort->node_offset = cpu_to_le32(iort_node_offset);
 
     /* ITS group node */
     node_size =  sizeof(*its) + sizeof(uint32_t);
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
         int irq =  vms->irqmap[VIRT_SMMU];
 
         /* SMMUv3 node */
-        smmu_offset = iort->node_offset + node_size;
+        smmu_offset = iort_node_offset + node_size;
         node_size = sizeof(*smmu) + sizeof(*idmap);
         iort_length += node_size;
         smmu = acpi_data_push(table_data, node_size);
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
         idmap->id_count = cpu_to_le32(0xFFFF);
         idmap->output_base = 0;
         /* output IORT node is the ITS group node (the first node) */
-        idmap->output_reference = cpu_to_le32(iort->node_offset);
+        idmap->output_reference = cpu_to_le32(iort_node_offset);
     }
 
     /* Root Complex Node */
@@ -XXX,XX +XXX,XX @@ build_iort(GArray *table_data, BIOSLinker *linker, VirtMachineState *vms)
         idmap->output_reference = cpu_to_le32(smmu_offset);
     } else {
         /* output IORT node is the ITS group node (the first node) */
-        idmap->output_reference = cpu_to_le32(iort->node_offset);
+        idmap->output_reference = cpu_to_le32(iort_node_offset);
     }
 
+    /*
+     * Update the pointer address in case table_data->data moves during above
+     * acpi_data_push operations.
+     */
+    iort = (AcpiIortTable *)(table_data->data + iort_start);
     iort->length = cpu_to_le32(iort_length);
 
     build_header(linker, table_data, (void *)(table_data->data + iort_start),
-- 
2.17.1

From: Shannon Zhao <zhaoshenglong@huawei.com>

kvm_irqchip_create called by kvm_init will call kvm_init_irq_routing to
initialize global capability variables. If we call kvm_init_irq_routing in
GIC realize function, previous allocated memory will leak.

Fix this by deleting the unnecessary call.

Signed-off-by: Shannon Zhao <zhaoshenglong@huawei.com>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Message-id: 1527750994-14360-1-git-send-email-zhaoshenglong@huawei.com
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 hw/intc/arm_gic_kvm.c   | 1 -
 hw/intc/arm_gicv3_kvm.c | 1 -
 2 files changed, 2 deletions(-)

diff --git a/hw/intc/arm_gic_kvm.c b/hw/intc/arm_gic_kvm.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/intc/arm_gic_kvm.c
+++ b/hw/intc/arm_gic_kvm.c
@@ -XXX,XX +XXX,XX @@ static void kvm_arm_gic_realize(DeviceState *dev, Error **errp)
 
     if (kvm_has_gsi_routing()) {
         /* set up irq routing */
-        kvm_init_irq_routing(kvm_state);
         for (i = 0; i < s->num_irq - GIC_INTERNAL; ++i) {
             kvm_irqchip_add_irq_route(kvm_state, i, 0, i);
         }
diff --git a/hw/intc/arm_gicv3_kvm.c b/hw/intc/arm_gicv3_kvm.c
index XXXXXXX..XXXXXXX 100644
--- a/hw/intc/arm_gicv3_kvm.c
+++ b/hw/intc/arm_gicv3_kvm.c
@@ -XXX,XX +XXX,XX @@ static void kvm_arm_gicv3_realize(DeviceState *dev, Error **errp)
 
     if (kvm_has_gsi_routing()) {
         /* set up irq routing */
-        kvm_init_irq_routing(kvm_state);
         for (i = 0; i < s->num_irq - GIC_INTERNAL; ++i) {
             kvm_irqchip_add_irq_route(kvm_state, i, 0, i);
         }
-- 
2.17.1

The following changes since commit 61fee7f45955cd0bf9b79be9fa9c7ebabb5e6a85:

Merge remote-tracking branch 'remotes/philmd-gitlab/tags/acceptance-testing-20200622' into staging (2020-06-22 20:50:10 +0100)

are available in the Git repository at:

https://git.linaro.org/people/pmaydell/qemu-arm.git tags/pull-target-arm-20200623

for you to fetch changes up to 539533b85fbd269f777bed931de8ccae1dd837e9:

arm/virt: Add memory hot remove support (2020-06-23 11:39:48 +0100)

----------------------------------------------------------------
target-arm queue:
 * util/oslib-posix : qemu_init_exec_dir implementation for Mac
 * target/arm: Last parts of neon decodetree conversion
 * hw/arm/virt: Add 5.0 HW compat props
 * hw/watchdog/cmsdk-apb-watchdog: Add trace event for lock status
 * mps2: Add CMSDK APB watchdog, FPGAIO block, S2I devices and I2C devices
 * mps2: Add some unimplemented-device stubs for audio and GPIO
 * mps2-tz: Use the ARM SBCon two-wire serial bus interface
 * target/arm: Check supported KVM features globally (not per vCPU)
 * tests/qtest/arm-cpu-features: Add feature setting tests
 * arm/virt: Add memory hot remove support

----------------------------------------------------------------
Andrew Jones (2):
      hw/arm/virt: Add 5.0 HW compat props
      tests/qtest/arm-cpu-features: Add feature setting tests

David CARLIER (1):
      util/oslib-posix : qemu_init_exec_dir implementation for Mac

Peter Maydell (23):
      target/arm: Convert Neon 2-reg-misc VREV64 to decodetree
      target/arm: Convert Neon 2-reg-misc pairwise ops to decodetree
      target/arm: Convert VZIP, VUZP to decodetree
      target/arm: Convert Neon narrowing moves to decodetree
      target/arm: Convert Neon 2-reg-misc VSHLL to decodetree
      target/arm: Convert Neon VCVT f16/f32 insns to decodetree
      target/arm: Convert vectorised 2-reg-misc Neon ops to decodetree
      target/arm: Convert Neon 2-reg-misc crypto operations to decodetree
      target/arm: Rename NeonGenOneOpFn to NeonGenOne64OpFn
      target/arm: Fix capitalization in NeonGenTwo{Single, Double}OPFn typedefs
      target/arm: Make gen_swap_half() take separate src and dest
      target/arm: Convert Neon 2-reg-misc VREV32 and VREV16 to decodetree
      target/arm: Convert remaining simple 2-reg-misc Neon ops
      target/arm: Convert Neon VQABS, VQNEG to decodetree
      target/arm: Convert simple fp Neon 2-reg-misc insns
      target/arm: Convert Neon 2-reg-misc fp-compare-with-zero insns to decodetree
      target/arm: Convert Neon 2-reg-misc VRINT insns to decodetree
      target/arm: Convert Neon 2-reg-misc VCVT insns to decodetree
      target/arm: Convert Neon VSWP to decodetree
      target/arm: Convert Neon VTRN to decodetree
      target/arm: Move some functions used only in translate-neon.inc.c to that file
      target/arm: Remove unnecessary gen_io_end() calls
      target/arm: Remove dead code relating to SABA and UABA

Philippe Mathieu-Daudé (15):
      hw/watchdog/cmsdk-apb-watchdog: Add trace event for lock status
      hw/i2c/versatile_i2c: Add definitions for register addresses
      hw/i2c/versatile_i2c: Add SCL/SDA definitions
      hw/i2c: Add header for ARM SBCon two-wire serial bus interface
      hw/arm: Use TYPE_VERSATILE_I2C instead of hardcoded string
      hw/arm/mps2: Document CMSDK/FPGA APB subsystem sections
      hw/arm/mps2: Rename CMSDK AHB peripheral region
      hw/arm/mps2: Add CMSDK APB watchdog device
      hw/arm/mps2: Add CMSDK AHB GPIO peripherals as unimplemented devices
      hw/arm/mps2: Map the FPGA I/O block
      hw/arm/mps2: Add SPI devices
      hw/arm/mps2: Add I2C devices
      hw/arm/mps2: Add audio I2S interface as unimplemented device
      hw/arm/mps2-tz: Use the ARM SBCon two-wire serial bus interface
      target/arm: Check supported KVM features globally (not per vCPU)

Shameer Kolothum (1):
      arm/virt: Add memory hot remove support

From: David CARLIER <devnexen@gmail.com>

From 3025a0ce3fdf7d3559fc35a52c659f635f5c750c Mon Sep 17 00:00:00 2001
From: David Carlier <devnexen@gmail.com>
Date: Tue, 26 May 2020 21:35:27 +0100
Subject: [PATCH] util/oslib-posix : qemu_init_exec_dir implementation for Mac

Using dyld API to get the full path of the current process.

Signed-off-by: David Carlier <devnexen@gmail.com>
Message-id: CA+XhMqxwC10XHVs4Z-JfE0-WLAU3ztDuU9QKVi31mjr59HWCxg@mail.gmail.com
Reviewed-by: Peter Maydell <peter.maydell@linaro.org>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
---
 util/oslib-posix.c | 15 +++++++++++++++
 1 file changed, 15 insertions(+)

diff --git a/util/oslib-posix.c b/util/oslib-posix.c
index XXXXXXX..XXXXXXX 100644
--- a/util/oslib-posix.c
+++ b/util/oslib-posix.c
@@ -XXX,XX +XXX,XX @@
 #include <lwp.h>
 #endif
 
+#ifdef __APPLE__
+#include <mach-o/dyld.h>
+#endif
+
 #include "qemu/mmap-alloc.h"
 
 #ifdef CONFIG_DEBUG_STACK_USAGE
@@ -XXX,XX +XXX,XX @@ void qemu_init_exec_dir(const char *argv0)
             p = buf;
         }
     }
+#elif defined(__APPLE__)
+    {
+        char fpath[PATH_MAX];
+        uint32_t len = sizeof(fpath);
+        if (_NSGetExecutablePath(fpath, &len) == 0) {
+            p = realpath(fpath, buf);
+            if (!p) {
+                return;
+            }
+        }
+    }
 #endif
     /* If we don't have any way of figuring out the actual executable
        location then try argv[0].  */
-- 
2.20.1

Convert the Neon VREV64 insn from the 2-reg-misc grouping to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-2-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       | 12 ++++++++
 target/arm/translate-neon.inc.c | 50 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 24 ++--------------
 3 files changed, 64 insertions(+), 22 deletions(-)

Convert the pairwise ops VPADDL and VPADAL in the 2-reg-misc grouping
to decodetree.

At this point we can get rid of the weird CPU_V001 #define that was
used to avoid having to explicitly list all the arguments being
passed to some TCG gen/helper functions.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-3-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |   6 ++
 target/arm/translate-neon.inc.c | 149 ++++++++++++++++++++++++++++++++
 target/arm/translate.c          |  35 +-------
 3 files changed, 157 insertions(+), 33 deletions(-)

Convert the Neon VZIP and VUZP insns in the 2-reg-misc group to
decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-4-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  3 ++
 target/arm/translate-neon.inc.c | 74 ++++++++++++++++++++++++++
 target/arm/translate.c          | 92 +--------------------------------
 3 files changed, 79 insertions(+), 90 deletions(-)

Convert the Neon narrowing moves VMQNV, VQMOVN, VQMOVUN in the 2-reg-misc
group to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-5-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  9 ++++
 target/arm/translate-neon.inc.c | 59 ++++++++++++++++++++++++
 target/arm/translate.c          | 81 +--------------------------------
 3 files changed, 70 insertions(+), 79 deletions(-)

Convert the VSHLL insn in the 2-reg-misc Neon group to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-6-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  2 ++
 target/arm/translate-neon.inc.c | 52 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 35 +---------------------
 3 files changed, 55 insertions(+), 34 deletions(-)

Convert the Neon insns in the 2-reg-misc group which are
VCVT between f32 and f16 to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-7-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  3 ++
 target/arm/translate-neon.inc.c | 96 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 65 ++--------------------
 3 files changed, 102 insertions(+), 62 deletions(-)

Convert to decodetree the insns in the Neon 2-reg-misc grouping which
we implement using gvec.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-8-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       | 11 +++++++
 target/arm/translate-neon.inc.c | 55 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 35 +++++----------------
 3 files changed, 74 insertions(+), 27 deletions(-)

Convert the Neon-2-reg misc crypto ops (AESE, AESMC, SHA1H, SHA1SU1)
to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-9-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       | 12 ++++++++
 target/arm/translate-neon.inc.c | 42 ++++++++++++++++++++++++++
 target/arm/translate.c          | 52 +++------------------------------
 3 files changed, 58 insertions(+), 48 deletions(-)

The NeonGenOneOpFn typedef breaks with the pattern of the other
NeonGen*Fn typedefs, because it is a TCGv_i64 -> TCGv_i64 operation
but it does not have '64' in its name. Rename it to NeonGenOne64OpFn,
so that the old name is available for a TCGv_i32 -> TCGv_i32 operation
(which we will need in a subsequent commit).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-10-peter.maydell@linaro.org
---
 target/arm/translate.h     | 2 +-
 target/arm/translate-a64.c | 4 ++--
 2 files changed, 3 insertions(+), 3 deletions(-)

All the other typedefs like these spell "Op" with a lowercase 'p';
remane the NeonGenTwoSingleOPFn and NeonGenTwoDoubleOPFn typedefs to
match.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-11-peter.maydell@linaro.org
---
 target/arm/translate.h          | 4 ++--
 target/arm/translate-a64.c      | 4 ++--
 target/arm/translate-neon.inc.c | 2 +-
 3 files changed, 5 insertions(+), 5 deletions(-)

Make gen_swap_half() take a source and destination TCGv_i32 rather
than modifying the input TCGv_i32; we're going to want to be able to
use it with the more flexible function signature, and this also
brings it into line with other functions like gen_rev16() and
gen_revsh().

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-12-peter.maydell@linaro.org
---
 target/arm/translate-neon.inc.c |  2 +-
 target/arm/translate.c          | 10 +++++-----
 2 files changed, 6 insertions(+), 6 deletions(-)

diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.inc.c
+++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VREV64(DisasContext *s, arg_VREV64 *a)
                 tcg_gen_bswap32_i32(tmp[half], tmp[half]);
                 break;
             case 1:
-                gen_swap_half(tmp[half]);
+                gen_swap_half(tmp[half], tmp[half]);
                 break;
             case 2:
                 break;
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_revsh(TCGv_i32 dest, TCGv_i32 var)
 }
 
 /* Swap low and high halfwords.  */
-static void gen_swap_half(TCGv_i32 var)
+static void gen_swap_half(TCGv_i32 dest, TCGv_i32 var)
 {
-    tcg_gen_rotri_i32(var, var, 16);
+    tcg_gen_rotri_i32(dest, var, 16);
 }
 
 /* Dual 16-bit add.  Result placed in t0 and t1 is marked as dead.
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                         case NEON_2RM_VREV32:
                             switch (size) {
                             case 0: tcg_gen_bswap32_i32(tmp, tmp); break;
-                            case 1: gen_swap_half(tmp); break;
+                            case 1: gen_swap_half(tmp, tmp); break;
                             default: abort();
                             }
                             break;
@@ -XXX,XX +XXX,XX @@ static bool op_smlad(DisasContext *s, arg_rrrr *a, bool m_swap, bool sub)
     t1 = load_reg(s, a->rn);
     t2 = load_reg(s, a->rm);
     if (m_swap) {
-        gen_swap_half(t2);
+        gen_swap_half(t2, t2);
     }
     gen_smul_dual(t1, t2);
 
@@ -XXX,XX +XXX,XX @@ static bool op_smlald(DisasContext *s, arg_rrrr *a, bool m_swap, bool sub)
     t1 = load_reg(s, a->rn);
     t2 = load_reg(s, a->rm);
     if (m_swap) {
-        gen_swap_half(t2);
+        gen_swap_half(t2, t2);
     }
     gen_smul_dual(t1, t2);
 
-- 
2.20.1

Convert the VREV32 and VREV16 insns in the Neon 2-reg-misc group
to decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-13-peter.maydell@linaro.org
---
 target/arm/translate.h          |  1 +
 target/arm/neon-dp.decode       |  2 ++
 target/arm/translate-neon.inc.c | 55 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 12 ++-----
 4 files changed, 60 insertions(+), 10 deletions(-)

diff --git a/target/arm/translate.h b/target/arm/translate.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.h
+++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ typedef void GVecGen4Fn(unsigned, uint32_t, uint32_t, uint32_t,
                         uint32_t, uint32_t, uint32_t);
 
 /* Function prototype for gen_ functions for calling Neon helpers */
+typedef void NeonGenOneOpFn(TCGv_i32, TCGv_i32);
 typedef void NeonGenOneOpEnvFn(TCGv_i32, TCGv_ptr, TCGv_i32);
 typedef void NeonGenTwoOpFn(TCGv_i32, TCGv_i32, TCGv_i32);
 typedef void NeonGenTwoOpEnvFn(TCGv_i32, TCGv_ptr, TCGv_i32, TCGv_i32);
diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/neon-dp.decode
+++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
                  &2misc vm=%vm_dp vd=%vd_dp q=1
 
     VREV64       1111 001 11 . 11 .. 00 .... 0 0000 . . 0 .... @2misc
+    VREV32       1111 001 11 . 11 .. 00 .... 0 0001 . . 0 .... @2misc
+    VREV16       1111 001 11 . 11 .. 00 .... 0 0010 . . 0 .... @2misc
 
     VPADDL_S     1111 001 11 . 11 .. 00 .... 0 0100 . . 0 .... @2misc
     VPADDL_U     1111 001 11 . 11 .. 00 .... 0 0101 . . 0 .... @2misc
diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.inc.c
+++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ DO_2M_CRYPTO(AESIMC, aa32_aes, 0)
 DO_2M_CRYPTO(SHA1H, aa32_sha1, 2)
 DO_2M_CRYPTO(SHA1SU1, aa32_sha1, 2)
 DO_2M_CRYPTO(SHA256SU0, aa32_sha2, 2)
+
+static bool do_2misc(DisasContext *s, arg_2misc *a, NeonGenOneOpFn *fn)
+{
+    int pass;
+
+    /* Handle a 2-reg-misc operation by iterating 32 bits at a time */
+    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
+        return false;
+    }
+
+    /* UNDEF accesses to D16-D31 if they don't exist. */
+    if (!dc_isar_feature(aa32_simd_r32, s) &&
+        ((a->vd | a->vm) & 0x10)) {
+        return false;
+    }
+
+    if (!fn) {
+        return false;
+    }
+
+    if ((a->vd | a->vm) & a->q) {
+        return false;
+    }
+
+    if (!vfp_access_check(s)) {
+        return true;
+    }
+
+    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
+        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
+        fn(tmp, tmp);
+        neon_store_reg(a->vd, pass, tmp);
+    }
+
+    return true;
+}
+
+static bool trans_VREV32(DisasContext *s, arg_2misc *a)
+{
+    static NeonGenOneOpFn * const fn[] = {
+        tcg_gen_bswap32_i32,
+        gen_swap_half,
+        NULL,
+        NULL,
+    };
+    return do_2misc(s, a, fn[a->size]);
+}
+
+static bool trans_VREV16(DisasContext *s, arg_2misc *a)
+{
+    if (a->size != 0) {
+        return false;
+    }
+    return do_2misc(s, a, gen_rev16);
+}
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                 case NEON_2RM_AESE: case NEON_2RM_AESMC:
                 case NEON_2RM_SHA1H:
                 case NEON_2RM_SHA1SU1:
+                case NEON_2RM_VREV32:
+                case NEON_2RM_VREV16:
                     /* handled by decodetree */
                     return 1;
                 case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                     for (pass = 0; pass < (q ? 4 : 2); pass++) {
                         tmp = neon_load_reg(rm, pass);
                         switch (op) {
-                        case NEON_2RM_VREV32:
-                            switch (size) {
-                            case 0: tcg_gen_bswap32_i32(tmp, tmp); break;
-                            case 1: gen_swap_half(tmp, tmp); break;
-                            default: abort();
-                            }
-                            break;
-                        case NEON_2RM_VREV16:
-                            gen_rev16(tmp, tmp);
-                            break;
                         case NEON_2RM_VCLS:
                             switch (size) {
                             case 0: gen_helper_neon_cls_s8(tmp, tmp); break;
-- 
2.20.1

Convert the remaining ops in the Neon 2-reg-misc group which
can be implemented simply with our do_2misc() helper.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-14-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       | 10 +++++
 target/arm/translate-neon.inc.c | 69 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 38 ++++--------------
 3 files changed, 86 insertions(+), 31 deletions(-)

Convert the Neon VQABS and VQNEG insns to decodetree.
Since these are the only ones which need cpu_env passing to
the helper, we wrap the helper rather than creating a whole
new do_2misc_env() function.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-15-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  3 +++
 target/arm/translate-neon.inc.c | 35 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 30 ++--------------------------
 3 files changed, 40 insertions(+), 28 deletions(-)

Convert the Neon 2-reg-misc insns which are implemented with
simple calls to functions that take the input, output and
fpstatus pointer.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-16-peter.maydell@linaro.org
---
 target/arm/translate.h          |  1 +
 target/arm/neon-dp.decode       |  8 +++++
 target/arm/translate-neon.inc.c | 62 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 56 ++++-------------------------
 4 files changed, 78 insertions(+), 49 deletions(-)

diff --git a/target/arm/translate.h b/target/arm/translate.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.h
+++ b/target/arm/translate.h
@@ -XXX,XX +XXX,XX @@ typedef void NeonGenNarrowFn(TCGv_i32, TCGv_i64);
 typedef void NeonGenNarrowEnvFn(TCGv_i32, TCGv_ptr, TCGv_i64);
 typedef void NeonGenWidenFn(TCGv_i64, TCGv_i32);
 typedef void NeonGenTwoOpWidenFn(TCGv_i64, TCGv_i32, TCGv_i32);
+typedef void NeonGenOneSingleOpFn(TCGv_i32, TCGv_i32, TCGv_ptr);
 typedef void NeonGenTwoSingleOpFn(TCGv_i32, TCGv_i32, TCGv_i32, TCGv_ptr);
 typedef void NeonGenTwoDoubleOpFn(TCGv_i64, TCGv_i64, TCGv_i64, TCGv_ptr);
 typedef void NeonGenOne64OpFn(TCGv_i64, TCGv_i64);
diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/neon-dp.decode
+++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
     SHA1SU1      1111 001 11 . 11 .. 10 .... 0 0111 0 . 0 .... @2misc_q1
     SHA256SU0    1111 001 11 . 11 .. 10 .... 0 0111 1 . 0 .... @2misc_q1
 
+    VRINTX       1111 001 11 . 11 .. 10 .... 0 1001 . . 0 .... @2misc
+
     VCVT_F16_F32 1111 001 11 . 11 .. 10 .... 0 1100 0 . 0 .... @2misc_q0
     VCVT_F32_F16 1111 001 11 . 11 .. 10 .... 0 1110 0 . 0 .... @2misc_q0
 
     VRECPE       1111 001 11 . 11 .. 11 .... 0 1000 . . 0 .... @2misc
     VRSQRTE      1111 001 11 . 11 .. 11 .... 0 1001 . . 0 .... @2misc
+    VRECPE_F     1111 001 11 . 11 .. 11 .... 0 1010 . . 0 .... @2misc
+    VRSQRTE_F    1111 001 11 . 11 .. 11 .... 0 1011 . . 0 .... @2misc
+    VCVT_FS      1111 001 11 . 11 .. 11 .... 0 1100 . . 0 .... @2misc
+    VCVT_FU      1111 001 11 . 11 .. 11 .... 0 1101 . . 0 .... @2misc
+    VCVT_SF      1111 001 11 . 11 .. 11 .... 0 1110 . . 0 .... @2misc
+    VCVT_UF      1111 001 11 . 11 .. 11 .... 0 1111 . . 0 .... @2misc
   ]
 
   # Subgroup for size != 0b11
diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.inc.c
+++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VQNEG(DisasContext *s, arg_2misc *a)
     };
     return do_2misc(s, a, fn[a->size]);
 }
+
+static bool do_2misc_fp(DisasContext *s, arg_2misc *a,
+                        NeonGenOneSingleOpFn *fn)
+{
+    int pass;
+    TCGv_ptr fpst;
+
+    /* Handle a 2-reg-misc operation by iterating 32 bits at a time */
+    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
+        return false;
+    }
+
+    /* UNDEF accesses to D16-D31 if they don't exist. */
+    if (!dc_isar_feature(aa32_simd_r32, s) &&
+        ((a->vd | a->vm) & 0x10)) {
+        return false;
+    }
+
+    if (a->size != 2) {
+        /* TODO: FP16 will be the size == 1 case */
+        return false;
+    }
+
+    if ((a->vd | a->vm) & a->q) {
+        return false;
+    }
+
+    if (!vfp_access_check(s)) {
+        return true;
+    }
+
+    fpst = get_fpstatus_ptr(1);
+    for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
+        TCGv_i32 tmp = neon_load_reg(a->vm, pass);
+        fn(tmp, tmp, fpst);
+        neon_store_reg(a->vd, pass, tmp);
+    }
+    tcg_temp_free_ptr(fpst);
+
+    return true;
+}
+
+#define DO_2MISC_FP(INSN, FUNC)                                 \
+    static bool trans_##INSN(DisasContext *s, arg_2misc *a)     \
+    {                                                           \
+        return do_2misc_fp(s, a, FUNC);                         \
+    }
+
+DO_2MISC_FP(VRECPE_F, gen_helper_recpe_f32)
+DO_2MISC_FP(VRSQRTE_F, gen_helper_rsqrte_f32)
+DO_2MISC_FP(VCVT_FS, gen_helper_vfp_sitos)
+DO_2MISC_FP(VCVT_FU, gen_helper_vfp_uitos)
+DO_2MISC_FP(VCVT_SF, gen_helper_vfp_tosizs)
+DO_2MISC_FP(VCVT_UF, gen_helper_vfp_touizs)
+
+static bool trans_VRINTX(DisasContext *s, arg_2misc *a)
+{
+    if (!arm_dc_feature(s, ARM_FEATURE_V8)) {
+        return false;
+    }
+    return do_2misc_fp(s, a, gen_helper_rints_exact);
+}
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                 case NEON_2RM_VRSQRTE:
                 case NEON_2RM_VQABS:
                 case NEON_2RM_VQNEG:
+                case NEON_2RM_VRECPE_F:
+                case NEON_2RM_VRSQRTE_F:
+                case NEON_2RM_VCVT_FS:
+                case NEON_2RM_VCVT_FU:
+                case NEON_2RM_VCVT_SF:
+                case NEON_2RM_VCVT_UF:
+                case NEON_2RM_VRINTX:
                     /* handled by decodetree */
                     return 1;
                 case NEON_2RM_VTRN:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                             tcg_temp_free_i32(tcg_rmode);
                             break;
                         }
-                        case NEON_2RM_VRINTX:
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_rints_exact(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
                         case NEON_2RM_VCVTAU:
                         case NEON_2RM_VCVTAS:
                         case NEON_2RM_VCVTNU:
@@ -XXX,XX +XXX,XX @@ static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
                             tcg_temp_free_ptr(fpst);
                             break;
                         }
-                        case NEON_2RM_VRECPE_F:
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_recpe_f32(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
-                        case NEON_2RM_VRSQRTE_F:
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_rsqrte_f32(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
-                        case NEON_2RM_VCVT_FS: /* VCVT.F32.S32 */
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_vfp_sitos(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
-                        case NEON_2RM_VCVT_FU: /* VCVT.F32.U32 */
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_vfp_uitos(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
-                        case NEON_2RM_VCVT_SF: /* VCVT.S32.F32 */
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_vfp_tosizs(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
-                        case NEON_2RM_VCVT_UF: /* VCVT.U32.F32 */
-                        {
-                            TCGv_ptr fpstatus = get_fpstatus_ptr(1);
-                            gen_helper_vfp_touizs(tmp, tmp, fpstatus);
-                            tcg_temp_free_ptr(fpstatus);
-                            break;
-                        }
                         default:
                             /* Reserved op values were caught by the
                              * neon_2rm_sizes[] check earlier.
-- 
2.20.1

Convert the fp-compare-with-zero insns in the Neon 2-reg-misc group to
decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-17-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  6 ++++
 target/arm/translate-neon.inc.c | 28 ++++++++++++++++++
 target/arm/translate.c          | 50 ++++-----------------------------
 3 files changed, 39 insertions(+), 45 deletions(-)

Convert the Neon 2-reg-misc VRINT insns to decodetree.
Giving these insns their own do_vrint() function allows us
to change the rounding mode just once at the start and end
rather than doing it for every element in the vector.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-18-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  8 +++++
 target/arm/translate-neon.inc.c | 61 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 31 +++--------------
 3 files changed, 74 insertions(+), 26 deletions(-)

Convert the VCVT instructions in the 2-reg-misc grouping to
decodetree.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-19-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  9 +++++
 target/arm/translate-neon.inc.c | 70 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 70 ++++-----------------------------
 3 files changed, 87 insertions(+), 62 deletions(-)

Convert the Neon VSWP insn to decodetree. Since the new implementation
doesn't have to share a pass-loop with the other 2-reg-misc operations
we can implement the swap with 64-bit accesses rather than 32-bits
(which brings us into line with the pseudocode and is more efficient).

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-20-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |  2 ++
 target/arm/translate-neon.inc.c | 41 +++++++++++++++++++++++++++++++++
 target/arm/translate.c          |  5 +---
 3 files changed, 44 insertions(+), 4 deletions(-)

Convert the Neon VTRN insn to decodetree. This is the last insn in the
Neon data-processing group, so we can remove all the now-unused old
decoder framework.

It's possible that there's a more efficient implementation of
VTRN, but for this conversion we just copy the existing approach.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-21-peter.maydell@linaro.org
---
 target/arm/neon-dp.decode       |   2 +-
 target/arm/translate-neon.inc.c |  90 ++++++++
 target/arm/translate.c          | 363 +-------------------------------
 3 files changed, 93 insertions(+), 362 deletions(-)

diff --git a/target/arm/neon-dp.decode b/target/arm/neon-dp.decode
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/neon-dp.decode
+++ b/target/arm/neon-dp.decode
@@ -XXX,XX +XXX,XX @@ Vimm_1r          1111 001 . 1 . 000 ... .... cmode:4 0 . op:1 1 .... @1reg_imm
     VNEG_F       1111 001 11 . 11 .. 01 .... 0 1111 . . 0 .... @2misc
 
     VSWP         1111 001 11 . 11 .. 10 .... 0 0000 . . 0 .... @2misc
-
+    VTRN         1111 001 11 . 11 .. 10 .... 0 0001 . . 0 .... @2misc
     VUZP         1111 001 11 . 11 .. 10 .... 0 0010 . . 0 .... @2misc
     VZIP         1111 001 11 . 11 .. 10 .... 0 0011 . . 0 .... @2misc
 
diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.inc.c
+++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static bool trans_VSWP(DisasContext *s, arg_2misc *a)
 
     return true;
 }
+static void gen_neon_trn_u8(TCGv_i32 t0, TCGv_i32 t1)
+{
+    TCGv_i32 rd, tmp;
+
+    rd = tcg_temp_new_i32();
+    tmp = tcg_temp_new_i32();
+
+    tcg_gen_shli_i32(rd, t0, 8);
+    tcg_gen_andi_i32(rd, rd, 0xff00ff00);
+    tcg_gen_andi_i32(tmp, t1, 0x00ff00ff);
+    tcg_gen_or_i32(rd, rd, tmp);
+
+    tcg_gen_shri_i32(t1, t1, 8);
+    tcg_gen_andi_i32(t1, t1, 0x00ff00ff);
+    tcg_gen_andi_i32(tmp, t0, 0xff00ff00);
+    tcg_gen_or_i32(t1, t1, tmp);
+    tcg_gen_mov_i32(t0, rd);
+
+    tcg_temp_free_i32(tmp);
+    tcg_temp_free_i32(rd);
+}
+
+static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
+{
+    TCGv_i32 rd, tmp;
+
+    rd = tcg_temp_new_i32();
+    tmp = tcg_temp_new_i32();
+
+    tcg_gen_shli_i32(rd, t0, 16);
+    tcg_gen_andi_i32(tmp, t1, 0xffff);
+    tcg_gen_or_i32(rd, rd, tmp);
+    tcg_gen_shri_i32(t1, t1, 16);
+    tcg_gen_andi_i32(tmp, t0, 0xffff0000);
+    tcg_gen_or_i32(t1, t1, tmp);
+    tcg_gen_mov_i32(t0, rd);
+
+    tcg_temp_free_i32(tmp);
+    tcg_temp_free_i32(rd);
+}
+
+static bool trans_VTRN(DisasContext *s, arg_2misc *a)
+{
+    TCGv_i32 tmp, tmp2;
+    int pass;
+
+    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
+        return false;
+    }
+
+    /* UNDEF accesses to D16-D31 if they don't exist. */
+    if (!dc_isar_feature(aa32_simd_r32, s) &&
+        ((a->vd | a->vm) & 0x10)) {
+        return false;
+    }
+
+    if ((a->vd | a->vm) & a->q) {
+        return false;
+    }
+
+    if (a->size == 3) {
+        return false;
+    }
+
+    if (!vfp_access_check(s)) {
+        return true;
+    }
+
+    if (a->size == 2) {
+        for (pass = 0; pass < (a->q ? 4 : 2); pass += 2) {
+            tmp = neon_load_reg(a->vm, pass);
+            tmp2 = neon_load_reg(a->vd, pass + 1);
+            neon_store_reg(a->vm, pass, tmp2);
+            neon_store_reg(a->vd, pass + 1, tmp);
+        }
+    } else {
+        for (pass = 0; pass < (a->q ? 4 : 2); pass++) {
+            tmp = neon_load_reg(a->vm, pass);
+            tmp2 = neon_load_reg(a->vd, pass);
+            if (a->size == 0) {
+                gen_neon_trn_u8(tmp, tmp2);
+            } else {
+                gen_neon_trn_u16(tmp, tmp2);
+            }
+            neon_store_reg(a->vm, pass, tmp2);
+            neon_store_reg(a->vd, pass, tmp);
+        }
+    }
+    return true;
+}
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static void gen_exception_return(DisasContext *s, TCGv_i32 pc)
     gen_rfe(s, pc, load_cpu_field(spsr));
 }
 
-static void gen_neon_trn_u8(TCGv_i32 t0, TCGv_i32 t1)
-{
-    TCGv_i32 rd, tmp;
-
-    rd = tcg_temp_new_i32();
-    tmp = tcg_temp_new_i32();
-
-    tcg_gen_shli_i32(rd, t0, 8);
-    tcg_gen_andi_i32(rd, rd, 0xff00ff00);
-    tcg_gen_andi_i32(tmp, t1, 0x00ff00ff);
-    tcg_gen_or_i32(rd, rd, tmp);
-
-    tcg_gen_shri_i32(t1, t1, 8);
-    tcg_gen_andi_i32(t1, t1, 0x00ff00ff);
-    tcg_gen_andi_i32(tmp, t0, 0xff00ff00);
-    tcg_gen_or_i32(t1, t1, tmp);
-    tcg_gen_mov_i32(t0, rd);
-
-    tcg_temp_free_i32(tmp);
-    tcg_temp_free_i32(rd);
-}
-
-static void gen_neon_trn_u16(TCGv_i32 t0, TCGv_i32 t1)
-{
-    TCGv_i32 rd, tmp;
-
-    rd = tcg_temp_new_i32();
-    tmp = tcg_temp_new_i32();
-
-    tcg_gen_shli_i32(rd, t0, 16);
-    tcg_gen_andi_i32(tmp, t1, 0xffff);
-    tcg_gen_or_i32(rd, rd, tmp);
-    tcg_gen_shri_i32(t1, t1, 16);
-    tcg_gen_andi_i32(tmp, t0, 0xffff0000);
-    tcg_gen_or_i32(t1, t1, tmp);
-    tcg_gen_mov_i32(t0, rd);
-
-    tcg_temp_free_i32(tmp);
-    tcg_temp_free_i32(rd);
-}
-
-/* Symbolic constants for op fields for Neon 2-register miscellaneous.
- * The values correspond to bits [17:16,10:7]; see the ARM ARM DDI0406B
- * table A7-13.
- */
-#define NEON_2RM_VREV64 0
-#define NEON_2RM_VREV32 1
-#define NEON_2RM_VREV16 2
-#define NEON_2RM_VPADDL 4
-#define NEON_2RM_VPADDL_U 5
-#define NEON_2RM_AESE 6 /* Includes AESD */
-#define NEON_2RM_AESMC 7 /* Includes AESIMC */
-#define NEON_2RM_VCLS 8
-#define NEON_2RM_VCLZ 9
-#define NEON_2RM_VCNT 10
-#define NEON_2RM_VMVN 11
-#define NEON_2RM_VPADAL 12
-#define NEON_2RM_VPADAL_U 13
-#define NEON_2RM_VQABS 14
-#define NEON_2RM_VQNEG 15
-#define NEON_2RM_VCGT0 16
-#define NEON_2RM_VCGE0 17
-#define NEON_2RM_VCEQ0 18
-#define NEON_2RM_VCLE0 19
-#define NEON_2RM_VCLT0 20
-#define NEON_2RM_SHA1H 21
-#define NEON_2RM_VABS 22
-#define NEON_2RM_VNEG 23
-#define NEON_2RM_VCGT0_F 24
-#define NEON_2RM_VCGE0_F 25
-#define NEON_2RM_VCEQ0_F 26
-#define NEON_2RM_VCLE0_F 27
-#define NEON_2RM_VCLT0_F 28
-#define NEON_2RM_VABS_F 30
-#define NEON_2RM_VNEG_F 31
-#define NEON_2RM_VSWP 32
-#define NEON_2RM_VTRN 33
-#define NEON_2RM_VUZP 34
-#define NEON_2RM_VZIP 35
-#define NEON_2RM_VMOVN 36 /* Includes VQMOVN, VQMOVUN */
-#define NEON_2RM_VQMOVN 37 /* Includes VQMOVUN */
-#define NEON_2RM_VSHLL 38
-#define NEON_2RM_SHA1SU1 39 /* Includes SHA256SU0 */
-#define NEON_2RM_VRINTN 40
-#define NEON_2RM_VRINTX 41
-#define NEON_2RM_VRINTA 42
-#define NEON_2RM_VRINTZ 43
-#define NEON_2RM_VCVT_F16_F32 44
-#define NEON_2RM_VRINTM 45
-#define NEON_2RM_VCVT_F32_F16 46
-#define NEON_2RM_VRINTP 47
-#define NEON_2RM_VCVTAU 48
-#define NEON_2RM_VCVTAS 49
-#define NEON_2RM_VCVTNU 50
-#define NEON_2RM_VCVTNS 51
-#define NEON_2RM_VCVTPU 52
-#define NEON_2RM_VCVTPS 53
-#define NEON_2RM_VCVTMU 54
-#define NEON_2RM_VCVTMS 55
-#define NEON_2RM_VRECPE 56
-#define NEON_2RM_VRSQRTE 57
-#define NEON_2RM_VRECPE_F 58
-#define NEON_2RM_VRSQRTE_F 59
-#define NEON_2RM_VCVT_FS 60
-#define NEON_2RM_VCVT_FU 61
-#define NEON_2RM_VCVT_SF 62
-#define NEON_2RM_VCVT_UF 63
-
-/* Each entry in this array has bit n set if the insn allows
- * size value n (otherwise it will UNDEF). Since unallocated
- * op values will have no bits set they always UNDEF.
- */
-static const uint8_t neon_2rm_sizes[] = {
-    [NEON_2RM_VREV64] = 0x7,
-    [NEON_2RM_VREV32] = 0x3,
-    [NEON_2RM_VREV16] = 0x1,
-    [NEON_2RM_VPADDL] = 0x7,
-    [NEON_2RM_VPADDL_U] = 0x7,
-    [NEON_2RM_AESE] = 0x1,
-    [NEON_2RM_AESMC] = 0x1,
-    [NEON_2RM_VCLS] = 0x7,
-    [NEON_2RM_VCLZ] = 0x7,
-    [NEON_2RM_VCNT] = 0x1,
-    [NEON_2RM_VMVN] = 0x1,
-    [NEON_2RM_VPADAL] = 0x7,
-    [NEON_2RM_VPADAL_U] = 0x7,
-    [NEON_2RM_VQABS] = 0x7,
-    [NEON_2RM_VQNEG] = 0x7,
-    [NEON_2RM_VCGT0] = 0x7,
-    [NEON_2RM_VCGE0] = 0x7,
-    [NEON_2RM_VCEQ0] = 0x7,
-    [NEON_2RM_VCLE0] = 0x7,
-    [NEON_2RM_VCLT0] = 0x7,
-    [NEON_2RM_SHA1H] = 0x4,
-    [NEON_2RM_VABS] = 0x7,
-    [NEON_2RM_VNEG] = 0x7,
-    [NEON_2RM_VCGT0_F] = 0x4,
-    [NEON_2RM_VCGE0_F] = 0x4,
-    [NEON_2RM_VCEQ0_F] = 0x4,
-    [NEON_2RM_VCLE0_F] = 0x4,
-    [NEON_2RM_VCLT0_F] = 0x4,
-    [NEON_2RM_VABS_F] = 0x4,
-    [NEON_2RM_VNEG_F] = 0x4,
-    [NEON_2RM_VSWP] = 0x1,
-    [NEON_2RM_VTRN] = 0x7,
-    [NEON_2RM_VUZP] = 0x7,
-    [NEON_2RM_VZIP] = 0x7,
-    [NEON_2RM_VMOVN] = 0x7,
-    [NEON_2RM_VQMOVN] = 0x7,
-    [NEON_2RM_VSHLL] = 0x7,
-    [NEON_2RM_SHA1SU1] = 0x4,
-    [NEON_2RM_VRINTN] = 0x4,
-    [NEON_2RM_VRINTX] = 0x4,
-    [NEON_2RM_VRINTA] = 0x4,
-    [NEON_2RM_VRINTZ] = 0x4,
-    [NEON_2RM_VCVT_F16_F32] = 0x2,
-    [NEON_2RM_VRINTM] = 0x4,
-    [NEON_2RM_VCVT_F32_F16] = 0x2,
-    [NEON_2RM_VRINTP] = 0x4,
-    [NEON_2RM_VCVTAU] = 0x4,
-    [NEON_2RM_VCVTAS] = 0x4,
-    [NEON_2RM_VCVTNU] = 0x4,
-    [NEON_2RM_VCVTNS] = 0x4,
-    [NEON_2RM_VCVTPU] = 0x4,
-    [NEON_2RM_VCVTPS] = 0x4,
-    [NEON_2RM_VCVTMU] = 0x4,
-    [NEON_2RM_VCVTMS] = 0x4,
-    [NEON_2RM_VRECPE] = 0x4,
-    [NEON_2RM_VRSQRTE] = 0x4,
-    [NEON_2RM_VRECPE_F] = 0x4,
-    [NEON_2RM_VRSQRTE_F] = 0x4,
-    [NEON_2RM_VCVT_FS] = 0x4,
-    [NEON_2RM_VCVT_FU] = 0x4,
-    [NEON_2RM_VCVT_SF] = 0x4,
-    [NEON_2RM_VCVT_UF] = 0x4,
-};
-
 static void gen_gvec_fn3_qc(uint32_t rd_ofs, uint32_t rn_ofs, uint32_t rm_ofs,
                             uint32_t opr_sz, uint32_t max_sz,
                             gen_helper_gvec_3_ptr *fn)
@@ -XXX,XX +XXX,XX @@ void gen_gvec_uaba(unsigned vece, uint32_t rd_ofs, uint32_t rn_ofs,
     tcg_gen_gvec_3(rd_ofs, rn_ofs, rm_ofs, opr_sz, max_sz, &ops[vece]);
 }
 
-/* Translate a NEON data processing instruction.  Return nonzero if the
-   instruction is invalid.
-   We process data in a mixture of 32-bit and 64-bit chunks.
-   Mostly we use 32-bit chunks so we can use normal scalar instructions.  */
-
-static int disas_neon_data_insn(DisasContext *s, uint32_t insn)
-{
-    int op;
-    int q;
-    int rd, rm;
-    int size;
-    int pass;
-    int u;
-    TCGv_i32 tmp, tmp2;
-
-    if (!arm_dc_feature(s, ARM_FEATURE_NEON)) {
-        return 1;
-    }
-
-    /* FIXME: this access check should not take precedence over UNDEF
-     * for invalid encodings; we will generate incorrect syndrome information
-     * for attempts to execute invalid vfp/neon encodings with FP disabled.
-     */
-    if (s->fp_excp_el) {
-        gen_exception_insn(s, s->pc_curr, EXCP_UDEF,
-                           syn_simd_access_trap(1, 0xe, false), s->fp_excp_el);
-        return 0;
-    }
-
-    if (!s->vfp_enabled)
-      return 1;
-    q = (insn & (1 << 6)) != 0;
-    u = (insn >> 24) & 1;
-    VFP_DREG_D(rd, insn);
-    VFP_DREG_M(rm, insn);
-    size = (insn >> 20) & 3;
-
-    if ((insn & (1 << 23)) == 0) {
-        /* Three register same length: handled by decodetree */
-        return 1;
-    } else if (insn & (1 << 4)) {
-        /* Two registers and shift or reg and imm: handled by decodetree */
-        return 1;
-    } else { /* (insn & 0x00800010 == 0x00800000) */
-        if (size != 3) {
-            /*
-             * Three registers of different lengths, or two registers and
-             * a scalar: handled by decodetree
-             */
-            return 1;
-        } else { /* size == 3 */
-            if (!u) {
-                /* Extract: handled by decodetree */
-                return 1;
-            } else if ((insn & (1 << 11)) == 0) {
-                /* Two register misc.  */
-                op = ((insn >> 12) & 0x30) | ((insn >> 7) & 0xf);
-                size = (insn >> 18) & 3;
-                /* UNDEF for unknown op values and bad op-size combinations */
-                if ((neon_2rm_sizes[op] & (1 << size)) == 0) {
-                    return 1;
-                }
-                if (q && ((rm | rd) & 1)) {
-                    return 1;
-                }
-                switch (op) {
-                case NEON_2RM_VREV64:
-                case NEON_2RM_VPADDL: case NEON_2RM_VPADDL_U:
-                case NEON_2RM_VPADAL: case NEON_2RM_VPADAL_U:
-                case NEON_2RM_VUZP:
-                case NEON_2RM_VZIP:
-                case NEON_2RM_VMOVN: case NEON_2RM_VQMOVN:
-                case NEON_2RM_VSHLL:
-                case NEON_2RM_VCVT_F16_F32:
-                case NEON_2RM_VCVT_F32_F16:
-                case NEON_2RM_VMVN:
-                case NEON_2RM_VNEG:
-                case NEON_2RM_VABS:
-                case NEON_2RM_VCEQ0:
-                case NEON_2RM_VCGT0:
-                case NEON_2RM_VCLE0:
-                case NEON_2RM_VCGE0:
-                case NEON_2RM_VCLT0:
-                case NEON_2RM_AESE: case NEON_2RM_AESMC:
-                case NEON_2RM_SHA1H:
-                case NEON_2RM_SHA1SU1:
-                case NEON_2RM_VREV32:
-                case NEON_2RM_VREV16:
-                case NEON_2RM_VCLS:
-                case NEON_2RM_VCLZ:
-                case NEON_2RM_VCNT:
-                case NEON_2RM_VABS_F:
-                case NEON_2RM_VNEG_F:
-                case NEON_2RM_VRECPE:
-                case NEON_2RM_VRSQRTE:
-                case NEON_2RM_VQABS:
-                case NEON_2RM_VQNEG:
-                case NEON_2RM_VRECPE_F:
-                case NEON_2RM_VRSQRTE_F:
-                case NEON_2RM_VCVT_FS:
-                case NEON_2RM_VCVT_FU:
-                case NEON_2RM_VCVT_SF:
-                case NEON_2RM_VCVT_UF:
-                case NEON_2RM_VRINTX:
-                case NEON_2RM_VCGT0_F:
-                case NEON_2RM_VCGE0_F:
-                case NEON_2RM_VCEQ0_F:
-                case NEON_2RM_VCLE0_F:
-                case NEON_2RM_VCLT0_F:
-                case NEON_2RM_VRINTN:
-                case NEON_2RM_VRINTA:
-                case NEON_2RM_VRINTM:
-                case NEON_2RM_VRINTP:
-                case NEON_2RM_VRINTZ:
-                case NEON_2RM_VCVTAU:
-                case NEON_2RM_VCVTAS:
-                case NEON_2RM_VCVTNU:
-                case NEON_2RM_VCVTNS:
-                case NEON_2RM_VCVTPU:
-                case NEON_2RM_VCVTPS:
-                case NEON_2RM_VCVTMU:
-                case NEON_2RM_VCVTMS:
-                case NEON_2RM_VSWP:
-                    /* handled by decodetree */
-                    return 1;
-                case NEON_2RM_VTRN:
-                    if (size == 2) {
-                        int n;
-                        for (n = 0; n < (q ? 4 : 2); n += 2) {
-                            tmp = neon_load_reg(rm, n);
-                            tmp2 = neon_load_reg(rd, n + 1);
-                            neon_store_reg(rm, n, tmp2);
-                            neon_store_reg(rd, n + 1, tmp);
-                        }
-                    } else {
-                        goto elementwise;
-                    }
-                    break;
-
-                default:
-                elementwise:
-                    for (pass = 0; pass < (q ? 4 : 2); pass++) {
-                        tmp = neon_load_reg(rm, pass);
-                        switch (op) {
-                        case NEON_2RM_VTRN:
-                            tmp2 = neon_load_reg(rd, pass);
-                            switch (size) {
-                            case 0: gen_neon_trn_u8(tmp, tmp2); break;
-                            case 1: gen_neon_trn_u16(tmp, tmp2); break;
-                            default: abort();
-                            }
-                            neon_store_reg(rm, pass, tmp2);
-                            break;
-                        default:
-                            /* Reserved op values were caught by the
-                             * neon_2rm_sizes[] check earlier.
-                             */
-                            abort();
-                        }
-                        neon_store_reg(rd, pass, tmp);
-                    }
-                    break;
-                }
-            } else {
-                /* VTBL, VTBX, VDUP: handled by decodetree */
-                return 1;
-            }
-        }
-    }
-    return 0;
-}
-
 static int disas_coproc_insn(DisasContext *s, uint32_t insn)
 {
     int cpnum, is64, crn, crm, opc1, opc2, isread, rt, rt2;
@@ -XXX,XX +XXX,XX @@ static void disas_arm_insn(DisasContext *s, unsigned int insn)
         }
         /* fall back to legacy decoder */
 
-        if (((insn >> 25) & 7) == 1) {
-            /* NEON Data processing.  */
-            if (disas_neon_data_insn(s, insn)) {
-                goto illegal_op;
-            }
-            return;
-        }
         if ((insn & 0x0e000f00) == 0x0c000100) {
             if (arm_dc_feature(s, ARM_FEATURE_IWMMXT)) {
                 /* iWMMXt register transfer.  */
@@ -XXX,XX +XXX,XX @@ static void disas_thumb2_insn(DisasContext *s, uint32_t insn)
             break;
         }
         if (((insn >> 24) & 3) == 3) {
-            /* Translate into the equivalent ARM encoding.  */
-            insn = (insn & 0xe2ffffff) | ((insn & (1 << 28)) >> 4) | (1 << 28);
-            if (disas_neon_data_insn(s, insn)) {
-                goto illegal_op;
-            }
+            /* Neon DP, but failed disas_neon_dp() */
+            goto illegal_op;
         } else if (((insn >> 8) & 0xe) == 10) {
             /* VFP, but failed disas_vfp.  */
             goto illegal_op;
-- 
2.20.1

The functions neon_element_offset(), neon_load_element(),
neon_load_element64(), neon_store_element() and
neon_store_element64() are used only in the translate-neon.inc.c
file, so move their definitions there.

Since the .inc.c file is #included in translate.c this doesn't make
much difference currently, but it's a more logical place to put the
functions and it might be helpful if we ever decide to try to make
the .inc.c files genuinely separate compilation units.

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200616170844.13318-22-peter.maydell@linaro.org
---
 target/arm/translate-neon.inc.c | 101 ++++++++++++++++++++++++++++++++
 target/arm/translate.c          | 101 --------------------------------
 2 files changed, 101 insertions(+), 101 deletions(-)

diff --git a/target/arm/translate-neon.inc.c b/target/arm/translate-neon.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-neon.inc.c
+++ b/target/arm/translate-neon.inc.c
@@ -XXX,XX +XXX,XX @@ static inline int rsub_8(DisasContext *s, int x)
 #include "decode-neon-ls.inc.c"
 #include "decode-neon-shared.inc.c"
 
+/* Return the offset of a 2**SIZE piece of a NEON register, at index ELE,
+ * where 0 is the least significant end of the register.
+ */
+static inline long
+neon_element_offset(int reg, int element, MemOp size)
+{
+    int element_size = 1 << size;
+    int ofs = element * element_size;
+#ifdef HOST_WORDS_BIGENDIAN
+    /* Calculate the offset assuming fully little-endian,
+     * then XOR to account for the order of the 8-byte units.
+     */
+    if (element_size < 8) {
+        ofs ^= 8 - element_size;
+    }
+#endif
+    return neon_reg_offset(reg, 0) + ofs;
+}
+
+static void neon_load_element(TCGv_i32 var, int reg, int ele, MemOp mop)
+{
+    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
+
+    switch (mop) {
+    case MO_UB:
+        tcg_gen_ld8u_i32(var, cpu_env, offset);
+        break;
+    case MO_UW:
+        tcg_gen_ld16u_i32(var, cpu_env, offset);
+        break;
+    case MO_UL:
+        tcg_gen_ld_i32(var, cpu_env, offset);
+        break;
+    default:
+        g_assert_not_reached();
+    }
+}
+
+static void neon_load_element64(TCGv_i64 var, int reg, int ele, MemOp mop)
+{
+    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
+
+    switch (mop) {
+    case MO_UB:
+        tcg_gen_ld8u_i64(var, cpu_env, offset);
+        break;
+    case MO_UW:
+        tcg_gen_ld16u_i64(var, cpu_env, offset);
+        break;
+    case MO_UL:
+        tcg_gen_ld32u_i64(var, cpu_env, offset);
+        break;
+    case MO_Q:
+        tcg_gen_ld_i64(var, cpu_env, offset);
+        break;
+    default:
+        g_assert_not_reached();
+    }
+}
+
+static void neon_store_element(int reg, int ele, MemOp size, TCGv_i32 var)
+{
+    long offset = neon_element_offset(reg, ele, size);
+
+    switch (size) {
+    case MO_8:
+        tcg_gen_st8_i32(var, cpu_env, offset);
+        break;
+    case MO_16:
+        tcg_gen_st16_i32(var, cpu_env, offset);
+        break;
+    case MO_32:
+        tcg_gen_st_i32(var, cpu_env, offset);
+        break;
+    default:
+        g_assert_not_reached();
+    }
+}
+
+static void neon_store_element64(int reg, int ele, MemOp size, TCGv_i64 var)
+{
+    long offset = neon_element_offset(reg, ele, size);
+
+    switch (size) {
+    case MO_8:
+        tcg_gen_st8_i64(var, cpu_env, offset);
+        break;
+    case MO_16:
+        tcg_gen_st16_i64(var, cpu_env, offset);
+        break;
+    case MO_32:
+        tcg_gen_st32_i64(var, cpu_env, offset);
+        break;
+    case MO_64:
+        tcg_gen_st_i64(var, cpu_env, offset);
+        break;
+    default:
+        g_assert_not_reached();
+    }
+}
+
 static bool trans_VCMLA(DisasContext *s, arg_VCMLA *a)
 {
     int opr_sz;
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ neon_reg_offset (int reg, int n)
     return vfp_reg_offset(0, sreg);
 }
 
-/* Return the offset of a 2**SIZE piece of a NEON register, at index ELE,
- * where 0 is the least significant end of the register.
- */
-static inline long
-neon_element_offset(int reg, int element, MemOp size)
-{
-    int element_size = 1 << size;
-    int ofs = element * element_size;
-#ifdef HOST_WORDS_BIGENDIAN
-    /* Calculate the offset assuming fully little-endian,
-     * then XOR to account for the order of the 8-byte units.
-     */
-    if (element_size < 8) {
-        ofs ^= 8 - element_size;
-    }
-#endif
-    return neon_reg_offset(reg, 0) + ofs;
-}
-
 static TCGv_i32 neon_load_reg(int reg, int pass)
 {
     TCGv_i32 tmp = tcg_temp_new_i32();
@@ -XXX,XX +XXX,XX @@ static TCGv_i32 neon_load_reg(int reg, int pass)
     return tmp;
 }
 
-static void neon_load_element(TCGv_i32 var, int reg, int ele, MemOp mop)
-{
-    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
-
-    switch (mop) {
-    case MO_UB:
-        tcg_gen_ld8u_i32(var, cpu_env, offset);
-        break;
-    case MO_UW:
-        tcg_gen_ld16u_i32(var, cpu_env, offset);
-        break;
-    case MO_UL:
-        tcg_gen_ld_i32(var, cpu_env, offset);
-        break;
-    default:
-        g_assert_not_reached();
-    }
-}
-
-static void neon_load_element64(TCGv_i64 var, int reg, int ele, MemOp mop)
-{
-    long offset = neon_element_offset(reg, ele, mop & MO_SIZE);
-
-    switch (mop) {
-    case MO_UB:
-        tcg_gen_ld8u_i64(var, cpu_env, offset);
-        break;
-    case MO_UW:
-        tcg_gen_ld16u_i64(var, cpu_env, offset);
-        break;
-    case MO_UL:
-        tcg_gen_ld32u_i64(var, cpu_env, offset);
-        break;
-    case MO_Q:
-        tcg_gen_ld_i64(var, cpu_env, offset);
-        break;
-    default:
-        g_assert_not_reached();
-    }
-}
-
 static void neon_store_reg(int reg, int pass, TCGv_i32 var)
 {
     tcg_gen_st_i32(var, cpu_env, neon_reg_offset(reg, pass));
     tcg_temp_free_i32(var);
 }
 
-static void neon_store_element(int reg, int ele, MemOp size, TCGv_i32 var)
-{
-    long offset = neon_element_offset(reg, ele, size);
-
-    switch (size) {
-    case MO_8:
-        tcg_gen_st8_i32(var, cpu_env, offset);
-        break;
-    case MO_16:
-        tcg_gen_st16_i32(var, cpu_env, offset);
-        break;
-    case MO_32:
-        tcg_gen_st_i32(var, cpu_env, offset);
-        break;
-    default:
-        g_assert_not_reached();
-    }
-}
-
-static void neon_store_element64(int reg, int ele, MemOp size, TCGv_i64 var)
-{
-    long offset = neon_element_offset(reg, ele, size);
-
-    switch (size) {
-    case MO_8:
-        tcg_gen_st8_i64(var, cpu_env, offset);
-        break;
-    case MO_16:
-        tcg_gen_st16_i64(var, cpu_env, offset);
-        break;
-    case MO_32:
-        tcg_gen_st32_i64(var, cpu_env, offset);
-        break;
-    case MO_64:
-        tcg_gen_st_i64(var, cpu_env, offset);
-        break;
-    default:
-        g_assert_not_reached();
-    }
-}
-
 static inline void neon_load_reg64(TCGv_i64 var, int reg)
 {
     tcg_gen_ld_i64(var, cpu_env, vfp_reg_offset(1, reg));
-- 
2.20.1

Since commit ba3e7926691ed3 it has been unnecessary for target code
to call gen_io_end() after an IO instruction in icount mode; it is
sufficient to call gen_io_start() before it and to force the end of
the TB.

Many now-unnecessary calls to gen_io_end() were removed in commit
9e9b10c6491153b, but some were missed or accidentally added later.
Remove unneeded calls from the arm target:

* the call in the handling of exception-return-via-LDM is
   unnecessary, and the code is already forcing end-of-TB
 * the call in the VFP access check code is more complicated:
   we weren't ending the TB, so we need to add the code to
   force that by setting DISAS_UPDATE
 * the doc comment for ARM_CP_IO doesn't need to mention
   gen_io_end() any more

Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Alex Bennée <alex.bennee@linaro.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Reviewed-by: Pavel Dovgalyuk <Pavel.Dovgaluk@ispras.ru>
Message-id: 20200619170324.12093-1-peter.maydell@linaro.org
---
 target/arm/cpu.h               | 2 +-
 target/arm/translate-vfp.inc.c | 7 +++----
 target/arm/translate.c         | 3 ---
 3 files changed, 4 insertions(+), 8 deletions(-)

diff --git a/target/arm/cpu.h b/target/arm/cpu.h
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/cpu.h
+++ b/target/arm/cpu.h
@@ -XXX,XX +XXX,XX @@ static inline uint64_t cpreg_to_kvm_id(uint32_t cpregid)
  * migration or KVM state synchronization. (Typically this is for "registers"
  * which are actually used as instructions for cache maintenance and so on.)
  * IO indicates that this register does I/O and therefore its accesses
- * need to be surrounded by gen_io_start()/gen_io_end(). In particular,
+ * need to be marked with gen_io_start() and also end the TB. In particular,
  * registers which implement clocks or timers require this.
  * RAISES_EXC is for when the read or write hook might raise an exception;
  * the generated code will synchronize the CPU state before calling the hook
diff --git a/target/arm/translate-vfp.inc.c b/target/arm/translate-vfp.inc.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-vfp.inc.c
+++ b/target/arm/translate-vfp.inc.c
@@ -XXX,XX +XXX,XX @@ static bool full_vfp_access_check(DisasContext *s, bool ignore_vfp_enabled)
         if (s->v7m_lspact) {
             /*
              * Lazy state saving affects external memory and also the NVIC,
-             * so we must mark it as an IO operation for icount.
+             * so we must mark it as an IO operation for icount (and cause
+             * this to be the last insn in the TB).
              */
             if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
+                s->base.is_jmp = DISAS_UPDATE;
                 gen_io_start();
             }
             gen_helper_v7m_preserve_fp_state(cpu_env);
-            if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
-                gen_io_end();
-            }
             /*
              * If the preserve_fp_state helper doesn't throw an exception
              * then it will clear LSPACT; we don't need to repeat this for
diff --git a/target/arm/translate.c b/target/arm/translate.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate.c
+++ b/target/arm/translate.c
@@ -XXX,XX +XXX,XX @@ static bool do_ldm(DisasContext *s, arg_ldst_block *a, int min_n)
             gen_io_start();
         }
         gen_helper_cpsr_write_eret(cpu_env, tmp);
-        if (tb_cflags(s->base.tb) & CF_USE_ICOUNT) {
-            gen_io_end();
-        }
         tcg_temp_free_i32(tmp);
         /* Must exit loop to check un-masked IRQs */
         s->base.is_jmp = DISAS_EXIT;
-- 
2.20.1

In commit cfdb2c0c95ae9205b0 ("target/arm: Vectorize SABA/UABA") we
replaced the old handling of SABA/UABA with a vectorized implementation
which returns early rather than falling into the loop-ever-elements
code. We forgot to delete the part of the old looping code that
did the accumulate step, and Coverity correctly warns (CID 1428955)
that this code is now dead. Delete it.

Fixes: cfdb2c0c95ae9205b0
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
Reviewed-by: Philippe Mathieu-Daudé <f4bug@amsat.org>
Reviewed-by: Richard Henderson <richard.henderson@linaro.org>
Message-id: 20200619171547.29780-1-peter.maydell@linaro.org
---
 target/arm/translate-a64.c | 12 ------------
 1 file changed, 12 deletions(-)

diff --git a/target/arm/translate-a64.c b/target/arm/translate-a64.c
index XXXXXXX..XXXXXXX 100644
--- a/target/arm/translate-a64.c
+++ b/target/arm/translate-a64.c
@@ -XXX,XX +XXX,XX @@ static void disas_simd_3same_int(DisasContext *s, uint32_t insn)
                 genfn(tcg_res, tcg_op1, tcg_op2);
             }
 
-            if (opcode == 0xf) {
-                /* SABA, UABA: accumulating ops */
-                static NeonGenTwoOpFn * const fns[3] = {
-                    gen_helper_neon_add_u8,
-                    gen_helper_neon_add_u16,
-                    tcg_gen_add_i32,
-                };
-
-                read_vec_element_i32(s, tcg_op1, rd, pass, MO_32);
-                fns[size](tcg_res, tcg_op1, tcg_res);
-            }
-
             write_vec_element_i32(s, tcg_res, rd, pass, MO_32);
 
             tcg_temp_free_i32(tcg_res);
-- 
2.20.1