From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pg1-f202.google.com (mail-pg1-f202.google.com [209.85.215.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01660376A1A for ; Fri, 10 Jul 2026 22:06:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.202 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721215; cv=none; b=Wdb8mz4PRF6rokY1UwZM7f3z4hSSkgjKFzZfQtV800E60YMiSEI58LpHOaHHbfhfS25RcvEq4s/3I5zBQHvxZH4KOHbSlS1mOTkF+M1qT321VoBUgmmofhQJPR2P/j2t33tcj7NjulWcUoukFbP/ryogdPVty5JzoF9HUCkF6NE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721215; c=relaxed/simple; bh=jsVQzkioRSLN2d9kRY421N6xZ4da3YTmA54HIOg25AM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=W6suFlkNyD8zLK8ELQbgHEq2zC2sreXXRbIh6o5F5CmM72e/H3GwrPT36cOftoukf8+wrkeLVOz8HMHm6ShgJqrYP34OVNQ+u7sEyqQH5E8247HmVebiVjfYOxHK7BhmNWH0ywcYEQ0itwCIICaT1ptfrarzsCWkFcH4OzgO9LQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=I8k7J5qc; arc=none smtp.client-ip=209.85.215.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="I8k7J5qc" Received: by mail-pg1-f202.google.com with SMTP id 41be03b00d2f7-c889d1eebafso329373a12.0 for ; Fri, 10 Jul 2026 15:06:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721210; x=1784326010; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=O5+xgnrzOMBdXqacb4txyXlwUxflXxozadFsGyCU5lw=; b=I8k7J5qcq6qbXdDfxS3e4qwEgnqhMZG7GXpdShPwhHAxQTCCyO92cBePUebBAgLaXC WhzJc5dMf2bEBVjjZL+1EbmZKv7+xvB+lzgGO9nc9X5k54mpBBPvOcFKZBmfOkwgl+8I ZfAkBXKNzvwg30YOfGQPv3aokglGW5lohtoyif7Q7Op6KaZjWXnrrLSvAQJ0oRk7H/F/ 6l3icFIjB50ol1zsorYEZKbqsfl8uJ6G/aDNG79o2EPv2bLlfrE04pxSPz7PfsSSlTy6 7WzCuyP5k+PjPkTZS7NVzMEeCCOeFcuKmUw5Ojh2BaTgffu7HGcDSgNlq0HXHU3ZVrTD XrvA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721210; x=1784326010; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=O5+xgnrzOMBdXqacb4txyXlwUxflXxozadFsGyCU5lw=; b=TH8AclUaAHTuRczj8k9Fv2jBiirLWDpIyyTqzarRMjcjYfWWrbW8I/F7Gab5jn1j3Q RVVA/+jAwnSY22Bt43jbhzHMUIHNgk1z9QHxNYXGykkDAlegSEiD81AqewLf18HmcN9w 2gFxSi+ZnSkLkzRDeJvDnBiUUOLVs8l9AzM4pua7BtbKGWxo+RNmngoSCXWD790V8quW WJzdgETpIpRzIyLowzcHS5qBb3rDVi/Aw2uTrvDSp018tUGH+r8cbasMbNjF7xHy3ZfR oyXQIppYkVUN5K772TdIMWm5heGtsDiJkRibKh8T3VJUKUWg/FLToLyzIai+Kw/Wr2W/ VLBg== X-Forwarded-Encrypted: i=1; AHgh+RoI747UbYcBPPnYTQbOtb6naOcSbvkPaHwOo5ekrPvwny7z7TEHO3KKth5d7UFoFPbUff2LZNIC/8MhN2g=@vger.kernel.org X-Gm-Message-State: AOJu0YxCJkl2F9z2OyHlLC3Dh5TiKvNHPyRcn5/MAt7VxNFSDXxjwfv4 LTWiwHjBs51Q99U7qBxk6EfkFaj5h999uAcOh6ZLY+C9d0iZ2klah0fLaOhdpM8bvEN/RdKipCg Bt2VvHS3t X-Received: from pghx11.prod.google.com ([2002:a63:f70b:0:b0:c82:7761:9936]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:6817:b0:3c0:9c1b:d0ba with SMTP id adf61e73a8af0-3c110a12332mr820314637.69.1783721209572; Fri, 10 Jul 2026 15:06:49 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:45 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=16242; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=jsVQzkioRSLN2d9kRY421N6xZ4da3YTmA54HIOg25AM=; b=+4ZKpbbOkexXOvJRuPv2c1mN+StKLlAWSj/RcOGgxi0HyniYrcsvEUp4oPa/vdajkWLrjSZk5 /EbJe3yBfowBBGBStP7LXOY5c3Rf6/edP29iWsZ5GJx++aep/HLHs0S X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-1-56e7e2576cc1@google.com> Subject: [PATCH v4 1/9] vfio: selftests: igb: Add driver for IGB QEMU device From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Add a VFIO selftest driver for the Intel Gigabit Ethernet controller (IGB). IGB is fully virtualized in QEMU [1] which makes it easy to run VFIO selftests without needing any specific hardware. This initial driver only implements the bare minimum configuration required to run successfully under QEMU's emulated IGB device, and does not support physical hardware. Subsequent patches add support for the physical device. IGB does not have a default memcpy operation, and memcpy is required for all VFIO selftest drivers. However, IGB has a "loopback" mode which is used in this driver to implement memcpy. When IGB is in loopback mode, it doesn't send data out to the network. Instead, it sends data from the Tx queue directly to the Rx queue which points to the memcpy destination, instead of sending the data out to the network. This driver passes all of the VFIO selftests in tools/testing/selftests/vfio/ when running in QEMU. The driver has not been tested using a real IGB device since the main goal of writing this driver is to run VFIO selftests in QEMU without requiring any hardware. This command is used to test the driver. It runs ./tools/testing/selftests/vfio/vfio_pci_driver_test using virtme-ng [2], which runs a kernel in a virtualized environment using QEMU that has a copy-on-write snapshot the host filesystem. vng \ --run arch/x86/boot/bzImage \ --user root \ --disable-microvm \ --memory 32G \ --cpus 8 \ --qemu-opts=3D"-M q35,accel=3Dkvm,kernel-irqchip=3Dsplit" \ --qemu-opts=3D"-device intel-iommu,intremap=3Don,caching-mode=3Don,device= -iotlb=3Don" \ --qemu-opts=3D"-netdev user,id=3Dnet0 -device igb,netdev=3Dnet0,addr=3D09= .0" \ --append "console=3DttyS0 earlyprintk=3DttyS0 intel_iommu=3Don iommu=3Dpt= " \ --exec "modprobe vfio-pci && \ ./tools/testing/selftests/vfio/scripts/setup.sh 0000:00:09.0 && \ ./tools/testing/selftests/vfio/scripts/run.sh ./tools/testing/sel= ftests/vfio/vfio_pci_driver_test" Code was written entirely by AI (Gemini) by feeding it the Intel IGB specification, the code for the real IGB driver, and the QEMU implementation of the IGB device. It took many iterations of prompting to get the driver to pass all of the tests, and lot's of de-slopping to remove unnecessary code and make the code readable. Code comments are also written by Gemini through iterative prompting. This patch is based on the kvm/queue branch. [1] https://www.qemu.org/docs/master/system/devices/igb.html [2] https://github.com/arighi/virtme-ng Assisted-by: Gemini:gemini-3-pro-preview Signed-off-by: Josh Hilke --- .../selftests/vfio/lib/drivers/igb/e1000_82575.h | 1 + .../selftests/vfio/lib/drivers/igb/e1000_defines.h | 1 + .../selftests/vfio/lib/drivers/igb/e1000_regs.h | 1 + tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 381 +++++++++++++++++= ++++ tools/testing/selftests/vfio/lib/libvfio.mk | 1 + tools/testing/selftests/vfio/lib/vfio_pci_driver.c | 2 + 6 files changed, 387 insertions(+) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h b/t= ools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h new file mode 120000 index 000000000000..b84affdec559 --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_82575.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h b= /tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h new file mode 120000 index 000000000000..9f97f4330086 --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_defines.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h b/to= ols/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h new file mode 120000 index 000000000000..c733634171bb --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_regs.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c new file mode 100644 index 000000000000..923b2341abad --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -0,0 +1,381 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "e1000_regs.h" +#include "e1000_defines.h" +#include "e1000_82575.h" + +#define PCI_DEVICE_ID_INTEL_82576 0x10C9 +#define IGB_MAX_CHUNK_SIZE 1024 +#define MSIX_VECTOR 0 +#define RING_SIZE 4096 /* Number of descriptors in ring */ + +struct igb_tx_desc { + union { + struct { + u64 buffer_addr; /* Address of descriptor's data buffer */ + u32 cmd_type_len; /* Command/Type/Length */ + u32 olinfo_status; /* Context/Buffer info */ + } read; + + struct { + u64 rsvd; /* Reserved */ + u32 nxtseq_seed; /* Next sequence seed */ + u32 status; /* Descriptor status */ + } wb; + }; +}; + +struct igb_rx_desc { + union { + struct { + u64 pkt_addr; /* Packet buffer address */ + u64 hdr_addr; /* Header buffer address */ + } read; + struct { + u16 pkt_info; /* RSS type, Packet type */ + u16 hdr_info; /* Split Head, buf len */ + u32 rss; /* RSS Hash */ + u32 status_error; /* ext status/error */ + u16 length; /* Packet length */ + u16 vlan; /* VLAN tag */ + } wb; /* writeback */ + }; +}; + +struct igb { + void *bar0; + u32 tx_tail; + u32 rx_tail; + struct igb_tx_desc tx_ring[RING_SIZE] __attribute__((aligned(128))); + struct igb_rx_desc rx_ring[RING_SIZE] __attribute__((aligned(128))); +}; + +static inline struct igb *to_igb_state(struct vfio_pci_device *device) +{ + return (struct igb *)device->driver.region.vaddr; +} + +static inline void igb_write32(struct igb *igb, u32 reg, u32 val) +{ + writel(val, igb->bar0 + reg); +} + +static inline u32 igb_read32(struct igb *igb, u32 reg) +{ + return readl(igb->bar0 + reg); +} + +static int igb_write_phy(struct igb *igb, u32 offset, u16 data) +{ + u32 mdic; + int i; + + /* + * Write a PHY register over MDIO. + * + * A production driver would hold the SW/FW semaphore (SWSM.SWESMBI + the + * SW_FW_SYNC PHY bit) across the MDIO transaction to serialize against t= he + * device's management firmware. The selftest owns the assigned function + * exclusively on a dedicated test device with no active manageability + * contending for the PHY, so the sync is omitted; it should be added here + * if this ever needs to run on a manageability-enabled NIC. + */ + mdic =3D (((u32)data) | + (offset << E1000_MDIC_REG_SHIFT) | + (1 << E1000_MDIC_PHY_SHIFT) | + E1000_MDIC_OP_WRITE); + + igb_write32(igb, E1000_MDIC, mdic); + + for (i =3D 0; i < 1000; i++) { + usleep(50); + mdic =3D igb_read32(igb, E1000_MDIC); + if (mdic & E1000_MDIC_READY) + break; + } + + if (!(mdic & E1000_MDIC_READY)) + return -1; + + if (mdic & E1000_MDIC_ERROR) + return -1; + + return 0; +} + +static int igb_read_phy(struct igb *igb, u32 offset, u16 *data) +{ + u32 mdic; + int i; + + mdic =3D ((offset << E1000_MDIC_REG_SHIFT) | + (1 << E1000_MDIC_PHY_SHIFT) | + E1000_MDIC_OP_READ); + + igb_write32(igb, E1000_MDIC, mdic); + + for (i =3D 0; i < 1000; i++) { + usleep(50); + mdic =3D igb_read32(igb, E1000_MDIC); + if (mdic & E1000_MDIC_READY) + break; + } + + if (!(mdic & E1000_MDIC_READY)) + return -1; + + if (mdic & E1000_MDIC_ERROR) + return -1; + + *data =3D (u16)mdic; + return 0; +} + +static void igb_phy_setup_autoneg(struct igb *igb) +{ + int timeout_ms =3D 1000; + bool success =3D false; + u16 phy_status; + int ret; + int i; + + /* Trigger auto-negotiation */ + ret =3D igb_write_phy(igb, MII_BMCR, + BMCR_ANENABLE | BMCR_ANRESTART); + VFIO_ASSERT_EQ(ret, 0, "Failed to write PHY control register"); + + for (i =3D 0; i < timeout_ms; i++) { + if (igb_read_phy(igb, MII_BMSR, &phy_status) =3D=3D 0) { + success =3D !!(phy_status & BMSR_ANEGCOMPLETE); + if (success) + break; + } + usleep(1000); + } + + VFIO_ASSERT_TRUE(success, "Auto-negotiation did not complete in time"); +} + +static int igb_probe(struct vfio_pci_device *device) +{ + if (!vfio_pci_device_match(device, PCI_VENDOR_ID_INTEL, PCI_DEVICE_ID_INT= EL_82576)) + return -EINVAL; + + return 0; +} + +static void igb_init(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + u64 iova_tx, iova_rx; + u32 ctrl, rctl; + u16 cmd_reg; + int retries; + + VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); + + /* Set up rings and calculate IOVAs */ + igb->bar0 =3D device->bars[0].vaddr; + + iova_tx =3D to_iova(device, igb->tx_ring); + iova_rx =3D to_iova(device, igb->rx_ring); + + igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); + /* + * Must wait at least 1 millisecond after setting the reset bit before + * checking if this device is ready to be used (82576 datasheet section + * 4.2.1.6.1). + */ + usleep(1000); + while (igb_read32(igb, E1000_CTRL) & E1000_CTRL_RST) + usleep(10); + igb_write32(igb, E1000_IMC, 0xFFFFFFFF); + + /* Signal that the driver is loaded */ + ctrl =3D igb_read32(igb, E1000_CTRL_EXT); + ctrl |=3D E1000_CTRL_EXT_DRV_LOAD; + ctrl &=3D ~E1000_CTRL_EXT_LINK_MODE_MASK; + igb_write32(igb, E1000_CTRL_EXT, ctrl); + + /* Enable PCI Bus Master. */ + cmd_reg =3D vfio_pci_config_readw(device, PCI_COMMAND); + if ((cmd_reg & (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY)) !=3D + (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY)) { + cmd_reg |=3D (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY); + vfio_pci_config_writew(device, PCI_COMMAND, cmd_reg); + } + + /* Trigger autonegotiation. This enables IGB to transmit data. */ + igb_phy_setup_autoneg(igb); + + /* Configure TX and RX descriptor rings */ + igb_write32(igb, E1000_TDBAL(0), (u32)iova_tx); + igb_write32(igb, E1000_TDBAH(0), (u32)(iova_tx >> 32)); + igb_write32(igb, E1000_TDLEN(0), RING_SIZE * sizeof(struct igb_tx_desc)); + igb_write32(igb, E1000_TDH(0), 0); + igb_write32(igb, E1000_TDT(0), 0); + igb_write32(igb, E1000_TXDCTL(0), E1000_TXDCTL_QUEUE_ENABLE); + + igb_write32(igb, E1000_RDBAL(0), (u32)iova_rx); + igb_write32(igb, E1000_RDBAH(0), (u32)(iova_rx >> 32)); + igb_write32(igb, E1000_RDLEN(0), RING_SIZE * sizeof(struct igb_rx_desc)); + igb_write32(igb, E1000_RDH(0), 0); + igb_write32(igb, E1000_RDT(0), 0); + igb_write32(igb, E1000_RXDCTL(0), E1000_RXDCTL_QUEUE_ENABLE); + + /* Wait for TX and RX queues to be enabled */ + retries =3D 2000; + while (retries-- > 0) { + if ((igb_read32(igb, E1000_TXDCTL(0)) & E1000_TXDCTL_QUEUE_ENABLE) && + (igb_read32(igb, E1000_RXDCTL(0)) & E1000_RXDCTL_QUEUE_ENABLE)) + break; + usleep(10); + } + + /* Enable Receiver and Transmitter */ + rctl =3D E1000_RCTL_EN | /* Receiver Enable */ + E1000_RCTL_UPE | /* Unicast Promiscuous (for dummy MAC) */ + E1000_RCTL_LBM_MAC | /* MAC Loopback Mode */ + E1000_RCTL_SECRC; /* Strip CRC (needed for memcmp) */ + igb_write32(igb, E1000_RCTL, rctl); + igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); + + /* Enable MSI-X with 1 vector for the test */ + vfio_pci_msix_enable(device, MSIX_VECTOR, 1); + + /* Enable auto-masking of interrupts to avoid storms without a real ISR */ + igb_write32(igb, E1000_GPIE, E1000_GPIE_EIAME | E1000_GPIE_MSIX_MODE); + + /* Map vector 0 to interrupt cause 0 and mark it valid */ + igb_write32(igb, E1000_IVAR0, E1000_IVAR_VALID); + + /* Enable interrupts on vector 0 */ + igb_write32(igb, E1000_EIMS, 1); + + /* Initialize driver state and capability limits */ + igb->tx_tail =3D 0; + igb->rx_tail =3D 0; + + device->driver.max_memcpy_size =3D IGB_MAX_CHUNK_SIZE; + device->driver.max_memcpy_count =3D RING_SIZE - 1; + device->driver.msi =3D MSIX_VECTOR; +} + +static void igb_remove(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + vfio_pci_msix_disable(device); + igb_write32(igb, E1000_RCTL, 0); + igb_write32(igb, E1000_TCTL, 0); + igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); +} + +static void igb_irq_disable(struct igb *igb) +{ + igb_write32(igb, E1000_EIMC, 1); +} + +static void igb_irq_enable(struct igb *igb) +{ + igb_write32(igb, E1000_EIMS, 1); +} + +static void igb_irq_clear(struct igb *igb) +{ + igb_read32(igb, E1000_EICR); +} + +static void igb_memcpy_start(struct vfio_pci_device *device, iova_t src, + iova_t dst, u64 size, u64 count) +{ + struct igb *igb =3D to_igb_state(device); + struct igb_rx_desc *rx; + struct igb_tx_desc *tx; + u32 i; + + igb_irq_disable(igb); + + for (i =3D 0; i < count; i++) { + tx =3D &igb->tx_ring[igb->tx_tail]; + rx =3D &igb->rx_ring[igb->rx_tail]; + + memset(tx, 0, sizeof(struct igb_tx_desc)); + memset(rx, 0, sizeof(struct igb_rx_desc)); + + rx->read.pkt_addr =3D cpu_to_le64(dst); + tx->read.buffer_addr =3D cpu_to_le64(src); + tx->read.cmd_type_len =3D cpu_to_le32((u32)size | E1000_TXD_CMD_EOP | E1= 000_TXD_CMD_IFCS); + + /* Set to 0 to disable offloads and avoid needing a context descriptor */ + tx->read.olinfo_status =3D cpu_to_le32(0); + + igb->tx_tail =3D (igb->tx_tail + 1) % RING_SIZE; + igb->rx_tail =3D (igb->rx_tail + 1) % RING_SIZE; + } + + igb_write32(igb, E1000_RDT(0), igb->rx_tail); + igb_write32(igb, E1000_TDT(0), igb->tx_tail); +} + +static int igb_memcpy_wait(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + struct igb_rx_desc *rx; + u32 status =3D 0; + u32 prev_tail; + int retries; + + prev_tail =3D (igb->rx_tail + RING_SIZE - 1) % RING_SIZE; + rx =3D &igb->rx_ring[prev_tail]; + + retries =3D 100; + while (retries-- > 0) { + status =3D le32_to_cpu(READ_ONCE(rx->wb.status_error)); + if (status & 1) + break; + usleep(10); + } + + if (status & 1) + /* + * Ensure the test code doesn't speculatively read the DMA + * destination buffer before we have verified that the + * descriptor writeback is complete. + */ + rmb(); + + igb_irq_clear(igb); + + igb_irq_enable(igb); + + return (status & 1) ? 0 : -ETIMEDOUT; +} + +static void igb_send_msi(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + igb_write32(igb, E1000_EICS, 1); +} + +const struct vfio_pci_driver_ops igb_ops =3D { + .name =3D "igb", + .probe =3D igb_probe, + .init =3D igb_init, + .remove =3D igb_remove, + .memcpy_start =3D igb_memcpy_start, + .memcpy_wait =3D igb_memcpy_wait, + .send_msi =3D igb_send_msi, +}; diff --git a/tools/testing/selftests/vfio/lib/libvfio.mk b/tools/testing/se= lftests/vfio/lib/libvfio.mk index 9f47bceed16f..1f13cca04348 100644 --- a/tools/testing/selftests/vfio/lib/libvfio.mk +++ b/tools/testing/selftests/vfio/lib/libvfio.mk @@ -12,6 +12,7 @@ LIBVFIO_C +=3D vfio_pci_driver.c ifeq ($(ARCH:x86_64=3Dx86),x86) LIBVFIO_C +=3D drivers/ioat/ioat.c LIBVFIO_C +=3D drivers/dsa/dsa.c +LIBVFIO_C +=3D drivers/igb/igb.c endif =20 LIBVFIO_OUTPUT :=3D $(OUTPUT)/libvfio diff --git a/tools/testing/selftests/vfio/lib/vfio_pci_driver.c b/tools/tes= ting/selftests/vfio/lib/vfio_pci_driver.c index 6827f4a6febe..a5d0547132c4 100644 --- a/tools/testing/selftests/vfio/lib/vfio_pci_driver.c +++ b/tools/testing/selftests/vfio/lib/vfio_pci_driver.c @@ -5,12 +5,14 @@ #ifdef __x86_64__ extern struct vfio_pci_driver_ops dsa_ops; extern struct vfio_pci_driver_ops ioat_ops; +extern struct vfio_pci_driver_ops igb_ops; #endif =20 static struct vfio_pci_driver_ops *driver_ops[] =3D { #ifdef __x86_64__ &dsa_ops, &ioat_ops, + &igb_ops, #endif }; =20 --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pf1-f201.google.com (mail-pf1-f201.google.com [209.85.210.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D4C113CCFDE for ; Fri, 10 Jul 2026 22:06:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721216; cv=none; b=mQYDcc8Prm3ZvDGqKUXGnL5WpaQfFWKVm0/jsb0C9lViNtBBVyeV7cwsEyRuEmGmNmeoaoe3LtI5//XBGcZYdImg2W6JqiNmLRRlAWJv33rgrIyRlDPBV8wWSWu6Q2KmvrUe1MqAfZywHPXyf6FkQbLZ+VOBF3w3g4rQ6M8j0kY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721216; c=relaxed/simple; bh=7JOS48I3NEUU/p6/YNWCZsTgsMDs9hkv7k7Q3FzIzig=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=h0rZvwVzZxc81aOx+yxAji1JLqVdx1DtNvCHPw4944SMU9XJfFjPCXIptLIk6xB0hxw6A0qyojR++UfQ+/T7kF3SaOy70YAOXOmNDpKo+Ag1cKrzJT47BmYWXLUKwcP0eh8HTozo5khkJwRSCEd96t26/49I5Ukpgb6e94+4ckQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Iq9FQ8mG; arc=none smtp.client-ip=209.85.210.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Iq9FQ8mG" Received: by mail-pf1-f201.google.com with SMTP id d2e1a72fcca58-8483e038efeso1394810b3a.2 for ; Fri, 10 Jul 2026 15:06:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721210; x=1784326010; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SWTiX2PkmTKlxw5Cf9kZZ45LEhjcTBibm2DK4Kwk3HM=; b=Iq9FQ8mGbnWDS+8a0w8B6sexKPcJT+dCKGTAm11NsI+3cmzyBxNnBieYy6Mv2+QcWp vBP4lgvhm2wRSDAK2+ylkjm1adD0kVivAJ7TJWG+CmbwWVNdNZ69f/WM9/he9PuGqAjC DR32uGn8LenW6DY/gj5ySwcN/ocZ+/uNS0jlebpYFBY13uaJafBwLWnPRChWelF25yB5 vtB2qGwUDRehUJXaho3hh3ZqpGyvOwAITCPCyGEC39FZWM52bRzOjVGdOIm1pExQCdDE 31e0B9oeK3KzOQhwe1zrR7EiFQHSuSAhmvGJtwyGmINfwSb68ZHKMOpL1qQYcdqfmXD1 F0pw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721210; x=1784326010; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SWTiX2PkmTKlxw5Cf9kZZ45LEhjcTBibm2DK4Kwk3HM=; b=nCy1paDkaOsES0aUNU1/TDjYhzUoRIDE8jkujHpDAOxW8FXusnfcCO69w3I2VHEy// GshE6t3/aPYhtM/Mj506QwvsaByoqSp+0yUMliq1xSKZst8+YrcM2HNLtKKIzIqwuDR2 tSjuocnFlV0kHSE5gSSkMcg2BiUGlUYE0iHNDH+ljfMFb5DPltdsv0QlQ7abEZ/ojpWj 1jSrld6c4hEaSiOt0NFgiOvAgrZMGjy/FSQiJM3fEvaVwJJXl+um4P5nREeABfwtubZk zHh+chiDSCo+lHixAknl/kG3TuWbBCKlTEzJZguFKUJniY8lkIHCOYSp5/PQch25S2Ho 5zPg== X-Forwarded-Encrypted: i=1; AHgh+RobBp1bFtsmYXZAV7Q7Og99P1j7nUFqW43aVU3sY6VGPFkEWMYiOPEz3Wp3uJD+Cwwu5Db+a9jG7PVtVV0=@vger.kernel.org X-Gm-Message-State: AOJu0YwR/c96KrBWnWQVcP9T/BPqG3qujKoW1Kf8u6lvRvFf106FGM6S 2Y4m7MoGbnTdcoKx3XHhRkn7tCU2KjSlt7chvpUxPUNoGDm8uBYEu22/ItgdaT9cghdJF65qVe1 0yYI6tOqo X-Received: from pfbln8.prod.google.com ([2002:a05:6a00:3cc8:b0:848:48a0:41e6]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:6c8f:b0:82d:62ed:b01d with SMTP id d2e1a72fcca58-848896bb05emr723565b3a.45.1783721210396; Fri, 10 Jul 2026 15:06:50 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:46 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=8817; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=SIgpWdDVQxVOygHfKVKsfIRiQ88wk46H3v6Ry1SUnxo=; b=d3VVWv2XY5UgtBcn4wpVRWVr5NuJKHrYR0LN7oFtQKV9q/7X5kztyQ2bHy5pqNXSmWaqk5v/e 8a2AvfpOOBhBWKsx8eU+2zSTy4Qwk8DsHP5IU4sFzrwKCNc42TExD3E X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-2-56e7e2576cc1@google.com> Subject: [PATCH v4 2/9] vfio: selftests: igb: Use PHY internal loopback on 82576 From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The submitted driver waits for PHY autonegotiation and then enables MAC loopback via RCTL.LBM_MAC. QEMU's emulated igb tolerates this, but the 82576 datasheet rejects the MAC loopback path on real hardware: Section 3.5.6.1 (Loopback Support / General): "Use PHY Loopback instead of MAC Loopback on the 82576." Section 3.5.6.2 (MAC Loopback): "MAC Loopback is not used on this device." Section 3.5.6.3.1 (Setting the 82576 to PHY loopback Mode): set PHY control register bits 8 (duplex), 14 (loopback), clear bit 12 (autoneg enable), set the speed via bits 6 and 13. For 1Gb/s the register value is 0x4140. Section 8.10.1 (RCTL register): "When using the internal PHY, LBM should remain set to 00b and the PHY instead configured for loopback through the MDIO interface." Replace igb_phy_setup_autoneg() with igb_setup_loopback() which: - writes PHY register 0 with LOOPBACK | SPEED_1000 | FULL_DUPLEX - forces the MAC into 1Gb/s full duplex via CTRL.FRCSPD, CTRL.FRCDPX, CTRL.SPD_1000, CTRL.FD, CTRL.SLU; without forcing the MAC link state, the descriptor engine does not run on real hardware PHY internal loopback (section 3.5.6.3) wraps data at the end of the PHY datapath before the MDI, so the physical link state and cable speed are irrelevant. This matches the kernel ethtool selftest path in igb_integrated_phy_loopback() (drivers/net/ethernet/intel/igb/igb_ethtool.c). QEMU's igb emulation drives STATUS.LU exclusively from its autoneg-done timer or a network-backend link-state change; the guest cannot set STATUS.LU through CTRL.SLU. Its receive path checks STATUS.LU (e1000x_hw_rx_enabled in hw/net/e1000x_common.c) and drops every loopback frame until LU is set. Issue a one-shot autoneg-restart PHY write at the top of igb_setup_loopback() to kick the timer; the subsequent PHY write clears autoneg-enable, so on real hardware autoneg never starts and the write is a no-op. QEMU's igb also does not honor PHY register 0 bit 14 (PHY internal loopback) and relies on RCTL.LBM_MAC to wrap TX descriptors back to the RX queue. Datasheet 8.10.1 advises that LBM remain 00b when using the internal PHY, but empirically setting LBM_MAC has no observable effect on real 82576 (MAC loopback is not implemented per 3.5.6.2), so set it alongside PHY loopback as the actual loopback mechanism under QEMU. With these two QEMU-only accommodations the selftest works in both environments without environment-specific code paths. Remove igb_read_phy() which becomes unused. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 122 +++++++++++++----= ---- 1 file changed, 74 insertions(+), 48 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 923b2341abad..cccaedf86e32 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -114,57 +114,70 @@ static int igb_write_phy(struct igb *igb, u32 offset,= u16 data) return 0; } =20 -static int igb_read_phy(struct igb *igb, u32 offset, u16 *data) +/* + * Configure the device for PHY internal loopback per 82576 datasheet + * section 3.5.6.3.1. Force the PHY to 1Gb/s full duplex with loopback + * enabled, then force the MAC link state to match. Internal loopback + * wraps data at the end of the PHY datapath (section 3.5.6.3), so the + * physical link state is irrelevant. + * + * Section 3.5.6.1 directs to "Use PHY Loopback instead of MAC Loopback + * on the 82576", and section 3.5.6.2 states "MAC Loopback is not used + * on this device." RCTL.LBM_MAC is still set elsewhere as a QEMU-only + * accommodation; see the RCTL programming in the caller for the + * rationale. + */ +static void igb_setup_loopback(struct igb *igb) { - u32 mdic; - int i; - - mdic =3D ((offset << E1000_MDIC_REG_SHIFT) | - (1 << E1000_MDIC_PHY_SHIFT) | - E1000_MDIC_OP_READ); - - igb_write32(igb, E1000_MDIC, mdic); - - for (i =3D 0; i < 1000; i++) { - usleep(50); - mdic =3D igb_read32(igb, E1000_MDIC); - if (mdic & E1000_MDIC_READY) - break; - } - - if (!(mdic & E1000_MDIC_READY)) - return -1; - - if (mdic & E1000_MDIC_ERROR) - return -1; - - *data =3D (u16)mdic; - return 0; -} - -static void igb_phy_setup_autoneg(struct igb *igb) -{ - int timeout_ms =3D 1000; - bool success =3D false; - u16 phy_status; + u32 ctrl; int ret; - int i; =20 - /* Trigger auto-negotiation */ - ret =3D igb_write_phy(igb, MII_BMCR, + /* + * Kick the autoneg machinery solely to bring STATUS.LU up under + * QEMU's igb emulation: QEMU only updates STATUS.LU via its + * autoneg-done timer, and without LU set its receive path + * (e1000x_hw_rx_enabled) drops every loopback frame. On real + * hardware autoneg cannot complete before the next PHY write + * below clears the autoneg-enable bit, so this is effectively a + * no-op there. + */ + (void)igb_write_phy(igb, MII_BMCR, BMCR_ANENABLE | BMCR_ANRESTART); + + /* PHY control: loopback + 1Gb/s full duplex, autoneg disabled. */ + ret =3D igb_write_phy(igb, MII_BMCR, + BMCR_LOOPBACK | + BMCR_SPEED1000 | + BMCR_FULLDPLX); VFIO_ASSERT_EQ(ret, 0, "Failed to write PHY control register"); =20 - for (i =3D 0; i < timeout_ms; i++) { - if (igb_read_phy(igb, MII_BMSR, &phy_status) =3D=3D 0) { - success =3D !!(phy_status & BMSR_ANEGCOMPLETE); - if (success) - break; - } - usleep(1000); - } + /* + * Brief delay before forcing the MAC, mirroring the kernel ethtool + * selftest in igb_integrated_phy_loopback(). Not specified by the + * datasheet, but empirically required by the kernel driver. + */ + usleep(50000); + + /* + * Force the MAC to 1Gb/s full duplex with link up. Without forcing + * the link state the descriptor engine does not run, since the chip + * normally waits for a real negotiated link. + */ + ctrl =3D igb_read32(igb, E1000_CTRL); + ctrl &=3D ~E1000_CTRL_SPD_SEL; + ctrl |=3D E1000_CTRL_FRCSPD | + E1000_CTRL_FRCDPX | + E1000_CTRL_SPD_1000 | + E1000_CTRL_FD | + E1000_CTRL_SLU; + igb_write32(igb, E1000_CTRL, ctrl); =20 - VFIO_ASSERT_TRUE(success, "Auto-negotiation did not complete in time"); + /* + * Settling delay matching the kernel ethtool selftest's msleep(500) + * at the tail of igb_integrated_phy_loopback(). Not specified by + * the datasheet; empirical, and inherited from the kernel driver. + */ + usleep(500000); } =20 static int igb_probe(struct vfio_pci_device *device) @@ -191,6 +204,7 @@ static void igb_init(struct vfio_pci_device *device) iova_tx =3D to_iova(device, igb->tx_ring); iova_rx =3D to_iova(device, igb->rx_ring); =20 + /* Reset device and disable all interrupts */ igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); /* * Must wait at least 1 millisecond after setting the reset bit before @@ -216,8 +230,8 @@ static void igb_init(struct vfio_pci_device *device) vfio_pci_config_writew(device, PCI_COMMAND, cmd_reg); } =20 - /* Trigger autonegotiation. This enables IGB to transmit data. */ - igb_phy_setup_autoneg(igb); + /* Configure PHY internal loopback for testing. */ + igb_setup_loopback(igb); =20 /* Configure TX and RX descriptor rings */ igb_write32(igb, E1000_TDBAL(0), (u32)iova_tx); @@ -243,10 +257,22 @@ static void igb_init(struct vfio_pci_device *device) usleep(10); } =20 - /* Enable Receiver and Transmitter */ + /* + * Enable Receiver and Transmitter. RCTL.LBM_MAC is set in addition + * to PHY loopback as a QEMU-only accommodation: QEMU's emulated igb + * does not honor PHY register 0 bit 14 (PHY internal loopback) and + * relies on RCTL.LBM_MAC to wrap TX descriptors back to the RX + * queue. Datasheet 8.10.1 (RCTL register) advises "When using the + * internal PHY, LBM should remain set to 00b", so setting LBM_MAC + * here deviates from datasheet guidance; empirically the bit has + * no observable effect on real 82576 hardware because MAC loopback + * is not implemented (datasheet 3.5.6.2). Setting both lets the + * selftest work on both real hardware and QEMU without conditional + * code paths. + */ rctl =3D E1000_RCTL_EN | /* Receiver Enable */ E1000_RCTL_UPE | /* Unicast Promiscuous (for dummy MAC) */ - E1000_RCTL_LBM_MAC | /* MAC Loopback Mode */ + E1000_RCTL_LBM_MAC | /* MAC Loopback - for QEMU emulation only */ E1000_RCTL_SECRC; /* Strip CRC (needed for memcmp) */ igb_write32(igb, E1000_RCTL, rctl); igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pl1-f201.google.com (mail-pl1-f201.google.com [209.85.214.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 624173CC314 for ; Fri, 10 Jul 2026 22:06:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721217; cv=none; b=o3a885yLfoG6x/dstybxvjSvdRKP3cKG4w/FV9sWsEHRPJBAYD/FCdzD6Qrj26euNzExLFQZytfq7BHNUjmcUsRP6y771cp5sKrguBtiPNbCI//4PpIkatLsk4FQER4spAtau+GBQP4ZXumnYcveZXlCoOKx5skDeYXEaDZTfrE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721217; c=relaxed/simple; bh=Rsta8vLB8hK9R4p5vcxEA3vQvmZfrYl7dCQRQrmn55g=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=avCqwBU2GtQZWD4wwre23BQCelAGQJ1xCc54rA0qsHq7qiwUOwzQBLaXnB5RfwU4/SskvoWFzNkFE7Em6tvAtV6CCXPt4otWK+jU0Sni8HACUxrg39w+jpIDemWsJP74ZrRjoAZA0IgkLkYsP7tA+xCNecIa8ysFMAr+joyW9y4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=QXyu9J44; arc=none smtp.client-ip=209.85.214.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="QXyu9J44" Received: by mail-pl1-f201.google.com with SMTP id d9443c01a7336-2c804e38c65so26680155ad.2 for ; Fri, 10 Jul 2026 15:06:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721211; x=1784326011; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=E1VBUyR4cccUA9/1mGNxW9MUPChjoTBoxCQDwJfrrNQ=; b=QXyu9J44fHlX0PE1NKW2KnlNCAmgpz83UjU2oA920k5nlzJTg5Glwa5JnMVL8Qbs34 HrGGGmALXjTtyYSY0eb5Y+fzLxLkuOdqgJTns1p0S35qzxjiFyK0tR5jqwB4zC7ULkmW oOyOVQ+gD+v6gMcgytV/JTmkaSbIdlhC8phrUXHWvNu5hzfddA87S4OtNARkoCBT8x3o r/3GvZ2lAway7coVbybIkTWXAwdZHPHYqovmY1W0HTnqJNh7L++bScQDvJ3qNP2upOT1 CZqj6WK8P5+dMEYTlbXHCGBC92S38HknoHcfb5ITojZ8RccTWVnPMpxzy6k4vKddcyIc Od/A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721211; x=1784326011; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=E1VBUyR4cccUA9/1mGNxW9MUPChjoTBoxCQDwJfrrNQ=; b=Ob8hunoiWsBz+HSZHV8kzfcRKtpiije0aHx2sCH4z9+zxlJxHYdN/pn+7Y98EKyXKz eR+0ttHaKPG/CsThcarv2oz6Z98g+2RRnm840Fv17LmmCLQG9py/Vxc64gv8egIWRLEw kMKLOI6hSQh7PWWLO7HTkvOtmw18knfsCv8R8sZHe/SV8d7c1xf5NHGyJpSqLOrJOKPa F+Z/YxzdGUI3X/HmwxO1gZuayOETKgnT9GmMeyLGxJuCxpfE789Gjsc1Xlib/n79fut2 kf2v+ptI6UY/K9x6m+nZWKMX7Nn+4BLiQuUL83mmlQLFhU7fMR5jIWgncPp1fUbdxaxH Jhbg== X-Forwarded-Encrypted: i=1; AHgh+RqGTYIQhnGzITpYl9o7iGktUvUFZj2F/9AQ4eZulsBOIFigLMVrjuQ3LnF4BTsTU8NH0JnYpZmSNZ8Hklg=@vger.kernel.org X-Gm-Message-State: AOJu0Yx/v3jZG6DIkZKh5IRM88mpiBmw0AenY6/vAHGNiUahmUrUOQq4 awHm/Tyv9mNiJ9kqBvX6yZQuD8c68nFKrVJwVe2XRNS0qxBvDsAAqnIyradjLH08SJc1cZwZGrM IZKhSDsTg X-Received: from plnm6.prod.google.com ([2002:a17:902:fda6:b0:2ca:b48c:5a92]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:2410:b0:2c9:e6d7:fbb4 with SMTP id d9443c01a7336-2ce9ee163e3mr7804965ad.31.1783721211262; Fri, 10 Jul 2026 15:06:51 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:47 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=4221; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=+v4iP4weLlHAsUthkQ6zLzOJqWm3vZ17kiKiO993AlY=; b=JMgox1Ej3OO54RInUmySHh8kP6O6ci8pddhJ52t3tQadEOkoVJkI/pX5MXcCqzgtSGqtlLYwK gqXG52xG7+ACKoFiR8QIgofX/LJtDhI7r9R3N66iMtOxigNlvXX5NL1 X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-3-56e7e2576cc1@google.com> Subject: [PATCH v4 3/9] vfio: selftests: igb: Use advanced TX and RX descriptors From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The submitted driver builds a partial legacy TX descriptor (just DTALEN | CMD_EOP) and never programs SRRCTL.DESCTYPE. QEMU's emulated igb tolerates this by treating descriptors as advanced regardless of DESCTYPE, but real 82576 hardware does not. For receive, 82576 datasheet section 7.1.5.2 states: "SRRCTL[n].DESCTYPE must be set to a value other than 000b for the 82576 to write back the special descriptors." struct igb_rx_desc matches the advanced one-buffer writeback layout, so the test polls rx.wb.status_error, which is only written in that layout. Section 8.10.2 places DESCTYPE in SRRCTL bits 27:25; program it with 001b (advanced one-buffer). For transmit, datasheet section 7.2.2.3 describes the advanced data descriptor with DEXT (DCMD bit 5) marking the descriptor as advanced, DTYP=3D0011b selecting the data descriptor, IFCS (DCMD bit 1) asking the MAC to append the Ethernet FCS (without it the frame is dropped as malformed), EOP (DCMD bit 0) marking end of packet, and PAYLEN in olinfo_status[31:14] carrying the total payload size. Build this descriptor in igb_memcpy_start(). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 42 ++++++++++++++++++= +--- 1 file changed, 38 insertions(+), 4 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index cccaedf86e32..033997ad45e0 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -246,6 +246,22 @@ static void igb_init(struct vfio_pci_device *device) igb_write32(igb, E1000_RDLEN(0), RING_SIZE * sizeof(struct igb_rx_desc)); igb_write32(igb, E1000_RDH(0), 0); igb_write32(igb, E1000_RDT(0), 0); + + /* + * Select the advanced one-buffer descriptor format. Per 82576 + * datasheet section 7.1.5.2: "SRRCTL[n].DESCTYPE must be set to a + * value other than 000b for the 82576 to write back the special + * descriptors." struct igb_rx_desc matches the advanced one-buffer + * writeback layout (section 7.1.5.2), so polling rx.wb.status_error + * requires this format. Section 8.10.2 specifies DESCTYPE[27:25]. + * + * The direct write also zeroes SRRCTL.BSIZEPACKET, which is + * intentional: per section 7.1.3.1 a zero BSIZEPACKET falls back to + * the RCTL.BSIZE buffer size, whose reset default (00b) is 2048 + * bytes -- ample for the loopback frames here. + */ + igb_write32(igb, E1000_SRRCTL(0), E1000_SRRCTL_DESCTYPE_ADV_ONEBUF); + igb_write32(igb, E1000_RXDCTL(0), E1000_RXDCTL_QUEUE_ENABLE); =20 /* Wait for TX and RX queues to be enabled */ @@ -341,11 +357,29 @@ static void igb_memcpy_start(struct vfio_pci_device *= device, iova_t src, memset(rx, 0, sizeof(struct igb_rx_desc)); =20 rx->read.pkt_addr =3D cpu_to_le64(dst); - tx->read.buffer_addr =3D cpu_to_le64(src); - tx->read.cmd_type_len =3D cpu_to_le32((u32)size | E1000_TXD_CMD_EOP | E1= 000_TXD_CMD_IFCS); + rx->read.hdr_addr =3D cpu_to_le64(0); =20 - /* Set to 0 to disable offloads and avoid needing a context descriptor */ - tx->read.olinfo_status =3D cpu_to_le32(0); + tx->read.buffer_addr =3D cpu_to_le64(src); + /* + * Build an advanced data descriptor per 82576 datasheet + * section 7.2.2.3. DEXT marks the descriptor as advanced + * (required by hardware); DTYP=3Ddata selects the data + * descriptor; IFCS asks the MAC to append the Ethernet + * FCS (without it the frame is dropped as malformed); + * EOP marks end of packet. DTALEN is the buffer length + * in bits 15:0 of cmd_type_len. + */ + tx->read.cmd_type_len =3D cpu_to_le32((uint32_t)size | + E1000_ADVTXD_DTYP_DATA | + E1000_ADVTXD_DCMD_DEXT | + E1000_ADVTXD_DCMD_IFCS | + E1000_ADVTXD_DCMD_EOP); + /* + * PAYLEN (section 7.2.2.3.11) is the total payload size + * in olinfo_status[31:14]. + */ + tx->read.olinfo_status =3D + cpu_to_le32((uint32_t)size << E1000_ADVTXD_PAYLEN_SHIFT); =20 igb->tx_tail =3D (igb->tx_tail + 1) % RING_SIZE; igb->rx_tail =3D (igb->rx_tail + 1) % RING_SIZE; --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pl1-f202.google.com (mail-pl1-f202.google.com [209.85.214.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6C2C53D301A for ; Fri, 10 Jul 2026 22:06:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.202 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721218; cv=none; b=aID5QjOKAqWcIEvuUUQqp8Y//0c71j37Ctc/w9VEllCVGPWHV++gkcJSsYpveiWAni+24NtPHTVnoFbAsNseBDwVmPcgskdw8k+R2w6t6zNPiDoN8jyn7wJcc7gtfnISdUxPPVhbwvipTl4FdiQXtHZWD7+8BQD472VeBNs7KKg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721218; c=relaxed/simple; bh=NagcJVpSawnrGB5aTa7laT6+AHk/7WT/b5cOVYFNLMs=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=V/y9XZlfas4WngEyV3YD5ls6dj3GqCn/6wkPOKca5c4auuCb6I4XHaCEq58Jr27iVEgzB6sC+e7Gkum9fzZssGBbwINe+sQ1Y9FTwyJvRdJsbagvp+wHWztaHKCScfXjoNRyavoCy6UCjHhcWYB0PbkR/zAwx+LFqS5yNq9+C20= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=hSa7kks2; arc=none smtp.client-ip=209.85.214.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="hSa7kks2" Received: by mail-pl1-f202.google.com with SMTP id d9443c01a7336-2cc77a6943eso31240195ad.0 for ; Fri, 10 Jul 2026 15:06:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721213; x=1784326013; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=xGZ8xblzsGAm9fHuvl78rZ7mmqIkbGhLb4jrriYJ9GM=; b=hSa7kks20OGMwc8jFuf1HMCBceTsS1TIo8NzC0JdCuAKLqhT/PPWs7A4vSOoPivfv0 oySp4eRlKBVkKqfD6jpXa0KtjnswShPzGAK5vs/E09yaQi6ZFxVHcgYBlqn0QCQB58ZX dQz4rV3YWe6Do2ERi6w213K5awV2Z17qXx5rvs9c9lXGpv8oj1xZHQCWkm0B4yXTAEdS yMyI3P+mxdFcwtnuydbXQYl033dBWiQ9Mxn+25P5Xi/ZeSNxqQ8qPc/B64Luidm51sBO oj3DPdA/9w3NC+C4WL6qh4YLa8KhgIXq3AqctkzAhS4j5jO+QVznebe2NPcJfSw8I3AH OuZw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721213; x=1784326013; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=xGZ8xblzsGAm9fHuvl78rZ7mmqIkbGhLb4jrriYJ9GM=; b=d4R0kBmfjCy9P7Cz/CYJ02Bl7ZzRA1JqzhrbIerEsE3CwBa6KSz8ODOrYkHVdIScLP KyuRchBhceXpdWY/0fh6G1/ixcB7tDT0EZJMfhxYRkFtPE5W3J50hMr3gl7yFuyUH4Wm 0/Lfqm2vNzXqA5XcP2UHhXURVhsWP5o2ey1bfcOUyUOlntgk1cyPNnORi9nEOXMUv4Gl 9iF85jeJZfYNyO8bp7i/OfY2dY2aqDCUM3vRMf+ulhDAg+vfCo6U9gpVJVjbnflyTxAz iv/+WSeHILIRxjLcIf0M0YcK/zL2SVFZb8P2ne/hF70rg7Ohdd6tdG5Yu1Q3kUfNsP71 ra2A== X-Forwarded-Encrypted: i=1; AHgh+RqPXqzgEqeMI9vVkw4Fk4dNQyVliz7dMOywnNaTWr3Xjvv3HA/cNiWsVmwJyQpzTh0Zjni7XN2KQL5T2vw=@vger.kernel.org X-Gm-Message-State: AOJu0YzDWezMt9j1E8SBbRnp7EzRC3CICqbKgbpiMr1VP9ZMJGMalqGQ NTEqmo5zns7Lx50nZQ/CLBsR5Wxt8hi6q8U3eSq+8xBcMe3teNL1Xl4tU8LSNsA5vKiZ9Jd0PYD oPnOyERAX X-Received: from plnm6.prod.google.com ([2002:a17:902:fda6:b0:2ca:b48c:5a92]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:1a0c:b0:2c9:e5ff:995d with SMTP id d9443c01a7336-2ce9ec0f09emr9029395ad.31.1783721212677; Fri, 10 Jul 2026 15:06:52 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:48 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=4613; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=nNNDGNJyFJDlOeB1he63eGYVs6CuWCKpVYovsf0zL60=; b=9R3nJNBsbirCZ4QZxsSBHZXAQa1UGt/eBuSwkrAzHrxhD1LsIhNtNL6Oz0rvTVXhQMgTtF6bE qni9q3JbcRjASVb9OtYiaREbOyexj4xDl8NTVD3+usBrkhWP0mwLmOR X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-4-56e7e2576cc1@google.com> Subject: [PATCH v4 4/9] vfio: selftests: igb: Program MSI-X interrupt routing From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The submitted driver writes only GPIE.EIAME (with a register value of 0x10, which is actually GPIE.Multiple_MSIX, bit 4) and clears EICR by reading it. On QEMU this works because the emulated loopback path is synchronous and EICR is implemented as read-to-clear unconditionally. Real 82576 hardware needs the full MSI-X programming sequence. Per 82576 datasheet section 7.3.2.11 Table 7-47, MSI-X mode requires: GPIE.Multiple_MSIX (bit 4): route causes through IVAR. GPIE.EIAME (bit 30): apply EIAM on MSI-X assertion. Without EIAME, section 7.3.2.11 specifies EIAM only takes effect on EICR read/write, which is not the path used here. Configure auto-clear and auto-mask for vector 0: EIAC (section 8.8.5): auto-clear of EICR cause bit on MSI-X assertion. EIAM (section 8.8.6): with EIAME set, auto-mask of EIMS on MSI-X assertion. This guarantees one interrupt per memcpy batch and prevents repeat delivery if the cause re-asserts before EIMS is restored. Replace the read-to-clear of EICR with write-to-clear. Section 8.8.5 states "If any bits are set in EIAC, the EICR register should not be read", and section 7.3.4.3 cautions against read-to-clear in MSI-X mode in general. Write-to-clear (section 7.3.4.2) is unconditional. Replace the magic '1' values written to EIMS/EIMC with the standard E1000_EICR_RX_QUEUE0 macro. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 40 ++++++++++++++++++= ---- 1 file changed, 34 insertions(+), 6 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 033997ad45e0..cd51a72b0e7f 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -296,14 +296,35 @@ static void igb_init(struct vfio_pci_device *device) /* Enable MSI-X with 1 vector for the test */ vfio_pci_msix_enable(device, MSIX_VECTOR, 1); =20 - /* Enable auto-masking of interrupts to avoid storms without a real ISR */ - igb_write32(igb, E1000_GPIE, E1000_GPIE_EIAME | E1000_GPIE_MSIX_MODE); + /* + * Program MSI-X interrupt routing per 82576 datasheet: + * + * GPIE (section 7.3.2.11, Table 7-47): set Multiple_MSIX (bit 4) to + * route interrupt causes through IVAR mapping, and EIAME (bit 30) + * to apply EIAM on MSI-X assertion (without EIAME, EIAM only + * applies on EICR read/write). + * + * EIAC (section 8.8.5): enable auto-clear of EICR for vector 0. + * Without auto-clear the cause stays set after delivery and the + * test can see spurious interrupts on the next memcpy batch. + * + * EIAM (section 8.8.6): enable auto-mask of EIMS for vector 0 on + * MSI-X assertion (effective because EIAME is set), so a single + * interrupt is delivered per memcpy batch even if the cause + * re-asserts before software re-enables the mask. + * + * IVAR (section 7.3.1.2, register definition in 8.8.13): map RX + * cause 0 to MSI-X vector 0 and mark the entry valid. + */ + igb_write32(igb, E1000_GPIE, E1000_GPIE_MSIX_MODE | E1000_GPIE_EIAME); + igb_write32(igb, E1000_EIAC, E1000_EICR_RX_QUEUE0); + igb_write32(igb, E1000_EIAM, E1000_EICR_RX_QUEUE0); =20 /* Map vector 0 to interrupt cause 0 and mark it valid */ igb_write32(igb, E1000_IVAR0, E1000_IVAR_VALID); =20 /* Enable interrupts on vector 0 */ - igb_write32(igb, E1000_EIMS, 1); + igb_write32(igb, E1000_EIMS, E1000_EICR_RX_QUEUE0); =20 /* Initialize driver state and capability limits */ igb->tx_tail =3D 0; @@ -326,17 +347,24 @@ static void igb_remove(struct vfio_pci_device *device) =20 static void igb_irq_disable(struct igb *igb) { - igb_write32(igb, E1000_EIMC, 1); + igb_write32(igb, E1000_EIMC, E1000_EICR_RX_QUEUE0); } =20 static void igb_irq_enable(struct igb *igb) { - igb_write32(igb, E1000_EIMS, 1); + igb_write32(igb, E1000_EIMS, E1000_EICR_RX_QUEUE0); } =20 static void igb_irq_clear(struct igb *igb) { - igb_read32(igb, E1000_EICR); + /* + * Use write-to-clear (datasheet 7.3.4.2). In MSI-X mode with EIAC + * programmed, section 8.8.5 explicitly states "If any bits are set + * in EIAC, the EICR register should not be read", which rules out + * the read-to-clear path in 7.3.4.3. Bits not in EIAC are still + * cleared by writing 1. + */ + igb_write32(igb, E1000_EICR, 0xFFFFFFFF); } =20 static void igb_memcpy_start(struct vfio_pci_device *device, iova_t src, --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pg1-f202.google.com (mail-pg1-f202.google.com [209.85.215.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 10006368D66 for ; Fri, 10 Jul 2026 22:06:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.202 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721220; cv=none; b=hJIV9T65CWiNWqErykAH523VsuGxSND+/Mt5et5ikFNoG9R+O3nej6wIEdzUl5srxW8nGvKIrdFPSt+dq7prjmx+JTmS4h4DEG6uP4ANfsVeTPmtewB9R8kW2hE1NlhdfvNWa4jC+KM5KgDL07HG8Y8xYQzdjhyUf2bvbDGQdbw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721220; c=relaxed/simple; bh=SXAt81smYsMjpJ1Hk7xiJa+jsCj2BWcuvU8qe14+cQA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=aIXgI2pKb+786zEiL8QQzytkDucGtuVQ58CsYCVIcfaq75WevkmAIJ6OeKs89cNW4K9IRhVPqR6QdeLjRLFwm5s4zBQh995QydeAr1ZmSojwPCwmsbxqZ3EnXrk5ebE9dJUY6A1CtLSYTfhK8PIUDrTqm+KZYdMh8PNVNbWnQNU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=vC1FtfGL; arc=none smtp.client-ip=209.85.215.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="vC1FtfGL" Received: by mail-pg1-f202.google.com with SMTP id 41be03b00d2f7-c88ad1558f4so2752665a12.2 for ; Fri, 10 Jul 2026 15:06:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721214; x=1784326014; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=yAmPzMMi8nVYfxKQwUu5xyjvBwQx6jMp227xZZBxUtM=; b=vC1FtfGLUwbhp16EFnYEuFujOOJk3ebo6xbvo2i+0S8l+izjINWLGhvJUUB+806h/T hKVMll+uG9pYW1pRUmA8bUXtE2l57ltL2HZxusTR8Vtz5J9ZpSsjoT/unTAr88c/72qa xHaLuOYZan5bHXXw5HU+TM80xIxHmZEgCSaitBWwpTPvzwUXiuZPifdpTOPGoXe8U2Il bmaL/N8MpQl73zgRt+K1t9PURzGWgeXXVi/+76pnMyABClJFgph/zIUlNaarekEiwTSu aKj/lnymceqja07Z0Upu3oC4SmwrTGZp7lyOw/bsYrYszDcrcM3iTjVDNeEpEewuF7+D XmVg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721214; x=1784326014; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=yAmPzMMi8nVYfxKQwUu5xyjvBwQx6jMp227xZZBxUtM=; b=Lk5ZisEFBIdCmoCzcSg+i4G3B195FeuElhDMankd1gJ/+JTy8QXRdM3368kdNkfNw6 MFFnyLpczpxzXjfHmq4UT8S7rBaKi2P6bw6tlwaEpMSJS+ICuuRKcW9sgeYibRQrcwgJ P+zqlH96U5dwYHjrmtRluySt6o/kS9WJpMgCI0b/82Q15fFgbVMyg2vjtwFepb8Lek/I YCbHz+vA+TCtm9vJeCgjbGX2QXDDQj6yGo78fWuujHMhrQrlWSJ1Z+dWPqfKmkm6YSyE AzwPUoQWNIez74am5d236k6p1jlKAxcHAjVIkEtR0LWPTNkwWgDx+hJ/whEuLWX4p/PV qiFQ== X-Forwarded-Encrypted: i=1; AHgh+RpcDZv/M9HfJFjjxL1EoY/T8rEyDxsSGXOXZyiDt4yojbQ89Z1YFbArdKB6MdnBA+666qf0wZHHw2c4NbU=@vger.kernel.org X-Gm-Message-State: AOJu0Yzi8WKrHbOtgwb7QH2WOwanJwvp2oL0VrrLiPIyy6tM9k2sy45I Zb9tDQ2DBYSlBDNSw4sU+oW4/0Kokipty6lnX2aCt3A348bT464Qq/1IsOL2w1Gll9pTjFKJDc/ DzNZ6fXU7 X-Received: from pjnu12.prod.google.com ([2002:a17:90a:890c:b0:382:a16a:19ad]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:4a4b:b0:37f:e8d6:72c0 with SMTP id 98e67ed59e1d1-38dc73c3657mr700004a91.1.1783721213576; Fri, 10 Jul 2026 15:06:53 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:49 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=2461; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=r1Tao9/LBmC1l3hkg5OeTapQHKhzjrz5/X4dLEtJza0=; b=9gpQVJ3/qOt3GGhj675mIHg27FlGJA4S2fNW+p4SxmalDB4d4r2lKHsONY5BHCeBbGxPk8w4k kNfZT5M7w1PBErVq3bvZbp69PBaA1r+RY/hPMR/Voqv0CTweFx6D1/L X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-5-56e7e2576cc1@google.com> Subject: [PATCH v4 5/9] vfio: selftests: igb: Extend memcpy completion timeout for line-rate hardware From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The submitted driver waits at most 1 ms (100 * 10 us) for the last RX descriptor to be written back. QEMU's emulated loopback is synchronous: by the time igb_memcpy_wait() runs, the receive descriptors are already written back. Real 82576 hardware processes the descriptor ring at line rate. max_memcpy_size is (RING_SIZE - 1) * IGB_MAX_CHUNK_SIZE, approximately 4 MB, split into 4095 1 KB frames. At 1 Gb/s line rate (~125 MB/s), 4 MB takes ~32 ms on the wire, plus per-frame preamble, SFD, inter-frame gap and FCS overhead (~3%) and descriptor fetch/writeback latency. The 1 ms cap times out long before any valid transfer can complete. Wait up to ~200 ms (200 iterations * 1 ms) for descriptor writeback before returning -ETIMEDOUT. ~6x the line-rate floor leaves comfortable headroom for host scheduling jitter while keeping the intentional invalid-DMA tests (mix_and_match) so they recover quickly on real faults. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 15 +++++++++++++-- 1 file changed, 13 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index cd51a72b0e7f..172c95cea3c8 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -428,12 +428,23 @@ static int igb_memcpy_wait(struct vfio_pci_device *de= vice) prev_tail =3D (igb->rx_tail + RING_SIZE - 1) % RING_SIZE; rx =3D &igb->rx_ring[prev_tail]; =20 - retries =3D 100; + /* + * Real 82576 hardware processes the descriptor ring at line rate. + * max_memcpy_size =3D (RING_SIZE - 1) * IGB_MAX_CHUNK_SIZE ~=3D 4 MB, + * split into 4095 1 KB frames. At 1 Gb/s (~125 MB/s) the worst + * valid memcpy takes ~32 ms on the wire, plus per-frame preamble, + * SFD, IFG and FCS overhead (~3%) and descriptor fetch/writeback + * latency. Wait up to ~200 ms before declaring the device hung; + * ~6x the line-rate floor leaves comfortable headroom for host + * scheduling jitter while keeping the intentional invalid-DMA + * tests bounded. + */ + retries =3D 200; while (retries-- > 0) { status =3D le32_to_cpu(READ_ONCE(rx->wb.status_error)); if (status & 1) break; - usleep(10); + usleep(1000); } =20 if (status & 1) --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pl1-f201.google.com (mail-pl1-f201.google.com [209.85.214.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9D69E3CBE6D for ; Fri, 10 Jul 2026 22:06:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721222; cv=none; b=sMJI50WrfdW1viQAQ71B0zFx2iwKscyRAl1x680ovPeHRq3FdIWJ47Y4sYnAyAjssh73VzegedGtjZKZgR/DTIrSv+do86YYSvrV8AgFIZTntWgZW4gPQP2Wl/m/8Ty81wTK/K7mO2fF05qJ0qPCzH1HwyOzGzFhlmeAg8t5+sU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721222; c=relaxed/simple; bh=shDmtApnZvdyo7L5TZAFGjMrm9vqx2VXSgijpuIkVbc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=k9LTODZoxbDFHxFR1JBlpdKo3W/11EAfyFN+Vof2E4ry09KRg4u40sGFgl9JJlwQATyrRNI0E29T6H5Wy4nkS5e0GxiEaNji3qOVMXpaXsp7MxSIWbm0cugeYpcqy5gVRcswyJWwbYBH+gbiCYQiRfZX7mSJvbfeAf21F8X+hHc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=GkVrxyz9; arc=none smtp.client-ip=209.85.214.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="GkVrxyz9" Received: by mail-pl1-f201.google.com with SMTP id d9443c01a7336-2cc7e86e7c5so26025705ad.3 for ; Fri, 10 Jul 2026 15:06:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721215; x=1784326015; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=tQns3HevZIzRhw7H0Nuhca8MxJp9vTEMOhRtrWCAyX8=; b=GkVrxyz9ajBJh5WbyHLI/BScIE/Vk/hATQ6746H/ZoQxtfn4057+EiHRUHQNXksPXS lVGYGF3QwSu/Kj/JkXHWqdFy2yCnhG6gdFxGDgjFPqAQQOCf8+GOHyWOy5JOrOiBL3zN fM+uOcrlpjr5mihIJeu4eh5LomoGmzHMdYu/lljujSXfhOFdq6ARgAipg+wEnCStAfFo MBw80LUK409kK4DvY7NZrmZf2raxdK9oDRvRBJ1P/sXV0MvAuv0pTGty7eWqxHxBFmKc HpVJ34YCf/YqXG4Tw8QMh9LpDZBqmz00HlDcOFW6/Rdz528PxOTVyCuMg9D9kjHkw/tm wEuw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721215; x=1784326015; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tQns3HevZIzRhw7H0Nuhca8MxJp9vTEMOhRtrWCAyX8=; b=C1/P9HzHvzxdt5rbuFw1Af5Zl2Jd0wdw4G+tpgpTAYA4Jb7dTgsj4SyVbdKkywSPm8 7AqFncAbbmLXGascoPQUeBiZ68MCT8rTYCmW2TbazW4XEV1G0u4t0c+JMCbfOLvWVmF6 jbS79w0XywUAM+kE4iTliyHL12lsAwwih+W1JeAHC8vZdX4daxJovuTG7IXJWQCI+Fos B9ZCpCwyWlTGGIf4soPalDfxVYoCYZGyBPHAE8hswYFT8l1bvjjpDpjcr4xgzuS0ydLV FMwj5pP4IrO+6azdJs7tQIWkf97WKCtCuoP95SPEg9s8RCPOVZlfRi1aY8IHXGeBrtvI 0fyQ== X-Forwarded-Encrypted: i=1; AHgh+RotG3UtjNxLuZxPIckzeN7C7/CMlVHoDmFtFrVOklDB0fkakyzwV7ffkHV/1HZNb3to/EuCNX8IPDcWTaE=@vger.kernel.org X-Gm-Message-State: AOJu0YxPNjPlu1p+f7UbH/RuKl0ERh4McinJ2ZzZkclggbAy6YBikdq8 BHnCPhl/eueBDu2uA64Ygc08avMpyh2VHWHhzckhrjAiGrMblrLxxUdNeTUW/kmS7bTd6/ywCUQ 321mVM2vu X-Received: from plhq14.prod.google.com ([2002:a17:903:11ce:b0:2ca:d246:c7c8]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:110f:b0:2cc:fa08:eebc with SMTP id d9443c01a7336-2ce9e7a544cmr8374465ad.1.1783721214508; Fri, 10 Jul 2026 15:06:54 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:50 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=1853; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=ULyYN7S4rWudzktl97NB9xV0EBoQIE8YAYNFU9RV/IU=; b=faKEKwlLwFC21rtzYiGCpQGJmOlryRDxBoYYhFyHp8njegao69icprtRlf8nh38PY+sZzwK2l DR3dNnTn1LyAkg2ZZPsDGJ0KklKobjPJMUPaj986jiAhVElN5+cUsJj X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-6-56e7e2576cc1@google.com> Subject: [PATCH v4 6/9] vfio: selftests: igb: Disable PCIe completion timeout retries From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The mix_and_match test intentionally submits TX descriptors with an unmapped source IOVA so that the DMA read fails. By default the 82576 re-sends the request after a PCIe completion timeout (datasheet section 8.6.1, GCR.Completion_Timeout_Resend, bit 16, initial value 1b). On real hardware this turns a single fault into a stream of retried reads, keeping PCIe AER and IOMMU error handling busy and interfering with reset recovery. Clear GCR.Completion_Timeout_Resend during device initialization so a failed read fails once and stays failed. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 172c95cea3c8..ac116834bbf6 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -230,6 +230,18 @@ static void igb_init(struct vfio_pci_device *device) vfio_pci_config_writew(device, PCI_COMMAND, cmd_reg); } =20 + /* + * Disable DMA re-send on PCIe completion timeout (82576 datasheet + * section 8.6.1, GCR.Completion_Timeout_Resend, bit 16). The + * mix_and_match test intentionally submits descriptors targeting + * unmapped IOVAs; with the default (set) value, the device keeps + * retrying the failed read indefinitely, which keeps PCIe AER and + * IOMMU error handling busy and interferes with reset recovery. + */ + ctrl =3D igb_read32(igb, E1000_GCR); + ctrl &=3D ~E1000_GCR_CMPL_TMOUT_RESEND; + igb_write32(igb, E1000_GCR, ctrl); + /* Configure PHY internal loopback for testing. */ igb_setup_loopback(igb); =20 --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pg1-f201.google.com (mail-pg1-f201.google.com [209.85.215.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DDA3C3B3C18 for ; Fri, 10 Jul 2026 22:06:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721222; cv=none; b=YR2QVQRJKMKnFLjPMMdWIurZGFOPQbF0vM5o+d6Ldf5aF5ukG8rmKTBzQVUPFVHfY9v7K4ovtza5s2Q4kL2481sBDKPh58G6X3LpJimFx4Hh0viLPOU9WWglSpzHKO8khBHVVL/ZkRar2y7BAX63DJ6yL6kezrdLksvmvnxPniE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721222; c=relaxed/simple; bh=YW5RyQU5o9OBhh9jhOLqy2BLQ4mW4jnepu221Af6evA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=DiNFPug04WfMN7/XxkTNl6ru5/0VkV2eH9ATxhHdNS8oR63FvMUJAGSJg0hJjWDqNI4Vu4qY3CvJYnCgyR5lUkk3iDDIUTXdZNfDlGbS7EUxCYRYDzQC2uMe/trSaYUhiDCw/w0Ambs0YJEhp3+OJPsdjR1OTljZqXWlSXW0cWg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=OjRD1ytD; arc=none smtp.client-ip=209.85.215.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="OjRD1ytD" Received: by mail-pg1-f201.google.com with SMTP id 41be03b00d2f7-c88ab059052so1325113a12.1 for ; Fri, 10 Jul 2026 15:06:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721216; x=1784326016; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=cm9aCzqOXkPalEf3rO5p1tnL2oSXVIGbLiKDcdTayQI=; b=OjRD1ytDyWB6MDC5JYD7/xRA7iSLeyXkLBI2DSnWmgi7g13TLml8ZJKiM2XZYobtY2 1kZo8WhvHn1zSMJLUIecsY+NWUn70ER/1TJzc7bxzF3R3almcqZdtHxSmMFxqD45H2wc xQPt0pocZPwhCT/m/SSEdpM7BAmf+Y0ekN6ipbSXueih3unXrl1jRr6IG0jjgjPlRpM4 zC3O2/N2wt3bxqiyEn+4O3ED/j1stFNC989WmdiifixQ1DOQwyTNB1Z/F/90aOxkER55 YHpSDp48iSNo8XF5V9bOIRVsjUfLnLKapVwu8KgUKXhZYHA4aqxKb9Sme9dgaTVNmt/p 5ytA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721216; x=1784326016; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cm9aCzqOXkPalEf3rO5p1tnL2oSXVIGbLiKDcdTayQI=; b=DgkGritqb5V9U7cZX2+SGU/qQ4Lq4hucRGPllrRckd3C3iapbALGYZ6oH3F/T/vN9J j07ZtIH3qJGtp5Ek7l/QkamBN2X3hOfGDtNPrXHtrjnbvJJoZSfywcU2H44G5pDEcL+w QYkeIa9uBdZbv2vpjNpQEN6RRjmOLiXKUbN1FE+GfnLxRjpH7/ZpiNZFdHQyuWxY9LVK IOGkqZJMHwtbvFvP1plsauiW8sxhpCxObZIBK5Ps/uP+E7ocRqU47f2IiWSCaTuDeRfG 7wBkIG8zr6yyS1EUaPFIQnApgT7GTHIrJhZkhDHPejSVZf3xbNTNLIfZ5eUQPvZlkacX 8yMA== X-Forwarded-Encrypted: i=1; AHgh+Rq88UI6HGDPs8zibOV3u+5REE9bY6u9KmpxaTNrZbtFpO59C5BNxOE7ei74cSLt/iZgJUvzQVrrQjb2Cbs=@vger.kernel.org X-Gm-Message-State: AOJu0YximJLRdxCnwv2qfbfVtKzuE/yIzE3ZbCC/AU4M46BYGuG6Db4o ua1Y5GD3ycEcY0OAgif+jA7mg3g0TlZirANsZz55ceS0Jj5X0rg3WaWPpvhy8SxObttnD7S86Zv cmTNUw4CZ X-Received: from pfev14-n2.prod.google.com ([2002:a05:6a00:c20e:20b0:848:7f21:167c]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:2290:b0:845:cad1:d691 with SMTP id d2e1a72fcca58-848703454d6mr4801069b3a.5.1783721215357; Fri, 10 Jul 2026 15:06:55 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:51 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=4367; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=DNJx8mJSNM15h+y28JiieS7tQ/YHFwGsaY13DBSJOYU=; b=gRFh46xEOPMrtazvQX+KBJ5CH/Grz5VuxsjQrCoiR/sC59vgO0S2z23AhGF+bXRoHPtCjXaJb NcPaDg4Mh7GCcMQv3CaaZMp0evKWXBMP0u183D4YNfIfC9vPvgk8VZF X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-7-56e7e2576cc1@google.com> Subject: [PATCH v4 7/9] vfio: selftests: Add helpers to re-enable interrupts From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson Selftest drivers that recover from a fault by issuing VFIO_DEVICE_RESET need to re-arm device interrupts afterwards. VFIO_DEVICE_RESET tears down the kernel-side IRQ trigger so a subsequent VFIO_DEVICE_SET_IRQS is required, but the user-side eventfds (and any fd cached in a test fixture) are still valid and must be preserved. vfio_pci_irq_enable() refuses to be called for vectors that already have an eventfd (VFIO_ASSERT_LT), and vfio_pci_irq_disable() closes all eventfds before resetting the trigger, so neither is suitable. Add vfio_pci_irq_reenable(device, index, vector, count) which asserts that the requested range has existing eventfds and re-issues VFIO_DEVICE_SET_IRQS using them. Signature mirrors vfio_pci_irq_enable(). Add vfio_pci_msi{,x}_reenable() wrappers around vfio_pci_irq_reenable() for additional ease of use and readability. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson --- .../vfio/lib/include/libvfio/vfio_pci_device.h | 14 ++++++++++++++ tools/testing/selftests/vfio/lib/vfio_pci_device.c | 22 ++++++++++++++++++= ++++ 2 files changed, 36 insertions(+) diff --git a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_devi= ce.h b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h index 2858885a89bb..2e67afc0d580 100644 --- a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h +++ b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h @@ -68,6 +68,8 @@ void vfio_pci_config_access(struct vfio_pci_device *devic= e, bool write, void vfio_pci_irq_enable(struct vfio_pci_device *device, u32 index, u32 vector, int count); void vfio_pci_irq_disable(struct vfio_pci_device *device, u32 index); +void vfio_pci_irq_reenable(struct vfio_pci_device *device, u32 index, + u32 vector, int count); void vfio_pci_irq_trigger(struct vfio_pci_device *device, u32 index, u32 v= ector); =20 static inline void fcntl_set_nonblock(int fd) @@ -92,6 +94,12 @@ static inline void vfio_pci_msi_disable(struct vfio_pci_= device *device) vfio_pci_irq_disable(device, VFIO_PCI_MSI_IRQ_INDEX); } =20 +static inline void vfio_pci_msi_reenable(struct vfio_pci_device *device, + u32 vector, int count) +{ + vfio_pci_irq_reenable(device, VFIO_PCI_MSI_IRQ_INDEX, vector, count); +} + static inline void vfio_pci_msix_enable(struct vfio_pci_device *device, u32 vector, int count) { @@ -103,6 +111,12 @@ static inline void vfio_pci_msix_disable(struct vfio_p= ci_device *device) vfio_pci_irq_disable(device, VFIO_PCI_MSIX_IRQ_INDEX); } =20 +static inline void vfio_pci_msix_reenable(struct vfio_pci_device *device, + u32 vector, int count) +{ + vfio_pci_irq_reenable(device, VFIO_PCI_MSIX_IRQ_INDEX, vector, count); +} + static inline int __to_iova(struct vfio_pci_device *device, void *vaddr, i= ova_t *iova) { return __iommu_hva2iova(device->iommu, vaddr, iova); diff --git a/tools/testing/selftests/vfio/lib/vfio_pci_device.c b/tools/tes= ting/selftests/vfio/lib/vfio_pci_device.c index fc75e04ef010..7b8394d0ac50 100644 --- a/tools/testing/selftests/vfio/lib/vfio_pci_device.c +++ b/tools/testing/selftests/vfio/lib/vfio_pci_device.c @@ -106,6 +106,28 @@ void vfio_pci_irq_disable(struct vfio_pci_device *devi= ce, u32 index) vfio_pci_irq_set(device, index, 0, 0, NULL); } =20 +/* + * Re-issue VFIO_DEVICE_SET_IRQS for an already-enabled vector range using + * the existing eventfds. Intended for drivers that need to re-arm device + * interrupts after a VFIO_DEVICE_RESET, which tears down the kernel-side + * IRQ trigger but leaves user-side eventfds intact. Recreating the + * eventfds would invalidate any test-fixture cache of the fd, so this + * helper deliberately preserves them. + */ +void vfio_pci_irq_reenable(struct vfio_pci_device *device, u32 index, + u32 vector, int count) +{ + int i; + + check_supported_irq_index(index); + + for (i =3D vector; i < vector + count; i++) + VFIO_ASSERT_GE(device->msi_eventfds[i], 0, + "vector %d eventfd not allocated\n", i); + + vfio_pci_irq_set(device, index, vector, count, device->msi_eventfds + vec= tor); +} + static void vfio_pci_irq_get(struct vfio_pci_device *device, u32 index, struct vfio_irq_info *irq_info) { --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pg1-f201.google.com (mail-pg1-f201.google.com [209.85.215.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CACD33D905F for ; Fri, 10 Jul 2026 22:06:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721223; cv=none; b=l8O0i5WpbNesRPE92K3MgEG73Xl+l6C0DhjBUHuyZ2SQ1gwd6gUlHd6emn1TOSBlYvR/ova64QstMmRbD1Uo0O0Eamzvnel23PlnvzCowLpS7fH3LCOB+0vEtms3njVA0etBJbuhBtQIjH6RJKzB9Mwjz20C+uWOoqnMl+UyZSU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721223; c=relaxed/simple; bh=5XZ+blBPN7UkWKijZIA5svjbA/3Bw9cDBIfFFagMloE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ovVL5XddQ/jshQeytP+deptH0jYXtaIX9TNRiDabdrk1NFHIMOZa0TCCm683t5anRSFXocfR4MWr6+3ZoU+6k19AsAzzwodk6hvP4KW3yTeYnG2qf9yz96Qgxckq2IbN1BYZQuii7NI4lbCxHJBsaEreuMK0gz13RVGiRz6oBHc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=rVvEzbD1; arc=none smtp.client-ip=209.85.215.201 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="rVvEzbD1" Received: by mail-pg1-f201.google.com with SMTP id 41be03b00d2f7-c894c1c4aa9so1780746a12.0 for ; Fri, 10 Jul 2026 15:06:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721216; x=1784326016; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=RTfqEjSO/zvcn0XSIJP07HCV2qOV7Xme/M8SMmLWtWU=; b=rVvEzbD1emavg0ZQwvUHhngQqrKY/UVd377+vTCExCq7Xc+3negef+SIgWQDpGo4bi JFiah1uegbUrxziR9U6T0gsoqqqz6PSSufe35ga4a9w2yQvTmfbQJE9+EDmEQHhnLesJ 8bSDsl45MrR2HVNTwGhG7IubDcoUSn70S4bIKpuL4qnIIit4j7+1fBVOMptrhzWsMMja q35t/yfFzsB9Dgq1HEfB4s3WN1qJTZLYSFgJqAMQRJxwWJa1aXA53P+yVtuc3qn3PxBF VJpVpP64kYfJQQlpqFvTqgktdunWFkbCLo3MQCdLUOx14Ofe5+mazNEHUJo63pGX8IFF a8SQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721216; x=1784326016; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=RTfqEjSO/zvcn0XSIJP07HCV2qOV7Xme/M8SMmLWtWU=; b=eTbb8pSUWs0svzf1bpY8syws/i4ov5C6jtH9jKpIgSXtk9aQH2AqxQl2zZwTrvDFTK de/okTb0DWTJa/1fc9m647+h50km/Z83qQpORbR++lnzzQkrZZmn7IzfXe/AKxVm31l5 GHTQ8DlcNyJhCTqZxu7Ozr+GSqdDKebJKGJHC+rc0uNjD96S3wWa075vAC/Q+laWpuod CahWwRZFE4kMsmfRHjCFjDvR7QfIkIZzV+6e+rgJCft5Vgjqi2Z90y3+IFB/CIUgmrne Q0x+H2iMRVGvSEIemMpIh5F/HCAwrYMtCDNsKLhmgYgJ1gR9L9FaelPgjACdYsn0Q+4H 6wIg== X-Forwarded-Encrypted: i=1; AHgh+Row3C0O9QPesAXqe8QFCf/1uynX1j47DzN9uUF7FBjwenOpNOvyGQdScW4Gh9ODFPWWoHDeQ7bs6I3ZiT8=@vger.kernel.org X-Gm-Message-State: AOJu0Ywv1Uem6kPihlpqd04uYjXq6RpjniwJO3iJT9LPU+bY62ZGjuh3 ydBEQEYFt0trAaugsO6swSjdeouuKUmTn811zf1UD1hCDYKfNEQhG6yvDNieTv+MwgC9cf3jFN8 Z0ygf+bXH X-Received: from pgmj9.prod.google.com ([2002:a63:5949:0:b0:c8b:2b50:846b]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:a10e:b0:3bf:b705:4b9d with SMTP id adf61e73a8af0-3c110a7a105mr768481637.28.1783721216316; Fri, 10 Jul 2026 15:06:56 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:52 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=4909; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=l2upaY11NNXoOWkPHwt4VKr5HfO5SZwlcbFOEbuFak4=; b=3IOSypgsWvEENQB/lo0rbusjaVmdMbswoFfQGcOyxCC9iq79Vifre6hBXVPA81sWTAPyVATut p5VJdJcNqHqAVnyhU42Nn4C19Q3bLYAnPLGoEc9BvYF/ZBmBpM6QW6p X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-8-56e7e2576cc1@google.com> Subject: [PATCH v4 8/9] vfio: selftests: igb: Factor hardware programming into igb_hw_init() From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson Split the device register programming out of igb_init() into a new igb_hw_init() helper so that the same sequence can be re-run after a VFIO_DEVICE_RESET to restore the registers that CTRL.RST clears. No functional change for the initial path. igb_init() now performs the one-shot setup: region size assertion, BAR mapping, CTRL.RST + IMC mask-all to put the device into a known state, and vfio_pci_msix_enable() to set up the kernel-side IRQ trigger. igb_hw_init() does the rest: ring pointer setup and IOVA calc, CTRL_EXT, PCI bus master, GCR, PHY loopback, descriptor rings, RCTL, TCTL, GPIE/EIAC/EIAM/EIMS/IVAR, and driver-state initialization. vfio_pci_msix_enable() moves from after RCTL/TCTL to before all device-side programming. Its only side effects are the VFIO kernel IRQ trigger setup and the PCI MSI-X capability bits in config space; neither has any ordering dependency on the 82576 device register writes performed in igb_hw_init(). Performing it once in igb_init() keeps igb_hw_init() reusable from the reset recovery path (which uses vfio_pci_irq_reenable() to re-arm the existing trigger). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 61 ++++++++++++++----= ---- 1 file changed, 40 insertions(+), 21 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index ac116834bbf6..23d924b791d3 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -188,7 +188,13 @@ static int igb_probe(struct vfio_pci_device *device) return 0; } =20 -static void igb_init(struct vfio_pci_device *device) +/* + * Program the device into a usable state. Split out of igb_init() so it + * can be reused after a device reset to re-program the registers that + * CTRL.RST clears. Expects bar0 to be mapped and MSI-X already enabled + * via VFIO. + */ +static void igb_hw_init(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); u64 iova_tx, iova_rx; @@ -196,26 +202,9 @@ static void igb_init(struct vfio_pci_device *device) u16 cmd_reg; int retries; =20 - VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); - - /* Set up rings and calculate IOVAs */ - igb->bar0 =3D device->bars[0].vaddr; - iova_tx =3D to_iova(device, igb->tx_ring); iova_rx =3D to_iova(device, igb->rx_ring); =20 - /* Reset device and disable all interrupts */ - igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); - /* - * Must wait at least 1 millisecond after setting the reset bit before - * checking if this device is ready to be used (82576 datasheet section - * 4.2.1.6.1). - */ - usleep(1000); - while (igb_read32(igb, E1000_CTRL) & E1000_CTRL_RST) - usleep(10); - igb_write32(igb, E1000_IMC, 0xFFFFFFFF); - /* Signal that the driver is loaded */ ctrl =3D igb_read32(igb, E1000_CTRL_EXT); ctrl |=3D E1000_CTRL_EXT_DRV_LOAD; @@ -305,9 +294,6 @@ static void igb_init(struct vfio_pci_device *device) igb_write32(igb, E1000_RCTL, rctl); igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); =20 - /* Enable MSI-X with 1 vector for the test */ - vfio_pci_msix_enable(device, MSIX_VECTOR, 1); - /* * Program MSI-X interrupt routing per 82576 datasheet: * @@ -347,6 +333,39 @@ static void igb_init(struct vfio_pci_device *device) device->driver.msi =3D MSIX_VECTOR; } =20 +static void igb_init(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); + + igb->bar0 =3D device->bars[0].vaddr; + + /* Reset device and disable all interrupts. */ + igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); + /* + * Must wait at least 1 millisecond after setting the reset bit before + * checking if this device is ready to be used (82576 datasheet section + * 4.2.1.6.1). + */ + usleep(1000); + while (igb_read32(igb, E1000_CTRL) & E1000_CTRL_RST) + usleep(10); + igb_write32(igb, E1000_IMC, 0xFFFFFFFF); + + /* + * Enable MSI-X via VFIO before device-side register programming. + * vfio_pci_msix_enable() only touches the VFIO IRQ machinery and the + * PCI MSI-X capability via config space; it has no ordering + * dependency on the device-side writes performed by igb_hw_init(). + * Placing it here keeps igb_hw_init() reusable from the reset + * recovery path (which calls vfio_pci_irq_reenable() instead). + */ + vfio_pci_msix_enable(device, MSIX_VECTOR, 1); + + igb_hw_init(device); +} + static void igb_remove(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); --=20 2.55.0.795.g602f6c329a-goog From nobody Sun Jul 26 00:21:52 2026 Received: from mail-pl1-f202.google.com (mail-pl1-f202.google.com [209.85.214.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 54FD9346FCA for ; Fri, 10 Jul 2026 22:07:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.202 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721225; cv=none; b=tATWf1Dwh7REFfT8MCyXEXfFpWIyz0lSnlJ7vvbWYV2S2YQor8z0ZFw93iNdDEe8KfZjtfiagOFgC6+SBJE99Ly5UtohJ1QoKe/gnymgeMaguTznBgQMi3wYjOt1+kvo6fxikpKew9Qvc3dO0/x49xBv1jiS1tnhMwprSltbf7A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783721225; c=relaxed/simple; bh=VQFmUAlim5jQu8woDrKlC6bOAbFrCCqDsh3j5t6QW9I=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=lCg2o9WQDmIrFZTUrpeH6D9SOf1oc2WBkyYdiC6yejkZV8ly/lnBvw3E7EmLLW1VcamakjAVgZzrt8CjSC4rlisU+zqlAnoKUSsS+W8IhuER6u4j/iCLHenoPbk9Rp2hipZ4tK7Qk4VbZbX4E1xEykBzTSH1ZT7gJ27d+s7vMyU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=uTj1N3yl; arc=none smtp.client-ip=209.85.214.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="uTj1N3yl" Received: by mail-pl1-f202.google.com with SMTP id d9443c01a7336-2cc86a9ef97so25248105ad.3 for ; Fri, 10 Jul 2026 15:07:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783721217; x=1784326017; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=7jWCHK+Mcfe3aW0kz1paSAUghYLWj0BILc9NRBI8+lI=; b=uTj1N3ylB04TB6YyXQhzbg4FjTCp+7HDQCTPrZf/D2mty341uWGLlrssznW9HnK05z 7tlnKYBfINyCTENUqQwV6aM3d19qCp62e1Qco6gXCqytIctKAo154apC34URQG6cI/h9 hCCeky+iUrJKPXOWsSx+wmJyRPj1qPBDHVT6HG+sIbE0tYd0+TuwhFbAGfBWqML7rkyh IQDMXqZsphxklkXQMGacQAeh9kD6HeUr+m5Nlb3orR8vNHVBRqSgiNxq3WETPQgGi30c VbaxxzEbuHmY2d9pxXbYnwN5o8IDs8ezOxodZEP3hZRk2eFhI3uAw6eor2+e8CY5iJei nJSQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783721217; x=1784326017; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7jWCHK+Mcfe3aW0kz1paSAUghYLWj0BILc9NRBI8+lI=; b=hUVkV92NlFHOaKqch6WrkXSjfWP10cnF2GwSw4t2xKwHmaRAMm9rp1ZUUygFHV6sYP AvLoFfNYWQnBV/L4OqTCleMIoHEQnxSLy9TF/NTjiM4ONsSkoxkJMK0DyGyKbbyZY/Iu vk7vYxDI2nL/zBPj2OsCPkny9LLkNSimYZu891QwiidQwy45+ONAOEHfaT+320EWuvPR hIQF18PJp7B+J/qdDX6AsDcbgsJAn1Wd6ytdkhdN4CNA1L4hXnj/r5CfgJmgHhNeQM4D eItq8kQNxrWNGsrOu5XKBqWx9h1+ZghmeU0VSkX1YyzuGlSZ9a4Ni9I2OBj6Psbv89PQ IHhA== X-Forwarded-Encrypted: i=1; AHgh+RoJnStL+LEzVwzQeu2bwIv0Cn9FiNfl529VC7TIa28NppKE1Guaa3n4PVW6HQbXeq+FvVXAL9tqVNfqNZQ=@vger.kernel.org X-Gm-Message-State: AOJu0Yw6J9O6ErIw0RV2B+8FmexNa072GH6Ex1hUJdeo4cRAdc6j21RK LwX7OX3ON+GxCBSO4zq6ky/iM8wLS3yL2kERdy2F0Ae1vHyxnRXC+J8yOj77rnueR5qG5k8D7Ak PalcQ0oXt X-Received: from plgv1.prod.google.com ([2002:a17:902:e8c1:b0:2cc:7044:4346]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:2983:b0:2ca:d666:df72 with SMTP id d9443c01a7336-2ce9ead1556mr8008955ad.21.1783721217228; Fri, 10 Jul 2026 15:06:57 -0700 (PDT) Date: Fri, 10 Jul 2026 22:06:53 +0000 In-Reply-To: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260710-igb_v3_b4-v4-0-56e7e2576cc1@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1783721208; l=4873; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=SybNIe800vS/azPYYOHHfOsE4B9hb1uiVaknYMjZzyw=; b=ylIW6ZAisGRB0cPzbwFGJTZ7ARz05XJiKCCX+flh9NIpfGkjb55JbgdLAW4MmdxUZNAtSLOtz 7OoeDATb+JzDwCNGeZTH5JQm6vuhhrNbdCN9xZWPq2cbbYI+4HE2fDY X-Mailer: b4 0.14.3 Message-ID: <20260710-igb_v3_b4-v4-9-56e7e2576cc1@google.com> Subject: [PATCH v4 9/9] vfio: selftests: igb: Recover after DMA-read faults From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The mix_and_match test intentionally submits a TX descriptor with an unmapped source IOVA so that the DMA read fails. On real 82576 hardware the resulting fault leaves the descriptor engine unable to service subsequent valid descriptors, so the next memcpy in the same test iteration times out. The 82576 datasheet (section 4.2.1.6.1) describes CTRL.RST as the software mechanism to recover from a hung device. Empirically CTRL.RST alone is not sufficient in this state: the visible queue registers are reinitialized, but the next valid memcpy still posts descriptors without any TDH/TDT progress in the same process. A fresh device open after the failure works, which points to a reset scope broader than CTRL.RST being required. The 82576 advertises PCIe FLR; VFIO_DEVICE_RESET drives FLR and supplies that scope while preserving the selftest process and its DMA mappings. Add igb_error_reset_and_reinit() implementing the recovery sequence: issue VFIO_DEVICE_RESET, re-arm the kernel-side MSI-X trigger against the still-valid eventfd via vfio_pci_irq_reenable() (this does not touch the eventfd, which test fixtures may have cached), and re-program the device via igb_hw_init(). FLR clears EICR and leaves EIMS=3D0, so no explicit interrupt mask or cause writes are needed. igb_hw_init() resets tx_tail/rx_tail to 0 and igb_memcpy_start() zeros each descriptor before submission, so no ring memset is needed either. Call this from igb_memcpy_wait() on completion timeout, preceded by a 10 ms delay so that PCIe/IOMMU/AER error handling triggered by the just-observed DMA fault can release the device lock VFIO_DEVICE_RESET contends for. The delay is heuristic and tied to the fault path, so it lives at the call site rather than inside the reset helper. The failed memcpy still returns -ETIMEDOUT; reset recovery only ensures the next operation starts from a usable device state. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 45 ++++++++++++++++++= +++- 1 file changed, 44 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 23d924b791d3..206a1e6489c2 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -448,6 +448,28 @@ static void igb_memcpy_start(struct vfio_pci_device *d= evice, iova_t src, igb_write32(igb, E1000_TDT(0), igb->tx_tail); } =20 +/* + * Reset the device via VFIO_DEVICE_RESET (PCIe FLR on the 82576) and + * re-program it. VFIO_DEVICE_RESET tears down the kernel-side MSI-X + * trigger but leaves user-side eventfds intact, so re-arm the trigger + * via vfio_pci_irq_reenable() before reprogramming so any caller-cached + * eventfd remains valid. + * + * FLR clears device-side state to power-on reset values (datasheet + * 4.2.1.5.1: a PF FLR is "equivalent to a D0->D3->D0 transition"), so + * EIMS and EICR come back as 0 from their register-defined initial + * values, and igb_hw_init() resets tx_tail/rx_tail to 0. The next + * igb_memcpy_start() will memset each descriptor it touches before + * submission, so no explicit IMC/EICR writes or ring memsets are + * needed here. + */ +static void igb_error_reset_and_reinit(struct vfio_pci_device *device) +{ + vfio_pci_device_reset(device); + vfio_pci_msix_reenable(device, MSIX_VECTOR, 1); + igb_hw_init(device); +} + static int igb_memcpy_wait(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); @@ -490,7 +512,28 @@ static int igb_memcpy_wait(struct vfio_pci_device *dev= ice) =20 igb_irq_enable(igb); =20 - return (status & 1) ? 0 : -ETIMEDOUT; + if (status & 1) + return 0; + + /* + * The descriptor never completed. On real 82576 hardware this + * typically follows a DMA-read fault from one of the intentional + * unmapped-IOVA tests; the fault leaves the descriptor engine + * unable to service subsequent valid descriptors. CTRL.RST alone + * reinitializes the queue registers but leaves the engine wedged + * for the current process, so a broader VFIO_DEVICE_RESET (FLR) + * is required. + * + * Delay before requesting reset so PCIe/IOMMU/AER error handling + * triggered by the just-observed DMA fault can release the device + * lock VFIO_DEVICE_RESET contends for. The 10 ms value is + * heuristic. The current memcpy still fails with -ETIMEDOUT; + * recovery only ensures the next memcpy starts from a usable state. + */ + usleep(10000); + igb_error_reset_and_reinit(device); + + return -ETIMEDOUT; } =20 static void igb_send_msi(struct vfio_pci_device *device) --=20 2.55.0.795.g602f6c329a-goog