From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 62B8B30F548 for ; Wed, 22 Jul 2026 00:15:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679349; cv=none; b=tJKYffOk5V7JVD9P497ULSCfxQ0lKRZTzM7oZICo4pCDEJQSE8klgfcc9TlUNXqUSyeGUjOiukZL8Mc2Xo7wM5JgTCZIHOLskSbz+Z+dUwLlWKyBoQxZsKUkFl7+GjZexKPd6RpiOVJTpa9NXnBeBfPsi9J1A72aVCfSwpN/8aM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679349; c=relaxed/simple; bh=jTtHVqRUbIZnPNC0wsBlFw8Ojc92HvjTzXQklUzrLcA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=AsBI2SYZuMKcBpkowaK/4KW4Pt1u6HewAnEqU3kp4TvIMZvsnSeD79/jIDW3LyytE+XTfKh+9Jbu/CNOJcrdQSvM6Y/7UJ+zu0haRn7NhqGqpJTeUKbX/xnXfarBP6tkwAKPTVr8xvS8VxFEIxac4me/ic/ZRntWLcoHd6HLc/4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=eqweJjmN; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="eqweJjmN" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-ca124bf0189so3899621a12.0 for ; Tue, 21 Jul 2026 17:15:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679340; x=1785284140; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=h440Bt12cqzEoynODVhXCA/nTddru9+uYmkcNbY8bXI=; b=eqweJjmNt233AHH97OUr4XWOFDFtdKnov6UugFPKY9EHX3WjhXEDF+6nFSxL+N6D/l GOFJdtq5hLljAqWjz323DXxWS7sD8XqAj7dK4Yhwl+q+0gGzrvBvv+s8jm9D3Mnc58Cz qw1lQHldlIwePkNAaWQfKC3W2rvreU0nS42+axeSLi0YLGobcPglF2g4jZu5/yIe4Yhq aZc3ZJmQUf36LdzYwO41Otbsde3SlKxFThS0x4cPcTrvONxsbYw6j8m1+bQRJnFuA3iM xUkqsqd1ZFPoAiCG6xdpCosGIbT5lsIE6asTCWNIlgjKvsv9vdv9+hdBBsiTrPVGKXUf Dm2w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679340; x=1785284140; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=h440Bt12cqzEoynODVhXCA/nTddru9+uYmkcNbY8bXI=; b=q7tjJQUwsbYLY51xqjoRfdEUh8Y8QedK81Mk9/VfWzoXVAtHieFJGyf5RjbsuTeEuq z+TfsZLHyYyzv3X+xa/k9iCFO9Jg8Y+4zENIOByuVPf6T0PoyqSZusSBlDKHkP7nSTFr zst1g3l6hzOCqzA///niBSlXr7ciBJeB2sLHknfyG1Z8lbdh65jM32wv9KkiyOvqChHa W4Wx2JOgNiqLWRPSdxZULGE1sQ8VYndYtp7bblUtXjMwXHAwpZAR+WeXiJ4tLyGkccwm 4vu75TnzEwzqu6p/05l666jhtHLYVJily2Nr7GTzAtmuF6vbkAlh9l9BaXniOdHEMUy7 p7wA== X-Forwarded-Encrypted: i=1; AHgh+Rp7acfdz93U8cHYPWpOTeknj/YdFXiv4gXinLOJ0wUIYaQ/EuKOrE4urCd5q+00Cx4avt1IMyVYk8UCb4s=@vger.kernel.org X-Gm-Message-State: AOJu0Ywq4w+Km4Rg8iQSswuu8nkIyJ0jNV3Vzdn4MJ9BzMtjy4r5qigY GYvUv30rbz6nv240h5yTkJXGI0uU462AxVwymRUpH3h0gT+3IehehcQYZ2AvJbuIR1wfkOxxuae S0KHgTPG6 X-Received: from pjwt15.prod.google.com ([2002:a17:90a:d14f:b0:38e:74f4:7e63]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:2151:b0:38e:4e61:c9e with SMTP id 98e67ed59e1d1-38ea3730563mr943006a91.21.1784679339966; Tue, 21 Jul 2026 17:15:39 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:36 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=19492; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=jTtHVqRUbIZnPNC0wsBlFw8Ojc92HvjTzXQklUzrLcA=; b=+FFMhVBCU0xpDXMNMfThmMwjh00uYhAWnmPqfY0B9eMW/0pxTNlZsiWlAtEkGW5R3Q965fLrx 6GO7RfPdwtAC9I3MSsAY3jeK3DYIGf6SQ2hEUVpwsFBtOyt0CZ/4HqM X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-1-0be1d29918d4@google.com> Subject: [PATCH v6 1/6] vfio: selftests: igb: Add driver for Intel 82576 device From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Add a VFIO selftest driver for the Intel Gigabit Ethernet controller (IGB), specifically targeting the 82576 device. IGB is fully virtualized in QEMU which makes it easy to run VFIO selftests without needing any specific hardware. The driver uses advanced descriptors for transmit and receive, configuring SRRCTL.DESCTYPE to the advanced one-buffer layout for receive and advanced data descriptor fields in igb_memcpy_start() for transmit. It also programs the full MSI-X routing sequence: - GPIE.Multiple_MSIX to route causes through IVAR. - GPIE.EIAME to apply EIAM on MSI-X assertion. - EIAC to enable auto-clear of EICR for vector 0. - EIAM to enable auto-mask of EIMS for vector 0 on MSI-X assertion. - IVAR to map RX cause 0 to MSI-X vector 0. Write-to-clear is used for EICR as required when EIAC is enabled. To support real 82576 hardware which processes the descriptor ring at line rate, the memcpy completion timeout in igb_memcpy_wait() is set to 200 ms (200 retries of 1ms). At 1 Gb/s (~125 MB/s) the worst valid memcpy (~4 MB) takes ~32 ms on the wire, plus overhead (~3%) and latency. A 200 ms timeout leaves comfortable headroom for host scheduling jitter while keeping intentional invalid-DMA tests bounded. Additionally, DMA re-send on PCIe completion timeout is disabled by clearing GCR.Completion_Timeout_Resend (datasheet section 8.6.1, bit 16). The mix_and_match test intentionally submits descriptors targeting unmapped IOVAs; with the default value, the device retries the failed read indefinitely, which keeps PCIe AER and IOMMU error handling busy and interferes with reset recovery. Co-developed-by: Alex Williamson Signed-off-by: Alex Williamson Signed-off-by: Josh Hilke Reviewed-by: David Matlack --- .../selftests/vfio/lib/drivers/igb/e1000_82575.h | 1 + .../selftests/vfio/lib/drivers/igb/e1000_defines.h | 1 + .../selftests/vfio/lib/drivers/igb/e1000_regs.h | 1 + tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 475 +++++++++++++++++= ++++ tools/testing/selftests/vfio/lib/libvfio.mk | 1 + tools/testing/selftests/vfio/lib/vfio_pci_driver.c | 2 + 6 files changed, 481 insertions(+) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h b/t= ools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h new file mode 120000 index 000000000000..b84affdec559 --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_82575.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_82575.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h b= /tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h new file mode 120000 index 000000000000..9f97f4330086 --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_defines.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_defines.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h b/to= ols/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h new file mode 120000 index 000000000000..c733634171bb --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/e1000_regs.h @@ -0,0 +1 @@ +../../../../../../../drivers/net/ethernet/intel/igb/e1000_regs.h \ No newline at end of file diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c new file mode 100644 index 000000000000..a59b95303092 --- /dev/null +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -0,0 +1,475 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "e1000_regs.h" +#include "e1000_defines.h" +#include "e1000_82575.h" + +#define PCI_DEVICE_ID_INTEL_82576 0x10C9 +#define IGB_MAX_CHUNK_SIZE 1024 +#define MSIX_VECTOR 0 +#define MSIX_VECTOR_MASK (1 << MSIX_VECTOR) +#define RING_SIZE 4096 /* Number of descriptors in ring */ + +struct igb_tx_desc { + union { + struct { + u64 buffer_addr; /* Address of descriptor's data buffer */ + u32 cmd_type_len; /* Command/Type/Length */ + u32 olinfo_status; /* Context/Buffer info */ + } read; + + struct { + u64 rsvd; /* Reserved */ + u32 nxtseq_seed; /* Next sequence seed */ + u32 status; /* Descriptor status */ + } wb; + }; +}; + +struct igb_rx_desc { + union { + struct { + u64 pkt_addr; /* Packet buffer address */ + u64 hdr_addr; /* Header buffer address */ + } read; + struct { + u16 pkt_info; /* RSS type, Packet type */ + u16 hdr_info; /* Split Head, buf len */ + u32 rss; /* RSS Hash */ + u32 status_error; /* ext status/error */ + u16 length; /* Packet length */ + u16 vlan; /* VLAN tag */ + } wb; /* writeback */ + }; +}; + +struct igb { + void *bar0; + u32 tx_tail; + u32 rx_tail; + struct igb_tx_desc tx_ring[RING_SIZE] __attribute__((aligned(128))); + struct igb_rx_desc rx_ring[RING_SIZE] __attribute__((aligned(128))); +}; + +static inline struct igb *to_igb_state(struct vfio_pci_device *device) +{ + return (struct igb *)device->driver.region.vaddr; +} + +static inline void igb_write32(struct igb *igb, u32 reg, u32 val) +{ + writel(val, igb->bar0 + reg); +} + +static inline u32 igb_read32(struct igb *igb, u32 reg) +{ + return readl(igb->bar0 + reg); +} + +static int igb_write_phy(struct igb *igb, u32 offset, u16 data) +{ + u32 mdic; + int i; + + /* + * Write a PHY register over MDIO. + * + * A production driver would hold the SW/FW semaphore (SWSM.SWESMBI + the + * SW_FW_SYNC PHY bit) across the MDIO transaction to serialize against t= he + * device's management firmware. The selftest owns the assigned function + * exclusively on a dedicated test device with no active manageability + * contending for the PHY, so the sync is omitted; it should be added here + * if this ever needs to run on a manageability-enabled NIC. + */ + mdic =3D (((u32)data) | + (offset << E1000_MDIC_REG_SHIFT) | + (1 << E1000_MDIC_PHY_SHIFT) | + E1000_MDIC_OP_WRITE); + + igb_write32(igb, E1000_MDIC, mdic); + + for (i =3D 0; i < 1000; i++) { + usleep(50); + mdic =3D igb_read32(igb, E1000_MDIC); + if (mdic & E1000_MDIC_READY) + break; + } + + if (!(mdic & E1000_MDIC_READY)) + return -1; + + if (mdic & E1000_MDIC_ERROR) + return -1; + + return 0; +} + +static int igb_read_phy(struct igb *igb, u32 offset, u16 *data) +{ + u32 mdic; + int i; + + mdic =3D ((offset << E1000_MDIC_REG_SHIFT) | + (1 << E1000_MDIC_PHY_SHIFT) | + E1000_MDIC_OP_READ); + + igb_write32(igb, E1000_MDIC, mdic); + + for (i =3D 0; i < 1000; i++) { + usleep(50); + mdic =3D igb_read32(igb, E1000_MDIC); + if (mdic & E1000_MDIC_READY) + break; + } + + if (!(mdic & E1000_MDIC_READY)) + return -1; + + if (mdic & E1000_MDIC_ERROR) + return -1; + + *data =3D (u16)mdic; + return 0; +} + +static void igb_phy_setup_autoneg(struct igb *igb) +{ + int timeout_ms =3D 1000; + bool success =3D false; + u16 phy_status; + int ret; + int i; + + /* Trigger auto-negotiation */ + ret =3D igb_write_phy(igb, MII_BMCR, + BMCR_ANENABLE | BMCR_ANRESTART); + VFIO_ASSERT_EQ(ret, 0, "Failed to write PHY control register"); + + for (i =3D 0; i < timeout_ms; i++) { + if (igb_read_phy(igb, MII_BMSR, &phy_status) =3D=3D 0) { + success =3D !!(phy_status & BMSR_ANEGCOMPLETE); + if (success) + break; + } + usleep(1000); + } + + VFIO_ASSERT_TRUE(success, "Auto-negotiation did not complete in time"); +} + +static int igb_probe(struct vfio_pci_device *device) +{ + if (!vfio_pci_device_match(device, PCI_VENDOR_ID_INTEL, PCI_DEVICE_ID_INT= EL_82576)) + return -EINVAL; + + return 0; +} + +static void igb_reset(struct igb *igb) +{ + igb_write32(igb, E1000_CTRL, igb_read32(igb, E1000_CTRL) | E1000_CTRL_RST= ); + /* + * Must wait at least 1 millisecond after setting the reset bit before + * checking if this device is ready to be used (82576 datasheet section + * 4.2.1.6.1). + */ + usleep(2000); + VFIO_ASSERT_EQ(igb_read32(igb, E1000_CTRL) & E1000_CTRL_RST, 0); + igb_write32(igb, E1000_IMC, 0xFFFFFFFF); +} + +static void igb_init(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + u64 iova_tx, iova_rx; + u32 ctrl, rctl; + u16 cmd_reg; + int retries; + + VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); + + /* Set up rings and calculate IOVAs */ + igb->bar0 =3D device->bars[0].vaddr; + + iova_tx =3D to_iova(device, igb->tx_ring); + iova_rx =3D to_iova(device, igb->rx_ring); + + igb_reset(igb); + + /* Signal that the driver is loaded */ + ctrl =3D igb_read32(igb, E1000_CTRL_EXT); + ctrl |=3D E1000_CTRL_EXT_DRV_LOAD; + ctrl &=3D ~E1000_CTRL_EXT_LINK_MODE_MASK; + igb_write32(igb, E1000_CTRL_EXT, ctrl); + + /* Enable PCI Bus Master. */ + cmd_reg =3D vfio_pci_config_readw(device, PCI_COMMAND); + if ((cmd_reg & (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY)) !=3D + (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY)) { + cmd_reg |=3D (PCI_COMMAND_MASTER | PCI_COMMAND_MEMORY); + vfio_pci_config_writew(device, PCI_COMMAND, cmd_reg); + } + + /* Trigger autonegotiation. This enables IGB to transmit data. */ + igb_phy_setup_autoneg(igb); + + /* + * Disable DMA re-send on PCIe completion timeout (82576 datasheet + * section 8.6.1, GCR.Completion_Timeout_Resend, bit 16). The + * mix_and_match test intentionally submits descriptors targeting + * unmapped IOVAs; with the default (set) value, the device keeps + * retrying the failed read indefinitely, which keeps PCIe AER and + * IOMMU error handling busy and interferes with reset recovery. + */ + ctrl =3D igb_read32(igb, E1000_GCR); + ctrl &=3D ~E1000_GCR_CMPL_TMOUT_RESEND; + igb_write32(igb, E1000_GCR, ctrl); + + /* Configure TX and RX descriptor rings */ + igb_write32(igb, E1000_TDBAL(0), (u32)iova_tx); + igb_write32(igb, E1000_TDBAH(0), (u32)(iova_tx >> 32)); + igb_write32(igb, E1000_TDLEN(0), RING_SIZE * sizeof(struct igb_tx_desc)); + igb_write32(igb, E1000_TDH(0), 0); + igb_write32(igb, E1000_TDT(0), 0); + igb_write32(igb, E1000_TXDCTL(0), E1000_TXDCTL_QUEUE_ENABLE); + + igb_write32(igb, E1000_RDBAL(0), (u32)iova_rx); + igb_write32(igb, E1000_RDBAH(0), (u32)(iova_rx >> 32)); + igb_write32(igb, E1000_RDLEN(0), RING_SIZE * sizeof(struct igb_rx_desc)); + igb_write32(igb, E1000_RDH(0), 0); + igb_write32(igb, E1000_RDT(0), 0); + + /* + * Select the advanced one-buffer descriptor format. Per 82576 + * datasheet section 7.1.5.2: "SRRCTL[n].DESCTYPE must be set to a + * value other than 000b for the 82576 to write back the special + * descriptors." struct igb_rx_desc matches the advanced one-buffer + * writeback layout (section 7.1.5.2), so polling rx.wb.status_error + * requires this format. Section 8.10.2 specifies DESCTYPE[27:25]. + * + * The direct write also zeroes SRRCTL.BSIZEPACKET, which is + * intentional: per section 7.1.3.1 a zero BSIZEPACKET falls back to + * the RCTL.BSIZE buffer size, whose reset default (00b) is 2048 + * bytes -- ample for the loopback frames here. + */ + igb_write32(igb, E1000_SRRCTL(0), E1000_SRRCTL_DESCTYPE_ADV_ONEBUF); + + igb_write32(igb, E1000_RXDCTL(0), E1000_RXDCTL_QUEUE_ENABLE); + + /* Wait for TX and RX queues to be enabled */ + retries =3D 2000; + while (retries-- > 0) { + if ((igb_read32(igb, E1000_TXDCTL(0)) & E1000_TXDCTL_QUEUE_ENABLE) && + (igb_read32(igb, E1000_RXDCTL(0)) & E1000_RXDCTL_QUEUE_ENABLE)) + break; + usleep(10); + } + VFIO_ASSERT_GE(retries, 0); + + /* Enable Receiver and Transmitter */ + rctl =3D E1000_RCTL_EN | /* Receiver Enable */ + E1000_RCTL_UPE | /* Unicast Promiscuous (for dummy MAC) */ + E1000_RCTL_MPE | /* Multicast Promiscuous */ + E1000_RCTL_BAM | /* Broadcast Accept Mode */ + E1000_RCTL_LBM_MAC | /* MAC Loopback Mode */ + E1000_RCTL_SECRC; /* Strip CRC (needed for memcmp) */ + igb_write32(igb, E1000_RCTL, rctl); + igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); + + /* Enable MSI-X with 1 vector for the test */ + vfio_pci_msix_enable(device, MSIX_VECTOR, 1); + + /* + * Program MSI-X interrupt routing per 82576 datasheet: + * + * GPIE (section 7.3.2.11, Table 7-47): set Multiple_MSIX (bit 4) to + * route interrupt causes through IVAR mapping, and EIAME (bit 30) + * to apply EIAM on MSI-X assertion (without EIAME, EIAM only + * applies on EICR read/write). + * + * EIAC (section 8.8.5): enable auto-clear of EICR for vector 0. + * Without auto-clear the cause stays set after delivery and the + * test can see spurious interrupts on the next memcpy batch. + * + * EIAM (section 8.8.6): enable auto-mask of EIMS for vector 0 on + * MSI-X assertion (effective because EIAME is set), so a single + * interrupt is delivered per memcpy batch even if the cause + * re-asserts before software re-enables the mask. + * + * IVAR (section 7.3.1.2, register definition in 8.8.13): map RX + * cause 0 to MSI-X vector 0 and mark the entry valid. + */ + igb_write32(igb, E1000_GPIE, E1000_GPIE_MSIX_MODE | E1000_GPIE_EIAME); + igb_write32(igb, E1000_EIAC, MSIX_VECTOR_MASK); + igb_write32(igb, E1000_EIAM, MSIX_VECTOR_MASK); + + /* Map vector 0 to interrupt cause 0 and mark it valid */ + igb_write32(igb, E1000_IVAR0, E1000_IVAR_VALID); + + /* Enable interrupts on vector 0 */ + igb_write32(igb, E1000_EIMS, MSIX_VECTOR_MASK); + + /* Initialize driver state and capability limits */ + igb->tx_tail =3D 0; + igb->rx_tail =3D 0; + + device->driver.max_memcpy_size =3D IGB_MAX_CHUNK_SIZE; + device->driver.max_memcpy_count =3D RING_SIZE - 1; + device->driver.msi =3D MSIX_VECTOR; +} + +static void igb_remove(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + igb_write32(igb, E1000_RCTL, 0); + igb_write32(igb, E1000_TCTL, 0); + igb_reset(igb); + + vfio_pci_msix_disable(device); +} + +static void igb_irq_disable(struct igb *igb) +{ + igb_write32(igb, E1000_EIMC, MSIX_VECTOR_MASK); +} + +static void igb_irq_enable(struct igb *igb) +{ + igb_write32(igb, E1000_EIMS, MSIX_VECTOR_MASK); +} + +static void igb_irq_clear(struct igb *igb) +{ + /* + * Use write-to-clear (datasheet 7.3.4.2). In MSI-X mode with EIAC + * programmed, section 8.8.5 explicitly states "If any bits are set + * in EIAC, the EICR register should not be read", which rules out + * the read-to-clear path in 7.3.4.3. Bits not in EIAC are still + * cleared by writing 1. + */ + igb_write32(igb, E1000_EICR, 0xFFFFFFFF); +} + +static void igb_memcpy_start(struct vfio_pci_device *device, iova_t src, + iova_t dst, u64 size, u64 count) +{ + struct igb *igb =3D to_igb_state(device); + struct igb_rx_desc *rx; + struct igb_tx_desc *tx; + u32 i; + + igb_irq_disable(igb); + + for (i =3D 0; i < count; i++) { + tx =3D &igb->tx_ring[igb->tx_tail]; + rx =3D &igb->rx_ring[igb->rx_tail]; + + memset(tx, 0, sizeof(struct igb_tx_desc)); + memset(rx, 0, sizeof(struct igb_rx_desc)); + + rx->read.pkt_addr =3D cpu_to_le64(dst); + rx->read.hdr_addr =3D cpu_to_le64(0); + + tx->read.buffer_addr =3D cpu_to_le64(src); + /* + * Build an advanced data descriptor per 82576 datasheet + * section 7.2.2.3. DEXT marks the descriptor as advanced + * (required by hardware); DTYP=3Ddata selects the data + * descriptor; IFCS asks the MAC to append the Ethernet + * FCS (without it the frame is dropped as malformed); + * EOP marks end of packet. DTALEN is the buffer length + * in bits 15:0 of cmd_type_len. + */ + tx->read.cmd_type_len =3D cpu_to_le32((uint32_t)size | + E1000_ADVTXD_DTYP_DATA | + E1000_ADVTXD_DCMD_DEXT | + E1000_ADVTXD_DCMD_IFCS | + E1000_ADVTXD_DCMD_EOP); + /* + * PAYLEN (section 7.2.2.3.11) is the total payload size + * in olinfo_status[31:14]. + */ + tx->read.olinfo_status =3D + cpu_to_le32((uint32_t)size << E1000_ADVTXD_PAYLEN_SHIFT); + + igb->tx_tail =3D (igb->tx_tail + 1) % RING_SIZE; + igb->rx_tail =3D (igb->rx_tail + 1) % RING_SIZE; + } + + igb_write32(igb, E1000_RDT(0), igb->rx_tail); + igb_write32(igb, E1000_TDT(0), igb->tx_tail); +} + +static int igb_memcpy_wait(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + struct igb_rx_desc *rx; + u32 status =3D 0; + u32 prev_tail; + int retries; + + prev_tail =3D (igb->rx_tail + RING_SIZE - 1) % RING_SIZE; + rx =3D &igb->rx_ring[prev_tail]; + + /* + * Real 82576 hardware processes the descriptor ring at line rate. + * max_memcpy_size =3D (RING_SIZE - 1) * IGB_MAX_CHUNK_SIZE ~=3D 4 MB, + * split into 4095 1 KB frames. At 1 Gb/s (~125 MB/s) the worst + * valid memcpy takes ~32 ms on the wire, plus per-frame preamble, + * SFD, IFG and FCS overhead (~3%) and descriptor fetch/writeback + * latency. Wait up to ~200 ms before declaring the device hung; + * ~6x the line-rate floor leaves comfortable headroom for host + * scheduling jitter while keeping the intentional invalid-DMA + * tests bounded. + */ + retries =3D 200; + while (retries-- > 0) { + status =3D le32_to_cpu(READ_ONCE(rx->wb.status_error)); + if (status & 1) + break; + usleep(1000); + } + + if (status & 1) + /* + * Ensure the test code doesn't speculatively read the DMA + * destination buffer before we have verified that the + * descriptor writeback is complete. + */ + rmb(); + + igb_irq_clear(igb); + + igb_irq_enable(igb); + + return (status & 1) ? 0 : -ETIMEDOUT; +} + +static void igb_send_msi(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + igb_write32(igb, E1000_EICS, MSIX_VECTOR_MASK); +} + +const struct vfio_pci_driver_ops igb_ops =3D { + .name =3D "igb", + .probe =3D igb_probe, + .init =3D igb_init, + .remove =3D igb_remove, + .memcpy_start =3D igb_memcpy_start, + .memcpy_wait =3D igb_memcpy_wait, + .send_msi =3D igb_send_msi, +}; diff --git a/tools/testing/selftests/vfio/lib/libvfio.mk b/tools/testing/se= lftests/vfio/lib/libvfio.mk index 9f47bceed16f..1f13cca04348 100644 --- a/tools/testing/selftests/vfio/lib/libvfio.mk +++ b/tools/testing/selftests/vfio/lib/libvfio.mk @@ -12,6 +12,7 @@ LIBVFIO_C +=3D vfio_pci_driver.c ifeq ($(ARCH:x86_64=3Dx86),x86) LIBVFIO_C +=3D drivers/ioat/ioat.c LIBVFIO_C +=3D drivers/dsa/dsa.c +LIBVFIO_C +=3D drivers/igb/igb.c endif =20 LIBVFIO_OUTPUT :=3D $(OUTPUT)/libvfio diff --git a/tools/testing/selftests/vfio/lib/vfio_pci_driver.c b/tools/tes= ting/selftests/vfio/lib/vfio_pci_driver.c index 6827f4a6febe..a5d0547132c4 100644 --- a/tools/testing/selftests/vfio/lib/vfio_pci_driver.c +++ b/tools/testing/selftests/vfio/lib/vfio_pci_driver.c @@ -5,12 +5,14 @@ #ifdef __x86_64__ extern struct vfio_pci_driver_ops dsa_ops; extern struct vfio_pci_driver_ops ioat_ops; +extern struct vfio_pci_driver_ops igb_ops; #endif =20 static struct vfio_pci_driver_ops *driver_ops[] =3D { #ifdef __x86_64__ &dsa_ops, &ioat_ops, + &igb_ops, #endif }; =20 --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 173AA310645 for ; Wed, 22 Jul 2026 00:15:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679348; cv=none; b=kUfPG6mZ46XbYgpsnmq9LPdtcpvP942NgS4J4eRwQ+LTjvxBEd8rO/t5Dw3G9+n9laCrm6GQmp/qN6+Xi6SI4WZysXAa73lq2tCiGB+9hlOf0abyBlMglCoS18WNaS0XDaNazBBgf2cHGYCYaB9lZtFq8JN90v/0eCsQsMGKR38= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679348; c=relaxed/simple; bh=D5CB7KCNOjXbUnqfcG1gEDz35xe9CIrcyj4jFdGxWq8=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=KuaZTqUgygjq7o9s3+CPj7H8NU5q516QCMNfSN8XLQsz4oZXS4D7R2iSxfQU/mniviA6yZoiofmCnhtmNUaS/BooJywfMiPsXpzhO9Bnu7U6rp4PpoFOcN7ndaXz07baczOXP3zIUxwReHSQ5bnI9ufXabMz1+Ci7AErbIGkeMI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=YR1uCo+U; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="YR1uCo+U" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-cb5cc1e139bso2838870a12.3 for ; Tue, 21 Jul 2026 17:15:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679341; x=1785284141; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=TzRoo9k7XnIWz4Y6r+tLwSZ0Adb1W51xHlr3qTd/Bqk=; b=YR1uCo+UCUmY3KP44yNsm6SnZkbOg896CdK+WJewzCvmD3UP6AkPCFRavFwwXFZunL DLxymR1+8YPUL2KW1f9VJ6fLnIuOYpzY+OWs+yFeWfwR96+rAe7j74nFmbyuufBOQk1b d2VlM+mNwEYIQyCCFxrWm1EVOQ/CYNLxEMU1ukLjU4ee0oiT1aDm6LZeNouYUwpK4Xto 4M99dGomizh1GwzJlBWc7UnqwjjNiSoYyqdB2y6M8wNVeOmG90wwPpfNQDtsH21H3ZZI GxGCuSsNw4+NeU2JdPTJM9ohLvNnmm51u1EVqSbJ3gCQ1xh/x0WfNGfxD1UG5hNq2Dud JJSg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679341; x=1785284141; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=TzRoo9k7XnIWz4Y6r+tLwSZ0Adb1W51xHlr3qTd/Bqk=; b=Uf5FQmE//vm3mYIARLusglZ/YxTvUUN+yHU3rn2oDhIkmBmK3xQZqjUepe0YYkp52U 0J2UJxXLidYgu4ZKc99kSvWgg7HQf5dAdRsRwi8/9GtIF4ddI683e0TQ8plpmzFqhsnf 8IueeeBip+B8aXZl8FZ9cLewroFgXkD0Getmo+W4VqR45yRbKeaj0Tn2NErSg6hQ6u/5 7uj3vqIRu/aN85XsLZfhZuTd+F7o0f5Xuw/rSt+pPGJZcc0iRS2cdVJDvxcmuGcVHaJk MqweL3BwnfBOiX+UK1rGJZHLQKLNIO6zbQVj/NkKSV5IEg+mNVUZ3xzZfCMkcfeSyNeW kXQA== X-Forwarded-Encrypted: i=1; AHgh+Rqr2/7+0u/J+bn6xPgUvZjSwwR8SNmB2MwX2IXVOpPhFYYXOb5l8rG/FCHzGFDYLm2hwciryNqRhza8aVA=@vger.kernel.org X-Gm-Message-State: AOJu0YwUg/gmYz4MEdxo+13UCSGCN/ZkHZXqzQH4RJTk0ipQH8FrKdTi IIZa73zd6JUE6n53uqTVmW9zYGPW8sEDlhbDA4ct0CQn6mmH7ffC6Q5ecmxd5uEqQGNDbMJ4rgw 5AaShdw+G X-Received: from pgbr21.prod.google.com ([2002:a63:5d15:0:b0:c94:9604:c8da]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:3983:b0:3c0:b766:751b with SMTP id adf61e73a8af0-3c3ad6777damr20212107637.14.1784679341018; Tue, 21 Jul 2026 17:15:41 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:37 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=8558; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=OGd8khMJFUVJe0i+6kv3miFK/tHU66vYeXc7UKHP5TY=; b=ufUkEU704J8quhs0BU4NfXeokqMtHqnE0F+gnQDxVD8MSrXvsAy+rW9eYtLHJlBHfAw8aykKB YuR/wyBSiaFBRFkhdIzcDM81+ZnMwCGwHM3xc+NXRCnwX9WGUEtYOEg X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-2-0be1d29918d4@google.com> Subject: [PATCH v6 2/6] vfio: selftests: igb: Use PHY internal loopback on 82576 From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The submitted driver waits for PHY autonegotiation and then enables MAC loopback via RCTL.LBM_MAC. QEMU's emulated igb tolerates this, but the 82576 datasheet rejects the MAC loopback path on real hardware: Section 3.5.6.1 (Loopback Support / General): "Use PHY Loopback instead of MAC Loopback on the 82576." Section 3.5.6.2 (MAC Loopback): "MAC Loopback is not used on this device." Section 3.5.6.3.1 (Setting the 82576 to PHY loopback Mode): set PHY control register bits 8 (duplex), 14 (loopback), clear bit 12 (autoneg enable), set the speed via bits 6 and 13. For 1Gb/s the register value is 0x4140. Section 8.10.1 (RCTL register): "When using the internal PHY, LBM should remain set to 00b and the PHY instead configured for loopback through the MDIO interface." Replace igb_phy_setup_autoneg() with igb_setup_loopback() which: - writes PHY register 0 with LOOPBACK | SPEED_1000 | FULL_DUPLEX - forces the MAC into 1Gb/s full duplex via CTRL.FRCSPD, CTRL.FRCDPX, CTRL.SPD_1000, CTRL.FD, CTRL.SLU; without forcing the MAC link state, the descriptor engine does not run on real hardware PHY internal loopback (section 3.5.6.3) wraps data at the end of the PHY datapath before the MDI, so the physical link state and cable speed are irrelevant. This matches the kernel ethtool selftest path in igb_integrated_phy_loopback() (drivers/net/ethernet/intel/igb/igb_ethtool.c). QEMU's igb emulation drives STATUS.LU exclusively from its autoneg-done timer or a network-backend link-state change; the guest cannot set STATUS.LU through CTRL.SLU. Its receive path checks STATUS.LU (e1000x_hw_rx_enabled in hw/net/e1000x_common.c) and drops every loopback frame until LU is set. Issue a one-shot autoneg-restart PHY write at the top of igb_setup_loopback() to kick the timer; the subsequent PHY write clears autoneg-enable, so on real hardware autoneg never starts and the write is a no-op. QEMU's igb also does not honor PHY register 0 bit 14 (PHY internal loopback) and relies on RCTL.LBM_MAC to wrap TX descriptors back to the RX queue. Datasheet 8.10.1 advises that LBM remain 00b when using the internal PHY, but empirically setting LBM_MAC has no observable effect on real 82576 (MAC loopback is not implemented per 3.5.6.2), so set it alongside PHY loopback as the actual loopback mechanism under QEMU. With these two QEMU-only accommodations the selftest works in both environments without environment-specific code paths. Remove igb_read_phy() which becomes unused. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 121 +++++++++++++----= ---- 1 file changed, 73 insertions(+), 48 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index a59b95303092..2b444be4bdfe 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -115,57 +115,70 @@ static int igb_write_phy(struct igb *igb, u32 offset,= u16 data) return 0; } =20 -static int igb_read_phy(struct igb *igb, u32 offset, u16 *data) +/* + * Configure the device for PHY internal loopback per 82576 datasheet + * section 3.5.6.3.1. Force the PHY to 1Gb/s full duplex with loopback + * enabled, then force the MAC link state to match. Internal loopback + * wraps data at the end of the PHY datapath (section 3.5.6.3), so the + * physical link state is irrelevant. + * + * Section 3.5.6.1 directs to "Use PHY Loopback instead of MAC Loopback + * on the 82576", and section 3.5.6.2 states "MAC Loopback is not used + * on this device." RCTL.LBM_MAC is still set elsewhere as a QEMU-only + * accommodation; see the RCTL programming in the caller for the + * rationale. + */ +static void igb_setup_loopback(struct igb *igb) { - u32 mdic; - int i; - - mdic =3D ((offset << E1000_MDIC_REG_SHIFT) | - (1 << E1000_MDIC_PHY_SHIFT) | - E1000_MDIC_OP_READ); - - igb_write32(igb, E1000_MDIC, mdic); - - for (i =3D 0; i < 1000; i++) { - usleep(50); - mdic =3D igb_read32(igb, E1000_MDIC); - if (mdic & E1000_MDIC_READY) - break; - } - - if (!(mdic & E1000_MDIC_READY)) - return -1; - - if (mdic & E1000_MDIC_ERROR) - return -1; - - *data =3D (u16)mdic; - return 0; -} - -static void igb_phy_setup_autoneg(struct igb *igb) -{ - int timeout_ms =3D 1000; - bool success =3D false; - u16 phy_status; + u32 ctrl; int ret; - int i; =20 - /* Trigger auto-negotiation */ - ret =3D igb_write_phy(igb, MII_BMCR, + /* + * Kick the autoneg machinery solely to bring STATUS.LU up under + * QEMU's igb emulation: QEMU only updates STATUS.LU via its + * autoneg-done timer, and without LU set its receive path + * (e1000x_hw_rx_enabled) drops every loopback frame. On real + * hardware autoneg cannot complete before the next PHY write + * below clears the autoneg-enable bit, so this is effectively a + * no-op there. + */ + (void)igb_write_phy(igb, MII_BMCR, BMCR_ANENABLE | BMCR_ANRESTART); + + /* PHY control: loopback + 1Gb/s full duplex, autoneg disabled. */ + ret =3D igb_write_phy(igb, MII_BMCR, + BMCR_LOOPBACK | + BMCR_SPEED1000 | + BMCR_FULLDPLX); VFIO_ASSERT_EQ(ret, 0, "Failed to write PHY control register"); =20 - for (i =3D 0; i < timeout_ms; i++) { - if (igb_read_phy(igb, MII_BMSR, &phy_status) =3D=3D 0) { - success =3D !!(phy_status & BMSR_ANEGCOMPLETE); - if (success) - break; - } - usleep(1000); - } + /* + * Brief delay before forcing the MAC, mirroring the kernel ethtool + * selftest in igb_integrated_phy_loopback(). Not specified by the + * datasheet, but empirically required by the kernel driver. + */ + usleep(50000); + + /* + * Force the MAC to 1Gb/s full duplex with link up. Without forcing + * the link state the descriptor engine does not run, since the chip + * normally waits for a real negotiated link. + */ + ctrl =3D igb_read32(igb, E1000_CTRL); + ctrl &=3D ~E1000_CTRL_SPD_SEL; + ctrl |=3D E1000_CTRL_FRCSPD | + E1000_CTRL_FRCDPX | + E1000_CTRL_SPD_1000 | + E1000_CTRL_FD | + E1000_CTRL_SLU; + igb_write32(igb, E1000_CTRL, ctrl); =20 - VFIO_ASSERT_TRUE(success, "Auto-negotiation did not complete in time"); + /* + * Settling delay matching the kernel ethtool selftest's msleep(500) + * at the tail of igb_integrated_phy_loopback(). Not specified by + * the datasheet; empirical, and inherited from the kernel driver. + */ + usleep(500000); } =20 static int igb_probe(struct vfio_pci_device *device) @@ -221,8 +234,8 @@ static void igb_init(struct vfio_pci_device *device) vfio_pci_config_writew(device, PCI_COMMAND, cmd_reg); } =20 - /* Trigger autonegotiation. This enables IGB to transmit data. */ - igb_phy_setup_autoneg(igb); + /* Configure PHY internal loopback for testing. */ + igb_setup_loopback(igb); =20 /* * Disable DMA re-send on PCIe completion timeout (82576 datasheet @@ -277,12 +290,24 @@ static void igb_init(struct vfio_pci_device *device) } VFIO_ASSERT_GE(retries, 0); =20 - /* Enable Receiver and Transmitter */ + /* + * Enable Receiver and Transmitter. RCTL.LBM_MAC is set in addition + * to PHY loopback as a QEMU-only accommodation: QEMU's emulated igb + * does not honor PHY register 0 bit 14 (PHY internal loopback) and + * relies on RCTL.LBM_MAC to wrap TX descriptors back to the RX + * queue. Datasheet 8.10.1 (RCTL register) advises "When using the + * internal PHY, LBM should remain set to 00b", so setting LBM_MAC + * here deviates from datasheet guidance; empirically the bit has + * no observable effect on real 82576 hardware because MAC loopback + * is not implemented (datasheet 3.5.6.2). Setting both lets the + * selftest work on both real hardware and QEMU without conditional + * code paths. + */ rctl =3D E1000_RCTL_EN | /* Receiver Enable */ E1000_RCTL_UPE | /* Unicast Promiscuous (for dummy MAC) */ E1000_RCTL_MPE | /* Multicast Promiscuous */ E1000_RCTL_BAM | /* Broadcast Accept Mode */ - E1000_RCTL_LBM_MAC | /* MAC Loopback Mode */ + E1000_RCTL_LBM_MAC | /* MAC Loopback - for QEMU emulation only */ E1000_RCTL_SECRC; /* Strip CRC (needed for memcmp) */ igb_write32(igb, E1000_RCTL, rctl); igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pf1-f197.google.com (mail-pf1-f197.google.com [209.85.210.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0B2C1313535 for ; Wed, 22 Jul 2026 00:15:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679350; cv=none; b=vGi4M6rGGD4KuuUY+zygkLIaaL8+j/hmfZqJGNsbI3iXlsZTz6LsEkOCXTlzJIKn6XjPk8XArSb+jjA3uanRpMsJHNT/70vDloHtmdTF2s0BibdSaWODU04VolqwwTGVxJGltRUXQ9o1K2oCHQ4FJJz9jt1H0ijME50GRJrtGJg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679350; c=relaxed/simple; bh=9dzLlS9AuJaT6xRIQC67AQ/17hTq/GvFHM3M5jcaNRE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=RKVERYIxtVyF0uA7mkULwjovU4S1SbSGzS5PsXkIpOumn/ydWXV2TMPYovTT+Pxgc85EuVmm6iSDII+oi/ZC+/0+Md5Wf6loKhV6GUKjuTai19FtsfmsYdx/moHc67huC+qhpVzszx4kSQAC2lJ2yc/qpUC70iiyat4v/SAJfLE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ULT9WZCr; arc=none smtp.client-ip=209.85.210.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ULT9WZCr" Received: by mail-pf1-f197.google.com with SMTP id d2e1a72fcca58-84885a4fcabso9402581b3a.3 for ; Tue, 21 Jul 2026 17:15:44 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679342; x=1785284142; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=TNalfHPNu/4hggAaxfIKgXNij2eI570gq4T5V2Wg4yw=; b=ULT9WZCr9m4y9Pck0qrYCRMe+Yepy10J3GZo5mWmh97+EVHenX3afmc78bOtEHVzSY IucinAkdf3uoyYZhs9aJK9C85HN2NqBrCjxHQEBICDbJQa0VCCHNo/P00mVVRt1IwQwg C7UgpC+xrE2z5kjtIAMkt0bucMtAkGX4GRsEkE4ZNirlofXoWOLlTjGy1cEJQqx1sI0y 2SbsLcuNqG9HU3JTzCXnvjuwUrRHFqU/yhYmY3x8p426DFkt+54K4PddScQvRMPZRZmy cEKyRYnqQT3uvxD8ROlaH6Jt6gOQnbBWyQ6jKtEp7//dAptGH9kOw8pW/NA/dlamK5Tv 0T0A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679342; x=1785284142; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=TNalfHPNu/4hggAaxfIKgXNij2eI570gq4T5V2Wg4yw=; b=nNnHK0dm0+CtnYXJJuKWzROl6c3w58ycEzZyPZG16xh1whPmX1wITMAWiPBhRetUio fpWjX1pg2SU0nlKrGBT4un/YhuPjK7HKHA1ghVHLUijcXklR99Ss7p2lUODtU5L1YX5J j/D8GcfsjbQTynLDHcij5RT+dqYmTlnTtj4143NV4LsyPej7PN+/NGSl7IrFGpob5H/Q hOo9pMuYQBa35oT2pAkPLrxwoKuj38y0l0Pcxuw426uWDz7yTYvlbC1d//5hzuyOGhEk tRhm0vpgYLa+b0RmETrSqoTMcjFqN/5ZB0214sdDV2aFdqDbsN377hI/Lfpt85to4Utp mPRg== X-Forwarded-Encrypted: i=1; AHgh+RrvTxTD60CFGpLebbCvvSFWMZeYk/lXV8aZX4zoB58YgFaRFc5q8F3FPiziqEg+RORHsL7r9ZK7ipnUadE=@vger.kernel.org X-Gm-Message-State: AOJu0YzsLoJvyoaWj3HwEwnfefY7AkHEG/KulX6Gq7KDt5cKzE3uuOhh VNR6+i2/mFD02ieDxk4UgQlb0HHOueBooyXw5X7/CLIl40hyq2+fMogf6xPWbAr60u4L6TOsUhC OY09nf6aF X-Received: from pfem16.prod.google.com ([2002:a05:6a00:c090:b0:847:99d5:ef10]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:1943:b0:847:759e:f61c with SMTP id d2e1a72fcca58-84c295013fdmr21103131b3a.47.1784679341869; Tue, 21 Jul 2026 17:15:41 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:38 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=4367; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=AJ0ZiImWQV3xIvutXWoT8SQkfP3+vR/ectign90lbW8=; b=Ma1Ca3e8tybbqM2WNpPBXIYN+xtY78hltH796r3rOkE+EJM5O2/kOmRwFvOL4KkAuhoTrcfJa Ztf4iGedGlcAKY2nsPUHfNqw8mgVpAVHckB8QkOpvh+LA990bo+Gur2 X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-3-0be1d29918d4@google.com> Subject: [PATCH v6 3/6] vfio: selftests: Add helpers to re-enable interrupts From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson Selftest drivers that recover from a fault by issuing VFIO_DEVICE_RESET need to re-arm device interrupts afterwards. VFIO_DEVICE_RESET tears down the kernel-side IRQ trigger so a subsequent VFIO_DEVICE_SET_IRQS is required, but the user-side eventfds (and any fd cached in a test fixture) are still valid and must be preserved. vfio_pci_irq_enable() refuses to be called for vectors that already have an eventfd (VFIO_ASSERT_LT), and vfio_pci_irq_disable() closes all eventfds before resetting the trigger, so neither is suitable. Add vfio_pci_irq_reenable(device, index, vector, count) which asserts that the requested range has existing eventfds and re-issues VFIO_DEVICE_SET_IRQS using them. Signature mirrors vfio_pci_irq_enable(). Add vfio_pci_msi{,x}_reenable() wrappers around vfio_pci_irq_reenable() for additional ease of use and readability. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- .../vfio/lib/include/libvfio/vfio_pci_device.h | 14 ++++++++++++++ tools/testing/selftests/vfio/lib/vfio_pci_device.c | 22 ++++++++++++++++++= ++++ 2 files changed, 36 insertions(+) diff --git a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_devi= ce.h b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h index 2858885a89bb..2e67afc0d580 100644 --- a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h +++ b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h @@ -68,6 +68,8 @@ void vfio_pci_config_access(struct vfio_pci_device *devic= e, bool write, void vfio_pci_irq_enable(struct vfio_pci_device *device, u32 index, u32 vector, int count); void vfio_pci_irq_disable(struct vfio_pci_device *device, u32 index); +void vfio_pci_irq_reenable(struct vfio_pci_device *device, u32 index, + u32 vector, int count); void vfio_pci_irq_trigger(struct vfio_pci_device *device, u32 index, u32 v= ector); =20 static inline void fcntl_set_nonblock(int fd) @@ -92,6 +94,12 @@ static inline void vfio_pci_msi_disable(struct vfio_pci_= device *device) vfio_pci_irq_disable(device, VFIO_PCI_MSI_IRQ_INDEX); } =20 +static inline void vfio_pci_msi_reenable(struct vfio_pci_device *device, + u32 vector, int count) +{ + vfio_pci_irq_reenable(device, VFIO_PCI_MSI_IRQ_INDEX, vector, count); +} + static inline void vfio_pci_msix_enable(struct vfio_pci_device *device, u32 vector, int count) { @@ -103,6 +111,12 @@ static inline void vfio_pci_msix_disable(struct vfio_p= ci_device *device) vfio_pci_irq_disable(device, VFIO_PCI_MSIX_IRQ_INDEX); } =20 +static inline void vfio_pci_msix_reenable(struct vfio_pci_device *device, + u32 vector, int count) +{ + vfio_pci_irq_reenable(device, VFIO_PCI_MSIX_IRQ_INDEX, vector, count); +} + static inline int __to_iova(struct vfio_pci_device *device, void *vaddr, i= ova_t *iova) { return __iommu_hva2iova(device->iommu, vaddr, iova); diff --git a/tools/testing/selftests/vfio/lib/vfio_pci_device.c b/tools/tes= ting/selftests/vfio/lib/vfio_pci_device.c index fc75e04ef010..7b8394d0ac50 100644 --- a/tools/testing/selftests/vfio/lib/vfio_pci_device.c +++ b/tools/testing/selftests/vfio/lib/vfio_pci_device.c @@ -106,6 +106,28 @@ void vfio_pci_irq_disable(struct vfio_pci_device *devi= ce, u32 index) vfio_pci_irq_set(device, index, 0, 0, NULL); } =20 +/* + * Re-issue VFIO_DEVICE_SET_IRQS for an already-enabled vector range using + * the existing eventfds. Intended for drivers that need to re-arm device + * interrupts after a VFIO_DEVICE_RESET, which tears down the kernel-side + * IRQ trigger but leaves user-side eventfds intact. Recreating the + * eventfds would invalidate any test-fixture cache of the fd, so this + * helper deliberately preserves them. + */ +void vfio_pci_irq_reenable(struct vfio_pci_device *device, u32 index, + u32 vector, int count) +{ + int i; + + check_supported_irq_index(index); + + for (i =3D vector; i < vector + count; i++) + VFIO_ASSERT_GE(device->msi_eventfds[i], 0, + "vector %d eventfd not allocated\n", i); + + vfio_pci_irq_set(device, index, vector, count, device->msi_eventfds + vec= tor); +} + static void vfio_pci_irq_get(struct vfio_pci_device *device, u32 index, struct vfio_irq_info *irq_info) { --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A436D312837 for ; Wed, 22 Jul 2026 00:15:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679354; cv=none; b=s4Fq9V3EoOtUhj+gxWfDH3TneI05i4iUCMvgg4ItD3cTExAevbVfHkWhaxTvgfug64Izd4WeAj7lBebPFXFdZtZUVmR2n7BbKIgAwLLjc1y8pG2p1T55blvzfNFpx6lMFQ2+jR/j1UaO9Vzl4xT1HokhPa0YDMt8bND3/1ayUz0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679354; c=relaxed/simple; bh=+u92+edGzJUogIivEKHvjJM5J721kuk7QGsl9GsyiVA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=fWvbNejil3iuRj7DZd1B9nfw5BjZ+TM/LBgcaR5OiBtLoRzKL4l9tGkeIm5r7+bZJAjGlp2tGIBzp7bndkLZ90WCuzxEamXO1y6wlzKY9+qstgGMhRHw1pybRYDPSXaDC2DpBJg6ZIh6hl/kykFH+jAh6BCZMuhuSTl+P0Z2aMk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Hb+af4dw; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Hb+af4dw" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2cc6dd43737so224664595ad.2 for ; Tue, 21 Jul 2026 17:15:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679346; x=1785284146; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=RPmtr/TNEphr6iJROeMyjAAqcZVr5NPU0l3kEo6G8jw=; b=Hb+af4dwxJ2MONl287btlm2st8bDJahde+XGVcdNKEcPwq/Agm8Jw/B5huCKMO0E08 ioCwwEnNjk4Fa0F70TiGHbuHKfpVV1nRrzWYU1uOW+sr4BgGxJQkZVuX65q6UwqqJ37t HVsMaQY4rWRzvDvRY2Sos+8fnFvYYSmsl6b1bb4cN0H389Tisw4WUrw9WeGruNCM2/xM Sm+kMyDTLZ/+fe+0jwsQiH6/xBpjUrRC/DdPbcHCOrBOe9BA8/VbritSkM+lAotwMT/G Fuc84+97ZiNmA7boq0XalNv/WeJpiVYZWDqWZTwMDmZW71r3JzA4OJtksEcdOgXL935V zJDw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679346; x=1785284146; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=RPmtr/TNEphr6iJROeMyjAAqcZVr5NPU0l3kEo6G8jw=; b=qHy5Q1+PskMGjfGzEyiPE6sQG6hSp4MR6bFTkbPLtJHZRGYxkocEdQik9lcANb2XoY bYAha7Y0dlP3sNqfzxAanEZDf4wPiTsvRxEGBxoroJsGsMxHTkBAbsfHtfEZ8HWMgje7 hCHV5wgOgdFDTS0R34Y1TsP+qi//+IIpXmXjdeAu3PSKb4ncUxSS+J5MmebTAkcvyWeY CXZWwY3KQRE7DTVe/KLna2CE38Hwg30Jt4dd9TJNRF6+Z54aX9FxwTpM0ZV8kbP6qSqH CHvpsl9P8XfV7YsB6jQ/tpX4gQIXltcAlTC/YF/tTvgX42DkAc/XUfHm4mWnO6TizwpT E1/Q== X-Forwarded-Encrypted: i=1; AHgh+RpFgNav3MjNilMu31ohUbgGBgNFRhkoEuEVCDKtXoo1FNnArbk80FAMoZeLGIlz17RcammlxrOnTRRy1Mw=@vger.kernel.org X-Gm-Message-State: AOJu0YzR12M0KhyrIGOvewxpyTiA857+pPHkqnJEb/ahnhc9M59OcVpf yhaRR3WM2mmZZc6IuoW4NZB9iv2MvHKnJs/Xqo+73ZmLb6j49XkFuzB4iDuh4hO1xXLu5ttjPGq U563XbtUM X-Received: from pgcp18.prod.google.com ([2002:a63:7412:0:b0:c9a:4ffc:3456]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:a103:b0:3c0:9c1a:8941 with SMTP id adf61e73a8af0-3c3ada59cd5mr21741694637.73.1784679345554; Tue, 21 Jul 2026 17:15:45 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:39 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=4049; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=3HaO+LlIH5EQwYBLepHXt1276V0e3otp2TZx5sNJQxA=; b=4EgBrBNIcrg+NHGZ5nkD3MnsDofc8PMsdapr/842JIP2bDaGYrH15xDznsSmgyiLoO8X7z9El aD1bo0i0JvQDXiyvjPK8mzZfohuPAiyLS8MVb3xgODFbOwmrpYZPAP1 X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-4-0be1d29918d4@google.com> Subject: [PATCH v6 4/6] vfio: selftests: igb: Factor hardware programming into igb_hw_init() From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson Split the device register programming out of igb_init() into a new igb_hw_init() helper so that the same sequence can be re-run after a VFIO_DEVICE_RESET to restore the registers that CTRL.RST clears. No functional change for the initial path. igb_init() now performs the one-shot setup: region size assertion, BAR mapping, CTRL.RST + IMC mask-all to put the device into a known state, and vfio_pci_msix_enable() to set up the kernel-side IRQ trigger. igb_hw_init() does the rest: ring pointer setup and IOVA calc, CTRL_EXT, PCI bus master, GCR, PHY loopback, descriptor rings, RCTL, TCTL, GPIE/EIAC/EIAM/EIMS/IVAR, and driver-state initialization. vfio_pci_msix_enable() moves from after RCTL/TCTL to before all device-side programming. Its only side effects are the VFIO kernel IRQ trigger setup and the PCI MSI-X capability bits in config space; neither has any ordering dependency on the 82576 device register writes performed in igb_hw_init(). Performing it once in igb_init() keeps igb_hw_init() reusable from the reset recovery path (which uses vfio_pci_irq_reenable() to re-arm the existing trigger). Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 41 ++++++++++++++++--= ---- 1 file changed, 31 insertions(+), 10 deletions(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 2b444be4bdfe..8eaf120330dc 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -202,7 +202,13 @@ static void igb_reset(struct igb *igb) igb_write32(igb, E1000_IMC, 0xFFFFFFFF); } =20 -static void igb_init(struct vfio_pci_device *device) +/* + * Program the device into a usable state. Split out of igb_init() so it + * can be reused after a device reset to re-program the registers that + * CTRL.RST clears. Expects bar0 to be mapped and MSI-X already enabled + * via VFIO. + */ +static void igb_hw_init(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); u64 iova_tx, iova_rx; @@ -210,15 +216,10 @@ static void igb_init(struct vfio_pci_device *device) u16 cmd_reg; int retries; =20 - VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); - - /* Set up rings and calculate IOVAs */ - igb->bar0 =3D device->bars[0].vaddr; - iova_tx =3D to_iova(device, igb->tx_ring); iova_rx =3D to_iova(device, igb->rx_ring); =20 - igb_reset(igb); + =20 /* Signal that the driver is loaded */ ctrl =3D igb_read32(igb, E1000_CTRL_EXT); @@ -312,9 +313,6 @@ static void igb_init(struct vfio_pci_device *device) igb_write32(igb, E1000_RCTL, rctl); igb_write32(igb, E1000_TCTL, E1000_TCTL_EN); =20 - /* Enable MSI-X with 1 vector for the test */ - vfio_pci_msix_enable(device, MSIX_VECTOR, 1); - /* * Program MSI-X interrupt routing per 82576 datasheet: * @@ -354,6 +352,29 @@ static void igb_init(struct vfio_pci_device *device) device->driver.msi =3D MSIX_VECTOR; } =20 +static void igb_init(struct vfio_pci_device *device) +{ + struct igb *igb =3D to_igb_state(device); + + VFIO_ASSERT_GE(device->driver.region.size, sizeof(struct igb)); + + igb->bar0 =3D device->bars[0].vaddr; + + igb_reset(igb); + + /* + * Enable MSI-X via VFIO before device-side register programming. + * vfio_pci_msix_enable() only touches the VFIO IRQ machinery and the + * PCI MSI-X capability via config space; it has no ordering + * dependency on the device-side writes performed by igb_hw_init(). + * Placing it here keeps igb_hw_init() reusable from the reset + * recovery path (which calls vfio_pci_irq_reenable() instead). + */ + vfio_pci_msix_enable(device, MSIX_VECTOR, 1); + + igb_hw_init(device); +} + static void igb_remove(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3D63F31960A for ; Wed, 22 Jul 2026 00:15:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679352; cv=none; b=UVNDfs6zdKgDy7gHpt8T4U6/PH6SLAAJX9tH5eXDyR5qE3B5KLRt7UZjNeZbvYrpvW1z39t7mOmX5hV7xrGX5PFepHXumCpdnbZjLdjAJ+kY49rZHR2tLDXtSQTGLLkdxX42LuR/H2FmQ8cYwGbzGgvbyx/vY1DbwnTBB5vJXow= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679352; c=relaxed/simple; bh=lfTItr5Vg+lZepk1YTtimNxtQxX2xxsuhXX8/E8FyOA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=kR0I248jS4MezXQnLwEkN2aOS9XCHeZwEEew20GnIuHEI0MXo+RXwHWyvPLmClTZ9GFwpHJbfPgpATK6eEvrIqrnrG06s8+NKgpSS1bp7m55QcE/6pwsdl8W0+kVUN37lcpiBtz+B5kDj5hMTNXkorUmhRaD8/ZPkkewira5aQU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=hmBqsgrU; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="hmBqsgrU" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-c860544c077so19990006a12.3 for ; Tue, 21 Jul 2026 17:15:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679347; x=1785284147; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=eZ+JYUl+bJF8E+CsVPDDHP5xTQJTCrW/WMk4UkvfNvM=; b=hmBqsgrUjnFUFOOyBjzxhfhOcRgrE6pvFKFhgW8D+E9YTV53QBstvn4DbBFWvU3vu3 ItNu32xBTfX3eD1qShImmIfeeZ25yI5eZYW9Yyqu5jnolrNKUKRBbcXJ0CHRp6oQvBUM l+u8WYf9GFHaCCZ1rWrpzrDKaSLCkyYwly1XebBvFserPsUFCBlMb8MYYMNdTUFH0y6s nIXS6ma1hDMsmxP24B4g64KhQMzk4BEAYzOhipqYE6xArr4tYuoU1fI42lQ6uap3Y7lF n2FIthZjsLqfcCvIfMwDaJQgCy3s6iShBth83WwJBeWsFpiUdyM4vHyz7cHtypMwBFrq eGIA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679347; x=1785284147; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=eZ+JYUl+bJF8E+CsVPDDHP5xTQJTCrW/WMk4UkvfNvM=; b=QJfhc0RBfaMyWt8zAvYsJYsM1A2TIIycZ68cbaquPphCK+AMXm/+A0UlZ4VR/Np0Tb 4tkg083gQC6OxokRuWTvzUN71MWKpdMbBp8FysY4d6jSxiy4kpzEuDYAURC34Bmjgn8m evf/F7KefXZi0zca2OpVRJmS2fm6ifqJcoybeOvhV9fWwK7KclsB7ZENNWfeEX4xTHk6 BZ+B2UmhiOlYwopO3qIrxidahEAZs5ozu858E6GJelzzxZxlBkYjcmOjHwtbgW1czs3P BCr+h/vikKRjwPZluWes6IH0WY39SE2Dx2dZZGquRrsOAD7xk3gBm3sL+6S+cgzCgw8u 1SHg== X-Forwarded-Encrypted: i=1; AHgh+RrYybihhGEOGYTqgUAd2Ilk/LIBLBZmdg/Wgeih4MYUkijqXvmraK2gUmN2zQPWLUomCqhH8EvKPAoyRe8=@vger.kernel.org X-Gm-Message-State: AOJu0Yz56R3Yjw8tCgpHdK2RbYnNzHXWvmd5iYICJyJkwFwB1cESFttY E2VQLOzh/CG35Bv71Nc+GcQ15NGww3J6aRN+dLjdIoG9ezGBy+xKkPSIXetUhNjNf4sgcRBes51 nMg8vOk8D X-Received: from pgww13-n1.prod.google.com ([2002:a05:6a02:2c8d:10b0:c8a:c79:47e7]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:c70b:b0:3c3:7427:5ed8 with SMTP id adf61e73a8af0-3c3ad5fc1cdmr23350917637.8.1784679346553; Tue, 21 Jul 2026 17:15:46 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:40 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=2385; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=lfTItr5Vg+lZepk1YTtimNxtQxX2xxsuhXX8/E8FyOA=; b=FwaSkyLKFRNCDrxuhKdnvvYY8lP/k/mmM0g9bHg2OIqLQqw3ScD4PobkBcMqhKdOSAoROja30 F/ZYfWzQN6UA1YEKPiTvjn4KNqM8wBlII40LcWweQbDp5vIgwzQYMmB X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-5-0be1d29918d4@google.com> Subject: [PATCH v6 5/6] vfio: selftests: Retry on EAGAIN during device reset From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Add retry logic to vfio_pci_device_reset() to handle the case where PCI resets fail due to lock contention, in which case pci_try_reset_function() returns -EAGAIN. Suggested-by: Sashiko Suggested-by: David Matlack Signed-off-by: Josh Hilke Reviewed-by: David Matlack --- .../vfio/lib/include/libvfio/vfio_pci_device.h | 1 + tools/testing/selftests/vfio/lib/vfio_pci_device.c | 17 +++++++++++++= +++- 2 files changed, 17 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_devi= ce.h b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h index 2e67afc0d580..27bdf561925f 100644 --- a/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h +++ b/tools/testing/selftests/vfio/lib/include/libvfio/vfio_pci_device.h @@ -41,6 +41,7 @@ struct vfio_pci_device { struct vfio_pci_device *vfio_pci_device_init(const char *bdf, struct iommu= *iommu); void vfio_pci_device_cleanup(struct vfio_pci_device *device); =20 +int __vfio_pci_device_reset(struct vfio_pci_device *device); void vfio_pci_device_reset(struct vfio_pci_device *device); =20 void vfio_pci_config_access(struct vfio_pci_device *device, bool write, diff --git a/tools/testing/selftests/vfio/lib/vfio_pci_device.c b/tools/tes= ting/selftests/vfio/lib/vfio_pci_device.c index 7b8394d0ac50..1b29cef96b04 100644 --- a/tools/testing/selftests/vfio/lib/vfio_pci_device.c +++ b/tools/testing/selftests/vfio/lib/vfio_pci_device.c @@ -1,5 +1,6 @@ // SPDX-License-Identifier: GPL-2.0-only #include +#include #include #include #include @@ -221,9 +222,23 @@ void vfio_pci_config_access(struct vfio_pci_device *de= vice, bool write, write ? "write to" : "read from", config); } =20 +int __vfio_pci_device_reset(struct vfio_pci_device *device) +{ + if (ioctl(device->fd, VFIO_DEVICE_RESET, NULL)) + return -errno; + + return 0; +} + void vfio_pci_device_reset(struct vfio_pci_device *device) { - ioctl_assert(device->fd, VFIO_DEVICE_RESET, NULL); + int r; + + do { + r =3D __vfio_pci_device_reset(device); + } while (r =3D=3D -EAGAIN); + + VFIO_ASSERT_EQ(r, 0, "ioctl(device->fd, VFIO_DEVICE_RESET) failed\n"); } =20 static unsigned int vfio_pci_get_group_from_dev(const char *bdf) --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 00:16:11 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 84C7E310784 for ; Wed, 22 Jul 2026 00:15:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679353; cv=none; b=F3xzFfjNRdAdSVXs/daobm7T4xBENVvpeOoOG00jwR690tCWCKz8ncxcElkXS4YQVnmUSMu4ymHcKONeI2qD9J5HIO/KpDJbP5dzLxDhJ1y+oZw8x8YfQ3xXEhtTrZRcwXnzZt2z31g1rEnMsg5lu7DCupjQnrMLyWgEXwzFJpE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784679353; c=relaxed/simple; bh=CyqVWFJKhdjZZeW6oCD79hj9ow+t23S9RMWtKThbXt4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=hHK87oWpfzABWizKYSmSCDpxK51VFbLOseVPuLbLjjy1yZwQ8Q7eS/48KAT9jzaCjDORNbCQTpmF+/+Ty8jVPwD/SLdFBmKf2hWSotrecFPlp0rhUu/KSYzpLjkhwLUrKvJO2bVcMerXopMnwoKgzSeKAVxwSLBzTDTPl/OrHjY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KNZqhKa3; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jrhilke.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KNZqhKa3" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-ca8aee88725so17934882a12.3 for ; Tue, 21 Jul 2026 17:15:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784679348; x=1785284148; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=OcSVBvE6dCwYoJRa7BsGgfHeCjoSQwGXcD7G6E4aFuA=; b=KNZqhKa3LKcxsZRzKBlboFenzwzkaP/i+p+KY0O0rKfTWMeJsIWPP0ECYk/om74HEl ZVFFquBowQ8x6Uv56IgVdD1ptQmgGclYg0rOhFXJPGR/j5te8LVWsTwfqcCguMetI7sQ fhCX63oXeUpxLyoGbb+7X0TuQ0wViWaNRmryo+lXYLRCqRdqd4y2jXbxJnuD9Wa1hTeP qp77b2HuNvxnMJIAscwTOwPzvy61aBdcxHgB9LBnC1x1mCfAI1b/QhM+JhxgwXAwT7Pu 4Ua9DLphZGgw+EjPf3GftZrzTQRZ/zoX4qic5pAKO/E0WOObXSbYvbe+PcP1xKxTQYM7 XWxw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784679348; x=1785284148; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OcSVBvE6dCwYoJRa7BsGgfHeCjoSQwGXcD7G6E4aFuA=; b=P1WXGMWtJhz7ccW6ps8WHD9jpPBz97k3jr0wYpKAQOBiNuiS7KTnJ4pG/yOSJseR7t yGHAQQOg+7pbqRChyBNO8Sm99hYurm7OVnoFAURCvJBQuyW9yG8B5qynIhYarOxJLEGG vCb5RW42QvN2TKh7tE6SPr6ceJATWTz0czvpBDALqj+N5h1NxIzGBa4l/f9FNkPIuG72 77nFkaxmHOkvMN8SADhyll2S5eQKTU0yrB3ANMcrh/c0DJTAHUO9s4NwYieysDIl41SE G3bFz2CHtZJH3LJ3ZrxjbHRyioYjbVOArGnrE26NBghb0LjWEhg9I3PfvsWdhd+LRZLE abDg== X-Forwarded-Encrypted: i=1; AHgh+Roz2lHbqnl4kLc+4L/IIAWGhJg9vd8DpLq3on2ay8NxAei6P4F9GlTOIaAxfSECpGmyQxtnjyXqAFKSPD8=@vger.kernel.org X-Gm-Message-State: AOJu0YzofGwFG9uYjpRuftW3tJa1j+u+OgpY9kydRGFfmuMwNEfVCSrK lPZDxOzaN1QswltYGhmcoi0Em8XCvasZ2FxXC5DIE6yw0qpD3+Bxgjr0sIBIP/LVTJ4TkNuuw5k fiEpEXOmZ X-Received: from pgcv13.prod.google.com ([2002:a05:6a02:530d:b0:c89:7f9:d93a]) (user=jrhilke job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:9397:b0:3c3:83e8:c211 with SMTP id adf61e73a8af0-3c3ada1fe89mr22261006637.63.1784679347472; Tue, 21 Jul 2026 17:15:47 -0700 (PDT) Date: Wed, 22 Jul 2026 00:15:41 +0000 In-Reply-To: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260722-igb_v3_b4-v6-0-0be1d29918d4@google.com> X-Developer-Key: i=jrhilke@google.com; a=ed25519; pk=cK+Urrh214DoCHywJiTwhMP+zcs1ofpGrewToO/i/3A= X-Developer-Signature: v=1; a=ed25519-sha256; t=1784679337; l=4509; i=jrhilke@google.com; s=20260710; h=from:subject:message-id; bh=VGmU399xNoo4pH5/NvUhjWUHFyMHFl2ICfhg1+qYqo8=; b=OAQZ/jJfOS302Re6gqYzhquV0gOAZ495QGOKzVfxJFtjVKDAOXcYRRwC3335kbzrNkPCd9Bp2 aS+xlYxYnBWAvEy5fPypR6USp6uGOreGqxUDLOUbBLmCD2yVY0Sp9iD X-Mailer: b4 0.14.3 Message-ID: <20260722-igb_v3_b4-v6-6-0be1d29918d4@google.com> Subject: [PATCH v6 6/6] vfio: selftests: igb: Recover after DMA-read faults From: Josh Hilke To: David Matlack , Alex Williamson Cc: Shuah Khan , linux-kernel@vger.kernel.org, kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, Vipin Sharma , Josh Hilke , Alex Williamson Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Alex Williamson The mix_and_match test intentionally submits a TX descriptor with an unmapped source IOVA so that the DMA read fails. On real 82576 hardware the resulting fault leaves the descriptor engine unable to service subsequent valid descriptors, so the next memcpy in the same test iteration times out. The 82576 datasheet (section 4.2.1.6.1) describes CTRL.RST as the software mechanism to recover from a hung device. Empirically CTRL.RST alone is not sufficient in this state: the visible queue registers are reinitialized, but the next valid memcpy still posts descriptors without any TDH/TDT progress in the same process. A fresh device open after the failure works, which points to a reset scope broader than CTRL.RST being required. The 82576 advertises PCIe FLR; VFIO_DEVICE_RESET drives FLR and supplies that scope while preserving the selftest process and its DMA mappings. Add igb_error_reset_and_reinit() implementing the recovery sequence: issue VFIO_DEVICE_RESET, re-arm the kernel-side MSI-X trigger against the still-valid eventfd via vfio_pci_irq_reenable() (this does not touch the eventfd, which test fixtures may have cached), and re-program the device via igb_hw_init(). FLR clears EICR and leaves EIMS=3D0, so no explicit interrupt mask or cause writes are needed. igb_hw_init() resets tx_tail/rx_tail to 0 and igb_memcpy_start() zeros each descriptor before submission, so no ring memset is needed either. Call this from igb_memcpy_wait() on completion timeout, preceded by a 10 ms delay so that PCIe/IOMMU/AER error handling triggered by the just-observed DMA fault can release the device lock VFIO_DEVICE_RESET contends for. The delay is heuristic and tied to the fault path, so it lives at the call site rather than inside the reset helper. The failed memcpy still returns -ETIMEDOUT; reset recovery only ensures the next operation starts from a usable device state. Assisted-by: Claude:claude-opus-4-7 Signed-off-by: Alex Williamson Reviewed-by: David Matlack --- tools/testing/selftests/vfio/lib/drivers/igb/igb.c | 38 ++++++++++++++++++= +++- 1 file changed, 37 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c b/tools/tes= ting/selftests/vfio/lib/drivers/igb/igb.c index 8eaf120330dc..25cac663de1d 100644 --- a/tools/testing/selftests/vfio/lib/drivers/igb/igb.c +++ b/tools/testing/selftests/vfio/lib/drivers/igb/igb.c @@ -458,6 +458,28 @@ static void igb_memcpy_start(struct vfio_pci_device *d= evice, iova_t src, igb_write32(igb, E1000_TDT(0), igb->tx_tail); } =20 +/* + * Reset the device via VFIO_DEVICE_RESET (PCIe FLR on the 82576) and + * re-program it. VFIO_DEVICE_RESET tears down the kernel-side MSI-X + * trigger but leaves user-side eventfds intact, so re-arm the trigger + * via vfio_pci_irq_reenable() before reprogramming so any caller-cached + * eventfd remains valid. + * + * FLR clears device-side state to power-on reset values (datasheet + * 4.2.1.5.1: a PF FLR is "equivalent to a D0->D3->D0 transition"), so + * EIMS and EICR come back as 0 from their register-defined initial + * values, and igb_hw_init() resets tx_tail/rx_tail to 0. The next + * igb_memcpy_start() will memset each descriptor it touches before + * submission, so no explicit IMC/EICR writes or ring memsets are + * needed here. + */ +static void igb_error_reset_and_reinit(struct vfio_pci_device *device) +{ + vfio_pci_device_reset(device); + vfio_pci_msix_reenable(device, MSIX_VECTOR, 1); + igb_hw_init(device); +} + static int igb_memcpy_wait(struct vfio_pci_device *device) { struct igb *igb =3D to_igb_state(device); @@ -500,7 +522,21 @@ static int igb_memcpy_wait(struct vfio_pci_device *dev= ice) =20 igb_irq_enable(igb); =20 - return (status & 1) ? 0 : -ETIMEDOUT; + if (status & 1) + return 0; + + /* + * The descriptor never completed. On real 82576 hardware this + * typically follows a DMA-read fault from one of the intentional + * unmapped-IOVA tests; the fault leaves the descriptor engine + * unable to service subsequent valid descriptors. CTRL.RST alone + * reinitializes the queue registers but leaves the engine wedged + * for the current process, so a broader VFIO_DEVICE_RESET (FLR) + * is required. + */ + igb_error_reset_and_reinit(device); + + return -ETIMEDOUT; } =20 static void igb_send_msi(struct vfio_pci_device *device) --=20 2.55.0.229.g6434b31f56-goog