From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130811; cv=none; d=zohomail.com; s=zohoarc; b=k7GiRq8XghwSZUZo+QS3gkzD482A0NxyUKw2372JZBXna19ywcrNQrvjPFV8IDt/oAiFav8Ox2JEyGzd3BTlkpLuEHWurarQsfkH5K3CsbJaxT/0oNr1nCmkHx/Sqks00e9svbv2wmCiDl5/mHe2ZnDSAB/5PSBy4FMww74bTtU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130811; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=NlvkrgsEA9eAyK8sO+gEUPMsIddzXNN5H6LFRBn5jHc=; b=HwlU0uKPMP1cekj8a4ys9PqgoUR1coIYwT9bw5NlVygyTWvHvBpdC/6O4Kp+BXzlyeyLjAesGQVASGksnhyvb1wDfgC8TE99TCYicGkcJ6PIAzbcsRsYcogjJbIx2eFv+Vd14Uv8xIruXH9VAPTOg1BhXYaJNCPuVSrtQn00hLE= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130811828387.2557828896739; Sun, 26 Jul 2026 22:40:11 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4I-00037p-1T; Mon, 27 Jul 2026 01:40:02 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4A-00035y-5g for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:39:54 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE48-0003St-H7 for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:39:53 -0400 Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-284-xZEqcayeM9yEDjlCmbhlxQ-1; Mon, 27 Jul 2026 01:39:47 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 17BB11955F26; Mon, 27 Jul 2026 05:39:46 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 452141956088; Mon, 27 Jul 2026 05:39:43 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130791; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=NlvkrgsEA9eAyK8sO+gEUPMsIddzXNN5H6LFRBn5jHc=; b=c0H009XRraTENkh7c96uUcsdrctjykLs3HxGiQLXTgdDUbFhs5e5KZyYBOX0KB9M/pPIFf TPh5x6kMpRytb/QBy8OakJYKKUK3QO/Ciw7LAgK6Im5nJoZPPk23d6etYMlBekifoXqxjC VpauE1i5NqGupcuJNyaxwdHcNAD/i90= X-MC-Unique: xZEqcayeM9yEDjlCmbhlxQ-1 X-Mimecast-MFC-AGG-ID: xZEqcayeM9yEDjlCmbhlxQ_1785130786 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 01/11] pci: Add PCI_BASE_ADDRESS_MEM_ALWAYS_ON BAR flag Date: Mon, 27 Jul 2026 07:39:25 +0200 Message-ID: <20260727053935.1392269-2-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130812986158500 When VFIO opens a VF, it issues a Function Level Reset which clears PCI_COMMAND_MEMORY. This unmaps all BARs, including the migration BAR. Add a QEMU-internal BAR type flag that keeps a memory BAR mapped regardless of PCI_COMMAND_MEMORY. This is needed for host-only control regions (e.g. migration BARs) that are accessed by a VFIO variant driver, not by the guest. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- include/hw/pci/pci.h | 6 ++++++ hw/pci/pci.c | 6 ++++-- 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/include/hw/pci/pci.h b/include/hw/pci/pci.h index f2448e941a0b..923c8f0e15bc 100644 --- a/include/hw/pci/pci.h +++ b/include/hw/pci/pci.h @@ -164,6 +164,12 @@ typedef struct PCIIORegion { MemoryRegion *address_space; } PCIIORegion; =20 +/* + * QEMU-internal: keep this BAR mapped regardless of PCI_COMMAND_MEMORY. + * Stripped before writing to config space. + */ +#define PCI_BASE_ADDRESS_MEM_ALWAYS_ON 0x10 + #define PCI_ROM_SLOT 6 #define PCI_NUM_REGIONS 7 =20 diff --git a/hw/pci/pci.c b/hw/pci/pci.c index d3191609e283..1394132c1472 100644 --- a/hw/pci/pci.c +++ b/hw/pci/pci.c @@ -1544,7 +1544,8 @@ void pci_register_bar(PCIDevice *pci_dev, int region_= num, } =20 addr =3D pci_bar(pci_dev, region_num); - pci_set_long(pci_dev->config + addr, type); + pci_set_long(pci_dev->config + addr, + type & ~PCI_BASE_ADDRESS_MEM_ALWAYS_ON); =20 if (!(r->type & PCI_BASE_ADDRESS_SPACE_IO) && r->type & PCI_BASE_ADDRESS_MEM_TYPE_64) { @@ -1686,7 +1687,8 @@ pcibus_t pci_bar_address(PCIDevice *d, return new_addr; } =20 - if (!(cmd & PCI_COMMAND_MEMORY)) { + if (!(cmd & PCI_COMMAND_MEMORY) && + !(type & PCI_BASE_ADDRESS_MEM_ALWAYS_ON)) { return PCI_BAR_UNMAPPED; } new_addr =3D pci_config_get_bar_addr(d, reg, type, size); --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130855; cv=none; d=zohomail.com; s=zohoarc; b=NgSyvYBjW3BuLDX3cFNYb/Uhma0mcpv1PeBCTVdpILB8akwGB2yOoC1thZCOc3mimLeBnFZOGcSlYAO0TOXlS0L551D+qDShCPXcrE0R+XsfnD6WqvsnTdS32x/Z+nV2HViTwKcbT7EBis8HoOgOzyjiuh3Ok5+lWbW+cHwb20g= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130855; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=QxrGdQ0nIXwDaXRl/lGKYpcwFMhthuZPnFSAeFGQRn4=; b=LYjYCR56S3ahLGG1pysu8ANqs6pAGi6HXiob0q55JAMuK69Xvs4vvbaGWu7r49B5UvNd7byEryQei+gr03k7dNTFxShYc1Hure3cjSEHwXddNo3gel0Y4PWicPGSGMW+wCohzVjny9peqHfVFASuoRZ/Epsra4SBFEZvTGs8uy0= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130855452749.1176562293126; Sun, 26 Jul 2026 22:40:55 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4N-0003B3-1D; Mon, 27 Jul 2026 01:40:07 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4C-00037Q-QZ for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:01 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4A-0003Ur-PO for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:39:56 -0400 Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-474-_sSkTS18M1ucnfoaP7jurg-1; Mon, 27 Jul 2026 01:39:50 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 30F6C195605A; Mon, 27 Jul 2026 05:39:49 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 888231956088; Mon, 27 Jul 2026 05:39:46 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130794; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=QxrGdQ0nIXwDaXRl/lGKYpcwFMhthuZPnFSAeFGQRn4=; b=iOqZ7ywjCxQlTC+/hey2r7Ck8dJNXEStcZvnivGoEeeAk1FTnst4K2Z3gwNe3JfIydFO1L 91aRmtzZhhOMAc9OUgrJ8zKHM2LumMPsmh/Em2Q7yDtwfeKNAewYTrcJR1cfSvE3eJZOVM Q+yCjNdG0Ld4b+bY5JUm4OBuVJt1b0c= X-MC-Unique: _sSkTS18M1ucnfoaP7jurg-1 X-Mimecast-MFC-AGG-ID: _sSkTS18M1ucnfoaP7jurg_1785130789 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 02/11] igb: Add x-vf-migration property and vendor-specific capability for IGBVF Date: Mon, 27 Jul 2026 07:39:26 +0200 Message-ID: <20260727053935.1392269-3-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130856649158500 Add a PF-level "x-vf-migration" boolean property (default off) that, when enabled, causes each emulated VF to advertise a vendor-specific PCI capability in its config space. The variant driver igb-vfio-pci probes for this capability at bind time and uses its presence to enable migration support. The capability is a 16-byte vendor-specific capability (PCI_CAP_ID_VNDR) containing a magic value ("MIGB"), the BAR index where the migration region will be mapped, and feature flags indicating which migration features are supported. Assisted-by: Claude Suggested-by: Alex Williamson Signed-off-by: C=C3=A9dric Le Goater --- MAINTAINERS | 6 +++ docs/system/device-emulation.rst | 1 + docs/system/devices/igb-migration.rst | 20 ++++++++ docs/system/devices/igb.rst | 6 +++ hw/net/igb_migration.h | 43 ++++++++++++++++ hw/net/igb.c | 12 +++++ hw/net/igb_migration.c | 72 +++++++++++++++++++++++++++ hw/net/igbvf.c | 11 ++++ hw/net/meson.build | 2 +- hw/net/trace-events | 1 + 10 files changed, 173 insertions(+), 1 deletion(-) create mode 100644 docs/system/devices/igb-migration.rst create mode 100644 hw/net/igb_migration.h create mode 100644 hw/net/igb_migration.c diff --git a/MAINTAINERS b/MAINTAINERS index a28935c89866..6d29b905c47c 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -2765,6 +2765,12 @@ F: tests/functional/x86_64/test_netdev_ethtool.py F: tests/qtest/igb-test.c F: tests/qtest/libqos/igb.c =20 +igb VF migration +M: C=C3=A9dric Le Goater +S: Maintained +F: hw/net/igb_migration.* +F: docs/system/devices/igb-migration.rst + eepro100 M: Stefan Weil S: Maintained diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulatio= n.rst index 40054bb7dfcc..75f423b795cd 100644 --- a/docs/system/device-emulation.rst +++ b/docs/system/device-emulation.rst @@ -90,6 +90,7 @@ Emulated Devices devices/cxl.rst devices/emmc.rst devices/igb.rst + devices/igb-migration.rst devices/ivshmem-flat.rst devices/ivshmem.rst devices/keyboard.rst diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/ig= b-migration.rst new file mode 100644 index 000000000000..017b47c4544e --- /dev/null +++ b/docs/system/devices/igb-migration.rst @@ -0,0 +1,20 @@ +.. SPDX-License-Identifier: GPL-2.0-or-later +.. _igb-migration: + +igb VF Migration +---------------- + +The igb device supports an experimental VF migration interface that allows +the ``igb-vfio-pci`` variant driver to migrate VF state during live +migration. This is enabled with the ``x-vf-migration`` property:: + + -device igb,x-vf-migration=3Don,... + +When enabled, each emulated VF advertises a vendor-specific PCI capability +(cap id 0x09) with a magic signature (``0x4D494742`` / "MIGB") that the +variant driver probes at bind time. The capability contains an interface +version number, the BAR index hosting the migration register region, and +feature flags indicating which migration features are supported. + +This feature is experimental and the ``x-`` prefix indicates the interface +may change. diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst index 50f625fd77e4..00271dbc92c3 100644 --- a/docs/system/devices/igb.rst +++ b/docs/system/devices/igb.rst @@ -64,6 +64,12 @@ command: =20 pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool =20 +VF Migration (experimental) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D + +See :ref:`igb-migration` for details on the experimental VF live migration +interface. + References =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =20 diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h new file mode 100644 index 000000000000..e79892436b21 --- /dev/null +++ b/hw/net/igb_migration.h @@ -0,0 +1,43 @@ +/* + * QEMU Intel 82576 SR/IOV VF Migration Support + * + * Copyright (c) 2026 Red Hat, Inc. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ + +#ifndef IGB_MIGRATION_H +#define IGB_MIGRATION_H + +#include "hw/pci/pci_device.h" + +/* Migration BAR definitions */ +#define IGB_MIG_BAR_IDX (2) +#define IGB_MIG_BAR_SIZE (64 * 1024) + +/* + * Vendor-specific PCI capability for migration discovery. + * + * The igb-vfio-pci variant driver probes for this at bind time. + * If present, the driver knows the emulated VF supports migration. + */ +#define IGB_MIG_CAP_MAGIC 0x4D494742 /* "MIGB" */ +#define IGB_MIG_CAP_VERSION 1 + +#define IGB_MIG_CAP_F_STATE (1u << 0) /* device state serialization */ +#define IGB_MIG_CAP_F_DIRTY (1u << 1) /* dirty page tracking */ + +#define IGB_MIG_CAP_SIZE 16 +#define IGB_MIG_CAP_OFF_MAGIC 4 /* offset within cap for magic field */ +#define IGB_MIG_CAP_OFF_BARID 8 /* offset within cap for BAR id */ +#define IGB_MIG_CAP_OFF_FLAGS 12 /* offset within cap for feature flags= */ + +typedef struct IgbVfMigState { + bool migration_cap; + MemoryRegion mig_bar; +} IgbVfMigState; + +void igb_pf_init_migration_bar(PCIDevice *dev); +bool igbvf_add_migration_cap(PCIDevice *dev, Error **errp); + +#endif diff --git a/hw/net/igb.c b/hw/net/igb.c index c076807e7110..b43235996db2 100644 --- a/hw/net/igb.c +++ b/hw/net/igb.c @@ -57,6 +57,7 @@ =20 #include "igb_common.h" #include "igb_core.h" +#include "igb_migration.h" =20 #include "trace.h" #include "qapi/error.h" @@ -69,6 +70,7 @@ struct IGBState { PCIDevice parent_obj; NICState *nic; NICConf conf; + bool vf_migration; =20 MemoryRegion mmio; MemoryRegion flash; @@ -459,6 +461,15 @@ static void igb_pci_realize(PCIDevice *pci_dev, Error = **errp) pcie_sriov_pf_init_vf_bar(pci_dev, IGBVF_MSIX_BAR_IDX, PCI_BASE_ADDRESS_MEM_TYPE_64 | PCI_BASE_ADDRESS_MEM_PREFETCH, IGBVF_MSIX_SIZE); + /* + * When VF migration support is enabled, register an additional VF + * BAR for the migration register region. The variant driver + * discovers this via a vendor-specific PCI capability that points + * to this BAR. + */ + if (s->vf_migration) { + igb_pf_init_migration_bar(pci_dev); + } =20 igb_init_net_peer(s, pci_dev, macaddr); =20 @@ -597,6 +608,7 @@ static const VMStateDescription igb_vmstate =3D { static const Property igb_properties[] =3D { DEFINE_NIC_PROPERTIES(IGBState, conf), DEFINE_PROP_BOOL("x-pcie-flr-init", IGBState, has_flr, true), + DEFINE_PROP_BOOL("x-vf-migration", IGBState, vf_migration, false), }; =20 static void igb_class_init(ObjectClass *class, const void *data) diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c new file mode 100644 index 000000000000..794f1217d3e5 --- /dev/null +++ b/hw/net/igb_migration.c @@ -0,0 +1,72 @@ +/* + * QEMU Intel 82576 SR/IOV VF Migration Support + * + * Copyright (c) 2026 Red Hat, Inc. + * + * SPDX-License-Identifier: GPL-2.0-or-later + */ + +#include "qemu/osdep.h" +#include "hw/pci/pci_device.h" +#include "hw/pci/pcie.h" +#include "igb_common.h" +#include "igb_migration.h" +#include "trace.h" + + + +/* + * 32-bit prefetchable BAR. A 64-bit BAR2 would consume BAR2+BAR3, but + * BAR3 is already used for MSI-X (IGBVF_MSIX_BAR_IDX =3D 3). + * + * BAR Index Type Size Purpose + * BAR0 0+1 64-bit prefetchable 16 KB MMIO registers + * BAR2 2 32-bit prefetchable 64 KB Migration + * BAR3 3+4 64-bit prefetchable 16 KB MSI-X table + PBA + * BAR5 5 - - Unused + */ +void igb_pf_init_migration_bar(PCIDevice *dev) +{ + pcie_sriov_pf_init_vf_bar(dev, IGB_MIG_BAR_IDX, + PCI_BASE_ADDRESS_MEM_PREFETCH, + IGB_MIG_BAR_SIZE); +} + +/* + * Add vendor-specific PCI capability that the variant driver probes for. + * + * Layout (16 bytes): + * [0] cap_id (PCI_CAP_ID_VNDR =3D 0x09) + * [1] next_cap + * [2] cap_len (16) + * [3] version (IGB_MIG_CAP_VERSION) + * [4-7] magic (IGB_MIG_CAP_MAGIC, little-endian) + * [8-11] bar_id (IGB_MIG_BAR_IDX, little-endian) + * [12-15] flags (feature flags, little-endian) + */ +bool igbvf_add_migration_cap(PCIDevice *dev, Error **errp) +{ + int offset; + + offset =3D pci_add_capability(dev, PCI_CAP_ID_VNDR, 0, + IGB_MIG_CAP_SIZE, errp); + if (offset < 0) { + return false; + } + + /* Length and version in the standard cap flags word */ + pci_set_byte(dev->config + offset + PCI_CAP_FLAGS, + IGB_MIG_CAP_SIZE); + pci_set_byte(dev->config + offset + PCI_CAP_FLAGS + 1, + IGB_MIG_CAP_VERSION); + + pci_set_long(dev->config + offset + IGB_MIG_CAP_OFF_MAGIC, + IGB_MIG_CAP_MAGIC); + pci_set_long(dev->config + offset + IGB_MIG_CAP_OFF_BARID, + IGB_MIG_BAR_IDX); + pci_set_long(dev->config + offset + IGB_MIG_CAP_OFF_FLAGS, + IGB_MIG_CAP_F_STATE); + + trace_igbvf_mig_cap_add(pcie_sriov_vf_number(dev), offset); + return true; +} diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c index 9a165c7063ee..94c9739cd58c 100644 --- a/hw/net/igbvf.c +++ b/hw/net/igbvf.c @@ -47,6 +47,7 @@ #include "net/net.h" #include "igb_common.h" #include "igb_core.h" +#include "igb_migration.h" #include "trace.h" #include "qapi/error.h" =20 @@ -57,6 +58,8 @@ struct IgbVfState { =20 MemoryRegion mmio; MemoryRegion msix; + + IgbVfMigState mig; }; =20 static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write) @@ -272,6 +275,14 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **= errp) hw_error("Failed to initialize PCIe capability"); } =20 + s->mig.migration_cap =3D object_property_get_bool(OBJECT(pcie_sriov_ge= t_pf(dev)), + "x-vf-migration", &error_a= bort); + if (s->mig.migration_cap) { + if (!igbvf_add_migration_cap(dev, errp)) { + return; + } + } + if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)), "x-pcie-flr-init", &error_abort)) { pcie_cap_flr_init(dev); diff --git a/hw/net/meson.build b/hw/net/meson.build index 84f142df222a..bb4b449b25ba 100644 --- a/hw/net/meson.build +++ b/hw/net/meson.build @@ -11,7 +11,7 @@ system_ss.add(when: 'CONFIG_E1000_PCI', if_true: files('e= 1000.c', 'e1000x_common system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('net_tx_pk= t.c', 'net_rx_pkt.c')) system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('e1000e.c'= , 'e1000e_core.c', 'e1000x_common.c')) system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('net_tx_pkt.c= ', 'net_rx_pkt.c')) -system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igb= vf.c', 'igb_core.c')) +system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igb= vf.c', 'igb_core.c', 'igb_migration.c')) system_ss.add(when: 'CONFIG_RTL8139_PCI', if_true: files('rtl8139.c')) system_ss.add(when: 'CONFIG_TULIP', if_true: files('tulip.c')) system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('net_tx_pkt.c', '= net_rx_pkt.c')) diff --git a/hw/net/trace-events b/hw/net/trace-events index 001a20b0e2ac..06d8848023e5 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -294,6 +294,7 @@ igb_wrn_rx_desc_modes_not_supp(int desc_type) "Not supp= orted descriptor type: %d =20 # igbvf.c igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64 +igbvf_mig_cap_add(uint16_t vfn, int offset) "VF%u: added migration vendor = cap at config offset 0x%x" =20 # spapr_llan.c spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_= bufs) "pool=3D%d count=3D%"PRId32" rxbufs=3D%"PRIu32 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130816; cv=none; d=zohomail.com; s=zohoarc; b=mB+/L+I111KHlxMEM0SjGd+6OtzTy7dSnC9hCB0VjAFpMpXpMPWv8dtoF/HbY93F5Z8yK9VkdnMnyq6egF2SUrYKPqRYxqsDwQDzpT2syTb/5FakquIYwwAFXeYhr29iT+YFJxJV7jreAa4jbQeZBGnW6eo6YYpbB4Gx++9BIYA= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130816; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=AIHHs32s1QVOrS+3x8DSNomksyIAyXwqLOWhx7sbx6k=; b=Hg9nH7ydu6xqjF6i9fHFECG/Og0ZWgreEydv0G5zdjBqDKmJ1UXTB4icLiImDWaaehQW4u/iMg0wzG+6vheQ6tBi3Db1nGTdK98Caq8fNqxMiJuLsfXA9fI/ZvFjDlMd/ENuOu6mgezsbL4ZlOp2RX/wMGMeGfO9b3EZXZOqswE= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130816075672.1589380280627; Sun, 26 Jul 2026 22:40:16 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4P-0003Js-LK; Mon, 27 Jul 2026 01:40:11 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4E-00037W-Vp for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:01 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4C-0003V4-Ea for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:39:58 -0400 Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-232-9TUn0N-rOTG84aagrt-OCA-1; Mon, 27 Jul 2026 01:39:54 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id CB81A1955F1E; Mon, 27 Jul 2026 05:39:52 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id A29CB1956088; Mon, 27 Jul 2026 05:39:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130795; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=AIHHs32s1QVOrS+3x8DSNomksyIAyXwqLOWhx7sbx6k=; b=JGIS9bTt7PB2BowbpUQVE6a+F4spyUnQ20t1EAQhC29UIobQnUOtmcEviyJ2I1UVgwHlhi UuTQ7BSCoNTaztxjy2Gbzk7h+AY47EtzG87Qg7a5NtGR8oi0NhsFTwKHX6+u6+xUsFaHdi taQRi5LQDJM6/B/jXJ4QDlZPfNzpl78= X-MC-Unique: 9TUn0N-rOTG84aagrt-OCA-1 X-Mimecast-MFC-AGG-ID: 9TUn0N-rOTG84aagrt-OCA_1785130793 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 03/11] igb: Add migration BAR with state machine Date: Mon, 27 Jul 2026 07:39:27 +0200 Message-ID: <20260727053935.1392269-4-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130818512158500 When x-vf-migration=3Don, register a 64KB migration BAR (BAR2) on each emulated VF. This BAR implements a VFIO-like migration state machine that the igb-vfio-pci variant driver uses to serialize/deserialize VF device state during live migration. The migration register region is laid out as: 0x000 DEVICE_STATE (RW) - migration state machine control 0x004 MIG_STATUS (RO) - status flags (DATA_AVAIL, ERROR, QUIESCED) + error code in [15:8] 0x008 MIG_CAPS (RO) - advertised capabilities 0x00C MIG_VERSION (RO) - interface version 0x010 DATA_SIZE (RW) - max state size at reset, actual after s= ave 0x014 DATA_XFER (WO) - trigger DMA save or DMA load 0x018 DATA_BUF_ADDR_LO (WO) - low 32 bits of state DMA buffer 0x01C DATA_BUF_ADDR_HI (WO) - high 32 bits of state DMA buffer State data is transferred via a driver-provided DMA buffer. The driver writes its PF DMA address to DATA_BUF_ADDR_LO/HI and triggers the transfer with DATA_XFER. The device DMA-writes the serialized state on save and DMA-reads it on restore. DMA is performed through the PF device (pcie_sriov_get_pf) because VFIO owns the VF's IOMMU domain. VF state serialization is added in the next patch. Assisted-by: Claude Suggested-by: Alex Williamson Signed-off-by: C=C3=A9dric Le Goater --- docs/system/devices/igb-migration.rst | 43 ++++ hw/net/igb_common.h | 11 + hw/net/igb_migration.h | 52 +++++ hw/net/igb_migration.c | 293 ++++++++++++++++++++++++++ hw/net/igbvf.c | 18 +- hw/net/trace-events | 8 + 6 files changed, 416 insertions(+), 9 deletions(-) diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/ig= b-migration.rst index 017b47c4544e..04c44ef072af 100644 --- a/docs/system/devices/igb-migration.rst +++ b/docs/system/devices/igb-migration.rst @@ -18,3 +18,46 @@ feature flags indicating which migration features are su= pported. =20 This feature is experimental and the ``x-`` prefix indicates the interface may change. + +Migration BAR layout +~~~~~~~~~~~~~~~~~~~~ + +The migration BAR (BAR2, 64 KB) implements a VFIO-like state machine with +the following register layout:: + + Offset Name Access Description + 0x000 DEVICE_STATE RW Migration state (RUNNING=3D2, STOP= =3D1, + STOP_COPY=3D3, RESUMING=3D4, PRE_COP= Y=3D5) + 0x004 STATUS RO Flags[2:0]: DATA_AVAIL, ERROR, QUIES= CED + Error code[15:8] (when ERROR is set) + 0x008 CAPS RO F_STATE, F_DIRTY, max_ranges[11:8], + pgsizes[31:12] + 0x00C VERSION RO Interface version (1) + 0x010 DATA_SIZE RW Max state size at reset, actual afte= r save + 0x014 DATA_XFER WO Trigger DMA save or DMA load + 0x018 DATA_BUF_ADDR_LO WO Low 32 bits of state DMA buffer addr= ess + 0x01C DATA_BUF_ADDR_HI WO High 32 bits of state DMA buffer add= ress + +State transitions follow the VFIO migration state machine: the driver +writes to ``DEVICE_STATE`` to move between states and reads ``STATUS`` +to check for completion. + +State data is transferred via a driver-provided DMA buffer. The driver +writes its PF DMA address to ``DATA_BUF_ADDR_LO/HI`` and triggers the +transfer with ``DATA_XFER``. The device DMA-writes the serialized state +on save and DMA-reads it on restore. DMA is performed through the PF +device because VFIO owns the VF's IOMMU domain. + +The state blob is a versioned sequence of register (offset, value) +pairs. + +When ``STATUS`` has the ``ERROR`` bit set, bits [15:8] contain an error +code identifying the failure:: + + 0 (none) No error + 1 BAD_MAGIC State blob magic mismatch + 2 BAD_VERSION State blob version mismatch + 3 BAD_SIZE State blob too large or empty + 4 BAD_VFN VF number mismatch (source !=3D destination) + 5 DMA_FAILED DMA transfer to/from state buffer failed + 6 NO_BUFFER DATA_XFER without buffer address set diff --git a/hw/net/igb_common.h b/hw/net/igb_common.h index b316a5bcfa5c..01816e002c23 100644 --- a/hw/net/igb_common.h +++ b/hw/net/igb_common.h @@ -27,6 +27,7 @@ #define HW_NET_IGB_COMMON_H =20 #include "igb_regs.h" +#include "igb_migration.h" =20 #define TYPE_IGBVF "igbvf" =20 @@ -36,6 +37,16 @@ #define IGBVF_MMIO_SIZE (16 * 1024) #define IGBVF_MSIX_SIZE (16 * 1024) =20 +struct IgbVfState { + PCIDevice parent_obj; + uint16_t vfn; + + MemoryRegion mmio; + MemoryRegion msix; + + IgbVfMigState mig; +}; + #define defreg(x) x =3D (E1000_##x >> 2) #define defreg_indexed(x, i) x##i =3D (E1000_##x(i) >> 2) #define defreg_indexeda(x, i) x##i##_A =3D (E1000_##x##_A(i) >> 2) diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h index e79892436b21..739a189810b0 100644 --- a/hw/net/igb_migration.h +++ b/hw/net/igb_migration.h @@ -32,12 +32,64 @@ #define IGB_MIG_CAP_OFF_BARID 8 /* offset within cap for BAR id */ #define IGB_MIG_CAP_OFF_FLAGS 12 /* offset within cap for feature flags= */ =20 +/* + * Maximum serialized VF state size, sized to hold all per-VF + * registers plus TX context descriptors with room to spare. + */ +#define IGB_VF_STATE_MAX_SIZE 4096 + +/* + * Migration BAR register offsets. + */ +#define IGB_MIG_DEVICE_STATE 0x000 +#define IGB_MIG_STATUS 0x004 +#define IGB_MIG_CAPS 0x008 +#define IGB_MIG_VERSION 0x00C +#define IGB_MIG_DATA_SIZE 0x010 +#define IGB_MIG_DATA_XFER 0x014 +#define IGB_MIG_DATA_BUF_ADDR_LO 0x018 +#define IGB_MIG_DATA_BUF_ADDR_HI 0x01C + +/* DEVICE_STATE values - mirrors VFIO migration states */ +#define IGB_MIG_STATE_ERROR 0 +#define IGB_MIG_STATE_STOP 1 +#define IGB_MIG_STATE_RUNNING 2 +#define IGB_MIG_STATE_STOP_COPY 3 +#define IGB_MIG_STATE_RESUMING 4 +#define IGB_MIG_STATE_PRE_COPY 5 + +/* MIG_STATUS bits */ +#define IGB_MIG_STATUS_DATA_AVAIL (1u << 0) +#define IGB_MIG_STATUS_ERROR (1u << 1) + +/* MIG_STATUS error codes in bits [15:8], valid when ERROR bit is set */ +#define IGB_MIG_STATUS_ERR_SHIFT 8 +#define IGB_MIG_STATUS_ERR_MASK (0xffu << IGB_MIG_STATUS_ERR_SHIFT) +#define IGB_MIG_STATUS_ERR(code) (IGB_MIG_STATUS_ERROR | \ + ((uint32_t)(code) << IGB_MIG_STATUS_E= RR_SHIFT)) + +#define IGB_MIG_ERR_BAD_MAGIC 1 +#define IGB_MIG_ERR_BAD_VERSION 2 +#define IGB_MIG_ERR_BAD_SIZE 3 +#define IGB_MIG_ERR_BAD_VFN 4 +#define IGB_MIG_ERR_DMA_FAILED 5 +#define IGB_MIG_ERR_NO_BUFFER 6 + typedef struct IgbVfMigState { bool migration_cap; MemoryRegion mig_bar; + + uint32_t mig_state; + uint8_t mig_error; + uint8_t mig_data[IGB_VF_STATE_MAX_SIZE]; + uint32_t mig_data_size; + uint64_t mig_data_buf_addr; } IgbVfMigState; =20 +typedef struct IgbVfState IgbVfState; void igb_pf_init_migration_bar(PCIDevice *dev); bool igbvf_add_migration_cap(PCIDevice *dev, Error **errp); +void igbvf_mig_bar_init(IgbVfState *s); +void igbvf_mig_state_reset(IgbVfState *s); =20 #endif diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index 794f1217d3e5..a0d044815fd9 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -7,9 +7,13 @@ */ =20 #include "qemu/osdep.h" +#include "qemu/log.h" #include "hw/pci/pci_device.h" #include "hw/pci/pcie.h" +#include "net/eth.h" +#include "net/net.h" #include "igb_common.h" +#include "igb_core.h" #include "igb_migration.h" #include "trace.h" =20 @@ -70,3 +74,292 @@ bool igbvf_add_migration_cap(PCIDevice *dev, Error **er= rp) trace_igbvf_mig_cap_add(pcie_sriov_vf_number(dev), offset); return true; } + +/* + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * Per-VF state serialization / deserialization + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + */ + +static int igb_core_vf_save_state(IgbVfState *s, + void *buf, size_t buf_size) +{ + int size =3D 0; + + trace_igbvf_mig_save_state(s->vfn, size); + return size; +} + +static int igb_core_vf_max_data_size(IgbVfState *s) +{ + int size =3D igb_core_vf_save_state(s, NULL, 0); + + g_assert(size > 0 && size <=3D IGB_VF_STATE_MAX_SIZE); + return size; +} + +static int igb_core_vf_load_state(IgbVfState *s, + const void *buf, size_t size) +{ + trace_igbvf_mig_load_state(s->vfn, (uint32_t)size); + return 0; +} + +static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size) +{ + int ret; + + ret =3D igb_core_vf_load_state(s, buf, size); + if (ret < 0) { + return ret; + } + + return 0; +} + +/* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * Migration BAR register read/write handlers + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D */ + +static bool igbvf_mig_set_state(IgbVfState *s, uint32_t new_state) +{ + IgbVfMigState *ms =3D &s->mig; + uint32_t old =3D ms->mig_state; + int ret; + + switch (new_state) { + case IGB_MIG_STATE_STOP: + if (old !=3D IGB_MIG_STATE_RUNNING && + old !=3D IGB_MIG_STATE_STOP_COPY && + old !=3D IGB_MIG_STATE_RESUMING && + old !=3D IGB_MIG_STATE_ERROR) { + return false; + } + /* Restore DATA_SIZE to max, same as at reset */ + ms->mig_data_size =3D igb_core_vf_max_data_size(s); + break; + + case IGB_MIG_STATE_RUNNING: + if (old !=3D IGB_MIG_STATE_STOP) { + return false; + } + break; + + case IGB_MIG_STATE_STOP_COPY: + if (old !=3D IGB_MIG_STATE_STOP) { + return false; + } + ret =3D igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_dat= a)); + if (ret < 0) { + ms->mig_error =3D -ret; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + return false; + } + ms->mig_data_size =3D ret; + break; + + case IGB_MIG_STATE_RESUMING: + if (old !=3D IGB_MIG_STATE_STOP) { + return false; + } + memset(ms->mig_data, 0, sizeof(ms->mig_data)); + ms->mig_data_size =3D 0; + break; + + default: + trace_igbvf_mig_set_state_err(s->vfn, old, new_state); + return false; + } + + ms->mig_state =3D new_state; + trace_igbvf_mig_set_state(s->vfn, old, new_state); + return true; +} + +static uint32_t igbvf_mig_get_status(IgbVfState *s) +{ + IgbVfMigState *ms =3D &s->mig; + uint32_t status =3D 0; + + if (ms->mig_state =3D=3D IGB_MIG_STATE_ERROR) { + status |=3D IGB_MIG_STATUS_ERR(ms->mig_error); + } + if (ms->mig_state =3D=3D IGB_MIG_STATE_STOP_COPY && ms->mig_data_size = > 0) { + status |=3D IGB_MIG_STATUS_DATA_AVAIL; + } + + return status; +} + +static void igbvf_mig_data_xfer(IgbVfState *s, uint32_t val) +{ + IgbVfMigState *ms =3D &s->mig; + MemTxResult r; + int ret; + + if (!ms->mig_data_buf_addr) { + ms->mig_error =3D IGB_MIG_ERR_NO_BUFFER; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + return; + } + + switch (ms->mig_state) { + case IGB_MIG_STATE_STOP_COPY: + /* Save: DMA-write serialized state to driver buffer */ + r =3D pci_dma_write(pcie_sriov_get_pf(PCI_DEVICE(s)), + ms->mig_data_buf_addr, + ms->mig_data, ms->mig_data_size); + if (r !=3D MEMTX_OK) { + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u state write failed at 0x%" PRIx64 "= \n", + s->vfn, ms->mig_data_buf_addr); + ms->mig_error =3D IGB_MIG_ERR_DMA_FAILED; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + } + break; + + case IGB_MIG_STATE_RESUMING: + /* Restore: DMA-read state from driver buffer and deserialize */ + if (ms->mig_data_size =3D=3D 0 || + ms->mig_data_size > sizeof(ms->mig_data)) { + ms->mig_error =3D IGB_MIG_ERR_BAD_SIZE; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + break; + } + + r =3D pci_dma_read(pcie_sriov_get_pf(PCI_DEVICE(s)), + ms->mig_data_buf_addr, + ms->mig_data, ms->mig_data_size); + if (r !=3D MEMTX_OK) { + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u state read failed at 0x%" PRIx64 "\= n", + s->vfn, ms->mig_data_buf_addr); + ms->mig_error =3D IGB_MIG_ERR_DMA_FAILED; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + break; + } + + ret =3D igbvf_mig_load(s, ms->mig_data, ms->mig_data_size); + if (ret < 0) { + ms->mig_error =3D -ret; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + } + break; + + default: + break; + } +} + +static uint64_t igbvf_mig_read(void *opaque, hwaddr addr, unsigned size) +{ + IgbVfState *s =3D opaque; + IgbVfMigState *ms =3D &s->mig; + uint64_t val =3D 0; + + switch (addr) { + case IGB_MIG_DEVICE_STATE: + val =3D ms->mig_state; + break; + case IGB_MIG_STATUS: + val =3D igbvf_mig_get_status(s); + break; + case IGB_MIG_CAPS: + val =3D IGB_MIG_CAP_F_STATE; + break; + case IGB_MIG_VERSION: + val =3D IGB_MIG_CAP_VERSION; + break; + case IGB_MIG_DATA_SIZE: + val =3D ms->mig_data_size; + break; + default: + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u bad migration BAR read at 0x%" + HWADDR_PRIx "\n", s->vfn, addr); + break; + } + + trace_igbvf_mig_bar_read(s->vfn, addr, val); + + return val; +} + +static void igbvf_mig_write(void *opaque, hwaddr addr, uint64_t val, + unsigned size) +{ + IgbVfState *s =3D opaque; + IgbVfMigState *ms =3D &s->mig; + + trace_igbvf_mig_bar_write(s->vfn, addr, val); + + switch (addr) { + case IGB_MIG_DEVICE_STATE: + igbvf_mig_set_state(s, (uint32_t)val); + break; + case IGB_MIG_DATA_SIZE: + if (val <=3D sizeof(ms->mig_data)) { + ms->mig_data_size =3D (uint32_t)val; + } else { + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u DATA_SIZE %" PRIu64 " exceeds max %= zu\n", + s->vfn, val, sizeof(ms->mig_data)); + } + break; + case IGB_MIG_DATA_XFER: + igbvf_mig_data_xfer(s, (uint32_t)val); + break; + case IGB_MIG_DATA_BUF_ADDR_LO: + ms->mig_data_buf_addr =3D + deposit64(ms->mig_data_buf_addr, 0, 32, val); + break; + case IGB_MIG_DATA_BUF_ADDR_HI: + ms->mig_data_buf_addr =3D + deposit64(ms->mig_data_buf_addr, 32, 32, val); + break; + default: + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u bad migration BAR write at 0x%" + HWADDR_PRIx "\n", s->vfn, addr); + break; + } +} + +static const MemoryRegionOps mig_bar_ops =3D { + .read =3D igbvf_mig_read, + .write =3D igbvf_mig_write, + .endianness =3D DEVICE_LITTLE_ENDIAN, + .impl =3D { + .min_access_size =3D 4, + .max_access_size =3D 4, + }, +}; + +/* + * Use the QEM-internal PCI_BASE_ADDRESS_MEM_ALWAYS_ON BAR type flag + * to keep the memory BAR always mapped. + */ +void igbvf_mig_bar_init(IgbVfState *s) +{ + IgbVfMigState *ms =3D &s->mig; + + memory_region_init_io(&ms->mig_bar, OBJECT(s), &mig_bar_ops, s, + "igbvf-mig", IGB_MIG_BAR_SIZE); + pci_register_bar(PCI_DEVICE(s), IGB_MIG_BAR_IDX, + PCI_BASE_ADDRESS_MEM_PREFETCH | + PCI_BASE_ADDRESS_MEM_ALWAYS_ON, + &ms->mig_bar); + trace_igbvf_mig_bar_init(s->vfn); +} + +void igbvf_mig_state_reset(IgbVfState *s) +{ + IgbVfMigState *ms =3D &s->mig; + + ms->mig_state =3D IGB_MIG_STATE_RUNNING; + ms->mig_error =3D 0; + ms->mig_data_size =3D igb_core_vf_max_data_size(s); + ms->mig_data_buf_addr =3D 0; + memset(ms->mig_data, 0, sizeof(ms->mig_data)); + trace_igbvf_mig_reset(s->vfn); +} diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c index 94c9739cd58c..e9f9fc3369d8 100644 --- a/hw/net/igbvf.c +++ b/hw/net/igbvf.c @@ -53,15 +53,6 @@ =20 OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF) =20 -struct IgbVfState { - PCIDevice parent_obj; - - MemoryRegion mmio; - MemoryRegion msix; - - IgbVfMigState mig; -}; - static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write) { switch (addr) { @@ -281,6 +272,10 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **= errp) if (!igbvf_add_migration_cap(dev, errp)) { return; } + + s->vfn =3D pcie_sriov_vf_number(dev); + igbvf_mig_bar_init(s); + igbvf_mig_state_reset(s); } =20 if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)), @@ -298,8 +293,13 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **= errp) static void igbvf_qdev_reset_hold(Object *obj, ResetType type) { PCIDevice *vf =3D PCI_DEVICE(obj); + IgbVfState *s =3D IGBVF(obj); =20 igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf)); + + if (s->mig.migration_cap) { + igbvf_mig_state_reset(s); + } } =20 static void igbvf_pci_uninit(PCIDevice *dev) diff --git a/hw/net/trace-events b/hw/net/trace-events index 06d8848023e5..0b13a99b3f32 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -295,6 +295,14 @@ igb_wrn_rx_desc_modes_not_supp(int desc_type) "Not sup= ported descriptor type: %d # igbvf.c igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64 igbvf_mig_cap_add(uint16_t vfn, int offset) "VF%u: added migration vendor = cap at config offset 0x%x" +igbvf_mig_bar_init(uint16_t vfn) "VF%u: migration BAR initialized" +igbvf_mig_bar_read(uint16_t vfn, uint64_t addr, uint64_t val) "VF%u: BAR r= ead addr=3D0x%"PRIx64" val=3D0x%"PRIx64 +igbvf_mig_bar_write(uint16_t vfn, uint64_t addr, uint64_t val) "VF%u: BAR = write addr=3D0x%"PRIx64" val=3D0x%"PRIx64 +igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) = "VF%u: state %u -> %u" +igbvf_mig_set_state_err(uint16_t vfn, uint32_t old_state, uint32_t new_sta= te) "VF%u: invalid transition %u -> %u" +igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of= device state" +igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes o= f device state" +igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset" =20 # spapr_llan.c spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_= bufs) "pool=3D%d count=3D%"PRId32" rxbufs=3D%"PRIu32 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130825; cv=none; d=zohomail.com; s=zohoarc; b=On77+a7IX2wSKrN9yMdNOCd8dvrCqEfIA4veXxwNFsBkBTcitk7G98BDXmHVW3mXeTItxlgpGtZ8oVBUL/WQZrPG1MM1mG5djuuudKwT7uOf3NfbgBB7Af/o6JsqJ0017TF2wOOBrh99bjzESxe1slXlrYu/yg/Y73nKBo27SEw= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130825; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=olDliIW7mfe1ldZof51Y2tBein2BaeyJp0eYsLkADPU=; b=fFsdd5UIXvY4TU1XODlzCQksffz7HCFjiUP3yEGp9uh9NoJuKNG4jiaUMgKXFLuM085MS05a6JDmWv955qSKXNE/osAnobLAC9VYEb/GhrcNhqLTtx1tTG2q2+ct+UyWCJsqEQYs1GImdNoi5t5VYejveZFsjbplKDDCswIoA2U= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130825868565.8858158283103; Sun, 26 Jul 2026 22:40:25 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4S-0003TF-BY; Mon, 27 Jul 2026 01:40:12 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4K-00039t-B0 for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:04 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4H-0003WF-PH for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:03 -0400 Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-582-6bv20D-MNiy9gpngUDaoyQ-1; Mon, 27 Jul 2026 01:39:57 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 038CD1955F2D; Mon, 27 Jul 2026 05:39:56 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 3042F1956088; Mon, 27 Jul 2026 05:39:52 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130801; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=olDliIW7mfe1ldZof51Y2tBein2BaeyJp0eYsLkADPU=; b=bUOqjfXQiL8qzrEsDdZsCgjeb6L+gP6MXv92qwqTuRJETxwnZJ1iWR40F+BlgefBuMKlkv qnRDNPGAcmGbGzv3B8Th5Ej2ai7jjmXU47pXUaCH/10+E8aadqOC7J6DDyElKNXs0wRGb6 FZ1g6DA7bJnLbruCfius24GEIMMWn00= X-MC-Unique: 6bv20D-MNiy9gpngUDaoyQ-1 X-Mimecast-MFC-AGG-ID: 6bv20D-MNiy9gpngUDaoyQ_1785130796 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 04/11] igb: Add VF state serialization for live migration Date: Mon, 27 Jul 2026 07:39:28 +0200 Message-ID: <20260727053935.1392269-5-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130826852158500 Add igb_pf_get_core() so migration code can reach the PF's IGBCore from a VF device and implement igb_core_vf_save_state() and igb_core_vf_load_state() to serialize and restore per-VF device state through the migration BAR. The wire format is a versioned blob: header (magic, version, VF number, register count), offset/value pairs for per-VF registers, dynamically scanned RA/RA2 entries, and TX context descriptors. PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) are saved instead of the PF aggregates which the L1 driver may have transiently cleared. The load path validates the header, restores registers to mac[], syncs EITR to eitr_guest_value[], and restores TX context. MSI-X table/PBA is not saved - L1's VFIO reprograms it after migration. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- hw/net/igb_core.h | 1 + hw/net/igb.c | 6 + hw/net/igb_migration.c | 277 ++++++++++++++++++++++++++++++++++++++++- 3 files changed, 283 insertions(+), 1 deletion(-) diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index d70b54e318f1..58d4f57c99bb 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -143,4 +143,5 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov,= int iovcnt); void igb_start_recv(IGBCore *core); =20 +IGBCore *igb_pf_get_core(void *pf); #endif diff --git a/hw/net/igb.c b/hw/net/igb.c index b43235996db2..222413dd237a 100644 --- a/hw/net/igb.c +++ b/hw/net/igb.c @@ -134,6 +134,12 @@ void igb_vf_reset(void *opaque, uint16_t vfn) igb_core_vf_reset(&s->core, vfn); } =20 +IGBCore *igb_pf_get_core(void *pf) +{ + IGBState *s =3D IGB(pf); + return &s->core; +} + static bool igb_io_get_reg_index(IGBState *s, uint32_t *idx) { diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index a0d044815fd9..61cf155a188d 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -75,16 +75,226 @@ bool igbvf_add_migration_cap(PCIDevice *dev, Error **e= rrp) return true; } =20 +static IGBCore *igbvf_get_core(IgbVfState *s) +{ + return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s))); +} + /* * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D * Per-VF state serialization / deserialization * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * + * Wire format: + * uint32_t magic (IGB_MIG_CAP_MAGIC) + * uint32_t version (1) + * uint32_t vfn (VF number) + * uint32_t num_regs (total register pairs, fixed + RA) + * { uint32_t offset; uint32_t value; } regs[num_regs] + * uint32_t num_tx_ctx (number of TX queue context blocks) + * { raw struct igb_tx data } tx_ctx[num_tx_ctx] */ =20 +/* Maximum number of registers in the VF state slice */ +#define IGB_VF_MAX_REGS 128 + +/* Register offsets that constitute a VF's state slice */ +static void igb_vf_reg_list(uint16_t vfn, uint32_t *offsets, int *count) +{ + int n =3D 0; + int q0 =3D vfn; + int q1 =3D vfn + IGB_NUM_VM_POOLS; + + /* Per-VF control and interrupt registers */ + offsets[n++] =3D E1000_PVTCTRL(vfn) >> 2; + offsets[n++] =3D E1000_PVTEICS(vfn) >> 2; + offsets[n++] =3D E1000_PVTEIMS(vfn) >> 2; + offsets[n++] =3D E1000_PVTEIMC(vfn) >> 2; + offsets[n++] =3D E1000_PVTEIAC(vfn) >> 2; + offsets[n++] =3D E1000_PVTEIAM(vfn) >> 2; + offsets[n++] =3D E1000_PVTEICR(vfn) >> 2; + + /* Per-VF statistics */ + offsets[n++] =3D E1000_PVFGPRC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGPTC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGORC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGOTC(vfn) >> 2; + offsets[n++] =3D E1000_PVFMPRC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGPRLBC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGPTLBC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGORLBC(vfn) >> 2; + offsets[n++] =3D E1000_PVFGOTLBC(vfn) >> 2; + + /* Mailbox */ + offsets[n++] =3D E1000_V2PMAILBOX(vfn) >> 2; + offsets[n++] =3D E1000_P2VMAILBOX(vfn) >> 2; + + /* Per-VF config */ + offsets[n++] =3D E1000_VMOLR(vfn) >> 2; + offsets[n++] =3D E1000_VMVIR(vfn) >> 2; + offsets[n++] =3D E1000_PSRTYPE(vfn) >> 2; + + /* + * VF receive addresses (RA/RA2) are saved dynamically in + * igb_core_vf_save_state by scanning for entries whose pool + * bits match this VF - the PF driver chooses the RA slot. + */ + + /* Interrupt routing */ + offsets[n++] =3D (E1000_VTIVAR + vfn * 4) >> 2; + offsets[n++] =3D (E1000_VTIVAR_MISC + vfn * 4) >> 2; + + /* + * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF. + * Each VF has 3 MSI-X vectors, each with its own EITR controlling + * interrupt coalescing. Without saving these, interrupt + * throttling resets to zero after migration which can cause + * interrupt storms or latency changes. VF N uses PF EITR indices + * (22 - N*3) .. (24 - N*3). + */ + { + int eitr_base =3D 22 - vfn * 3; + offsets[n++] =3D E1000_EITR(eitr_base) >> 2; + offsets[n++] =3D E1000_EITR(eitr_base + 1) >> 2; + offsets[n++] =3D E1000_EITR(eitr_base + 2) >> 2; + } + + /* RX and TX queue registers for queues q0 and q1 */ +#define ADD_QUEUE_REGS(q) do { \ + offsets[n++] =3D E1000_RDBAL(q) >> 2; \ + offsets[n++] =3D E1000_RDBAH(q) >> 2; \ + offsets[n++] =3D E1000_RDLEN(q) >> 2; \ + offsets[n++] =3D E1000_SRRCTL(q) >> 2; \ + offsets[n++] =3D E1000_RDH(q) >> 2; \ + offsets[n++] =3D E1000_RDT(q) >> 2; \ + offsets[n++] =3D E1000_RXDCTL(q) >> 2; \ + offsets[n++] =3D E1000_RXCTL(q) >> 2; \ + offsets[n++] =3D E1000_RQDPC(q) >> 2; \ + offsets[n++] =3D E1000_TDBAL(q) >> 2; \ + offsets[n++] =3D E1000_TDBAH(q) >> 2; \ + offsets[n++] =3D E1000_TDLEN(q) >> 2; \ + offsets[n++] =3D E1000_TDH(q) >> 2; \ + offsets[n++] =3D E1000_TDT(q) >> 2; \ + offsets[n++] =3D E1000_TXDCTL(q) >> 2; \ + offsets[n++] =3D E1000_TXCTL(q) >> 2; \ + offsets[n++] =3D E1000_TDWBAL(q) >> 2; \ + offsets[n++] =3D E1000_TDWBAH(q) >> 2; \ +} while (0) + + ADD_QUEUE_REGS(q0); + ADD_QUEUE_REGS(q1); +#undef ADD_QUEUE_REGS + + g_assert(n <=3D IGB_VF_MAX_REGS); + *count =3D n; +} + +/* + * Scan RA and RA2 arrays for receive address entries assigned to + * this VF. The PF driver picks the RA slot, so we cannot use a + * fixed index - instead check each entry's pool bits. + */ +static uint32_t *igb_core_vf_save_ra(IGBCore *core, uint16_t vfn, + uint32_t *p, int *total_regs) +{ + uint32_t vf_pool_bit =3D E1000_RAH_POOL_1 << vfn; + static const struct { + uint32_t base; + int count; + } ra_banks[] =3D { + { RA, 16 }, + { RA2, 8 }, + }; + int i, j; + + for (i =3D 0; i < ARRAY_SIZE(ra_banks); i++) { + for (j =3D 0; j < ra_banks[i].count; j++) { + uint32_t ral_off =3D ra_banks[i].base + j * 2; + uint32_t rah_off =3D ra_banks[i].base + j * 2 + 1; + uint32_t rah_val =3D core->mac[rah_off]; + + if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) { + *p++ =3D cpu_to_le32(ral_off); + *p++ =3D cpu_to_le32(core->mac[ral_off]); + *p++ =3D cpu_to_le32(rah_off); + *p++ =3D cpu_to_le32(rah_val); + *total_regs +=3D 2; + } + } + } + return p; +} + +static uint32_t *igb_core_vf_save_tx_ctx(IGBCore *core, int queue, + uint32_t *p) +{ + memcpy(p, &core->tx[queue], sizeof(struct igb_tx)); + return (uint32_t *)((uint8_t *)p + sizeof(struct igb_tx)); +} + +static size_t igb_core_vf_state_max_size(int num_fixed_regs) +{ + int max_ra_entries =3D 16 + 8; /* RA bank (16) + RA2 bank (8) */ + int max_ra_regs =3D max_ra_entries * 2; /* RAL + RAH per entry */ + + return 4 * sizeof(uint32_t) /* header */ + + num_fixed_regs * 2 * sizeof(uint32_t) /* fixed reg pairs */ + + max_ra_regs * 2 * sizeof(uint32_t) /* RA reg pairs */ + + sizeof(uint32_t) /* num_tx_ctx */ + + 2 * sizeof(struct igb_tx); /* TX context */ +} + static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size) { - int size =3D 0; + IGBCore *core =3D igbvf_get_core(s); + uint32_t offsets[IGB_VF_MAX_REGS]; + int num_regs, total_regs; + uint32_t *p =3D buf; + uint32_t *num_regs_p; + int i, size; + int q0 =3D s->vfn; + int q1 =3D s->vfn + IGB_NUM_VM_POOLS; + + /* + * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of + * extracting from PF aggregates - the L1 PF driver may have + * transiently cleared EIMS via EIMC. The load path ORs them back. + */ + igb_vf_reg_list(s->vfn, offsets, &num_regs); + + if (!buf) { + return igb_core_vf_state_max_size(num_regs); + } + + if (igb_core_vf_state_max_size(num_regs) > buf_size) { + return -IGB_MIG_ERR_BAD_SIZE; + } + + /* Header: magic, version, vfn, num_regs (updated below) */ + *p++ =3D cpu_to_le32(IGB_MIG_CAP_MAGIC); + *p++ =3D cpu_to_le32(1); /* version */ + *p++ =3D cpu_to_le32(s->vfn); + num_regs_p =3D p; + *p++ =3D cpu_to_le32(num_regs); + + for (i =3D 0; i < num_regs; i++) { + *p++ =3D cpu_to_le32(offsets[i]); + *p++ =3D cpu_to_le32(core->mac[offsets[i]]); + } + + total_regs =3D num_regs; + + p =3D igb_core_vf_save_ra(core, s->vfn, p, &total_regs); + + *num_regs_p =3D cpu_to_le32(total_regs); + + /* TX context descriptors for this VF's two queues */ + *p++ =3D cpu_to_le32(2); /* num_tx_ctx */ + p =3D igb_core_vf_save_tx_ctx(core, q0, p); + p =3D igb_core_vf_save_tx_ctx(core, q1, p); + + size =3D (uint8_t *)p - (uint8_t *)buf; =20 trace_igbvf_mig_save_state(s->vfn, size); return size; @@ -98,9 +308,74 @@ static int igb_core_vf_max_data_size(IgbVfState *s) return size; } =20 +static const void *igb_core_vf_load_tx_ctx(IGBCore *core, int queue, + const void *data) +{ + struct NetTxPkt *saved_pkt =3D core->tx[queue].tx_pkt; + + memcpy(&core->tx[queue], data, sizeof(struct igb_tx)); + core->tx[queue].tx_pkt =3D saved_pkt; + return (const uint8_t *)data + sizeof(struct igb_tx); +} + static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size) { + IGBCore *core =3D igbvf_get_core(s); + const uint32_t *p =3D buf; + uint32_t magic, version, saved_vfn, num_regs, num_tx; + int i; + int q0 =3D s->vfn; + int q1 =3D s->vfn + IGB_NUM_VM_POOLS; + + magic =3D le32_to_cpu(*p++); + version =3D le32_to_cpu(*p++); + saved_vfn =3D le32_to_cpu(*p++); + num_regs =3D le32_to_cpu(*p++); + + if (magic !=3D IGB_MIG_CAP_MAGIC) { + return -IGB_MIG_ERR_BAD_MAGIC; + } + if (version !=3D IGB_MIG_CAP_VERSION) { + return -IGB_MIG_ERR_BAD_VERSION; + } + if (saved_vfn !=3D s->vfn) { + return -IGB_MIG_ERR_BAD_VFN; + } + if (num_regs > IGB_VF_MAX_REGS) { + return -IGB_MIG_ERR_BAD_SIZE; + } + + for (i =3D 0; i < num_regs; i++) { + uint32_t offset =3D le32_to_cpu(*p++); + uint32_t value =3D le32_to_cpu(*p++); + + if (offset < E1000E_MAC_SIZE) { + core->mac[offset] =3D value; + + /* + * Sync EITR to eitr_guest_value[] shadow array, stripping + * E1000_EITR_CNT_IGNR so guest register readback returns + * the correct value. + */ + if (offset >=3D EITR0 && offset < EITR0 + IGB_INTR_NUM) { + core->eitr_guest_value[offset - EITR0] =3D + value & ~E1000_EITR_CNT_IGNR; + } + } + } + + num_tx =3D le32_to_cpu(*p++); + if (num_tx =3D=3D 2) { + p =3D igb_core_vf_load_tx_ctx(core, q0, p); + p =3D igb_core_vf_load_tx_ctx(core, q1, p); + } + + /* + * MSI-X table/PBA is not saved - L1's VFIO reprograms it with + * destination-specific IRTE references after migration. + */ + trace_igbvf_mig_load_state(s->vfn, (uint32_t)size); return 0; } --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130866; cv=none; d=zohomail.com; s=zohoarc; b=lmno3AaiM3A2PzagCe0+huDwqRoPZ9vE3btrEFKeCBfz6yS2lqBO79Qvv+6/ssNSjdUVZumAINBIY5Zy2Zem3MhDOibW89twY/nt5nsgFGQ9vrMKjKndVpW1wi6AkRbzsLkotXNDmZJ8luv8nNY+VOVCdIvoGKoz1LKtJ1JtlOo= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130866; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=k0a1wivckEfS8wGNJKXHzTIYJDleTc2ip6NMFUsCIMc=; b=ffwGsQPWAeWVv/ZlvpU8SDKFMBWzi5OLQKHb7FFlKh+IekYcXafqEHYmIoVvVSpoHMVPzMXaL5BeWu+2oPIGfV7M6DZ8ozHkzMvMyqQTdcOVyfpt9qiijG3XIGFreaXElGnF/YUsshakD1DTCHe6PJQVlXq1vdDmPStJMcCYaBg= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130866550110.51564170840459; Sun, 26 Jul 2026 22:41:06 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4U-0003ZU-Fp; Mon, 27 Jul 2026 01:40:14 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4N-0003G8-Tp for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:09 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4M-0003jR-8I for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:07 -0400 Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-353-FZ_DUatdMBypug5zSgj6tw-1; Mon, 27 Jul 2026 01:40:01 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id D95871955F22; Mon, 27 Jul 2026 05:39:59 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 71EC31956088; Mon, 27 Jul 2026 05:39:56 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130805; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=k0a1wivckEfS8wGNJKXHzTIYJDleTc2ip6NMFUsCIMc=; b=iUshBrrXsWa5ID9BmlqIJA6qvRjB2+DRCz9tL51VKXO5z2XWixgqx9jEUr3kqTPodiyIb2 lJ/gkIw4jKt1T0bTCNFPh7KbWoxv3VPrrqwH4X/qbvUewrAZYUVhDQu6OBVKzCSUQHkwhS ZrdBzU393CO/VPEUfV262N1f71fA6fs= X-MC-Unique: FZ_DUatdMBypug5zSgj6tw-1 X-Mimecast-MFC-AGG-ID: FZ_DUatdMBypug5zSgj6tw_1785130800 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 05/11] igb: Add VF post-load fixups for live migration Date: Mon, 27 Jul 2026 07:39:29 +0200 Message-ID: <20260727053935.1392269-6-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.129.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H2=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130868732158500 Add post-load fixups in igbvf_mig_load() to propagate VF interrupt state that the register load path bypasses by writing directly to mac[] without triggering register handler side effects. igb_core_vf_propagate_irqs() ORs PVT shadow values back into the PF aggregates (EIMS/EIAC/EIAM) and clears stale VF bits from EICR. igb_core_vf_propagate_ivar() re-applies VTIVAR routing to the shared IVAR0 entries that the L1 PF driver may have overwritten after L0 vmstate restore. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- hw/net/igb_core.h | 2 ++ hw/net/igb_core.c | 56 ++++++++++++++++++++++++++++++++++++++++++ hw/net/igb_migration.c | 7 ++++++ 3 files changed, 65 insertions(+) diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index 58d4f57c99bb..150567eb346d 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -144,4 +144,6 @@ void igb_start_recv(IGBCore *core); =20 IGBCore *igb_pf_get_core(void *pf); +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn); +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn); #endif diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c index 45d8fd795b84..2565dd7f96d3 100644 --- a/hw/net/igb_core.c +++ b/hw/net/igb_core.c @@ -4550,3 +4550,59 @@ igb_core_post_load(IGBCore *core) =20 return 0; } + +/* + * Propagate VF interrupt state to PF aggregates after loading VF + * registers. The load path writes directly to mac[] bypassing the + * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear + * stale VF bits in EICR that may have been set by packets arriving + * between PF vmstate restore and VF state load. + */ +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn) +{ + uint32_t shift =3D 22 - vfn * IGBVF_MSIX_VEC_NUM; + uint32_t pvt_idx; + + pvt_idx =3D PVTEIMS0 + vfn * 0x40; + core->mac[EIMS] |=3D (core->mac[pvt_idx] & 0x7) << shift; + pvt_idx =3D PVTEIAC0 + vfn * 0x40; + core->mac[EIAC] |=3D (core->mac[pvt_idx] & 0x7) << shift; + pvt_idx =3D PVTEIAM0 + vfn * 0x40; + core->mac[EIAM] |=3D (core->mac[pvt_idx] & 0x7) << shift; + + core->mac[EICR] &=3D ~(0x7 << shift); +} + +/* + * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver + * may have overwritten the shared IVAR0 entries with its own + * queue routing after L0 vmstate restore. + */ +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn) +{ + uint32_t vtivar =3D core->mac[VTIVAR + vfn]; + int n; + uint8_t ent; + uint32_t mask; + + if (vtivar & E1000_IVAR_VALID) { + n =3D igb_ivar_entry_rx(vfn); + ent =3D E1000_IVAR_VALID | + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7))); + mask =3D 0xffU << (8 * (n % 4)); + core->mac[IVAR0 + n / 4] =3D + (core->mac[IVAR0 + n / 4] & ~mask) | + ((uint32_t)ent << (8 * (n % 4))); + } + + ent =3D vtivar >> 8; + if (ent & E1000_IVAR_VALID) { + n =3D igb_ivar_entry_tx(vfn); + ent =3D E1000_IVAR_VALID | + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7))); + mask =3D 0xffU << (8 * (n % 4)); + core->mac[IVAR0 + n / 4] =3D + (core->mac[IVAR0 + n / 4] & ~mask) | + ((uint32_t)ent << (8 * (n % 4))); + } +} diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index 61cf155a188d..4f123df6795d 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -382,6 +382,7 @@ static int igb_core_vf_load_state(IgbVfState *s, =20 static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size) { + IGBCore *core =3D igbvf_get_core(s); int ret; =20 ret =3D igb_core_vf_load_state(s, buf, size); @@ -389,6 +390,12 @@ static int igbvf_mig_load(IgbVfState *s, const void *b= uf, size_t size) return ret; } =20 + /* + * Post-load: sync VF interrupt and routing state to PF aggregates + */ + igb_core_vf_propagate_irqs(core, s->vfn); + igb_core_vf_propagate_ivar(core, s->vfn); + return 0; } =20 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130899; cv=none; d=zohomail.com; s=zohoarc; b=SrixDdKDokuEWzTzJ+wFC8EZ/lMaRDDQRkZTzShiLBPoa6O1KZtebKR9dQCd4RrjyL8maI/pGpvrlMbYEco8oVlLWGrNZ8QRaTR2zFi4YKBwXtDYOXWFcbbcq5j6KQSvHc1mYea5hsfbghAV4a3KNUXJ7CPCBvfPuWKaA6KUgIg= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130899; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=EhtBfEWfcH0eCIwUfpQ2j9BpkmYL2KbQaPZ7lhlJ9Ok=; b=WXsML3ygWRIVDm24ro2dCbNA9lJQd3Imjumal5/kwaRuGRk98bLcJcuHWaO99bxwzj8Pil8nKxTlyUWC+la+nwfl2WSZOOvq2NtjgGjZ42u3vX7CJ1WgOQ16LIhfUblwBC8xpN5Avn9Z++rwZRGuCOBHvYnO1+nghYo3xtHTQTU= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130899427685.6048998282516; Sun, 26 Jul 2026 22:41:39 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4V-0003bl-EQ; Mon, 27 Jul 2026 01:40:15 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4S-0003WV-UR for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:13 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4O-0003k8-TM for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:12 -0400 Received: from mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-480-OMazQusDOVaf88AaEFwGew-1; Mon, 27 Jul 2026 01:40:04 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-06.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 022541800605; Mon, 27 Jul 2026 05:40:03 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 5A2151956088; Mon, 27 Jul 2026 05:40:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130807; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=EhtBfEWfcH0eCIwUfpQ2j9BpkmYL2KbQaPZ7lhlJ9Ok=; b=OITKzrO/lif84/WOvHu5/RhlaY6RQ3ayCXkchPUI+EkVlJ7SztWHm9BrUKMVpjdMPIBSnM fMJwTalkZWnPV2if7NuIT5+Rd/GfWxBi9gH5iR9QAT4fI+bSc+hPZnHjxm725vHSfjWGYW 0tsR5xX8L0m1mh6JVXDMpSyu8Df6GZo= X-MC-Unique: OMazQusDOVaf88AaEFwGew-1 X-Mimecast-MFC-AGG-ID: OMazQusDOVaf88AaEFwGew_1785130803 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 06/11] igb: Add dirty page tracking for IGBVF migration Date: Mon, 27 Jul 2026 07:39:30 +0200 Message-ID: <20260727053935.1392269-7-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -36 X-Spam_score: -3.7 X-Spam_bar: --- X-Spam_report: (-3.7 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130901123158500 Extend the migration BAR with dirty page tracking registers: 0x020 DIRTY_PGSIZE (RW) - dirty tracking page granularity 0x024 DIRTY_CTRL (WO) - 0=3DDISABLE, 1=3DENABLE, 2=3DQUERY 0x028 DIRTY_RANGE_IOVA_LO (WO) - low 32 bits of tracked range start 0x02C DIRTY_RANGE_IOVA_HI (WO) - high 32 bits of tracked range start 0x030 DIRTY_RANGE_SIZE (WO) - tracked range size in bytes 0x034 DIRTY_BUF_ADDR_LO (WO) - low 32 bits of shared buffer address 0x038 DIRTY_BUF_ADDR_HI (WO) - high 32 bits of shared buffer address 0x03C DIRTY_STATUS (RO) - result of last DIRTY_CTRL The CAPS register advertises the maximum number of ranges and supported page sizes. The driver enables tracking on IOVA ranges via DIRTY_CTRL and queries dirty bitmaps through a shared buffer. As for state transfers, dirty buffer DMA is performed through the PF device (pcie_sriov_get_pf). Each tracked range maintains its own bitmap scoped to its boundaries. DMA paths (TX/RX data, descriptor writeback) are instrumented to record the touched pages. This enables the PRE_COPY state where the driver iterates on dirty pages while the VM continues to run. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- docs/system/devices/igb-migration.rst | 73 +++++ hw/net/igb_core.h | 2 + hw/net/igb_migration.h | 82 ++++++ hw/net/igb_core.c | 48 ++-- hw/net/igb_migration.c | 367 +++++++++++++++++++++++++- hw/net/igbvf.c | 4 + hw/net/trace-events | 7 + 7 files changed, 562 insertions(+), 21 deletions(-) diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/ig= b-migration.rst index 04c44ef072af..4a3a93f01dfc 100644 --- a/docs/system/devices/igb-migration.rst +++ b/docs/system/devices/igb-migration.rst @@ -37,6 +37,14 @@ the following register layout:: 0x014 DATA_XFER WO Trigger DMA save or DMA load 0x018 DATA_BUF_ADDR_LO WO Low 32 bits of state DMA buffer addr= ess 0x01C DATA_BUF_ADDR_HI WO High 32 bits of state DMA buffer add= ress + 0x020 DIRTY_PGSIZE RW Dirty tracking page granularity + 0x024 DIRTY_CTRL WO 0=3DDISABLE, 1=3DENABLE, 2=3DQUERY + 0x028 DIRTY_RANGE_IOVA_LO WO Low 32 bits of tracked range start + 0x02C DIRTY_RANGE_IOVA_HI WO High 32 bits of tracked range start + 0x030 DIRTY_RANGE_SIZE WO Tracked range size in bytes + 0x034 DIRTY_BUF_ADDR_LO WO Low 32 bits of shared buffer address + 0x038 DIRTY_BUF_ADDR_HI WO High 32 bits of shared buffer address + 0x03C DIRTY_STATUS RO Result of last DIRTY_CTRL (0=3DOK, 1= -5=3Derror) =20 State transitions follow the VFIO migration state machine: the driver writes to ``DEVICE_STATE`` to move between states and reads ``STATUS`` @@ -61,3 +69,68 @@ code identifying the failure:: 4 BAD_VFN VF number mismatch (source !=3D destination) 5 DMA_FAILED DMA transfer to/from state buffer failed 6 NO_BUFFER DATA_XFER without buffer address set + +Dirty page tracking +~~~~~~~~~~~~~~~~~~~ + +The migration interface supports per-VF dirty page tracking, advertised +by the ``F_DIRTY`` flag in ``CAPS``. This allows the variant driver to +enter ``PRE_COPY`` state (``DEVICE_STATE`` =3D 5) while the VM continues +to run, iterating on dirty pages to reduce the final stop-and-copy +window. + +The device maintains one dirty tracking engine per range, each with its +own bitmap scoped to the range boundaries. The ``CAPS`` register +advertises the maximum number of ranges (``max_ranges`` in bits [11:8]) +and supported page sizes (bits [31:12]). + +Dirty tracking is controlled through the ``DIRTY_CTRL`` register: + +- **ENABLE** (1): the driver programs a tracked range via + ``DIRTY_RANGE_IOVA_LO/HI`` + ``DIRTY_RANGE_SIZE`` then writes + ``DIRTY_CTRL=3DENABLE``. The device allocates a fixed-size bitmap for + the range and begins recording pages touched by DMA (TX data, RX + data, descriptor writeback). The page granularity is set by + ``DIRTY_PGSIZE`` (default 4096). After each ENABLE the driver reads + ``DIRTY_STATUS`` to check for errors. +- **DISABLE** (0): tears down all ranges and stops tracking. +- **QUERY** (2): the driver writes (IOVA, size, page_size) into a + shared buffer registered via ``DIRTY_BUF_ADDR_LO/HI`` (PF DMA + address, as for state transfers), then writes + ``DIRTY_CTRL=3DQUERY``. The device finds the matching range, copies + the dirty bitmap into the shared buffer, clears the tracked bits, + and sets the buffer's completion status. + +``DIRTY_STATUS`` values after each ``DIRTY_CTRL`` write:: + + 0 OK Success + 1 TOO_MANY_RANGES Exceeds max_ranges from CAPS + 2 BAD_RANGE Invalid range (zero size) + 3 BAD_PGSIZE Invalid or misaligned page size + 4 NOT_ENABLED Query without prior enable + 5 NO_BUFFER Query without shared buffer + +Dirty query shared buffer +~~~~~~~~~~~~~~~~~~~~~~~~~ + +The shared buffer used for dirty queries is registered via +``DIRTY_BUF_ADDR_LO/HI`` (PF DMA address). It is cache-line aligned +(64 bytes) to separate driver-written request fields from +device-written completion fields:: + + Offset Field Written by Description + 0x00 iova driver Query range start + 0x08 size driver Query range size + 0x10 page_size driver Page granularity + 0x14 flags driver Reserved (must be 0) + 0x18 reserved[10] - Pad to 64-byte cache line + 0x40 status device 0 =3D pending, 1 =3D complete + 0x44 bitmap_size device Bytes written to bitmap + 0x48 dirty_page_count device Number of set bits in bitmap + 0x4C reserved[12] - Pad to 64-byte cache line + 0x80 bitmap[] device Dirty page bitmap + +The driver fills the request fields, issues ``DIRTY_CTRL=3DQUERY``, and +polls ``status`` for completion. The device reads the request, writes +the dirty bitmap and completion fields via DMA, then sets +``status =3D 1``. diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index 150567eb346d..50cee2c7683c 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -83,6 +83,8 @@ struct IGBCore { struct NetTxPkt *tx_pkt; } tx[IGB_NUM_QUEUES]; =20 + IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS]; + struct NetRxPkt *rx_pkt; =20 bool has_vnet; diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h index 739a189810b0..a2e7b363cb04 100644 --- a/hw/net/igb_migration.h +++ b/hw/net/igb_migration.h @@ -27,6 +27,12 @@ #define IGB_MIG_CAP_F_STATE (1u << 0) /* device state serialization */ #define IGB_MIG_CAP_F_DIRTY (1u << 1) /* dirty page tracking */ =20 +/* MIG_CAPS register layout (read-only, offset 0x008) */ +#define IGB_MIG_CAPS_MAX_RANGES_SHIFT 8 +#define IGB_MIG_CAPS_MAX_RANGES_MASK (0xfu << 8) /* bits [11:8] */ +#define IGB_MIG_CAPS_MAX_RANGES 4 +#define IGB_MIG_CAPS_PGSIZES_MASK 0xfffff000u /* bits [31:12] */ + #define IGB_MIG_CAP_SIZE 16 #define IGB_MIG_CAP_OFF_MAGIC 4 /* offset within cap for magic field */ #define IGB_MIG_CAP_OFF_BARID 8 /* offset within cap for BAR id */ @@ -49,6 +55,14 @@ #define IGB_MIG_DATA_XFER 0x014 #define IGB_MIG_DATA_BUF_ADDR_LO 0x018 #define IGB_MIG_DATA_BUF_ADDR_HI 0x01C +#define IGB_MIG_DIRTY_PGSIZE 0x020 +#define IGB_MIG_DIRTY_CTRL 0x024 +#define IGB_MIG_DIRTY_RANGE_IOVA_LO 0x028 +#define IGB_MIG_DIRTY_RANGE_IOVA_HI 0x02C +#define IGB_MIG_DIRTY_RANGE_SIZE 0x030 +#define IGB_MIG_DIRTY_BUF_ADDR_LO 0x034 +#define IGB_MIG_DIRTY_BUF_ADDR_HI 0x038 +#define IGB_MIG_DIRTY_STATUS 0x03C =20 /* DEVICE_STATE values - mirrors VFIO migration states */ #define IGB_MIG_STATE_ERROR 0 @@ -75,6 +89,50 @@ #define IGB_MIG_ERR_DMA_FAILED 5 #define IGB_MIG_ERR_NO_BUFFER 6 =20 +/* + * DIRTY_CTRL register values. Write one of these to control + * the dirty page tracking state machine. + */ +#define IGB_MIG_DIRTY_CTRL_DISABLE 0 +#define IGB_MIG_DIRTY_CTRL_ENABLE 1 +#define IGB_MIG_DIRTY_CTRL_QUERY 2 /* query-and-clear */ + +/* DIRTY_STATUS register values (read-only, cleared on next DIRTY_CTRL wri= te) */ +#define IGB_MIG_DIRTY_STATUS_OK 0 +#define IGB_MIG_DIRTY_STATUS_TOO_MANY_RANGES 1 +#define IGB_MIG_DIRTY_STATUS_BAD_RANGE 2 +#define IGB_MIG_DIRTY_STATUS_BAD_PGSIZE 3 +#define IGB_MIG_DIRTY_STATUS_NOT_ENABLED 4 +#define IGB_MIG_DIRTY_STATUS_NO_BUFFER 5 + +#define IGB_MIG_DIRTY_DEFAULT_PGSIZE 4096 + +/* + * Dirty query shared buffer layout - DMA between driver and device. + * The driver writes request fields, kicks DIRTY_CTRL=3DQUERY, and the + * device reads the request, fills bitmap + completion via DMA. Each + * section is cache-line aligned (64 bytes). + */ +struct igb_mig_dirty_query { + /* Cache line 0: request (written by driver) */ + uint64_t iova; + uint64_t size; + uint32_t page_size; + uint32_t flags; + uint32_t reserved0[10]; + + /* Cache line 1: completion (written by device) */ + uint32_t status; + uint32_t bitmap_size; + uint32_t dirty_page_count; + uint32_t reserved1[12]; + + /* Cache line 2+: bitmap (written by device) */ + uint8_t bitmap[]; +}; + +#define IGB_MIG_DIRTY_STATUS_COMPLETE 1 + typedef struct IgbVfMigState { bool migration_cap; MemoryRegion mig_bar; @@ -84,6 +142,12 @@ typedef struct IgbVfMigState { uint8_t mig_data[IGB_VF_STATE_MAX_SIZE]; uint32_t mig_data_size; uint64_t mig_data_buf_addr; + + uint32_t mig_dirty_pgsize; + uint64_t mig_dirty_range_iova; + uint32_t mig_dirty_range_size; + uint64_t mig_dirty_buf_addr; + uint32_t mig_dirty_status; } IgbVfMigState; =20 typedef struct IgbVfState IgbVfState; @@ -92,4 +156,22 @@ bool igbvf_add_migration_cap(PCIDevice *dev, Error **er= rp); void igbvf_mig_bar_init(IgbVfState *s); void igbvf_mig_state_reset(IgbVfState *s); =20 +typedef struct IGBVfDirtyRange { + uint64_t iova; + uint64_t size; + uint64_t page_size; + unsigned long *bitmap; + uint64_t nbits; +} IGBVfDirtyRange; + +typedef struct IGBVfDirtyState { + uint32_t num_ranges; + IGBVfDirtyRange ranges[IGB_MIG_CAPS_MAX_RANGES]; +} IGBVfDirtyState; + +typedef struct IGBCore IGBCore; +void igb_core_dirty_track_dma(IGBCore *core, int vfn, + dma_addr_t addr, dma_addr_t len); +void igb_core_vf_dirty_disable(IgbVfState *s); + #endif diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c index 2565dd7f96d3..03c349575c0a 100644 --- a/hw/net/igb_core.c +++ b/hw/net/igb_core.c @@ -824,6 +824,14 @@ igb_rx_ring_init(IGBCore *core, E1000E_RxRing *rxr, in= t idx) rxr->i =3D &i[idx]; } =20 +static inline void +igb_pci_dma_write(IGBCore *core, PCIDevice *dev, int vfn, + dma_addr_t addr, const void *buf, dma_addr_t len) +{ + pci_dma_write(dev, addr, buf, len); + igb_core_dirty_track_dma(core, vfn, addr, len); +} + static uint32_t igb_txdesc_writeback(IGBCore *core, dma_addr_t base, union e1000_adv_tx_desc *tx_desc, @@ -847,13 +855,15 @@ igb_txdesc_writeback(IGBCore *core, dma_addr_t base, =20 if (tdwba & 1) { uint32_t buffer =3D cpu_to_le32(core->mac[txi->dh]); - pci_dma_write(d, tdwba & ~3, &buffer, sizeof(buffer)); + igb_pci_dma_write(core, d, txi->idx % IGB_NUM_VM_POOLS, + tdwba & ~3, &buffer, sizeof(buffer)); } else { uint32_t status =3D le32_to_cpu(tx_desc->wb.status) | E1000_TXD_ST= AT_DD; =20 tx_desc->wb.status =3D cpu_to_le32(status); - pci_dma_write(d, base + offsetof(union e1000_adv_tx_desc, wb), - &tx_desc->wb, sizeof(tx_desc->wb)); + igb_pci_dma_write(core, d, txi->idx % IGB_NUM_VM_POOLS, + base + offsetof(union e1000_adv_tx_desc, wb), + &tx_desc->wb, sizeof(tx_desc->wb)); } =20 return igb_tx_wb_eic(core, txi->idx); @@ -1264,6 +1274,7 @@ typedef struct IGBPacketRxDMAState { size_t iov_ofs; bool do_ps; bool is_first; + int vfn; IGBBAState bastate; hwaddr ba[IGB_MAX_PS_BUFFERS]; IGBSplitDescriptorData ps_desc_data; @@ -1590,7 +1601,8 @@ igb_write_rx_descr(IGBCore *core, =20 static inline void igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice *dev, dma_addr_t addr, - union e1000_rx_desc_union *desc, dma_addr_t len) + union e1000_rx_desc_union *desc, dma_addr_t len, + int vfn) { if (igb_rx_use_legacy_descriptor(core)) { struct e1000_rx_desc *d =3D &desc->legacy; @@ -1598,11 +1610,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice = *dev, dma_addr_t addr, uint8_t status =3D d->status; =20 d->status &=3D ~E1000_RXD_STAT_DD; - pci_dma_write(dev, addr, desc, len); + igb_pci_dma_write(core, dev, vfn, addr, desc, len); =20 if (status & E1000_RXD_STAT_DD) { d->status =3D status; - pci_dma_write(dev, addr + offset, &status, sizeof(status)); + igb_pci_dma_write(core, dev, vfn, + addr + offset, &status, sizeof(status)); } } else { union e1000_adv_rx_desc *d =3D &desc->adv; @@ -1611,11 +1624,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice = *dev, dma_addr_t addr, uint32_t status =3D d->wb.upper.status_error; =20 d->wb.upper.status_error &=3D ~E1000_RXD_STAT_DD; - pci_dma_write(dev, addr, desc, len); + igb_pci_dma_write(core, dev, vfn, addr, desc, len); =20 if (status & E1000_RXD_STAT_DD) { d->wb.upper.status_error =3D status; - pci_dma_write(dev, addr + offset, &status, sizeof(status)); + igb_pci_dma_write(core, dev, vfn, + addr + offset, &status, sizeof(status)); } } } @@ -1737,9 +1751,9 @@ igb_write_hdr_frag_to_rx_buffers(IGBCore *core, { assert(data_len <=3D pdma_st->rx_desc_header_buf_size - pdma_st->bastate.written[0]); - pci_dma_write(d, - pdma_st->ba[0] + pdma_st->bastate.written[0], - data, data_len); + igb_pci_dma_write(core, d, pdma_st->vfn, + pdma_st->ba[0] + pdma_st->bastate.written[0], + data, data_len); pdma_st->bastate.written[0] +=3D data_len; pdma_st->bastate.cur_idx =3D 1; } @@ -1804,10 +1818,10 @@ igb_write_payload_frag_to_rx_buffers(IGBCore *core, data, bytes_to_write); =20 - pci_dma_write(d, - pdma_st->ba[pdma_st->bastate.cur_idx] + - pdma_st->bastate.written[pdma_st->bastate.cur_idx], - data, bytes_to_write); + igb_pci_dma_write(core, d, pdma_st->vfn, + pdma_st->ba[pdma_st->bastate.cur_idx] + + pdma_st->bastate.written[pdma_st->bastate.cur_id= x], + data, bytes_to_write); =20 pdma_st->bastate.written[pdma_st->bastate.cur_idx] +=3D bytes_to_w= rite; data +=3D bytes_to_write; @@ -1908,6 +1922,7 @@ igb_write_packet_to_guest(IGBCore *core, struct NetRx= Pkt *pkt, =20 rxi =3D rxr->i; rx_desc_len =3D core->rx_desc_len; + pdma_st.vfn =3D rxi->idx % IGB_NUM_VM_POOLS; pdma_st.rx_desc_packet_buf_size =3D igb_rxbufsize(core, rxi); pdma_st.rx_desc_header_buf_size =3D igb_rxhdrbufsize(core, rxi); pdma_st.iov =3D net_rx_pkt_get_iovec(pkt); @@ -1944,7 +1959,8 @@ igb_write_packet_to_guest(IGBCore *core, struct NetRx= Pkt *pkt, etqf, ts, &pdma_st, rxi); - igb_pci_dma_write_rx_desc(core, d, base, &desc, rx_desc_len); + igb_pci_dma_write_rx_desc(core, d, base, &desc, rx_desc_len, + rxi->idx % IGB_NUM_VM_POOLS); igb_ring_advance(core, rxi, rx_desc_len / E1000_MIN_RX_DESC_LEN); } while (pdma_st.desc_offset < pdma_st.total_size); =20 diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index 4f123df6795d..9a9ce902bb4a 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -8,6 +8,8 @@ =20 #include "qemu/osdep.h" #include "qemu/log.h" +#include "qemu/units.h" +#include "qemu/bitmap.h" #include "hw/pci/pci_device.h" #include "hw/pci/pcie.h" #include "net/eth.h" @@ -69,7 +71,7 @@ bool igbvf_add_migration_cap(PCIDevice *dev, Error **errp) pci_set_long(dev->config + offset + IGB_MIG_CAP_OFF_BARID, IGB_MIG_BAR_IDX); pci_set_long(dev->config + offset + IGB_MIG_CAP_OFF_FLAGS, - IGB_MIG_CAP_F_STATE); + IGB_MIG_CAP_F_STATE | IGB_MIG_CAP_F_DIRTY); =20 trace_igbvf_mig_cap_add(pcie_sriov_vf_number(dev), offset); return true; @@ -399,6 +401,161 @@ static int igbvf_mig_load(IgbVfState *s, const void *= buf, size_t size) return 0; } =20 +/* + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * Per-VF dirty page tracking + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * + * All VF DMA writes in igb_core.c go through igb_pci_dma_write(), + * which calls igb_core_dirty_track_dma() to mark the target page in a + * per-range bitmap before performing the actual DMA. + * + * The IGBCore::vf_dirty[] bitmaps live in IGBCore so they are easily + * accessible from the core TX and RX paths without reaching back into + * VF state. + */ + +void igb_core_dirty_track_dma(IGBCore *core, int vfn, + dma_addr_t addr, dma_addr_t len) +{ + IGBVfDirtyState *ds =3D &core->vf_dirty[vfn]; + bool matched =3D false; + uint32_t i; + + if (!ds->num_ranges) { + return; + } + + trace_igb_core_dirty_track_dma(vfn, addr, len); + + for (i =3D 0; i < ds->num_ranges; i++) { + IGBVfDirtyRange *r =3D &ds->ranges[i]; + uint64_t r_end =3D r->iova + r->size; + uint64_t dma_end =3D addr + len; + uint64_t start, end, start_page, end_page, page; + + if (addr >=3D r_end || dma_end <=3D r->iova) { + continue; + } + + matched =3D true; + start =3D MAX(addr, r->iova); + end =3D MIN(dma_end, r_end); + + start_page =3D (start - r->iova) / r->page_size; + end_page =3D (end - 1 - r->iova) / r->page_size; + + for (page =3D start_page; page <=3D end_page; page++) { + if (page < r->nbits) { + set_bit(page, r->bitmap); + } + } + } + + if (!matched) { + trace_igb_core_dirty_track_dma_drop(vfn, addr, len); + } +} + +static IGBVfDirtyState *igb_core_vf_dirty_state(IgbVfState *s) +{ + IGBCore *core =3D igbvf_get_core(s); + return &core->vf_dirty[s->vfn]; +} + +#define IGB_MIG_DIRTY_MAX_PAGES ((256 * GiB) / (4 * KiB)) + +static uint32_t igb_core_vf_dirty_enable(IgbVfState *s, uint64_t pgsize, + uint64_t range_iova, uint64_t range_size) +{ + IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); + IGBVfDirtyRange *r; + + if (ds->num_ranges >=3D IGB_MIG_CAPS_MAX_RANGES) { + return IGB_MIG_DIRTY_STATUS_TOO_MANY_RANGES; + } + + if (!range_size) { + return IGB_MIG_DIRTY_STATUS_BAD_RANGE; + } + + if (!pgsize || (range_iova % pgsize) || (range_size % pgsize)) { + return IGB_MIG_DIRTY_STATUS_BAD_PGSIZE; + } + + if (range_size / pgsize > IGB_MIG_DIRTY_MAX_PAGES) { + return IGB_MIG_DIRTY_STATUS_BAD_RANGE; + } + + r =3D &ds->ranges[ds->num_ranges]; + r->iova =3D range_iova; + r->size =3D range_size; + r->page_size =3D pgsize; + r->nbits =3D range_size / pgsize; + r->bitmap =3D bitmap_new(r->nbits); + ds->num_ranges++; + return IGB_MIG_DIRTY_STATUS_OK; +} + +void igb_core_vf_dirty_disable(IgbVfState *s) +{ + IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); + uint32_t i; + + for (i =3D 0; i < ds->num_ranges; i++) { + IGBVfDirtyRange *r =3D &ds->ranges[i]; + + g_free(r->bitmap); + r->bitmap =3D NULL; + r->nbits =3D 0; + } + ds->num_ranges =3D 0; +} + +static bool igb_core_vf_dirty_enabled(IgbVfState *s) +{ + IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); + return !!ds->num_ranges; +} + +static bool igb_core_vf_dirty_query(IgbVfState *s, + void *buf, size_t buf_size, size_t *out_size, + uint64_t range_iova, uint64_t range_size) +{ + IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); + uint32_t i; + + for (i =3D 0; i < ds->num_ranges; i++) { + IGBVfDirtyRange *r =3D &ds->ranges[i]; + uint64_t start_page, range_pages, count; + + if (range_iova < r->iova || + range_iova + range_size > r->iova + r->size) { + continue; + } + + start_page =3D (range_iova - r->iova) / r->page_size; + range_pages =3D (uint64_t)range_size / r->page_size; + count =3D MIN(range_pages, (uint64_t)buf_size * 8); + + memset(buf, 0, buf_size); + + if (start_page < r->nbits) { + uint64_t avail =3D r->nbits - start_page; + uint64_t n =3D MIN(count, avail); + + bitmap_copy_with_src_offset(buf, r->bitmap, start_page, n); + bitmap_clear(r->bitmap, start_page, n); + } + + *out_size =3D bitmap_empty(buf, count) ? 0 : DIV_ROUND_UP(count, 8= ); + return true; + } + + *out_size =3D 0; + return false; +} + /* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D * Migration BAR register read/write handlers * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D */ @@ -410,25 +567,53 @@ static bool igbvf_mig_set_state(IgbVfState *s, uint32= _t new_state) int ret; =20 switch (new_state) { + case IGB_MIG_STATE_PRE_COPY: + if (old !=3D IGB_MIG_STATE_RUNNING) { + return false; + } + /* + * Take an initial snapshot so DATA_SIZE reflects the actual + * state size and DATA_AVAIL is set for the driver. + */ + ret =3D igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_dat= a)); + if (ret < 0) { + ms->mig_error =3D -ret; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + return false; + } + ms->mig_data_size =3D ret; + break; + case IGB_MIG_STATE_STOP: if (old !=3D IGB_MIG_STATE_RUNNING && old !=3D IGB_MIG_STATE_STOP_COPY && + old !=3D IGB_MIG_STATE_PRE_COPY && old !=3D IGB_MIG_STATE_RESUMING && old !=3D IGB_MIG_STATE_ERROR) { return false; } /* Restore DATA_SIZE to max, same as at reset */ ms->mig_data_size =3D igb_core_vf_max_data_size(s); + if (old =3D=3D IGB_MIG_STATE_PRE_COPY || + old =3D=3D IGB_MIG_STATE_STOP_COPY || + old =3D=3D IGB_MIG_STATE_ERROR) { + igb_core_vf_dirty_disable(s); + } break; =20 case IGB_MIG_STATE_RUNNING: - if (old !=3D IGB_MIG_STATE_STOP) { + if (old !=3D IGB_MIG_STATE_STOP && + old !=3D IGB_MIG_STATE_PRE_COPY) { return false; } + if (old =3D=3D IGB_MIG_STATE_PRE_COPY) { + igb_core_vf_dirty_disable(s); + } break; =20 case IGB_MIG_STATE_STOP_COPY: - if (old !=3D IGB_MIG_STATE_STOP) { + if (old !=3D IGB_MIG_STATE_STOP && + old !=3D IGB_MIG_STATE_PRE_COPY) { return false; } ret =3D igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_dat= a)); @@ -466,7 +651,8 @@ static uint32_t igbvf_mig_get_status(IgbVfState *s) if (ms->mig_state =3D=3D IGB_MIG_STATE_ERROR) { status |=3D IGB_MIG_STATUS_ERR(ms->mig_error); } - if (ms->mig_state =3D=3D IGB_MIG_STATE_STOP_COPY && ms->mig_data_size = > 0) { + if ((ms->mig_state =3D=3D IGB_MIG_STATE_STOP_COPY || + ms->mig_state =3D=3D IGB_MIG_STATE_PRE_COPY) && ms->mig_data_size= > 0) { status |=3D IGB_MIG_STATUS_DATA_AVAIL; } =20 @@ -486,6 +672,30 @@ static void igbvf_mig_data_xfer(IgbVfState *s, uint32_= t val) } =20 switch (ms->mig_state) { + case IGB_MIG_STATE_PRE_COPY: + /* + * Re-snapshot device state so the driver can read a fresh + * copy on each pre-copy iteration. + */ + ret =3D igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_dat= a)); + if (ret < 0) { + ms->mig_error =3D -ret; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + return; + } + ms->mig_data_size =3D ret; + r =3D pci_dma_write(pcie_sriov_get_pf(PCI_DEVICE(s)), + ms->mig_data_buf_addr, + ms->mig_data, ms->mig_data_size); + if (r !=3D MEMTX_OK) { + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u state write failed at 0x%" PRIx64 "= \n", + s->vfn, ms->mig_data_buf_addr); + ms->mig_error =3D IGB_MIG_ERR_DMA_FAILED; + ms->mig_state =3D IGB_MIG_STATE_ERROR; + } + break; + case IGB_MIG_STATE_STOP_COPY: /* Save: DMA-write serialized state to driver buffer */ r =3D pci_dma_write(pcie_sriov_get_pf(PCI_DEVICE(s)), @@ -547,7 +757,9 @@ static uint64_t igbvf_mig_read(void *opaque, hwaddr add= r, unsigned size) val =3D igbvf_mig_get_status(s); break; case IGB_MIG_CAPS: - val =3D IGB_MIG_CAP_F_STATE; + val =3D IGB_MIG_CAP_F_STATE | IGB_MIG_CAP_F_DIRTY | + (IGB_MIG_CAPS_MAX_RANGES << IGB_MIG_CAPS_MAX_RANGES_SHIFT) | + (1u << 12); /* 4K page size supported */ break; case IGB_MIG_VERSION: val =3D IGB_MIG_CAP_VERSION; @@ -555,6 +767,19 @@ static uint64_t igbvf_mig_read(void *opaque, hwaddr ad= dr, unsigned size) case IGB_MIG_DATA_SIZE: val =3D ms->mig_data_size; break; + case IGB_MIG_DIRTY_PGSIZE: + val =3D ms->mig_dirty_pgsize ? ms->mig_dirty_pgsize + : IGB_MIG_DIRTY_DEFAULT_PGSIZE; + break; + case IGB_MIG_DIRTY_BUF_ADDR_LO: + val =3D (uint32_t)ms->mig_dirty_buf_addr; + break; + case IGB_MIG_DIRTY_BUF_ADDR_HI: + val =3D (uint32_t)(ms->mig_dirty_buf_addr >> 32); + break; + case IGB_MIG_DIRTY_STATUS: + val =3D ms->mig_dirty_status; + break; default: qemu_log_mask(LOG_GUEST_ERROR, "igbvf: VF%u bad migration BAR read at 0x%" @@ -567,6 +792,106 @@ static uint64_t igbvf_mig_read(void *opaque, hwaddr a= ddr, unsigned size) return val; } =20 +static uint32_t igbvf_mig_dirty_count(const void *bitmap, size_t size) +{ + const unsigned long *p =3D bitmap; + size_t n =3D size / sizeof(unsigned long); + uint32_t count =3D 0; + + for (size_t i =3D 0; i < n; i++) { + count +=3D ctpopl(p[i]); + } + return count; +} + +static void igbvf_mig_dirty_query(IgbVfState *s, uint64_t pgsize) +{ + IgbVfMigState *ms =3D &s->mig; + PCIDevice *dev =3D pcie_sriov_get_pf(PCI_DEVICE(s)); + uint64_t buf_addr =3D ms->mig_dirty_buf_addr; + uint64_t range_iova =3D 0, range_size =3D 0; + uint32_t bmp_bytes; + size_t out_size; + g_autofree void *bitmap =3D NULL; + bool valid; + + ldq_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, iova), + &range_iova, MEMTXATTRS_UNSPECIFIED); + ldq_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, size), + &range_size, MEMTXATTRS_UNSPECIFIED); + + bmp_bytes =3D DIV_ROUND_UP(range_size / pgsize, 8); + bitmap =3D g_malloc0(bmp_bytes); + + valid =3D igb_core_vf_dirty_query(s, bitmap, bmp_bytes, &out_size, + range_iova, range_size); + + if (valid && out_size) { + if (pci_dma_write(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, bitmap= ), + bitmap, out_size)) { + qemu_log_mask(LOG_GUEST_ERROR, + "igbvf: VF%u dirty bitmap write failed at 0x%" PRIx64 "\n", + s->vfn, buf_addr); + valid =3D false; + } + } + + uint32_t dirty_pages =3D valid ? igbvf_mig_dirty_count(bitmap, out_siz= e) : 0; + + stl_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, bitmap_= size), + valid ? out_size : 0, MEMTXATTRS_UNSPECIFIED); + stl_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, dirty_p= age_count), + dirty_pages, MEMTXATTRS_UNSPECIFIED); + stl_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, status), + valid ? IGB_MIG_DIRTY_STATUS_COMPLETE : 0, + MEMTXATTRS_UNSPECIFIED); + + trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages); +} + +static void igbvf_mig_dirty_ctrl(IgbVfState *s, uint32_t val) +{ + IgbVfMigState *ms =3D &s->mig; + uint64_t pgsize =3D ms->mig_dirty_pgsize + ? ms->mig_dirty_pgsize : IGB_MIG_DIRTY_DEFAULT_PGSIZE; + + switch (val) { + case IGB_MIG_DIRTY_CTRL_ENABLE: + ms->mig_dirty_status =3D igb_core_vf_dirty_enable(s, pgsize, + ms->mig_dirty_range_iova, + ms->mig_dirty_range_size); + if (ms->mig_dirty_status) { + break; + } + trace_igbvf_mig_dirty_enable(s->vfn, pgsize, + ms->mig_dirty_range_size / pgsize); + break; + case IGB_MIG_DIRTY_CTRL_DISABLE: + igb_core_vf_dirty_disable(s); + ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_OK; + trace_igbvf_mig_dirty_disable(s->vfn); + break; + case IGB_MIG_DIRTY_CTRL_QUERY: + if (!igb_core_vf_dirty_enabled(s)) { + ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_NOT_ENABLED; + break; + } + if (!ms->mig_dirty_buf_addr) { + ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_NO_BUFFER; + break; + } + igbvf_mig_dirty_query(s, pgsize); + ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_OK; + break; + } +} + static void igbvf_mig_write(void *opaque, hwaddr addr, uint64_t val, unsigned size) { @@ -599,6 +924,31 @@ static void igbvf_mig_write(void *opaque, hwaddr addr,= uint64_t val, ms->mig_data_buf_addr =3D deposit64(ms->mig_data_buf_addr, 32, 32, val); break; + case IGB_MIG_DIRTY_PGSIZE: + ms->mig_dirty_pgsize =3D (uint32_t)val; + break; + case IGB_MIG_DIRTY_CTRL: + igbvf_mig_dirty_ctrl(s, (uint32_t)val); + break; + case IGB_MIG_DIRTY_RANGE_IOVA_LO: + ms->mig_dirty_range_iova =3D + deposit64(ms->mig_dirty_range_iova, 0, 32, val); + break; + case IGB_MIG_DIRTY_RANGE_IOVA_HI: + ms->mig_dirty_range_iova =3D + deposit64(ms->mig_dirty_range_iova, 32, 32, val); + break; + case IGB_MIG_DIRTY_RANGE_SIZE: + ms->mig_dirty_range_size =3D (uint32_t)val; + break; + case IGB_MIG_DIRTY_BUF_ADDR_LO: + ms->mig_dirty_buf_addr =3D + deposit64(ms->mig_dirty_buf_addr, 0, 32, val); + break; + case IGB_MIG_DIRTY_BUF_ADDR_HI: + ms->mig_dirty_buf_addr =3D + deposit64(ms->mig_dirty_buf_addr, 32, 32, val); + break; default: qemu_log_mask(LOG_GUEST_ERROR, "igbvf: VF%u bad migration BAR write at 0x%" @@ -643,5 +993,12 @@ void igbvf_mig_state_reset(IgbVfState *s) ms->mig_data_size =3D igb_core_vf_max_data_size(s); ms->mig_data_buf_addr =3D 0; memset(ms->mig_data, 0, sizeof(ms->mig_data)); + + igb_core_vf_dirty_disable(s); + ms->mig_dirty_pgsize =3D 0; + ms->mig_dirty_range_iova =3D 0; + ms->mig_dirty_range_size =3D 0; + ms->mig_dirty_buf_addr =3D 0; + ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_OK; trace_igbvf_mig_reset(s->vfn); } diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c index e9f9fc3369d8..398306d3bfc6 100644 --- a/hw/net/igbvf.c +++ b/hw/net/igbvf.c @@ -306,6 +306,10 @@ static void igbvf_pci_uninit(PCIDevice *dev) { IgbVfState *s =3D IGBVF(dev); =20 + if (s->mig.migration_cap) { + igb_core_vf_dirty_disable(s); + } + pcie_aer_exit(dev); pcie_cap_exit(dev); msix_unuse_all_vectors(dev); diff --git a/hw/net/trace-events b/hw/net/trace-events index 0b13a99b3f32..95fdea89d49c 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -303,6 +303,13 @@ igbvf_mig_set_state_err(uint16_t vfn, uint32_t old_sta= te, uint32_t new_state) "V igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of= device state" igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes o= f device state" igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset" +igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%= u: dirty tracking enabled pgsize=3D%"PRIu64" nbits=3D%"PRIu64 +igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled" +igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "= VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages" + +# igb_core.c - VF migration diagnostics +igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirt= y DMA addr=3D0x%"PRIx64" len=3D%"PRIu64 +igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d:= dirty DMA dropped addr=3D0x%"PRIx64" len=3D%"PRIu64" no matching range" =20 # spapr_llan.c spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_= bufs) "pool=3D%d count=3D%"PRId32" rxbufs=3D%"PRIu32 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130912; cv=none; d=zohomail.com; s=zohoarc; b=RxeQxuEbfRZP9vZldXxLu1C9kPPJpZURKSXEuuEjxbIr5ufseSj9b+9uM2qRkDEnHFymVr8TY1A78PmkXh465Erusi51hUHkHmjM9atwd9r+Gf50qzg+nOhTlPHQhYnvfpPogky9nIfaxPPrbkF6vwLcTDftLEjQHqGN4/2fUBI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130912; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=/DQi5EVb5djvOpoc59q1ru9zq7bNILZ8NCa1752G+d0=; b=jNzT7+0c4p3OsdZpMwQ7myfQ6Sy+JjzGpnhRV3j/k5UXXSwvbjb+8mXzj+j7cJoQ5vBSOO3ly2A5jv7UM0SMeZ/XfGDZR0l22LOfZTR4uH6j42qn6Q3KAbe2hb75iDsJO45tazHzEj5TnsOp8xAYOKSg88Do7zwhJy+VBrLV0yI= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 17851309122381004.6219121330018; Sun, 26 Jul 2026 22:41:52 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4W-0003fZ-FH; Mon, 27 Jul 2026 01:40:16 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4U-0003ax-S4 for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:14 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4S-0003kS-09 for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:14 -0400 Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-677-3CUxu2ukNIqFYg3hcwJnFA-1; Mon, 27 Jul 2026 01:40:07 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 1A34A1944DD9; Mon, 27 Jul 2026 05:40:06 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 718631956088; Mon, 27 Jul 2026 05:40:03 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130811; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=/DQi5EVb5djvOpoc59q1ru9zq7bNILZ8NCa1752G+d0=; b=ICoGICkX7iqEKDz3hegIFj8Rnu9gHo1MQwlkAb4eYzk92JDyXES5Iurc2uhy6FyLMczPTJ Y/Jzw4CyP//s3W6vIgIJZZn5gVUE/P6UiIVVLXquU/4iMUI8FZUnoF6xnw7Hh7Deouc+oK 4xfqj8BsvW0lqTZL34FpYWGFE4mMtuA= X-MC-Unique: 3CUxu2ukNIqFYg3hcwJnFA-1 X-Mimecast-MFC-AGG-ID: 3CUxu2ukNIqFYg3hcwJnFA_1785130806 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 07/11] igb: Quiesce VFs on STOP and include PF enable state in migration blob Date: Mon, 27 Jul 2026 07:39:31 +0200 Message-ID: <20260727053935.1392269-8-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.129.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H2=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130915021158500 Quiesce VF queues when entering the STOP state by disabling RX and TX at the PF level (VFRE/VFTE). This ensures no DMA activity occurs while the device state is being serialized. The unquiesce path re-enables queues when returning to RUNNING. The VFRE and VFTE registers are shared PF-level registers with one bit per VF controlling RX and TX enable respectively. Writing the full 32-bit register during restore would clobber other VFs' bits, so these bits cannot be included in the regular per-VF register list. Instead, save the per-VF VFRE/VFTE bit state at quiesce time - before the bits are cleared - and append it to the migration blob as two bytes. On the destination, extract these bits and restore them during unquiesce. For QEMU's emulated igb, the device state machine guarantees no DMA in these states. This register establishes the protocol for real hardware implementations where DMA drain has latency. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- hw/net/igb_migration.h | 4 ++ hw/net/igb_migration.c | 110 +++++++++++++++++++++++++++++++++++++++-- hw/net/trace-events | 6 ++- 3 files changed, 115 insertions(+), 5 deletions(-) diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h index a2e7b363cb04..414412f6392e 100644 --- a/hw/net/igb_migration.h +++ b/hw/net/igb_migration.h @@ -75,6 +75,7 @@ /* MIG_STATUS bits */ #define IGB_MIG_STATUS_DATA_AVAIL (1u << 0) #define IGB_MIG_STATUS_ERROR (1u << 1) +#define IGB_MIG_STATUS_QUIESCED (1u << 2) =20 /* MIG_STATUS error codes in bits [15:8], valid when ERROR bit is set */ #define IGB_MIG_STATUS_ERR_SHIFT 8 @@ -143,6 +144,9 @@ typedef struct IgbVfMigState { uint32_t mig_data_size; uint64_t mig_data_buf_addr; =20 + bool mig_saved_vfre; + bool mig_saved_vfte; + uint32_t mig_dirty_pgsize; uint64_t mig_dirty_range_iova; uint32_t mig_dirty_range_size; diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index 9a9ce902bb4a..187077c14c48 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -95,6 +95,7 @@ static IGBCore *igbvf_get_core(IgbVfState *s) * { uint32_t offset; uint32_t value; } regs[num_regs] * uint32_t num_tx_ctx (number of TX queue context blocks) * { raw struct igb_tx data } tx_ctx[num_tx_ctx] + * uint32_t vfre_vfte (VFRE in [7:0], VFTE in [15:8]) */ =20 /* Maximum number of registers in the VF state slice */ @@ -227,6 +228,13 @@ static uint32_t *igb_core_vf_save_ra(IGBCore *core, ui= nt16_t vfn, return p; } =20 +static uint32_t *igb_core_vf_save_vfre_vfte(uint32_t *p, + bool vfre, bool vfte) +{ + *p++ =3D cpu_to_le32((uint32_t)vfre | ((uint32_t)vfte << 8)); + return p; +} + static uint32_t *igb_core_vf_save_tx_ctx(IGBCore *core, int queue, uint32_t *p) { @@ -243,12 +251,14 @@ static size_t igb_core_vf_state_max_size(int num_fixe= d_regs) + num_fixed_regs * 2 * sizeof(uint32_t) /* fixed reg pairs */ + max_ra_regs * 2 * sizeof(uint32_t) /* RA reg pairs */ + sizeof(uint32_t) /* num_tx_ctx */ - + 2 * sizeof(struct igb_tx); /* TX context */ + + 2 * sizeof(struct igb_tx) /* TX context */ + + sizeof(uint32_t); /* VFRE + VFTE */ } =20 static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size) { + IgbVfMigState *ms =3D &s->mig; IGBCore *core =3D igbvf_get_core(s); uint32_t offsets[IGB_VF_MAX_REGS]; int num_regs, total_regs; @@ -296,9 +306,15 @@ static int igb_core_vf_save_state(IgbVfState *s, p =3D igb_core_vf_save_tx_ctx(core, q0, p); p =3D igb_core_vf_save_tx_ctx(core, q1, p); =20 + /* Per-VF VFRE/VFTE enable bits */ + p =3D igb_core_vf_save_vfre_vfte(p, ms->mig_saved_vfre, + ms->mig_saved_vfte); + size =3D (uint8_t *)p - (uint8_t *)buf; =20 - trace_igbvf_mig_save_state(s->vfn, size); + trace_igbvf_mig_save_state(s->vfn, size, ms->mig_saved_vfre, + ms->mig_saved_vfte, + core->mac[VFRE]); return size; } =20 @@ -310,6 +326,16 @@ static int igb_core_vf_max_data_size(IgbVfState *s) return size; } =20 +static const void *igb_core_vf_load_vfre_vfte(const void *p, + bool *out_vfre, bool *out_vf= te) +{ + uint32_t val =3D le32_to_cpu(*(const uint32_t *)p); + + *out_vfre =3D !!(val & 0xff); + *out_vfte =3D !!((val >> 8) & 0xff); + return (const uint8_t *)p + sizeof(uint32_t); +} + static const void *igb_core_vf_load_tx_ctx(IGBCore *core, int queue, const void *data) { @@ -323,6 +349,7 @@ static const void *igb_core_vf_load_tx_ctx(IGBCore *cor= e, int queue, static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size) { + IgbVfMigState *ms =3D &s->mig; IGBCore *core =3D igbvf_get_core(s); const uint32_t *p =3D buf; uint32_t magic, version, saved_vfn, num_regs, num_tx; @@ -378,7 +405,13 @@ static int igb_core_vf_load_state(IgbVfState *s, * destination-specific IRTE references after migration. */ =20 - trace_igbvf_mig_load_state(s->vfn, (uint32_t)size); + /* Per-VF VFRE/VFTE enable bits */ + p =3D igb_core_vf_load_vfre_vfte(p, &ms->mig_saved_vfre, + &ms->mig_saved_vfte); + + trace_igbvf_mig_load_state(s->vfn, (uint32_t)size, + ms->mig_saved_vfre, + ms->mig_saved_vfte); return 0; } =20 @@ -387,6 +420,12 @@ static int igbvf_mig_load(IgbVfState *s, const void *b= uf, size_t size) IGBCore *core =3D igbvf_get_core(s); int ret; =20 + /* + * Pre-load: Clear the VFLRE bit before restoring state so the PF + * watchdog does not overwrite what we are about to load. + */ + core->mac[VFLRE] &=3D ~BIT(s->vfn); + ret =3D igb_core_vf_load_state(s, buf, size); if (ret < 0) { return ret; @@ -556,6 +595,45 @@ static bool igb_core_vf_dirty_query(IgbVfState *s, return false; } =20 +/* Quiesce a VF by disabling its RX and TX at the PF level. */ +static void igb_core_vf_quiesce(IgbVfState *s) +{ + IgbVfMigState *ms =3D &s->mig; + IGBCore *core =3D igbvf_get_core(s); + + ms->mig_saved_vfre =3D !!(core->mac[VFRE] & BIT(s->vfn)); + ms->mig_saved_vfte =3D !!(core->mac[VFTE] & BIT(s->vfn)); + + core->mac[VFRE] &=3D ~BIT(s->vfn); + core->mac[VFTE] &=3D ~BIT(s->vfn); + trace_igbvf_mig_quiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]); +} + +static void igb_core_vf_unquiesce(IgbVfState *s) +{ + IgbVfMigState *ms =3D &s->mig; + IGBCore *core =3D igbvf_get_core(s); + bool re =3D ms->mig_saved_vfre; + bool te =3D ms->mig_saved_vfte; + + if (re) { + core->mac[VFRE] |=3D BIT(s->vfn); + } else { + core->mac[VFRE] &=3D ~BIT(s->vfn); + } + if (te) { + core->mac[VFTE] |=3D BIT(s->vfn); + } else { + core->mac[VFTE] &=3D ~BIT(s->vfn); + } + + trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]); + + if (re) { + igb_start_recv(core); + } +} + /* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D * Migration BAR register read/write handlers * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D */ @@ -599,6 +677,10 @@ static bool igbvf_mig_set_state(IgbVfState *s, uint32_= t new_state) old =3D=3D IGB_MIG_STATE_ERROR) { igb_core_vf_dirty_disable(s); } + if (old =3D=3D IGB_MIG_STATE_RUNNING || + old =3D=3D IGB_MIG_STATE_PRE_COPY) { + igb_core_vf_quiesce(s); + } break; =20 case IGB_MIG_STATE_RUNNING: @@ -609,6 +691,9 @@ static bool igbvf_mig_set_state(IgbVfState *s, uint32_t= new_state) if (old =3D=3D IGB_MIG_STATE_PRE_COPY) { igb_core_vf_dirty_disable(s); } + if (old =3D=3D IGB_MIG_STATE_STOP) { + igb_core_vf_unquiesce(s); + } break; =20 case IGB_MIG_STATE_STOP_COPY: @@ -616,6 +701,9 @@ static bool igbvf_mig_set_state(IgbVfState *s, uint32_t= new_state) old !=3D IGB_MIG_STATE_PRE_COPY) { return false; } + if (old =3D=3D IGB_MIG_STATE_PRE_COPY) { + igb_core_vf_quiesce(s); + } ret =3D igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_dat= a)); if (ret < 0) { ms->mig_error =3D -ret; @@ -656,6 +744,20 @@ static uint32_t igbvf_mig_get_status(IgbVfState *s) status |=3D IGB_MIG_STATUS_DATA_AVAIL; } =20 + /* + * QUIESCED tells the driver it is safe to read device state. + * STOP and STOP_COPY must guarantee that no VF DMA is in flight; + * igb_core_vf_quiesce() enforces this by disabling RX/TX for the + * VF. + * + * Under QEMU, DMA completes synchronously within the MMIO handler + * so there is nothing to drain, but real hardware would need this. + */ + if (ms->mig_state =3D=3D IGB_MIG_STATE_STOP || + ms->mig_state =3D=3D IGB_MIG_STATE_STOP_COPY) { + status |=3D IGB_MIG_STATUS_QUIESCED; + } + return status; } =20 @@ -1000,5 +1102,7 @@ void igbvf_mig_state_reset(IgbVfState *s) ms->mig_dirty_range_size =3D 0; ms->mig_dirty_buf_addr =3D 0; ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_OK; + ms->mig_saved_vfre =3D true; + ms->mig_saved_vfte =3D true; trace_igbvf_mig_reset(s->vfn); } diff --git a/hw/net/trace-events b/hw/net/trace-events index 95fdea89d49c..f4940e3e176d 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -300,8 +300,8 @@ igbvf_mig_bar_read(uint16_t vfn, uint64_t addr, uint64_= t val) "VF%u: BAR read ad igbvf_mig_bar_write(uint16_t vfn, uint64_t addr, uint64_t val) "VF%u: BAR = write addr=3D0x%"PRIx64" val=3D0x%"PRIx64 igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) = "VF%u: state %u -> %u" igbvf_mig_set_state_err(uint16_t vfn, uint32_t old_state, uint32_t new_sta= te) "VF%u: invalid transition %u -> %u" -igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of= device state" -igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes o= f device state" +igbvf_mig_save_state(uint16_t vfn, int size, bool vfre, bool vfte, uint32_= t reg_vfre) "VF%u: saved %d bytes vfre=3D%d vfte=3D%d VFRE=3D0x%x" +igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool vfre, bool vfte) "V= F%u: loaded %u bytes vfre=3D%d vfte=3D%d" igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset" igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%= u: dirty tracking enabled pgsize=3D%"PRIu64" nbits=3D%"PRIu64 igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled" @@ -310,6 +310,8 @@ igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint= 32_t dirty_pages) "VF%u: # igb_core.c - VF migration diagnostics igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirt= y DMA addr=3D0x%"PRIx64" len=3D%"PRIu64 igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d:= dirty DMA dropped addr=3D0x%"PRIx64" len=3D%"PRIu64" no matching range" +igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quies= ce VFRE=3D0x%x VFTE=3D0x%x" +igbvf_mig_unquiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: unq= uiesce VFRE=3D0x%x VFTE=3D0x%x" =20 # spapr_llan.c spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_= bufs) "pool=3D%d count=3D%"PRId32" rxbufs=3D%"PRIu32 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130844; cv=none; d=zohomail.com; s=zohoarc; b=denv8xHvnqBslWz6xsyreVf26pFYXQScaHcCJdDihrFG+pSk+4gt6WZTlo12I3t+Du1nMPPiwLJ72Oa8+K1z6PY21jKrhynptSh9eERJ+GDYBmJT/FJNMwTV3f5BIuQeNPLCnfnO2r/5I8RD6jKGUumQ9HDyc3aupBBYJYvnVYg= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130844; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=IBUiIomosN7SjHQH0P/1cm4lrkCiBywj3i72g0UpvMM=; b=D3HAJRLr5rB0xO6qghyd5yjXFdLkPyIDo3JTVq3wKujtZpjpdjNY4gp5qOEblx+9PmX+jNTyNiXppJyqGF9vCDOSE/DdNr6YGCLBk9rd4hz41vG4XhvcC+KvxE/H0x3NUYblMl2Qj5IlwXzO8Hy2XTcJaCrDR+2kqoxSBULkBAU= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130844467935.4315448783025; Sun, 26 Jul 2026 22:40:44 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4e-0003qb-34; Mon, 27 Jul 2026 01:40:24 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4Y-0003lt-OJ for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:18 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4X-0003lM-4w for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:18 -0400 Received: from mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-691-dMegtwUyMIqv7NDgodClOA-1; Mon, 27 Jul 2026 01:40:10 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 305531955F29; Mon, 27 Jul 2026 05:40:09 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 89B1E1955F74; Mon, 27 Jul 2026 05:40:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130816; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=IBUiIomosN7SjHQH0P/1cm4lrkCiBywj3i72g0UpvMM=; b=Fa3DvNGtT7nPyRlX1VEM8iOpng0ZfP7Q1PPSlwOJH5EacMNaVWD7GKttNG5Wj28aNfykw2 MzWL5OJ4OLBBvnA5LHSx5Y7HuKqpbXoe5WYsp8octmlN7dQjfc4Ggan3P5WlbXsRqRm92N MttYNuXoohNdCVdtGN1uBKCGkqZ19Vo= X-MC-Unique: dMegtwUyMIqv7NDgodClOA-1 X-Mimecast-MFC-AGG-ID: dMegtwUyMIqv7NDgodClOA_1785130809 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 08/11] igb: Fix post-migration RX ring deadlock Date: Mon, 27 Jul 2026 07:39:32 +0200 Message-ID: <20260727053935.1392269-9-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.129.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H2=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130846528158500 After migration, the VF unquiesce path calls igb_start_recv() to kick RX processing. This does not trigger an interrupt, so NAPI stays idle and the RX ring is never replenished. Replace igb_start_recv() with igb_core_vf_rearm_irqs(): re-apply the VF's PVT shadow register into the PF aggregates, then raise interrupt causes via igb_set_eics(). This follows the normal EICR -> EIMS -> MSI-X routing. NAPI wakes, replenishes the ring via igbvf_alloc_rx_buffers(), and writes RDT - breaking the deadlock. The PVT re-application is needed because the L1 PF driver's normal interrupt handling may have cleared EIMS between state load and unquiesce. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- hw/net/igb_core.h | 1 + hw/net/igb_core.c | 12 ++++++++++++ hw/net/igb_migration.c | 2 +- 3 files changed, 14 insertions(+), 1 deletion(-) diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index 50cee2c7683c..3db520048732 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -147,5 +147,6 @@ igb_start_recv(IGBCore *core); =20 IGBCore *igb_pf_get_core(void *pf); void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn); +void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn); void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn); #endif diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c index 03c349575c0a..1621ee6602d5 100644 --- a/hw/net/igb_core.c +++ b/hw/net/igb_core.c @@ -4589,6 +4589,18 @@ void igb_core_vf_propagate_irqs(IGBCore *core, uint1= 6_t vfn) core->mac[EICR] &=3D ~(0x7 << shift); } =20 +/* + * Re-apply VF interrupt enables to PF aggregates and raise interrupt + * causes to wake NAPI so it replenishes the ring after migration. + */ +void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn) +{ + uint32_t shift =3D 22 - vfn * IGBVF_MSIX_VEC_NUM; + + igb_core_vf_propagate_irqs(core, vfn); + igb_set_eics(core, EICS, 0x7 << shift); +} + /* * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver * may have overwritten the shared IVAR0 entries with its own diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index 187077c14c48..a0b4a52c865f 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -630,7 +630,7 @@ static void igb_core_vf_unquiesce(IgbVfState *s) trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]); =20 if (re) { - igb_start_recv(core); + igb_core_vf_rearm_irqs(core, s->vfn); } } =20 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130851; cv=none; d=zohomail.com; s=zohoarc; b=LXgLyA07ijF2Q0NfdvPcDxsJly8qaFyMQ/FAZAE8jLUBqagOVuRERaPpQbT+lLpNN5muFh1/EJXE5Vgfk61ugrr1YBgONCACb0ojp7mr1esoIimJaarGckDbok4p2gq4nthx+CAJDnU/OEsgt87gyweKJ5tvJf+L0Q8+ZTlIsqo= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130851; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=sEWKxZjddL/Yv5HQH67qrxFGud3IuDmuA7QtqmGJdwc=; b=WkCFL/nRahkX82tIXYJdMFX/9cUKkGmUfe+gSsPI2iXZq68JPuw0Br+sOnEtrYIK5EBbkp1Usoqh+Hx8A5EusCoR94J/oXm/gv9sKIbanslsbH9cp/cBBhy5e9AUdLOPq9wIW/31NyoQ0Ho7w4DXDb1OecYakE465B/M2YlpECY= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130851193898.0355714096195; Sun, 26 Jul 2026 22:40:51 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4c-0003mw-U2; Mon, 27 Jul 2026 01:40:22 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4Y-0003lI-EV for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:18 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4W-0003l8-IX for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:18 -0400 Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-522-m661FR3HNtCMy01qbYKGUw-1; Mon, 27 Jul 2026 01:40:13 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 4E0E81955F04; Mon, 27 Jul 2026 05:40:12 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 9FFB21956088; Mon, 27 Jul 2026 05:40:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130815; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=sEWKxZjddL/Yv5HQH67qrxFGud3IuDmuA7QtqmGJdwc=; b=Jv6sqv689iaKMYKr+lYi2QWIZ72aGVTvx8DEFsI07Le9ny/6JxtXMYJml5lG/rlRRF6myG fmotmC7+q1CNI0NDqQ0CGwLjz4ZUQbXGNJ0QSalOGWGGdE9BLZ2zNCCe4C/dYZUtc/9uRq DLECecR/zldUZdfgmkpJM5VuKsxbiVU= X-MC-Unique: m661FR3HNtCMy01qbYKGUw-1 X-Mimecast-MFC-AGG-ID: m661FR3HNtCMy01qbYKGUw_1785130812 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 09/11] igb: Send RARP after VF migration to update bridge FDB Date: Mon, 27 Jul 2026 07:39:33 +0200 Message-ID: <20260727053935.1392269-10-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.133.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H3=0.001, RCVD_IN_MSPIKE_WL=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130854593158500 After VF migration the L0 bridge FDB (Forwarding Database) still maps the VF's MAC to the old tap port, causing a ~30 s connectivity stall while peer ARP caches expire. Add igb_core_vf_get_mac() to look up a VF's MAC address from the PF's Receive Address registers, and use it to send a RARP broadcast from the PF's network backend during VF unquiesce so the bridge relearns the correct port immediately. Assisted-by: Claude Signed-off-by: C=C3=A9dric Le Goater --- hw/net/igb_core.h | 1 + hw/net/igb_core.c | 28 ++++++++++++++++++++++++++++ hw/net/igb_migration.c | 33 +++++++++++++++++++++++++++++++++ hw/net/trace-events | 2 ++ 4 files changed, 64 insertions(+) diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index 3db520048732..07b4270708c4 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -149,4 +149,5 @@ IGBCore *igb_pf_get_core(void *pf); void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn); void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn); void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn); +bool igb_core_vf_get_mac(IGBCore *core, uint16_t vfn, uint8_t *mac); #endif diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c index 1621ee6602d5..281f3a474392 100644 --- a/hw/net/igb_core.c +++ b/hw/net/igb_core.c @@ -4634,3 +4634,31 @@ void igb_core_vf_propagate_ivar(IGBCore *core, uint1= 6_t vfn) ((uint32_t)ent << (8 * (n % 4))); } } + +bool igb_core_vf_get_mac(IGBCore *core, uint16_t vfn, uint8_t *mac) +{ + uint32_t vf_pool_bit =3D E1000_RAH_POOL_1 << vfn; + static const struct { + uint32_t base; + int count; + } ra_banks[] =3D { + { RA, 16 }, + { RA2, 8 }, + }; + int i, j; + + for (i =3D 0; i < ARRAY_SIZE(ra_banks); i++) { + for (j =3D 0; j < ra_banks[i].count; j++) { + uint32_t ral_off =3D ra_banks[i].base + j * 2; + uint32_t rah_off =3D ra_banks[i].base + j * 2 + 1; + uint32_t rah_val =3D core->mac[rah_off]; + + if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) { + stl_le_p(mac, core->mac[ral_off]); + stw_le_p(mac + 4, rah_val & 0xffff); + return true; + } + } + } + return false; +} diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index a0b4a52c865f..fce661f68e3d 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -609,6 +609,36 @@ static void igb_core_vf_quiesce(IgbVfState *s) trace_igbvf_mig_quiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]); } =20 +/* + * Send a RARP broadcast so the network bridge relearns which port + * carries this VF's MAC after migration. + */ +static void igb_core_vf_send_rarp(IGBCore *core, uint16_t vfn) +{ + uint8_t mac[ETH_ALEN]; + uint8_t buf[60]; + + if (!igb_core_vf_get_mac(core, vfn, mac)) { + trace_igbvf_mig_no_mac(vfn); + return; + } + + trace_igbvf_mig_send_rarp(vfn, mac[0], mac[1], mac[2], + mac[3], mac[4], mac[5]); + + memset(buf, 0xff, ETH_ALEN); + memcpy(buf + 6, mac, ETH_ALEN); + stw_be_p(buf + 12, 0x8035); /* ETH_P_RARP */ + stw_be_p(buf + 14, 1); /* hw addr space: ethernet */ + stw_be_p(buf + 16, ETH_P_IP); /* protocol addr space */ + buf[18] =3D 6; buf[19] =3D 4; /* hw/proto addr lengths */ + stw_be_p(buf + 20, 3); /* opcode: RARP request */ + memcpy(buf + 22, mac, ETH_ALEN); + memset(buf + 28, 0, 32); + + qemu_send_packet_raw(qemu_get_queue(core->owner_nic), buf, sizeof(buf)= ); +} + static void igb_core_vf_unquiesce(IgbVfState *s) { IgbVfMigState *ms =3D &s->mig; @@ -632,6 +662,9 @@ static void igb_core_vf_unquiesce(IgbVfState *s) if (re) { igb_core_vf_rearm_irqs(core, s->vfn); } + + /* TODO : RARP should be sent only if resumed */ + igb_core_vf_send_rarp(core, s->vfn); } =20 /* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D diff --git a/hw/net/trace-events b/hw/net/trace-events index f4940e3e176d..1e38d9ab3697 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -312,6 +312,8 @@ igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64= _t len) "VF%d: dirty DMA igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d:= dirty DMA dropped addr=3D0x%"PRIx64" len=3D%"PRIu64" no matching range" igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quies= ce VFRE=3D0x%x VFTE=3D0x%x" igbvf_mig_unquiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: unq= uiesce VFRE=3D0x%x VFTE=3D0x%x" +igbvf_mig_no_mac(uint16_t vfn) "VF%u: no MAC address found, skipping RARP" +igbvf_mig_send_rarp(uint16_t vfn, uint8_t b0, uint8_t b1, uint8_t b2, uint= 8_t b3, uint8_t b4, uint8_t b5) "VF%u: RARP %02x:%02x:%02x:%02x:%02x:%02x" =20 # spapr_llan.c spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_= bufs) "pool=3D%d count=3D%"PRId32" rxbufs=3D%"PRIu32 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130846; cv=none; d=zohomail.com; s=zohoarc; b=e6JIeY6oSlaWo6J3pZjWfwClL7gGwuThYRl8UQ85LtjfE1skCXHOcgsRvmLLts/RUnA4BaHXTeSsBfhDqrxyHkTtmjmyZVpU6kF7Wx8smZ2cPgAW4P185OnTcTPsjaR1+kTMIs3AzWuiYSHHZqq0PTjNQKDmnR3c1TJV1YGuDN4= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130846; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=PNpE2nNq2DkmFOVkbIPR6qFBZzODAefD/FRCjOwoQx8=; b=alvnzDV7UJSumRMSN4C++zgjWDAD0mN09E26+iZ5qM0wdg1z+cZxDnuLVTnd/xMqCl71ajY+Rg82N6b1cOo7ucULUa90Le/L82Jpd9jdn6jmV8jDWb47iS2WAG9TdYdrrL2XMHyn8VQMMj2CTc9lQsCZReMe75Q71fpiOm7uB3Q= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130846283823.1881428074846; Sun, 26 Jul 2026 22:40:46 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4u-00048f-8M; Mon, 27 Jul 2026 01:40:40 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4f-0003rA-24 for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:26 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4c-0003lf-QB for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:24 -0400 Received: from mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (ec2-35-165-154-97.us-west-2.compute.amazonaws.com [35.165.154.97]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-207-T9iYyycTNHKL2ocymMtqLw-1; Mon, 27 Jul 2026 01:40:17 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-08.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id A7580180134A; Mon, 27 Jul 2026 05:40:15 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id BE0481956088; Mon, 27 Jul 2026 05:40:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130820; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=PNpE2nNq2DkmFOVkbIPR6qFBZzODAefD/FRCjOwoQx8=; b=aX7oUkBllZrroOAuwDDledtOXN2lpWCz4gEvRgKOB0bVBBOLkhX7Usv2s0/oSV3TNQa00F McIyfUpq8MRaMHpykg4kTeAgkf1aycyjnJMyHxfsMeOLT9ei7u/8kyHx49meztmy6ipeqS 2W7RCj0vzmPGWS89d8MzQK/vbjbq+hw= X-MC-Unique: T9iYyycTNHKL2ocymMtqLw-1 X-Mimecast-MFC-AGG-ID: T9iYyycTNHKL2ocymMtqLw_1785130815 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 10/11] docs: Add igb VF migration testing setup guide Date: Mon, 27 Jul 2026 07:39:34 +0200 Message-ID: <20260727053935.1392269-11-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.129.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -36 X-Spam_score: -3.7 X-Spam_bar: --- X-Spam_report: (-3.7 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H2=0.001, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130848571158500 Document the NetworkManager configuration and network topology needed to test igb VF live migration in a nested virtualization environment. Signed-off-by: C=C3=A9dric Le Goater --- docs/system/devices/igb-migration.rst | 155 ++++++++++++++++++++++++++ 1 file changed, 155 insertions(+) diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/ig= b-migration.rst index 4a3a93f01dfc..0e34e4b9fc68 100644 --- a/docs/system/devices/igb-migration.rst +++ b/docs/system/devices/igb-migration.rst @@ -134,3 +134,158 @@ The driver fills the request fields, issues ``DIRTY_C= TRL=3DQUERY``, and polls ``status`` for completion. The device reads the request, writes the dirty bitmap and completion fields via DMA, then sets ``status =3D 1``. + +Testing setup +~~~~~~~~~~~~~ + +The target scenario is nested virtualization:: + + L0 QEMU + igb PF with x-vf-migration=3Don + =E2=94=94=E2=94=80=E2=94=80 VFs with migration BAR + vendor cap + + L1 kernel + igb-vfio-pci variant driver + translates VFIO migration v2 ioctls =E2=86=92 BAR2 MMIO + + L1 QEMU (stock, unmodified) + vfio-pci device model, standard migration fd + + L2 guest + standard igbvf driver, unaware of migration + +NetworkManager configuration +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +In a nested setup, the L1 VMs (source and destination) have emulated +igb PFs connected to the L0 bridge. By default, NetworkManager +acquires DHCP leases on those PF interfaces and on any igb VFs created +later. This causes the VF MAC address to be learned on the L0 bridge, +which can misdirect iperf3 traffic after migration. + +To prevent this, configure NetworkManager on **both L1 VMs** and on the +**L2 guest disk image**. + +L1 VMs (source and destination) +............................... + +1. Prevent NetworkManager from managing igbvf interfaces: + +.. code-block:: bash + + cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <" ipv4.method disabled ipv6= .method disabled + +3. Reload NetworkManager: + +.. code-block:: bash + + nmcli general reload + +L2 guest disk image +................... + +Use ``virt-customize`` to add the igbvf unmanaged config to the guest +image (offline, before any test run): + +.. code-block:: bash + + virt-customize -a /srv/migration/rhel10.qcow2 \ + --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile] + unmanaged-devices=3Ddriver:igbvf' + +Network diagram +^^^^^^^^^^^^^^^ + +The diagram below shows the nested setup where the source and +destination hosts are themselves VMs (L1) running on a physical +host (L0) that emulates the igb NIC:: + + =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=90 + =E2=94=82 L0: physical host = =E2=94=82 + =E2=94=82 = =E2=94=82 + =E2=94=82 virbr0 192.168.199.1/24 = =E2=94=82 + =E2=94=82 =E2=94=9C=E2=94=80=E2=94=80 NFS server: /srv/migration = =E2=94=82 + =E2=94=82 =E2=94=94=E2=94=80=E2=94=80 iperf3 client: iperf3 -c 192.168.= 199.200 -t 60 -i 1 =E2=94=82 + =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=82 L0 virbr0 bridge (192.168.199.0/24) = =E2=94=82 + =E2=94=82 =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=B4=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=90 =E2=94=8C= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=B4=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=90 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 virtio =E2=94=82 =E2=94= =82 emulated=E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82c0:ff:ee:=E2=94=82 =E2=94= =82 igb PF =E2=94=82 L0 QEMU (vm6) =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 :00:06 =E2=94=82 =E2=94= =82 + igbvf =E2=94=82 tracks DMA dirty pages =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=98 =E2=94=94= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=98 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=90 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 L1: vm6 (source) =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 enp1s0: 192.168.199.6 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 (management) =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 enp8s0 (igb PF, = no IP) =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 igb VF0 =E2=94= =80=E2=94=80=E2=96=BA igb-vfio-pci (VFIO) =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = dirty_sync =E2=86=92 L0 igbvf =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 virbr0 =E2=94=82 V= FIO passthrough =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 192.168.200.1/24 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 =E2= =94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=90 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 L2: rhel10 gue= st =E2=94=82 =E2=94=82 =E2=94=82 =E2=94= =82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 =E2=94=82 =E2=94= =82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 virtio NIC igb VF= (enp7s0) =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 192.168.200.130/24 192.16= 8.199.200/24 =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 (SSH login) (iperf= 3 data path) =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 =E2= =94=82 =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 iperf3 -s -D =E2= =94=82 (listens on 0.0.0.0) =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=98 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 virsh migrate --live =E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=96=BA vm7 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=98 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=82 iperf3 traffic =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=98 = =E2=94=82 + =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80 = =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=90 =E2=94=82 =E2=94=8C=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=90 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 virtio =E2=94=82 =E2=94=82 = =E2=94=82emulated =E2=94=82 L0 QEMU (vm7) =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82c0:ff:ee=E2=94=82 =E2=94=82 = =E2=94=82 igb PF =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 :00:07 =E2=94=82 =E2=94=82 = =E2=94=82 + igbvf =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2=94=98 =E2=94=82 =E2=94=94=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=AC=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=98 =E2=94=82 + =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=90 =E2=94=82 + =E2=94=82 =E2=94=82 L1: vm7 (destination) =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 enp1s0: 192.168.199.7 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 (management) enp8s0 (igb PF, no IP) = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 igb VF0 =E2=94=80=E2=94=80= =E2=96=BA igb-vfio-pci (VFIO) =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 virbr0 =E2=94=82 VFIO passth= rough =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 192.168.200.1/24 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=90 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 L2: rhel10 (after migratio= n) =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 =E2= =94=82 =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 virtio NIC igb VF (enp7s0) = =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 192.168.200.130/24 192.168.199.200/= 24 =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 =E2=94=82 = =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=82 iperf3 -s -D =E2=94=82 (co= nnection survives) =E2=94=82 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=BC=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=98 =E2=94=82 =E2=94=82 + =E2=94=82 =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=98 =E2=94=82 + =E2=94=82 =E2=94=82 = =E2=94=82 + =E2=94=82 iperf3 traffic resumes =E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=98 =E2=94=82 + =E2=94=82 (same IP, same MAC, same L2 segment =E2=86=92 transparent= to client) =E2=94=82 + =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =80=E2=94=80=E2=94=80=E2=94=98 --=20 2.55.0 From nobody Mon Sep 28 02:06:11 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=quarantine dis=none) header.from=redhat.com ARC-Seal: i=1; a=rsa-sha256; t=1785130863; cv=none; d=zohomail.com; s=zohoarc; b=VzYNs1Ht0PKjx7YxqVL4xAojs0kxCjEPwH40NMP7kM2Z/MAIguG+7VutpZMB5S3GIFIYYTDVteAdd1JObSfeuf4MI87BkIb5zaEL3C9dyy0S+EXRLHjnsoxI7K86LXQsoM4mACGriKin1zdKbufWWnnnRtF8Y+GhpI/oTXBfNPY= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1785130863; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=PHRXSzNCTNiD560sOJzof2nGp13HSfDabGJCc6oJh3o=; b=NrKyxjxL4erGVjX4s+okwGy6HyPTFkmGuUVy0MVBLelCEerKLCh5MTp4JKjijEGnzQNIEUn/H/DXNjvdJA/dXnOaoHTOMv+5XErRaHnG5TDrxNKJMWzf/7iBrH46yRrvWMaCEPNn+9pflgAuTR/i0tHYgngpSzTakeLZsqhhiqY= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=quarantine dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1785130863765864.2123337556968; Sun, 26 Jul 2026 22:41:03 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1woE4u-0004F0-Hb; Mon, 27 Jul 2026 01:40:40 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4g-0003uM-Ul for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:28 -0400 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1woE4e-0003lo-RJ for qemu-devel@nongnu.org; Mon, 27 Jul 2026 01:40:26 -0400 Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-695-DFZmoylRNZOUohgeFLEL4A-1; Mon, 27 Jul 2026 01:40:20 -0400 Received: from mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.12]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id E1A551944D2A; Mon, 27 Jul 2026 05:40:18 +0000 (UTC) Received: from corto.redhat.com (unknown [10.44.32.30]) by mx-prod-int-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 240991956088; Mon, 27 Jul 2026 05:40:15 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785130824; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=PHRXSzNCTNiD560sOJzof2nGp13HSfDabGJCc6oJh3o=; b=KxtjYB+iSI/PTJ4V9nVT9HIiaiI/aqHvOVJsBWJH0Etb1HnNCf3wreEwJtN8EB6Gtfxyq6 MfjS5Twf28qgYraNh2nKGFtLLqZiiTBejMagOW1Ii7Rzy/wObT0pBXbANfPcNSaI883NfC ItiMpq3pOK5tdY9Gv9IVff19R8U1vIs= X-MC-Unique: DFZmoylRNZOUohgeFLEL4A-1 X-Mimecast-MFC-AGG-ID: DFZmoylRNZOUohgeFLEL4A_1785130819 From: =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= To: qemu-devel@nongnu.org Cc: Akihiko Odaki , Sriram Yagnaraman , Jason Wang , Alex Williamson , "Michael S . Tsirkin" , Peter Xu , Avihai Horon , =?UTF-8?q?C=C3=A9dric=20Le=20Goater?= Subject: [RFC PATCH 11/11] igb: Add migration statistics registers to VF migration BAR Date: Mon, 27 Jul 2026 07:39:35 +0200 Message-ID: <20260727053935.1392269-12-clg@redhat.com> In-Reply-To: <20260727053935.1392269-1-clg@redhat.com> References: <20260727053935.1392269-1-clg@redhat.com> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.0 on 10.30.177.12 Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=170.10.129.124; envelope-from=clg@redhat.com; helo=us-smtp-delivery-124.mimecast.com X-Spam_score_int: -2 X-Spam_score: -0.3 X-Spam_bar: / X-Spam_report: (-0.3 / 5.0 requ) BAYES_00=-1.9, DKIMWL_WL_HIGH=-1.58, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, RCVD_IN_DNSWL_NONE=-0.0001, RCVD_IN_MSPIKE_H2=0.001, RCVD_IN_SBL_CSS=3.335, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @redhat.com) X-ZM-MESSAGEID: 1785130864796158500 Add read-only statistics registers to the migration BAR at 0x100-0x11C for monitoring dirty page tracking and DMA activity per VF. The dma_writes counter is also reported in the dirty query buffer so the driver gets it alongside the bitmap without an extra MMIO read. Counters are reset on first DIRTY_CTRL=3DENABLE or device reset, so the driver can read final values after DIRTY_CTRL=3DDISABLE. Signed-off-by: C=C3=A9dric Le Goater --- docs/system/devices/igb-migration.rst | 24 ++++++++++- hw/net/igb_core.h | 1 + hw/net/igb_migration.h | 22 +++++++++- hw/net/igb_migration.c | 59 +++++++++++++++++++++++++-- hw/net/trace-events | 2 +- 5 files changed, 102 insertions(+), 6 deletions(-) diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/ig= b-migration.rst index 0e34e4b9fc68..3f861eed59cd 100644 --- a/docs/system/devices/igb-migration.rst +++ b/docs/system/devices/igb-migration.rst @@ -127,7 +127,9 @@ device-written completion fields:: 0x40 status device 0 =3D pending, 1 =3D complete 0x44 bitmap_size device Bytes written to bitmap 0x48 dirty_page_count device Number of set bits in bitmap - 0x4C reserved[12] - Pad to 64-byte cache line + 0x4C reserved - Padding + 0x50 dma_writes device Total DMA writes since enable + 0x58 reserved[10] - Pad to 64-byte cache line 0x80 bitmap[] device Dirty page bitmap =20 The driver fills the request fields, issues ``DIRTY_CTRL=3DQUERY``, and @@ -135,6 +137,26 @@ polls ``status`` for completion. The device reads the = request, writes the dirty bitmap and completion fields via DMA, then sets ``status =3D 1``. =20 +Migration statistics +~~~~~~~~~~~~~~~~~~~~ + +The migration BAR includes read-only statistics counters in a dedicated +aperture at offset ``0x100``. These counters provide a live +per-migration-cycle view of activity without requiring a +``DIRTY_CTRL=3DQUERY``:: + + Offset Name Description + 0x100 MIG_STAT_DMA_WRITES Total DMA write operations tracked + 0x104 MIG_STAT_DMA_BYTES_LO Total DMA bytes written (low 32 bits) + 0x108 MIG_STAT_DMA_BYTES_HI Total DMA bytes written (high 32 bits) + 0x10C MIG_STAT_DIRTY_PAGES_SET Dirty pages marked since enable + 0x110 MIG_STAT_DIRTY_PAGES_CLR Dirty pages cleared by queries + 0x114 MIG_STAT_DIRTY_PAGE_COUNT Current dirty pages (set minus cleare= d) + 0x118 MIG_STAT_DIRTY_QUERY_CNT Number of QUERY operations + +All counters are reset on first ``DIRTY_CTRL=3DENABLE`` or device reset, +so the driver can read final values after ``DIRTY_CTRL=3DDISABLE``. + Testing setup ~~~~~~~~~~~~~ =20 diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h index 07b4270708c4..91796aec8d7a 100644 --- a/hw/net/igb_core.h +++ b/hw/net/igb_core.h @@ -84,6 +84,7 @@ struct IGBCore { } tx[IGB_NUM_QUEUES]; =20 IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS]; + IgbVfMigStats vf_mig_stats[IGB_MAX_VF_FUNCTIONS]; =20 struct NetRxPkt *rx_pkt; =20 diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h index 414412f6392e..6d2e3442781f 100644 --- a/hw/net/igb_migration.h +++ b/hw/net/igb_migration.h @@ -64,6 +64,15 @@ #define IGB_MIG_DIRTY_BUF_ADDR_HI 0x038 #define IGB_MIG_DIRTY_STATUS 0x03C =20 +/* Migration statistics registers (read-only) */ +#define IGB_MIG_STAT_DMA_WRITES 0x100 +#define IGB_MIG_STAT_DMA_BYTES_LO 0x104 +#define IGB_MIG_STAT_DMA_BYTES_HI 0x108 +#define IGB_MIG_STAT_DIRTY_PAGES_SET 0x10C +#define IGB_MIG_STAT_DIRTY_PAGES_CLR 0x110 +#define IGB_MIG_STAT_DIRTY_PAGE_COUNT 0x114 +#define IGB_MIG_STAT_DIRTY_QUERY_CNT 0x118 + /* DEVICE_STATE values - mirrors VFIO migration states */ #define IGB_MIG_STATE_ERROR 0 #define IGB_MIG_STATE_STOP 1 @@ -126,7 +135,9 @@ struct igb_mig_dirty_query { uint32_t status; uint32_t bitmap_size; uint32_t dirty_page_count; - uint32_t reserved1[12]; + uint32_t reserved1; + uint64_t dma_writes; + uint32_t reserved2[10]; =20 /* Cache line 2+: bitmap (written by device) */ uint8_t bitmap[]; @@ -154,6 +165,15 @@ typedef struct IgbVfMigState { uint32_t mig_dirty_status; } IgbVfMigState; =20 +typedef struct IgbVfMigStats { + uint32_t dma_writes; + uint64_t dma_bytes; + uint32_t dirty_pages_set; + uint32_t dirty_pages_cleared; + uint32_t dirty_page_count; + uint32_t dirty_query_count; +} IgbVfMigStats; + typedef struct IgbVfState IgbVfState; void igb_pf_init_migration_bar(PCIDevice *dev); bool igbvf_add_migration_cap(PCIDevice *dev, Error **errp); diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c index fce661f68e3d..7fce8d535909 100644 --- a/hw/net/igb_migration.c +++ b/hw/net/igb_migration.c @@ -458,6 +458,7 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn, dma_addr_t addr, dma_addr_t len) { IGBVfDirtyState *ds =3D &core->vf_dirty[vfn]; + IgbVfMigStats *stats =3D &core->vf_mig_stats[vfn]; bool matched =3D false; uint32_t i; =20 @@ -465,6 +466,8 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn, return; } =20 + stats->dma_writes++; + stats->dma_bytes +=3D len; trace_igb_core_dirty_track_dma(vfn, addr, len); =20 for (i =3D 0; i < ds->num_ranges; i++) { @@ -486,7 +489,10 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn, =20 for (page =3D start_page; page <=3D end_page; page++) { if (page < r->nbits) { - set_bit(page, r->bitmap); + if (!test_and_set_bit(page, r->bitmap)) { + stats->dirty_pages_set++; + stats->dirty_page_count++; + } } } } @@ -510,6 +516,15 @@ static uint32_t igb_core_vf_dirty_enable(IgbVfState *s= , uint64_t pgsize, IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); IGBVfDirtyRange *r; =20 + /* + * Reset stats on first enable so the driver can read them after + * disable + */ + if (ds->num_ranges =3D=3D 0) { + memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0, + sizeof(IgbVfMigStats)); + } + if (ds->num_ranges >=3D IGB_MIG_CAPS_MAX_RANGES) { return IGB_MIG_DIRTY_STATUS_TOO_MANY_RANGES; } @@ -561,7 +576,9 @@ static bool igb_core_vf_dirty_query(IgbVfState *s, void *buf, size_t buf_size, size_t *out_size, uint64_t range_iova, uint64_t range_size) { - IGBVfDirtyState *ds =3D igb_core_vf_dirty_state(s); + IGBCore *core =3D igbvf_get_core(s); + IGBVfDirtyState *ds =3D &core->vf_dirty[s->vfn]; + IgbVfMigStats *stats =3D &core->vf_mig_stats[s->vfn]; uint32_t i; =20 for (i =3D 0; i < ds->num_ranges; i++) { @@ -587,6 +604,13 @@ static bool igb_core_vf_dirty_query(IgbVfState *s, bitmap_clear(r->bitmap, start_page, n); } =20 + { + uint32_t cleared =3D bitmap_count_one(buf, count); + stats->dirty_pages_cleared +=3D cleared; + stats->dirty_page_count -=3D MIN(stats->dirty_page_count, clea= red); + stats->dirty_query_count++; + } + *out_size =3D bitmap_empty(buf, count) ? 0 : DIV_ROUND_UP(count, 8= ); return true; } @@ -881,6 +905,7 @@ static void igbvf_mig_data_xfer(IgbVfState *s, uint32_t= val) static uint64_t igbvf_mig_read(void *opaque, hwaddr addr, unsigned size) { IgbVfState *s =3D opaque; + IgbVfMigStats *stats =3D &igbvf_get_core(s)->vf_mig_stats[s->vfn]; IgbVfMigState *ms =3D &s->mig; uint64_t val =3D 0; =20 @@ -915,6 +940,27 @@ static uint64_t igbvf_mig_read(void *opaque, hwaddr ad= dr, unsigned size) case IGB_MIG_DIRTY_STATUS: val =3D ms->mig_dirty_status; break; + case IGB_MIG_STAT_DMA_WRITES: + val =3D stats->dma_writes; + break; + case IGB_MIG_STAT_DMA_BYTES_LO: + val =3D (uint32_t)stats->dma_bytes; + break; + case IGB_MIG_STAT_DMA_BYTES_HI: + val =3D (uint32_t)(stats->dma_bytes >> 32); + break; + case IGB_MIG_STAT_DIRTY_PAGES_SET: + val =3D stats->dirty_pages_set; + break; + case IGB_MIG_STAT_DIRTY_PAGES_CLR: + val =3D stats->dirty_pages_cleared; + break; + case IGB_MIG_STAT_DIRTY_PAGE_COUNT: + val =3D stats->dirty_page_count; + break; + case IGB_MIG_STAT_DIRTY_QUERY_CNT: + val =3D stats->dirty_query_count; + break; default: qemu_log_mask(LOG_GUEST_ERROR, "igbvf: VF%u bad migration BAR read at 0x%" @@ -942,6 +988,7 @@ static uint32_t igbvf_mig_dirty_count(const void *bitma= p, size_t size) static void igbvf_mig_dirty_query(IgbVfState *s, uint64_t pgsize) { IgbVfMigState *ms =3D &s->mig; + IgbVfMigStats *stats =3D &igbvf_get_core(s)->vf_mig_stats[s->vfn]; PCIDevice *dev =3D pcie_sriov_get_pf(PCI_DEVICE(s)); uint64_t buf_addr =3D ms->mig_dirty_buf_addr; uint64_t range_iova =3D 0, range_size =3D 0; @@ -982,12 +1029,17 @@ static void igbvf_mig_dirty_query(IgbVfState *s, uin= t64_t pgsize) stl_le_pci_dma(dev, buf_addr + offsetof(struct igb_mig_dirty_query, dirty_p= age_count), dirty_pages, MEMTXATTRS_UNSPECIFIED); + stq_le_pci_dma(dev, + buf_addr + offsetof(struct igb_mig_dirty_query, dma_wri= tes), + stats->dma_writes, + MEMTXATTRS_UNSPECIFIED); stl_le_pci_dma(dev, buf_addr + offsetof(struct igb_mig_dirty_query, status), valid ? IGB_MIG_DIRTY_STATUS_COMPLETE : 0, MEMTXATTRS_UNSPECIFIED); =20 - trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages); + trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages, + stats->dma_writes); } =20 static void igbvf_mig_dirty_ctrl(IgbVfState *s, uint32_t val) @@ -1135,6 +1187,7 @@ void igbvf_mig_state_reset(IgbVfState *s) ms->mig_dirty_range_size =3D 0; ms->mig_dirty_buf_addr =3D 0; ms->mig_dirty_status =3D IGB_MIG_DIRTY_STATUS_OK; + memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0, sizeof(IgbVfMigSta= ts)); ms->mig_saved_vfre =3D true; ms->mig_saved_vfte =3D true; trace_igbvf_mig_reset(s->vfn); diff --git a/hw/net/trace-events b/hw/net/trace-events index 1e38d9ab3697..76584b43bf19 100644 --- a/hw/net/trace-events +++ b/hw/net/trace-events @@ -305,7 +305,7 @@ igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool = vfre, bool vfte) "VF%u: l igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset" igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%= u: dirty tracking enabled pgsize=3D%"PRIu64" nbits=3D%"PRIu64 igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled" -igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "= VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages" +igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages, u= int64_t dma_writes) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty p= ages (dma_writes=3D%"PRIu64")" =20 # igb_core.c - VF migration diagnostics igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirt= y DMA addr=3D0x%"PRIx64" len=3D%"PRIu64 --=20 2.55.0