From nobody Sat Jul 25 09:34:06 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1784639272; cv=none; d=zohomail.com; s=zohoarc; b=l9FlXJ7k4Q9LqnskkmitwLitj04uvKbIGJNmNCmNWkWEDP98063TdMDoSfbGFdHB5AEjuD1rELV8ffHC9yl4BX32bpbo8SPYYozLm6voyihCoFb8GBVV6MzWhJS1+2l+UT4mE8zeFIAPyWq9ZH1WQoqN6l7pK+/10xs3ahD0qZU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1784639272; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=iGQXlDF8BLnTNvzRXb+T1jikVulOYpNI8Qmce8fxuHI=; b=O3FWgAwSl0xK5YqbuhMWb3+2O/48JdFD8CdG1HIXCC/qnaPdac6sTKvBrKYBqSZbUumC3IErxdjFRSKOI65uBzY4T4qKyx7gW69xLTNOFmfrUvt7Q1OeuK64xvMcdqM2QYE62ZDdsNSn1BLxZ4aWjABsF1oAVuV0+KmkHTjXcJ0= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1784639272649854.5459999859154; Tue, 21 Jul 2026 06:07:52 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1wmAC8-0006nF-ED; Tue, 21 Jul 2026 09:07:36 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1wmAC7-0006n7-5n for qemu-devel@nongnu.org; Tue, 21 Jul 2026 09:07:35 -0400 Received: from mail-pj1-x1032.google.com ([2607:f8b0:4864:20::1032]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1wmAC4-0003Aa-GG for qemu-devel@nongnu.org; Tue, 21 Jul 2026 09:07:34 -0400 Received: by mail-pj1-x1032.google.com with SMTP id 98e67ed59e1d1-38deea72eebso9298695a91.1 for ; Tue, 21 Jul 2026 06:07:31 -0700 (PDT) Received: from gmail.com ([220.248.78.34]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-38e9232c4ccsm1553737a91.7.2026.07.21.06.07.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 06:07:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784639250; x=1785244050; darn=nongnu.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=iGQXlDF8BLnTNvzRXb+T1jikVulOYpNI8Qmce8fxuHI=; b=ZB0Eb4t0F1xbWSvAs3ea4WJpgPQ3yGb2141UlyQKjnW3TrX5hPANdSN7MmmrFMFiw2 iahtH/jfoArCcytYsUzxiDF0oK6Avp16gcuC28IEKL3TLiFbkSdLgch3ndhGZAXkUUiw lKlF99xRoXaWR+F55Y1GkqcadDuVCtAxhAPDGrL8p9dThngu8OfSmNbLCKtVDUpw80kb gd2sYsGn3JuHyexINLXoXCzxu3lNCq1lB8Nh/O3GxFb7396Mb7Ee7KobAIjNpFbJ2b6E b0hHxM0ZjAwJK5FiiiC0BE7DGDTUzLRzPmnC4VxUpRtSNOjf2w3J4Pht0wn3Okejd2TA K3GA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784639250; x=1785244050; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=iGQXlDF8BLnTNvzRXb+T1jikVulOYpNI8Qmce8fxuHI=; b=Asw3kRujJJPUDKQWcnwPLlmRa1w9IBNUHTV2Cvz6Av+Yi64zF9HR+hNBhZpn2gOGFU GChw4Wq9yAagU98FwnpR8dxSqfVmfchFmjOGxRzlADyHUgsobPk7MVK3cXdXPt04/6hI HdbtiLJ0dWya4vLxhiuW++GfMQhBHhxN7x12XUj1yISyHrpjb3DFJp5zI83D9AK12aj5 qa8VqdDqlI9I6bCrvKsVY2a6nzJpVQfoOOuDVdn56K9fJZ/MYsA6Vpn7XwX6QjEw7wQG pw4+r4AOOg+Sd3ffRVlamAZA0uZxqoHAP3wa4jQQ1VKtgREa7etsyL0932wR9tDhPenz 7Fjw== X-Gm-Message-State: AOJu0Yy5wjzN3uCfjnWR+LmApmRaJLvBu/cKbb4pytJ7TAGoWdyRmV+H rGbJzccq8+80Ng3+tz8XVPfnjJWgmlJhP7hxdCtXRpUwvMwm4Z7B9jajCexxOP5M+uAa4A== X-Gm-Gg: AR+sD13/FEd38MHqcjBv9BMl3beZr931tec75SadSsX1gYpx9eXGPw4Jw55t4bEflre pdKiHW5utAKmWjTM18961XciUg/69y1855aYNNKR4lY7fb5uCUfT8IRu+o5RnxWq0b5F3nrwg9b A8pb0nRQKyMQ4xX7OXYYmHayC+cs+e6cTypmNnfdfgmHmPr8CWWOVlo8vtfP8Idc5gty2YLP0JO NPbjoaQghlc/nxir3yktLUuKuchhYQDB2z9Kzo7xN1GaHJSRQoyOB0xoEQtyX7rSz5bsxR1IjLe lwkaKY9OWh3tRKAgv/XuZhV2aJCaDQgwX4yr8sXKRMAEYMFFrLG53ipVz6dN+SB6GBzLYfZH5yt cznGbFy85GftA+pKXzHlZ/+W3sAe2FPnn41AzLs3T0eYBghaUc0ykkmsd/WdFQ/AvwwzgF3qGbd m/otlZFc8= X-Received: by 2002:a17:90b:1f8b:b0:38d:f710:63f0 with SMTP id 98e67ed59e1d1-38e4b5d3ecfmr17946365a91.43.1784639249859; Tue, 21 Jul 2026 06:07:29 -0700 (PDT) From: Yang Wencheng To: qemu-devel@nongnu.org Cc: alex@shazbot.org, clg@redhat.com, pbonzini@redhat.com, peterx@redhat.com, philmd@mailo.com, mst@redhat.com, YangWencheng , Yangwencheng Subject: [PATCH] hw/vfio: Coalesce repeated PCI_COMMAND memory-decode DMA (re)map for passthrough BARs Date: Tue, 21 Jul 2026 21:06:11 +0800 Message-ID: <20260721130611.539089-1-east.moutain.yang@gmail.com> X-Mailer: git-send-email 2.43.0 MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2607:f8b0:4864:20::1032; envelope-from=east.moutain.yang@gmail.com; helo=mail-pj1-x1032.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1784639273478158500 Content-Type: text/plain; charset="utf-8" From: YangWencheng A passthrough device's own MMIO/BAR range is registered with the IOMMU (VFIO_IOMMU_MAP_DMA) when its memory region becomes part of the guest address space, to support peer-to-peer DMA into that BAR from other devices. Toggling the guest's PCI_COMMAND memory-decode-enable bit off and back on -- something PCI enumeration/attribute code does routinely, sometimes several times for the same device (e.g. disable before reprogramming a BAR then re-enable, or generic driver probing during guest OS boot) -- currently tears down and rebuilds this mapping every single time via vfio_listener_region_del()/region_add(), unconditionally. For devices with very large BARs this is expensive: mapping a 64GB BAR into IOMMU page tables measured at ~8.5s per call on this hardware, and was observed being repeated 3-5 times for the same BAR during a single guest boot, because nothing changed about the underlying mapping between the disable and the following re-enable. Defer the actual VFIO_IOMMU_UNMAP_DMA for "ram device" regions (i.e. passthrough device BARs used for P2P DMA -- explicitly not regular guest RAM or RamDiscardManager-backed regions, so migration/ballooning/ virtio-mem are unaffected) instead of issuing it immediately in region_del(). If a matching region_add() for the exact same region (same MemoryRegion pointer, iova, size, vaddr, readonly) follows, cancel the deferred unmap and skip the VFIO_IOMMU_MAP_DMA entirely -- the host-side mapping was never actually removed, so there is nothing to redo. A region_add() for anything that doesn't match maps normally, as before. Nothing is leaked: any mapping still on the pending list is flushed with a real unmap in two places -- vfio_container_instance_finalize(), before the container's fd is closed (VM shutdown / last device in a group removed), and vfio_bars_finalize(), matched by MemoryRegion pointer to just that device's own BARs, so a device hot-unplugged while the container stays alive for other devices can't leave a dangling deferred entry either. This does not weaken the PCI_COMMAND memory-decode security boundary: every guest write to PCI_COMMAND is still forwarded synchronously and unconditionally to the real device's config space in vfio_pci_write_config() (an entirely separate code path, untouched by this change), which is what actually gates whether the physical device claims/responds to transactions targeting its BAR, independent of whatever the IOMMU's routing table still contains. This change only avoids redundant bookkeeping of that routing table when nothing about the mapping has changed. Signed-off-by: Yangwencheng --- hw/vfio/container.c | 53 ++++++++++++++++++++++ hw/vfio/listener.c | 76 ++++++++++++++++++++++++++++++++ hw/vfio/pci.c | 11 +++++ include/hw/vfio/vfio-container.h | 28 ++++++++++++ 4 files changed, 168 insertions(+) diff --git a/hw/vfio/container.c b/hw/vfio/container.c index d09a663732..00676cadc9 100644 --- a/hw/vfio/container.c +++ b/hw/vfio/container.c @@ -298,6 +298,56 @@ GList *vfio_container_get_iova_ranges(const VFIOContai= ner *bcontainer) return g_list_copy_deep(bcontainer->iova_ranges, copy_iova_range, NULL= ); } =20 +/* + * Actually unmap (and free the bookkeeping for) any deferred "ram + * device" unmaps still pending on this container. Must be called + * before the container's fd is closed/reused, and is safe to call at + * any time (e.g. also from a specific device's exit path, to avoid + * leaking a mapping if that device is hot-unplugged while the + * container otherwise stays alive for other devices). + */ +void vfio_flush_pending_ram_device_unmaps(VFIOContainer *bcontainer) +{ + VFIOPendingRamDeviceUnmap *pending, *tmp; + + QLIST_FOREACH_SAFE(pending, &bcontainer->pending_ram_device_unmap_list, + next, tmp) { + int ret =3D vfio_container_dma_unmap(bcontainer, pending->iova, + pending->size, NULL, false); + if (ret) { + error_report("vfio_container_dma_unmap(%p, 0x%"HWADDR_PRIx", " + "0x%"HWADDR_PRIx") =3D %d (%s)", + bcontainer, pending->iova, pending->size, ret, + strerror(-ret)); + } + QLIST_REMOVE(pending, next); + g_free(pending); + } +} + +void vfio_flush_pending_ram_device_unmaps_for_mr(VFIOContainer *bcontainer, + MemoryRegion *mr) +{ + VFIOPendingRamDeviceUnmap *pending, *tmp; + + QLIST_FOREACH_SAFE(pending, &bcontainer->pending_ram_device_unmap_list, + next, tmp) { + if (pending->mr !=3D mr) { + continue; + } + int ret =3D vfio_container_dma_unmap(bcontainer, pending->iova, + pending->size, NULL, false); + if (ret) { + error_report("vfio_container_dma_unmap(%p, 0x%"HWADDR_PRIx", " + "0x%"HWADDR_PRIx") =3D %d (%s)", + bcontainer, pending->iova, pending->size, ret, + strerror(-ret)); + } + QLIST_REMOVE(pending, next); + g_free(pending); + } +} + static void vfio_container_instance_finalize(Object *obj) { VFIOContainer *bcontainer =3D VFIO_IOMMU(obj); @@ -312,6 +362,8 @@ static void vfio_container_instance_finalize(Object *ob= j) g_free(giommu); } =20 + vfio_flush_pending_ram_device_unmaps(bcontainer); + g_list_free_full(bcontainer->iova_ranges, g_free); } =20 @@ -325,6 +377,7 @@ static void vfio_container_instance_init(Object *obj) bcontainer->iova_ranges =3D NULL; QLIST_INIT(&bcontainer->giommu_list); QLIST_INIT(&bcontainer->vrdl_list); + QLIST_INIT(&bcontainer->pending_ram_device_unmap_list); } =20 static const TypeInfo types[] =3D { diff --git a/hw/vfio/listener.c b/hw/vfio/listener.c index c19600e980..e6bcfd028e 100644 --- a/hw/vfio/listener.c +++ b/hw/vfio/listener.c @@ -49,6 +49,55 @@ * Device state interfaces */ =20 +/* + * Defer the VFIO_IOMMU_UNMAP_DMA for a "ram device" (passthrough device + * MMIO/BAR) region instead of doing it immediately. See the comment on + * VFIOPendingRamDeviceUnmap in vfio-container.h for why. + */ +static void vfio_defer_ram_device_unmap(VFIOContainer *bcontainer, + MemoryRegion *mr, hwaddr iova, + hwaddr size, void *vaddr, + bool readonly) +{ + VFIOPendingRamDeviceUnmap *pending =3D g_malloc0(sizeof(*pending)); + + pending->mr =3D mr; + pending->iova =3D iova; + pending->size =3D size; + pending->vaddr =3D vaddr; + pending->readonly =3D readonly; + QLIST_INSERT_HEAD(&bcontainer->pending_ram_device_unmap_list, pending, + next); +} + +/* + * If a deferred unmap exactly matching this (mr, iova, size, vaddr, + * readonly) is pending, cancel it (drop it without ever issuing the + * VFIO_IOMMU_UNMAP_DMA) and report success -- the caller should skip + * mapping, since the host-side mapping was never actually removed. + * Returns false if there was no matching pending unmap, in which case + * the caller must map normally. + */ +static bool vfio_cancel_pending_ram_device_unmap(VFIOContainer *bcontainer, + MemoryRegion *mr, + hwaddr iova, hwaddr size, + void *vaddr, bool readonl= y) +{ + VFIOPendingRamDeviceUnmap *pending; + + QLIST_FOREACH(pending, &bcontainer->pending_ram_device_unmap_list, nex= t) { + if (pending->mr =3D=3D mr && pending->iova =3D=3D iova && + pending->size =3D=3D size && pending->vaddr =3D=3D vaddr && + pending->readonly =3D=3D readonly) { + QLIST_REMOVE(pending, next); + g_free(pending); + return true; + } + } + + return false; +} + =20 static bool vfio_log_sync_needed(const VFIOContainer *bcontainer) { @@ -608,6 +657,22 @@ void vfio_container_region_add(VFIOContainer *bcontain= er, pgmask + 1); return; } + + /* + * If region_del deferred an identical unmap for this exact region + * (same MemoryRegion, iova, size, vaddr, readonly), the underlying + * VFIO_IOMMU_MAP_DMA mapping is still live host-side -- cancel the + * deferred unmap and skip re-mapping. This is what makes toggling + * a passthrough device's PCI_COMMAND memory-decode bit off and + * back on (which normal PCI enumeration/attribute code does, + * sometimes several times per device) cheap instead of repeating + * a possibly multi-second VFIO_IOMMU_MAP_DMA for huge BARs. + */ + if (vfio_cancel_pending_ram_device_unmap(bcontainer, section->mr, + iova, int128_get64(llsize= ), + vaddr, section->readonly)= ) { + return; + } } =20 if (memory_region_skip_iommu_map(section->mr)) { @@ -714,6 +779,17 @@ static void vfio_listener_region_del(MemoryListener *l= istener, =20 pgmask =3D (1ULL << ctz64(bcontainer->pgsizes)) - 1; try_unmap =3D !((iova & pgmask) || (int128_get64(llsize) & pgmask)= ); + + if (try_unmap) { + void *vaddr =3D memory_region_get_ram_ptr(section->mr) + + section->offset_within_region + + (iova - section->offset_within_address_space); + + vfio_defer_ram_device_unmap(bcontainer, section->mr, iova, + int128_get64(llsize), vaddr, + section->readonly); + try_unmap =3D false; + } } else if (memory_region_has_ram_discard_manager(section->mr)) { vfio_ram_discard_unregister_listener(bcontainer, section); /* Unregistering will trigger an unmap. */ diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index c204706e63..86d61891fe 100644 --- a/hw/vfio/pci.c +++ b/hw/vfio/pci.c @@ -30,6 +30,7 @@ #include "hw/core/qdev-properties.h" #include "hw/core/qdev-properties-system.h" #include "hw/vfio/vfio-cpr.h" +#include "hw/vfio/vfio-container.h" #include "migration/vmstate.h" #include "migration/cpr.h" #include "qobject/qdict.h" @@ -2042,6 +2043,16 @@ static void vfio_bars_finalize(VFIOPCIDevice *vdev) vfio_region_finalize(&bar->region); if (bar->mr) { assert(bar->size); + /* + * If a PCI_COMMAND memory-decode toggle deferred this BAR's + * unmap (see vfio_defer_ram_device_unmap()), this device is + * now really going away -- flush it for real so the mapping + * isn't leaked for the remaining lifetime of the container. + */ + if (vdev->vbasedev.bcontainer) { + vfio_flush_pending_ram_device_unmaps_for_mr( + vdev->vbasedev.bcontainer, bar->region.mem); + } g_free(bar->mr); bar->mr =3D NULL; } diff --git a/include/hw/vfio/vfio-container.h b/include/hw/vfio/vfio-contai= ner.h index a15ee2df2b..24525683fb 100644 --- a/include/hw/vfio/vfio-container.h +++ b/include/hw/vfio/vfio-container.h @@ -48,6 +48,7 @@ struct VFIOContainer { bool dirty_pages_started; /* Protected by BQL */ QLIST_HEAD(, VFIOGuestIOMMU) giommu_list; QLIST_HEAD(, VFIORamDiscardListener) vrdl_list; + QLIST_HEAD(, VFIOPendingRamDeviceUnmap) pending_ram_device_unmap_list; QLIST_ENTRY(VFIOContainer) next; QLIST_HEAD(, VFIODevice) device_list; GList *iova_ranges; @@ -76,6 +77,29 @@ typedef struct VFIORamDiscardListener { QLIST_ENTRY(VFIORamDiscardListener) next; } VFIORamDiscardListener; =20 +/* + * A "ram device" region (a passthrough device's own MMIO/BAR range, + * mapped into the IOMMU for peer-to-peer DMA) that region_del wants to + * unmap. The actual VFIO_IOMMU_UNMAP_DMA is deferred: if a matching + * region_add for the identical region shows up again (e.g. the guest + * toggled PCI_COMMAND memory-decode off then back on, which is common + * and can happen several times per device during firmware/OS PCI + * enumeration), the map is still valid host-side and both the unmap and + * the re-map can be skipped entirely. This avoids repeating the + * (potentially multi-second, for huge BARs) VFIO_IOMMU_MAP_DMA ioctl for + * no functional reason. Any mapping still pending here when the + * container is finally torn down gets a real unmap first, so nothing + * is ever leaked. + */ +typedef struct VFIOPendingRamDeviceUnmap { + MemoryRegion *mr; + hwaddr iova; + hwaddr size; + void *vaddr; + bool readonly; + QLIST_ENTRY(VFIOPendingRamDeviceUnmap) next; +} VFIOPendingRamDeviceUnmap; + VFIOAddressSpace *vfio_address_space_get(AddressSpace *as); void vfio_address_space_put(VFIOAddressSpace *space); void vfio_address_space_insert(VFIOAddressSpace *space, @@ -267,4 +291,8 @@ VFIORamDiscardListener *vfio_find_ram_discard_listener( void vfio_container_region_add(VFIOContainer *bcontainer, MemoryRegionSection *section, bool cpr_rema= p); =20 +void vfio_flush_pending_ram_device_unmaps(VFIOContainer *bcontainer); +void vfio_flush_pending_ram_device_unmaps_for_mr(VFIOContainer *bcontainer, + MemoryRegion *mr); + #endif /* HW_VFIO_VFIO_CONTAINER_H */ --=20 2.43.0