From nobody Fri Sep 25 13:54:28 2026 Received: from pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com [34.218.115.239]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 792C647D466; Fri, 11 Sep 2026 12:59:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=34.218.115.239 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131573; cv=none; b=lAK94cBrXNAIDYI/u+fFL4+9vqHvIFVPtOz7uDbUFTfnhENM9UWcIuFAv/N5anT0limGZSFnjP835YgWSOegt1fAJLM57i+GxqnD/8wn182M5P3EaimJEOqgbtNWDw5kdMwRbEjdynTYayju5UzuWR57JQ5jqaTiEIp+Fg5UuOY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131573; c=relaxed/simple; bh=CboY5wvVpRytI8cl2m3ayTAWBXACTQ8PugRCHjW4z7w=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=C5XyMrzRybtiVZIR8cImMuTWi/HhlUGurHFFhizxJcS1WA+auJS8xOCgqdPU23b99hZcqtv5g3YxwAsWHQoejek2MOp2IP/6AGrZnpbUa9z88dGCE83YmNPk1fJLowxFX2MJsp/2CT0bMa6mkhiRTXYqGPv9WokOQLTVdc0VH+w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de; spf=pass smtp.mailfrom=amazon.de; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b=lzXKxTJc; arc=none smtp.client-ip=34.218.115.239 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b="lzXKxTJc" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.de; i=@amazon.de; q=dns/txt; s=amazoncorp2; t=1789131560; x=1820667560; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=0j7orxQTtOZt3x4KoixyDnG0kz3UItRqwWiC4q3s8g4=; b=lzXKxTJc8rG/iGZC858s+f40OBoa6WV2dGzaYi5b19KbJpzrkSV1LEmZ 21vonmjM/aBtE57qC5ylt3Lu/v1FHan9IwKzG7cfZ03WoyLLAaCPltnBm j3Z/I0xH2ZXFBz2xBaBHouSEgowA+hzcrwyhFsvnzmqFLLLCdHkq4BDCt hHoffD7D6dYo2ghFVlQe6l41enOaWx3GkAS2sZqDVbZuvQ78bK7+6Buzd mhir3prI6lquUma3rJRA161yLVl8OEBkRf8Sanunu7FmhCF6lR9HMx4Po zwLuEifq0oPCecDRjwmNcyaU2HK+4R4XynYjDeNEbQ4D6nhbF+oGt2nbt w==; X-CSE-ConnectionGUID: d22A1/TuR0S1FmMpnqCYMA== X-CSE-MsgGUID: 5sbjRQTwQ5aZ3caN2MbuCw== X-IronPort-AV: E=Sophos;i="6.27,97,1787011200"; d="scan'208";a="28213593" Received: from ip-10-5-6-203.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.6.203]) by internal-pdx-out-013.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 11 Sep 2026 12:59:11 +0000 Received: from EX19MTAUWA001.ant.amazon.com [205.251.233.236:5640] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.62.198:2525] with esmtp (Farcaster) id 575257ac-a3ea-49b7-8157-9228a933ad10; Fri, 11 Sep 2026 12:59:11 +0000 (UTC) X-Farcaster-Flow-ID: 575257ac-a3ea-49b7-8157-9228a933ad10 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWA001.ant.amazon.com (10.250.64.217) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 11 Sep 2026 12:59:10 +0000 Received: from dev-dsk-sakacpav-1a-480d1124.eu-west-1.amazon.com (172.19.96.155) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.46; Fri, 11 Sep 2026 12:59:09 +0000 From: Pavol Sakac To: Joerg Roedel , Will Deacon CC: Robin Murphy , , , Bjorn Helgaas , , Subject: [RFC PATCH 1/3] iommu: split sysfs link publication out of iommu_group_alloc_device() Date: Fri, 11 Sep 2026 14:58:32 +0200 Message-ID: <20260911125907.67105-1-sakacpav@amazon.de> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> References: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D036UWC001.ant.amazon.com (10.13.139.233) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" iommu_group_alloc_device() both allocates a group_device and publishes it in sysfs (the "iommu_group" link and the group "devices/" member link with its .%d rename loop), while the iommu instance links are published separately in iommu_init_device() and removed in iommu_deinit_device(). Gather all per-member publication into one helper, iommu_group_link_device(), run under group->mutex, and record success in a new group_device::linked flag. Teardown moves next to the member link removal in __iommu_group_free_device(), which consults the flag so an unpublished member skips sysfs. No functional change intended; this gives publication a single seam so a later commit can move it out of iommu_probe_device_lock. Assisted-by: LLM Signed-off-by: Pavol Sakac --- drivers/iommu/iommu.c | 91 +++++++++++++++++++++++++++++++------------ 1 file changed, 67 insertions(+), 24 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index cd1bca7ede9a..5f92981d9f34 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -77,6 +77,8 @@ struct group_device { struct list_head list; struct device *dev; char *name; + /* Membership published in sysfs; set under group->mutex */ + bool linked; /* * Device is blocked for a pending recovery while its group->domain is * retained. This can happen when: @@ -166,6 +168,8 @@ static ssize_t iommu_group_store_type(struct iommu_grou= p *group, const char *buf, size_t count); static struct group_device *iommu_group_alloc_device(struct iommu_group *g= roup, struct device *dev); +static int iommu_group_link_device(struct iommu_group *group, + struct group_device *device); static void __iommu_group_free_device(struct iommu_group *group, struct group_device *grp_dev); static void iommu_domain_init(struct iommu_domain *domain, unsigned int ty= pe, @@ -515,16 +519,12 @@ static int iommu_init_device(struct device *dev) } dev->iommu->iommu_dev =3D iommu_dev; =20 - ret =3D iommu_device_link(iommu_dev, dev); - if (ret) - goto err_release; - group =3D ops->device_group(dev); if (WARN_ON_ONCE(group =3D=3D NULL)) group =3D ERR_PTR(-EINVAL); if (IS_ERR(group)) { ret =3D PTR_ERR(group); - goto err_unlink; + goto err_release; } dev->iommu_group =3D group; =20 @@ -533,8 +533,6 @@ static int iommu_init_device(struct device *dev) dev->iommu->attach_deferred =3D ops->is_attach_deferred(dev); return 0; =20 -err_unlink: - iommu_device_unlink(iommu_dev, dev); err_release: if (ops->release_device) ops->release_device(dev); @@ -553,8 +551,6 @@ static void iommu_deinit_device(struct device *dev) =20 lockdep_assert_held(&group->mutex); =20 - iommu_device_unlink(dev->iommu->iommu_dev, dev); - /* * release_device() must stop using any attached domain on the device. * If there are still other devices in the group, they are not affected @@ -667,6 +663,11 @@ static int __iommu_probe_device(struct device *dev, st= ruct list_head *group_list */ list_add_tail(&gdev->list, &group->devices); WARN_ON(group->default_domain && !group->domain); + + ret =3D iommu_group_link_device(group, gdev); + if (ret) + goto err_remove_gdev; + if (group->default_domain) iommu_create_device_direct_mappings(group->default_domain, dev); if (group->domain) { @@ -729,10 +730,14 @@ static void __iommu_group_free_device(struct iommu_gr= oup *group, { struct device *dev =3D grp_dev->dev; =20 - sysfs_remove_link(group->devices_kobj, grp_dev->name); - sysfs_remove_link(&dev->kobj, "iommu_group"); + if (grp_dev->linked) { + if (dev_has_iommu(dev)) + iommu_device_unlink(dev->iommu->iommu_dev, dev); + sysfs_remove_link(group->devices_kobj, grp_dev->name); + sysfs_remove_link(&dev->kobj, "iommu_group"); =20 - trace_remove_device_from_group(group->id, dev); + trace_remove_device_from_group(group->id, dev); + } =20 /* * If the group has become empty then ownership must have been @@ -1266,7 +1271,6 @@ static int iommu_create_device_direct_mappings(struct= iommu_domain *domain, static struct group_device *iommu_group_alloc_device(struct iommu_group *g= roup, struct device *dev) { - int ret, i =3D 0; struct group_device *device; =20 device =3D kzalloc_obj(*device); @@ -1275,11 +1279,40 @@ static struct group_device *iommu_group_alloc_devic= e(struct iommu_group *group, =20 device->dev =3D dev; =20 + device->name =3D kasprintf(GFP_KERNEL, "%s", kobject_name(&dev->kobj)); + if (!device->name) { + kfree(device); + dev_err(dev, "Failed to add to iommu group %d: %d\n", + group->id, -ENOMEM); + return ERR_PTR(-ENOMEM); + } + + return device; +} + +/* + * Publish a member's sysfs links under group->mutex. All-or-nothing: + * on failure nothing is left behind and gdev->linked stays false, so + * __iommu_group_free_device() removes only what was published. + */ +static int iommu_group_link_device(struct iommu_group *group, + struct group_device *device) +{ + struct device *dev =3D device->dev; + int ret, i =3D 0; + + lockdep_assert_held(&group->mutex); + + if (dev_has_iommu(dev)) { + ret =3D iommu_device_link(dev->iommu->iommu_dev, dev); + if (ret) + goto err_out; + } + ret =3D sysfs_create_link(&dev->kobj, &group->kobj, "iommu_group"); if (ret) - goto err_free_device; + goto err_unlink_iommu; =20 - device->name =3D kasprintf(GFP_KERNEL, "%s", kobject_name(&dev->kobj)); rename: if (!device->name) { ret =3D -ENOMEM; @@ -1299,23 +1332,25 @@ static struct group_device *iommu_group_alloc_devic= e(struct iommu_group *group, kobject_name(&dev->kobj), i++); goto rename; } - goto err_free_name; + goto err_remove_link; } =20 + device->linked =3D true; + trace_add_device_to_group(group->id, dev); =20 dev_info(dev, "Adding to iommu group %d\n", group->id); =20 - return device; + return 0; =20 -err_free_name: - kfree(device->name); err_remove_link: sysfs_remove_link(&dev->kobj, "iommu_group"); -err_free_device: - kfree(device); +err_unlink_iommu: + if (dev_has_iommu(dev)) + iommu_device_unlink(dev->iommu->iommu_dev, dev); +err_out: dev_err(dev, "Failed to add to iommu group %d: %d\n", group->id, ret); - return ERR_PTR(ret); + return ret; } =20 /** @@ -1329,15 +1364,23 @@ static struct group_device *iommu_group_alloc_devic= e(struct iommu_group *group, int iommu_group_add_device(struct iommu_group *group, struct device *dev) { struct group_device *gdev; + int ret; =20 gdev =3D iommu_group_alloc_device(group, dev); if (IS_ERR(gdev)) return PTR_ERR(gdev); =20 + mutex_lock(&group->mutex); + ret =3D iommu_group_link_device(group, gdev); + if (ret) { + mutex_unlock(&group->mutex); + kfree(gdev->name); + kfree(gdev); + return ret; + } + iommu_group_ref_get(group); dev->iommu_group =3D group; - - mutex_lock(&group->mutex); list_add_tail(&gdev->list, &group->devices); mutex_unlock(&group->mutex); return 0; --=20 2.47.3 From nobody Fri Sep 25 13:54:28 2026 Received: from pdx-out-014.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-014.esa.us-west-2.outbound.mail-perimeter.amazon.com [35.83.148.184]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A781D47D93F; Fri, 11 Sep 2026 12:59:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=35.83.148.184 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131617; cv=none; b=p1jScqBCjFTNg9PlMPvHlFHMIv1dfK956rw7fgbrr9sBrjYWKh74m4sI/9yxtgiM48tm1oXOv2KHktNFvmaN9JxxF2dU2bcsP8z0GY1rLTkZDl/o1bN5cl5cdkdBvcGze/Qya4uc5lCv7k0Cqimhpt2cL5wJEnLyBjY66i4QDwY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131617; c=relaxed/simple; bh=8T5MONSZ8Js55NmVWx8mM2Wfxi/Kbz4ZEP3YDs8/jR8=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=tb4HsN4AMGMtRPXIbYBLJH5A0OEz0us2oVZ7M0HnyQWHKykkWBb8lHvJvgVNhmhXznEte4j1DZrWohav+ZoS2WTRsSTrWYN5FvzYzVE4yRTVHbl4FF/ewiWoQ8gHFlX5kpNzLO9F2AFxc94NcmFNg41acGwRxYh+2eOJNcBA/es= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de; spf=pass smtp.mailfrom=amazon.de; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b=c8o5l+sK; arc=none smtp.client-ip=35.83.148.184 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b="c8o5l+sK" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.de; i=@amazon.de; q=dns/txt; s=amazoncorp2; t=1789131591; x=1820667591; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=2L9S+3Dh+JSZV/kkL2FGWTo1pnT3Slzf5NoDHwIP3q8=; b=c8o5l+sKsdG5ylzjkFEfzl/U1Nikjsz2KSB8zZgvsomCwuAeAjraUDom Xdh73i7gnKJ8w1MPtg09xRtbZKh0U6Gtz3K4Ui+8BVof222Wb7QiW0tNJ i8Haoc6W+jO+hD1rn7jwnsqRtPuLTXNW8MFx0uVTlMQW04P7BRcgqaWz9 W0MOSLRmz4u30xAn1MtH7cOezv9zh+bkbzMC8g2rlhBpthUvRzFFqfY8i 1NWuZq9TDT6x6flLARBipMJ4gCo1jiNMDR4HtUvRJhFML7uYBON7UPL57 nO6RI+98PoWykRGl+eGJWxc7NLMTe4pwpzOh2YrjcmtbaSTNT5nh2MjgJ Q==; X-CSE-ConnectionGUID: j7VRNDKfTU+1yq2BGLqxdQ== X-CSE-MsgGUID: KeS1QwBPSayXD+ZMCPjPtA== X-IronPort-AV: E=Sophos;i="6.27,97,1787011200"; d="scan'208";a="28217848" Received: from ip-10-5-0-115.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.0.115]) by internal-pdx-out-014.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 11 Sep 2026 12:59:44 +0000 Received: from EX19MTAUWC001.ant.amazon.com [205.251.233.105:30085] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.1.54:2525] with esmtp (Farcaster) id af5fa79d-4d72-4a3b-97c5-76066859d93e; Fri, 11 Sep 2026 12:59:44 +0000 (UTC) X-Farcaster-Flow-ID: af5fa79d-4d72-4a3b-97c5-76066859d93e Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWC001.ant.amazon.com (10.250.64.174) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 11 Sep 2026 12:59:43 +0000 Received: from dev-dsk-sakacpav-1a-480d1124.eu-west-1.amazon.com (172.19.96.155) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.46; Fri, 11 Sep 2026 12:59:42 +0000 From: Pavol Sakac To: Joerg Roedel , Will Deacon CC: Robin Murphy , , , Bjorn Helgaas , , Subject: [RFC PATCH 2/3] iommu: create device sysfs links outside iommu_probe_device_lock Date: Fri, 11 Sep 2026 14:58:33 +0200 Message-ID: <20260911125907.67105-2-sakacpav@amazon.de> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> References: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D031UWA003.ant.amazon.com (10.13.139.47) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Parallel device probes convoy on iommu_probe_device_lock, and parallel VF initialisation shows it as a top contention source, with per-member sysfs publication the dominant term under the hold. Publication needs no global ordering: removal already does the reverse under group->mutex alone in __iommu_group_free_device(). Pin the group under the global lock, drop it, and publish every member still lacking links under group->mutex. A device joining a group that already has a domain is published under the global lock instead, because no later pass is guaranteed to visit it; members created by probe_iommu_group() are published by bus_iommu_probe()'s existing drain. The group reference is needed because not every caller pins the device: iommu_add_device() on powerpc, the __init sweep in probe_acpi_namespace_devices() and the of_dma_configure() replay run neither inside device_add() nor under device_lock(). A sibling whose publication failed stays unpublished, retried by any later finalisation or drain; in the worst case it waits until a new member joins, the same terminal shape a failed bus_iommu_probe() drain leaves today. On the hotplug path the links exist before device_add() emits KOBJ_ADD, so udev cannot tell the difference. A failed call also removes the probing device itself from the group: iommu_probe_device() early-exits on membership, so a member left behind would turn a retried probe into a no-op success while the group's members remain unpublished. Assisted-by: LLM Signed-off-by: Pavol Sakac --- drivers/iommu/iommu.c | 70 +++++++++++++++++++++++++++++++++++++++---- 1 file changed, 65 insertions(+), 5 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index 5f92981d9f34..a5e3327aeaca 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -663,11 +663,12 @@ static int __iommu_probe_device(struct device *dev, s= truct list_head *group_list */ list_add_tail(&gdev->list, &group->devices); WARN_ON(group->default_domain && !group->domain); - - ret =3D iommu_group_link_device(group, gdev); - if (ret) - goto err_remove_gdev; - + if (group->domain || group->default_domain) { + /* No later pass is guaranteed to publish this member. */ + ret =3D iommu_group_link_device(group, gdev); + if (ret) + goto err_remove_gdev; + } if (group->default_domain) iommu_create_device_direct_mappings(group->default_domain, dev); if (group->domain) { @@ -710,19 +711,68 @@ static int __iommu_probe_device(struct device *dev, s= truct list_head *group_list int iommu_probe_device(struct device *dev) { const struct iommu_ops *ops; + struct group_device *gdev, *own; + struct iommu_group *group; int ret; =20 mutex_lock(&iommu_probe_device_lock); + /* Already probed; publication belongs to the creating call. */ + if (dev->iommu_group) { + mutex_unlock(&iommu_probe_device_lock); + goto probe_finalize; + } ret =3D __iommu_probe_device(dev, NULL); + if (!ret) { + /* + * Not every caller pins the device, so pin the group across + * the unlock. + */ + group =3D dev->iommu_group; + iommu_group_ref_get(group); + } mutex_unlock(&iommu_probe_device_lock); if (ret) return ret; =20 + mutex_lock(&group->mutex); + /* iommu_group_link_device() is all-or-nothing. */ + for_each_group_device(group, gdev) { + if (gdev->linked) + continue; + ret =3D iommu_group_link_device(group, gdev); + if (ret) + goto err_remove_device; + } + mutex_unlock(&group->mutex); + iommu_group_put(group); + +probe_finalize: ops =3D dev_iommu_ops(dev); if (ops->probe_finalize) ops->probe_finalize(dev); =20 return 0; + +err_remove_device: + /* Retried probes early-exit on membership; failure must undo it. */ + own =3D NULL; + for_each_group_device(group, gdev) { + if (gdev->dev =3D=3D dev) { + own =3D gdev; + break; + } + } + if (own) { + list_del(&own->list); + __iommu_group_free_device(group, own); + iommu_deinit_device(dev); + } + mutex_unlock(&group->mutex); + if (own) + iommu_group_put(group); /* iommu_init_device()'s reference */ + iommu_group_put(group); /* the reference taken above */ + + return ret; } =20 static void __iommu_group_free_device(struct iommu_group *group, @@ -2012,6 +2062,16 @@ static int bus_iommu_probe(const struct bus_type *bu= s) /* Remove item from the list */ list_del_init(&group->entry); =20 + for_each_group_device(group, gdev) { + if (gdev->linked) + continue; + ret =3D iommu_group_link_device(group, gdev); + if (ret) { + mutex_unlock(&group->mutex); + return ret; + } + } + /* * We go to the trouble of deferred default domain creation so * that the cross-group default domain type and the setup of the --=20 2.47.3 From nobody Fri Sep 25 13:54:28 2026 Received: from pdx-out-002.esa.us-west-2.outbound.mail-perimeter.amazon.com (pdx-out-002.esa.us-west-2.outbound.mail-perimeter.amazon.com [44.246.1.125]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D0EEC476070; Fri, 11 Sep 2026 13:00:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=44.246.1.125 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131634; cv=none; b=ss2m+0uEQYj695MlDOld9k0CerDGIw3Bvb45ugIfaWclL/7E7pmIdKOGLmoUKIBXTWjbHAtdi5VZhjaCg6sP0TmUNVM19OoIs1Q0FyrDaQDWEte7/VE+U29g95jpjR7FSgiTETGNEqAk5s+v8QQoEEWinr3F7Uca9C2CjNKsruc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789131634; c=relaxed/simple; bh=S6mB6tAfA18u0SeIBQ79vmZ4kKuamVt3f/eGRJ+jtnY=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=XEPWOp9plsPoajhPllFj9Sv9L5CNksFMwvDwOS/Q0zC2r/qs/M3OpqcQ5A1BGER0lesoKzCAlC3+sBljh23aGwNyFRJeeuiyVWbPHCC/EEycYMzdDFNhdWYI7fulb4RhOC96CBA/UFe+d/WPBwnVp9HSbhg1jIIELZrnj6hp6WU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de; spf=pass smtp.mailfrom=amazon.de; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b=pCBTaPZn; arc=none smtp.client-ip=44.246.1.125 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amazon.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=amazon.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=amazon.de header.i=@amazon.de header.b="pCBTaPZn" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amazon.de; i=@amazon.de; q=dns/txt; s=amazoncorp2; t=1789131623; x=1820667623; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=0tN/bx7Tag7WLT7gLKiEvgdR33AXYplSodqqSmnbwXs=; b=pCBTaPZnbmEYO354f5lwWUp1+SVaYBudha2qfMhPV8a/8/qxpv6bNjPJ pwHZyanXxQZr9TGqXcNDzrYS//Ku7GdjF22gYhhpibLxRN5+mPwn2/a9u Pclp3VAAmM/RTuiD0R+doE7aVUOTfTqcSpvuDzDOkuaAGZixXFYoY61yG bKvrZIIYFBT8RX07cx/pXCC4STzRFVAeu32ih28uMlvkzrZObaXgA5U0B doXWJQI6Hy/7oOZQ4k/1UpGyQq5YCKglvbH3pSf2nRurO1Vjq9dqt9Tmf xNZHFLHwSLYhEv5LCQ4jl2MWRT9m/eql48hnqkBrlAbuqbfBCrmbjnChh Q==; X-CSE-ConnectionGUID: f4KebPnHSOKcqD4LJ3FNJw== X-CSE-MsgGUID: Pw6GVIvFSYiYEhG0pJ72mA== X-IronPort-AV: E=Sophos;i="6.27,97,1787011200"; d="scan'208";a="28428211" Received: from ip-10-5-12-219.us-west-2.compute.internal (HELO smtpout.naws.us-west-2.prod.farcaster.email.amazon.dev) ([10.5.12.219]) by internal-pdx-out-002.esa.us-west-2.outbound.mail-perimeter.amazon.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 11 Sep 2026 13:00:17 +0000 Received: from EX19MTAUWB002.ant.amazon.com [205.251.233.111:16618] by smtpin.naws.us-west-2.prod.farcaster.email.amazon.dev [10.0.43.236:2525] with esmtp (Farcaster) id 9a453679-8488-4fa8-871e-1c9c62cab044; Fri, 11 Sep 2026 13:00:17 +0000 (UTC) X-Farcaster-Flow-ID: 9a453679-8488-4fa8-871e-1c9c62cab044 Received: from EX19D001UWA001.ant.amazon.com (10.13.138.214) by EX19MTAUWB002.ant.amazon.com (10.250.64.231) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.45; Fri, 11 Sep 2026 13:00:17 +0000 Received: from dev-dsk-sakacpav-1a-480d1124.eu-west-1.amazon.com (172.19.96.155) by EX19D001UWA001.ant.amazon.com (10.13.138.214) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA) id 15.2.2562.46; Fri, 11 Sep 2026 13:00:15 +0000 From: Pavol Sakac To: Joerg Roedel , Will Deacon CC: Robin Murphy , , , Bjorn Helgaas , , Subject: [RFC PATCH 3/3] iommu: set up the default domain outside iommu_probe_device_lock Date: Fri, 11 Sep 2026 14:58:34 +0200 Message-ID: <20260911125907.67105-3-sakacpav@amazon.de> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> References: <20260911-vfopt-s2-v1-0-fff3db7e01c2@amazon.de> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: EX19D039UWA002.ant.amazon.com (10.13.139.32) To EX19D001UWA001.ant.amazon.com (10.13.138.214) Content-Type: text/plain; charset="utf-8" Default domain allocation, attach and direct-mapping setup for a group's first device still run inline under iommu_probe_device_lock, though they need only group->mutex: iommu_group_store_type() does this at runtime under group->mutex alone, and bus_iommu_probe() already defers it via the group_list. With publication moved out it is the largest term left under the hold, and parallel VF initialisation convoys behind it. Use the group_list deferral for the singleton probe path too: iommu_probe_device() passes a list and runs the setup after the global unlock, under group->mutex. The global lock's original purpose - commit 01657bc14a39 ("iommu: Avoid races around device probe") - kept two devices of one group from double-initialising it; that invariant moves to group->mutex, where the setup is rechecked: of two racing siblings the first to find the group unset sets it up for every member, and the rest skip. A failed setup unwinds only the calling device, like a failed publication; siblings wait in -EPROBE_DEFER until a later probe finalises the group, or until a new member joins if the failing call had claimed a bus-scanned group's queued entry - the same terminal shape a failed bus_iommu_probe() drain leaves today. A domain the failed setup published stays attached (the first attach is IOMMU_SET_DOMAIN_MUST_SUCCEED), so members left behind get the DMA ops it owed them. bus_iommu_probe() rechecks group->default_domain, as a hotplug probe can finalise a queued group before the drain takes the mutex. Assisted-by: LLM Signed-off-by: Pavol Sakac --- drivers/iommu/iommu.c | 48 ++++++++++++++++++++++++++++++------------- 1 file changed, 34 insertions(+), 14 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index a5e3327aeaca..48fe22bbd1bc 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -657,10 +657,7 @@ static int __iommu_probe_device(struct device *dev, st= ruct list_head *group_list goto err_put_group; } =20 - /* - * The gdev must be in the list before calling - * iommu_setup_default_domain() - */ + /* List the gdev first: the finaliser covers every listed member. */ list_add_tail(&gdev->list, &group->devices); WARN_ON(group->default_domain && !group->domain); if (group->domain || group->default_domain) { @@ -676,15 +673,10 @@ static int __iommu_probe_device(struct device *dev, s= truct list_head *group_list 0); if (ret) goto err_remove_gdev; - } else if (!group->default_domain && !group_list) { - ret =3D iommu_setup_default_domain(group, 0); - if (ret) - goto err_remove_gdev; } else if (!group->default_domain) { /* - * With a group_list argument we defer the default_domain setup - * to the caller by providing a de-duplicated list of groups - * that need further setup. + * Defer setup to the caller draining group_list; both in-tree + * callers pass one. */ if (list_empty(&group->entry)) list_add_tail(&group->entry, group_list); @@ -713,6 +705,7 @@ int iommu_probe_device(struct device *dev) const struct iommu_ops *ops; struct group_device *gdev, *own; struct iommu_group *group; + LIST_HEAD(group_list); int ret; =20 mutex_lock(&iommu_probe_device_lock); @@ -721,7 +714,7 @@ int iommu_probe_device(struct device *dev) mutex_unlock(&iommu_probe_device_lock); goto probe_finalize; } - ret =3D __iommu_probe_device(dev, NULL); + ret =3D __iommu_probe_device(dev, &group_list); if (!ret) { /* * Not every caller pins the device, so pin the group across @@ -735,6 +728,9 @@ int iommu_probe_device(struct device *dev) return ret; =20 mutex_lock(&group->mutex); + /* Only the caller that queued the entry may unlink it. */ + if (!list_empty(&group_list)) + list_del_init(&group->entry); /* iommu_group_link_device() is all-or-nothing. */ for_each_group_device(group, gdev) { if (gdev->linked) @@ -743,6 +739,20 @@ int iommu_probe_device(struct device *dev) if (ret) goto err_remove_device; } + /* + * First thread to find the group unset sets it up for every member; + * binding needs a default domain, so no DMA races the window. + */ + if (!list_empty(&group->devices) && + !group->domain && !group->default_domain) { + ret =3D iommu_setup_default_domain(group, 0); + if (ret) + goto err_remove_device; + for_each_group_device(group, gdev) + if (dev_has_iommu(gdev->dev)) + iommu_setup_dma_ops(gdev->dev, + group->default_domain); + } mutex_unlock(&group->mutex); iommu_group_put(group); =20 @@ -765,8 +775,15 @@ int iommu_probe_device(struct device *dev) if (own) { list_del(&own->list); __iommu_group_free_device(group, own); - iommu_deinit_device(dev); } + /* Members left behind stay attached; give them the DMA ops owed. */ + if (group->default_domain) + for_each_group_device(group, gdev) + if (dev_has_iommu(gdev->dev)) + iommu_setup_dma_ops(gdev->dev, + group->default_domain); + if (own) + iommu_deinit_device(dev); mutex_unlock(&group->mutex); if (own) iommu_group_put(group); /* iommu_init_device()'s reference */ @@ -2077,7 +2094,10 @@ static int bus_iommu_probe(const struct bus_type *bu= s) * that the cross-group default domain type and the setup of the * IOMMU_RESV_DIRECT will work correctly in non-hotpug scenarios. */ - ret =3D iommu_setup_default_domain(group, 0); + ret =3D 0; + /* A hotplug probe may have finalised this group meanwhile. */ + if (!group->default_domain) + ret =3D iommu_setup_default_domain(group, 0); if (ret) { mutex_unlock(&group->mutex); return ret; --=20 2.47.3